今日论文合集:cs.SD语音14篇,eess.AS音频处理17篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】LVNS-RAVE: Diversified audio generation with RAVE and Latent Vector  Novelty Search

标题:LVNS-RAVE:通过RAVE和潜在载体新颖搜索实现多元化音频生成
链接:https://arxiv.org/abs/2404.14063
作者:Jinyue Guo,Anna-Maria Christodoulou,Balint Laczko,Kyrre Glette
备注:Accepted to GECCO 24 Companion
摘要:进化算法和生成式深度学习是声音生成任务中最强大的两种工具。然而,它们也有局限性:进化算法需要复杂的设计,在控制和实现逼真的声音生成方面提出了挑战。生成式深度学习模型通常从数据集复制,缺乏创造力。在本文中,我们提出了LVNS-RAVE,这是一种结合进化算法和生成式深度学习来产生逼真和新颖声音的方法。我们使用的RAVE模型作为声音发生器和VGGish模型作为新颖性评价的潜在矢量新颖性搜索(LVNS)算法。实验结果表明,该方法可以成功地产生多样化的,新颖的音频样本在不同的突变设置使用不同的预训练的RAVE模型。利用突变参数可以很容易地控制生成过程的特性。该算法可以为声音艺术家和音乐家提供一个创造性的工具。
摘要:Evolutionary Algorithms and Generative Deep Learning have been two of the most powerful tools for sound generation tasks. However, they have limitations: Evolutionary Algorithms require complicated designs, posing challenges in control and achieving realistic sound generation. Generative Deep Learning models often copy from the dataset and lack creativity. In this paper, we propose LVNS-RAVE, a method to combine Evolutionary Algorithms and Generative Deep Learning to produce realistic and novel sounds. We use the RAVE model as the sound generator and the VGGish model as a novelty evaluator in the Latent Vector Novelty Search (LVNS) algorithm. The reported experiments show that the method can successfully generate diversified, novel audio samples under different mutation setups using different pre-trained RAVE models. The characteristics of the generation process can be easily controlled with the mutation parameters. The proposed algorithm can be a creative tool for sound artists and musicians.


【2】 Audio Anti-Spoofing Detection: A Survey
标题:音频反欺骗检测:调查
链接:https://arxiv.org/abs/2404.13914
作者:Menglu Li,Yasaman Ahmadiadli,Xiao-Ping Zhang
备注:submitted to ACM Computing Surveys
摘要:智能设备的可用性导致多媒体内容呈指数级增长。然而,深度学习的快速发展已经产生了能够操纵或创建多媒体虚假内容的复杂算法,称为Deepfake。音频Deepfakes通过产生高度逼真的声音构成了重大威胁,从而促进了错误信息的传播。为了解决这个问题,已经组织了许多音频反欺骗检测挑战,以促进反欺骗对策的发展。这份调查报告全面回顾了检测管道中的每个组件,包括算法架构、优化技术、应用程序通用性、评估指标、性能比较、可用数据集和开源可用性。对于每个方面,我们都对最近的进展进行了系统的评估,并讨论了现有的挑战。此外,我们还探讨了音频反欺骗的新兴研究课题,包括部分欺骗检测,跨数据集评估和对抗性攻击防御,同时为未来的工作提出了一些有前途的研究方向。这篇调查论文不仅确定了当前最先进的技术,为未来的实验建立了强有力的基线,而且还为未来的研究人员提供了一条清晰的道路,以理解和增强音频反欺骗检测机制。
摘要:The availability of smart devices leads to an exponential increase in multimedia content. However, the rapid advancements in deep learning have given rise to sophisticated algorithms capable of manipulating or creating multimedia fake content, known as Deepfake. Audio Deepfakes pose a significant threat by producing highly realistic voices, thus facilitating the spread of misinformation. To address this issue, numerous audio anti-spoofing detection challenges have been organized to foster the development of anti-spoofing countermeasures. This survey paper presents a comprehensive review of every component within the detection pipeline, including algorithm architectures, optimization techniques, application generalizability, evaluation metrics, performance comparisons, available datasets, and open-source availability. For each aspect, we conduct a systematic evaluation of the recent advancements, along with discussions on existing challenges. Additionally, we also explore emerging research topics on audio anti-spoofing, including partial spoofing detection, cross-dataset evaluation, and adversarial attack defence, while proposing some promising research directions for future work. This survey paper not only identifies the current state-of-the-art to establish strong baselines for future experiments but also guides future researchers on a clear path for understanding and enhancing the audio anti-spoofing detection mechanisms.


【3】 Retrieval-Augmented Audio Deepfake Detection
标题:检索增强音频Deepfake检测
链接:https://arxiv.org/abs/2404.13892
作者:Zuheng Kang,Yayun He,Botao Zhao,Xiaoyang Qu,Junqing Peng,Jing Xiao,Jianzong Wang
备注:Accepted by the 2024 International Conference on Multimedia Retrieval (ICMR 2024)
摘要:随着语音合成的最新进展,包括文本到语音(TTS)和语音转换(VC)系统,可以生成超现实的音频deepfake,人们越来越担心它们的潜在滥用。然而,大多数Deepfake(DF)检测方法仅依赖于单个模型学习的模糊知识,导致性能瓶颈和透明度问题。受检索增强生成(RAG)的启发,我们提出了一个检索增强检测(RAD)框架,该框架使用相似的检索样本增强测试样本以增强检测。我们还扩展了多融合注意分类器,将其与我们提出的RAD框架相结合。广泛的实验表明,拟议的RAD框架的性能优于基线方法,在ASVspoof 2021 DF集上实现了最先进的结果,在2019年和2021年LA集上实现了竞争性结果。进一步的样本分析表明,检索器一致地检索样本大多来自同一扬声器与声学特性高度一致的查询音频,从而提高检测性能。
摘要:With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growing concern about their potential misuse. However, most deepfake (DF) detection methods rely solely on the fuzzy knowledge learned by a single model, resulting in performance bottlenecks and transparency issues. Inspired by retrieval-augmented generation (RAG), we propose a retrieval-augmented detection (RAD) framework that augments test samples with similar retrieved samples for enhanced detection. We also extend the multi-fusion attentive classifier to integrate it with our proposed RAD framework. Extensive experiments show the superior performance of the proposed RAD framework over baseline methods, achieving state-of-the-art results on the ASVspoof 2021 DF set and competitive results on the 2019 and 2021 LA sets. Further sample analysis indicates that the retriever consistently retrieves samples mostly from the same speaker with acoustic characteristics highly consistent with the query audio, thereby improving detection performance.

【4】 Robotic Blended Sonification: Consequential Robot Sound as Creative  Material for Human-Robot Interaction
标题:机器人混合发声:作为人机交互创意材料的必然机器人声音
链接:https://arxiv.org/abs/2404.13821
作者:Stine S. Johansen,Yanto Browning,Anthony Brumpton,Jared Donovan,Markus Rittenbruch
备注:Paper accepted at ISEA 24, The 29th International Symposium on Electronic Art, Brisbane, Australia, 21-29 June 2024
摘要:目前对机器人声音的研究通常集中在掩盖机器人产生的相应声音或对机器人的声音化数据以创建合成机器人声音。我们建议捕捉,修改和利用,而不是掩盖机器人已经产生的声音。简而言之,这种方法依赖于捕捉机器人的声音,根据上下文信息(例如,协作者的接近度或特定工作序列),并回放修改后的声音。以前的研究表明,非语义的,甚至是机械的,声音作为传达机器人的影响和功能的通信工具是有用的。除此之外,本文提出了一种新的方法,它有两个关键的贡献:(1)一种实时捕获和处理相应的机器人声音的技术,以及(2)一种通过直接人机交互来探索这些声音的方法。从设计,人机交互和创造性实践的方法,由此产生的“机器人混合声化”是一个概念,将相应的机器人声音转化为创造性的材料,可以探索艺术和应用为基础的研究。
摘要:Current research in robotic sounds generally focuses on either masking the consequential sound produced by the robot or on sonifying data about the robot to create a synthetic robot sound. We propose to capture, modify, and utilise rather than mask the sounds that robots are already producing. In short, this approach relies on capturing a robot's sounds, processing them according to contextual information (e.g., collaborators' proximity or particular work sequences), and playing back the modified sound. Previous research indicates the usefulness of non-semantic, and even mechanical, sounds as a communication tool for conveying robotic affect and function. Adding to this, this paper presents a novel approach which makes two key contributions: (1) a technique for real-time capture and processing of consequential robot sounds, and (2) an approach to explore these sounds through direct human-robot interaction. Drawing on methodologies from design, human-robot interaction, and creative practice, the resulting 'Robotic Blended Sonification' is a concept which transforms the consequential robot sounds into a creative material that can be explored artistically and within application-based studies.


【5】 Anchor-aware Deep Metric Learning for Audio-visual Retrieval
标题:用于视听检索的锚感知深度度量学习
链接:https://arxiv.org/abs/2404.13789
作者:Donghuo Zeng,Yanan Wang,Kazushi Ikeda,Yi Yu
备注:9 pages, 5 figures. Accepted by ACM ICMR 2024
摘要:度量学习最大限度地减少了相似(正)数据点对之间的差距,并增加了不同(负)对的分离,旨在捕获底层数据结构并提高视听跨模态检索(AV-CMR)等任务的性能。最近的工作采用采样方法在训练期间从嵌入空间中选择有影响力的数据点。然而,由于训练数据点的稀缺性,模型训练未能充分探索空间,导致整体正负分布的不完整表示。在本文中,我们提出了一种创新的锚点感知深度度量学习(AADML)方法,通过揭示现有数据点之间的潜在相关性来解决这一挑战,从而提高共享嵌入空间的质量。具体来说,我们的方法建立了一个基于相关图的流形结构,考虑每个样本之间的依赖关系作为锚和语义相似的样本。通过使用注意力驱动机制对该底层流形结构内的相关性进行动态加权,获得每个锚点的锚点感知(AA)分数。这些AA分数作为数据代理来计算度量学习方法中的相对距离。在两个视听基准数据集上进行的大量实验证明了我们提出的AADML方法的有效性,显着超过了最先进的模型。此外,我们研究了AA代理与各种度量学习方法的集成,进一步突出了我们方法的有效性。
摘要:Metric learning minimizes the gap between similar (positive) pairs of data points and increases the separation of dissimilar (negative) pairs, aiming at capturing the underlying data structure and enhancing the performance of tasks like audio-visual cross-modal retrieval (AV-CMR). Recent works employ sampling methods to select impactful data points from the embedding space during training. However, the model training fails to fully explore the space due to the scarcity of training data points, resulting in an incomplete representation of the overall positive and negative distributions. In this paper, we propose an innovative Anchor-aware Deep Metric Learning (AADML) method to address this challenge by uncovering the underlying correlations among existing data points, which enhances the quality of the shared embedding space. Specifically, our method establishes a correlation graph-based manifold structure by considering the dependencies between each sample as the anchor and its semantically similar samples. Through dynamic weighting of the correlations within this underlying manifold structure using an attention-driven mechanism, Anchor Awareness (AA) scores are obtained for each anchor. These AA scores serve as data proxies to compute relative distances in metric learning approaches. Extensive experiments conducted on two audio-visual benchmark datasets demonstrate the effectiveness of our proposed AADML method, significantly surpassing state-of-the-art models. Furthermore, we investigate the integration of AA proxies with various metric learning methods, further highlighting the efficacy of our approach.


【6】 Musical Word Embedding for Music Tagging and Retrieval
标题:用于音乐标记和检索的音乐词嵌入
链接:https://arxiv.org/abs/2404.13569
作者:SeungHeon Doh,Jongpil Lee,Dasaem Jeong,Juhan Nam
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:词嵌入已成为基于文本的信息检索的重要手段。通常,词嵌入是从大量的一般和非结构化文本数据中学习的。然而,在音乐领域,单词嵌入可能难以理解音乐上下文或识别音乐相关的实体,如艺术家和曲目。为了解决这个问题,我们提出了一种新的方法,称为音乐词嵌入(MWE),它涉及从各种类型的文本,包括日常和音乐相关的词汇学习。我们将MWE集成到一个音频-单词联合表示框架中,用于标记和检索音乐,使用标签,艺术家和轨道等具有不同级别的音乐特异性的单词。我们的实验表明,使用一个更具体的音乐词,如轨道的结果在更好的检索性能,而使用一个不太具体的术语,如标签,导致更好的标记性能。为了平衡这种妥协,我们建议多原型训练,使用不同级别的音乐特异性的话联合。我们在两个数据集(百万歌曲数据集和MTG-Jamendo)上评估了四个任务(标签排名预测,音乐标签,按标签查询和按曲目查询)的单词嵌入和音频单词联合嵌入。我们的研究结果表明,建议的MWE是更有效和强大的比传统的词嵌入。
摘要:Word embedding has become an essential means for text-based information retrieval. Typically, word embeddings are learned from large quantities of general and unstructured text data. However, in the domain of music, the word embedding may have difficulty understanding musical contexts or recognizing music-related entities like artists and tracks. To address this issue, we propose a new approach called Musical Word Embedding (MWE), which involves learning from various types of texts, including both everyday and music-related vocabulary. We integrate MWE into an audio-word joint representation framework for tagging and retrieving music, using words like tag, artist, and track that have different levels of musical specificity. Our experiments show that using a more specific musical word like track results in better retrieval performance, while using a less specific term like tag leads to better tagging performance. To balance this compromise, we suggest multi-prototype training that uses words with different levels of musical specificity jointly. We evaluate both word embedding and audio-word joint embedding on four tasks (tag rank prediction, music tagging, query-by-tag, and query-by-track) across two datasets (Million Song Dataset and MTG-Jamendo). Our findings show that the suggested MWE is more efficient and robust than the conventional word embedding.


【7】 Sparse Direction of Arrival Estimation Method Based on Vector Signal  Reconstruction with a Single Vector Sensor
标题:基于单载体传感器载体信号重建的稀疏波达方向估计方法
链接:https://arxiv.org/abs/2404.13568
作者:Jiabin Guo
备注:20 pages
摘要:本文研究了单矢量水听器在水声信号波达方向估计中的应用。针对传统DOA估计方法在多源环境和噪声干扰下的局限性,提出了一种矢量信号重构(VSR)技术。该方法通过复数运算和矢量信号重构,将单个矢量水听器信号的协方差矩阵变换为适合于无网格稀疏方法的Toeplitz结构。在此基础上,介绍了两种基于矢量信号重构的稀疏DOA估计算法。理论分析和仿真实验表明,与传统算法相比,本文算法在多源信号和低信噪比环境下显著提高了DOA估计的精度和分辨率。本研究的贡献在于为复杂环境下的单矢量水听器波达方向估计提供了一种有效的新方法,为矢量水听器信号处理领域引入了新的研究方向和解决方案。
摘要:This study investigates the application of single vector hydrophones in underwater acoustic signal processing for Direction of Arrival (DOA) estimation. Addressing the limitations of traditional DOA estimation methods in multi-source environments and under noise interference, this research proposes a Vector Signal Reconstruction (VSR) technique. This technique transforms the covariance matrix of single vector hydrophone signals into a Toeplitz structure suitable for gridless sparse methods through complex calculations and vector signal reconstruction. Furthermore, two sparse DOA estimation algorithms based on vector signal reconstruction are introduced. Theoretical analysis and simulation experiments demonstrate that the proposed algorithms significantly improve the accuracy and resolution of DOA estimation in multi-source signals and low Signal-to-Noise Ratio (SNR) environments compared to traditional algorithms. The contribution of this study lies in providing an effective new method for DOA estimation with single vector hydrophones in complex environments, introducing new research directions and solutions in the field of vector hydrophone signal processing.


【8】 AudioRepInceptionNeXt: A lightweight single-stream architecture for  efficient audio recognition
标题:AudioRepInceptionNeXt:用于高效音频识别的轻量级单流架构
链接:https://arxiv.org/abs/2404.13551
作者:Kin Wai Lau,Yasar Abbas Ur Rehman,Lai-Man Po
摘要:最近的研究已经成功地将基于视觉的卷积神经网络(CNN)架构用于使用Mel-频谱图的音频识别任务。然而,这些CNN具有较高的计算成本和内存需求,限制了它们在低端边缘设备上的部署。受InceptionNeXt和ConvNeXt等高效视觉模型的成功启发,我们提出了AudioRepInceptionNeXt,一种单流架构。它的基本构建块将具有k x k内核的递减尺度的并行多分支深度方向卷积分解为两个多分支深度方向卷积的级联。第一多分支由并行多尺度1 x k深度方向卷积层组成,随后是采用并行多尺度k x 1深度方向卷积层的类似多分支。这减少了计算和内存占用,同时分离了梅尔频谱图的时间和频率处理。大的内核捕获全局频率和长活动,而小的内核捕获局部频率和短活动。我们还在推理过程中重新参数化了多分支设计,以进一步提高速度而不损失准确性。实验表明,AudioRepInceptionNeXt将参数和计算量减少了50%以上,并将推理速度提高了1.28倍,同时保持了相当的准确性。它还可以在各种音频识别任务中进行鲁棒学习。代码可在https://github.com/StevenLauHKHK/AudioRepInceptionNeXt上获得。
摘要:Recent research has successfully adapted vision-based convolutional neural network (CNN) architectures for audio recognition tasks using Mel-Spectrograms. However, these CNNs have high computational costs and memory requirements, limiting their deployment on low-end edge devices. Motivated by the success of efficient vision models like InceptionNeXt and ConvNeXt, we propose AudioRepInceptionNeXt, a single-stream architecture. Its basic building block breaks down the parallel multi-branch depth-wise convolutions with descending scales of k x k kernels into a cascade of two multi-branch depth-wise convolutions. The first multi-branch consists of parallel multi-scale 1 x k depth-wise convolutional layers followed by a similar multi-branch employing parallel multi-scale k x 1 depth-wise convolutional layers. This reduces computational and memory footprint while separating time and frequency processing of Mel-Spectrograms. The large kernels capture global frequencies and long activities, while small kernels get local frequencies and short activities. We also reparameterize the multi-branch design during inference to further boost speed without losing accuracy. Experiments show that AudioRepInceptionNeXt reduces parameters and computations by 50%+ and improves inference speed 1.28x over state-of-the-art CNNs like the Slow-Fast while maintaining comparable accuracy. It also learns robustly across a variety of audio recognition tasks. Codes are available at https://github.com/StevenLauHKHK/AudioRepInceptionNeXt.

【9】 MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and  Hierarchical Cooperative Attention
标题:MFHCA:通过多空间融合和分层合作注意增强语音情感识别
链接:https://arxiv.org/abs/2404.13509
作者:Xinxin Jiao,Liejun Wang,Yinfeng Yu
备注:Main paper (5 pages). Accepted for publication by ICME 2024
摘要:语音情感识别在人机交互中至关重要,但从音频中提取和使用情感线索带来了挑战。本文介绍了MFHCA,一种新的方法,语音情感识别的多空间融合和层次合作注意力的频谱和原始音频。我们采用多空间融合模块(MF),以有效地识别情感相关的频谱图区域,并整合更高层次的声学信息的休伯特功能。我们的方法还包括一个层次合作注意模块(HCA)合并功能,从不同的听觉水平。在IEMOCAP数据集上的实验结果表明,该方法在加权精度和非加权精度上分别提高了2.6%和1.87%.大量的实验证明了该方法的有效性。
摘要:Speech emotion recognition is crucial in human-computer interaction, but extracting and using emotional cues from audio poses challenges. This paper introduces MFHCA, a novel method for Speech Emotion Recognition using Multi-Spatial Fusion and Hierarchical Cooperative Attention on spectrograms and raw audio. We employ the Multi-Spatial Fusion module (MF) to efficiently identify emotion-related spectrogram regions and integrate Hubert features for higher-level acoustic information. Our approach also includes a Hierarchical Cooperative Attention module (HCA) to merge features from various auditory levels. We evaluate our method on the IEMOCAP dataset and achieve 2.6\% and 1.87\% improvements on the weighted accuracy and unweighted accuracy, respectively. Extensive experiments demonstrate the effectiveness of the proposed method.


【10】 Text-dependent Speaker Verification (TdSV) Challenge 2024: Challenge  Evaluation Plan
标题:2024年文本相关说话者验证(TdSV)挑战:挑战评估计划
链接:https://arxiv.org/abs/2404.13428
作者:Zeinali Hossein,Lee Kong Aik,Alam Jahangir,Burget Lukas
摘要:本文概述了2024年文本相关说话人验证(TdSV)挑战赛,该挑战赛的重点是分析和探索文本相关说话人验证的新方法。这项挑战的主要目标是激励参与者开发单一但有竞争力的系统,进行彻底的分析,并探索创新的概念,如多任务学习,自我监督学习,Few-Shot学习等,用于文本相关的说话人验证。
摘要:This document outlines the Text-dependent Speaker Verification (TdSV) Challenge 2024, which centers on analyzing and exploring novel approaches for text-dependent speaker verification. The primary goal of this challenge is to motive participants to develop single yet competitive systems, conduct thorough analyses, and explore innovative concepts such as multi-task learning, self-supervised learning, few-shot learning, and others, for text-dependent speaker verification.

【11】 Music Consistency Models
标题:音乐一致性模型
链接:https://arxiv.org/abs/2404.13358
作者:Zhengcong Fei,Mingyuan Fan,Junshi Huang
摘要:一致性模型在促进有效的图像/视频生成方面表现出显着的能力,从而能够以最少的采样步骤进行合成。它已被证明是有利的,在减轻与扩散模型相关的计算负担。然而,一致性模型在音乐生成中的应用在很大程度上仍然是未开发的。为了解决这个问题,我们提出了音乐一致性模型(\texttt{MusicCM}),它利用一致性模型的概念来有效地合成音乐片段的梅尔频谱图,保持高质量,同时最大限度地减少采样步骤的数量。建立在现有的文本到音乐的扩散模型,\texttt{MusicCM}模型结合了一致性蒸馏和对抗性训练。此外,我们发现它有利于产生扩展的连贯音乐,将多个扩散过程与共享的约束。实验结果表明,我们的模型在计算效率,保真度和自然度方面的有效性。值得注意的是,\texttt{MusicCM}仅用四个采样步骤就实现了无缝音乐合成,例如,每分钟只有一秒的音乐片段,展示了实时应用的潜力。
摘要:Consistency models have exhibited remarkable capabilities in facilitating efficient image/video generation, enabling synthesis with minimal sampling steps. It has proven to be advantageous in mitigating the computational burdens associated with diffusion models. Nevertheless, the application of consistency models in music generation remains largely unexplored. To address this gap, we present Music Consistency Models (\texttt{MusicCM}), which leverages the concept of consistency models to efficiently synthesize mel-spectrogram for music clips, maintaining high quality while minimizing the number of sampling steps. Building upon existing text-to-music diffusion models, the \texttt{MusicCM} model incorporates consistency distillation and adversarial discriminator training. Moreover, we find it beneficial to generate extended coherent music by incorporating multiple diffusion processes with shared constraints. Experimental results reveal the effectiveness of our model in terms of computational efficiency, fidelity, and naturalness. Notable, \texttt{MusicCM} achieves seamless music synthesis with a mere four sampling steps, e.g., only one second per minute of the music clip, showcasing the potential for real-time application.


【12】 Double Mixture: Towards Continual Event Detection from Speech
标题:双重混合:从语音中实现连续事件检测
链接:https://arxiv.org/abs/2404.13289
作者:Jingqi Kang,Tongtong Wu,Jinming Zhao,Guitao Wang,Yinwei Wei,Hao Yang,Guilin Qi,Yuan-Fang Li,Gholamreza Haffari
备注:The first two authors contributed equally to this work
摘要:语音事件检测是多媒体检索的关键,涉及语义和声学事件的标记。传统的ASR系统往往忽视这些事件之间的相互作用,只关注内容,即使对话的解释可能会因环境背景而异。本文解决了语音事件检测中的两个主要挑战:新事件的不断整合,而不会忘记以前的,从声学事件的语义的解开。我们引入了一个新的任务,从语音中连续事件检测,为此我们还提供了两个基准数据集。为了解决灾难性遗忘和有效解纠缠的挑战,我们提出了一种新的方法,“双重混合”。这种方法将语音专业知识与强大的记忆机制相结合,以增强适应性并防止遗忘。我们的综合实验表明,这项任务提出了重大挑战,目前最先进的计算机视觉或自然语言处理方法无法有效解决这些挑战。我们的方法实现了最低的遗忘率和最高水平的泛化,证明了在各种持续学习序列中的鲁棒性。我们的代码和数据可在https://anonymous.4open.science/status/Continual-SpeechED-6461上获得。
摘要:Speech event detection is crucial for multimedia retrieval, involving the tagging of both semantic and acoustic events. Traditional ASR systems often overlook the interplay between these events, focusing solely on content, even though the interpretation of dialogue can vary with environmental context. This paper tackles two primary challenges in speech event detection: the continual integration of new events without forgetting previous ones, and the disentanglement of semantic from acoustic events. We introduce a new task, continual event detection from speech, for which we also provide two benchmark datasets. To address the challenges of catastrophic forgetting and effective disentanglement, we propose a novel method, 'Double Mixture.' This method merges speech expertise with robust memory mechanisms to enhance adaptability and prevent forgetting. Our comprehensive experiments show that this task presents significant challenges that are not effectively addressed by current state-of-the-art methods in either computer vision or natural language processing. Our approach achieves the lowest rates of forgetting and the highest levels of generalization, proving robust across various continual learning sequences. Our code and data are available at https://anonymous.4open.science/status/Continual-SpeechED-6461.


【13】 Track Role Prediction of Single-Instrumental Sequences
标题:单乐器序列的音轨角色预测
链接:https://arxiv.org/abs/2404.13286
作者:Changheon Han,Suhyun Lee,Minsam Ko
备注:ISMIR LBD 2023
摘要:在创作过程中,选择合适的单器乐序列并分配其轨道角色是一项不可或缺的工作。然而,手动确定大量音乐样本的音轨角色可能是耗时且劳动密集的。这项研究介绍了一种深度学习模型,旨在自动预测单乐器音乐序列的轨道角色。我们的评估显示,在符号域和音频域的预测准确率分别为87%和84%。所提出的轨道角色预测方法有望在人工智能音乐生成和分析中应用。
摘要:In the composition process, selecting appropriate single-instrumental music sequences and assigning their track-role is an indispensable task. However, manually determining the track-role for a myriad of music samples can be time-consuming and labor-intensive. This study introduces a deep learning model designed to automatically predict the track-role of single-instrumental music sequences. Our evaluations show a prediction accuracy of 87% in the symbolic domain and 84% in the audio domain. The proposed track-role prediction methods hold promise for future applications in AI music generation and analysis.


【14】 Intro to Quantum Harmony: Chords in Superposition
标题:量子和谐简介:叠加的和弦
链接:https://arxiv.org/abs/2404.13140
作者:Christopher Dobrian,Omar Costa Hamido
摘要:量子理论和音乐理论之间的相关性-特别是量子计算原理和音乐和声之间的相关性-可以为音乐理论家和作曲家带来新的理解和新的方法。量子叠加原理被证明与音乐意义的不同解释密切相关。叠加直接在作者的量子计算模拟中实现,应用于计算机生成音乐创作的决策过程。
摘要:Correlations between quantum theory and music theory - specifically between principles of quantum computing and musical harmony - can lead to new understandings and new methodologies for music theorists and composers. The quantum principle of superposition is shown to be closely related to different interpretations of musical meaning. Superposition is implemented directly in the authors' simulations of quantum computing, as applied in the decision-making processes of computer-generated music composition.


eess.AS音频处理
【1】 On fusing active and passive acoustic sensing for simultaneous  localization and mapping
标题:融合主动和被动声学传感以同时定位和绘图
链接:https://arxiv.org/abs/2404.13116
作者:Aidan J. Bradley,Nicole Abaid
备注:14 pages, 13 figures, 2 tables, journal submission
摘要:对蝙蝠社会行为的研究表明,它们有能力窃听附近同类发出的信号。它们可以将这种"被动”数据与从自己的信号中主动收集的数据融合在一起,以获得更多关于其环境的信息,使它们能够更有效地飞行和狩猎,并在争夺猎物时避免或造成干扰。声学传感器也有类似的功能,但通常一次只能用于主动或被动功能。在同一阵列中同时使用主动和被动传感是否有好处?在这项工作中,我们定义了一个家庭的主动,被动和融合传感系统的模型来测量范围和轴承数据从一个环境中定义的基于点的地标。这些测量被用来解决问题的同时定位和地图(SLAM)与扩展卡尔曼滤波器(EKF)和FastSLAM 2.0方法。我们的研究结果与以前的研究结果一致。具体而言,当主动感测限于窄角度范围时,融合感测即使不能更好,也可以同样准确地执行,同时还允许传感器感知更多的周围环境。
摘要:Studies on the social behaviors of bats show that they have the ability to eavesdrop on the signals emitted by conspecifics in their vicinity. They can fuse this ``passive" data with actively collected data from their own signals to get more information about their environment, allowing them to fly and hunt more efficiently and to avoid or cause jamming when competing for prey. Acoustic sensors are capable of similar feats but are generally used in only an active or passive capacity at one time. Is there a benefit to using both active and passive sensing simultaneously in the same array? In this work we define a family of models for active, passive, and fused sensing systems to measure range and bearing data from an environment defined by point-based landmarks. These measurements are used to solve the problem of simultaneous localization and mapping (SLAM) with extended Kalman filter (EKF) and FastSLAM 2.0 approaches. Our results show agreement with previous findings. Specifically, when active sensing is limited to a narrow angular range, fused sensing can perform just as accurately if not better, while also allowing the sensor to perceive more of the surrounding environment.


【2】 LVNS-RAVE: Diversified audio generation with RAVE and Latent Vector  Novelty Search
标题:LVNS-RAVE:通过RAVE和潜在载体新颖搜索实现多元化音频生成
链接:https://arxiv.org/abs/2404.14063
作者:Jinyue Guo,Anna-Maria Christodoulou,Balint Laczko,Kyrre Glette
备注:Accepted to GECCO 24 Companion
摘要:进化算法和生成式深度学习是声音生成任务中最强大的两种工具。然而,它们也有局限性:进化算法需要复杂的设计,在控制和实现逼真的声音生成方面提出了挑战。生成式深度学习模型通常从数据集复制,缺乏创造力。在本文中,我们提出了LVNS-RAVE,这是一种结合进化算法和生成式深度学习来产生逼真和新颖声音的方法。我们使用的RAVE模型作为声音发生器和VGGish模型作为新颖性评价的潜在矢量新颖性搜索(LVNS)算法。实验结果表明,该方法可以成功地产生多样化的,新颖的音频样本在不同的突变设置使用不同的预训练的RAVE模型。利用突变参数可以很容易地控制生成过程的特性。该算法可以为声音艺术家和音乐家提供一个创造性的工具。
摘要:Evolutionary Algorithms and Generative Deep Learning have been two of the most powerful tools for sound generation tasks. However, they have limitations: Evolutionary Algorithms require complicated designs, posing challenges in control and achieving realistic sound generation. Generative Deep Learning models often copy from the dataset and lack creativity. In this paper, we propose LVNS-RAVE, a method to combine Evolutionary Algorithms and Generative Deep Learning to produce realistic and novel sounds. We use the RAVE model as the sound generator and the VGGish model as a novelty evaluator in the Latent Vector Novelty Search (LVNS) algorithm. The reported experiments show that the method can successfully generate diversified, novel audio samples under different mutation setups using different pre-trained RAVE models. The characteristics of the generation process can be easily controlled with the mutation parameters. The proposed algorithm can be a creative tool for sound artists and musicians.


【3】 Audio Anti-Spoofing Detection: A Survey
标题:音频反欺骗检测:调查
链接:https://arxiv.org/abs/2404.13914
作者:Menglu Li,Yasaman Ahmadiadli,Xiao-Ping Zhang
备注:submitted to ACM Computing Surveys
摘要:智能设备的可用性导致多媒体内容呈指数级增长。然而,深度学习的快速发展已经产生了能够操纵或创建多媒体虚假内容的复杂算法,称为Deepfake。音频Deepfakes通过产生高度逼真的声音构成了重大威胁,从而促进了错误信息的传播。为了解决这个问题,已经组织了许多音频反欺骗检测挑战,以促进反欺骗对策的发展。这份调查报告全面回顾了检测管道中的每个组件,包括算法架构、优化技术、应用程序通用性、评估指标、性能比较、可用数据集和开源可用性。对于每个方面,我们都对最近的进展进行了系统的评估,并讨论了现有的挑战。此外,我们还探讨了音频反欺骗的新兴研究课题,包括部分欺骗检测,跨数据集评估和对抗性攻击防御,同时为未来的工作提出了一些有前途的研究方向。这篇调查论文不仅确定了当前最先进的技术,为未来的实验建立了强有力的基线,而且还为未来的研究人员提供了一条清晰的道路,以理解和增强音频反欺骗检测机制。
摘要:The availability of smart devices leads to an exponential increase in multimedia content. However, the rapid advancements in deep learning have given rise to sophisticated algorithms capable of manipulating or creating multimedia fake content, known as Deepfake. Audio Deepfakes pose a significant threat by producing highly realistic voices, thus facilitating the spread of misinformation. To address this issue, numerous audio anti-spoofing detection challenges have been organized to foster the development of anti-spoofing countermeasures. This survey paper presents a comprehensive review of every component within the detection pipeline, including algorithm architectures, optimization techniques, application generalizability, evaluation metrics, performance comparisons, available datasets, and open-source availability. For each aspect, we conduct a systematic evaluation of the recent advancements, along with discussions on existing challenges. Additionally, we also explore emerging research topics on audio anti-spoofing, including partial spoofing detection, cross-dataset evaluation, and adversarial attack defence, while proposing some promising research directions for future work. This survey paper not only identifies the current state-of-the-art to establish strong baselines for future experiments but also guides future researchers on a clear path for understanding and enhancing the audio anti-spoofing detection mechanisms.


【4】 Retrieval-Augmented Audio Deepfake Detection
标题:检索增强音频Deepfake检测
链接:https://arxiv.org/abs/2404.13892
作者:Zuheng Kang,Yayun He,Botao Zhao,Xiaoyang Qu,Junqing Peng,Jing Xiao,Jianzong Wang
备注:Accepted by the 2024 International Conference on Multimedia Retrieval (ICMR 2024)
摘要:随着语音合成的最新进展,包括文本到语音(TTS)和语音转换(VC)系统,可以生成超现实的音频deepfake,人们越来越担心它们的潜在滥用。然而,大多数Deepfake(DF)检测方法仅依赖于单个模型学习的模糊知识,导致性能瓶颈和透明度问题。受检索增强生成(RAG)的启发,我们提出了一个检索增强检测(RAD)框架,该框架使用相似的检索样本增强测试样本以增强检测。我们还扩展了多融合注意分类器,将其与我们提出的RAD框架相结合。广泛的实验表明,拟议的RAD框架的性能优于基线方法,在ASVspoof 2021 DF集上实现了最先进的结果,在2019年和2021年LA集上实现了竞争性结果。进一步的样本分析表明,检索器一致地检索样本大多来自同一扬声器与声学特性高度一致的查询音频,从而提高检测性能。
摘要:With recent advances in speech synthesis including text-to-speech (TTS) and voice conversion (VC) systems enabling the generation of ultra-realistic audio deepfakes, there is growing concern about their potential misuse. However, most deepfake (DF) detection methods rely solely on the fuzzy knowledge learned by a single model, resulting in performance bottlenecks and transparency issues. Inspired by retrieval-augmented generation (RAG), we propose a retrieval-augmented detection (RAD) framework that augments test samples with similar retrieved samples for enhanced detection. We also extend the multi-fusion attentive classifier to integrate it with our proposed RAD framework. Extensive experiments show the superior performance of the proposed RAD framework over baseline methods, achieving state-of-the-art results on the ASVspoof 2021 DF set and competitive results on the 2019 and 2021 LA sets. Further sample analysis indicates that the retriever consistently retrieves samples mostly from the same speaker with acoustic characteristics highly consistent with the query audio, thereby improving detection performance.

【5】 Robotic Blended Sonification: Consequential Robot Sound as Creative  Material for Human-Robot Interaction
标题:机器人混合发声:作为人机交互创意材料的必然机器人声音
链接:https://arxiv.org/abs/2404.13821
作者:Stine S. Johansen,Yanto Browning,Anthony Brumpton,Jared Donovan,Markus Rittenbruch
备注:Paper accepted at ISEA 24, The 29th International Symposium on Electronic Art, Brisbane, Australia, 21-29 June 2024
摘要:目前对机器人声音的研究通常集中在掩盖机器人产生的相应声音或对机器人的声音化数据以创建合成机器人声音。我们建议捕捉,修改和利用,而不是掩盖机器人已经产生的声音。简而言之,这种方法依赖于捕捉机器人的声音,根据上下文信息(例如,协作者的接近度或特定工作序列),并回放修改后的声音。以前的研究表明,非语义的,甚至是机械的,声音作为传达机器人的影响和功能的通信工具是有用的。除此之外,本文提出了一种新的方法,它有两个关键的贡献:(1)一种实时捕获和处理相应的机器人声音的技术,以及(2)一种通过直接人机交互来探索这些声音的方法。从设计,人机交互和创造性实践的方法,由此产生的“机器人混合声化”是一个概念,将相应的机器人声音转化为创造性的材料,可以探索艺术和应用为基础的研究。
摘要:Current research in robotic sounds generally focuses on either masking the consequential sound produced by the robot or on sonifying data about the robot to create a synthetic robot sound. We propose to capture, modify, and utilise rather than mask the sounds that robots are already producing. In short, this approach relies on capturing a robot's sounds, processing them according to contextual information (e.g., collaborators' proximity or particular work sequences), and playing back the modified sound. Previous research indicates the usefulness of non-semantic, and even mechanical, sounds as a communication tool for conveying robotic affect and function. Adding to this, this paper presents a novel approach which makes two key contributions: (1) a technique for real-time capture and processing of consequential robot sounds, and (2) an approach to explore these sounds through direct human-robot interaction. Drawing on methodologies from design, human-robot interaction, and creative practice, the resulting 'Robotic Blended Sonification' is a concept which transforms the consequential robot sounds into a creative material that can be explored artistically and within application-based studies.


【6】 Anchor-aware Deep Metric Learning for Audio-visual Retrieval
标题:用于视听检索的锚感知深度度量学习
链接:https://arxiv.org/abs/2404.13789
作者:Donghuo Zeng,Yanan Wang,Kazushi Ikeda,Yi Yu
备注:9 pages, 5 figures. Accepted by ACM ICMR 2024
摘要:度量学习最大限度地减少了相似(正)数据点对之间的差距,并增加了不同(负)对的分离,旨在捕获底层数据结构并提高视听跨模态检索(AV-CMR)等任务的性能。最近的工作采用采样方法在训练期间从嵌入空间中选择有影响力的数据点。然而,由于训练数据点的稀缺性,模型训练未能充分探索空间,导致整体正负分布的不完整表示。在本文中,我们提出了一种创新的锚点感知深度度量学习(AADML)方法,通过揭示现有数据点之间的潜在相关性来解决这一挑战,从而提高共享嵌入空间的质量。具体来说,我们的方法建立了一个基于相关图的流形结构,考虑每个样本之间的依赖关系作为锚和语义相似的样本。通过使用注意力驱动机制对该底层流形结构内的相关性进行动态加权,获得每个锚点的锚点感知(AA)分数。这些AA分数作为数据代理来计算度量学习方法中的相对距离。在两个视听基准数据集上进行的大量实验证明了我们提出的AADML方法的有效性,显着超过了最先进的模型。此外,我们研究了AA代理与各种度量学习方法的集成,进一步突出了我们方法的有效性。
摘要:Metric learning minimizes the gap between similar (positive) pairs of data points and increases the separation of dissimilar (negative) pairs, aiming at capturing the underlying data structure and enhancing the performance of tasks like audio-visual cross-modal retrieval (AV-CMR). Recent works employ sampling methods to select impactful data points from the embedding space during training. However, the model training fails to fully explore the space due to the scarcity of training data points, resulting in an incomplete representation of the overall positive and negative distributions. In this paper, we propose an innovative Anchor-aware Deep Metric Learning (AADML) method to address this challenge by uncovering the underlying correlations among existing data points, which enhances the quality of the shared embedding space. Specifically, our method establishes a correlation graph-based manifold structure by considering the dependencies between each sample as the anchor and its semantically similar samples. Through dynamic weighting of the correlations within this underlying manifold structure using an attention-driven mechanism, Anchor Awareness (AA) scores are obtained for each anchor. These AA scores serve as data proxies to compute relative distances in metric learning approaches. Extensive experiments conducted on two audio-visual benchmark datasets demonstrate the effectiveness of our proposed AADML method, significantly surpassing state-of-the-art models. Furthermore, we investigate the integration of AA proxies with various metric learning methods, further highlighting the efficacy of our approach.


【7】 Musical Word Embedding for Music Tagging and Retrieval
标题:用于音乐标记和检索的音乐词嵌入
链接:https://arxiv.org/abs/2404.13569
作者:SeungHeon Doh,Jongpil Lee,Dasaem Jeong,Juhan Nam
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing (TASLP)
摘要:None
摘要:Word embedding has become an essential means for text-based information retrieval. Typically, word embeddings are learned from large quantities of general and unstructured text data. However, in the domain of music, the word embedding may have difficulty understanding musical contexts or recognizing music-related entities like artists and tracks. To address this issue, we propose a new approach called Musical Word Embedding (MWE), which involves learning from various types of texts, including both everyday and music-related vocabulary. We integrate MWE into an audio-word joint representation framework for tagging and retrieving music, using words like tag, artist, and track that have different levels of musical specificity. Our experiments show that using a more specific musical word like track results in better retrieval performance, while using a less specific term like tag leads to better tagging performance. To balance this compromise, we suggest multi-prototype training that uses words with different levels of musical specificity jointly. We evaluate both word embedding and audio-word joint embedding on four tasks (tag rank prediction, music tagging, query-by-tag, and query-by-track) across two datasets (Million Song Dataset and MTG-Jamendo). Our findings show that the suggested MWE is more efficient and robust than the conventional word embedding.

【8】 Sparse Direction of Arrival Estimation Method Based on Vector Signal  Reconstruction with a Single Vector Sensor
标题:基于单载体传感器载体信号重建的稀疏波达方向估计方法
链接:https://arxiv.org/abs/2404.13568
作者:Jiabin Guo
备注:20 pages
摘要:本文研究了单矢量水听器在水声信号波达方向估计中的应用。针对传统DOA估计方法在多源环境和噪声干扰下的局限性,提出了一种矢量信号重构(VSR)技术。该方法通过复数运算和矢量信号重构,将单个矢量水听器信号的协方差矩阵变换为适合于无网格稀疏方法的Toeplitz结构。在此基础上,介绍了两种基于矢量信号重构的稀疏DOA估计算法。理论分析和仿真实验表明,与传统算法相比,本文算法在多源信号和低信噪比环境下显著提高了DOA估计的精度和分辨率。本研究的贡献在于为复杂环境下的单矢量水听器波达方向估计提供了一种有效的新方法,为矢量水听器信号处理领域引入了新的研究方向和解决方案。
摘要:This study investigates the application of single vector hydrophones in underwater acoustic signal processing for Direction of Arrival (DOA) estimation. Addressing the limitations of traditional DOA estimation methods in multi-source environments and under noise interference, this research proposes a Vector Signal Reconstruction (VSR) technique. This technique transforms the covariance matrix of single vector hydrophone signals into a Toeplitz structure suitable for gridless sparse methods through complex calculations and vector signal reconstruction. Furthermore, two sparse DOA estimation algorithms based on vector signal reconstruction are introduced. Theoretical analysis and simulation experiments demonstrate that the proposed algorithms significantly improve the accuracy and resolution of DOA estimation in multi-source signals and low Signal-to-Noise Ratio (SNR) environments compared to traditional algorithms. The contribution of this study lies in providing an effective new method for DOA estimation with single vector hydrophones in complex environments, introducing new research directions and solutions in the field of vector hydrophone signal processing.


【9】 AudioRepInceptionNeXt: A lightweight single-stream architecture for  efficient audio recognition
标题:AudioRepInceptionNeXt:用于高效音频识别的轻量级单流架构
链接:https://arxiv.org/abs/2404.13551
作者:Kin Wai Lau,Yasar Abbas Ur Rehman,Lai-Man Po
摘要:最近的研究已经成功地将基于视觉的卷积神经网络(CNN)架构用于使用Mel-频谱图的音频识别任务。然而,这些CNN具有较高的计算成本和内存需求,限制了它们在低端边缘设备上的部署。受InceptionNeXt和ConvNeXt等高效视觉模型的成功启发,我们提出了AudioRepInceptionNeXt,一种单流架构。它的基本构建块将具有k x k内核的递减尺度的并行多分支深度方向卷积分解为两个多分支深度方向卷积的级联。第一多分支由并行多尺度1 x k深度方向卷积层组成,随后是采用并行多尺度k x 1深度方向卷积层的类似多分支。这减少了计算和内存占用,同时分离了梅尔频谱图的时间和频率处理。大的内核捕获全局频率和长活动,而小的内核捕获局部频率和短活动。我们还在推理过程中重新参数化了多分支设计,以进一步提高速度而不损失准确性。实验表明,AudioRepInceptionNeXt将参数和计算量减少了50%以上,并将推理速度提高了1.28倍,同时保持了相当的准确性。它还可以在各种音频识别任务中进行鲁棒学习。代码可在https://github.com/StevenLauHKHK/AudioRepInceptionNeXt上获得。
摘要:Recent research has successfully adapted vision-based convolutional neural network (CNN) architectures for audio recognition tasks using Mel-Spectrograms. However, these CNNs have high computational costs and memory requirements, limiting their deployment on low-end edge devices. Motivated by the success of efficient vision models like InceptionNeXt and ConvNeXt, we propose AudioRepInceptionNeXt, a single-stream architecture. Its basic building block breaks down the parallel multi-branch depth-wise convolutions with descending scales of k x k kernels into a cascade of two multi-branch depth-wise convolutions. The first multi-branch consists of parallel multi-scale 1 x k depth-wise convolutional layers followed by a similar multi-branch employing parallel multi-scale k x 1 depth-wise convolutional layers. This reduces computational and memory footprint while separating time and frequency processing of Mel-Spectrograms. The large kernels capture global frequencies and long activities, while small kernels get local frequencies and short activities. We also reparameterize the multi-branch design during inference to further boost speed without losing accuracy. Experiments show that AudioRepInceptionNeXt reduces parameters and computations by 50%+ and improves inference speed 1.28x over state-of-the-art CNNs like the Slow-Fast while maintaining comparable accuracy. It also learns robustly across a variety of audio recognition tasks. Codes are available at https://github.com/StevenLauHKHK/AudioRepInceptionNeXt.


【10】 MFHCA: Enhancing Speech Emotion Recognition Via Multi-Spatial Fusion and  Hierarchical Cooperative Attention
标题:MFHCA:通过多空间融合和分层合作注意增强语音情感识别
链接:https://arxiv.org/abs/2404.13509
作者:Xinxin Jiao,Liejun Wang,Yinfeng Yu
备注:Main paper (5 pages). Accepted for publication by ICME 2024
摘要:语音情感识别在人机交互中至关重要,但从音频中提取和使用情感线索带来了挑战。本文介绍了MFHCA,一种新的方法,语音情感识别的多空间融合和层次合作注意力的频谱和原始音频。我们采用多空间融合模块(MF),以有效地识别情感相关的频谱图区域,并整合更高层次的声学信息的休伯特功能。我们的方法还包括一个层次合作注意模块(HCA)合并功能,从不同的听觉水平。在IEMOCAP数据集上的实验结果表明,该方法在加权精度和非加权精度上分别提高了2.6%和1.87%.大量的实验证明了该方法的有效性。
摘要:Speech emotion recognition is crucial in human-computer interaction, but extracting and using emotional cues from audio poses challenges. This paper introduces MFHCA, a novel method for Speech Emotion Recognition using Multi-Spatial Fusion and Hierarchical Cooperative Attention on spectrograms and raw audio. We employ the Multi-Spatial Fusion module (MF) to efficiently identify emotion-related spectrogram regions and integrate Hubert features for higher-level acoustic information. Our approach also includes a Hierarchical Cooperative Attention module (HCA) to merge features from various auditory levels. We evaluate our method on the IEMOCAP dataset and achieve 2.6\% and 1.87\% improvements on the weighted accuracy and unweighted accuracy, respectively. Extensive experiments demonstrate the effectiveness of the proposed method.

【11】 Text-dependent Speaker Verification (TdSV) Challenge 2024: Challenge  Evaluation Plan
标题:2024年文本相关说话者验证(TdSV)挑战:挑战评估计划
链接:https://arxiv.org/abs/2404.13428
作者:Zeinali Hossein,Lee Kong Aik,Alam Jahangir,Burget Lukas
摘要:本文概述了2024年文本相关说话人验证(TdSV)挑战赛,该挑战赛的重点是分析和探索文本相关说话人验证的新方法。这项挑战的主要目标是激励参与者开发单一但有竞争力的系统,进行彻底的分析,并探索创新的概念,如多任务学习,自我监督学习,Few-Shot学习等,用于文本相关的说话人验证。
摘要:This document outlines the Text-dependent Speaker Verification (TdSV) Challenge 2024, which centers on analyzing and exploring novel approaches for text-dependent speaker verification. The primary goal of this challenge is to motive participants to develop single yet competitive systems, conduct thorough analyses, and explore innovative concepts such as multi-task learning, self-supervised learning, few-shot learning, and others, for text-dependent speaker verification.

【12】 Interactive tools for making temporally variable, multiple-attributes,  and multiple-instances morphing accessible: Flexible manipulation of  divergent speech instances for explorational research and education
标题:用于使时间可变、多属性和多实例变形易于访问的交互工具:灵活操纵不同语音实例以进行探索性研究和教育
链接:https://arxiv.org/abs/2404.13418
作者:Hideki Kawahara,Masanori Morise
备注:5 pages, 7 figures, submitted to Acoustical Science and Technology of Acoustical Society of Japan
摘要:我们推广了一种语音变形算法,能够处理时间变量,多属性,多实例。广义Morphing为研究语音多样性提供了一种新的策略。然而,过度的复杂性和准备的困难使研究人员和学生无法享受其好处。为了解决这个问题,我们引入了一组交互式工具,使准备和测试不那么繁琐。这些工具集成到我们以前报告的交互式工具作为扩展。在研究生教育课程中引入扩展工具是成功的。最后,我们概述了进一步的扩展,以探索过于复杂的变形参数设置。
摘要:We generalized a voice morphing algorithm capable of handling temporally variable, multiple-attributes, and multiple instances. The generalized morphing provides a new strategy for investigating speech diversity. However, excessive complexity and the difficulty of preparation have prevented researchers and students from enjoying its benefits. To address this issue, we introduced a set of interactive tools to make preparation and tests less cumbersome. These tools are integrated into our previously reported interactive tools as extensions. The introduction of the extended tools in lessons in graduate education was successful. Finally, we outline further extensions to explore excessively complex morphing parameter settings.


【13】 Semantically Corrected Amharic Automatic Speech Recognition
标题:语义纠正的阿姆哈拉语自动语音识别
链接:https://arxiv.org/abs/2404.13362
作者:Samuael Adnew,Paul Pu Liang
摘要:自动语音识别(ASR)可以在提高全球口语的可访问性方面发挥关键作用。在本文中,我们为阿姆哈拉语构建了一套ASR工具,阿姆哈拉语是一种主要在东非有5000多万人使用的语言。阿姆哈拉语是用Geez文字书写的,这是一系列带有空格的字素,表示单词的边界。这使得阿姆哈拉语的计算处理具有挑战性,因为间距的位置可以显着影响所形成的句子的含义。我们发现,阿姆哈拉语ASR的现有基准不考虑这些间距,只测量单个字素错误率,导致显着膨胀的测量在野外的性能。在本文中,我们首先发布了现有阿姆哈拉语ASR测试数据集的修正translation,使社区能够准确评估进展。此外,我们引入了一个后处理方法,使用一个Transformer编码器-解码器架构,将原始ASR输出组织成一个语法上完整和语义上有意义的阿姆哈拉语句子。通过在修正后的测试数据集上的实验,我们的模型提高了阿姆哈拉语语音识别系统的语义正确率,实现了5.5%的字符错误率(CER)和23.3%的单词错误率(WER)。
摘要:Automatic Speech Recognition (ASR) can play a crucial role in enhancing the accessibility of spoken languages worldwide. In this paper, we build a set of ASR tools for Amharic, a language spoken by more than 50 million people primarily in eastern Africa. Amharic is written in the Ge'ez script, a sequence of graphemes with spacings denoting word boundaries. This makes computational processing of Amharic challenging since the location of spacings can significantly impact the meaning of formed sentences. We find that existing benchmarks for Amharic ASR do not account for these spacings and only measure individual grapheme error rates, leading to significantly inflated measurements of in-the-wild performance. In this paper, we first release corrected transcriptions of existing Amharic ASR test datasets, enabling the community to accurately evaluate progress. Furthermore, we introduce a post-processing approach using a transformer encoder-decoder architecture to organize raw ASR outputs into a grammatically complete and semantically meaningful Amharic sentence. Through experiments on the corrected test dataset, our model enhances the semantic correctness of Amharic speech recognition systems, achieving a Character Error Rate (CER) of 5.5\% and a Word Error Rate (WER) of 23.3\%.


【14】 Music Consistency Models
标题:音乐一致性模型
链接:https://arxiv.org/abs/2404.13358
作者:Zhengcong Fei,Mingyuan Fan,Junshi Huang
摘要:一致性模型在促进有效的图像/视频生成方面表现出显着的能力,从而能够以最少的采样步骤进行合成。它已被证明是有利的,在减轻与扩散模型相关的计算负担。然而,一致性模型在音乐生成中的应用在很大程度上仍然是未开发的。为了解决这个问题,我们提出了音乐一致性模型(\texttt{MusicCM}),它利用一致性模型的概念来有效地合成音乐片段的梅尔频谱图,保持高质量,同时最大限度地减少采样步骤的数量。建立在现有的文本到音乐的扩散模型,\texttt{MusicCM}模型结合了一致性蒸馏和对抗性训练。此外,我们发现它有利于产生扩展的连贯音乐,将多个扩散过程与共享的约束。实验结果表明,我们的模型在计算效率,保真度和自然度方面的有效性。值得注意的是,\texttt{MusicCM}仅用四个采样步骤就实现了无缝音乐合成,例如,每分钟只有一秒的音乐片段,展示了实时应用的潜力。
摘要:Consistency models have exhibited remarkable capabilities in facilitating efficient image/video generation, enabling synthesis with minimal sampling steps. It has proven to be advantageous in mitigating the computational burdens associated with diffusion models. Nevertheless, the application of consistency models in music generation remains largely unexplored. To address this gap, we present Music Consistency Models (\texttt{MusicCM}), which leverages the concept of consistency models to efficiently synthesize mel-spectrogram for music clips, maintaining high quality while minimizing the number of sampling steps. Building upon existing text-to-music diffusion models, the \texttt{MusicCM} model incorporates consistency distillation and adversarial discriminator training. Moreover, we find it beneficial to generate extended coherent music by incorporating multiple diffusion processes with shared constraints. Experimental results reveal the effectiveness of our model in terms of computational efficiency, fidelity, and naturalness. Notable, \texttt{MusicCM} achieves seamless music synthesis with a mere four sampling steps, e.g., only one second per minute of the music clip, showcasing the potential for real-time application.


【15】 Double Mixture: Towards Continual Event Detection from Speech
标题:双重混合:从语音中实现连续事件检测
链接:https://arxiv.org/abs/2404.13289
作者:Jingqi Kang,Tongtong Wu,Jinming Zhao,Guitao Wang,Yinwei Wei,Hao Yang,Guilin Qi,Yuan-Fang Li,Gholamreza Haffari
备注:The first two authors contributed equally to this work
摘要:语音事件检测是多媒体检索的关键,涉及语义和声学事件的标记。传统的ASR系统往往忽视这些事件之间的相互作用,只关注内容,即使对话的解释可能会因环境背景而异。本文解决了语音事件检测中的两个主要挑战:新事件的不断整合,而不会忘记以前的,从声学事件的语义的解开。我们引入了一个新的任务,从语音中连续事件检测,为此我们还提供了两个基准数据集。为了解决灾难性遗忘和有效解纠缠的挑战,我们提出了一种新的方法,“双重混合”。这种方法将语音专业知识与强大的记忆机制相结合,以增强适应性并防止遗忘。我们的综合实验表明,这项任务提出了重大挑战,目前最先进的计算机视觉或自然语言处理方法无法有效解决这些挑战。我们的方法实现了最低的遗忘率和最高水平的泛化,证明了在各种持续学习序列中的鲁棒性。我们的代码和数据可在https://anonymous.4open.science/status/Continual-SpeechED-6461上获得。
摘要:Speech event detection is crucial for multimedia retrieval, involving the tagging of both semantic and acoustic events. Traditional ASR systems often overlook the interplay between these events, focusing solely on content, even though the interpretation of dialogue can vary with environmental context. This paper tackles two primary challenges in speech event detection: the continual integration of new events without forgetting previous ones, and the disentanglement of semantic from acoustic events. We introduce a new task, continual event detection from speech, for which we also provide two benchmark datasets. To address the challenges of catastrophic forgetting and effective disentanglement, we propose a novel method, 'Double Mixture.' This method merges speech expertise with robust memory mechanisms to enhance adaptability and prevent forgetting. Our comprehensive experiments show that this task presents significant challenges that are not effectively addressed by current state-of-the-art methods in either computer vision or natural language processing. Our approach achieves the lowest rates of forgetting and the highest levels of generalization, proving robust across various continual learning sequences. Our code and data are available at https://anonymous.4open.science/status/Continual-SpeechED-6461.


【16】 Track Role Prediction of Single-Instrumental Sequences
标题:单乐器序列的音轨角色预测
链接:https://arxiv.org/abs/2404.13286
作者:Changheon Han,Suhyun Lee,Minsam Ko
备注:ISMIR LBD 2023
摘要:在创作过程中,选择合适的单器乐序列并分配其轨道角色是一项不可或缺的工作。然而,手动确定大量音乐样本的音轨角色可能是耗时且劳动密集的。这项研究介绍了一种深度学习模型,旨在自动预测单乐器音乐序列的轨道角色。我们的评估显示,在符号域和音频域的预测准确率分别为87%和84%。所提出的轨道角色预测方法有望在人工智能音乐生成和分析中得到应用。
摘要:In the composition process, selecting appropriate single-instrumental music sequences and assigning their track-role is an indispensable task. However, manually determining the track-role for a myriad of music samples can be time-consuming and labor-intensive. This study introduces a deep learning model designed to automatically predict the track-role of single-instrumental music sequences. Our evaluations show a prediction accuracy of 87% in the symbolic domain and 84% in the audio domain. The proposed track-role prediction methods hold promise for future applications in AI music generation and analysis.


【17】 Intro to Quantum Harmony: Chords in Superposition
标题:量子和谐简介:叠加的和弦
链接:https://arxiv.org/abs/2404.13140
作者:Christopher Dobrian,Omar Costa Hamido
摘要:量子理论和音乐理论之间的相关性-特别是量子计算原理和音乐和声之间的相关性-可以为音乐理论家和作曲家带来新的理解和新的方法。量子叠加原理被证明与音乐意义的不同解释密切相关。叠加直接在作者的量子计算模拟中实现,应用于计算机生成音乐创作的决策过程。
摘要:Correlations between quantum theory and music theory - specifically between principles of quantum computing and musical harmony - can lead to new understandings and new methodologies for music theorists and composers. The quantum principle of superposition is shown to be closely related to different interpretations of musical meaning. Superposition is implemented directly in the authors' simulations of quantum computing, as applied in the decision-making processes of computer-generated music composition.

机器翻译由腾讯交互翻译提供,仅供参考