今日论文合集:cs.SD语音13篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language
标题: 推进尼泊尔语音克隆:利用低资源语言的迁移学习
作者:Manjil Karki,Pratik Shakya,Sandesh Acharya,Ravi Pandit,Dinesh Gothe
备注:7 pages, 10 figures
链接:点击下载PDF文件
摘要:语音克隆是个性化语音界面的一个重要特征。一个神经声音克隆系统可以只用几个音频样本来模仿某人的声音。说话人编码和说话人自适应都是语音克隆领域的研究课题。说话人自适应依赖于微调多说话人生成模型,其涉及训练单独的模型以推断用于说话人编码的新说话人嵌入。这两种方法都可以实现良好的性能,即使有少量的克隆音频,在语音的自然度和相似性的原始扬声器。扬声器编码方法更适合于低资源部署,因为它们需要的内存显着减少,并且比扬声器自适应具有更快的克隆时间,可以提供更大的自然度和相似性。主要目标是创建一个语音克隆系统,产生带有尼泊尔口音或听起来像尼泊尔语的音频输出。为了进一步推进TTS,迁移学习的思想被有效地用于解决该系统开发过程中遇到的几个问题,包括音频质量差和缺乏可用数据。摘要:Voice cloning is a prominent feature in personalized speech interfaces. A neural vocal cloning system can mimic someone's voice using just a few audio samples. Both speaker encoding and speaker adaptation are topics of research in the field of voice cloning. Speaker adaptation relies on fine-tuning a multi-speaker generative model, which involves training a separate model to infer a new speaker embedding used for speaker encoding. Both methods can achieve excellent performance, even with a small number of cloning audios, in terms of the speech's naturalness and similarity to the original speaker. Speaker encoding approaches are more appropriate for low-resource deployment since they require significantly less memory and have a faster cloning time than speaker adaption, which can offer slightly greater naturalness and similarity. The main goal is to create a vocal cloning system that produces audio output with a Nepali accent or that sounds like Nepali. For the further advancement of TTS, the idea of transfer learning was effectively used to address several issues that were encountered in the development of this system, including the poor audio quality and the lack of available data.

【2】 Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
标题: 转换和说话:Zero-Shot口音转换,最少监督
作者:Zhijun Jia,Huaying Xue,Xiulian Peng,Yan Lu
备注:9 pages, 4 figures, conference
链接:点击下载PDF文件
摘要:重音转换是一个语音单元和韵律模式都需要转换的问题,其关键是并行数据资源少。我们提出了一个两阶段的生成框架“转换和说话”,其中转换只操作的语义标记水平和语音合成的条件下,转换后的语义标记与语音生成模型在目标口音域。解耦设计使“说话”模块能够使用大量的目标口音语音,并减轻了“转换”模块所需的并行数据。以语义标记为桥梁的转换也减轻了对文本转换数据的需求,并解锁了语言预训练技术的使用,以进一步有效地减少对并行口音语音数据的需求。为了降低“说话”的复杂度和延迟,设计了一个单级AR生成模型,以实现良好的质量以及较低的计算成本。对印度英语到一般美国英语转换的实验表明,该框架实现了最先进的性能,口音相似性,语音质量和扬声器维护只有15分钟的弱并行数据,不限于同一个扬声器。对不同口音类型的广泛实验表明,该框架具有高度的适应性,使其易于扩展,以适应其他口音与低资源数据。音频样本可在https: www.microsoft.com en-us research project convert-and-speak-zero-shot-accent-conversion-with-minimumsupervision 上获得。摘要:Low resource of parallel data is the key challenge of accent conversion(AC) problem in which both the pronunciation units and prosody pattern need to be converted. We propose a two-stage generative framework "convert-and-speak" in which the conversion is only operated on the semantic token level and the speech is synthesized conditioned on the converted semantic token with a speech generative model in target accent domain. The decoupling design enables the "speaking" module to use massive amount of target accent speech and relieves the parallel data required for the "conversion" module. Conversion with the bridge of semantic token also relieves the requirement for the data with text transcriptions and unlocks the usage of language pre-training technology to further efficiently reduce the need of parallel accent speech data. To reduce the complexity and latency of "speaking", a single-stage AR generative model is designed to achieve good quality as well as lower computation cost. Experiments on Indian-English to general American-English conversion show that the proposed framework achieves state-of-the-art performance in accent similarity, speech quality, and speaker maintenance with only 15 minutes of weakly parallel data which is not constrained to the same speaker. Extensive experimentation with diverse accent types suggests that this framework possesses a high degree of adaptability, making it readily scalable to accommodate other accents with low-resource data. Audio samples are available at https: www.microsoft.com en-us research project convert-and-speak-zero-shot-accent-conversion-with-minimumsupervision .

【3】 SZU-AFS Antispoofing System for the ASVspoof 5 Challenge
标题: ASVspoof 5挑战赛的SZU-ATL反恶搞系统
作者:Yuxiong Xu,Jiafeng Zhong,Sengui Zheng,Zefeng Liu,Bin Li
备注:8 pages, 2 figures, ASVspoof 5 Workshop (Interspeech2024 Satellite)
链接:点击下载PDF文件
摘要:本文介绍了SZU-AFS反欺骗系统,在开放条件下的ASVspoof 5挑战赛的轨道1设计。该系统分为四个阶段:选择基线模型,探索有效的数据增强(DA)方法进行微调,应用基于梯度范数感知最小化(GAM)的协同增强策略进行二次微调,并融合两个性能最佳的微调模型的logits分数。该系统利用Wav 2 Vec 2前端特征提取器和AASIST后端分类器作为基线模型。在模型的微调,三个不同的DA政策进行了研究:单DA,随机DA,级联DA。此外,采用基于GAM的协同增强策略,旨在微调的增强模型在数据和优化器的水平,帮助亚当优化器找到平坦的最小值,从而提高模型的泛化。总体而言,最终融合系统在评估集中实现了0.115的minDCF和4.04%的EER。摘要:This paper presents the SZU-AFS anti-spoofing system, designed for Track 1 of the ASVspoof 5 Challenge under open conditions. The system is built with four stages: selecting a baseline model, exploring effective data augmentation (DA) methods for fine-tuning, applying a co-enhancement strategy based on gradient norm aware minimization (GAM) for secondary fine-tuning, and fusing logits scores from the two best-performing fine-tuned models. The system utilizes the Wav2Vec2 front-end feature extractor and the AASIST back-end classifier as the baseline model. During model fine-tuning, three distinct DA policies have been investigated: single-DA, random-DA, and cascade-DA. Moreover, the employed GAM-based co-enhancement strategy, designed to fine-tune the augmented model at both data and optimizer levels, helps the Adam optimizer find flatter minima, thereby boosting model generalization. Overall, the final fusion system achieves a minDCF of 0.115 and an EER of 4.04% on the evaluation set.

【4】 Hear Your Face: Face-based voice conversion with F0 estimation
标题: 听到你的脸:基于脸的语音转换,具有F0估计
作者:Jaejun Lee,Yoori Oh,Injune Hwang,Kyogu Lee
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:本文深入研究了新兴的基于人脸的语音转换领域,利用个人的面部特征和他们的声音特征之间的独特关系。我们提出了一种新的基于人脸的语音转换框架,特别是利用目标扬声器的平均基频,完全来自他们的面部图像。通过广泛的分析,我们的框架展示了卓越的语音生成质量和将面部特征与语音特征对齐的能力,包括跟踪目标说话者的基频。摘要:This paper delves into the emerging field of face-based voice conversion, leveraging the unique relationship between an individual's facial features and their vocal characteristics. We present a novel face-based voice conversion framework that particularly utilizes the average fundamental frequency of the target speaker, derived solely from their facial images. Through extensive analysis, our framework demonstrates superior speech generation quality and the ability to align facial features with voice characteristics, including tracking of the target speaker's fundamental frequency.

【5】 Unsupervised Composable Representations for Audio
标题: 音频的无监督可组合表示
作者:Giovanni Bindi,Philippe Esling
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:目前的生成模型能够生成高质量的人工制品,但已被证明与组合推理,这可以被定义为从简单的元素生成复杂结构的能力斗争。在本文中,我们专注于音乐数据的组成表示学习问题,特别是针对完全无监督的设置。我们提出了一个简单且可扩展的框架,该框架利用了一个明确的组合归纳偏差,该偏差由一个灵活的自动编码目标定义,该目标可以利用任何当前最先进的生成模型。我们证明了我们的框架,与扩散模型一起使用,自然地解决了无监督音频源分离的任务,表明我们的模型能够执行高质量的分离。我们的研究结果表明,我们的建议实现了可比或优于其他盲源分离方法的性能,此外,它甚至超过了目前最先进的监督基线的信号干扰比指标。此外,通过学习一个后验掩蔽扩散模型在空间的组合表示,我们实现了一个系统能够无缝执行无监督源分离,无条件生成和变化生成。最后,由于我们的建议在预训练的神经音频编解码器的潜在空间中工作,因此相对于其他神经基线,它还提供了更低的计算成本。摘要:Current generative models are able to generate high-quality artefacts but have been shown to struggle with compositional reasoning, which can be defined as the ability to generate complex structures from simpler elements. In this paper, we focus on the problem of compositional representation learning for music data, specifically targeting the fully-unsupervised setting. We propose a simple and extensible framework that leverages an explicit compositional inductive bias, defined by a flexible auto-encoding objective that can leverage any of the current state-of-art generative models. We demonstrate that our framework, used with diffusion models, naturally addresses the task of unsupervised audio source separation, showing that our model is able to perform high-quality separation. Our findings reveal that our proposal achieves comparable or superior performance with respect to other blind source separation methods and, furthermore, it even surpasses current state-of-art supervised baselines on signal-to-interference ratio metrics. Additionally, by learning an a-posteriori masking diffusion model in the space of composable representations, we achieve a system capable of seamlessly performing unsupervised source separation, unconditional generation, and variation generation. Finally, as our proposal works in the latent space of pre-trained neural audio codecs, it also provides a lower computational cost with respect to other neural baselines.

【6】 A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
标题: 基于转录预算的高效音频大语言模型用于鲁棒语音识别
作者:Yangze Li,Xiong Wang,Songjun Cao,Yike Zhang,Long Ma,Lei Xie
链接:点击下载PDF文件
摘要:Audio-LLM将音频模态引入大型语言模型(LLM),使强大的LLM能够识别、理解和生成音频。然而,在嘈杂环境中的语音识别过程中,我们观察到音频LLM中存在错觉和重复问题,导致替换和插入错误。针对上述问题,提出了一种基于语音识别专家的音频LLM算法,该算法引入语音识别专家作为语音识别标记器,并采用混合自回归(AR)和非自回归(NAR)解码方法。在10 kh的WenetSpeech中文语料上的实验表明,该方法在Test_Net和Test_Meeting两个评测集上的CER分别比基线降低了12.2%和9.6%。值得注意的是,我们降低了评估集上的解码重复率为零,这表明解码重复问题已从根本上得到解决。摘要:Audio-LLM introduces audio modality into a large language model (LLM) to enable a powerful LLM to recognize, understand, and generate audio. However, during speech recognition in noisy environments, we observed the presence of illusions and repetition issues in audio-LLM, leading to substitution and insertion errors. This paper proposes a transcription prompt-based audio-LLM by introducing an ASR expert as a transcription tokenizer and a hybrid Autoregressive (AR) Non-autoregressive (NAR) decoding approach to solve the above problems. Experiments on 10k-hour WenetSpeech Mandarin corpus show that our approach decreases 12.2% and 9.6% CER relatively on Test_Net and Test_Meeting evaluation sets compared with baseline. Notably, we reduce the decoding repetition rate on the evaluation set to zero, showing that the decoding repetition problem has been solved fundamentally.

【7】 Enhancing Modal Fusion by Alignment and Label Matching for Multimodal Emotion Recognition
标题: 通过对齐和标签匹配增强多模式情感识别的模式融合
作者:Qifei Li,Yingming Gao,Yuhua Wen,Cong Wang,Ya Li
备注:The paper has been accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:为了解决跨模态信息融合引起的多模态情感识别(MER)性能的限制,我们提出了一种基于多任务学习的新型MER框架,其中融合在对齐后发生,称为Foal-Net。该框架旨在提高模态融合的有效性,并包括两个辅助任务:音视频情感对齐(AVEL)和跨模态情感标签匹配(MEM)。首先,AVEL通过对比学习实现音频-视频表示中的情感信息的对齐。然后,模态融合网络集成对齐的特征。同时,MEM评估当前样本对的情感是否相同,为模态信息融合提供帮助,引导模型更多地关注情感信息。在IEMOCAP语料库上的实验结果表明,Foal-Net的性能优于现有的方法,并且在模态融合之前必须进行情感对齐。摘要:To address the limitation in multimodal emotion recognition (MER) performance arising from inter-modal information fusion, we propose a novel MER framework based on multitask learning where fusion occurs after alignment, called Foal-Net. The framework is designed to enhance the effectiveness of modality fusion and includes two auxiliary tasks: audio-video emotion alignment (AVEL) and cross-modal emotion label matching (MEM). First, AVEL achieves alignment of emotional information in audio-video representations through contrastive learning. Then, a modal fusion network integrates the aligned features. Meanwhile, MEM assesses whether the emotions of the current sample pair are the same, providing assistance for modal information fusion and guiding the model to focus more on emotional information. The experimental results conducted on IEMOCAP corpus show that Foal-Net outperforms the state-of-the-art methods and emotion alignment is necessary before modal fusion.

【8】 Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
标题: 元学习赋予元面孔:个性化说话风格改编,用于音频驱动的3D会说话的面孔动画
作者:Xukun Zhou,Fengxin Li,Ziqiao Peng,Kejian Wu,Jun He,Biao Qin,Zhaoxin Fan,Hongyan Liu
链接:点击下载PDF文件
摘要:音频驱动的3D人脸动画在实时流媒体和增强现实应用中越来越重要。虽然已经观察到了显着的进展,大多数现有的方法是专为特定的个人与预定义的说话风格,从而忽略了不同的说话风格的适应性。为了解决这个问题,本文介绍了MetaFace,一种新的方法精心制作的讲话风格适应。基于元学习的新概念,MetaFace由几个关键组件组成:用于基本说话风格适应的鲁棒Meta学习阶段(RMIS),用于在观察到的和未观察到的说话风格之间建立联系的动态关系挖掘神经过程(DRMN),以及用于提高模型优化效率和学习风格细节的低秩矩阵内存缩减方法。利用这些新颖的设计,MetaFace不仅显著优于现有的强大基线,而且还建立了一个新的最先进的,正如我们的实验结果所证实的那样。摘要:Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with predefined speaking styles, thus neglecting the adaptability to varied speaking styles. To address this limitation, this paper introduces MetaFace, a novel methodology meticulously crafted for speaking style adaptation. Grounded in the novel concept of meta-learning, MetaFace is composed of several key components: the Robust Meta Initialization Stage (RMIS) for fundamental speaking style adaptation, the Dynamic Relation Mining Neural Process (DRMN) for forging connections between observed and unobserved speaking styles, and the Low-rank Matrix Memory Reduction Approach to enhance the efficiency of model optimization as well as learning style details. Leveraging these novel designs, MetaFace not only significantly outperforms robust existing baselines but also establishes a new state-of-the-art, as substantiated by our experimental results.

【9】 Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
标题: 优化:延展实境的空间音频线索的最佳放置
作者:Hyunsung Cho,Alexander Wang,Divya Kartik,Emily Liying Xie,Yukang Yan,David Lindlbauer
备注:UIST 2024
链接:点击下载PDF文件
摘要:延展实境(XR)中的空间音频为用户提供了更好的虚拟元素放置位置意识,并有效地引导他们的事件,如通知,来自不同窗口的系统警报,或接近化身。然而,人类在定位声音线索方面并不准确,尤其是在多个来源的情况下,这是由于人类听觉感知的局限性,例如角度辨别误差和前后混淆。这会降低XR界面的效率,因为用户会错误识别声音来自哪个XR元素。为了解决这个问题,我们提出了Auptimize,一种新的计算方法,用于放置XR声源,通过利用口技效应来减轻这种定位误差。Auptimize将声源位置从视觉元素中分离出来,并将声源重新定位到最佳位置,以明确识别声音提示,避免由于源间接近和前后混淆而导致的错误。我们的评估表明,Auptimize减少空间音频为基础的源识别错误相比,在成对的视觉-声音位置播放声音线索。我们证明了Auptimize对于各种基于空间音频的交互式XR场景的适用性。摘要:Spatial audio in Extended Reality (XR) provides users with better awareness of where virtual elements are placed, and efficiently guides them to events such as notifications, system alerts from different windows, or approaching avatars. Humans, however, are inaccurate in localizing sound cues, especially with multiple sources due to limitations in human auditory perception such as angular discrimination error and front-back confusion. This decreases the efficiency of XR interfaces because users misidentify from which XR element a sound is coming. To address this, we propose Auptimize, a novel computational approach for placing XR sound sources, which mitigates such localization errors by utilizing the ventriloquist effect. Auptimize disentangles the sound source locations from the visual elements and relocates the sound sources to optimal positions for unambiguous identification of sound cues, avoiding errors due to inter-source proximity and front-back confusion. Our evaluation shows that Auptimize decreases spatial audio-based source identification errors compared to playing sound cues at the paired visual-sound locations. We demonstrate the applicability of Auptimize for diverse spatial audio-based interactive XR scenarios.

【10】 Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
标题: 通过文本-音频对的自我监督后训练增强音频语言模型
作者:Anshuman Sinha,Camille Migozzi,Aubin Rey,Chao Zhang
备注:31 pages, 11 figures
链接:点击下载PDF文件
摘要:音频和文本的多模态对比学习策略的研究迅速引起了人们的兴趣。对比训练的音频语言模型(ALM),如CLAP,建立了跨音频和语言模态的统一表示,通过提供良好的文本对齐音频编码器,提高了各种后续任务的效率,反之亦然。这些改进在诸如zero-shot音频分类和音频检索等领域中是明显的。然而,这些模型理解自然语言和时间关系的能力在很大程度上仍然是一个未开发和开放的研究领域。在本文中,我们建议配备的多模态ALMs的时间理解,而不失去其固有的先前能力的音频语言任务的时间灌输方法TeminAL。我们实现了一个两阶段的训练方案TeminAL A $ $ B,其中模型首先学习区分TeminAL A中的多个声音,然后是一个灌输时间感的阶段,从而增强其在TeminAL B中的时间理解。这种方法导致在ESC-50数据集上的时间理解的平均性能增益为5.28 %$,而该模型在AudioCap Clotho数据集上的zero-shot检索和分类任务中仍然具有竞争力。我们还注意到,缺乏适当的评价技术,对比ALM,并提出了一个战略评估ALM在zero-shot设置。通用的zero-shot模型评估策略ZSTE,用于评估各种先验模型。ZSTE展示了一个评估所有ZS对比模型的通用策略。使用TeminAL训练的模型在大多数下游任务上成功地优于当前模型。摘要:Research on multi-modal contrastive learning strategies for audio and text has rapidly gained interest. Contrastively trained Audio-Language Models (ALMs), such as CLAP, which establish a unified representation across audio and language modalities, have enhanced the efficacy in various subsequent tasks by providing good text aligned audio encoders and vice versa. These improvements are evident in areas like zero-shot audio classification and audio retrieval, among others. However, the ability of these models to understand natural language and temporal relations is still a largely unexplored and open field for research. In this paper, we propose to equip the multi-modal ALMs with temporal understanding without loosing their inherent prior capabilities of audio-language tasks with a temporal instillation method TeminAL. We implement a two-stage training scheme TeminAL A $ &$ B, where the model first learns to differentiate between multiple sounds in TeminAL A, followed by a phase that instills a sense of time, thereby enhancing its temporal understanding in TeminAL B. This approach results in an average performance gain of $5.28 %$ in temporal understanding on the ESC-50 dataset, while the model remains competitive in zero-shot retrieval and classification tasks on the AudioCap Clotho datasets. We also note the lack of proper evaluation techniques for contrastive ALMs and propose a strategy for evaluating ALMs in zero-shot settings. The general-purpose zero-shot model evaluation strategy ZSTE, is used to evaluate various prior models. ZSTE demonstrates a general strategy to evaluate all ZS contrastive models. The model trained with TeminAL successfully outperforms current models on most downstream tasks.

【11】 Efficient Autoregressive Audio Modeling via Next-Scale Prediction
标题: 通过下一规模预测的高效自回归音频建模
作者:Kai Qiu,Xiang Li,Hao Chen,Jie Sun,Jinglu Wang,Zhe Lin,Marios Savvides,Bhiksha Raj
备注:7 pages, 6 figures, 7 tables
链接:点击下载PDF文件
摘要:随着扩散模型(DM)和自回归(AR)模型等复杂生成模型的发展,音频生成技术取得了显著的进步。然而,由于音频的自然显著序列长度,音频生成的效率仍然是一个需要解决的重要问题,特别是对于合并在大型语言模型(LLM)中的AR模型。本文分析了音频标记化的标记长度,提出了一种新的 textbf{S}标度级 textbf{A}音频 textbf{T}标记器(SAT),改进了残差量化。在SAT的基础上,进一步提出了尺度级的 textbf{A}自回归 textbf {A}uto textbf{R}渐进(AAR)建模框架,将下一个标记的AR预测转移到下一个尺度的AR预测,显著降低了训练成本和推理时间。为了验证所提出的方法的有效性,我们全面分析了设计选择,并证明了所提出的AAR框架实现了显着的 textbf{35}$ times$更快的推理速度和+ textbf{1.33} Fr 'echet音频距离(FAD)对AudioSet基准的基线。代码: url{https: github.com qiuk2 AAR}。摘要:Audio generation has achieved remarkable progress with the advance of sophisticated generative models, such as diffusion models (DMs) and autoregressive (AR) models. However, due to the naturally significant sequence length of audio, the efficiency of audio generation remains an essential issue to be addressed, especially for AR models that are incorporated in large language models (LLMs). In this paper, we analyze the token length of audio tokenization and propose a novel textbf{S}cale-level textbf{A}udio textbf{T}okenizer (SAT), with improved residual quantization. Based on SAT, a scale-level textbf{A}coustic textbf{A}uto textbf{R}egressive (AAR) modeling framework is further proposed, which shifts the next-token AR prediction to next-scale AR prediction, significantly reducing the training cost and inference time. To validate the effectiveness of the proposed approach, we comprehensively analyze design choices and demonstrate the proposed AAR framework achieves a remarkable textbf{35}$ times$ faster inference speed and + textbf{1.33} Fr 'echet Audio Distance (FAD) against baselines on the AudioSet benchmark. Code: url{https: github.com qiuk2 AAR}.

【12】 Efficient Area-based and Speaker-Agnostic Source Separation
标题: 高效的基于区域和说话者不可知的源分离
作者:Martin Strauss,Okan Köpüklü
备注:Preprint. Accepted to the International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
摘要:本文介绍了一种基于区域的虚拟会议场景的源分离方法。其目的是保留来自线性麦克风阵列前面的定义空间区域内的未指定数量的源的语音信号,同时抑制所有其他声音。因此,我们采用了一个高效的神经网络架构,适用于多通道输入,以涵盖预定义的目标区域。为了评估该方法,模拟了训练数据和特定的测试场景,包括多个目标和干扰扬声器,以及背景噪声。所有型号都根据DNSMOS和标度不变信号失真比进行评级。我们的实验表明,所提出的方法分离语音从多个扬声器内的目标区域以及,除了是非常低的复杂度,用于实时处理。此外,功率降低热图用于展示网络识别位于目标区域内的源的能力。我们把我们的方法的背景下,一个完善的基线扬声器扬声器分离,并讨论其优势和挑战。摘要:This paper introduces an area-based source separation method designed for virtual meeting scenarios. The aim is to preserve speech signals from an unspecified number of sources within a defined spatial area in front of a linear microphone array, while suppressing all other sounds. Therefore, we employ an efficient neural network architecture adapted for multi-channel input to encompass the predefined target area. To evaluate the approach, training data and specific test scenarios including multiple target and interfering speakers, as well as background noise are simulated. All models are rated according to DNSMOS and scale-invariant signal-to-distortion ratio. Our experiments show that the proposed method separates speech from multiple speakers within the target area well, besides being of very low complexity, intended for real-time processing. In addition, a power reduction heatmap is used to demonstrate the networks' ability to identify sources located within the target area. We put our approach in context with a well-established baseline for speaker-speaker separation and discuss its strengths and challenges.

【13】 Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
标题: 使用文本到语音和大语言模型生成数据以进行对话语音识别
作者:Samuele Cornell,Jordan Darefsky,Zhiyao Duan,Shinji Watanabe
备注:To appear at SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:目前,许多语音处理任务中的一种常见方法是通过针对特定应用的域内数据对大规模预训练模型进行微调来利用它们。然而,由于隐私问题和注释成本,即使获得少量这样的数据也可能是有问题的,特别是对于敏感领域和会话语音场景。为了解决这个问题,已经采用了使用单个说话者数据集的合成数据生成。然而,对于多说话者的情况,这种方法通常需要大量的手动工作,并且容易出现域不匹配。在这项工作中,我们提出了一个用于多说话者会话ASR的合成数据生成管道,利用大型语言模型(LLM)进行内容创建,并利用会话多说话者文本到语音(TTS)模型进行语音合成。我们使用域内数据和生成的合成数据,通过微调电话和远程会话语音设置的Whisper ASR模型进行评估。我们的研究结果表明,该方法能够显着优于经典的多说话人生成方法,使用外部,非会话语音数据集。摘要:Currently, a common approach in many speech processing tasks is to leverage large scale pre-trained models by fine-tuning them on in-domain data for a particular application. Yet obtaining even a small amount of such data can be problematic, especially for sensitive domains and conversational speech scenarios, due to both privacy issues and annotation costs. To address this, synthetic data generation using single speaker datasets has been employed. Yet, for multi-speaker cases, such an approach often requires extensive manual effort and is prone to domain mismatches. In this work, we propose a synthetic data generation pipeline for multi-speaker conversational ASR, leveraging a large language model (LLM) for content creation and a conversational multi-speaker text-to-speech (TTS) model for speech synthesis. We conduct evaluation by fine-tuning the Whisper ASR model for telephone and distant conversational speech settings, using both in-domain data and generated synthetic data. Our results show that the proposed method is able to significantly outperform classical multi-speaker generation approaches that use external, non-conversational speech datasets.


eess.AS音频处理
【1】 Efficient Area-based and Speaker-Agnostic Source Separation
标题: 高效的基于区域和说话者不可知的源分离
作者:Martin Strauss,Okan Köpüklü
备注:Preprint. Accepted to the International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
摘要:本文介绍了一种基于区域的虚拟会议场景的源分离方法。其目的是保留来自线性麦克风阵列前面的定义空间区域内的未指定数量的源的语音信号,同时抑制所有其他声音。因此,我们采用了一个高效的神经网络架构,适用于多通道输入,以涵盖预定义的目标区域。为了评估该方法,模拟了训练数据和特定的测试场景,包括多个目标和干扰扬声器,以及背景噪声。所有型号都根据DNSMOS和标度不变信号失真比进行评级。我们的实验表明,所提出的方法分离语音从多个扬声器内的目标区域以及,除了是非常低的复杂度,用于实时处理。此外,功率降低热图用于展示网络识别位于目标区域内的源的能力。我们把我们的方法的背景下,一个完善的基线扬声器扬声器分离,并讨论其优势和挑战。摘要:This paper introduces an area-based source separation method designed for virtual meeting scenarios. The aim is to preserve speech signals from an unspecified number of sources within a defined spatial area in front of a linear microphone array, while suppressing all other sounds. Therefore, we employ an efficient neural network architecture adapted for multi-channel input to encompass the predefined target area. To evaluate the approach, training data and specific test scenarios including multiple target and interfering speakers, as well as background noise are simulated. All models are rated according to DNSMOS and scale-invariant signal-to-distortion ratio. Our experiments show that the proposed method separates speech from multiple speakers within the target area well, besides being of very low complexity, intended for real-time processing. In addition, a power reduction heatmap is used to demonstrate the networks' ability to identify sources located within the target area. We put our approach in context with a well-established baseline for speaker-speaker separation and discuss its strengths and challenges.

【2】 Malacopula: adversarial automatic speaker verification attacks using a neural-based generalised Hammerstein model
标题: Malacopula:使用基于神经的广义Hammerstein模型的对抗性自动说话人验证攻击
作者:Massimiliano Todisco,Michele Panariello,Xin Wang,Héctor Delgado,Kong Aik Lee,Nicholas Evans
备注:Accepted at ASVspoof Workshop 2024
链接:点击下载PDF文件
摘要:我们提出了Malacopula,一个基于神经的广义Hammerstein模型,旨在引入对抗性扰动的欺骗性语音话语,使他们更好地欺骗自动说话人验证(ASV)系统。Malacopula使用非线性过程来修改语音,增强了欺骗攻击的有效性。该模型包括多项式函数的并行分支,然后是线性时不变滤波器。对抗优化过程用于最小化从欺骗和真实话语中提取的说话人嵌入之间的余弦距离。使用三个最近的ASV系统和ASVspoof 2019数据集进行的实验表明,Malacopula大幅增加了漏洞。然而,语音质量降低,攻击可以有效地检测在受控条件下。研究结果强调了识别新漏洞和设计防御措施的必要性,以保护ASV系统免受野外对抗性攻击。摘要:We present Malacopula, a neural-based generalised Hammerstein model designed to introduce adversarial perturbations to spoofed speech utterances so that they better deceive automatic speaker verification (ASV) systems. Using non-linear processes to modify speech utterances, Malacopula enhances the effectiveness of spoofing attacks. The model comprises parallel branches of polynomial functions followed by linear time-invariant filters. The adversarial optimisation procedure acts to minimise the cosine distance between speaker embeddings extracted from spoofed and bona fide utterances. Experiments, performed using three recent ASV systems and the ASVspoof 2019 dataset, show that Malacopula increases vulnerabilities by a substantial margin. However, speech quality is reduced and attacks can be detected effectively under controlled conditions. The findings emphasise the need to identify new vulnerabilities and design defences to protect ASV systems from adversarial attacks in the wild.

【3】 Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
标题: 使用文本到语音和大语言模型生成数据以进行对话语音识别
作者:Samuele Cornell,Jordan Darefsky,Zhiyao Duan,Shinji Watanabe
备注:To appear at SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:目前,许多语音处理任务中的一种常见方法是通过针对特定应用的域内数据对大规模预训练模型进行微调来利用它们。然而,由于隐私问题和注释成本,即使获得少量这样的数据也可能是有问题的,特别是对于敏感领域和会话语音场景。为了解决这个问题,已经采用了使用单个说话者数据集的合成数据生成。然而,对于多说话者的情况,这种方法通常需要大量的手动工作,并且容易出现域不匹配。在这项工作中,我们提出了一个用于多说话者会话ASR的合成数据生成管道,利用大型语言模型(LLM)进行内容创建,并利用会话多说话者文本到语音(TTS)模型进行语音合成。我们使用域内数据和生成的合成数据,通过微调电话和远程会话语音设置的Whisper ASR模型进行评估。我们的研究结果表明,该方法能够显着优于经典的多说话人生成方法,使用外部,非会话语音数据集。摘要:Currently, a common approach in many speech processing tasks is to leverage large scale pre-trained models by fine-tuning them on in-domain data for a particular application. Yet obtaining even a small amount of such data can be problematic, especially for sensitive domains and conversational speech scenarios, due to both privacy issues and annotation costs. To address this, synthetic data generation using single speaker datasets has been employed. Yet, for multi-speaker cases, such an approach often requires extensive manual effort and is prone to domain mismatches. In this work, we propose a synthetic data generation pipeline for multi-speaker conversational ASR, leveraging a large language model (LLM) for content creation and a conversational multi-speaker text-to-speech (TTS) model for speech synthesis. We conduct evaluation by fine-tuning the Whisper ASR model for telephone and distant conversational speech settings, using both in-domain data and generated synthetic data. Our results show that the proposed method is able to significantly outperform classical multi-speaker generation approaches that use external, non-conversational speech datasets.

【4】 Advancing Voice Cloning for Nepali: Leveraging Transfer Learning in a Low-Resource Language
标题: 推进尼泊尔语音克隆:利用低资源语言的迁移学习
作者:Manjil Karki,Pratik Shakya,Sandesh Acharya,Ravi Pandit,Dinesh Gothe
备注:7 pages, 10 figures
链接:点击下载PDF文件
摘要:语音克隆是个性化语音界面的一个重要特征。一个神经声音克隆系统可以只用几个音频样本来模仿某人的声音。说话人编码和说话人自适应都是语音克隆领域的研究课题。说话人自适应依赖于微调多说话人生成模型,这涉及训练单独的模型来推断用于说话人编码的新说话人嵌入。这两种方法都可以实现良好的性能,即使有少量的克隆音频,在语音的自然度和相似性的原始扬声器。扬声器编码方法更适合于低资源部署,因为它们需要的内存显着减少,并且比扬声器自适应具有更快的克隆时间,可以提供更大的自然度和相似性。主要目标是创建一个语音克隆系统,产生带有尼泊尔口音或听起来像尼泊尔语的音频输出。为了进一步推进TTS,迁移学习的思想被有效地用于解决该系统开发过程中遇到的几个问题,包括音频质量差和缺乏可用数据。摘要:Voice cloning is a prominent feature in personalized speech interfaces. A neural vocal cloning system can mimic someone's voice using just a few audio samples. Both speaker encoding and speaker adaptation are topics of research in the field of voice cloning. Speaker adaptation relies on fine-tuning a multi-speaker generative model, which involves training a separate model to infer a new speaker embedding used for speaker encoding. Both methods can achieve excellent performance, even with a small number of cloning audios, in terms of the speech's naturalness and similarity to the original speaker. Speaker encoding approaches are more appropriate for low-resource deployment since they require significantly less memory and have a faster cloning time than speaker adaption, which can offer slightly greater naturalness and similarity. The main goal is to create a vocal cloning system that produces audio output with a Nepali accent or that sounds like Nepali. For the further advancement of TTS, the idea of transfer learning was effectively used to address several issues that were encountered in the development of this system, including the poor audio quality and the lack of available data.

【5】 Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
标题: 转换和说话:Zero-Shot口音转换,最少监督
作者:Zhijun Jia,Huaying Xue,Xiulian Peng,Yan Lu
备注:9 pages, 4 figures, conference
链接:点击下载PDF文件
摘要:重音转换是一个语音单元和韵律模式都需要转换的问题,其关键是并行数据资源少。我们提出了一个两阶段的生成框架“转换和说话”,其中转换只操作的语义标记水平和语音合成的条件下,转换后的语义标记与语音生成模型在目标口音域。解耦设计使“说话”模块能够使用大量的目标口音语音,并减轻了“转换”模块所需的并行数据。以语义标记为桥梁的转换也减轻了对文本转换数据的需求,并解锁了语言预训练技术的使用,以进一步有效地减少对并行口音语音数据的需求。为了降低“说话”的复杂度和延迟,设计了一个单级AR生成模型,以实现良好的质量以及较低的计算成本。对印度英语到一般美国英语转换的实验表明,该框架实现了最先进的性能,口音相似性,语音质量和扬声器维护只有15分钟的弱并行数据,不限于同一个扬声器。对不同口音类型的广泛实验表明,该框架具有高度的适应性,使其易于扩展,以适应其他口音与低资源数据。音频样本可在https: www.microsoft.com en-us research project convert-and-speak-zero-shot-accent-conversion-with-minimumsupervision 上获得。摘要:Low resource of parallel data is the key challenge of accent conversion(AC) problem in which both the pronunciation units and prosody pattern need to be converted. We propose a two-stage generative framework "convert-and-speak" in which the conversion is only operated on the semantic token level and the speech is synthesized conditioned on the converted semantic token with a speech generative model in target accent domain. The decoupling design enables the "speaking" module to use massive amount of target accent speech and relieves the parallel data required for the "conversion" module. Conversion with the bridge of semantic token also relieves the requirement for the data with text transcriptions and unlocks the usage of language pre-training technology to further efficiently reduce the need of parallel accent speech data. To reduce the complexity and latency of "speaking", a single-stage AR generative model is designed to achieve good quality as well as lower computation cost. Experiments on Indian-English to general American-English conversion show that the proposed framework achieves state-of-the-art performance in accent similarity, speech quality, and speaker maintenance with only 15 minutes of weakly parallel data which is not constrained to the same speaker. Extensive experimentation with diverse accent types suggests that this framework possesses a high degree of adaptability, making it readily scalable to accommodate other accents with low-resource data. Audio samples are available at https: www.microsoft.com en-us research project convert-and-speak-zero-shot-accent-conversion-with-minimumsupervision .

【6】 SZU-AFS Antispoofing System for the ASVspoof 5 Challenge
标题: ASVspoof 5挑战赛的SZU-ATL反恶搞系统
作者:Yuxiong Xu,Jiafeng Zhong,Sengui Zheng,Zefeng Liu,Bin Li
备注:8 pages, 2 figures, ASVspoof 5 Workshop (Interspeech2024 Satellite)
链接:点击下载PDF文件
摘要:本文介绍了SZU-AFS反欺骗系统,在开放条件下的ASVspoof 5挑战赛的轨道1设计。该系统分为四个阶段:选择基线模型,探索有效的数据增强(DA)方法进行微调,应用基于梯度范数感知最小化(GAM)的协同增强策略进行二次微调,并融合两个性能最佳的微调模型的logits分数。该系统利用Wav 2 Vec 2前端特征提取器和AASIST后端分类器作为基线模型。在模型微调期间,研究了三种不同的DA策略:单一DA、随机DA和级联DA。此外,采用的基于GAM的协同增强策略旨在在数据和优化器级别对增强模型进行微调,帮助Adam优化器找到更平坦的最小值,从而提高模型的泛化能力。总体而言,最终融合系统在评价集上实现了0.115的minDCF和4.04%的EER。摘要:This paper presents the SZU-AFS anti-spoofing system, designed for Track 1 of the ASVspoof 5 Challenge under open conditions. The system is built with four stages: selecting a baseline model, exploring effective data augmentation (DA) methods for fine-tuning, applying a co-enhancement strategy based on gradient norm aware minimization (GAM) for secondary fine-tuning, and fusing logits scores from the two best-performing fine-tuned models. The system utilizes the Wav2Vec2 front-end feature extractor and the AASIST back-end classifier as the baseline model. During model fine-tuning, three distinct DA policies have been investigated: single-DA, random-DA, and cascade-DA. Moreover, the employed GAM-based co-enhancement strategy, designed to fine-tune the augmented model at both data and optimizer levels, helps the Adam optimizer find flatter minima, thereby boosting model generalization. Overall, the final fusion system achieves a minDCF of 0.115 and an EER of 4.04% on the evaluation set.

【7】 Hear Your Face: Face-based voice conversion with F0 estimation
标题: 听到你的脸:基于脸的语音转换,具有F0估计
作者:Jaejun Lee,Yoori Oh,Injune Hwang,Kyogu Lee
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:本文深入研究了新兴的基于人脸的语音转换领域,利用个人的面部特征和他们的声音特征之间的独特关系。我们提出了一种新的基于人脸的语音转换框架,特别是利用目标扬声器的平均基频,完全来自他们的面部图像。通过广泛的分析,我们的框架展示了卓越的语音生成质量以及将面部特征与语音特征对齐的能力,包括跟踪目标说话者的基本频率。摘要:This paper delves into the emerging field of face-based voice conversion, leveraging the unique relationship between an individual's facial features and their vocal characteristics. We present a novel face-based voice conversion framework that particularly utilizes the average fundamental frequency of the target speaker, derived solely from their facial images. Through extensive analysis, our framework demonstrates superior speech generation quality and the ability to align facial features with voice characteristics, including tracking of the target speaker's fundamental frequency.

【8】 Unsupervised Composable Representations for Audio
标题: 音频的无监督可组合表示
作者:Giovanni Bindi,Philippe Esling
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:目前的生成模型能够生成高质量的人工制品,但已被证明与组合推理,这可以被定义为从简单的元素生成复杂结构的能力斗争。在本文中,我们专注于音乐数据的组成表示学习问题,特别是针对完全无监督的设置。我们提出了一个简单且可扩展的框架,该框架利用了一个明确的组合归纳偏差,该偏差由一个灵活的自动编码目标定义,该目标可以利用任何当前最先进的生成模型。我们证明了我们的框架,与扩散模型一起使用,自然地解决了无监督音频源分离的任务,表明我们的模型能够执行高质量的分离。我们的研究结果表明,我们的建议实现了可比或优于其他盲源分离方法的性能,此外,它甚至超过了目前最先进的监督基线的信号干扰比指标。此外,通过学习一个后验掩蔽扩散模型在空间的组合表示,我们实现了一个系统能够无缝执行无监督源分离,无条件生成和变化生成。最后,由于我们的建议在预训练的神经音频编解码器的潜在空间中工作,因此相对于其他神经基线,它还提供了更低的计算成本。摘要:Current generative models are able to generate high-quality artefacts but have been shown to struggle with compositional reasoning, which can be defined as the ability to generate complex structures from simpler elements. In this paper, we focus on the problem of compositional representation learning for music data, specifically targeting the fully-unsupervised setting. We propose a simple and extensible framework that leverages an explicit compositional inductive bias, defined by a flexible auto-encoding objective that can leverage any of the current state-of-art generative models. We demonstrate that our framework, used with diffusion models, naturally addresses the task of unsupervised audio source separation, showing that our model is able to perform high-quality separation. Our findings reveal that our proposal achieves comparable or superior performance with respect to other blind source separation methods and, furthermore, it even surpasses current state-of-art supervised baselines on signal-to-interference ratio metrics. Additionally, by learning an a-posteriori masking diffusion model in the space of composable representations, we achieve a system capable of seamlessly performing unsupervised source separation, unconditional generation, and variation generation. Finally, as our proposal works in the latent space of pre-trained neural audio codecs, it also provides a lower computational cost with respect to other neural baselines.

【9】 A Transcription Prompt-based Efficient Audio Large Language Model for Robust Speech Recognition
标题: 基于转录预算的高效音频大语言模型用于鲁棒语音识别
作者:Yangze Li,Xiong Wang,Songjun Cao,Yike Zhang,Long Ma,Lei Xie
链接:点击下载PDF文件
摘要:Audio-LLM将音频模态引入到大型语言模型(LLM)中,使强大的LLM能够识别,理解和生成音频。然而,在嘈杂环境中的语音识别过程中,我们观察到音频LLM中存在错觉和重复问题,导致替换和插入错误。针对上述问题,提出了一种基于语音识别专家的音频LLM算法,该算法引入语音识别专家作为语音识别标记器,并采用混合自回归(AR)和非自回归(NAR)解码方法。在10 kh的WenetSpeech中文语料上的实验表明,该方法在Test_Net和Test_Meeting两个评测集上的CER分别比基线降低了12.2%和9.6%。值得注意的是,我们降低了评估集上的解码重复率为零,这表明解码重复问题已从根本上得到解决。摘要:Audio-LLM introduces audio modality into a large language model (LLM) to enable a powerful LLM to recognize, understand, and generate audio. However, during speech recognition in noisy environments, we observed the presence of illusions and repetition issues in audio-LLM, leading to substitution and insertion errors. This paper proposes a transcription prompt-based audio-LLM by introducing an ASR expert as a transcription tokenizer and a hybrid Autoregressive (AR) Non-autoregressive (NAR) decoding approach to solve the above problems. Experiments on 10k-hour WenetSpeech Mandarin corpus show that our approach decreases 12.2% and 9.6% CER relatively on Test_Net and Test_Meeting evaluation sets compared with baseline. Notably, we reduce the decoding repetition rate on the evaluation set to zero, showing that the decoding repetition problem has been solved fundamentally.

【10】 Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face Animation
标题: 元学习赋予元面孔:个性化说话风格改编,用于音频驱动的3D会说话的面孔动画
作者:Xukun Zhou,Fengxin Li,Ziqiao Peng,Kejian Wu,Jun He,Biao Qin,Zhaoxin Fan,Hongyan Liu
链接:点击下载PDF文件
摘要:音频驱动的3D人脸动画在实时流媒体和增强现实应用中越来越重要。虽然已经观察到了显着的进展,大多数现有的方法是专为特定的个人与预定义的说话风格,从而忽略了不同的说话风格的适应性。为了解决这个问题,本文介绍了MetaFace,一种新的方法精心制作的讲话风格适应。基于元学习的新概念,MetaFace由几个关键组件组成:用于基本说话风格适应的鲁棒Meta学习阶段(RMIS),用于在观察到的和未观察到的说话风格之间建立联系的动态关系挖掘神经过程(DRMN),以及用于提高模型优化效率和学习风格细节的低秩矩阵内存缩减方法。利用这些新颖的设计,MetaFace不仅显著优于现有的强大基线,而且还建立了一个新的最先进的,正如我们的实验结果所证实的那样。摘要:Audio-driven 3D face animation is increasingly vital in live streaming and augmented reality applications. While remarkable progress has been observed, most existing approaches are designed for specific individuals with predefined speaking styles, thus neglecting the adaptability to varied speaking styles. To address this limitation, this paper introduces MetaFace, a novel methodology meticulously crafted for speaking style adaptation. Grounded in the novel concept of meta-learning, MetaFace is composed of several key components: the Robust Meta Initialization Stage (RMIS) for fundamental speaking style adaptation, the Dynamic Relation Mining Neural Process (DRMN) for forging connections between observed and unobserved speaking styles, and the Low-rank Matrix Memory Reduction Approach to enhance the efficiency of model optimization as well as learning style details. Leveraging these novel designs, MetaFace not only significantly outperforms robust existing baselines but also establishes a new state-of-the-art, as substantiated by our experimental results.

【11】 Auptimize: Optimal Placement of Spatial Audio Cues for Extended Reality
标题: 优化:延展实境的空间音频线索的最佳放置
作者:Hyunsung Cho,Alexander Wang,Divya Kartik,Emily Liying Xie,Yukang Yan,David Lindlbauer
备注:UIST 2024
链接:点击下载PDF文件
摘要:延展实境(XR)中的空间音频为用户提供了更好的虚拟元素放置位置意识,并有效地引导他们的事件,如通知,来自不同窗口的系统警报,或接近化身。然而,人类在定位声音线索方面是不准确的,特别是由于人类听觉感知的限制(例如角度辨别误差和前后混淆)而具有多个源。这会降低XR界面的效率,因为用户会错误识别声音来自哪个XR元素。为了解决这个问题,我们提出了Auptimize,一种新的计算方法,用于放置XR声源,通过利用口技效应来减轻这种定位误差。Auptimize将声源位置从视觉元素中分离出来,并将声源重新定位到最佳位置,以明确识别声音提示,避免由于源间接近和前后混淆而导致的错误。我们的评估表明,Auptimize减少空间音频为基础的源识别错误相比,在成对的视觉-声音位置播放声音线索。我们证明了Auptimize对于基于空间音频的交互式XR场景的适用性。摘要:Spatial audio in Extended Reality (XR) provides users with better awareness of where virtual elements are placed, and efficiently guides them to events such as notifications, system alerts from different windows, or approaching avatars. Humans, however, are inaccurate in localizing sound cues, especially with multiple sources due to limitations in human auditory perception such as angular discrimination error and front-back confusion. This decreases the efficiency of XR interfaces because users misidentify from which XR element a sound is coming. To address this, we propose Auptimize, a novel computational approach for placing XR sound sources, which mitigates such localization errors by utilizing the ventriloquist effect. Auptimize disentangles the sound source locations from the visual elements and relocates the sound sources to optimal positions for unambiguous identification of sound cues, avoiding errors due to inter-source proximity and front-back confusion. Our evaluation shows that Auptimize decreases spatial audio-based source identification errors compared to playing sound cues at the paired visual-sound locations. We demonstrate the applicability of Auptimize for diverse spatial audio-based interactive XR scenarios.

【12】 Enhancing Audio-Language Models through Self-Supervised Post-Training with Text-Audio Pairs
标题: 通过文本-音频对的自我监督后训练增强音频语言模型
作者:Anshuman Sinha,Camille Migozzi,Aubin Rey,Chao Zhang
备注:31 pages, 11 figures
链接:点击下载PDF文件
摘要:音频和文本的多模态对比学习策略的研究迅速引起了人们的兴趣。对比训练的音频语言模型(ALM),如CLAP,建立了跨音频和语言模态的统一表示,通过提供良好的文本对齐音频编码器,提高了各种后续任务的效率,反之亦然。这些改进在诸如zero-shot音频分类和音频检索等领域中是明显的。然而,这些模型理解自然语言和时间关系的能力在很大程度上仍然是一个未开发和开放的研究领域。在本文中,我们建议配备的多模态ALMs的时间理解,而不失去其固有的先前能力的音频语言任务的时间灌输方法TeminAL。我们实现了一个两阶段的训练方案TeminAL A $ $ B,其中模型首先学习区分TeminAL A中的多个声音,然后是一个灌输时间感的阶段,从而增强其在TeminAL B中的时间理解。这种方法导致在ESC-50数据集上的时间理解的平均性能增益为5.28 %$,而该模型在AudioCap Clotho数据集上的zero-shot检索和分类任务中仍然具有竞争力。我们还注意到,缺乏适当的评价技术,对比ALM,并提出了一个战略评估ALM在zero-shot设置。通用的zero-shot模型评估策略ZSTE,用于评估各种先验模型。ZSTE展示了一个评估所有ZS对比模型的通用策略。使用TeminAL训练的模型在大多数下游任务上成功地优于当前模型。摘要:Research on multi-modal contrastive learning strategies for audio and text has rapidly gained interest. Contrastively trained Audio-Language Models (ALMs), such as CLAP, which establish a unified representation across audio and language modalities, have enhanced the efficacy in various subsequent tasks by providing good text aligned audio encoders and vice versa. These improvements are evident in areas like zero-shot audio classification and audio retrieval, among others. However, the ability of these models to understand natural language and temporal relations is still a largely unexplored and open field for research. In this paper, we propose to equip the multi-modal ALMs with temporal understanding without loosing their inherent prior capabilities of audio-language tasks with a temporal instillation method TeminAL. We implement a two-stage training scheme TeminAL A $ &$ B, where the model first learns to differentiate between multiple sounds in TeminAL A, followed by a phase that instills a sense of time, thereby enhancing its temporal understanding in TeminAL B. This approach results in an average performance gain of $5.28 %$ in temporal understanding on the ESC-50 dataset, while the model remains competitive in zero-shot retrieval and classification tasks on the AudioCap Clotho datasets. We also note the lack of proper evaluation techniques for contrastive ALMs and propose a strategy for evaluating ALMs in zero-shot settings. The general-purpose zero-shot model evaluation strategy ZSTE, is used to evaluate various prior models. ZSTE demonstrates a general strategy to evaluate all ZS contrastive models. The model trained with TeminAL successfully outperforms current models on most downstream tasks.

【13】 Efficient Autoregressive Audio Modeling via Next-Scale Prediction
标题: 通过下一规模预测的高效自回归音频建模
作者:Kai Qiu,Xiang Li,Hao Chen,Jie Sun,Jinglu Wang,Zhe Lin,Marios Savvides,Bhiksha Raj
备注:7 pages, 6 figures, 7 tables
链接:点击下载PDF文件
摘要:随着扩散模型(DM)和自回归(AR)模型等复杂生成模型的发展,音频生成技术取得了显著的进步。然而,由于音频的自然显著序列长度,音频生成的效率仍然是一个需要解决的重要问题,特别是对于合并在大型语言模型(LLM)中的AR模型。本文分析了音频标记化的标记长度,提出了一种新的 textbf{S}标度级 textbf{A}音频 textbf{T}标记器(SAT),改进了残差量化。在SAT的基础上,进一步提出了尺度级的 textbf{A}自回归 textbf {A}uto textbf{R}渐进(AAR)建模框架,将下一个标记的AR预测转移到下一个尺度的AR预测,显著降低了训练成本和推理时间。为了验证所提出的方法的有效性,我们全面分析了设计选择,并证明了所提出的AAR框架实现了显着的 textbf{35}$ times$更快的推理速度和+ textbf{1.33} Fr 'echet音频距离(FAD)对AudioSet基准的基线。代码: url{https: github.com qiuk2 AAR}。摘要:Audio generation has achieved remarkable progress with the advance of sophisticated generative models, such as diffusion models (DMs) and autoregressive (AR) models. However, due to the naturally significant sequence length of audio, the efficiency of audio generation remains an essential issue to be addressed, especially for AR models that are incorporated in large language models (LLMs). In this paper, we analyze the token length of audio tokenization and propose a novel textbf{S}cale-level textbf{A}udio textbf{T}okenizer (SAT), with improved residual quantization. Based on SAT, a scale-level textbf{A}coustic textbf{A}uto textbf{R}egressive (AAR) modeling framework is further proposed, which shifts the next-token AR prediction to next-scale AR prediction, significantly reducing the training cost and inference time. To validate the effectiveness of the proposed approach, we comprehensively analyze design choices and demonstrate the proposed AAR framework achieves a remarkable textbf{35}$ times$ faster inference speed and + textbf{1.33} Fr 'echet Audio Distance (FAD) against baselines on the AudioSet benchmark. Code: url{https: github.com qiuk2 AAR}.


机器翻译,仅供参考