今日论文合集:cs.SD语音19篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Towards Enhanced Classification of Abnormal Lung sound in Multi-breath: A Light Weight Multi-label and Multi-head Attention Classification Method
标题: 多次呼吸中异常肺音的增强分类:一种轻量级多标签和多头注意力分类方法
作者:Yi-Wei Chua,Yun-Chien Cheng
链接:点击下载PDF文件
摘要:本研究旨在开发一种肺呼吸音异常分类的辅助诊断系统,通过创新的多标签学习方法和多头注意机制,提高异常呼吸音自动分类的准确性。为了解决现有呼吸声数据集的类别不平衡和缺乏多样性的问题,我们的研究采用了一个轻量级和高度准确的模型,使用二维标签集来表示多个呼吸声特征。我们的方法在ICBHI2017数据集上的四类任务中获得了59.2%的ICBHI分数,证明了其在轻量级和高准确性方面的优势。该研究不仅提高了肺呼吸音异常自动诊断的准确性,而且为临床应用开辟了新的可能性。摘要:This study aims to develop an auxiliary diagnostic system for classifying abnormal lung respiratory sounds, enhancing the accuracy of automatic abnormal breath sound classification through an innovative multi-label learning approach and multi-head attention mechanism. Addressing the issue of class imbalance and lack of diversity in existing respiratory sound datasets, our study employs a lightweight and highly accurate model, using a two-dimensional label set to represent multiple respiratory sound characteristics. Our method achieved a 59.2% ICBHI score in the four-category task on the ICBHI2017 dataset, demonstrating its advantages in terms of lightweight and high accuracy. This study not only improves the accuracy of automatic diagnosis of lung respiratory sound abnormalities but also opens new possibilities for clinical applications.

【2】 Towards zero-shot amplifier modeling: One-to-many amplifier modeling via tone embedding control
标题: 走向零触发放大器建模:通过音调嵌入控制进行一对多放大器建模
作者:Yu-Hua Chen,Yen-Tung Yeh,Yuan-Chiao Cheng,Jui-Te Wu,Yu-Hsiang Ho,Jyh-Shing Roger Jang,Yi-Hsuan Yang
备注:ISMIR 2024
链接:点击下载PDF文件
摘要:近年来,通过神经音频效果建模来复制模拟设备电路已经引起了越来越多的兴趣。现有的工作主要集中在一对一的仿真策略,单独建模特定的设备。在本文中,我们解决了较少探索的一对多仿真场景,利用条件反射机制通过单个神经模型来仿真多个吉他放大器。对于条件表示,我们使用对比学习来构建一个音调嵌入编码器,该编码器提取各种放大器的风格相关特征,利用全面放大器设置的数据集。针对zero-shot应用场景,我们还研究了音调嵌入表示的各种策略,针对训练时间内看不到的放大器的两种基于检索的嵌入方法评估参考音调嵌入。我们的研究结果展示了所提出的方法在实现多功能一对多放大器建模方面的功效和潜力,为zero-shot音频建模应用迈出了基础性的一步。摘要:Replicating analog device circuits through neural audio effect modeling has garnered increasing interest in recent years. Existing work has predominantly focused on a one-to-one emulation strategy, modeling specific devices individually. In this paper, we tackle the less-explored scenario of one-to-many emulation, utilizing conditioning mechanisms to emulate multiple guitar amplifiers through a single neural model. For condition representation, we use contrastive learning to build a tone embedding encoder that extracts style-related features of various amplifiers, leveraging a dataset of comprehensive amplifier settings. Targeting zero-shot application scenarios, we also examine various strategies for tone embedding representation, evaluating referenced tone embedding against two retrieval-based embedding methods for amplifiers unseen in the training time. Our findings showcase the efficacy and potential of the proposed methods in achieving versatile one-to-many amplifier modeling, contributing a foundational step towards zero-shot audio modeling applications.

【3】 GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio Synthesis
标题: GRAOT:为基于扩散模型的音频合成生成稳健水印
作者:Weizhi Liu,Yue Li,Dongdong Lin,Hui Tian,Haizhou Li
链接:点击下载PDF文件
摘要:在扩散模型等生成模型的蓬勃发展中,将合成音频与其自然对应物区分开来的任务变得更加艰巨。Deepfake检测为应对这一挑战提供了可行的解决方案。然而,这种防御性措施无意中推动了生成模型的不断完善。水印作为一种积极主动的可持续策略出现,先发制人地规范合成内容的创建和传播。因此,本文,作为先驱,提出了生成鲁棒音频水印方法(Groot),提出了一个范例,主动监督合成音频及其源扩散模型。在这个范例中,水印生成和音频合成的过程同时发生,通过配备专用编码器的参数固定扩散模型来促进。嵌入在音频中的水印随后可以由轻量级解码器检索。实验结果突出了Groot的出色表现,特别是在鲁棒性方面,超过了领先的最先进的方法。除了对单个后处理攻击具有令人印象深刻的恢复能力外,Groot在面对复合攻击时表现出出色的鲁棒性,保持了约95%的平均水印提取准确率。摘要:Amid the burgeoning development of generative models like diffusion models, the task of differentiating synthesized audio from its natural counterpart grows more daunting. Deepfake detection offers a viable solution to combat this challenge. Yet, this defensive measure unintentionally fuels the continued refinement of generative models. Watermarking emerges as a proactive and sustainable tactic, preemptively regulating the creation and dissemination of synthesized content. Thus, this paper, as a pioneer, proposes the generative robust audio watermarking method (Groot), presenting a paradigm for proactively supervising the synthesized audio and its source diffusion models. In this paradigm, the processes of watermark generation and audio synthesis occur simultaneously, facilitated by parameter-fixed diffusion models equipped with a dedicated encoder. The watermark embedded within the audio can subsequently be retrieved by a lightweight decoder. The experimental results highlight Groot's outstanding performance, particularly in terms of robustness, surpassing that of the leading state-of-the-art methods. Beyond its impressive resilience against individual post-processing attacks, Groot exhibits exceptional robustness when facing compound attacks, maintaining an average watermark extraction accuracy of around 95%.

【4】 LiteFocus: Accelerated Diffusion Inference for Long Audio Synthesis
标题: LiteFocus:长音频合成的加速扩散推理
作者:Zhenxiong Tan,Xinyin Ma,Gongfan Fang,Xinchao Wang
备注:Interspeech 2024; Code: this https URL
链接:点击下载PDF文件
摘要:潜在扩散模型在音频生成中显示出有希望的结果,与传统方法相比取得了显着进步。然而,他们的表现,虽然令人印象深刻的短音频剪辑,面临挑战时,扩展到更长的音频序列。这些挑战是由于模型的自我注意机制和主要在10秒剪辑上的训练,这使得在没有适应的情况下扩展到更长的音频变得复杂。针对这些问题,我们引入了一种新的方法,LiteFocus,增强了现有的音频潜在的扩散模型在长音频合成的推断。通过观察自注意中的注意模式,我们采用了一种双稀疏形式的注意计算,即同频聚焦和跨频补偿,它在同频约束下减少了注意计算,同时通过跨频再滤波提高了音频质量。LiteFocus在合成80秒音频片段时,使用基于扩散的TTA模型将推理时间大幅减少了1.99倍,同时还获得了更高的音频质量。摘要:Latent diffusion models have shown promising results in audio generation, making notable advancements over traditional methods. However, their performance, while impressive with short audio clips, faces challenges when extended to longer audio sequences. These challenges are due to model's self-attention mechanism and training predominantly on 10-second clips, which complicates the extension to longer audio without adaptation. In response to these issues, we introduce a novel approach, LiteFocus that enhances the inference of existing audio latent diffusion models in long audio synthesis. Observed the attention pattern in self-attention, we employ a dual sparse form for attention calculation, designated as same-frequency focus and cross-frequency compensation, which curtails the attention computation under same-frequency constraints, while enhancing audio quality through cross-frequency refillment. LiteFocus demonstrates substantial reduction on inference time with diffusion-based TTA model by 1.99x in synthesizing 80-second audio clips while also obtaining improved audio quality.

【5】 BandControlNet: Parallel Transformers-based Steerable Popular Music Generation with Fine-Grained Spatiotemporal Features
标题: BandControlNet:基于并行转换器的可操纵流行音乐生成,具有细粒度时空特征
作者:Jing Luo,Xinyu Yang,Dorien Herremans
备注:Demo page: this https URL
链接:点击下载PDF文件
摘要:可控音乐生成通过将用户的意图投射到他们想要的音乐上来促进人类和作曲系统之间的交互。引入可控性的挑战是符号音乐生成领域越来越重要的问题。在构建可控生成流行多乐器音乐系统时,通常存在两个主要挑战,即弱可控性和差的音乐质量。为了解决这些问题,我们首先提出时空功能强大,细粒度的控制,以提高生成模型的可控性。此外,一个高效的音乐表示称为REMI_Track被设计成多个并行的音乐序列,并缩短每个轨道的序列长度与字节对编码(BPE)技术。随后,我们发布了BandControlNet,一个基于并行Transformers的条件模型,来处理多个音乐序列,并生成高质量的音乐样本,这些音乐样本是根据给定的时空控制特征进行调节的。更具体地说,BandControlNet的两个专门设计的模块,即结构增强的自我注意力(SE-SA)和跨轨道Transformer(CTT),分别用于加强所产生的音乐结构和轨道间的和谐建模。在两个不同长度的流行音乐数据集上的实验结果表明,所提出的BandControlNet在保真度和推理速度方面的最客观指标上优于其他条件音乐生成模型,并且在生成长音乐样本时表现出很强的鲁棒性。主观评估显示,在短数据集上训练的BandControlNet可以生成与最先进的模型相当的音乐质量,同时使用较长的数据集显着优于它们。摘要:Controllable music generation promotes the interaction between humans and composition systems by projecting the users' intent on their desired music. The challenge of introducing controllability is an increasingly important issue in the symbolic music generation field. When building controllable generative popular multi-instrument music systems, two main challenges typically present themselves, namely weak controllability and poor music quality. To address these issues, we first propose spatiotemporal features as powerful and fine-grained controls to enhance the controllability of the generative model. In addition, an efficient music representation called REMI_Track is designed to convert multitrack music into multiple parallel music sequences and shorten the sequence length of each track with Byte Pair Encoding (BPE) techniques. Subsequently, we release BandControlNet, a conditional model based on parallel Transformers, to tackle the multiple music sequences and generate high-quality music samples that are conditioned to the given spatiotemporal control features. More concretely, the two specially designed modules of BandControlNet, namely structure-enhanced self-attention (SE-SA) and Cross-Track Transformer (CTT), are utilized to strengthen the resulting musical structure and inter-track harmony modeling respectively. Experimental results tested on two popular music datasets of different lengths demonstrate that the proposed BandControlNet outperforms other conditional music generation models on most objective metrics in terms of fidelity and inference speed and shows great robustness in generating long music samples. The subjective evaluations show BandControlNet trained on short datasets can generate music with comparable quality to state-of-the-art models, while outperforming them significantly using longer datasets.

【6】 DDFAD: Dataset Distillation Framework for Audio Data
标题: DDFAD:音频数据的数据集蒸馏框架
作者:Wenbo Jiang,Rui Zhang,Hongwei Li,Xiaoyuan Liu,Haomiao Yang,Shui Yu
链接:点击下载PDF文件
摘要:深度神经网络(DNN)在许多应用中取得了巨大的成功。DNN的卓越性能在很大程度上归功于大量高质量训练数据集的可用性。然而,处理如此庞大的训练数据需要巨大的计算和存储资源。数据集蒸馏是解决这个问题的一个很有前途的解决方案,它提供了将大型数据集压缩成较小的蒸馏数据集的能力。在提取数据集上训练的模型可以实现与在整个数据集上训练的模型相当的性能。 虽然已经在图像数据中展示了数据集蒸馏,但是没有人探索用于音频数据的数据集蒸馏。在这项工作中,我们第一次提出了一个音频数据的数据集蒸馏框架(DDFAD)。具体来说,我们首先提出融合差分MFCC(FD-MFCC)作为音频数据的提取特征。然后,通过匹配训练轨迹提取方法提取FD-MFCC。最后,我们提出了一种基于Griffin-Lim算法的音频信号重建算法,从提取的FD-MFCC中重建音频信号。大量的实验证明了DDFAD在各种音频数据集上的有效性。此外,我们表明,DDFAD在许多应用中,如持续学习和神经结构搜索具有良好的应用前景。摘要:Deep neural networks (DNNs) have achieved significant success in numerous applications. The remarkable performance of DNNs is largely attributed to the availability of massive, high-quality training datasets. However, processing such massive training data requires huge computational and storage resources. Dataset distillation is a promising solution to this problem, offering the capability to compress a large dataset into a smaller distilled dataset. The model trained on the distilled dataset can achieve comparable performance to the model trained on the whole dataset. While dataset distillation has been demonstrated in image data, none have explored dataset distillation for audio data. In this work, for the first time, we propose a Dataset Distillation Framework for Audio Data (DDFAD). Specifically, we first propose the Fused Differential MFCC (FD-MFCC) as extracted features for audio data. After that, the FD-MFCC is distilled through the matching training trajectory distillation method. Finally, we propose an audio signal reconstruction algorithm based on the Griffin-Lim Algorithm to reconstruct the audio signal from the distilled FD-MFCC. Extensive experiments demonstrate the effectiveness of DDFAD on various audio datasets. In addition, we show that DDFAD has promising application prospects in many applications, such as continual learning and neural architecture search.

【7】 Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity
标题: 具有增强同步性的屏蔽生成视频到音频Transformer
作者:Santiago Pascual,Chunghsin Yeh,Ioannis Tsiamas,Joan Serrà
备注:Accepted to ECCV 2024
链接:点击下载PDF文件
摘要:视频到音频(V2A)生成利用仅视觉视频特征来呈现与场景匹配的合理声音。重要的是,生成的声音起始应该匹配与它们对齐的视觉动作,否则会出现不自然的同步伪像。最近的工作已经探索了在静止图像上调节声音发生器的进展,然后是视频特征,专注于质量和语义匹配而忽略同步,或者通过牺牲一些质量来专注于提高同步。在这项工作中,我们提出了一个V2A生成模型,名为MaskVAT,它将全频带高质量的通用音频编解码器与序列到序列掩蔽生成模型互连。这种组合允许同时对高音频质量、语义匹配和时间同步性进行建模。我们的研究结果表明,通过将高质量的编解码器与适当的预训练的视听特征和序列到序列的并行结构相结合,我们一方面能够产生高度同步的结果,同时与非编解码器生成音频模型的最新技术水平相竞争。示例视频和生成的音频可在https: maskvat.github.io上获得。摘要:Video-to-audio (V2A) generation leverages visual-only video features to render plausible sounds that match the scene. Importantly, the generated sound onsets should match the visual actions that are aligned with them, otherwise unnatural synchronization artifacts arise. Recent works have explored the progression of conditioning sound generators on still images and then video features, focusing on quality and semantic matching while ignoring synchronization, or by sacrificing some amount of quality to focus on improving synchronization only. In this work, we propose a V2A generative model, named MaskVAT, that interconnects a full-band high-quality general audio codec with a sequence-to-sequence masked generative model. This combination allows modeling both high audio quality, semantic matching, and temporal synchronicity at the same time. Our results show that, by combining a high-quality codec with the proper pre-trained audio-visual features and a sequence-to-sequence parallel structure, we are able to yield highly synchronized results on one hand, whilst being competitive with the state of the art of non-codec generative audio models. Sample videos and generated audios are available at https: maskvat.github.io .

【8】 Mutual Learning for Acoustic Matching and Dereverberation via Visual Scene-driven Diffusion
标题: 通过视觉场景驱动扩散进行声学匹配和去回响的相互学习
作者:Jian Ma,Wenguan Wang,Yi Yang,Feng Zheng
备注:ECCV 2024; Project page: this https URL
链接:点击下载PDF文件
摘要:视觉声学匹配(VAM)是增强沉浸式体验的关键,而去混响的任务是有效地提高音频可懂度。现有的方法单独对待每项任务,忽略了它们之间固有的相互作用。此外,这些方法依赖于成对的训练数据,这是具有挑战性的获取,阻碍了大量的未配对数据的利用。在本文中,我们介绍MVSD,一个基于扩散模型的相互学习框架。MVSD认为这两个任务是对称的,利用互惠关系,以促进从反向任务中学习,并克服数据稀缺。此外,我们采用扩散模型作为基础条件转换器,以规避传统GAN架构的训练不稳定性和过度平滑的缺点。具体而言,MVSD采用两个转换器:一个用于VAM,称为混响器,另一个用于去混响,称为去混响器。去混响器判断混响器产生的混响音频是否听起来像是在条件视觉场景中,反之亦然。通过形成闭环,这两个转换器可以生成信息反馈信号来优化反向任务,即使使用轻松获取的单向未配对数据也是如此。对两个标准基准进行了广泛的实验,即,SoundSpaces-Speech和Acoustic AVSpeech表明我们的框架可以提高混响器和去混响器的性能,并更好地匹配指定的视觉场景。摘要:Visual acoustic matching (VAM) is pivotal for enhancing the immersive experience, and the task of dereverberation is effective in improving audio intelligibility. Existing methods treat each task independently, overlooking the inherent reciprocity between them. Moreover, these methods depend on paired training data, which is challenging to acquire, impeding the utilization of extensive unpaired data. In this paper, we introduce MVSD, a mutual learning framework based on diffusion models. MVSD considers the two tasks symmetrically, exploiting the reciprocal relationship to facilitate learning from inverse tasks and overcome data scarcity. Furthermore, we employ the diffusion model as foundational conditional converters to circumvent the training instability and over-smoothing drawbacks of conventional GAN architectures. Specifically, MVSD employs two converters: one for VAM called reverberator and one for dereverberation called dereverberator. The dereverberator judges whether the reverberation audio generated by reverberator sounds like being in the conditional visual scenario, and vice versa. By forming a closed loop, these two converters can generate informative feedback signals to optimize the inverse tasks, even with easily acquired one-way unpaired data. Extensive experiments on two standard benchmarks, i.e., SoundSpaces-Speech and Acoustic AVSpeech, exhibit that our framework can improve the performance of the reverberator and dereverberator and better match specified visual scenarios.

【9】 The Interpretation Gap in Text-to-Music Generation Models
标题: 文本到音乐生成模型中的解释差距
作者:Yongyi Zang,Yixiao Zhang
备注:Under review
链接:点击下载PDF文件
摘要:大规模的文本到音乐生成模型大大增强了音乐创作能力,提供了前所未有的创作自由。然而,他们与人类音乐家有效合作的能力仍然有限。在本文中,我们提出了一个框架来描述音乐的互动过程,其中包括表达,解释和执行的控制。根据这个框架,我们认为,现有的文本到音乐的模型和音乐家之间的主要差距在于解释阶段,模型缺乏解释音乐家的控制的能力。我们还提出了两种策略来解决这一差距,并呼吁音乐信息检索社区解决解释挑战,以改善人类与人工智能的音乐合作。摘要:Large-scale text-to-music generation models have significantly enhanced music creation capabilities, offering unprecedented creative freedom. However, their ability to collaborate effectively with human musicians remains limited. In this paper, we propose a framework to describe the musical interaction process, which includes expression, interpretation, and execution of controls. Following this framework, we argue that the primary gap between existing text-to-music models and musicians lies in the interpretation stage, where models lack the ability to interpret controls from musicians. We also propose two strategies to address this gap and call on the music information retrieval community to tackle the interpretation challenge to improve human-AI musical collaboration.

【10】 CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
标题: CUSIDE-T:基于传感器的流媒体ASB的分块、模拟未来和解码
作者:Wenbo Zhao,Ziwei Li,Chuan Yu,Zhijian Ou
链接:点击下载PDF文件
摘要:流式自动语音识别(ASR)对于许多实际的ASR应用是非常重要的。然而,流传输ASR系统的一个显著挑战在于平衡操作性能与延迟约束。最近,一种分块,模拟未来的上下文和解码的方法,称为CUSIDE,已被提出用于基于连接主义时间分类(CTC)的流式ASR,它在减少延迟和高识别准确率之间获得了很好的平衡。在本文中,我们提出了CUSIDE-T,它成功地适应了CUSIDE方法的递归神经网络转换器(RNN-T)ASR架构,而不是基于CTC架构。我们还在CUSIDE-T中加入了语言模型重评分,以进一步提高准确性,同时只带来很小的额外延迟。在AISHELL-1、WenetSpeech和SpeechIO数据集上进行了大量实验,比较了CUSIDE-T和U2++(两者都基于RNN-T)。U2++是基于块的流式ASR方法的现有对应物。结果表明,CUSIDE-T实现了卓越的准确性性能的流ASR,具有相同的设置的延迟。摘要:Streaming automatic speech recognition (ASR) is very important for many real-world ASR applications. However, a notable challenge for streaming ASR systems lies in balancing operational performance against latency constraint. Recently, a method of chunking, simulating future context and decoding, called CUSIDE, has been proposed for connectionist temporal classification (CTC) based streaming ASR, which obtains a good balance between reduced latency and high recognition accuracy. In this paper, we present CUSIDE-T, which successfully adapts the CUSIDE method over the recurrent neural network transducer (RNN-T) ASR architecture, instead of being based on the CTC architecture. We also incorporate language model rescoring in CUSIDE-T to further enhance accuracy, while only bringing a small additional latency. Extensive experiments are conducted over the AISHELL-1, WenetSpeech and SpeechIO datasets, comparing CUSIDE-T and U2++ (both based on RNN-T). U2++ is an existing counterpart of chunk based streaming ASR method. It is shown that CUSIDE-T achieves superior accuracy performance for streaming ASR, with equal settings of latency.

【11】 Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
标题: 利用帧级嵌入学习系统进行Few-Shot生物声学事件检测
作者:PengYuan Zhao,ChengWei Lu,Liang Zou
链接:点击下载PDF文件
摘要:本技术报告介绍了我们的帧级嵌入学习系统,用于Few-Shot生物声事件检测的DCASE 2024挑战(任务5)。在这项工作中,我们使用log-mel和PCEN对输入音频进行特征提取,Netmamba Encoder作为信息交互网络,并采用数据增强策略来提高训练模型的泛化能力以及多种后处理方法。我们的最终系统实现了56.4%的F-测量分数,在2024年声学场景和事件检测和分类挑战赛的Few-Shot生物声学事件检测类别中获得第二名。摘要:This technical report presents our frame-level embedding learning system for the DCASE2024 challenge for few-shot bioacoustic event detection (Task 5).In this work, we used log-mel and PCEN for feature extraction of the input audio, Netmamba Encoder as the information interaction network, and adopted data augmentation strategies to improve the generalizability of the trained model as well as multiple post-processing methods. Our final system achieved an F-measure score of 56.4%, securing the 2nd rank in the few-shot bioacoustic event detection category of the Detection and Classification of Acoustic Scenes and Events Challenge 2024.

【12】 Whisper-SV: Adapting Whisper for Low-data-resource Speaker Verification
标题: Whisper-SV:将Whisper调整为低数据资源说话者验证
作者:Li Zhang,Ning Jiang,Qing Wang,Yue Li,Quan Lu,Lei Xie
链接:点击下载PDF文件
摘要:经过680,000小时的海量语音数据训练,Whisper是一个多任务,多语言语音基础模型,在自动语音识别,翻译和语言识别方面表现出卓越的性能。然而,它在说话人确认(SV)任务中的适用性仍然没有被探索,特别是在低数据资源的情况下,在特定领域的标记说话人数据是有限的。为了填补这一空白,我们提出了一个轻量级的适配器框架,以提高SV与耳语,即耳语SV。鉴于Whisper不是专门针对SV任务进行优化的,我们引入了一个表示选择模块来量化Whisper每层中包含的特定于说话者的特征,并选择具有突出区分性说话者特征的前k层。为了聚合关键的说话者相关特征,同时减少Whisper中选定的前k个不同层的非说话者冗余,我们在Whisper-SV中设计了一个多层聚合模块,将多层表示集成到SV的单一紧凑表示中。在多层聚合模块中,我们采用不同层之间具有快捷连接的卷积层来细化从Whisper的多层表示中获得的扬声器特征。此外,注意力聚合层用于减少非扬声器干扰并放大SV任务的扬声器特定线索。最后,一个简单的分类模块用于说话人分类。在VoxCeleb 1、FFSVC和IMSV数据集上的实验表明,Whisper-SV分别实现了2.22% 0.307、6.14% 0.488和7.50% 0.582的EER minDCF,在低数据资源SV场景中表现出优异的性能。摘要:Trained on 680,000 hours of massive speech data, Whisper is a multitasking, multilingual speech foundation model demonstrating superior performance in automatic speech recognition, translation, and language identification. However, its applicability in speaker verification (SV) tasks remains unexplored, particularly in low-data-resource scenarios where labeled speaker data in specific domains are limited. To fill this gap, we propose a lightweight adaptor framework to boost SV with Whisper, namely Whisper-SV. Given that Whisper is not specifically optimized for SV tasks, we introduce a representation selection module to quantify the speaker-specific characteristics contained in each layer of Whisper and select the top-k layers with prominent discriminative speaker features. To aggregate pivotal speaker-related features while diminishing non-speaker redundancies across the selected top-k distinct layers of Whisper, we design a multi-layer aggregation module in Whisper-SV to integrate multi-layer representations into a singular, compacted representation for SV. In the multi-layer aggregation module, we employ convolutional layers with shortcut connections among different layers to refine speaker characteristics derived from multi-layer representations from Whisper. In addition, an attention aggregation layer is used to reduce non-speaker interference and amplify speaker-specific cues for SV tasks. Finally, a simple classification module is used for speaker classification. Experiments on VoxCeleb1, FFSVC, and IMSV datasets demonstrate that Whisper-SV achieves EER minDCF of 2.22% 0.307, 6.14% 0.488, and 7.50% 0.582, respectively, showing superior performance in low-data-resource SV scenarios.

【13】 Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
标题: 赋予Whisper作为联合多说话者和目标说话者语音识别系统的能力
作者:Lingwei Meng,Jiawen Kang,Yuejiao Wang,Zengrui Jin,Xixin Wu,Xunying Liu,Helen Meng
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:多说话者语音识别和目标说话者语音识别都涉及多说话者上下文中的转录,仍然是重大的挑战。然而,现有的方法很少尝试同时解决这两个任务。在这项研究中,我们提出了一种开创性的方法来授权Whisper,这是一个语音基础模型,以解决联合多说话者和目标说话者语音识别任务。具体而言,(i)我们冻结Whisper并将Sidecar分离器插入其编码器以分离多个说话者的混合嵌入;(ii)引入目标说话者标识符以实时识别目标说话者的嵌入流,仅需要三秒的注册语音作为提示;(iii)探索解码器的软提示调谐以实现更好的任务适应。我们的方法在两个和三个说话者LibriMix和LibriSpeechMix数据集上的性能优于以前的方法,并在AishellMix普通话数据集上的多说话者ASR上提供了可接受的zero-shot性能。摘要:Multi-talker speech recognition and target-talker speech recognition, both involve transcription in multi-talker contexts, remain significant challenges. However, existing methods rarely attempt to simultaneously address both tasks. In this study, we propose a pioneering approach to empower Whisper, which is a speech foundation model, to tackle joint multi-talker and target-talker speech recognition tasks. Specifically, (i) we freeze Whisper and plug a Sidecar separator into its encoder to separate mixed embedding for multiple talkers; (ii) a Target Talker Identifier is introduced to identify the embedding flow of the target talker on the fly, requiring only three-second enrollment speech as a cue; (iii) soft prompt tuning for decoder is explored for better task adaptation. Our method outperforms previous methods on two- and three-talker LibriMix and LibriSpeechMix datasets for both tasks, and delivers acceptable zero-shot performance on multi-talker ASR on AishellMix Mandarin dataset.

【14】 ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
标题: ecVoice:基于习语相似度替换的音频文本提取和视频优化
作者:Jinwei Lin
备注:APSIPA ASC 2023
链接:点击下载PDF文件
摘要:从视频中提取音频文本在多媒体编辑和处理中有着重要的作用。作为一个流行的开源工具包,Whisper在人类语音识别中执行速度很快。然而,识别性能依赖于计算资源,这使得低计算内存运行Whisper变得困难。本文提出了一种有效的解决方案,从视频中提取人的语音,并获得高质量的文本生成的语音。生成的语音可用于视频语言翻译和翻译语音模拟。为了提高语音文本的提取和转换质量,提出了一种基于习语相似度计算和分析的语音文本提取方法ecVoice。通过相关实验验证,ecVoice能将成语语法正确率平均提高到90%。该方法简单快速,在提高语音识别率的同时,不会造成计算资源消耗的不良影响。我们的方法和解决方案可以显着提高耳语识别与低计算内存。摘要:The Text Extraction of the Audio from the Video plays an important role in multimedia editing and processing. As a popular open source toolkit, Whisper performs fast in human voice recognition. However, the recognition performance is dependent on the computing resource, which makes the low computing memory running Whisper become difficult. Our paper presents an available solution to extract the human voice from the video and gain the high quality text generation from the voice. The generated voice can be used in video language translation and translated voice simulation. To improve the extraction and transform quality of human voice, we present ecVoice, a method using the idioms similarity computation and analysis to improve the quality of audio text extraction. Relative experiments are held to verify that the ecVoice can improve the idiom grammar correction rate to 90 % on average. The method is simple but fast which means this method will cause less bad influence of consuming computing resources when improving the voice recognition rate. Our method and solution can significantly enhance the Whisper recognition with low computing memory.

【15】 Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
标题: 利用多分支深度卷积网络和LSTM-CNN进行心弦分类
作者:Seyed Amir Latifi,Hassan Ghassemian,Maryam Imani
备注:19 pages
链接:点击下载PDF文件
摘要:本文提出了一种快速和成本效益的方法,用于诊断心脏异常,具有高准确性和可靠性,使用低成本的系统在诊所。心脏疾病的自动诊断的主要限制是正确和可接受的标记样品的稀缺性,这可能是昂贵的准备。为了解决这个问题,在这项工作中提出了两种方法。第一种方法是受人类听觉处理启发的独特多分支深度卷积神经网络(MBDCN)架构,专门设计用于通过采用各种大小的卷积滤波器和音频信号功率谱作为输入来优化特征提取。在第二种方法中,称为长短期记忆-卷积神经(LSCN)模型,此外,网络架构包括长短期记忆(LSTM)网络块,以改善时域中的特征提取。将由一维卷积层组成的多个并行分支与LSTM块相结合的创新方法有助于在音频信号处理任务中实现卓越的结果。实验结果表明,所提出的方法优于国家的最先进的技术。LSCN网络对心音的整体分类准确率超过96%。该网络的效率是显着相比,常见的特征提取方法,如梅尔频率倒谱系数(MFCC)和小波变换。因此,该方法在心音信号的自动分析中显示了良好的效果,在心血管疾病的诊断和早期检测中具有潜在的应用价值。摘要:This paper presents a fast and cost-effective method for diagnosing cardiac abnormalities with high accuracy and reliability using low-cost systems in clinics. The primary limitation of automatic diagnosing of cardiac diseases is the rarity of correct and acceptable labeled samples, which can be expensive to prepare. To address this issue, two methods are proposed in this work. The first method is a unique Multi-Branch Deep Convolutional Neural Network (MBDCN) architecture inspired by human auditory processing, specifically designed to optimize feature extraction by employing various sizes of convolutional filters and audio signal power spectrum as input. In the second method, called as Long short-term memory-Convolutional Neural (LSCN) model, Additionally, the network architecture includes Long Short-Term Memory (LSTM) network blocks to improve feature extraction in the time domain. The innovative approach of combining multiple parallel branches consisting of the one-dimensional convolutional layers along with LSTM blocks helps in achieving superior results in audio signal processing tasks. The experimental results demonstrate superiority of the proposed methods over the state-of-the-art techniques. The overall classification accuracy of heart sounds with the LSCN network is more than 96%. The efficiency of this network is significant compared to common feature extraction methods such as Mel Frequency Cepstral Coefficients (MFCC) and wavelet transform. Therefore, the proposed method shows promising results in the automatic analysis of heart sounds and has potential applications in the diagnosis and early detection of cardiovascular diseases.

【16】 Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
标题: 知识提炼过程中不留下任何知识:使用现实数据实现代码转换ASB的实用有效知识提炼
作者:Liang-Hsuan Tseng,Zih-Ching Chen,Wei-Shun Chang,Cheng-Kuang Lee,Tsung-Ren Huang,Hung-yi Lee
链接:点击下载PDF文件
摘要:自动语音识别(ASR)的最新进展通常依赖于大型语音基础模型来生成高质量的译文。然而,由于有限的计算资源,这些模型可能不切实际。在更现实或更困难的情况下,如语码转换ASR(CS-ASR),情况甚至更加严重。为了解决这个问题,我们提出了一个框架,通过知识蒸馏使用现实的语音数据开发更有效的模型CS-ASR。我们提出的方法,在知识蒸馏过程中不留知识(K$^2$D),利用了教师模型的知识和来自一个小辅助模型的额外见解。我们在两个域内和两个域外数据集上评估了我们的方法,证明了K$^2$D是有效的。通过对未标记的真实数据进行K$^2$D,我们成功地获得了一个2倍更小的模型,生成速度快5倍,同时在所有测试集上都优于基线方法和教师模型。我们已经在Hugging Face(https: huggingface.co andybi7676 k2d-whisper.zh-en)上公开了我们的模型。摘要:Recent advances in automatic speech recognition (ASR) often rely on large speech foundation models for generating high-quality transcriptions. However, these models can be impractical due to limited computing resources. The situation is even more severe in terms of more realistic or difficult scenarios, such as code-switching ASR (CS-ASR). To address this, we present a framework for developing more efficient models for CS-ASR through knowledge distillation using realistic speech-only data. Our proposed method, Leave No Knowledge Behind During Knowledge Distillation (K$^2$D), leverages both the teacher model's knowledge and additional insights from a small auxiliary model. We evaluate our approach on two in-domain and two out-domain datasets, demonstrating that K$^2$D is effective. By conducting K$^2$D on the unlabeled realistic data, we have successfully obtained a 2-time smaller model with 5-time faster generation speed while outperforming the baseline methods and the teacher model on all the testing sets. We have made our model publicly available on Hugging Face (https: huggingface.co andybi7676 k2d-whisper.zh-en).

【17】 Advancing Continual Learning for Robust Deepfake Audio Classification
标题: 推进持续学习以实现稳健的Deepfake音频分类
作者:Feiyi Dong,Qingchen Tang,Yichen Bai,Zihan Wang
备注:Submitted to IEEE Tencon. 5 pages
链接:点击下载PDF文件
摘要:新的欺骗攻击的出现对音频安全提出了越来越大的挑战。当前的检测方法在面对看不见的欺骗攻击时往往会动摇。传统的策略,如使用新数据进行再培训,由于存储量大,并不总是可行的。介绍了一种新的连续学习方法--连续音频防御增强器(CADE).首先,通过利用固定的内存大小来存储从以前的数据集中随机选择的样本,我们的方法节省了资源,并遵守隐私约束。此外,我们还在CADE中应用了两个蒸馏损失。通过在分类器中进行蒸馏,CADE确保学生模型与教师模型非常相似。这种相似性有助于模型在面对不可见数据时保留旧信息。我们进一步改进了我们的模型的性能,使用了一种新的嵌入相似性损失,这种相似性损失扩展到多个深度层,促进了优秀的正样本对齐。在ASVspoof 2019数据集上进行的实验表明,我们提出的方法优于基线方法。摘要:The emergence of new spoofing attacks poses an increasing challenge to audio security. Current detection methods often falter when faced with unseen spoofing attacks. Traditional strategies, such as retraining with new data, are not always feasible due to extensive storage. This paper introduces a novel continual learning method Continual Audio Defense Enhancer (CADE). First, by utilizing a fixed memory size to store randomly selected samples from previous datasets, our approach conserves resources and adheres to privacy constraints. Additionally, we also apply two distillation losses in CADE. By distillation in classifiers, CADE ensures that the student model closely resembles that of the teacher model. This resemblance helps the model retain old information while facing unseen data. We further refine our model's performance with a novel embedding similarity loss that extends across multiple depth layers, facilitating superior positive sample alignment. Experiments conducted on the ASVspoof2019 dataset show that our proposed method outperforms the baseline methods.

【18】 Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
标题: Speech-Copilot:通过任务分解、模块化和程序生成利用大型语言模型进行语音处理
作者:Chun-Yi Kuan,Chih-Kai Yang,Wei-Ping Huang,Ke-Han Lu,Hung-yi Lee
备注:8 pages, 2 figures
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了Speech-Copilot,一个模块化的框架,面向语音处理任务,最大限度地减少人类的努力,在工具集建设。与使用大型音频语言模型的端到端方法不同,Speech-Copilot通过分析预先收集的任务指令并将任务分解为可管理的子任务来构建特定于语音处理的工具集。它具有基于大型语言模型的灵活代理,通过程序生成执行任务。我们的方法在Dynamic-SUPERB基准测试中实现了最先进的性能,证明了其在各种语音处理任务中的有效性。主要贡献包括:1)开发一个创新的框架,用于语音处理特定的工具集建设,2)建立一个基于大型语言模型的高性能代理,3)提供一个新的视角来解决具有挑战性的面向语音处理任务。无需端到端方法所需的额外训练过程,我们的方法为广泛的语音处理应用提供了灵活且可扩展的解决方案。摘要:In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot builds speech processing-specific toolsets by analyzing pre-collected task instructions and breaking tasks into manageable sub-tasks. It features a flexible agent based on large language models that performs tasks through program generation. Our approach achieves state-of-the-art performance on the Dynamic-SUPERB benchmark, demonstrating its effectiveness across diverse speech-processing tasks. Key contributions include: 1) developing an innovative framework for speech processing-specific toolset construction, 2) establishing a high-performing agent based on large language models, and 3) offering a new perspective on addressing challenging instruction-oriented speech-processing tasks. Without additional training processes required by end-to-end approaches, our method provides a flexible and extendable solution for a wide range of speech-processing applications.

【19】 Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
标题: 语音斯莱特林:检查曼巴语音分离、识别和合成的性能和效率
作者:Xilin Jiang,Yinghao Aaron Li,Adrian Nicolas Florea,Cong Han,Nima Mesgarani
链接:点击下载PDF文件
摘要:在比较Mamba和Transformers在多个语音相关任务中的性能和效率之前,得出Mamba是Transformers的更好替代品的结论还为时过早。为了得出这一结论,我们提出并评估了三个模型的三个任务:Mamba-TasNet语音分离,ConMamba语音识别,和VALL-M语音合成。我们将它们与性能、内存和速度类似大小的Transformers进行比较。我们的Mamba或Mamba-Transformer混合型号显示出与其Transformer同类产品(Sepformer、Conformer和VALL-E)相当或更高的性能。它们在存储器和速度上比Transformers更有效,用于长于阈值持续时间的语音,与语音令牌的分辨率成反比。用于分离的Mamba算法效率最高,用于识别的Mamba算法效率最低。此外,我们表明,曼巴是不是更有效的比Transformer的语音短于阈值持续时间和执行更差的模型,需要联合建模的文本和语音,如交叉或掩蔽注意两个输入。因此,我们认为,Mamba或Transformer的优越性取决于特定的问题和模型。代码可在https: github.com xi-j Mamba-TasNet和https: github.com xi-j Mamba-ASR上获得。摘要:It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks. To reach this conclusion, we propose and evaluate three models for three tasks: Mamba-TasNet for speech separation, ConMamba for speech recognition, and VALL-M for speech synthesis. We compare them with transformers of similar sizes in performance, memory, and speed. Our Mamba or Mamba-transformer hybrid models show comparable or higher performance than their transformer counterparts: Sepformer, Conformer, and VALL-E. They are more efficient than transformers in memory and speed for speech longer than a threshold duration, inversely related to the resolution of a speech token. Mamba for separation is the most efficient, and Mamba for recognition is the least. Further, we show that Mamba is not more efficient than transformer for speech shorter than the threshold duration and performs worse in models that require joint modeling of text and speech, such as cross or masked attention of two inputs. Therefore, we argue that the superiority of Mamba or transformer depends on particular problems and models. Code available at https: github.com xi-j Mamba-TasNet and https: github.com xi-j Mamba-ASR.


eess.AS音频处理
【1】 Qwen2-Audio Technical Report
标题: Qwen 2-音频技术报告
作者:Yunfei Chu,Jin Xu,Qian Yang,Haojie Wei,Xipin Wei,Zhifang Guo,Yichong Leng,Yuanjun Lv,Jinzheng He,Junyang Lin,Chang Zhou,Jingren Zhou
备注:this https URL Checkpoints, codes and scripts will be opensoursed soon
链接:点击下载PDF文件
摘要:None摘要:We introduce the latest progress of Qwen-Audio, a large-scale audio-language model called Qwen2-Audio, which is capable of accepting various audio signal inputs and performing audio analysis or direct textual responses with regard to speech instructions. In contrast to complex hierarchical tags, we have simplified the pre-training process by utilizing natural language prompts for different data and tasks, and have further expanded the data volume. We have boosted the instruction-following capability of Qwen2-Audio and implemented two distinct audio interaction modes for voice chat and audio analysis. In the voice chat mode, users can freely engage in voice interactions with Qwen2-Audio without text input. In the audio analysis mode, users could provide audio and text instructions for analysis during the interaction. Note that we do not use any system prompts to switch between voice chat and audio analysis modes. Qwen2-Audio is capable of intelligently comprehending the content within audio and following voice commands to respond appropriately. For instance, in an audio segment that simultaneously contains sounds, multi-speaker conversations, and a voice command, Qwen2-Audio can directly understand the command and provide an interpretation and response to the audio. Additionally, DPO has optimized the model's performance in terms of factuality and adherence to desired behavior. According to the evaluation results from AIR-Bench, Qwen2-Audio outperformed previous SOTAs, such as Gemini-1.5-pro, in tests focused on audio-centric instruction-following capabilities. Qwen2-Audio is open-sourced with the aim of fostering the advancement of the multi-modal language community.

【2】 Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
标题: 利用多分支深度卷积网络和LSTM-CNN进行心弦分类
作者:Seyed Amir Latifi,Hassan Ghassemian,Maryam Imani
备注:19 pages
链接:点击下载PDF文件
摘要:本文提出了一种快速和成本效益的方法,用于诊断心脏异常,具有高准确性和可靠性,使用低成本的系统在诊所。心脏疾病的自动诊断的主要限制是正确和可接受的标记样品的稀缺性,这可能是昂贵的准备。为了解决这个问题,在这项工作中提出了两种方法。第一种方法是受人类听觉处理启发的独特多分支深度卷积神经网络(MBDCN)架构,专门设计用于通过采用各种大小的卷积滤波器和音频信号功率谱作为输入来优化特征提取。在第二种方法中,称为长短期记忆-卷积神经(LSCN)模型,此外,网络架构包括长短期记忆(LSTM)网络块,以改善时域中的特征提取。将由一维卷积层组成的多个并行分支与LSTM块相结合的创新方法有助于在音频信号处理任务中实现卓越的结果。实验结果表明,所提出的方法优于国家的最先进的技术。LSCN网络对心音的整体分类准确率超过96%。该网络的效率是显着相比,常见的特征提取方法,如梅尔频率倒谱系数(MFCC)和小波变换。因此,该方法在心音信号的自动分析中显示了良好的效果,在心血管疾病的诊断和早期检测中具有潜在的应用价值。摘要:This paper presents a fast and cost-effective method for diagnosing cardiac abnormalities with high accuracy and reliability using low-cost systems in clinics. The primary limitation of automatic diagnosing of cardiac diseases is the rarity of correct and acceptable labeled samples, which can be expensive to prepare. To address this issue, two methods are proposed in this work. The first method is a unique Multi-Branch Deep Convolutional Neural Network (MBDCN) architecture inspired by human auditory processing, specifically designed to optimize feature extraction by employing various sizes of convolutional filters and audio signal power spectrum as input. In the second method, called as Long short-term memory-Convolutional Neural (LSCN) model, Additionally, the network architecture includes Long Short-Term Memory (LSTM) network blocks to improve feature extraction in the time domain. The innovative approach of combining multiple parallel branches consisting of the one-dimensional convolutional layers along with LSTM blocks helps in achieving superior results in audio signal processing tasks. The experimental results demonstrate superiority of the proposed methods over the state-of-the-art techniques. The overall classification accuracy of heart sounds with the LSCN network is more than 96%. The efficiency of this network is significant compared to common feature extraction methods such as Mel Frequency Cepstral Coefficients (MFCC) and wavelet transform. Therefore, the proposed method shows promising results in the automatic analysis of heart sounds and has potential applications in the diagnosis and early detection of cardiovascular diseases.

【3】 Leave No Knowledge Behind During Knowledge Distillation: Towards Practical and Effective Knowledge Distillation for Code-Switching ASR Using Realistic Data
标题: 知识提炼过程中不留下任何知识:使用现实数据实现代码转换ASB的实用有效知识提炼
作者:Liang-Hsuan Tseng,Zih-Ching Chen,Wei-Shun Chang,Cheng-Kuang Lee,Tsung-Ren Huang,Hung-yi Lee
链接:点击下载PDF文件
摘要:自动语音识别(ASR)的最新进展通常依赖于大型语音基础模型来生成高质量的译文。然而,由于有限的计算资源,这些模型可能不切实际。在更现实或更困难的情况下,如语码转换ASR(CS-ASR),情况甚至更加严重。为了解决这个问题,我们提出了一个框架,通过知识蒸馏使用现实的语音数据开发更有效的模型CS-ASR。我们提出的方法,在知识蒸馏过程中不留知识(K$^2$D),利用了教师模型的知识和来自一个小辅助模型的额外见解。我们在两个域内和两个域外数据集上评估了我们的方法,证明了K$^2$D是有效的。通过对未标记的真实数据进行K$^2$D,我们成功地获得了一个2倍更小的模型,生成速度快5倍,同时在所有测试集上都优于基线方法和教师模型。我们已经在Hugging Face(https: huggingface.co andybi7676 k2d-whisper.zh-en)上公开了我们的模型。摘要:Recent advances in automatic speech recognition (ASR) often rely on large speech foundation models for generating high-quality transcriptions. However, these models can be impractical due to limited computing resources. The situation is even more severe in terms of more realistic or difficult scenarios, such as code-switching ASR (CS-ASR). To address this, we present a framework for developing more efficient models for CS-ASR through knowledge distillation using realistic speech-only data. Our proposed method, Leave No Knowledge Behind During Knowledge Distillation (K$^2$D), leverages both the teacher model's knowledge and additional insights from a small auxiliary model. We evaluate our approach on two in-domain and two out-domain datasets, demonstrating that K$^2$D is effective. By conducting K$^2$D on the unlabeled realistic data, we have successfully obtained a 2-time smaller model with 5-time faster generation speed while outperforming the baseline methods and the teacher model on all the testing sets. We have made our model publicly available on Hugging Face (https: huggingface.co andybi7676 k2d-whisper.zh-en).

【4】 Improving Neural Biasing for Contextual Speech Recognition by Early Context Injection and Text Perturbation
标题: 通过早期上下文注入和文本扰动改善上下文语音识别的神经偏置
作者:Ruizhe Huang,Mahsa Yarmohammadi,Sanjeev Khudanpur,Daniel Povey
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:现有的研究表明,自动语音识别(ASR)模型可以受益于额外的上下文(例如,联系人列表、用户指定词汇表)。通过上下文可以更好地识别稀有词和命名实体。在这项工作中,我们提出了两个简单而有效的技术,以提高上下文感知的ASR模型。首先,我们在早期阶段将上下文注入编码器,而不仅仅是在最后一层。其次,为了在训练过程中强制模型利用上下文,我们用替代拼写扰乱参考转录,以便模型学习依赖上下文来做出正确的预测。在LibriSpeech上,与无偏置和浅融合相比,我们的技术一起将罕见字错误率降低了60%和25%,从而实现了新的最先进的性能。在SPGISpeech和真实世界的数据集ConEC上,我们的技术也比基线有了很好的改进。摘要:Existing research suggests that automatic speech recognition (ASR) models can benefit from additional contexts (e.g., contact lists, user specified vocabulary). Rare words and named entities can be better recognized with contexts. In this work, we propose two simple yet effective techniques to improve context-aware ASR models. First, we inject contexts into the encoders at an early stage instead of merely at their last layers. Second, to enforce the model to leverage the contexts during training, we perturb the reference transcription with alternative spellings so that the model learns to rely on the contexts to make correct predictions. On LibriSpeech, our techniques together reduce the rare word error rate by 60% and 25% relatively compared to no biasing and shallow fusion, making the new state-of-the-art performance. On SPGISpeech and a real-world dataset ConEC, our techniques also yield good improvements over the baselines.

【5】 Advancing Continual Learning for Robust Deepfake Audio Classification
标题: 推进持续学习以实现稳健的Deepfake音频分类
作者:Feiyi Dong,Qingchen Tang,Yichen Bai,Zihan Wang
备注:Submitted to IEEE Tencon. 5 pages
链接:点击下载PDF文件
摘要:新的欺骗攻击的出现对音频安全提出了越来越大的挑战。当前的检测方法在面对看不见的欺骗攻击时往往会动摇。传统的策略,如使用新数据进行再培训,由于存储量大,并不总是可行的。介绍了一种新的连续学习方法--连续音频防御增强器(CADE).首先,通过利用固定的内存大小来存储从以前的数据集中随机选择的样本,我们的方法节省了资源,并遵守隐私约束。此外,我们还在CADE中应用了两个蒸馏损失。通过在分类器中进行蒸馏,CADE确保学生模型与教师模型非常相似。这种相似性有助于模型在面对不可见数据时保留旧信息。我们进一步改进了我们的模型的性能,使用了一种新的嵌入相似性损失,这种相似性损失扩展到多个深度层,促进了优秀的正样本对齐。在ASVspoof 2019数据集上进行的实验表明,我们提出的方法优于基线方法。摘要:The emergence of new spoofing attacks poses an increasing challenge to audio security. Current detection methods often falter when faced with unseen spoofing attacks. Traditional strategies, such as retraining with new data, are not always feasible due to extensive storage. This paper introduces a novel continual learning method Continual Audio Defense Enhancer (CADE). First, by utilizing a fixed memory size to store randomly selected samples from previous datasets, our approach conserves resources and adheres to privacy constraints. Additionally, we also apply two distillation losses in CADE. By distillation in classifiers, CADE ensures that the student model closely resembles that of the teacher model. This resemblance helps the model retain old information while facing unseen data. We further refine our model's performance with a novel embedding similarity loss that extends across multiple depth layers, facilitating superior positive sample alignment. Experiments conducted on the ASVspoof2019 dataset show that our proposed method outperforms the baseline methods.

【6】 The feasibility of sound zone control using an array of parametric array loudspeakers
标题: 使用参数阵列扬声器阵列进行音区控制的可行性
作者:Tao Zhuang,Jia-Xin Zhong,Jing Lu
链接:点击下载PDF文件
摘要:参数阵列扬声器(PAL)已知用于产生高度定向的音频波束,这是用传统电动扬声器(EDL)实现的更具挑战性的壮举。由于其固有的物理机制,PAL在虚拟现实(VR)等空间音频应用中具有很大的潜力。然而,使用PAL阵列进行声区控制(SZC)的可行性仍然没有被探索,这主要是由于PAL中固有的非线性解调过程的复杂性。利用PAL建模的最新进展,这项工作提出了一种优化算法,以实现使用PAL阵列的两个目标区域之间的声学对比度控制(ACC)。通过仿真研究了所提出的基于ACC的使用PAL阵列的SZC的性能和鲁棒性,并将结果与使用EDL阵列获得的结果进行了比较。结果表明,PAL阵列在较高频率和较低信噪比下的SZC性能和鲁棒性优于EDL阵列,而在其他条件下相当。这项工作为使用PAL阵列进行高对比度声学控制铺平了道路。摘要:Parametric array loudspeakers (PALs) are known for producing highly directional audio beams, a feat more challenging to achieve with conventional electro-dynamic loudspeakers (EDLs). Due to their intrinsic physical mechanisms, PALs hold promising potential for spatial audio applications such as virtual reality (VR). However, the feasibility of using an array of PALs for sound zone control (SZC) has remained unexplored, mainly due to the complexity of the nonlinear demodulation process inherent in PALs. Leveraging recent advancements in PAL modeling, this work proposes an optimization algorithm to achieve the acoustic contrast control (ACC) between two target areas using a PAL array. The performance and robustness of the proposed ACC-based SZC using PAL arrays are investigated through simulations, and the results are compared with those obtained using EDL arrays. The results show that the PAL array outperforms the EDL array in SZC performance and robustness at higher frequencies and lower signal-to-noise ratio, while being comparable under other conditions. This work paves the way for high-contrast acoustic control using PAL arrays.

【7】 Speech-Copilot: Leveraging Large Language Models for Speech Processing via Task Decomposition, Modularization, and Program Generation
标题: Speech-Copilot:通过任务分解、模块化和程序生成利用大型语言模型进行语音处理
作者:Chun-Yi Kuan,Chih-Kai Yang,Wei-Ping Huang,Ke-Han Lu,Hung-yi Lee
备注:8 pages, 2 figures
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了Speech-Copilot,一个模块化的框架,面向语音处理任务,最大限度地减少人类的努力,在工具集建设。与使用大型音频语言模型的端到端方法不同,Speech-Copilot通过分析预先收集的任务指令并将任务分解为可管理的子任务来构建特定于语音处理的工具集。它具有基于大型语言模型的灵活代理,通过程序生成执行任务。我们的方法在Dynamic-SUPERB基准测试中实现了最先进的性能,证明了其在各种语音处理任务中的有效性。主要贡献包括:1)开发一个创新的框架,用于语音处理特定的工具集建设,2)建立一个基于大型语言模型的高性能代理,3)提供一个新的视角来解决具有挑战性的面向语音处理任务。无需端到端方法所需的额外训练过程,我们的方法为广泛的语音处理应用提供了灵活且可扩展的解决方案。摘要:In this work, we introduce Speech-Copilot, a modular framework for instruction-oriented speech-processing tasks that minimizes human effort in toolset construction. Unlike end-to-end methods using large audio-language models, Speech-Copilot builds speech processing-specific toolsets by analyzing pre-collected task instructions and breaking tasks into manageable sub-tasks. It features a flexible agent based on large language models that performs tasks through program generation. Our approach achieves state-of-the-art performance on the Dynamic-SUPERB benchmark, demonstrating its effectiveness across diverse speech-processing tasks. Key contributions include: 1) developing an innovative framework for speech processing-specific toolset construction, 2) establishing a high-performing agent based on large language models, and 3) offering a new perspective on addressing challenging instruction-oriented speech-processing tasks. Without additional training processes required by end-to-end approaches, our method provides a flexible and extendable solution for a wide range of speech-processing applications.

【8】 A Streaming Multi-Channel End-to-End Speech Recognition System with Realistic Evaluations
标题: 具有真实评估的流媒体多通道端到端语音识别系统
作者:Xiangzhu Kong,Tianqi Ning,Hao Huang,Zhijian Ou
链接:点击下载PDF文件
摘要:最近出现了多通道端到端(ME2E)ASR系统。虽然流式单通道端到端ASR已被广泛研究,但流式ME2E ASR在探索方面受到限制。此外,最近的研究呼吁关注分销(ID)和分销(OOD)测试之间的差距,并进行现实的评估。本文主要研究两个问题:实现流媒体ME2E ASR和提高面向对象设计的通用性。我们提出了CUSIDE阵列方法,它集成了最近的CUSIDE方法(分块,模拟未来上下文和解码)到ME2E ASR的神经波束形成器方法。它支持前端和后端的流处理,总延迟为402 ms。CUSIDE阵列的ME2E模型显示,以实现优异的流结果在ID和OOD测试。现实的评估证实了CUSIDE阵列的优势,它能够消耗单通道数据,通过后端预训练和ME2E微调来提高OOD的泛化能力。摘要:Recently multi-channel end-to-end (ME2E) ASR systems have emerged. While streaming single-channel end-to-end ASR has been extensively studied, streaming ME2E ASR is limited in exploration. Additionally, recent studies call attention to the gap between in-distribution (ID) and out-of-distribution (OOD) tests and doing realistic evaluations. This paper focuses on two research problems: realizing streaming ME2E ASR and improving OOD generalization. We propose the CUSIDE-array method, which integrates the recent CUSIDE methodology (Chunking, Simulating Future Context and Decoding) into the neural beamformer approach of ME2E ASR. It enables streaming processing of both front-end and back-end with a total latency of 402ms. The CUSIDE-array ME2E models are shown to achieve superior streaming results in both ID and OOD tests. Realistic evaluations confirm the advantage of CUSIDE-array in its capability to consume single-channel data to improve OOD generalization via back-end pre-training and ME2E fine-tuning.

【9】 Speech Slytherin: Examining the Performance and Efficiency of Mamba for Speech Separation, Recognition, and Synthesis
标题: 语音斯莱特林:检查曼巴语音分离、识别和合成的性能和效率
作者:Xilin Jiang,Yinghao Aaron Li,Adrian Nicolas Florea,Cong Han,Nima Mesgarani
链接:点击下载PDF文件
摘要:在比较Mamba和Transformers在多个语音相关任务中的性能和效率之前,得出Mamba是Transformers的更好替代品的结论还为时过早。为了得出这一结论,我们提出并评估了三个模型的三个任务:Mamba-TasNet语音分离,ConMamba语音识别,和VALL-M语音合成。我们将它们与性能、内存和速度类似大小的Transformers进行比较。我们的Mamba或Mamba-Transformer混合型号显示出与其Transformer同类产品(Sepformer、Conformer和VALL-E)相当或更高的性能。它们在存储器和速度上比Transformers更有效,用于长于阈值持续时间的语音,与语音令牌的分辨率成反比。用于分离的Mamba算法效率最高,用于识别的Mamba算法效率最低。此外,我们表明,曼巴是不是更有效的比Transformer的语音短于阈值持续时间和执行更差的模型,需要联合建模的文本和语音,如交叉或掩蔽注意两个输入。因此,我们认为,Mamba或Transformer的优越性取决于特定的问题和模型。代码可在https: github.com xi-j Mamba-TasNet和https: github.com xi-j Mamba-ASR上获得。摘要:It is too early to conclude that Mamba is a better alternative to transformers for speech before comparing Mamba with transformers in terms of both performance and efficiency in multiple speech-related tasks. To reach this conclusion, we propose and evaluate three models for three tasks: Mamba-TasNet for speech separation, ConMamba for speech recognition, and VALL-M for speech synthesis. We compare them with transformers of similar sizes in performance, memory, and speed. Our Mamba or Mamba-transformer hybrid models show comparable or higher performance than their transformer counterparts: Sepformer, Conformer, and VALL-E. They are more efficient than transformers in memory and speed for speech longer than a threshold duration, inversely related to the resolution of a speech token. Mamba for separation is the most efficient, and Mamba for recognition is the least. Further, we show that Mamba is not more efficient than transformer for speech shorter than the threshold duration and performs worse in models that require joint modeling of text and speech, such as cross or masked attention of two inputs. Therefore, we argue that the superiority of Mamba or transformer depends on particular problems and models. Code available at https: github.com xi-j Mamba-TasNet and https: github.com xi-j Mamba-ASR.

【10】 CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
标题: CUSIDE-T:基于传感器的流媒体ASB的分块、模拟未来和解码
作者:Wenbo Zhao,Ziwei Li,Chuan Yu,Zhijian Ou
链接:点击下载PDF文件
摘要:流式自动语音识别(ASR)对于许多实际的ASR应用是非常重要的。然而,流传输ASR系统的一个显著挑战在于平衡操作性能与延迟约束。最近,一种分块,模拟未来的上下文和解码的方法,称为CUSIDE,已被提出用于基于连接主义时间分类(CTC)的流式ASR,它在减少延迟和高识别准确率之间获得了很好的平衡。在本文中,我们提出了CUSIDE-T,它成功地适应了CUSIDE方法的递归神经网络转换器(RNN-T)ASR架构,而不是基于CTC架构。我们还在CUSIDE-T中加入了语言模型重评分,以进一步提高准确性,同时只带来很小的额外延迟。在AISHELL-1、WenetSpeech和SpeechIO数据集上进行了大量实验,比较了CUSIDE-T和U2++(两者都基于RNN-T)。U2++是基于块的流式ASR方法的现有对应物。结果表明,CUSIDE-T实现了卓越的准确性性能的流ASR,具有相同的设置的延迟。摘要:Streaming automatic speech recognition (ASR) is very important for many real-world ASR applications. However, a notable challenge for streaming ASR systems lies in balancing operational performance against latency constraint. Recently, a method of chunking, simulating future context and decoding, called CUSIDE, has been proposed for connectionist temporal classification (CTC) based streaming ASR, which obtains a good balance between reduced latency and high recognition accuracy. In this paper, we present CUSIDE-T, which successfully adapts the CUSIDE method over the recurrent neural network transducer (RNN-T) ASR architecture, instead of being based on the CTC architecture. We also incorporate language model rescoring in CUSIDE-T to further enhance accuracy, while only bringing a small additional latency. Extensive experiments are conducted over the AISHELL-1, WenetSpeech and SpeechIO datasets, comparing CUSIDE-T and U2++ (both based on RNN-T). U2++ is an existing counterpart of chunk based streaming ASR method. It is shown that CUSIDE-T achieves superior accuracy performance for streaming ASR, with equal settings of latency.

【11】 Few-Shot Bioacoustic Event Detection with Frame-Level Embedding Learning System
标题: 利用帧级嵌入学习系统进行Few-Shot生物声学事件检测
作者:PengYuan Zhao,ChengWei Lu,Liang Zou
链接:点击下载PDF文件
摘要:本技术报告介绍了我们的帧级嵌入学习系统,用于Few-Shot生物声事件检测的DCASE 2024挑战(任务5)。在这项工作中,我们使用log-mel和PCEN对输入音频进行特征提取,Netmamba Encoder作为信息交互网络,并采用数据增强策略来提高训练模型的泛化能力以及多种后处理方法。我们的最终系统实现了56.4%的F-测量分数,在2024年声学场景和事件检测和分类挑战赛的Few-Shot生物声学事件检测类别中获得第二名。摘要:This technical report presents our frame-level embedding learning system for the DCASE2024 challenge for few-shot bioacoustic event detection (Task 5).In this work, we used log-mel and PCEN for feature extraction of the input audio, Netmamba Encoder as the information interaction network, and adopted data augmentation strategies to improve the generalizability of the trained model as well as multiple post-processing methods. Our final system achieved an F-measure score of 56.4%, securing the 2nd rank in the few-shot bioacoustic event detection category of the Detection and Classification of Acoustic Scenes and Events Challenge 2024.

【12】 Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
标题: 赋予Whisper作为联合多说话者和目标说话者语音识别系统的能力
作者:Lingwei Meng,Jiawen Kang,Yuejiao Wang,Zengrui Jin,Xixin Wu,Xunying Liu,Helen Meng
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:多说话者语音识别和目标说话者语音识别都涉及多说话者上下文中的转录,仍然是重大的挑战。然而,现有的方法很少尝试同时解决这两个任务。在这项研究中,我们提出了一种开创性的方法来授权Whisper,这是一个语音基础模型,以解决联合多说话者和目标说话者语音识别任务。具体而言,(i)我们冻结Whisper并将Sidecar分离器插入其编码器以分离多个说话者的混合嵌入;(ii)引入目标说话者标识符以实时识别目标说话者的嵌入流,仅需要三秒的注册语音作为提示;(iii)探索解码器的软提示调谐以实现更好的任务适应。我们的方法在两个和三个说话者LibriMix和LibriSpeechMix数据集上的性能优于以前的方法,并在AishellMix普通话数据集上的多说话者ASR上提供了可接受的zero-shot性能。摘要:Multi-talker speech recognition and target-talker speech recognition, both involve transcription in multi-talker contexts, remain significant challenges. However, existing methods rarely attempt to simultaneously address both tasks. In this study, we propose a pioneering approach to empower Whisper, which is a speech foundation model, to tackle joint multi-talker and target-talker speech recognition tasks. Specifically, (i) we freeze Whisper and plug a Sidecar separator into its encoder to separate mixed embedding for multiple talkers; (ii) a Target Talker Identifier is introduced to identify the embedding flow of the target talker on the fly, requiring only three-second enrollment speech as a cue; (iii) soft prompt tuning for decoder is explored for better task adaptation. Our method outperforms previous methods on two- and three-talker LibriMix and LibriSpeechMix datasets for both tasks, and delivers acceptable zero-shot performance on multi-talker ASR on AishellMix Mandarin dataset.

【13】 ecVoice: Audio Text Extraction and Optimization of Video Based on Idioms Similarity Replacement
标题: ecVoice:基于习语相似度替换的音频文本提取和视频优化
作者:Jinwei Lin
备注:APSIPA ASC 2023
链接:点击下载PDF文件
摘要:从视频中提取音频文本在多媒体编辑和处理中有着重要的作用。作为一个流行的开源工具包,Whisper在人类语音识别中执行速度很快。然而,识别性能依赖于计算资源,这使得低计算内存运行Whisper变得困难。本文提出了一种有效的解决方案,从视频中提取人的语音,并获得高质量的文本生成的语音。生成的语音可用于视频语言翻译和翻译语音模拟。为了提高语音文本的提取和转换质量,提出了一种基于习语相似度计算和分析的语音文本提取方法ecVoice。通过相关实验验证,ecVoice能将成语语法正确率平均提高到90%。该方法简单快速,在提高语音识别率的同时,不会造成计算资源消耗的不良影响。我们的方法和解决方案可以显着提高耳语识别与低计算内存。摘要:The Text Extraction of the Audio from the Video plays an important role in multimedia editing and processing. As a popular open source toolkit, Whisper performs fast in human voice recognition. However, the recognition performance is dependent on the computing resource, which makes the low computing memory running Whisper become difficult. Our paper presents an available solution to extract the human voice from the video and gain the high quality text generation from the voice. The generated voice can be used in video language translation and translated voice simulation. To improve the extraction and transform quality of human voice, we present ecVoice, a method using the idioms similarity computation and analysis to improve the quality of audio text extraction. Relative experiments are held to verify that the ecVoice can improve the idiom grammar correction rate to 90 % on average. The method is simple but fast which means this method will cause less bad influence of consuming computing resources when improving the voice recognition rate. Our method and solution can significantly enhance the Whisper recognition with low computing memory.


机器翻译,仅供参考