今日论文合集:cs.SD语音32篇,eess.AS音频处理47篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps
标题: 通过量化感知训练和二进制激活地图实现资源高效的语音质量预测
作者:Mattias Nilsson,Riccardo Miccini,Clément Laroche,Tobias Piechowiak,Friedemann Zenke
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:随着移动和边缘设备中的语音处理系统变得越来越普遍,对非侵入式语音质量监控的需求也在增加。深度学习方法提供了对客观和主观语音质量指标的高质量估计。然而,其显著的计算要求通常在资源受限的设备上是禁止的。为了解决这个问题,我们研究了基于DNSMOS的卷积架构上的语音质量预测的二进制激活映射(BAM)。我们表明,量化感知训练的二进制激活模型与基线模型的预测性能相匹配。它还允许使用其他压缩技术。结合8位权重量化,我们的方法在推理过程中减少了25倍的内存,同时用求和代替了几乎所有的点积。我们的研究结果显示了一条通过在硬件和软件中支持混合精度二进制乘法来节省大量资源的途径。摘要:As speech processing systems in mobile and edge devices become more commonplace, the demand for unintrusive speech quality monitoring increases. Deep learning methods provide high-quality estimates of objective and subjective speech quality metrics. However, their significant computational requirements are often prohibitive on resource-constrained devices. To address this issue, we investigated binary activation maps (BAMs) for speech quality prediction on a convolutional architecture based on DNSMOS. We show that the binary activation model with quantization aware training matches the predictive performance of the baseline model. It further allows using other compression techniques. Combined with 8-bit weight quantization, our approach results in a 25-fold memory reduction during inference, while replacing almost all dot products with summations. Our findings show a path toward substantial resource savings by supporting mixed-precision binary multiplication in hard- and software.

【2】 Real-time Timbre Remapping with Differentiable DSP
标题: 利用差异化的DSP实现实时音色重映射
作者:Jordie Shier,Charalampos Saitis,Andrew Robertson,Andrew McPherson
备注:Accepted for publication at the 24th International Conference on New Interfaces for Musical Expression in Utrecht, Netherlands
链接:点击下载PDF文件
摘要:音色是不同音乐背景下的主要表达方式。然而,流行的音频驱动合成方法主要依赖于音调和响度包络,有效地平整了输入的音色表达。我们的方法借鉴了音色类比的概念,并研究了如何将输入信号的音色表达映射到合成器的控制上。利用可微分数字信号处理,我们的方法通过新颖的特征差异损失促进了合成器参数的直接优化。这个损失函数旨在学习音乐事件之间的相对音色差异,优先考虑短语内分级音色调制的微妙之处,允许在音色空间中进行有意义的翻译。使用小军鼓表演作为一个案例研究,其中音色表达是中央,我们展示了实时音色重新映射从声学小军鼓到罗兰TR-808后建模的可区分合成器。摘要:Timbre is a primary mode of expression in diverse musical contexts. However, prevalent audio-driven synthesis methods predominantly rely on pitch and loudness envelopes, effectively flattening timbral expression from the input. Our approach draws on the concept of timbre analogies and investigates how timbral expression from an input signal can be mapped onto controls for a synthesizer. Leveraging differentiable digital signal processing, our method facilitates direct optimization of synthesizer parameters through a novel feature difference loss. This loss function, designed to learn relative timbral differences between musical events, prioritizes the subtleties of graded timbre modulations within phrases, allowing for meaningful translations in a timbre space. Using snare drum performances as a case study, where timbral expression is central, we demonstrate real-time timbre remapping from acoustic snare drums to a differentiable synthesizer modeled after the Roland TR-808.

【3】 Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
标题: 突尼斯方言低资源SL和ASB语音编码器的性能分析
作者:Salima Mdhaffar,Haroun Elleuch,Fethi Bougares,Yannick Estève
备注:Accepted in ArabicNLP 2024
链接:点击下载PDF文件
摘要:通过自监督学习(SSL)预训练的语音编码器在各种下游任务中表现出卓越的性能,包括口语理解(SLU)和自动语音识别(ASR)。例如,针对此类任务微调SSL模型已显示出巨大的潜力,从而提高了SOTA在具有挑战性的数据集上的性能。与现有的研究相比,本文的贡献是通过比较SSL方法的有效性的上下文中(一)低资源的突尼斯阿拉伯语方言口语和(二)其与低资源SLU和ASR的情况下,只有少数语义注释可用于微调相结合。我们使用许多SSL语音编码器在TARIC-SLU数据集上进行实验。我们使用的语音编码器是在单语或多语言语音数据上预先训练的。其中一些也被完善,没有在域,也没有突尼斯的数据,通过多模态监督师生范式。这项研究产生了许多重要的发现,我们在本文中讨论。摘要:Speech encoders pretrained through self-supervised learning (SSL) have demonstrated remarkable performance in various downstream tasks, including Spoken Language Understanding (SLU) and Automatic Speech Recognition (ASR). For instance, fine-tuning SSL models for such tasks has shown significant potential, leading to improvements in the SOTA performance across challenging datasets. In contrast to existing research, this paper contributes by comparing the effectiveness of SSL approaches in the context of (i) the low-resource spoken Tunisian Arabic dialect and (ii) its combination with a low-resource SLU and ASR scenario, where only a few semantic annotations are available for fine-tuning. We conduct experiments using many SSL speech encoders on the TARIC-SLU dataset. We use speech encoders that were pre-trained on either monolingual or multilingual speech data. Some of them have also been refined without in-domain nor Tunisian data through multimodal supervised teacher-student paradigm. This study yields numerous significant findings that we are discussing in this paper.

【4】 Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
标题: 控制耳语:控制语音基础模型的通用声学对抗攻击
作者:Vyas Raina,Mark Gales
链接:点击下载PDF文件
摘要:支持语音的基础模型,无论是基于灵活语音识别的系统还是音频提示的大型语言模型(LLM),都变得越来越流行。这些模型的一个有趣的方面是,它们能够使用适当的提示执行自动语音识别(ASR)以外的任务。例如,OpenAI Whisper模型可以执行语音转录和语音翻译。随着音频提示LLM的发展,有可能提供更大的控制选项。在这项工作中,我们证明了这种更大的灵活性,系统可以容易受到模型控制对抗攻击。在没有对模型提示的任何访问的情况下,可以通过适当地改变音频输入来修改系统的行为。为了说明这种风险,我们证明,它是可能的prepend一个短的通用对抗性的声学段的任何输入语音信号覆盖的ASR基础模型的提示设置。具体来说,我们成功地使用了一个通用的对抗性声学段来控制Whisper始终执行语音翻译,尽管它被设置为执行语音转录。总的来说,这项工作展示了一种新形式的对抗性攻击,对支持多任务语音的基础模型进行攻击,需要在部署这种形式的模型之前加以考虑。摘要:Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these models is their ability to perform tasks other than automatic speech recognition (ASR) using an appropriate prompt. For example, the OpenAI Whisper model can perform both speech transcription and speech translation. With the development of audio-prompted LLMs there is the potential for even greater control options. In this work we demonstrate that with this greater flexibility the systems can be susceptible to model-control adversarial attacks. Without any access to the model prompt it is possible to modify the behaviour of the system by appropriately changing the audio input. To illustrate this risk, we demonstrate that it is possible to prepend a short universal adversarial acoustic segment to any input speech signal to override the prompt setting of an ASR foundation model. Specifically, we successfully use a universal adversarial acoustic segment to control Whisper to always perform speech translation, despite being set to perform speech transcription. Overall, this work demonstrates a new form of adversarial attack on multi-tasking speech enabled foundation models that needs to be considered prior to the deployment of this form of model.

【5】 TokenVerse: Unifying Speech and NLP Tasks via Transducer-based ASR
标题: TokenVerse:通过基于传感器的ASB统一语音和NLP任务
作者:Shashi Kumar,Srikanth Madikeri,Juan Zuluaga-Gomez,Iuliia Nigmatulina,Esaú Villatoro-Tello,Sergio Burdisso,Petr Motlicek,Karthik Pandia,Aravind Ganapathiraju
备注:5 pages, double column
链接:点击下载PDF文件
摘要:在传统的语音会话智能中,使用级联管道,涉及语音活动检测,日记,转录等任务,以及针对语义端点和命名实体识别(NER)等任务的不同NLP模型的后续处理。我们的论文介绍了TokenVerse,这是一个基于单个传感器的模型,旨在处理多个任务。这是通过在ASR模型训练期间将特定于任务的标记集成到参考文本中来实现的,简化了推理并消除了对单独的NLP模型的需求。除了ASR,我们进行实验3个不同的任务:说话人变化检测,端点,和NER。我们在公共和私有数据集上的实验表明,该方法在相对WER上将ASR提高了7.7%,同时在单个任务性能上优于级联管道方法。此外,我们提出了任务迁移学习到现有TokenVerse中的新任务。摘要:In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent processing with different NLP models for tasks like semantic endpointing and named entity recognition (NER). Our paper introduces TokenVerse, a single Transducer-based model designed to handle multiple tasks. This is achieved by integrating task-specific tokens into the reference text during ASR model training, streamlining the inference and eliminating the need for separate NLP models. In addition to ASR, we conduct experiments on 3 different tasks: speaker change detection, endpointing, and NER. Our experiments on a public and a private dataset show that the proposed method improves ASR by up to 7.7% in relative WER while outperforming the cascaded pipeline approach in individual task performance. Additionally, we present task transfer learning to a new task within an existing TokenVerse.

【6】 Improving Audio Generation with Visual Enhanced Caption
标题: 使用视觉增强字幕改进音频生成
作者:Yi Yuan,Dongya Jia,Xiaobin Zhuang,Yuanzhe Chen,Zhengxi Liu,Zhuo Chen,Yuping Wang,Yuxuan Wang,Xubo Liu,Mark D. Plumbley,Wenwu Wang
备注:5 pages with 1 appendix
链接:点击下载PDF文件
摘要:生成模型已经在音频生成任务中显示出显著的成就。然而,现有的模型难以处理复杂而详细的提示,导致潜在的性能下降。我们假设这个问题源于低质量和相对少量的训练数据。在这项工作中,我们的目标是创建一个具有丰富字幕的大规模音频数据集,以改进音频生成模型。我们开发了一个自动化的管道,通过使用大型语言模型(LLM)将预测的视觉字幕,音频字幕和标记标签转换为全面的描述,为视听数据集生成详细的字幕。我们引入Sound-VECaps,这是一个包含1.66 M高质量音频字幕对的数据集,其中包含丰富的细节,包括音频事件顺序,发生地点和环境信息。我们证明,使用Sound-VECaps进行训练可以显着增强文本到音频生成模型的能力,以便从复杂的输入提示中理解和生成音频,从而提高整体系统性能。此外,我们在几个音频语言任务中进行声音VECaps的消融研究,表明其在推进音频文本表征学习中的潜力。我们的数据集和模型可在线获取。摘要:Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the low quality and relatively small quantity of training data. In this work, we aim to create a large-scale audio dataset with rich captions for improving audio generation models. We develop an automated pipeline to generate detailed captions for audio-visual datasets by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). We introduce Sound-VECaps, a dataset comprising 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We demonstrate that training with Sound-VECaps significantly enhances the capability of text-to-audio generation models to comprehend and generate audio from complex input prompts, improving overall system performance. Furthermore, we conduct ablation studies of Sound-VECaps across several audio-language tasks, suggesting its potential in advancing audio-text representation learning. Our dataset and models are available online.

【7】 A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
标题: 使用艺术材料与潜在音频合成交互的映射策略
作者:Shuoyang Zheng,Anna Xambó Sedó,Nick Bryan-Kinns
链接:点击下载PDF文件
摘要:本文提出了一种与生成式AI模型的潜在空间交互的映射策略。我们的方法涉及使用无监督特征学习来编码人类控制空间,并将其映射到音频合成模型的潜在空间。为了演示这种映射策略如何将高维传感器数据转化为深度生成模型的控制机制,我们提出了一个概念验证系统,该系统使用视觉草图来控制音频合成模型。我们借鉴XAIxArts中新兴的话语,讨论这种方法如何在艺术和创意背景下为XAI做出贡献,我们还讨论了它目前的局限性,并提出了未来的研究方向。摘要:This paper presents a mapping strategy for interacting with the latent spaces of generative AI models. Our approach involves using unsupervised feature learning to encode a human control space and mapping it to an audio synthesis model's latent space. To demonstrate how this mapping strategy can turn high-dimensional sensor data into control mechanisms of a deep generative model, we present a proof-of-concept system that uses visual sketches to control an audio synthesis model. We draw on emerging discourses in XAIxArts to discuss how this approach can contribute to XAI in artistic and creative contexts, we also discuss its current limitations and propose future research directions.

【8】 Romanization Encoding For Multilingual ASR
标题: 多语言ASB的罗马化编码
作者:Wen Ding,Fei Jia,Hainan Xu,Yu Xi,Junjie Lai,Boris Ginsburg
链接:点击下载PDF文件
摘要:我们为脚本密集的语言引入罗马化编码,以优化多语言和代码切换自动语音识别(ASR)系统。通过在配备Roman 2Char模块的FastConformer-RNNT框架中采用罗马化编码以及平衡的级联标记器,我们显着减少了词汇量和输出维度,从而实现了更大的训练批次并减少了内存消耗。该方法将声学建模和语言建模相结合,增强了系统的灵活性和适应性。在我们的研究中,将这种方法应用于普通话-英语ASR导致了显着的63.51%的词汇减少和显着的性能提高13.72%和15.03%的SEAME代码转换基准。对普通话-韩语和普通话-日语的消融研究突出了我们的方法解决其他脚本繁重语言的复杂性的强大能力,为更通用和有效的多语言ASR系统铺平了道路。摘要:We introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated tokenizer within a FastConformer-RNNT framework equipped with a Roman2Char module, we significantly reduce vocabulary and output dimensions, enabling larger training batches and reduced memory consumption. Our method decouples acoustic modeling and language modeling, enhancing the flexibility and adaptability of the system. In our study, applying this method to Mandarin-English ASR resulted in a remarkable 63.51% vocabulary reduction and notable performance gains of 13.72% and 15.03% on SEAME code-switching benchmarks. Ablation studies on Mandarin-Korean and Mandarin-Japanese highlight our method's strong capability to address the complexities of other script-heavy languages, paving the way for more versatile and effective multilingual ASR systems.

【9】 PAGURI: a user experience study of creative interaction with text-to-music models
标题: PAGuri:与文本到音乐模型的创造性交互的用户体验研究
作者:Francesca Ronchini,Luca Comanducci,Gabriele Perego,Fabio Antonacci
链接:点击下载PDF文件
摘要:近年来,文本到音乐模型是自动音乐生成的最大突破。虽然它们无疑是技术进步的展示,但目前尚不清楚它们如何切实融入音乐家和音乐从业者的艺术实践。本文旨在通过即时音频生成用户研究调查(PAGURI)来解决这个问题,PAGURI是一项用户体验研究,我们利用最近的文本到音乐的发展来研究音乐家和从业者如何与这些系统进行交互,评估他们的满意度。我们开发了一个在线工具,用户可以通过它生成音乐样本和 或应用最近提出的个性化技术,基于微调,使文本到音乐模型生成更接近他们的需求和偏好的声音。使用问卷调查,我们分析了参与者如何与所提出的工具进行交互,以了解文本到音乐模型在增强用户创造力方面的有效性。结果表明,即使生成的音频样本及其质量可能并不总是满足用户的期望,大多数参与者将在他们的创作过程中使用该工具。此外,他们还深入了解了该系统的潜在增强功能及其与音乐实践的整合。摘要:In recent years, text-to-music models have been the biggest breakthrough in automatic music generation. While they are unquestionably a showcase of technological progress, it is not clear yet how they can be realistically integrated into the artistic practice of musicians and music practitioners. This paper aims to address this question via Prompt Audio Generation User Research Investigation (PAGURI), a user experience study where we leverage recent text-to-music developments to study how musicians and practitioners interact with these systems, evaluating their satisfaction levels. We developed an online tool through which users can generate music samples and or apply recently proposed personalization techniques, based on fine-tuning, to make the text-to-music model generate sounds closer to their needs and preferences. Using questionnaires, we analyzed how participants interacted with the proposed tool, to understand the effectiveness of text-to-music models in enhancing users' creativity. Results show that even if the audio samples generated and their quality may not always meet user expectations, the majority of the participants would incorporate the tool in their creative process. Furthermore, they provided insights into potential enhancements for the system and its integration into their music practice.

【10】 MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
标题: MuseBarControl:通过预训练和反事实损失增强象征性音乐生成中的细粒度控制
作者:Yangyang Shu,Haiming Xu,Ziqin Zhou,Anton van den Hengel,Lingqiao Liu
备注:Demo is available at: this https URL
链接:点击下载PDF文件
摘要:自动生成符号音乐--根据人类特定需求定制的乐谱--对音乐家和爱好者来说是非常有益的。最近的研究表明,使用广泛的数据集和先进的Transformer架构,结果很有希望。然而,这些最先进的模型通常只提供对整个作品的节奏和风格等方面的基本控制,缺乏管理更精细细节的能力,例如在单个小节级别的控制。虽然微调预训练的符号音乐生成模型似乎是实现这种更精细控制的简单方法,但我们的研究表明这种方法存在挑战。该模型往往不能充分响应新的,细粒度的酒吧级控制信号。为此,我们提出了两个创新的解决方案。首先,我们引入了一个预训练任务,旨在将控制信号直接与相应的音乐令牌联系起来,这有助于实现更有效的初始化,以便随后进行微调。其次,我们实现了一种新的反事实损失,促进生成的音乐和控制提示之间更好的对齐。总之,这些技术显着提高了我们的能力,控制音乐生成在酒吧的水平,显示了13.06%的改进,比传统的方法。我们的主观评估也证实了这种增强的控制不会损害原始预训练生成模型的音乐质量。摘要:Automatically generating symbolic music-music scores tailored to specific human needs-can be highly beneficial for musicians and enthusiasts. Recent studies have shown promising results using extensive datasets and advanced transformer architectures. However, these state-of-the-art models generally offer only basic control over aspects like tempo and style for the entire composition, lacking the ability to manage finer details, such as control at the level of individual bars. While fine-tuning a pre-trained symbolic music generation model might seem like a straightforward method for achieving this finer control, our research indicates challenges in this approach. The model often fails to respond adequately to new, fine-grained bar-level control signals. To address this, we propose two innovative solutions. First, we introduce a pre-training task designed to link control signals directly with corresponding musical tokens, which helps in achieving a more effective initialization for subsequent fine-tuning. Second, we implement a novel counterfactual loss that promotes better alignment between the generated music and the control prompts. Together, these techniques significantly enhance our ability to control music generation at the bar level, showing a 13.06 % improvement over conventional methods. Our subjective evaluations also confirm that this enhanced control does not compromise the musical quality of the original pre-trained generative model.

【11】 Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
标题: 在线发言人拨号系统的延迟性系统评估
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages
链接:点击下载PDF文件
摘要:在本文中,不同的在线发言人日记系统进行了评估,在相同的硬件与相同的测试数据方面的延迟。延迟是从音频输入到对应扬声器标签的输出的时间跨度。作为评估的一部分,DIART框架内的各种模型组合,基于在线聚类算法UIS-RNN-SML的日记系统,和端到端的在线日记系统FS-EEND进行了比较。使用嵌入模型pyannote embedding和分割模型pyannote segmentation,DIART管道实现了最低的延迟。FS-EEND系统显示出类似的良好延迟。一般来说,目前还没有发表的研究,比较几个在线日记系统的延迟。这使得这项工作更加相关。摘要:In this paper, different online speaker diarization systems are evaluated on the same hardware with the same test data with regard to their latency. The latency is the time span from audio input to the output of the corresponding speaker label. As part of the evaluation, various model combinations within the DIART framework, a diarization system based on the online clustering algorithm UIS-RNN-SML, and the end-to-end online diarization system FS-EEND are compared. The lowest latency is achieved for the DIART-pipeline with the embedding model pyannote embedding and the segmentation model pyannote segmentation. The FS-EEND system shows a similarly good latency. In general there is currently no published research that compares several online diarization systems in terms of their latency. This makes this work even more relevant.

【12】 LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
标题: LearnerVoice:非英语母语学习者自发言语的数据集
作者:Haechan Kim,Junho Myung,Seoyoung Kim,Sungpah Lee,Dongyeop Kang,Juho Kim
备注:Accepted for INTERSPEECH 2024
链接:点击下载PDF文件
摘要:第二语言学习者自发言语中普遍存在的不符合语法的表达和不流利现象给自动语音识别系统带来了独特的挑战。然而,很少有数据集是针对L2学习者语音的。我们公开发布了LearnerVoice,这是一个由50.04小时的音频和L2学习者自发语音组成的数据集。我们的语言学分析表明,我们的数据集中的transanterior包含L2 S(L2学习者的自发语音)特征,包括不符合语法的表达和不流利(例如,填充词、单词重复、自我修复、错误开始),显著多于母语数据集。使用LearnerVoice微调whisper-small.en实现了10.26%的WER,比vanilla whisper-small. en低44.2%。此外,我们的定性分析表明,54.2%的错误从香草模型的LearnerVoice归因于L2 S功能,其中48.1%被减少在微调模型。摘要:Prevalent ungrammatical expressions and disfluencies in spontaneous speech from second language (L2) learners pose unique challenges to Automatic Speech Recognition (ASR) systems. However, few datasets are tailored to L2 learner speech. We publicly release LearnerVoice, a dataset consisting of 50.04 hours of audio and transcriptions of L2 learners' spontaneous speech. Our linguistic analysis reveals that transcriptions in our dataset contain L2S (L2 learner's Spontaneous speech) features, consisting of ungrammatical expressions and disfluencies (e.g., filler words, word repetitions, self-repairs, false starts), significantly more than native speech datasets. Fine-tuning whisper-small.en with LearnerVoice achieves a WER of 10.26%, 44.2% lower than vanilla whisper-small.en. Furthermore, our qualitative analysis indicates that 54.2% of errors from the vanilla model on LearnerVoice are attributable to L2S features, with 48.1% of them being reduced in the fine-tuned model.

【13】 BiosERC: Integrating Biography Speakers Supported by LLMs for ERC Tasks
标题: BiosERC:集成由LLM支持的传记演讲者来执行ERC任务
作者:Jieying Xue,Minh Phuong Nguyen,Blake Matheny,Le Minh Nguyen
备注:Accepted in the 33rd International Conference on Artificial Neural Networks (ICANN 2024)
链接:点击下载PDF文件
摘要:在会话中的情感识别任务中,最近的研究利用注意机制探索来自内部和内部说话者的话语之间的关系,以建模它们之间的情感交互。然而,属性,如扬声器的个性特征仍然未被探索,并提出了挑战,他们的适用性,以其他任务或兼容性与不同的模型架构。因此,这项工作引入了一个新的框架名为BiosERC,它调查说话人的特点在对话中。通过采用大型语言模型(LLM),我们提取的“传记信息”的谈话中的扬声器作为补充知识注入到模型中,为每个话语的情感标签进行分类。我们提出的方法在三个著名的基准数据集上取得了最先进的(SOTA)结果:IEMOCAP,MELD和EmoryNLP,证明了我们模型的有效性和通用性,并展示了其适应各种会话分析任务的潜力。我们的源代码可以在https: github.com yingjie7 BiosERC上找到。摘要:In the Emotion Recognition in Conversation task, recent investigations have utilized attention mechanisms exploring relationships among utterances from intra- and inter-speakers for modeling emotional interaction between them. However, attributes such as speaker personality traits remain unexplored and present challenges in terms of their applicability to other tasks or compatibility with diverse model architectures. Therefore, this work introduces a novel framework named BiosERC, which investigates speaker characteristics in a conversation. By employing Large Language Models (LLMs), we extract the "biographical information" of the speaker within a conversation as supplementary knowledge injected into the model to classify emotional labels for each utterance. Our proposed method achieved state-of-the-art (SOTA) results on three famous benchmark datasets: IEMOCAP, MELD, and EmoryNLP, demonstrating the effectiveness and generalization of our model and showcasing its potential for adaptation to various conversation analysis tasks. Our source code is available at https: github.com yingjie7 BiosERC.

【14】 FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
标题: FunAudioLLM:人类与LLM之间自然互动的语音理解和生成基础模型
作者:Tongyi SpeechTeam
备注:work in progress
链接:点击下载PDF文件
摘要:本报告介绍了FunAudioLLM,一个旨在增强人类与大型语言模型(LLM)之间的自然语音交互的模型系列。其核心是两个创新模型:SenseVoice,处理多语言语音识别,情感识别和音频事件检测; CosyVoice,通过控制多种语言,音色,说话风格和扬声器身份来促进自然语音生成。SenseVoice-Small可为5种语言提供极低延迟的ASR,SenseVoice-Large支持50多种语言的高精度ASR,而CosyVoice在多语言语音生成、zero-shot上下文学习、跨语言语音克隆和语音跟踪功能方面表现出色。与SenseVoice和CosyVoice相关的模型已经在Modelscope和Huggingface上开源,同时在GitHub上发布了相应的训练,推理和微调代码。通过将这些模型与LLM集成,FunAudioLLM支持语音到语音翻译、情感语音聊天、交互式播客和富有表现力的有声读物叙事等应用,从而推动了语音交互技术的发展。演示可在https: fun-audio-llm.github.io上获得,代码可在https: github.com FunAudioLLM上访问。摘要:This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https: fun-audio-llm.github.io, and the code can be accessed at https: github.com FunAudioLLM.

【15】 Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
标题: 使用基于无监督文本到语音合成的数据增强来改进重读语音识别
作者:Cong-Thanh Do,Shuhei Imai,Rama Doddipatla,Thomas Hain
备注:Accepted to EUSIPCO 2024
链接:点击下载PDF文件
摘要:本文研究了使用无监督的文本到语音合成(TTS)作为一种数据增强方法,以提高口音语音识别。TTS系统是用少量的重音语音训练数据及其伪标签而不是手动训练来训练的,因此是无监督的。这种方法使得能够使用带口音的语音数据,而无需手动转录来执行带口音的语音识别的数据增强。通过使用TTS系统从文本提示生成的合成重音语音数据然后与可用的非重音语音数据组合以训练自动语音识别(ASR)系统。ASR实验在自监督学习框架中使用Wav2vec2.0模型进行,该模型在大量无监督口音语音数据上进行预训练。用于训练无监督TTS的口音语音数据是从L2-ARCTIC和不列颠群岛语料库中选择的朗读语音,而来自爱丁堡国际口音英语语料库的自发会话语音被用作评估数据。实验结果表明,Wav2vec2.0模型,微调到下游ASR任务与合成重音语音数据,由无监督TTS生成,产生高达6.1%的相对字错误率降低相比,Wav2vec2.0基线,这是微调与非重音语音数据从Librispeech语料库。摘要:This paper investigates the use of unsupervised text-to-speech synthesis (TTS) as a data augmentation method to improve accented speech recognition. TTS systems are trained with a small amount of accented speech training data and their pseudo-labels rather than manual transcriptions, and hence unsupervised. This approach enables the use of accented speech data without manual transcriptions to perform data augmentation for accented speech recognition. Synthetic accented speech data, generated from text prompts by using the TTS systems, are then combined with available non-accented speech data to train automatic speech recognition (ASR) systems. ASR experiments are performed in a self-supervised learning framework using a Wav2vec2.0 model which was pre-trained on large amount of unsupervised accented speech data. The accented speech data for training the unsupervised TTS are read speech, selected from L2-ARCTIC and British Isles corpora, while spontaneous conversational speech from the Edinburgh international accents of English corpus are used as the evaluation data. Experimental results show that Wav2vec2.0 models which are fine-tuned to downstream ASR task with synthetic accented speech data, generated by the unsupervised TTS, yield up to 6.1% relative word error rate reductions compared to a Wav2vec2.0 baseline which is fine-tuned with the non-accented speech data from Librispeech corpus.

【16】 Serialized Output Training by Learned Dominance
标题: 通过习得优势进行系列输出训练
作者:Ying Shi,Lantian Li,Shi Yin,Dong Wang,Jiqing Han
备注:accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:串行输出训练(SOT)通过顺序解码单个说话人的语音,在多说话人语音识别中展示了最先进的性能。为了解决具有挑战性的标签排列问题,现有方法依赖于排列不变训练(PIT)或基于时间的先进先出(FIFO)规则。本研究提出了一种基于模型的序列化策略,将一个辅助模块的注意力编码器-解码器架构,自主识别的关键因素,以订购的输出序列的语音组件在多说话人的语音。在LibriSpeech和LibriMix数据库上进行的实验表明,我们的方法在2-mix和3-mix场景中的性能明显优于PIT和FIFO基线。进一步的分析表明,序列化模块识别的因素,包括响度和性别的混合中的主导语音成分,并根据优势得分的语音成分的顺序。摘要:Serialized Output Training (SOT) has showcased state-of-the-art performance in multi-talker speech recognition by sequentially decoding the speech of individual speakers. To address the challenging label-permutation issue, prior methods have relied on either the Permutation Invariant Training (PIT) or the time-based First-In-First-Out (FIFO) rule. This study presents a model-based serialization strategy that incorporates an auxiliary module into the Attention Encoder-Decoder architecture, autonomously identifying the crucial factors to order the output sequence of the speech components in multi-talker speech. Experiments conducted on the LibriSpeech and LibriMix databases reveal that our approach significantly outperforms the PIT and FIFO baselines in both 2-mix and 3-mix scenarios. Further analysis shows that the serialization module identifies dominant speech components in a mixture by factors including loudness and gender, and orders speech components based on the dominance score.

【17】 On the Effectiveness of Acoustic BPE in Decoder-Only TTS
标题: 声学BPE在纯解码器的TTC中的有效性
作者:Bohan Li,Feiyu Shen,Yiwei Guo,Shuai Wang,Xie Chen,Kai Yu
备注:5 pages, 3 tables, 1 figures. accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:将语音离散化并由解码器模型生成语音标记是文本到语音(TTS)和口语建模(SLM)的一个很有前途的方向。为了缩短语音标记的序列长度,声学字节对编码(BPE)已经出现在SLM中,其将来自自监督语义表示的语音标记作为字符来进一步压缩标记序列。但是TTS中的增益还没有得到充分的研究,并且声学BPE的正确选择仍然不清楚。在这项工作中,我们进行了全面的研究,各种设置的声学BPE,以探讨其有效性解码器只有TTS模型与语义语音令牌。LibriTTS上的实验验证了声学BPE均匀地增加了合成语音的可懂度和多样性,同时在BPE设置中显示出不同的特征。因此,声学BPE是用于仅解码器TTS的有利工具。摘要:Discretizing speech into tokens and generating them by a decoder-only model have been a promising direction for text-to-speech (TTS) and spoken language modeling (SLM). To shorten the sequence length of speech tokens, acoustic byte-pair encoding (BPE) has emerged in SLM that treats speech tokens from self-supervised semantic representations as characters to further compress the token sequence. But the gain in TTS has not been fully investigated, and the proper choice of acoustic BPE remains unclear. In this work, we conduct a comprehensive study on various settings of acoustic BPE to explore its effectiveness in decoder-only TTS models with semantic speech tokens. Experiments on LibriTTS verify that acoustic BPE uniformly increases the intelligibility and diversity of synthesized speech, while showing different features across BPE settings. Hence, acoustic BPE is a favorable tool for decoder-only TTS.

【18】 Unsupervised speech enhancement with spectral kurtosis and double deep priors
标题: 具有谱峰度和双重深先验的无监督语音增强
作者:Hien Ohnaka,Ryoichi Miyazaki
备注:11 pages, 12 figures, and 2 Tables, submitted to Acoustical Science and Technology
链接:点击下载PDF文件
摘要:本文提出了一种基于深度先验(DP)的无监督DNN语音增强方法。这里,DP表示DNN比噪声更倾向于产生干净的语音信号。基于DP的常规方法通常涉及使用随机噪声特征作为输入对有噪声的语音信号进行训练,停止训练,仅生成干净的语音信号。然而,此些常规方法在确定最佳停止时序方面遇到挑战,经历归因于环境背景噪声的性能降级,且遭受干净语音信号的失真与噪声减少性能之间的折衷。为了解决这些挑战,我们利用两个DNN:一个用于生成干净的语音信号,另一个用于生成噪声。这些网络的组合输出非常接近有噪语音信号,并利用基于谱峭度的损失项将有噪语音信号分离为干净的语音信号和噪声。该方法的关键优势在于它能够规避权衡和早期停止问题,因为信号被足够的步骤分解。通过评估实验,我们证明该方法在高斯白噪声和环境噪声的情况下优于传统方法,同时有效地减轻了早期停止问题。摘要:This paper proposes an unsupervised DNN-based speech enhancement approach founded on deep priors (DPs). Here, DP signifies that DNNs are more inclined to produce clean speech signals than noises. Conventional methods based on DP typically involve training on a noisy speech signal using a random noise feature as input, stopping training only a clean speech signal is generated. However, such conventional approaches encounter challenges in determining the optimal stop timing, experience performance degradation due to environmental background noise, and suffer a trade-off between distortion of the clean speech signal and noise reduction performance. To address these challenges, we utilize two DNNs: one to generate a clean speech signal and the other to generate noise. The combined output of these networks closely approximates the noisy speech signal, with a loss term based on spectral kurtosis utilized to separate the noisy speech signal into a clean speech signal and noise. The key advantage of this method lies in its ability to circumvent trade-offs and early stopping problems, as the signal is decomposed by enough steps. Through evaluation experiments, we demonstrate that the proposed method outperforms conventional methods in the case of white Gaussian and environmental noise while effectively mitigating early stopping problems.

【19】 Semantic Grouping Network for Audio Source Separation
标题: 用于音频源分离的语义队列网络
作者:Shentong Mo,Yapeng Tian
链接:点击下载PDF文件
摘要:最近,视听分离方法已经利用两种模态之间的自然同步来提高音频源分离性能。他们从视觉输入中提取高级语义作为指导,以帮助解开单个来源的声音表示。我们能直接学会从声音本身中分离出个体语义吗?困境在于,多个声源在原始空间中混合在一起。为了解决这个问题,在本文中,我们提出了一种新的语义网络,称为SGN,可以直接解开声音表示和提取高层次的语义信息,为每个源从输入的音频混合。具体来说,SGN通过声音的可学习类令牌聚合类别方面的源特征。然后,聚合的语义特征可以用作从混合物中分离相应音频源的指导。我们对纯音乐和通用声音分离基准进行了广泛的实验:MUSIC,FUSS,MUSDB 18和VGG-Sound。结果表明,我们的SGN显着优于以前的音频方法和视听模型,而不利用额外的视觉线索。摘要:Recently, audio-visual separation approaches have taken advantage of the natural synchronization between the two modalities to boost audio source separation performance. They extracted high-level semantics from visual inputs as the guidance to help disentangle sound representation for individual sources. Can we directly learn to disentangle the individual semantics from the sound itself? The dilemma is that multiple sound sources are mixed together in the original space. To tackle the difficulty, in this paper, we present a novel Semantic Grouping Network, termed as SGN, that can directly disentangle sound representations and extract high-level semantic information for each source from input audio mixture. Specifically, SGN aggregates category-wise source features through learnable class tokens of sounds. Then, the aggregated semantic features can be used as the guidance to separate the corresponding audio sources from the mixture. We conducted extensive experiments on music-only and universal sound separation benchmarks: MUSIC, FUSS, MUSDB18, and VGG-Sound. The results demonstrate that our SGN significantly outperforms previous audio-only methods and audio-visual models without utilizing additional visual cues.

【20】 Improving Self-supervised Pre-training using Accent-Specific Codebooks
标题: 使用特定口音的代码簿改进自我监督的预训练
作者:Darshan Prabhu,Abhishek Gupta,Omkar Nitsure,Preethi Jyothi,Sriram Ganapathy
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音口音对最先进的端到端自动语音识别(ASR)系统的性能提出了严峻的挑战。即使使用ASR模型的自监督学习和预训练,也很少实现重音不变性。在这项工作中,我们提出了一种口音感知自适应技术的自监督学习,引入了一组可训练的口音特定的码本的自监督架构。这些可学习的码本使模型能够在预训练期间捕获口音特定信息,并在ASR微调期间进一步细化。在Mozilla Common Voice数据集上,我们提出的方法在可见和不可见的英语口音上都优于所有其他口音适应方法,单词错误率(WER)相对降低了9%。摘要:Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER).

【21】 Multi-Convformer: Extending Conformer with Multiple Convolution Kernels
标题: Multi-Conformer:用多个卷积核扩展Conformer
作者:Darshan Prabhu,Yifan Peng,Preethi Jyothi,Shinji Watanabe
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:卷积由于其对局部上下文的有效建模而成为最先进的端到端自动语音识别(ASR)系统中必不可少的。值得注意的是,它在Conformers中的使用导致了与基于vanilla Transformer的ASR系统相比更优越的性能。虽然Conformer中卷积模块以外的组件已经被重新检查,但更改卷积模块本身的探索要少得多。为此,我们引入了Multi-Convformer,它在Conformer的卷积模块中使用多个卷积内核,并结合门控。这有助于改进不同粒度的局部依赖关系的建模。我们的模型在性能上可与现有的Conformer变体(如CgMLP和E-Branchformer)相媲美,同时具有更高的参数效率。我们在四个不同的数据集和三个不同的建模范式上实证比较了我们的方法与Conformer及其变体,并显示出高达8%的相对单词错误率(WER)改善。摘要:Convolutions have become essential in state-of-the-art end-to-end Automatic Speech Recognition~(ASR) systems due to their efficient modelling of local context. Notably, its use in Conformers has led to superior performance compared to vanilla Transformer-based ASR systems. While components other than the convolution module in the Conformer have been reexamined, altering the convolution module itself has been far less explored. Towards this, we introduce Multi-Convformer that uses multiple convolution kernels within the convolution module of the Conformer in conjunction with gating. This helps in improved modeling of local dependencies at varying granularities. Our model rivals existing Conformer variants such as CgMLP and E-Branchformer in performance, while being more parameter efficient. We empirically compare our approach with Conformer and its variants across four different datasets and three different modelling paradigms and show up to 8% relative word error rate~(WER) improvements.

【22】 Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
标题: 多语言ASB系统自回归解码器的持续学习优化
作者:Chin Yuen Kwok,Jia Qi Yip,Eng Siong Chng
链接:点击下载PDF文件
摘要:持续学习(CL)涉及使用新数据微调预训练模型,同时保持预训练数据的性能。这对于扩展多语言ASR(MASR)功能尤其重要。然而,现有的CL方法,主要是为计算机视觉和强化学习任务设计的,当直接应用于MASR时,往往会产生次优的结果。我们假设这是因为自回归解码器在MASR模型中的CL是困难的。为了验证这一点,我们提出了四个优化的解码器。它们包括解码器层梯度手术,冻结未使用的令牌嵌入,抑制新添加的令牌的输出,以及学习率重新缩放。我们将Whisper适配到Common Voice数据集中的10种看不见的语言的实验表明,与Experience Replay相比,这些优化将预训练语言的平均单词错误率(AWER)从14.2%降低到12.4%,而不会影响新语言的AWER。摘要:Continual Learning (CL) involves fine-tuning pre-trained models with new data while maintaining the performance on the pre-trained data. This is particularly relevant for expanding multilingual ASR (MASR) capabilities. However, existing CL methods, mainly designed for computer vision and reinforcement learning tasks, often yield sub-optimal results when directly applied to MASR. We hypothesise that this is because CL of the auto-regressive decoder in the MASR model is difficult. To verify this, we propose four optimizations on the decoder. They include decoder-layer gradient surgery, freezing unused token embeddings, suppressing output of newly added tokens, and learning rate re-scaling. Our experiments on adapting Whisper to 10 unseen languages from the Common Voice dataset demonstrate that these optimizations reduce the Average Word Error Rate (AWER) of pretrained languages from 14.2% to 12.4% compared with Experience Replay, without compromising the AWER of new languages.

【23】 Towards Attention-based Contrastive Learning for Audio Spoof Detection
标题: 基于注意力的对比学习用于音频欺骗检测
作者:Chirag Goel,Surya Koppisetti,Ben Colman,Ali Shahriyari,Gaurav Bharaj
备注:Proc. INTERSPEECH 2023
链接:点击下载PDF文件
摘要:Vision Transformers(ViT)在计算机视觉分类任务方面取得了实质性进展。最近,Gong et.等人'21,引入了用于若干音频任务的基于注意力的建模。然而,相对未开发的是使用ViT进行音频欺骗检测任务。我们弥合了这一差距,并为此任务引入了ViT。建立在微调SSAST基础上的香草基线(Gong et.’22)音频ViT模型实现了次优等错误率(EER)。为了提高性能,我们提出了一种新的基于注意力的对比学习框架(SSAST-CL),使用交叉注意力来帮助表征学习。实验表明,我们的框架成功地解开的善意和欺骗类,并帮助学习更好的分类任务。通过适当的数据增强策略,在我们的框架上训练的模型在ASVSpoof 2021挑战中取得了有竞争力的表现。我们提供了比较和消融研究来证明我们的说法。摘要:Vision transformers (ViT) have made substantial progress for classification tasks in computer vision. Recently, Gong et. al. '21, introduced attention-based modeling for several audio tasks. However, relatively unexplored is the use of a ViT for audio spoof detection task. We bridge this gap and introduce ViTs for this task. A vanilla baseline built on fine-tuning the SSAST (Gong et. al. '22) audio ViT model achieves sub-optimal equal error rates (EERs). To improve performance, we propose a novel attention-based contrastive learning framework (SSAST-CL) that uses cross-attention to aid the representation learning. Experiments show that our framework successfully disentangles the bonafide and spoof classes and helps learn better classifiers for the task. With appropriate data augmentations policy, a model trained on our framework achieves competitive performance on the ASVSpoof 2021 challenge. We provide comparisons and ablation studies to justify our claim.

【24】 Prosody-Driven Privacy-Preserving Dementia Detection
标题: 韵律驱动的隐私保护痴呆症检测
作者:Dominika Woszczyk,Ranya Aloufi,Soteris Demetriou
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:从语音记录中提取的说话人嵌入已被证明是有价值的痴呆症检测。然而,就其性质而言,这些嵌入包含可识别的信息,这引起了隐私问题。在这项工作中,我们的目标是匿名嵌入,同时保留痴呆症检测的诊断实用程序。以前的研究依赖于对抗性学习和在有限资源环境中训练目标属性和斗争的模型。我们提出了一种新的方法,利用领域知识来解开韵律功能相关的痴呆症从扬声器嵌入,而不依赖于痴呆症分类。我们的实验表明,我们的方法在保护说话人隐私(说话人识别F1-得分.01%),同时保持高痴呆症检测分数F1-得分74%的ADReSS数据集的有效性。我们的结果也与ADReSSo上更受约束的分类器依赖系统(0.01%和0.66%)相当,并且对合成语音自然度没有影响。摘要:Speaker embeddings extracted from voice recordings have been proven valuable for dementia detection. However, by their nature, these embeddings contain identifiable information which raises privacy concerns. In this work, we aim to anonymize embeddings while preserving the diagnostic utility for dementia detection. Previous studies rely on adversarial learning and models trained on the target attribute and struggle in limited-resource settings. We propose a novel approach that leverages domain knowledge to disentangle prosody features relevant to dementia from speaker embeddings without relying on a dementia classifier. Our experiments show the effectiveness of our approach in preserving speaker privacy (speaker recognition F1-score .01%) while maintaining high dementia detection score F1-score of 74% on the ADReSS dataset. Our results are also on par with a more constrained classifier-dependent system on ADReSSo (.01% and .66%), and have no impact on synthesized speech naturalness.

【25】 Advanced Framework for Animal Sound Classification With Features Optimization
标题: 具有特征优化的动物声音分类高级框架
作者:Qiang Yang,Xiuying Chen,Changsheng Ma,Carlos M. Duarte,Xiangliang Zhang
链接:点击下载PDF文件
摘要:动物声音的自动分类在生物声学中提出了一个持久的挑战,由于声音信号的不同统计特性,记录设备的变化,以及普遍的低信噪比(SNR)条件。像卷积神经网络(CNN)和长短期记忆(LSTM)这样的深度学习模型在人类语音识别方面表现出色,但尚未有效地针对动物声音的复杂性质进行定制,即使在同一领域内也表现出很大的多样性。我们提出了一个自动分类框架适用于一般动物的声音分类。我们的方法首先从梅尔频率倒谱系数(MFCC)优化音频特征,包括特征重排和特征约简。然后,它使用深度学习模型的优化特征,即,基于注意力的双向LSTM(Bi-LSTM),用于提取声音分类的深层语义特征。我们还提供了一个动物声音基准数据集,包括海洋动物和鸟类1。对真实世界数据集的广泛实验表明,我们的方法在精确度、召回率和准确度方面始终优于基线方法25%以上,这在动物声音分类方面取得了可喜的进步。摘要:The automatic classification of animal sounds presents an enduring challenge in bioacoustics, owing to the diverse statistical properties of sound signals, variations in recording equipment, and prevalent low Signal-to-Noise Ratio (SNR) conditions. Deep learning models like Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) have excelled in human speech recognition but have not been effectively tailored to the intricate nature of animal sounds, which exhibit substantial diversity even within the same domain. We propose an automated classification framework applicable to general animal sound classification. Our approach first optimizes audio features from Mel-frequency cepstral coefficients (MFCC) including feature rearrangement and feature reduction. It then uses the optimized features for the deep learning model, i.e., an attention-based Bidirectional LSTM (Bi-LSTM), to extract deep semantic features for sound classification. We also contribute an animal sound benchmark dataset encompassing oceanic animals and birds1. Extensive experimentation with real-world datasets demonstrates that our approach consistently outperforms baseline methods by over 25% in precision, recall, and accuracy, promising advancements in animal sound classification.

【26】 PianoBART: Symbolic Piano Music Generation and Understanding with Large-Scale Pre-Training
标题: PianoBART:通过大规模预训练进行象征性钢琴音乐的生成和理解
作者:Xiao Liang,Zijian Zhao,Weichao Zeng,Yutong He,Fupeng He,Yiyi Wang,Chengying Gao
链接:点击下载PDF文件
摘要:学习音乐结构和作曲模式对于音乐生成和理解都是必要的,但是目前的方法没有统一使用学习到的特征来同时生成和理解音乐。在本文中,我们提出了PianoBART,这是一个预先训练的模型,它使用BART来生成和理解象征性的钢琴音乐。针对PianoBART的不同预训练任务,设计了一种多层次的对象选择策略,可以防止信息泄漏或丢失,提高学习能力。在预训练中捕获的音乐语义针对音乐生成和理解任务进行了微调。实验表明,PianoBART有效地学习音乐模式,并在生成高质量的连贯作品和理解音乐方面取得了出色的表现。我们的代码和补充材料可在https: github.com RS2002 PianoBart上获得。摘要:Learning musical structures and composition patterns is necessary for both music generation and understanding, but current methods do not make uniform use of learned features to generate and comprehend music simultaneously. In this paper, we propose PianoBART, a pre-trained model that uses BART for both symbolic piano music generation and understanding. We devise a multi-level object selection strategy for different pre-training tasks of PianoBART, which can prevent information leakage or loss and enhance learning ability. The musical semantics captured in pre-training are fine-tuned for music generation and understanding tasks. Experiments demonstrate that PianoBART efficiently learns musical patterns and achieves outstanding performance in generating high-quality coherent pieces and comprehending music. Our code and supplementary material are available at https: github.com RS2002 PianoBart.

【27】 Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
标题: Seed-ASB:利用基于LLM的语音识别来理解不同的语音和上下文
作者:Ye Bai,Jingping Chen,Jitong Chen,Wei Chen,Zhuo Chen,Chen Ding,Linhao Dong,Qianqian Dong,Yujiao Du,Kepan Gao,Lu Gao,Yi Guo,Minglun Han,Ting Han,Wenchao Hu,Xinying Hu,Yuxiang Hu,Deyu Hua,Lu Huang,Mingkun Huang,Youjia Huang,Jishuo Jin,Fanliu Kong,Zongwei Lan,Tianyu Li,Xiaoyang Li,Zeyang Li,Zehua Lin,Rui Liu,Shouda Liu,Lu Lu,Yizhou Lu,Jingting Ma,Shengtao Ma,Yulin Pei,Chen Shen,Tian Tan,Xiaogang Tian,Ming Tu,Bo Wang,Hao Wang,Yuping Wang,Yuxuan Wang,Hanzhang Xia,Rui Xia,Shuangyi Xie,Hongmin Xu,Meng Yang,Bihong Zhang,Jun Zhang,Wanyi Zhang,Yang Zhang,Yawei Zhang,Yijie Zheng,Ming Zou
链接:点击下载PDF文件
摘要:现代自动语音识别(ASR)模型需要在各种应用场景中准确地转录给定特定上下文信息的各种语音信号(来自不同领域、语言、口音等)。融合了额外语言模型的经典端到端模型表现良好,但主要是在数据匹配场景中,并逐渐接近瓶颈。在这项工作中,我们介绍了种子ASR,一个大的语言模型(LLM)为基础的语音识别模型。Seed-ASR是基于音频条件LLM(AcLLM)的框架开发的,通过将连续语音表示与上下文信息一起输入到LLM中来利用LLM的能力。通过分阶段的大规模训练和LLM中上下文感知能力的启发,Seed-ASR在综合评估集(包括多个领域,口音 方言和语言)上展示了端到端模型的显着改进。此外,Seed-ASR可以进一步部署,以支持各种场景中的特定需求,而无需额外的语言模型。与最近发布的大型ASR模型相比,Seed-ASR在中文和英文公共测试集上的单词(或字符,对于中文)错误率降低了10%-40%,进一步证明了其强大的性能。摘要:Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.

【28】 Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
标题: 谁觉得这个声音很有吸引力?使用野外数据的大规模实验
作者:Hitoshi Suda,Aya Watanabe,Shinnosuke Takamichi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了CocoNut-Humoresque,一个开源的大规模语音喜欢度语料库,包括语音片段和每个听众的喜欢度分数。评价声音的可爱性对于为语音系统(如对话或公告系统)设计更好的声音至关重要。在这项研究中,我们让885名听众对1800个演讲片段进行了评价,这些演讲片段来自各种各样的演讲者。在构建语料库时,我们还收集了多个说话者属性:性别,年龄和最喜欢的YouTube视频。因此,语料库使大规模的统计分析的声音喜欢关于扬声器和听众的因素。本文描述了构建方法和初步的数据分析,以揭示性别和年龄偏见的声音喜欢。此外,讨喜度和两个声学特征,基频和给定的话语的x-向量之间的关系,也进行了研究。摘要:This paper introduces CocoNut-Humoresque, an open-source large-scale speech likability corpus that includes speech segments and their per-listener likability scores. Evaluating voice likability is essential to designing preferable voices for speech systems, such as dialogue or announcement systems. In this study, we let 885 listeners rate 1800 speech segments of a wide range of speakers regarding their likability. When constructing the corpus, we also collected the multiple speaker attributes: genders, ages, and favorite YouTube videos. Therefore, the corpus enables the large-scale statistical analysis of voice likability regarding both speaker and listener factors. This paper describes the construction methodology and preliminary data analysis to reveal the gender and age biases in voice likability. In addition, the relationship between the likability and two acoustic features, the fundamental frequencies and the x-vectors of given utterances, is also investigated.

【29】 Configurable DOA Estimation using Incremental Learning
标题: 使用增量学习的可配置波达方向估计
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to DCASE WS 2024
链接:点击下载PDF文件
摘要:本文介绍了一种用于波达方向(DOA)估计的渐进式神经网络(PNN)模型,即DOA-PNN,以解决适应动态声学环境中由于灾难性遗忘所带来的挑战。虽然传统的方法,如GCC,MUSIC和SRP-PHAT在静态环境中是有效的,但它们在嘈杂的混响条件下表现较差。深度学习模型,特别是CNN,提供了改进,但在训练和推理阶段之间存在不匹配的配置。所提出的DOA PNN通过将任务增量学习与持续学习相结合来克服这些限制,从而允许在不同的声学场景中进行适应,并且较少忘记先前学习的知识。DOA-PNN具有特定于任务的子网络和缩放机制,可有效管理参数增长,确保增量麦克风配置的高性能。我们研究了DOA PNN在各种麦克风距离的麦克风设置下的模拟数据。研究表明,该方法能够以最小的参数增加保持性能,为DOA估计提供了一种有效的解决方案。摘要:This study introduces a progressive neural network (PNN) model for direction of arrival (DOA) estimation, DOA-PNN, addressing the challenge due to catastrophic forgetting in adapting dynamic acoustic environments. While traditional methods such as GCC, MUSIC, and SRP-PHAT are effective in static settings, they perform worse in noisy, reverberant conditions. Deep learning models, particularly CNNs, offer improvements but struggle with a mismatch configuration between the training and inference phases. The proposed DOA-PNN overcomes these limitations by incorporating task incremental learning of continual learning, allowing for adaptation across varying acoustic scenarios with less forgetting of previously learned knowledge. Featuring task-specific sub-networks and a scaling mechanism, DOA-PNN efficiently manages parameter growth, ensuring high performance across incremental microphone configurations. We study DOA-PNN on a simulated data under various mic distance based microphone settings. The studies reveal its capability to maintain performance with minimal parameter increase, presenting an efficient solution for DOA estimation.

【30】 UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event Detection
标题: UCIL:一种用于声音事件检测的无监督类增量学习方法
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to DCASE WS 2024
链接:点击下载PDF文件
摘要:这项工作探索了用于声音事件检测(SED)的类增量学习(CIL),提高了对现实世界场景的适应性。CIL在计算机视觉等领域的成功启发了我们的SED定制方法,解决了多样化和复杂音频环境的独特挑战。我们的方法采用具有蒸馏损失函数的独立无监督学习框架来集成新的声音类,同时在增量任务中保持SED模型的一致性。我们通过针对未标记数据的样本选择策略和平衡的样本更新机制进一步增强了该框架,确保多样化且具有说明性的合理表示。通过对DCASE 2023 Task 4数据集的各种持续学习方法进行评估,我们发现我们的研究可以深入了解每种方法对实际SED系统的适用性,这些系统可以添加新的声音类。研究结果还描绘了未来的方向CIL在动态音频设置。摘要:This work explores class-incremental learning (CIL) for sound event detection (SED), advancing adaptability towards real-world scenarios. CIL's success in domains like computer vision inspired our SED-tailored method, addressing the unique challenges of diverse and complex audio environments. Our approach employs an independent unsupervised learning framework with a distillation loss function to integrate new sound classes while preserving the SED model consistency across incremental tasks. We further enhance this framework with a sample selection strategy for unlabeled data and a balanced exemplar update mechanism, ensuring varied and illustrative sound representations. Evaluating various continual learning methods on the DCASE 2023 Task 4 dataset, we find that our research offers insights into each method's applicability for real-world SED systems that can have newly added sound classes. The findings also delineate future directions of CIL in dynamic audio settings.

【31】 WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System
标题: WildDEMED:一个用于野外家庭环境声音事件检测系统的LLM支持数据集
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to DCASE WS 2024
链接:点击下载PDF文件
摘要:这项工作的目的是通过提出一个新的大语言模型(LLM)驱动的数据集,即野生国内环境声音事件检测(WildDESED),以推进声音事件检测(SED)的研究。它是对原始DESED数据集的扩展,以反映家庭环境中的各种声学变化和复杂噪声。我们利用LLM根据DESED数据集的目标声音类别生成了八种不同的国内场景。然后,我们用从AudioSet中选择的精心定制的噪声混合物丰富了场景,并确保与目标声音没有重叠。我们考虑广泛流行的卷积神经递归网络来研究WildDESED数据集,这描述了其具有挑战性的性质。然后,我们通过逐渐增加噪声复杂度来应用课程学习,以增强模型在各种噪声水平下的泛化能力。我们使用这种方法的结果显示了噪声环境中的改进,验证了WildDESED数据集的有效性,促进了噪声鲁棒性SED的进步。摘要:This work aims to advance sound event detection (SED) research by presenting a new large language model (LLM)-powered dataset namely wild domestic environment sound event detection (WildDESED). It is crafted as an extension to the original DESED dataset to reflect diverse acoustic variability and complex noises in home settings. We leveraged LLMs to generate eight different domestic scenarios based on target sound categories of the DESED dataset. Then we enriched the scenarios with a carefully tailored mixture of noises selected from AudioSet and ensured no overlap with target sound. We consider widely popular convolutional neural recurrent network to study WildDESED dataset, which depicts its challenging nature. We then apply curriculum learning by gradually increasing noise complexity to enhance the model's generalization capabilities across various noise levels. Our results with this approach show improvements within the noisy environment, validating the effectiveness on the WildDESED dataset promoting noise-robust SED advancements.

【32】 High Fidelity Text-Guided Music Generation and Editing via Single-Stage Flow Matching
标题: 通过单级流匹配实现高保真文本引导音乐生成和编辑
作者:Gael Le Lan,Bowen Shi,Zhaoheng Ni,Sidd Srinivasan,Anurag Kumar,Brian Ellis,David Kant,Varun Nagaraja,Ernie Chang,Wei-Ning Hsu,Yangyang Shi,Vikas Chandra
链接:点击下载PDF文件
摘要:本文提出了一种简单高效的文本可控高保真音乐生成与编辑模型。它对来自低帧速率48 kHz立体声变分自动编码器编解码器的连续潜在表示序列进行操作,该编解码器消除了离散表示的信息丢失缺点。基于在流匹配目标上训练的扩散Transformer架构,该模型可以生成和编辑具有简单文本描述的可变持续时间的各种高质量立体声样本。我们还探索了一种新的正则化潜在反演方法,用于zero-shot测试时间文本引导编辑,并证明了其优于朴素去噪扩散隐式模型(DDIM)反演的各种音乐编辑提示的性能。对客观和主观指标进行了评估,并表明所提出的模型不仅在标准的文本到音乐基准(质量和效率方面)上与评估基线具有竞争力,而且在与我们提出的潜在反转相结合时,还优于以前的音乐编辑技术。样品可在https: melodyflow.github.io上获得。摘要:We introduce a simple and efficient text-controllable high-fidelity music generation and editing model. It operates on sequences of continuous latent representations from a low frame rate 48 kHz stereo variational auto encoder codec that eliminates the information loss drawback of discrete representations. Based on a diffusion transformer architecture trained on a flow-matching objective the model can generate and edit diverse high quality stereo samples of variable duration, with simple text descriptions. We also explore a new regularized latent inversion method for zero-shot test-time text-guided editing and demonstrate its superior performance over naive denoising diffusion implicit model (DDIM) inversion for variety of music editing prompts. Evaluations are conducted on both objective and subjective metrics and demonstrate that the proposed model is not only competitive to the evaluated baselines on a standard text-to-music benchmark - quality and efficiency-wise - but also outperforms previous state of the art for music editing when combined with our proposed latent inversion. Samples are available at https: melodyflow.github.io.


eess.AS音频处理
【1】 Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
标题: Seed-ASB:利用基于LLM的语音识别来理解不同的语音和上下文
作者:Ye Bai,Jingping Chen,Jitong Chen,Wei Chen,Zhuo Chen,Chen Ding,Linhao Dong,Qianqian Dong,Yujiao Du,Kepan Gao,Lu Gao,Yi Guo,Minglun Han,Ting Han,Wenchao Hu,Xinying Hu,Yuxiang Hu,Deyu Hua,Lu Huang,Mingkun Huang,Youjia Huang,Jishuo Jin,Fanliu Kong,Zongwei Lan,Tianyu Li,Xiaoyang Li,Zeyang Li,Zehua Lin,Rui Liu,Shouda Liu,Lu Lu,Yizhou Lu,Jingting Ma,Shengtao Ma,Yulin Pei,Chen Shen,Tian Tan,Xiaogang Tian,Ming Tu,Bo Wang,Hao Wang,Yuping Wang,Yuxuan Wang,Hanzhang Xia,Rui Xia,Shuangyi Xie,Hongmin Xu,Meng Yang,Bihong Zhang,Jun Zhang,Wanyi Zhang,Yang Zhang,Yawei Zhang,Yijie Zheng,Ming Zou
链接:点击下载PDF文件
摘要:现代自动语音识别(ASR)模型需要在各种应用场景中准确地转录给定特定上下文信息的各种语音信号(来自不同领域、语言、口音等)。融合了额外语言模型的经典端到端模型表现良好,但主要是在数据匹配场景中,并逐渐接近瓶颈。在这项工作中,我们介绍了种子ASR,一个大的语言模型(LLM)为基础的语音识别模型。Seed-ASR是基于音频条件LLM(AcLLM)的框架开发的,通过将连续语音表示与上下文信息一起输入到LLM中来利用LLM的能力。通过分阶段的大规模训练和LLM中上下文感知能力的启发,Seed-ASR在综合评估集(包括多个领域,口音 方言和语言)上展示了端到端模型的显着改进。此外,Seed-ASR可以进一步部署,以支持各种场景中的特定需求,而无需额外的语言模型。与最近发布的大型ASR模型相比,Seed-ASR在中文和英文公共测试集上的单词(或字符,对于中文)错误率降低了10%-40%,进一步证明了其强大的性能。摘要:Modern automatic speech recognition (ASR) model is required to accurately transcribe diverse speech signals (from different domains, languages, accents, etc) given the specific contextual information in various application scenarios. Classic end-to-end models fused with extra language models perform well, but mainly in data matching scenarios and are gradually approaching a bottleneck. In this work, we introduce Seed-ASR, a large language model (LLM) based speech recognition model. Seed-ASR is developed based on the framework of audio conditioned LLM (AcLLM), leveraging the capabilities of LLMs by inputting continuous speech representations together with contextual information into the LLM. Through stage-wise large-scale training and the elicitation of context-aware capabilities in LLM, Seed-ASR demonstrates significant improvement over end-to-end models on comprehensive evaluation sets, including multiple domains, accents dialects and languages. Additionally, Seed-ASR can be further deployed to support specific needs in various scenarios without requiring extra language models. Compared to recently released large ASR models, Seed-ASR achieves 10%-40% reduction in word (or character, for Chinese) error rates on Chinese and English public test sets, further demonstrating its powerful performance.

【2】 Multitaper mel-spectrograms for keyword spotting
标题: 用于关键词识别的多锥梅尔光谱图
作者:Douglas Baptista de Souza,Khaled Jamal Bakri,Fernanda Ferreira,Juliana Inacio
链接:点击下载PDF文件
摘要:关键词识别(KWS)是语音识别中对特征表示质量最敏感的任务之一。然而,传统上对KWS的研究主要集中在新的模型拓扑结构上,很少关注其他方面,如特征提取。本文研究了使用多锥度技术来创建KWS的改进功能。实验研究进行了不同的测试场景,窗口和参数,数据集,和神经网络中常用的嵌入式KWS应用程序。实验结果证实了使用所提出的改进功能的优势。摘要:Keyword spotting (KWS) is one of the speech recognition tasks most sensitive to the quality of the feature representation. However, the research on KWS has traditionally focused on new model topologies, putting little emphasis on other aspects like feature extraction. This paper investigates the use of the multitaper technique to create improved features for KWS. The experimental study is carried out for different test scenarios, windows and parameters, datasets, and neural networks commonly used in embedded KWS applications. Experiment results confirm the advantages of using the proposed improved features.

【3】 Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units
标题: 使用自动发现的声学单元预训练端到端关键字搜索
作者:Bolaji Yusuf,Jan "Honza" Černocký,Murat Saraçlar
备注:Interspeech 2024. KWS code at: this https URL; AUD code at this https URL
链接:点击下载PDF文件
摘要:端到端(E2E)关键字搜索(KWS)已经作为依赖于自动语音识别(ASR)系统的输出的传统关键字搜索的替代和补充方法而出现。虽然E2E方法极大地简化了KWS管道,但它们的性能通常比基于ASR的方法差,后者可以从使用未转录数据的预训练中受益。在这项工作中,我们提出了一种方法,用于预训练E2E KWS系统与未转录的数据,其中涉及使用声学单元发现(AUD),以获得未转录的数据的离散单元,然后学习定位这些单元的序列的语音。我们跨语言和AUD系统进行实验:我们表明,微调这样的模型显着优于从头开始训练的模型,并且性能改进通常与用于预训练的AUD系统的质量相关。摘要:End-to-end (E2E) keyword search (KWS) has emerged as an alternative and complimentary approach to conventional keyword search which depends on the output of automatic speech recognition (ASR) systems. While E2E methods greatly simplify the KWS pipeline, they generally have worse performance than their ASR-based counterparts, which can benefit from pretraining with untranscribed data. In this work, we propose a method for pretraining E2E KWS systems with untranscribed data, which involves using acoustic unit discovery (AUD) to obtain discrete units for untranscribed data and then learning to locate sequences of such units in the speech. We conduct experiments across languages and AUD systems: we show that finetuning such a model significantly outperforms a model trained from scratch, and the performance improvements are generally correlated with the quality of the AUD system used for pretraining.

【4】 Speculative Speech Recognition by Audio-Prefixed Low-Rank Adaptation of Language Models
标题: 通过语言模型的音频前置低等级自适应进行推测性语音识别
作者:Bolaji Yusuf,Murali Karthick Baskar,Andrew Rosenberg,Bhuvana Ramabhadran
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:本文探讨了投机性语音识别(SSR),我们赋予传统的自动语音识别(ASR)的投机能力,允许识别器运行在音频之前。我们引入了一个衡量SSR性能的指标,并提出了一个模型,该模型通过将基于RNN传感器的ASR系统与音频前缀语言模型(LM)相结合来进行SSR。ASR系统转录正在进行的音频,并将得到的转录本连同依赖于音频的前缀一起馈送到LM,LM推测转录的可能完成。我们用各种ASR数据集进行了实验,结果表明我们的方法有效,SSR作为减少ASR延迟的方法是可行的。摘要:This paper explores speculative speech recognition (SSR), where we empower conventional automatic speech recognition (ASR) with speculation capabilities, allowing the recognizer to run ahead of audio. We introduce a metric for measuring SSR performance and we propose a model which does SSR by combining a RNN-Transducer-based ASR system with an audio-prefixed language model (LM). The ASR system transcribes ongoing audio and feeds the resulting transcripts, along with an audio-dependent prefix, to the LM, which speculates likely completions for the transcriptions. We experiment with a variety of ASR datasets on which show the efficacy our method and the feasibility of SSR as a method of reducing ASR latency.

【5】 Written Term Detection Improves Spoken Term Detection
标题: 书面术语检测改进口语术语检测
作者:Bolaji Yusuf,Murat Saraçlar
Journal-ref:in IEEEACM Transactions on Audio, Speech, and Language Processing, vol. 32, pp. 3213-3223, 2024
链接:点击下载PDF文件
摘要:与使用自动语音识别(ASR)系统输出的方法相比,关键字搜索(KWS)的端到端(E2E)方法在训练和索引复杂性方面要简单得多。然而,这种简化由于模块化的损失而具有缺点。特别是,在基于ASR的KWS系统可以通过语言模型从外部未配对文本中受益的情况下,E2E KWS系统的当前公式没有这样的机制。因此,在本文中,我们提出了一个多任务训练目标,它允许不成对的文本集成到E2E KWS,而不复杂的索引和搜索。除了训练E2E KWS模型从口语文档中检索文本查询外,我们还联合训练它从掩蔽的书面文档中检索文本查询。我们的经验表明,这种方法可以有效地利用未配对的文本KWS,在各种语言的搜索性能显着改善。我们进行的分析表明,这些改进是实现,因为所提出的方法提高了文档表示的单词在不成对的文本。最后,我们表明,所提出的方法可以用于域自适应设置在域配对数据是稀缺或不存在的。摘要:End-to-end (E2E) approaches to keyword search (KWS) are considerably simpler in terms of training and indexing complexity when compared to approaches which use the output of automatic speech recognition (ASR) systems. This simplification however has drawbacks due to the loss of modularity. In particular, where ASR-based KWS systems can benefit from external unpaired text via a language model, current formulations of E2E KWS systems have no such mechanism. Therefore, in this paper, we propose a multitask training objective which allows unpaired text to be integrated into E2E KWS without complicating indexing and search. In addition to training an E2E KWS model to retrieve text queries from spoken documents, we jointly train it to retrieve text queries from masked written documents. We show empirically that this approach can effectively leverage unpaired text for KWS, with significant improvements in search performance across a wide variety of languages. We conduct analysis which indicates that these improvements are achieved because the proposed method improves document representations for words in the unpaired text. Finally, we show that the proposed method can be used for domain adaptation in settings where in-domain paired data is scarce or nonexistent.

【6】 FA-GAN: Artifacts-free and Phase-aware High-fidelity GAN-based Vocoder
标题: FA-GAN:无伪影和相感知的高保真基于GAN的声码器
作者:Rubing Shen,Yanzhen Ren,Zongkun Sun
链接:点击下载PDF文件
摘要:基于生成对抗网络(GAN)的声码器以其高质量和快速的推理速度在语音合成中获得了极大的关注。然而,合成语音中仍然存在许多明显的谱失真,导致合成语音质量下降。在这项工作中,我们采用了一种新的基于GAN的声码器,设计用于少伪影和高保真度,称为FA-GAN。为了抑制高频分量中非理想上采样层引起的混叠伪影,我们在发生器中引入了抗混叠的双反卷积模块。为了减轻模糊伪影和丰富的频谱细节的重建,我们提出了一种新的细粒度多分辨率实部和虚部损失,以协助相位信息的建模。实验结果表明,FA-GAN优于比较的方法,在提高音频质量和减轻频谱伪影,并表现出优越的性能时,应用于看不见的扬声器场景。摘要:Generative adversarial network (GAN) based vocoders have achieved significant attention in speech synthesis with high quality and fast inference speed. However, there still exist many noticeable spectral artifacts, resulting in the quality decline of synthesized speech. In this work, we adopt a novel GAN-based vocoder designed for few artifacts and high fidelity, called FA-GAN. To suppress the aliasing artifacts caused by non-ideal upsampling layers in high-frequency components, we introduce the anti-aliased twin deconvolution module in the generator. To alleviate blurring artifacts and enrich the reconstruction of spectral details, we propose a novel fine-grained multi-resolution real and imaginary loss to assist in the modeling of phase information. Experimental results reveal that FA-GAN outperforms the compared approaches in promoting audio quality and alleviating spectral artifacts, and exhibits superior performance when applied to unseen speaker scenarios.

【7】 From Audio Encoders to Piano Judges: Benchmarking Performance Understanding for Solo Piano
标题: 从音频编码员到钢琴评委:钢琴独奏的表演理解基准
作者:Huan Zhang,Jinhua Liang,Simon Dixon
备注:Accepted by the 25th International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:我们的研究探讨了通过音频编码模型的镜头理解音乐表演的方法,重点放在独奏西方古典钢琴音乐领域。与作曲级别的属性理解(如基调或流派)相比,我们确定了演奏级别音乐理解的知识差距,并解决了三个关键任务:专业知识排名,难度估计和钢琴技术检测,为此目的引入了一个全面的钢琴标记数据集(PLD)。我们利用预先训练的音频编码器,特别是Juventure,Audio-MAE,MERT和DAC,展示了处理下游任务的各种能力,以探索特定领域的微调是否增强了捕捉性能细微差别的能力。我们的最佳方法在专业知识排名中达到了93.6%的准确率,在难度估计中达到了33.7%,在技术检测中达到了46.7%,Audio-MAE是总体上最有效的编码器。最后,我们对肖邦钢琴比赛数据进行了案例研究,使用经过训练的模型进行专业排名,这突出了准确评估顶级表演的挑战。摘要:Our study investigates an approach for understanding musical performances through the lens of audio encoding models, focusing on the domain of solo Western classical piano music. Compared to composition-level attribute understanding such as key or genre, we identify a knowledge gap in performance-level music understanding, and address three critical tasks: expertise ranking, difficulty estimation, and piano technique detection, introducing a comprehensive Pianism-Labelling Dataset (PLD) for this purpose. We leverage pre-trained audio encoders, specifically Jukebox, Audio-MAE, MERT, and DAC, demonstrating varied capabilities in tackling downstream tasks, to explore whether domain-specific fine-tuning enhances capability in capturing performance nuances. Our best approach achieved 93.6 % accuracy in expertise ranking, 33.7 % in difficulty estimation, and 46.7 % in technique detection, with Audio-MAE as the overall most effective encoder. Finally, we conducted a case study on Chopin Piano Competition data using trained models for expertise ranking, which highlights the challenge of accurately assessing top-tier performances.

【8】 XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
标题: XLSR传感器:用于自我监督预训练模型的流传输ASB
作者:Shashi Kumar,Srikanth Madikeri,Juan Zuluaga-Gomez,Esaú Villatoro-Tello,Iuliia Nigmatulina,Petr Motlicek,Manjunath K E,Aravind Ganapathiraju
备注:5 pages, double column
链接:点击下载PDF文件
摘要:自监督预训练模型在自动语音识别中表现出有竞争力的性能,即使在有限的域内监督数据进行训练。然而,流行的预训练模型不适合流式ASR,因为它们是用完全注意力上下文训练的。在本文中,我们介绍了XLSR传感器,其中XLSR-53模型被用作传感器设置中的编码器。我们在AMI数据集上的实验表明,XLSR-传感器实现了4%的绝对WER改善Whisper large-v2和8%的Zipformer传感器模型从头开始训练。为了实现流功能,我们研究了XLSR-53模型内Transformer层的自注意力计算中的不同注意力掩蔽模式。在低资源场景下,我们在AMI和CommonVoice的5种语言上验证了XLSR换能器。最后,通过引入注意力汇,我们将左上下文减少了一半,同时实现了WER的相对12%的改善。摘要:Self-supervised pretrained models exhibit competitive performance in automatic speech recognition on finetuning, even with limited in-domain supervised data for training. However, popular pretrained models are not suitable for streaming ASR because they are trained with full attention context. In this paper, we introduce XLSR-Transducer, where the XLSR-53 model is used as encoder in transducer setup. Our experiments on the AMI dataset reveal that the XLSR-Transducer achieves 4% absolute WER improvement over Whisper large-v2 and 8% over a Zipformer transducer model trained from scratch.To enable streaming capabilities, we investigate different attention masking patterns in the self-attention computation of transformer layers within the XLSR-53 model. We validate XLSR-Transducer on AMI and 5 languages from CommonVoice under low-resource scenarios. Finally, with the introduction of attention sinks, we reduce the left context by half while achieving a relative 12% improvement in WER.

【9】 Sound Field Estimation Using Deep Kernel Learning Regularized by the Wave Equation
标题: 使用由波动方程正规化的深度核学习进行声学估计
作者:David Sundström,Shoichi Koyama,Andreas Jakobsson
备注:Accepted for IWAENC 2024
链接:点击下载PDF文件
摘要:在这项工作中,我们引入了一个时空内核高斯过程(GP)回归为基础的声场估计。值得注意的是,GP具有吸引人的属性,即声场是测量的线性函数,允许从分布式麦克风测量有效地估计场。然而,为了确保分析的易处理性,大多数现有的用于声场估计的内核已经在频域中制定,针对每个频率独立地形成。为了解决时空内核的分析难题,我们在这里建议通过深度内核学习直接从数据中学习内核。此外,为了提高深度核的泛化能力,我们提出了一种使用波动方程正则化学习过程的方法。通过数值模拟说明了深核的代表性优势和通过使用波动方程正则化得到的改进的推广。摘要:In this work, we introduce a spatio-temporal kernel for Gaussian process (GP) regression-based sound field estimation. Notably, GPs have the attractive property that the sound field is a linear function of the measurements, allowing the field to be estimated efficiently from distributed microphone measurements. However, to ensure analytical tractability, most existing kernels for sound field estimation have been formulated in the frequency domain, formed independently for each frequency. To address the analytical intractability of spatio-temporal kernels, we here propose to instead learn the kernel directly from data by the means of deep kernel learning. Furthermore, to improve the generalization of the deep kernel, we propose a method for regularizing the learning process using the wave equation. The representational advantages of the deep kernel and the improved generalization obtained by using the wave equation regularization are illustrated using numerical simulations.

【10】 We Need Variations in Speech Synthesis: Sub-center Modelling for Speaker Embeddings
标题: 我们需要语音合成的变体:说话者嵌入的子中心建模
作者:Ismail Rasim Ulgen,Carlos Busso,John H. L. Hansen,Berrak Sisman
备注:Submitted to IEEE Signal Processing Letters
链接:点击下载PDF文件
摘要:在语音合成中,对人类语音中丰富的情感和韵律变化进行建模是合成自然语音的关键。虽然说话人嵌入作为条件输入已被广泛应用于个性化语音合成中,但它们被设计为丢失变化以优化说话人识别准确性。因此,在对输出语音分布处的丰富变化进行建模方面,它们对于语音合成是次优的。在这项工作中,我们提出了一种新的说话人嵌入网络,它利用多个类中心的说话人分类训练,而不是一个单一的类中心作为传统的嵌入。所提出的方法在保留说话人识别性能的同时引入了说话人嵌入的变化,因为模型不必将说话人的所有话语映射到单个类中心。我们应用我们提出的嵌入在语音转换任务,并表明我们的方法提供了更好的自然度和韵律合成语音。摘要:In speech synthesis, modeling of rich emotions and prosodic variations present in human voice are crucial to synthesize natural speech. Although speaker embeddings have been widely used in personalized speech synthesis as conditioning inputs, they are designed to lose variation to optimize speaker recognition accuracy. Thus, they are suboptimal for speech synthesis in terms of modeling the rich variations at the output speech distribution. In this work, we propose a novel speaker embedding network which utilizes multiple class centers in the speaker classification training rather than a single class center as traditional embeddings. The proposed approach introduces variations in the speaker embedding while retaining the speaker recognition performance since model does not have to map all of the utterances of a speaker into a single class center. We apply our proposed embedding in voice conversion task and show that our method provides better naturalness and prosody in synthesized speech.

【11】 Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
标题: 谁觉得这个声音很有吸引力?使用野外数据的大规模实验
作者:Hitoshi Suda,Aya Watanabe,Shinnosuke Takamichi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了CocoNut-Humoresque,一个开源的大规模语音喜欢度语料库,包括语音片段和每个听众的喜欢度分数。评价声音的可爱性对于为语音系统(如对话或公告系统)设计更好的声音至关重要。在这项研究中,我们让885名听众对1800个演讲片段进行了评价,这些演讲片段来自各种各样的演讲者。在构建语料库时,我们还收集了多个说话者属性:性别,年龄和最喜欢的YouTube视频。因此,语料库使大规模的统计分析的声音喜欢关于扬声器和听众的因素。本文描述了构建方法和初步的数据分析,以揭示性别和年龄偏见的声音喜欢。此外,讨喜度和两个声学特征,基频和给定的话语的x-向量之间的关系,也进行了研究。摘要:This paper introduces CocoNut-Humoresque, an open-source large-scale speech likability corpus that includes speech segments and their per-listener likability scores. Evaluating voice likability is essential to designing preferable voices for speech systems, such as dialogue or announcement systems. In this study, we let 885 listeners rate 1800 speech segments of a wide range of speakers regarding their likability. When constructing the corpus, we also collected the multiple speaker attributes: genders, ages, and favorite YouTube videos. Therefore, the corpus enables the large-scale statistical analysis of voice likability regarding both speaker and listener factors. This paper describes the construction methodology and preliminary data analysis to reveal the gender and age biases in voice likability. In addition, the relationship between the likability and two acoustic features, the fundamental frequencies and the x-vectors of given utterances, is also investigated.

【12】 Semi-supervised Learning for Code-Switching ASR with Large Language Model Filter
标题: 具有大语言模型过滤器的代码转换ASB的半监督学习
作者:Yu Xi,Wen Ding,Kai Yu,Junjie Lai
链接:点击下载PDF文件
摘要:语码转换是指不同语言的词语或短语在同一个句子中交替出现的现象。由于数据稀缺,构建有效的CS自动语音识别(ASR)系统仍然具有挑战性。在本文中,我们建议通过在半监督学习框架内利用丰富的无监督单语语音数据来增强CS-ASR系统,特别是当对CS数据的访问受到限制时。为了实现这一目标,我们建立了一个一般的范例应用嘈杂的学生训练(NST)的CS-ASR任务。具体来说,我们引入了LLM过滤器,它利用精心设计的提示模板来激活大型语言模型(LLM)的纠正能力,用于NST期间的单语数据选择和伪标签细化。我们在有监督的ASRU-CS和无监督的AISHELL-2和LibriSpeech数据集上的实验表明,我们的方法不仅在CS任务的监督和半监督学习基线上取得了显着的改进,而且在CS英语部分上与完全监督的Oracle上界相比也取得了更好的性能。此外,我们进一步研究了口音对AESRC数据集的影响,并证明了当单语数据包含相关语言特征时,我们的方法可以获得额外的好处。摘要:Code-switching (CS) phenomenon occurs when words or phrases from different languages are alternated in a single sentence. Due to data scarcity, building an effective CS Automatic Speech Recognition (ASR) system remains challenging. In this paper, we propose to enhance CS-ASR systems by utilizing rich unsupervised monolingual speech data within a semi-supervised learning framework, particularly when access to CS data is limited. To achieve this, we establish a general paradigm for applying noisy student training (NST) to the CS-ASR task. Specifically, we introduce the LLM-Filter, which leverages well-designed prompt templates to activate the correction capability of large language models (LLMs) for monolingual data selection and pseudo-labels refinement during NST. Our experiments on the supervised ASRU-CS and unsupervised AISHELL-2 and LibriSpeech datasets show that our method not only achieves significant improvements over supervised and semi-supervised learning baselines for the CS task, but also attains better performance compared with the fully-supervised oracle upper-bound on the CS English part. Additionally, we further investigate the influence of accent on AESRC dataset and demonstrate that our method can get achieve additional benefits when the monolingual data contains relevant linguistic characteristic.

【13】 DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners
标题: DASS:蒸馏音频状态空间模型更强大、持续时间更可扩展
作者:Saurabhchand Bhati,Yuan Gong,Leonid Karlinsky,Hilde Kuehne,Rogerio Feris,James Glass
链接:点击下载PDF文件
摘要:状态空间模型(SSM)由于其在长输入下的高计算效率而成为用于音频建模的Transformers的替代方案。虽然最近对音频SSM的努力报告了令人鼓舞的结果,但仍然存在两个主要限制:首先,在10秒的短音频标记任务中,音频SSM与基于transformer的模型(如音频频谱图Transformer(AST))相比仍然表现不佳。其次,虽然音频SSM理论上支持长音频输入,但其长音频的实际性能尚未得到彻底评估。为了解决这些局限性,在本文中,1)我们将知识蒸馏应用于音频空间模型训练,产生了一个称为知识蒸馏音频SSM(DASS)的模型。据我们所知,这是第一个在AudioSet上优于Transformers的SSM,并达到了47.6的mAP; 2)我们设计了一个名为Audio Needle In A Haystack(Audio NIAH)的新测试。我们发现,仅用10秒音频片段训练的DASS可以检索长达2.5小时的音频记录中的声音事件,而AST模型在输入仅为50秒时失败,这表明SSM确实具有更大的持续时间可扩展性。摘要:State-space models (SSMs) have emerged as an alternative to Transformers for audio modeling due to their high computational efficiency with long inputs. While recent efforts on Audio SSMs have reported encouraging results, two main limitations remain: First, in 10-second short audio tagging tasks, Audio SSMs still underperform compared to Transformer-based models such as Audio Spectrogram Transformer (AST). Second, although Audio SSMs theoretically support long audio inputs, their actual performance with long audio has not been thoroughly evaluated. To address these limitations, in this paper, 1) We applied knowledge distillation in audio space model training, resulting in a model called Knowledge Distilled Audio SSM (DASS). To the best of our knowledge, it is the first SSM that outperforms the Transformers on AudioSet and achieves an mAP of 47.6; and 2) We designed a new test called Audio Needle In A Haystack (Audio NIAH). We find that DASS, trained with only 10-second audio clips, can retrieve sound events in audio recordings up to 2.5 hours long, while the AST model fails when the input is just 50 seconds, demonstrating SSMs are indeed more duration scalable.

【14】 Optimizing a-DCF for Spoofing-Robust Speaker Verification
标题: 优化a-DCF以实现欺骗稳健的说话人验证
作者:Oğuzhan Kurnaz,Jagabandhu Mishra,Tomi H. Kinnunen,Cemal Hanilçi
链接:点击下载PDF文件
摘要:自动说话人确认(ASV)系统容易受到欺骗攻击,如文本到语音。在这项研究中,我们提出了一种新的欺骗强大的ASV后端分类器,直接优化最近推出的,架构不可知的检测成本函数(a-DCF)。我们结合了a-DCF和二进制交叉熵(BCE)损失来优化网络权重,并结合了一种新的,简单的检测阈值优化技术。在ASVspoof 2019数据库上的实验表明,与仅使用BCE优化的基线相比有相当大的改进(从最小a-DCF 0.1445到0.1254),代表13%的相对改进。这些初步的有希望的结果表明,有可能调整ASV系统,以在用户便利性和安全性这两个相互矛盾的目标之间找到适当的平衡。摘要:Automatic speaker verification (ASV) systems are vulnerable to spoofing attacks such as text-to-speech. In this study, we propose a novel spoofing-robust ASV back-end classifier, optimized directly for the recently introduced, architecture-agnostic detection cost function (a-DCF). We combine a-DCF and binary cross-entropy (BCE) losses to optimize the network weights, combined by a novel, straightforward detection threshold optimization technique. Experiments on the ASVspoof2019 database demonstrate considerable improvement over the baseline optimized using BCE only (from minimum a-DCF of 0.1445 to 0.1254), representing 13% relative improvement. These initial promising results demonstrate that it is possible to adjust an ASV system to find appropriate balance across the contradicting aims of user convenience and security against adversaries.

【15】 Configurable DOA Estimation using Incremental Learning
标题: 使用增量学习的可配置波达方向估计
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to DCASE WS 2024
链接:点击下载PDF文件
摘要:本文介绍了一种用于波达方向(DOA)估计的渐进式神经网络(PNN)模型,即DOA-PNN,以解决适应动态声学环境中由于灾难性遗忘所带来的挑战。虽然传统的方法,如GCC,MUSIC和SRP-PHAT在静态环境中是有效的,但它们在嘈杂的混响条件下表现较差。深度学习模型,特别是CNN,提供了改进,但在训练和推理阶段之间存在不匹配的配置。所提出的DOA PNN通过将任务增量学习与持续学习相结合来克服这些限制,从而允许在不同的声学场景中进行适应,并且较少忘记先前学习的知识。DOA-PNN具有特定于任务的子网络和缩放机制,可有效管理参数增长,确保增量麦克风配置的高性能。我们研究了DOA PNN在各种麦克风距离的麦克风设置下的模拟数据。研究表明,该方法能够以最小的参数增加保持性能,为DOA估计提供了一种有效的解决方案。摘要:This study introduces a progressive neural network (PNN) model for direction of arrival (DOA) estimation, DOA-PNN, addressing the challenge due to catastrophic forgetting in adapting dynamic acoustic environments. While traditional methods such as GCC, MUSIC, and SRP-PHAT are effective in static settings, they perform worse in noisy, reverberant conditions. Deep learning models, particularly CNNs, offer improvements but struggle with a mismatch configuration between the training and inference phases. The proposed DOA-PNN overcomes these limitations by incorporating task incremental learning of continual learning, allowing for adaptation across varying acoustic scenarios with less forgetting of previously learned knowledge. Featuring task-specific sub-networks and a scaling mechanism, DOA-PNN efficiently manages parameter growth, ensuring high performance across incremental microphone configurations. We study DOA-PNN on a simulated data under various mic distance based microphone settings. The studies reveal its capability to maintain performance with minimal parameter increase, presenting an efficient solution for DOA estimation.

【16】 UCIL: An Unsupervised Class Incremental Learning Approach for Sound Event Detection
标题: UCIL:一种用于声音事件检测的无监督类增量学习方法
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to DCASE WS 2024
链接:点击下载PDF文件
摘要:这项工作探讨了声音事件检测(SED)的类增量学习(CIL),提高了对现实世界场景的适应性。CIL在计算机视觉等领域的成功启发了我们的SED定制方法,解决了多样化和复杂音频环境的独特挑战。我们的方法采用了一个独立的无监督学习框架与蒸馏损失函数,以整合新的声音类,同时保持SED模型的一致性,在增量任务。我们进一步增强了这个框架的样本选择策略,为未标记的数据和平衡的样本更新机制,确保不同的和说明性的声音表示。通过对DCASE 2023 Task 4数据集的各种持续学习方法进行评估,我们发现我们的研究可以深入了解每种方法对实际SED系统的适用性,这些系统可以添加新的声音类。研究结果还描绘了未来的方向CIL在动态音频设置。摘要:This work explores class-incremental learning (CIL) for sound event detection (SED), advancing adaptability towards real-world scenarios. CIL's success in domains like computer vision inspired our SED-tailored method, addressing the unique challenges of diverse and complex audio environments. Our approach employs an independent unsupervised learning framework with a distillation loss function to integrate new sound classes while preserving the SED model consistency across incremental tasks. We further enhance this framework with a sample selection strategy for unlabeled data and a balanced exemplar update mechanism, ensuring varied and illustrative sound representations. Evaluating various continual learning methods on the DCASE 2023 Task 4 dataset, we find that our research offers insights into each method's applicability for real-world SED systems that can have newly added sound classes. The findings also delineate future directions of CIL in dynamic audio settings.

【17】 WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System
标题: WildDEMED:一个用于野外家庭环境声音事件检测系统的LLM支持数据集
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to DCASE WS 2024
链接:点击下载PDF文件
摘要:这项工作的目的是通过提出一个新的大语言模型(LLM)驱动的数据集,即野生国内环境声音事件检测(WildDESED),以推进声音事件检测(SED)的研究。它是对原始DESED数据集的扩展,以反映家庭环境中的各种声学变化和复杂噪声。我们利用LLM根据DESED数据集的目标声音类别生成了八种不同的国内场景。然后,我们用从AudioSet中选择的精心定制的噪声混合物丰富了场景,并确保与目标声音没有重叠。我们考虑广泛流行的卷积神经递归网络来研究WildDESED数据集,这描述了其具有挑战性的性质。然后,我们通过逐渐增加噪声复杂度来应用课程学习,以增强模型在各种噪声水平下的泛化能力。我们使用这种方法的结果显示了噪声环境中的改进,验证了WildDESED数据集的有效性,促进了噪声鲁棒性SED的进步。摘要:This work aims to advance sound event detection (SED) research by presenting a new large language model (LLM)-powered dataset namely wild domestic environment sound event detection (WildDESED). It is crafted as an extension to the original DESED dataset to reflect diverse acoustic variability and complex noises in home settings. We leveraged LLMs to generate eight different domestic scenarios based on target sound categories of the DESED dataset. Then we enriched the scenarios with a carefully tailored mixture of noises selected from AudioSet and ensured no overlap with target sound. We consider widely popular convolutional neural recurrent network to study WildDESED dataset, which depicts its challenging nature. We then apply curriculum learning by gradually increasing noise complexity to enhance the model's generalization capabilities across various noise levels. Our results with this approach show improvements within the noisy environment, validating the effectiveness on the WildDESED dataset promoting noise-robust SED advancements.

【18】 Mixstyle based Domain Generalization for Sound Event Detection with Heterogeneous Training Data
标题: 基于Mixstyle的具有异类训练数据的声音事件检测领域概括
作者:Yang Xiao,Han Yin,Jisheng Bai,Rohan Kumar Das
备注:Sumbitted to DCASE WS 2024. 5 pages. arXiv admin note: text overlap with arXiv:2407.00291
链接:点击下载PDF文件
摘要:这项工作探讨了声音事件检测(SED)的域泛化(DG),提高了对现实世界场景的适应性。我们的方法采用了一个平均教师框架与域泛化来整合异构的训练数据,同时保持SED模型在数据集上的性能。具体来说,我们首先将mixstyle应用于频率维度,以适应来自不同域的梅尔频谱图。接下来,我们使用自适应残差归一化方法,通过在频率维度上应用实例归一化来概括多个域的特征。最后,我们使用声音事件包围盒方法进行后处理。我们的方法集成了来自音频Transformers和卷积递归神经网络的双向编码器表示的功能。我们在DCASE 2024 Challenge Task 4数据集上评估了所提出的方法,在DESED数据集上测量了复调SED评分(PSDS),在MAESTRO数据集上测量了宏观平均pAUC。结果表明,与挑战基线相比,所提出的基于DG的方法改善了PSDS和宏观平均pAUC。摘要:This work explores domain generalization (DG) for sound event detection (SED), advancing adaptability towards real-world scenarios. Our approach employs a mean-teacher framework with domain generalization to integrate heterogeneous training data, while preserving the SED model performance across the datasets. Specifically, we first apply mixstyle to the frequency dimension to adapt the mel-spectrograms from different domains. Next, we use the adaptive residual normalization method to generalize features across multiple domains by applying instance normalization in the frequency dimension. Lastly, we use the sound event bounding boxes method for post-processing. Our approach integrates features from bidirectional encoder representations from audio transformers and a convolutional recurrent neural network. We evaluate the proposed approach on DCASE 2024 Challenge Task 4 dataset, measuring polyphonic SED score (PSDS) on the DESED dataset and macro-average pAUC on the MAESTRO dataset. The results indicate that the proposed DG-based method improves both PSDS and macro-average pAUC compared to the challenge baseline.

【19】 High Fidelity Text-Guided Music Generation and Editing via Single-Stage Flow Matching
标题: 通过单级流匹配实现高保真文本引导音乐生成和编辑
作者:Gael Le Lan,Bowen Shi,Zhaoheng Ni,Sidd Srinivasan,Anurag Kumar,Brian Ellis,David Kant,Varun Nagaraja,Ernie Chang,Wei-Ning Hsu,Yangyang Shi,Vikas Chandra
链接:点击下载PDF文件
摘要:本文提出了一种简单高效的文本可控高保真音乐生成与编辑模型。它对来自低帧速率48 kHz立体声变分自动编码器编解码器的连续潜在表示序列进行操作,该编解码器消除了离散表示的信息丢失缺点。基于在流匹配目标上训练的扩散Transformer架构,该模型可以生成和编辑具有简单文本描述的可变持续时间的各种高质量立体声样本。我们还探索了一种新的正则化潜在反演方法,用于zero-shot测试时间文本引导编辑,并证明了其优于朴素去噪扩散隐式模型(DDIM)反演的各种音乐编辑提示的性能。对客观和主观指标进行了评估,并表明所提出的模型不仅在标准的文本到音乐基准(质量和效率方面)上与评估基线具有竞争力,而且在与我们提出的潜在反转相结合时,还优于以前的音乐编辑技术。样品可在https: melodyflow.github.io上获得。摘要:We introduce a simple and efficient text-controllable high-fidelity music generation and editing model. It operates on sequences of continuous latent representations from a low frame rate 48 kHz stereo variational auto encoder codec that eliminates the information loss drawback of discrete representations. Based on a diffusion transformer architecture trained on a flow-matching objective the model can generate and edit diverse high quality stereo samples of variable duration, with simple text descriptions. We also explore a new regularized latent inversion method for zero-shot test-time text-guided editing and demonstrate its superior performance over naive denoising diffusion implicit model (DDIM) inversion for variety of music editing prompts. Evaluations are conducted on both objective and subjective metrics and demonstrate that the proposed model is not only competitive to the evaluated baselines on a standard text-to-music benchmark - quality and efficiency-wise - but also outperforms previous state of the art for music editing when combined with our proposed latent inversion. Samples are available at https: melodyflow.github.io.

【20】 Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
标题: 具有跨模式注意力的视频时间动态学习以实现稳健的视听语音识别
作者:Sungnyun Kim,Kangwook Jang,Sangmin Bae,Hoirin Kim,Se-Young Yun
链接:点击下载PDF文件
摘要:视听语音识别(AVSR)旨在使用音频和视频模态转录人类语音。在具有噪声破坏的音频的实际环境中,视频信息的作用变得至关重要。然而,先前的工作主要集中在增强AVSR中的音频特征,而忽略了视频特征的重要性。在这项研究中,我们通过学习视频数据中的三个时间动态来加强视频特征:上下文顺序,播放方向和视频帧的速度。引入跨模态注意模块来丰富视频特征与音频信息,以便在训练视频时间动态时可以考虑语音可变性。基于我们的方法,我们实现了最先进的性能LRS2和LRS3 AVSR基准的噪声占主导地位的设置。我们的方法在某些场景中表现出色,尤其是在胡言乱语和语音噪音方面,这表明我们有能力区分应识别的语音信号与视频模式中的嘴唇运动。我们通过提供时间动态损失和跨模态注意力架构设计的消融实验来支持我们方法的有效性。摘要:Audio-visual speech recognition (AVSR) aims to transcribe human speech using both audio and video modalities. In practical environments with noise-corrupted audio, the role of video information becomes crucial. However, prior works have primarily focused on enhancing audio features in AVSR, overlooking the importance of video features. In this study, we strengthen the video features by learning three temporal dynamics in video data: context order, playback direction, and the speed of video frames. Cross-modal attention modules are introduced to enrich video features with audio information so that speech variability can be taken into account when training on the video temporal dynamics. Based on our approach, we achieve the state-of-the-art performance on the LRS2 and LRS3 AVSR benchmarks for the noise-dominant settings. Our approach excels in scenarios especially for babble and speech noise, indicating the ability to distinguish the speech signal that should be recognized from lip movements in the video modality. We support the validity of our methodology by offering the ablation experiments for the temporal dynamics losses and the cross-modal attention architecture design.

【21】 Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations
标题: Codec-ASB:用离散语音表示训练高性能的自动语音识别系统
作者:Kunal Dhawan,Nithin Rao Koluguri,Ante Jukić,Ryan Langman,Jagadeesh Balam,Boris Ginsburg
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:离散语音表示最近因其在训练用于各种语音相关任务的基于transformer的模型中的功效而受到关注,所述语音相关任务例如自动语音识别(ASR)、翻译、说话人验证和联合语音文本基础模型。在这项工作中,我们提出了一个全面的分析建设ASR系统的离散代码。我们研究了不同的编解码器训练方法,如量化方案和时域与频谱特征编码。我们进一步探索ASR训练技术,旨在提高性能,训练效率和噪声鲁棒性。利用我们的研究结果,我们引入了一个编解码器ASR管道,在类似的比特率优于Encodec。值得注意的是,它还超过了强大的自监督模型在143种语言ML-SUPERB基准测试中取得的最先进的结果,尽管它的尺寸更小,并且在更少的数据上进行了预训练。摘要:Discrete speech representations have garnered recent attention for their efficacy in training transformer-based models for various speech-related tasks such as automatic speech recognition (ASR), translation, speaker verification, and joint speech-text foundational models. In this work, we present a comprehensive analysis on building ASR systems with discrete codes. We investigate different methods for codec training such as quantization schemes and time-domain vs spectral feature encodings. We further explore ASR training techniques aimed at enhancing performance, training efficiency, and noise robustness. Drawing upon our findings, we introduce a codec ASR pipeline that outperforms Encodec at similar bit-rate. Remarkably, it also surpasses the state-of-the-art results achieved by strong self-supervised models on the 143 languages ML-SUPERB benchmark despite being smaller in size and pretrained on significantly less data.

【22】 Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps
标题: 通过量化感知训练和二进制激活地图实现资源高效的语音质量预测
作者:Mattias Nilsson,Riccardo Miccini,Clément Laroche,Tobias Piechowiak,Friedemann Zenke
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:随着移动和边缘设备中的语音处理系统变得越来越普遍,对非侵入式语音质量监控的需求也在增加。深度学习方法提供了对客观和主观语音质量指标的高质量估计。然而,其显著的计算要求通常在资源受限的设备上是禁止的。为了解决这个问题,我们研究了基于DNSMOS的卷积架构上的语音质量预测的二进制激活映射(BAM)。我们表明,量化感知训练的二进制激活模型与基线模型的预测性能相匹配。它还允许使用其他压缩技术。结合8位权重量化,我们的方法在推理过程中减少了25倍的内存,同时用求和代替了几乎所有的点积。我们的研究结果显示了一条通过在硬件和软件中支持混合精度二进制乘法来节省大量资源的途径。摘要:As speech processing systems in mobile and edge devices become more commonplace, the demand for unintrusive speech quality monitoring increases. Deep learning methods provide high-quality estimates of objective and subjective speech quality metrics. However, their significant computational requirements are often prohibitive on resource-constrained devices. To address this issue, we investigated binary activation maps (BAMs) for speech quality prediction on a convolutional architecture based on DNSMOS. We show that the binary activation model with quantization aware training matches the predictive performance of the baseline model. It further allows using other compression techniques. Combined with 8-bit weight quantization, our approach results in a 25-fold memory reduction during inference, while replacing almost all dot products with summations. Our findings show a path toward substantial resource savings by supporting mixed-precision binary multiplication in hard- and software.

【23】 Real-time Timbre Remapping with Differentiable DSP
标题: 利用差异化的DSP实现实时音色重映射
作者:Jordie Shier,Charalampos Saitis,Andrew Robertson,Andrew McPherson
备注:Accepted for publication at the 24th International Conference on New Interfaces for Musical Expression in Utrecht, Netherlands
链接:点击下载PDF文件
摘要:音色是不同音乐环境中的主要表达方式。然而,流行的音频驱动的合成方法主要依赖于音高和响度包络,有效地使来自输入的音色表达平坦化。我们的方法借鉴了音色类比的概念,并探讨如何从输入信号的音色表达可以映射到控制合成器。利用可微数字信号处理,我们的方法有利于直接优化的合成器参数,通过一个新的功能差异损失。这个损失函数旨在学习音乐事件之间的相对音色差异,优先考虑短语内分级音色调制的微妙之处,允许在音色空间中进行有意义的翻译。使用小军鼓表演作为一个案例研究,其中音色表达是中央,我们展示了实时音色重新映射从声学小军鼓到罗兰TR-808后建模的可区分合成器。摘要:Timbre is a primary mode of expression in diverse musical contexts. However, prevalent audio-driven synthesis methods predominantly rely on pitch and loudness envelopes, effectively flattening timbral expression from the input. Our approach draws on the concept of timbre analogies and investigates how timbral expression from an input signal can be mapped onto controls for a synthesizer. Leveraging differentiable digital signal processing, our method facilitates direct optimization of synthesizer parameters through a novel feature difference loss. This loss function, designed to learn relative timbral differences between musical events, prioritizes the subtleties of graded timbre modulations within phrases, allowing for meaningful translations in a timbre space. Using snare drum performances as a case study, where timbral expression is central, we demonstrate real-time timbre remapping from acoustic snare drums to a differentiable synthesizer modeled after the Roland TR-808.

【24】 Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
标题: 突尼斯方言低资源SL和ASB语音编码器的性能分析
作者:Salima Mdhaffar,Haroun Elleuch,Fethi Bougares,Yannick Estève
备注:Accepted in ArabicNLP 2024
链接:点击下载PDF文件
摘要:通过自监督学习(SSL)预训练的语音编码器在各种下游任务中表现出卓越的性能,包括口语理解(SLU)和自动语音识别(ASR)。例如,针对此类任务微调SSL模型已显示出巨大的潜力,从而提高了SOTA在具有挑战性的数据集上的性能。与现有的研究相比,本文的贡献是通过比较SSL方法的有效性的上下文中(一)低资源的突尼斯阿拉伯语方言口语和(二)其与低资源SLU和ASR的情况下,只有少数语义注释可用于微调相结合。我们使用许多SSL语音编码器在TARIC-SLU数据集上进行实验。我们使用的语音编码器是在单语或多语言语音数据上预先训练的。其中一些也被完善,没有在域,也没有突尼斯的数据,通过多模态监督师生范式。这项研究产生了许多重要的发现,我们在本文中讨论。摘要:Speech encoders pretrained through self-supervised learning (SSL) have demonstrated remarkable performance in various downstream tasks, including Spoken Language Understanding (SLU) and Automatic Speech Recognition (ASR). For instance, fine-tuning SSL models for such tasks has shown significant potential, leading to improvements in the SOTA performance across challenging datasets. In contrast to existing research, this paper contributes by comparing the effectiveness of SSL approaches in the context of (i) the low-resource spoken Tunisian Arabic dialect and (ii) its combination with a low-resource SLU and ASR scenario, where only a few semantic annotations are available for fine-tuning. We conduct experiments using many SSL speech encoders on the TARIC-SLU dataset. We use speech encoders that were pre-trained on either monolingual or multilingual speech data. Some of them have also been refined without in-domain nor Tunisian data through multimodal supervised teacher-student paradigm. This study yields numerous significant findings that we are discussing in this paper.

【25】 Controlling Whisper: Universal Acoustic Adversarial Attacks to Control Speech Foundation Models
标题: 控制耳语:控制语音基础模型的通用声学对抗攻击
作者:Vyas Raina,Mark Gales
链接:点击下载PDF文件
摘要:支持语音的基础模型,无论是基于灵活语音识别的系统还是音频提示的大型语言模型(LLM),都变得越来越流行。这些模型的一个有趣的方面是,它们能够使用适当的提示执行自动语音识别(ASR)以外的任务。例如,OpenAI Whisper模型可以执行语音转录和语音翻译。随着音频提示LLM的发展,有可能提供更大的控制选项。在这项工作中,我们证明了这种更大的灵活性,系统可以容易受到模型控制对抗攻击。在没有对模型提示的任何访问的情况下,可以通过适当地改变音频输入来修改系统的行为。为了说明这种风险,我们证明,它是可能的prepend一个短的通用对抗性的声学段的任何输入语音信号覆盖的ASR基础模型的提示设置。具体来说,我们成功地使用了一个通用的对抗性声学段来控制Whisper始终执行语音翻译,尽管它被设置为执行语音转录。总的来说,这项工作展示了一种新形式的对抗性攻击,对支持多任务语音的基础模型进行攻击,需要在部署这种形式的模型之前加以考虑。摘要:Speech enabled foundation models, either in the form of flexible speech recognition based systems or audio-prompted large language models (LLMs), are becoming increasingly popular. One of the interesting aspects of these models is their ability to perform tasks other than automatic speech recognition (ASR) using an appropriate prompt. For example, the OpenAI Whisper model can perform both speech transcription and speech translation. With the development of audio-prompted LLMs there is the potential for even greater control options. In this work we demonstrate that with this greater flexibility the systems can be susceptible to model-control adversarial attacks. Without any access to the model prompt it is possible to modify the behaviour of the system by appropriately changing the audio input. To illustrate this risk, we demonstrate that it is possible to prepend a short universal adversarial acoustic segment to any input speech signal to override the prompt setting of an ASR foundation model. Specifically, we successfully use a universal adversarial acoustic segment to control Whisper to always perform speech translation, despite being set to perform speech transcription. Overall, this work demonstrates a new form of adversarial attack on multi-tasking speech enabled foundation models that needs to be considered prior to the deployment of this form of model.

【26】 TokenVerse: Unifying Speech and NLP Tasks via Transducer-based ASR
标题: TokenVerse:通过基于传感器的ASB统一语音和NLP任务
作者:Shashi Kumar,Srikanth Madikeri,Juan Zuluaga-Gomez,Iuliia Nigmatulina,Esaú Villatoro-Tello,Sergio Burdisso,Petr Motlicek,Karthik Pandia,Aravind Ganapathiraju
备注:5 pages, double column
链接:点击下载PDF文件
摘要:在传统的语音会话智能中,使用级联管道,涉及语音活动检测,日记,转录等任务,以及针对语义端点和命名实体识别(NER)等任务的不同NLP模型的后续处理。我们的论文介绍了TokenVerse,这是一个基于单个传感器的模型,旨在处理多个任务。这是通过在ASR模型训练期间将特定于任务的标记集成到参考文本中来实现的,简化了推理并消除了对单独的NLP模型的需求。除了ASR,我们进行实验3个不同的任务:说话人变化检测,端点,和NER。我们在公共和私有数据集上的实验表明,该方法在相对WER上将ASR提高了7.7%,同时在单个任务性能上优于级联管道方法。此外,我们提出了任务迁移学习到现有TokenVerse中的新任务。摘要:In traditional conversational intelligence from speech, a cascaded pipeline is used, involving tasks such as voice activity detection, diarization, transcription, and subsequent processing with different NLP models for tasks like semantic endpointing and named entity recognition (NER). Our paper introduces TokenVerse, a single Transducer-based model designed to handle multiple tasks. This is achieved by integrating task-specific tokens into the reference text during ASR model training, streamlining the inference and eliminating the need for separate NLP models. In addition to ASR, we conduct experiments on 3 different tasks: speaker change detection, endpointing, and NER. Our experiments on a public and a private dataset show that the proposed method improves ASR by up to 7.7% in relative WER while outperforming the cascaded pipeline approach in individual task performance. Additionally, we present task transfer learning to a new task within an existing TokenVerse.

【27】 Improving Audio Generation with Visual Enhanced Caption
标题: 使用视觉增强字幕改进音频生成
作者:Yi Yuan,Dongya Jia,Xiaobin Zhuang,Yuanzhe Chen,Zhengxi Liu,Zhuo Chen,Yuping Wang,Yuxuan Wang,Xubo Liu,Mark D. Plumbley,Wenwu Wang
备注:5 pages with 1 appendix
链接:点击下载PDF文件
摘要:生成模型已经在音频生成任务中显示出显著的成就。然而,现有的模型难以处理复杂而详细的提示,导致潜在的性能下降。我们假设这个问题源于低质量和相对少量的训练数据。在这项工作中,我们的目标是创建一个具有丰富字幕的大规模音频数据集,以改进音频生成模型。我们开发了一个自动化的管道,通过使用大型语言模型(LLM)将预测的视觉字幕,音频字幕和标记标签转换为全面的描述,为视听数据集生成详细的字幕。我们引入Sound-VECaps,这是一个包含1.66 M高质量音频字幕对的数据集,其中包含丰富的细节,包括音频事件顺序,发生地点和环境信息。我们证明,使用Sound-VECaps进行训练可以显着增强文本到音频生成模型的能力,以便从复杂的输入提示中理解和生成音频,从而提高整体系统性能。此外,我们在几个音频语言任务中进行声音VECaps的消融研究,表明其在推进音频文本表征学习中的潜力。我们的数据集和模型可在线获取。摘要:Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the low quality and relatively small quantity of training data. In this work, we aim to create a large-scale audio dataset with rich captions for improving audio generation models. We develop an automated pipeline to generate detailed captions for audio-visual datasets by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). We introduce Sound-VECaps, a dataset comprising 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We demonstrate that training with Sound-VECaps significantly enhances the capability of text-to-audio generation models to comprehend and generate audio from complex input prompts, improving overall system performance. Furthermore, we conduct ablation studies of Sound-VECaps across several audio-language tasks, suggesting its potential in advancing audio-text representation learning. Our dataset and models are available online.

【28】 A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
标题: 使用艺术材料与潜在音频合成交互的映射策略
作者:Shuoyang Zheng,Anna Xambó Sedó,Nick Bryan-Kinns
链接:点击下载PDF文件
摘要:本文提出了一种与生成式AI模型的潜在空间交互的映射策略。我们的方法涉及使用无监督特征学习来编码人类控制空间,并将其映射到音频合成模型的潜在空间。为了演示这种映射策略如何将高维传感器数据转化为深度生成模型的控制机制,我们提出了一个概念验证系统,该系统使用视觉草图来控制音频合成模型。我们借鉴XAIxArts中新兴的话语,讨论这种方法如何在艺术和创意背景下为XAI做出贡献,我们还讨论了它目前的局限性,并提出了未来的研究方向。摘要:This paper presents a mapping strategy for interacting with the latent spaces of generative AI models. Our approach involves using unsupervised feature learning to encode a human control space and mapping it to an audio synthesis model's latent space. To demonstrate how this mapping strategy can turn high-dimensional sensor data into control mechanisms of a deep generative model, we present a proof-of-concept system that uses visual sketches to control an audio synthesis model. We draw on emerging discourses in XAIxArts to discuss how this approach can contribute to XAI in artistic and creative contexts, we also discuss its current limitations and propose future research directions.

【29】 Romanization Encoding For Multilingual ASR
标题: 多语言ASB的罗马化编码
作者:Wen Ding,Fei Jia,Hainan Xu,Yu Xi,Junjie Lai,Boris Ginsburg
链接:点击下载PDF文件
摘要:我们为脚本密集的语言引入罗马化编码,以优化多语言和代码切换自动语音识别(ASR)系统。通过在配备Roman 2Char模块的FastConformer-RNNT框架中采用罗马化编码以及平衡的级联标记器,我们显着减少了词汇量和输出维度,从而实现了更大的训练批次并减少了内存消耗。该方法将声学建模和语言建模相结合,增强了系统的灵活性和适应性。在我们的研究中,将这种方法应用于普通话-英语ASR导致了显着的63.51%的词汇减少和显着的性能提高13.72%和15.03%的SEAME代码转换基准。对普通话-韩语和普通话-日语的消融研究突出了我们的方法解决其他脚本繁重语言的复杂性的强大能力,为更通用和有效的多语言ASR系统铺平了道路。摘要:We introduce romanization encoding for script-heavy languages to optimize multilingual and code-switching Automatic Speech Recognition (ASR) systems. By adopting romanization encoding alongside a balanced concatenated tokenizer within a FastConformer-RNNT framework equipped with a Roman2Char module, we significantly reduce vocabulary and output dimensions, enabling larger training batches and reduced memory consumption. Our method decouples acoustic modeling and language modeling, enhancing the flexibility and adaptability of the system. In our study, applying this method to Mandarin-English ASR resulted in a remarkable 63.51% vocabulary reduction and notable performance gains of 13.72% and 15.03% on SEAME code-switching benchmarks. Ablation studies on Mandarin-Korean and Mandarin-Japanese highlight our method's strong capability to address the complexities of other script-heavy languages, paving the way for more versatile and effective multilingual ASR systems.

【30】 PAGURI: a user experience study of creative interaction with text-to-music models
标题: PAGuri:与文本到音乐模型的创造性交互的用户体验研究
作者:Francesca Ronchini,Luca Comanducci,Gabriele Perego,Fabio Antonacci
链接:点击下载PDF文件
摘要:近年来,文本到音乐模型是自动音乐生成的最大突破。虽然它们无疑是技术进步的展示,但目前尚不清楚它们如何切实融入音乐家和音乐从业者的艺术实践。本文旨在通过即时音频生成用户研究调查(PAGURI)来解决这个问题,PAGURI是一项用户体验研究,我们利用最近的文本到音乐的发展来研究音乐家和从业者如何与这些系统进行交互,评估他们的满意度。我们开发了一个在线工具,用户可以通过它生成音乐样本和 或应用最近提出的个性化技术,基于微调,使文本到音乐模型生成更接近他们的需求和偏好的声音。使用问卷调查,我们分析了参与者如何与所提出的工具进行交互,以了解文本到音乐模型在增强用户创造力方面的有效性。结果表明,即使生成的音频样本及其质量可能并不总是满足用户的期望,大多数参与者将在他们的创作过程中使用该工具。此外,他们还深入了解了该系统的潜在增强功能及其与音乐实践的整合。摘要:In recent years, text-to-music models have been the biggest breakthrough in automatic music generation. While they are unquestionably a showcase of technological progress, it is not clear yet how they can be realistically integrated into the artistic practice of musicians and music practitioners. This paper aims to address this question via Prompt Audio Generation User Research Investigation (PAGURI), a user experience study where we leverage recent text-to-music developments to study how musicians and practitioners interact with these systems, evaluating their satisfaction levels. We developed an online tool through which users can generate music samples and or apply recently proposed personalization techniques, based on fine-tuning, to make the text-to-music model generate sounds closer to their needs and preferences. Using questionnaires, we analyzed how participants interacted with the proposed tool, to understand the effectiveness of text-to-music models in enhancing users' creativity. Results show that even if the audio samples generated and their quality may not always meet user expectations, the majority of the participants would incorporate the tool in their creative process. Furthermore, they provided insights into potential enhancements for the system and its integration into their music practice.

【31】 MuseBarControl: Enhancing Fine-Grained Control in Symbolic Music Generation through Pre-Training and Counterfactual Loss
标题: MuseBarControl:通过预训练和反事实损失增强象征性音乐生成中的细粒度控制
作者:Yangyang Shu,Haiming Xu,Ziqin Zhou,Anton van den Hengel,Lingqiao Liu
备注:Demo is available at: this https URL
链接:点击下载PDF文件
摘要:自动生成符号音乐--根据人类特定需求定制的乐谱--对音乐家和爱好者来说是非常有益的。最近的研究表明,使用广泛的数据集和先进的Transformer架构,结果很有希望。然而,这些最先进的模型通常只提供对整个作品的节奏和风格等方面的基本控制,缺乏管理更精细细节的能力,例如在单个小节级别的控制。虽然微调预训练的符号音乐生成模型似乎是实现这种更精细控制的简单方法,但我们的研究表明这种方法存在挑战。该模型往往不能充分响应新的,细粒度的酒吧级控制信号。为此,我们提出了两个创新的解决方案。首先,我们引入了一个预训练任务,旨在将控制信号直接与相应的音乐令牌联系起来,这有助于实现更有效的初始化,以便随后进行微调。其次,我们实现了一种新的反事实损失,促进生成的音乐和控制提示之间更好的对齐。总之,这些技术显着提高了我们的能力,控制音乐生成在酒吧的水平,显示了13.06%的改进,比传统的方法。我们的主观评估也证实了这种增强的控制不会损害原始预训练生成模型的音乐质量。摘要:Automatically generating symbolic music-music scores tailored to specific human needs-can be highly beneficial for musicians and enthusiasts. Recent studies have shown promising results using extensive datasets and advanced transformer architectures. However, these state-of-the-art models generally offer only basic control over aspects like tempo and style for the entire composition, lacking the ability to manage finer details, such as control at the level of individual bars. While fine-tuning a pre-trained symbolic music generation model might seem like a straightforward method for achieving this finer control, our research indicates challenges in this approach. The model often fails to respond adequately to new, fine-grained bar-level control signals. To address this, we propose two innovative solutions. First, we introduce a pre-training task designed to link control signals directly with corresponding musical tokens, which helps in achieving a more effective initialization for subsequent fine-tuning. Second, we implement a novel counterfactual loss that promotes better alignment between the generated music and the control prompts. Together, these techniques significantly enhance our ability to control music generation at the bar level, showing a 13.06 % improvement over conventional methods. Our subjective evaluations also confirm that this enhanced control does not compromise the musical quality of the original pre-trained generative model.

【32】 Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
标题: 在线发言人拨号系统的延迟性系统评估
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages
链接:点击下载PDF文件
摘要:在本文中,不同的在线发言人日记系统进行了评估,在相同的硬件与相同的测试数据方面的延迟。延迟是从音频输入到对应扬声器标签的输出的时间跨度。作为评估的一部分,DIART框架内的各种模型组合,基于在线聚类算法UIS-RNN-SML的日记系统,和端到端的在线日记系统FS-EEND进行了比较。使用嵌入模型pyannote embedding和分割模型pyannote segmentation,DIART管道实现了最低的延迟。FS-EEND系统显示出类似的良好延迟。一般来说,目前还没有发表的研究,比较几个在线日记系统的延迟。这使得这项工作更加相关。摘要:In this paper, different online speaker diarization systems are evaluated on the same hardware with the same test data with regard to their latency. The latency is the time span from audio input to the output of the corresponding speaker label. As part of the evaluation, various model combinations within the DIART framework, a diarization system based on the online clustering algorithm UIS-RNN-SML, and the end-to-end online diarization system FS-EEND are compared. The lowest latency is achieved for the DIART-pipeline with the embedding model pyannote embedding and the segmentation model pyannote segmentation. The FS-EEND system shows a similarly good latency. In general there is currently no published research that compares several online diarization systems in terms of their latency. This makes this work even more relevant.

【33】 LearnerVoice: A Dataset of Non-Native English Learners' Spontaneous Speech
标题: LearnerVoice:非英语母语学习者自发言语的数据集
作者:Haechan Kim,Junho Myung,Seoyoung Kim,Sungpah Lee,Dongyeop Kang,Juho Kim
备注:Accepted for INTERSPEECH 2024
链接:点击下载PDF文件
摘要:第二语言学习者自发言语中普遍存在的不符合语法的表达和不流利现象给自动语音识别系统带来了独特的挑战。然而,很少有数据集是针对L2学习者语音的。我们公开发布了LearnerVoice,这是一个由50.04小时的音频和L2学习者自发语音组成的数据集。我们的语言学分析表明,我们的数据集中的transanterior包含L2 S(L2学习者的自发语音)特征,包括不符合语法的表达和不流利(例如,填充词、单词重复、自我修复、错误开始),显著多于母语数据集。使用LearnerVoice微调whisper-small.en实现了10.26%的WER,比vanilla whisper-small. en低44.2%。此外,我们的定性分析表明,54.2%的错误从香草模型的LearnerVoice归因于L2 S功能,其中48.1%被减少在微调模型。摘要:Prevalent ungrammatical expressions and disfluencies in spontaneous speech from second language (L2) learners pose unique challenges to Automatic Speech Recognition (ASR) systems. However, few datasets are tailored to L2 learner speech. We publicly release LearnerVoice, a dataset consisting of 50.04 hours of audio and transcriptions of L2 learners' spontaneous speech. Our linguistic analysis reveals that transcriptions in our dataset contain L2S (L2 learner's Spontaneous speech) features, consisting of ungrammatical expressions and disfluencies (e.g., filler words, word repetitions, self-repairs, false starts), significantly more than native speech datasets. Fine-tuning whisper-small.en with LearnerVoice achieves a WER of 10.26%, 44.2% lower than vanilla whisper-small.en. Furthermore, our qualitative analysis indicates that 54.2% of errors from the vanilla model on LearnerVoice are attributable to L2S features, with 48.1% of them being reduced in the fine-tuned model.

【34】 BiosERC: Integrating Biography Speakers Supported by LLMs for ERC Tasks
标题: BiosERC:集成由LLM支持的传记演讲者来执行ERC任务
作者:Jieying Xue,Minh Phuong Nguyen,Blake Matheny,Le Minh Nguyen
备注:Accepted in the 33rd International Conference on Artificial Neural Networks (ICANN 2024)
链接:点击下载PDF文件
摘要:在会话中的情感识别任务中,最近的研究利用注意机制探索来自内部和内部说话者的话语之间的关系,以建模它们之间的情感交互。然而,属性,如扬声器的个性特征仍然未被探索,并提出了挑战,在其适用于其他任务或兼容性与不同的模型架构。因此,这项工作引入了一个新的框架名为BiosERC,它调查说话人的特点在对话中。通过采用大语言模型(LLM),我们提取的“传记信息”的谈话中的扬声器作为补充知识注入到模型中,为每个话语的情感标签进行分类。我们提出的方法在三个著名的基准数据集上取得了最先进的(SOTA)结果:IEMOCAP,MELD和EmoryNLP,证明了我们模型的有效性和通用性,并展示了其适应各种会话分析任务的潜力。我们的源代码可以在https: github.com yingjie7 BiosERC上找到。摘要:In the Emotion Recognition in Conversation task, recent investigations have utilized attention mechanisms exploring relationships among utterances from intra- and inter-speakers for modeling emotional interaction between them. However, attributes such as speaker personality traits remain unexplored and present challenges in terms of their applicability to other tasks or compatibility with diverse model architectures. Therefore, this work introduces a novel framework named BiosERC, which investigates speaker characteristics in a conversation. By employing Large Language Models (LLMs), we extract the "biographical information" of the speaker within a conversation as supplementary knowledge injected into the model to classify emotional labels for each utterance. Our proposed method achieved state-of-the-art (SOTA) results on three famous benchmark datasets: IEMOCAP, MELD, and EmoryNLP, demonstrating the effectiveness and generalization of our model and showcasing its potential for adaptation to various conversation analysis tasks. Our source code is available at https: github.com yingjie7 BiosERC.

【35】 FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
标题: FunAudioLLM:人类与LLM之间自然互动的语音理解和生成基础模型
作者:Tongyi SpeechTeam
备注:work in progress
链接:点击下载PDF文件
摘要:本报告介绍了FunAudioLLM,一个旨在增强人类与大型语言模型(LLM)之间的自然语音交互的模型系列。其核心是两个创新模型:SenseVoice,处理多语言语音识别,情感识别和音频事件检测; CosyVoice,通过控制多种语言,音色,说话风格和扬声器身份来促进自然语音生成。SenseVoice-Small可为5种语言提供极低延迟的ASR,SenseVoice-Large支持50多种语言的高精度ASR,而CosyVoice在多语言语音生成、zero-shot上下文学习、跨语言语音克隆和语音跟踪功能方面表现出色。与SenseVoice和CosyVoice相关的模型已经在Modelscope和Huggingface上开源,同时在GitHub上发布了相应的训练,推理和微调代码。通过将这些模型与LLM集成,FunAudioLLM支持语音到语音翻译、情感语音聊天、交互式播客和富有表现力的有声读物叙事等应用,从而推动了语音交互技术的发展。演示可在https: fun-audio-llm.github.io上获得,代码可在https: github.com FunAudioLLM上访问。摘要:This report introduces FunAudioLLM, a model family designed to enhance natural voice interactions between humans and large language models (LLMs). At its core are two innovative models: SenseVoice, which handles multilingual speech recognition, emotion recognition, and audio event detection; and CosyVoice, which facilitates natural speech generation with control over multiple languages, timbre, speaking style, and speaker identity. SenseVoice-Small delivers exceptionally low-latency ASR for 5 languages, and SenseVoice-Large supports high-precision ASR for over 50 languages, while CosyVoice excels in multi-lingual voice generation, zero-shot in-context learning, cross-lingual voice cloning, and instruction-following capabilities. The models related to SenseVoice and CosyVoice have been open-sourced on Modelscope and Huggingface, along with the corresponding training, inference, and fine-tuning codes released on GitHub. By integrating these models with LLMs, FunAudioLLM enables applications such as speech-to-speech translation, emotional voice chat, interactive podcasts, and expressive audiobook narration, thereby pushing the boundaries of voice interaction technology. Demos are available at https: fun-audio-llm.github.io, and the code can be accessed at https: github.com FunAudioLLM.

【36】 Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
标题: 使用基于无监督文本到语音合成的数据增强来改进重读语音识别
作者:Cong-Thanh Do,Shuhei Imai,Rama Doddipatla,Thomas Hain
备注:Accepted to EUSIPCO 2024
链接:点击下载PDF文件
摘要:本文研究了使用无监督的文本到语音合成(TTS)作为一种数据增强方法,以提高口音语音识别。TTS系统是用少量的重音语音训练数据及其伪标签而不是手动训练来训练的,因此是无监督的。这种方法使得能够使用带口音的语音数据,而无需手动转录来执行带口音的语音识别的数据增强。通过使用TTS系统从文本提示生成的合成重音语音数据然后与可用的非重音语音数据组合以训练自动语音识别(ASR)系统。ASR实验在自监督学习框架中使用Wav2vec2.0模型进行,该模型在大量无监督口音语音数据上进行预训练。用于训练无监督TTS的口音语音数据是从L2-ARCTIC和不列颠群岛语料库中选择的朗读语音,而来自爱丁堡国际英语口音语料库的自发对话语音则用作评估数据。实验结果表明,与使用无监督TTS生成的合成口音语音数据进行微调的Wav2vec2.0基线相比,Wav2vec2.0模型针对下游ASR任务进行了微调,产生了高达6.1%的相对单词错误率降低。来自Librisepeech语料库的非口音语音数据。摘要:This paper investigates the use of unsupervised text-to-speech synthesis (TTS) as a data augmentation method to improve accented speech recognition. TTS systems are trained with a small amount of accented speech training data and their pseudo-labels rather than manual transcriptions, and hence unsupervised. This approach enables the use of accented speech data without manual transcriptions to perform data augmentation for accented speech recognition. Synthetic accented speech data, generated from text prompts by using the TTS systems, are then combined with available non-accented speech data to train automatic speech recognition (ASR) systems. ASR experiments are performed in a self-supervised learning framework using a Wav2vec2.0 model which was pre-trained on large amount of unsupervised accented speech data. The accented speech data for training the unsupervised TTS are read speech, selected from L2-ARCTIC and British Isles corpora, while spontaneous conversational speech from the Edinburgh international accents of English corpus are used as the evaluation data. Experimental results show that Wav2vec2.0 models which are fine-tuned to downstream ASR task with synthetic accented speech data, generated by the unsupervised TTS, yield up to 6.1% relative word error rate reductions compared to a Wav2vec2.0 baseline which is fine-tuned with the non-accented speech data from Librispeech corpus.

【37】 Serialized Output Training by Learned Dominance
标题: 通过习得优势进行系列输出训练
作者:Ying Shi,Lantian Li,Shi Yin,Dong Wang,Jiqing Han
备注:accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:串行输出训练(SOT)通过顺序解码单个说话人的语音,在多说话人语音识别中展示了最先进的性能。为了解决具有挑战性的标签排列问题,现有方法依赖于排列不变训练(PIT)或基于时间的先进先出(FIFO)规则。本研究提出了一种基于模型的序列化策略,将一个辅助模块的注意力编码器-解码器架构,自主识别的关键因素,以订购的输出序列的语音组件在多说话人的语音。在LibriSpeech和LibriMix数据库上进行的实验表明,我们的方法在2-mix和3-mix场景中的性能明显优于PIT和FIFO基线。进一步的分析表明,序列化模块识别的因素,包括响度和性别的混合中的主导语音成分,并根据优势得分的语音成分的顺序。摘要:Serialized Output Training (SOT) has showcased state-of-the-art performance in multi-talker speech recognition by sequentially decoding the speech of individual speakers. To address the challenging label-permutation issue, prior methods have relied on either the Permutation Invariant Training (PIT) or the time-based First-In-First-Out (FIFO) rule. This study presents a model-based serialization strategy that incorporates an auxiliary module into the Attention Encoder-Decoder architecture, autonomously identifying the crucial factors to order the output sequence of the speech components in multi-talker speech. Experiments conducted on the LibriSpeech and LibriMix databases reveal that our approach significantly outperforms the PIT and FIFO baselines in both 2-mix and 3-mix scenarios. Further analysis shows that the serialization module identifies dominant speech components in a mixture by factors including loudness and gender, and orders speech components based on the dominance score.

【38】 On the Effectiveness of Acoustic BPE in Decoder-Only TTS
标题: 声学BPE在纯解码器的TTC中的有效性
作者:Bohan Li,Feiyu Shen,Yiwei Guo,Shuai Wang,Xie Chen,Kai Yu
备注:5 pages, 3 tables, 1 figures. accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:将语音离散化并由解码器模型生成语音标记是文本到语音(TTS)和口语建模(SLM)的一个很有前途的方向。为了缩短语音标记的序列长度,声学字节对编码(BPE)已经出现在SLM中,其将来自自监督语义表示的语音标记作为字符来进一步压缩标记序列。但是TTS中的增益还没有得到充分的研究,并且声学BPE的正确选择仍然不清楚。在这项工作中,我们进行了全面的研究,各种设置的声学BPE,以探讨其有效性解码器只有TTS模型与语义语音令牌。LibriTTS上的实验验证了声学BPE均匀地增加了合成语音的可懂度和多样性,同时在BPE设置中显示出不同的特征。因此,声学BPE是用于仅解码器TTS的有利工具。摘要:Discretizing speech into tokens and generating them by a decoder-only model have been a promising direction for text-to-speech (TTS) and spoken language modeling (SLM). To shorten the sequence length of speech tokens, acoustic byte-pair encoding (BPE) has emerged in SLM that treats speech tokens from self-supervised semantic representations as characters to further compress the token sequence. But the gain in TTS has not been fully investigated, and the proper choice of acoustic BPE remains unclear. In this work, we conduct a comprehensive study on various settings of acoustic BPE to explore its effectiveness in decoder-only TTS models with semantic speech tokens. Experiments on LibriTTS verify that acoustic BPE uniformly increases the intelligibility and diversity of synthesized speech, while showing different features across BPE settings. Hence, acoustic BPE is a favorable tool for decoder-only TTS.

【39】 Unsupervised speech enhancement with spectral kurtosis and double deep priors
标题: 具有谱峰度和双重深先验的无监督语音增强
作者:Hien Ohnaka,Ryoichi Miyazaki
备注:11 pages, 12 figures, and 2 Tables, submitted to Acoustical Science and Technology
链接:点击下载PDF文件
摘要:本文提出了一种基于深度先验(DP)的无监督DNN语音增强方法。这里,DP表示DNN比噪声更倾向于产生干净的语音信号。基于DP的常规方法通常涉及使用随机噪声特征作为输入对有噪声的语音信号进行训练,停止训练,仅生成干净的语音信号。然而,此些常规方法在确定最佳停止时序方面遇到挑战,经历归因于环境背景噪声的性能降级,且遭受干净语音信号的失真与噪声减少性能之间的折衷。为了解决这些挑战,我们利用两个DNN:一个用于生成干净的语音信号,另一个用于生成噪声。这些网络的组合输出非常接近有噪语音信号,并利用基于谱峭度的损失项将有噪语音信号分离为干净的语音信号和噪声。该方法的关键优势在于它能够规避权衡和早期停止问题,因为信号被足够的步骤分解。通过评估实验,我们证明该方法在高斯白噪声和环境噪声的情况下优于传统方法,同时有效地减轻了早期停止问题。摘要:This paper proposes an unsupervised DNN-based speech enhancement approach founded on deep priors (DPs). Here, DP signifies that DNNs are more inclined to produce clean speech signals than noises. Conventional methods based on DP typically involve training on a noisy speech signal using a random noise feature as input, stopping training only a clean speech signal is generated. However, such conventional approaches encounter challenges in determining the optimal stop timing, experience performance degradation due to environmental background noise, and suffer a trade-off between distortion of the clean speech signal and noise reduction performance. To address these challenges, we utilize two DNNs: one to generate a clean speech signal and the other to generate noise. The combined output of these networks closely approximates the noisy speech signal, with a loss term based on spectral kurtosis utilized to separate the noisy speech signal into a clean speech signal and noise. The key advantage of this method lies in its ability to circumvent trade-offs and early stopping problems, as the signal is decomposed by enough steps. Through evaluation experiments, we demonstrate that the proposed method outperforms conventional methods in the case of white Gaussian and environmental noise while effectively mitigating early stopping problems.

【40】 Finetuning End-to-End Models for Estonian Conversational Spoken Language Translation
标题: 爱沙尼亚对话口语翻译的端到端模型微调
作者:Tiia Sildam,Andra Velve,Tanel Alumäe
备注:Accepted to LoResMT 2024 (ACL workshop)
链接:点击下载PDF文件
摘要:本文研究了双向爱沙尼亚语-英语和爱沙尼亚语-俄语会话语音到文本翻译的端到端模型的微调。由于爱沙尼亚语的语音翻译数据有限,我们通过网络抓取和使用机器翻译从语音识别数据集中合成数据来创建额外的训练数据。我们评估了三个公开的端到端模型:Whisper,OWSM 3.1和无源M4 T。我们的研究结果表明,使用合成数据进行微调可以大幅提高翻译准确性,并且无障碍M4 T匹配或超越使用最先进的语音识别和机器翻译模型的级联语音翻译系统。摘要:This paper investigates the finetuning of end-to-end models for bidirectional Estonian-English and Estonian-Russian conversational speech-to-text translation. Due to the limited availability of speech translation data for Estonian, we created additional training data by web scraping and synthesizing data from speech recognition datasets using machine translation. We evaluated three publicly available end-to-end models: Whisper, OWSM 3.1, and SeamlessM4T. Our results indicate that fine-tuning with synthetic data enhances translation accuracy by a large margin, with SeamlessM4T matching or surpassing cascaded speech translation systems that use state-of-the-art speech recognition and machine translation models.

【41】 Improving Self-supervised Pre-training using Accent-Specific Codebooks
标题: 使用特定口音的代码簿改进自我监督的预训练
作者:Darshan Prabhu,Abhishek Gupta,Omkar Nitsure,Preethi Jyothi,Sriram Ganapathy
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音口音对最先进的端到端自动语音识别(ASR)系统的性能提出了严峻的挑战。即使使用ASR模型的自监督学习和预训练,也很少实现重音不变性。在这项工作中,我们提出了一种口音感知自适应技术的自监督学习,引入了一组可训练的口音特定的码本的自监督架构。这些可学习的码本使模型能够在预训练期间捕获口音特定信息,并在ASR微调期间进一步细化。在Mozilla Common Voice数据集上,我们提出的方法在可见和不可见的英语口音上都优于所有其他口音适应方法,单词错误率(WER)相对降低了9%。摘要:Speech accents present a serious challenge to the performance of state-of-the-art end-to-end Automatic Speech Recognition (ASR) systems. Even with self-supervised learning and pre-training of ASR models, accent invariance is seldom achieved. In this work, we propose an accent-aware adaptation technique for self-supervised learning that introduces a trainable set of accent-specific codebooks to the self-supervised architecture. These learnable codebooks enable the model to capture accent specific information during pre-training, that is further refined during ASR finetuning. On the Mozilla Common Voice dataset, our proposed approach outperforms all other accent-adaptation approaches on both seen and unseen English accents, with up to 9% relative reduction in word error rate (WER).

【42】 Multi-Convformer: Extending Conformer with Multiple Convolution Kernels
标题: Multi-Conformer:用多个卷积核扩展Conformer
作者:Darshan Prabhu,Yifan Peng,Preethi Jyothi,Shinji Watanabe
备注:Accepted to INTERSPEECH 2024
链接:点击下载PDF文件
摘要:卷积由于其对局部上下文的有效建模而成为最先进的端到端自动语音识别(ASR)系统中必不可少的。值得注意的是,它在Conformers中的使用导致了与基于vanilla Transformer的ASR系统相比更优越的性能。虽然Conformer中卷积模块以外的组件已经被重新检查,但更改卷积模块本身的探索要少得多。为此,我们引入了Multi-Convformer,它在Conformer的卷积模块中使用多个卷积内核,并结合门控。这有助于改进不同粒度的局部依赖关系的建模。我们的模型在性能上可与现有的Conformer变体(如CgMLP和E-Branchformer)相媲美,同时具有更高的参数效率。我们在四个不同的数据集和三个不同的建模范式上实证比较了我们的方法与Conformer及其变体,并显示出高达8%的相对单词错误率(WER)改善。摘要:Convolutions have become essential in state-of-the-art end-to-end Automatic Speech Recognition~(ASR) systems due to their efficient modelling of local context. Notably, its use in Conformers has led to superior performance compared to vanilla Transformer-based ASR systems. While components other than the convolution module in the Conformer have been reexamined, altering the convolution module itself has been far less explored. Towards this, we introduce Multi-Convformer that uses multiple convolution kernels within the convolution module of the Conformer in conjunction with gating. This helps in improved modeling of local dependencies at varying granularities. Our model rivals existing Conformer variants such as CgMLP and E-Branchformer in performance, while being more parameter efficient. We empirically compare our approach with Conformer and its variants across four different datasets and three different modelling paradigms and show up to 8% relative word error rate~(WER) improvements.

【43】 Continual Learning Optimizations for Auto-regressive Decoder of Multilingual ASR systems
标题: 多语言ASB系统自回归解码器的持续学习优化
作者:Chin Yuen Kwok,Jia Qi Yip,Eng Siong Chng
链接:点击下载PDF文件
摘要:持续学习(CL)涉及使用新数据微调预训练模型,同时保持预训练数据的性能。这对于扩展多语言ASR(MASR)功能尤其重要。然而,现有的CL方法,主要是为计算机视觉和强化学习任务设计的,当直接应用于MASR时,往往会产生次优的结果。我们假设这是因为自回归解码器在MASR模型中的CL是困难的。为了验证这一点,我们提出了四个优化的解码器。它们包括解码器层梯度手术,冻结未使用的令牌嵌入,抑制新添加的令牌的输出,以及学习率重新缩放。我们将Whisper适配到Common Voice数据集中的10种看不见的语言的实验表明,与Experience Replay相比,这些优化将预训练语言的平均单词错误率(AWER)从14.2%降低到12.4%,而不会影响新语言的AWER。摘要:Continual Learning (CL) involves fine-tuning pre-trained models with new data while maintaining the performance on the pre-trained data. This is particularly relevant for expanding multilingual ASR (MASR) capabilities. However, existing CL methods, mainly designed for computer vision and reinforcement learning tasks, often yield sub-optimal results when directly applied to MASR. We hypothesise that this is because CL of the auto-regressive decoder in the MASR model is difficult. To verify this, we propose four optimizations on the decoder. They include decoder-layer gradient surgery, freezing unused token embeddings, suppressing output of newly added tokens, and learning rate re-scaling. Our experiments on adapting Whisper to 10 unseen languages from the Common Voice dataset demonstrate that these optimizations reduce the Average Word Error Rate (AWER) of pretrained languages from 14.2% to 12.4% compared with Experience Replay, without compromising the AWER of new languages.

【44】 Towards Attention-based Contrastive Learning for Audio Spoof Detection
标题: 基于注意力的对比学习用于音频欺骗检测
作者:Chirag Goel,Surya Koppisetti,Ben Colman,Ali Shahriyari,Gaurav Bharaj
备注:Proc. INTERSPEECH 2023
链接:点击下载PDF文件
摘要:Vision Transformers(ViT)在计算机视觉分类任务方面取得了实质性进展。最近,Gong et.等人'21,引入了用于若干音频任务的基于注意力的建模。然而,相对未开发的是使用ViT进行音频欺骗检测任务。我们弥合了这一差距,并为此任务引入了ViT。建立在微调SSAST基础上的香草基线(Gong et.’22)音频ViT模型实现了次优等错误率(EER)。为了提高性能,我们提出了一种新的基于注意力的对比学习框架(SSAST-CL),使用交叉注意力来帮助表征学习。实验表明,我们的框架成功地解开的善意和欺骗类,并帮助学习更好的分类任务。通过适当的数据增强策略,在我们的框架上训练的模型在ASVSpoof 2021挑战中取得了有竞争力的表现。我们提供了比较和消融研究来证明我们的说法。摘要:Vision transformers (ViT) have made substantial progress for classification tasks in computer vision. Recently, Gong et. al. '21, introduced attention-based modeling for several audio tasks. However, relatively unexplored is the use of a ViT for audio spoof detection task. We bridge this gap and introduce ViTs for this task. A vanilla baseline built on fine-tuning the SSAST (Gong et. al. '22) audio ViT model achieves sub-optimal equal error rates (EERs). To improve performance, we propose a novel attention-based contrastive learning framework (SSAST-CL) that uses cross-attention to aid the representation learning. Experiments show that our framework successfully disentangles the bonafide and spoof classes and helps learn better classifiers for the task. With appropriate data augmentations policy, a model trained on our framework achieves competitive performance on the ASVSpoof 2021 challenge. We provide comparisons and ablation studies to justify our claim.

【45】 Prosody-Driven Privacy-Preserving Dementia Detection
标题: 韵律驱动的隐私保护痴呆症检测
作者:Dominika Woszczyk,Ranya Aloufi,Soteris Demetriou
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:从语音记录中提取的说话人嵌入已被证明是有价值的痴呆症检测。然而,就其性质而言,这些嵌入包含可识别的信息,这引起了隐私问题。在这项工作中,我们的目标是匿名嵌入,同时保留痴呆症检测的诊断实用程序。以前的研究依赖于对抗性学习和在有限资源环境中训练目标属性和斗争的模型。我们提出了一种新的方法,该方法利用领域知识从说话者嵌入中解开与痴呆症相关的韵律特征,而不依赖于痴呆症分类器。我们的实验表明,我们的方法在保护说话者隐私(说话者识别F1-评分0.01%)方面有效,同时在ADReSS数据集中保持高痴呆检测评分F1-评分74%。我们的结果也与ADReSSo上约束更严格的分类器依赖系统相当(0.01%和0.66%),并且对合成语音自然度没有影响。摘要:Speaker embeddings extracted from voice recordings have been proven valuable for dementia detection. However, by their nature, these embeddings contain identifiable information which raises privacy concerns. In this work, we aim to anonymize embeddings while preserving the diagnostic utility for dementia detection. Previous studies rely on adversarial learning and models trained on the target attribute and struggle in limited-resource settings. We propose a novel approach that leverages domain knowledge to disentangle prosody features relevant to dementia from speaker embeddings without relying on a dementia classifier. Our experiments show the effectiveness of our approach in preserving speaker privacy (speaker recognition F1-score .01%) while maintaining high dementia detection score F1-score of 74% on the ADReSS dataset. Our results are also on par with a more constrained classifier-dependent system on ADReSSo (.01% and .66%), and have no impact on synthesized speech naturalness.

【46】 Advanced Framework for Animal Sound Classification With Features Optimization
标题: 具有特征优化的动物声音分类高级框架
作者:Qiang Yang,Xiuying Chen,Changsheng Ma,Carlos M. Duarte,Xiangliang Zhang
链接:点击下载PDF文件
摘要:动物声音的自动分类在生物声学中提出了一个持久的挑战,由于声音信号的不同统计特性,记录设备的变化,以及普遍的低信噪比(SNR)条件。像卷积神经网络(CNN)和长短期记忆(LSTM)这样的深度学习模型在人类语音识别方面表现出色,但尚未有效地针对动物声音的复杂性质进行定制,即使在同一领域内也表现出很大的多样性。我们提出了一个自动分类框架适用于一般动物的声音分类。我们的方法首先从梅尔频率倒谱系数(MFCC)优化音频特征,包括特征重排和特征约简。然后,它使用深度学习模型的优化特征,即,基于注意力的双向LSTM(Bi-LSTM),用于提取声音分类的深层语义特征。我们还提供了一个动物声音基准数据集,包括海洋动物和鸟类1。对真实世界数据集的广泛实验表明,我们的方法在精确度、召回率和准确度方面始终优于基线方法25%以上,这在动物声音分类方面取得了可喜的进步。摘要:The automatic classification of animal sounds presents an enduring challenge in bioacoustics, owing to the diverse statistical properties of sound signals, variations in recording equipment, and prevalent low Signal-to-Noise Ratio (SNR) conditions. Deep learning models like Convolutional Neural Networks (CNN) and Long Short-Term Memory (LSTM) have excelled in human speech recognition but have not been effectively tailored to the intricate nature of animal sounds, which exhibit substantial diversity even within the same domain. We propose an automated classification framework applicable to general animal sound classification. Our approach first optimizes audio features from Mel-frequency cepstral coefficients (MFCC) including feature rearrangement and feature reduction. It then uses the optimized features for the deep learning model, i.e., an attention-based Bidirectional LSTM (Bi-LSTM), to extract deep semantic features for sound classification. We also contribute an animal sound benchmark dataset encompassing oceanic animals and birds1. Extensive experimentation with real-world datasets demonstrates that our approach consistently outperforms baseline methods by over 25% in precision, recall, and accuracy, promising advancements in animal sound classification.

【47】 PianoBART: Symbolic Piano Music Generation and Understanding with Large-Scale Pre-Training
标题: PianoBART:通过大规模预训练进行象征性钢琴音乐的生成和理解
作者:Xiao Liang,Zijian Zhao,Weichao Zeng,Yutong He,Fupeng He,Yiyi Wang,Chengying Gao
链接:点击下载PDF文件
摘要:学习音乐结构和作曲模式对于音乐生成和理解都是必要的,但是目前的方法没有统一使用学习到的特征来同时生成和理解音乐。在本文中,我们提出了PianoBART,这是一个预先训练的模型,它使用BART来生成和理解象征性的钢琴音乐。针对PianoBART的不同预训练任务,设计了一种多层次的对象选择策略,可以防止信息泄漏或丢失,提高学习能力。在预训练中捕获的音乐语义针对音乐生成和理解任务进行了微调。实验表明,PianoBART有效地学习音乐模式,并在生成高质量的连贯作品和理解音乐方面取得了出色的表现。我们的代码和补充材料可在https: github.com RS2002 PianoBart上获得。摘要:Learning musical structures and composition patterns is necessary for both music generation and understanding, but current methods do not make uniform use of learned features to generate and comprehend music simultaneously. In this paper, we propose PianoBART, a pre-trained model that uses BART for both symbolic piano music generation and understanding. We devise a multi-level object selection strategy for different pre-training tasks of PianoBART, which can prevent information leakage or loss and enhance learning ability. The musical semantics captured in pre-training are fine-tuned for music generation and understanding tasks. Experiments demonstrate that PianoBART efficiently learns musical patterns and achieves outstanding performance in generating high-quality coherent pieces and comprehending music. Our code and supplementary material are available at https: github.com RS2002 PianoBart.


机器翻译,仅供参考