本文经arXiv每日学术速递授权转载
【1】 Language Model Can Listen While Speaking
标题: 语言模型可以边听边说
作者:Ziyang Ma,Yakun Song,Chenpeng Du,Jian Cong,Zhuo Chen,Yuping Wang,Yuxuan Wang,Xie Chen
备注:Demo can be found at this https URL
链接:点击下载PDF文件
【2】 Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
标题: 集群和挖掘强调语音以实现包容性和公平的语音识别
作者:Jaeyoung Kim,Han Lu,Soheil Khorram,Anshuman Tripathi,Qian Zhang,Hasim Sak
链接:点击下载PDF文件
【3】 Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
标题: Stem-JEPA:用于音乐茎兼容性估计的联合嵌入预测架构
作者:Alain Riou,Stefan Lattner,Gaëtan Hadjeres,Michael Anslow,Geoffroy Peeters
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
【4】 Steer-by-prior Editing of Symbolic Music Loops
标题: 符号音乐循环的优先引导编辑
作者:Nicolas Jonason,Luca Casini,Bob L. T. Sturm
备注:Accepted to MML 2024
链接:点击下载PDF文件
【5】 An approach to optimize inference of the DIART speaker diarization pipeline
标题: 一种优化DIART扬声器拨号流水线推理的方法
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages, 3 figures
链接:点击下载PDF文件
【6】 Diseño de sonido para producciones audiovisuales e historias sonoras en el aula. Hacia una docencia creativa mediante el uso de herramientas inteligentes
标题: 《索尼多》是为了制作视听和历史而创作的。Hacia una docencia mediana el uso de herramientas intelligentes
作者:Miguel Civit,Francisco Cuadrado
备注:11 pages, in Spanish language. 1 figure. In La nueva era del p'odcast
链接:点击下载PDF文件
【7】 Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
标题: 基于对比学习的多语言声控关联链簇
作者:Wuyang Chen,Yanjie Sun,Kele Xu,Yong Dou
链接:点击下载PDF文件
【8】 Joint Learning of Emotions in Music and Generalized Sounds
标题: 音乐和广义声音中情感的联合学习
作者:Simonetta Federico,Certo Francesca,Ntalampiras Stavros
备注:Accepted at Audio Mostly 2024, Milan
链接:点击下载PDF文件
【9】 Why Perturbing Symbolic Music is Necessary: Fitting the Distribution of Never-used Notes through a Joint Probabilistic Diffusion Model
标题: 为什么有必要扰乱象征性音乐:通过联合概率扩散模型匹配从未使用过的音符的分布
作者:Shipei Liu,Xiaoya Fan,Guowei Wu
链接:点击下载PDF文件
【10】 ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic Features
标题: ALIF:使用语言特征对黑匣子语音平台进行低成本对抗性音频攻击
作者:Peng Cheng,Yuwei Wang,Peng Huang,Zhongjie Ba,Xiaodong Lin,Feng Lin,Li Lu,Kui Ren
备注:Published in the 2024 IEEE Symposium on Security and Privacy (SP)
链接:点击下载PDF文件
【11】 Generating High-quality Symbolic Music Using Fine-grained Discriminators
标题: 使用细粒度鉴别器生成高质量的象征音乐
作者:Zhedong Zhang,Liang Li,Jiehua Zhang,Zhenghui Hu,Hongkui Wang,Chenggang Yan,Jian Yang,Yuankai Qi
备注:Accepted by ICPR2024
链接:点击下载PDF文件
【12】 PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
标题: PiCoGen 2:采用迁移学习方法和弱对齐数据的钢琴封面生成
作者:Chih-Pin Tan,Hsin Ai,Yi-Hsin Chang,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
【13】 Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
标题: 视听Deepfake检测和定位的上下文跨模式注意力
作者:Vinaya Sree Katamneni,Ajita Rattani
链接:点击下载PDF文件
【14】 StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
标题: StreamVoice+:演变为端到端流媒体Zero-Shot语音转换
作者:Zhichao Wang,Yuanzhe Chen,Xinsheng Wang,Lei Xie,Yuping Wang
链接:点击下载PDF文件
标题: 超越正音:阿拉伯语短元音和方言音的自动恢复
作者:Yassine El Kheir,Hamdy Mubarak,Ahmed Ali,Shammur Absar Chowdhury
备注:Accepted ACL 2024 Main Conference
链接:点击下载PDF文件
【2】 StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
标题: StreamVoice+:演变为端到端流媒体Zero-Shot语音转换
作者:Zhichao Wang,Yuanzhe Chen,Xinsheng Wang,Lei Xie,Yuping Wang
链接:点击下载PDF文件
【3】 Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
标题: Re-ENACT:使用演员评论家策略进行情感语音生成的强化学习
作者:Ravi Shankar,Archana Venkataraman
备注:7 pages, 10 figures
链接:点击下载PDF文件
【4】 Language Model Can Listen While Speaking
标题: 语言模型可以边听边说
作者:Ziyang Ma,Yakun Song,Chenpeng Du,Jian Cong,Zhuo Chen,Yuping Wang,Yuxuan Wang,Xie Chen
备注:Demo can be found at this https URL
链接:点击下载PDF文件
【5】 Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
标题: 集群和挖掘强调语音以实现包容性和公平的语音识别
作者:Jaeyoung Kim,Han Lu,Soheil Khorram,Anshuman Tripathi,Qian Zhang,Hasim Sak
链接:点击下载PDF文件
【6】 Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
标题: Stem-JEPA:用于音乐茎兼容性估计的联合嵌入预测架构
作者:Alain Riou,Stefan Lattner,Gaëtan Hadjeres,Michael Anslow,Geoffroy Peeters
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
【7】 Steer-by-prior Editing of Symbolic Music Loops
标题: 符号音乐循环的优先引导编辑
作者:Nicolas Jonason,Luca Casini,Bob L. T. Sturm
备注:Accepted to MML 2024
链接:点击下载PDF文件
【8】 An approach to optimize inference of the DIART speaker diarization pipeline
标题: 一种优化DIART扬声器拨号流水线推理的方法
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages, 3 figures
链接:点击下载PDF文件
【9】 Diseño de sonido para producciones audiovisuales e historias sonoras en el aula. Hacia una docencia creativa mediante el uso de herramientas inteligentes
标题: 《索尼多》是为了制作视听和历史而创作的。Hacia una docencia mediana el uso de herramientas intelligentes
作者:Miguel Civit,Francisco Cuadrado
备注:11 pages, in Spanish language. 1 figure. In La nueva era del p'odcast
链接:点击下载PDF文件
【10】 Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
标题: 基于对比学习的多语言声控关联链簇
作者:Wuyang Chen,Yanjie Sun,Kele Xu,Yong Dou
链接:点击下载PDF文件
【11】 Joint Learning of Emotions in Music and Generalized Sounds
标题: 音乐和广义声音中情感的联合学习
作者:Simonetta Federico,Certo Francesca,Ntalampiras Stavros
备注:Accepted at Audio Mostly 2024, Milan
链接:点击下载PDF文件
【12】 Why Perturbing Symbolic Music is Necessary: Fitting the Distribution of Never-used Notes through a Joint Probabilistic Diffusion Model
标题: 为什么有必要扰乱象征性音乐:通过联合概率扩散模型匹配从未使用过的音符的分布
作者:Shipei Liu,Xiaoya Fan,Guowei Wu
链接:点击下载PDF文件
【13】 ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic Features
标题: ALIF:使用语言特征对黑匣子语音平台进行低成本对抗性音频攻击
作者:Peng Cheng,Yuwei Wang,Peng Huang,Zhongjie Ba,Xiaodong Lin,Feng Lin,Li Lu,Kui Ren
备注:Published in the 2024 IEEE Symposium on Security and Privacy (SP)
链接:点击下载PDF文件
【14】 Generating High-quality Symbolic Music Using Fine-grained Discriminators
标题: 使用细粒度鉴别器生成高质量的象征音乐
作者:Zhedong Zhang,Liang Li,Jiehua Zhang,Zhenghui Hu,Hongkui Wang,Chenggang Yan,Jian Yang,Yuankai Qi
备注:Accepted by ICPR2024
链接:点击下载PDF文件
【15】 PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
标题: PiCoGen 2:采用迁移学习方法和弱对齐数据的钢琴封面生成
作者:Chih-Pin Tan,Hsin Ai,Yi-Hsin Chang,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
【16】 Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
标题: 视听Deepfake检测和定位的上下文跨模式注意力
作者:Vinaya Sree Katamneni,Ajita Rattani
链接:点击下载PDF文件
标题: 语言模型可以边听边说
作者:Ziyang Ma,Yakun Song,Chenpeng Du,Jian Cong,Zhuo Chen,Yuping Wang,Yuxuan Wang,Xie Chen
备注:Demo can be found at this https URL
链接:点击下载PDF文件
摘要:对话是最自然的人机交互方式。语音语言模型(SLM)的最新进展显着增强了基于语音的会话AI。然而,这些模型仅限于基于回合的对话,缺乏在实时口语场景中与人类交互的能力,例如,当生成的内容不令人满意时被中断。为了解决这些限制,我们探讨全双工建模(FDM)在交互式语音语言模型(iSLM),重点是增强实时交互,更明确地说,探索中断的本质能力。我们介绍了一种新的模型设计,即边听边说的语言模型(LSLM),一个端到端的系统配备了听和说通道。我们的LSLM采用基于令牌的解码器的语音生成和流自监督学习(SSL)编码器的实时音频输入的TTS。LSLM融合两个通道进行自回归生成,并实时检测话轮转换。三种融合策略-早期融合,中期融合,后期融合-探索,中期融合实现语音生成和实时交互之间的最佳平衡。两个实验设置,基于命令的FDM和基于语音的FDM,证明LSLM的鲁棒性噪声和灵敏度的不同的指令。我们的研究结果突出了LSLM的能力,实现双工通信,对现有系统的影响最小。本研究的目的是促进交互式语音对话系统的发展,提高其在现实世界中的适用性。摘要:Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based conversation, lacking the ability to interact with humans in real-time spoken scenarios, for example, being interrupted when the generated content is not satisfactory. To address these limitations, we explore full duplex modeling (FDM) in interactive speech language models (iSLM), focusing on enhancing real-time interaction and, more explicitly, exploring the quintessential ability of interruption. We introduce a novel model design, namely listening-while-speaking language model (LSLM), an end-to-end system equipped with both listening and speaking channels. Our LSLM employs a token-based decoder-only TTS for speech generation and a streaming self-supervised learning (SSL) encoder for real-time audio input. LSLM fuses both channels for autoregressive generation and detects turn-taking in real time. Three fusion strategies -- early fusion, middle fusion, and late fusion -- are explored, with middle fusion achieving an optimal balance between speech generation and real-time interaction. Two experimental settings, command-based FDM and voice-based FDM, demonstrate LSLM's robustness to noise and sensitivity to diverse instructions. Our results highlight LSLM's capability to achieve duplex communication with minimal impact on existing systems. This study aims to advance the development of interactive speech dialogue systems, enhancing their applicability in real-world contexts.
【2】 Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
标题: 集群和挖掘强调语音以实现包容性和公平的语音识别
作者:Jaeyoung Kim,Han Lu,Soheil Khorram,Anshuman Tripathi,Qian Zhang,Hasim Sak
链接:点击下载PDF文件
摘要:现代自动语音识别(ASR)系统通常在超过数万小时的语音数据上进行训练,这是其取得巨大成功的主要因素之一。然而,这样的数据的分布通常偏向于常见的口音或典型的语音模式。因此,这些系统通常对非典型口音语音表现不佳。在本文中,我们提出了公平的语音识别系统,可以表现得同样好,代表口音语音口音聚类和挖掘计划。对于口音识别,我们应用了三种方案来克服监督口音数据的有限大小:监督或无监督预训练,分布鲁棒优化(DRO)和无监督聚类。三种方案都能显著改善口音识别模型,尤其是对不平衡和小口音语音的识别效果。使用建议的监督或无监督聚类方案对挖掘的印度口音语音进行微调ASR,与对随机采样语音进行微调相比,分别显示出10.0%和5.3%的相对改善。摘要:Modern automatic speech recognition (ASR) systems are typically trained on more than tens of thousands hours of speech data, which is one of the main factors for their great success. However, the distribution of such data is typically biased towards common accents or typical speech patterns. As a result, those systems often poorly perform on atypical accented speech. In this paper, we present accent clustering and mining schemes for fair speech recognition systems which can perform equally well on under-represented accented speech. For accent recognition, we applied three schemes to overcome limited size of supervised accent data: supervised or unsupervised pre-training, distributionally robust optimization (DRO) and unsupervised clustering. Three schemes can significantly improve the accent recognition model especially for unbalanced and small accented speech. Fine-tuning ASR on the mined Indian accent speech using the proposed supervised or unsupervised clustering schemes showed 10.0% and 5.3% relative improvements compared to fine-tuning on the randomly sampled speech, respectively.
【3】 Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
标题: Stem-JEPA:用于音乐茎兼容性估计的联合嵌入预测架构
作者:Alain Riou,Stefan Lattner,Gaëtan Hadjeres,Michael Anslow,Geoffroy Peeters
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
摘要:本文探讨了确定干兼容性的自动化过程中,通过识别音频记录的单一乐器,以及与给定的音乐背景融合。为了应对这一挑战,我们提出了Stem-JEPA,这是一种新型的联合嵌入预测架构(JEPA),使用自监督学习方法在多轨道数据集上进行训练。 我们的模型包括两个网络:一个编码器和一个预测器,它们被联合训练以预测来自给定上下文的嵌入的兼容词干的嵌入,通常是几种乐器的混合。以这种方式训练模型允许其用于估计干兼容性-检索,对齐或生成干以匹配给定的混合-或用于下游任务,例如流派或基调估计,因为训练范例需要模型学习与音色,和声和节奏相关的信息。 我们评估了我们的模型在MUSDB 18数据集上的检索任务上的性能,测试了它从混合数据中找到缺失干的能力,并通过主观用户研究。我们还展示了学习的嵌入捕获时间对齐信息,最后,评估了我们的模型在几个下游任务中学习的表示,强调它们有效地捕获了有意义的音乐特征。摘要:This paper explores the automated process of determining stem compatibility by identifying audio recordings of single instruments that blend well with a given musical context. To tackle this challenge, we present Stem-JEPA, a novel Joint-Embedding Predictive Architecture (JEPA) trained on a multi-track dataset using a self-supervised learning approach. Our model comprises two networks: an encoder and a predictor, which are jointly trained to predict the embeddings of compatible stems from the embeddings of a given context, typically a mix of several instruments. Training a model in this manner allows its use in estimating stem compatibility - retrieving, aligning, or generating a stem to match a given mix - or for downstream tasks such as genre or key estimation, as the training paradigm requires the model to learn information related to timbre, harmony, and rhythm. We evaluate our model's performance on a retrieval task on the MUSDB18 dataset, testing its ability to find the missing stem from a mix and through a subjective user study. We also show that the learned embeddings capture temporal alignment information and, finally, evaluate the representations learned by our model on several downstream tasks, highlighting that they effectively capture meaningful musical features.
【4】 Steer-by-prior Editing of Symbolic Music Loops
标题: 符号音乐循环的优先引导编辑
作者:Nicolas Jonason,Luca Casini,Bob L. T. Sturm
备注:Accepted to MML 2024
链接:点击下载PDF文件
摘要:与建立一个系统的目标,能够控制的符号音乐循环生成和编辑,本文探讨了一个概括的掩蔽语言建模,我们称之为叠加语言建模。叠加语言模型不是已知或未知的输入标记,而是将序列的先验信息作为输入,使我们能够在推理时对生成应用各种约束。在详细介绍了我们的方法后,我们展示了我们的模型在多仪器循环领域的各种编辑任务。最后,我们强调的方法和未来工作的途径的一些局限性。我们在https: erl-j.github.io slm-mml-demo 上提供了SLM跨多个生成和编辑任务的示例。摘要:With the goal of building a system capable of controllable symbolic music loop generation and editing, this paper explores a generalisation of Masked Language Modelling we call Superposed Language Modelling. Rather than input tokens being known or unknown, a Superposed Language Model takes priors over the sequence as input, enabling us to apply various constraints to the generation at inference time. After detailing our approach, we demonstrate our model across various editing tasks in the domain of multi-instrument MIDI loops. We end by highlighting some limitations of the approach and avenues for future work. We provides examples from the SLM across multiple generation and editing tasks at https: erl-j.github.io slm-mml-demo .
【5】 An approach to optimize inference of the DIART speaker diarization pipeline
标题: 一种优化DIART扬声器拨号流水线推理的方法
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages, 3 figures
链接:点击下载PDF文件
摘要:发言人日记回答了音频文件的“谁在什么时候发言”的问题。在某些日志化场景中,转录需要低延迟。具有低延迟的说话人日记被称为在线说话人日记。DIART管道是一个在线发言人日记系统。它由分割和嵌入模型组成。嵌入模型在整个延迟中所占的份额最大。本文的目的是优化DIART流水线的推理延迟。在流水线嵌入模型中采用了知识提取、剪枝、量化、层次融合等推理优化方法。事实证明,知识蒸馏优化了延迟,但对准确性有负面影响。量化和层融合也对延迟有积极的影响,而不会恶化准确性。另一方面,修剪不会改善延迟。摘要:Speaker diarization answers the question "who spoke when" for an audio file. In some diarization scenarios, low latency is required for transcription. Speaker diarization with low latency is referred to as online speaker diarization. The DIART pipeline is an online speaker diarization system. It consists of a segmentation and an embedding model. The embedding model has the largest share of the overall latency. The aim of this paper is to optimize the inference latency of the DIART pipeline. Different inference optimization methods such as knowledge distilation, pruning, quantization and layer fusion are applied to the embedding model of the pipeline. It turns out that knowledge distillation optimizes the latency, but has a negative effect on the accuracy. Quantization and layer fusion also have a positive influence on the latency without worsening the accuracy. Pruning, on the other hand, does not improve latency.
【6】 Diseño de sonido para producciones audiovisuales e historias sonoras en el aula. Hacia una docencia creativa mediante el uso de herramientas inteligentes
标题: 《索尼多》是为了制作视听和历史而创作的。Hacia una docencia mediana el uso de herramientas intelligentes
作者:Miguel Civit,Francisco Cuadrado
备注:11 pages, in Spanish language. 1 figure. In La nueva era del p'odcast
链接:点击下载PDF文件
摘要:本研究旨在分享教学经验的教学声音设计的视听产品和比较不同的项目,学生处理。本报告的目的不是对不同类型的教学进行比较分析,而是分析在不同年级学习该科目的不同学生中观察到的不同问题。对于很大一部分学生来说,音频世界可能非常有趣,无论是那些具有创意倾向还是技术倾向的学生。音乐创作和制作、图像同步、配音等。这些学科通常很有趣,但由于技术复杂,进入门槛很高。有时,新手可能需要几周甚至几个月的时间才能开始轻松地使用音频编辑程序,这对学生来说并不总是特别直观。根据我们的经验,通过使用PBL方法进行学习,其结果远远优于通过使用其他教学方法(如大师班)所观察到的结果。学生在开发他们亲自参与的创造性项目的同时获得技术技能。尽管有上述内容,教师和学生之间的大多数互动都集中在技术纠正方面。从混响中的不同参数(如预延迟,衰减,调制.)如何正确调整压缩机、噪声门等;处理音频的工具数量非常广泛,其许多功能也会因制造商的不同而存在严重差异。摘要:This study aims to share a teaching experience teaching sound design for audiovisual productions and compares different projects tackled by students. It is not intended to be a comparative analysis of different types of teaching but rather an analysis of different problems observed in different profiles of students of the subject who study it in different grades. The world of audio can be very interesting for a large part of the students, both those with creative and technical inclinations. Musical creation and production, synchronization with images, dubbing, etc. They are disciplines that are generally interesting but can have a very high barrier to entry due to their great technical complexity. Sometimes it can take weeks or even months for the uninitiated to begin to use audio editing programs with the necessary ease, which are not always particularly intuitive for students. Learning through the use of PBL methodologies generates, in our experience, results much superior to those that can be observed through the use of other teaching methods such as master classes. Students acquire technical skills while developing creative projects in which they get personally involved. Despite everything mentioned above, most interactions between teachers and students focus on aspects of technical correction. From different parameters in reverbs (such as pre-delay, decay, modulation...) to how to correctly adjust compressors, noise gates, etc.; The number of tools with which to work with audio is incredibly extensive, as well as many of its features that can present serious differences depending on their manufacturers.
【7】 Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
标题: 基于对比学习的多语言声控关联链簇
作者:Wuyang Chen,Yanjie Sun,Kele Xu,Yong Dou
链接:点击下载PDF文件
摘要:一个人的脸和声音之间的内在相关性最近已经成为一个引人注目的研究领域,特别是在多语言环境中。本文介绍了我们的新解决方案,在多语言环境(FAME)2024年的挑战,重点是一个基于对比学习的链式聚类方法,以提高脸声关联。这项任务涉及在听觉和视觉模态线索之间建立生物识别关系的挑战,以及对不同语言之间的韵律相互依赖性进行建模,同时解决数据中存在的内在和外在变异性。为了应对这些重大挑战,我们的方法采用了监督交叉对比(SCC)学习,以在多语言场景中建立语音和人脸之间的强大关联。在此之后,我们专门设计了一个基于链簇的后处理步骤,以减轻在野生数据中经常发现的离群值的影响。我们进行了大量的实验,以调查语言的影响,面孔-声音的联想。在FAME公共评估平台上对整体结果进行了评估,我们获得了第二名。结果表明,我们的方法的优越性能,我们验证了我们提出的方法的鲁棒性和有效性。代码可在https: github.com colaudiolab FAME24_solution上获得。摘要:The innate correlation between a person's face and voice has recently emerged as a compelling area of study, especially within the context of multilingual environments. This paper introduces our novel solution to the Face-Voice Association in Multilingual Environments (FAME) 2024 challenge, focusing on a contrastive learning-based chaining-cluster method to enhance face-voice association. This task involves the challenges of building biometric relations between auditory and visual modality cues and modelling the prosody interdependence between different languages while addressing both intrinsic and extrinsic variability present in the data. To handle these non-trivial challenges, our method employs supervised cross-contrastive (SCC) learning to establish robust associations between voices and faces in multi-language scenarios. Following this, we have specifically designed a chaining-cluster-based post-processing step to mitigate the impact of outliers often found in unconstrained in the wild data. We conducted extensive experiments to investigate the impact of language on face-voice association. The overall results were evaluated on the FAME public evaluation platform, where we achieved 2nd place. The results demonstrate the superior performance of our method, and we validate the robustness and effectiveness of our proposed approach. Code is available at https: github.com colaudiolab FAME24_solution.
【8】 Joint Learning of Emotions in Music and Generalized Sounds
标题: 音乐和广义声音中情感的联合学习
作者:Simonetta Federico,Certo Francesca,Ntalampiras Stavros
备注:Accepted at Audio Mostly 2024, Milan
链接:点击下载PDF文件
摘要:在这项研究中,我们的目标是确定广义的声音和音乐是否可以共享一个共同的情感空间,提高情绪的唤醒和效价方面的预测。我们建议使用多个数据集作为多域学习技术。我们的方法包括创建一个共同的空间,包含概括的声音和音乐的特征,因为它们可以以类似的方式唤起情感。为了实现这一目标,我们利用了两个公开的数据集,即IADS-E和PMEmo,遵循标准化的实验方案。我们采用了各种各样的功能,捕捉音频结构的各个方面,包括频谱,能量和发声的关键参数。随后,我们利用异构模型架构对公共特征空间进行联合学习。有趣的是,这种协同方案在声音和音乐情感预测方面都优于最先进的方法。实现所呈现的实验流水线的完全复制的代码可在https: github.com LIMUNIMI MusicSoundEmotions获得。摘要:In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning technique. Our approach involves creating a common space encompassing features that characterize both generalized sounds and music, as they can evoke emotions in a similar manner. To achieve this, we utilized two publicly available datasets, namely IADS-E and PMEmo, following a standardized experimental protocol. We employed a wide variety of features that capture diverse aspects of the audio structure including key parameters of spectrum, energy, and voicing. Subsequently, we performed joint learning on the common feature space, leveraging heterogeneous model architectures. Interestingly, this synergistic scheme outperforms the state-of-the-art in both sound and music emotion prediction. The code enabling full replication of the presented experimental pipeline is available at https: github.com LIMUNIMI MusicSoundEmotions.
【9】 Why Perturbing Symbolic Music is Necessary: Fitting the Distribution of Never-used Notes through a Joint Probabilistic Diffusion Model
标题: 为什么有必要扰乱象征性音乐:通过联合概率扩散模型匹配从未使用过的音符的分布
作者:Shipei Liu,Xiaoya Fan,Guowei Wu
链接:点击下载PDF文件
摘要:现有的音乐生成模型大多是基于语言的,忽略了音符的频率连续性,导致对稀有或从未使用过的音符的拟合不足,从而降低了生成样本的多样性。我们认为,音符的分布可以通过平移不变性和周期性来建模,特别是使用扩散模型通过注入频域高斯噪声来概括音符。然而,由于音乐符号的低密度性质,估计高密度解空间中潜在的音符分布带来了重大挑战。为了解决这个问题,我们引入了音乐差异架构,它适合一个联合分布的音符和伴随的语义信息,以产生符号音乐条件。我们首先增强了碎片模块提取语义通过使用基于事件的符号和结构相似性指数,从而防止边界模糊。作为多变量扰动的先决条件,我们引入了一种联合预训练方法来构建音符和音乐语义之间的进行,同时避免了对低密度音符的直接建模。最后,我们通过一个多分支去噪器来恢复扰动的音符,该去噪器通过帕累托优化来适应多个噪声目标。我们的实验表明,与语言模型相比,联合概率扩散模型在音符和语义水平上的扰动可以提供更多的样本多样性和成分规律性。该案例研究通过分析自相似性度量中表达的层次结构,强调了我们的模型相对于基于语言和DDPM的模型的节奏优势。摘要:Existing music generation models are mostly language-based, neglecting the frequency continuity property of notes, resulting in inadequate fitting of rare or never-used notes and thus reducing the diversity of generated samples. We argue that the distribution of notes can be modeled by translational invariance and periodicity, especially using diffusion models to generalize notes by injecting frequency-domain Gaussian noise. However, due to the low-density nature of music symbols, estimating the distribution of notes latent in the high-density solution space poses significant challenges. To address this problem, we introduce the Music-Diff architecture, which fits a joint distribution of notes and accompanying semantic information to generate symbolic music conditionally. We first enhance the fragmentation module for extracting semantics by using event-based notations and the structural similarity index, thereby preventing boundary blurring. As a prerequisite for multivariate perturbation, we introduce a joint pre-training method to construct the progressions between notes and musical semantics while avoiding direct modeling of low-density notes. Finally, we recover the perturbed notes by a multi-branch denoiser that fits multiple noise objectives via Pareto optimization. Our experiments suggest that in contrast to language models, joint probability diffusion models perturbing at both note and semantic levels can provide more sample diversity and compositional regularity. The case study highlights the rhythmic advantages of our model over language- and DDPMs-based models by analyzing the hierarchical structure expressed in the self-similarity metrics.
【10】 ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic Features
标题: ALIF:使用语言特征对黑匣子语音平台进行低成本对抗性音频攻击
作者:Peng Cheng,Yuwei Wang,Peng Huang,Zhongjie Ba,Xiaodong Lin,Feng Lin,Li Lu,Kui Ren
备注:Published in the 2024 IEEE Symposium on Security and Privacy (SP)
链接:点击下载PDF文件
摘要:广泛的研究表明,对抗性示例(AE)对语音控制的智能设备构成了重大威胁。最近的研究提出了黑盒对抗攻击,只需要自动语音识别(ASR)系统的最终转录。然而,这些攻击通常涉及对ASR的许多查询,导致大量成本。此外,基于AE的对抗性音频样本容易受到ASR更新的影响。在本文中,我们确定了这些限制的根本原因,即无法直接围绕深度学习(DL)模型的决策边界构建AE攻击样本。基于这一观察,我们提出了ALIF,这是第一个基于黑盒对抗语言特征的攻击管道。我们利用文本到语音(TTS)和ASR模型的相互作用过程,在决策边界所在的语言嵌入空间中产生扰动。基于ALIF流水线,我们提出了ALIF-OTL和ALIF-OTA方案,用于在数字域和物理播放环境中对四个商业ASR和语音助手发起攻击。广泛的评估表明,ALIF-OTL和OTA显着提高查询效率分别为97.7%和73.3%,同时实现竞争力的性能相比,现有的方法。值得注意的是,ALIF-OTL可以仅使用一个查询生成攻击样本。此外,我们的时间测试实验验证了我们的方法对ASR更新的鲁棒性。摘要:Extensive research has revealed that adversarial examples (AE) pose a significant threat to voice-controllable smart devices. Recent studies have proposed black-box adversarial attacks that require only the final transcription from an automatic speech recognition (ASR) system. However, these attacks typically involve many queries to the ASR, resulting in substantial costs. Moreover, AE-based adversarial audio samples are susceptible to ASR updates. In this paper, we identify the root cause of these limitations, namely the inability to construct AE attack samples directly around the decision boundary of deep learning (DL) models. Building on this observation, we propose ALIF, the first black-box adversarial linguistic feature-based attack pipeline. We leverage the reciprocal process of text-to-speech (TTS) and ASR models to generate perturbations in the linguistic embedding space where the decision boundary resides. Based on the ALIF pipeline, we present the ALIF-OTL and ALIF-OTA schemes for launching attacks in both the digital domain and the physical playback environment on four commercial ASRs and voice assistants. Extensive evaluations demonstrate that ALIF-OTL and -OTA significantly improve query efficiency by 97.7% and 73.3%, respectively, while achieving competitive performance compared to existing methods. Notably, ALIF-OTL can generate an attack sample with only one query. Furthermore, our test-of-time experiment validates the robustness of our approach against ASR updates.
【11】 Generating High-quality Symbolic Music Using Fine-grained Discriminators
标题: 使用细粒度鉴别器生成高质量的象征音乐
作者:Zhedong Zhang,Liang Li,Jiehua Zhang,Zhenghui Hu,Hongkui Wang,Chenggang Yan,Jian Yang,Yuankai Qi
备注:Accepted by ICPR2024
链接:点击下载PDF文件
摘要:现有的符号化音乐生成方法通常通过对音乐的全局感知来利用符号化来提高生成音乐的质量。但是,考虑到音乐中信息的复杂性,如节奏和旋律,单一的旋律不能完全反映音乐这两个主要维度的差异。在这项工作中,我们提出了从音乐中分离旋律和节奏,并设计相应的细粒度判别器来解决上述问题。具体而言,配备了音高增强策略,旋律识别器辨别由所生成的样本呈现的旋律变化。相比之下,用小节级相对位置编码增强的节奏感集中在所生成音符的速度上。这样的设计允许生成器更明确地知道在生成的音乐中应该调整哪些方面,从而更容易模仿人类创作的音乐。POP909基准测试的实验结果表明,该方法的良好性能相比,几个国家的最先进的方法在客观和主观指标。摘要:Existing symbolic music generation methods usually utilize discriminator to improve the quality of generated music via global perception of music. However, considering the complexity of information in music, such as rhythm and melody, a single discriminator cannot fully reflect the differences in these two primary dimensions of music. In this work, we propose to decouple the melody and rhythm from music, and design corresponding fine-grained discriminators to tackle the aforementioned issues. Specifically, equipped with a pitch augmentation strategy, the melody discriminator discerns the melody variations presented by the generated samples. By contrast, the rhythm discriminator, enhanced with bar-level relative positional encoding, focuses on the velocity of generated notes. Such a design allows the generator to be more explicitly aware of which aspects should be adjusted in the generated music, making it easier to mimic human-composed music. Experimental results on the POP909 benchmark demonstrate the favorable performance of the proposed method compared to several state-of-the-art methods in terms of both objective and subjective metrics.
【12】 PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
标题: PiCoGen 2:采用迁移学习方法和弱对齐数据的钢琴封面生成
作者:Chih-Pin Tan,Hsin Ai,Yi-Hsin Chang,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
摘要:钢琴翻唱生成的目的是从流行歌曲中创建钢琴翻唱。现有的方法主要采用监督学习,训练需要通过将钢琴音符重新映射到歌曲音频来构建的高度对齐和配对的歌曲到钢琴数据。然而,这将导致钢琴信息的丢失,从而导致原始钢琴版本和重新映射的钢琴版本之间的不一致。为了克服这一限制,我们提出了一种迁移学习方法,该方法在仅钢琴数据上预训练我们的模型,并在没有音符重新映射的情况下构建的弱对齐配对数据上对其进行微调。在预训练过程中,为了引导模型学习钢琴作曲概念,而不仅仅是转录音频,我们使用现有的铅板转录模型作为编码器,从钢琴录音中提取高级特征。然后,预训练的模型在配对的歌曲-钢琴数据上进行微调,以将学习到的作曲知识转移到流行歌曲领域。我们的评估表明,这种训练策略使我们的模型PiCoGen 2能够获得高质量的结果,在五种流行音乐类型的客观和主观指标上都优于基线。摘要:Piano cover generation aims to create a piano cover from a pop song. Existing approaches mainly employ supervised learning and the training demands strongly-aligned and paired song-to-piano data, which is built by remapping piano notes to song audio. This would, however, result in the loss of piano information and accordingly cause inconsistencies between the original and remapped piano versions. To overcome this limitation, we propose a transfer learning approach that pre-trains our model on piano-only data and fine-tunes it on weakly-aligned paired data constructed without note remapping. During pre-training, to guide the model to learn piano composition concepts instead of merely transcribing audio, we use an existing lead sheet transcription model as the encoder to extract high-level features from the piano recordings. The pre-trained model is then fine-tuned on the paired song-piano data to transfer the learned composition knowledge to the pop song domain. Our evaluation shows that this training strategy enables our model, named PiCoGen2, to attain high-quality results, outperforming baselines on both objective and subjective metrics across five pop genres.
【13】 Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
标题: 视听Deepfake检测和定位的上下文跨模式注意力
作者:Vinaya Sree Katamneni,Ajita Rattani
链接:点击下载PDF文件
摘要:在数字时代,deepfakes和合成媒体的出现对社会和政治诚信构成了重大威胁。基于多模态操纵(如视听)的Deepfake更真实,威胁更大。目前的多模态深度伪造检测器通常基于对来自多个模态的异构数据流的基于注意力的融合。然而,数据的异质性(如音频和视觉信号)造成了分布模态差距,并对有效融合以及多模态深度伪造检测提出了重大挑战。在本文中,我们提出了一种基于递归神经网络(RNN)的新型多模态注意力框架,该框架利用上下文信息进行视听深度伪造检测。所提出的方法关注多模态多序列表示,并学习其中的贡献特征,以进行深度伪造检测和定位。在视听deepfake数据集(即FakeAVCeleb,AV-Deepfake 1 M,TVIL和LAV-DF数据集)上进行的彻底实验验证证明了我们方法的有效性。与已发表研究的交叉比较表明,我们的方法在deepfake检测和定位方面的准确度和精度分别提高了3.47%和2.05%。从而获得最先进的性能。为了便于再现,代码和数据集信息可在https: github.com vcbsl audiovisual-deepfake 上获得。摘要:In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic and pose a greater threat. Current multi-modal deepfake detectors are often based on the attention-based fusion of heterogeneous data streams from multiple modalities. However, the heterogeneous nature of the data (such as audio and visual signals) creates a distributional modality gap and poses a significant challenge in effective fusion and hence multi-modal deepfake detection. In this paper, we propose a novel multi-modal attention framework based on recurrent neural networks (RNNs) that leverages contextual information for audio-visual deepfake detection. The proposed approach applies attention to multi-modal multi-sequence representations and learns the contributing features among them for deepfake detection and localization. Thorough experimental validations on audio-visual deepfake datasets, namely FakeAVCeleb, AV-Deepfake1M, TVIL, and LAV-DF datasets, demonstrate the efficacy of our approach. Cross-comparison with the published studies demonstrates superior performance of our approach with an improved accuracy and precision by 3.47% and 2.05% in deepfake detection and localization, respectively. Thus, obtaining state-of-the-art performance. To facilitate reproducibility, the code and the datasets information is available at https: github.com vcbsl audiovisual-deepfake .
【14】 StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
标题: StreamVoice+:演变为端到端流媒体Zero-Shot语音转换
作者:Zhichao Wang,Yuanzhe Chen,Xinsheng Wang,Lei Xie,Yuping Wang
链接:点击下载PDF文件
摘要:StreamVoice最近推动了流媒体领域中zero-shot语音转换(VC)的边界。它使用一个流语言模型(LM)与上下文感知的方法来转换语义特征自动语音识别(ASR)到声学特征与所需的扬声器音色。尽管它的创新,StreamVoice面临的挑战,由于其依赖于流ASR在级联框架,这复杂的系统部署和优化,影响VC系统的设计和性能的基础上选择的ASR,并与转换稳定性的斗争时,面对低质量的语义输入。为了克服这些限制,我们引入了StreamVoice+,这是一种增强的基于LM的端到端流媒体框架,独立于流媒体ASR运行。StreamVoice+将语义编码器和连接器与原始的StreamVoice框架集成在一起,现在使用非流ASR进行训练。该模型经历了两个阶段的训练过程:首先,StreamVoice骨干进行语音转换的预训练,语义编码器进行鲁棒的语义提取。随后,该系统进行端到端微调,纳入LoRA矩阵以激活全面的流媒体功能。此外,StreamVoice+主要引入了两个战略增强来提高转换质量:连接器中的剩余补偿机制,以确保有效的语义传输;以及自细化策略,利用转换主干生成的伪并行语音对来改善语音解耦。实验表明,StreamVoice+不仅在语音转换中实现了比前代更高的自然度和说话人相似度,而且对流媒体和非流媒体转换场景都提供了多功能支持。摘要:StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to convert semantic features from automatic speech recognition (ASR) into acoustic features with the desired speaker timbre. Despite its innovations, StreamVoice faces challenges due to its dependency on a streaming ASR within a cascaded framework, which complicates system deployment and optimization, affects VC system's design and performance based on the choice of ASR, and struggles with conversion stability when faced with low-quality semantic inputs. To overcome these limitations, we introduce StreamVoice+, an enhanced LM-based end-to-end streaming framework that operates independently of streaming ASR. StreamVoice+ integrates a semantic encoder and a connector with the original StreamVoice framework, now trained using a non-streaming ASR. This model undergoes a two-stage training process: initially, the StreamVoice backbone is pre-trained for voice conversion and the semantic encoder for robust semantic extraction. Subsequently, the system is fine-tuned end-to-end, incorporating a LoRA matrix to activate comprehensive streaming functionality. Furthermore, StreamVoice+ mainly introduces two strategic enhancements to boost conversion quality: a residual compensation mechanism in the connector to ensure effective semantic transmission and a self-refinement strategy that leverages pseudo-parallel speech pairs generated by the conversion backbone to improve speech decoupling. Experiments demonstrate that StreamVoice+ not only achieves higher naturalness and speaker similarity in voice conversion than its predecessor but also provides versatile support for both streaming and non-streaming conversion scenarios.
eess.AS音频处理
【1】 Beyond Orthography: Automatic Recovery of Short Vowels and Dialectal Sounds in Arabic标题: 超越正音:阿拉伯语短元音和方言音的自动恢复
作者:Yassine El Kheir,Hamdy Mubarak,Ahmed Ali,Shammur Absar Chowdhury
备注:Accepted ACL 2024 Main Conference
链接:点击下载PDF文件
摘要:本文提出了一种新的方言语音和元音化恢复框架,旨在识别借用和方言语音在语音多样性和方言丰富的语言,超出其标准的正字法的声音集。所提出的框架利用了一个量化的输入序列与(出)连续的预训练的自监督表示。我们使用阿拉伯语的有限数据显示了管道的有效性,阿拉伯语是一种方言丰富的语言,包含超过22种主要方言。发音正确的阿拉伯方言转录语音资源是稀缺的。因此,我们推出了ArabVoice 15,这是首个同类的精心策划的测试集,包含15个阿拉伯国家的5个小时的方言语音,语音准确,包括借来的和方言特有的声音。我们详细地描述了注释指南以及方言混淆对的分析。我们广泛的评估包括主观-人类感知测试和客观措施。我们的经验结果,报告了三个测试集,表明只有一个半小时的训练数据,我们的模型在ArabVoice 15中的字符错误率比基线提高了约7%。摘要:This paper presents a novel Dialectal Sound and Vowelization Recovery framework, designed to recognize borrowed and dialectal sounds within phonologically diverse and dialect-rich languages, that extends beyond its standard orthographic sound sets. The proposed framework utilized a quantized sequence of input with(out) continuous pretrained self-supervised representation. We show the efficacy of the pipeline using limited data for Arabic, a dialect-rich language containing more than 22 major dialects. Phonetically correct transcribed speech resources for dialectal Arabic are scarce. Therefore, we introduce ArabVoice15, a first-of-its-kind, curated test set featuring 5 hours of dialectal speech across 15 Arab countries, with phonetically accurate transcriptions, including borrowed and dialect-specific sounds. We described in detail the annotation guideline along with the analysis of the dialectal confusion pairs. Our extensive evaluation includes both subjective -- human perception tests and objective measures. Our empirical results, reported with three test sets, show that with only one and half hours of training data, our model improve character error rate by ~ 7 % in ArabVoice15 compared to the baseline.
【2】 StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
标题: StreamVoice+:演变为端到端流媒体Zero-Shot语音转换
作者:Zhichao Wang,Yuanzhe Chen,Xinsheng Wang,Lei Xie,Yuping Wang
链接:点击下载PDF文件
摘要:StreamVoice最近推动了流媒体领域中zero-shot语音转换(VC)的边界。它使用一个流语言模型(LM)与上下文感知的方法来转换语义特征自动语音识别(ASR)到声学特征与所需的扬声器音色。尽管它的创新,StreamVoice面临的挑战,由于其依赖于流ASR在级联框架,这复杂的系统部署和优化,影响VC系统的设计和性能的基础上选择的ASR,并与转换稳定性的斗争时,面对低质量的语义输入。为了克服这些限制,我们引入了StreamVoice+,这是一种增强的基于LM的端到端流媒体框架,独立于流媒体ASR运行。StreamVoice+将语义编码器和连接器与原始的StreamVoice框架集成在一起,现在使用非流ASR进行训练。该模型经历了两个阶段的训练过程:首先,StreamVoice骨干进行语音转换的预训练,语义编码器进行鲁棒的语义提取。随后,该系统进行端到端微调,纳入LoRA矩阵以激活全面的流媒体功能。此外,StreamVoice+主要引入了两个战略增强来提高转换质量:连接器中的剩余补偿机制,以确保有效的语义传输;以及自细化策略,利用转换主干生成的伪并行语音对来改善语音解耦。实验表明,StreamVoice+不仅在语音转换中实现了比前代更高的自然度和说话人相似度,而且对流媒体和非流媒体转换场景都提供了多功能支持。摘要:StreamVoice has recently pushed the boundaries of zero-shot voice conversion (VC) in the streaming domain. It uses a streamable language model (LM) with a context-aware approach to convert semantic features from automatic speech recognition (ASR) into acoustic features with the desired speaker timbre. Despite its innovations, StreamVoice faces challenges due to its dependency on a streaming ASR within a cascaded framework, which complicates system deployment and optimization, affects VC system's design and performance based on the choice of ASR, and struggles with conversion stability when faced with low-quality semantic inputs. To overcome these limitations, we introduce StreamVoice+, an enhanced LM-based end-to-end streaming framework that operates independently of streaming ASR. StreamVoice+ integrates a semantic encoder and a connector with the original StreamVoice framework, now trained using a non-streaming ASR. This model undergoes a two-stage training process: initially, the StreamVoice backbone is pre-trained for voice conversion and the semantic encoder for robust semantic extraction. Subsequently, the system is fine-tuned end-to-end, incorporating a LoRA matrix to activate comprehensive streaming functionality. Furthermore, StreamVoice+ mainly introduces two strategic enhancements to boost conversion quality: a residual compensation mechanism in the connector to ensure effective semantic transmission and a self-refinement strategy that leverages pseudo-parallel speech pairs generated by the conversion backbone to improve speech decoupling. Experiments demonstrate that StreamVoice+ not only achieves higher naturalness and speaker similarity in voice conversion than its predecessor but also provides versatile support for both streaming and non-streaming conversion scenarios.
【3】 Re-ENACT: Reinforcement Learning for Emotional Speech Generation using Actor-Critic Strategy
标题: Re-ENACT:使用演员评论家策略进行情感语音生成的强化学习
作者:Ravi Shankar,Archana Venkataraman
备注:7 pages, 10 figures
链接:点击下载PDF文件
摘要:在本文中,我们提出了第一种方法来修改一个给定的语音信号的韵律特征,使用actor-critic强化学习策略。我们的方法使用贝叶斯框架,以确定连续段的重要性,链接段的给定的话语在人类的情感感知。我们训练一个神经网络来产生一组伯努利随机变量的变分后验;我们的模型在其上应用马尔可夫先验来确保连续性。来自该分布的样本用于下游情感预测。此外,我们训练神经网络来预测作为目标变量的情感类别的软分配。在接下来的步骤中,我们修改的韵律特征(音高,强度和节奏)的掩蔽段,以增加目标情绪的分数。我们采用演员-评论家强化学习训练的韵律修改器离散化的修改空间。此外,它提供了一个简单的解决方案的梯度计算问题,通过WSOLA操作的节奏操纵。我们的实验表明,这个框架改变了感知的情绪,一个给定的语音话语的目标。此外,我们还证明了我们的统一技术与来自监督和无监督领域的最先进的情感转换模型不相上下,这些模型需要成对训练。摘要:In this paper, we propose the first method to modify the prosodic features of a given speech signal using actor-critic reinforcement learning strategy. Our approach uses a Bayesian framework to identify contiguous segments of importance that links segments of the given utterances to perception of emotions in humans. We train a neural network to produce the variational posterior of a collection of Bernoulli random variables; our model applies a Markov prior on it to ensure continuity. A sample from this distribution is used for downstream emotion prediction. Further, we train the neural network to predict a soft assignment over emotion categories as the target variable. In the next step, we modify the prosodic features (pitch, intensity, and rhythm) of the masked segment to increase the score of target emotion. We employ an actor-critic reinforcement learning to train the prosody modifier by discretizing the space of modifications. Further, it provides a simple solution to the problem of gradient computation through WSOLA operation for rhythm manipulation. Our experiments demonstrate that this framework changes the perceived emotion of a given speech utterance to the target. Further, we show that our unified technique is on par with state-of-the-art emotion conversion models from supervised and unsupervised domains that require pairwise training.
【4】 Language Model Can Listen While Speaking
标题: 语言模型可以边听边说
作者:Ziyang Ma,Yakun Song,Chenpeng Du,Jian Cong,Zhuo Chen,Yuping Wang,Yuxuan Wang,Xie Chen
备注:Demo can be found at this https URL
链接:点击下载PDF文件
摘要:对话是最自然的人机交互方式。语音语言模型(SLM)的最新进展显着增强了基于语音的会话AI。然而,这些模型仅限于基于回合的对话,缺乏在实时口语场景中与人类交互的能力,例如,当生成的内容不令人满意时被中断。为了解决这些限制,我们探讨全双工建模(FDM)在交互式语音语言模型(iSLM),重点是增强实时交互,更明确地说,探索中断的本质能力。我们介绍了一种新的模型设计,即边听边说的语言模型(LSLM),一个端到端的系统配备了听和说通道。我们的LSLM采用基于令牌的解码器的语音生成和流自监督学习(SSL)编码器的实时音频输入的TTS。LSLM融合两个通道进行自回归生成,并实时检测话轮转换。三种融合策略-早期融合,中期融合,后期融合-探索,中期融合实现语音生成和实时交互之间的最佳平衡。两个实验设置,基于命令的FDM和基于语音的FDM,证明LSLM的鲁棒性噪声和灵敏度的不同的指令。我们的研究结果突出了LSLM的能力,实现双工通信,对现有系统的影响最小。本研究的目的是促进交互式语音对话系统的发展,提高其在现实世界中的适用性。摘要:Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM) have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based conversation, lacking the ability to interact with humans in real-time spoken scenarios, for example, being interrupted when the generated content is not satisfactory. To address these limitations, we explore full duplex modeling (FDM) in interactive speech language models (iSLM), focusing on enhancing real-time interaction and, more explicitly, exploring the quintessential ability of interruption. We introduce a novel model design, namely listening-while-speaking language model (LSLM), an end-to-end system equipped with both listening and speaking channels. Our LSLM employs a token-based decoder-only TTS for speech generation and a streaming self-supervised learning (SSL) encoder for real-time audio input. LSLM fuses both channels for autoregressive generation and detects turn-taking in real time. Three fusion strategies -- early fusion, middle fusion, and late fusion -- are explored, with middle fusion achieving an optimal balance between speech generation and real-time interaction. Two experimental settings, command-based FDM and voice-based FDM, demonstrate LSLM's robustness to noise and sensitivity to diverse instructions. Our results highlight LSLM's capability to achieve duplex communication with minimal impact on existing systems. This study aims to advance the development of interactive speech dialogue systems, enhancing their applicability in real-world contexts.
【5】 Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
标题: 集群和挖掘强调语音以实现包容性和公平的语音识别
作者:Jaeyoung Kim,Han Lu,Soheil Khorram,Anshuman Tripathi,Qian Zhang,Hasim Sak
链接:点击下载PDF文件
摘要:现代自动语音识别(ASR)系统通常在超过数万小时的语音数据上进行训练,这是其取得巨大成功的主要因素之一。然而,这样的数据的分布通常偏向于常见的口音或典型的语音模式。因此,这些系统通常对非典型口音语音表现不佳。在本文中,我们提出了公平的语音识别系统,可以表现得同样好,代表口音语音口音聚类和挖掘计划。对于口音识别,我们应用了三种方案来克服监督口音数据的有限大小:监督或无监督预训练,分布鲁棒优化(DRO)和无监督聚类。三种方案都能显著改善口音识别模型,尤其是对不平衡和小口音语音的识别效果。使用建议的监督或无监督聚类方案对挖掘的印度口音语音进行微调ASR,与对随机采样语音进行微调相比,分别显示出10.0%和5.3%的相对改善。摘要:Modern automatic speech recognition (ASR) systems are typically trained on more than tens of thousands hours of speech data, which is one of the main factors for their great success. However, the distribution of such data is typically biased towards common accents or typical speech patterns. As a result, those systems often poorly perform on atypical accented speech. In this paper, we present accent clustering and mining schemes for fair speech recognition systems which can perform equally well on under-represented accented speech. For accent recognition, we applied three schemes to overcome limited size of supervised accent data: supervised or unsupervised pre-training, distributionally robust optimization (DRO) and unsupervised clustering. Three schemes can significantly improve the accent recognition model especially for unbalanced and small accented speech. Fine-tuning ASR on the mined Indian accent speech using the proposed supervised or unsupervised clustering schemes showed 10.0% and 5.3% relative improvements compared to fine-tuning on the randomly sampled speech, respectively.
【6】 Stem-JEPA: A Joint-Embedding Predictive Architecture for Musical Stem Compatibility Estimation
标题: Stem-JEPA:用于音乐茎兼容性估计的联合嵌入预测架构
作者:Alain Riou,Stefan Lattner,Gaëtan Hadjeres,Michael Anslow,Geoffroy Peeters
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
摘要:本文探讨了确定干兼容性的自动化过程中,通过识别音频记录的单一乐器,以及与给定的音乐背景融合。为了应对这一挑战,我们提出了Stem-JEPA,这是一种新型的联合嵌入预测架构(JEPA),使用自监督学习方法在多轨道数据集上进行训练。 我们的模型包括两个网络:一个编码器和一个预测器,它们被联合训练以预测来自给定上下文的嵌入的兼容词干的嵌入,通常是几种乐器的混合。以这种方式训练模型允许其用于估计干兼容性-检索,对齐或生成干以匹配给定的混合-或用于下游任务,例如流派或基调估计,因为训练范例需要模型学习与音色,和声和节奏相关的信息。 我们评估了我们的模型在MUSDB 18数据集上的检索任务上的性能,测试了它从混合数据中找到缺失干的能力,并通过主观用户研究。我们还展示了学习的嵌入捕获时间对齐信息,最后,评估了我们的模型在几个下游任务中学习的表示,强调它们有效地捕获了有意义的音乐特征。摘要:This paper explores the automated process of determining stem compatibility by identifying audio recordings of single instruments that blend well with a given musical context. To tackle this challenge, we present Stem-JEPA, a novel Joint-Embedding Predictive Architecture (JEPA) trained on a multi-track dataset using a self-supervised learning approach. Our model comprises two networks: an encoder and a predictor, which are jointly trained to predict the embeddings of compatible stems from the embeddings of a given context, typically a mix of several instruments. Training a model in this manner allows its use in estimating stem compatibility - retrieving, aligning, or generating a stem to match a given mix - or for downstream tasks such as genre or key estimation, as the training paradigm requires the model to learn information related to timbre, harmony, and rhythm. We evaluate our model's performance on a retrieval task on the MUSDB18 dataset, testing its ability to find the missing stem from a mix and through a subjective user study. We also show that the learned embeddings capture temporal alignment information and, finally, evaluate the representations learned by our model on several downstream tasks, highlighting that they effectively capture meaningful musical features.
【7】 Steer-by-prior Editing of Symbolic Music Loops
标题: 符号音乐循环的优先引导编辑
作者:Nicolas Jonason,Luca Casini,Bob L. T. Sturm
备注:Accepted to MML 2024
链接:点击下载PDF文件
摘要:与建立一个系统的目标,能够控制的符号音乐循环生成和编辑,本文探讨了一个概括的掩蔽语言建模,我们称之为叠加语言建模。叠加语言模型不是已知或未知的输入标记,而是将序列的先验信息作为输入,使我们能够在推理时对生成应用各种约束。在详细介绍了我们的方法后,我们展示了我们的模型在多仪器循环领域的各种编辑任务。最后,我们强调的方法和未来工作的途径的一些局限性。我们在https: erl-j.github.io slm-mml-demo 上提供了SLM跨多个生成和编辑任务的示例。摘要:With the goal of building a system capable of controllable symbolic music loop generation and editing, this paper explores a generalisation of Masked Language Modelling we call Superposed Language Modelling. Rather than input tokens being known or unknown, a Superposed Language Model takes priors over the sequence as input, enabling us to apply various constraints to the generation at inference time. After detailing our approach, we demonstrate our model across various editing tasks in the domain of multi-instrument MIDI loops. We end by highlighting some limitations of the approach and avenues for future work. We provides examples from the SLM across multiple generation and editing tasks at https: erl-j.github.io slm-mml-demo .
【8】 An approach to optimize inference of the DIART speaker diarization pipeline
标题: 一种优化DIART扬声器拨号流水线推理的方法
作者:Roman Aperdannier,Sigurd Schacht,Alexander Piazza
备注:6 pages, 3 figures
链接:点击下载PDF文件
摘要:发言人日记回答了音频文件的“谁在什么时候发言”的问题。在某些日志化场景中,转录需要低延迟。具有低延迟的说话人日记被称为在线说话人日记。DIART管道是一个在线发言人日记系统。它由分割和嵌入模型组成。嵌入模型在整个延迟中所占的份额最大。本文的目的是优化DIART流水线的推理延迟。在流水线嵌入模型中采用了知识提取、剪枝、量化、层次融合等推理优化方法。事实证明,知识蒸馏优化了延迟,但对准确性有负面影响。量化和层融合也对延迟有积极的影响,而不会恶化准确性。另一方面,修剪不会改善延迟。摘要:Speaker diarization answers the question "who spoke when" for an audio file. In some diarization scenarios, low latency is required for transcription. Speaker diarization with low latency is referred to as online speaker diarization. The DIART pipeline is an online speaker diarization system. It consists of a segmentation and an embedding model. The embedding model has the largest share of the overall latency. The aim of this paper is to optimize the inference latency of the DIART pipeline. Different inference optimization methods such as knowledge distilation, pruning, quantization and layer fusion are applied to the embedding model of the pipeline. It turns out that knowledge distillation optimizes the latency, but has a negative effect on the accuracy. Quantization and layer fusion also have a positive influence on the latency without worsening the accuracy. Pruning, on the other hand, does not improve latency.
【9】 Diseño de sonido para producciones audiovisuales e historias sonoras en el aula. Hacia una docencia creativa mediante el uso de herramientas inteligentes
标题: 《索尼多》是为了制作视听和历史而创作的。Hacia una docencia mediana el uso de herramientas intelligentes
作者:Miguel Civit,Francisco Cuadrado
备注:11 pages, in Spanish language. 1 figure. In La nueva era del p'odcast
链接:点击下载PDF文件
摘要:本研究旨在分享教学经验的教学声音设计的视听产品和比较不同的项目,学生处理。本报告的目的不是对不同类型的教学进行比较分析,而是分析在不同年级学习该科目的不同学生中观察到的不同问题。音频的世界可以是非常有趣的大部分学生,无论是那些有创意和技术倾向。音乐创作和制作、图像同步、配音等。这些学科通常很有趣,但由于技术复杂,进入门槛很高。有时,新手可能需要几周甚至几个月的时间才能开始轻松地使用音频编辑程序,这对学生来说并不总是特别直观。根据我们的经验,通过使用PBL方法进行学习所产生的结果远优于通过使用大师班等其他教学方法所观察到的结果。学生在开发他们亲自参与的创造性项目的同时获得技术技能。尽管有上述内容,教师和学生之间的大多数互动都集中在技术纠正方面。从混响中的不同参数(如预延迟,衰减,调制.)如何正确调整压缩机、噪声门等;处理音频的工具数量非常广泛,其许多功能也会因制造商的不同而存在严重差异。摘要:This study aims to share a teaching experience teaching sound design for audiovisual productions and compares different projects tackled by students. It is not intended to be a comparative analysis of different types of teaching but rather an analysis of different problems observed in different profiles of students of the subject who study it in different grades. The world of audio can be very interesting for a large part of the students, both those with creative and technical inclinations. Musical creation and production, synchronization with images, dubbing, etc. They are disciplines that are generally interesting but can have a very high barrier to entry due to their great technical complexity. Sometimes it can take weeks or even months for the uninitiated to begin to use audio editing programs with the necessary ease, which are not always particularly intuitive for students. Learning through the use of PBL methodologies generates, in our experience, results much superior to those that can be observed through the use of other teaching methods such as master classes. Students acquire technical skills while developing creative projects in which they get personally involved. Despite everything mentioned above, most interactions between teachers and students focus on aspects of technical correction. From different parameters in reverbs (such as pre-delay, decay, modulation...) to how to correctly adjust compressors, noise gates, etc.; The number of tools with which to work with audio is incredibly extensive, as well as many of its features that can present serious differences depending on their manufacturers.
【10】 Contrastive Learning-based Chaining-Cluster for Multilingual Voice-Face Association
标题: 基于对比学习的多语言声控关联链簇
作者:Wuyang Chen,Yanjie Sun,Kele Xu,Yong Dou
链接:点击下载PDF文件
摘要:一个人的脸和声音之间的内在相关性最近已经成为一个引人注目的研究领域,特别是在多语言环境中。本文介绍了我们的新解决方案,在多语言环境(FAME)2024年的挑战,重点是一个基于对比学习的链式聚类方法,以提高脸声关联。这项任务涉及的挑战,建立听觉和视觉模态线索之间的生物识别关系和建模不同语言之间的韵律相互依存关系,同时解决内在和外在的变化存在于数据中。为了应对这些重大挑战,我们的方法采用了监督交叉对比(SCC)学习,以在多语言场景中建立语音和人脸之间的强大关联。在此之后,我们专门设计了一个基于链簇的后处理步骤,以减轻在野生数据中经常发现的离群值的影响。我们进行了大量的实验,以调查语言的影响,面孔-声音的联想。在FAME公共评估平台上对整体结果进行了评估,我们获得了第二名。结果表明,我们的方法的优越性能,我们验证了我们提出的方法的鲁棒性和有效性。代码可从https: github.com colaudiolab FAME24_solution获得。摘要:The innate correlation between a person's face and voice has recently emerged as a compelling area of study, especially within the context of multilingual environments. This paper introduces our novel solution to the Face-Voice Association in Multilingual Environments (FAME) 2024 challenge, focusing on a contrastive learning-based chaining-cluster method to enhance face-voice association. This task involves the challenges of building biometric relations between auditory and visual modality cues and modelling the prosody interdependence between different languages while addressing both intrinsic and extrinsic variability present in the data. To handle these non-trivial challenges, our method employs supervised cross-contrastive (SCC) learning to establish robust associations between voices and faces in multi-language scenarios. Following this, we have specifically designed a chaining-cluster-based post-processing step to mitigate the impact of outliers often found in unconstrained in the wild data. We conducted extensive experiments to investigate the impact of language on face-voice association. The overall results were evaluated on the FAME public evaluation platform, where we achieved 2nd place. The results demonstrate the superior performance of our method, and we validate the robustness and effectiveness of our proposed approach. Code is available at https: github.com colaudiolab FAME24_solution.
【11】 Joint Learning of Emotions in Music and Generalized Sounds
标题: 音乐和广义声音中情感的联合学习
作者:Simonetta Federico,Certo Francesca,Ntalampiras Stavros
备注:Accepted at Audio Mostly 2024, Milan
链接:点击下载PDF文件
摘要:在这项研究中,我们的目标是确定广义的声音和音乐是否可以共享一个共同的情感空间,提高情绪的唤醒和效价方面的预测。我们建议使用多个数据集作为多域学习技术。我们的方法包括创建一个共同的空间,包含概括的声音和音乐的特征,因为它们可以以类似的方式唤起情感。为了实现这一目标,我们利用了两个公开的数据集,即IADS-E和PMEmo,遵循标准化的实验方案。我们采用了各种各样的功能,捕捉音频结构的各个方面,包括频谱,能量和发声的关键参数。随后,我们利用异构模型架构对公共特征空间进行联合学习。有趣的是,这种协同方案在声音和音乐情感预测方面都优于最先进的方法。实现所呈现的实验流水线的完全复制的代码可在https: github.com LIMUNIMI MusicSoundEmotions获得。摘要:In this study, we aim to determine if generalized sounds and music can share a common emotional space, improving predictions of emotion in terms of arousal and valence. We propose the use of multiple datasets as a multi-domain learning technique. Our approach involves creating a common space encompassing features that characterize both generalized sounds and music, as they can evoke emotions in a similar manner. To achieve this, we utilized two publicly available datasets, namely IADS-E and PMEmo, following a standardized experimental protocol. We employed a wide variety of features that capture diverse aspects of the audio structure including key parameters of spectrum, energy, and voicing. Subsequently, we performed joint learning on the common feature space, leveraging heterogeneous model architectures. Interestingly, this synergistic scheme outperforms the state-of-the-art in both sound and music emotion prediction. The code enabling full replication of the presented experimental pipeline is available at https: github.com LIMUNIMI MusicSoundEmotions.
【12】 Why Perturbing Symbolic Music is Necessary: Fitting the Distribution of Never-used Notes through a Joint Probabilistic Diffusion Model
标题: 为什么有必要扰乱象征性音乐:通过联合概率扩散模型匹配从未使用过的音符的分布
作者:Shipei Liu,Xiaoya Fan,Guowei Wu
链接:点击下载PDF文件
摘要:现有的音乐生成模型大多是基于语言的,忽略了音符的频率连续性,导致对稀有或从未使用过的音符的拟合不足,从而降低了生成样本的多样性。我们认为,音符的分布可以通过平移不变性和周期性来建模,特别是使用扩散模型通过注入频域高斯噪声来概括音符。然而,由于音乐符号的低密度性质,估计高密度解空间中潜在的音符分布带来了重大挑战。为了解决这个问题,我们引入了Music-Diff架构,它适合音符和附带的语义信息的联合分布,以有条件地生成符号音乐。我们首先增强了碎片模块提取语义通过使用基于事件的符号和结构相似性指数,从而防止边界模糊。作为多变量扰动的先决条件,我们引入了一种联合预训练方法来构建音符和音乐语义之间的进行,同时避免了对低密度音符的直接建模。最后,我们通过一个多分支去噪器来恢复扰动的音符,该去噪器通过帕累托优化来适应多个噪声目标。我们的实验表明,与语言模型相比,联合概率扩散模型在音符和语义水平上的扰动可以提供更多的样本多样性和成分规律性。该案例研究通过分析自相似性度量中表达的层次结构,强调了我们的模型相对于基于语言和DDPM的模型的节奏优势。摘要:Existing music generation models are mostly language-based, neglecting the frequency continuity property of notes, resulting in inadequate fitting of rare or never-used notes and thus reducing the diversity of generated samples. We argue that the distribution of notes can be modeled by translational invariance and periodicity, especially using diffusion models to generalize notes by injecting frequency-domain Gaussian noise. However, due to the low-density nature of music symbols, estimating the distribution of notes latent in the high-density solution space poses significant challenges. To address this problem, we introduce the Music-Diff architecture, which fits a joint distribution of notes and accompanying semantic information to generate symbolic music conditionally. We first enhance the fragmentation module for extracting semantics by using event-based notations and the structural similarity index, thereby preventing boundary blurring. As a prerequisite for multivariate perturbation, we introduce a joint pre-training method to construct the progressions between notes and musical semantics while avoiding direct modeling of low-density notes. Finally, we recover the perturbed notes by a multi-branch denoiser that fits multiple noise objectives via Pareto optimization. Our experiments suggest that in contrast to language models, joint probability diffusion models perturbing at both note and semantic levels can provide more sample diversity and compositional regularity. The case study highlights the rhythmic advantages of our model over language- and DDPMs-based models by analyzing the hierarchical structure expressed in the self-similarity metrics.
【13】 ALIF: Low-Cost Adversarial Audio Attacks on Black-Box Speech Platforms using Linguistic Features
标题: ALIF:使用语言特征对黑匣子语音平台进行低成本对抗性音频攻击
作者:Peng Cheng,Yuwei Wang,Peng Huang,Zhongjie Ba,Xiaodong Lin,Feng Lin,Li Lu,Kui Ren
备注:Published in the 2024 IEEE Symposium on Security and Privacy (SP)
链接:点击下载PDF文件
摘要:广泛的研究表明,对抗性示例(AE)对语音控制的智能设备构成了重大威胁。最近的研究提出了黑盒对抗性攻击,仅需要自动语音识别(ASR)系统的最终转录。然而,这些攻击通常涉及对ASR的许多查询,导致大量成本。此外,基于AE的对抗性音频样本容易受到ASR更新的影响。在本文中,我们确定了这些限制的根本原因,即无法直接围绕深度学习(DL)模型的决策边界构建AE攻击样本。基于这一观察,我们提出了ALIF,这是第一个基于黑盒对抗语言特征的攻击管道。我们利用文本到语音(TTS)和ASR模型的相互作用过程,在决策边界所在的语言嵌入空间中产生扰动。基于ALIF流水线,我们提出了ALIF-OTL和ALIF-OTA方案,用于在数字域和物理播放环境中对四个商业ASR和语音助手发起攻击。广泛的评估表明,ALIF-OTL和OTA显着提高查询效率分别为97.7%和73.3%,同时实现竞争力的性能相比,现有的方法。值得注意的是,ALIF-OTL可以仅使用一个查询生成攻击样本。此外,我们的时间测试实验验证了我们的方法对ASR更新的鲁棒性。摘要:Extensive research has revealed that adversarial examples (AE) pose a significant threat to voice-controllable smart devices. Recent studies have proposed black-box adversarial attacks that require only the final transcription from an automatic speech recognition (ASR) system. However, these attacks typically involve many queries to the ASR, resulting in substantial costs. Moreover, AE-based adversarial audio samples are susceptible to ASR updates. In this paper, we identify the root cause of these limitations, namely the inability to construct AE attack samples directly around the decision boundary of deep learning (DL) models. Building on this observation, we propose ALIF, the first black-box adversarial linguistic feature-based attack pipeline. We leverage the reciprocal process of text-to-speech (TTS) and ASR models to generate perturbations in the linguistic embedding space where the decision boundary resides. Based on the ALIF pipeline, we present the ALIF-OTL and ALIF-OTA schemes for launching attacks in both the digital domain and the physical playback environment on four commercial ASRs and voice assistants. Extensive evaluations demonstrate that ALIF-OTL and -OTA significantly improve query efficiency by 97.7% and 73.3%, respectively, while achieving competitive performance compared to existing methods. Notably, ALIF-OTL can generate an attack sample with only one query. Furthermore, our test-of-time experiment validates the robustness of our approach against ASR updates.
【14】 Generating High-quality Symbolic Music Using Fine-grained Discriminators
标题: 使用细粒度鉴别器生成高质量的象征音乐
作者:Zhedong Zhang,Liang Li,Jiehua Zhang,Zhenghui Hu,Hongkui Wang,Chenggang Yan,Jian Yang,Yuankai Qi
备注:Accepted by ICPR2024
链接:点击下载PDF文件
摘要:现有的符号化音乐生成方法通常通过对音乐的全局感知来利用符号化来提高生成音乐的质量。但是,考虑到音乐中信息的复杂性,如节奏和旋律,单一的旋律不能完全反映音乐这两个主要维度的差异。在这项工作中,我们提出了从音乐中分离旋律和节奏,并设计相应的细粒度判别器来解决上述问题。具体而言,配备了音高增强策略,旋律识别器辨别由所生成的样本呈现的旋律变化。相比之下,用小节级相对位置编码增强的节奏感集中在所生成音符的速度上。这样的设计允许生成器更明确地知道在生成的音乐中应该调整哪些方面,从而更容易模仿人类创作的音乐。POP909基准测试的实验结果表明,该方法的良好性能相比,几个国家的最先进的方法在客观和主观指标。摘要:Existing symbolic music generation methods usually utilize discriminator to improve the quality of generated music via global perception of music. However, considering the complexity of information in music, such as rhythm and melody, a single discriminator cannot fully reflect the differences in these two primary dimensions of music. In this work, we propose to decouple the melody and rhythm from music, and design corresponding fine-grained discriminators to tackle the aforementioned issues. Specifically, equipped with a pitch augmentation strategy, the melody discriminator discerns the melody variations presented by the generated samples. By contrast, the rhythm discriminator, enhanced with bar-level relative positional encoding, focuses on the velocity of generated notes. Such a design allows the generator to be more explicitly aware of which aspects should be adjusted in the generated music, making it easier to mimic human-composed music. Experimental results on the POP909 benchmark demonstrate the favorable performance of the proposed method compared to several state-of-the-art methods in terms of both objective and subjective metrics.
【15】 PiCoGen2: Piano cover generation with transfer learning approach and weakly aligned data
标题: PiCoGen 2:采用迁移学习方法和弱对齐数据的钢琴封面生成
作者:Chih-Pin Tan,Hsin Ai,Yi-Hsin Chang,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
摘要:钢琴翻唱生成的目的是从流行歌曲中创建钢琴翻唱。现有的方法主要采用监督学习,训练需要通过将钢琴音符重新映射到歌曲音频来构建的高度对齐和配对的歌曲到钢琴数据。然而,这将导致钢琴信息的丢失,从而导致原始钢琴版本和重新映射的钢琴版本之间的不一致。为了克服这一限制,我们提出了一种迁移学习方法,该方法在仅钢琴数据上预训练我们的模型,并在没有音符重新映射的情况下构建的弱对齐配对数据上对其进行微调。在预训练过程中,为了引导模型学习钢琴作曲概念,而不仅仅是转录音频,我们使用现有的铅板转录模型作为编码器,从钢琴录音中提取高级特征。然后,预训练的模型在配对的歌曲-钢琴数据上进行微调,以将学习到的作曲知识转移到流行歌曲领域。我们的评估表明,这种训练策略使我们的模型PiCoGen 2能够获得高质量的结果,在五种流行音乐类型的客观和主观指标上都优于基线。摘要:Piano cover generation aims to create a piano cover from a pop song. Existing approaches mainly employ supervised learning and the training demands strongly-aligned and paired song-to-piano data, which is built by remapping piano notes to song audio. This would, however, result in the loss of piano information and accordingly cause inconsistencies between the original and remapped piano versions. To overcome this limitation, we propose a transfer learning approach that pre-trains our model on piano-only data and fine-tunes it on weakly-aligned paired data constructed without note remapping. During pre-training, to guide the model to learn piano composition concepts instead of merely transcribing audio, we use an existing lead sheet transcription model as the encoder to extract high-level features from the piano recordings. The pre-trained model is then fine-tuned on the paired song-piano data to transfer the learned composition knowledge to the pop song domain. Our evaluation shows that this training strategy enables our model, named PiCoGen2, to attain high-quality results, outperforming baselines on both objective and subjective metrics across five pop genres.
【16】 Contextual Cross-Modal Attention for Audio-Visual Deepfake Detection and Localization
标题: 视听Deepfake检测和定位的上下文跨模式注意力
作者:Vinaya Sree Katamneni,Ajita Rattani
链接:点击下载PDF文件
摘要:在数字时代,deepfakes和合成媒体的出现对社会和政治诚信构成了重大威胁。基于多模态操纵(如视听)的Deepfake更真实,威胁更大。目前的多模态深度伪造检测器通常基于对来自多个模态的异构数据流的基于注意力的融合。然而,数据的异质性(如音频和视觉信号)造成了分布模态差距,并对有效融合以及多模态深度伪造检测提出了重大挑战。在本文中,我们提出了一种基于递归神经网络(RNN)的新型多模态注意力框架,该框架利用上下文信息进行视听深度伪造检测。所提出的方法关注多模态多序列表示,并学习其中的贡献特征,以进行深度伪造检测和定位。在视听deepfake数据集(即FakeAVCeleb,AV-Deepfake 1 M,TVIL和LAV-DF数据集)上进行的彻底实验验证证明了我们方法的有效性。与已发表研究的交叉比较表明,我们的方法在deepfake检测和定位方面的准确度和精度分别提高了3.47%和2.05%。从而获得最先进的性能。为了便于再现,代码和数据集信息可在https: github.com vcbsl audiovisual-deepfake 上获得。摘要:In the digital age, the emergence of deepfakes and synthetic media presents a significant threat to societal and political integrity. Deepfakes based on multi-modal manipulation, such as audio-visual, are more realistic and pose a greater threat. Current multi-modal deepfake detectors are often based on the attention-based fusion of heterogeneous data streams from multiple modalities. However, the heterogeneous nature of the data (such as audio and visual signals) creates a distributional modality gap and poses a significant challenge in effective fusion and hence multi-modal deepfake detection. In this paper, we propose a novel multi-modal attention framework based on recurrent neural networks (RNNs) that leverages contextual information for audio-visual deepfake detection. The proposed approach applies attention to multi-modal multi-sequence representations and learns the contributing features among them for deepfake detection and localization. Thorough experimental validations on audio-visual deepfake datasets, namely FakeAVCeleb, AV-Deepfake1M, TVIL, and LAV-DF datasets, demonstrate the efficacy of our approach. Cross-comparison with the published studies demonstrates superior performance of our approach with an improved accuracy and precision by 3.47% and 2.05% in deepfake detection and localization, respectively. Thus, obtaining state-of-the-art performance. To facilitate reproducibility, the code and the datasets information is available at https: github.com vcbsl audiovisual-deepfake .
机器翻译,仅供参考
