今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音
【1】Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
标题:索尼克:肖像动画中将焦点转向全球音频感知
链接:https://arxiv.org/abs/2411.16331
作者:Xiaozhong Ji,  Xiaobin Hu,  Zhihong Xu,  Junwei Zhu,  Chuming Lin,  Qingdong He,  Jiangning Zhang,  Donghao Luo,  Yi Chen,  Qin Lin,  Qinglin Lu,  Chengjie Wang
备注:refer to our main-page \url{this https URL}
摘要:说话脸生成的研究主要探讨了同步面部运动和制作视觉上吸引人的,时间连贯的动画的复杂性。然而,由于对整体听觉感知的探索有限,目前的方法主要采用辅助的视觉和空间知识来稳定动作,这往往会导致自然度和时间不一致性的恶化。考虑到音频驱动动画的本质,音频信号作为理想的和唯一的先验来调整面部表情和嘴唇运动,而不借助于任何视觉信号的干扰。基于这一动机,我们提出了一种新的范式,称为Sonic,以{s}转移f{o}cus探索全局音频per{c}ept{i}o{n}.为了有效地利用全局音频知识,我们将其分解为片段内和片段间音频感知,并将两个方面合作以增强整体感知. \textbf{上下文增强音频学习},其中提取长距离片段内时间音频知识以提供隐含地表达为语音的音调和速度的面部表情和嘴唇运动先验。2)的情况。\textbf{运动解耦控制器},其中头部的运动和表情运动被分解并由内部音频剪辑独立控制。最重要的是,对于片段间音频感知,作为连接片段内音频的桥梁,实现全局感知,\textbf{Time-aware position shift fusion},其中考虑并融合全局片段间音频信息,通过连续的时间感知移位窗口进行长音频推理。大量的实验表明,新的音频驱动的范例优于现有的SOTA方法的视频质量,时间一致性,唇同步精度,和运动多样性。
摘要:The study of talking face generation mainly explores the intricacies ofsynchronizing facial movements and crafting visually appealing,temporally-coherent animations. However, due to the limited exploration ofglobal audio perception, current approaches predominantly employ auxiliaryvisual and spatial knowledge to stabilize the movements, which often results inthe deterioration of the naturalness and temporal inconsistencies.Consideringthe essence of audio-driven animation, the audio signal serves as the ideal andunique priors to adjust facial expressions and lip movements, without resortingto interference of any visual signals. Based on this motivation, we propose anovel paradigm, dubbed as Sonic, to {s}hift f{o}cus on the exploration ofglobal audio per{c}ept{i}o{n}.To effectively leverage global audio knowledge,we disentangle it into intra- and inter-clip audio perception and collaboratewith both aspects to enhance overall perception.For the intra-clip audioperception, 1). \textbf{Context-enhanced audio learning}, in which long-rangeintra-clip temporal audio knowledge is extracted to provide facial expressionand lip motion priors implicitly expressed as the tone and speed of speech. 2).\textbf{Motion-decoupled controller}, in which the motion of the head andexpression movement are disentangled and independently controlled byintra-audio clips. Most importantly, for inter-clip audio perception, as abridge to connect the intra-clips to achieve the global perception,\textbf{Time-aware position shift fusion}, in which the global inter-clip audioinformation is considered and fused for long-audio inference via throughconsecutively time-aware shifted windows. Extensive experiments demonstratethat the novel audio-driven paradigm outperform existing SOTA methodologies interms of video quality, temporally consistency, lip synchronization precision,and motion diversity.

【2】 The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC  Challenge 2024
标题:用于文本相关说话人验证(TdSV)的SVSVR系统AAIC挑战2024
链接:https://arxiv.org/abs/2411.16276
作者:Mohammadreza Molavi,  Reza Khodadadi
摘要:本文介绍了一种高效,准确的流水线文本相关的说话人验证(TDSV),旨在解决高性能的生物识别系统的需求。所提出的系统采用了一个基于快速一致性的ASR模块来验证语音内容,过滤掉目标错误(TW)和冒名顶替者错误(IW)的试验。对于说话人验证,我们提出了一种特征融合方法,该方法结合了从wav 2 vec-BERT和ReDimNet模型中提取的说话人嵌入,以创建统一的说话人表示。该系统在TDSV 2024 Challenge测试集上取得了有竞争力的结果,归一化min-DCF为0.0452(排名2),突出了其在平衡准确性和鲁棒性方面的有效性。
摘要:This paper introduces an efficient and accurate pipeline for text-dependentspeaker verification (TDSV), designed to address the need for high-performancebiometric systems. The proposed system incorporates a Fast-Conformer-based ASRmodule to validate speech content, filtering out Target-Wrong (TW) andImpostor-Wrong (IW) trials. For speaker verification, we propose a featurefusion approach that combines speaker embeddings extracted from wav2vec-BERTand ReDimNet models to create a unified speaker representation. This systemachieves competitive results on the TDSV 2024 Challenge test set, with anormalized min-DCF of 0.0452 (rank 2), highlighting its effectiveness inbalancing accuracy and robustness.

【3】 SKQVC: One-Shot Voice Conversion by K-Means Quantization with  Self-Supervised Speech Representations
标题:SKQVC:通过K-Means量化和自我监督语音表示的单次语音转换
链接:https://arxiv.org/abs/2411.16147
作者:Youngjun Sim,  Jinsung Yoon,  Young-Joo Suh
备注:5 pages
摘要:单次语音转换(VC)是一种仅使用单个目标说话人话语实现任意两个说话人之间的转换的方法。现有的方法通常依赖于复杂的架构和预先训练的说话人验证(SV)模型,以提高转换语音的保真度。最近的作品利用K均值量化(KQ)与自我监督学习(SSL)功能已被证明能够捕捉语音的内容信息。然而,它们通常难以保持说话的变化,例如韵律细节和语音变化,特别是对于较小的码本。在这项工作中,我们提出了一个简单而有效的一次性VC模型,利用SSL功能和语音属性的特点。我们的方法解决了丢失说话变化的问题,使高保真语音转换训练只重建损失,而不需要外部扬声器嵌入。我们展示了我们的模型在6个评估指标上的性能,结果突出了说话变化补偿方法的好处。
摘要:One-shot voice conversion (VC) is a method that enables the transformationbetween any two speakers using only a single target speaker utterance. Existingmethods often rely on complex architectures and pre-trained speakerverification (SV) models to improve the fidelity of converted speech. Recentworks utilizing K-means quantization (KQ) with self-supervised learning (SSL)features have proven capable of capturing content information from speech.However, they often struggle to preserve speaking variation, such as prosodicdetail and phonetic variation, particularly with smaller codebooks. In thiswork, we propose a simple yet effective one-shot VC model that utilizes thecharacteristics of SSL features and speech attributes. Our approach addressesthe issue of losing speaking variation, enabling high-fidelity voice conversiontrained with only reconstruction losses, without requiring external speakerembeddings. We demonstrate the performance of our model across 6 evaluationmetrics, with results highlighting the benefits of the speaking variationcompensation method.

【4】 A Training-Free Approach for Music Style Transfer with Latent Diffusion  Models
标题:具有潜在扩散模型的音乐风格转移免训练方法
链接:https://arxiv.org/abs/2411.15913
作者:Sooyoung Kim,  Joonwoo Kwon,  Heehwan Wang,  Shinjae Yoo,  Yuewei Lin,  Jiook Cha
备注:Codes will be released upon acceptance
摘要:音乐风格转换虽然为个性化音乐生成提供了令人兴奋的可能性,但通常需要大量的培训或详细的文本描述。本文介绍了一种利用预训练的潜在扩散模型(LDMs)的新的免训练方法。通过操纵LDM的自我注意特征,我们有效地将参考音乐的风格转移到内容音乐上,而无需额外的训练。我们的方法实现了优越的风格转移和旋律保存相比,现有的方法。这项工作为个性化音乐的生成开辟了新的创作途径。
摘要:Music style transfer, while offering exciting possibilities for personalizedmusic generation, often requires extensive training or detailed textualdescriptions. This paper introduces a novel training-free approach leveragingpre-trained Latent Diffusion Models (LDMs). By manipulating the self-attentionfeatures of the LDM, we effectively transfer the style of reference music ontocontent music without additional training. Our method achieves superior styletransfer and melody preservation compared to existing methods. This work opensnew creative avenues for personalized music generation.

【5】 Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video  Deepfake Dataset
标题:印地语音频视频-Deepfake(HAV-DF):基于印地语的音频视频Deepfake数据集
链接:https://arxiv.org/abs/2411.15457
作者:Sukhandeep Kaur,  Mubashir Buhari,  Naman Khandelwal,  Priyansh Tyagi,  Kiran Sharma
摘要:Deepfakes为创新和创造力提供了巨大的潜力,但它们也对隐私、信任和安全构成了重大风险。印度有大量讲印地语的人口,因此特别容易受到深度伪造驱动的错误信息运动的影响。印地语的假视频或演讲可能会对农村和半城市社区产生巨大影响,这些社区的数字素养往往较低,人们更倾向于信任视频内容。开发有效的框架和检测工具来打击deepfake滥用需要高质量,多样化和广泛的数据集。现有的流行数据集,如FF-DF(FaceForensics++)和DFDC(DeepFake Detection Challenge),都是基于英语的。因此,本文旨在创建第一个新的印地语深度假数据集,名为“印地语音频-视频-Deepfake”(HAV-DF)。该数据集是使用faceswap,lipsyn和语音克隆方法生成的。这个多步骤的过程使我们能够创建一个丰富多样的数据集,捕捉印地语语音和面部表情的细微差别,为在印地语环境中训练和评估deepfake检测模型提供了坚实的基础。它是独一无二的,因为所有以前的数据集都包含deepfake视频或合成音频。这种类型的deepfake数据集可用于训练deepfake视频和音频数据集的检测器。值得注意的是,新引入的HAV-DF数据集与其他知名数据集FF-DF和DFDC相比,在现有的检测方法(如Headpose,Xception-c40等)中显示出较低的检测准确性。这一趋势表明,HAV-DF数据集对检测提出了更深层次的挑战,可能是由于其专注于印地语内容和不同的操作技术。HAV-DF数据集填补了印地语特定的deepfake数据集的空白,有助于多语言deepfake检测开发。
摘要:Deepfakes offer great potential for innovation and creativity, but they alsopose significant risks to privacy, trust, and security. With a vastHindi-speaking population, India is particularly vulnerable to deepfake-drivenmisinformation campaigns. Fake videos or speeches in Hindi can have an enormousimpact on rural and semi-urban communities, where digital literacy tends to belower and people are more inclined to trust video content. The development ofeffective frameworks and detection tools to combat deepfake misuse requireshigh-quality, diverse, and extensive datasets. The existing popular datasetslike FF-DF (FaceForensics++), and DFDC (DeepFake Detection Challenge) are basedon English language.. Hence, this paper aims to create a first novel Hindi deepfake dataset, named ``Hindi audio-video-Deepfake'' (HAV-DF). The dataset hasbeen generated using the faceswap, lipsyn and voice cloning methods. Thismulti-step process allows us to create a rich, varied dataset that captures thenuances of Hindi speech and facial expressions, providing a robust foundationfor training and evaluating deepfake detection models in a Hindi languagecontext. It is unique of its kind as all of the previous datasets containeither deepfake videos or synthesized audio. This type of deepfake dataset canbe used for training a detector for both deepfake video and audio datasets.Notably, the newly introduced HAV-DF dataset demonstrates lower detectionaccuracy's across existing detection methods like Headpose, Xception-c40, etc.Compared to other well-known datasets FF-DF, and DFDC. This trend suggests thatthe HAV-DF dataset presents deeper challenges to detect, possibly due to itsfocus on Hindi language content and diverse manipulation techniques. The HAV-DFdataset fills the gap in Hindi-specific deepfake datasets, aiding multilingualdeepfake detection development.

【6】 Gotta Hear Them All: Sound Source Aware Vision to Audio Generation
标题:必须听到所有内容:从声音源感知视觉到音频生成
链接:https://arxiv.org/abs/2411.15447
作者:Wei Guo,  Heng Wang,  Weidong Cai,  Jianbo Ma
备注:16 pages, 9 figures, source code released at this https URL
摘要:视频到音频(V2A)合成在多媒体中有着广泛的应用。V2A方法的最新进展使得可以从视频或静止图像的输入生成相关音频。然而,这一代人的沉浸感和表现力是有限的。一个可能的问题是现有方法仅仅依赖于全局场景而忽略了局部探测对象的细节(即,声源)。为了解决这个问题,我们提出了一个声源感知V2A(SSV2A)发生器。SSV2A能够通过视觉检测和跨模态转换从场景中局部感知多模态声源。然后,它对比学习跨模态声源(CMSS)流形语义消歧每个源。最后,我们用心地将它们的CMSS语义混合到丰富的音频表示中,预训练的音频生成器从中输出声音。为了对CMSS流形进行建模,我们从VGGSound中策划了一个新的单声源视听数据集VGGS3。我们还设计了一个声源匹配分数来衡量本地化的音频相关性。据我们所知,这是解决声源级V2A生成的第一项工作。大量的实验表明,SSV2A超越国家的最先进的方法在生成保真度和相关性。我们进一步证明了SSV2A的能力,实现直观的V2A控制的合成视觉,文本和音频条件。我们的SSV2A一代可以在https://ssv2a.github.io/SSV2A-demo上试用和收听。
摘要:Vision-to-audio (V2A) synthesis has broad applications in multimedia. Recentadvancements of V2A methods have made it possible to generate relevant audiosfrom inputs of videos or still images. However, the immersiveness andexpressiveness of the generation are limited. One possible problem is thatexisting methods solely rely on the global scene and overlook details of localsounding objects (i.e., sound sources). To address this issue, we propose aSound Source-Aware V2A (SSV2A) generator. SSV2A is able to locally perceivemultimodal sound sources from a scene with visual detection and cross-modalitytranslation. It then contrastively learns a Cross-Modal Sound Source (CMSS)Manifold to semantically disambiguate each source. Finally, we attentively mixtheir CMSS semantics into a rich audio representation, from which a pretrainedaudio generator outputs the sound. To model the CMSS manifold, we curate anovel single-sound-source visual-audio dataset VGGS3 from VGGSound. We alsodesign a Sound Source Matching Score to measure localized audio relevance. Thisis to our knowledge the first work to address V2A generation at thesound-source level. Extensive experiments show that SSV2A surpassesstate-of-the-art methods in both generation fidelity and relevance. We furtherdemonstrate SSV2A's ability to achieve intuitive V2A control by compositingvision, text, and audio conditions. Our SSV2A generation can be tried and heardat https://ssv2a.github.io/SSV2A-demo .

eess.AS音频处理

【1】 State-Space Large Audio Language Models
标题:状态空间大型音频语言模型
链接:https://arxiv.org/abs/2411.15685
作者:Saurabhchand Bhati,  Yuan Gong,  Leonid Karlinsky,  Hilde Kuehne,  Rogerio Feris,  James Glass
摘要:大型音频语言模型(LALM)结合了音频感知模型和大型语言模型(LLM),并显示出对输入音频进行推理,推断含义和理解意图的显着能力。然而,这些系统依赖于Transformers,其与输入序列长度成二次方地缩放,这在存储器和时间受限的场景中部署这些系统时带来了计算挑战。最近,状态空间模型(SSM)已成为替代Transformer网络。  虽然已经成功地尝试用状态空间模型取代基于变压器的音频感知模型,但基于状态空间的LALM仍然未被探索。首先,我们首先替换基于transformer的音频感知模块,然后替换基于transformer的LLM,并提出第一个基于状态空间的LALM。实验结果表明,空间为基础的LALM,尽管有一个显着较低的参数数量进行竞争与基于transformer的LALM在各种数据集上的封闭式任务。
摘要:Large Audio Language Models (LALM) combine the audio perception models andthe Large Language Models (LLM) and show a remarkable ability to reason aboutthe input audio, infer the meaning, and understand the intent. However, thesesystems rely on Transformers which scale quadratically with the input sequencelengths which poses computational challenges in deploying these systems inmemory and time-constrained scenarios. Recently, the state-space models (SSMs)have emerged as an alternative to transformer networks. While there have been successful attempts to replace transformer-based audioperception models with state-space ones, state-space-based LALMs remainunexplored. First, we begin by replacing the transformer-based audio perceptionmodule and then replace the transformer-based LLM and propose the firststate-space-based LALM. Experimental results demonstrate that space-based LALMdespite having a significantly lower number of parameters performscompetitively with transformer-based LALMs on close-ended tasks on a variety ofdatasets.

【2】 Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
标题:索尼克:肖像动画中将焦点转向全球音频感知
链接:https://arxiv.org/abs/2411.16331
作者:Xiaozhong Ji,  Xiaobin Hu,  Zhihong Xu,  Junwei Zhu,  Chuming Lin,  Qingdong He,  Jiangning Zhang,  Donghao Luo,  Yi Chen,  Qin Lin,  Qinglin Lu,  Chengjie Wang
备注:refer to our main-page \url{this https URL}
摘要:说话脸生成的研究主要探讨了同步面部运动和制作视觉上吸引人的,时间连贯的动画的复杂性。然而,由于对整体听觉感知的探索有限,目前的方法主要采用辅助的视觉和空间知识来稳定动作,这往往会导致自然度和时间不一致性的恶化。考虑到音频驱动动画的本质,音频信号作为理想的和唯一的先验来调整面部表情和嘴唇运动,而不借助于任何视觉信号的干扰。基于这一动机,我们提出了一种新的范式,称为Sonic,以{s}转移f{o}cus探索全局音频per{c}ept{i}o{n}.为了有效地利用全局音频知识,我们将其分解为片段内和片段间音频感知,并将两个方面合作以增强整体感知. \textbf{上下文增强音频学习},其中提取长距离片段内时间音频知识以提供隐含地表达为语音的音调和速度的面部表情和嘴唇运动先验。2)的情况。\textbf{运动解耦控制器},其中头部的运动和表情运动被分解并由内部音频剪辑独立控制。最重要的是,对于片段间音频感知,作为连接片段内以实现全局感知的桥梁,\textbf{时间感知位置偏移融合},其中考虑全局片段间音频信息并通过连续的时间感知偏移窗口融合用于长音频推断。大量的实验表明,新的音频驱动的范例优于现有的SOTA方法的视频质量,时间一致性,唇同步精度,和运动多样性。
摘要:The study of talking face generation mainly explores the intricacies ofsynchronizing facial movements and crafting visually appealing,temporally-coherent animations. However, due to the limited exploration ofglobal audio perception, current approaches predominantly employ auxiliaryvisual and spatial knowledge to stabilize the movements, which often results inthe deterioration of the naturalness and temporal inconsistencies.Consideringthe essence of audio-driven animation, the audio signal serves as the ideal andunique priors to adjust facial expressions and lip movements, without resortingto interference of any visual signals. Based on this motivation, we propose anovel paradigm, dubbed as Sonic, to {s}hift f{o}cus on the exploration ofglobal audio per{c}ept{i}o{n}.To effectively leverage global audio knowledge,we disentangle it into intra- and inter-clip audio perception and collaboratewith both aspects to enhance overall perception.For the intra-clip audioperception, 1). \textbf{Context-enhanced audio learning}, in which long-rangeintra-clip temporal audio knowledge is extracted to provide facial expressionand lip motion priors implicitly expressed as the tone and speed of speech. 2).\textbf{Motion-decoupled controller}, in which the motion of the head andexpression movement are disentangled and independently controlled byintra-audio clips. Most importantly, for inter-clip audio perception, as abridge to connect the intra-clips to achieve the global perception,\textbf{Time-aware position shift fusion}, in which the global inter-clip audioinformation is considered and fused for long-audio inference via throughconsecutively time-aware shifted windows. Extensive experiments demonstratethat the novel audio-driven paradigm outperform existing SOTA methodologies interms of video quality, temporally consistency, lip synchronization precision,and motion diversity.

【3】 The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC  Challenge 2024
标题:用于文本相关说话人验证(TdSV)的SVSVR系统AAIC挑战2024
链接:https://arxiv.org/abs/2411.16276
作者:Mohammadreza Molavi,  Reza Khodadadi
摘要:本文介绍了一种高效,准确的流水线文本相关的说话人验证(TDSV),旨在解决高性能的生物识别系统的需求。所提出的系统采用了一个基于快速一致性的ASR模块来验证语音内容,过滤掉目标错误(TW)和冒名顶替者错误(IW)的试验。对于说话人验证,我们提出了一种特征融合方法,该方法结合了从wav 2 vec-BERT和ReDimNet模型中提取的说话人嵌入,以创建统一的说话人表示。该系统在TDSV 2024 Challenge测试集上取得了有竞争力的结果,归一化min-DCF为0.0452(排名2),突出了其在平衡准确性和鲁棒性方面的有效性。
摘要:This paper introduces an efficient and accurate pipeline for text-dependentspeaker verification (TDSV), designed to address the need for high-performancebiometric systems. The proposed system incorporates a Fast-Conformer-based ASRmodule to validate speech content, filtering out Target-Wrong (TW) andImpostor-Wrong (IW) trials. For speaker verification, we propose a featurefusion approach that combines speaker embeddings extracted from wav2vec-BERTand ReDimNet models to create a unified speaker representation. This systemachieves competitive results on the TDSV 2024 Challenge test set, with anormalized min-DCF of 0.0452 (rank 2), highlighting its effectiveness inbalancing accuracy and robustness.

【4】 SKQVC: One-Shot Voice Conversion by K-Means Quantization with  Self-Supervised Speech Representations
标题:SKQVC:通过K-Means量化和自我监督语音表示的单次语音转换
链接:https://arxiv.org/abs/2411.16147
作者:Youngjun Sim,  Jinsung Yoon,  Young-Joo Suh
备注:5 pages
摘要:单次语音转换(VC)是一种仅使用单个目标说话人话语实现任意两个说话人之间的转换的方法。现有的方法通常依赖于复杂的架构和预先训练的说话人验证(SV)模型,以提高转换语音的保真度。最近的作品利用K均值量化(KQ)与自我监督学习(SSL)功能已被证明能够捕捉语音的内容信息。然而,它们通常难以保持说话的变化,例如韵律细节和语音变化,特别是对于较小的码本。在这项工作中,我们提出了一个简单而有效的一次性VC模型,利用SSL功能和语音属性的特点。我们的方法解决了丢失说话变化的问题,使高保真语音转换训练只重建损失,而不需要外部扬声器嵌入。我们展示了我们的模型在6个评估指标上的性能,结果突出了说话变化补偿方法的好处。
摘要:One-shot voice conversion (VC) is a method that enables the transformationbetween any two speakers using only a single target speaker utterance. Existingmethods often rely on complex architectures and pre-trained speakerverification (SV) models to improve the fidelity of converted speech. Recentworks utilizing K-means quantization (KQ) with self-supervised learning (SSL)features have proven capable of capturing content information from speech.However, they often struggle to preserve speaking variation, such as prosodicdetail and phonetic variation, particularly with smaller codebooks. In thiswork, we propose a simple yet effective one-shot VC model that utilizes thecharacteristics of SSL features and speech attributes. Our approach addressesthe issue of losing speaking variation, enabling high-fidelity voice conversiontrained with only reconstruction losses, without requiring external speakerembeddings. We demonstrate the performance of our model across 6 evaluationmetrics, with results highlighting the benefits of the speaking variationcompensation method.

【5】 A Training-Free Approach for Music Style Transfer with Latent Diffusion  Models
标题:具有潜在扩散模型的音乐风格转移免训练方法
链接:https://arxiv.org/abs/2411.15913
作者:Sooyoung Kim,  Joonwoo Kwon,  Heehwan Wang,  Shinjae Yoo,  Yuewei Lin,  Jiook Cha
备注:Codes will be released upon acceptance
摘要:音乐风格转换虽然为个性化音乐生成提供了令人兴奋的可能性,但通常需要大量的培训或详细的文本描述。本文介绍了一种利用预训练的潜在扩散模型(LDMs)的新的免训练方法。通过操纵LDM的自我注意特征,我们有效地将参考音乐的风格转移到内容音乐上,而无需额外的训练。我们的方法实现了优越的风格转移和旋律保存相比,现有的方法。这项工作为个性化音乐的生成开辟了新的创作途径。
摘要:Music style transfer, while offering exciting possibilities for personalizedmusic generation, often requires extensive training or detailed textualdescriptions. This paper introduces a novel training-free approach leveragingpre-trained Latent Diffusion Models (LDMs). By manipulating the self-attentionfeatures of the LDM, we effectively transfer the style of reference music ontocontent music without additional training. Our method achieves superior styletransfer and melody preservation compared to existing methods. This work opensnew creative avenues for personalized music generation.

【6】 Hindi audio-video-Deepfake (HAV-DF): A Hindi language-based Audio-video  Deepfake Dataset
标题:印地语音频视频-Deepfake(HAV-DF):基于印地语的音频视频Deepfake数据集
链接:https://arxiv.org/abs/2411.15457
作者:Sukhandeep Kaur,  Mubashir Buhari,  Naman Khandelwal,  Priyansh Tyagi,  Kiran Sharma
摘要:Deepfake为创新和创造力提供了巨大的潜力,但也对隐私、信任和安全构成了重大风险。印度有大量讲印地语的人口,因此特别容易受到深度伪造驱动的错误信息运动的影响。印地语的假视频或演讲可能会对农村和半城市社区产生巨大影响,这些社区的数字素养往往较低,人们更倾向于信任视频内容。开发有效的框架和检测工具来打击deepfake滥用需要高质量,多样化和广泛的数据集。现有的流行数据集,如FF-DF(FaceForensics++)和DFDC(DeepFake Detection Challenge),都是基于英语的。因此,本文旨在创建第一个新的印地语深度假数据集,名为“印地语音频-视频-Deepfake”(HAV-DF)。该数据集是使用faceswap,lipsyn和语音克隆方法生成的。这个多步骤的过程使我们能够创建一个丰富多样的数据集,捕捉印地语语音和面部表情的细微差别,为在印地语环境中训练和评估deepfake检测模型提供了坚实的基础。它是独一无二的,因为所有以前的数据集都包含deepfake视频或合成音频。这种类型的deepfake数据集可用于训练deepfake视频和音频数据集的检测器。值得注意的是,新引入的HAV-DF数据集与其他知名数据集FF-DF和DFDC相比,在现有的检测方法(如Headpose,Xception-c40等)中显示出较低的检测准确性。这一趋势表明,HAV-DF数据集对检测提出了更深层次的挑战,可能是由于其专注于印地语内容和不同的操作技术。HAV-DF数据集填补了印地语特定深度伪造数据集的空白,有助于多语言深度伪造检测开发。
摘要:Deepfakes offer great potential for innovation and creativity, but they alsopose significant risks to privacy, trust, and security. With a vastHindi-speaking population, India is particularly vulnerable to deepfake-drivenmisinformation campaigns. Fake videos or speeches in Hindi can have an enormousimpact on rural and semi-urban communities, where digital literacy tends to belower and people are more inclined to trust video content. The development ofeffective frameworks and detection tools to combat deepfake misuse requireshigh-quality, diverse, and extensive datasets. The existing popular datasetslike FF-DF (FaceForensics++), and DFDC (DeepFake Detection Challenge) are basedon English language.. Hence, this paper aims to create a first novel Hindi deepfake dataset, named ``Hindi audio-video-Deepfake'' (HAV-DF). The dataset hasbeen generated using the faceswap, lipsyn and voice cloning methods. Thismulti-step process allows us to create a rich, varied dataset that captures thenuances of Hindi speech and facial expressions, providing a robust foundationfor training and evaluating deepfake detection models in a Hindi languagecontext. It is unique of its kind as all of the previous datasets containeither deepfake videos or synthesized audio. This type of deepfake dataset canbe used for training a detector for both deepfake video and audio datasets.Notably, the newly introduced HAV-DF dataset demonstrates lower detectionaccuracy's across existing detection methods like Headpose, Xception-c40, etc.Compared to other well-known datasets FF-DF, and DFDC. This trend suggests thatthe HAV-DF dataset presents deeper challenges to detect, possibly due to itsfocus on Hindi language content and diverse manipulation techniques. The HAV-DFdataset fills the gap in Hindi-specific deepfake datasets, aiding multilingualdeepfake detection development.

【7】 Gotta Hear Them All: Sound Source Aware Vision to Audio Generation
标题:必须听到所有内容:从声音源感知视觉到音频生成
链接:https://arxiv.org/abs/2411.15447
作者:Wei Guo,  Heng Wang,  Weidong Cai,  Jianbo Ma
备注:16 pages, 9 figures, source code released at this https URL
摘要:视频到音频(V2A)合成在多媒体中有着广泛的应用。V2A方法的最新进展使得可以从视频或静止图像的输入生成相关音频。然而,这一代人的沉浸感和表现力是有限的。一个可能的问题是现有方法仅仅依赖于全局场景而忽略了局部探测对象的细节(即,声源)。为了解决这个问题,我们提出了一个声源感知V2A(SSV2A)发生器。SSV2A能够通过视觉检测和跨模态转换从场景中局部感知多模态声源。然后,它对比学习跨模态声源(CMSS)流形语义消歧每个源。最后,我们用心地将它们的CMSS语义混合到丰富的音频表示中,预训练的音频生成器从中输出声音。为了对CMSS流形进行建模,我们从VGGSound中策划了一个新的单声源视听数据集VGGS3。我们还设计了一个声源匹配分数来衡量本地化的音频相关性。据我们所知,这是解决声源级V2A生成的第一项工作。大量的实验表明,SSV2A超越国家的最先进的方法在生成保真度和相关性。我们进一步证明了SSV2A的能力,实现直观的V2A控制的合成视觉,文本和音频条件。我们的SSV2A一代可以在https://ssv2a.github.io/SSV2A-demo上试用和收听。
摘要:Vision-to-audio (V2A) synthesis has broad applications in multimedia. Recentadvancements of V2A methods have made it possible to generate relevant audiosfrom inputs of videos or still images. However, the immersiveness andexpressiveness of the generation are limited. One possible problem is thatexisting methods solely rely on the global scene and overlook details of localsounding objects (i.e., sound sources). To address this issue, we propose aSound Source-Aware V2A (SSV2A) generator. SSV2A is able to locally perceivemultimodal sound sources from a scene with visual detection and cross-modalitytranslation. It then contrastively learns a Cross-Modal Sound Source (CMSS)Manifold to semantically disambiguate each source. Finally, we attentively mixtheir CMSS semantics into a rich audio representation, from which a pretrainedaudio generator outputs the sound. To model the CMSS manifold, we curate anovel single-sound-source visual-audio dataset VGGS3 from VGGSound. We alsodesign a Sound Source Matching Score to measure localized audio relevance. Thisis to our knowledge the first work to address V2A generation at thesound-source level. Extensive experiments show that SSV2A surpassesstate-of-the-art methods in both generation fidelity and relevance. We furtherdemonstrate SSV2A's ability to achieve intuitive V2A control by compositingvision, text, and audio conditions. Our SSV2A generation can be tried and heardat https://ssv2a.github.io/SSV2A-demo .

机器翻译由腾讯交互翻译提供,仅供参考