今日论文合集:cs.SD语音8篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】As Good As A Coin Toss Human detection of AI-generated images, videos,  audio, and audiovisual stimuli

标题:人类检测人工智能生成的图像、视频、音频和视听刺激

链接:https://arxiv.org/abs/2403.16760

作者:Di Cooke,Abigail Edwards,Sophia Barkoff,Kathryn Kelly

备注:For study pre-registration, see this https URL

摘要:随着合成媒体变得越来越现实,使用它的障碍不断降低,这项技术越来越多地被用于恶意目的,从金融欺诈到未经同意的色情。今天,防止被合成媒体误导的主要防御手段依赖于人类观察者在视觉上和视觉上辨别真假的能力。然而,目前还不清楚人们在日常生活中对欺骗性合成媒体的脆弱程度。我们对1276名参与者进行了一项感知研究,以评估人们在区分合成图像、仅音频、仅视频和视听刺激与真实刺激时的准确性。为了反映人们在野外可能遇到合成媒体的情况,测试条件和刺激模拟了典型的在线平台,而调查中使用的所有合成媒体都来自可公开访问的生成人工智能技术。我们发现,总体而言,参与者很难有意义地区分合成内容和真实内容。我们还发现,当刺激包含合成内容相比,真实的内容,具有人脸的图像相比,非人脸对象,单一模态相比,多模态刺激,混合真实性相比,完全合成的视听刺激,和功能外语相比,观察者是流利的语言。最后,我们还发现,合成媒体的先验知识不会有意义地影响其检测性能。总的来说,这些结果表明,人们在日常生活中非常容易被合成媒体欺骗,人类的感知检测能力不再是一种有效的防御手段。

摘要:As synthetic media becomes progressively more realistic and barriers to using it continue to lower, the technology has been increasingly utilized for malicious purposes, from financial fraud to nonconsensual pornography. Today, the principal defense against being misled by synthetic media relies on the ability of the human observer to visually and auditorily discern between real and fake. However, it remains unclear just how vulnerable people actually are to deceptive synthetic media in the course of their day to day lives. We conducted a perceptual study with 1276 participants to assess how accurate people were at distinguishing synthetic images, audio only, video only, and audiovisual stimuli from authentic. To reflect the circumstances under which people would likely encounter synthetic media in the wild, testing conditions and stimuli emulated a typical online platform, while all synthetic media used in the survey was sourced from publicly accessible generative AI technology.  We find that overall, participants struggled to meaningfully discern between synthetic and authentic content. We also find that detection performance worsens when the stimuli contains synthetic content as compared to authentic content, images featuring human faces as compared to non face objects, a single modality as compared to multimodal stimuli, mixed authenticity as compared to being fully synthetic for audiovisual stimuli, and features foreign languages as compared to languages the observer is fluent in. Finally, we also find that prior knowledge of synthetic media does not meaningfully impact their detection performance. Collectively, these results indicate that people are highly susceptible to being tricked by synthetic media in their daily lives and that human perceptual detection capabilities can no longer be relied upon as an effective counterdefense.


【2】 Training Generative Adversarial Network-Based Vocoder with Limited Data  Using Augmentation-Conditional Discriminator
标题:用增广条件鉴别器训练有限数据生成对抗网络声码器
链接:https://arxiv.org/abs/2403.16464
作者:Takuhiro Kaneko,Hirokazu Kameoka,Kou Tanaka
备注:Accepted to ICASSP 2024. Project page: this https URL
摘要:基于生成对抗网络(GAN)的声码器由于其快速、轻量级和高质量的特性而通常用于语音合成。然而,这种数据驱动的模型需要大量的训练数据,从而导致高昂的数据收集成本。这一事实促使我们在有限的数据上训练基于GAN的声码器。一个有前途的解决方案是增加训练数据以避免过度拟合。然而,一个标准的分布是无条件的,对数据扩充引起的分布变化不敏感。因此,增强语音(其可以是非凡的)可以被认为是真实语音。为了解决这个问题,我们提出了一个增强条件的扩展(AugCondD),除了语音接收的增强状态作为输入,从而评估输入语音根据增强状态,而不抑制学习的原始非增强分布。实验结果表明,AugCondD提高语音质量在有限的数据条件下,同时实现足够的数据条件下的可比语音质量。音频样本可在https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/augcondd/上获得。
摘要:A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model requires a large amount of training data incurring high data-collection costs. This fact motivates us to train a GAN-based vocoder on limited data. A promising solution is to augment the training data to avoid overfitting. However, a standard discriminator is unconditional and insensitive to distributional changes caused by data augmentation. Thus, augmented speech (which can be extraordinary) may be considered real speech. To address this issue, we propose an augmentation-conditional discriminator (AugCondD) that receives the augmentation state as input in addition to speech, thereby assessing the input speech according to the augmentation state, without inhibiting the learning of the original non-augmented distribution. Experimental results indicate that AugCondD improves speech quality under limited data conditions while achieving comparable speech quality under sufficient data conditions. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/augcondd/.


【3】 Modeling Analog Dynamic Range Compressors using Deep Learning and  State-space Models
标题:基于深度学习和状态空间模型的模拟动态范围压缩器建模
链接:https://arxiv.org/abs/2403.16331
作者:Hanzhi Yin,Gang Cheng,Christian J. Steinmetz,Ruibin Yuan,Richard M. Stern,Roger B. Dannenberg
摘要:我们描述了一种新的方法来开发现实的数字模型的动态范围压缩器的数字音频制作,通过分析他们的模拟原型。虽然现实的数字动态压缩器对于许多应用是潜在有用的,但是设计过程是具有挑战性的,因为压缩器在长时间尺度上非线性地操作。我们的方法是基于结构化的状态空间序列模型(S4),实现状态空间模型(SSM)已被证明是有效的学习长程依赖关系,是有前途的建模动态范围压缩机。本文提出了一种具有S4层的深度学习模型,用于对Teletronix LA-2A模拟动态范围压缩器进行建模。该模型是因果关系的,实时有效地执行,并达到与以前的深度学习模型大致相同的质量,但参数更少。
摘要:We describe a novel approach for developing realistic digital models of dynamic range compressors for digital audio production by analyzing their analog prototypes. While realistic digital dynamic compressors are potentially useful for many applications, the design process is challenging because the compressors operate nonlinearly over long time scales. Our approach is based on the structured state space sequence model (S4), as implementing the state-space model (SSM) has proven to be efficient at learning long-range dependencies and is promising for modeling dynamic range compressors. We present in this paper a deep learning model with S4 layers to model the Teletronix LA-2A analog dynamic range compressor. The model is causal, executes efficiently in real time, and achieves roughly the same quality as previous deep-learning models but with fewer parameters.

【4】 Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover  Strategy
标题:基于预训练AV—HuBERT和掩蔽—恢复策略的目标语音提取
链接:https://arxiv.org/abs/2403.16078
作者:Wenxuan Wu,Xueyuan Chen,Xixin Wu,Haizhou Li,Helen Meng
备注:Accepted by IJCNN 2024
摘要:视听目标语音提取(AV-TSE)是机器人技术和许多视听应用的使能技术之一。AV-TSE的挑战之一是如何有效地利用过程中的视听同步信息。AV-HuBERT可以是一个有用的唇读预训练模型,这还没有被AV-TSE采用。在本文中,我们想探索如何将预训练的AV-HuBERT集成到我们的AV-TSE系统中。我们有充分的理由期待业绩的改善。为了从模态间和模态内的相关性中受益,我们还提出了一种新的自监督学习的Mask-And-Recover(MAR)策略。在VoxCeleb 2数据集上的实验结果表明,我们提出的模型在主观和客观指标方面都优于基线,这表明预训练的AV-HuBERT模型为目标语音提取提供了更多信息的视觉线索。此外,通过比较研究,我们证实了所提出的屏蔽和恢复策略是显着有效的。
摘要:Audio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audio-visual applications. One of the challenges of AV-TSE is how to effectively utilize audio-visual synchronization information in the process. AV-HuBERT can be a useful pre-trained model for lip-reading, which has not been adopted by AV-TSE. In this paper, we would like to explore the way to integrate a pre-trained AV-HuBERT into our AV-TSE system. We have good reasons to expect an improved performance. To benefit from the inter and intra-modality correlations, we also propose a novel Mask-And-Recover (MAR) strategy for self-supervised learning. The experimental results on the VoxCeleb2 dataset show that our proposed model outperforms the baselines both in terms of subjective and objective metrics, suggesting that the pre-trained AV-HuBERT model provides more informative visual cues for target speech extraction. Furthermore, through a comparative study, we confirm that the proposed Mask-And-Recover strategy is significantly effective.

【5】 Music to Dance as Language Translation using Sequence Models
标题:利用序列模型实现音乐到舞蹈的语言翻译
链接:https://arxiv.org/abs/2403.15569
作者:André Correia,Luís A. Alexandre
摘要:从音乐中合成合适的舞蹈编排仍然是一个悬而未决的问题。我们介绍MDLT,一种新的方法,框架的编排生成问题的翻译任务。我们的方法利用现有的数据集来学习将音频序列翻译成相应的舞蹈姿势。我们提出了两个变种的MDLT:一个利用Transformer架构和其他采用Mamba架构。我们训练我们的方法在AIST++和AntoTomDance数据集教机器人手臂跳舞,但我们的方法可以应用于一个完整的人形机器人。评估指标,包括平均关节误差和Frechet初始距离,一致表明,当给定一段音乐时,MDLT擅长制作逼真和高质量的编舞。代码可以在github.com/meowatthemoon/MDLT上找到。
摘要:Synthesising appropriate choreographies from music remains an open problem. We introduce MDLT, a novel approach that frames the choreography generation problem as a translation task. Our method leverages an existing data set to learn to translate sequences of audio into corresponding dance poses. We present two variants of MDLT: one utilising the Transformer architecture and the other employing the Mamba architecture. We train our method on AIST++ and PhantomDance data sets to teach a robotic arm to dance, but our method can be applied to a full humanoid robot. Evaluation metrics, including Average Joint Error and Frechet Inception Distance, consistently demonstrate that, when given a piece of music, MDLT excels at producing realistic and high-quality choreography. The code can be found at github.com/meowatthemoon/MDLT.

【6】 VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
标题:VoiceCraft:Zero-Shot语音编辑和文本语音转换
链接:https://arxiv.org/abs/2403.16973
作者:Puyuan Peng,Po-Yao Huang,Daniel Li,Abdelrahman Mohamed,David Harwath
备注:Data, code, and model weights are available at this https URL
摘要:我们介绍了VoiceCraft,一个令牌填充神经编解码器语言模型,它在有声读物,互联网视频和播客上的语音编辑和zero-shot文本到语音(TTS)方面都达到了最先进的性能。VoiceCraft采用了一个Transformer解码器架构,并引入了一个令牌重排过程,该过程结合了因果掩蔽和延迟堆叠,以实现在现有序列中的生成。在语音编辑任务中,VoiceCraft产生编辑后的语音,在自然度方面与未经编辑的录音几乎无法区分,正如人类所评估的那样;对于zero-shot TTS,我们的模型优于之前的SotA模型,包括VALLE和流行的商业模型XTTS-v2。至关重要的是,这些模型是在具有挑战性和现实性的数据集上进行评估的,这些数据集包括不同的口音、说话风格、录音条件、背景噪音和音乐,与其他模型和真实录音相比,我们的模型表现一贯良好。特别是,对于语音编辑评估,我们引入了一个高质量,具有挑战性和现实的数据集名为RealEdit。我们鼓励读者在https://jasonppy.github.io/VoiceCraft_web上收听演示。
摘要:We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts. VoiceCraft employs a Transformer decoder architecture and introduces a token rearrangement procedure that combines causal masking and delayed stacking to enable generation within an existing sequence. On speech editing tasks, VoiceCraft produces edited speech that is nearly indistinguishable from unedited recordings in terms of naturalness, as evaluated by humans; for zero-shot TTS, our model outperforms prior SotA models including VALLE and the popular commercial model XTTS-v2. Crucially, the models are evaluated on challenging and realistic datasets, that consist of diverse accents, speaking styles, recording conditions, and background noise and music, and our model performs consistently well compared to other models and real recordings. In particular, for speech editing evaluation, we introduce a high quality, challenging, and realistic dataset named RealEdit. We encourage readers to listen to the demos at https://jasonppy.github.io/VoiceCraft_web.

【7】 Distributed collaborative anomalous sound detection by embedding sharing
标题:基于嵌入共享的分布式协同异常声音检测
链接:https://arxiv.org/abs/2403.16610
作者:Kota Dohi,Yohei Kawaguchi
摘要:为了开发机器声音监测系统,提出了一种异常声音检测方法。在本文中,我们探索了一种多个客户端协作学习异常声音检测模型的方法,同时保持他们的原始数据彼此之间的私密性。在工业机器异常声音检测的背景下,每个客户端都拥有来自不同机器或不同操作状态的数据,这使得通过联合学习或分裂学习进行学习变得具有挑战性。在我们提出的方法中,每个客户端使用为声音数据分类开发的通用预训练模型计算嵌入,并且这些计算的嵌入在服务器上聚合,以通过离群值暴露执行异常声音检测。实验表明,我们提出的方法提高了异常声音检测的AUC平均为6.8%。
摘要:To develop a machine sound monitoring system, a method for detecting anomalous sound is proposed. In this paper, we explore a method for multiple clients to collaboratively learn an anomalous sound detection model while keeping their raw data private from each other. In the context of industrial machine anomalous sound detection, each client possesses data from different machines or different operational states, making it challenging to learn through federated learning or split learning. In our proposed method, each client calculates embeddings using a common pre-trained model developed for sound data classification, and these calculated embeddings are aggregated on the server to perform anomalous sound detection through outlier exposure. Experiments showed that our proposed method improves the AUC of anomalous sound detection by an average of 6.8%.

【8】 Towards auditory attention decoding with noise-tagging: A pilot study
标题:带噪声标签的听觉注意解码研究
链接:https://arxiv.org/abs/2403.15523
作者:H. A. Scheppink,S. Ahmadi,P. Desain,M. Tangermann,J. Thielen
备注:6 pages, 2 figures, 9th Graz Brain-Computer Interface Conference 2024
摘要:听觉注意解码(AAD)旨在从大脑活动中提取候选说话人中的被关注说话人,为神经操纵的听力设备和脑机接口提供有前途的应用。这项试点研究使第一步AAD使用噪声标记刺激协议,它唤起可靠的代码调制诱发电位,但最低限度地探讨了听觉模态。参与者依次呈现两个荷兰语的语音刺激,振幅调制与一个独特的二进制伪随机噪声码,有效地标记这些额外的可解码的信息。我们比较了未调制的音频与用各种调制深度调制的音频的解码,以及传统的AAD方法与解码噪声代码的标准方法。我们的试点研究显示,与未调制的音频相比,传统方法的性能更高,调制深度为70%至100%。噪声码解码器没有进一步改善这些结果。这些基本的见解突出了潜在的语音集成噪声代码,以提高听觉扬声器检测时,多个扬声器同时出现。
摘要:Auditory attention decoding (AAD) aims to extract from brain activity the attended speaker amidst candidate speakers, offering promising applications for neuro-steered hearing devices and brain-computer interfacing. This pilot study makes a first step towards AAD using the noise-tagging stimulus protocol, which evokes reliable code-modulated evoked potentials, but is minimally explored in the auditory modality. Participants were sequentially presented with two Dutch speech stimuli that were amplitude modulated with a unique binary pseudo-random noise-code, effectively tagging these with additional decodable information. We compared the decoding of unmodulated audio against audio modulated with various modulation depths, and a conventional AAD method against a standard method to decode noise-codes. Our pilot study revealed higher performances for the conventional method with 70 to 100 percent modulation depths compared to unmodulated audio. The noise-code decoder did not further improve these results. These fundamental insights highlight the potential of integrating noise-codes in speech to enhance auditory speaker detection when multiple speakers are presented simultaneously.

eess.AS音频处理
【1】 VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild
标题:VoiceCraft:Zero-Shot语音编辑和文本语音转换
链接:https://arxiv.org/abs/2403.16973
作者:Puyuan Peng,Po-Yao Huang,Daniel Li,Abdelrahman Mohamed,David Harwath
备注:Data, code, and model weights are available at this https URL
摘要:我们介绍了VoiceCraft,一个令牌填充神经编解码器语言模型,它在有声读物,互联网视频和播客上的语音编辑和zero-shot文本到语音(TTS)方面都达到了最先进的性能。VoiceCraft采用了一个Transformer解码器架构,并引入了一个令牌重排过程,该过程结合了因果掩蔽和延迟堆叠,以实现在现有序列中的生成。在语音编辑任务中,VoiceCraft产生编辑后的语音,在自然度方面与未经编辑的录音几乎无法区分,正如人类所评估的那样;对于zero-shot TTS,我们的模型优于之前的SotA模型,包括VALLE和流行的商业模型XTTS-v2。至关重要的是,这些模型是在具有挑战性和现实性的数据集上进行评估的,这些数据集包括不同的口音、说话风格、录音条件、背景噪音和音乐,与其他模型和真实录音相比,我们的模型表现一贯良好。特别是,对于语音编辑评估,我们引入了一个高质量,具有挑战性和现实的数据集名为RealEdit。我们鼓励读者在https://jasonppy.github.io/VoiceCraft_web上收听演示。
摘要:We introduce VoiceCraft, a token infilling neural codec language model, that achieves state-of-the-art performance on both speech editing and zero-shot text-to-speech (TTS) on audiobooks, internet videos, and podcasts. VoiceCraft employs a Transformer decoder architecture and introduces a token rearrangement procedure that combines causal masking and delayed stacking to enable generation within an existing sequence. On speech editing tasks, VoiceCraft produces edited speech that is nearly indistinguishable from unedited recordings in terms of naturalness, as evaluated by humans; for zero-shot TTS, our model outperforms prior SotA models including VALLE and the popular commercial model XTTS-v2. Crucially, the models are evaluated on challenging and realistic datasets, that consist of diverse accents, speaking styles, recording conditions, and background noise and music, and our model performs consistently well compared to other models and real recordings. In particular, for speech editing evaluation, we introduce a high quality, challenging, and realistic dataset named RealEdit. We encourage readers to listen to the demos at https://jasonppy.github.io/VoiceCraft_web.


【2】 Distributed collaborative anomalous sound detection by embedding sharing
标题:基于嵌入共享的分布式协同异常声音检测
链接:https://arxiv.org/abs/2403.16610
作者:Kota Dohi,Yohei Kawaguchi
摘要:为了开发机器声音监测系统,提出了一种异常声音检测方法。在本文中,我们探索了一种多个客户端协作学习异常声音检测模型的方法,同时保持他们的原始数据彼此之间的私密性。在工业机器异常声音检测的背景下,每个客户端都拥有来自不同机器或不同操作状态的数据,这使得通过联合学习或分裂学习进行学习变得具有挑战性。在我们提出的方法中,每个客户端使用为声音数据分类开发的通用预训练模型计算嵌入,并且这些计算的嵌入在服务器上聚合,以通过离群值暴露执行异常声音检测。实验表明,我们提出的方法提高了异常声音检测的AUC平均为6.8%。
摘要:To develop a machine sound monitoring system, a method for detecting anomalous sound is proposed. In this paper, we explore a method for multiple clients to collaboratively learn an anomalous sound detection model while keeping their raw data private from each other. In the context of industrial machine anomalous sound detection, each client possesses data from different machines or different operational states, making it challenging to learn through federated learning or split learning. In our proposed method, each client calculates embeddings using a common pre-trained model developed for sound data classification, and these calculated embeddings are aggregated on the server to perform anomalous sound detection through outlier exposure. Experiments showed that our proposed method improves the AUC of anomalous sound detection by an average of 6.8%.


【3】 Advanced Artificial Intelligence Algorithms in Cochlear Implants: Review  of Healthcare Strategies, Challenges, and Perspectives
标题:人工智能算法在人工耳蜗植入物:医疗保健战略,挑战和前景综述
链接:https://arxiv.org/abs/2403.15442
作者:Billel Essaid,Hamza Kheddar,Noureddine Batel,Abderrahmane Lakas,Muhammad E. H. Chowdhury
摘要:自动语音识别(ASR)在我们的日常生活中起着至关重要的作用,它不仅可以与机器进行交互,还可以为部分或严重听力障碍的个人提供沟通便利。该过程包括接收模拟形式的语音信号,然后通过各种信号处理算法使其与有限容量的设备(如人工耳蜗(CI))兼容。不幸的是,这些植入物配备了有限数量的电极,通常会导致合成过程中的语音失真。尽管研究人员努力使用各种最先进的信号处理技术来增强接收到的语音质量,但挑战仍然存在,特别是在涉及多个语音源、环境噪声和其他情况的场景中。新的人工智能(AI)方法的出现带来了尖端的策略,以解决与专用于CI的传统信号处理技术相关的限制和困难。本文旨在全面回顾基于CI的ASR和语音增强等相关方面的进展。主要目标是提供指标和数据集的全面概述,探索AI算法在生物医学领域的能力,总结和评论所获得的最佳结果。此外,审查将深入研究潜在的应用,并提出未来的方向,以弥补这一领域现有的研究差距。
摘要:Automatic speech recognition (ASR) plays a pivotal role in our daily lives, offering utility not only for interacting with machines but also for facilitating communication for individuals with either partial or profound hearing impairments. The process involves receiving the speech signal in analogue form, followed by various signal processing algorithms to make it compatible with devices of limited capacity, such as cochlear implants (CIs). Unfortunately, these implants, equipped with a finite number of electrodes, often result in speech distortion during synthesis. Despite efforts by researchers to enhance received speech quality using various state-of-the-art signal processing techniques, challenges persist, especially in scenarios involving multiple sources of speech, environmental noise, and other circumstances. The advent of new artificial intelligence (AI) methods has ushered in cutting-edge strategies to address the limitations and difficulties associated with traditional signal processing techniques dedicated to CIs. This review aims to comprehensively review advancements in CI-based ASR and speech enhancement, among other related aspects. The primary objective is to provide a thorough overview of metrics and datasets, exploring the capabilities of AI algorithms in this biomedical field, summarizing and commenting on the best results obtained. Additionally, the review will delve into potential applications and suggest future directions to bridge existing research gaps in this domain.

【4】 Encoding of lexical tone in self-supervised models of spoken language
标题:口语自我监督模型中的词汇声调编码
链接:https://arxiv.org/abs/2403.16865
作者:Gaofei Shen,Michaela Watkins,Afra Alishahi,Arianna Bisazza,Grzegorz Chrupała
备注:Accepted to NAACL 2024
摘要:可解释性研究表明,自监督口语语言模型(SLMs)编码的人类语音从声学,语音,音韵,句法和语义水平,说话人的特点,各种各样的功能。大量的语音表征的先前研究集中在音段特征,如音素;编码的超音段语音(如音调和重音模式)在SLM尚未得到很好的理解。声调是一种超音段特征,存在于世界上一半以上的语言中。本文以汉语和越南语为例,分析了SLMs的声调编码能力。我们发现,SLM编码的词汇音调显着的程度,即使当他们从非音调语言的数据进行训练。我们进一步发现,SLM的表现类似于本地和非本地的人类参与者的音调和辅音感知研究,但他们不遵循相同的发展轨迹。
摘要:Interpretability research has shown that self-supervised Spoken Language Models (SLMs) encode a wide variety of features in human speech from the acoustic, phonetic, phonological, syntactic and semantic levels, to speaker characteristics. The bulk of prior research on representations of phonology has focused on segmental features such as phonemes; the encoding of suprasegmental phonology (such as tone and stress patterns) in SLMs is not yet well understood. Tone is a suprasegmental feature that is present in more than half of the world's languages. This paper aims to analyze the tone encoding capabilities of SLMs, using Mandarin and Vietnamese as case studies. We show that SLMs encode lexical tone to a significant degree even when they are trained on data from non-tonal languages. We further find that SLMs behave similarly to native and non-native human participants in tone and consonant perception studies, but they do not follow the same developmental trajectory.


【5】 As Good As A Coin Toss Human detection of AI-generated images, videos,  audio, and audiovisual stimuli
标题:人工智能生成的图像、视频、音频和视听刺激的检测就像扔硬币一样好
链接:https://arxiv.org/abs/2403.16760
作者:Di Cooke,Abigail Edwards,Sophia Barkoff,Kathryn Kelly
备注:For study pre-registration, see this https URL
摘要:随着合成媒体变得越来越现实,使用它的障碍不断降低,这项技术越来越多地被用于恶意目的,从金融欺诈到未经同意的色情。今天,防止被合成媒体误导的主要防御手段依赖于人类观察者在视觉上和视觉上辨别真假的能力。然而,目前还不清楚人们在日常生活中对欺骗性合成媒体的脆弱程度。我们对1276名参与者进行了一项感知研究,以评估人们在区分合成图像、仅音频、仅视频和视听刺激与真实刺激时的准确性。为了反映人们在野外可能遇到合成媒体的情况,测试条件和刺激模拟了典型的在线平台,而调查中使用的所有合成媒体都来自可公开访问的生成人工智能技术。我们发现,总体而言,参与者很难有意义地区分合成内容和真实内容。我们还发现,当刺激包含合成内容相比,真实的内容,具有人脸的图像相比,非人脸对象,单一模态相比,多模态刺激,混合真实性相比,完全合成的视听刺激,和功能外语相比,观察者是流利的语言。最后,我们还发现,合成媒体的先验知识不会有意义地影响其检测性能。总的来说,这些结果表明,人们在日常生活中非常容易被合成媒体欺骗,人类的感知检测能力不再是一种有效的防御手段。
摘要:As synthetic media becomes progressively more realistic and barriers to using it continue to lower, the technology has been increasingly utilized for malicious purposes, from financial fraud to nonconsensual pornography. Today, the principal defense against being misled by synthetic media relies on the ability of the human observer to visually and auditorily discern between real and fake. However, it remains unclear just how vulnerable people actually are to deceptive synthetic media in the course of their day to day lives. We conducted a perceptual study with 1276 participants to assess how accurate people were at distinguishing synthetic images, audio only, video only, and audiovisual stimuli from authentic. To reflect the circumstances under which people would likely encounter synthetic media in the wild, testing conditions and stimuli emulated a typical online platform, while all synthetic media used in the survey was sourced from publicly accessible generative AI technology.  We find that overall, participants struggled to meaningfully discern between synthetic and authentic content. We also find that detection performance worsens when the stimuli contains synthetic content as compared to authentic content, images featuring human faces as compared to non face objects, a single modality as compared to multimodal stimuli, mixed authenticity as compared to being fully synthetic for audiovisual stimuli, and features foreign languages as compared to languages the observer is fluent in. Finally, we also find that prior knowledge of synthetic media does not meaningfully impact their detection performance. Collectively, these results indicate that people are highly susceptible to being tricked by synthetic media in their daily lives and that human perceptual detection capabilities can no longer be relied upon as an effective counterdefense.

【6】 Training Generative Adversarial Network-Based Vocoder with Limited Data  Using Augmentation-Conditional Discriminator
标题:利用增强-条件判别器训练有限数据的生成性对抗性网络声码器
链接:https://arxiv.org/abs/2403.16464
作者:Takuhiro Kaneko,Hirokazu Kameoka,Kou Tanaka
备注:Accepted to ICASSP 2024. Project page: this https URL
摘要:基于生成对抗网络(GAN)的声码器由于其快速、轻量级和高质量的特性而通常用于语音合成。然而,这种数据驱动的模型需要大量的训练数据,从而导致高昂的数据收集成本。这一事实促使我们在有限的数据上训练基于GAN的声码器。一个有前途的解决方案是增加训练数据以避免过度拟合。然而,一个标准的分布是无条件的,对数据扩充引起的分布变化不敏感。因此,增强语音(其可以是非凡的)可以被认为是真实语音。为了解决这个问题,我们提出了一个增强条件的扩展(AugCondD),除了语音接收的增强状态作为输入,从而评估输入语音根据增强状态,而不抑制学习的原始非增强分布。实验结果表明,AugCondD提高语音质量在有限的数据条件下,同时实现足够的数据条件下的可比语音质量。音频样本可在https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/augcondd/上获得。
摘要:A generative adversarial network (GAN)-based vocoder trained with an adversarial discriminator is commonly used for speech synthesis because of its fast, lightweight, and high-quality characteristics. However, this data-driven model requires a large amount of training data incurring high data-collection costs. This fact motivates us to train a GAN-based vocoder on limited data. A promising solution is to augment the training data to avoid overfitting. However, a standard discriminator is unconditional and insensitive to distributional changes caused by data augmentation. Thus, augmented speech (which can be extraordinary) may be considered real speech. To address this issue, we propose an augmentation-conditional discriminator (AugCondD) that receives the augmentation state as input in addition to speech, thereby assessing the input speech according to the augmentation state, without inhibiting the learning of the original non-augmented distribution. Experimental results indicate that AugCondD improves speech quality under limited data conditions while achieving comparable speech quality under sufficient data conditions. Audio samples are available at https://www.kecl.ntt.co.jp/people/kaneko.takuhiro/projects/augcondd/.

【7】 Music to Dance as Language Translation using Sequence Models
标题:利用序列模型实现音乐到舞蹈的语言翻译
链接:https://arxiv.org/abs/2403.15569
作者:André Correia,Luís A. Alexandre
摘要:从音乐中合成合适的舞蹈编排仍然是一个悬而未决的问题。我们介绍MDLT,一种新的方法,框架的编排生成问题的翻译任务。我们的方法利用现有的数据集来学习将音频序列翻译成相应的舞蹈姿势。我们提出了两个变种的MDLT:一个利用Transformer架构和其他采用Mamba架构。我们训练我们的方法在AIST++和AntoTomDance数据集教机器人手臂跳舞,但我们的方法可以应用于一个完整的人形机器人。评估指标,包括平均关节误差和Frechet初始距离,一致表明,当给定一段音乐时,MDLT擅长制作逼真和高质量的编舞。代码可以在github.com/meowatthemoon/MDLT上找到。
摘要:Synthesising appropriate choreographies from music remains an open problem. We introduce MDLT, a novel approach that frames the choreography generation problem as a translation task. Our method leverages an existing data set to learn to translate sequences of audio into corresponding dance poses. We present two variants of MDLT: one utilising the Transformer architecture and the other employing the Mamba architecture. We train our method on AIST++ and PhantomDance data sets to teach a robotic arm to dance, but our method can be applied to a full humanoid robot. Evaluation metrics, including Average Joint Error and Frechet Inception Distance, consistently demonstrate that, when given a piece of music, MDLT excels at producing realistic and high-quality choreography. The code can be found at github.com/meowatthemoon/MDLT.

【8】 Towards auditory attention decoding with noise-tagging: A pilot study
标题:利用噪声标记实现听觉注意解码:一项初步研究
链接:https://arxiv.org/abs/2403.15523
作者:H. A. Scheppink,S. Ahmadi,P. Desain,M. Tangermann,J. Thielen
备注:6 pages, 2 figures, 9th Graz Brain-Computer Interface Conference 2024
摘要:听觉注意解码(AAD)旨在从大脑活动中提取候选说话人中的被关注说话人,为神经操纵的听力设备和脑机接口提供有前途的应用。这项试点研究使第一步AAD使用噪声标记刺激协议,它唤起可靠的代码调制诱发电位,但最低限度地探讨了听觉模态。参与者依次呈现两个荷兰语的语音刺激,振幅调制与一个独特的二进制伪随机噪声码,有效地标记这些额外的可解码的信息。我们比较了未调制的音频与用各种调制深度调制的音频的解码,以及传统的AAD方法与解码噪声代码的标准方法。我们的试点研究显示,与未调制的音频相比,传统方法的性能更高,调制深度为70%至100%。噪声码解码器没有进一步改善这些结果。这些基本的见解突出了潜在的语音集成噪声代码,以提高听觉扬声器检测时,多个扬声器同时出现。
摘要:Auditory attention decoding (AAD) aims to extract from brain activity the attended speaker amidst candidate speakers, offering promising applications for neuro-steered hearing devices and brain-computer interfacing. This pilot study makes a first step towards AAD using the noise-tagging stimulus protocol, which evokes reliable code-modulated evoked potentials, but is minimally explored in the auditory modality. Participants were sequentially presented with two Dutch speech stimuli that were amplitude modulated with a unique binary pseudo-random noise-code, effectively tagging these with additional decodable information. We compared the decoding of unmodulated audio against audio modulated with various modulation depths, and a conventional AAD method against a standard method to decode noise-codes. Our pilot study revealed higher performances for the conventional method with 70 to 100 percent modulation depths compared to unmodulated audio. The noise-code decoder did not further improve these results. These fundamental insights highlight the potential of integrating noise-codes in speech to enhance auditory speaker detection when multiple speakers are presented simultaneously.

【9】 Privacy-Preserving End-to-End Spoken Language Understanding
标题:保护隐私的端到端口语理解
链接:https://arxiv.org/abs/2403.15510
作者:Yinggui Wang,Wei Huang,Le Yang
备注:Accepted by IJCAI
摘要:口语理解(SLU)是物联网设备中人机交互的关键技术之一,提供了易于使用的用户界面。人类语音可能包含大量用户敏感信息,如性别、身份和敏感内容。因此,出现了新型的安全和隐私侵犯行为。用户不希望将其个人敏感信息暴露给不受信任的第三方的恶意攻击。因此,SLU系统需要确保潜在的恶意攻击者不能推断出用户的敏感属性,同时应避免极大地损害SLU准确性。针对上述问题,提出了一种新的SLU多任务隐私保护模型,能够同时防止语音识别(ASR)和身份识别(IR)攻击。该模型采用隐层分离技术,使得SLU信息只分布在隐层的特定部分,而其他两类信息被移除,从而获得一个隐私安全的隐层。为了在效率和隐私之间取得良好的平衡,我们引入了一种新的模型预训练机制,即联合对抗训练,以进一步增强用户的隐私。在两个SLU数据集上的实验表明,该方法可以将ASR和IR攻击的准确性降低到接近随机猜测的准确性,而SLU性能基本不受影响。
摘要:Spoken language understanding (SLU), one of the key enabling technologies for human-computer interaction in IoT devices, provides an easy-to-use user interface. Human speech can contain a lot of user-sensitive information, such as gender, identity, and sensitive content. New types of security and privacy breaches have thus emerged. Users do not want to expose their personal sensitive information to malicious attacks by untrusted third parties. Thus, the SLU system needs to ensure that a potential malicious attacker cannot deduce the sensitive attributes of the users, while it should avoid greatly compromising the SLU accuracy. To address the above challenge, this paper proposes a novel SLU multi-task privacy-preserving model to prevent both the speech recognition (ASR) and identity recognition (IR) attacks. The model uses the hidden layer separation technique so that SLU information is distributed only in a specific portion of the hidden layer, and the other two types of information are removed to obtain a privacy-secure hidden layer. In order to achieve good balance between efficiency and privacy, we introduce a new mechanism of model pre-training, namely joint adversarial training, to further enhance the user privacy. Experiments over two SLU datasets show that the proposed method can reduce the accuracy of both the ASR and IR attacks close to that of a random guess, while leaving the SLU performance largely unaffected.

【10】 Isometric Neural Machine Translation using Phoneme Count Ratio  Reward-based Reinforcement Learning
标题:基于音素计数比奖励的强化学习等距神经机器翻译
链接:https://arxiv.org/abs/2403.15469
作者:Shivam Ratnakant Mhaskar,Nirmesh J. Shah,Mohammadi Zaki,Ashishkumar P. Gudmalwar,Pankaj Wasnik,Rajiv Ratn Shah
备注:Accepted in NAACL2024 Findings
摘要:传统的自动视频配音(AVD)流水线由三个关键模块组成,即自动语音识别(ASR),神经机器翻译(NMT)和文本到语音(TTS)。在AVD流水线中,采用等距NMT算法来调节合成输出文本的长度。这样做是为了保证在配音过程之后关于视频和音频的对准的同步。以前的方法集中在对齐机器翻译模型的源语言和目标语言文本中的字符和单词的数量。然而,我们的方法旨在调整音素的数量,因为它们与语音持续时间密切相关。在本文中,我们提出了使用强化学习(RL)的等距NMT系统的开发,重点是优化源语言和目标语言句子对中的音素计数的对齐。为了评估我们的模型,我们提出了音素计数依从性(PCC)分数,这是长度依从性的度量。我们的方法表明,在PCC分数相比,应用于英语-印地语语言对的最先进的模型,约36%的大幅改善。此外,我们提出了一个学生-教师架构的框架内,我们的RL方法,以保持音素计数和翻译质量之间的权衡。
摘要:Traditional Automatic Video Dubbing (AVD) pipeline consists of three key modules, namely, Automatic Speech Recognition (ASR), Neural Machine Translation (NMT), and Text-to-Speech (TTS). Within AVD pipelines, isometric-NMT algorithms are employed to regulate the length of the synthesized output text. This is done to guarantee synchronization with respect to the alignment of video and audio subsequent to the dubbing process. Previous approaches have focused on aligning the number of characters and words in the source and target language texts of Machine Translation models. However, our approach aims to align the number of phonemes instead, as they are closely associated with speech duration. In this paper, we present the development of an isometric NMT system using Reinforcement Learning (RL), with a focus on optimizing the alignment of phoneme counts in the source and target language sentence pairs. To evaluate our models, we propose the Phoneme Count Compliance (PCC) score, which is a measure of length compliance. Our approach demonstrates a substantial improvement of approximately 36% in the PCC score compared to the state-of-the-art models when applied to English-Hindi language pairs. Moreover, we propose a student-teacher architecture within the framework of our RL approach to maintain a trade-off between the phoneme count and translation quality.

机器翻译由腾讯交互翻译提供,仅供参考