本文经arXiv每日学术速递授权转载
【1】AdvWave: Stealthy Adversarial Jailbreak Attack against Large Audio-Language Models
链接:https://arxiv.org/abs/2412.08608
摘要:大型音频语言模型(LALM)的最新进展已经实现了基于语音的用户交互,显著增强了用户体验并加速了LALM在现实世界应用中的部署。然而,确保LALM的安全性对于防止可能引起社会关注或违反人工智能法规的风险输出至关重要。尽管这个问题很重要,但对越狱LALM的研究仍然有限,因为它们最近才出现,而且与对基于DNN的音频模型的攻击相比,它们带来了额外的技术挑战。具体来说,LALM中的音频编码器涉及离散化操作,通常会导致梯度破碎,从而阻碍了依赖于基于梯度的优化的攻击的有效性。LALM的行为可变性进一步使有效(对抗性)优化目标的识别复杂化。此外,对对抗性音频波形实施隐形约束引入了减小的非凸可行解空间,进一步加剧了优化过程的挑战。为了克服这些挑战,我们开发了AdvWave,这是第一个针对LALM的越狱框架。我们提出了一种双阶段优化方法,解决梯度破碎,实现有效的端到端的基于梯度的优化。此外,我们开发了一种自适应对抗性目标搜索算法,该算法根据LALM对特定查询的响应模式动态调整对抗性优化目标。为了确保对抗性音频对人类听众保持感知自然,我们设计了一种分类器引导的优化方法,该方法生成类似于常见城市声音的对抗性噪声。对多个高级LALM的广泛评估表明,AdvWave优于基准方法,平均越狱攻击成功率提高了40%。
摘要:Recent advancements in large audio-language models (LALMs) have enabledspeech-based user interactions, significantly enhancing user experience andaccelerating the deployment of LALMs in real-world applications. However,ensuring the safety of LALMs is crucial to prevent risky outputs that may raisesocietal concerns or violate AI regulations. Despite the importance of thisissue, research on jailbreaking LALMs remains limited due to their recentemergence and the additional technical challenges they present compared toattacks on DNN-based audio models. Specifically, the audio encoders in LALMs,which involve discretization operations, often lead to gradient shattering,hindering the effectiveness of attacks relying on gradient-based optimizations.The behavioral variability of LALMs further complicates the identification ofeffective (adversarial) optimization targets. Moreover, enforcing stealthinessconstraints on adversarial audio waveforms introduces a reduced, non-convexfeasible solution space, further intensifying the challenges of theoptimization process. To overcome these challenges, we develop AdvWave, thefirst jailbreak framework against LALMs. We propose a dual-phase optimizationmethod that addresses gradient shattering, enabling effective end-to-endgradient-based optimization. Additionally, we develop an adaptive adversarialtarget search algorithm that dynamically adjusts the adversarial optimizationtarget based on the response patterns of LALMs for specific queries. To ensurethat adversarial audio remains perceptually natural to human listeners, wedesign a classifier-guided optimization approach that generates adversarialnoise resembling common urban sounds. Extensive evaluations on multipleadvanced LALMs demonstrate that AdvWave outperforms baseline methods, achievinga 40% higher average jailbreak attack success rate.
标题:Mel-Refine:一种即插即用方法在音频生成中完善Mel-Spectrum
链接:https://arxiv.org/abs/2412.08577
摘要:文本到音频(TTA)模型能够从文本提示生成各种音频。然而,大多数主流TTA模型主要依赖于Mel频谱图,在产生具有丰富内容的音频方面仍然面临挑战。梅尔声谱图对此类音频所需的复杂细节和纹理往往超过模型的能力,导致输出模糊或缺乏连贯性。在本文中,我们首先调查的关键作用,U-网络在梅尔频谱生成。我们的分析表明,在U-Net结构中,跳跃连接和主干中的高频分量影响纹理和细节,而主干中的低频分量对扩散去噪过程至关重要。我们进一步提出了“梅尔精炼”,即插即用的方法,增强梅尔频谱图纹理和细节,通过调整不同的分量权重在推理。我们的方法不需要额外的训练或微调,并且与任何基于扩散的TTA架构完全兼容。实验结果表明,我们的方法提高了最新的TTA模型Tango 2的性能指标的25%,证明了其有效性。
摘要:Text-to-audio (TTA) model is capable of generating diverse audio from textualprompts. However, most mainstream TTA models, which predominantly rely onMel-spectrograms, still face challenges in producing audio with rich content.The intricate details and texture required in Mel-spectrograms for such audiooften surpass the models' capacity, leading to outputs that are blurred or lackcoherence. In this paper, we begin by investigating the critical role of U-Netin Mel-spectrogram generation. Our analysis shows that in U-Net structure,high-frequency components in skip-connections and the backbone influencetexture and detail, while low-frequency components in the backbone are criticalfor the diffusion denoising process. We further propose ``Mel-Refine'', aplug-and-play approach that enhances Mel-spectrogram texture and detail byadjusting different component weights during inference. Our method requires noadditional training or fine-tuning and is fully compatible with anydiffusion-based TTA architecture. Experimental results show that our approachboosts performance metrics of the latest TTA model Tango2 by 25\%,demonstrating its effectiveness.
标题:Sketch 2Sound:通过时变信号和索尼克模仿实现可控音频生成
链接:https://arxiv.org/abs/2412.08550
摘要:我们提出了Sketch2Sound,这是一个生成音频模型,能够从一组可解释的时变控制信号创建高质量的声音:响度,亮度和音高,以及文本提示。Sketch2Sound可以从声音模仿中合成任意声音(即,声音模仿或参考声音形状)。Sketch2Sound可以在任何文本到音频的潜在扩散Transformer(DiT)之上实现,并且每个控件只需要40k步的微调和单个线性层,使其比ControlNet等现有方法更轻量级。为了从草图般的声音模仿合成,我们建议在训练期间将随机中值滤波器应用于控制信号,允许Sketch2Sound使用具有灵活水平的时间特异性的控制来提示。我们表明,Sketch2Sound可以合成的声音,按照输入控制的要点,从一个声乐模仿,同时保留遵守输入文本提示和音频质量相比,一个纯文本的基线。Sketch2Sound允许声音艺术家使用文本提示的语义灵活性和声音手势或声音模仿的表现力和精确度来创建声音。在https://hugofloresgarcia.art/sketch2sound/上可以找到正确的例子。
摘要:We present Sketch2Sound, a generative audio model capable of creatinghigh-quality sounds from a set of interpretable time-varying control signals:loudness, brightness, and pitch, as well as text prompts. Sketch2Sound cansynthesize arbitrary sounds from sonic imitations (i.e.,~a vocal imitation or areference sound-shape). Sketch2Sound can be implemented on top of anytext-to-audio latent diffusion transformer (DiT), and requires only 40k stepsof fine-tuning and a single linear layer per control, making it morelightweight than existing methods like ControlNet. To synthesize fromsketchlike sonic imitations, we propose applying random median filters to thecontrol signals during training, allowing Sketch2Sound to be prompted usingcontrols with flexible levels of temporal specificity. We show thatSketch2Sound can synthesize sounds that follow the gist of input controls froma vocal imitation while retaining the adherence to an input text prompt andaudio quality compared to a text-only baseline. Sketch2Sound allows soundartists to create sounds with the semantic flexibility of text prompts and theexpressivity and precision of a sonic gesture or vocal imitation. Soundexamples are available at https://hugofloresgarcia.art/sketch2sound/.
标题:对音乐生成模型的训练数据进行水印
链接:https://arxiv.org/abs/2412.08549
摘要:生成式人工智能(Gen-AI)模型越来越多地用于跨领域生成内容,包括文本、图像和音频。虽然这些模型代表了一项重大的技术突破,但它们的生成能力是通过对大量人类生成的内容进行训练而获得的,这些内容通常包括受版权保护的材料。在这项工作中,我们调查音频水印技术是否可以用来检测未经授权的使用内容训练的音乐生成模型。我们将在水印数据上训练的模型生成的输出与在非水印数据上训练的模型生成的输出进行比较。我们研究影响模型的生成行为的因素:水印技术,在训练集中的水印样本的比例,以及水印技术对模型的标记器的鲁棒性。我们的研究结果表明,音频水印技术,包括一些人类无法察觉的,可以导致模型的输出明显的变化。我们还研究了一个国家的最先进的水印技术的鲁棒性去除技术。
摘要:Generative Artificial Intelligence (Gen-AI) models are increasingly used toproduce content across domains, including text, images, and audio. While thesemodels represent a major technical breakthrough, they gain their generativecapabilities from being trained on enormous amounts of human-generated content,which often includes copyrighted material. In this work, we investigate whetheraudio watermarking techniques can be used to detect an unauthorized usage ofcontent to train a music generation model. We compare outputs generated by amodel trained on watermarked data to a model trained on non-watermarked data.We study factors that impact the model's generation behaviour: the watermarkingtechnique, the proportion of watermarked samples in the training set, and therobustness of the watermarking technique against the model's tokenizer. Ourresults show that audio watermarking techniques, including some that areimperceptible to humans, can lead to noticeable shifts in the model's outputs.We also study the robustness of a state-of-the-art watermarking technique toremoval techniques.
标题:PointTalk:音频驱动的动态唇点云,用于3D基于高斯的会说话的头部合成
链接:https://arxiv.org/abs/2412.08504
备注:9 pages, accepted by AAAI 2025
摘要:具有任意语音音频的说话头部合成是数字人领域的一个关键挑战。最近,基于辐射场的方法受到越来越多的关注,因为它们能够从几分钟的训练视频合成高保真和身份一致的说话头。然而,由于训练数据的规模有限,这些方法往往表现出较差的性能在音频唇同步和视觉质量。在本文中,我们提出了一种新的基于3D高斯的方法称为PointTalk,它构建了一个静态的头部3D高斯场,并使其与音频同步变形。它还采用了音频驱动的动态唇点云作为条件信息的关键组成部分,从而促进了说话的头部的有效合成。具体地,初始步骤涉及从音频信号生成对应的嘴唇点云并捕获其拓扑结构。动态差异编码器的设计旨在更有效地捕捉动态嘴唇运动中固有的细微差别。此外,我们集成了音频点增强模块,这不仅确保了音频信号与特征空间内相应的嘴唇点云的同步,而且有助于更深入地理解跨模态条件特征之间的相互关系。大量的实验表明,我们的方法实现了优越的高保真度和音频唇同步说话头合成相比,以前的方法。
摘要:Talking head synthesis with arbitrary speech audio is a crucial challenge inthe field of digital humans. Recently, methods based on radiance fields havereceived increasing attention due to their ability to synthesize high-fidelityand identity-consistent talking heads from just a few minutes of trainingvideo. However, due to the limited scale of the training data, these methodsoften exhibit poor performance in audio-lip synchronization and visual quality.In this paper, we propose a novel 3D Gaussian-based method called PointTalk,which constructs a static 3D Gaussian field of the head and deforms it in syncwith the audio. It also incorporates an audio-driven dynamic lip point cloud asa critical component of the conditional information, thereby facilitating theeffective synthesis of talking heads. Specifically, the initial step involvesgenerating the corresponding lip point cloud from the audio signal andcapturing its topological structure. The design of the dynamic differenceencoder aims to capture the subtle nuances inherent in dynamic lip movementsmore effectively. Furthermore, we integrate the audio-point enhancement module,which not only ensures the synchronization of the audio signal with thecorresponding lip point cloud within the feature space, but also facilitates adeeper understanding of the interrelations among cross-modal conditionalfeatures. Extensive experiments demonstrate that our method achieves superiorhigh-fidelity and audio-lip synchronization in talking head synthesis comparedto previous methods.
标题:Zero-Shot单到双耳语音合成
链接:https://arxiv.org/abs/2412.08356
摘要:我们提出了ZeroBAS,一种神经方法,可以从单声道音频记录和位置信息合成双耳音频,而无需对任何双耳数据进行训练。据我们所知,这是第一个发表的zero-shot神经方法单到双耳音频合成。具体来说,我们表明,一个无参数的几何时间扭曲和幅度缩放源位置的基础上,足以得到一个初始的双耳合成,可以通过迭代地应用一个预先训练的去噪声码器细化。此外,我们发现这导致了整个房间条件的泛化,我们通过引入一个新的数据集TUT Mono-to-Binaural来测量,以评估在看不见的条件下最先进的单声道到双耳合成方法。我们的zero-shot方法在感知上与标准单声道到双耳数据集上的监督方法的性能相当,甚至在我们的分布外TUT单声道到双耳数据集上超过它们。我们的研究结果突出了预训练的生成音频模型和zero-shot学习解锁强大的双耳音频合成的潜力。
摘要:We present ZeroBAS, a neural method to synthesize binaural audio frommonaural audio recordings and positional information without training on anybinaural data. To our knowledge, this is the first published zero-shot neuralapproach to mono-to-binaural audio synthesis. Specifically, we show that aparameter-free geometric time warping and amplitude scaling based on sourcelocation suffices to get an initial binaural synthesis that can be refined byiteratively applying a pretrained denoising vocoder. Furthermore, we find thisleads to generalization across room conditions, which we measure by introducinga new dataset, TUT Mono-to-Binaural, to evaluate state-of-the-artmonaural-to-binaural synthesis methods on unseen conditions. Our zero-shotmethod is perceptually on-par with the performance of supervised methods on thestandard mono-to-binaural dataset, and even surpasses them on ourout-of-distribution TUT Mono-to-Binaural dataset. Our results highlight thepotential of pretrained generative audio models and zero-shot learning tounlock robust binaural audio synthesis.
标题:SyncViolist:基于鞠躬和手指的音乐导向小提琴动作生成
链接:https://arxiv.org/abs/2412.08343
备注:10 pages, 7 figures, 6 tables, WACV 2025
摘要:自动生成逼真的音乐表演动作可以极大地增强数字媒体制作,通常涉及专业人士和音乐家之间的合作。然而,捕捉精确的音乐表演所需的复杂的身体,手和手指运动是具有挑战性的。现有的方法通常由于音频和运动之间的复杂映射而不足,通常需要额外的输入,如分数或运动数据。在这项工作中,我们提出了SyncViolinist,一个多阶段的端到端的框架,生成同步的小提琴演奏运动完全从音频输入。我们的方法通过两个关键模块克服了捕获全局和细粒度性能特征的挑战:弓/指法模块和运动生成模块。弓法/指法模块从音频中提取详细的演奏信息,运动生成模块使用该信息来创建反映小提琴演奏的时间粒度和性质的精确、协调的身体运动。我们展示了SyncViolinist的有效性,从看不见的小提琴演奏音频中获得了显著改善的定性和定量结果,优于最先进的方法。涉及专业小提琴家的广泛的主观评价进一步验证了我们的方法。代码和数据集可在https://github.com/Kakanat/SyncViolinist上获得。
摘要:Automatically generating realistic musical performance motion can greatlyenhance digital media production, often involving collaboration betweenprofessionals and musicians. However, capturing the intricate body, hand, andfinger movements required for accurate musical performances is challenging.Existing methods often fall short due to the complex mapping between audio andmotion, typically requiring additional inputs like scores or MIDI data. In thiswork, we present SyncViolinist, a multi-stage end-to-end framework thatgenerates synchronized violin performance motion solely from audio input. Ourmethod overcomes the challenge of capturing both global and fine-grainedperformance features through two key modules: a bowing/fingering module and amotion generation module. The bowing/fingering module extracts detailed playinginformation from the audio, which the motion generation module uses to createprecise, coordinated body motions reflecting the temporal granularity andnature of the violin performance. We demonstrate the effectiveness ofSyncViolinist with significantly improved qualitative and quantitative resultsfrom unseen violin performance audio, outperforming state-of-the-art methods.Extensive subjective evaluations involving professional violinists furthervalidate our approach. The code and dataset are available athttps://github.com/Kakanat/SyncViolinist.
标题:使用自我监督学习和特征提取的语音和歌唱中语音和口音转换的统一模型
链接:https://arxiv.org/abs/2412.08312
备注:7 pages, 5 figures, 2 tables
摘要:本文提出了一种新的语音转换模型,能够转换说话和唱歌的声音。它解决了当前系统中的关键挑战,例如传达情感,管理发音和口音变化,以及再现非语言声音。该模型的突出功能之一是能够对包含语音和唱歌的混合语音样本执行口音转换,从而在保留原始内容和韵律的同时改变说话者的口音。该模型使用编码器-解码器架构:编码器基于HuBERT来处理语音的声学和语言内容,而HiFi-GAN解码器音频匹配目标说话者的声音。该模型结合了基频(f0)特征和歌手嵌入,以提高性能,同时确保在转换过程中保持音高和音调的准确性和声音的身份。这种方法提高了语音风格转换的自然性和灵活性,在语音配音、内容创建以及文本到语音(TTS)和交互式语音响应(IVR)系统等技术中显示出强大的应用潜力。
摘要:This paper presents a new voice conversion model capable of transforming bothspeaking and singing voices. It addresses key challenges in current systems,such as conveying emotions, managing pronunciation and accent changes, andreproducing non-verbal sounds. One of the model's standout features is itsability to perform accent conversion on hybrid voice samples that encompassboth speech and singing, allowing it to change the speaker's accent whilepreserving the original content and prosody. The proposed model uses anencoder-decoder architecture: the encoder is based on HuBERT to process thespeech's acoustic and linguistic content, while the HiFi-GAN decoder audiomatches the target speaker's voice. The model incorporates fundamentalfrequency (f0) features and singer embeddings to enhance performance whileensuring the pitch & tone accuracy and vocal identity are preserved duringtransformation. This approach improves how naturally and flexibly voice stylecan be transformed, showing strong potential for applications in voice dubbing,content creation, and technologies like Text-to-Speech (TTS) and InteractiveVoice Response (IVR) systems.
标题:MoMuSE:视觉线索受损的实时场景的动量多模式目标说话人提取
链接:https://arxiv.org/abs/2412.08247
摘要:视听目标说话者提取(AV-TSE)旨在使用时间同步的视觉线索从音频混合中分离出特定目标说话者的语音。在现实世界的场景中,由于各种损伤,视觉线索并不总是可用的,这破坏了AV-TSE的稳定性。尽管存在这一挑战,但人类可以随着时间的推移保持注意力的势头,即使目标说话者不可见。在本文中,我们介绍了动量多模态目标说话人提取(MoMuSE),它保留在内存中的说话人身份动量,使模型能够连续跟踪目标说话人。MoMuSE专为实时推理而设计,它在视觉提示和动态更新的说话者动量的指导下提取当前的语音窗口。实验结果表明,MoMuSE表现出显着的改善,特别是在视觉线索严重受损的情况下。
摘要:Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech ofa specific target speaker from an audio mixture using time-synchronized visualcues. In real-world scenarios, visual cues are not always available due tovarious impairments, which undermines the stability of AV-TSE. Despite thischallenge, humans can maintain attentional momentum over time, even when thetarget speaker is not visible. In this paper, we introduce the MomentumMulti-modal target Speaker Extraction (MoMuSE), which retains a speakeridentity momentum in memory, enabling the model to continuously track thetarget speaker. Designed for real-time inference, MoMuSE extracts the currentspeech window with guidance from both visual cues and dynamically updatedspeaker momentum. Experimental results demonstrate that MoMuSE exhibitssignificant improvement, particularly in scenarios with severe impairment ofvisual cues.
标题:TouchTTC:一个令人尴尬的简单的TTC框架,每个人都可以触摸
链接:https://arxiv.org/abs/2412.08237
备注:Technical Report
摘要:众所周知,基于LLM的系统是数据饥渴的。最近基于LLM的TTS作品通常采用复杂的数据处理管道来获得高质量的训练数据。这些复杂的流水线在每个阶段都需要优秀的模型(例如,语音去噪、语音增强、说话人日志化和标点符号模型),它们本身需要高质量的训练数据,并且很少开源。即使使用最先进的模型,问题仍然存在,例如不完全的背景噪声去除以及标点符号和实际语音停顿之间的不一致。此外,严格的过滤策略通常仅保留原始数据的10- 30%,这显著阻碍了数据缩放工作。在这项工作中,我们利用噪声鲁棒的音频tokenizer(S3 Tokenizer)设计一个简化而有效的TTS数据处理管道,保持数据质量,同时大大降低数据采集成本,实现超过50%的数据留存率。除了数据扩展挑战之外,与传统方法相比,基于LLM的TTS系统也会产生更高的部署成本。当前的系统通常仅使用LLM用于文本到令牌生成,而需要单独的模型(例如,流匹配模型),其不能由LLM推理引擎直接执行,进一步使部署复杂化。为了应对这些挑战,我们消除了LLM和流组件中的冗余模块,用LLM架构取代了流模型主干。在此基础上,简化的流骨干,我们提出了一个统一的体系结构,流和非流推理,显着降低部署成本。最后,我们探索了使用相同的数据进行训练来统一TTS和ASR任务的可行性,这要归功于简化的管道和S3 Tokenizer,它降低了对TTS训练数据的质量要求。
摘要:It is well known that LLM-based systems are data-hungry. Recent LLM-based TTSworks typically employ complex data processing pipelines to obtain high-qualitytraining data. These sophisticated pipelines require excellent models at eachstage (e.g., speech denoising, speech enhancement, speaker diarization, andpunctuation models), which themselves demand high-quality training data and arerarely open-sourced. Even with state-of-the-art models, issues persist, such asincomplete background noise removal and misalignment between punctuation andactual speech pauses. Moreover, the stringent filtering strategies often retainonly 10-30\% of the original data, significantly impeding data scaling efforts.In this work, we leverage a noise-robust audio tokenizer (S3Tokenizer) todesign a simplified yet effective TTS data processing pipeline that maintainsdata quality while substantially reducing data acquisition costs, achieving adata retention rate of over 50\%. Beyond data scaling challenges, LLM-based TTSsystems also incur higher deployment costs compared to conventional approaches.Current systems typically use LLMs solely for text-to-token generation, whilerequiring separate models (e.g., flow matching models) for token-to-waveformgeneration, which cannot be directly executed by LLM inference engines, furthercomplicating deployment. To address these challenges, we eliminate redundantmodules in both LLM and flow components, replacing the flow model backbone withan LLM architecture. Building upon this simplified flow backbone, we propose aunified architecture for both streaming and non-streaming inference,significantly reducing deployment costs. Finally, we explore the feasibility ofunifying TTS and ASR tasks using the same data for training, thanks to thesimplified pipeline and the S3Tokenizer that reduces the quality requirementsfor TTS training data.
标题:视听分割中时间失调的协作混合运算器
链接:https://arxiv.org/abs/2412.08161
摘要:视听视频分割(AVVS)旨在生成与相应音频准确对齐的声音产生对象的像素级映射。然而,现有的方法往往面临时间错位,其中音频线索和分割结果在时间上不协调。音频提供了两个关键信息:i)目标对象级别的细节,以及ii)对象开始和停止产生声音的时间。目前的方法更多地关注对象级信息,但忽略了音频语义变化的边界,导致时间错位。为了解决这个问题,我们提出了一个协作混合调度框架~(Co-Prop)。该框架包括两个主要步骤:初步音频边界划分和逐帧音频插入传播。为了锚定音频边界,我们使用Qwen大型语言模型的检索辅助提示来识别音频语义变化的控制点。这些控制点将音频分割成语义一致的音频部分。在获得控制点列表后,我们提出了音频插入解码器,使用逐帧音频插入传播和匹配方法来处理每个音频部分。我们策划了一个包含不同来源转换案例的紧凑数据集,并设计了一个评估对齐率的指标。与传统的同步处理方法相比,我们的方法减少了内存需求,有利于帧对齐。实验结果表明,我们的方法在三个数据集和两个骨干的有效性。此外,我们的方法可以与现有的AVVS方法集成,提供即插即用功能,以提高其性能。
摘要:Audio-visual video segmentation (AVVS) aims to generate pixel-level maps ofsound-producing objects that accurately align with the corresponding audio.However, existing methods often face temporal misalignment, where audio cuesand segmentation results are not temporally coordinated. Audio provides twocritical pieces of information: i) target object-level details and ii) thetiming of when objects start and stop producing sounds. Current methods focusmore on object-level information but neglect the boundaries of audio semanticchanges, leading to temporal misalignment. To address this issue, we propose aCollaborative Hybrid Propagator Framework~(Co-Prop). This framework includestwo main steps: Preliminary Audio Boundary Anchoring and Frame-by-FrameAudio-Insert Propagation. To Anchor the audio boundary, we employretrieval-assist prompts with Qwen large language models to identify controlpoints of audio semantic changes. These control points split the audio intosemantically consistent audio portions. After obtaining the control pointlists, we propose the Audio Insertion Propagator to process each audio portionusing a frame-by-frame audio insertion propagation and matching approach. Wecurated a compact dataset comprising diverse source conversion cases anddevised a metric to assess alignment rates. Compared to traditionalsimultaneous processing methods, our approach reduces memory requirements andfacilitates frame alignment. Experimental results demonstrate the effectivenessof our approach across three datasets and two backbones. Furthermore, ourmethod can be integrated with existing AVVS approaches, offering plug-and-playfunctionality to enhance their performance.
标题:LatentSpeech:文本到语音生成的潜在扩散
链接:https://arxiv.org/abs/2412.08117
摘要:基于扩散的生成式AI因其优于其他生成技术(如生成对抗网络和变分自动编码器)的性能而备受关注。虽然它在计算机视觉和自然语言处理等领域取得了显着的进步,但它们在语音生成中的应用仍然没有得到充分的探索。主流的文本到语音系统主要将输出映射到频谱空间中的Mel-Spectrograms,由于MelSpecs的稀疏性导致高计算负载。为了解决这些限制,我们提出了LatentSpeech,一种新的TTS生成方法,利用潜在的扩散模型。通过使用潜在嵌入作为中间表示,LatentSpeech将目标维度降低到MelSpecs所需的5%,简化了TTS编码器和声码器的处理,并实现了高效的高质量语音生成。该研究首次将潜在扩散模型集成到TTS中,提高了生成语音的准确性和自然度。在基准数据集上的实验结果表明,与现有模型相比,LatentSpeech在单词错误率方面实现了25%的改善,在梅尔倒谱系数失真方面实现了24%的改善,在额外的训练数据下,进一步的改善分别上升到49.5%和26%。这些发现突出了潜在的语音推进国家的最先进的TTS技术
摘要:Diffusion-based Generative AI gains significant attention for its superiorperformance over other generative techniques like Generative AdversarialNetworks and Variational Autoencoders. While it has achieved notableadvancements in fields such as computer vision and natural language processing,their application in speech generation remains under-explored. MainstreamText-to-Speech systems primarily map outputs to Mel-Spectrograms in thespectral space, leading to high computational loads due to the sparsity ofMelSpecs. To address these limitations, we propose LatentSpeech, a novel TTSgeneration approach utilizing latent diffusion models. By using latentembeddings as the intermediate representation, LatentSpeech reduces the targetdimension to 5% of what is required for MelSpecs, simplifying the processingfor the TTS encoder and vocoder and enabling efficient high-quality speechgeneration. This study marks the first integration of latent diffusion modelsin TTS, enhancing the accuracy and naturalness of generated speech.Experimental results on benchmark datasets demonstrate that LatentSpeechachieves a 25% improvement in Word Error Rate and a 24% improvement in MelCepstral Distortion compared to existing models, with further improvementsrising to 49.5% and 26%, respectively, with additional training data. Thesefindings highlight the potential of LatentSpeech to advance thestate-of-the-art in TTS technology
标题:对齐器引导的训练范式:通过对齐器引导持续时间推进文本到语音模型
链接:https://arxiv.org/abs/2412.08112
摘要:文本到语音(TTS)系统的最新进展,如FastSpeech和StyleSpeech,显着提高了语音生成质量。然而,这些模型通常依赖于蒙特利尔强制调整器等外部工具生成的持续时间,这可能非常耗时且缺乏灵活性。准确的音长的重要性往往被低估,尽管它们在实现自然韵律和可理解性方面发挥着至关重要的作用。为了解决这些局限性,我们提出了一种新的对齐器引导的训练范式,通过在TTS模型之前训练对齐器来优先考虑准确的持续时间标签。这种方法减少了对外部工具的依赖,并提高了对准精度。我们进一步探讨了不同的声学特征,包括梅尔频谱图,MFCC和潜在的功能,TTS模型的性能的影响。我们的实验结果表明,对齐器引导的持续时间标记可以实现高达16%的字错误率的改善,并显着提高音素和音调对齐。这些发现突出了我们的方法在优化TTS系统更自然和更清晰的语音生成的有效性。
摘要:Recent advancements in text-to-speech (TTS) systems, such as FastSpeech andStyleSpeech, have significantly improved speech generation quality. However,these models often rely on duration generated by external tools like theMontreal Forced Aligner, which can be time-consuming and lack flexibility. Theimportance of accurate duration is often underestimated, despite their crucialrole in achieving natural prosody and intelligibility. To address theselimitations, we propose a novel Aligner-Guided Training Paradigm thatprioritizes accurate duration labelling by training an aligner before the TTSmodel. This approach reduces dependence on external tools and enhancesalignment accuracy. We further explore the impact of different acousticfeatures, including Mel-Spectrograms, MFCCs, and latent features, on TTS modelperformance. Our experimental results show that aligner-guided durationlabelling can achieve up to a 16\% improvement in word error rate andsignificantly enhance phoneme and tone alignment. These findings highlight theeffectiveness of our approach in optimizing TTS systems for more natural andintelligible speech generation.
标题:弗雷切特音乐距离:生成符号音乐评估的指标
链接:https://arxiv.org/abs/2412.07948
摘要:在本文中,我们介绍了Frechet音乐距离(FMD),一种新的生成符号音乐模型的评估指标,灵感来自计算机视觉中的Frechet初始距离(FID)和生成音频中的Frechet音频距离(FAD)。FMD计算参考和生成的符号音乐嵌入分布之间的距离,捕获抽象的音乐特征。我们在多个数据集和模型上验证FMD。结果表明,FMD有效地区分模型质量,提供了一个特定领域的指标,用于评估符号音乐生成,并建立了一个可重复的标准,为未来的研究在符号音乐建模。
摘要:In this paper we introduce the Frechet Music Distance (FMD), a novelevaluation metric for generative symbolic music models, inspired by the FrechetInception Distance (FID) in computer vision and Frechet Audio Distance (FAD) ingenerative audio. FMD calculates the distance between distributions ofreference and generated symbolic music embeddings, capturing abstract musicalfeatures. We validate FMD across several datasets and models. Results indicatethat FMD effectively differentiates model quality, providing a domain-specificmetric for evaluating symbolic music generation, and establishing areproducible standard for future research in symbolic music modeling.
标题:评估区分性和生成性E2 E语音增强模型对音节重读保留的影响
链接:https://arxiv.org/abs/2412.08306
摘要:音节重音自动检测是计算机辅助语言学习系统中的一个重要组成部分。当前的压力检测模型通常是在干净的语音上训练的,这在背景噪声普遍存在的现实世界场景中可能不鲁棒。为了解决这个问题,可以采用旨在通过去除噪声来增强语音的语音增强(SE)模型,但它们对保留音节重音模式的影响还没有得到很好的研究。本研究探讨了不同的SE模型,代表歧视性和生成建模方法,影响音节重音检测噪声条件下。我们评估这些模型,将它们应用到语音数据与不同的信噪比(SNR)从0到20 dB,并评估其有效性,在保持压力模式。此外,我们探索不同的特征集,以确定哪些是最有效的捕捉噪声中的压力模式。为了进一步了解SE模型的影响,进行了一项基于人类的感知研究,以比较SE增强语音中的感知重音模式与干净语音中的感知重音模式,从而深入了解这些模型如何保持音节重音。实验进行英语语音数据从非母语的德语和意大利语。结果表明,当使用启发式特征时,生成式SE模型的应力检测性能是鲁棒的。此外,知觉研究的观察结果与所有SE模型下的压力检测结果一致。
摘要:Automatic syllable stress detection is a crucial component inComputer-Assisted Language Learning (CALL) systems for language learners.Current stress detection models are typically trained on clean speech, whichmay not be robust in real-world scenarios where background noise is prevalent.To address this, speech enhancement (SE) models, designed to enhance speech byremoving noise, might be employed, but their impact on preserving syllablestress patterns is not well studied. This study examines how different SEmodels, representing discriminative and generative modeling approaches, affectsyllable stress detection under noisy conditions. We assess these models byapplying them to speech data with varying signal-to-noise ratios (SNRs) from 0to 20 dB, and evaluating their effectiveness in maintaining stress patterns.Additionally, we explore different feature sets to determine which ones aremost effective for capturing stress patterns amidst noise. To furtherunderstand the impact of SE models, a human-based perceptual study is conductedto compare the perceived stress patterns in SE-enhanced speech with those inclean speech, providing insights into how well these models preserve syllablestress as perceived by listeners. Experiments are performed on English speechdata from non-native speakers of German and Italian. And the results revealthat the stress detection performance is robust with the generative SE modelswhen heuristic features are used. Also, the observations from the perceptualstudy are consistent with the stress detection outcomes under all SE models.
标题:评估区分性和生成性E2 E语音增强模型对音节重读保留的影响
链接:https://arxiv.org/abs/2412.08306
摘要:音节重音自动检测是计算机辅助语言学习系统中的一个重要组成部分。当前的压力检测模型通常是在干净的语音上训练的,这在背景噪声普遍存在的现实世界场景中可能不鲁棒。为了解决这个问题,可以采用旨在通过去除噪声来增强语音的语音增强(SE)模型,但它们对保留音节重音模式的影响还没有得到很好的研究。本研究探讨了不同的SE模型,代表歧视性和生成建模方法,影响音节重音检测噪声条件下。我们评估这些模型,将它们应用到语音数据与不同的信噪比(SNR)从0到20 dB,并评估其有效性,在保持压力模式。此外,我们探索不同的特征集,以确定哪些是最有效的捕捉噪声中的压力模式。为了进一步了解SE模型的影响,进行了一项基于人类的感知研究,将SE增强语音中的感知重读模式与干净语音中的感知重读模式进行比较,以深入了解这些模型如何保留听众感知的音节重读。实验进行英语语音数据从非母语的德语和意大利语。结果表明,当使用启发式特征时,生成式SE模型的应力检测性能是鲁棒的。此外,知觉研究的观察结果与所有SE模型下的压力检测结果一致。
摘要:Automatic syllable stress detection is a crucial component inComputer-Assisted Language Learning (CALL) systems for language learners.Current stress detection models are typically trained on clean speech, whichmay not be robust in real-world scenarios where background noise is prevalent.To address this, speech enhancement (SE) models, designed to enhance speech byremoving noise, might be employed, but their impact on preserving syllablestress patterns is not well studied. This study examines how different SEmodels, representing discriminative and generative modeling approaches, affectsyllable stress detection under noisy conditions. We assess these models byapplying them to speech data with varying signal-to-noise ratios (SNRs) from 0to 20 dB, and evaluating their effectiveness in maintaining stress patterns.Additionally, we explore different feature sets to determine which ones aremost effective for capturing stress patterns amidst noise. To furtherunderstand the impact of SE models, a human-based perceptual study is conductedto compare the perceived stress patterns in SE-enhanced speech with those inclean speech, providing insights into how well these models preserve syllablestress as perceived by listeners. Experiments are performed on English speechdata from non-native speakers of German and Italian. And the results revealthat the stress detection performance is robust with the generative SE modelswhen heuristic features are used. Also, the observations from the perceptualstudy are consistent with the stress detection outcomes under all SE models.
标题:AdvWave:针对大型音频语言模型的隐形对抗越狱攻击
链接:https://arxiv.org/abs/2412.08608
摘要:大型音频语言模型(LALM)的最新进展已经实现了基于语音的用户交互,显著增强了用户体验并加速了LALM在现实世界应用中的部署。然而,确保LALM的安全性对于防止可能引起社会关注或违反人工智能法规的风险输出至关重要。尽管这个问题很重要,但对越狱LALM的研究仍然有限,因为它们最近才出现,而且与对基于DNN的音频模型的攻击相比,它们带来了额外的技术挑战。具体来说,LALM中的音频编码器涉及离散化操作,通常会导致梯度破碎,从而阻碍了依赖于基于梯度的优化的攻击的有效性。LALM的行为可变性进一步使有效(对抗性)优化目标的识别复杂化。此外,对对抗性音频波形实施隐形约束引入了减小的非凸可行解空间,进一步加剧了优化过程的挑战。为了克服这些挑战,我们开发了AdvWave,这是第一个针对LALM的越狱框架。我们提出了一种双阶段优化方法,解决梯度破碎,实现有效的端到端的基于梯度的优化。此外,我们开发了一种自适应对抗性目标搜索算法,该算法根据LALM对特定查询的响应模式动态调整对抗性优化目标。为了确保对抗性音频对人类听众保持感知自然,我们设计了一种分类器引导的优化方法,该方法生成类似于常见城市声音的对抗性噪声。对多个高级LALM的广泛评估表明,AdvWave优于基准方法,平均越狱攻击成功率提高了40%。
摘要:Recent advancements in large audio-language models (LALMs) have enabledspeech-based user interactions, significantly enhancing user experience andaccelerating the deployment of LALMs in real-world applications. However,ensuring the safety of LALMs is crucial to prevent risky outputs that may raisesocietal concerns or violate AI regulations. Despite the importance of thisissue, research on jailbreaking LALMs remains limited due to their recentemergence and the additional technical challenges they present compared toattacks on DNN-based audio models. Specifically, the audio encoders in LALMs,which involve discretization operations, often lead to gradient shattering,hindering the effectiveness of attacks relying on gradient-based optimizations.The behavioral variability of LALMs further complicates the identification ofeffective (adversarial) optimization targets. Moreover, enforcing stealthinessconstraints on adversarial audio waveforms introduces a reduced, non-convexfeasible solution space, further intensifying the challenges of theoptimization process. To overcome these challenges, we develop AdvWave, thefirst jailbreak framework against LALMs. We propose a dual-phase optimizationmethod that addresses gradient shattering, enabling effective end-to-endgradient-based optimization. Additionally, we develop an adaptive adversarialtarget search algorithm that dynamically adjusts the adversarial optimizationtarget based on the response patterns of LALMs for specific queries. To ensurethat adversarial audio remains perceptually natural to human listeners, wedesign a classifier-guided optimization approach that generates adversarialnoise resembling common urban sounds. Extensive evaluations on multipleadvanced LALMs demonstrate that AdvWave outperforms baseline methods, achievinga 40% higher average jailbreak attack success rate.
标题:Mel-Refine:一种即插即用方法在音频生成中完善Mel-Spectrum
链接:https://arxiv.org/abs/2412.08577
摘要:文本到音频(TTA)模型能够从文本提示生成各种音频。然而,大多数主流TTA模型主要依赖于Mel频谱图,在产生具有丰富内容的音频方面仍然面临挑战。梅尔声谱图对这种音频所需的复杂细节和纹理往往超过了模型的能力,导致输出模糊或缺乏连贯性。在本文中,我们首先调查的关键作用,U-网络在梅尔频谱生成。我们的分析表明,在U-Net结构中,跳跃连接和主干中的高频分量影响纹理和细节,而主干中的低频分量对扩散去噪过程至关重要。我们进一步提出了“梅尔精炼”,即插即用的方法,增强梅尔频谱图纹理和细节,通过调整不同的分量权重在推理。我们的方法不需要额外的训练或微调,并且与任何基于扩散的TTA架构完全兼容。实验结果表明,我们的方法提高了最新的TTA模型Tango 2的性能指标的25%,证明了其有效性。
摘要:Text-to-audio (TTA) model is capable of generating diverse audio from textualprompts. However, most mainstream TTA models, which predominantly rely onMel-spectrograms, still face challenges in producing audio with rich content.The intricate details and texture required in Mel-spectrograms for such audiooften surpass the models' capacity, leading to outputs that are blurred or lackcoherence. In this paper, we begin by investigating the critical role of U-Netin Mel-spectrogram generation. Our analysis shows that in U-Net structure,high-frequency components in skip-connections and the backbone influencetexture and detail, while low-frequency components in the backbone are criticalfor the diffusion denoising process. We further propose ``Mel-Refine'', aplug-and-play approach that enhances Mel-spectrogram texture and detail byadjusting different component weights during inference. Our method requires noadditional training or fine-tuning and is fully compatible with anydiffusion-based TTA architecture. Experimental results show that our approachboosts performance metrics of the latest TTA model Tango2 by 25\%,demonstrating its effectiveness.
标题:Sketch 2Sound:通过时变信号和索尼克模仿实现可控音频生成
链接:https://arxiv.org/abs/2412.08550
摘要:我们提出了Sketch2Sound,这是一个生成音频模型,能够从一组可解释的时变控制信号创建高质量的声音:响度,亮度和音高,以及文本提示。Sketch2Sound可以从声音模仿中合成任意声音(即,声音模仿或参考声音形状)。Sketch2Sound可以在任何文本到音频的潜在扩散Transformer(DiT)之上实现,并且每个控件只需要40k步的微调和单个线性层,使其比ControlNet等现有方法更轻量级。为了从草图般的声音模仿合成,我们建议在训练期间将随机中值滤波器应用于控制信号,允许Sketch2Sound使用具有灵活水平的时间特异性的控制来提示。我们表明,Sketch2Sound可以合成的声音,按照输入控制的要点,从一个声乐模仿,同时保留遵守输入文本提示和音频质量相比,一个纯文本的基线。Sketch2Sound允许声音艺术家使用文本提示的语义灵活性和声音手势或声音模仿的表现力和精确度来创建声音。在https://hugofloresgarcia.art/sketch2sound/上可以找到正确的例子。
摘要:We present Sketch2Sound, a generative audio model capable of creatinghigh-quality sounds from a set of interpretable time-varying control signals:loudness, brightness, and pitch, as well as text prompts. Sketch2Sound cansynthesize arbitrary sounds from sonic imitations (i.e.,~a vocal imitation or areference sound-shape). Sketch2Sound can be implemented on top of anytext-to-audio latent diffusion transformer (DiT), and requires only 40k stepsof fine-tuning and a single linear layer per control, making it morelightweight than existing methods like ControlNet. To synthesize fromsketchlike sonic imitations, we propose applying random median filters to thecontrol signals during training, allowing Sketch2Sound to be prompted usingcontrols with flexible levels of temporal specificity. We show thatSketch2Sound can synthesize sounds that follow the gist of input controls froma vocal imitation while retaining the adherence to an input text prompt andaudio quality compared to a text-only baseline. Sketch2Sound allows soundartists to create sounds with the semantic flexibility of text prompts and theexpressivity and precision of a sonic gesture or vocal imitation. Soundexamples are available at https://hugofloresgarcia.art/sketch2sound/.
标题:对音乐生成模型的训练数据进行水印
链接:https://arxiv.org/abs/2412.08549
摘要:生成式人工智能(Gen-AI)模型越来越多地用于跨领域生成内容,包括文本、图像和音频。虽然这些模型代表了一项重大的技术突破,但它们的生成能力是通过对大量人类生成的内容进行训练而获得的,这些内容通常包括受版权保护的材料。在这项工作中,我们调查音频水印技术是否可以用来检测未经授权的使用内容训练的音乐生成模型。我们将在水印数据上训练的模型生成的输出与在非水印数据上训练的模型生成的输出进行比较。我们研究影响模型的生成行为的因素:水印技术,在训练集中的水印样本的比例,以及水印技术对模型的标记器的鲁棒性。我们的研究结果表明,音频水印技术,包括一些人类无法察觉的,可以导致模型的输出明显的变化。我们还研究了一个国家的最先进的水印技术的鲁棒性去除技术。
摘要:Generative Artificial Intelligence (Gen-AI) models are increasingly used toproduce content across domains, including text, images, and audio. While thesemodels represent a major technical breakthrough, they gain their generativecapabilities from being trained on enormous amounts of human-generated content,which often includes copyrighted material. In this work, we investigate whetheraudio watermarking techniques can be used to detect an unauthorized usage ofcontent to train a music generation model. We compare outputs generated by amodel trained on watermarked data to a model trained on non-watermarked data.We study factors that impact the model's generation behaviour: the watermarkingtechnique, the proportion of watermarked samples in the training set, and therobustness of the watermarking technique against the model's tokenizer. Ourresults show that audio watermarking techniques, including some that areimperceptible to humans, can lead to noticeable shifts in the model's outputs.We also study the robustness of a state-of-the-art watermarking technique toremoval techniques.
标题:PointTalk:音频驱动的动态唇点云,用于3D基于高斯的会说话的头部合成
链接:https://arxiv.org/abs/2412.08504
备注:9 pages, accepted by AAAI 2025
摘要:具有任意语音音频的说话头部合成是数字人领域的一个关键挑战。最近,基于辐射场的方法受到越来越多的关注,因为它们能够从几分钟的训练视频合成高保真和身份一致的说话头。然而,由于训练数据的规模有限,这些方法往往表现出较差的性能在音频唇同步和视觉质量。在本文中,我们提出了一种新的基于3D高斯的方法称为PointTalk,它构建了一个静态的头部3D高斯场,并使其与音频同步变形。它还集成了音频驱动的动态唇点云作为条件信息的关键组成部分,从而促进了说话头部的有效合成。具体地,初始步骤涉及从音频信号生成对应的嘴唇点云并捕获其拓扑结构。动态差分编码器的设计旨在更有效地捕捉动态嘴唇运动中固有的细微差别。此外,我们集成了音频点增强模块,这不仅确保了音频信号与特征空间内相应的唇点云的同步,而且有助于更深入地理解跨模态条件特征之间的相互关系。大量的实验表明,我们的方法实现了优越的高保真度和音频唇同步说话头合成相比,以前的方法。
摘要:Talking head synthesis with arbitrary speech audio is a crucial challenge inthe field of digital humans. Recently, methods based on radiance fields havereceived increasing attention due to their ability to synthesize high-fidelityand identity-consistent talking heads from just a few minutes of trainingvideo. However, due to the limited scale of the training data, these methodsoften exhibit poor performance in audio-lip synchronization and visual quality.In this paper, we propose a novel 3D Gaussian-based method called PointTalk,which constructs a static 3D Gaussian field of the head and deforms it in syncwith the audio. It also incorporates an audio-driven dynamic lip point cloud asa critical component of the conditional information, thereby facilitating theeffective synthesis of talking heads. Specifically, the initial step involvesgenerating the corresponding lip point cloud from the audio signal andcapturing its topological structure. The design of the dynamic differenceencoder aims to capture the subtle nuances inherent in dynamic lip movementsmore effectively. Furthermore, we integrate the audio-point enhancement module,which not only ensures the synchronization of the audio signal with thecorresponding lip point cloud within the feature space, but also facilitates adeeper understanding of the interrelations among cross-modal conditionalfeatures. Extensive experiments demonstrate that our method achieves superiorhigh-fidelity and audio-lip synchronization in talking head synthesis comparedto previous methods.
标题:Zero-Shot单到双耳语音合成
链接:https://arxiv.org/abs/2412.08356
摘要:我们提出了ZeroBAS,一种神经方法,可以从单声道音频记录和位置信息合成双耳音频,而无需对任何双耳数据进行训练。据我们所知,这是第一个发表的zero-shot神经方法单到双耳音频合成。具体来说,我们表明,一个无参数的几何时间扭曲和幅度缩放源位置的基础上,足以得到一个初始的双耳合成,可以通过迭代地应用一个预先训练的去噪声码器细化。此外,我们发现这导致了整个房间条件的泛化,我们通过引入一个新的数据集TUT Mono-to-Binaural来测量,以评估在看不见的条件下最先进的单声道到双耳合成方法。我们的zero-shot方法在感知上与标准单声道到双耳数据集上的监督方法的性能相当,甚至在我们的分布外TUT单声道到双耳数据集上超过它们。我们的研究结果突出了预训练的生成音频模型和zero-shot学习解锁强大的双耳音频合成的潜力。
摘要:We present ZeroBAS, a neural method to synthesize binaural audio frommonaural audio recordings and positional information without training on anybinaural data. To our knowledge, this is the first published zero-shot neuralapproach to mono-to-binaural audio synthesis. Specifically, we show that aparameter-free geometric time warping and amplitude scaling based on sourcelocation suffices to get an initial binaural synthesis that can be refined byiteratively applying a pretrained denoising vocoder. Furthermore, we find thisleads to generalization across room conditions, which we measure by introducinga new dataset, TUT Mono-to-Binaural, to evaluate state-of-the-artmonaural-to-binaural synthesis methods on unseen conditions. Our zero-shotmethod is perceptually on-par with the performance of supervised methods on thestandard mono-to-binaural dataset, and even surpasses them on ourout-of-distribution TUT Mono-to-Binaural dataset. Our results highlight thepotential of pretrained generative audio models and zero-shot learning tounlock robust binaural audio synthesis.
标题:SyncViolist:基于鞠躬和手指的音乐导向小提琴动作生成
链接:https://arxiv.org/abs/2412.08343
备注:10 pages, 7 figures, 6 tables, WACV 2025
摘要:自动生成逼真的音乐表演动作可以极大地增强数字媒体制作,通常涉及专业人士和音乐家之间的合作。然而,捕捉精确的音乐表演所需的复杂的身体,手和手指运动是具有挑战性的。现有的方法通常由于音频和运动之间的复杂映射而不足,通常需要额外的输入,如分数或运动数据。在这项工作中,我们提出了SyncViolinist,一个多阶段的端到端的框架,生成同步的小提琴演奏运动完全从音频输入。我们的方法通过两个关键模块克服了捕获全局和细粒度性能特征的挑战:弓/指法模块和运动生成模块。弓法/指法模块从音频中提取详细的演奏信息,运动生成模块使用该信息来创建反映小提琴演奏的时间粒度和性质的精确、协调的身体运动。我们展示了SyncViolinist的有效性,从看不见的小提琴演奏音频中获得了显著改善的定性和定量结果,优于最先进的方法。涉及专业小提琴家的广泛的主观评价进一步验证了我们的方法。代码和数据集可在https://github.com/Kakanat/SyncViolinist上获得。
摘要:Automatically generating realistic musical performance motion can greatlyenhance digital media production, often involving collaboration betweenprofessionals and musicians. However, capturing the intricate body, hand, andfinger movements required for accurate musical performances is challenging.Existing methods often fall short due to the complex mapping between audio andmotion, typically requiring additional inputs like scores or MIDI data. In thiswork, we present SyncViolinist, a multi-stage end-to-end framework thatgenerates synchronized violin performance motion solely from audio input. Ourmethod overcomes the challenge of capturing both global and fine-grainedperformance features through two key modules: a bowing/fingering module and amotion generation module. The bowing/fingering module extracts detailed playinginformation from the audio, which the motion generation module uses to createprecise, coordinated body motions reflecting the temporal granularity andnature of the violin performance. We demonstrate the effectiveness ofSyncViolinist with significantly improved qualitative and quantitative resultsfrom unseen violin performance audio, outperforming state-of-the-art methods.Extensive subjective evaluations involving professional violinists furthervalidate our approach. The code and dataset are available athttps://github.com/Kakanat/SyncViolinist.
标题:使用自我监督学习和特征提取的语音和歌唱中语音和口音转换的统一模型
链接:https://arxiv.org/abs/2412.08312
备注:7 pages, 5 figures, 2 tables
摘要:本文提出了一种新的语音转换模型,能够转换说话和唱歌的声音。它解决了当前系统中的关键挑战,例如传达情感、管理发音和口音变化以及再现非语言声音。该模型的突出功能之一是能够对包含语音和唱歌的混合语音样本执行口音转换,从而在保留原始内容和韵律的同时改变说话者的口音。该模型使用编码器-解码器架构:编码器基于HuBERT来处理语音的声学和语言内容,而HiFi-GAN解码器音频匹配目标说话者的声音。该模型结合了基频(f0)特征和歌手嵌入,以提高性能,同时确保在转换过程中保持音高和音调的准确性和声音的身份。这种方法提高了语音风格转换的自然性和灵活性,在语音配音、内容创建以及文本到语音(TTS)和交互式语音响应(IVR)系统等技术中显示出强大的应用潜力。
摘要:This paper presents a new voice conversion model capable of transforming bothspeaking and singing voices. It addresses key challenges in current systems,such as conveying emotions, managing pronunciation and accent changes, andreproducing non-verbal sounds. One of the model's standout features is itsability to perform accent conversion on hybrid voice samples that encompassboth speech and singing, allowing it to change the speaker's accent whilepreserving the original content and prosody. The proposed model uses anencoder-decoder architecture: the encoder is based on HuBERT to process thespeech's acoustic and linguistic content, while the HiFi-GAN decoder audiomatches the target speaker's voice. The model incorporates fundamentalfrequency (f0) features and singer embeddings to enhance performance whileensuring the pitch & tone accuracy and vocal identity are preserved duringtransformation. This approach improves how naturally and flexibly voice stylecan be transformed, showing strong potential for applications in voice dubbing,content creation, and technologies like Text-to-Speech (TTS) and InteractiveVoice Response (IVR) systems.
标题:具有文本到语音韵律嵌入的非母语语音中自动单词和音节突出检测的初步分析
链接:https://arxiv.org/abs/2412.08283
摘要:在单词和音节水平上自动检测突出度对于构建计算机辅助语言学习系统至关重要。研究表明,由当前最先进的(SOTA)文本到语音(TTS)系统学习的韵律嵌入可以在合成语音中像在母语中一样自然地生成单词和音节级的突出。为了了解语音合成中韵律嵌入在非母语语境下用于显著性检测的有效性,本文对从母语和非母语语音中提取的韵律嵌入进行了比较分析,并考虑了与韵律相关的嵌入:时长、能量和音高。这些嵌入是在两种情况下提取的:1)仅考虑文本,2)语音和文本。对于第一种情况,嵌入直接从TTS推理模式中提取,而对于第二种情况,我们建议在训练模式下从TTS中提取。实验在本族语语料库Tatoeba和非本族语语料库ISLE上进行。为了进行实验,对两个语料库手动注释单词级别的显著位置。与基于启发式的特征和自监督Wav 2 Vec-2.0表示相比,TTS嵌入的单词和音节级显著性检测准确率的最高相对提高分别为13.7%和5.9%以及16.2%和6.9%。
摘要:Automatic detection of prominence at the word and syllable-levels is criticalfor building computer-assisted language learning systems. It has been shownthat prosody embeddings learned by the current state-of-the-art (SOTA)text-to-speech (TTS) systems could generate word- and syllable-level prominencein the synthesized speech as natural as in native speech. To understand theeffectiveness of prosody embeddings from TTS for prominence detection undernonnative context, a comparative analysis is conducted on the embeddingsextracted from native and non-native speech considering the prominence-relatedembeddings: duration, energy, and pitch from a SOTA TTS named FastSpeech2.These embeddings are extracted under two conditions considering: 1) only text,2) both speech and text. For the first condition, the embeddings are extracteddirectly from the TTS inference mode, whereas for the second condition, wepropose to extract from the TTS under training mode. Experiments are conductedon native speech corpus: Tatoeba, and non-native speech corpus: ISLE. Forexperimentation, word-level prominence locations are manually annotated forboth corpora. The highest relative improvement on word \& syllable-levelprominence detection accuracies with the TTS embeddings are found to be 13.7% &5.9% and 16.2% & 6.9% compared to those with the heuristic-based features andself-supervised Wav2Vec-2.0 representations, respectively.
标题:MoMuSE:视觉线索受损的实时场景的动量多模式目标说话人提取
链接:https://arxiv.org/abs/2412.08247
摘要:视听目标说话者提取(AV-TSE)旨在使用时间同步的视觉线索从音频混合中分离出特定目标说话者的语音。在现实世界的场景中,由于各种损伤,视觉线索并不总是可用的,这破坏了AV-TSE的稳定性。尽管存在这一挑战,但人类可以随着时间的推移保持注意力的势头,即使目标说话者不可见。在本文中,我们介绍了动量多模态目标说话人提取(MoMuSE),它保留在内存中的说话人身份动量,使模型能够连续跟踪目标说话人。MoMuSE专为实时推理而设计,它在视觉提示和动态更新的说话者动量的指导下提取当前的语音窗口。实验结果表明,MoMuSE表现出显着的改善,特别是在视觉线索严重受损的情况下。
摘要:Audio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech ofa specific target speaker from an audio mixture using time-synchronized visualcues. In real-world scenarios, visual cues are not always available due tovarious impairments, which undermines the stability of AV-TSE. Despite thischallenge, humans can maintain attentional momentum over time, even when thetarget speaker is not visible. In this paper, we introduce the MomentumMulti-modal target Speaker Extraction (MoMuSE), which retains a speakeridentity momentum in memory, enabling the model to continuously track thetarget speaker. Designed for real-time inference, MoMuSE extracts the currentspeech window with guidance from both visual cues and dynamically updatedspeaker momentum. Experimental results demonstrate that MoMuSE exhibitssignificant improvement, particularly in scenarios with severe impairment ofvisual cues.
标题:TouchTTC:一个令人尴尬的简单的TTC框架,每个人都可以触摸
链接:https://arxiv.org/abs/2412.08237
备注:Technical Report
摘要:众所周知,基于LLM的系统是数据饥渴的。最近基于LLM的TTS作品通常采用复杂的数据处理管道来获得高质量的训练数据。这些复杂的流水线在每个阶段都需要优秀的模型(例如,语音去噪、语音增强、说话人日志化和标点符号模型),它们本身需要高质量的训练数据,并且很少开源。即使使用最先进的模型,问题仍然存在,例如不完全的背景噪声去除以及标点符号和实际语音停顿之间的不一致。此外,严格的过滤策略通常仅保留原始数据的10- 30%,这显著阻碍了数据缩放工作。在这项工作中,我们利用噪声鲁棒的音频tokenizer(S3 Tokenizer)设计一个简化而有效的TTS数据处理管道,保持数据质量,同时大大降低数据采集成本,实现超过50%的数据留存率。除了数据扩展挑战之外,与传统方法相比,基于LLM的TTS系统也会产生更高的部署成本。当前的系统通常仅使用LLM用于文本到令牌生成,而需要单独的模型(例如,流匹配模型),其不能由LLM推理引擎直接执行,进一步使部署复杂化。为了应对这些挑战,我们消除了LLM和流组件中的冗余模块,用LLM架构取代了流模型主干。在此基础上,简化的流骨干,我们提出了一个统一的体系结构,流和非流推理,显着降低部署成本。最后,我们探索了使用相同的数据进行训练来统一TTS和ASR任务的可行性,这要归功于简化的管道和S3 Tokenizer,它降低了对TTS训练数据的质量要求。
摘要:It is well known that LLM-based systems are data-hungry. Recent LLM-based TTSworks typically employ complex data processing pipelines to obtain high-qualitytraining data. These sophisticated pipelines require excellent models at eachstage (e.g., speech denoising, speech enhancement, speaker diarization, andpunctuation models), which themselves demand high-quality training data and arerarely open-sourced. Even with state-of-the-art models, issues persist, such asincomplete background noise removal and misalignment between punctuation andactual speech pauses. Moreover, the stringent filtering strategies often retainonly 10-30\% of the original data, significantly impeding data scaling efforts.In this work, we leverage a noise-robust audio tokenizer (S3Tokenizer) todesign a simplified yet effective TTS data processing pipeline that maintainsdata quality while substantially reducing data acquisition costs, achieving adata retention rate of over 50\%. Beyond data scaling challenges, LLM-based TTSsystems also incur higher deployment costs compared to conventional approaches.Current systems typically use LLMs solely for text-to-token generation, whilerequiring separate models (e.g., flow matching models) for token-to-waveformgeneration, which cannot be directly executed by LLM inference engines, furthercomplicating deployment. To address these challenges, we eliminate redundantmodules in both LLM and flow components, replacing the flow model backbone withan LLM architecture. Building upon this simplified flow backbone, we propose aunified architecture for both streaming and non-streaming inference,significantly reducing deployment costs. Finally, we explore the feasibility ofunifying TTS and ASR tasks using the same data for training, thanks to thesimplified pipeline and the S3Tokenizer that reduces the quality requirementsfor TTS training data.
标题:视听分割中时间失调的协作混合运算器
链接:https://arxiv.org/abs/2412.08161
摘要:视听视频分割(AVVS)旨在生成与相应音频准确对齐的声音产生对象的像素级映射。然而,现有的方法往往面临时间错位,其中音频线索和分割结果在时间上不协调。音频提供了两个关键信息:i)目标对象级别的细节,以及ii)对象开始和停止产生声音的时间。目前的方法更多地关注对象级信息,但忽略了音频语义变化的边界,导致时间错位。为了解决这个问题,我们提出了一个协作混合调度框架~(Co-Prop)。该框架包括两个主要步骤:初步音频边界划分和逐帧音频插入传播。为了锚定音频边界,我们使用Qwen大型语言模型的检索辅助提示来识别音频语义变化的控制点。这些控制点将音频分成语义一致的音频部分。在获得控制点列表后,我们提出了音频插入解码器,使用逐帧音频插入传播和匹配方法来处理每个音频部分。我们策划了一个包含不同来源转换案例的紧凑数据集,并设计了一个评估对齐率的指标。与传统的同步处理方法相比,我们的方法减少了内存需求,有利于帧对齐。实验结果表明,我们的方法在三个数据集和两个骨干的有效性。此外,我们的方法可以与现有的AVVS方法集成,提供即插即用功能,以提高其性能。
摘要:Audio-visual video segmentation (AVVS) aims to generate pixel-level maps ofsound-producing objects that accurately align with the corresponding audio.However, existing methods often face temporal misalignment, where audio cuesand segmentation results are not temporally coordinated. Audio provides twocritical pieces of information: i) target object-level details and ii) thetiming of when objects start and stop producing sounds. Current methods focusmore on object-level information but neglect the boundaries of audio semanticchanges, leading to temporal misalignment. To address this issue, we propose aCollaborative Hybrid Propagator Framework~(Co-Prop). This framework includestwo main steps: Preliminary Audio Boundary Anchoring and Frame-by-FrameAudio-Insert Propagation. To Anchor the audio boundary, we employretrieval-assist prompts with Qwen large language models to identify controlpoints of audio semantic changes. These control points split the audio intosemantically consistent audio portions. After obtaining the control pointlists, we propose the Audio Insertion Propagator to process each audio portionusing a frame-by-frame audio insertion propagation and matching approach. Wecurated a compact dataset comprising diverse source conversion cases anddevised a metric to assess alignment rates. Compared to traditionalsimultaneous processing methods, our approach reduces memory requirements andfacilitates frame alignment. Experimental results demonstrate the effectivenessof our approach across three datasets and two backbones. Furthermore, ourmethod can be integrated with existing AVVS approaches, offering plug-and-playfunctionality to enhance their performance.
标题:LatentSpeech:文本到语音生成的潜在扩散
链接:https://arxiv.org/abs/2412.08117
摘要:基于扩散的生成式AI因其优于其他生成技术(如生成对抗网络和变分自动编码器)的性能而备受关注。虽然它在计算机视觉和自然语言处理等领域取得了显着的进步,但它们在语音生成中的应用仍然没有得到充分的探索。主流的文本到语音系统主要将输出映射到频谱空间中的Mel-Spectrograms,由于MelSpecs的稀疏性导致高计算负载。为了解决这些限制,我们提出了LatentSpeech,一种新的TTS生成方法,利用潜在的扩散模型。通过使用潜在嵌入作为中间表示,LatentSpeech将目标维度降低到MelSpecs所需的5%,简化了TTS编码器和声码器的处理,并实现了高效的高质量语音生成。该研究首次将潜在扩散模型集成到TTS中,提高了生成语音的准确性和自然度。在基准数据集上的实验结果表明,与现有模型相比,LatentSpeech在单词错误率方面实现了25%的改善,在梅尔倒谱系数失真方面实现了24%的改善,在额外的训练数据下,进一步的改善分别上升到49.5%和26%。这些发现突出了潜在的语音推进国家的最先进的TTS技术
摘要:Diffusion-based Generative AI gains significant attention for its superiorperformance over other generative techniques like Generative AdversarialNetworks and Variational Autoencoders. While it has achieved notableadvancements in fields such as computer vision and natural language processing,their application in speech generation remains under-explored. MainstreamText-to-Speech systems primarily map outputs to Mel-Spectrograms in thespectral space, leading to high computational loads due to the sparsity ofMelSpecs. To address these limitations, we propose LatentSpeech, a novel TTSgeneration approach utilizing latent diffusion models. By using latentembeddings as the intermediate representation, LatentSpeech reduces the targetdimension to 5% of what is required for MelSpecs, simplifying the processingfor the TTS encoder and vocoder and enabling efficient high-quality speechgeneration. This study marks the first integration of latent diffusion modelsin TTS, enhancing the accuracy and naturalness of generated speech.Experimental results on benchmark datasets demonstrate that LatentSpeechachieves a 25% improvement in Word Error Rate and a 24% improvement in MelCepstral Distortion compared to existing models, with further improvementsrising to 49.5% and 26%, respectively, with additional training data. Thesefindings highlight the potential of LatentSpeech to advance thestate-of-the-art in TTS technology
标题:对齐器引导的训练范式:通过对齐器引导持续时间推进文本到语音模型
链接:https://arxiv.org/abs/2412.08112
摘要:文本到语音(TTS)系统的最新进展,如FastSpeech和StyleSpeech,显着提高了语音生成质量。然而,这些模型通常依赖于蒙特利尔强制调整器等外部工具生成的持续时间,这可能非常耗时且缺乏灵活性。尽管准确时长在实现自然韵律和可理解性方面发挥着关键作用,但其重要性往往被低估。为了解决这些局限性,我们提出了一种新的对齐器引导的训练范式,通过在TTS模型之前训练对齐器来优先考虑准确的持续时间标签。这种方法减少了对外部工具的依赖,并提高了对准精度。我们进一步探讨了不同的声学特征,包括梅尔频谱图,MFCC和潜在的功能,TTS模型的性能的影响。我们的实验结果表明,对齐器引导的持续时间标记可以实现高达16%的字错误率的改善,并显着提高音素和音调对齐。这些发现突出了我们的方法在优化TTS系统更自然和更清晰的语音生成的有效性。
摘要:Recent advancements in text-to-speech (TTS) systems, such as FastSpeech andStyleSpeech, have significantly improved speech generation quality. However,these models often rely on duration generated by external tools like theMontreal Forced Aligner, which can be time-consuming and lack flexibility. Theimportance of accurate duration is often underestimated, despite their crucialrole in achieving natural prosody and intelligibility. To address theselimitations, we propose a novel Aligner-Guided Training Paradigm thatprioritizes accurate duration labelling by training an aligner before the TTSmodel. This approach reduces dependence on external tools and enhancesalignment accuracy. We further explore the impact of different acousticfeatures, including Mel-Spectrograms, MFCCs, and latent features, on TTS modelperformance. Our experimental results show that aligner-guided durationlabelling can achieve up to a 16\% improvement in word error rate andsignificantly enhance phoneme and tone alignment. These findings highlight theeffectiveness of our approach in optimizing TTS systems for more natural andintelligible speech generation.
标题:弗雷切特音乐距离:生成符号音乐评估的指标
链接:https://arxiv.org/abs/2412.07948
摘要:在本文中,我们介绍了Frechet音乐距离(FMD),一种新的生成符号音乐模型的评估指标,灵感来自计算机视觉中的Frechet初始距离(FID)和生成音频中的Frechet音频距离(FAD)。FMD计算参考和生成的符号音乐嵌入分布之间的距离,捕获抽象的音乐特征。我们在多个数据集和模型上验证FMD。结果表明,FMD有效地区分模型质量,提供了一个特定领域的指标,用于评估符号音乐生成,并建立了一个可重复的标准,为未来的研究在符号音乐建模。
摘要:In this paper we introduce the Frechet Music Distance (FMD), a novelevaluation metric for generative symbolic music models, inspired by the FrechetInception Distance (FID) in computer vision and Frechet Audio Distance (FAD) ingenerative audio. FMD calculates the distance between distributions ofreference and generated symbolic music embeddings, capturing abstract musicalfeatures. We validate FMD across several datasets and models. Results indicatethat FMD effectively differentiates model quality, providing a domain-specificmetric for evaluating symbolic music generation, and establishing areproducible standard for future research in symbolic music modeling.
标题:我应该使用哪种增强功能?自我监督的音素图代表学习增强的实证研究
链接:https://arxiv.org/abs/2312.00502
备注:Accepted in IEEE ACCESS
摘要:尽管深度学习最近取得了进展,但其在现实世界医疗环境中的应用,如心音图(PCG)分类,仍然有限。一个重要的障碍是缺乏高质量的注释数据集,这阻碍了开发健壮的、可推广的模型,这些模型可以在新收集的、非分布(OOD)数据上表现良好。自监督学习(SSL)对比学习,通过使用未标记的数据来增强模型的鲁棒性,在缓解数据稀缺问题方面表现出了希望。尽管SSL方法已经在其他领域被提出和研究,但专注于数据增强对PCG分类模型鲁棒性的影响的工作是有限的。特别是,虽然增强是SSL中的关键组件,但在训练期间选择最合适的策略具有很大的挑战性。不适当的增强会导致性能大幅下降,甚至会阻碍网络学习有意义的表示的能力。为了解决这一差距,我们的研究旨在探索和评估各种基于音频的增强,并发现增强SSL模型在PCG分类中性能的组合。我们对多个数据集进行了全面的比较分析,评估了各种增强对模型性能的影响。我们的研究结果表明,根据训练分布,增强选择显着影响模型的鲁棒性,当对看不见的数据进行评估时,完全监督模型的有效性下降了32%,而SSL模型表现出更大的弹性,仅损失10%,甚至在某些情况下有所改善。这项研究还强调了最有前途和适当的PCG信号处理的增强,通过计算其对训练的影响大小。这些见解为研究人员在PCG信号处理中开发可靠模型提供了宝贵的指导。
摘要:Despite recent advancements in deep learning, its application in real-worldmedical settings, such as phonocardiogram (PCG) classification, remainslimited. A significant barrier is the lack of high-quality annotated datasets,which hampers the development of robust, generalizable models that can performwell on newly collected, out-of-distribution (OOD) data. Self-SupervisedLearning (SSL) contrastive learning, has shown promise in mitigating the issueof data scarcity by using unlabeled data to enhance model robustness. Eventhough SSL methods have been proposed and researched in other domains, worksfocusing on the impact of data augmentations on model robustness for PCGclassification are limited. In particular, while augmentations are a keycomponent in SSL, selecting the most suitable policy during training is highlychallenging. Improper augmentations can lead to substantial performancedegradation and even hinder a network's ability to learn meaningfulrepresentations. Addressing this gap, our research aims to explore and evaluatea wide range of audio-based augmentations and uncover combinations that enhanceSSL model performance in PCG classification. We conduct a comprehensivecomparative analysis across multiple datasets, assessing the impact of variousaugmentations on model performance. Our findings reveal that depending on thetraining distribution, augmentation choice significantly influences modelrobustness, with fully-supervised models experiencing up to a 32\% drop ineffectiveness when evaluated on unseen data, while SSL models demonstrategreater resilience, losing only 10\% or even improving in some cases. Thisstudy also highlights the most promising and appropriate augmentations for PCGsignal processing, by calculating their effect size on training. These insightsequip researchers with valuable guidelines for developing reliable models inPCG signal processing.
