本文经arXiv每日学术速递授权转载
微信公众号:arXiv_Daily
cs.SD语音
链接:https://arxiv.org/abs/2503.10522
备注:The code and datasets will be available at this https URL
摘要:音频和音乐生成已成为许多应用中的关键任务,但现有方法面临着重大限制:它们孤立地运行,没有跨模态的统一功能,缺乏高质量的多模态训练数据,并且难以有效地整合不同的输入。在这项工作中,我们提出了AudioX,一个统一的扩散Transformer模型的任何音频和音乐生成。与以前的特定领域模型不同,AudioX可以生成高质量的通用音频和音乐,同时提供灵活的自然语言控制和各种形式的无缝处理,包括文本,视频,图像,音乐和音频。它的关键创新是一种多模态掩蔽训练策略,该策略掩蔽了跨模态的输入,并迫使模型从掩蔽的输入中学习,从而产生鲁棒和统一的跨模态表示。为了解决数据稀缺问题,我们策划了两个综合数据集:基于VGGSound数据集的19万音频字幕的vggsound-caps,以及来自V2 M数据集的600万音乐字幕的V2 M-caps。大量的实验表明,AudioX不仅匹配或优于最先进的专业模型,而且在处理不同的输入方式和生成任务的统一架构中提供了显着的多功能性。代码和数据集将在https://zeyuet.github.io/AudioX/上提供
摘要:Audio and music generation have emerged as crucial tasks in manyapplications, yet existing approaches face significant limitations: theyoperate in isolation without unified capabilities across modalities, sufferfrom scarce high-quality, multi-modal training data, and struggle toeffectively integrate diverse inputs. In this work, we propose AudioX, aunified Diffusion Transformer model for Anything-to-Audio and Music Generation.Unlike previous domain-specific models, AudioX can generate both general audioand music with high quality, while offering flexible natural language controland seamless processing of various modalities including text, video, image,music, and audio. Its key innovation is a multi-modal masked training strategythat masks inputs across modalities and forces the model to learn from maskedinputs, yielding robust and unified cross-modal representations. To addressdata scarcity, we curate two comprehensive datasets: vggsound-caps with 190Kaudio captions based on the VGGSound dataset, and V2M-caps with 6 million musiccaptions derived from the V2M dataset. Extensive experiments demonstrate thatAudioX not only matches or outperforms state-of-the-art specialized models, butalso offers remarkable versatility in handling diverse input modalities andgeneration tasks within a unified architecture. The code and datasets will beavailable at https://zeyuet.github.io/AudioX/
标题:Whisper说话人识别:利用预先训练的多语言变形器实现稳健的说话人嵌入
链接:https://arxiv.org/abs/2503.10446
备注:6 pages
摘要:多语言环境中的说话人识别提出了独特的挑战,特别是当传统模型主要是在英语数据上训练时。在本文中,我们提出了WSI(Whisper Speaker Identification),这是一个框架,它重新利用了在大量多语言数据上预先训练的Whisper自动语音识别模型的编码器,通过利用在线硬三元组挖掘和自监督归一化温度尺度交叉熵损失的联合损失优化策略来生成鲁棒的扬声器嵌入。通过利用Whisper语言不可知的声学表示,我们的方法有效地区分了不同语言和录音条件下的扬声器。对多个语料库的广泛评估,包括VoxTube(多语言),JVS(日语),CallHome(德语,西班牙语,中文和日语)和Voxconverse(英语),表明WSI在较低的相等错误率和较高的AUC分数方面始终优于最先进的基线,即Pyannote Embedding,ECAPA TDNN和Xvector。这些结果验证了我们的假设,即多语言预训练的ASR编码器,结合联合损失优化,大大提高了非英语语言的说话人识别性能。
摘要:Speaker identification in multilingual settings presents unique challenges,particularly when conventional models are predominantly trained on Englishdata. In this paper, we propose WSI (Whisper Speaker Identification), aframework that repurposes the encoder of the Whisper automatic speechrecognition model pre trained on extensive multilingual data to generate robustspeaker embeddings via a joint loss optimization strategy that leverages onlinehard triplet mining and self supervised Normalized Temperature-scaled CrossEntropy loss. By capitalizing on Whisper language-agnostic acousticrepresentations, our approach effectively distinguishes speakers across diverselanguages and recording conditions. Extensive evaluations on multiple corpora,including VoxTube (multilingual), JVS (Japanese), CallHome (German, Spanish,Chinese, and Japanese), and Voxconverse (English), demonstrate that WSIconsistently outperforms state-of-the-art baselines, namely Pyannote Embedding,ECAPA TDNN, and Xvector, in terms of lower equal error rates and higher AUCscores. These results validate our hypothesis that a multilingual pre-trainedASR encoder, combined with joint loss optimization, substantially improvesspeaker identification performance in non-English languages.
标题:MACS:具有上下文意义和语义对齐的多源音频到图像生成
链接:https://arxiv.org/abs/2503.10287
摘要:在深度生成模型突破的推动下,音频到图像生成已经成为一个关键的跨模型任务,将复杂的听觉信号转换为丰富的视觉表示。然而,以往的工作只集中在单源音频输入的图像生成,忽略了多源特性的自然听觉场景,从而限制了在生成全面的视觉内容的性能。为了弥补这一差距,提出了一种称为MACS的方法来进行多源音频到图像的生成。这是第一个明确分离多源音频以在图像生成之前捕获丰富音频分量的工作。MACS是一种两阶段方法。在第一阶段,多源音频输入通过弱监督方法分离,其中音频和文本标签通过使用大型预训练CLAP模型投射到公共空间中而在语义上对齐。我们引入了一个排名损失考虑分离的音频信号的上下文意义。在第二阶段中,通过仅使用可训练适配器和MLP层将分离的音频信号映射到生成条件来实现有效的图像生成。我们将LLP数据集作为第一个完整的多源音频到图像生成基准进行预处理。实验在多源、混合源和单源音频到图像生成任务上进行。所提出的MACS在所有任务的21个评价指标中的17个方面优于当前最先进的方法,并提供卓越的视觉质量。该代码将公开发布。
摘要:Propelled by the breakthrough in deep generative models, audio-to-imagegeneration has emerged as a pivotal cross-model task that converts complexauditory signals into rich visual representations. However, previous works onlyfocus on single-source audio inputs for image generation, ignoring themulti-source characteristic in natural auditory scenes, thus limiting theperformance in generating comprehensive visual content. To bridge this gap, amethod called MACS is proposed to conduct multi-source audio-to-imagegeneration. This is the first work that explicitly separates multi-source audioto capture the rich audio components before image generation. MACS is atwo-stage method. In the first stage, multi-source audio inputs are separatedby a weakly supervised method, where the audio and text labels are semanticallyaligned by casting into a common space using the large pre-trained CLAP model.We introduce a ranking loss to consider the contextual significance of theseparated audio signals. In the second stage, efficient image generation isachieved by mapping the separated audio signals to the generation conditionusing only a trainable adapter and a MLP layer. We preprocess the LLP datasetas the first full multi-source audio-to-image generation benchmark. Theexperiments are conducted on multi-source, mixed-source, and single-sourceaudio-to-image generation tasks. The proposed MACS outperforms the currentstate-of-the-art methods in 17 of the 21 evaluation indexes on all tasks anddelivers superior visual quality. The code will be publicly available.
标题:基于LLM的语音翻译的自适应内部语音-文本对齐
链接:https://arxiv.org/abs/2503.10211
备注:12 pages, 7 figures
摘要:大型语言模型(LLM)的最新进展导致了各种任务的重大突破,为基于LLM的语音翻译系统的开发奠定了基础。现有的方法主要集中在跨模态对齐输入和输出,而忽略了模型表示内更深层次的语义对齐。为了解决这一限制,我们提出了一种自适应内部语音-文本对齐(AI-STA)方法,通过明确对齐LLM内选定层的语音和文本表示来弥合模态差距。为了实现这一目标,我们利用最优传输(OT)理论来量化语音和文本之间的细粒度表示差异。此外,我们利用跨模态检索技术来识别最适合对齐的层,并对这些层进行联合训练。语音翻译(ST)任务的实验结果表明,AI-STA显着提高了大型语音文本模型(LSM)的翻译性能,优于以前的最先进的方法。我们的研究结果强调了LLM中内层语音文本对齐的重要性,并为增强跨模态学习提供了新的见解。
摘要:Recent advancement of large language models (LLMs) has led to significantbreakthroughs across various tasks, laying the foundation for the developmentof LLM-based speech translation systems. Existing methods primarily focus onaligning inputs and outputs across modalities while overlooking deeper semanticalignment within model representations. To address this limitation, we proposean Adaptive Inner Speech-Text Alignment (AI-STA) method to bridge the modalitygap by explicitly aligning speech and text representations at selected layerswithin LLMs. To achieve this, we leverage the optimal transport (OT) theory toquantify fine-grained representation discrepancies between speech and text.Furthermore, we utilize the cross-modal retrieval technique to identify thelayers that are best suited for alignment and perform joint training on theselayers. Experimental results on speech translation (ST) tasks demonstrate thatAI-STA significantly improves the translation performance of large speech-textmodels (LSMs), outperforming previous state-of-the-art approaches. Our findingshighlight the importance of inner-layer speech-text alignment in LLMs andprovide new insights into enhancing cross-modal learning.
标题:具有自我监督学习功能的高效适配器调整,用于联合歌唱声音节拍和低沉节拍跟踪
链接:https://arxiv.org/abs/2503.10086
备注:Accepted by ISMIR2024
摘要:由于缺乏通常包含稳健的节奏和和声模式的音乐伴奏,歌唱语音节拍跟踪是一项具有挑战性的任务,大多数现有的节拍跟踪系统利用这些模式并且对于估计节拍是必不可少的。本文提出了一种新的基于时间卷积网络的节拍跟踪方法,该方法具有自监督学习(SSL)表示和适配器调整,以联合跟踪歌唱声音的节拍和强拍。SSL DistilHuBERT表示被用来捕获歌唱声音的语义信息,并进一步与通用频谱特征融合,以促进节拍估计。通过有效的适配器调谐减少了对于非均匀歌唱声音数据特别突出的可变性的来源。大量的实验表明,特征融合和适配器调整分别提高了性能,两者的结合导致比未适应的基线系统更好的性能,在节拍和强拍跟踪上分别有高达31.6%和42.4%的绝对F1分数提高。
摘要:Singing voice beat tracking is a challenging task, due to the lack of musicalaccompaniment that often contains robust rhythmic and harmonic patterns,something most existing beat tracking systems utilize and can be essential forestimating beats. In this paper, a novel temporal convolutional network-basedbeat-tracking approach featuring self-supervised learning (SSL) representationsand adapter tuning is proposed to track the beat and downbeat of singing voicesjointly. The SSL DistilHuBERT representations are utilized to capture thesemantic information of singing voices and are further fused with the genericspectral features to facilitate beat estimation. Sources of variabilities thatare particularly prominent with the non-homogeneous singing voice data arereduced by the efficient adapter tuning. Extensive experiments show thatfeature fusion and adapter tuning improve the performance individually, and thecombination of both leads to significantly better performances than theun-adapted baseline system, with up to 31.6% and 42.4% absolute F1-scoreimprovements on beat and downbeat tracking, respectively.
标题:OpenAI Whisper模型的量化:比较分析
链接:https://arxiv.org/abs/2503.09905
备注:7 pages
摘要:自动语音识别(ASR)模型在字幕、语音翻译和实时转录等应用中已经获得了显著的地位。本文研究Whisper和两个模型变体:一个针对实时语音流进行了优化,另一个针对离线转录进行了优化。值得注意的是,这些模型被发现会产生幻觉内容,降低转录的可靠性。此外,更大的模型变体表现出增加的延迟,并对资源受限设备上的部署提出挑战。本研究分析了三种Whisper模型之间的异同,定性地考察了它们的不同能力。接下来,这项研究量化了模型量化对延迟的影响,并评估了其在边缘部署中的可行性。使用开源LibriSpeech数据集,本文使用3种量化方法(INT 4,INT 5,INT 8)评估了单词错误率(WER)以及whispermpp的延迟分析。结果表明,量化减少了19%的延迟和45%的模型大小,同时保持转录的准确性。这些发现为不同Whisper模型和边缘设备部署可能性的最佳用例提供了见解。所有代码、数据集和实现细节都可以在公共GitHub存储库中找到:https://github.com/allisonandreyev/WhisperQuantization.git
摘要:Automated speech recognition (ASR) models have gained prominence forapplications such as captioning, speech translation, and live transcription.This paper studies Whisper and two model variants: one optimized for livespeech streaming and another for offline transcription. Notably, these modelshave been found to generate hallucinated content, reducing transcriptionreliability. Furthermore, larger model variants exhibit increased latency andpose challenges for deployment on resource-constrained devices. This studyanalyzes the similarities and differences between three Whisper models,qualitatively examining their distinct capabilities. Next, this studyquantifies the impact of model quantization on latency and evaluates itsviability for edge deployment. Using the open source LibriSpeech dataset, thispaper evaluates the word error rate (WER) along with latency analysis ofwhispercpp using 3 quantization methods (INT4, INT5, INT8). Results show thatquantization reduces latency by 19\% and model size by 45\%, while preservingtranscription accuracy. These findings provide insights into the optimal usecases of different Whisper models and edge device deployment possibilities. Allcode, datasets, and implementation details are available in a public GitHubrepository: https://github.com/allisonandreyev/WhisperQuantization.git
标题:处理异常声音检测的域转移:DCASE相关工作回顾
链接:https://arxiv.org/abs/2503.10435
摘要:在复杂环境中检测异常声音时,主要困难之一是训练的模型必须对监测目标信号的细微差异敏感,而许多实际应用还要求它们对声学域的变化不敏感。这种域偏移的示例包括改变麦克风的类型或声学传感器的位置,这可能对声学信号产生比细微异常本身更强的影响。此外,用户通常旨在仅在源域数据上训练模型,他们可能具有相对大的源域数据集合,并且他们希望这样的经训练的模型将能够通过仅提供最小数量的样本来表征该域中的声学信号而很好地推广到看不见的目标域。在这项工作中,我们审查和讨论最近的出版物,专注于这一领域的泛化问题的异常声音检测的背景下,DCASE的挑战声机状态监测。
摘要:When detecting anomalous sounds in complex environments, one of the maindifficulties is that trained models must be sensitive to subtle differences inmonitored target signals, while many practical applications also require themto be insensitive to changes in acoustic domains. Examples of such domainshifts include changing the type of microphone or the location of acousticsensors, which can have a much stronger impact on the acoustic signal thansubtle anomalies themselves. Moreover, users typically aim to train a modelonly on source domain data, which they may have a relatively large collectionof, and they hope that such a trained model will be able to generalize well toan unseen target domain by providing only a minimal number of samples tocharacterize the acoustic signals in that domain. In this work, we review anddiscuss recent publications focusing on this domain generalization problem foranomalous sound detection in the context of the DCASE challenges on acousticmachine condition monitoring.
标题:ValSub:对验证数据进行二次抽样,以减少ASB个性化期间的遗忘
链接:https://arxiv.org/abs/2503.09906
备注:Accepted at ICASSP 2025
摘要:自动语音识别(ASR)被广泛用于诸如移动电话之类的消费者设备中。最近,个性化或设备上的模型微调已经表明,ASR模型对目标用户语音的适应提高了它们在罕见单词或口音语音上的性能。尽管有这些收益,但对用户数据(目标域)的微调有可能使个性化模型忘记关于其原始训练分布(源域)的知识,即灾难性遗忘,从而导致低于标准的一般ASR性能。一个简单而有效的方法来对抗灾难性遗忘是通过一个验证集,代表源域分布来衡量遗忘。然而,这样的验证集对于移动设备是大的并且不切实际的。为此,我们提出了一种新的方法来子采样一个相当大的验证集到一个较小的,同时保持估计遗忘的能力。我们证明了这样的数据集在减轻遗忘的有效性,利用它来动态地确定理想的微调时期的数量。当针对50倍大的验证集(oracle)测量每个用户微调时期的偏差时,与随机选择的相同大小的子集(3.78-8.65)相比,我们的方法实现了更低的平均绝对误差(3.39)。与随机基线不同,我们的方法在三个不同的遗忘阈值上始终跟踪预言机的行为。
摘要:Automatic Speech Recognition (ASR) is widely used within consumer devicessuch as mobile phones. Recently, personalization or on-device model fine-tuninghas shown that adaptation of ASR models towards target user speech improvestheir performance over rare words or accented speech. Despite these gains,fine-tuning on user data (target domain) risks the personalized model to forgetknowledge about its original training distribution (source domain) i.e.catastrophic forgetting, leading to subpar general ASR performance. A simpleand efficient approach to combat catastrophic forgetting is to measureforgetting via a validation set that represents the source domain distribution.However, such validation sets are large and impractical for mobile devices.Towards this, we propose a novel method to subsample a substantially largevalidation set into a smaller one while maintaining the ability to estimateforgetting. We demonstrate the efficacy of such a dataset in mitigatingforgetting by utilizing it to dynamically determine the number of idealfine-tuning epochs. When measuring the deviations in per user fine-tuningepochs against a 50x larger validation set (oracle), our method achieves alower mean-absolute-error (3.39) compared to randomly selected subsets of thesame size (3.78-8.65). Unlike random baselines, our method consistently tracksthe oracle's behaviour across three different forgetting thresholds.
标题:处理异常声音检测的域转移:DCASE相关工作回顾
链接:https://arxiv.org/abs/2503.10435
摘要:在复杂环境中检测异常声音时,主要困难之一是训练的模型必须对监测目标信号的细微差异敏感,而许多实际应用还要求它们对声学域的变化不敏感。这种域偏移的示例包括改变麦克风的类型或声学传感器的位置,这可能对声学信号产生比细微异常本身更强的影响。此外,用户通常旨在仅在源域数据上训练模型,他们可能具有相对大的源域数据集合,并且他们希望这样的经训练的模型将能够通过仅提供最小数量的样本来表征该域中的声学信号而很好地推广到看不见的目标域。在这项工作中,我们审查和讨论最近的出版物,专注于这一领域的泛化问题的异常声音检测的背景下,DCASE的挑战声机状态监测。
摘要:When detecting anomalous sounds in complex environments, one of the maindifficulties is that trained models must be sensitive to subtle differences inmonitored target signals, while many practical applications also require themto be insensitive to changes in acoustic domains. Examples of such domainshifts include changing the type of microphone or the location of acousticsensors, which can have a much stronger impact on the acoustic signal thansubtle anomalies themselves. Moreover, users typically aim to train a modelonly on source domain data, which they may have a relatively large collectionof, and they hope that such a trained model will be able to generalize well toan unseen target domain by providing only a minimal number of samples tocharacterize the acoustic signals in that domain. In this work, we review anddiscuss recent publications focusing on this domain generalization problem foranomalous sound detection in the context of the DCASE challenges on acousticmachine condition monitoring.
标题:从语音检测帕金森病的双语双头深度模型
链接:https://arxiv.org/abs/2503.10301
备注:Accepted at ICASSP 2025 - Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses
摘要:这项工作旨在通过提出一种用于基于类型的二进制分类的ad-hoc双头深度神经架构,解决双语环境中语音信号中的帕金森病(PD)检测问题。其中一个头专门用于旋转运动模式。另一个头寻找自然的语音模式存在于连续的口语话语。根据输入的性质,两个磁头中只有一个磁头工作。语音表示提取自监督学习(SSL)模型和小波变换。利用自适应层、卷积瓶颈和对比学习来减少语言之间的差异。我们的解决方案是针对两个不同的数据集,EWA-DB和PC-GITA,分别涵盖斯洛伐克语和西班牙语,进行评估。结果表明,在单一语言数据集上训练的传统模型难以进行跨语言泛化,而数据集的朴素组合是次优的。相比之下,我们的模型同时提高了两种语言的泛化能力。
摘要:This work aims to tackle the Parkinson's disease (PD) detection problem fromthe speech signal in a bilingual setting by proposing an ad-hoc dual-head deepneural architecture for type-based binary classification. One head isspecialized for diadochokinetic patterns. The other head looks for naturalspeech patterns present in continuous spoken utterances. Only one of the twoheads is operative accordingly to the nature of the input. Speechrepresentations are extracted from self-supervised learning (SSL) models andwavelet transforms. Adaptive layers, convolutional bottlenecks, and contrastivelearning are exploited to reduce variations across languages. Our solution isassessed against two distinct datasets, EWA-DB, and PC-GITA, which cover Slovakand Spanish languages, respectively. Results indicate that conventional modelstrained on a single language dataset struggle with cross-linguisticgeneralization, and naive combinations of datasets are suboptimal. In contrast,our model improves generalization on both languages, simultaneously.
标题:声学场估计:理论与应用
链接:https://arxiv.org/abs/2503.10016
备注:published in Foundations and Trends in Signal Processing, vol. 19, no. 1; see https://www.nowpublishers.com/article/Details/SIG-121
摘要:声音的空间信息在从日常活动到先进工程技术的各种情况下都起着至关重要的作用。为了充分利用其潜力,在文献中已经进行了许多关于空间音频信号处理的研究。声场估计是可以应用于广泛的声学信号处理技术的关键基础技术之一,包括使用扬声器的声场再现和通过耳机的双耳回放。本文的目的是提出一个概述的声场估计方法。在提供必要的数学背景之后,将解释声场估计的两种不同方法。本文着重阐明每种方法的基本理论,同时也参考了最新的发展。最后,将讨论几种声信号处理技术作为声场估计应用的例子。
摘要:The spatial information of sound plays a crucial role in various situations,ranging from daily activities to advanced engineering technologies. To fullyutilize its potential, numerous research studies on spatial audio signalprocessing have been carried out in the literature. Sound field estimation isone of the key foundational technologies that can be applied to a wide range ofacoustic signal processing techniques, including sound field reproduction usingloudspeakers and binaural playback through headphones. The purpose of thispaper is to present an overview of sound field estimation methods. Afterproviding the necessary mathematical background, two different approaches tosound field estimation will be explained. This paper focuses on clarifying theessential theories of each approach, while also referencing state-of-the-artdevelopments. Finally, several acoustic signal processing technologies will bediscussed as examples of the application of sound field estimation.
标题:ValSub:对验证数据进行二次抽样,以减少ASB个性化期间的遗忘
链接:https://arxiv.org/abs/2503.09906
备注:Accepted at ICASSP 2025
摘要:自动语音识别(ASR)被广泛用于诸如移动电话之类的消费者设备中。最近,个性化或设备上的模型微调已经表明,ASR模型对目标用户语音的适应提高了它们在罕见单词或口音语音上的性能。尽管有这些收益,但对用户数据(目标域)的微调有可能使个性化模型忘记关于其原始训练分布(源域)的知识,即灾难性遗忘,从而导致低于标准的一般ASR性能。一个简单而有效的方法来对抗灾难性遗忘是通过一个验证集,代表源域分布来衡量遗忘。然而,这样的验证集对于移动设备是大的并且不切实际的。为此,我们提出了一种新的方法来子采样一个相当大的验证集到一个较小的,同时保持估计遗忘的能力。我们通过利用这样的数据集动态确定理想微调时期的数量,证明了它在减轻遗忘方面的功效。当针对50倍大的验证集(oracle)测量每个用户微调时期的偏差时,与随机选择的相同大小的子集(3.78-8.65)相比,我们的方法实现了更低的平均绝对误差(3.39)。与随机基线不同,我们的方法在三个不同的遗忘阈值上始终跟踪预言机的行为。
摘要:Automatic Speech Recognition (ASR) is widely used within consumer devicessuch as mobile phones. Recently, personalization or on-device model fine-tuninghas shown that adaptation of ASR models towards target user speech improvestheir performance over rare words or accented speech. Despite these gains,fine-tuning on user data (target domain) risks the personalized model to forgetknowledge about its original training distribution (source domain) i.e.catastrophic forgetting, leading to subpar general ASR performance. A simpleand efficient approach to combat catastrophic forgetting is to measureforgetting via a validation set that represents the source domain distribution.However, such validation sets are large and impractical for mobile devices.Towards this, we propose a novel method to subsample a substantially largevalidation set into a smaller one while maintaining the ability to estimateforgetting. We demonstrate the efficacy of such a dataset in mitigatingforgetting by utilizing it to dynamically determine the number of idealfine-tuning epochs. When measuring the deviations in per user fine-tuningepochs against a 50x larger validation set (oracle), our method achieves alower mean-absolute-error (3.39) compared to randomly selected subsets of thesame size (3.78-8.65). Unlike random baselines, our method consistently tracksthe oracle's behaviour across three different forgetting thresholds.
标题:AudioX:用于任何事物到音频生成的扩散Transformer
链接:https://arxiv.org/abs/2503.10522
备注:The code and datasets will be available at this https URL
摘要:音频和音乐生成已成为许多应用中的关键任务,但现有方法面临着重大限制:它们孤立地运行,没有跨模态的统一功能,缺乏高质量的多模态训练数据,并且难以有效地整合不同的输入。在这项工作中,我们提出了AudioX,一个统一的扩散Transformer模型的任何音频和音乐生成。与以前的特定领域模型不同,AudioX可以生成高质量的通用音频和音乐,同时提供灵活的自然语言控制和各种形式的无缝处理,包括文本,视频,图像,音乐和音频。它的关键创新是一种多模态掩蔽训练策略,该策略掩蔽了跨模态的输入,并迫使模型从掩蔽的输入中学习,从而产生鲁棒和统一的跨模态表示。为了解决数据稀缺问题,我们策划了两个综合数据集:基于VGGSound数据集的19万音频字幕的vggsound-caps,以及来自V2 M数据集的600万音乐字幕的V2 M-caps。大量的实验表明,AudioX不仅匹配或优于最先进的专业模型,而且在处理不同的输入方式和生成任务的统一架构中提供了显着的多功能性。代码和数据集将在https://zeyuet.github.io/AudioX/上提供
摘要:Audio and music generation have emerged as crucial tasks in manyapplications, yet existing approaches face significant limitations: theyoperate in isolation without unified capabilities across modalities, sufferfrom scarce high-quality, multi-modal training data, and struggle toeffectively integrate diverse inputs. In this work, we propose AudioX, aunified Diffusion Transformer model for Anything-to-Audio and Music Generation.Unlike previous domain-specific models, AudioX can generate both general audioand music with high quality, while offering flexible natural language controland seamless processing of various modalities including text, video, image,music, and audio. Its key innovation is a multi-modal masked training strategythat masks inputs across modalities and forces the model to learn from maskedinputs, yielding robust and unified cross-modal representations. To addressdata scarcity, we curate two comprehensive datasets: vggsound-caps with 190Kaudio captions based on the VGGSound dataset, and V2M-caps with 6 million musiccaptions derived from the V2M dataset. Extensive experiments demonstrate thatAudioX not only matches or outperforms state-of-the-art specialized models, butalso offers remarkable versatility in handling diverse input modalities andgeneration tasks within a unified architecture. The code and datasets will beavailable at https://zeyuet.github.io/AudioX/
标题:Whisper说话人识别:利用预先训练的多语言变形器实现稳健的说话人嵌入
链接:https://arxiv.org/abs/2503.10446
备注:6 pages
摘要:多语言环境中的说话人识别提出了独特的挑战,特别是当传统模型主要是在英语数据上训练时。在本文中,我们提出了WSI(Whisper Speaker Identification),这是一个框架,它重新利用了在大量多语言数据上预先训练的Whisper自动语音识别模型的编码器,通过利用在线硬三元组挖掘和自监督归一化温度尺度交叉熵损失的联合损失优化策略来生成鲁棒的扬声器嵌入。通过利用Whisper语言不可知的声学表示,我们的方法有效地区分了不同语言和录音条件下的扬声器。对多个语料库的广泛评估,包括VoxTube(多语言),JVS(日语),CallHome(德语,西班牙语,中文和日语)和Voxconverse(英语),表明WSI在较低的相等错误率和较高的AUC分数方面始终优于最先进的基线,即Pyannote Embedding,ECAPA TDNN和Xvector。这些结果验证了我们的假设,即多语言预训练的ASR编码器,结合联合损失优化,大大提高了非英语语言的说话人识别性能。
摘要:Speaker identification in multilingual settings presents unique challenges,particularly when conventional models are predominantly trained on Englishdata. In this paper, we propose WSI (Whisper Speaker Identification), aframework that repurposes the encoder of the Whisper automatic speechrecognition model pre trained on extensive multilingual data to generate robustspeaker embeddings via a joint loss optimization strategy that leverages onlinehard triplet mining and self supervised Normalized Temperature-scaled CrossEntropy loss. By capitalizing on Whisper language-agnostic acousticrepresentations, our approach effectively distinguishes speakers across diverselanguages and recording conditions. Extensive evaluations on multiple corpora,including VoxTube (multilingual), JVS (Japanese), CallHome (German, Spanish,Chinese, and Japanese), and Voxconverse (English), demonstrate that WSIconsistently outperforms state-of-the-art baselines, namely Pyannote Embedding,ECAPA TDNN, and Xvector, in terms of lower equal error rates and higher AUCscores. These results validate our hypothesis that a multilingual pre-trainedASR encoder, combined with joint loss optimization, substantially improvesspeaker identification performance in non-English languages.
标题:MACS:具有上下文意义和语义对齐的多源音频到图像生成
链接:https://arxiv.org/abs/2503.10287
摘要:在深度生成模型突破的推动下,音频到图像生成已经成为一个关键的跨模型任务,将复杂的听觉信号转换为丰富的视觉表示。然而,以往的工作只集中在单源音频输入的图像生成,忽略了多源特性的自然听觉场景,从而限制了在生成全面的视觉内容的性能。为了弥补这一差距,提出了一种称为MACS的方法来进行多源音频到图像的生成。这是第一个明确分离多源音频以在图像生成之前捕获丰富音频分量的工作。MACS是一种两阶段方法。在第一阶段,多源音频输入通过弱监督方法分离,其中音频和文本标签通过使用大型预训练CLAP模型投射到公共空间中而在语义上对齐。我们引入了一个排名损失考虑分离的音频信号的上下文意义。在第二阶段中,通过仅使用可训练适配器和MLP层将分离的音频信号映射到生成条件来实现有效的图像生成。我们将LLP数据集作为第一个完整的多源音频到图像生成基准进行预处理。实验在多源、混合源和单源音频到图像生成任务上进行。所提出的MACS在所有任务的21个评价指标中的17个方面优于当前最先进的方法,并提供卓越的视觉质量。该代码将公开发布。
摘要:Propelled by the breakthrough in deep generative models, audio-to-imagegeneration has emerged as a pivotal cross-model task that converts complexauditory signals into rich visual representations. However, previous works onlyfocus on single-source audio inputs for image generation, ignoring themulti-source characteristic in natural auditory scenes, thus limiting theperformance in generating comprehensive visual content. To bridge this gap, amethod called MACS is proposed to conduct multi-source audio-to-imagegeneration. This is the first work that explicitly separates multi-source audioto capture the rich audio components before image generation. MACS is atwo-stage method. In the first stage, multi-source audio inputs are separatedby a weakly supervised method, where the audio and text labels are semanticallyaligned by casting into a common space using the large pre-trained CLAP model.We introduce a ranking loss to consider the contextual significance of theseparated audio signals. In the second stage, efficient image generation isachieved by mapping the separated audio signals to the generation conditionusing only a trainable adapter and a MLP layer. We preprocess the LLP datasetas the first full multi-source audio-to-image generation benchmark. Theexperiments are conducted on multi-source, mixed-source, and single-sourceaudio-to-image generation tasks. The proposed MACS outperforms the currentstate-of-the-art methods in 17 of the 21 evaluation indexes on all tasks anddelivers superior visual quality. The code will be publicly available.
标题:基于LLM的语音翻译的自适应内部语音-文本对齐
链接:https://arxiv.org/abs/2503.10211
备注:12 pages, 7 figures
摘要:大型语言模型(LLM)的最新进展导致了各种任务的重大突破,为基于LLM的语音翻译系统的开发奠定了基础。现有的方法主要集中在跨模态对齐输入和输出,而忽略了模型表示内更深层次的语义对齐。为了解决这一限制,我们提出了一种自适应内部语音-文本对齐(AI-STA)方法,通过明确对齐LLM内选定层的语音和文本表示来弥合模态差距。为了实现这一目标,我们利用最优传输(OT)理论来量化语音和文本之间的细粒度表示差异。此外,我们利用跨模态检索技术来识别最适合对齐的层,并对这些层进行联合训练。语音翻译(ST)任务的实验结果表明,AI-STA显着提高了大型语音文本模型(LSM)的翻译性能,优于以前的最先进的方法。我们的研究结果强调了LLM中内层语音文本对齐的重要性,并为增强跨模态学习提供了新的见解。
摘要:Recent advancement of large language models (LLMs) has led to significantbreakthroughs across various tasks, laying the foundation for the developmentof LLM-based speech translation systems. Existing methods primarily focus onaligning inputs and outputs across modalities while overlooking deeper semanticalignment within model representations. To address this limitation, we proposean Adaptive Inner Speech-Text Alignment (AI-STA) method to bridge the modalitygap by explicitly aligning speech and text representations at selected layerswithin LLMs. To achieve this, we leverage the optimal transport (OT) theory toquantify fine-grained representation discrepancies between speech and text.Furthermore, we utilize the cross-modal retrieval technique to identify thelayers that are best suited for alignment and perform joint training on theselayers. Experimental results on speech translation (ST) tasks demonstrate thatAI-STA significantly improves the translation performance of large speech-textmodels (LSMs), outperforming previous state-of-the-art approaches. Our findingshighlight the importance of inner-layer speech-text alignment in LLMs andprovide new insights into enhancing cross-modal learning.
标题:具有自我监督学习功能的高效适配器调整,用于联合歌唱声音节拍和低沉节拍跟踪
链接:https://arxiv.org/abs/2503.10086
备注:Accepted by ISMIR2024
摘要:由于缺乏通常包含稳健的节奏和和声模式的音乐伴奏,歌唱语音节拍跟踪是一项具有挑战性的任务,大多数现有的节拍跟踪系统利用这些模式并且对于估计节拍是必不可少的。本文提出了一种新的基于时间卷积网络的节拍跟踪方法,该方法具有自监督学习(SSL)表示和适配器调整,以联合跟踪歌唱声音的节拍和强拍。SSL DistilHuBERT表示被用来捕获歌唱声音的语义信息,并进一步与通用频谱特征融合,以促进节拍估计。通过有效的适配器调谐减少了对于非均匀歌唱声音数据特别突出的可变性的来源。大量的实验表明,特征融合和适配器调整分别提高了性能,两者的结合导致比未适应的基线系统更好的性能,在节拍和强拍跟踪上分别有高达31.6%和42.4%的绝对F1分数提高。
摘要:Singing voice beat tracking is a challenging task, due to the lack of musicalaccompaniment that often contains robust rhythmic and harmonic patterns,something most existing beat tracking systems utilize and can be essential forestimating beats. In this paper, a novel temporal convolutional network-basedbeat-tracking approach featuring self-supervised learning (SSL) representationsand adapter tuning is proposed to track the beat and downbeat of singing voicesjointly. The SSL DistilHuBERT representations are utilized to capture thesemantic information of singing voices and are further fused with the genericspectral features to facilitate beat estimation. Sources of variabilities thatare particularly prominent with the non-homogeneous singing voice data arereduced by the efficient adapter tuning. Extensive experiments show thatfeature fusion and adapter tuning improve the performance individually, and thecombination of both leads to significantly better performances than theun-adapted baseline system, with up to 31.6% and 42.4% absolute F1-scoreimprovements on beat and downbeat tracking, respectively.
标题:OpenAI Whisper模型的量化:比较分析
链接:https://arxiv.org/abs/2503.09905
备注:7 pages
摘要:自动语音识别(ASR)模型在字幕、语音翻译和实时转录等应用中已经获得了显著的地位。本文研究Whisper和两个模型变体:一个针对实时语音流进行了优化,另一个针对离线转录进行了优化。值得注意的是,这些模型被发现会产生幻觉内容,降低转录的可靠性。此外,更大的模型变体表现出增加的延迟,并对资源受限设备上的部署提出挑战。本研究分析了三种Whisper模型之间的异同,定性地考察了它们的不同能力。接下来,这项研究量化了模型量化对延迟的影响,并评估了其在边缘部署中的可行性。使用开源LibriSpeech数据集,本文使用3种量化方法(INT 4,INT 5,INT 8)评估了单词错误率(WER)以及whispermpp的延迟分析。结果表明,量化将延迟减少了19%,模型大小减少了45%,同时保持了转录准确性。这些发现为不同Whisper模型和边缘设备部署可能性的最佳用例提供了见解。所有代码、数据集和实现细节都可以在公共GitHub存储库中找到:https://github.com/allisonandreyev/WhisperQuantization.git
摘要:Automated speech recognition (ASR) models have gained prominence forapplications such as captioning, speech translation, and live transcription.This paper studies Whisper and two model variants: one optimized for livespeech streaming and another for offline transcription. Notably, these modelshave been found to generate hallucinated content, reducing transcriptionreliability. Furthermore, larger model variants exhibit increased latency andpose challenges for deployment on resource-constrained devices. This studyanalyzes the similarities and differences between three Whisper models,qualitatively examining their distinct capabilities. Next, this studyquantifies the impact of model quantization on latency and evaluates itsviability for edge deployment. Using the open source LibriSpeech dataset, thispaper evaluates the word error rate (WER) along with latency analysis ofwhispercpp using 3 quantization methods (INT4, INT5, INT8). Results show thatquantization reduces latency by 19\% and model size by 45\%, while preservingtranscription accuracy. These findings provide insights into the optimal usecases of different Whisper models and edge device deployment possibilities. Allcode, datasets, and implementation details are available in a public GitHubrepository: https://github.com/allisonandreyev/WhisperQuantization.git
