今日论文合集:cs.SD语音5篇,eess.AS音频处理10篇。本文经arXiv每日学术速递授权转载
【1】Images that Sound: Composing Images and Sounds on a Single Canvas
作者:Ziyang Chen,Daniel Geng,Andrew Owens备注:Project site: this https URL摘要:声谱图是声音的2D表示,看起来与我们视觉世界中的图像非常不同。而自然的图像,当作为声谱图播放时,会发出不自然的声音。在本文中,我们表明,它是可以合成的频谱图,同时看起来像自然的图像和声音像自然的音频。我们称这些声谱图为声音的图像。我们的方法是简单和zero-shot的,它利用了在共享潜在空间中操作的预训练的文本到图像和文本到频谱图扩散模型。在反向过程中,我们并行地使用音频和图像扩散模型对噪声潜伏期进行降噪,从而产生可能在两种模型下的样本。通过定量评估和感知研究,我们发现,我们的方法成功地生成与所需的音频提示对齐的频谱图,同时也采取了所需的图像提示的视觉外观。请参阅我们的项目页面视频结果:https://ificl.github.io/images-that-sound/摘要:Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/
【2】 Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification标题:具有渐进通道融合的邻居注意力Transformer用于说话者验证备注:8 pages, 2 figures, 3 tables摘要:基于变换器的说话人验证架构通常需要比ECAPA-TDNN更多的训练数据。因此,最近的工作通常是在VoxCeleb1&2上训练的。我们提出了一个基于自我注意力的骨干网络,当单独在VoxCeleb 2上训练时,它可以达到有竞争力的结果。该网络在邻域关注和全局关注之间交替,以捕获局部和全局特征,然后聚合不同层次的特征,最后执行关注的统计池。此外,我们采用了渐进的通道融合策略,以扩大在通道维的感受野作为网络的深化。我们在VoxCeleb 2上训练了所提出的PCF-NAT模型,并在VoxCeleb 1和VoxSRC的验证集上对其进行了评估。浅PCF-NAT的EER和minDCF平均比类似大小的ECAPA-TDNN低20%以上。深度PCF-NAT在VoxCeleb 1-O上实现了低于0.5%的EER。摘要:Transformer-based architectures for speaker verification typically require more training data than ECAPA-TDNN. Therefore, recent work has generally been trained on VoxCeleb1&2. We propose a backbone network based on self-attention, which can achieve competitive results when trained on VoxCeleb2 alone. The network alternates between neighborhood attention and global attention to capture local and global features, then aggregates features of different hierarchical levels, and finally performs attentive statistical pooling. Additionally, we employ a progressive channel fusion strategy to expand the receptive field in the channel dimension as the network deepens. We trained the proposed PCF-NAT model on VoxCeleb2 and evaluated it on VoxCeleb1 and the validation sets of VoxSRC. The EER and minDCF of the shallow PCF-NAT are on average more than 20% lower than those of similarly sized ECAPA-TDNN. Deep PCF-NAT achieves an EER lower than 0.5% on VoxCeleb1-O.【3】 DAC-JAX: A JAX Implementation of the Descript Audio Codec标题:DAC-JAX:描述音频编解码器的JAX实现备注:5 pages, 3 figures, 2 tables摘要:我们使用Google的JAX生态系统Flax,Optax,Orbax,AUX和CLU来实现Descept Audio Codec(DAC)的开源实现。我们的代码库可以重用原始PyTorch DAC的模型权重,并且我们确认,如果给定相同的输入,这两种实现会产生等效的令牌序列和解码音频。我们提供了一个支持设备并行性的训练和微调脚本,尽管我们只使用小数据集的简短训练运行对其进行了验证。即使GPU内存有限,原始的DAC也可以通过将长音频文件作为一系列重叠的“块”进行处理来压缩或重编。“我们在JAX中实现了这一功能,并在两种类型的GPU上进行了性能测试。在消费级GPU上,DAC-JAX在所有块大小的压缩和解压缩方面都优于原始DAC。然而,在高性能的基于集群的GPU上,DAC-JAX在小块大小的情况下性能优于原始DAC,但在大块的情况下性能较差。摘要:We present an open-source implementation of the Descript Audio Codec (DAC) using Google's JAX ecosystem of Flax, Optax, Orbax, AUX, and CLU. Our codebase enables the reuse of model weights from the original PyTorch DAC, and we confirm that the two implementations produce equivalent token sequences and decoded audio if given the same input. We provide a training and fine-tuning script which supports device parallelism, although we have only verified it using brief training runs with a small dataset. Even with limited GPU memory, the original DAC can compress or decompress a long audio file by processing it as a sequence of overlapping "chunks." We implement this feature in JAX and benchmark the performance on two types of GPUs. On a consumer-grade GPU, DAC-JAX outperforms the original DAC for compression and decompression at all chunk sizes. However, on a high-performance, cluster-based GPU, DAC-JAX outperforms the original DAC for small chunk sizes but performs worse for large chunks.【4】 Multi-speaker Text-to-speech Training with Speaker Anonymized Data作者:Wen-Chin Huang,Yi-Chiao Wu,Tomoki Toda备注:5 pages. Submitted to Signal Processing Letters. Audio sample page: this https URL摘要:扩大语音生成模型的趋势构成了训练数据中语音身份的生物识别信息泄漏的威胁,从而引发了隐私和安全问题。在本文中,我们研究训练多说话人的文本到语音(TTS)模型使用的数据进行扬声器匿名化(SA),一个过程,往往隐藏输入语音的扬声器身份,同时保持其他属性。使用两种基于信号处理的SA方法和三种基于深度神经网络的SA方法来匿名化VCTK(一种多说话者TTS数据集),其进一步用于训练端到端TTS模型VITS,以在测试阶段执行看不见的说话者TTS。我们进行了广泛的客观和主观实验来评估匿名训练数据,以及使用这些数据训练的下游TTS模型的性能。重要的是,我们发现,UTMOS,一个数据驱动的主观评级预测模型,和GVD,一个衡量语音独特性增益的指标,是下游TTS性能的良好指标。我们总结的见解,希望能帮助未来的研究人员确定多说话人TTS训练的SA系统的优点。摘要:The trend of scaling up speech generation models poses a threat of biometric information leakage of the identities of the voices in the training data, raising privacy and security concerns. In this paper, we investigate training multi-speaker text-to-speech (TTS) models using data that underwent speaker anonymization (SA), a process that tends to hide the speaker identity of the input speech while maintaining other attributes. Two signal processing-based and three deep neural network-based SA methods were used to anonymize VCTK, a multi-speaker TTS dataset, which is further used to train an end-to-end TTS model, VITS, to perform unseen speaker TTS during the testing phase. We conducted extensive objective and subjective experiments to evaluate the anonymized training data, as well as the performance of the downstream TTS model trained using those data. Importantly, we found that UTMOS, a data-driven subjective rating predictor model, and GVD, a metric that measures the gain of voice distinctiveness, are good indicators of the downstream TTS performance. We summarize insights in the hope of helping future researchers determine the goodness of the SA system for multi-speaker TTS training.
【5】 AudioSetMix: Enhancing Audio-Language Datasets with LLM-Assisted Augmentations标题:AudioSetMix:通过法学硕士辅助增强音频语言数据集摘要:近年来,多模态学习在听觉语言领域取得了重大进展。然而,与图像语言任务相比,由于数据有限且质量较低,音频语言学习面临挑战。现有的音频语言数据集非常小,并且手动标记由于需要收听整个音频片段以进行准确标记而受到阻碍。 我们的方法系统地生成音频字幕对增强音频剪辑与自然语言标签和相应的音频信号处理操作。利用一个大的语言模型,我们生成增强的音频片段的描述与提示模板。这种可扩展的方法产生AudioSetMix,这是一个用于文本和音频相关模型的高质量训练数据集。 我们的数据集的集成通过提供多样化和更好的对齐示例来提高模型在基准测试中的性能。值得注意的是,我们的数据集解决了现有数据集中缺少修饰语(形容词和副词)的问题。通过使模型能够学习这些概念,并在训练过程中生成困难的负面示例,我们在多个基准测试中实现了最先进的性能。摘要:Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing audio-language datasets are notably smaller, and manual labeling is hindered by the need to listen to entire audio clips for accurate labeling. Our method systematically generates audio-caption pairs by augmenting audio clips with natural language labels and corresponding audio signal processing operations. Leveraging a Large Language Model, we generate descriptions of augmented audio clips with a prompt template. This scalable method produces AudioSetMix, a high-quality training dataset for text-and-audio related models. Integration of our dataset improves models performance on benchmarks by providing diversified and better-aligned examples. Notably, our dataset addresses the absence of modifiers (adjectives and adverbs) in existing datasets. By enabling models to learn these concepts, and generating hard negative examples during training, we achieve state-of-the-art performance on multiple benchmarks.
【1】 SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model标题:SSAMBA:使用Mamba状态空间模型的自我监督音频表示学习作者:Siavash Shams,Sukru Samet Dindar,Xilin Jiang,Nima Mesgarani备注:Code at this https URL摘要:Transformers凭借其强大的建模功能,彻底改变了各种任务的深度学习,包括音频表示学习。然而,它们通常在GPU内存使用和计算推理时间方面遭受二次复杂性,从而影响其效率。最近,像Mamba这样的状态空间模型(SSM)已经成为一种有前途的替代方案,通过避免这些复杂性提供了一种更有效的方法。鉴于这些优势,我们探索基于SSM的模型在音频任务中的潜力。在本文中,我们介绍了自我监督音频曼巴(SSAMBA),第一个自我监督,无注意力,基于SSM的音频表示学习模型。SSAMBA利用双向Mamba有效地捕获复杂的音频模式。我们结合了一个自我监督的预训练框架,优化了区分和生成目标,使模型能够从大规模的未标记数据集中学习鲁棒的音频表示。我们评估了SSAMBA的各种任务,如音频分类,关键字定位和说话人识别。我们的研究结果表明,SSAMBA优于自我监督音频频谱图Transformer(SSAST)在大多数任务。值得注意的是,SSAMBA在批量推理速度方面比SSAST快约92.7%,并且对于输入令牌大小为22k的微小模型大小,其内存效率比SSAST高95.4%。这些效率的提高,加上卓越的性能,强调了SSAMBA的架构创新的有效性,使其成为一个令人信服的选择,为广泛的音频处理应用。摘要:Transformers have revolutionized deep learning across various tasks, including audio representation learning, due to their powerful modeling capabilities. However, they often suffer from quadratic complexity in both GPU memory usage and computational inference time, affecting their efficiency. Recently, state space models (SSMs) like Mamba have emerged as a promising alternative, offering a more efficient approach by avoiding these complexities. Given these advantages, we explore the potential of SSM-based models in audio tasks. In this paper, we introduce Self-Supervised Audio Mamba (SSAMBA), the first self-supervised, attention-free, and SSM-based model for audio representation learning. SSAMBA leverages the bidirectional Mamba to capture complex audio patterns effectively. We incorporate a self-supervised pretraining framework that optimizes both discriminative and generative objectives, enabling the model to learn robust audio representations from large-scale, unlabeled datasets. We evaluated SSAMBA on various tasks such as audio classification, keyword spotting, and speaker identification. Our results demonstrate that SSAMBA outperforms the Self-Supervised Audio Spectrogram Transformer (SSAST) in most tasks. Notably, SSAMBA is approximately 92.7% faster in batch inference speed and 95.4% more memory-efficient than SSAST for the tiny model size with an input token size of 22k. These efficiency gains, combined with superior performance, underscore the effectiveness of SSAMBA's architectural innovation, making it a compelling choice for a wide range of audio processing applications.
【2】 Source Localization by Multidimensional Steered Response Power Mapping with Sparse Bayesian Learning标题:基于稀疏Bayesian学习的多维定向响应功率映射的源定位作者:Wei-Ting Lai,Lachlan Birnie,Xingyu Chen,Amy Bastine,Thushara D. Abhayapala,Prasanga N. Samarasinghe摘要:我们提出了一个先进的转向响应功率(SRP)方法定位多个源。虽然传统的SRP在不利条件下表现良好,但它仍然在与紧密相邻的源的场景中挣扎,导致SRP地图模糊。我们通过在SRP中应用稀疏优化来解决这个问题,以获得高分辨率的地图。我们的方法将SRP映射表示为多维矩阵,以保留时频信息,并进一步提高在不利条件下的性能。我们使用多字典稀疏贝叶斯学习来定位源,而不需要事先知道它们的数量。我们验证了我们的方法,通过实际的实验与16通道平面麦克风阵列,并与其他三个SRP和基于稀疏性的方法进行比较。我们的多维SRP方法优于传统的SRP和当前国家的最先进的稀疏SRP方法定位在混响室中的紧密间隔的源。摘要:We propose an advance Steered Response Power (SRP) method for localizing multiple sources. While conventional SRP performs well in adverse conditions, it remains to struggle in scenarios with closely neighboring sources, resulting in ambiguous SRP maps. We address this issue by applying sparsity optimization in SRP to obtain high-resolution maps. Our approach represents SRP maps as multidimensional matrices to preserve time-frequency information and further improve performance in unfavorable conditions. We use multi-dictionary Sparse Bayesian Learning to localize sources without needing prior knowledge of their quantity. We validate our method through practical experiments with a 16-channel planar microphone array and compare against three other SRP and sparsity-based methods. Our multidimensional SRP approach outperforms conventional SRP and the current state-of-the-art sparse SRP methods for localizing closely spaced sources in a reverberant room.
【3】 Multi-speaker Text-to-speech Training with Speaker Anonymized Data作者:Wen-Chin Huang,Yi-Chiao Wu,Tomoki Toda备注:5 pages. Submitted to Signal Processing Letters. Audio sample page: this https URL摘要:扩大语音生成模型的趋势构成了训练数据中语音身份的生物识别信息泄漏的威胁,从而引发了隐私和安全问题。在本文中,我们研究训练多说话人的文本到语音(TTS)模型使用的数据进行扬声器匿名化(SA),一个过程,往往隐藏输入语音的扬声器身份,同时保持其他属性。使用两种基于信号处理的SA方法和三种基于深度神经网络的SA方法来匿名化VCTK(一种多说话者TTS数据集),其进一步用于训练端到端TTS模型VITS,以在测试阶段执行看不见的说话者TTS。我们进行了广泛的客观和主观实验来评估匿名训练数据,以及使用这些数据训练的下游TTS模型的性能。重要的是,我们发现,UTMOS,一个数据驱动的主观评级预测模型,和GVD,一个衡量语音独特性增益的指标,是下游TTS性能的良好指标。我们总结的见解,希望能帮助未来的研究人员确定多说话人TTS训练的SA系统的优点。摘要:The trend of scaling up speech generation models poses a threat of biometric information leakage of the identities of the voices in the training data, raising privacy and security concerns. In this paper, we investigate training multi-speaker text-to-speech (TTS) models using data that underwent speaker anonymization (SA), a process that tends to hide the speaker identity of the input speech while maintaining other attributes. Two signal processing-based and three deep neural network-based SA methods were used to anonymize VCTK, a multi-speaker TTS dataset, which is further used to train an end-to-end TTS model, VITS, to perform unseen speaker TTS during the testing phase. We conducted extensive objective and subjective experiments to evaluate the anonymized training data, as well as the performance of the downstream TTS model trained using those data. Importantly, we found that UTMOS, a data-driven subjective rating predictor model, and GVD, a metric that measures the gain of voice distinctiveness, are good indicators of the downstream TTS performance. We summarize insights in the hope of helping future researchers determine the goodness of the SA system for multi-speaker TTS training.
【4】 Speech-dependent Data Augmentation for Own Voice Reconstruction with Hearable Microphones in Noisy Environments标题:在噪音环境中使用可听麦克风重建自己的语音依赖数据增强作者:Mattes Ohlenbusch,Christian Rollwage,Simon Doclo摘要:在嘈杂的环境中使用外部麦克风和耳内麦克风,可以使助听器在闭塞的耳朵外部和内部获得自己的语音拾取。由于在两个麦克风处记录的环境噪声、以及在低频率处自身语音的放大和在入耳式麦克风处的频带限制,需要自身语音重建系统来实现通信。需要大量的自身语音信号来训练基于监督式深度学习的自身语音重建系统。训练数据可以通过用特定设备记录不同说话者的大量自己的语音信号来获得,这是昂贵的,或者通过增强可用的语音数据来获得。可以通过假设每个音素的可听麦克风之间的线性时不变相对传递函数来模拟自己的语音信号,称为自己的语音传递特性。在本文中,我们提出了数据增强技术训练自己的语音重建系统的语音依赖模型的基础上自己的语音传输特性之间的可听麦克风。所提出的技术使用少量记录的自己的语音信号来估计传输特性,然后可以用于基于单通道语音信号来模拟大量自己的语音信号。实验结果表明,所提出的语音相关的个人数据增强技术导致更好的性能相比,其他数据增强技术或相比,仅在可用的记录自己的语音信号的训练,和可用的记录信号上的额外的微调可以进一步提高性能。摘要:Own voice pickup for hearables in noisy environments benefits from using both an outer and an in-ear microphone outside and inside the occluded ear. Due to environmental noise recorded at both microphones, and amplification of the own voice at low frequencies and band-limitation at the in-ear microphone, an own voice reconstruction system is needed to enable communication. A large amount of own voice signals is required to train a supervised deep learning-based own voice reconstruction system. Training data can either be obtained by recording a large amount of own voice signals of different talkers with a specific device, which is costly, or through augmentation of available speech data. Own voice signals can be simulated by assuming a linear time-invariant relative transfer function between hearable microphones for each phoneme, referred to as own voice transfer characteristics. In this paper, we propose data augmentation techniques for training an own voice reconstruction system based on speech-dependent models of own voice transfer characteristics between hearable microphones. The proposed techniques use few recorded own voice signals to estimate transfer characteristics and can then be used to simulate a large amount of own voice signals based on single-channel speech signals. Experimental results show that the proposed speech-dependent individual data augmentation technique leads to better performance compared to other data augmentation techniques or compared to training only on the available recorded own voice signals, and additional fine-tuning on the available recorded signals can improve performance further.【5】 Exploring speech style spaces with language models: Emotional TTS without emotion labels标题:用语言模型探索言语风格空间:没有情感标签的情感TTC作者:Shreeram Suresh Chandra,Zongyang Du,Berrak Sisman备注:Accepted at Speaker Odyssey 2024摘要:许多用于情感文本到语音(E—TTS)的框架依赖于人类注释的情感标签,这些标签通常不准确且难以获得。由于情绪的主观性,学习情绪韵律隐含地提出了一个艰巨的挑战。在这项研究中,我们提出了一种新的方法,利用文本意识来获得情感风格,而不需要明确的情感标签或文本提示。我们提出了TEMOTTS,一个两阶段的框架,E—TTS的训练没有情感标签,是能够推理没有辅助输入。我们所提出的方法进行知识转移之间的语言空间学习的BERT和情感风格空间构造的全球风格令牌。我们的实验结果证明了我们提出的框架的有效性,展示了情感准确性和自然性的改善。这是第一个研究,以利用情绪的语音合成的口语内容和表达之间的情感相关性。摘要:Many frameworks for emotional text-to-speech (E-TTS) rely on human-annotated emotion labels that are often inaccurate and difficult to obtain. Learning emotional prosody implicitly presents a tough challenge due to the subjective nature of emotions. In this study, we propose a novel approach that leverages text awareness to acquire emotional styles without the need for explicit emotion labels or text prompts. We present TEMOTTS, a two-stage framework for E-TTS that is trained without emotion labels and is capable of inference without auxiliary inputs. Our proposed method performs knowledge transfer between the linguistic space learned by BERT and the emotional style space constructed by global style tokens. Our experimental results demonstrate the effectiveness of our proposed framework, showcasing improvements in emotional accuracy and naturalness. This is one of the first studies to leverage the emotional correlation between spoken content and expressive delivery for emotional TTS.【6】 AudioSetMix: Enhancing Audio-Language Datasets with LLM-Assisted Augmentations标题:AudioSetMix:通过法学硕士辅助增强音频语言数据集摘要:近年来,多模态学习在听觉语言领域取得了重大进展。然而,与图像语言任务相比,由于数据有限且质量较低,音频语言学习面临挑战。现有的音频语言数据集非常小,并且手动标记由于需要收听整个音频片段以进行准确标记而受到阻碍。 我们的方法系统地生成音频字幕对增强音频剪辑与自然语言标签和相应的音频信号处理操作。利用一个大的语言模型,我们生成增强的音频片段的描述与提示模板。这种可扩展的方法产生AudioSetMix,这是一个用于文本和音频相关模型的高质量训练数据集。 我们的数据集的集成通过提供多样化和更好的对齐示例来提高模型在基准测试中的性能。值得注意的是,我们的数据集解决了现有数据集中缺少修饰语(形容词和副词)的问题。通过使模型能够学习这些概念,并在训练过程中生成困难的负面示例,我们在多个基准测试中实现了最先进的性能。摘要:Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing audio-language datasets are notably smaller, and manual labeling is hindered by the need to listen to entire audio clips for accurate labeling. Our method systematically generates audio-caption pairs by augmenting audio clips with natural language labels and corresponding audio signal processing operations. Leveraging a Large Language Model, we generate descriptions of augmented audio clips with a prompt template. This scalable method produces AudioSetMix, a high-quality training dataset for text-and-audio related models. Integration of our dataset improves models performance on benchmarks by providing diversified and better-aligned examples. Notably, our dataset addresses the absence of modifiers (adjectives and adverbs) in existing datasets. By enabling models to learn these concepts, and generating hard negative examples during training, we achieve state-of-the-art performance on multiple benchmarks.
【7】 Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System标题:重叠语音识别的声学建模:JHU Chime-5挑战系统作者:Vimal Manohar,Szu-Jui Chen,Zhiqi Wang,Yusuke Fujita,Shinji Watanabe,Sanjeev Khudanpur摘要:本文总结了我们在约翰霍普金斯大学的语音识别系统的CHiME—5的挑战,识别高度重叠的晚宴语音记录多个麦克风阵列的声学建模工作。我们探索数据增强方法,神经网络架构,前端语音去混响,波束成形和鲁棒的i向量提取与我们的内部实现和公开可用的工具的比较。我们最终在开发集上实现了69.4%的单词错误率,这比之前的基线81.1%绝对改善了11.7%,并将此改进的基线与改进的技术/工具一起发布为高级CHiME—5配方。摘要:This paper summarizes our acoustic modeling efforts in the Johns Hopkins University speech recognition system for the CHiME-5 challenge to recognize highly-overlapped dinner party speech recorded by multiple microphone arrays. We explore data augmentation approaches, neural network architectures, front-end speech dereverberation, beamforming and robust i-vector extraction with comparisons of our in-house implementations and publicly available tools. We finally achieved a word error rate of 69.4% on the development set, which is a 11.7% absolute improvement over the previous baseline of 81.1%, and release this improved baseline with refined techniques/tools as an advanced CHiME-5 recipe.
【8】 Images that Sound: Composing Images and Sounds on a Single Canvas作者:Ziyang Chen,Daniel Geng,Andrew Owens备注:Project site: this https URL摘要:声谱图是声音的2D表示,看起来与我们视觉世界中的图像非常不同。而自然的图像,当作为声谱图播放时,会发出不自然的声音。在本文中,我们表明,它是可以合成的频谱图,同时看起来像自然的图像和声音像自然的音频。我们称这些声谱图为声音的图像。我们的方法是简单和zero-shot的,它利用了在共享潜在空间中操作的预训练的文本到图像和文本到频谱图扩散模型。在反向过程中,我们并行地使用音频和图像扩散模型对噪声潜伏期进行降噪,从而产生可能在两种模型下的样本。通过定量评估和感知研究,我们发现,我们的方法成功地生成与所需的音频提示对齐的频谱图,同时也采取了所需的图像提示的视觉外观。请参阅我们的项目页面视频结果:https://ificl.github.io/images-that-sound/摘要:Spectrograms are 2D representations of sound that look very different from the images found in our visual world. And natural images, when played as spectrograms, make unnatural sounds. In this paper, we show that it is possible to synthesize spectrograms that simultaneously look like natural images and sound like natural audio. We call these spectrograms images that sound. Our approach is simple and zero-shot, and it leverages pre-trained text-to-image and text-to-spectrogram diffusion models that operate in a shared latent space. During the reverse process, we denoise noisy latents with both the audio and image diffusion models in parallel, resulting in a sample that is likely under both models. Through quantitative evaluations and perceptual studies, we find that our method successfully generates spectrograms that align with a desired audio prompt while also taking the visual appearance of a desired image prompt. Please see our project page for video results: https://ificl.github.io/images-that-sound/
【9】 Neighborhood Attention Transformer with Progressive Channel Fusion for Speaker Verification标题:具有渐进通道融合的邻居注意力Transformer用于说话者验证备注:8 pages, 2 figures, 3 tables摘要:基于变换器的说话人验证架构通常需要比ECAPA-TDNN更多的训练数据。因此,最近的工作通常是在VoxCeleb1&2上训练的。我们提出了一个基于自我注意力的骨干网络,当单独在VoxCeleb 2上训练时,它可以达到有竞争力的结果。该网络在邻域关注和全局关注之间交替,以捕获局部和全局特征,然后聚合不同层次的特征,最后执行关注的统计池。此外,我们采用了渐进的通道融合策略,以扩大在通道维的感受野作为网络的深化。我们在VoxCeleb 2上训练了所提出的PCF-NAT模型,并在VoxCeleb 1和VoxSRC的验证集上对其进行了评估。浅PCF-NAT的EER和minDCF平均比类似大小的ECAPA-TDNN低20%以上。深度PCF-NAT在VoxCeleb 1-O上实现了低于0.5%的EER。摘要:Transformer-based architectures for speaker verification typically require more training data than ECAPA-TDNN. Therefore, recent work has generally been trained on VoxCeleb1&2. We propose a backbone network based on self-attention, which can achieve competitive results when trained on VoxCeleb2 alone. The network alternates between neighborhood attention and global attention to capture local and global features, then aggregates features of different hierarchical levels, and finally performs attentive statistical pooling. Additionally, we employ a progressive channel fusion strategy to expand the receptive field in the channel dimension as the network deepens. We trained the proposed PCF-NAT model on VoxCeleb2 and evaluated it on VoxCeleb1 and the validation sets of VoxSRC. The EER and minDCF of the shallow PCF-NAT are on average more than 20% lower than those of similarly sized ECAPA-TDNN. Deep PCF-NAT achieves an EER lower than 0.5% on VoxCeleb1-O.
【10】 DAC-JAX: A JAX Implementation of the Descript Audio Codec标题:DAC-JAX:描述音频编解码器的JAX实现备注:5 pages, 3 figures, 2 tables摘要:我们使用Google的JAX生态系统Flax,Optax,Orbax,AUX和CLU来实现Descept Audio Codec(DAC)的开源实现。我们的代码库可以重用原始PyTorch DAC的模型权重,并且我们确认,如果给定相同的输入,这两种实现会产生等效的令牌序列和解码音频。我们提供了一个支持设备并行性的训练和微调脚本,尽管我们只使用小数据集的简短训练运行对其进行了验证。即使GPU内存有限,原始的DAC也可以通过将长音频文件作为一系列重叠的“块”进行处理来压缩或重编。“我们在JAX中实现了这一功能,并在两种类型的GPU上进行了性能测试。在消费级GPU上,DAC-JAX在所有块大小的压缩和解压缩方面都优于原始DAC。然而,在高性能的基于集群的GPU上,DAC-JAX在小块大小的情况下性能优于原始DAC,但在大块的情况下性能较差。摘要:We present an open-source implementation of the Descript Audio Codec (DAC) using Google's JAX ecosystem of Flax, Optax, Orbax, AUX, and CLU. Our codebase enables the reuse of model weights from the original PyTorch DAC, and we confirm that the two implementations produce equivalent token sequences and decoded audio if given the same input. We provide a training and fine-tuning script which supports device parallelism, although we have only verified it using brief training runs with a small dataset. Even with limited GPU memory, the original DAC can compress or decompress a long audio file by processing it as a sequence of overlapping "chunks." We implement this feature in JAX and benchmark the performance on two types of GPUs. On a consumer-grade GPU, DAC-JAX outperforms the original DAC for compression and decompression at all chunk sizes. However, on a high-performance, cluster-based GPU, DAC-JAX outperforms the original DAC for small chunk sizes but performs worse for large chunks.