今日论文合集:cs.SD语音7篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
标题: Diret-MusicGen:通过指令调优解锁音乐语言模型的文本到音乐编辑
作者:Yixiao Zhang,Yukara Ikemiya,Woosung Choi,Naoki Murata,Marco A. Martínez-Ramírez,Liwei Lin,Gus Xia,Wei-Hsiang Liao,Yuki Mitsufuji,Simon Dixon
备注:Demo are available at: this https URL
链接:点击下载PDF文件
摘要:文本到音乐编辑的最新进展,它采用文本查询来修改音乐(例如,通过改变其风格或调整乐器组件),为人工智能辅助音乐创作带来了独特的挑战和机遇。在这个领域中,以前的方法受到从头开始训练特定编辑模型的必要性的限制,这是资源密集型和效率低下的;其他研究使用大型语言模型来预测编辑的音乐,导致不精确的音频重建。为了结合这些优势并解决这些限制,我们引入了Instruct-MusicGen,这是一种新颖的方法,可以微调预训练的MusicGen模型,以有效地遵循编辑指令,例如添加,删除或分离词干。我们的方法涉及到修改的原始MusicGen架构,将文本融合模块和音频融合模块,允许该模型同时处理指令文本和音频输入,并产生所需的编辑音乐。值得注意的是,Instruct-MusicGen只为原始MusicGen模型引入了8%的新参数,并且只训练了5 K步,但与现有基线相比,它在所有任务中都实现了卓越的性能,并展示了与针对特定任务训练的模型相当的性能。这一进步不仅提高了文本到音乐编辑的效率,而且拓宽了音乐语言模型在动态音乐制作环境中的适用性。摘要:Recent advances in text-to-music editing, which employ text queries to modify music (e.g. by changing its style or adjusting instrumental components), present unique challenges and opportunities for AI-assisted music creation. Previous approaches in this domain have been constrained by the necessity to train specific editing models from scratch, which is both resource-intensive and inefficient; other research uses large language models to predict edited music, resulting in imprecise audio reconstruction. To Combine the strengths and address these limitations, we introduce Instruct-MusicGen, a novel approach that finetunes a pretrained MusicGen model to efficiently follow editing instructions such as adding, removing, or separating stems. Our approach involves a modification of the original MusicGen architecture by incorporating a text fusion module and an audio fusion module, which allow the model to process instruction texts and audio inputs concurrently and yield the desired edited music. Remarkably, Instruct-MusicGen only introduces 8% new parameters to the original MusicGen model and only trains for 5K steps, yet it achieves superior performance across all tasks compared to existing baselines, and demonstrates performance comparable to the models trained for specific tasks. This advancement not only enhances the efficiency of text-to-music editing but also broadens the applicability of music language models in dynamic music production environments.

【2】 NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
标题: NeRAF:3D场景注入神经辐射和声学场
作者:Amandine Brunetto,Sascha Hornauer,Fabien Moutarde
备注:Project Page: this https URL
链接:点击下载PDF文件
摘要:声音在人类感知中起着重要作用,为理解我们的环境提供重要的场景信息以及视觉。尽管在神经内隐表征方面取得了进展,但学习与视觉场景相匹配的声学仍然具有挑战性。我们提出了NeRAF,一种联合学习声场和辐射场的方法。NeRAF被设计成一个Nerfstudio模块,用于方便地访问逼真的视听生成。它在新的位置合成新的视图和空间化音频,利用辐射场功能来调节具有3D场景信息的声场。在推理时,每种模态都可以独立地在空间上分离的位置呈现,从而提供更大的通用性。我们在SoundSpaces数据集上展示了我们方法的优势。NeRAF在数据效率更高的同时,比以前的作品实现了实质性的性能改进。此外,NeRAF通过交叉模态学习增强了用稀疏数据训练的复杂场景的新颖视图合成。摘要:Sound plays a major role in human perception, providing essential scene information alongside vision for understanding our environment. Despite progress in neural implicit representations, learning acoustics that match a visual scene is still challenging. We propose NeRAF, a method that jointly learns acoustic and radiance fields. NeRAF is designed as a Nerfstudio module for convenient access to realistic audio-visual generation. It synthesizes both novel views and spatialized audio at new positions, leveraging radiance field capabilities to condition the acoustic field with 3D scene information. At inference, each modality can be rendered independently and at spatially separated positions, providing greater versatility. We demonstrate the advantages of our method on the SoundSpaces dataset. NeRAF achieves substantial performance improvements over previous works while being more data-efficient. Furthermore, NeRAF enhances novel view synthesis of complex scenes trained with sparse data through cross-modal learning.

【3】 Practical aspects for the creation of an audio dataset from field recordings with optimized labeling budget with AI-assisted strategy
标题: 通过人工智能辅助策略通过现场录音创建音频数据集的实用方面优化了标签预算
作者:Javier Naranjo-Alcazar,Jordi Grau-Haro,Ruben Ribes-Serrano,Pedro Zuccarello
备注:Submitted to ICML 2024 Workshop on Data-Centric Machine Learning Research
链接:点击下载PDF文件
摘要:Machine Listening专注于开发从音频信号中提取相关信息的技术。这些项目的一个关键方面是获取和标记上下文数据,这本身就很复杂,需要特定的资源和策略。尽管有一些音频数据集可用,但许多不适合商业应用。本文强调了使用专家标签器的主动学习(AL)的重要性,而不是众包,后者通常缺乏对数据集结构的详细了解。人工智能是一个迭代过程,结合了人工标签和人工智能模型,通过智能选择样本供人工审查来优化标签预算。这种方法解决了处理超过可用计算资源和内存的大型、不断增长的数据集的挑战。本文提出了一个全面的以数据为中心的机器监听项目框架,详细介绍了资源受限场景下的录音节点配置、数据库结构和标签预算优化。该框架应用于西班牙瓦伦西亚的一个工业港口,在五个月的时间里,一个小团队成功地标记了6540个10秒的音频样本,证明了其有效性和对各种资源可用性情况的适应性。摘要:Machine Listening focuses on developing technologies to extract relevant information from audio signals. A critical aspect of these projects is the acquisition and labeling of contextualized data, which is inherently complex and requires specific resources and strategies. Despite the availability of some audio datasets, many are unsuitable for commercial applications. The paper emphasizes the importance of Active Learning (AL) using expert labelers over crowdsourcing, which often lacks detailed insights into dataset structures. AL is an iterative process combining human labelers and AI models to optimize the labeling budget by intelligently selecting samples for human review. This approach addresses the challenge of handling large, constantly growing datasets that exceed available computational resources and memory. The paper presents a comprehensive data-centric framework for Machine Listening projects, detailing the configuration of recording nodes, database structure, and labeling budget optimization in resource-constrained scenarios. Applied to an industrial port in Valencia, Spain, the framework successfully labeled 6540 ten-second audio samples over five months with a small team, demonstrating its effectiveness and adaptability to various resource availability situations.

【4】 Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
标题: 用于联合音频和视频生成的鉴别器引导的合作扩散
作者:Akio Hayakawa,Masato Ishii,Takashi Shibuya,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:在这项研究中,我们的目标是通过利用预训练的音频和视频单模态生成模型来构建一个具有最小计算成本的音频-视频生成模型。为了实现这一目标,我们提出了一种新的方法,指导每个单模态模型合作生成跨模态对齐良好的样本。具体来说,给定两个预先训练的基础扩散模型,我们训练一个轻量级的联合制导模块来调整由基础模型分别估计的分数,以匹配音频和视频上的联合分布的分数。我们从理论上表明,这种指导可以通过最佳的梯度来计算,区分真实的音频-视频对与由基础模型独立生成的假音频-视频对。在此基础上,通过训练该神经网络,构建了联合制导模块。此外,我们采用了一个损失函数,使梯度的噪声估计工作,在标准的扩散模型,稳定梯度的梯度。几个基准数据集的实证评估表明,我们的方法提高了单模态保真度和多模态对齐与相对较少的参数。摘要:In this study, we aim to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides each single-modal model to cooperatively generate well-aligned samples across modalities. Specifically, given two pre-trained base diffusion models, we train a lightweight joint guidance module to adjust scores separately estimated by the base models to match the score of joint distribution over audio and video. We theoretically show that this guidance can be computed through the gradient of the optimal discriminator distinguishing real audio-video pairs from fake ones independently generated by the base models. On the basis of this analysis, we construct the joint guidance module by training this discriminator. Additionally, we adopt a loss function to make the gradient of the discriminator work as a noise estimator, as in standard diffusion models, stabilizing the gradient of the discriminator. Empirical evaluations on several benchmark datasets demonstrate that our method improves both single-modal fidelity and multi-modal alignment with a relatively small number of parameters.

【5】 TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
标题: TransVIP:具有语音和等时性保留的语音到语音翻译系统
作者:Chenyang Le,Yao Qian,Dongmei Wang,Long Zhou,Shujie Liu,Xiaofei Wang,Midia Yousefi,Yanmin Qian,Jinyu Li,Sheng Zhao,Michael Zeng
备注:Work in progress
链接:点击下载PDF文件
摘要:直接将语音从一种语言翻译成另一种语言,称为端到端语音到语音翻译,这一研究的兴趣和趋势不断增加。然而,大多数端到端模型都难以超越级联模型,即,一个流水线框架,通过连接语音识别,机器翻译和文本到语音模型。主要挑战来自直接翻译任务的固有复杂性和数据的稀缺性。在这项研究中,我们引入了一个新的模型框架transVIP,利用不同的数据集在级联的方式,但有利于端到端的推理,通过联合概率。此外,我们提出了两个独立的编码器,以保持扬声器的语音特征和isochrony从源语音在翻译过程中,使其非常适合的场景,如视频配音。我们在法语-英语语言对上的实验表明,我们的模型优于当前最先进的语音到语音翻译模型。摘要:There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition, machine translation and text-to-speech models. The primary challenges stem from the inherent complexities involved in direct translation tasks and the scarcity of data. In this study, we introduce a novel model framework TransVIP that leverages diverse datasets in a cascade fashion yet facilitates end-to-end inference through joint probability. Furthermore, we propose two separated encoders to preserve the speaker's voice characteristics and isochrony from the source speech during the translation process, making it highly suitable for scenarios such as video dubbing. Our experiments on the French-English language pair demonstrate that our model outperforms the current state-of-the-art speech-to-speech translation model.

【6】 Listenable Maps for Zero-Shot Audio Classifiers
标题: Zero-Shot音频分类器的可听地图
作者:Francesco Paissan,Luca Della Libera,Mirco Ravanelli,Cem Subakan
链接:点击下载PDF文件
摘要:解释深度学习模型(包括音频分类器)的决策对于确保这项技术的透明度和可信度至关重要。在本文中,我们介绍LMAC-ZS(Listenable Maps for Audio Classifiers in the Zero-Shot Context),据我们所知,这是第一个基于解码器的事后解释方法,用于解释zero-shot音频分类器的决策。该方法利用了一种新的损失函数,最大限度地忠实于一个给定的文本和音频对之间的原始相似性。我们提供了一个广泛的评估,使用对比存储音频预训练(CLAP)模型,以展示我们的解释器仍然忠实于在zero-shot分类上下文中的决定。此外,我们定性地表明,我们的方法产生有意义的解释,以及与不同的文本提示。摘要:Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Audio Classifiers in the Zero-Shot context), which, to the best of our knowledge, is the first decoder-based post-hoc interpretation method for explaining the decisions of zero-shot audio classifiers. The proposed method utilizes a novel loss function that maximizes the faithfulness to the original similarity between a given text-and-audio pair. We provide an extensive evaluation using the Contrastive Language-Audio Pretraining (CLAP) model to showcase that our interpreter remains faithful to the decisions in a zero-shot classification context. Moreover, we qualitatively show that our method produces meaningful explanations that correlate well with different text prompts.

【7】 Discriminant audio properties in deep learning based respiratory insufficiency detection in Brazilian Portuguese
标题: 巴西葡萄牙语中基于深度学习的呼吸功能不全检测中的区分性音频属性
作者:Marcelo Matheus Gauy,Larissa Cristina Berti,Arnaldo Cândido Jr,Augusto Camargo Neto,Alfredo Goldman,Anna Sara Shafferman Levin,Marcus Martins,Beatriz Raposo de Medeiros,Marcelo Queiroz,Ester Cerdeira Sabino,Flaviane Romani Fernandes Svartman,Marcelo Finger
Journal-ref:Artificial Intellingence in Medicine Proceedings 2023, page 271-275
链接:点击下载PDF文件
摘要:这项工作研究了通过分析语音音频检测呼吸功能不全(RI)的人工智能(AI)系统,从而将语音视为RI生物标志物。之前的工作在大流行的第一阶段收集了COVID-19患者的RI数据(P1),并训练了现代AI模型,如CNN和Transformers,其准确率达到96.5%,显示了通过AI进行RI检测的可行性。在这里,我们收集了RI患者数据(P2),除了COVID-19之外还有几个原因,旨在扩展基于AI的RI检测。我们还收集了无RI的住院患者的对照数据。我们发现,当在P1上训练时,所考虑的模型不会推广到P2,这表明COVID-19 RI具有可能在所有RI类型中找不到的特征。摘要:This work investigates Artificial Intelligence (AI) systems that detect respiratory insufficiency (RI) by analyzing speech audios, thus treating speech as a RI biomarker. Previous works collected RI data (P1) from COVID-19 patients during the first phase of the pandemic and trained modern AI models, such as CNNs and Transformers, which achieved $96.5 %$ accuracy, showing the feasibility of RI detection via AI. Here, we collect RI patient data (P2) with several causes besides COVID-19, aiming at extending AI-based RI detection. We also collected control data from hospital patients without RI. We show that the considered models, when trained on P1, do not generalize to P2, indicating that COVID-19 RI has features that may not be found in all RI types.


eess.AS音频处理

【1】 Instruct-MusicGen: Unlocking Text-to-Music Editing for Music Language Models via Instruction Tuning
标题: Diret-MusicGen:通过指令调优解锁音乐语言模型的文本到音乐编辑
作者:Yixiao Zhang,Yukara Ikemiya,Woosung Choi,Naoki Murata,Marco A. Martínez-Ramírez,Liwei Lin,Gus Xia,Wei-Hsiang Liao,Yuki Mitsufuji,Simon Dixon
备注:Demo are available at: this https URL
链接:点击下载PDF文件
摘要:文本到音乐编辑的最新进展,它采用文本查询来修改音乐(例如,通过改变其风格或调整乐器组件),为人工智能辅助音乐创作带来了独特的挑战和机遇。在这个领域中,以前的方法受到从头开始训练特定编辑模型的必要性的限制,这是资源密集型和效率低下的;其他研究使用大型语言模型来预测编辑的音乐,导致不精确的音频重建。为了结合这些优势并解决这些限制,我们引入了Instruct-MusicGen,这是一种新颖的方法,可以微调预训练的MusicGen模型,以有效地遵循编辑指令,例如添加,删除或分离词干。我们的方法涉及到修改的原始MusicGen架构,将文本融合模块和音频融合模块,允许该模型同时处理指令文本和音频输入,并产生所需的编辑音乐。值得注意的是,Instruct-MusicGen只为原始MusicGen模型引入了8%的新参数,并且只训练了5 K步,但与现有基线相比,它在所有任务中都实现了卓越的性能,并展示了与针对特定任务训练的模型相当的性能。这一进步不仅提高了文本到音乐编辑的效率,而且拓宽了音乐语言模型在动态音乐制作环境中的适用性。摘要:Recent advances in text-to-music editing, which employ text queries to modify music (e.g. by changing its style or adjusting instrumental components), present unique challenges and opportunities for AI-assisted music creation. Previous approaches in this domain have been constrained by the necessity to train specific editing models from scratch, which is both resource-intensive and inefficient; other research uses large language models to predict edited music, resulting in imprecise audio reconstruction. To Combine the strengths and address these limitations, we introduce Instruct-MusicGen, a novel approach that finetunes a pretrained MusicGen model to efficiently follow editing instructions such as adding, removing, or separating stems. Our approach involves a modification of the original MusicGen architecture by incorporating a text fusion module and an audio fusion module, which allow the model to process instruction texts and audio inputs concurrently and yield the desired edited music. Remarkably, Instruct-MusicGen only introduces 8% new parameters to the original MusicGen model and only trains for 5K steps, yet it achieves superior performance across all tasks compared to existing baselines, and demonstrates performance comparable to the models trained for specific tasks. This advancement not only enhances the efficiency of text-to-music editing but also broadens the applicability of music language models in dynamic music production environments.

【2】 NeRAF: 3D Scene Infused Neural Radiance and Acoustic Fields
标题: NeRAF:3D场景注入神经辐射和声学场
作者:Amandine Brunetto,Sascha Hornauer,Fabien Moutarde
备注:Project Page: this https URL
链接:点击下载PDF文件
摘要:声音在人类感知中起着重要作用,为理解我们的环境提供重要的场景信息以及视觉。尽管在神经内隐表征方面取得了进展,但学习与视觉场景相匹配的声学仍然具有挑战性。我们提出了NeRAF,一种联合学习声场和辐射场的方法。NeRAF被设计成一个Nerfstudio模块,用于方便地访问逼真的视听生成。它在新的位置合成新的视图和空间化音频,利用辐射场功能来调节具有3D场景信息的声场。在推理时,每种模态都可以独立地在空间上分离的位置呈现,从而提供更大的通用性。我们在SoundSpaces数据集上展示了我们方法的优势。NeRAF在数据效率更高的同时,比以前的作品实现了实质性的性能改进。此外,NeRAF通过交叉模态学习增强了用稀疏数据训练的复杂场景的新颖视图合成。摘要:Sound plays a major role in human perception, providing essential scene information alongside vision for understanding our environment. Despite progress in neural implicit representations, learning acoustics that match a visual scene is still challenging. We propose NeRAF, a method that jointly learns acoustic and radiance fields. NeRAF is designed as a Nerfstudio module for convenient access to realistic audio-visual generation. It synthesizes both novel views and spatialized audio at new positions, leveraging radiance field capabilities to condition the acoustic field with 3D scene information. At inference, each modality can be rendered independently and at spatially separated positions, providing greater versatility. We demonstrate the advantages of our method on the SoundSpaces dataset. NeRAF achieves substantial performance improvements over previous works while being more data-efficient. Furthermore, NeRAF enhances novel view synthesis of complex scenes trained with sparse data through cross-modal learning.

【3】 Practical aspects for the creation of an audio dataset from field recordings with optimized labeling budget with AI-assisted strategy
标题: 通过人工智能辅助策略通过现场录音创建音频数据集的实用方面优化了标签预算
作者:Javier Naranjo-Alcazar,Jordi Grau-Haro,Ruben Ribes-Serrano,Pedro Zuccarello
备注:Submitted to ICML 2024 Workshop on Data-Centric Machine Learning Research
链接:点击下载PDF文件
摘要:Machine Listening专注于开发从音频信号中提取相关信息的技术。这些项目的一个关键方面是获取和标记上下文数据,这本身就很复杂,需要特定的资源和策略。尽管有一些音频数据集可用,但许多不适合商业应用。本文强调了使用专家标签器的主动学习(AL)的重要性,而不是众包,后者通常缺乏对数据集结构的详细了解。人工智能是一个迭代过程,结合了人工标签和人工智能模型,通过智能选择样本供人工审查来优化标签预算。这种方法解决了处理超过可用计算资源和内存的大型、不断增长的数据集的挑战。本文提出了一个全面的以数据为中心的机器监听项目框架,详细介绍了资源受限场景下的录音节点配置、数据库结构和标签预算优化。该框架应用于西班牙瓦伦西亚的一个工业港口,在五个月的时间里,一个小团队成功地标记了6540个10秒的音频样本,证明了其有效性和对各种资源可用性情况的适应性。摘要:Machine Listening focuses on developing technologies to extract relevant information from audio signals. A critical aspect of these projects is the acquisition and labeling of contextualized data, which is inherently complex and requires specific resources and strategies. Despite the availability of some audio datasets, many are unsuitable for commercial applications. The paper emphasizes the importance of Active Learning (AL) using expert labelers over crowdsourcing, which often lacks detailed insights into dataset structures. AL is an iterative process combining human labelers and AI models to optimize the labeling budget by intelligently selecting samples for human review. This approach addresses the challenge of handling large, constantly growing datasets that exceed available computational resources and memory. The paper presents a comprehensive data-centric framework for Machine Listening projects, detailing the configuration of recording nodes, database structure, and labeling budget optimization in resource-constrained scenarios. Applied to an industrial port in Valencia, Spain, the framework successfully labeled 6540 ten-second audio samples over five months with a small team, demonstrating its effectiveness and adaptability to various resource availability situations.

【4】 The Evolution of Multimodal Model Architectures
标题: 多模式模型架构的演变
作者:Shakti N. Wadekar,Abhishek Chaurasia,Aman Chadha,Eugenio Culurciello
备注:30 pages, 6 tables, 7 figures
链接:点击下载PDF文件
摘要:这项工作独特地识别和表征了当代多模态景观中四种流行的多模态模型建筑模式。通过架构类型系统地对模型进行分类有助于监控多模式领域的发展。不同于最近的调查文件,目前的一般信息多模态建筑,这项研究进行了全面的探索,建筑细节,并确定了四个具体的建筑类型。这些类型通过将多模态输入集成到深度神经网络模型中的各自方法来区分。前两种类型(A型和B型)在模型的内部层中深度融合多模态输入,而以下两种类型(C型和D型)则有助于在输入阶段进行早期融合。类型A采用标准的交叉注意,而类型B利用定制设计的层在内部层内进行模态融合。另一方面,Type-C利用特定于模态的编码器,而Type-D利用标记器在模型的输入阶段处理模态。所识别的架构类型有助于监控任意多模式模型开发。值得注意的是,Type-C和Type-D目前在构建任意多模态模型中受到青睐。Type-C以其非标记化的多模态模型架构而闻名,正在成为Type-D的可行替代方案,Type-D利用了输入标记化技术。为了帮助模型选择,这项工作突出了每个架构类型的优点和缺点,基于数据和计算要求,架构复杂性,可扩展性,简化添加模态,训练目标和任何对任何多模态生成能力。摘要:This work uniquely identifies and characterizes four prevalent multimodal model architectural patterns in the contemporary multimodal landscape. Systematically categorizing models by architecture type facilitates monitoring of developments in the multimodal domain. Distinct from recent survey papers that present general information on multimodal architectures, this research conducts a comprehensive exploration of architectural details and identifies four specific architectural types. The types are distinguished by their respective methodologies for integrating multimodal inputs into the deep neural network model. The first two types (Type A and B) deeply fuses multimodal inputs within the internal layers of the model, whereas the following two types (Type C and D) facilitate early fusion at the input stage. Type-A employs standard cross-attention, whereas Type-B utilizes custom-designed layers for modality fusion within the internal layers. On the other hand, Type-C utilizes modality-specific encoders, while Type-D leverages tokenizers to process the modalities at the model's input stage. The identified architecture types aid the monitoring of any-to-any multimodal model development. Notably, Type-C and Type-D are currently favored in the construction of any-to-any multimodal models. Type-C, distinguished by its non-tokenizing multimodal model architecture, is emerging as a viable alternative to Type-D, which utilizes input-tokenizing techniques. To assist in model selection, this work highlights the advantages and disadvantages of each architecture type based on data and compute requirements, architecture complexity, scalability, simplification of adding modalities, training objectives, and any-to-any multimodal generation capability.

【5】 Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
标题: 用于联合音频和视频生成的鉴别器引导的合作扩散
作者:Akio Hayakawa,Masato Ishii,Takashi Shibuya,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:在这项研究中,我们的目标是通过利用预训练的音频和视频单模态生成模型来构建一个具有最小计算成本的音频-视频生成模型。为了实现这一目标,我们提出了一种新的方法,指导每个单模态模型合作生成跨模态对齐良好的样本。具体来说,给定两个预先训练的基础扩散模型,我们训练一个轻量级的联合制导模块来调整由基础模型分别估计的分数,以匹配音频和视频上的联合分布的分数。我们从理论上表明,这种指导可以通过最佳的梯度来计算,区分真实的音频-视频对与由基础模型独立生成的假音频-视频对。在此基础上,通过训练该神经网络,构建了联合制导模块。此外,我们采用了一个损失函数,使梯度的噪声估计工作,在标准的扩散模型,稳定梯度的梯度。几个基准数据集的实证评估表明,我们的方法提高了单模态保真度和多模态对齐与相对较少的参数。摘要:In this study, we aim to construct an audio-video generative model with minimal computational cost by leveraging pre-trained single-modal generative models for audio and video. To achieve this, we propose a novel method that guides each single-modal model to cooperatively generate well-aligned samples across modalities. Specifically, given two pre-trained base diffusion models, we train a lightweight joint guidance module to adjust scores separately estimated by the base models to match the score of joint distribution over audio and video. We theoretically show that this guidance can be computed through the gradient of the optimal discriminator distinguishing real audio-video pairs from fake ones independently generated by the base models. On the basis of this analysis, we construct the joint guidance module by training this discriminator. Additionally, we adopt a loss function to make the gradient of the discriminator work as a noise estimator, as in standard diffusion models, stabilizing the gradient of the discriminator. Empirical evaluations on several benchmark datasets demonstrate that our method improves both single-modal fidelity and multi-modal alignment with a relatively small number of parameters.

【6】 TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
标题: TransVIP:具有语音和等时性保留的语音到语音翻译系统
作者:Chenyang Le,Yao Qian,Dongmei Wang,Long Zhou,Shujie Liu,Xiaofei Wang,Midia Yousefi,Yanmin Qian,Jinyu Li,Sheng Zhao,Michael Zeng
备注:Work in progress
链接:点击下载PDF文件
摘要:直接将语音从一种语言翻译成另一种语言,称为端到端语音到语音翻译,这一研究的兴趣和趋势不断增加。然而,大多数端到端模型都难以超越级联模型,即,一个流水线框架,通过连接语音识别,机器翻译和文本到语音模型。主要挑战来自直接翻译任务的固有复杂性和数据的稀缺性。在这项研究中,我们引入了一个新的模型框架transVIP,利用不同的数据集在级联的方式,但有利于端到端的推理,通过联合概率。此外,我们提出了两个独立的编码器,以保持扬声器的语音特征和isochrony从源语音在翻译过程中,使其非常适合的场景,如视频配音。我们在法语-英语语言对上的实验表明,我们的模型优于当前最先进的语音到语音翻译模型。摘要:There is a rising interest and trend in research towards directly translating speech from one language to another, known as end-to-end speech-to-speech translation. However, most end-to-end models struggle to outperform cascade models, i.e., a pipeline framework by concatenating speech recognition, machine translation and text-to-speech models. The primary challenges stem from the inherent complexities involved in direct translation tasks and the scarcity of data. In this study, we introduce a novel model framework TransVIP that leverages diverse datasets in a cascade fashion yet facilitates end-to-end inference through joint probability. Furthermore, we propose two separated encoders to preserve the speaker's voice characteristics and isochrony from the source speech during the translation process, making it highly suitable for scenarios such as video dubbing. Our experiments on the French-English language pair demonstrate that our model outperforms the current state-of-the-art speech-to-speech translation model.

【7】 Listenable Maps for Zero-Shot Audio Classifiers
标题: Zero-Shot音频分类器的可听地图
作者:Francesco Paissan,Luca Della Libera,Mirco Ravanelli,Cem Subakan
链接:点击下载PDF文件
摘要:解释深度学习模型(包括音频分类器)的决策对于确保这项技术的透明度和可信度至关重要。在本文中,我们介绍LMAC-ZS(Listenable Maps for Audio Classifiers in the Zero-Shot Context),据我们所知,这是第一个基于解码器的事后解释方法,用于解释zero-shot音频分类器的决策。该方法利用了一种新的损失函数,最大限度地忠实于一个给定的文本和音频对之间的原始相似性。我们提供了一个广泛的评估,使用对比存储音频预训练(CLAP)模型,以展示我们的解释器仍然忠实于在zero-shot分类上下文中的决定。此外,我们定性地表明,我们的方法产生有意义的解释,以及与不同的文本提示。摘要:Interpreting the decisions of deep learning models, including audio classifiers, is crucial for ensuring the transparency and trustworthiness of this technology. In this paper, we introduce LMAC-ZS (Listenable Maps for Audio Classifiers in the Zero-Shot context), which, to the best of our knowledge, is the first decoder-based post-hoc interpretation method for explaining the decisions of zero-shot audio classifiers. The proposed method utilizes a novel loss function that maximizes the faithfulness to the original similarity between a given text-and-audio pair. We provide an extensive evaluation using the Contrastive Language-Audio Pretraining (CLAP) model to showcase that our interpreter remains faithful to the decisions in a zero-shot classification context. Moreover, we qualitatively show that our method produces meaningful explanations that correlate well with different text prompts.

【8】 Discriminant audio properties in deep learning based respiratory insufficiency detection in Brazilian Portuguese
标题: 巴西葡萄牙语中基于深度学习的呼吸功能不全检测中的区分性音频属性
作者:Marcelo Matheus Gauy,Larissa Cristina Berti,Arnaldo Cândido Jr,Augusto Camargo Neto,Alfredo Goldman,Anna Sara Shafferman Levin,Marcus Martins,Beatriz Raposo de Medeiros,Marcelo Queiroz,Ester Cerdeira Sabino,Flaviane Romani Fernandes Svartman,Marcelo Finger
Journal-ref:Artificial Intellingence in Medicine Proceedings 2023, page 271-275
链接:点击下载PDF文件
摘要:这项工作研究了通过分析语音音频检测呼吸功能不全(RI)的人工智能(AI)系统,从而将语音视为RI生物标志物。之前的工作在大流行的第一阶段收集了COVID-19患者的RI数据(P1),并训练了现代AI模型,如CNN和Transformers,其准确率达到96.5%,显示了通过AI进行RI检测的可行性。在这里,我们收集了RI患者数据(P2),除了COVID-19之外还有几个原因,旨在扩展基于AI的RI检测。我们还收集了无RI的住院患者的对照数据。我们发现,当在P1上训练时,所考虑的模型不会推广到P2,这表明COVID-19 RI具有可能在所有RI类型中找不到的特征。摘要:This work investigates Artificial Intelligence (AI) systems that detect respiratory insufficiency (RI) by analyzing speech audios, thus treating speech as a RI biomarker. Previous works collected RI data (P1) from COVID-19 patients during the first phase of the pandemic and trained modern AI models, such as CNNs and Transformers, which achieved $96.5 %$ accuracy, showing the feasibility of RI detection via AI. Here, we collect RI patient data (P2) with several causes besides COVID-19, aiming at extending AI-based RI detection. We also collected control data from hospital patients without RI. We show that the considered models, when trained on P1, do not generalize to P2, indicating that COVID-19 RI has features that may not be found in all RI types.


机器翻译,仅供参考