今日论文合集:cs.SD语音16篇,eess.AS音频处理27篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
标题: MA-AVT:参数高效视听变形机的模式对齐
作者:Tanvir Mahmud,Shentong Mo,Yapeng Tian,Diana Marculescu
备注:Accepted in Efficient Deep Learning for Computer Vision CVPR Workshop 2024
链接:点击下载PDF文件
摘要:预训练的Vision Transformers的最新进展表明,在没有音频预训练的情况下,参数高效的视听学习是有希望的。然而,很少有研究调查的有效方法,调整多模态功能参数高效的视听Transformers。在本文中,我们提出了MA-AVT,一个新的参数高效的视听Transformer采用深模态对齐相应的多模态语义特征。具体来说,我们引入了联合单峰和多模态令牌学习,用于将两种模态与冻结的模态共享Transformer对齐。这允许模型为每个模态学习单独的表示,同时也关注它们之间的跨模态关系。此外,与之前只从单峰编码器的输出中对齐粗特征的工作不同,我们引入了分块对比学习来在整个编码阶段对齐粗到细粒度的分层特征。此外,为了从前景匹配的视听特征中抑制每个模态中的背景特征,我们引入了一个鲁棒的判别式前景挖掘方案。通过对基准AVE,VGGSound和CREMA-D数据集的广泛实验,我们实现了SOTA方法的相当大的性能改进。摘要:Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimodal features in parameter-efficient audio-visual transformers. In this paper, we propose MA-AVT, a new parameter-efficient audio-visual transformer employing deep modality alignment for corresponding multimodal semantic features. Specifically, we introduce joint unimodal and multimodal token learning for aligning the two modalities with a frozen modality-shared transformer. This allows the model to learn separate representations for each modality, while also attending to the cross-modal relationships between them. In addition, unlike prior work that only aligns coarse features from the output of unimodal encoders, we introduce blockwise contrastive learning to align coarse-to-fine-grain hierarchical features throughout the encoding phase. Furthermore, to suppress the background features in each modality from foreground matched audio-visual features, we introduce a robust discriminative foreground mining scheme. Through extensive experiments on benchmark AVE, VGGSound, and CREMA-D datasets, we achieve considerable performance improvements over SOTA methods.

【2】 TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
标题: TraceableSpeech:通过水印实现主动可追溯的文本到语音
作者:Junzuo Zhou,Jiangyan Yi,Tao Wang,Jianhua Tao,Ye Bai,Chu Yuan Zhang,Yong Ren,Zhengqi Wen
备注:acceped by interspeech 2024
链接:点击下载PDF文件
摘要:文语转换(TTS)技术的发展带来的各种威胁促使人们需要可靠地跟踪合成语音。然而,目前的方法是在生成音频后单独添加水印,这一过程会损害语音质量和水印的不可感知性。此外,这些方法在鲁棒性和灵活性方面受到限制。为了解决这些问题,我们提出了TraceableSpeech,一种新的TTS模型,直接产生水印的语音,提高水印的不可见性和语音质量。此外,我们设计了分帧的水印嵌入和提取算法,实现了对重拼接攻击的鲁棒性和时间灵活性。实验结果表明,TraceableSpeech算法在水印不可见性、语音质量和抗重拼接攻击能力等方面均优于VALL-E或HiFicodec单独使用WavMark的强基线算法。它也可以应用于各种持续时间的语音。摘要:Various threats posed by the progress in text-to-speech (TTS) have prompted the need to reliably trace synthesized speech. However, contemporary approaches to this task involve adding watermarks to the audio separately after generation, a process that hurts both speech quality and watermark imperceptibility. In addition, these approaches are limited in robustness and flexibility. To address these problems, we propose TraceableSpeech, a novel TTS model that directly generates watermarked speech, improving watermark imperceptibility and speech quality. Furthermore, We design the frame-wise imprinting and extraction of watermarks, achieving higher robustness against resplicing attacks and temporal flexibility in operation. Experimental results show that TraceableSpeech outperforms the strong baseline where VALL-E or HiFicodec individually uses WavMark in watermark imperceptibility, speech quality and resilience against resplicing attacks. It also can apply to speech of various durations.

【3】 Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
标题: 端到端ASB的扬声器平滑kNN扬声器自适应
作者:Shaojun Li,Daimeng Wei,Jiaxin Guo,ZongYao Li,Zhanglin Wu,Zhiqiang Rao,Yuanchang Luo,Xianghui He,Hao Yang
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:尽管最近在端到端自动语音识别(E2 E ASR)系统中有所改进,但是由于训练数据和测试数据之间的声音特征不匹配,特别是在有限的目标说话人自适应数据的情况下,性能可能会下降。我们提出了一种新的说话人自适应方法Speaker-Smoothed kNN,该方法利用k-最近邻(kNN)检索技术,通过在解码阶段从其预建模型中找到正确发音的令牌来提高模型输出。此外,我们利用x向量动态调整kNN插值参数的数据稀疏性问题。该方法在域内和全域环境下使用KeSpeech和MagicData语料库进行了验证。我们的方法始终执行微调,而没有相关的性能下降,在扬声器的变化。此外,在全域设置中,我们的方法实现了最先进的结果,降低了CER在单扬声器和多扬声器测试场景。摘要:Despite recent improvements in End-to-End Automatic Speech Recognition (E2E ASR) systems, the performance can degrade due to vocal characteristic mismatches between training and testing data, particularly with limited target speaker adaptation data. We propose a novel speaker adaptation approach Speaker-Smoothed kNN that leverages k-Nearest Neighbors (kNN) retrieval techniques to improve model output by finding correctly pronounced tokens from its pre-built datastore during the decoding phase. Moreover, we utilize x-vector to dynamically adjust kNN interpolation parameters for data sparsity issue. This approach was validated using KeSpeech and MagicData corpora under in-domain and all-domain settings. Our method consistently performs comparably to fine-tuning without the associated performance degradation during speaker changes. Furthermore, in the all-domain setting, our method achieves state-of-the-art results, reducing the CER in both single speaker and multi-speaker test scenarios.

【4】 PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
标题: PPPR:用于文本到音频生成的便携式插件提示细化器
作者:Shuchen Shi,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Tao Wang,Chunyu Qiang,Yi Lu,Xin Qi,Xuefei Liu,Yukun Liu,Yongwei Li,Zhiyong Wang,Xiaopeng Wang
备注:accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:文本到音频(TTA)旨在生成与给定文本描述相对应的音频,在媒体制作中起着至关重要的作用。TTA数据集中的文本描述缺乏丰富的变化和多样性,导致TTA模型在面对复杂文本时性能下降。为了解决这个问题,我们提出了一种方法,称为便携式插件提示细化,它利用丰富的知识,在大型语言模型中固有的文本描述,有效地提高TTA声学模型的鲁棒性,而不改变声学训练集。此外,还引入了模仿人工验证的Chain-of-Thought来提高音频描述的准确性,从而提高了实际应用中生成内容的准确性。实验表明,我们的方法实现了最先进的初始得分(IS)为8.72,超过AudioGen,AudioLDM和Tango。摘要:Text-to-Audio (TTA) aims to generate audio that corresponds to the given text description, playing a crucial role in media production. The text descriptions in TTA datasets lack rich variations and diversity, resulting in a drop in TTA model performance when faced with complex text. To address this issue, we propose a method called Portable Plug-in Prompt Refiner, which utilizes rich knowledge about textual descriptions inherent in large language models to effectively enhance the robustness of TTA acoustic models without altering the acoustic training set. Furthermore, a Chain-of-Thought that mimics human verification is introduced to enhance the accuracy of audio descriptions, thereby improving the accuracy of generated content in practical applications. The experiments show that our method achieves a state-of-the-art Inception Score (IS) of 8.72, surpassing AudioGen, AudioLDM and Tango.

【5】 Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis
标题: 具有音调感知的RNN-T用于普通话发音错误检测和诊断
作者:Xintong Wang,Mingqian Shi,Ye Wang
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:利用自动语音识别(ASR)的发音错误检测和诊断(MDD)系统在汉语普通话中面临两个主要挑战:1)两阶段模型在音素或声调分类阶段和MDD阶段之间产生信息间隙。2)Mandarin MDD数据集的稀缺限制了模型训练。在本文中,我们介绍了一个无状态的RNN-T模型,用于普通话MDD,利用HuBERT特征,通过音高融合块进行音高嵌入。我们的模型,仅在母语数据上训练,在非母语场景中,电话错误率提高了3%,错误接受率提高了7%,超过了最先进的基线摘要:Mispronunciation Detection and Diagnosis (MDD) systems, leveraging Automatic Speech Recognition (ASR), face two main challenges in Mandarin Chinese: 1) The two-stage models create an information gap between the phoneme or tone classification stage and the MDD stage. 2) The scarcity of Mandarin MDD datasets limits model training. In this paper, we introduce a stateless RNN-T model for Mandarin MDD, utilizing HuBERT features with pitch embedding through a Pitch Fusion Block. Our model, trained solely on native speaker data, shows a 3% improvement in Phone Error Rate and a 7% increase in False Acceptance Rate over the state-of-the-art baseline in non-native scenarios

【6】 MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
标题: MUSE:灵活的声纹接收场和多路径融合增强泰勒Transformer,用于基于U-Net的语音增强
作者:Zizhen Lin,Xiaoting Chen,Junyu Wang
链接:点击下载PDF文件
摘要:实现轻量化设计和高性能之间的平衡仍然是语音增强的一项具有挑战性的任务。在本文中,我们介绍了多路径增强泰勒(MET)Transformer为基础的U-网络语音增强(MUSE),一个轻量级的语音增强网络建立在Unet架构。我们的方法采用了一种新的多路径增强泰勒(MET)Transformer块,它集成了可变形嵌入(DE),使灵活的声纹感受野。MET Transformer被独特地设计为融合通道和空间注意力(CSA)分支,促进通道信息交换并解决泰勒-Transformer框架内的空间注意力缺陷。通过在VoiceBank+DEMAND数据集上进行的大量实验,我们证明了MUSE在显著降低训练和部署成本的同时,实现了具有竞争力的性能,仅拥有0.51 M参数。摘要:Achieving a balance between lightweight design and high performance remains a challenging task for speech enhancement. In this paper, we introduce Multi-path Enhanced Taylor (MET) Transformer based U-net for Speech Enhancement (MUSE), a lightweight speech enhancement network built upon the Unet architecture. Our approach incorporates a novel Multi-path Enhanced Taylor (MET) Transformer block, which integrates Deformable Embedding (DE) to enable flexible receptive fields for voiceprints. The MET Transformer is uniquely designed to fuse Channel and Spatial Attention (CSA) branches, facilitating channel information exchange and addressing spatial attention deficits within the Taylor-Transformer framework. Through extensive experiments conducted on the VoiceBank+DEMAND dataset, we demonstrate that MUSE achieves competitive performance while significantly reducing both training and deployment costs, boasting a mere 0.51M parameters.

【7】 To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
标题: 蒸馏还是不蒸馏?论稳健知识蒸馏的鲁棒性
作者:Abdul Waheed,Karima Kadaoui,Muhammad Abdul-Mageed
备注:Accepted at ACL'24 main
链接:点击下载PDF文件
摘要:众所周知,阿拉伯语对自动语音识别(ASR)提出了独特的挑战。一方面,其丰富的语言多样性和广泛的方言使发展强大的包容性模型变得复杂。另一方面,目前的多语言ASR模型是计算密集型的,缺乏适当的全面评估。鉴于这些挑战,我们将知识从大型教师模型中提取到更小的学生变量中,这些变量更有效。我们还介绍了一个新的人类注释数据集,涵盖五个代表性不足的阿拉伯方言进行评估。我们进一步评估我们的模型和现有的SoTA多语言模型的标准可用基准和我们的新方言数据。我们最好的蒸馏模型的整体性能(45.0 $ % WER)超过了SoTA模型的两倍大小(无障碍M4 T-large-v2,WER= 47.0 $ %)及其教师模型(Whisper-large-v2,WER= 55.1 $ %),其在我们新方言数据上的平均性能(56.9 $ % WER)优于所有其他模型。为了更深入地了解这些模型对方言数据的不良表现,我们进行了错误分析,并报告了不同模型倾向于犯的主要错误类型。该项目的GitHub存储库位于 url{https: github.com UBC-NLP UBC-whisper-ar}。摘要:Arabic is known to present unique challenges for Automatic Speech Recognition (ASR). On one hand, its rich linguistic diversity and wide range of dialects complicate the development of robust, inclusive models. On the other, current multilingual ASR models are compute-intensive and lack proper comprehensive evaluations. In light of these challenges, we distill knowledge from large teacher models into smaller student variants that are more efficient. We also introduce a novel human-annotated dataset covering five under-represented Arabic dialects for evaluation. We further evaluate both our models and existing SoTA multilingual models on both standard available benchmarks and our new dialectal data. Our best-distilled model's overall performance ($45.0$ % WER) surpasses that of a SoTA model twice its size (SeamlessM4T-large-v2, WER=$47.0$ %) and its teacher model (Whisper-large-v2, WER=$55.1$ %), and its average performance on our new dialectal data ($56.9$ % WER) outperforms all other models. To gain more insight into the poor performance of these models on dialectal data, we conduct an error analysis and report the main types of errors the different models tend to make. The GitHub repository for the project is available at url{https: github.com UBC-NLP distill-whisper-ar}.

【8】 Prompt-guided Precise Audio Editing with Diffusion Models
标题: 使用扩散模型的预算引导精确音频编辑
作者:Manjie Xu,Chenxing Li,Duzhen zhang,Dan Su,Wei Liang,Dong Yu
备注:Accepted by ICML 2024
链接:点击下载PDF文件
摘要:音频编辑涉及通过精确控制对音频内容的任意操作。虽然文本引导扩散模型在文本到音频生成方面取得了显着的进步,但它们仍然面临着寻找灵活和精确的方式来修改音轨内的目标事件的挑战。我们提出了一种新的方法,称为PPAE,它作为一个通用的扩散模型模块,并实现精确的音频编辑。编辑仅基于输入的文本提示,并且完全不需要训练。我们利用扩散模型的交叉注意图来促进准确的局部编辑,并采用分层的局部-全局管道来确保更平滑的编辑过程。实验结果突出了我们的方法在各种编辑任务的有效性。摘要:Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as PPAE, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.

【9】 Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
标题: 端到端综合分析背景下的可区分时变线性预测
作者:Chin-Yun Yu,György Fazekas
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:在现代深度学习框架中,端到端训练线性预测(LP)算子进行音频合成是缓慢的,这是由于其递归公式化。此外,逐帧近似作为一种加速方法不能很好地推广到LP是逐样本计算的测试时间条件。用于端到端训练的高效可微样本LP是消除这一障碍的关键。我们概括了有效的时不变LP实现从GOLF声码器随时间变化的情况下。将其与经典的源滤波器模型相结合,我们表明改进的GOLF学习LP系数并比逐帧同行更好地重建语音。此外,在我们的听力测试中,GOLF的合成输出在质量评级上比最先进的可区分WORLD声码器得分更高。摘要:Training the linear prediction (LP) operator end-to-end for audio synthesis in modern deep learning frameworks is slow due to its recursive formulation. In addition, frame-wise approximation as an acceleration method cannot generalise well to test time conditions where the LP is computed sample-wise. Efficient differentiable sample-wise LP for end-to-end training is the key to removing this barrier. We generalise the efficient time-invariant LP implementation from the GOLF vocoder to time-varying cases. Combining this with the classic source-filter model, we show that the improved GOLF learns LP coefficients and reconstructs the voice better than its frame-wise counterparts. Moreover, in our listening test, synthesised outputs from GOLF scored higher in quality ratings than the state-of-the-art differentiable WORLD vocoder.

【10】 XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
标题: XTTC:一种大规模多语言Zero-Shot文本到语音模型
作者:Edresson Casanova,Kelly Davis,Eren Gölge,Görkem Göknar,Iulian Gulea,Logan Hart,Aya Aljafari,Joshua Meyer,Reuben Morais,Samuel Olayemi,Julian Weber
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:大多数Zero-shot多扬声器TTS(Zero-shot Multi-speaker TTS)系统只支持一种语言。虽然像YourTTS、VALL-E X、Mega-TTS 2和Voicebox这样的模型探索了多语言TTS,但它们仅限于少数高 中资源语言,限制了这些模型在大多数低 中资源语言中的应用。在本文中,我们的目标是缓解这个问题,提出并公开提供XTTS系统。我们的方法建立在Toronto模型的基础上,并添加了几个新的修改,以实现多语言训练,改进语音克隆,并实现更快的训练和推理。XTTS接受了16种语言的培训,并在其中大多数语言中取得了最先进的成果。摘要:Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they are limited to just a few high medium resource languages, limiting the applications of these models in most of the low medium resource languages. In this paper, we aim to alleviate this issue by proposing and making publicly available the XTTS system. Our method builds upon the Tortoise model and adds several novel modifications to enable multilingual training, improve voice cloning, and enable faster training and inference. XTTS was trained in 16 languages and achieved state-of-the-art (SOTA) results in most of them.

【11】 URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
标题: 紧急挑战:语音增强的通用性、鲁棒性和通用性
作者:Wangyou Zhang,Robin Scheibler,Kohei Saijo,Samuele Cornell,Chenda Li,Zhaoheng Ni,Anurag Kumar,Jan Pirklbauer,Marvin Sach,Shinji Watanabe,Tim Fingscheidt,Yanmin Qian
备注:6 pages, 3 figures, 3 tables. Accepted by Interspeech 2024. An extended version of the accepted manuscript with appendix
链接:点击下载PDF文件
摘要:在过去的十年中,基于深度学习的语音增强(SE)取得了重大进展。然而,大多数现有的SE研究的覆盖面SE子任务,数据的多样性和数量,以及评估指标的局限性。为了填补这一空白,促进研究走向普遍的SE,我们建立了一个新的SE挑战,名为紧急,专注于SE的普遍性,鲁棒性和概括性。我们的目标是扩展SE定义,以涵盖不同的子任务,以探索SE模型的限制,从去噪,去混响,带宽扩展和去唇。提出了一种新的框架,统一所有这些子任务在一个单一的模型,允许使用所有现有的SE方法。我们收集了来自不同领域的公共语音和噪声数据,以构建多样化的评估数据。最后,我们讨论了从我们的初步基线实验中获得的见解,这些实验基于生成和判别SE方法,具有12个策划指标。摘要:The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To fill this gap and promote research toward universal SE, we establish a new SE challenge, named URGENT, to focus on the universality, robustness, and generalizability of SE. We aim to extend the SE definition to cover different sub-tasks to explore the limits of SE models, starting from denoising, dereverberation, bandwidth extension, and declipping. A novel framework is proposed to unify all these sub-tasks in a single model, allowing the use of all existing SE approaches. We collected public speech and noise data from different domains to construct diverse evaluation data. Finally, we discuss the insights gained from our preliminary baseline experiments based on both generative and discriminative SE methods with 12 curated metrics.

【12】 What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models
标题: MLLM听到了什么?在多模式大型语言模型中检查文本和声音成分的推理
作者:Enis Berk Çoban,Michael I. Mandel,Johanna Devaney
备注:9 pages
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已展示出非凡的推理能力,特别是在连接思想和遵守逻辑规则来解决问题方面。这些模型已经发展到适应各种数据形态,包括声音和图像,称为多模式LLM(MLLM),能够描述图像或录音。之前的工作已经证明,当MLLM中的LLM组件被冻结时,音频或视觉编码器用于为声音或图像输入添加字幕,从而促进使用LLM组件进行基于文本的推理。我们有兴趣使用LLM的推理能力来促进分类。在本文中,我们通过一个字幕 分类实验表明,音频MLLM不能充分利用其LLM的基于文本的推理时,生成音频字幕。我们还考虑了这可能是由于MLLM分别表示听觉和文本信息,使得它切断了从LLM到音频编码器的推理路径。摘要:Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, notably in connecting ideas and adhering to logical rules to solve problems. These models have evolved to accommodate various data modalities, including sound and images, known as multimodal LLMs (MLLMs), which are capable of describing images or sound recordings. Previous work has demonstrated that when the LLM component in MLLMs is frozen, the audio or visual encoder serves to caption the sound or image input facilitating text-based reasoning with the LLM component. We are interested in using the LLM's reasoning capabilities in order to facilitate classification. In this paper, we demonstrate through a captioning classification experiment that an audio MLLM cannot fully leverage its LLM's text-based reasoning when generating audio captions. We also consider how this may be due to MLLMs separately representing auditory and textual information such that it severs the reasoning pathway from the LLM to the audio encoder.

【13】 Neural Codec-based Adversarial Sample Detection for Speaker Verification
标题: 基于神经编解码器的对抗样本检测用于说话人验证
作者:Xuanjun Chen,Jiawei Du,Haibin Wu,Jyh-Shing Roger Jang,Hung-yi Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自动说话人验证(ASV)越来越多地用于安全关键型应用程序,面临着不断上升的对抗性攻击的漏洞,几乎没有有效的防御措施。本文提出了一种基于神经编解码器的ASV对抗样本检测方法。该方法利用编解码器的能力,以丢弃冗余的扰动和保留必要的信息。具体来说,我们通过比较原始音频和重新合成音频之间的ASV得分差异(通过编解码器模型)来区分真实样本和对抗样本。这项全面的研究探索了所有开源神经编解码器及其变体模型。描述音频编解码器模型在15种神经编解码器中具有最高的检测率,并超过了7种现有的最先进的(SOTA)检测方法。请注意,我们的单模型方法甚至比SOTA集成方法性能更好。摘要:Automatic Speaker Verification (ASV), increasingly used in security-critical applications, faces vulnerabilities from rising adversarial attacks, with few effective defenses available. In this paper, we propose a neural codec-based adversarial sample detection method for ASV. The approach leverages the codec's ability to discard redundant perturbations and retain essential information. Specifically, we distinguish between genuine and adversarial samples by comparing ASV score differences between original and re-synthesized audio (by codec models). This comprehensive study explores all open-source neural codecs and their variant models for experiments. The Descript-audio-codec model stands out by delivering the highest detection rate among 15 neural codecs and surpassing seven prior state-of-the-art (SOTA) detection methods. Note that, our single-model method even outperforms a SOTA ensemble method by a large margin.

【14】 Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
标题: Small-E:具有线性注意力的小型语言模型,用于高效语音合成
作者:Théodor Lemerle,Nicolas Obin,Axel Roebel
备注:Interspeech
链接:点击下载PDF文件
摘要:由语言模型驱动的文本到语音(TTS)的最新进展展示了在实现自然度和zero-shot语音克隆方面的卓越能力。值得注意的是,解码器专用的Transformer是该领域的主要架构。然而,Transformers面临着序列长度的二次复杂性带来的挑战,阻碍了长序列和资源受限硬件的训练。此外,他们缺乏具体的归纳偏见方面的单调性质的TTS对齐。作为回应,我们建议用新兴的循环架构取代Transformers,并引入专门的交叉注意机制来减少重复和跳过问题。因此,我们的架构可以有效地训练长样本,并实现最先进的zero-shot语音克隆对基线的可比大小。摘要:Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning. Notably, the decoder-only transformer is the prominent architecture in this domain. However, transformers face challenges stemming from their quadratic complexity in sequence length, impeding training on lengthy sequences and resource-constrained hardware. Moreover they lack specific inductive bias with regards to the monotonic nature of TTS alignments. In response, we propose to replace transformers with emerging recurrent architectures and introduce specialized cross-attention mechanisms for reducing repeating and skipping issues. Consequently our architecture can be efficiently trained on long samples and achieve state-of-the-art zero-shot voice cloning against baselines of comparable size.

【15】 InaGVAD : a Challenging French TV and Radio Corpus Annotated for Speech Activity Detection and Speaker Gender Segmentation
标题: InaGVAR:一个令人惊叹的法国电视和广播数据库,用于语音活动检测和说话者性别分割
作者:David Doukhan,Christine Maertens,William Le Personnic,Ludovic Speroni,Reda Dehak
Journal-ref:Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 8963-8974, Torino, Italia. ELRA and ICCL
链接:点击下载PDF文件
摘要:InaGVAD是从10个法国广播和18个电视频道收集的音频语料库,分为4组:通才广播,音乐广播,新闻电视和通才电视。它包含277个1分钟长的注释录音,旨在代表法国视听节目的声音多样性,主要是为了建立能够监测媒体中男子和妇女发言时间的系统。inaGVAD提供有语音活动检测(VAD)和扬声器性别分割(SGS)注释,其扩展有重叠、扬声器特征(性别、年龄、语音质量)和10个非语音事件类别。每个通道类别的注释分布都有详细说明。该数据集被划分为1h开发和3 h37测试子集,允许公平和可重现的系统评估。一个基准的6个免费提供的VAD软件,显示不同的能力的基础上,渠道和非语音事件类别。两个现有的SGS系统的语料库进行评估,并与基线的X向量迁移学习策略,训练的发展子集。结果表明,我们的建议,在一个单一的,但不同的数据小时训练,取得了有竞争力的SGS结果。整个inaGVAD包;包括语料库、注释、评估脚本和基线训练代码;可以免费访问,促进该领域的未来发展。摘要:InaGVAD is an audio corpus collected from 10 French radio and 18 TV channels categorized into 4 groups: generalist radio, music radio, news TV, and generalist TV. It contains 277 1-minute-long annotated recordings aimed at representing the acoustic diversity of French audiovisual programs and was primarily designed to build systems able to monitor men's and women's speaking time in media. inaGVAD is provided with Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) annotations extended with overlap, speaker traits (gender, age, voice quality), and 10 non-speech event categories. Annotation distributions are detailed for each channel category. This dataset is partitioned into a 1h development and a 3h37 test subset, allowing fair and reproducible system evaluation. A benchmark of 6 freely available VAD software is presented, showing diverse abilities based on channel and non-speech event categories. Two existing SGS systems are evaluated on the corpus and compared against a baseline X-vector transfer learning strategy, trained on the development subset. Results demonstrate that our proposal, trained on a single - but diverse - hour of data, achieved competitive SGS results. The entire inaGVAD package; including corpus, annotations, evaluation scripts, and baseline training code; is made freely accessible, fostering future advancement in the domain.

【16】 Introducing the Brand New QiandaoEar22 Dataset for Specific Ship Identification Using Ship-Radiated Noise
标题: 推出全新QiandaoEar22数据集,用于利用船舶辐射噪音识别特定船舶
作者:Xiaoyang Du,Feng Hong
链接:点击下载PDF文件
摘要:舰船辐射噪声的目标识别是水下目标识别的一个重要研究领域。然而,目前缺乏多目标船舶数据集,准确地代表现实世界的水声条件。为了解决这一问题,我们发布了QiandaoEar 22 textmdash一个水声多目标数据集,可以在https: ieee-dataport.org documents qiandaoear22上下载。该数据集包含9小时28分钟的真实船舶辐射噪声数据和21小时58分钟的背景噪声数据。我们通过从多个目标中识别特定船舶的实验,验证了QiandaoEar 22的可用性。以不同的特征作为输入,六个深度学习网络作为分类器,我们评估了不同方法的基线性能。实验结果表明,将UUV的特定目标与其他目标区分开来,识别率最高可达97.78%,并且发现采用光谱和MFCC作为特征输入,DenseNet作为分类器可以获得更好的识别性能。我们的工作不仅为数据集建立了基准,而且有助于进一步开发水下声目标检测(UATD)和水下声目标识别(UATR)任务的创新方法。摘要:Target identification of ship-radiated noise is a crucial area in underwater target recognition. However, there is currently a lack of multi-target ship datasets that accurately represent real-world underwater acoustic conditions. To ntackle this issue, we release QiandaoEar22 textemdash an underwater acoustic multi-target dataset, which can be download on https: ieee-dataport.org documents qiandaoear22. This dataset encompasses 9 hours and 28 minutes of real-world ship-radiated noise data and 21 hours and 58 minutes of background noise data. We demonstrate the availability of QiandaoEar22 by conducting an experiment of identifying specific ship from the multiple targets. Taking different features as the input and six deep learning networks as classifier, we evaluate the baseline performance of different methods. The experimental results reveal that identifying the specific target of UUV from others can achieve the optimal recognition accuracy of 97.78 %, and we find using spectrum and MFCC as feature inputs and DenseNet as the classifier can achieve better recognition performance. Our work not only establishes a benchmark for the dataset but helps the further development of innovative methods for the tasks of underwater acoustic target detection (UATD) and underwater acoustic target recognition(UATR).


eess.AS音频处理
【1】 Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
标题: 端到端综合分析背景下的可区分时变线性预测
作者:Chin-Yun Yu,György Fazekas
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:在现代深度学习框架中,端到端训练线性预测(LP)算子进行音频合成是缓慢的,这是由于其递归公式化。此外,逐帧近似作为一种加速方法不能很好地推广到LP是逐样本计算的测试时间条件。用于端到端训练的高效可微样本LP是消除这一障碍的关键。我们概括了有效的时不变LP实现从GOLF声码器随时间变化的情况下。结合这与经典的源滤波器模型,我们表明,改进的GOLF学习LP系数和重建的声音比它的帧明智的同行。此外,在我们的听力测试中,GOLF的合成输出在质量评级上比最先进的可区分WORLD声码器得分更高。摘要:Training the linear prediction (LP) operator end-to-end for audio synthesis in modern deep learning frameworks is slow due to its recursive formulation. In addition, frame-wise approximation as an acceleration method cannot generalise well to test time conditions where the LP is computed sample-wise. Efficient differentiable sample-wise LP for end-to-end training is the key to removing this barrier. We generalise the efficient time-invariant LP implementation from the GOLF vocoder to time-varying cases. Combining this with the classic source-filter model, we show that the improved GOLF learns LP coefficients and reconstructs the voice better than its frame-wise counterparts. Moreover, in our listening test, synthesised outputs from GOLF scored higher in quality ratings than the state-of-the-art differentiable WORLD vocoder.

【2】 Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
标题: 性别偏见:言语情感识别社会偏见的大规模评估
作者:Yi-Cheng Lin,Haibin Wu,Huang-Cheng Chou,Chi-Chun Lee,Hung-yi Lee
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音情感识别(SER)的快速增长具有多种全球应用,从改善人机交互到帮助心理健康诊断。然而,SER模型可能包含对性别的社会偏见,导致不公平的结果。本研究分析了大规模自我监督学习(SSL)训练的SER模型中的性别偏见,探讨了影响性别偏见的因素。我们的研究开创性地从上游模型和数据角度研究SER中的性别偏见。我们的研究结果表明,女性表现出略高于男性的整体SER性能。修正CPC和XLS-R两个著名的SSL模型,显着表现出显着的偏差。此外,用普通话数据集训练的模型显示出明显的偏向效价。最后,我们发现训练数据中的性别情感分布差异显著影响性别偏见,而上游模型表示的影响有限。摘要:The rapid growth of Speech Emotion Recognition (SER) has diverse global applications, from improving human-computer interactions to aiding mental health diagnostics. However, SER models might contain social bias toward gender, leading to unfair outcomes. This study analyzes gender bias in SER models trained with Self-Supervised Learning (SSL) at scale, exploring factors influencing it. SSL-based SER models are chosen for their cutting-edge performance. Our research pioneering research gender bias in SER from both upstream model and data perspectives. Our findings reveal that females exhibit slightly higher overall SER performance than males. Modified CPC and XLS-R, two well-known SSL models, notably exhibit significant bias. Moreover, models trained with Mandarin datasets display a pronounced bias toward valence. Lastly, we find that gender-wise emotion distribution differences in training data significantly affect gender bias, while upstream model representation has a limited impact.

【3】 On the social bias of speech self-supervised models
标题: 言语自我监督模型的社会偏见
作者:Yi-Cheng Lin,Tzu-Quan Lin,Hsi-Che Lin,Andy T. Liu,Hung-yi Lee
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:自监督学习(SSL)语音模型在各种任务中取得了显着的性能,但有偏见的结果,特别是影响边缘化群体,引起了重大关注。社会偏见是指算法可能放大用于训练的数据中存在的社会群体之间的不同属性的现象。SSL模型中的偏见可以通过自动化歧视模式和强化不公平的系统来延续不公正。这项工作表明,流行的SSL模型无意中获得偏见的协会。我们探讨了各种因素,如模型架构,大小和训练方法,如何影响这些模型中的社会偏见的传播。最后,我们探索了通过正则化技术,特别是通过模型压缩来消除SSL模型偏置的有效性。我们的研究结果表明,采用行修剪和训练更宽,更浅的模型等技术可以有效地减轻SSL模型中的社会偏见。摘要:Self-supervised learning (SSL) speech models have achieved remarkable performance in various tasks, yet the biased outcomes, especially affecting marginalized groups, raise significant concerns. Social bias refers to the phenomenon where algorithms potentially amplify disparate properties between social groups present in the data used for training. Bias in SSL models can perpetuate injustice by automating discriminatory patterns and reinforcing inequitable systems. This work reveals that prevalent SSL models inadvertently acquire biased associations. We probe how various factors, such as model architecture, size, and training methodologies, influence the propagation of social bias within these models. Finally, we explore the efficacy of debiasing SSL models through regularization techniques, specifically via model compression. Our findings reveal that employing techniques such as row-pruning and training wider, shallower models can effectively mitigate social bias within SSL model.

【4】 The Database and Benchmark for Source Speaker Verification Against Voice Conversion
标题: 针对语音转换的源说话者验证的数据库和基准
作者:Ze Li,Yuke Lin,Tian Yao,Hongbin Suo,Ming Li
链接:点击下载PDF文件
摘要:语音转换系统可以将音频转换为模仿另一个说话者的语音,从而攻击说话者验证系统。然而,正在进行的研究,对源语者确认受到有限的数据和方法的限制。本文基于MFA-Conformer架构,建立了一个大规模的转换语音库,并训练了一批基线系统,以推进源说话人确认任务。此外,我们还介绍了一个相关的任务,称为转换方法识别。采用基于自适应器的多任务学习方法,在不影响源说话人确认性能的前提下实现有效的转换方法识别。此外,我们调查和有效地解决开集转换方法识别问题,通过实施一个开集最近邻方法。摘要:Voice conversion systems can transform audio to mimic another speaker's voice, thereby attacking speaker verification systems. However, ongoing studies on source speaker verification are hindered by limited data availability and methodological constraints. In this paper, we generate a large-scale converted speech database and train a batch of baseline systems based on the MFA-Conformer architecture to promote the source speaker verification task. In addition, we introduce a related task called conversion method recognition. An adapter-based multi-task learning approach is employed to achieve effective conversion method recognition without compromising source speaker verification performance. Additionally, we investigate and effectively address the open-set conversion method recognition problem through the implementation of an open-set nearest neighbor approach.

【5】 LLM-based speaker diarization correction: A generalizable approach
标题: 基于LLM的说话者日记化纠正:一种可推广的方法
作者:Georgios Efstathiadis,Vijay Yadav,Anzar Abbas
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:说话人日志化是口译使用自动语音识别(ASR)工具转录的对话所必需的。尽管在日记化方法上有了重大的发展,但日记化的准确性仍然是一个问题。在这里,我们研究了使用大型语言模型(LLM)作为后处理步骤进行日志化校正。LLM使用Fisher语料库进行了微调,Fisher语料库是一个大型的转录对话数据集。测量了模型在保持数据集中提高日志化准确性的能力。我们报告说,微调LLM可以显着提高diarization准确性。然而,模型性能受限于使用与用于微调的转录本相同的ASR工具产生的转录本,从而限制了可推广性。为了解决这一限制,通过组合来自三个单独模型的权重来开发集成模型,每个模型使用来自不同ASR工具的转录本进行微调。集成模型表现出更好的整体性能比每个ASR特定的模型,这表明一个可推广的和ASR不可知的方法是可以实现的。我们希望通过面向公众的API来访问这些模型,以供第三方应用程序使用。摘要:Speaker diarization is necessary for interpreting conversations transcribed using automated speech recognition (ASR) tools. Despite significant developments in diarization methods, diarization accuracy remains an issue. Here, we investigate the use of large language models (LLMs) for diarization correction as a post-processing step. LLMs were fine-tuned using the Fisher corpus, a large dataset of transcribed conversations. The ability of the models to improve diarization accuracy in a holdout dataset was measured. We report that fine-tuned LLMs can markedly improve diarization accuracy. However, model performance is constrained to transcripts produced using the same ASR tool as the transcripts used for fine-tuning, limiting generalizability. To address this constraint, an ensemble model was developed by combining weights from three separate models, each fine-tuned using transcripts from a different ASR tool. The ensemble model demonstrated better overall performance than each of the ASR-specific models, suggesting that a generalizable and ASR-agnostic approach may be achievable. We hope to make these models accessible through public-facing APIs for use by third-party applications.

【6】 XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model
标题: XTTC:一种大规模多语言Zero-Shot文本到语音模型
作者:Edresson Casanova,Kelly Davis,Eren Gölge,Görkem Göknar,Iulian Gulea,Logan Hart,Aya Aljafari,Joshua Meyer,Reuben Morais,Samuel Olayemi,Julian Weber
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:大多数Zero-shot多扬声器TTS(Zero-shot Multi-speaker TTS)系统只支持一种语言。虽然像YourTTS、VALL-E X、Mega-TTS 2和Voicebox这样的模型探索了多语言TTS,但它们仅限于少数高 中资源语言,限制了这些模型在大多数低 中资源语言中的应用。在本文中,我们的目标是缓解这个问题,提出并公开提供XTTS系统。我们的方法建立在Toronto模型的基础上,并添加了几个新的修改,以实现多语言训练,改进语音克隆,并实现更快的训练和推理。XTTS接受了16种语言的培训,并在其中大多数语言中取得了最先进的成果。摘要:Most Zero-shot Multi-speaker TTS (ZS-TTS) systems support only a single language. Although models like YourTTS, VALL-E X, Mega-TTS 2, and Voicebox explored Multilingual ZS-TTS they are limited to just a few high medium resource languages, limiting the applications of these models in most of the low medium resource languages. In this paper, we aim to alleviate this issue by proposing and making publicly available the XTTS system. Our method builds upon the Tortoise model and adds several novel modifications to enable multilingual training, improve voice cloning, and enable faster training and inference. XTTS was trained in 16 languages and achieved state-of-the-art (SOTA) results in most of them.

【7】 URGENT Challenge: Universality, Robustness, and Generalizability For Speech Enhancement
标题: 紧急挑战:语音增强的通用性、鲁棒性和通用性
作者:Wangyou Zhang,Robin Scheibler,Kohei Saijo,Samuele Cornell,Chenda Li,Zhaoheng Ni,Anurag Kumar,Jan Pirklbauer,Marvin Sach,Shinji Watanabe,Tim Fingscheidt,Yanmin Qian
备注:6 pages, 3 figures, 3 tables. Accepted by Interspeech 2024. An extended version of the accepted manuscript with appendix
链接:点击下载PDF文件
摘要:在过去的十年中,基于深度学习的语音增强(SE)取得了重大进展。然而,大多数现有的SE研究的覆盖面SE子任务,数据的多样性和数量,以及评估指标的局限性。为了填补这一空白,促进研究走向普遍的SE,我们建立了一个新的SE挑战,名为紧急,专注于SE的普遍性,鲁棒性和概括性。我们的目标是扩展SE定义,以涵盖不同的子任务,以探索SE模型的限制,从去噪,去混响,带宽扩展和去唇。提出了一种新的框架,统一所有这些子任务在一个单一的模型,允许使用所有现有的SE方法。我们收集了来自不同领域的公共语音和噪声数据,以构建多样化的评估数据。最后,我们讨论了从我们的初步基线实验中获得的见解,这些实验基于生成和判别SE方法,具有12个策划指标。摘要:The last decade has witnessed significant advancements in deep learning-based speech enhancement (SE). However, most existing SE research has limitations on the coverage of SE sub-tasks, data diversity and amount, and evaluation metrics. To fill this gap and promote research toward universal SE, we establish a new SE challenge, named URGENT, to focus on the universality, robustness, and generalizability of SE. We aim to extend the SE definition to cover different sub-tasks to explore the limits of SE models, starting from denoising, dereverberation, bandwidth extension, and declipping. A novel framework is proposed to unify all these sub-tasks in a single model, allowing the use of all existing SE approaches. We collected public speech and noise data from different domains to construct diverse evaluation data. Finally, we discuss the insights gained from our preliminary baseline experiments based on both generative and discriminative SE methods with 12 curated metrics.

【8】 Boosting Diffusion Model for Spectrogram Up-sampling in Text-to-speech: An Empirical Study
标题: 文本到语音中频谱图上采样的增强扩散模型:一项实证研究
作者:Chong Zhang,Yanqing Liu,Yang Zheng,Sheng Zhao
链接:点击下载PDF文件
摘要:基于自回归语言模型(LM)的文本到语音转换(TTS)技术,通过将波形量化为离散的语音符号,在获取人类语音的多样性和表现力方面取得了很大的进展,但离散语音符号的语音重建质量远不能满足要求,这取决于压缩后的语音符号压缩比。用分数匹配损失训练的生成扩散模型和用流匹配损失训练的连续归一化流在图像和语音的生成中已经变得突出。基于LM的TTS系统通常将语音分解为离散的令牌并自回归生成这些令牌,最后使用扩散模型将粗粒度的语音令牌上采样为细粒度的编解码器特征或梅尔频谱图,然后用声码器重建成波形,这具有高延迟并且对于实时语音应用是不现实的。本文系统地研究了上采样阶段的各种扩散模型,这是LM和基于扩散的体系结构的流合成的主要瓶颈,我们给出了模型体系结构,客观和主观度量来显示质量和效率的提高。摘要:Scaling text-to-speech (TTS) with autoregressive language model (LM) to large-scale datasets by quantizing waveform into discrete speech tokens is making great progress to capture the diversity and expressiveness in human speech, but the speech reconstruction quality from discrete speech token is far from satisfaction depending on the compressed speech token compression ratio. Generative diffusion models trained with score-matching loss and continuous normalized flow trained with flow-matching loss have become prominent in generation of images as well as speech. LM based TTS systems usually quantize speech into discrete tokens and generate these tokens autoregressively, and finally use a diffusion model to up sample coarse-grained speech tokens into fine-grained codec features or mel-spectrograms before reconstructing into waveforms with vocoder, which has a high latency and is not realistic for real time speech applications. In this paper, we systematically investigate varied diffusion models for up sampling stage, which is the main bottleneck for streaming synthesis of LM and diffusion-based architecture, we present the model architecture, objective and subjective metrics to show quality and efficiency improvement.

【9】 What do MLLMs hear? Examining reasoning with text and sound components in Multimodal Large Language Models
标题: MLLM听到了什么?在多模式大型语言模型中检查文本和声音成分的推理
作者:Enis Berk Çoban,Michael I. Mandel,Johanna Devaney
备注:9 pages
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已展示出非凡的推理能力,特别是在连接思想和遵守逻辑规则来解决问题方面。这些模型已经发展到适应各种数据形态,包括声音和图像,称为多模式LLM(MLLM),能够描述图像或录音。之前的工作已经证明,当MLLM中的LLM组件被冻结时,音频或视觉编码器用于为声音或图像输入添加字幕,从而促进使用LLM组件进行基于文本的推理。我们有兴趣使用LLM的推理能力来促进分类。在本文中,我们通过一个字幕 分类实验表明,音频MLLM不能充分利用其LLM的基于文本的推理时,生成音频字幕。我们还考虑了这可能是由于MLLM分别表示听觉和文本信息,使得它切断了从LLM到音频编码器的推理路径。摘要:Large Language Models (LLMs) have demonstrated remarkable reasoning capabilities, notably in connecting ideas and adhering to logical rules to solve problems. These models have evolved to accommodate various data modalities, including sound and images, known as multimodal LLMs (MLLMs), which are capable of describing images or sound recordings. Previous work has demonstrated that when the LLM component in MLLMs is frozen, the audio or visual encoder serves to caption the sound or image input facilitating text-based reasoning with the LLM component. We are interested in using the LLM's reasoning capabilities in order to facilitate classification. In this paper, we demonstrate through a captioning classification experiment that an audio MLLM cannot fully leverage its LLM's text-based reasoning when generating audio captions. We also consider how this may be due to MLLMs separately representing auditory and textual information such that it severs the reasoning pathway from the LLM to the audio encoder.

【10】 Neural Codec-based Adversarial Sample Detection for Speaker Verification
标题: 基于神经编解码器的对抗样本检测用于说话人验证
作者:Xuanjun Chen,Jiawei Du,Haibin Wu,Jyh-Shing Roger Jang,Hung-yi Lee
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:自动说话人验证(ASV)越来越多地用于安全关键型应用程序,面临着不断上升的对抗性攻击的漏洞,几乎没有有效的防御措施。本文提出了一种基于神经编解码器的ASV对抗样本检测方法。该方法利用编解码器的能力,以丢弃冗余的扰动和保留必要的信息。具体来说,我们通过比较原始音频和重新合成音频之间的ASV得分差异(通过编解码器模型)来区分真实样本和对抗样本。这项全面的研究探索了所有开源神经编解码器及其变体模型。描述音频编解码器模型在15种神经编解码器中具有最高的检测率,并超过了7种现有的最先进的(SOTA)检测方法。请注意,我们的单模型方法甚至比SOTA集成方法性能更好。摘要:Automatic Speaker Verification (ASV), increasingly used in security-critical applications, faces vulnerabilities from rising adversarial attacks, with few effective defenses available. In this paper, we propose a neural codec-based adversarial sample detection method for ASV. The approach leverages the codec's ability to discard redundant perturbations and retain essential information. Specifically, we distinguish between genuine and adversarial samples by comparing ASV score differences between original and re-synthesized audio (by codec models). This comprehensive study explores all open-source neural codecs and their variant models for experiments. The Descript-audio-codec model stands out by delivering the highest detection rate among 15 neural codecs and surpassing seven prior state-of-the-art (SOTA) detection methods. Note that, our single-model method even outperforms a SOTA ensemble method by a large margin.

【11】 Flexible Multichannel Speech Enhancement for Noise-Robust Frontend
标题: 灵活的多通道语音增强,实现噪音稳健前端
作者:Ante Jukić,Jagadeesh Balam,Boris Ginsburg
Journal-ref:WASPAA 2023
链接:点击下载PDF文件
摘要:本文提出了一种灵活的多通道语音增强系统,主要目标是提高噪声条件下的自动语音识别(ASR)的鲁棒性。该系统结合了一个灵活的神经掩模估计适用于不同的通道计数和配置和多通道滤波器与自动参考选择。一个transform-attend-concatenate层提出来处理跨信道信息的掩模估计,这是有效的任意麦克风配置。所提出的评估证明了几个看不见的紧凑阵列几何形状的灵活系统的有效性,匹配固定的配置特定的系统的性能。此外,对于具有随机放置的麦克风的配置,观察到显著改善的ASR性能。摘要:This paper proposes a flexible multichannel speech enhancement system with the main goal of improving robustness of automatic speech recognition (ASR) in noisy conditions. The proposed system combines a flexible neural mask estimator applicable to different channel counts and configurations and a multichannel filter with automatic reference selection. A transform-attend-concatenate layer is proposed to handle cross-channel information in the mask estimator, which is shown to be effective for arbitrary microphone configurations. The presented evaluation demonstrates the effectiveness of the flexible system for several seen and unseen compact array geometries, matching the performance of fixed configuration-specific systems. Furthermore, a significantly improved ASR performance is observed for configurations with randomly-placed microphones.

【12】 Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
标题: 走向自然主义声音转换:具有自动处理管道的NaturalVoices数据集
作者:Ali N. Salman,Zongyang Du,Shreeram Suresh Chandra,Ismail Rasim Ulgen,Carlos Busso,Berrak Sisman
链接:点击下载PDF文件
摘要:语音转换(VC)研究传统上依赖于脚本或表演的语音,这缺乏现实生活中对话的自然自发性。虽然自然语音数据是有限的VC,我们的研究重点是填补这一空白。我们引入了一个新的数据源管道,它为VC发布了一个名为NaturalVoices的自然语音数据集。该管道利用最新的深度学习方法,从原始播客数据中提取语音中的丰富信息,如情感和信噪比(SNR),并提供灵活性和易用性。NaturalVoices是一个大规模的、自发的、富有表现力和情感的语音数据集,包括来自MSP播客数据集中原始播客的超过3,800小时的语音。客观和主观评估表明,使用我们的管道为VC提供自然和富有表现力的数据是有效的,这表明NaturalVoices在更广泛的语音生成任务中具有潜力。摘要:Voice conversion (VC) research traditionally depends on scripted or acted speech, which lacks the natural spontaneity of real-life conversations. While natural speech data is limited for VC, our study focuses on filling in this gap. We introduce a novel data-sourcing pipeline that makes the release of a natural speech dataset for VC, named NaturalVoices. The pipeline extracts rich information in speech such as emotion and signal-to-noise ratio (SNR) from raw podcast data, utilizing recent deep learning methods and providing flexibility and ease of use. NaturalVoices marks a large-scale, spontaneous, expressive, and emotional speech dataset, comprising over 3,800 hours speech sourced from the original podcasts in the MSP-Podcast dataset. Objective and subjective evaluations demonstrate the effectiveness of using our pipeline for providing natural and expressive data for VC, suggesting the potential of NaturalVoices for broader speech generation tasks.

【13】 Small-E: Small Language Model with Linear Attention for Efficient Speech Synthesis
标题: Small-E:具有线性注意力的小型语言模型,用于高效语音合成
作者:Théodor Lemerle,Nicolas Obin,Axel Roebel
备注:Interspeech
链接:点击下载PDF文件
摘要:由语言模型驱动的文本到语音(TTS)的最新进展展示了在实现自然度和zero-shot语音克隆方面的卓越能力。值得注意的是,解码器专用的Transformer是该领域的主要架构。然而,Transformers面临着序列长度的二次复杂性带来的挑战,阻碍了长序列和资源受限硬件的训练。此外,他们缺乏具体的归纳偏见方面的单调性质的TTS对齐。作为回应,我们建议用新兴的循环架构取代Transformers,并引入专门的交叉注意机制来减少重复和跳过问题。因此,我们的架构可以有效地训练长样本,并实现最先进的zero-shot语音克隆对基线的可比大小。摘要:Recent advancements in text-to-speech (TTS) powered by language models have showcased remarkable capabilities in achieving naturalness and zero-shot voice cloning. Notably, the decoder-only transformer is the prominent architecture in this domain. However, transformers face challenges stemming from their quadratic complexity in sequence length, impeding training on lengthy sequences and resource-constrained hardware. Moreover they lack specific inductive bias with regards to the monotonic nature of TTS alignments. In response, we propose to replace transformers with emerging recurrent architectures and introduce specialized cross-attention mechanisms for reducing repeating and skipping issues. Consequently our architecture can be efficiently trained on long samples and achieve state-of-the-art zero-shot voice cloning against baselines of comparable size.

【14】 LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
标题: LipGER:用于鲁棒自动语音识别的视觉条件生成错误纠正
作者:Sreyan Ghosh,Sonal Kumar,Ashish Seth,Purva Chiniya,Utkarsh Tyagi,Ramani Duraiswami,Dinesh Manocha
备注:InterSpeech 2024. Code and Data: this https URL
链接:点击下载PDF文件
摘要:视觉提示,如嘴唇运动,已被证明可以提高噪声环境中自动语音识别(ASR)系统的性能。我们提出了LipGER(唇运动辅助生成错误校正),一种利用视觉线索进行噪声鲁棒ASR的新框架。我们让LLM学习视觉条件(生成)ASR纠错的任务,而不是学习音频和视觉模态之间的跨模态相关性。具体来说,我们指示LLM从使用ASR波束搜索生成的N个最佳假设预测转录。这进一步取决于嘴唇运动。这种方法解决了传统AVSR学习中的关键挑战,例如缺乏大规模配对数据集以及难以适应新领域。我们在不同设置的4个数据集上进行了实验,结果表明LipGER在1.1%-49.2%的范围内改善了单词错误率。我们还发布了LipHyp,这是一个带有假设-转录对的大规模数据集,它还配备了嘴唇运动线索,以促进这一领域的进一步研究摘要:Visual cues, like lip motion, have been shown to improve the performance of Automatic Speech Recognition (ASR) systems in noisy environments. We propose LipGER (Lip Motion aided Generative Error Correction), a novel framework for leveraging visual cues for noise-robust ASR. Instead of learning the cross-modal correlation between the audio and visual modalities, we make an LLM learn the task of visually-conditioned (generative) ASR error correction. Specifically, we instruct an LLM to predict the transcription from the N-best hypotheses generated using ASR beam-search. This is further conditioned on lip motions. This approach addresses key challenges in traditional AVSR learning, such as the lack of large-scale paired datasets and difficulties in adapting to new domains. We experiment on 4 datasets in various settings and show that LipGER improves the Word Error Rate in the range of 1.1%-49.2%. We also release LipHyp, a large-scale dataset with hypothesis-transcription pairs that is additionally equipped with lip motion cues to promote further research in this space

【15】 InaGVAD : a Challenging French TV and Radio Corpus Annotated for Speech Activity Detection and Speaker Gender Segmentation
标题: InaGVAR:一个令人惊叹的法国电视和广播数据库,用于语音活动检测和说话者性别分割
作者:David Doukhan,Christine Maertens,William Le Personnic,Ludovic Speroni,Reda Dehak
Journal-ref:Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), pages 8963-8974, Torino, Italia. ELRA and ICCL
链接:点击下载PDF文件
摘要:InaGVAD是从10个法国广播和18个电视频道收集的音频语料库,分为4组:通才广播,音乐广播,新闻电视和通才电视。它包含277个1分钟长的注释录音,旨在代表法国视听节目的声音多样性,主要是为了建立能够监测媒体中男子和妇女发言时间的系统。inaGVAD提供有语音活动检测(VAD)和扬声器性别分割(SGS)注释,其扩展有重叠、扬声器特征(性别、年龄、语音质量)和10个非语音事件类别。每个通道类别的注释分布都有详细说明。该数据集被划分为1h开发和3 h37测试子集,允许公平和可重现的系统评估。一个基准的6个免费提供的VAD软件,显示不同的能力的基础上,渠道和非语音事件类别。两个现有的SGS系统的语料库进行评估,并与基线的X向量迁移学习策略,训练的发展子集。结果表明,我们的建议,在一个单一的,但不同的数据小时训练,取得了有竞争力的SGS结果。整个inaGVAD包;包括语料库、注释、评估脚本和基线训练代码;可以免费访问,促进该领域的未来发展。摘要:InaGVAD is an audio corpus collected from 10 French radio and 18 TV channels categorized into 4 groups: generalist radio, music radio, news TV, and generalist TV. It contains 277 1-minute-long annotated recordings aimed at representing the acoustic diversity of French audiovisual programs and was primarily designed to build systems able to monitor men's and women's speaking time in media. inaGVAD is provided with Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) annotations extended with overlap, speaker traits (gender, age, voice quality), and 10 non-speech event categories. Annotation distributions are detailed for each channel category. This dataset is partitioned into a 1h development and a 3h37 test subset, allowing fair and reproducible system evaluation. A benchmark of 6 freely available VAD software is presented, showing diverse abilities based on channel and non-speech event categories. Two existing SGS systems are evaluated on the corpus and compared against a baseline X-vector transfer learning strategy, trained on the development subset. Results demonstrate that our proposal, trained on a single - but diverse - hour of data, achieved competitive SGS results. The entire inaGVAD package; including corpus, annotations, evaluation scripts, and baseline training code; is made freely accessible, fostering future advancement in the domain.

【16】 QiandaoEar22: A high quality noise dataset for identifying specific ship from multiple underwater acoustic targets using ship-radiated noise
标题: QiandaoEar22:一个高质量的噪音数据集,用于利用船舶辐射噪音从多个水下声学目标中识别特定船舶
作者:Xiaoyang Du,Feng Hong
链接:点击下载PDF文件
摘要:舰船辐射噪声的目标识别是水下目标识别的一个重要研究领域。然而,目前缺乏多目标船舶数据集,准确地代表现实世界的水声条件。为了解决这个问题,我们进行了实验数据采集,从而发布了QiandaoEar22 textemdash一个全面的水声多目标数据集。该数据集包含9小时28分钟的真实船舶辐射噪声数据和21小时58分钟的背景噪声数据。为了证明QiandaoEar22的可用性,我们执行了两个实验任务。第一个任务的重点是评估船舶辐射噪声的存在,而第二个任务涉及识别特定的船舶在多船混合数据中的已识别目标。在后一项任务中,我们从数据中提取了8个特征,并使用了6个深度学习网络进行分类,旨在评估和比较各种特征和网络的性能。实验结果表明,在99%以上的情况下,舰船辐射噪声可以成功地从背景噪声中识别出来。对于具体的单船识别,最优识别率达到99.56%。最后,根据我们的研究结果,我们提供了选择合适的特征和深度学习网络的建议,这可能会为相关研究提供有价值的见解。我们的工作不仅建立了算法评估的基准,而且还激发了创新方法的发展,以增强UATD和UATR系统。摘要:Target identification of ship-radiated noise is a crucial area in underwater target recognition. However, there is currently a lack of multi-target ship datasets that accurately represent real-world underwater acoustic conditions. To tackle this issue, we conducted experimental data acquisition, resulting in the release of QiandaoEar22 textemdash a comprehensive underwater acoustic multi-target dataset. This dataset encompasses 9 hours and 28 minutes of real-world ship-radiated noise data and 21 hours and 58 minutes of background noise data. To demonstrate the availability of QiandaoEar22, we executed two experimental tasks. The first task focuses on assessing the presence of ship-radiated noise, while the second task involves identifying specific ships within the recognized targets in the multi-ship mixed data. In the latter task, we extracted eight features from the data and employed six deep learning networks for classification, aiming to evaluate and compare the performance of various features and networks. The experimental results reveal that ship-radiated noise can be successfully identified from background noise in over 99 % of cases. Additionally, for the specific identification of individual ships, the optimal recognition accuracy achieves 99.56 %. Finally, based on our findings, we provide advice on selecting appropriate features and deep learning networks, which may offer valuable insights for related research. Our work not only establishes a benchmark for algorithm evaluation but also inspires the development of innovative methods to enhance UATD and UATR systems.

【17】 Introducing the Brand New QiandaoEar22 Dataset for Specific Ship Identification Using Ship-Radiated Noise
标题: 推出全新QiandaoEar22数据集,用于利用船舶辐射噪音识别特定船舶
作者:Xiaoyang Du,Feng Hong
链接:点击下载PDF文件
摘要:舰船辐射噪声的目标识别是水下目标识别的一个重要研究领域。然而,目前缺乏多目标船舶数据集,准确地代表现实世界的水声条件。为了解决这一问题,我们发布了QiandaoEar 22 textmdash一个水声多目标数据集,可以在https: ieee-dataport.org documents qiandaoear22上下载。该数据集包含9小时28分钟的真实船舶辐射噪声数据和21小时58分钟的背景噪声数据。我们通过从多个目标中识别特定船舶的实验,验证了QiandaoEar 22的可用性。以不同的特征作为输入,六个深度学习网络作为分类器,我们评估了不同方法的基线性能。实验结果表明,将UUV的特定目标与其他目标区分开来,识别率最高可达97.78%,并且发现采用光谱和MFCC作为特征输入,DenseNet作为分类器可以获得更好的识别性能。我们的工作不仅为数据集建立了基准,而且有助于进一步开发水下声目标检测(UATD)和水下声目标识别(UATR)任务的创新方法。摘要:Target identification of ship-radiated noise is a crucial area in underwater target recognition. However, there is currently a lack of multi-target ship datasets that accurately represent real-world underwater acoustic conditions. To ntackle this issue, we release QiandaoEar22 textemdash an underwater acoustic multi-target dataset, which can be download on https: ieee-dataport.org documents qiandaoear22. This dataset encompasses 9 hours and 28 minutes of real-world ship-radiated noise data and 21 hours and 58 minutes of background noise data. We demonstrate the availability of QiandaoEar22 by conducting an experiment of identifying specific ship from the multiple targets. Taking different features as the input and six deep learning networks as classifier, we evaluate the baseline performance of different methods. The experimental results reveal that identifying the specific target of UUV from others can achieve the optimal recognition accuracy of 97.78 %, and we find using spectrum and MFCC as feature inputs and DenseNet as the classifier can achieve better recognition performance. Our work not only establishes a benchmark for the dataset but helps the further development of innovative methods for the tasks of underwater acoustic target detection (UATD) and underwater acoustic target recognition(UATR).

【18】 MA-AVT: Modality Alignment for Parameter-Efficient Audio-Visual Transformers
标题: MA-AVT:参数高效视听变形机的模式对齐
作者:Tanvir Mahmud,Shentong Mo,Yapeng Tian,Diana Marculescu
备注:Accepted in Efficient Deep Learning for Computer Vision CVPR Workshop 2024
链接:点击下载PDF文件
摘要:预训练的Vision Transformers的最新进展表明,在没有音频预训练的情况下,参数高效的视听学习是有希望的。然而,很少有研究调查的有效方法,调整多模态功能参数高效的视听Transformers。在本文中,我们提出了MA-AVT,一个新的参数高效的视听Transformer采用深模态对齐相应的多模态语义特征。具体来说,我们引入了联合单峰和多模态令牌学习,用于将两种模态与冻结的模态共享Transformer对齐。这允许模型为每个模态学习单独的表示,同时也关注它们之间的跨模态关系。此外,与之前只从单峰编码器的输出中对齐粗特征的工作不同,我们引入了分块对比学习来在整个编码阶段对齐粗到细粒度的分层特征。此外,为了从前景匹配的视听特征中抑制每个模态中的背景特征,我们引入了一个鲁棒的判别式前景挖掘方案。通过对基准AVE,VGGSound和CREMA-D数据集的广泛实验,我们实现了SOTA方法的相当大的性能改进。摘要:Recent advances in pre-trained vision transformers have shown promise in parameter-efficient audio-visual learning without audio pre-training. However, few studies have investigated effective methods for aligning multimodal features in parameter-efficient audio-visual transformers. In this paper, we propose MA-AVT, a new parameter-efficient audio-visual transformer employing deep modality alignment for corresponding multimodal semantic features. Specifically, we introduce joint unimodal and multimodal token learning for aligning the two modalities with a frozen modality-shared transformer. This allows the model to learn separate representations for each modality, while also attending to the cross-modal relationships between them. In addition, unlike prior work that only aligns coarse features from the output of unimodal encoders, we introduce blockwise contrastive learning to align coarse-to-fine-grain hierarchical features throughout the encoding phase. Furthermore, to suppress the background features in each modality from foreground matched audio-visual features, we introduce a robust discriminative foreground mining scheme. Through extensive experiments on benchmark AVE, VGGSound, and CREMA-D datasets, we achieve considerable performance improvements over SOTA methods.

【19】 TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
标题: TraceableSpeech:通过水印实现主动可追溯的文本到语音
作者:Junzuo Zhou,Jiangyan Yi,Tao Wang,Jianhua Tao,Ye Bai,Chu Yuan Zhang,Yong Ren,Zhengqi Wen
备注:acceped by interspeech 2024
链接:点击下载PDF文件
摘要:文语转换(TTS)技术的发展带来的各种威胁促使人们需要可靠地跟踪合成语音。然而,目前的方法是在生成音频后单独添加水印,这一过程会损害语音质量和水印的不可感知性。此外,这些方法在鲁棒性和灵活性方面受到限制。为了解决这些问题,我们提出了TraceableSpeech,一种新的TTS模型,直接产生水印的语音,提高水印的不可见性和语音质量。此外,我们设计了分帧的水印嵌入和提取算法,实现了对重拼接攻击的鲁棒性和时间灵活性。实验结果表明,TraceableSpeech算法在水印不可见性、语音质量和抗重拼接攻击能力等方面均优于VALL-E或HiFicodec单独使用WavMark的强基线算法。它也可以应用于各种持续时间的语音。摘要:Various threats posed by the progress in text-to-speech (TTS) have prompted the need to reliably trace synthesized speech. However, contemporary approaches to this task involve adding watermarks to the audio separately after generation, a process that hurts both speech quality and watermark imperceptibility. In addition, these approaches are limited in robustness and flexibility. To address these problems, we propose TraceableSpeech, a novel TTS model that directly generates watermarked speech, improving watermark imperceptibility and speech quality. Furthermore, We design the frame-wise imprinting and extraction of watermarks, achieving higher robustness against resplicing attacks and temporal flexibility in operation. Experimental results show that TraceableSpeech outperforms the strong baseline where VALL-E or HiFicodec individually uses WavMark in watermark imperceptibility, speech quality and resilience against resplicing attacks. It also can apply to speech of various durations.

【20】 Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
标题: 端到端ASB的扬声器平滑kNN扬声器自适应
作者:Shaojun Li,Daimeng Wei,Jiaxin Guo,ZongYao Li,Zhanglin Wu,Zhiqiang Rao,Yuanchang Luo,Xianghui He,Hao Yang
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:尽管最近在端到端自动语音识别(E2 E ASR)系统中有所改进,但是由于训练数据和测试数据之间的声音特征不匹配,特别是在有限的目标说话人自适应数据的情况下,性能可能会下降。我们提出了一种新的说话人自适应方法Speaker-Smoothed kNN,该方法利用k-最近邻(kNN)检索技术,通过在解码阶段从其预建模型中找到正确发音的令牌来提高模型输出。此外,我们利用x向量动态调整kNN插值参数的数据稀疏性问题。该方法在域内和全域环境下使用KeSpeech和MagicData语料库进行了验证。我们的方法始终执行微调,而没有相关的性能下降,在扬声器的变化。此外,在全域设置中,我们的方法实现了最先进的结果,降低了CER在单扬声器和多扬声器测试场景。摘要:Despite recent improvements in End-to-End Automatic Speech Recognition (E2E ASR) systems, the performance can degrade due to vocal characteristic mismatches between training and testing data, particularly with limited target speaker adaptation data. We propose a novel speaker adaptation approach Speaker-Smoothed kNN that leverages k-Nearest Neighbors (kNN) retrieval techniques to improve model output by finding correctly pronounced tokens from its pre-built datastore during the decoding phase. Moreover, we utilize x-vector to dynamically adjust kNN interpolation parameters for data sparsity issue. This approach was validated using KeSpeech and MagicData corpora under in-domain and all-domain settings. Our method consistently performs comparably to fine-tuning without the associated performance degradation during speaker changes. Furthermore, in the all-domain setting, our method achieves state-of-the-art results, reducing the CER in both single speaker and multi-speaker test scenarios.

【21】 PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
标题: PPPR:用于文本到音频生成的便携式插件提示细化器
作者:Shuchen Shi,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Tao Wang,Chunyu Qiang,Yi Lu,Xin Qi,Xuefei Liu,Yukun Liu,Yongwei Li,Zhiyong Wang,Xiaopeng Wang
备注:accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:文本到音频(TTA)旨在生成与给定文本描述相对应的音频,在媒体制作中起着至关重要的作用。TTA数据集中的文本描述缺乏丰富的变化和多样性,导致TTA模型在面对复杂文本时性能下降。为了解决这个问题,我们提出了一种方法,称为便携式插件提示细化,它利用丰富的知识,在大型语言模型中固有的文本描述,有效地提高TTA声学模型的鲁棒性,而不改变声学训练集。此外,还引入了模仿人工验证的Chain-of-Thought来提高音频描述的准确性,从而提高了实际应用中生成内容的准确性。实验表明,我们的方法实现了最先进的初始得分(IS)为8.72,超过AudioGen,AudioLDM和Tango。摘要:Text-to-Audio (TTA) aims to generate audio that corresponds to the given text description, playing a crucial role in media production. The text descriptions in TTA datasets lack rich variations and diversity, resulting in a drop in TTA model performance when faced with complex text. To address this issue, we propose a method called Portable Plug-in Prompt Refiner, which utilizes rich knowledge about textual descriptions inherent in large language models to effectively enhance the robustness of TTA acoustic models without altering the acoustic training set. Furthermore, a Chain-of-Thought that mimics human verification is introduced to enhance the accuracy of audio descriptions, thereby improving the accuracy of generated content in practical applications. The experiments show that our method achieves a state-of-the-art Inception Score (IS) of 8.72, surpassing AudioGen, AudioLDM and Tango.

【22】 MeLFusion: Synthesizing Music from Image and Language Cues using Diffusion Models
标题: MeLFusion:使用扩散模型从图像和语言线索合成音乐
作者:Sanjoy Chowdhury,Sayan Nag,K J Joseph,Balaji Vasan Srinivasan,Dinesh Manocha
备注:Accepted at CVPR 2024 as Highlight paper. Webpage: this https URL
链接:点击下载PDF文件
摘要:音乐是一种通用的语言,可以传达情感和感受。它是整个创意媒体的重要组成部分,从电影到社交媒体帖子。可以合成音乐的机器学习模型主要取决于它的文本描述。受到音乐家如何不仅从电影剧本中创作音乐,而且还通过可视化来创作音乐的启发,我们提出了MeLFusion,这是一种可以有效使用文本描述和相应图像中的线索来合成音乐的模型。MeLFusion是一个文本到音乐的扩散模型,它具有一种新颖的“视觉突触”,可以有效地将视觉模态的语义注入到生成的音乐中。为了方便这方面的研究,我们引入了一个新的数据集MeLBench,并提出了一个新的评价指标IMSM。我们详尽的实验评估表明,将视觉信息添加到音乐合成管道中可以显着提高所生成音乐的质量,无论是客观还是主观测量,FAD分数的相对增益高达67.98%。我们希望,我们的工作将引起人们对这一务实的,但相对不足的研究领域的关注。摘要:Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media, ranging from movies to social media posts. Machine learning models that can synthesize music are predominantly conditioned on textual descriptions of it. Inspired by how musicians compose music not just from a movie script, but also through visualizations, we propose MeLFusion, a model that can effectively use cues from a textual description and the corresponding image to synthesize music. MeLFusion is a text-to-music diffusion model with a novel "visual synapse", which effectively infuses the semantics from the visual modality into the generated music. To facilitate research in this area, we introduce a new dataset MeLBench, and propose a new evaluation metric IMSM. Our exhaustive experimental evaluation suggests that adding visual information to the music synthesis pipeline significantly improves the quality of generated music, measured both objectively and subjectively, with a relative gain of up to 67.98% on the FAD score. We hope that our work will gather attention to this pragmatic, yet relatively under-explored research area.

【23】 Pitch-Aware RNN-T for Mandarin Chinese Mispronunciation Detection and Diagnosis
标题: 具有音调感知的RNN-T用于普通话发音错误检测和诊断
作者:Xintong Wang,Mingqian Shi,Ye Wang
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:利用自动语音识别(ASR)的发音错误检测和诊断(MDD)系统在汉语普通话中面临两个主要挑战:1)两阶段模型在音素或声调分类阶段和MDD阶段之间产生信息间隙。2)Mandarin MDD数据集的稀缺限制了模型训练。在本文中,我们介绍了一个无状态的RNN-T模型,用于普通话MDD,利用HuBERT特征,通过音高融合块进行音高嵌入。我们的模型,仅在母语数据上训练,在非母语场景中,电话错误率提高了3%,错误接受率提高了7%,超过了最先进的基线摘要:Mispronunciation Detection and Diagnosis (MDD) systems, leveraging Automatic Speech Recognition (ASR), face two main challenges in Mandarin Chinese: 1) The two-stage models create an information gap between the phoneme or tone classification stage and the MDD stage. 2) The scarcity of Mandarin MDD datasets limits model training. In this paper, we introduce a stateless RNN-T model for Mandarin MDD, utilizing HuBERT features with pitch embedding through a Pitch Fusion Block. Our model, trained solely on native speaker data, shows a 3% improvement in Phone Error Rate and a 7% increase in False Acceptance Rate over the state-of-the-art baseline in non-native scenarios

【24】 MUSE: Flexible Voiceprint Receptive Fields and Multi-Path Fusion Enhanced Taylor Transformer for U-Net-based Speech Enhancement
标题: MUSE:灵活的声纹接收场和多路径融合增强泰勒Transformer,用于基于U-Net的语音增强
作者:Zizhen Lin,Xiaoting Chen,Junyu Wang
链接:点击下载PDF文件
摘要:实现轻量化设计和高性能之间的平衡仍然是语音增强的一项具有挑战性的任务。在本文中,我们介绍了多路径增强泰勒(MET)Transformer为基础的U-网络语音增强(MUSE),一个轻量级的语音增强网络建立在Unet架构。我们的方法采用了一种新的多路径增强泰勒(MET)Transformer块,它集成了可变形嵌入(DE),使灵活的声纹感受野。MET Transformer被独特地设计为融合通道和空间注意力(CSA)分支,促进通道信息交换并解决泰勒-Transformer框架内的空间注意力缺陷。通过在VoiceBank+DEMAND数据集上进行的大量实验,我们证明了MUSE在显著降低训练和部署成本的同时,实现了具有竞争力的性能,仅拥有0.51 M参数。摘要:Achieving a balance between lightweight design and high performance remains a challenging task for speech enhancement. In this paper, we introduce Multi-path Enhanced Taylor (MET) Transformer based U-net for Speech Enhancement (MUSE), a lightweight speech enhancement network built upon the Unet architecture. Our approach incorporates a novel Multi-path Enhanced Taylor (MET) Transformer block, which integrates Deformable Embedding (DE) to enable flexible receptive fields for voiceprints. The MET Transformer is uniquely designed to fuse Channel and Spatial Attention (CSA) branches, facilitating channel information exchange and addressing spatial attention deficits within the Taylor-Transformer framework. Through extensive experiments conducted on the VoiceBank+DEMAND dataset, we demonstrate that MUSE achieves competitive performance while significantly reducing both training and deployment costs, boasting a mere 0.51M parameters.

【25】 Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
标题: 用于E2 E同步语音翻译的标签同步神经传感器
作者:Keqi Deng,Philip C. Woodland
备注:Accepted by ACL 2024 Main Conference
链接:点击下载PDF文件
摘要:虽然神经传感器在在线语音识别中很受欢迎,但同步语音翻译(SST)需要流和重新排序功能。本文介绍了LS-换能器-SST,标签同步神经传感器SST,它自然具有这两个属性。LS-Transducer-SST基于自回归积分触发(AIF)机制动态决定何时发出翻译令牌。本文还提出了一种时延可控的AIF,它可以只在解码过程中控制质量-时延的权衡,也可以同时用于解码和训练。LS-Transducer-SST可以通过其预测网络自然地利用单语纯文本数据,这有助于缓解E2 E SST数据稀疏的关键问题。在解码过程中,设计了一种基于块的增量式联合解码技术来细化和扩展搜索空间。在Fisher-CallHome Spanish(Es-En)和MuST-C En-De数据上的实验表明,LS-Transducer-SST比现有的流行方法提供了更好的质量-延迟权衡。例如,LS-传感器-SST在相似的延迟下相对于CAAT给出了3.1 2.9点的BLEU增加(Es-En En-De),并且平均滞后延迟减少1.4 s,相对于Wait-k具有相似的BLEU评分。摘要:While the neural transducer is popular for online speech recognition, simultaneous speech translation (SST) requires both streaming and re-ordering capabilities. This paper presents the LS-Transducer-SST, a label-synchronous neural transducer for SST, which naturally possesses these two properties. The LS-Transducer-SST dynamically decides when to emit translation tokens based on an Auto-regressive Integrate-and-Fire (AIF) mechanism. A latency-controllable AIF is also proposed, which can control the quality-latency trade-off either only during decoding, or it can be used in both decoding and training. The LS-Transducer-SST can naturally utilise monolingual text-only data via its prediction network which helps alleviate the key issue of data sparsity for E2E SST. During decoding, a chunk-based incremental joint decoding technique is designed to refine and expand the search space. Experiments on the Fisher-CallHome Spanish (Es-En) and MuST-C En-De data show that the LS-Transducer-SST gives a better quality-latency trade-off than existing popular methods. For example, the LS-Transducer-SST gives a 3.1 2.9 point BLEU increase (Es-En En-De) relative to CAAT at a similar latency and a 1.4 s reduction in average lagging latency with similar BLEU scores relative to Wait-k.

【26】 To Distill or Not to Distill? On the Robustness of Robust Knowledge Distillation
标题: 蒸馏还是不蒸馏?论稳健知识蒸馏的鲁棒性
作者:Abdul Waheed,Karima Kadaoui,Muhammad Abdul-Mageed
备注:Accepted at ACL'24 main
链接:点击下载PDF文件
摘要:众所周知,阿拉伯语对自动语音识别(ASR)提出了独特的挑战。一方面,其丰富的语言多样性和广泛的方言使发展强大的包容性模型变得复杂。另一方面,目前的多语言ASR模型是计算密集型的,缺乏适当的全面评估。鉴于这些挑战,我们将知识从大型教师模型中提取到更小的学生变量中,这些变量更有效。我们还介绍了一个新的人类注释数据集,涵盖五个代表性不足的阿拉伯方言进行评估。我们进一步评估我们的模型和现有的SoTA多语言模型的标准可用基准和我们的新方言数据。我们最好的蒸馏模型的整体性能(45.0 $ % WER)超过了SoTA模型的两倍大小(无障碍M4 T-large-v2,WER= 47.0 $ %)及其教师模型(Whisper-large-v2,WER= 55.1 $ %),其在我们新方言数据上的平均性能(56.9 $ % WER)优于所有其他模型。为了更深入地了解这些模型对方言数据的不良表现,我们进行了错误分析,并报告了不同模型倾向于犯的主要错误类型。该项目的GitHub存储库位于 url{https: github.com UBC-NLP UBC-whisper-ar}。摘要:Arabic is known to present unique challenges for Automatic Speech Recognition (ASR). On one hand, its rich linguistic diversity and wide range of dialects complicate the development of robust, inclusive models. On the other, current multilingual ASR models are compute-intensive and lack proper comprehensive evaluations. In light of these challenges, we distill knowledge from large teacher models into smaller student variants that are more efficient. We also introduce a novel human-annotated dataset covering five under-represented Arabic dialects for evaluation. We further evaluate both our models and existing SoTA multilingual models on both standard available benchmarks and our new dialectal data. Our best-distilled model's overall performance ($45.0$ % WER) surpasses that of a SoTA model twice its size (SeamlessM4T-large-v2, WER=$47.0$ %) and its teacher model (Whisper-large-v2, WER=$55.1$ %), and its average performance on our new dialectal data ($56.9$ % WER) outperforms all other models. To gain more insight into the poor performance of these models on dialectal data, we conduct an error analysis and report the main types of errors the different models tend to make. The GitHub repository for the project is available at url{https: github.com UBC-NLP distill-whisper-ar}.

【27】 Prompt-guided Precise Audio Editing with Diffusion Models
标题: 使用扩散模型的预算引导精确音频编辑
作者:Manjie Xu,Chenxing Li,Duzhen zhang,Dan Su,Wei Liang,Dong Yu
备注:Accepted by ICML 2024
链接:点击下载PDF文件
摘要:音频编辑涉及通过精确控制对音频内容的任意操作。虽然文本引导扩散模型在文本到音频生成方面取得了显着的进步,但它们仍然面临着寻找灵活和精确的方式来修改音轨内的目标事件的挑战。我们提出了一种新的方法,称为PPAE,它作为一个通用的扩散模型模块,并实现精确的音频编辑。编辑仅基于输入的文本提示,并且完全不需要训练。我们利用扩散模型的交叉注意图来促进准确的局部编辑,并采用分层的局部-全局管道来确保更平滑的编辑过程。实验结果突出了我们的方法在各种编辑任务的有效性。摘要:Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as PPAE, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.


机器翻译,仅供参考