今日论文合集:cs.SD语音21篇,eess.AS音频处理24篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
标题: SimpleSpeech:利用纯量潜在Transformer扩散模型实现简单有效的文本到语音
作者:Dongchao Yang,Dingdong Wang,Haohan Guo,Xueyuan Chen,Xixin Wu,Helen Meng
备注:Accepted by InterSpeech 2024
链接:点击下载PDF文件
摘要:在这项研究中,我们提出了一个简单而有效的非自回归(NAR)的文本到语音(TTS)系统的基础上扩散,命名为SimpleSpeech。它的简单性表现在三个方面:(1)它可以在纯语音数据集上训练,没有任何对齐信息;(2)它直接以纯文本作为输入,通过NAR方式生成语音;(3)它试图在有限且紧凑的潜在空间中对语音进行建模,这减轻了扩散建模的难度。更具体地说,我们提出了一种新的语音编解码模型(SQ-Codec)与标量量化,SQ-Codec有效地将复杂的语音信号映射到一个有限的和紧凑的潜在空间,命名为标量潜在空间。借鉴SQ-Codec的思想,我们在SQ-Codec的标量潜在空间中应用了一种新的Transformer扩散模型。我们在4k小时的纯语音数据集上对SimpleSpeech进行了训练,它表现出了自然的韵律和语音克隆能力。与以往的大规模TTS模型相比,该模型在语音质量和生成速度上都有显著的提高。Demos已发布。摘要:In this study, we propose a simple and efficient Non-Autoregressive (NAR) text-to-speech (TTS) system based on diffusion, named SimpleSpeech. Its simpleness shows in three aspects: (1) It can be trained on the speech-only dataset, without any alignment information; (2) It directly takes plain text as input and generates speech through an NAR way; (3) It tries to model speech in a finite and compact latent space, which alleviates the modeling difficulty of diffusion. More specifically, we propose a novel speech codec model (SQ-Codec) with scalar quantization, SQ-Codec effectively maps the complex speech signal into a finite and compact latent space, named scalar latent space. Benefits from SQ-Codec, we apply a novel transformer diffusion model in the scalar latent space of SQ-Codec. We train SimpleSpeech on 4k hours of a speech-only dataset, it shows natural prosody and voice cloning ability. Compared with previous large-scale TTS models, it presents significant speech quality and generation speed improvement. Demos are released.

【2】 An Independence-promoting Loss for Music Generation with Language Models
标题: 语言模型音乐生成的独立性丧失
作者:Jean-Marie Lemercier,Simon Rouard,Jade Copet,Yossi Adi,Alexandre Déffosez
备注:Accepted to ICML 2024
链接:点击下载PDF文件
摘要:使用语言建模的音乐生成方案依赖于音频令牌的词汇表,通常作为由自动编码器学习的离散潜在空间中的代码提供。通常采用多级量化器来产生这些令牌,因此用于令牌预测的解码策略必须适应多个码本:它应该对所有码本上的联合分布进行建模,或者拟合码本边缘分布的乘积。对联合分布进行建模需要增加自回归步骤的数量,而拟合边际的乘积会产生不精确的模型,除非码本相互独立。在这项工作中,我们引入了一个独立的促进损失正则化的自动编码器作为标记的语言模型的音乐生成。建议的损失是基于最大平均差异原则的互信息的代理,应用于可再生核希尔伯特空间。该准则易于实现和训练,并可推广到其他多流编解码器。我们表明,它减少了自动编码过程中码本之间的统计依赖性。这导致在对边际分布的乘积进行建模时所生成的音乐质量的增加,同时比联合分布模型更快地生成音频。摘要:Music generation schemes using language modeling rely on a vocabulary of audio tokens, generally provided as codes in a discrete latent space learnt by an auto-encoder. Multi-stage quantizers are often employed to produce these tokens, therefore the decoding strategy used for token prediction must be adapted to account for multiple codebooks: either it should model the joint distribution over all codebooks, or fit the product of the codebook marginal distributions. Modelling the joint distribution requires a costly increase in the number of auto-regressive steps, while fitting the product of the marginals yields an inexact model unless the codebooks are mutually independent. In this work, we introduce an independence-promoting loss to regularize the auto-encoder used as the tokenizer in language models for music generation. The proposed loss is a proxy for mutual information based on the maximum mean discrepancy principle, applied in reproducible kernel Hilbert spaces. Our criterion is simple to implement and train, and it is generalizable to other multi-stream codecs. We show that it reduces the statistical dependence between codebooks during auto-encoding. This leads to an increase in the generated music quality when modelling the product of the marginal distributions, while generating audio much faster than the joint distribution model.

【3】 Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
标题: 音频曼巴:自我监督音频表示的选择性状态空间
作者:Sarthak Yadav,Zheng-Hua Tan
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:尽管Transformer作为突出的神经架构被广泛采用,但它已经激发了几个独立的工作线来解决其局限性。其中一种方法是选择性状态空间模型,它已经证明了语言建模的有前途的结果。然而,他们的可行性学习自我监督,通用的音频表示还有待研究。这项工作提出了音频曼巴,一个选择性的状态空间模型学习通用的音频表示随机掩蔽频谱补丁通过自我监督。十个不同的音频识别下游任务的实证结果表明,所提出的模型,在AudioSet数据集上进行预训练,始终优于可比的自监督音频频谱图Transformer(SSAST)基线相当大的幅度,并在数据集大小,序列长度和模型大小比较中表现出更好的性能。摘要:Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which have demonstrated promising results for language modelling. However, their feasibility for learning self-supervised, general-purpose audio representations is yet to be investigated. This work proposes Audio Mamba, a selective state space model for learning general-purpose audio representations from randomly masked spectrogram patches through self-supervision. Empirical results on ten diverse audio recognition downstream tasks show that the proposed models, pretrained on the AudioSet dataset, consistently outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by a considerable margin and demonstrate better performance in dataset size, sequence length and model size comparisons.

【4】 Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
标题: Whistle:通过弱语音监督实现数据高效的多语言和跨语言语音识别
作者:Saierdaer Yusuyin,Te Ma,Hao Huang,Wenbo Zhao,Zhijian Ou
链接:点击下载PDF文件
摘要:多语言和跨语言自动语音识别(MCL-ASR)有三种方法-语音或字形转录的监督预训练和自我监督预训练。我们发现,到目前为止,语音监督的预训练对于MCL-ASR来说一直没有得到充分的重视,而从概念上讲,它更有利于不同语言之间的信息共享。本文探讨了一种基于弱语音监督的预训练方法,即Whistle。我们放宽了黄金标准的人类验证的语音成绩单的要求,并获得国际音标(IPA)的基础上的转录,通过利用philageNet字素到音素(G2 P)模型。我们基于CommonVoice数据集构建了一个通用的实验设置,称为CV-Lang 10,其中包含10种可见语言和2种不可见语言。在CV-Lang 10上进行了一组实验,以尽可能公平地比较MCL-ASR常见设置下的三种方法。实验证明了基于音素的MCL-ASR模型(Whistle)的优势,在已知语言的语音识别、不同数量的Few-Shot数据的未知语言的跨语言性能、克服灾难性遗忘和训练效率方面,发现当训练数据更有限时,音素监督可以获得比子词监督和自我监督更好的结果,从而提供更高的数据效率。为了支持可重复性并促进沿着这个方向的未来研究,我们将在发布后在https: github.com thu-spmi CAT上发布整个Whistle管道的代码,模型和数据。摘要:There exist three approaches for multilingual and crosslingual automatic speech recognition (MCL-ASR) - supervised pre-training with phonetic or graphemic transcription, and self-supervised pre-training. We find that pre-training with phonetic supervision has been underappreciated so far for MCL-ASR, while conceptually it is more advantageous for information sharing between different languages. This paper explores the approach of pre-training with weakly phonetic supervision towards data-efficient MCL-ASR, which is called Whistle. We relax the requirement of gold-standard human-validated phonetic transcripts, and obtain International Phonetic Alphabet (IPA) based transcription by leveraging the LanguageNet grapheme-to-phoneme (G2P) models. We construct a common experimental setup based on the CommonVoice dataset, called CV-Lang10, with 10 seen languages and 2 unseen languages. A set of experiments are conducted on CV-Lang10 to compare, as fair as possible, the three approaches under the common setup for MCL-ASR. Experiments demonstrate the advantages of phoneme-based models (Whistle) for MCL-ASR, in terms of speech recognition for seen languages, crosslingual performance for unseen languages with different amounts of few-shot data, overcoming catastrophic forgetting, and training efficiency.It is found that when training data is more limited, phoneme supervision can achieve better results compared to subword supervision and self-supervision, thereby providing higher data-efficiency. To support reproducibility and promote future research along this direction, we will release the code, models and data for the whole pipeline of Whistle at https: github.com thu-spmi CAT upon publication.

【5】 MaskSR: Masked Language Model for Full-band Speech Restoration
标题: MaskSR:全频段语音恢复的掩蔽语言模型
作者:Xu Li,Qirui Wang,Xiaoyu Liu
备注:Accepted by INTERSPEECH 2024. Demo page: this https URL
链接:点击下载PDF文件
摘要:语音恢复的目的是在存在各种失真的情况下恢复高质量的语音。尽管已经研究了几种深度学习范式来完成这项任务,但最近新兴的语言模型的力量还没有得到充分的探索。在本文中,我们提出了MaskSR,一个掩蔽的语言模型,能够恢复全频带44.1 kHz的语音联合考虑噪声,混响,削波和低带宽。MaskSR与使用预训练的神经编解码器提取的离散声学标记一起工作。在训练过程中,MaskSR被优化以预测从高质量目标语音中提取的随机掩蔽令牌,条件是具有各种失真的受损语音。在推理过程中,MaskSR利用高效的迭代采样重构目标语音标记。大量的实验表明,MaskSR获得竞争力的结果,无论是全频带语音恢复任务,也对子任务相比,广泛的模型。摘要:Speech restoration aims at restoring high quality speech in the presence of a diverse set of distortions. Although several deep learning paradigms have been studied for this task, the power of the recently emerging language models has not been fully explored. In this paper, we propose MaskSR, a masked language model capable of restoring full-band 44.1 kHz speech jointly considering noise, reverb, clipping, and low bandwidth. MaskSR works with discrete acoustic tokens extracted using a pre-trained neural codec. During training, MaskSR is optimized to predict randomly masked tokens extracted from the high quality target speech, conditioned on the corrupted speech with various distortions. During inference, MaskSR reconstructs the target speech tokens with efficient iterative sampling. Extensive experiments show that MaskSR obtains competitive results on both the full-band speech restoration task and also on sub-tasks compared with a wide range of models.

【6】 Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping
标题: 有效训练可减少小型化并通过逐核剪裁表现更好的ASB模型
作者:Lun Wang,Om Thakkar,Zhong Meng,Nicole Rafidi,Rohit Prabhavalkar,Arun Narayanan
链接:点击下载PDF文件
摘要:梯度裁剪在大规模自动语音识别(ASR)模型的训练中起着至关重要的作用。它通常应用于小批量梯度以防止梯度爆炸,并应用于单个样本梯度以减轻无意记忆。这项工作系统地研究了特定粒度的梯度裁剪,即每核心裁剪(PCC),在训练各种ASR模型的影响。我们的经验表明,PCC可以有效地减轻非故意的记忆在ASR模型。令人惊讶的是,我们发现PCC积极影响ASR性能指标,从而提高收敛速度和降低字错误率。为了避免调整PCC引入的额外超参数,我们进一步提出了一种新的变体,自适应每核裁剪(APCC),用于流线型优化。我们的研究结果强调了PCC作为强大的隐私前瞻性ASR模型培训策略的多方面好处。摘要:Gradient clipping plays a vital role in training large-scale automatic speech recognition (ASR) models. It is typically applied to minibatch gradients to prevent gradient explosion, and to the individual sample gradients to mitigate unintended memorization. This work systematically investigates the impact of a specific granularity of gradient clipping, namely per-core clip-ping (PCC), across training a wide range of ASR models. We empirically demonstrate that PCC can effectively mitigate unintended memorization in ASR models. Surprisingly, we find that PCC positively influences ASR performance metrics, leading to improved convergence rates and reduced word error rates. To avoid tuning the additional hyperparameter introduced by PCC, we further propose a novel variant, adaptive per-core clipping (APCC), for streamlined optimization. Our findings highlight the multifaceted benefits of PCC as a strategy for robust, privacy-forward ASR model training.

【7】 TinySV: Speaker Verification in TinyML with On-device Learning
标题: TinySV:TinyML中的说话者验证,通过设备上学习
作者:Massimo Pavan,Gioele Mombelli,Francesco Sinacori,Manuel Roveri
链接:点击下载PDF文件
摘要:TinyML是机器学习的一个新领域,由于能够在微型设备(如物联网或嵌入式系统)上执行机器学习算法,在过去几年中获得了巨大的发展势头。有趣的是,这一领域的研究集中在微型设备上TinyML模型推理阶段的有效执行上,而由于学习算法引入的相关开销,文献中很少有TinyML模型的设备学习解决方案。 本文的目的是介绍一种新型的自适应TinyML解决方案,可用于任务,如提出的 textit{Tiny Speaker Verification}(TinySV),需要用设备上的学习算法来解决。实现这一目标需要(i)减少TinyML学习算法的内存和计算需求,以及(ii)设计一种使用少量且可能未标记的训练数据的TinyML学习算法。所提出的TinySV解决方案依赖于两层分层TinyML解决方案,包括关键字定位和自适应说话人验证模块。我们在专门为此任务收集的数据集上评估了所提出的TinySV解决方案的有效性和效率,并在真实的物联网设备(英飞凌PSoC 62 S2 Wi-Fi BT Pioneer Kit)上测试了所提出的解决方案。摘要:TinyML is a novel area of machine learning that gained huge momentum in the last few years thanks to the ability to execute machine learning algorithms on tiny devices (such as Internet-of-Things or embedded systems). Interestingly, research in this area focused on the efficient execution of the inference phase of TinyML models on tiny devices, while very few solutions for on-device learning of TinyML models are available in the literature due to the relevant overhead introduced by the learning algorithms. The aim of this paper is to introduce a new type of adaptive TinyML solution that can be used in tasks, such as the presented textit{Tiny Speaker Verification} (TinySV), that require to be tackled with an on-device learning algorithm. Achieving this goal required (i) reducing the memory and computational demand of TinyML learning algorithms, and (ii) designing a TinyML learning algorithm operating with few and possibly unlabelled training data. The proposed TinySV solution relies on a two-layer hierarchical TinyML solution comprising Keyword Spotting and Adaptive Speaker Verification module. We evaluated the effectiveness and efficiency of the proposed TinySV solution on a dataset collected expressly for the task and tested the proposed solution on a real-world IoT device (Infineon PSoC 62S2 Wi-Fi BT Pioneer Kit).

【8】 Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
标题: Zero-Shot多语言口语关键词识别的模糊通用语音属性建模
作者:Hao Yen,Pin-Jui Ku,Sabato Marco Siniscalchi,Chin-Hui Lee
链接:点击下载PDF文件
摘要:我们提出了一种新的语言通用方法来实现端到端自动口语关键词识别(SKR),该方法利用(i)自监督预训练模型,以及(ii)一组通用语音属性(发音方式和发音位置)。具体来说,Wav2Vec2.0用于生成鲁棒的语音表示,然后是线性输出层以生成属性序列。然后,不可训练的发音模型将属性序列映射到多语言设置中的口语关键字。在多语言口语语料库上的实验表明,与基于字符和音素的SKR在所见语言中的性能相当。包含域对抗训练(DAT)改进了所提出的框架,优于基于字符和音素的SKR方法,在可见语言中相对单词错误率(WER)降低了13.73%和17.22%,在zero-shot设置中,对于未见过的语言,WER降低了32.14%和19.92%。摘要:We propose a novel language-universal approach to end-to-end automatic spoken keyword recognition (SKR) leveraging upon (i) a self-supervised pre-trained model, and (ii) a set of universal speech attributes (manner and place of articulation). Specifically, Wav2Vec2.0 is used to generate robust speech representations, followed by a linear output layer to produce attribute sequences. A non-trainable pronunciation model then maps sequences of attributes into spoken keywords in a multilingual setting. Experiments on the Multilingual Spoken Words Corpus show comparable performances to character- and phoneme-based SKR in seen languages. The inclusion of domain adversarial training (DAT) improves the proposed framework, outperforming both character- and phoneme-based SKR approaches with 13.73% and 17.22% relative word error rate (WER) reduction in seen languages, and achieves 32.14% and 19.92% WER reduction for unseen languages in zero-shot settings.

【9】 How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
标题: 神经欺骗对策如何检测部分欺骗的音频?
作者:Tianchi Liu,Lin Zhang,Rohan Kumar Das,Yi Ma,Ruijie Tao,Haizhou Li
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:对一个句子进行部分的修饰可以大大改变它的意思。最近的工作表明,在部分欺骗音频上训练的对策(CM)可以有效地检测这种欺骗。然而,目前对CM的决策过程的理解是有限的。我们利用Grad-CAM并引入定量分析指标来解释CM的决策。我们发现,CM优先级的文物时,连接真正的和欺骗性的音频创建的过渡区。这一重点不同于在完全欺骗的音频上训练的CM,后者专注于真实和欺骗部分之间的模式差异。我们的进一步调查解释了CM在做出正确或不正确的预测时的不同性质。这些见解为CM模型的设计和数据集的创建提供了基础。此外,这项工作奠定了可解释性的基础,在该领域的部分欺骗音频检测,以前没有得到很好的探索。摘要:Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofing. However, the current understanding of the decision-making process of CMs is limited. We utilize Grad-CAM and introduce a quantitative analysis metric to interpret CMs' decisions. We find that CMs prioritize the artifacts of transition regions created when concatenating bona fide and spoofed audio. This focus differs from that of CMs trained on fully spoofed audio, which concentrate on the pattern differences between bona fide and spoofed parts. Our further investigation explains the varying nature of CMs' focus while making correct or incorrect predictions. These insights provide a basis for the design of CM models and the creation of datasets. Moreover, this work lays a foundation of interpretability in the field of partial spoofed audio detection that has not been well explored previously.

【10】 CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
标题: CtrSDDD:受控歌唱声音Deepfake检测的基准数据集和基线分析
作者:Yongyi Zang,Jiatong Shi,You Zhang,Ryuichi Yamamoto,Jionghao Han,Yuxun Tang,Shengyuan Xu,Wenxiao Zhao,Jing Guo,Tomoki Toda,Zhiyao Duan
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:最近的歌声合成和转换进步需要鲁棒的歌声深度假检测(SVDD)模型。由于可控性有限、deepfake方法的多样性和许可限制,目前的SVDD数据集面临挑战。为了解决这些差距,我们引入了CtrSVDD,这是一个大规模的,多样化的bonafide和deepfake人声集合。这些声音是使用最先进的方法从公开访问的歌声数据集合成的。CtrSVDD包括47.64小时的bonafide和260.34小时的deepfake演唱,跨越14种deepfake方法,涉及164个歌手身份。我们还提出了一个基线系统,灵活的前端功能,对结构化的训练 开发 评估分裂。实验表明了特征选择的重要性,并强调了对进一步偏离训练分布的deepfake方法进行泛化的必要性。CtrSVDD数据集和基线可公开访问。摘要:Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited controllability, diversity in deepfake methods, and licensing restrictions. Addressing these gaps, we introduce CtrSVDD, a large-scale, diverse collection of bonafide and deepfake singing vocals. These vocals are synthesized using state-of-the-art methods from publicly accessible singing voice datasets. CtrSVDD includes 47.64 hours of bonafide and 260.34 hours of deepfake singing vocals, spanning 14 deepfake methods and involving 164 singer identities. We also present a baseline system with flexible front-end features, evaluated against a structured train dev eval split. The experiments show the importance of feature selection and highlight a need for generalization towards deepfake methods that deviate further from training distribution. The CtrSVDD dataset and baselines are publicly accessible.

【11】 Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
标题: Seed-TTC:一系列高质量多功能语音生成模型
作者:Philip Anastassiou,Jiawei Chen,Jitong Chen,Yuanzhe Chen,Zhuo Chen,Ziyi Chen,Jian Cong,Lelai Deng,Chuang Ding,Lu Gao,Mingqing Gong,Peisong Huang,Qingqing Huang,Zhiying Huang,Yuanyuan Huo,Dongya Jia,Chumin Li,Feiya Li,Hui Li,Jiaxin Li,Xiaoyang Li,Xingxing Li,Lin Liu,Shouda Liu,Sichao Liu,Xudong Liu,Yuchen Liu,Zhengxi Liu,Lu Lu,Junjie Pan,Xin Wang,Yuping Wang,Yuxuan Wang,Zhen Wei,Jian Wu,Chao Yao,Yifeng Yang,Yuanhao Yi,Junteng Zhang,Qidi Zhang,Shuo Zhang,Wenjie Zhang,Yang Zhang,Zilin Zhao,Dejian Zhong,Xiaobin Zhuang
链接:点击下载PDF文件
摘要:我们介绍种子TTS,一个家庭的大规模自回归文本到语音(TTS)模型,能够生成语音,这是几乎无法区分从人类的语音。Seed-TTS作为语音生成的基础模型,在语音上下文学习方面表现出色,在客观和主观评估方面都实现了与地面真实人类语音相匹配的说话人相似性和自然度性能。通过微调,我们在这些指标上获得了更高的主观分数。Seed-TTS提供了对各种语音属性(如情感)的卓越可控性,并且能够为野外的说话者生成高度表达和多样化的语音。此外,我们提出了一种自蒸馏方法的语音分解,以及强化学习方法,以提高模型的鲁棒性,说话人相似性和可控性。我们还提出了一个非自回归(NAR)的种子TTS模型的变体,命名为$ text{Seed-TTS}_ text{DiT}$,它利用了一个完全基于扩散的架构。与以前基于NAR的TTS系统不同,$ text{Seed-TTS}_ text{DiT}$不依赖于预先估计的音素持续时间,而是通过端到端处理来执行语音生成。我们证明,这种变体实现了与基于语言模型的变体相当的性能,并展示了其在语音编辑中的有效性。我们鼓励读者在 url{https: bytedancespeech.github.io seedtts_tech_report}上收听演示。摘要:We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-TTS serves as a foundation model for speech generation and excels in speech in-context learning, achieving performance in speaker similarity and naturalness that matches ground truth human speech in both objective and subjective evaluations. With fine-tuning, we achieve even higher subjective scores across these metrics. Seed-TTS offers superior controllability over various speech attributes such as emotion and is capable of generating highly expressive and diverse speech for speakers in the wild. Furthermore, we propose a self-distillation method for speech factorization, as well as a reinforcement learning approach to enhance model robustness, speaker similarity, and controllability. We additionally present a non-autoregressive (NAR) variant of the Seed-TTS model, named $ text{Seed-TTS}_ text{DiT}$, which utilizes a fully diffusion-based architecture. Unlike previous NAR-based TTS systems, $ text{Seed-TTS}_ text{DiT}$ does not depend on pre-estimated phoneme durations and performs speech generation through end-to-end processing. We demonstrate that this variant achieves comparable performance to the language model-based variant and showcase its effectiveness in speech editing. We encourage readers to listen to demos at url{https: bytedancespeech.github.io seedtts_tech_report}.

【12】 Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
标题: 自我监督的歌唱声音预训练以实现语音到歌唱的转换
作者:Ruiqi Li,Rongjie Huang,Yongqi Wang,Zhiqing Hong,Zhou Zhao
备注:13 pages
链接:点击下载PDF文件
摘要:语音到歌声的转换(STS)任务,往往遭受数据稀缺,因为它需要成对的语音和歌唱数据。使这个问题更加复杂的是内容-音高对齐的挑战和生成的输出的次优质量,在STS研究中提出了重大障碍。本文提出了SVPT,一种STS方法,由一个自我监督的歌声预训练模型。我们利用口语模型技术来解决节奏对齐问题,并利用上下文学习能力来实现zero-shot转换。我们采用离散单元随机重传和音高腐败策略,使训练与未配对的歌唱数据,从而减轻数据稀缺的问题。SVPT还可以作为歌唱声音合成(SVS)的有效骨干,为扩展SVS模型提供见解。实验结果表明,SVPT提供了显着的改善,在STS和SVS的努力。音频样本可在https: speech2sing.github.io上获得。摘要:Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal quality of generated outputs, presenting significant hurdles in STS research. This paper presents SVPT, an STS approach boosted by a self-supervised singing voice pre-training model. We leverage spoken language model techniques to tackle the rhythm alignment problem and the in-context learning capability to achieve zero-shot conversion. We adopt discrete-unit random resampling and pitch corruption strategies, enabling training with unpaired singing data and thus mitigating the issue of data scarcity. SVPT also serves as an effective backbone for singing voice synthesis (SVS), offering insights into scaling up SVS models. Experimental results indicate that SVPT delivers notable improvements in both STS and SVS endeavors. Audio samples are available at https: speech2sing.github.io.

【13】 Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
标题: 通过利用大规模ASB模型,通过自我监督学习实现说话人验证的监督性能
作者:Victor Miara,Theo Lepage,Reda Dehak
备注:accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:自监督学习(SSL)的最新进展在说话人确认(SV)中显示出了有希望的结果。然而,缩小与监督系统的性能差距仍然是一个持续的挑战。一些研究已经观察到,来自大规模ASR模型的语音表示包含有价值的说话人信息。这项工作探讨了在端到端的方法中使用SSL对比目标微调SV的这些模型的局限性。然后,我们提出了一个框架,通过使用伪标签对预训练的WavLM进行微调,使用监督损失来学习SSL上下文中的说话人表示。初始伪标签来自基于SSL DINO的模型,并通过聚类模型嵌入来迭代细化。我们的方法在VoxCeleb 1-O上实现了0.99%的EER,在自监督SV上建立了新的最先进水平。由于该性能接近我们的监督基线0.94%EER,因此该贡献是向使用SSL的SV的监督性能迈出的一步。摘要:Recent advancements in Self-Supervised Learning (SSL) have shown promising results in Speaker Verification (SV). However, narrowing the performance gap with supervised systems remains an ongoing challenge. Several studies have observed that speech representations from large-scale ASR models contain valuable speaker information. This work explores the limitations of fine-tuning these models for SV using an SSL contrastive objective in an end-to-end approach. Then, we propose a framework to learn speaker representations in an SSL context by fine-tuning a pre-trained WavLM with a supervised loss using pseudo-labels. Initial pseudo-labels are derived from an SSL DINO-based model and are iteratively refined by clustering the model embeddings. Our method achieves 0.99% EER on VoxCeleb1-O, establishing the new state-of-the-art on self-supervised SV. As this performance is close to our supervised baseline of 0.94% EER, this contribution is a step towards supervised performance on SV with SSL.

【14】 MidiCaps -- A large-scale MIDI dataset with text captions
标题: MidiCaps --带文本字幕的大规模收件箱数据集
作者:Jan Melechovsky,Abhinaba Roy,Dorien Herremans
备注:Under review
链接:点击下载PDF文件
摘要:由文本提示引导的生成模型越来越受欢迎。然而,目前还不存在文本到文本的模型,主要是由于缺乏标题的文本数据集。这项工作的目的是通过提供第一个带有公开文本标题的大规模数字音乐数据集(MidiCaps)来实现将LLM与符号音乐相结合的研究。乐器数字接口(Musical Instrument Digital Interface,简称MIDI)是一种广泛使用的音乐信息编码格式。他们的结构化格式捕捉到了音乐作品的细微差别,并为音乐制作人、作曲家、音乐学家以及表演者提供了实际应用。受应用于各个领域的字幕技术的最新进展的启发,我们提出了一个大规模的策划数据集,其中包含超过168k的文本描述文件。每个字幕简洁地描述了音乐内容,包括节奏,和弦进行,节拍,乐器,流派和情绪,从而促进多模态探索和分析。该数据集包含各种类型,风格和复杂性的混合,为训练和评估音乐信息检索,音乐理解和跨模态翻译等任务的模型提供了丰富的资源。我们提供了有关数据集的详细统计数据,并在广泛的听力研究中评估了字幕的质量。我们预计,这一资源将刺激音乐和自然语言处理交叉领域的进一步研究,促进这两个领域的进步。摘要:Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist, mostly due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic music by presenting the first large-scale MIDI dataset with text captions that is openly available: MidiCaps. MIDI (Musical Instrument Digital Interface) files are a widely used format for encoding musical information. Their structured format captures the nuances of musical composition and has practical applications by music producers, composers, musicologists, as well as performers. Inspired by recent advancements in captioning techniques applied to various domains, we present a large-scale curated dataset of over 168k MIDI files accompanied by textual descriptions. Each MIDI caption succinctly describes the musical content, encompassing tempo, chord progression, time signature, instruments present, genre and mood; thereby facilitating multi-modal exploration and analysis. The dataset contains a mix of various genres, styles, and complexities, offering a rich source for training and evaluating models for tasks such as music information retrieval, music understanding and cross-modal translation. We provide detailed statistics about the dataset and have assessed the quality of the captions in an extensive listening study. We anticipate that this resource will stimulate further research in the intersection of music and natural language processing, fostering advancements in both fields.

【15】 Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
标题: 具有灵活采样率控制的多阶段语音带宽扩展
作者:Ye-Xin Lu,Yang Ai,Zheng-Yan Sheng,Zhen-Hua Ling
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:现有的语音带宽扩展(BWE)方法大多在固定的源和目标采样率的约束下工作,这限制了它们在实际应用中的灵活性。本文提出了一种多级语音BWE模型MS-BWE,它可以处理一组源和目标采样率对,并实现灵活的频带扩展。所提出的MS-BWE模型包括BWE块的级联,每个块具有双流架构以实现幅度和相位扩展,逐步逐阶段地描绘语音频带。教师强迫策略被用来缓解训练和推理之间的差异。实验结果表明,我们提出的MS-BWE是最先进的语音BWE方法的语音质量相媲美。在生成效率方面,MS-BWE的一级生成在GPU上可以达到1000倍以上的实时性,在CPU上可以达到60倍左右。摘要:The majority of existing speech bandwidth extension (BWE) methods operate under the constraint of fixed source and target sampling rates, which limits their flexibility in practical applications. In this paper, we propose a multi-stage speech BWE model named MS-BWE, which can handle a set of source and target sampling rate pairs and achieve flexible extensions of frequency bandwidth. The proposed MS-BWE model comprises a cascade of BWE blocks, with each block featuring a dual-stream architecture to realize amplitude and phase extension, progressively painting the speech frequency bands stage by stage. The teacher-forcing strategy is employed to mitigate the discrepancy between training and inference. Experimental results demonstrate that our proposed MS-BWE is comparable to state-of-the-art speech BWE methods in speech quality. Regarding generation efficiency, the one-stage generation of MS-BWE can achieve over one thousand times real-time on GPU and about sixty times on CPU.

【16】 BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
标题: BiVocoder:集成特征提取和波形生成的双向神经声码器
作者:Hui-Peng Du,Ye-Xin Lu,Yang Ai,Zhen-Hua Ling
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的双向神经声码器,命名为BiVocoder,能够在短时傅立叶变换(STFT)域的特征提取和反向波形生成。对于特征提取,BiVocoder将STFT中得到的幅度和相位谱作为输入,通过卷积神经网络将其转换为长帧移位和低维特征。提取的特征被证明是适合直接预测的声学模型,支持其应用在文本到语音(TTS)的任务。对于波形生成,BiVocoder通过对称网络从特征中恢复幅度和相位谱,然后通过逆STFT重建语音波形。实验结果表明,我们提出的BiVocoder实现了更好的性能相比,一些基线声码器,综合考虑合成语音质量和推理速度的分析合成和TTS任务。摘要:This paper proposes a novel bidirectional neural vocoder, named BiVocoder, capable both of feature extraction and reverse waveform generation within the short-time Fourier transform (STFT) domain. For feature extraction, the BiVocoder takes amplitude and phase spectra derived from STFT as inputs, transforms them into long-frame-shift and low-dimensional features through convolutional neural networks. The extracted features are demonstrated suitable for direct prediction by acoustic models, supporting its application in text-to-speech (TTS) task. For waveform generation, the BiVocoder restores amplitude and phase spectra from the features by a symmetric network, followed by inverse STFT to reconstruct the speech waveform. Experimental results show that our proposed BiVocoder achieves better performance compared to some baseline vocoders, by comprehensively considering both synthesized speech quality and inference speed for both analysis-synthesis and TTS tasks.

【17】 SimulTron: On-Device Simultaneous Speech to Speech Translation
标题: SimulTron:设备上同步语音到语音翻译
作者:Alex Agranovich,Eliya Nachmani,Oleg Rybakov,Yifan Ding,Ye Jia,Nadav Bar,Heiga Zen,Michelle Tadmor Ramanovich
链接:点击下载PDF文件
摘要:同步语音到语音翻译(S2 ST)有望打破沟通障碍,实现跨语言的流畅对话。然而,通过移动设备实现准确的实时翻译仍然是一个重大挑战。我们介绍SimulTron,一种新的S2 ST架构,旨在解决这一任务。SimulTron是一个轻量级的直接S2 ST模型,它使用了Translatotron框架的优势,同时结合了流操作的关键修改和可调的固定延迟。我们的实验表明,SimulTron在离线评估中超过了Translatotron 2。此外,实时评估显示,SimulTron在Translatotron 1的性能上有所改进。此外,与MuST-C数据集上以前的实时S2 ST方法相比,SimulTron实现了更好的BLEU分数和延迟。值得注意的是,我们已经成功地在Pixel 7 Pro设备上部署了SimulTron,显示了其在设备上同步S2 ST的潜力。摘要:Simultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST architecture designed to tackle this task. SimulTron is a lightweight direct S2ST model that uses the strengths of the Translatotron framework while incorporating key modifications for streaming operation, and an adjustable fixed delay. Our experiments show that SimulTron surpasses Translatotron 2 in offline evaluations. Furthermore, real-time evaluations reveal that SimulTron improves upon the performance achieved by Translatotron 1. Additionally, SimulTron achieves superior BLEU scores and latency compared to previous real-time S2ST method on the MuST-C dataset. Significantly, we have successfully deployed SimulTron on a Pixel 7 Pro device, show its potential for simultaneous S2ST on-device.

【18】 M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
标题: M2 D-CLAP:掩蔽建模二人组与CLAP会面,学习通用音频语言表示
作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Masahiro Yasuda,Shunsuke Tsubaki,Keisuke Imoto
备注:5 pages, 1 figure, 5 tables. Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:对比语言-音频预训练(CLAP)实现了音频的zero-shot(ZS)推断,并在几个分类任务中表现出良好的性能。然而,传统的音频表示对于ZS不适用的许多任务(例如,回归问题)。在这里,我们探索一种新的表示,一种通用的音频语言表示,在ZS和迁移学习中表现良好。为此,我们提出了一种新的方法,M2 D-CLAP,它结合了自监督学习Masked Modeling Duo(M2 D)和CLAP。M2 D学习有效的表示来对音频信号进行建模,CLAP将表示与文本嵌入对齐。因此,M2 D-CLAP学习了一种通用的表示,允许ZS和迁移学习。实验表明,M2 D-CLAP在线性评估、微调和ZS分类方面表现良好,GTZAN的最新水平为75.17%,从而实现了通用的音频语言表示。摘要:Contrastive language-audio pre-training (CLAP) enables zero-shot (ZS) inference of audio and exhibits promising performance in several classification tasks. However, conventional audio representations are still crucial for many tasks where ZS is not applicable (e.g., regression problems). Here, we explore a new representation, a general-purpose audio-language representation, that performs well in both ZS and transfer learning. To do so, we propose a new method, M2D-CLAP, which combines self-supervised learning Masked Modeling Duo (M2D) and CLAP. M2D learns an effective representation to model audio signals, and CLAP aligns the representation with text embedding. As a result, M2D-CLAP learns a versatile representation that allows for both ZS and transfer learning. Experiments show that M2D-CLAP performs well on linear evaluation, fine-tuning, and ZS classification with a GTZAN state-of-the-art of 75.17%, thus achieving a general-purpose audio-language representation.

【19】 Understanding Auditory Evoked Brain Signal via Physics-informed Embedding Network with Multi-Task Transformer
标题: 通过具有多任务Transformer的物理信息嵌入网络理解听觉诱发的大脑信号
作者:Wanli Ma,Xuegang Tang,Jin Gu,Ying Wang,Yuling Xia
链接:点击下载PDF文件
摘要:在脑机交互和认知神经科学领域,有效解码来自基于任务的功能磁共振成像(fMRI)的听觉信号是理解大脑如何处理复杂听觉信息的关键。虽然现有的方法具有增强的解码能力,但在信息利用和模型表示方面仍然存在限制。为了克服这些挑战,我们提出了一种创新的多任务学习模型,即具有多任务Transformer的物理信息嵌入网络(PEMT-Net),它通过物理信息嵌入和深度学习技术来增强解码性能。PEMT-Net由两个主要组成部分组成:特征增强和分类。对于特征增强,我们提出了一种新的方法,通过节点嵌入创建神经嵌入图,利用随机游走来模拟神经信息的物理扩散。该方法捕获局部和非局部信息溢出,并提出了一种基于相对物理坐标的位置编码。在分类部分,我们提出了自适应嵌入融合,以最大限度地捕捉线性和非线性特征。此外,我们提出了一个创新的参数共享机制,以优化提取的特征的保留和学习。在特定数据集上的实验证明了PEMT-Net在多任务听觉信号解码方面的显著性能,超越了现有方法,并为大脑处理复杂听觉信息的机制提供了新的见解。摘要:In the fields of brain-computer interaction and cognitive neuroscience, effective decoding of auditory signals from task-based functional magnetic resonance imaging (fMRI) is key to understanding how the brain processes complex auditory information. Although existing methods have enhanced decoding capabilities, limitations remain in information utilization and model representation. To overcome these challenges, we propose an innovative multi-task learning model, Physics-informed Embedding Network with Multi-Task Transformer (PEMT-Net), which enhances decoding performance through physics-informed embedding and deep learning techniques. PEMT-Net consists of two principal components: feature augmentation and classification. For feature augmentation, we propose a novel approach by creating neural embedding graphs via node embedding, utilizing random walks to simulate the physical diffusion of neural information. This method captures both local and non-local information overflow and proposes a position encoding based on relative physical coordinates. In the classification segment, we propose adaptive embedding fusion to maximally capture linear and non-linear characteristics. Furthermore, we propose an innovative parameter-sharing mechanism to optimize the retention and learning of extracted features. Experiments on a specific dataset demonstrate PEMT-Net's significant performance in multi-task auditory signal decoding, surpassing existing methods and offering new insights into the brain's mechanisms for processing complex auditory information.

【20】 Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
标题: 用于文本到语音合成的语音增强语言建模
作者:Kun Zhou,Shengkui Zhao,Yukun Ma,Chong Zhang,Hao Wang,Dianwen Ng,Chongjia Ni,Nguyen Trung Hieu,Jia Qi Yip,Bin Ma
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:最近基于语言模型的文本到语音(TTS)框架展示了可扩展性和上下文学习能力。然而,他们遭受鲁棒性问题,由于自回归语言建模过程中的语音单元预测中的错误的积累。在本文中,我们提出了一种语音增强的语言建模方法,以提高TTS模型的性能。我们利用语音丰富的自监督表示作为自回归语言模型的训练目标。随后,采用非自回归模型来预测包含细粒度声学细节的离散声学编解码器。TTS模型只关注自回归训练期间的语言建模,从而减少了非自回归训练中发生的错误传播。客观和主观评价都验证了我们所提出的方法的有效性。摘要:Recent language model-based text-to-speech (TTS) frameworks demonstrate scalability and in-context learning capabilities. However, they suffer from robustness issues due to the accumulation of errors in speech unit predictions during autoregressive language modeling. In this paper, we propose a phonetic enhanced language modeling method to improve the performance of TTS models. We leverage self-supervised representations that are phonetically rich as the training target for the autoregressive language model. Subsequently, a non-autoregressive model is employed to predict discrete acoustic codecs that contain fine-grained acoustic details. The TTS model focuses solely on linguistic modeling during autoregressive training, thereby reducing the error propagation that occurs in non-autoregressive training. Both objective and subjective evaluations validate the effectiveness of our proposed method.

【21】 Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
标题: 揭开隐藏因素:用于语音情感识别特征增强的可解释人工智能
作者:Alaa Nfissi,Wassim Bouachir,Nizar Bouguila,Brian Mishara
Journal-ref:Applied Intelligence (2024)
链接:点击下载PDF文件
摘要:语音情感识别由于其在心理健康、教育、人机交互等领域的广泛应用而受到广泛关注。然而,SER系统的准确性受到可能包含不相关和冗余信息的高维特征集的阻碍。为了克服这一挑战,本研究提出了一种迭代的SER特征提升方法,强调特征的相关性和可解释性,以提高机器学习模型的性能。我们的方法包括细致的特征选择和分析,以构建高效的SER系统。在通过模型可解释性来解决我们的主要问题时,我们采用了一个具有Shapley值的特征评估循环来迭代地细化特征集。这个过程在模型性能和透明度之间取得了平衡,从而能够全面了解模型的预测。所提出的方法提供了几个优点,包括识别和删除不相关和冗余的功能,导致一个更有效的模型。此外,它还提高了可解释性,促进了对模型预测的理解,并识别了情感确定的关键特征。在多伦多情感语音集(TESS)、柏林情感语音数据库(EMO-DB)、瑞尔森情感语音和歌曲视听数据库(RAVDESS)和萨里视听表达情感(SAVEE)数据集的SER基准测试上,验证了该方法的有效性,性能优于现有方法。这些结果突出了所提出的技术在开发准确和可解释的SER系统的潜力。据我们所知,这是第一个将模型可解释性纳入SER框架的工作。摘要:Speech emotion recognition (SER) has gained significant attention due to its several application fields, such as mental health, education, and human-computer interaction. However, the accuracy of SER systems is hindered by high-dimensional feature sets that may contain irrelevant and redundant information. To overcome this challenge, this study proposes an iterative feature boosting approach for SER that emphasizes feature relevance and explainability to enhance machine learning model performance. Our approach involves meticulous feature selection and analysis to build efficient SER systems. In addressing our main problem through model explainability, we employ a feature evaluation loop with Shapley values to iteratively refine feature sets. This process strikes a balance between model performance and transparency, which enables a comprehensive understanding of the model's predictions. The proposed approach offers several advantages, including the identification and removal of irrelevant and redundant features, leading to a more effective model. Additionally, it promotes explainability, facilitating comprehension of the model's predictions and the identification of crucial features for emotion determination. The effectiveness of the proposed method is validated on the SER benchmarks of the Toronto emotional speech set (TESS), Berlin Database of Emotional Speech (EMO-DB), Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS), and Surrey Audio-Visual Expressed Emotion (SAVEE) datasets, outperforming state-of-the-art methods. These results highlight the potential of the proposed technique in developing accurate and explainable SER systems. To the best of our knowledge, this is the first work to incorporate model explainability into an SER framework.


eess.AS音频处理
【1】 Language-Universal Speech Attributes Modeling for Zero-Shot Multilingual Spoken Keyword Recognition
标题: Zero-Shot多语言口语关键词识别的模糊通用语音属性建模
作者:Hao Yen,Pin-Jui Ku,Sabato Marco Siniscalchi,Chin-Hui Lee
链接:点击下载PDF文件
摘要:我们提出了一种新的语言通用方法来实现端到端自动口语关键词识别(SKR),该方法利用(i)自监督预训练模型,以及(ii)一组通用语音属性(发音方式和发音位置)。具体来说,Wav2Vec2.0用于生成鲁棒的语音表示,然后是线性输出层以生成属性序列。然后,不可训练的发音模型将属性序列映射到多语言设置中的口语关键字。在多语言口语语料库上的实验表明,与基于字符和音素的SKR在所见语言中的性能相当。包含域对抗训练(DAT)改进了所提出的框架,优于基于字符和音素的SKR方法,在可见语言中相对单词错误率(WER)降低了13.73%和17.22%,在zero-shot设置中,对于未见过的语言,WER降低了32.14%和19.92%。摘要:We propose a novel language-universal approach to end-to-end automatic spoken keyword recognition (SKR) leveraging upon (i) a self-supervised pre-trained model, and (ii) a set of universal speech attributes (manner and place of articulation). Specifically, Wav2Vec2.0 is used to generate robust speech representations, followed by a linear output layer to produce attribute sequences. A non-trainable pronunciation model then maps sequences of attributes into spoken keywords in a multilingual setting. Experiments on the Multilingual Spoken Words Corpus show comparable performances to character- and phoneme-based SKR in seen languages. The inclusion of domain adversarial training (DAT) improves the proposed framework, outperforming both character- and phoneme-based SKR approaches with 13.73% and 17.22% relative word error rate (WER) reduction in seen languages, and achieves 32.14% and 19.92% WER reduction for unseen languages in zero-shot settings.

【2】 How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
标题: 神经欺骗对策如何检测部分欺骗的音频?
作者:Tianchi Liu,Lin Zhang,Rohan Kumar Das,Yi Ma,Ruijie Tao,Haizhou Li
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:对一个句子进行部分的修饰可以大大改变它的意思。最近的工作表明,在部分欺骗音频上训练的对策(CM)可以有效地检测这种欺骗。然而,目前对CM的决策过程的理解是有限的。我们利用Grad-CAM并引入定量分析指标来解释CM的决策。我们发现,CM优先级的文物时,连接真正的和欺骗性的音频创建的过渡区。这一重点不同于在完全欺骗的音频上训练的CM,后者专注于真实和欺骗部分之间的模式差异。我们的进一步调查解释了CM在做出正确或不正确的预测时的不同性质。这些见解为CM模型的设计和数据集的创建提供了基础。此外,这项工作奠定了可解释性的基础,在该领域的部分欺骗音频检测,以前没有得到很好的探索。摘要:Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofing. However, the current understanding of the decision-making process of CMs is limited. We utilize Grad-CAM and introduce a quantitative analysis metric to interpret CMs' decisions. We find that CMs prioritize the artifacts of transition regions created when concatenating bona fide and spoofed audio. This focus differs from that of CMs trained on fully spoofed audio, which concentrate on the pattern differences between bona fide and spoofed parts. Our further investigation explains the varying nature of CMs' focus while making correct or incorrect predictions. These insights provide a basis for the design of CM models and the creation of datasets. Moreover, this work lays a foundation of interpretability in the field of partial spoofed audio detection that has not been well explored previously.

【3】 Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
标题: 印度艺术音乐中Raga识别的可解释深度学习分析
作者:Parampreet Singh,Vipul Arora
链接:点击下载PDF文件
摘要:在音乐信息检索中,歌曲识别是一个非常热门的研究课题。很少有研究探索这一任务,采用了各种方法,如信号处理,机器学习(ML)方法,以及最近基于深度学习(DL)的方法。然而,在所有这些工作中,一个关键问题仍然没有得到回答:这些ML DL方法是否以类似于人类专家的方式学习和解释Ragas?此外,这项研究中的一个重要障碍是缺乏丰富的标记数据集,这驱动了这些基于ML DL的方法。在本文中,我们介绍了“Prasarbharti Indian Music”version-1(PIM-v1),这是一个新的数据集,包含191小时精心标记的印度斯坦古典音乐(HCM)录音,据我们所知,这是HCM录音的最大标记数据集。我们的方法涉及进行消融研究,以找到使用PIM-v1数据集的自动Raga识别(ARI)的基准分类模型。对于12个Raga类的子集,我们实现了0.89的分块f1分数。随后,我们采用模型可解释性技术来评估分类器的预测,旨在确定它们是否符合人类对Ragas的理解,或者是由任意模式驱动的。我们通过比较两个ExAI模型与人类专家注释给出的解释来验证模型预测的正确性。在此之后,我们分析了各个测试示例的解释,以了解解释所强调的区域在模型做出的正确或不正确预测中的作用。摘要:The task of Raga Identification is a very popular research problem in Music Information Retrieval. Few studies that have explored this task employed various approaches, such as signal processing, Machine Learning (ML) methods, and more recently Deep Learning (DL) based methods. However, a key question remains unanswered in all of these works: do these ML DL methods learn and interpret Ragas in a manner similar to human experts? Besides, a significant roadblock in this research is the unavailability of ample supply of rich, labeled datasets, which drives these ML DL based methods. In this paper, we introduce "Prasarbharti Indian Music" version-1 (PIM-v1), a novel dataset comprising of 191 hours of meticulously labeled Hindustani Classical Music (HCM) recordings, which is the largest labeled dataset for HCM recordings to the best of our knowledge. Our approach involves conducting ablation studies to find the benchmark classification model for Automatic Raga Identification (ARI) using PIM-v1 dataset. We achieve a chunk-wise f1-score of 0.89 for a subset of 12 Raga classes. Subsequently, we employ model explainability techniques to evaluate the classifier's predictions, aiming to ascertain whether they align with human understanding of Ragas or are driven by arbitrary patterns. We validate the correctness of model's predictions by comparing the explanations given by two ExAI models with human expert annotations. Following this, we analyze explanations for individual test examples to understand the role of regions highlighted by explanations in correct or incorrect predictions made by the model.

【4】 CtrSVDD: A Benchmark Dataset and Baseline Analysis for Controlled Singing Voice Deepfake Detection
标题: CtrSDDD:受控歌唱声音Deepfake检测的基准数据集和基线分析
作者:Yongyi Zang,Jiatong Shi,You Zhang,Ryuichi Yamamoto,Jionghao Han,Yuxun Tang,Shengyuan Xu,Wenxiao Zhao,Jing Guo,Tomoki Toda,Zhiyao Duan
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:最近的歌声合成和转换进步需要鲁棒的歌声深度假检测(SVDD)模型。由于可控性有限、deepfake方法的多样性和许可限制,目前的SVDD数据集面临挑战。为了解决这些差距,我们引入了CtrSVDD,这是一个大规模的,多样化的bonafide和deepfake人声集合。这些声音是使用最先进的方法从公开访问的歌声数据集合成的。CtrSVDD包括47.64小时的bonafide和260.34小时的deepfake演唱,跨越14种deepfake方法,涉及164个歌手身份。我们还提出了一个基线系统,灵活的前端功能,对结构化的训练 开发 评估分裂。实验表明了特征选择的重要性,并强调了对进一步偏离训练分布的deepfake方法进行泛化的必要性。CtrSVDD数据集和基线可公开访问。摘要:Recent singing voice synthesis and conversion advancements necessitate robust singing voice deepfake detection (SVDD) models. Current SVDD datasets face challenges due to limited controllability, diversity in deepfake methods, and licensing restrictions. Addressing these gaps, we introduce CtrSVDD, a large-scale, diverse collection of bonafide and deepfake singing vocals. These vocals are synthesized using state-of-the-art methods from publicly accessible singing voice datasets. CtrSVDD includes 47.64 hours of bonafide and 260.34 hours of deepfake singing vocals, spanning 14 deepfake methods and involving 164 singer identities. We also present a baseline system with flexible front-end features, evaluated against a structured train dev eval split. The experiments show the importance of feature selection and highlight a need for generalization towards deepfake methods that deviate further from training distribution. The CtrSVDD dataset and baselines are publicly accessible.

【5】 Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
标题: Seed-TTC:一系列高质量多功能语音生成模型
作者:Philip Anastassiou,Jiawei Chen,Jitong Chen,Yuanzhe Chen,Zhuo Chen,Ziyi Chen,Jian Cong,Lelai Deng,Chuang Ding,Lu Gao,Mingqing Gong,Peisong Huang,Qingqing Huang,Zhiying Huang,Yuanyuan Huo,Dongya Jia,Chumin Li,Feiya Li,Hui Li,Jiaxin Li,Xiaoyang Li,Xingxing Li,Lin Liu,Shouda Liu,Sichao Liu,Xudong Liu,Yuchen Liu,Zhengxi Liu,Lu Lu,Junjie Pan,Xin Wang,Yuping Wang,Yuxuan Wang,Zhen Wei,Jian Wu,Chao Yao,Yifeng Yang,Yuanhao Yi,Junteng Zhang,Qidi Zhang,Shuo Zhang,Wenjie Zhang,Yang Zhang,Zilin Zhao,Dejian Zhong,Xiaobin Zhuang
链接:点击下载PDF文件
摘要:我们介绍种子TTS,一个家庭的大规模自回归文本到语音(TTS)模型,能够生成语音,这是几乎无法区分从人类的语音。Seed-TTS作为语音生成的基础模型,在语音上下文学习方面表现出色,在客观和主观评估方面都实现了与地面真实人类语音相匹配的说话人相似性和自然度性能。通过微调,我们在这些指标上获得了更高的主观分数。Seed-TTS提供了对各种语音属性(如情感)的卓越可控性,并且能够为野外的说话者生成高度表达和多样化的语音。此外,我们提出了一种自蒸馏方法的语音分解,以及强化学习方法,以提高模型的鲁棒性,说话人相似性和可控性。我们还提出了一个非自回归(NAR)的种子TTS模型的变体,命名为$ text{Seed-TTS}_ text{DiT}$,它利用了一个完全基于扩散的架构。与以前基于NAR的TTS系统不同,$ text{Seed-TTS}_ text{DiT}$不依赖于预先估计的音素持续时间,而是通过端到端处理来执行语音生成。我们证明,这种变体实现了与基于语言模型的变体相当的性能,并展示了其在语音编辑中的有效性。我们鼓励读者在 url{https: bytedancespeech.github.io seedtts_tech_report}上收听演示。摘要:We introduce Seed-TTS, a family of large-scale autoregressive text-to-speech (TTS) models capable of generating speech that is virtually indistinguishable from human speech. Seed-TTS serves as a foundation model for speech generation and excels in speech in-context learning, achieving performance in speaker similarity and naturalness that matches ground truth human speech in both objective and subjective evaluations. With fine-tuning, we achieve even higher subjective scores across these metrics. Seed-TTS offers superior controllability over various speech attributes such as emotion and is capable of generating highly expressive and diverse speech for speakers in the wild. Furthermore, we propose a self-distillation method for speech factorization, as well as a reinforcement learning approach to enhance model robustness, speaker similarity, and controllability. We additionally present a non-autoregressive (NAR) variant of the Seed-TTS model, named $ text{Seed-TTS}_ text{DiT}$, which utilizes a fully diffusion-based architecture. Unlike previous NAR-based TTS systems, $ text{Seed-TTS}_ text{DiT}$ does not depend on pre-estimated phoneme durations and performs speech generation through end-to-end processing. We demonstrate that this variant achieves comparable performance to the language model-based variant and showcase its effectiveness in speech editing. We encourage readers to listen to demos at url{https: bytedancespeech.github.io seedtts_tech_report}.

【6】 Self-Supervised Singing Voice Pre-Training towards Speech-to-Singing Conversion
标题: 自我监督的歌唱声音预训练以实现语音到歌唱的转换
作者:Ruiqi Li,Rongjie Huang,Yongqi Wang,Zhiqing Hong,Zhou Zhao
备注:13 pages
链接:点击下载PDF文件
摘要:语音到歌声的转换(STS)任务,往往遭受数据稀缺,因为它需要成对的语音和歌唱数据。使这个问题更加复杂的是内容-音高对齐的挑战和生成的输出的次优质量,在STS研究中提出了重大障碍。本文提出了SVPT,一种STS方法,由一个自我监督的歌声预训练模型。我们利用口语模型技术来解决节奏对齐问题,并利用上下文学习能力来实现zero-shot转换。我们采用离散单元随机重传和音高腐败策略,使训练与未配对的歌唱数据,从而减轻数据稀缺的问题。SVPT还可以作为歌唱声音合成(SVS)的有效骨干,为扩展SVS模型提供见解。实验结果表明,SVPT提供了显着的改善,在STS和SVS的努力。音频样本可在https: speech2sing.github.io上获得。摘要:Speech-to-singing voice conversion (STS) task always suffers from data scarcity, because it requires paired speech and singing data. Compounding this issue are the challenges of content-pitch alignment and the suboptimal quality of generated outputs, presenting significant hurdles in STS research. This paper presents SVPT, an STS approach boosted by a self-supervised singing voice pre-training model. We leverage spoken language model techniques to tackle the rhythm alignment problem and the in-context learning capability to achieve zero-shot conversion. We adopt discrete-unit random resampling and pitch corruption strategies, enabling training with unpaired singing data and thus mitigating the issue of data scarcity. SVPT also serves as an effective backbone for singing voice synthesis (SVS), offering insights into scaling up SVS models. Experimental results indicate that SVPT delivers notable improvements in both STS and SVS endeavors. Audio samples are available at https: speech2sing.github.io.

【7】 Towards Supervised Performance on Speaker Verification with Self-Supervised Learning by Leveraging Large-Scale ASR Models
标题: 通过利用大规模ASB模型,通过自我监督学习实现说话人验证的监督性能
作者:Victor Miara,Theo Lepage,Reda Dehak
备注:accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:自监督学习(SSL)的最新进展在说话人确认(SV)中显示出了有希望的结果。然而,缩小与监督系统的性能差距仍然是一个持续的挑战。一些研究已经观察到,来自大规模ASR模型的语音表示包含有价值的说话人信息。这项工作探讨了在端到端的方法中使用SSL对比目标微调SV的这些模型的局限性。然后,我们提出了一个框架,通过使用伪标签对预训练的WavLM进行微调,使用监督损失来学习SSL上下文中的说话人表示。初始伪标签来自基于SSL DINO的模型,并通过聚类模型嵌入来迭代细化。我们的方法在VoxCeleb 1-O上实现了0.99%的EER,在自监督SV上建立了新的最先进水平。由于该性能接近我们的监督基线0.94%EER,因此该贡献是向使用SSL的SV的监督性能迈出的一步。摘要:Recent advancements in Self-Supervised Learning (SSL) have shown promising results in Speaker Verification (SV). However, narrowing the performance gap with supervised systems remains an ongoing challenge. Several studies have observed that speech representations from large-scale ASR models contain valuable speaker information. This work explores the limitations of fine-tuning these models for SV using an SSL contrastive objective in an end-to-end approach. Then, we propose a framework to learn speaker representations in an SSL context by fine-tuning a pre-trained WavLM with a supervised loss using pseudo-labels. Initial pseudo-labels are derived from an SSL DINO-based model and are iteratively refined by clustering the model embeddings. Our method achieves 0.99% EER on VoxCeleb1-O, establishing the new state-of-the-art on self-supervised SV. As this performance is close to our supervised baseline of 0.94% EER, this contribution is a step towards supervised performance on SV with SSL.

【8】 MidiCaps -- A large-scale MIDI dataset with text captions
标题: MidiCaps --带文本字幕的大规模收件箱数据集
作者:Jan Melechovsky,Abhinaba Roy,Dorien Herremans
备注:Under review
链接:点击下载PDF文件
摘要:由文本提示引导的生成模型越来越受欢迎。然而,目前还不存在文本到文本的模型,主要是由于缺乏标题的文本数据集。这项工作的目的是通过提供第一个带有公开文本标题的大规模数字音乐数据集(MidiCaps)来实现将LLM与符号音乐相结合的研究。乐器数字接口(Musical Instrument Digital Interface,简称MIDI)是一种广泛使用的音乐信息编码格式。他们的结构化格式捕捉到了音乐作品的细微差别,并为音乐制作人、作曲家、音乐学家以及表演者提供了实际应用。受应用于各个领域的字幕技术的最新进展的启发,我们提出了一个大规模的策划数据集,其中包含超过168k的文本描述文件。每个字幕简洁地描述了音乐内容,包括节奏,和弦进行,节拍,乐器,流派和情绪,从而促进多模态探索和分析。该数据集包含各种类型,风格和复杂性的混合,为训练和评估音乐信息检索,音乐理解和跨模态翻译等任务的模型提供了丰富的资源。我们提供了有关数据集的详细统计数据,并在广泛的听力研究中评估了字幕的质量。我们预计,这一资源将刺激音乐和自然语言处理交叉领域的进一步研究,促进这两个领域的进步。摘要:Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist, mostly due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic music by presenting the first large-scale MIDI dataset with text captions that is openly available: MidiCaps. MIDI (Musical Instrument Digital Interface) files are a widely used format for encoding musical information. Their structured format captures the nuances of musical composition and has practical applications by music producers, composers, musicologists, as well as performers. Inspired by recent advancements in captioning techniques applied to various domains, we present a large-scale curated dataset of over 168k MIDI files accompanied by textual descriptions. Each MIDI caption succinctly describes the musical content, encompassing tempo, chord progression, time signature, instruments present, genre and mood; thereby facilitating multi-modal exploration and analysis. The dataset contains a mix of various genres, styles, and complexities, offering a rich source for training and evaluating models for tasks such as music information retrieval, music understanding and cross-modal translation. We provide detailed statistics about the dataset and have assessed the quality of the captions in an extensive listening study. We anticipate that this resource will stimulate further research in the intersection of music and natural language processing, fostering advancements in both fields.

【9】 Multi-Stage Speech Bandwidth Extension with Flexible Sampling Rate Control
标题: 具有灵活采样率控制的多阶段语音带宽扩展
作者:Ye-Xin Lu,Yang Ai,Zheng-Yan Sheng,Zhen-Hua Ling
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:现有的语音带宽扩展(BWE)方法大多在固定的源和目标采样率的约束下工作,这限制了它们在实际应用中的灵活性。本文提出了一种多级语音BWE模型MS-BWE,它可以处理一组源和目标采样率对,并实现灵活的频带扩展。所提出的MS-BWE模型包括BWE块的级联,每个块具有双流架构以实现幅度和相位扩展,逐步逐阶段地描绘语音频带。教师强迫策略被用来缓解训练和推理之间的差异。实验结果表明,我们提出的MS-BWE是最先进的语音BWE方法的语音质量相媲美。在生成效率方面,MS-BWE的一级生成在GPU上可以达到1000倍以上的实时性,在CPU上可以达到60倍左右。摘要:The majority of existing speech bandwidth extension (BWE) methods operate under the constraint of fixed source and target sampling rates, which limits their flexibility in practical applications. In this paper, we propose a multi-stage speech BWE model named MS-BWE, which can handle a set of source and target sampling rate pairs and achieve flexible extensions of frequency bandwidth. The proposed MS-BWE model comprises a cascade of BWE blocks, with each block featuring a dual-stream architecture to realize amplitude and phase extension, progressively painting the speech frequency bands stage by stage. The teacher-forcing strategy is employed to mitigate the discrepancy between training and inference. Experimental results demonstrate that our proposed MS-BWE is comparable to state-of-the-art speech BWE methods in speech quality. Regarding generation efficiency, the one-stage generation of MS-BWE can achieve over one thousand times real-time on GPU and about sixty times on CPU.

【10】 Towards Out-of-Distribution Detection in Vocoder Recognition via Latent Feature Reconstruction
标题: 通过潜在特征重建实现声码器识别中的分布外检测
作者:Renmingyue Du,Jixun Yao,Qiuqiang Kong,Yin Cao
备注:5 pages, 4 figures
链接:点击下载PDF文件
摘要:合成语音的进步带来了越来越大的模仿威胁,这使得开发deepfake算法识别变得至关重要。一个重要的方面是分布外(OOD)检测,由于其在Deepfake算法识别中的重要作用而受到了广泛关注。然而,目前用于在deepfake算法识别中检测OOD的大多数方法都依赖于概率分数或分类距离,这可能导致阈值边缘处样本的准确性受到限制。在这项研究中,我们提出了一种基于重建的检测方法,该方法采用自动编码器架构来压缩和重建从预训练的WavLM模型中提取的声学特征。属于特定声码器类的每个声学特征仅由其对应的解码器适当地重建。当没有一个解码器可以令人满意地重建一个特征时,它被分类为OOD样本。为了增强每个解码器的重构特征的独特性,我们结合了对比学习和辅助分类器来进一步约束重构特征。实验表明,我们提出的方法超过基线系统的相对利润率为10%的评估数据集。消融研究进一步验证了我们提出的方法中的对比约束和辅助分类器的有效性。摘要:Advancements in synthesized speech have created a growing threat of impersonation, making it crucial to develop deepfake algorithm recognition. One significant aspect is out-of-distribution (OOD) detection, which has gained notable attention due to its important role in deepfake algorithm recognition. However, most of the current approaches for detecting OOD in deepfake algorithm recognition rely on probability-score or classified-distance, which may lead to limitations in the accuracy of the sample at the edge of the threshold. In this study, we propose a reconstruction-based detection approach that employs an autoencoder architecture to compress and reconstruct the acoustic feature extracted from a pre-trained WavLM model. Each acoustic feature belonging to a specific vocoder class is only aptly reconstructed by its corresponding decoder. When none of the decoders can satisfactorily reconstruct a feature, it is classified as an OOD sample. To enhance the distinctiveness of the reconstructed features by each decoder, we incorporate contrastive learning and an auxiliary classifier to further constrain the reconstructed feature. Experiments demonstrate that our proposed approach surpasses baseline systems by a relative margin of 10 % in the evaluation dataset. Ablation studies further validate the effectiveness of both the contrastive constraint and the auxiliary classifier within our proposed approach.

【11】 ERes2NetV2: Boosting Short-Duration Speaker Verification Performance with Computational Efficiency
标题: ERes2NetV2:通过计算效率提高短期说话人验证性能
作者:Yafeng Chen,Siqi Zheng,Hui Wang,Luyao Cheng,Qian Chen,Shiliang Zhang,Junjie Li
链接:点击下载PDF文件
摘要:说话人确认系统在进行短时间的试验录音时会出现显著的性能下降。为了解决这一问题,提出了一种多尺度特征融合的方法来有效地从短话语中提取说话人特征。受模型大小的限制,结合全局和局部特征融合的鲁棒骨干增强型Res 2Net(ERes 2Net)在短时间说话人确认中表现出次优性能。为了进一步提高ERes 2Net的短时特征提取能力,我们在每个阶段内扩展了通道维度。然而,这种修改也增加了模型参数的数量和计算复杂度。为了缓解这个问题,我们提出了一个改进的ERes 2NetV 2通过修剪冗余结构,最终降低模型参数和计算成本。在VoxCeleb数据集上进行的一系列实验显示了ERes 2NetV 2的优越性,其在VoxCeleb 1-O上的全持续时间试验、3s持续时间试验和2s持续时间试验的EER分别为0.61%、0.98%和1.48%。摘要:Speaker verification systems experience significant performance degradation when tasked with short-duration trial recordings. To address this challenge, a multi-scale feature fusion approach has been proposed to effectively capture speaker characteristics from short utterances. Constrained by the model's size, a robust backbone Enhanced Res2Net (ERes2Net) combining global and local feature fusion demonstrates sub-optimal performance in short-duration speaker verification. To further improve the short-duration feature extraction capability of ERes2Net, we expand the channel dimension within each stage. However, this modification also increases the number of model parameters and computational complexity. To alleviate this problem, we propose an improved ERes2NetV2 by pruning redundant structures, ultimately reducing both the model parameters and its computational cost. A range of experiments conducted on the VoxCeleb datasets exhibits the superiority of ERes2NetV2, which achieves EER of 0.61% for the full-duration trial, 0.98% for the 3s-duration trial, and 1.48% for the 2s-duration trial on VoxCeleb1-O, respectively.

【12】 BiVocoder: A Bidirectional Neural Vocoder Integrating Feature Extraction and Waveform Generation
标题: BiVocoder:集成特征提取和波形生成的双向神经声码器
作者:Hui-Peng Du,Ye-Xin Lu,Yang Ai,Zhen-Hua Ling
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的双向神经声码器,命名为BiVocoder,能够在短时傅立叶变换(STFT)域的特征提取和反向波形生成。对于特征提取,BiVocoder将STFT中得到的幅度和相位谱作为输入,通过卷积神经网络将其转换为长帧移位和低维特征。提取的特征被证明是适合直接预测的声学模型,支持其应用在文本到语音(TTS)的任务。对于波形生成,BiVocoder通过对称网络从特征中恢复幅度和相位谱,然后通过逆STFT重建语音波形。实验结果表明,我们提出的BiVocoder实现了更好的性能相比,一些基线声码器,综合考虑合成语音质量和推理速度的分析合成和TTS任务。摘要:This paper proposes a novel bidirectional neural vocoder, named BiVocoder, capable both of feature extraction and reverse waveform generation within the short-time Fourier transform (STFT) domain. For feature extraction, the BiVocoder takes amplitude and phase spectra derived from STFT as inputs, transforms them into long-frame-shift and low-dimensional features through convolutional neural networks. The extracted features are demonstrated suitable for direct prediction by acoustic models, supporting its application in text-to-speech (TTS) task. For waveform generation, the BiVocoder restores amplitude and phase spectra from the features by a symmetric network, followed by inverse STFT to reconstruct the speech waveform. Experimental results show that our proposed BiVocoder achieves better performance compared to some baseline vocoders, by comprehensively considering both synthesized speech quality and inference speed for both analysis-synthesis and TTS tasks.

【13】 SimulTron: On-Device Simultaneous Speech to Speech Translation
标题: SimulTron:设备上同步语音到语音翻译
作者:Alex Agranovich,Eliya Nachmani,Oleg Rybakov,Yifan Ding,Ye Jia,Nadav Bar,Heiga Zen,Michelle Tadmor Ramanovich
链接:点击下载PDF文件
摘要:同步语音到语音翻译(S2 ST)有望打破沟通障碍,实现跨语言的流畅对话。然而,通过移动设备实现准确的实时翻译仍然是一个重大挑战。我们介绍SimulTron,一种新的S2 ST架构,旨在解决这一任务。SimulTron是一个轻量级的直接S2 ST模型,它使用了Translatotron框架的优势,同时结合了流操作的关键修改和可调的固定延迟。我们的实验表明,SimulTron在离线评估中超过了Translatotron 2。此外,实时评估显示,SimulTron在Translatotron 1的性能上有所改进。此外,与MuST-C数据集上以前的实时S2 ST方法相比,SimulTron实现了更好的BLEU分数和延迟。值得注意的是,我们已经成功地在Pixel 7 Pro设备上部署了SimulTron,显示了其在设备上同步S2 ST的潜力。摘要:Simultaneous speech-to-speech translation (S2ST) holds the promise of breaking down communication barriers and enabling fluid conversations across languages. However, achieving accurate, real-time translation through mobile devices remains a major challenge. We introduce SimulTron, a novel S2ST architecture designed to tackle this task. SimulTron is a lightweight direct S2ST model that uses the strengths of the Translatotron framework while incorporating key modifications for streaming operation, and an adjustable fixed delay. Our experiments show that SimulTron surpasses Translatotron 2 in offline evaluations. Furthermore, real-time evaluations reveal that SimulTron improves upon the performance achieved by Translatotron 1. Additionally, SimulTron achieves superior BLEU scores and latency compared to previous real-time S2ST method on the MuST-C dataset. Significantly, we have successfully deployed SimulTron on a Pixel 7 Pro device, show its potential for simultaneous S2ST on-device.

【14】 M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation
标题: M2 D-CLAP:掩蔽建模二人组与CLAP会面,学习通用音频语言表示
作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Masahiro Yasuda,Shunsuke Tsubaki,Keisuke Imoto
备注:5 pages, 1 figure, 5 tables. Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:对比语言-音频预训练(CLAP)实现了音频的zero-shot(ZS)推断,并在几个分类任务中表现出良好的性能。然而,传统的音频表示对于ZS不适用的许多任务(例如,回归问题)。在这里,我们探索一种新的表示,一种通用的音频语言表示,在ZS和迁移学习中表现良好。为此,我们提出了一种新的方法,M2D-CLAP,它结合了自监督学习Masked Modeling Duo(M2D)和CLAP。M2D学习有效的表示来对音频信号进行建模,CLAP将表示与文本嵌入对齐。因此,M2D-CLAP学习了一种通用的表示,允许ZS和迁移学习。实验表明,M2D-CLAP在线性评估、微调和ZS分类方面表现良好,GTZAN的最新水平为75.17%,从而实现了通用的音频语言表示。摘要:Contrastive language-audio pre-training (CLAP) enables zero-shot (ZS) inference of audio and exhibits promising performance in several classification tasks. However, conventional audio representations are still crucial for many tasks where ZS is not applicable (e.g., regression problems). Here, we explore a new representation, a general-purpose audio-language representation, that performs well in both ZS and transfer learning. To do so, we propose a new method, M2D-CLAP, which combines self-supervised learning Masked Modeling Duo (M2D) and CLAP. M2D learns an effective representation to model audio signals, and CLAP aligns the representation with text embedding. As a result, M2D-CLAP learns a versatile representation that allows for both ZS and transfer learning. Experiments show that M2D-CLAP performs well on linear evaluation, fine-tuning, and ZS classification with a GTZAN state-of-the-art of 75.17%, thus achieving a general-purpose audio-language representation.

【15】 Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
标题: 用于文本到语音合成的语音增强语言建模
作者:Kun Zhou,Shengkui Zhao,Yukun Ma,Chong Zhang,Hao Wang,Dianwen Ng,Chongjia Ni,Nguyen Trung Hieu,Jia Qi Yip,Bin Ma
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:最近基于语言模型的文本到语音(TTS)框架展示了可扩展性和上下文学习能力。然而,他们遭受鲁棒性问题,由于自回归语言建模过程中的语音单元预测中的错误的积累。在本文中,我们提出了一种语音增强的语言建模方法,以提高TTS模型的性能。我们利用语音丰富的自监督表示作为自回归语言模型的训练目标。随后,采用非自回归模型来预测包含细粒度声学细节的离散声学编解码器。TTS模型只关注自回归训练期间的语言建模,从而减少了非自回归训练中发生的错误传播。客观和主观评价都验证了我们所提出的方法的有效性。摘要:Recent language model-based text-to-speech (TTS) frameworks demonstrate scalability and in-context learning capabilities. However, they suffer from robustness issues due to the accumulation of errors in speech unit predictions during autoregressive language modeling. In this paper, we propose a phonetic enhanced language modeling method to improve the performance of TTS models. We leverage self-supervised representations that are phonetically rich as the training target for the autoregressive language model. Subsequently, a non-autoregressive model is employed to predict discrete acoustic codecs that contain fine-grained acoustic details. The TTS model focuses solely on linguistic modeling during autoregressive training, thereby reducing the error propagation that occurs in non-autoregressive training. Both objective and subjective evaluations validate the effectiveness of our proposed method.

【16】 Unveiling Hidden Factors: Explainable AI for Feature Boosting in Speech Emotion Recognition
标题: 揭开隐藏因素:用于语音情感识别特征增强的可解释人工智能
作者:Alaa Nfissi,Wassim Bouachir,Nizar Bouguila,Brian Mishara
Journal-ref:Applied Intelligence (2024)
链接:点击下载PDF文件
摘要:语音情感识别由于其在心理健康、教育、人机交互等领域的广泛应用而受到广泛关注。然而,SER系统的准确性受到可能包含不相关和冗余信息的高维特征集的阻碍。为了克服这一挑战,本研究提出了一种迭代的SER特征提升方法,强调特征的相关性和可解释性,以提高机器学习模型的性能。我们的方法包括细致的特征选择和分析,以构建高效的SER系统。在通过模型可解释性来解决我们的主要问题时,我们采用了一个具有Shapley值的特征评估循环来迭代地细化特征集。这个过程在模型性能和透明度之间取得了平衡,从而能够全面了解模型的预测。所提出的方法提供了几个优点,包括识别和删除不相关和冗余的功能,导致一个更有效的模型。此外,它还提高了可解释性,促进了对模型预测的理解,并识别了情感确定的关键特征。在多伦多情感语音集(TESS)、柏林情感语音数据库(EMO-DB)、瑞尔森情感语音和歌曲视听数据库(RAVDESS)和萨里视听表达情感(SAVEE)数据集的SER基准测试上,验证了该方法的有效性,性能优于现有方法。这些结果突出了所提出的技术在开发准确和可解释的SER系统的潜力。据我们所知,这是第一个将模型可解释性纳入SER框架的工作。摘要:Speech emotion recognition (SER) has gained significant attention due to its several application fields, such as mental health, education, and human-computer interaction. However, the accuracy of SER systems is hindered by high-dimensional feature sets that may contain irrelevant and redundant information. To overcome this challenge, this study proposes an iterative feature boosting approach for SER that emphasizes feature relevance and explainability to enhance machine learning model performance. Our approach involves meticulous feature selection and analysis to build efficient SER systems. In addressing our main problem through model explainability, we employ a feature evaluation loop with Shapley values to iteratively refine feature sets. This process strikes a balance between model performance and transparency, which enables a comprehensive understanding of the model's predictions. The proposed approach offers several advantages, including the identification and removal of irrelevant and redundant features, leading to a more effective model. Additionally, it promotes explainability, facilitating comprehension of the model's predictions and the identification of crucial features for emotion determination. The effectiveness of the proposed method is validated on the SER benchmarks of the Toronto emotional speech set (TESS), Berlin Database of Emotional Speech (EMO-DB), Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS), and Surrey Audio-Visual Expressed Emotion (SAVEE) datasets, outperforming state-of-the-art methods. These results highlight the potential of the proposed technique in developing accurate and explainable SER systems. To the best of our knowledge, this is the first work to incorporate model explainability into an SER framework.

【17】 SimpleSpeech: Towards Simple and Efficient Text-to-Speech with Scalar Latent Transformer Diffusion Models
标题: SimpleSpeech:利用纯量潜在Transformer扩散模型实现简单有效的文本到语音
作者:Dongchao Yang,Dingdong Wang,Haohan Guo,Xueyuan Chen,Xixin Wu,Helen Meng
备注:Accepted by InterSpeech 2024
链接:点击下载PDF文件
摘要:在这项研究中,我们提出了一个简单而有效的非自回归(NAR)的文本到语音(TTS)系统的基础上扩散,命名为SimpleSpeech。它的简单性表现在三个方面:(1)它可以在纯语音数据集上训练,没有任何对齐信息;(2)它直接以纯文本作为输入,通过NAR方式生成语音;(3)它试图在有限且紧凑的潜在空间中对语音进行建模,这减轻了扩散建模的难度。更具体地说,我们提出了一种新的语音编解码模型(SQ-Codec)与标量量化,SQ-Codec有效地将复杂的语音信号映射到一个有限的和紧凑的潜在空间,命名为标量潜在空间。借鉴SQ-Codec的思想,我们在SQ-Codec的标量潜在空间中应用了一种新的Transformer扩散模型。我们在4k小时的纯语音数据集上对SimpleSpeech进行了训练,它表现出了自然的韵律和语音克隆能力。与以往的大规模TTS模型相比,该模型在语音质量和生成速度上都有显著的提高。Demos已发布。摘要:In this study, we propose a simple and efficient Non-Autoregressive (NAR) text-to-speech (TTS) system based on diffusion, named SimpleSpeech. Its simpleness shows in three aspects: (1) It can be trained on the speech-only dataset, without any alignment information; (2) It directly takes plain text as input and generates speech through an NAR way; (3) It tries to model speech in a finite and compact latent space, which alleviates the modeling difficulty of diffusion. More specifically, we propose a novel speech codec model (SQ-Codec) with scalar quantization, SQ-Codec effectively maps the complex speech signal into a finite and compact latent space, named scalar latent space. Benefits from SQ-Codec, we apply a novel transformer diffusion model in the scalar latent space of SQ-Codec. We train SimpleSpeech on 4k hours of a speech-only dataset, it shows natural prosody and voice cloning ability. Compared with previous large-scale TTS models, it presents significant speech quality and generation speed improvement. Demos are released.

【18】 An Independence-promoting Loss for Music Generation with Language Models
标题: 语言模型音乐生成的独立性丧失
作者:Jean-Marie Lemercier,Simon Rouard,Jade Copet,Yossi Adi,Alexandre Déffosez
备注:Accepted to ICML 2024
链接:点击下载PDF文件
摘要:使用语言建模的音乐生成方案依赖于音频令牌的词汇表,通常作为由自动编码器学习的离散潜在空间中的代码提供。通常采用多级量化器来产生这些令牌,因此用于令牌预测的解码策略必须适应多个码本:它应该对所有码本上的联合分布进行建模,或者拟合码本边缘分布的乘积。对联合分布进行建模需要增加自回归步骤的数量,而拟合边际的乘积会产生不精确的模型,除非码本相互独立。在这项工作中,我们引入了一个独立的促进损失正则化的自动编码器作为标记的语言模型的音乐生成。建议的损失是基于最大平均差异原则的互信息的代理,应用于可再生核希尔伯特空间。该准则易于实现和训练,并可推广到其他多流编解码器。我们表明,它减少了自动编码过程中码本之间的统计依赖性。这导致在对边际分布的乘积进行建模时所生成的音乐质量的增加,同时比联合分布模型更快地生成音频。摘要:Music generation schemes using language modeling rely on a vocabulary of audio tokens, generally provided as codes in a discrete latent space learnt by an auto-encoder. Multi-stage quantizers are often employed to produce these tokens, therefore the decoding strategy used for token prediction must be adapted to account for multiple codebooks: either it should model the joint distribution over all codebooks, or fit the product of the codebook marginal distributions. Modelling the joint distribution requires a costly increase in the number of auto-regressive steps, while fitting the product of the marginals yields an inexact model unless the codebooks are mutually independent. In this work, we introduce an independence-promoting loss to regularize the auto-encoder used as the tokenizer in language models for music generation. The proposed loss is a proxy for mutual information based on the maximum mean discrepancy principle, applied in reproducible kernel Hilbert spaces. Our criterion is simple to implement and train, and it is generalizable to other multi-stream codecs. We show that it reduces the statistical dependence between codebooks during auto-encoding. This leads to an increase in the generated music quality when modelling the product of the marginal distributions, while generating audio much faster than the joint distribution model.

【19】 Audio Mamba: Selective State Spaces for Self-Supervised Audio Representations
标题: 音频曼巴:自我监督音频表示的选择性状态空间
作者:Sarthak Yadav,Zheng-Hua Tan
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:尽管Transformer作为突出的神经架构被广泛采用,但它已经激发了几个独立的工作线来解决其局限性。其中一种方法是选择性状态空间模型,它已经证明了语言建模的有前途的结果。然而,他们的可行性学习自我监督,通用的音频表示还有待研究。这项工作提出了音频曼巴,一个选择性的状态空间模型学习通用的音频表示随机掩蔽频谱补丁通过自我监督。十个不同的音频识别下游任务的实证结果表明,所提出的模型,在AudioSet数据集上进行预训练,始终优于可比的自监督音频频谱图Transformer(SSAST)基线相当大的幅度,并在数据集大小,序列长度和模型大小比较中表现出更好的性能。摘要:Despite its widespread adoption as the prominent neural architecture, the Transformer has spurred several independent lines of work to address its limitations. One such approach is selective state space models, which have demonstrated promising results for language modelling. However, their feasibility for learning self-supervised, general-purpose audio representations is yet to be investigated. This work proposes Audio Mamba, a selective state space model for learning general-purpose audio representations from randomly masked spectrogram patches through self-supervision. Empirical results on ten diverse audio recognition downstream tasks show that the proposed models, pretrained on the AudioSet dataset, consistently outperform comparable self-supervised audio spectrogram transformer (SSAST) baselines by a considerable margin and demonstrate better performance in dataset size, sequence length and model size comparisons.

【20】 Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
标题: Whistle:通过弱语音监督实现数据高效的多语言和跨语言语音识别
作者:Saierdaer Yusuyin,Te Ma,Hao Huang,Wenbo Zhao,Zhijian Ou
链接:点击下载PDF文件
摘要:多语言和跨语言自动语音识别(MCL-ASR)有三种方法-语音或字形转录的监督预训练和自我监督预训练。我们发现,到目前为止,语音监督的预训练对于MCL-ASR来说一直没有得到充分的重视,而从概念上讲,它更有利于不同语言之间的信息共享。本文探讨了一种基于弱语音监督的预训练方法,即Whistle。我们放宽了黄金标准的人类验证的语音成绩单的要求,并获得国际音标(IPA)的基础上的转录,通过利用philageNet字素到音素(G2 P)模型。我们基于CommonVoice数据集构建了一个通用的实验设置,称为CV-Lang 10,其中包含10种可见语言和2种不可见语言。在CV-Lang 10上进行了一组实验,以尽可能公平地比较MCL-ASR常见设置下的三种方法。实验证明了基于音素的MCL-ASR模型(Whistle)的优势,在已知语言的语音识别、不同数量的Few-Shot数据的未知语言的跨语言性能、克服灾难性遗忘和训练效率方面,发现当训练数据更有限时,音素监督可以获得比子词监督和自我监督更好的结果,从而提供更高的数据效率。为了支持可重复性并促进沿着这个方向的未来研究,我们将在发布后在https: github.com thu-spmi CAT上发布整个Whistle管道的代码,模型和数据。摘要:There exist three approaches for multilingual and crosslingual automatic speech recognition (MCL-ASR) - supervised pre-training with phonetic or graphemic transcription, and self-supervised pre-training. We find that pre-training with phonetic supervision has been underappreciated so far for MCL-ASR, while conceptually it is more advantageous for information sharing between different languages. This paper explores the approach of pre-training with weakly phonetic supervision towards data-efficient MCL-ASR, which is called Whistle. We relax the requirement of gold-standard human-validated phonetic transcripts, and obtain International Phonetic Alphabet (IPA) based transcription by leveraging the LanguageNet grapheme-to-phoneme (G2P) models. We construct a common experimental setup based on the CommonVoice dataset, called CV-Lang10, with 10 seen languages and 2 unseen languages. A set of experiments are conducted on CV-Lang10 to compare, as fair as possible, the three approaches under the common setup for MCL-ASR. Experiments demonstrate the advantages of phoneme-based models (Whistle) for MCL-ASR, in terms of speech recognition for seen languages, crosslingual performance for unseen languages with different amounts of few-shot data, overcoming catastrophic forgetting, and training efficiency.It is found that when training data is more limited, phoneme supervision can achieve better results compared to subword supervision and self-supervision, thereby providing higher data-efficiency. To support reproducibility and promote future research along this direction, we will release the code, models and data for the whole pipeline of Whistle at https: github.com thu-spmi CAT upon publication.

【21】 MaskSR: Masked Language Model for Full-band Speech Restoration
标题: MaskSR:全频段语音恢复的掩蔽语言模型
作者:Xu Li,Qirui Wang,Xiaoyu Liu
备注:Accepted by INTERSPEECH 2024. Demo page: this https URL
链接:点击下载PDF文件
摘要:语音恢复的目的是在存在各种失真的情况下恢复高质量的语音。尽管已经研究了几种深度学习范式来完成这项任务,但最近新兴的语言模型的力量还没有得到充分的探索。在本文中,我们提出了MaskSR,一个掩蔽的语言模型,能够恢复全频带44.1 kHz的语音联合考虑噪声,混响,削波和低带宽。MaskSR与使用预训练的神经编解码器提取的离散声学标记一起工作。在训练过程中,MaskSR被优化以预测从高质量目标语音中提取的随机掩蔽令牌,条件是具有各种失真的受损语音。在推理过程中,MaskSR利用高效的迭代采样重构目标语音标记。大量的实验表明,MaskSR获得竞争力的结果,无论是全频带语音恢复任务,也对子任务相比,广泛的模型。摘要:Speech restoration aims at restoring high quality speech in the presence of a diverse set of distortions. Although several deep learning paradigms have been studied for this task, the power of the recently emerging language models has not been fully explored. In this paper, we propose MaskSR, a masked language model capable of restoring full-band 44.1 kHz speech jointly considering noise, reverb, clipping, and low bandwidth. MaskSR works with discrete acoustic tokens extracted using a pre-trained neural codec. During training, MaskSR is optimized to predict randomly masked tokens extracted from the high quality target speech, conditioned on the corrupted speech with various distortions. During inference, MaskSR reconstructs the target speech tokens with efficient iterative sampling. Extensive experiments show that MaskSR obtains competitive results on both the full-band speech restoration task and also on sub-tasks compared with a wide range of models.

【22】 Understanding Auditory Evoked Brain Signal via Physics-informed Embedding Network with Multi-Task Transformer
标题: 通过具有多任务Transformer的物理信息嵌入网络理解听觉诱发的大脑信号
作者:Wanli Ma,Xuegang Tang,Jin Gu,Ying Wang,Yuling Xia
链接:点击下载PDF文件
摘要:在脑机交互和认知神经科学领域,有效解码来自基于任务的功能磁共振成像(fMRI)的听觉信号是理解大脑如何处理复杂听觉信息的关键。虽然现有的方法具有增强的解码能力,但在信息利用和模型表示方面仍然存在限制。为了克服这些挑战,我们提出了一种创新的多任务学习模型,即具有多任务Transformer的物理信息嵌入网络(PEMT-Net),它通过物理信息嵌入和深度学习技术来增强解码性能。PEMT-Net由两个主要组成部分组成:特征增强和分类。对于特征增强,我们提出了一种新的方法,通过节点嵌入创建神经嵌入图,利用随机游走来模拟神经信息的物理扩散。该方法捕获局部和非局部信息溢出,并提出了一种基于相对物理坐标的位置编码。在分类部分,我们提出了自适应嵌入融合,以最大限度地捕捉线性和非线性特征。此外,我们提出了一个创新的参数共享机制,以优化提取的特征的保留和学习。在特定数据集上的实验证明了PEMT-Net在多任务听觉信号解码方面的显著性能,超越了现有方法,并为大脑处理复杂听觉信息的机制提供了新的见解。摘要:In the fields of brain-computer interaction and cognitive neuroscience, effective decoding of auditory signals from task-based functional magnetic resonance imaging (fMRI) is key to understanding how the brain processes complex auditory information. Although existing methods have enhanced decoding capabilities, limitations remain in information utilization and model representation. To overcome these challenges, we propose an innovative multi-task learning model, Physics-informed Embedding Network with Multi-Task Transformer (PEMT-Net), which enhances decoding performance through physics-informed embedding and deep learning techniques. PEMT-Net consists of two principal components: feature augmentation and classification. For feature augmentation, we propose a novel approach by creating neural embedding graphs via node embedding, utilizing random walks to simulate the physical diffusion of neural information. This method captures both local and non-local information overflow and proposes a position encoding based on relative physical coordinates. In the classification segment, we propose adaptive embedding fusion to maximally capture linear and non-linear characteristics. Furthermore, we propose an innovative parameter-sharing mechanism to optimize the retention and learning of extracted features. Experiments on a specific dataset demonstrate PEMT-Net's significant performance in multi-task auditory signal decoding, surpassing existing methods and offering new insights into the brain's mechanisms for processing complex auditory information.

【23】 Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping
标题: 有效训练可减少小型化并通过逐核剪裁表现更好的ASB模型
作者:Lun Wang,Om Thakkar,Zhong Meng,Nicole Rafidi,Rohit Prabhavalkar,Arun Narayanan
链接:点击下载PDF文件
摘要:梯度裁剪在大规模自动语音识别(ASR)模型的训练中起着至关重要的作用。它通常应用于小批量梯度以防止梯度爆炸,并应用于单个样本梯度以减轻无意记忆。这项工作系统地研究了特定粒度的梯度裁剪,即每核心裁剪(PCC),在训练各种ASR模型的影响。我们的经验表明,PCC可以有效地减轻非故意的记忆在ASR模型。令人惊讶的是,我们发现PCC积极影响ASR性能指标,从而提高收敛速度和降低字错误率。为了避免调整PCC引入的额外超参数,我们进一步提出了一种新的变体,自适应每核裁剪(APCC),用于流线型优化。我们的研究结果强调了PCC作为强大的隐私前瞻性ASR模型培训策略的多方面好处。摘要:Gradient clipping plays a vital role in training large-scale automatic speech recognition (ASR) models. It is typically applied to minibatch gradients to prevent gradient explosion, and to the individual sample gradients to mitigate unintended memorization. This work systematically investigates the impact of a specific granularity of gradient clipping, namely per-core clip-ping (PCC), across training a wide range of ASR models. We empirically demonstrate that PCC can effectively mitigate unintended memorization in ASR models. Surprisingly, we find that PCC positively influences ASR performance metrics, leading to improved convergence rates and reduced word error rates. To avoid tuning the additional hyperparameter introduced by PCC, we further propose a novel variant, adaptive per-core clipping (APCC), for streamlined optimization. Our findings highlight the multifaceted benefits of PCC as a strategy for robust, privacy-forward ASR model training.

【24】 TinySV: Speaker Verification in TinyML with On-device Learning
标题: TinySV:TinyML中的说话者验证,通过设备上学习
作者:Massimo Pavan,Gioele Mombelli,Francesco Sinacori,Manuel Roveri
链接:点击下载PDF文件
摘要:TinyML是机器学习的一个新领域,由于能够在微型设备(如物联网或嵌入式系统)上执行机器学习算法,在过去几年中获得了巨大的发展势头。有趣的是,这一领域的研究集中在微型设备上TinyML模型推理阶段的有效执行上,而由于学习算法引入的相关开销,文献中很少有TinyML模型的设备学习解决方案。 本文的目的是介绍一种新型的自适应TinyML解决方案,可用于任务,如提出的 textit{Tiny Speaker Verification}(TinySV),需要用设备上的学习算法来解决。实现这一目标需要(i)减少TinyML学习算法的内存和计算需求,以及(ii)设计一种使用少量且可能未标记的训练数据的TinyML学习算法。所提出的TinySV解决方案依赖于两层分层TinyML解决方案,包括关键字定位和自适应说话人验证模块。我们在专门为此任务收集的数据集上评估了所提出的TinySV解决方案的有效性和效率,并在真实的物联网设备(英飞凌PSoC 62 S2 Wi-Fi BT Pioneer Kit)上测试了所提出的解决方案。摘要:TinyML is a novel area of machine learning that gained huge momentum in the last few years thanks to the ability to execute machine learning algorithms on tiny devices (such as Internet-of-Things or embedded systems). Interestingly, research in this area focused on the efficient execution of the inference phase of TinyML models on tiny devices, while very few solutions for on-device learning of TinyML models are available in the literature due to the relevant overhead introduced by the learning algorithms. The aim of this paper is to introduce a new type of adaptive TinyML solution that can be used in tasks, such as the presented textit{Tiny Speaker Verification} (TinySV), that require to be tackled with an on-device learning algorithm. Achieving this goal required (i) reducing the memory and computational demand of TinyML learning algorithms, and (ii) designing a TinyML learning algorithm operating with few and possibly unlabelled training data. The proposed TinySV solution relies on a two-layer hierarchical TinyML solution comprising Keyword Spotting and Adaptive Speaker Verification module. We evaluated the effectiveness and efficiency of the proposed TinySV solution on a dataset collected expressly for the task and tested the proposed solution on a real-world IoT device (Infineon PSoC 62S2 Wi-Fi BT Pioneer Kit).


机器翻译,仅供参考