今天跟大家分享一篇语音相关的论文合集:cs.SD语音18篇,eess.AS音频处理19篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily


cs.SD语音

【1】 Individualized Conditioning and Negative Distances for Speaker  Separation

标题:说话人分离的个别化条件作用和负距离

链接:https://arxiv.org/abs/2210.06368

作者:Tao Sun,Nidal Abuhajar,Shuyu Gong,Zhewei Wang,Charles D. Smith,Xianhui Wang,Li Xu,Jundong Liu
机构:∗School of Electrical Engineering and Computer Science, Ohio University, Athens, OH , †Department of Neurology, University of Kentucky, Lexington, KY , ‡Division of Communication Sciences, Ohio University, Athens, OH
备注:Accepted to ICMLA 2022
摘要:说话人分离的目的是从混合信号中提取多个语音。本文提出了两种说话人感知的说话人分离方案,以改进现有的说话人分离方案。第一种模型是说话人调节网络,其集成语音样本以生成个性化的说话人条件,然后个性化的说话人条件为分离模块提供有根据的指导以产生分离良好的输出。  第二种设计旨在减少分离语音中的非目标语音。为此,我们提出负距离来惩罚通道输出中任何非目标语音的出现,而正距离使分离后的语音更接近干净目标。我们探索了两种不同的设置,加权和和三重类,以整合这两个距离,形成分离网络的组合辅助损耗。在LibriMix上进行的实验验证了所提模型的有效性.
摘要:Speaker separation aims to extract multiple voices from a mixed signal. In this paper, we propose two speaker-aware designs to improve the existing speaker separation solutions. The first model is a speaker conditioning network that integrates speech samples to generate individualized speaker conditions, which then provide informed guidance for a separation module to produce well-separated outputs.  The second design aims to reduce non-target voices in the separated speech. To this end, we propose negative distances to penalize the appearance of any non-target voice in the channel outputs, and positive distances to drive the separated voices closer to the clean targets. We explore two different setups, weighted-sum and triplet-like, to integrate these two distances to form a combined auxiliary loss for the separation networks. Experiments conducted on LibriMix demonstrate the effectiveness of our proposed models.


【2】 Text-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption  Similarity

标题:基于文本-音频基础的音频字幕相似度评价新指标

链接:https://arxiv.org/abs/2210.06354

作者:Swapnil Bhosale,Rupayan Chakraborty,Sunil Kumar Kopparapu
机构:TCS Research, Tata Consultancy Services Limited, India.
备注:9 pages, 8 figures,
摘要:自动音频字幕(AAC)指的是将音频样本翻译成描述音频事件、事件源及其关系的自然语言(NL)文本的任务。与NL文本生成任务(其依赖于基于词汇语义的度量(如BLEU、ROUGE、METEOR)来进行评估)不同,AAC评估度量需要映射NL文本(短语)的能力,所述NL文本对应于除词汇语义之外的类似声音。当前用于评估AAC任务的度量缺乏对由文本表示的声音的感知属性的理解。本文提出了一种新的基于文本到音频基础(TAG)的评价指标,该指标对跨模态任务如AAC的评价非常有用。在公开的AAC数据集上的实验表明,与NL文本和图像字幕文献中使用的现有评价指标相比,本文提出的评价指标具有更好的性能。
摘要:Automatic Audio Captioning (AAC) refers to the task of translating an audio sample into a natural language (NL) text that describes the audio events, source of the events and their relationships. Unlike NL text generation tasks, which rely on metrics like BLEU, ROUGE, METEOR based on lexical semantics for evaluation, the AAC evaluation metric requires an ability to map NL text (phrases) that correspond to similar sounds in addition lexical semantics. Current metrics used for evaluation of AAC tasks lack an understanding of the perceived properties of sound represented by text. In this paper, wepropose a novel metric based on Text-to-Audio Grounding (TAG), which is, useful for evaluating cross modal tasks like AAC. Experiments on publicly available AAC data-set shows our evaluation metric to perform better compared to existing metrics used in NL text and image captioning literature.


【3】 SQuId: Measuring Speech Naturalness in Many Languages

标题:SQUID:测量多种语言的语言自然度

链接:https://arxiv.org/abs/2210.06324

作者:Thibault Sellam,Ankur Bapna,Joshua Camp,Diana Mackinnon,Ankur P. Parikh,Jason Riesa
机构:Google
摘要:许多文本到语音的研究依赖于人工评估,这导致了巨大的成本,并减缓了开发过程。这个问题在大量使用多种语言的应用程序中尤其严重,在这些应用程序中,招聘和投票法官可能需要数周的时间。我们介绍SQuId(语音质量识别),这是一个多语言自然度预测模型,它在超过100万个评分上进行了训练,并在65个地区进行了测试--这是迄今为止此类研究中最大的一次。主要的观点是,在多个语言环境上训练一个模型始终优于单一语言环境基线。我们介绍了我们的任务和模型,并表明它的性能比基于w2 v-BERT和VoiceMOS的竞争基准高50.0%。然后,我们证明了微调过程中跨区域转换的有效性,并强调了其对零炮点区域的影响,即:没有微调数据的语言环境。通过一系列的分析,我们强调了非语言效应如声音假象在跨语言环境迁移中的作用。最后,我们介绍了我们的设计决策的效果,例如,模型大小、预训练多样性和语言再平衡。
摘要:Much of text-to-speech research relies on human evaluation, which incurs heavy costs and slows down the development process. The problem is particularly acute in heavily multilingual applications, where recruiting and polling judges can take weeks. We introduce SQuId (Speech Quality Identification), a multilingual naturalness prediction model trained on over a million ratings and tested in 65 locales-the largest effort of this type to date. The main insight is that training one model on many locales consistently outperforms mono-locale baselines. We present our task, the model, and show that it outperforms a competitive baseline based on w2v-BERT and VoiceMOS by 50.0%. We then demonstrate the effectiveness of cross-locale transfer during fine-tuning and highlight its effect on zero-shot locales, i.e., locales for which there is no fine-tuning data. Through a series of analyses, we highlight the role of non-linguistic effects such as sound artifacts in cross-locale transfer. Finally, we present the effect of our design decision, e.g., model size, pre-training diversity, and language rebalancing with several ablation experiments.


【4】 A context-aware knowledge transferring strategy for CTC-based ASR

标题:一种基于CTC的ASR上下文感知知识转移策略

链接:https://arxiv.org/abs/2210.06244

作者:Ke-Han Lu,Kuan-Yu Chen
机构:National Taiwan University of Science and Technology, Taiwan
备注:Accepted by SLT 2022
摘要:非自回归自动语音识别(ASR)模型因其解码速度快、性能优越而受到越来越多的关注。其中,基于连接主义时态分类(CTC)的方法仍然是主流。然而,理论上的固有缺陷--符号之间的独立性假设--为作品派的表演设置了障碍。针对这一问题,本文提出了一种基于上下文感知的ASR知识转移策略,该策略由知识转移模块和上下文感知的训练策略组成。前者旨在从预先训练好的语言模型中提取语言信息,后者旨在调整条件独立假设所带来的局限性。在此基础上,本文提出了一种基于wav 2 vec 2. 0的知识注入的上下文感知的基于CTC的ASR。在AISHELL-1和AISHELL-2数据集上的一系列实验验证了该方法的有效性.
摘要:Non-autoregressive automatic speech recognition (ASR) modeling has received increasing attention recently because of its fast decoding speed and superior performance. Among representatives, methods based on the connectionist temporal classification (CTC) are still a dominating stream. However, the theoretically inherent flaw, the assumption of independence between tokens, creates a performance barrier for the school of works. To mitigate the challenge, we propose a context-aware knowledge transferring strategy, consisting of a knowledge transferring module and a context-aware training strategy, for CTC-based ASR. The former is designed to distill linguistic information from a pre-trained language model, and the latter is framed to modulate the limitations caused by the conditional independence assumption. As a result, a knowledge-injected context-aware CTC-based ASR built upon the wav2vec2.0 is presented in this paper. A series of experiments on the AISHELL-1 and AISHELL-2 datasets demonstrate the effectiveness of the proposed method.


【5】 Towards visually prompted keyword localisation for zero-resource spoken  languages

标题:面向零资源口语的视觉提示关键词本地化

链接:https://arxiv.org/abs/2210.06229

作者:Leanne Nortje,Herman Kamper
机构:MediaLab, Electrical & Electronic Engineering, Stellenbosch University, South Africa
备注:Accepted to IEEE SLT 2022
摘要:想象一下,能够向系统显示关键字的可视描述,并从零资源语音语料库中找到包含该关键字的口语话语。我们将此任务形式化,并称之为视觉提示关键字本地化(VPKL):给定关键字的图像,检测并预测关键字在话语中的何处出现。为了进行VPKL,我们提出了一个具有新颖的局部注意机制的语音-视觉模型,我们用一个新的关键词采样方案训练该模型。我们表明,这些创新在VPKL中提供了优于现有语音-视觉模型的改进。我们还比较了视觉词袋(BoW)模型,其中图像自动标记视觉标签,并与未标记的语音配对。尽管可以使用书面关键字直接查询该可视BoW(而我们的采用图像查询),但我们的新模型在检测和定位方面仍优于可视BoW,在定位F1方面相对提高了16%。
摘要:Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this task and call it visually prompted keyword localisation (VPKL): given an image of a keyword, detect and predict where in an utterance the keyword occurs. To do VPKL, we propose a speech-vision model with a novel localising attention mechanism which we train with a new keyword sampling scheme. We show that these innovations give improvements in VPKL over an existing speech-vision model. We also compare to a visual bag-of-words (BoW) model where images are automatically tagged with visual labels and paired with unlabelled speech. Although this visual BoW can be queried directly with a written keyword (while our's takes image queries), our new model still outperforms the visual BoW in both detection and localisation, giving a 16% relative improvement in localisation F1.


【6】 VCSE: Time-Domain Visual-Contextual Speaker Extraction Network

标题:VCSE:时间域视觉语境说话人提取网络

链接:https://arxiv.org/abs/2210.06177

作者:Junjie Li,Meng Ge,Zexu Pan,Longbiao Wang,Jianwu Dang
机构:Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University, Tianjin, China,  Department of Electrical and Computer Engineering, National University of Singapore, Singapore
摘要:说话人提取寻求在给定辅助参考的多说话人场景中提取目标语音。这样的参考可以是听觉的,即,预先记录的语音、视觉,即,嘴唇运动或上下文,即,语音序列。不同模态中的指称提供了不同的和互补的信息,这些信息可以被融合以形成对目标说话人的自上而下的注意。先前的研究已经在单一模型中引入了视觉和情境模态。本文提出了一种两级时域视觉语境说话人提取网络VCSE,该网络将视觉语境线索和自注册语境线索逐级融合,充分利用了各个模态的优势。在第一阶段,我们利用视觉线索预先撷取目标语音,并估计潜在的语音序列。在第二阶段,我们使用自注册的上下文线索对预提取的目标语音进行精炼。在真实的唇读句子3(LRS3)数据库上的实验结果表明,本文提出的VCSE网络的性能始终优于其他现有的基线.
摘要:Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence. References in different modalities provide distinct and complementary information that could be fused to form top-down attention on the target speaker. Previous studies have introduced visual and contextual modalities in a single model. In this paper, we propose a two-stage time-domain visual-contextual speaker extraction network named VCSE, which incorporates visual and self-enrolled contextual cues stage by stage to take full advantage of every modality. In the first stage, we pre-extract a target speech with visual cues and estimate the underlying phonetic sequence. In the second stage, we refine the pre-extracted target speech with the self-enrolled contextual cues. Experimental results on the real-world Lip Reading Sentences 3 (LRS3) database demonstrate that our proposed VCSE network consistently outperforms other state-of-the-art baselines.


【7】 THUEE system description for NIST 2020 SRE CTS challenge

标题:NIST 2020 SRE CTS挑战赛THUEE系统描述

链接:https://arxiv.org/abs/2210.06111

作者:Yu Zheng,Jinghan Peng,Miao Zhao,Yufeng Ma,Min Liu,Xinyue Ma,Tianyu Liang,Tianlong Kong,Liang He,Minqiang Xu
机构:SpeakIn Technologies Co. Ltd., ShangHai, China, Department of Electronic Engineering Tsinghua University, Beijing, China
备注:3 pages, 1 table; System desciption of NIST 2020 SRE CTS challenge
摘要:本文介绍了THUEE团队参加NIST 2020说话人识别评估(SRE)会话电话语音(CTS)挑战赛的系统描述。在本评估中,包括ResNet 74、ResNet 152和RepVGG-B2在内的子系统被开发为扬声器嵌入提取器。我们使用基于AM-Softmax和AAM-Softmax的组合损失函数,即CM-Softmax。我们采用了两阶段训练策略来进一步提高系统性能。我们融合了所有单个系统作为最终提交。我们的方法带来了出色的性能,并在挑战中排名第一。
摘要:This paper presents the system description of the THUEE team for the NIST 2020 Speaker Recognition Evaluation (SRE) conversational telephone speech (CTS) challenge. The subsystems including ResNet74, ResNet152, and RepVGG-B2 are developed as speaker embedding extractors in this evaluation. We used combined AM-Softmax and AAM-Softmax based loss functions, namely CM-Softmax. We adopted a two-staged training strategy to further improve system performance. We fused all individual systems as our final submission. Our approach leads to excellent performance and ranks 1st in the challenge.


【8】 SpecRNet: Towards Faster and More Accessible Audio DeepFake Detection

标题:SPECRNet:向更快、更易访问的音频DeepFake检测迈进

链接:https://arxiv.org/abs/2210.06105

作者:Piotr Kawa,Marcin Plata,Piotr Syga
机构:Department of Artificial Intelligence, Wrocław University of Science and Technology, Wrocław, Poland, ©, IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including
备注:Accepted by TrustCom 2022: The 21st IEEE International Conference on Trust, Security and Privacy in Computing and Communications
摘要:音频DeepFakes是使用深度神经网络生成的话语。它们具有高度误导性,并因用于假新闻、冒充或勒索而构成威胁。在这项工作中,我们通过提供SpecRNet(一种具有快速推理时间和低计算要求的神经网络架构),专注于增加音频DeepFake检测方法的可访问性。我们的基准测试表明,SpecRNet处理音频样本所需的时间最多可减少40%,其性能可与LCNN架构(最佳的音频DeepFake检测模型之一)相媲美。这种方法不仅可以被在线多媒体服务用来验证每天上传的大量内容,而且由于其低要求,还可以被普通公民用来评估他们设备上的材料。此外,我们还提供了三种独特设置的基准测试,以确认我们的模型的正确性。它们反映了低资源数据集、短话语检测和有限攻击基准的场景,其中我们更仔细地观察了特定攻击对给定体系结构的影响。
摘要:Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the audio DeepFake detection methods by providing SpecRNet, a neural network architecture characterized by a quick inference time and low computational requirements. Our benchmark shows that SpecRNet, requiring up to about 40% less time to process an audio sample, provides performance comparable to LCNN architecture - one of the best audio DeepFake detection models. Such a method can not only be used by online multimedia services to verify a large bulk of content uploaded daily but also, thanks to its low requirements, by average citizens to evaluate materials on their devices. In addition, we provide benchmarks in three unique settings that confirm the correctness of our model. They reflect scenarios of low-resource datasets, detection on short utterances and limited attacks benchmark in which we take a closer look at the influence of particular attacks on given architectures.


【9】 Summary on the ISCSLP 2022 Chinese-English Code-Switching ASR Challenge

标题:ISCSLP 2022汉英代码转换ASR挑战赛综述

链接:https://arxiv.org/abs/2210.06091

作者:Shuhao Deng,Chengfei Li,infeng Bai,Qingqing Zhang,Wei-Qiang Zhang,Runyan Yang,Gaofeng Cheng,Pengyuan Zhang,Yonghong Yan
机构:TAL Education Group, Beijing, China ,Magic Data ,Tsinghua University ,Institute of Acoustics, Chinese Academy of Sciences
备注:accepted by ISCSLP 2022
摘要:由于多种语言之间的语码转换现象以及日常生活中频繁发生的语码转换现象,使得语码转换自动语音识别成为自动语音识别中最具挑战性和最有价值的场景之一。ISCSLP 2022汉英语码转换自动语音识别(CSASR)挑战赛旨在推动语码转换自动语音识别的发展。ISCSLP 2022 CSASR挑战赛为参赛者提供了TAL_CSASR语料库和MagicData-RAMC语料库两个训练集、一个开发集和一个测试集,用于CSASR模型训练和评估。除了挑战之外,我们还提供了基准系统性能以供参考。因此,有40多个团队参与了此次挑战,获胜团队在测试集上实现了16.70%的混合错误率(MER)性能,与基线系统相比,MER绝对提升了9.8%。本文将描述数据集、相关基线系统和需求,并总结CSASR挑战结果和提交系统中使用的主要技术和技巧。
摘要:Code-switching automatic speech recognition becomes one of the most challenging and the most valuable scenarios of automatic speech recognition, due to the code-switching phenomenon between multilingual language and the frequent occurrence of code-switching phenomenon in daily life. The ISCSLP 2022 Chinese-English Code-Switching Automatic Speech Recognition (CSASR) Challenge aims to promote the development of code-switching automatic speech recognition. The ISCSLP 2022 CSASR challenge provided two training sets, TAL_CSASR corpus and MagicData-RAMC corpus, a development and a test set for participants, which are used for CSASR model training and evaluation. Along with the challenge, we also provide the baseline system performance for reference. As a result, more than 40 teams participated in this challenge, and the winner team achieved 16.70% Mixture Error Rate (MER) performance on the test set and has achieved 9.8% MER absolute improvement compared with the baseline system. In this paper, we will describe the datasets, the associated baselines system and the requirements, and summarize the CSASR challenge results and major techniques and tricks used in the submitted systems.


【10】 JukeDrummer: Conditional Beat-aware Audio-domain Drum Accompaniment  Generation via Transformer VQ-VA

标题:JukeDrummer:通过TransformerVQ-VA生成条件节拍感知音域鼓伴奏

链接:https://arxiv.org/abs/2210.06007

作者:Yueh-Kao Wu,Ching-Yu Chiu,Yi-Hsuan Yang
机构:Academia Sinica, National Cheng Kung University, Taiwan AI Labs
备注:Accepted at ISMIR 2022
摘要:本文提出了一种在音频域中生成鼓音轨的模型,以与用户提供的无鼓录音一起播放。具体来说,使用无鼓轨道和相应的人造鼓轨道的配对数据,我们训练一个Transformer模型来即兴演奏一个看不见的无鼓录音的鼓部分。我们结合两种方法来编码输入音频。首先,我们训练一个向量量化变分自动编码器(VQ-VAE),用离散码表示输入音频,然后可以在Transformer中使用。第二,使用音频域节拍跟踪模型,我们计算输入音频的节拍相关特征,并将其作为Transformer中的嵌入。我们使用单独的VQ-VAE将鼓轨道的Mel谱图编码为另一组离散码,并训练Transformer预测与鼓相关的离散码的序列,而不是将鼓轨道直接生成为波形。然后,用解码器将输出码转换成梅尔谱图,然后用声码器将其转换成波形。我们报告了对所提出的模型的变体的客观和主观评估,证明了具有节拍信息的模型生成在节奏和风格上与输入音频一致的鼓伴奏。
摘要:This paper proposes a model that generates a drum track in the audio domain to play along to a user-provided drum-free recording. Specifically, using paired data of drumless tracks and the corresponding human-made drum tracks, we train a Transformer model to improvise the drum part of an unseen drumless recording. We combine two approaches to encode the input audio. First, we train a vector-quantized variational autoencoder (VQ-VAE) to represent the input audio with discrete codes, which can then be readily used in a Transformer. Second, using an audio-domain beat tracking model, we compute beat-related features of the input audio and use them as embeddings in the Transformer. Instead of generating the drum track directly as waveforms, we use a separate VQ-VAE to encode the mel-spectrogram of a drum track into another set of discrete codes, and train the Transformer to predict the sequence of drum-related discrete codes. The output codes are then converted to a mel-spectrogram with a decoder, and then to the waveform with a vocoder. We report both objective and subjective evaluations of variants of the proposed model, demonstrating that the model with beat information generates drum accompaniment that is rhythmically and stylistically consistent with the input audio.


【11】 Enemy Spotted: in-game gun sound dataset for gunshot classification and  localization

标题:发现敌人:用于枪击分类和定位的游戏中枪声数据集

链接:https://arxiv.org/abs/2210.05917

作者:Junwoo Park,Youngwoo Cho,Gyuhyeon Sim,Hojoon Lee,Jaegul Choo
机构:Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
备注:Accepted at IEEE Conference on Games (GoG) 2022
摘要:近年来,基于深度学习的方法在声音分类和定位中由于其简单高效、无需领域知识等优点而受到广泛关注。然而,现有数据集中缺乏枪声一直是实施支持系统通过利用深度学习模型从枪声中发现罪犯的主要障碍。由于枪声的发生是罕见的和不可预测的,因此在现实世界中收集枪声是不切实际的。作为替代,可以从被设计为模拟真实世界战争的FPS游戏中获得枪声。最近的FPS游戏提供了一个现实的环境,我们可以安全地收集射击数据,同时模拟甚至危险的情况。利用游戏环境的优势,构建了一个枪械数据集BGG,用于枪械分类和枪械定位任务。BGG数据集包括37种不同类型的枪械、距离以及声源和接收器之间的方向。通过在BGG数据集上训练多个声音分类和定位基线,我们仔细验证了游戏中的枪声数据具有足够的信息来识别枪声的位置和类型。最后,我们证明了利用BGG数据集可以提高真实世界枪支分类和定位任务的准确性。
摘要:Recently, deep learning-based methods have drawn huge attention due to their simple yet high performance without domain knowledge in sound classification and localization tasks. However, a lack of gun sounds in existing datasets has been a major obstacle to implementing a support system to spot criminals from their gunshots by leveraging deep learning models. Since the occurrence of gunshot is rare and unpredictable, it is impractical to collect gun sounds in the real world. As an alternative, gun sounds can be obtained from an FPS game that is designed to mimic real-world warfare. The recent FPS game offers a realistic environment where we can safely collect gunshot data while simulating even dangerous situations. By exploiting the advantage of the game environment, we construct a gunshot dataset, namely BGG, for the firearm classification and gunshot localization tasks. The BGG dataset consists of 37 different types of firearms, distances, and directions between the sound source and a receiver. We carefully verify that the in-game gunshot data has sufficient information to identify the location and type of gunshots by training several sound classification and localization baselines on the BGG dataset. Afterward, we demonstrate that the accuracy of real-world firearm classification and localization tasks can be enhanced by utilizing the BGG dataset.


【12】 Comparison of Soft and Hard Target RNN-T Distillation for Large-scale  ASR

标题:大型ASR软硬靶RNN-T精馏的比较

链接:https://arxiv.org/abs/2210.05793

作者:Dongseong Hwang,Khe Chai Sim,Yu Zhang,Trevor Strohman
机构:Google LLC, USA
备注:8 pages, 1 figure
摘要:知识提炼是一种有效的机器学习技术,它可以将知识从教师模型转移到更小的学生模型,特别是在未标记数据的情况下。本文主要研究了在自动语音识别中广泛应用的RNN-T模型的知识提取问题。具体而言,我们比较了使用软目标和硬目标提取在LibriSpeech/LibriLight公共数据集(60 k小时)和我们的内部数据(600 k小时)上训练大规模RNN-T模型的情况。我们发现,当教师和学生的结构不同时,如大教师和小分流学生,硬目标更有效。另一方面,软目标蒸馏在自我培训场景中效果更好,比如迭代的大型教师培训。对于具有0.6B权重的大模型,我们使用带软目标蒸馏的噪声学生训练在LibriSpeech上实现了新的SoTA单词错误率(WER)(相对于dev-other提高8%)。它还允许我们的生产教师不断适应新的数据域。
摘要:Knowledge distillation is an effective machine learning technique to transfer knowledge from a teacher model to a smaller student model, especially with unlabeled data. In this paper, we focus on knowledge distillation for the RNN-T model, which is widely used in state-of-the-art (SoTA) automatic speech recognition (ASR). Specifically, we compared using soft and hard target distillation to train large-scaleRNN-T models on the LibriSpeech/LibriLight public dataset (60k hours) and our in-house data (600k hours). We found that hard tar-gets are more effective when the teacher and student have different architecture, such as large teacher and small streaming student. On the other hand, soft target distillation works better in self-training scenario like iterative large teacher training. For a large model with0.6B weights, we achieve a new SoTA word error rate (WER) on LibriSpeech (8% relative improvement on dev-other) using Noisy Student Training with soft target distillation. It also allows our production teacher to adapt new data domain continuously.


【13】 Scaling Up Deliberation for Multilingual ASR

标题:扩大对多语言ASR的审议

链接:https://arxiv.org/abs/2210.05785

作者:Ke Hu,Bo Li,Tara N. Sainath
机构:Google LLC, USA
摘要:多语言端到端自动语音识别模型由于其训练和部署简单而具有吸引力。与单语模型相比,最近对这种模型的大规模训练的工作已经显示出有希望的结果。然而,在单遍设置中,工作通常集中在多语言模型本身。本文主要研究多语言语音识别中的二次通过审议问题。我们提议的审议是多语种的,文本编码器对来自多种语言的假设文本进行编码,而解码器处理多语言文本和音频。研究了审议文本编码器和解码器的可伸缩性,并比较了审议解码器和第一遍级联编码器的可伸缩性。我们发现,与单遍模型相比,审议在9种语言上的平均WER提高了4%。通过将审议的大小增加到1B个参数,平均WER改善增加到9%,对于某些语言最高可达14%。我们的审议重新评分器基于Transformer层,并且可以在重新评分期间并行化。
摘要:Multilingual end-to-end automatic speech recognition models are attractive due to its simplicity in training and deployment. Recent work on large-scale training of such models has shown promising results compared to monolingual models. However, the work often focuses on multilingual models themselves in a single-pass setup. In this work, we investigate second-pass deliberation for multilingual speech recognition. Our proposed deliberation is multilingual, i.e., the text encoder encodes hypothesis text from multiple languages, and the decoder attends to multilingual text and audio. We investigate scaling the deliberation text encoder and decoder, and compare scaling the deliberation decoder and the first-pass cascaded encoder. We show that deliberation improves the average WER on 9 languages by 4% relative compared to the single-pass model. By increasing the size of the deliberation up to 1B parameters, the average WER improvement increases to 9%, with up to 14% for certain languages. Our deliberation rescorer is based on transformer layers and can be parallelized during rescoring.


【14】 An Ensemble Teacher-Student Learning Approach with Poisson Sub-sampling  to Differential Privacy Preserving Speech Recognition

标题:用于差分隐私保护语音识别的泊松次抽样集成师生学习方法

链接:https://arxiv.org/abs/2210.06382

作者:Chao-Han Huck Yang,Jun Qi,Sabato Marco Siniscalchi,Chin-Hui Lee
机构:Georgia Institute of Technology, USA and ,Kore University of Enna, Italy, Department of Electronic Systems, NTNU, Trondheim, Norway
备注:Accepted to ISCA, ISCSLP 2022, Singapore. 5 Pages
摘要:提出了一种基于Poisson子采样的集成学习框架,有效地训练一组教师模型,为训练数据提供一定的差分隐私(DP)保证.通过在DP下提升,从训练数据导出的学生模型相对于没有隐私保护的训练模型遭受很少的模型退化。我们建议的解决方案利用两种机制,即:(i)经由泊松子采样的隐私预算放大,以训练目标预测模型,该目标预测模型需要较少的噪声来实现相同级别的隐私预算,以及(ii)子采样技术与整体教师-学生学习框架的组合,该整体教师-学生学习框架在教师模型的输出处引入DP保持噪声,并且经由噪声标签来传递DP保持属性。然后利用噪声标签训练隐私保护的学生模型,从教师模型集合中学习具有DP保护的知识。在语音命令识别和汉语连续语音识别上的实验结果表明,该框架在这两种语音处理任务中的性能均优于现有的DP保持算法.
摘要:We propose an ensemble learning framework with Poisson sub-sampling to effectively train a collection of teacher models to issue some differential privacy (DP) guarantee for training data. Through boosting under DP, a student model derived from the training data suffers little model degradation from the models trained with no privacy protection. Our proposed solution leverages upon two mechanisms, namely: (i) a privacy budget amplification via Poisson sub-sampling to train a target prediction model that requires less noise to achieve a same level of privacy budget, and (ii) a combination of the sub-sampling technique and an ensemble teacher-student learning framework that introduces DP-preserving noise at the output of the teacher models and transfers DP-preserving properties via noisy labels. Privacy-preserving student models are then trained with the noisy labels to learn the knowledge with DP-protection from the teacher model ensemble. Experimental evidences on spoken command recognition and continuous speech recognition of Mandarin speech show that our proposed framework greatly outperforms existing DP-preserving algorithms in both speech processing tasks.


【15】 Can we use Common Voice to train a Multi-Speaker TTS system?

标题:我们可以使用通用语音来训练多说话人TTS系统吗?

链接:https://arxiv.org/abs/2210.06370

作者:Sewade Ogun,Vincent Colotte,Emmanuel Vincent
机构:Universit´e de Lorraine, CNRS, Inria, LORIA, F-, Nancy, France
备注:To appear in Proc. SLT 2022, Jan 09-12, 2023, Doha, Qatar
摘要:多说话者文本到语音(TTS)系统的训练依赖于基于高质量录音或有声读物的精选数据集。这种数据集通常缺乏说话者多样性,并且收集起来很昂贵。作为一种替代方案,最近的研究利用了大的、众包的自动语音识别(ASR)数据集的可用性。这种数据集的主要问题是存在噪声和/或失真样本,这降低了TTS质量。本文提出一种非侵入式平均意见得分(MOS)估计器WV-MOS,用于自动选择高质量的训练样本。我们展示了该方法在Common Voice English数据集上训练多说话人GlowTTS模型的可行性。我们的方法相对于在所有样本上的训练将生成的话语的总体质量提高了1.26MOS点,并且相对于在LibriTTS数据集上的训练将生成的话语的总体质量提高了0.35MOS点。这为更广泛的语言的自动TTS数据集管理打开了大门。
摘要:Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies have leveraged the availability of large, crowdsourced automatic speech recognition (ASR) datasets. A major problem with such datasets is the presence of noisy and/or distorted samples, which degrade TTS quality. In this paper, we propose to automatically select high-quality training samples using a non-intrusive mean opinion score (MOS) estimator, WV-MOS. We show the viability of this approach for training a multi-speaker GlowTTS model on the Common Voice English dataset. Our approach improves the overall quality of generated utterances by 1.26 MOS point with respect to training on all the samples and by 0.35 MOS point with respect to training on the LibriTTS dataset. This opens the door to automatic TTS dataset curation for a wider range of languages.


【16】 Exploring Efficient-tuning Methods in Self-supervised Speech Models

标题:探索自监督语音模型中的高效调谐方法

链接:https://arxiv.org/abs/2210.06175

作者:Zih-Ching Chen,Chin-Lun Fu,Chih-Ying Liu,Shang-Wen Li,Hung-yi Lee
机构:National Taiwan University,  Amazon AI
备注:SLT 2022
摘要:在本研究中,我们的目标是探索有效的语音自我监督学习的调整方法。近年来的研究表明,自监督学习(SSL)能够针对不同的语音任务学习到强有力的表征。然而,针对每个下游任务对预先训练的模型进行微调是参数低效的,因为SSL模型众所周知地具有数百万个参数。适配器是NLP中常用的轻量级模块,用于解决这个问题。在下游任务中,SSL模型的参数被冻结,并且仅训练适配器。由于缺乏对自监督语音任务的适配器有效性的研究,我们打算通过在预训练的语音SSL模型中添加各种适配器模块来填补这一空白。我们证明了在参数减少90%以上的情况下可以达到性能等价,并讨论了有效调优技术的优缺点。这是第一次对跨言语任务的各种适配器类型进行全面调查。
摘要:In this study, we aim to explore efficient tuning methods for speech self-supervised learning. Recent studies show that self-supervised learning (SSL) can learn powerful representations for different speech tasks. However, fine-tuning pre-trained models for each downstream task is parameter-inefficient since SSL models are notoriously large with millions of parameters. Adapters are lightweight modules commonly used in NLP to solve this problem. In downstream tasks, the parameters of SSL models are frozen, and only the adapters are trained. Given the lack of studies generally exploring the effectiveness of adapters for self-supervised speech tasks, we intend to fill this gap by adding various adapter modules in pre-trained speech SSL models. We show that the performance parity can be achieved with over 90% parameter reduction, and discussed the pros and cons of efficient tuning techniques. This is the first comprehensive investigation of various adapter types across speech tasks.


【17】 Adversarial Speaker-Consistency Learning Using Untranscribed Speech Data  for Zero-Shot Multi-Speaker Text-to-Speech

标题:基于未转录语音数据的对抗性说话人一致性学习

链接:https://arxiv.org/abs/2210.05979

作者:Byoung Jin Choi,Myeonghun Jeong,Minchan Kim,Sung Hwan Mun,Nam Soo Kim
机构:Department of Electrical and Computer Engineering and INMC, Seoul National University, Seoul, Korea备注:Accepted to APSIPA 2022
摘要:近年来提出的几种文语转换(TTS)模型在单说话人和多说话人情况下都能生成具有人级质量的语音样本。然而,用单个参考音频合成新说话人的语音,通常称为零激发多说话人文本到语音(ZSM-TTS),仍然是非常具有挑战性的任务。ZSM-TTS的主要挑战是在新说话人的语音生成时的说话人域转换问题。为了解决这个问题,我们提出了对抗式说话人一致性学习(ASCL)。该方法首先在每次训练迭代中使用外部未转录数据集生成查询说话人的附加语音。然后,该模型通过采用对抗学习方案来学习一致地生成同一说话人的语音样本作为对应的说话人嵌入向量。实验结果表明,该方法在ZSM-TTS中的语音识别质量和说话人相似度方面均优于基线方法。
摘要:Several recently proposed text-to-speech (TTS) models achieved to generate the speech samples with the human-level quality in the single-speaker and multi-speaker TTS scenarios with a set of pre-defined speakers. However, synthesizing a new speaker's voice with a single reference audio, commonly known as zero-shot multi-speaker text-to-speech (ZSM-TTS), is still a very challenging task. The main challenge of ZSM-TTS is the speaker domain shift problem upon the speech generation of a new speaker. To mitigate this problem, we propose adversarial speaker-consistency learning (ASCL). The proposed method first generates an additional speech of a query speaker using the external untranscribed datasets at each training iteration. Then, the model learns to consistently generate the speech sample of the same speaker as the corresponding speaker embedding vector by employing an adversarial learning scheme. The experimental results show that the proposed method is effective compared to the baseline in terms of the quality and speaker similarity in ZSM-TTS.


【18】 Cross-dataset COVID-19 Transfer Learning with Cough Detection, Cough  Segmentation, and Data Augmentation

标题:具有咳嗽检测、咳嗽分割和数据增强的跨数据集新冠肺炎转移学习

链接:https://arxiv.org/abs/2210.05843

作者:Bagus Tris Atmaja,Zanjabila,Suyanto,Akira Sasou
机构:• Three processing blocks were proposed to improve cough-based COVID-,  detection., • The study reports ablation studies on these three blocks and hyperpa-, rameters tuning., • It also summarizes previous studies on the same test set of COVID-
摘要:本文讨论了基于咳嗽的COVID-19检测问题。我们提出了一种跨数据集迁移学习方法,通过结合咳嗽检测、咳嗽分割和数据扩充来提高COVID-19检测的性能。第一种方法旨在去除非咳嗽信号和低概率咳嗽信号。第二个目的是将波形中的几次咳嗽分离成单独的咳嗽。第三个目标是增加深度学习模型的样本数量。这三个处理模块非常重要,因为我们的发现表明,相对于没有这些模块的基线方法,有很大的改进余地。进行烧蚀研究以优化超参数,并且发现α混合是通过这种增强方法改善模型性能的重要因素。对本研究与之前在同一评估集上的研究进行了总结,以深入了解基于咳嗽的COVID-19检测的不同方法。
摘要:This paper addresses issues on cough-based COVID-19 detection. We propose a cross-dataset transfer learning approach to improve the performance of COVID-19 detection by incorporating cough detection, cough segmentation, and data augmentation. The first aimed at removing non-cough signals and cough signals with low probability. The second aimed at segregating several coughs in a waveform into individual coughs. The third aimed at increasing the number of samples for the deep learning model. These three processing blocks are important as our finding revealed a large margin of improvement relative to the baseline methods without these blocks. An ablation study is conducted to optimize hyperparameters and it was found that alpha mixup is an important factor among others in improving the model performance via this augmentation method. A summary of this study with previous studies on the same evaluation set was given to gain insights into different methods of cough-based COVID-19 detection.


eess.AS音频处理

【1】 An Ensemble Teacher-Student Learning Approach with Poisson Sub-sampling  to Differential Privacy Preserving Speech Recognition

标题:用于差分隐私保护语音识别的泊松次抽样集成师生学习方法

链接:https://arxiv.org/abs/2210.06382

* 与cs.SD语音【14】为同一篇

作者:Chao-Han Huck Yang,Jun Qi,Sabato Marco Siniscalchi,Chin-Hui Lee
机构:Georgia Institute of Technology, USA and ,Kore University of Enna, Italy, Department of Electronic Systems, NTNU, Trondheim, Norway
备注:Accepted to ISCA, ISCSLP 2022, Singapore. 5 Pages
摘要:提出了一种基于Poisson子采样的集成学习框架,有效地训练一组教师模型,为训练数据提供一定的差分隐私(DP)保证.通过在DP下提升,从训练数据导出的学生模型相对于没有隐私保护的训练模型遭受很少的模型退化。我们建议的解决方案利用两种机制,即:(i)经由泊松子采样的隐私预算放大,以训练目标预测模型,该目标预测模型需要较少的噪声来实现相同级别的隐私预算,以及(ii)子采样技术与整体教师-学生学习框架的组合,该整体教师-学生学习框架在教师模型的输出处引入DP保持噪声,并且经由噪声标签来传递DP保持属性。然后利用噪声标签训练隐私保护的学生模型,从教师模型集合中学习具有DP保护的知识。在语音命令识别和汉语连续语音识别上的实验结果表明,该框架在这两种语音处理任务中的性能均优于现有的DP保持算法.
摘要:We propose an ensemble learning framework with Poisson sub-sampling to effectively train a collection of teacher models to issue some differential privacy (DP) guarantee for training data. Through boosting under DP, a student model derived from the training data suffers little model degradation from the models trained with no privacy protection. Our proposed solution leverages upon two mechanisms, namely: (i) a privacy budget amplification via Poisson sub-sampling to train a target prediction model that requires less noise to achieve a same level of privacy budget, and (ii) a combination of the sub-sampling technique and an ensemble teacher-student learning framework that introduces DP-preserving noise at the output of the teacher models and transfers DP-preserving properties via noisy labels. Privacy-preserving student models are then trained with the noisy labels to learn the knowledge with DP-protection from the teacher model ensemble. Experimental evidences on spoken command recognition and continuous speech recognition of Mandarin speech show that our proposed framework greatly outperforms existing DP-preserving algorithms in both speech processing tasks.


【2】 Can we use Common Voice to train a Multi-Speaker TTS system?

标题:我们可以使用通用语音来训练多说话人TTS系统吗?

链接:https://arxiv.org/abs/2210.06370

* 与cs.SD语音【15】为同一篇

作者:Sewade Ogun,Vincent Colotte,Emmanuel Vincent
机构:Universit´e de Lorraine, CNRS, Inria, LORIA, F-, Nancy, France
备注:To appear in Proc. SLT 2022, Jan 09-12, 2023, Doha, Qatar
摘要:多说话者文本到语音(TTS)系统的训练依赖于基于高质量录音或有声读物的精选数据集。这种数据集通常缺乏说话者多样性,并且收集起来很昂贵。作为一种替代方案,最近的研究利用了大的、众包的自动语音识别(ASR)数据集的可用性。这种数据集的主要问题是存在噪声和/或失真样本,这降低了TTS质量。本文提出一种非侵入式平均意见得分(MOS)估计器WV-MOS,用于自动选择高质量的训练样本。我们展示了该方法在Common Voice English数据集上训练多说话人GlowTTS模型的可行性。我们的方法相对于在所有样本上的训练将生成的话语的总体质量提高了1.26MOS点,并且相对于在LibriTTS数据集上的训练将生成的话语的总体质量提高了0.35MOS点。这为更广泛的语言的自动TTS数据集管理打开了大门。
摘要:Training of multi-speaker text-to-speech (TTS) systems relies on curated datasets based on high-quality recordings or audiobooks. Such datasets often lack speaker diversity and are expensive to collect. As an alternative, recent studies have leveraged the availability of large, crowdsourced automatic speech recognition (ASR) datasets. A major problem with such datasets is the presence of noisy and/or distorted samples, which degrade TTS quality. In this paper, we propose to automatically select high-quality training samples using a non-intrusive mean opinion score (MOS) estimator, WV-MOS. We show the viability of this approach for training a multi-speaker GlowTTS model on the Common Voice English dataset. Our approach improves the overall quality of generated utterances by 1.26 MOS point with respect to training on all the samples and by 0.35 MOS point with respect to training on the LibriTTS dataset. This opens the door to automatic TTS dataset curation for a wider range of languages.


【3】 Exploring Efficient-tuning Methods in Self-supervised Speech Models

标题:探索自监督语音模型中的高效调谐方法

链接:https://arxiv.org/abs/2210.06175

* 与cs.SD语音【16】为同一篇

作者:Zih-Ching Chen,Chin-Lun Fu,Chih-Ying Liu,Shang-Wen Li,Hung-yi Lee
机构:National Taiwan University,  Amazon AI备注:SLT 2022
摘要:在本研究中,我们的目标是探索有效的语音自我监督学习的调整方法。近年来的研究表明,自监督学习(SSL)能够针对不同的语音任务学习到强有力的表征。然而,针对每个下游任务对预先训练的模型进行微调是参数低效的,因为SSL模型众所周知地具有数百万个参数。适配器是NLP中常用的轻量级模块,用于解决这个问题。在下游任务中,SSL模型的参数被冻结,并且仅训练适配器。由于缺乏对自监督语音任务的适配器有效性的研究,我们打算通过在预训练的语音SSL模型中添加各种适配器模块来填补这一空白。我们证明了在参数减少90%以上的情况下可以达到性能等价,并讨论了有效调优技术的优缺点。这是第一次对跨言语任务的各种适配器类型进行全面调查。
摘要:In this study, we aim to explore efficient tuning methods for speech self-supervised learning. Recent studies show that self-supervised learning (SSL) can learn powerful representations for different speech tasks. However, fine-tuning pre-trained models for each downstream task is parameter-inefficient since SSL models are notoriously large with millions of parameters. Adapters are lightweight modules commonly used in NLP to solve this problem. In downstream tasks, the parameters of SSL models are frozen, and only the adapters are trained. Given the lack of studies generally exploring the effectiveness of adapters for self-supervised speech tasks, we intend to fill this gap by adding various adapter modules in pre-trained speech SSL models. We show that the performance parity can be achieved with over 90% parameter reduction, and discussed the pros and cons of efficient tuning techniques. This is the first comprehensive investigation of various adapter types across speech tasks.


【4】 Adversarial Speaker-Consistency Learning Using Untranscribed Speech Data  for Zero-Shot Multi-Speaker Text-to-Speech

标题:基于未转录语音数据的对抗性说话人一致性学习

链接:https://arxiv.org/abs/2210.05979

* 与cs.SD语音【17】为同一篇

作者:Byoung Jin Choi,Myeonghun Jeong,Minchan Kim,Sung Hwan Mun,Nam Soo Kim
机构:Department of Electrical and Computer Engineering and INMC, Seoul National University, Seoul, Korea
备注:Accepted to APSIPA 2022
摘要:近年来提出的几种文语转换(TTS)模型在单说话人和多说话人情况下都能生成具有人级质量的语音样本。然而,用单个参考音频合成新说话人的语音,通常称为zero-shot多说话人文本到语音(ZSM-TTS),仍然是非常具有挑战性的任务。ZSM-TTS的主要挑战是在新说话人的语音生成时的说话人域转换问题。为了解决这个问题,我们提出了对抗式说话人一致性学习(ASCL)。该方法首先在每次训练迭代中使用外部未转录数据集生成查询说话人的附加语音。然后,该模型通过采用对抗学习方案来学习一致地生成同一说话人的语音样本作为对应的说话人嵌入向量。实验结果表明,该方法在ZSM-TTS中的语音识别质量和说话人相似度方面均优于基线方法。
摘要:Several recently proposed text-to-speech (TTS) models achieved to generate the speech samples with the human-level quality in the single-speaker and multi-speaker TTS scenarios with a set of pre-defined speakers. However, synthesizing a new speaker's voice with a single reference audio, commonly known as zero-shot multi-speaker text-to-speech (ZSM-TTS), is still a very challenging task. The main challenge of ZSM-TTS is the speaker domain shift problem upon the speech generation of a new speaker. To mitigate this problem, we propose adversarial speaker-consistency learning (ASCL). The proposed method first generates an additional speech of a query speaker using the external untranscribed datasets at each training iteration. Then, the model learns to consistently generate the speech sample of the same speaker as the corresponding speaker embedding vector by employing an adversarial learning scheme. The experimental results show that the proposed method is effective compared to the baseline in terms of the quality and speaker similarity in ZSM-TTS.


【5】 Cross-dataset COVID-19 Transfer Learning with Cough Detection, Cough  Segmentation, and Data Augmentation

标题:具有咳嗽检测、咳嗽分割和数据增强的跨数据集新冠肺炎转移学习

链接:https://arxiv.org/abs/2210.05843

* 与cs.SD语音【18】为同一篇

作者:Bagus Tris Atmaja,Zanjabila,Suyanto,Akira Sasou
机构:• Three processing blocks were proposed to improve cough-based COVID-,  detection., • The study reports ablation studies on these three blocks and hyperpa-, rameters tuning., • It also summarizes previous studies on the same test set of COVID-
摘要:本文讨论了基于咳嗽的COVID-19检测问题。我们提出了一种跨数据集迁移学习方法,通过结合咳嗽检测、咳嗽分割和数据扩充来提高COVID-19检测的性能。第一种方法旨在去除非咳嗽信号和低概率咳嗽信号。第二个目的是将波形中的几次咳嗽分离成单独的咳嗽。第三个目标是增加深度学习模型的样本数量。这三个处理模块非常重要,因为我们的发现表明,相对于没有这些模块的基线方法,有很大的改进余地。进行烧蚀研究以优化超参数,并且发现α混合是通过这种增强方法改善模型性能的重要因素。对本研究与之前在同一评估集上的研究进行了总结,以深入了解基于咳嗽的COVID-19检测的不同方法。
摘要:This paper addresses issues on cough-based COVID-19 detection. We propose a cross-dataset transfer learning approach to improve the performance of COVID-19 detection by incorporating cough detection, cough segmentation, and data augmentation. The first aimed at removing non-cough signals and cough signals with low probability. The second aimed at segregating several coughs in a waveform into individual coughs. The third aimed at increasing the number of samples for the deep learning model. These three processing blocks are important as our finding revealed a large margin of improvement relative to the baseline methods without these blocks. An ablation study is conducted to optimize hyperparameters and it was found that alpha mixup is an important factor among others in improving the model performance via this augmentation method. A summary of this study with previous studies on the same evaluation set was given to gain insights into different methods of cough-based COVID-19 detection.


【6】 Individualized Conditioning and Negative Distances for Speaker  Separation

标题:说话人分离的个别化条件作用和负距离

链接:https://arxiv.org/abs/2210.06368

* 与cs.SD语音【1】为同一篇

作者:Tao Sun,Nidal Abuhajar,Shuyu Gong,Zhewei Wang,Charles D. Smith,Xianhui Wang,Li Xu,Jundong Liu
机构:∗School of Electrical Engineering and Computer Science, Ohio University, Athens, OH , †Department of Neurology, University of Kentucky, Lexington, KY , ‡Division of Communication Sciences, Ohio University, Athens, OH
备注:Accepted to ICMLA 2022
摘要:说话人分离的目的是从混合信号中提取多个语音。本文提出了两种说话人感知的说话人分离方案,以改进现有的说话人分离方案。第一种模型是说话人调节网络,其集成语音样本以生成个性化的说话人条件,然后个性化的说话人条件为分离模块提供有根据的指导以产生分离良好的输出。  第二种设计旨在减少分离语音中的非目标语音。为此,我们提出负距离来惩罚通道输出中任何非目标语音的出现,而正距离使分离后的语音更接近干净目标。我们探索了两种不同的设置,加权和和三重类,以整合这两个距离,形成分离网络的组合辅助损耗。在LibriMix上进行的实验验证了所提模型的有效性.
摘要:Speaker separation aims to extract multiple voices from a mixed signal. In this paper, we propose two speaker-aware designs to improve the existing speaker separation solutions. The first model is a speaker conditioning network that integrates speech samples to generate individualized speaker conditions, which then provide informed guidance for a separation module to produce well-separated outputs.  The second design aims to reduce non-target voices in the separated speech. To this end, we propose negative distances to penalize the appearance of any non-target voice in the channel outputs, and positive distances to drive the separated voices closer to the clean targets. We explore two different setups, weighted-sum and triplet-like, to integrate these two distances to form a combined auxiliary loss for the separation networks. Experiments conducted on LibriMix demonstrate the effectiveness of our proposed models.


【7】 Text-to-Audio Grounding Based Novel Metric for Evaluating Audio Caption  Similarity

标题:基于文本-音频基础的音频字幕相似度评价新指标

链接:https://arxiv.org/abs/2210.06354

* 与cs.SD语音【2】为同一篇

作者:Swapnil Bhosale,Rupayan Chakraborty,Sunil Kumar Kopparapu
机构:TCS Research, Tata Consultancy Services Limited, India.
备注:9 pages, 8 figures,
摘要:自动音频字幕(AAC)指的是将音频样本翻译成描述音频事件、事件源及其关系的自然语言(NL)文本的任务。与NL文本生成任务(其依赖于基于词汇语义的度量(如BLEU、ROUGE、METEOR)来进行评估)不同,AAC评估度量需要映射NL文本(短语)的能力,所述NL文本对应于除词汇语义之外的类似声音。当前用于评估AAC任务的度量缺乏对由文本表示的声音的感知属性的理解。本文提出了一种新的基于文本到音频基础(TAG)的评价指标,该指标对跨模态任务如AAC的评价非常有用。在公开的AAC数据集上的实验表明,与NL文本和图像字幕文献中使用的现有评价指标相比,本文提出的评价指标具有更好的性能。
摘要:Automatic Audio Captioning (AAC) refers to the task of translating an audio sample into a natural language (NL) text that describes the audio events, source of the events and their relationships. Unlike NL text generation tasks, which rely on metrics like BLEU, ROUGE, METEOR based on lexical semantics for evaluation, the AAC evaluation metric requires an ability to map NL text (phrases) that correspond to similar sounds in addition lexical semantics. Current metrics used for evaluation of AAC tasks lack an understanding of the perceived properties of sound represented by text. In this paper, wepropose a novel metric based on Text-to-Audio Grounding (TAG), which is, useful for evaluating cross modal tasks like AAC. Experiments on publicly available AAC data-set shows our evaluation metric to perform better compared to existing metrics used in NL text and image captioning literature.


【8】 TaskMix: Data Augmentation for Meta-Learning of Spoken Intent  Understanding

标题:TaskMix:口语意图理解元学习的数据增强

链接:https://arxiv.org/abs/2210.06341

作者:Surya Kant Sahu
机构:Skit.ai, The Learning Machines
备注:Accepted at Findings of AACL-IJCNLP 2022
摘要:元学习是一个研究方向,它旨在更好地将知识从相关的任务转移到看不见但相关的任务。然而,元学习需要许多训练任务来学习能够很好地转移到看不见的任务的表示;否则,会导致过拟合,性能退化到不如多任务学习。我们表明,当任务多样性较低时,一种最先进的数据扩增方法恶化了这种过拟合问题。提出了一种简单的任务合成方法TaskMix,通过对现有任务进行线性插值来合成新任务。我们将TaskMix与一个内部多语言意图分类数据集上的许多基线进行了比较,该数据集包含来自真实生活中人机电话话语的N-Best ASR假设,以及来自MTOP的两个数据集。我们证明TaskMix的性能优于基线,在任务多样性较低时缓解了过度拟合,并且即使在任务多样性较高时也不会降低性能。
摘要:Meta-Learning has emerged as a research direction to better transfer knowledge from related tasks to unseen but related tasks. However, Meta-Learning requires many training tasks to learn representations that transfer well to unseen tasks; otherwise, it leads to overfitting, and the performance degenerates to worse than Multi-task Learning. We show that a state-of-the-art data augmentation method worsens this problem of overfitting when the task diversity is low. We propose a simple method, TaskMix, which synthesizes new tasks by linearly interpolating existing tasks. We compare TaskMix against many baselines on an in-house multilingual intent classification dataset of N-Best ASR hypotheses derived from real-life human-machine telephony utterances and two datasets derived from MTOP. We show that TaskMix outperforms baselines, alleviates overfitting when task diversity is low, and does not degrade performance even when it is high.


【9】 SQuId: Measuring Speech Naturalness in Many Languages

标题:SQUID:测量多种语言的语言自然度

链接:https://arxiv.org/abs/2210.06324

* 与cs.SD语音【3】为同一篇

作者:Thibault Sellam,Ankur Bapna,Joshua Camp,Diana Mackinnon,Ankur P. Parikh,Jason Riesa
机构:Google
摘要:许多文本到语音的研究依赖于人工评估,这导致了巨大的成本,并减缓了开发过程。这个问题在大量使用多种语言的应用程序中尤其严重,在这些应用程序中,招聘和投票法官可能需要数周的时间。我们介绍SQuId(语音质量识别),这是一个多语言自然度预测模型,它在超过100万个评分上进行了训练,并在65个地区进行了测试--这是迄今为止此类研究中最大的一次。主要的观点是,在多个语言环境上训练一个模型始终优于单一语言环境基线。我们介绍了我们的任务和模型,并表明它的性能比基于w2 v-BERT和VoiceMOS的竞争基准高50.0%。然后,我们证明了微调过程中跨区域转换的有效性,并强调了其对zero-shot区域的影响,即:没有微调数据的语言环境。通过一系列的分析,我们强调了非语言效应如声音假象在跨语言环境迁移中的作用。最后,我们介绍了我们的设计决策的效果,例如,模型大小、预训练多样性和语言再平衡。
摘要:Much of text-to-speech research relies on human evaluation, which incurs heavy costs and slows down the development process. The problem is particularly acute in heavily multilingual applications, where recruiting and polling judges can take weeks. We introduce SQuId (Speech Quality Identification), a multilingual naturalness prediction model trained on over a million ratings and tested in 65 locales-the largest effort of this type to date. The main insight is that training one model on many locales consistently outperforms mono-locale baselines. We present our task, the model, and show that it outperforms a competitive baseline based on w2v-BERT and VoiceMOS by 50.0%. We then demonstrate the effectiveness of cross-locale transfer during fine-tuning and highlight its effect on zero-shot locales, i.e., locales for which there is no fine-tuning data. Through a series of analyses, we highlight the role of non-linguistic effects such as sound artifacts in cross-locale transfer. Finally, we present the effect of our design decision, e.g., model size, pre-training diversity, and language rebalancing with several ablation experiments.


【10】 A context-aware knowledge transferring strategy for CTC-based ASR

标题:一种基于CTC的ASR上下文感知知识转移策略

链接:https://arxiv.org/abs/2210.06244

* 与cs.SD语音【4】为同一篇

作者:Ke-Han Lu,Kuan-Yu Chen
机构:National Taiwan University of Science and Technology, Taiwan
备注:Accepted by SLT 2022
摘要:非自回归自动语音识别(ASR)模型因其解码速度快、性能优越而受到越来越多的关注。其中,基于连接主义时态分类(CTC)的方法仍然是主流。然而,理论上的固有缺陷--符号之间的独立性假设--为作品派的表演设置了障碍。针对这一问题,本文提出了一种基于上下文感知的ASR知识转移策略,该策略由知识转移模块和上下文感知的训练策略组成。前者旨在从预先训练好的语言模型中提取语言信息,后者旨在调整条件独立假设所带来的局限性。在此基础上,本文提出了一种基于wav 2 vec 2. 0的知识注入的上下文感知的基于CTC的ASR。在AISHELL-1和AISHELL-2数据集上的一系列实验验证了该方法的有效性.
摘要:Non-autoregressive automatic speech recognition (ASR) modeling has received increasing attention recently because of its fast decoding speed and superior performance. Among representatives, methods based on the connectionist temporal classification (CTC) are still a dominating stream. However, the theoretically inherent flaw, the assumption of independence between tokens, creates a performance barrier for the school of works. To mitigate the challenge, we propose a context-aware knowledge transferring strategy, consisting of a knowledge transferring module and a context-aware training strategy, for CTC-based ASR. The former is designed to distill linguistic information from a pre-trained language model, and the latter is framed to modulate the limitations caused by the conditional independence assumption. As a result, a knowledge-injected context-aware CTC-based ASR built upon the wav2vec2.0 is presented in this paper. A series of experiments on the AISHELL-1 and AISHELL-2 datasets demonstrate the effectiveness of the proposed method.


【11】 Towards visually prompted keyword localisation for zero-resource spoken  languages

标题:面向零资源口语的视觉提示关键词本地化

链接:https://arxiv.org/abs/2210.06229

* 与cs.SD语音【5】为同一篇

作者:Leanne Nortje,Herman Kamper
机构:MediaLab, Electrical & Electronic Engineering, Stellenbosch University, South Africa
备注:Accepted to IEEE SLT 2022
摘要:想象一下,能够向系统显示关键字的可视描述,并从零资源语音语料库中找到包含该关键字的口语话语。我们将此任务形式化,并称之为视觉提示关键字本地化(VPKL):给定关键字的图像,检测并预测关键字在话语中的何处出现。为了进行VPKL,我们提出了一个具有新颖的局部注意机制的语音-视觉模型,我们用一个新的关键词采样方案训练该模型。我们表明,这些创新在VPKL中提供了优于现有语音-视觉模型的改进。我们还比较了视觉词袋(BoW)模型,其中图像自动标记视觉标签,并与未标记的语音配对。尽管可以使用书面关键字直接查询该可视BoW(而我们的采用图像查询),但我们的新模型在检测和定位方面仍优于可视BoW,在定位F1方面相对提高了16%。
摘要:Imagine being able to show a system a visual depiction of a keyword and finding spoken utterances that contain this keyword from a zero-resource speech corpus. We formalise this task and call it visually prompted keyword localisation (VPKL): given an image of a keyword, detect and predict where in an utterance the keyword occurs. To do VPKL, we propose a speech-vision model with a novel localising attention mechanism which we train with a new keyword sampling scheme. We show that these innovations give improvements in VPKL over an existing speech-vision model. We also compare to a visual bag-of-words (BoW) model where images are automatically tagged with visual labels and paired with unlabelled speech. Although this visual BoW can be queried directly with a written keyword (while our's takes image queries), our new model still outperforms the visual BoW in both detection and localisation, giving a 16% relative improvement in localisation F1.


【12】 VCSE: Time-Domain Visual-Contextual Speaker Extraction Network

标题:VCSE:时间域视觉语境说话人提取网络

链接:https://arxiv.org/abs/2210.06177

* 与cs.SD语音【6】为同一篇

作者:Junjie Li,Meng Ge,Zexu Pan,Longbiao Wang,Jianwu Dang
机构:Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University, Tianjin, China,  Department of Electrical and Computer Engineering, National University of Singapore, Singapore
摘要:说话人提取寻求在给定辅助参考的多说话人场景中提取目标语音。这样的参考可以是听觉的,即,预先记录的语音、视觉,即,嘴唇运动或上下文,即,语音序列。不同模态中的指称提供了不同的和互补的信息,这些信息可以被融合以形成对目标说话人的自上而下的注意。先前的研究已经在单一模型中引入了视觉和情境模态。本文提出了一种两级时域视觉语境说话人提取网络VCSE,该网络将视觉语境线索和自注册语境线索逐级融合,充分利用了各个模态的优势。在第一阶段,我们利用视觉线索预先撷取目标语音,并估计潜在的语音序列。在第二阶段,我们使用自注册的上下文线索对预提取的目标语音进行精炼。在真实的唇读句子3(LRS3)数据库上的实验结果表明,本文提出的VCSE网络的性能始终优于其他现有的基线.
摘要:Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence. References in different modalities provide distinct and complementary information that could be fused to form top-down attention on the target speaker. Previous studies have introduced visual and contextual modalities in a single model. In this paper, we propose a two-stage time-domain visual-contextual speaker extraction network named VCSE, which incorporates visual and self-enrolled contextual cues stage by stage to take full advantage of every modality. In the first stage, we pre-extract a target speech with visual cues and estimate the underlying phonetic sequence. In the second stage, we refine the pre-extracted target speech with the self-enrolled contextual cues. Experimental results on the real-world Lip Reading Sentences 3 (LRS3) database demonstrate that our proposed VCSE network consistently outperforms other state-of-the-art baselines.


【13】 THUEE system description for NIST 2020 SRE CTS challenge

标题:NIST 2020 SRE CTS挑战赛THUEE系统描述

链接:https://arxiv.org/abs/2210.06111

* 与cs.SD语音【7】为同一篇

作者:Yu Zheng,Jinghan Peng,Miao Zhao,Yufeng Ma,Min Liu,Xinyue Ma,Tianyu Liang,Tianlong Kong,Liang He,Minqiang Xu
机构:SpeakIn Technologies Co. Ltd., ShangHai, China, Department of Electronic Engineering Tsinghua University, Beijing, China
备注:3 pages, 1 table; System desciption of NIST 2020 SRE CTS challenge
摘要:本文介绍了THUEE团队参加NIST 2020说话人识别评估(SRE)会话电话语音(CTS)挑战赛的系统描述。在本评估中,包括ResNet 74、ResNet 152和RepVGG-B2在内的子系统被开发为扬声器嵌入提取器。我们使用基于AM-Softmax和AAM-Softmax的组合损失函数,即CM-Softmax。我们采用了两阶段训练策略来进一步提高系统性能。我们融合了所有单个系统作为最终提交。我们的方法带来了出色的性能,并在挑战中排名第一。
摘要:This paper presents the system description of the THUEE team for the NIST 2020 Speaker Recognition Evaluation (SRE) conversational telephone speech (CTS) challenge. The subsystems including ResNet74, ResNet152, and RepVGG-B2 are developed as speaker embedding extractors in this evaluation. We used combined AM-Softmax and AAM-Softmax based loss functions, namely CM-Softmax. We adopted a two-staged training strategy to further improve system performance. We fused all individual systems as our final submission. Our approach leads to excellent performance and ranks 1st in the challenge.


【14】 SpecRNet: Towards Faster and More Accessible Audio DeepFake Detection

标题:SPECRNet:向更快、更易访问的音频DeepFake检测迈进

链接:https://arxiv.org/abs/2210.06105

* 与cs.SD语音【8】为同一篇

作者:Piotr Kawa,Marcin Plata,Piotr Syga
机构:Department of Artificial Intelligence, Wrocław University of Science and Technology, Wrocław, Poland, ©, IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including
备注:Accepted by TrustCom 2022: The 21st IEEE International Conference on Trust, Security and Privacy in Computing and Communications
摘要:音频DeepFakes是使用深度神经网络生成的话语。它们具有高度误导性,并因用于假新闻、冒充或勒索而构成威胁。在这项工作中,我们通过提供SpecRNet(一种具有快速推理时间和低计算要求的神经网络架构),专注于增加音频DeepFake检测方法的可访问性。我们的基准测试表明,SpecRNet处理音频样本所需的时间最多可减少40%,其性能可与LCNN架构(最佳的音频DeepFake检测模型之一)相媲美。这种方法不仅可以被在线多媒体服务用来验证每天上传的大量内容,而且由于其低要求,还可以被普通公民用来评估他们设备上的材料。此外,我们还提供了三种独特设置的基准测试,以确认我们的模型的正确性。它们反映了低资源数据集、短话语检测和有限攻击基准的场景,其中我们更仔细地观察了特定攻击对给定体系结构的影响。
摘要:Audio DeepFakes are utterances generated with the use of deep neural networks. They are highly misleading and pose a threat due to use in fake news, impersonation, or extortion. In this work, we focus on increasing accessibility to the audio DeepFake detection methods by providing SpecRNet, a neural network architecture characterized by a quick inference time and low computational requirements. Our benchmark shows that SpecRNet, requiring up to about 40% less time to process an audio sample, provides performance comparable to LCNN architecture - one of the best audio DeepFake detection models. Such a method can not only be used by online multimedia services to verify a large bulk of content uploaded daily but also, thanks to its low requirements, by average citizens to evaluate materials on their devices. In addition, we provide benchmarks in three unique settings that confirm the correctness of our model. They reflect scenarios of low-resource datasets, detection on short utterances and limited attacks benchmark in which we take a closer look at the influence of particular attacks on given architectures.


【15】 Summary on the ISCSLP 2022 Chinese-English Code-Switching ASR Challenge

标题:ISCSLP 2022汉英代码转换ASR挑战赛综述

链接:https://arxiv.org/abs/2210.06091

* 与cs.SD语音【9】为同一篇

作者:Shuhao Deng,Chengfei Li,infeng Bai,Qingqing Zhang,Wei-Qiang Zhang,Runyan Yang,Gaofeng Cheng,Pengyuan Zhang,Yonghong Yan
机构:TAL Education Group, Beijing, China ,Magic Data ,Tsinghua University ,Institute of Acoustics, Chinese Academy of Sciences
备注:accepted by ISCSLP 2022
摘要:由于多种语言之间的语码转换现象以及日常生活中频繁发生的语码转换现象,使得语码转换自动语音识别成为自动语音识别中最具挑战性和最有价值的场景之一。ISCSLP 2022汉英语码转换自动语音识别(CSASR)挑战赛旨在推动语码转换自动语音识别的发展。ISCSLP 2022 CSASR挑战赛为参赛者提供了TAL_CSASR语料库和MagicData-RAMC语料库两个训练集、一个开发集和一个测试集,用于CSASR模型训练和评估。除了挑战之外,我们还提供了基准系统性能以供参考。因此,有40多个团队参与了此次挑战,获胜团队在测试集上实现了16.70%的混合错误率(MER)性能,与基线系统相比,MER绝对提升了9.8%。本文将描述数据集、相关基线系统和需求,并总结CSASR挑战结果和提交系统中使用的主要技术和技巧。
摘要:Code-switching automatic speech recognition becomes one of the most challenging and the most valuable scenarios of automatic speech recognition, due to the code-switching phenomenon between multilingual language and the frequent occurrence of code-switching phenomenon in daily life. The ISCSLP 2022 Chinese-English Code-Switching Automatic Speech Recognition (CSASR) Challenge aims to promote the development of code-switching automatic speech recognition. The ISCSLP 2022 CSASR challenge provided two training sets, TAL_CSASR corpus and MagicData-RAMC corpus, a development and a test set for participants, which are used for CSASR model training and evaluation. Along with the challenge, we also provide the baseline system performance for reference. As a result, more than 40 teams participated in this challenge, and the winner team achieved 16.70% Mixture Error Rate (MER) performance on the test set and has achieved 9.8% MER absolute improvement compared with the baseline system. In this paper, we will describe the datasets, the associated baselines system and the requirements, and summarize the CSASR challenge results and major techniques and tricks used in the submitted systems.


【16】 JukeDrummer: Conditional Beat-aware Audio-domain Drum Accompaniment  Generation via Transformer VQ-VA

标题:JukeDrummer:通过TransformerVQ-VA生成条件节拍感知音域鼓伴奏

链接:https://arxiv.org/abs/2210.06007

* 与cs.SD语音【10】为同一篇

作者:Yueh-Kao Wu,Ching-Yu Chiu,Yi-Hsuan Yang
机构:Academia Sinica, National Cheng Kung University, Taiwan AI Labs
备注:Accepted at ISMIR 2022
摘要:本文提出了一种在音频域中生成鼓音轨的模型,以与用户提供的无鼓录音一起播放。具体来说,使用无鼓轨道和相应的人造鼓轨道的配对数据,我们训练一个Transformer模型来即兴演奏一个看不见的无鼓录音的鼓部分。我们结合两种方法来编码输入音频。首先,我们训练一个向量量化变分自动编码器(VQ-VAE),用离散码表示输入音频,然后可以在Transformer中使用。第二,使用音频域节拍跟踪模型,我们计算输入音频的节拍相关特征,并将其作为Transformer中的嵌入。我们使用单独的VQ-VAE将鼓轨道的Mel谱图编码为另一组离散码,并训练Transformer预测与鼓相关的离散码的序列,而不是将鼓轨道直接生成为波形。然后,用解码器将输出码转换成梅尔谱图,然后用声码器将其转换成波形。我们报告了对所提出的模型的变体的客观和主观评估,证明了具有节拍信息的模型生成在节奏和风格上与输入音频一致的鼓伴奏。
摘要:This paper proposes a model that generates a drum track in the audio domain to play along to a user-provided drum-free recording. Specifically, using paired data of drumless tracks and the corresponding human-made drum tracks, we train a Transformer model to improvise the drum part of an unseen drumless recording. We combine two approaches to encode the input audio. First, we train a vector-quantized variational autoencoder (VQ-VAE) to represent the input audio with discrete codes, which can then be readily used in a Transformer. Second, using an audio-domain beat tracking model, we compute beat-related features of the input audio and use them as embeddings in the Transformer. Instead of generating the drum track directly as waveforms, we use a separate VQ-VAE to encode the mel-spectrogram of a drum track into another set of discrete codes, and train the Transformer to predict the sequence of drum-related discrete codes. The output codes are then converted to a mel-spectrogram with a decoder, and then to the waveform with a vocoder. We report both objective and subjective evaluations of variants of the proposed model, demonstrating that the model with beat information generates drum accompaniment that is rhythmically and stylistically consistent with the input audio.


【17】 Enemy Spotted: in-game gun sound dataset for gunshot classification and  localization

标题:发现敌人:用于枪击分类和定位的游戏中枪声数据集

链接:https://arxiv.org/abs/2210.05917

* 与cs.SD语音【11】为同一篇

作者:Junwoo Park,Youngwoo Cho,Gyuhyeon Sim,Hojoon Lee,Jaegul Choo
机构:Kim Jaechul Graduate School of AI, Korea Advanced Institute of Science and Technology (KAIST), Daejeon, South Korea
备注:Accepted at IEEE Conference on Games (GoG) 2022
摘要:近年来,基于深度学习的方法在声音分类和定位中由于其简单高效、无需领域知识等优点而受到广泛关注。然而,现有数据集中缺乏枪声一直是实施支持系统通过利用深度学习模型从枪声中发现罪犯的主要障碍。由于枪声的发生是罕见的和不可预测的,因此在现实世界中收集枪声是不切实际的。作为替代,可以从被设计为模拟真实世界战争的FPS游戏中获得枪声。最近的FPS游戏提供了一个现实的环境,我们可以安全地收集射击数据,同时模拟甚至危险的情况。利用游戏环境的优势,构建了一个枪械数据集BGG,用于枪械分类和枪械定位任务。BGG数据集包括37种不同类型的枪械、距离以及声源和接收器之间的方向。通过在BGG数据集上训练多个声音分类和定位基线,我们仔细验证了游戏中的枪声数据具有足够的信息来识别枪声的位置和类型。最后,我们证明了利用BGG数据集可以提高真实世界枪支分类和定位任务的准确性。
摘要:Recently, deep learning-based methods have drawn huge attention due to their simple yet high performance without domain knowledge in sound classification and localization tasks. However, a lack of gun sounds in existing datasets has been a major obstacle to implementing a support system to spot criminals from their gunshots by leveraging deep learning models. Since the occurrence of gunshot is rare and unpredictable, it is impractical to collect gun sounds in the real world. As an alternative, gun sounds can be obtained from an FPS game that is designed to mimic real-world warfare. The recent FPS game offers a realistic environment where we can safely collect gunshot data while simulating even dangerous situations. By exploiting the advantage of the game environment, we construct a gunshot dataset, namely BGG, for the firearm classification and gunshot localization tasks. The BGG dataset consists of 37 different types of firearms, distances, and directions between the sound source and a receiver. We carefully verify that the in-game gunshot data has sufficient information to identify the location and type of gunshots by training several sound classification and localization baselines on the BGG dataset. Afterward, we demonstrate that the accuracy of real-world firearm classification and localization tasks can be enhanced by utilizing the BGG dataset.


【18】 Comparison of Soft and Hard Target RNN-T Distillation for Large-scale  ASR

标题:大型ASR软硬靶RNN-T精馏的比较

链接:https://arxiv.org/abs/2210.05793

* 与cs.SD语音【12】为同一篇

作者:Dongseong Hwang,Khe Chai Sim,Yu Zhang,Trevor Strohman
机构:Google LLC, USA
备注:8 pages, 1 figure
摘要:知识提炼是一种有效的机器学习技术,它可以将知识从教师模型转移到更小的学生模型,特别是在未标记数据的情况下。本文主要研究了在自动语音识别中广泛应用的RNN-T模型的知识提取问题。具体而言,我们比较了使用软目标和硬目标提取在LibriSpeech/LibriLight公共数据集(60 k小时)和我们的内部数据(600 k小时)上训练大规模RNN-T模型的情况。我们发现,当教师和学生的结构不同时,如大教师和小分流学生,硬目标更有效。另一方面,软目标蒸馏在自我培训场景中效果更好,比如迭代的大型教师培训。对于具有0.6B权重的大模型,我们使用带软目标蒸馏的噪声学生训练在LibriSpeech上实现了新的SoTA单词错误率(WER)(相对于dev-other提高8%)。它还允许我们的生产教师不断适应新的数据域。
摘要:Knowledge distillation is an effective machine learning technique to transfer knowledge from a teacher model to a smaller student model, especially with unlabeled data. In this paper, we focus on knowledge distillation for the RNN-T model, which is widely used in state-of-the-art (SoTA) automatic speech recognition (ASR). Specifically, we compared using soft and hard target distillation to train large-scaleRNN-T models on the LibriSpeech/LibriLight public dataset (60k hours) and our in-house data (600k hours). We found that hard tar-gets are more effective when the teacher and student have different architecture, such as large teacher and small streaming student. On the other hand, soft target distillation works better in self-training scenario like iterative large teacher training. For a large model with0.6B weights, we achieve a new SoTA word error rate (WER) on LibriSpeech (8% relative improvement on dev-other) using Noisy Student Training with soft target distillation. It also allows our production teacher to adapt new data domain continuously.


【19】 Scaling Up Deliberation for Multilingual ASR

标题:扩大对多语言ASR的审议

链接:https://arxiv.org/abs/2210.05785

* 与cs.SD语音【13】为同一篇

作者:Ke Hu,Bo Li,Tara N. Sainath
机构:Google LLC, USA
摘要:多语言端到端自动语音识别模型由于其训练和部署简单而具有吸引力。与单语模型相比,最近对这种模型的大规模训练的工作已经显示出有希望的结果。然而,在单遍设置中,工作通常集中在多语言模型本身。本文主要研究多语言语音识别中的二次通过审议问题。我们提议的审议是多语种的,文本编码器对来自多种语言的假设文本进行编码,而解码器处理多语言文本和音频。研究了审议文本编码器和解码器的可伸缩性,并比较了审议解码器和第一遍级联编码器的可伸缩性。我们发现,与单遍模型相比,审议在9种语言上的平均WER提高了4%。通过将审议的大小增加到1B个参数,平均WER改善增加到9%,对于某些语言最高可达14%。我们的审议重新评分器基于Transformer层,并且可以在重新评分期间并行化。
摘要:Multilingual end-to-end automatic speech recognition models are attractive due to its simplicity in training and deployment. Recent work on large-scale training of such models has shown promising results compared to monolingual models. However, the work often focuses on multilingual models themselves in a single-pass setup. In this work, we investigate second-pass deliberation for multilingual speech recognition. Our proposed deliberation is multilingual, i.e., the text encoder encodes hypothesis text from multiple languages, and the decoder attends to multilingual text and audio. We investigate scaling the deliberation text encoder and decoder, and compare scaling the deliberation decoder and the first-pass cascaded encoder. We show that deliberation improves the average WER on 9 languages by 4% relative compared to the single-pass model. By increasing the size of the deliberation up to 1B parameters, the average WER improvement increases to 9%, with up to 14% for certain languages. Our deliberation rescorer is based on transformer layers and can be parallelized during rescoring.


机器翻译,仅供参考