今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理14篇。

cs.SD语音

【1】 DrumGAN VST: A Plugin for Drum Sound Analysis/Synthesis With  Autoencoding Generative Adversarial Networks

标题:DrumGan VST:一个自动编码生成对抗网络的鼓声分析/合成插件

链接:https://arxiv.org/abs/2206.14723

作者:Javier Nistal,Cyran Aouameur,Ithan Velarde,Stefan Lattner
机构:Paris  2National Instituteof Applied Sciences
备注:7 pages, 2 figures, 3 tables, ICML2022 Machine Learning for Audio Synthesis (MLAS) Workshop, for sound examples visit this https URL
摘要:在当代流行音乐制作中,鼓的声音设计通常是通过在声音库中浏览和处理预先录制的样本来完成的。人们还可以使用专门的合成硬件,通常通过低水平、无音乐意义的参数进行控制。今天,深度学习领域提供了通过学习高级特征来控制合成过程的方法,并允许生成各种各样的声音。在本文中,我们介绍了DrumGAN VST,这是一个使用生成对抗网络合成鼓声的插件。DrumGAN VST以44.1 kHz的采样率音频运行,提供独立和连续的乐器类控制,并具有编码神经网络,将声音映射到GAN的潜在空间,从而能够重新合成和操纵预先存在的鼓声。我们提供了许多声音示例和建议的VST插件的演示。
摘要:In contemporary popular music production, drum sound design is commonly performed by cumbersome browsing and processing of pre-recorded samples in sound libraries. One can also use specialized synthesis hardware, typically controlled through low-level, musically meaningless parameters. Today, the field of Deep Learning offers methods to control the synthesis process via learned high-level features and allows generating a wide variety of sounds. In this paper, we present DrumGAN VST, a plugin for synthesizing drum sounds using a Generative Adversarial Network. DrumGAN VST operates on 44.1 kHz sample-rate audio, offers independent and continuous instrument class controls, and features an encoding neural network that maps sounds into the GAN's latent space, enabling resynthesis and manipulation of pre-existing drum sounds. We provide numerous sound examples and a demo of the proposed VST plugin.


【2】 Improving Deliberation by Text-Only and Semi-Supervised Training

标题:通过纯文本和半监督训练提高思考力

链接:https://arxiv.org/abs/2206.14716

作者:Ke Hu,Tara N. Sainath,Yanzhang He,Rohit Prabhavalkar,Trevor Strohman,Sepand Mavandadi,Weiran Wang
机构:Google LLC, USA
备注:Accepted by Interspeech 2022
摘要:由于未标记文本和语音数据的广泛可用性,基于纯音频数据的纯文本和半监督训练近年来得到了广泛的应用。在这项工作中,我们建议将纯文本和半监督训练纳入基于注意的审议模型。通过将纯文本数据用于训练来自transformer(BERT)的双向编码器表示用于审议文本编码器,以及使用联合声学和文本解码器(JATD)和半监督训练的大规模文本到语音和纯音频话语,我们在各种任务中比基线审议减少了4%-12%。与最先进的语言模型(LM)重新排序方法相比,审议模型相对减少了11%的谷歌语音搜索WER。我们表明,与具有合理终点延迟的最先进的LM rescorer相比,审议模型还实现了积极的人类并肩评估。
摘要:Text-only and semi-supervised training based on audio-only data has gained popularity recently due to the wide availability of unlabeled text and speech data. In this work, we propose incorporating text-only and semi-supervised training into an attention-based deliberation model. By incorporating text-only data in training a bidirectional encoder representation from transformer (BERT) for the deliberation text encoder, and large-scale text-to-speech and audio-only utterances using joint acoustic and text decoder (JATD) and semi-supervised training, we achieved 4%-12% WER reduction for various tasks compared to the baseline deliberation. Compared to a state-of-the-art language model (LM) rescoring method, the deliberation model reduces the Google Voice Search WER by 11% relative. We show that the deliberation model also achieves a positive human side-by-side evaluation compared to the state-of-the-art LM rescorer with reasonable endpointer latencies.


【3】 The THUEE System Description for the IARPA OpenASR21 Challenge

标题:IARPA OpenASR21挑战赛的THUEE系统描述

链接:https://arxiv.org/abs/2206.14660

作者:Jing Zhao,Haoyu Wang,Jinpeng Li,Shuzhou Chai,Guan-Bo Wang,Guoguo Chen,Wei-Qiang Zhang
机构:Beijing National Research Center for Information Science and Technology, Department of Electronic Engineering, Tsinghua University, Beijing , China
备注:accepted by INTERSPEECH 2022
摘要:本文描述了THUEE团队针对IARPA开放式自动语音识别挑战赛(OpenASR21)的语音识别系统,并进行了进一步的实验探索。我们在约束和约束+训练条件下都取得了优异的效果。对于受限的训练条件,我们构建了基于标准混合体系结构的基本ASR系统。为了缓解词汇表外(OOV)问题,我们使用字形到音素(G2P)技术对OOV和潜在新词的发音词典进行了扩展。采用CNN-TDNN-F和CNN-TDNN-F-A等标准声学模型结构。此外,还应用了多种数据扩充技术。对于约束+训练条件,我们使用自监督学习框架wav2vec2.0。我们在公开的预训练模型XLSR-53的基础上,使用连接主义时间分类(CTC)标准对各种微调技术进行了实验。我们发现,在将wav2vec2.0预训练模型应用于基于编码-解码器的CTC/Attention ASR体系结构时,前端特征提取器起着重要作用。通过使用目标语言中经过微调的CTC模型作为前端特征提取器,可以实现额外的改进。
摘要:This paper describes the THUEE team's speech recognition system for the IARPA Open Automatic Speech Recognition Challenge (OpenASR21), with further experiment explorations. We achieve outstanding results under both the Constrained and Constrained-plus training conditions. For the Constrained training condition, we construct our basic ASR system based on the standard hybrid architecture. To alleviate the Out-Of-Vocabulary (OOV) problem, we extend the pronunciation lexicon using Grapheme-to-Phoneme (G2P) techniques for both OOV and potential new words. Standard acoustic model structures such as CNN-TDNN-F and CNN-TDNN-F-A are adopted. In addition, multiple data augmentation techniques are applied. For the Constrained-plus training condition, we use the self-supervised learning framework wav2vec2.0. We experiment with various fine-tuning techniques with the Connectionist Temporal Classification (CTC) criterion on top of the publicly available pre-trained model XLSR-53. We find that the frontend feature extractor plays an important role when applying the wav2vec2.0 pre-trained model to the encoder-decoder based CTC/Attention ASR architecture. Extra improvements can be achieved by using the CTC model finetuned in the target language as the frontend feature extractor.


【4】 Language-Based Audio Retrieval with Converging Tied Layers and  Contrastive Loss

标题:基于语言的层级收敛和对比损失的音频检索

链接:https://arxiv.org/abs/2206.14659

作者:Andrew Koh,Eng Siong Chng
机构:Nanyang Technological University
摘要:在本文中,我们处理DCASE 2022中提出的新的基于语言的音频检索任务。首先,我们介绍了一种简单、可扩展的体系结构,它将音频和文本编码器连接在一起。其次,我们表明,使用这种架构以及对比损失可以使模型显著优于基线模型的性能。最后,除了极低的训练内存需求外,我们还可以使用预先训练的模型,而无需对其进行微调。我们对我们的方法进行了测试,结果表明,结合使用我们的方法可以显著优于基线得分。
摘要:In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and text encoder together. Secondly, we show that using this architecture along with contrastive loss allows the model to significantly beat the performance of the baseline model. Finally, in addition to having an extremely low training memory requirement, we are able to use pretrained models as it is without needing to finetune them. We test our methods and show that using a combination of our methods beats the baseline scores significantly.


【5】 Finstreder: Simple and fast Spoken Language Understanding with Finite  State Transducers using modern Speech-to-Text models

标题:Finstreder:使用现代语音到文本模型的有限状态转换器实现简单快速的口语理解

链接:https://arxiv.org/abs/2206.14589

作者:Daniel Bermuth,Alexander Poeppel,Wolfgang Reif
机构:University of Augsburg, Institute for Software & Systems Engineering
摘要:在口语理解(SLU)中,任务是从音频命令中提取重要信息,如用户希望系统做什么的意图以及位置或数字等特殊实体。本文提出了一种将意图和实体嵌入有限状态传感器的简单方法,并结合预训练的通用语音-文本模型,允许在无需任何额外训练的情况下构建SLU模型。构建这些模型非常快,只需几秒钟。它也是完全独立于语言的。通过对不同基准的比较,表明该方法的性能优于其他多种资源要求更高的SLU方法。
摘要:In Spoken Language Understanding (SLU) the task is to extract important information from audio commands, like the intent of what a user wants the system to do and special entities like locations or numbers. This paper presents a simple method for embedding intents and entities into Finite State Transducers, and, in combination with a pretrained general-purpose Speech-to-Text model, allows building SLU-models without any additional training. Building those models is very fast and only takes a few seconds. It is also completely language independent. With a comparison on different benchmarks it is shown that this method can outperform multiple other, more resource demanding SLU approaches.


【6】 DDKtor: Automatic Diadochokinetic Speech Analysis

标题:DDKtor:自动对偶动态语音分析

链接:https://arxiv.org/abs/2206.14639

作者:Yael Segal,Kasia Hitczenko,Matthew Goldrick,Adam Buchwald,Angela Roberts,Joseph Keshet
机构:Keshet, Technion–Israel Institute of Technology, Israel, Laboratoire de Sciences Cognitives et Psycholinguistique, D´epartement d’Etudes Cognitives, ENS, EHESS, CNRS, PSL University, France, Department of Linguistics, Northwestern University, IL, USA
备注:Accepted to Interspeech 2022
摘要:Diadochokinetic speech tasks(DDK),即参与者反复产生音节,通常被用作言语运动障碍评估的一部分。这些研究依赖于时间密集、主观的手动分析,并且只能提供粗粒度的语音图像。本文提出了两种深度神经网络模型,可以自动从未注释、未翻译的语音中分割辅音和元音。这两种模型都处理原始波形,并使用卷积层进行特征提取。第一个模型基于LSTM分类器,然后是完全连接的层,而第二个模型添加更多的卷积层,然后是完全连接的层。这些模型预测的分段用于获得语音速率和声音持续时间的度量。在一个年轻健康个体数据集上的结果表明,我们的LSTM模型优于当前最先进的系统,其性能与训练有素的人类注释员相当。此外,LSTM模型在使用帕金森病数据集对看不见的老年人进行评估时,也提供了与经过训练的人类注释员相当的结果。
摘要:Diadochokinetic speech tasks (DDK), in which participants repeatedly produce syllables, are commonly used as part of the assessment of speech motor impairments. These studies rely on manual analyses that are time-intensive, subjective, and provide only a coarse-grained picture of speech. This paper presents two deep neural network models that automatically segment consonants and vowels from unannotated, untranscribed speech. Both models work on the raw waveform and use convolutional layers for feature extraction. The first model is based on an LSTM classifier followed by fully connected layers, while the second model adds more convolutional layers followed by fully connected layers. These segmentations predicted by the models are used to obtain measures of speech rate and sound duration. Results on a young healthy individuals dataset show that our LSTM model outperforms the current state-of-the-art systems and performs comparably to trained human annotators. Moreover, the LSTM model also presents comparable results to trained human annotators when evaluated on unseen older individuals with Parkinson's Disease dataset.


【7】 A light-weight full-band speech enhancement model

标题:一种轻量级全频带语音增强模型

链接:https://arxiv.org/abs/2206.14524

作者:Qinwen Hu,Zhongshu Hou,Xiaohuai Le,Jing Lu
机构:Key Laboratory of Modern Acoustics, Nanjing University, Nanjing, China
摘要:基于深度神经网络的全频段语音增强系统面临着计算资源需求高和频率分布不平衡的挑战。本文提出了一种轻量级全波段模型,该模型具有两种专用策略,即可学习的光谱压缩映射用于更有效的高频光谱信息压缩,以及利用多头注意机制更有效地建模全局光谱模式。实验验证了所提策略的有效性,并表明所提模型仅需0.89M的参数即可获得有竞争力的性能。
摘要:Deep neural network based full-band speech enhancement systems face challenges of high demand of computational resources and imbalanced frequency distribution. In this paper, a light-weight full-band model is proposed with two dedicated strategies, i.e., a learnable spectral compression mapping for more effective high-band spectral information compression, and the utilization of the multi-head attention mechanism for more effective modeling of the global spectral pattern. Experiments validate the efficacy of the proposed strategies and show that the proposed model achieves competitive performance with only 0.89M parameters.


【8】 Comparing Conventional Pitch Detection Algorithms with a Neural Network  Approach

标题:传统基音检测算法与神经网络方法的比较

链接:https://arxiv.org/abs/2206.14357

作者:Anja Kroon
机构:ECSE , Speech Communications Final Project, Dept. of Electrical and Computer Engineering, McGill University, Montreal, Quebec, Canada
备注:6 pages, 11 figures
摘要:尽管进行了大量的研究,但传统的基音预测方法仍不完善。随着神经网络(NNs)的出现,研究人员希望创建一种性能优于传统方法的基于神经网络的基音预测器。本文比较了pYIN、YAAPT和CREPE三种基音检测算法。pYIN和YAAPT是考虑时域和频域处理的常规方法。CREPE利用经过数据训练的深度卷积神经网络来估计音高。它涉及6个紧密连接的卷积隐藏层,并确定给定输入信号的基音概率。将CREPE表示的神经网络基音预测器的性能与pYIN和YAAPT表示的更经典的方法进行了比较。优数(FOM)将包括清音到浊音错误、浊音到浊音错误、总音高错误和细音高错误的数量。
摘要:Despite much research, traditional methods to pitch prediction are still not perfect. With the emergence of neural networks (NNs), researchers hope to create a NN-based pitch predictor that outperforms traditional methods. Three pitch detection algorithms (PDAs), pYIN, YAAPT, and CREPE are compared in this paper. pYIN and YAAPT are conventional approaches considering time domain and frequency domain processing. CREPE utilizes a data-trained deep convolutional neural network to estimate pitch. It involves 6 densely connected convolutional hidden layers and determines pitch probabilities for a given input signal. The performance of CREPE representing neural network pitch predictors is compared to more classical approaches represented by pYIN and YAAPT. The figure of merit (FOM) will include the amount of unvoiced-to-voiced errors, voiced-to-voiced errors, gross pitch errors, and fine pitch errors.



eess.AS音频处理

【1】 Nextformer: A ConvNeXt Augmented Conformer For End-To-End Speech  Recognition

标题:Nextform:一种端到端语音识别的ConvNeXt增强型整形器

链接:https://arxiv.org/abs/2206.14747

作者:Yongjun Jiang,Jian Yu,Wenwen Yang,Bihong Zhang,Yanfeng Wang
机构:AI Interaction Division, Tencent PCG
备注:5 pages, 1 figure
摘要:一致性模型在端到端语音识别中取得了最先进的(SOTA)结果。然而Conformer主要关注时间建模,而对语音特征的时频特性关注较少。在本文中,我们用convenxt扩充Conformer,并提出Nextformer结构。为了利用时频语音特征中包含的信息,我们使用ConvNeXt块的堆栈来代替Conformer中常用的子采样模块。此外,我们在构象层中间插入了一个额外的下采样模块,以使我们的模型高效准确。我们在两个开放数据集AISHELL-1和WenetSpeech上进行了实验。在AISHELL-1上,与Conformer基线相比,Nextformer在非流模式和流模式下分别获得7.3%和6.3%的相对CER改进,而在更大的WenetSpeech数据集上,Nextformer在非流模式和流模式下分别获得5.0%~6.5%和7.5%~14.6%的相对改进,同时保持计算成本与Conformer相当。据我们所知,拟议的Nextformer模型在AISHELL-1(CER 4.06%)和WenetSpeech(CER 7.56%/11.29%)上实现了SOTA结果。
摘要:Conformer models have achieved state-of-the-art(SOTA) results in end-to-end speech recognition. However Conformer mainly focuses on temporal modeling while pays less attention on time-frequency property of speech feature. In this paper we augment Conformer with ConvNeXt and propose Nextformer structure. We use stacks of ConvNeXt block to replace the commonly used subsampling module in Conformer for utilizing the information contained in time-frequency speech feature. Besides, we insert an additional downsampling module in middle of Conformer layers to make our model efficient and accurate. We conduct experiments on two opening datasets, AISHELL-1 and WenetSpeech. On AISHELL-1, compared to Conformer baselines, Nextformer obtains 7.3% and 6.3% relative CER improvements in non-streaming and streaming mode respectively, and on a much larger WenetSpeech dataset, Nextformer gives 5.0%~6.5% and 7.5%~14.6% relative improvements in non-streaming and streaming mode, while keep the computational cost FLOPs comparable to Conformer. To the best of our knowledge, the proposed Nextformer model achieves SOTA results on AISHELL-1(CER 4.06%) and WenetSpeech(CER 7.56% /11.29%).


【2】 Simple and Effective Multi-sentence TTS with Expressive and Coherent  Prosody

标题:简单有效的多句语篇,韵律表达连贯

链接:https://arxiv.org/abs/2206.14643

作者:Peter Makarov,Ammar Abbas,Mateusz Łajszczak,Arnaud Joly,Sri Karlapati,Alexis Moinet,Thomas Drugman,Penny Karanasou
机构:Alexa AI, Amazon
备注:Accepted to be published in the Proceedings of InterSpeech 2022
摘要:对于现代文语转换(TTS)系统来说,生成富有表现力且符合语境的韵律仍然是一个挑战。这对于长时间、多句子输入尤其明显。在本文中,我们研究了基于变换器的类快速语音系统的简单扩展,目的是改进多句TTS的韵律。我们发现,长上下文、强大的文本特征以及对多说话人数据的训练都可以改善韵律。更有趣的是,它们产生了协同效应。长上下文消除韵律歧义,提高连贯性,发挥Transformer的优势。从功能强大的语言模型(如BERT)中微调单词级功能,似乎可以从更多的训练数据中获益,这些数据在多说话人环境中很容易获得。我们研究停顿和节奏的客观指标,并对语音自然度进行彻底的主观评估。我们的主系统整合了所有扩展,取得了始终如一的良好效果,包括在统计上显著改善了所有竞争对手的语音自然度。
摘要:Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs. In this paper, we examine simple extensions to a Transformer-based FastSpeech-like system, with the goal of improving prosody for multi-sentence TTS. We find that long context, powerful text features, and training on multi-speaker data all improve prosody. More interestingly, they result in synergies. Long context disambiguates prosody, improves coherence, and plays to the strengths of Transformers. Fine-tuning word-level features from a powerful language model, such as BERT, appears to profit from more training data, readily available in a multi-speaker setting. We look into objective metrics on pausing and pacing and perform thorough subjective evaluations for speech naturalness. Our main system, which incorporates all the extensions, achieves consistently strong results, including statistically significant improvements in speech naturalness over all its competitors.


【3】 DDKtor: Automatic Diadochokinetic Speech Analysis

标题:DDKtor:自动对偶动态语音分析

链接:https://arxiv.org/abs/2206.14639

* 与cs.SD语音【6】为同一篇

作者:Yael Segal,Kasia Hitczenko,Matthew Goldrick,Adam Buchwald,Angela Roberts,Joseph Keshet
机构:Keshet, Technion–Israel Institute of Technology, Israel, Laboratoire de Sciences Cognitives et Psycholinguistique, D´epartement d’Etudes Cognitives, ENS, EHESS, CNRS, PSL University, France, Department of Linguistics, Northwestern University, IL, USA
备注:Accepted to Interspeech 2022
摘要:Diadochokinetic speech tasks(DDK),即参与者反复产生音节,通常被用作言语运动障碍评估的一部分。这些研究依赖于时间密集、主观的手动分析,并且只能提供粗粒度的语音图像。本文提出了两种深度神经网络模型,可以自动从未注释、未翻译的语音中分割辅音和元音。这两种模型都处理原始波形,并使用卷积层进行特征提取。第一个模型基于LSTM分类器,然后是完全连接的层,而第二个模型添加更多的卷积层,然后是完全连接的层。这些模型预测的分段用于获得语音速率和声音持续时间的度量。在一个年轻健康个体数据集上的结果表明,我们的LSTM模型优于当前最先进的系统,其性能与训练有素的人类注释员相当。此外,LSTM模型在使用帕金森病数据集对看不见的老年人进行评估时,也提供了与经过训练的人类注释员相当的结果。
摘要:Diadochokinetic speech tasks (DDK), in which participants repeatedly produce syllables, are commonly used as part of the assessment of speech motor impairments. These studies rely on manual analyses that are time-intensive, subjective, and provide only a coarse-grained picture of speech. This paper presents two deep neural network models that automatically segment consonants and vowels from unannotated, untranscribed speech. Both models work on the raw waveform and use convolutional layers for feature extraction. The first model is based on an LSTM classifier followed by fully connected layers, while the second model adds more convolutional layers followed by fully connected layers. These segmentations predicted by the models are used to obtain measures of speech rate and sound duration. Results on a young healthy individuals dataset show that our LSTM model outperforms the current state-of-the-art systems and performs comparably to trained human annotators. Moreover, the LSTM model also presents comparable results to trained human annotators when evaluated on unseen older individuals with Parkinson's Disease dataset.


【4】 Contextual Density Ratio for Language Model Biasing of Sequence to  Sequence ASR Systems

标题:序列对序列ASR系统语言模型偏差的语境密度比

链接:https://arxiv.org/abs/2206.14623

作者:Jesús Andrés-Ferrer,Dario Albesano,Puming Zhan,Paul Vozila
机构:Nuance Communications, Inc.,Valencia, Spain,Torino, Italy,Burlington, MA, USA
备注:None
摘要:端2端(E2E)模型由于其性能和优势,在一些ASR任务中越来越流行。这些E2E模型直接近似于给定声学输入的标记的后验分布。因此,E2E系统在输出标记上隐式定义了一个语言模型(LM),这使得独立训练的语言模型的利用不如传统ASR系统那么简单。这使得E2E ASR系统很难动态地适应上下文概要文件,以便更好地识别命名实体等特殊单词。在这项工作中,我们提出了一种上下文密度比方法,用于训练上下文感知的E2E模型和使语言模型适应命名实体。我们将上述技术应用于E2E ASR系统,该系统转录医生和患者的对话,以便更好地使E2E系统适应对话中的姓名。我们提出的技术在不降低整个测试集的整体识别准确率的情况下,与E2E基线相比,姓名的相对改善率高达46.5%。此外,它还比上下文浅层融合基线高出22.1%。
摘要:End-2-end (E2E) models have become increasingly popular in some ASR tasks because of their performance and advantages. These E2E models directly approximate the posterior distribution of tokens given the acoustic inputs. Consequently, the E2E systems implicitly define a language model (LM) over the output tokens, which makes the exploitation of independently trained language models less straightforward than in conventional ASR systems. This makes it difficult to dynamically adapt E2E ASR system to contextual profiles for better recognizing special words such as named entities. In this work, we propose a contextual density ratio approach for both training a contextual aware E2E model and adapting the language model to named entities. We apply the aforementioned technique to an E2E ASR system, which transcribes doctor and patient conversations, for better adapting the E2E system to the names in the conversations. Our proposed technique achieves a relative improvement of up to 46.5% on the names over an E2E baseline without degrading the overall recognition accuracy of the whole test set. Moreover, it also surpasses a contextual shallow fusion baseline by 22.1 % relative.


【5】 On the Prediction Network Architecture in RNN-T for ASR

标题:面向ASR的RNN-T预测网络体系结构研究

链接:https://arxiv.org/abs/2206.14618

作者:Dario Albesano,Jesús Andrés-Ferrer,Nicola Ferri,Puming Zhan
机构:Nuance Communications, Inc.
备注:To appear at Interspeech 2022
摘要:RNN-T模型由于其在在线流模式下的竞争力和操作能力,在文献和商业系统中得到了普及。在这项工作中,我们对单调和原始RNN-T模型的几种预测网络结构进行了广泛的比较研究。我们比较了4种基于通用最先进Conformer编码器的预测网络,并报告了在Librispeech和内部医疗对话数据集上获得的结果。我们的研究涵盖了离线批量模式和在线流媒体场景。与之前的一些工作相比,我们的结果表明,当与Conformer编码器一起用作预测网络时,Transformer并不总是优于LSTM。受我们记分板的启发,我们提出了一种新的简单预测网络体系结构N-Concat,该体系结构在我们的在线流媒体基准测试中优于其他体系结构。Transformer和n-gram简化体系结构的性能非常相似,但在以前的上下文中有一些重要的不同行为。总体而言,与LSTM基线相比,我们获得了高达4.1%的相对WER改善,同时将预测网络参数减少了近一个数量级(8.4倍)。
摘要:RNN-T models have gained popularity in the literature and in commercial systems because of their competitiveness and capability of operating in online streaming mode. In this work, we conduct an extensive study comparing several prediction network architectures for both monotonic and original RNN-T models. We compare 4 types of prediction networks based on a common state-of-the-art Conformer encoder and report results obtained on Librispeech and an internal medical conversation data set. Our study covers both offline batch-mode and online streaming scenarios. In contrast to some previous works, our results show that Transformer does not always outperform LSTM when used as prediction network along with Conformer encoder. Inspired by our scoreboard, we propose a new simple prediction network architecture, N-Concat, that outperforms the others in our on-line streaming benchmark. Transformer and n-gram reduced architectures perform very similarly yet with some important distinct behaviour in terms of previous context. Overall we obtained up to 4.1 % relative WER improvement compared to our LSTM baseline, while reducing prediction network parameters by nearly an order of magnitude (8.4 times).


【6】 A light-weight full-band speech enhancement model

标题:一种轻量级全频带语音增强模型

链接:https://arxiv.org/abs/2206.14524

* 与cs.SD语音【7】为同一篇

作者:Qinwen Hu,Zhongshu Hou,Xiaohuai Le,Jing Lu
机构:Key Laboratory of Modern Acoustics, Nanjing University, Nanjing, China
摘要:基于深度神经网络的全频段语音增强系统面临着计算资源需求高和频率分布不平衡的挑战。本文提出了一种轻量级全波段模型,该模型具有两种专用策略,即可学习的光谱压缩映射用于更有效的高频光谱信息压缩,以及利用多头注意机制更有效地建模全局光谱模式。实验验证了所提策略的有效性,并表明所提模型仅需0.89M的参数即可获得有竞争力的性能。
摘要:Deep neural network based full-band speech enhancement systems face challenges of high demand of computational resources and imbalanced frequency distribution. In this paper, a light-weight full-band model is proposed with two dedicated strategies, i.e., a learnable spectral compression mapping for more effective high-band spectral information compression, and the utilization of the multi-head attention mechanism for more effective modeling of the global spectral pattern. Experiments validate the efficacy of the proposed strategies and show that the proposed model achieves competitive performance with only 0.89M parameters.


【7】 Comparing Conventional Pitch Detection Algorithms with a Neural Network  Approach

标题:传统基音检测算法与神经网络方法的比较

链接:https://arxiv.org/abs/2206.14357

* 与cs.SD语音【8】为同一篇

作者:Anja Kroon
机构:ECSE , Speech Communications Final Project, Dept. of Electrical and Computer Engineering, McGill University, Montreal, Quebec, Canada
备注:6 pages, 11 figures
摘要:尽管进行了大量的研究,但传统的基音预测方法仍不完善。随着神经网络(NNs)的出现,研究人员希望创建一种性能优于传统方法的基于神经网络的基音预测器。本文比较了pYIN、YAAPT和CREPE三种基音检测算法。pYIN和YAAPT是考虑时域和频域处理的常规方法。CREPE利用经过数据训练的深度卷积神经网络来估计音高。它涉及6个紧密连接的卷积隐藏层,并确定给定输入信号的基音概率。将CREPE表示的神经网络基音预测器的性能与pYIN和YAAPT表示的更经典的方法进行了比较。优数(FOM)将包括清音到浊音错误、浊音到浊音错误、总音高错误和细音高错误的数量。
摘要:Despite much research, traditional methods to pitch prediction are still not perfect. With the emergence of neural networks (NNs), researchers hope to create a NN-based pitch predictor that outperforms traditional methods. Three pitch detection algorithms (PDAs), pYIN, YAAPT, and CREPE are compared in this paper. pYIN and YAAPT are conventional approaches considering time domain and frequency domain processing. CREPE utilizes a data-trained deep convolutional neural network to estimate pitch. It involves 6 densely connected convolutional hidden layers and determines pitch probabilities for a given input signal. The performance of CREPE representing neural network pitch predictors is compared to more classical approaches represented by pYIN and YAAPT. The figure of merit (FOM) will include the amount of unvoiced-to-voiced errors, voiced-to-voiced errors, gross pitch errors, and fine pitch errors.


【8】 DrumGAN VST: A Plugin for Drum Sound Analysis/Synthesis With  Autoencoding Generative Adversarial Networks

标题:DrumGan VST:一个自动编码生成对抗网络的鼓声分析/合成插件

链接:https://arxiv.org/abs/2206.14723

* 与cs.SD语音【1】为同一篇

作者:Javier Nistal,Cyran Aouameur,Ithan Velarde,Stefan Lattner
机构:Paris  2National Instituteof Applied Sciences
备注:7 pages, 2 figures, 3 tables, ICML2022 Machine Learning for Audio Synthesis (MLAS) Workshop, for sound examples visit this https URL
摘要:在当代流行音乐制作中,鼓的声音设计通常是通过在声音库中浏览和处理预先录制的样本来完成的。人们还可以使用专门的合成硬件,通常通过低水平、无音乐意义的参数进行控制。今天,深度学习领域提供了通过学习高级特征来控制合成过程的方法,并允许生成各种各样的声音。在本文中,我们介绍了DrumGAN VST,这是一个使用生成对抗网络合成鼓声的插件。DrumGAN VST以44.1 kHz的采样率音频运行,提供独立和连续的乐器类控制,并具有编码神经网络,将声音映射到GAN的潜在空间,从而能够重新合成和操纵预先存在的鼓声。我们提供了许多声音示例和建议的VST插件的演示。
摘要:In contemporary popular music production, drum sound design is commonly performed by cumbersome browsing and processing of pre-recorded samples in sound libraries. One can also use specialized synthesis hardware, typically controlled through low-level, musically meaningless parameters. Today, the field of Deep Learning offers methods to control the synthesis process via learned high-level features and allows generating a wide variety of sounds. In this paper, we present DrumGAN VST, a plugin for synthesizing drum sounds using a Generative Adversarial Network. DrumGAN VST operates on 44.1 kHz sample-rate audio, offers independent and continuous instrument class controls, and features an encoding neural network that maps sounds into the GAN's latent space, enabling resynthesis and manipulation of pre-existing drum sounds. We provide numerous sound examples and a demo of the proposed VST plugin.


【9】 Improving Deliberation by Text-Only and Semi-Supervised Training

标题:通过纯文本和半监督训练提高思考力

链接:https://arxiv.org/abs/2206.14716

* 与cs.SD语音【2】为同一篇

作者:Ke Hu,Tara N. Sainath,Yanzhang He,Rohit Prabhavalkar,Trevor Strohman,Sepand Mavandadi,Weiran Wang
机构:Google LLC, USA
备注:Accepted by Interspeech 2022
摘要:由于未标记文本和语音数据的广泛可用性,基于纯音频数据的纯文本和半监督训练近年来得到了广泛的应用。在这项工作中,我们建议将纯文本和半监督训练纳入基于注意的审议模型。通过将纯文本数据用于训练来自transformer(BERT)的双向编码器表示用于审议文本编码器,以及使用联合声学和文本解码器(JATD)和半监督训练的大规模文本到语音和纯音频话语,我们在各种任务中比基线审议减少了4%-12%。与最先进的语言模型(LM)重新排序方法相比,审议模型相对减少了11%的谷歌语音搜索WER。我们表明,与具有合理终点延迟的最先进的LM rescorer相比,审议模型还实现了积极的人类并肩评估。
摘要:Text-only and semi-supervised training based on audio-only data has gained popularity recently due to the wide availability of unlabeled text and speech data. In this work, we propose incorporating text-only and semi-supervised training into an attention-based deliberation model. By incorporating text-only data in training a bidirectional encoder representation from transformer (BERT) for the deliberation text encoder, and large-scale text-to-speech and audio-only utterances using joint acoustic and text decoder (JATD) and semi-supervised training, we achieved 4%-12% WER reduction for various tasks compared to the baseline deliberation. Compared to a state-of-the-art language model (LM) rescoring method, the deliberation model reduces the Google Voice Search WER by 11% relative. We show that the deliberation model also achieves a positive human side-by-side evaluation compared to the state-of-the-art LM rescorer with reasonable endpointer latencies.


【10】 The THUEE System Description for the IARPA OpenASR21 Challenge

标题:IARPA OpenASR21挑战赛的THUEE系统描述

链接:https://arxiv.org/abs/2206.14660

* 与cs.SD语音【3】为同一篇

作者:Jing Zhao,Haoyu Wang,Jinpeng Li,Shuzhou Chai,Guan-Bo Wang,Guoguo Chen,Wei-Qiang Zhang
机构:Beijing National Research Center for Information Science and Technology, Department of Electronic Engineering, Tsinghua University, Beijing , China
备注:accepted by INTERSPEECH 2022
摘要:本文描述了THUEE团队针对IARPA开放式自动语音识别挑战赛(OpenASR21)的语音识别系统,并进行了进一步的实验探索。我们在约束和约束+训练条件下都取得了优异的效果。对于受限的训练条件,我们构建了基于标准混合体系结构的基本ASR系统。为了缓解词汇表外(OOV)问题,我们使用字形到音素(G2P)技术对OOV和潜在新词的发音词典进行了扩展。采用CNN-TDNN-F和CNN-TDNN-F-A等标准声学模型结构。此外,还应用了多种数据扩充技术。对于约束+训练条件,我们使用自监督学习框架wav2vec2.0。我们在公开的预训练模型XLSR-53的基础上,使用连接主义时间分类(CTC)标准对各种微调技术进行了实验。我们发现,在将wav2vec2.0预训练模型应用于基于编码-解码器的CTC/Attention ASR体系结构时,前端特征提取器起着重要作用。通过使用目标语言中经过微调的CTC模型作为前端特征提取器,可以实现额外的改进。
摘要:This paper describes the THUEE team's speech recognition system for the IARPA Open Automatic Speech Recognition Challenge (OpenASR21), with further experiment explorations. We achieve outstanding results under both the Constrained and Constrained-plus training conditions. For the Constrained training condition, we construct our basic ASR system based on the standard hybrid architecture. To alleviate the Out-Of-Vocabulary (OOV) problem, we extend the pronunciation lexicon using Grapheme-to-Phoneme (G2P) techniques for both OOV and potential new words. Standard acoustic model structures such as CNN-TDNN-F and CNN-TDNN-F-A are adopted. In addition, multiple data augmentation techniques are applied. For the Constrained-plus training condition, we use the self-supervised learning framework wav2vec2.0. We experiment with various fine-tuning techniques with the Connectionist Temporal Classification (CTC) criterion on top of the publicly available pre-trained model XLSR-53. We find that the frontend feature extractor plays an important role when applying the wav2vec2.0 pre-trained model to the encoder-decoder based CTC/Attention ASR architecture. Extra improvements can be achieved by using the CTC model finetuned in the target language as the frontend feature extractor.


【11】 Language-Based Audio Retrieval with Converging Tied Layers and  Contrastive Loss

标题:基于语言的层级收敛和对比损失的音频检索

链接:https://arxiv.org/abs/2206.14659

* 与cs.SD语音【4】为同一篇

作者:Andrew Koh,Eng Siong Chng
机构:Nanyang Technological University
摘要:在本文中,我们处理DCASE 2022中提出的新的基于语言的音频检索任务。首先,我们介绍了一种简单、可扩展的体系结构,它将音频和文本编码器连接在一起。其次,我们表明,使用这种架构以及对比损失可以使模型显著优于基线模型的性能。最后,除了极低的训练内存需求外,我们还可以使用预先训练的模型,而无需对其进行微调。我们对我们的方法进行了测试,结果表明,结合使用我们的方法可以显著优于基线得分。
摘要:In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and text encoder together. Secondly, we show that using this architecture along with contrastive loss allows the model to significantly beat the performance of the baseline model. Finally, in addition to having an extremely low training memory requirement, we are able to use pretrained models as it is without needing to finetune them. We test our methods and show that using a combination of our methods beats the baseline scores significantly.


【12】 Finstreder: Simple and fast Spoken Language Understanding with Finite  State Transducers using modern Speech-to-Text models

标题:Finstreder:使用现代语音到文本模型的有限状态转换器实现简单快速的口语理解

链接:https://arxiv.org/abs/2206.14589

* 与cs.SD语音【5】为同一篇

作者:Daniel Bermuth,Alexander Poeppel,Wolfgang Reif
机构:University of Augsburg, Institute for Software & Systems Engineering
摘要:在口语理解(SLU)中,任务是从音频命令中提取重要信息,如用户希望系统做什么的意图以及位置或数字等特殊实体。本文提出了一种将意图和实体嵌入有限状态传感器的简单方法,并结合预训练的通用语音-文本模型,允许在无需任何额外训练的情况下构建SLU模型。构建这些模型非常快,只需几秒钟。它也是完全独立于语言的。通过对不同基准的比较,表明该方法的性能优于其他多种资源要求更高的SLU方法。
摘要:In Spoken Language Understanding (SLU) the task is to extract important information from audio commands, like the intent of what a user wants the system to do and special entities like locations or numbers. This paper presents a simple method for embedding intents and entities into Finite State Transducers, and, in combination with a pretrained general-purpose Speech-to-Text model, allows building SLU-models without any additional training. Building those models is very fast and only takes a few seconds. It is also completely language independent. With a comparison on different benchmarks it is shown that this method can outperform multiple other, more resource demanding SLU approaches.


【13】 Language-specific Characteristic Assistance for Code-switching Speech  Recognition

标题:用于码型转换语音识别的特定语言特征辅助

链接:https://arxiv.org/abs/2206.14580

作者:Tongtong Song,Qiang Xu,Meng Ge,Longbiao Wang,Hao Shi,Yongjie Lv,Yuqin Lin,Jianwu Dang
机构:Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University, Tianjin, China, Department of Electrical and Computer Engineering, National University of Singapore, Singapore
备注:Accepted by Interspeech 2022
摘要:双编码器结构成功地利用了两个特定于语言的编码器(LSE)进行代码切换语音识别。由于LSE由两个预先训练的语言特定模型(LSM)初始化,因此双编码器结构可以利用足够的单语数据并捕获单个语言属性。然而,现有的方法对LSE没有语言约束,并且没有充分利用LSM的特定语言知识。在本文中,我们提出了一种特定于语言的特征辅助(LSCA)方法来缓解上述问题。具体来说,在训练期间,我们引入两种特定于语言的损失作为语言约束,并为它们生成相应的特定于语言的目标。在解码过程中,我们通过结合两个LSM的输出概率和混合模型来获得最终预测,从而考虑LSM的解码能力。实验表明,无论是LSCA的训练方法还是解码方法都可以提高模型的性能。此外,通过结合LSCA的训练和解码方法,在代码切换测试集上的最佳结果可以获得高达15.4%的相对错误减少。此外,该方法可以很好地处理码切换语音识别任务,无需额外的共享参数,甚至无需基于两个预先训练好的LSM进行再训练。
摘要:Dual-encoder structure successfully utilizes two language-specific encoders (LSEs) for code-switching speech recognition. Because LSEs are initialized by two pre-trained language-specific models (LSMs), the dual-encoder structure can exploit sufficient monolingual data and capture the individual language attributes. However, existing methods have no language constraints on LSEs and underutilize language-specific knowledge of LSMs. In this paper, we propose a language-specific characteristic assistance (LSCA) method to mitigate the above problems. Specifically, during training, we introduce two language-specific losses as language constraints and generate corresponding language-specific targets for them. During decoding, we take the decoding abilities of LSMs into account by combining the output probabilities of two LSMs and the mixture model to obtain the final predictions. Experiments show that either the training or decoding method of LSCA can improve the model's performance. Furthermore, the best result can obtain up to 15.4% relative error reduction on the code-switching test set by combining the training and decoding methods of LSCA. Moreover, the system can process code-switching speech recognition tasks well without extra shared parameters or even retraining based on two pre-trained LSMs by using our method.


【14】 Bottleneck Low-rank Transformers for Low-resource Spoken Language  Understanding

标题:低资源口语理解的瓶颈--低阶转换器

链接:https://arxiv.org/abs/2206.14318

作者:Pu Wang,Hugo Van hamme
机构:Department of Electrical Engineering-ESAT, KU Leuven, Belgium
备注:Accepted by Interspeech 2022
摘要:端到端口语理解(SLU)系统受益于对大型语料库进行预训练,然后对特定于应用程序的数据进行微调。生成的模型对于边缘应用程序来说太大了。例如,基于BERT的系统包含超过1.1亿个参数。观察到模型被过度参数化,我们提出了精益Transformer结构,其中注意机制的维度使用组稀疏性自动降低。我们提出了一种变体,将学习到的注意子空间转移到注意瓶颈层。在低资源环境下,无需预训练,生成的紧凑型SLU模型的精度可与预训练的大型模型相媲美。
摘要:End-to-end spoken language understanding (SLU) systems benefit from pretraining on large corpora, followed by fine-tuning on application-specific data. The resulting models are too large for on-edge applications. For instance, BERT-based systems contain over 110M parameters. Observing the model is overparameterized, we propose lean transformer structure where the dimension of the attention mechanism is automatically reduced using group sparsity. We propose a variant where the learned attention subspace is transferred to an attention bottleneck layer. In a low-resource setting and without pre-training, the resulting compact SLU model achieves accuracies competitive with pre-trained large models.