cs.SD语音,共计9篇,eess.AS音频处理,共计10篇


1.cs.SD语音:

【1】 On Building Spoken Language Understanding Systems for Low Resourced Languages

标题:构建面向低资源语言的口语理解系统

链接:https://arxiv.org/abs/2205.12818

作者:Akshat Gupta
机构:J.P.Morgan AI Research, New York, USA
摘要:口语对话系统由于其相对于文本界面的各种优势,正逐渐成为人类经验的组成部分。口语理解系统是口语对话系统的基本组成部分。但为资源不足的语言创建SLU系统仍然是一个挑战。在大量资源不足的语言中,我们无法获得足够的数据来构建自动语音识别(ASR)技术,这是任何SLU系统的基础。此外,基于ASR的SLU系统不会推广到非书面语言。在本文中,我们提供了一系列实验来探索资源极低的环境,在这些环境中,我们使用训练系统对每个意图进行一个数据点的训练,并且数据集中只有一个说话人。我们也在低资源环境下工作,不使用特定语言的ASR系统来转录输入语音,这就增加了构建SLU系统以模拟真正低资源环境的挑战。我们在比利时荷兰语(佛兰芒语)和英语上测试了我们的系统,发现在这种资源不足的情况下,使用语音转录来制作意图分类系统的性能明显优于使用语音特征。具体来说,当使用基于语音转录的系统而不是基于特征的系统时,我们看到,当在49种不同的实验设置上平均时,二进制和四类分类问题的平均改善率分别为12.37%和13.08%。
摘要:Spoken dialog systems are slowly becoming and integral part of the human experience due to their various advantages over textual interfaces. Spoken language understanding (SLU) systems are fundamental building blocks of spoken dialog systems. But creating SLU systems for low resourced languages is still a challenge. In a large number of low resourced language, we don't have access to enough data to build automatic speech recognition (ASR) technologies, which are fundamental to any SLU system. Also, ASR based SLU systems do not generalize to unwritten languages. In this paper, we present a series of experiments to explore extremely low-resourced settings where we perform intent classification with systems trained on as low as one data-point per intent and with only one speaker in the dataset. We also work in a low-resourced setting where we do not use language specific ASR systems to transcribe input speech, which compounds the challenge of building SLU systems to simulate a true low-resourced setting. We test our system on Belgian Dutch (Flemish) and English and find that using phonetic transcriptions to make intent classification systems in such low-resourced setting performs significantly better than using speech features. Specifically, when using a phonetic transcription based system over a feature based system, we see average improvements of 12.37% and 13.08% for binary and four-class classification problems respectively, when averaged over 49 different experimental settings.


【2】 Heterogeneous Reservoir Computing Models for Persian Speech Recognition

标题:用于波斯语音识别的非均质储层计算模型

链接:https://arxiv.org/abs/2205.12594

作者:Zohreh Ansari,Farzin Pourhoseini,Fatemeh Hadaeghi
机构:Biomedical Engineering Department, Meybod University, Meybod, Iran, Institute of Computational Neuroscience, University Medical Center Hamburg-Eppendorf (UKE), Hamburg, Germany
备注:This paper was accepted for oral presentation in IEEE WCCI 2022 + IJCNN 2022, special session on Reservoir Computing: algorithms, implementations and applications
摘要:在过去的十年中,深度学习方法已逐渐融入到传统的自动语音识别(ASR)框架中,以创建声学、发音和语言模型。尽管由于硬件要求(例如计算能力和内存使用)的严格限制,ASR的识别精度有了显著提高,但目前尚不清楚这些方法是否是嵌入式ASR应用程序中计算效率和能效最高的选择。另一方面,储层计算(RC)模型(如回声状态网络(ESN)和液体状态机(LSM))已被证明训练成本低廉,参数少得多,并且与应急硬件技术兼容。然而,它们在语音处理任务中的性能相对不如基于深度学习的模型。为了提高ASR应用中RC的准确性,我们提出了异构单层和多层ESN来创建输入的非线性转换,以捕获不同尺度下的时间上下文。为了测试我们的模型,我们在Farsdat波斯数据集上执行了语音识别任务。据我们所知,标准RC尚未用于执行任何波斯ASR任务,因此我们还训练了传统的单层和深层ESN,以提供比较基线。此外,我们还将RC性能与标准的长-短期记忆(LSTM)模型进行了比较。异构RC模型(1)的性能优于标准RC模型;(2) 在识别精度方面与LSTM相当,(3)大大减少训练时间。
摘要:Over the last decade, deep-learning methods have been gradually incorporated into conventional automatic speech recognition (ASR) frameworks to create acoustic, pronunciation, and language models. Although it led to significant improvements in ASRs' recognition accuracy, due to their hard constraints related to hardware requirements (e.g., computing power and memory usage), it is unclear if such approaches are the most computationally- and energy-efficient options for embedded ASR applications. Reservoir computing (RC) models (e.g., echo state networks (ESNs) and liquid state machines (LSMs)), on the other hand, have been proven inexpensive to train, have vastly fewer parameters, and are compatible with emergent hardware technologies. However, their performance in speech processing tasks is relatively inferior to that of the deep-learning-based models. To enhance the accuracy of the RC in ASR applications, we propose heterogeneous single and multi-layer ESNs to create non-linear transformations of the inputs that capture temporal context at different scales. To test our models, we performed a speech recognition task on the Farsdat Persian dataset. Since, to the best of our knowledge, standard RC has not yet been employed to conduct any Persian ASR tasks, we also trained conventional single-layer and deep ESNs to provide baselines for comparison. Besides, we compared the RC performance with a standard long-short-term memory (LSTM) model. Heterogeneous RC models (1) show improved performance to the standard RC models; (2) perform on par in terms of recognition accuracy with the LSTM, and (3) reduce the training time considerably.


【3】 TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

标题:TranSpeech:双边扰动下的语音到语音翻译

链接:https://arxiv.org/abs/2205.12523

作者:Rongjie Huang,Zhou Zhao,Jinglin Liu,Huadai Liu,Yi Ren,Lichao Zhang,Jinzheng He
机构:Zhejiang University
摘要:直接语音到语音翻译(S2ST)系统利用语音表示学习的最新进展,其中以自监督方式导出的离散表示(单元)序列从模型中预测,并传递给声码器进行语音合成,仍然面临以下挑战:1)声学多模态:由于声学特性(如节奏、音调和能量),来自相同内容语音的离散单元可能具有不确定性,从而导致翻译精度下降;2) 高延迟:当前的S2ST系统使用自回归模型,该模型根据先前生成的序列预测每个单元,未能充分利用并行性。在这项工作中,我们提出了TranSpeech,一种具有双边扰动的语音到语音翻译模型。为了缓解声学多模态问题,我们提出了双边扰动,包括风格规范化和信息增强阶段,仅从语音样本中学习语言信息,并生成更具确定性的表示。随着多模态的减少,我们向前迈进,成为第一个建立非自回归S2ST技术的人,该技术可以重复屏蔽和预测单元选择,并在短短几个周期内产生高精度结果。在三对语言上的实验结果表明,最先进的结果比最好的无文本S2ST基线高出2.5个BLEU点。此外,TranSpeech在推理延迟方面表现出了显著的改善,使速度比自回归技术提高了21.4倍。音频示例位于\url{https://TranSpeech.github.io/}
摘要:Direct speech-to-speech translation (S2ST) systems leverage recent progress in speech representation learning, where a sequence of discrete representations (units) derived in a self-supervised manner, are predicted from the model and passed to a vocoder for speech synthesis, still facing the following challenges: 1) Acoustic multimodality: the discrete units derived from speech with same content could be indeterministic due to the acoustic property (e.g., rhythm, pitch, and energy), which causes deterioration of translation accuracy; 2) high latency: current S2ST systems utilize autoregressive models which predict each unit conditioned on the sequence previously generated, failing to take full advantage of parallelism. In this work, we propose TranSpeech, a speech-to-speech translation model with bilateral perturbation. To alleviate the acoustic multimodal problem, we propose bilateral perturbation, which consists of the style normalization and information enhancement stages, to learn only the linguistic information from speech samples and generate more deterministic representations. With reduced multimodality, we step forward and become the first to establish a non-autoregressive S2ST technique, which repeatedly masks and predicts unit choices and produces high-accuracy results in just a few cycles. Experimental results on three language pairs demonstrate the state-of-the-art results by up to 2.5 BLEU points over the best publicly-available textless S2ST baseline. Moreover, TranSpeech shows a significant improvement in inference latency, enabling speedup up to 21.4x than autoregressive technique. Audio samples are available at \url{https://TranSpeech.github.io/}


【4】 Improving CTC-based ASR Models with Gated Interlayer Collaboration

标题:利用门控层间协作改进基于CTC的ASR模型

链接:https://arxiv.org/abs/2205.12462

作者:Yuting Yang,Yuke Li,Binbin Du
机构:NetEase Yidun AI Lab, Hangzhou, China
备注:Submitted to INTERSPEECH2022
摘要:对于自动语音识别(ASR),基于CTC的方法因其简单的结构和高效的非自回归推理方式而成为主流。然而,这些没有外部语言模型的方法通常缺乏对条件依赖和文本交互进行建模的能力。在这项工作中,我们提出了一种门控层间协作(GIC)机制,该机制将上下文信息引入到模型中,并放松了基于CTC的模型的条件独立性假设。具体地说,我们训练了一个中间CTC损耗的模型,该损耗由模型的层间输出计算,其中中间层的概率分布自然地作为软标签序列。GIC块包括一个嵌入层,用于在每个位置获得软标签的文本嵌入,以及一个门单元,用于融合文本嵌入和声学特征。在AISHELL-1和AIDATATANG基准上的实验表明,该方法优于最近发布的基于CTC的ASR模型。具体而言,我们的方法在AISHELL-1开发/测试集上实现了4.0%/4.4%的CER,在AIDATATANG开发/测试集上实现了3.8%/4.4%的CER,使用CTC贪婪搜索解码,无需外部语言模型。
摘要:For Automatic Speech Recognition (ASR), the CTC-based methods have become a dominant paradigm due to its simple architecture and efficient non-autoregressive inference manner. However, these methods without external language models usually lack the capacity of modeling the conditional dependencies and the textual interaction. In this work, we present a Gated Interlayer Collaboration (GIC) mechanism which introduces the contextual information into the models and relaxes the conditional independence assumption of the CTC-based models. Specifically, we train the model with intermediate CTC losses calculated by the interlayer outputs of the model, in which the probability distributions of the intermediate layers naturally serve as soft label sequences. The GIC block consists of an embedding layer to obtain the textual embedding of the soft label at each position, and a gate unit to fuse the textual embedding and the acoustic features. Experiments on AISHELL-1 and AIDATATANG benchmarks show that the proposed method outperforms the recently published CTC-based ASR models. Specifically, our method achieves CER of 4.0%/4.4% on AISHELL-1 dev/test sets and CER of 3.8%/4.4% on AIDATATANG dev/test sets using CTC greedy search decoding without external language models.


【5】 FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

标题:FLERS:通用语音表征的少试性学习评估

链接:https://arxiv.org/abs/2205.12446

作者:Alexis Conneau,Min Ma,Simran Khanuja,Yu Zhang,Vera Axelrod,Siddharth Dalmia,Jason Riesa,Clara Rivera,Ankur Bapna
机构:Meta AI Research,Google Research,Carnegie Mellon University
摘要:我们介绍了FLEURS,即语音基准的通用表示的少数镜头学习评估。FLEURS是基于机器翻译FLoRes-101基准测试构建的102种语言的n路并行语音数据集,每种语言大约有12小时的语音监控。FLEURS可用于各种语音任务,包括自动语音识别(ASR)、语音语言识别(speech LangID)、翻译和检索。在本文中,我们基于多语言预训练模型(如mSLAM)为任务提供基线。FLEURS的目标是在更多语言中实现语音技术,并促进低资源语音理解的研究。
摘要:We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic Speech Recognition (ASR), Speech Language Identification (Speech LangID), Translation and Retrieval. In this paper, we provide baselines for the tasks based on multilingual pre-trained models like mSLAM. The goal of FLEURS is to enable speech technology in more languages and catalyze research in low-resource speech understanding.


【6】 Adaptive multilingual speech recognition with pretrained models

标题:基于预训练模型的自适应多语言语音识别

链接:https://arxiv.org/abs/2205.12304

作者:Ngoc-Quan Pham,Alex Waibel,Jan Niehues
机构:Interactive Systems Lab, Karlsruhe Institute of Technology, Karlsruhe, Germany, Carnegie Mellon University, Pittsburgh PA, USA
备注:Submitted to INTERSPEECH 2022
摘要:最近的研究表明,基于监督学习的多语语音识别已经取得了很大的成果。随着音频和文本数据预训练方法的发展,迫切需要从无监督的多语言模型中转移知识,以便于识别,尤其是在数据有限的多种语言中。我们的工作研究了两种模式使用两种预训练模型的有效性:wav2vec 2.0用于音频,MBART50用于文本,以及自适应权重技术,以大幅提高包含CommonVoice和Europarl的公共数据集的识别质量。总的来说,我们注意到与纯监督学习相比,有44%的改进,更重要的是,每种技术在不同的语言中都提供了不同的强化。我们还探索了通过稍微增加对架构的深度或相对关注度来获得最佳模型的其他可能性。
摘要:Multilingual speech recognition with supervised learning has achieved great results as reflected in recent research. With the development of pretraining methods on audio and text data, it is imperative to transfer the knowledge from unsupervised multilingual models to facilitate recognition, especially in many languages with limited data. Our work investigated the effectiveness of using two pretrained models for two modalities: wav2vec 2.0 for audio and MBART50 for text, together with the adaptive weight techniques to massively improve the recognition quality on the public datasets containing CommonVoice and Europarl. Overall, we noticed an 44% improvement over purely supervised learning, and more importantly, each technique provides a different reinforcement in different languages. We also explore other possibilities to potentially obtain the best model by slightly adding either depth or relative attention to the architecture.


【7】 Boosting Tail Neural Network for Realtime Custom Keyword Spotting

标题:改进的尾部神经网络在实时关键词识别中的应用

链接:https://arxiv.org/abs/2205.12933

作者:Sihao Xue,Qianyao Shen,Guoqing Li
机构:Metawall
备注:4 pages, 8 figures, 2 tables
摘要:在本文中,我们提出了一种Boosting Tail神经网络(BTNN)来提高实时自定义关键字定位(RCKS)的性能,这对于在有限的计算资源下要求强大的分类能力仍然是一个行业挑战。受脑科学的启发,大脑只有部分激活才能进行神经模拟,许多机器学习算法被开发出来,以使用一批弱分类器来解决棘手的问题,这些问题通常被证明是有效的。我们证明了这种方法对RCKS问题是有帮助的。该方法在唤醒率和虚警率方面都取得了较好的性能。在我们的实验中,与那些只使用一个强分类器的传统算法相比,它得到了18%的相对改进。我们还指出,这种方法在未来的ASR勘探中很有前景。
摘要:In this paper, we propose a Boosting Tail Neural Network (BTNN) for improving the performance of Realtime Custom Keyword Spotting (RCKS) that is still an industrial challenge for demanding powerful classification ability with limited computation resources. Inspired by Brain Science that a brain is only partly activated for a nerve simulation and numerous machine learning algorithms are developed to use a batch of weak classifiers to resolve arduous problems, which are often proved to be effective. We show that this method is helpful to the RCKS problem. The proposed approach achieve better performances in terms of wakeup rate and false alarm. In our experiments compared with those traditional algorithms that use only one strong classifier, it gets 18\% relative improvement. We also point out that this approach may be promising in future ASR exploration.


【8】 Semantic-preserved Communication System for Highly Efficient Speech Transmission

标题:一种高效语音传输的语义保留通信系统

链接:https://arxiv.org/abs/2205.12727

作者:Tianxiao Han,Qianqian Yang,Zhiguo Shi,Shibo He,Zhaoyang Zhang
机构: Zhejiang University
备注:arXiv admin note: substantial text overlap with arXiv:2202.03211
摘要:近年来,人们探索了基于深度学习(DL)的语义通信方法,以实现图像、文本和语音的高效传输。与传统的侧重于抽象符号传输的无线通信方法相比,语义通信方法试图通过仅发送源数据的语义相关信息来实现更好的传输效率。在本文中,我们考虑面向语义的语音传输,该传输仅在语音识别任务的信道上传输语义相关信息,并为语音重建任务提供一组紧凑的附加语义无关信息。我们提出了一种新型的基于DL的端到端收发器,该收发器在发射机处从输入语音频谱中提取语义信息并进行编码,在接收机处从解码后的语义信息中输出相应的转录。对于语音到语音的传输,我们还包括一个CTC对齐模块,该模块提取少量额外的语义无关但与语音相关的信息,以便在接收器处更好地重建原始语音信号。仿真结果表明,我们提出的方法在预测文本的准确性和恢复语音信号的质量方面优于现有方法,并显著提高了传输效率。更具体地说,所提出的方法仅发送现有方法所需的传输符号量的16%,同时实现了语音到文本传输的约10%的WER减少。对于语音到语音传输,其在传输效率方面取得了更显著的改进,仅为现有方法所需传输符号量的0.2%。
摘要:Deep learning (DL) based semantic communication methods have been explored for the efficient transmission of images, text, and speech in recent years. In contrast to traditional wireless communication methods that focus on the transmission of abstract symbols, semantic communication approaches attempt to achieve better transmission efficiency by only sending the semantic-related information of the source data. In this paper, we consider semantic-oriented speech transmission which transmits only the semantic-relevant information over the channel for the speech recognition task, and a compact additional set of semantic-irrelevant information for the speech reconstruction task. We propose a novel end-to-end DL-based transceiver which extracts and encodes the semantic information from the input speech spectrums at the transmitter and outputs the corresponding transcriptions from the decoded semantic information at the receiver. For the speech to speech transmission, we further include a CTC alignment module that extracts a small number of additional semantic-irrelevant but speech-related information for the better reconstruction of the original speech signals at the receiver. The simulation results confirm that our proposed method outperforms current methods in terms of the accuracy of the predicted text for the speech to text transmission and the quality of the recovered speech signals for the speech to speech transmission, and significantly improves transmission efficiency. More specifically, the proposed method only sends 16% of the amount of the transmitted symbols required by the existing methods while achieving about 10% reduction in WER for the speech to text transmission. For the speech to speech transmission, it results in an even more remarkable improvement in terms of transmission efficiency with only 0.2% of the amount of the transmitted symbols required by the existing method.


【9】 An Investigation on Applying Acoustic Feature Conversion to ASR of Adult and Child Speech

标题:声学特征转换在成人和儿童语音ASR中的应用研究

链接:https://arxiv.org/abs/2205.12477

作者:Wei Liu,Jingyu Li,Tan Lee
机构:Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong
备注:5 pages, 4 figures, submitted to InterSpeech2022
摘要:由于训练数据量有限,儿童语音识别的性能通常不如成人语音令人满意。由于域不匹配,在将成人语音训练的自动语音识别(ASR)系统直接应用于儿童语音时,预期性能会显著下降。本研究的重点是成人到儿童的声学特征转换,以缓解这种不匹配。在一个公平的实验环境下,研究和比较了不同的声学特征转换方法,包括基于深度神经网络和基于信号处理的转换方法,其中使用相同数量的标记成人语音转换的声学特征从头训练ASR模型。实验结果表明,并非所有的转换方法都能提高ASR的性能。具体来说,统计匹配作为一种经典的无监督领域自适应方法,并没有显示出有效性。一个基于解纠缠的自动编码器(DAE)转换框架是有用的,F0归一化方法实现了最好的性能。值得注意的是,转换特征的F0分布是反映转换质量的一个重要属性,而利用成人-儿童深度分类模型进行判断是不合适的。
摘要:The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is expected when applying an automatic speech recognition (ASR) system trained on adult speech to child speech directly, as a result of domain mismatch. The present study is focused on adult-to-child acoustic feature conversion to alleviate this mismatch. Different acoustic feature conversion approaches, including deep neural network based and signal processing based, are investigated and compared under a fair experimental setting, in which converted acoustic features from the same amount of labeled adult speech are used to train the ASR models from scratch. Experimental results reveal that not all of the conversion methods lead to ASR performance gain. Specifically, as a classic unsupervised domain adaptation method, the statistic matching does not show an effectiveness. A disentanglement-based auto-encoder (DAE) conversion framework is found to be useful and the approach of F0 normalization achieves the best performance. It is noted that the F0 distribution of converted features is an important attribute to reflect the conversion quality, while utilizing an adult-child deep classification model to make judgment is shown to be inappropriate.


2.eess.AS音频处理:

【1】 Boosting Tail Neural Network for Realtime Custom Keyword Spotting

标题:改进的尾部神经网络在实时关键词识别中的应用

链接:https://arxiv.org/abs/2205.12933

作者:Sihao Xue,Qianyao Shen,Guoqing Li
机构:Metawall
备注:4 pages, 8 figures, 2 tables
摘要:在本文中,我们提出了一种Boosting Tail神经网络(BTNN)来提高实时自定义关键字定位(RCKS)的性能,这对于在有限的计算资源下要求强大的分类能力仍然是一个行业挑战。受脑科学的启发,大脑只有部分激活才能进行神经模拟,许多机器学习算法被开发出来,以使用一批弱分类器来解决棘手的问题,这些问题通常被证明是有效的。我们证明了这种方法对RCKS问题是有帮助的。该方法在唤醒率和虚警率方面都取得了较好的性能。在我们的实验中,与那些只使用一个强分类器的传统算法相比,它得到了18%的相对改进。我们还指出,这种方法在未来的ASR勘探中很有前景。
摘要:In this paper, we propose a Boosting Tail Neural Network (BTNN) for improving the performance of Realtime Custom Keyword Spotting (RCKS) that is still an industrial challenge for demanding powerful classification ability with limited computation resources. Inspired by Brain Science that a brain is only partly activated for a nerve simulation and numerous machine learning algorithms are developed to use a batch of weak classifiers to resolve arduous problems, which are often proved to be effective. We show that this method is helpful to the RCKS problem. The proposed approach achieve better performances in terms of wakeup rate and false alarm. In our experiments compared with those traditional algorithms that use only one strong classifier, it gets 18\% relative improvement. We also point out that this approach may be promising in future ASR exploration.


【2】 Compensation of Driving Signals for Soundfield Synthesis through Irregular Loudspeaker Arrays Based on Convolutional Neural Networks

标题:基于卷积神经网络的不规则扬声器阵列声场合成驱动信号补偿

链接:https://arxiv.org/abs/2205.12872

作者:Luca Comanducci,Fabio Antonacci,Augusto Sarti
摘要:在本文中,我们提出了一种基于深度学习的不规则扬声器阵列声场合成技术,即扬声器之间的间距不是恒定的。输入是通过基于平面波分解的技术获得的驱动信号。虽然所考虑的驱动信号能够用规则阵列正确再现声场,但当使用不规则设置时,它们的性能会下降。通过卷积神经网络(CNN)修改驱动信号,以补偿所需声场再现中的误差。由于补偿后的声场没有地面真实驱动信号,我们通过计算多个控制点处的期望声场与通过网络估计的驱动信号获得的声场之间的损失来训练模型。数值结果表明,与基于平面波分解的方法和压力匹配方法相比,该方法具有更好的性能。
摘要:In this article we propose a technique for soundfield synthesis for irregular loudspeaker arrays, i.e. where the spacing between loudspeakers is not constant, based on deep learning. The input are the driving signals obtained through a plane wave decomposition-based technique. While the considered driving signals are able to correctly reproduce the soundfield with a regular array, they show degraded performances when using irregular setups. Through a Convolutional Neural Network (CNN) we modify the driving signals in order to compensate the errors in the reproduction of the desired soundfield. Since no ground-truth driving signals are available for the compensated ones, we train the model by calculating the loss between the desired soundfield at a number of control points and the one obtained through the driving signals estimated by the network. Numerical results show better performances both with respect to the plane wave decomposition-based technique and the pressure-matching approach.


【3】 Semantic-preserved Communication System for Highly Efficient Speech Transmission

标题:一种高效语音传输的语义保留通信系统

链接:https://arxiv.org/abs/2205.12727

作者:Tianxiao Han,Qianqian Yang,Zhiguo Shi,Shibo He,Zhaoyang Zhang
机构: Zhejiang University
备注:arXiv admin note: substantial text overlap with arXiv:2202.03211
摘要:近年来,人们探索了基于深度学习(DL)的语义通信方法,以实现图像、文本和语音的高效传输。与传统的侧重于抽象符号传输的无线通信方法相比,语义通信方法试图通过仅发送源数据的语义相关信息来实现更好的传输效率。在本文中,我们考虑面向语义的语音传输,该传输仅在语音识别任务的信道上传输语义相关信息,并为语音重建任务提供一组紧凑的附加语义无关信息。我们提出了一种新型的基于DL的端到端收发器,该收发器在发射机处从输入语音频谱中提取语义信息并进行编码,在接收机处从解码后的语义信息中输出相应的转录。对于语音到语音的传输,我们还包括一个CTC对齐模块,该模块提取少量额外的语义无关但与语音相关的信息,以便在接收器处更好地重建原始语音信号。仿真结果表明,我们提出的方法在预测文本的准确性和恢复语音信号的质量方面优于现有方法,并显著提高了传输效率。更具体地说,所提出的方法仅发送现有方法所需的传输符号量的16%,同时实现了语音到文本传输的约10%的WER减少。对于语音到语音传输,其在传输效率方面取得了更显著的改进,仅为现有方法所需传输符号量的0.2%。
摘要:Deep learning (DL) based semantic communication methods have been explored for the efficient transmission of images, text, and speech in recent years. In contrast to traditional wireless communication methods that focus on the transmission of abstract symbols, semantic communication approaches attempt to achieve better transmission efficiency by only sending the semantic-related information of the source data. In this paper, we consider semantic-oriented speech transmission which transmits only the semantic-relevant information over the channel for the speech recognition task, and a compact additional set of semantic-irrelevant information for the speech reconstruction task. We propose a novel end-to-end DL-based transceiver which extracts and encodes the semantic information from the input speech spectrums at the transmitter and outputs the corresponding transcriptions from the decoded semantic information at the receiver. For the speech to speech transmission, we further include a CTC alignment module that extracts a small number of additional semantic-irrelevant but speech-related information for the better reconstruction of the original speech signals at the receiver. The simulation results confirm that our proposed method outperforms current methods in terms of the accuracy of the predicted text for the speech to text transmission and the quality of the recovered speech signals for the speech to speech transmission, and significantly improves transmission efficiency. More specifically, the proposed method only sends 16% of the amount of the transmitted symbols required by the existing methods while achieving about 10% reduction in WER for the speech to text transmission. For the speech to speech transmission, it results in an even more remarkable improvement in terms of transmission efficiency with only 0.2% of the amount of the transmitted symbols required by the existing method.


【4】 An Investigation on Applying Acoustic Feature Conversion to ASR of Adult and Child Speech

标题:声学特征转换在成人和儿童语音ASR中的应用研究

链接:https://arxiv.org/abs/2205.12477

作者:Wei Liu,Jingyu Li,Tan Lee
机构:Department of Electronic Engineering, The Chinese University of Hong Kong, Hong Kong
备注:5 pages, 4 figures, submitted to InterSpeech2022
摘要:由于训练数据量有限,儿童语音识别的性能通常不如成人语音令人满意。由于域不匹配,在将成人语音训练的自动语音识别(ASR)系统直接应用于儿童语音时,预期性能会显著下降。本研究的重点是成人到儿童的声学特征转换,以缓解这种不匹配。在一个公平的实验环境下,研究和比较了不同的声学特征转换方法,包括基于深度神经网络和基于信号处理的转换方法,其中使用相同数量的标记成人语音转换的声学特征从头训练ASR模型。实验结果表明,并非所有的转换方法都能提高ASR的性能。具体来说,统计匹配作为一种经典的无监督领域自适应方法,并没有显示出有效性。一个基于解纠缠的自动编码器(DAE)转换框架是有用的,F0归一化方法实现了最好的性能。值得注意的是,转换特征的F0分布是反映转换质量的一个重要属性,而利用成人-儿童深度分类模型进行判断是不合适的。
摘要:The performance of child speech recognition is generally less satisfactory compared to adult speech due to limited amount of training data. Significant performance degradation is expected when applying an automatic speech recognition (ASR) system trained on adult speech to child speech directly, as a result of domain mismatch. The present study is focused on adult-to-child acoustic feature conversion to alleviate this mismatch. Different acoustic feature conversion approaches, including deep neural network based and signal processing based, are investigated and compared under a fair experimental setting, in which converted acoustic features from the same amount of labeled adult speech are used to train the ASR models from scratch. Experimental results reveal that not all of the conversion methods lead to ASR performance gain. Specifically, as a classic unsupervised domain adaptation method, the statistic matching does not show an effectiveness. A disentanglement-based auto-encoder (DAE) conversion framework is found to be useful and the approach of F0 normalization achieves the best performance. It is noted that the F0 distribution of converted features is an important attribute to reflect the conversion quality, while utilizing an adult-child deep classification model to make judgment is shown to be inappropriate.


【5】 On Building Spoken Language Understanding Systems for Low Resourced Languages

标题:构建面向低资源语言的口语理解系统

链接:https://arxiv.org/abs/2205.12818

作者:Akshat Gupta
机构:J.P.Morgan AI Research, New York, USA
摘要:口语对话系统由于其相对于文本界面的各种优势,正逐渐成为人类经验的组成部分。口语理解系统是口语对话系统的基本组成部分。但为资源不足的语言创建SLU系统仍然是一个挑战。在大量资源不足的语言中,我们无法获得足够的数据来构建自动语音识别(ASR)技术,这是任何SLU系统的基础。此外,基于ASR的SLU系统不会推广到非书面语言。在本文中,我们提供了一系列实验来探索资源极低的环境,在这些环境中,我们使用训练系统对每个意图进行一个数据点的训练,并且数据集中只有一个说话人。我们也在低资源环境下工作,不使用特定语言的ASR系统来转录输入语音,这就增加了构建SLU系统以模拟真正低资源环境的挑战。我们在比利时荷兰语(佛兰芒语)和英语上测试了我们的系统,发现在这种资源不足的情况下,使用语音转录来制作意图分类系统的性能明显优于使用语音特征。具体来说,当使用基于语音转录的系统而不是基于特征的系统时,我们看到,当在49种不同的实验设置上平均时,二进制和四类分类问题的平均改善率分别为12.37%和13.08%。
摘要:Spoken dialog systems are slowly becoming and integral part of the human experience due to their various advantages over textual interfaces. Spoken language understanding (SLU) systems are fundamental building blocks of spoken dialog systems. But creating SLU systems for low resourced languages is still a challenge. In a large number of low resourced language, we don't have access to enough data to build automatic speech recognition (ASR) technologies, which are fundamental to any SLU system. Also, ASR based SLU systems do not generalize to unwritten languages. In this paper, we present a series of experiments to explore extremely low-resourced settings where we perform intent classification with systems trained on as low as one data-point per intent and with only one speaker in the dataset. We also work in a low-resourced setting where we do not use language specific ASR systems to transcribe input speech, which compounds the challenge of building SLU systems to simulate a true low-resourced setting. We test our system on Belgian Dutch (Flemish) and English and find that using phonetic transcriptions to make intent classification systems in such low-resourced setting performs significantly better than using speech features. Specifically, when using a phonetic transcription based system over a feature based system, we see average improvements of 12.37% and 13.08% for binary and four-class classification problems respectively, when averaged over 49 different experimental settings.


【6】 Heterogeneous Reservoir Computing Models for Persian Speech Recognition

标题:用于波斯语音识别的非均质储层计算模型

链接:https://arxiv.org/abs/2205.12594

作者:Zohreh Ansari,Farzin Pourhoseini,Fatemeh Hadaeghi
机构:Biomedical Engineering Department, Meybod University, Meybod, Iran, Institute of Computational Neuroscience, University Medical Center Hamburg-Eppendorf (UKE), Hamburg, Germany
备注:This paper was accepted for oral presentation in IEEE WCCI 2022 + IJCNN 2022, special session on Reservoir Computing: algorithms, implementations and applications
摘要:在过去的十年中,深度学习方法已逐渐融入到传统的自动语音识别(ASR)框架中,以创建声学、发音和语言模型。尽管由于硬件要求(例如计算能力和内存使用)的严格限制,ASR的识别精度有了显著提高,但目前尚不清楚这些方法是否是嵌入式ASR应用程序中计算效率和能效最高的选择。另一方面,储层计算(RC)模型(如回声状态网络(ESN)和液体状态机(LSM))已被证明训练成本低廉,参数少得多,并且与应急硬件技术兼容。然而,它们在语音处理任务中的性能相对不如基于深度学习的模型。为了提高ASR应用中RC的准确性,我们提出了异构单层和多层ESN来创建输入的非线性转换,以捕获不同尺度下的时间上下文。为了测试我们的模型,我们在Farsdat波斯数据集上执行了语音识别任务。据我们所知,标准RC尚未用于执行任何波斯ASR任务,因此我们还训练了传统的单层和深层ESN,以提供比较基线。此外,我们还将RC性能与标准的长-短期记忆(LSTM)模型进行了比较。异构RC模型(1)的性能优于标准RC模型;(2) 在识别精度方面与LSTM相当,(3)大大减少训练时间。
摘要:Over the last decade, deep-learning methods have been gradually incorporated into conventional automatic speech recognition (ASR) frameworks to create acoustic, pronunciation, and language models. Although it led to significant improvements in ASRs' recognition accuracy, due to their hard constraints related to hardware requirements (e.g., computing power and memory usage), it is unclear if such approaches are the most computationally- and energy-efficient options for embedded ASR applications. Reservoir computing (RC) models (e.g., echo state networks (ESNs) and liquid state machines (LSMs)), on the other hand, have been proven inexpensive to train, have vastly fewer parameters, and are compatible with emergent hardware technologies. However, their performance in speech processing tasks is relatively inferior to that of the deep-learning-based models. To enhance the accuracy of the RC in ASR applications, we propose heterogeneous single and multi-layer ESNs to create non-linear transformations of the inputs that capture temporal context at different scales. To test our models, we performed a speech recognition task on the Farsdat Persian dataset. Since, to the best of our knowledge, standard RC has not yet been employed to conduct any Persian ASR tasks, we also trained conventional single-layer and deep ESNs to provide baselines for comparison. Besides, we compared the RC performance with a standard long-short-term memory (LSTM) model. Heterogeneous RC models (1) show improved performance to the standard RC models; (2) perform on par in terms of recognition accuracy with the LSTM, and (3) reduce the training time considerably.


【7】 TranSpeech: Speech-to-Speech Translation With Bilateral Perturbation

标题:TranSpeech:双边扰动下的语音到语音翻译

链接:https://arxiv.org/abs/2205.12523

作者:Rongjie Huang,Zhou Zhao,Jinglin Liu,Huadai Liu,Yi Ren,Lichao Zhang,Jinzheng He
机构:Zhejiang University
摘要:直接语音到语音翻译(S2ST)系统利用语音表示学习的最新进展,其中以自监督方式导出的离散表示(单元)序列从模型中预测,并传递给声码器进行语音合成,仍然面临以下挑战:1)声学多模态:由于声学特性(如节奏、音调和能量),来自相同内容语音的离散单元可能具有不确定性,从而导致翻译精度下降;2) 高延迟:当前的S2ST系统使用自回归模型,该模型根据先前生成的序列预测每个单元,未能充分利用并行性。在这项工作中,我们提出了TranSpeech,一种具有双边扰动的语音到语音翻译模型。为了缓解声学多模态问题,我们提出了双边扰动,包括风格规范化和信息增强阶段,仅从语音样本中学习语言信息,并生成更具确定性的表示。随着多模态的减少,我们向前迈进,成为第一个建立非自回归S2ST技术的人,该技术可以重复屏蔽和预测单元选择,并在短短几个周期内产生高精度结果。在三对语言上的实验结果表明,最先进的结果比最好的无文本S2ST基线高出2.5个BLEU点。此外,TranSpeech在推理延迟方面表现出了显著的改善,使速度比自回归技术提高了21.4倍。音频示例位于\url{https://TranSpeech.github.io/}
摘要:Direct speech-to-speech translation (S2ST) systems leverage recent progress in speech representation learning, where a sequence of discrete representations (units) derived in a self-supervised manner, are predicted from the model and passed to a vocoder for speech synthesis, still facing the following challenges: 1) Acoustic multimodality: the discrete units derived from speech with same content could be indeterministic due to the acoustic property (e.g., rhythm, pitch, and energy), which causes deterioration of translation accuracy; 2) high latency: current S2ST systems utilize autoregressive models which predict each unit conditioned on the sequence previously generated, failing to take full advantage of parallelism. In this work, we propose TranSpeech, a speech-to-speech translation model with bilateral perturbation. To alleviate the acoustic multimodal problem, we propose bilateral perturbation, which consists of the style normalization and information enhancement stages, to learn only the linguistic information from speech samples and generate more deterministic representations. With reduced multimodality, we step forward and become the first to establish a non-autoregressive S2ST technique, which repeatedly masks and predicts unit choices and produces high-accuracy results in just a few cycles. Experimental results on three language pairs demonstrate the state-of-the-art results by up to 2.5 BLEU points over the best publicly-available textless S2ST baseline. Moreover, TranSpeech shows a significant improvement in inference latency, enabling speedup up to 21.4x than autoregressive technique. Audio samples are available at \url{https://TranSpeech.github.io/}


【8】 Improving CTC-based ASR Models with Gated Interlayer Collaboration

标题:利用门控层间协作改进基于CTC的ASR模型

链接:https://arxiv.org/abs/2205.12462

作者:Yuting Yang,Yuke Li,Binbin Du
机构:NetEase Yidun AI Lab, Hangzhou, China
备注:Submitted to INTERSPEECH2022
摘要:对于自动语音识别(ASR),基于CTC的方法因其简单的结构和高效的非自回归推理方式而成为主流。然而,这些没有外部语言模型的方法通常缺乏对条件依赖和文本交互进行建模的能力。在这项工作中,我们提出了一种门控层间协作(GIC)机制,该机制将上下文信息引入到模型中,并放松了基于CTC的模型的条件独立性假设。具体地说,我们训练了一个中间CTC损耗的模型,该损耗由模型的层间输出计算,其中中间层的概率分布自然地作为软标签序列。GIC块包括一个嵌入层,用于在每个位置获得软标签的文本嵌入,以及一个门单元,用于融合文本嵌入和声学特征。在AISHELL-1和AIDATATANG基准上的实验表明,该方法优于最近发布的基于CTC的ASR模型。具体而言,我们的方法在AISHELL-1开发/测试集上实现了4.0%/4.4%的CER,在AIDATATANG开发/测试集上实现了3.8%/4.4%的CER,使用CTC贪婪搜索解码,无需外部语言模型。
摘要:For Automatic Speech Recognition (ASR), the CTC-based methods have become a dominant paradigm due to its simple architecture and efficient non-autoregressive inference manner. However, these methods without external language models usually lack the capacity of modeling the conditional dependencies and the textual interaction. In this work, we present a Gated Interlayer Collaboration (GIC) mechanism which introduces the contextual information into the models and relaxes the conditional independence assumption of the CTC-based models. Specifically, we train the model with intermediate CTC losses calculated by the interlayer outputs of the model, in which the probability distributions of the intermediate layers naturally serve as soft label sequences. The GIC block consists of an embedding layer to obtain the textual embedding of the soft label at each position, and a gate unit to fuse the textual embedding and the acoustic features. Experiments on AISHELL-1 and AIDATATANG benchmarks show that the proposed method outperforms the recently published CTC-based ASR models. Specifically, our method achieves CER of 4.0%/4.4% on AISHELL-1 dev/test sets and CER of 3.8%/4.4% on AIDATATANG dev/test sets using CTC greedy search decoding without external language models.


【9】 FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

标题:FLERS:通用语音表征的少试性学习评估

链接:https://arxiv.org/abs/2205.12446

作者:Alexis Conneau,Min Ma,Simran Khanuja,Yu Zhang,Vera Axelrod,Siddharth Dalmia,Jason Riesa,Clara Rivera,Ankur Bapna
机构:Meta AI Research,Google Research,Carnegie Mellon University
摘要:我们介绍了FLEURS,即语音基准的通用表示的少数镜头学习评估。FLEURS是基于机器翻译FLoRes-101基准测试构建的102种语言的n路并行语音数据集,每种语言大约有12小时的语音监控。FLEURS可用于各种语音任务,包括自动语音识别(ASR)、语音语言识别(speech LangID)、翻译和检索。在本文中,我们基于多语言预训练模型(如mSLAM)为任务提供基线。FLEURS的目标是在更多语言中实现语音技术,并促进低资源语音理解的研究。
摘要:We introduce FLEURS, the Few-shot Learning Evaluation of Universal Representations of Speech benchmark. FLEURS is an n-way parallel speech dataset in 102 languages built on top of the machine translation FLoRes-101 benchmark, with approximately 12 hours of speech supervision per language. FLEURS can be used for a variety of speech tasks, including Automatic Speech Recognition (ASR), Speech Language Identification (Speech LangID), Translation and Retrieval. In this paper, we provide baselines for the tasks based on multilingual pre-trained models like mSLAM. The goal of FLEURS is to enable speech technology in more languages and catalyze research in low-resource speech understanding.


【10】 Adaptive multilingual speech recognition with pretrained models

标题:基于预训练模型的自适应多语言语音识别

链接:https://arxiv.org/abs/2205.12304

作者:Ngoc-Quan Pham,Alex Waibel,Jan Niehues
机构:Interactive Systems Lab, Karlsruhe Institute of Technology, Karlsruhe, Germany, Carnegie Mellon University, Pittsburgh PA, USA
备注:Submitted to INTERSPEECH 2022
摘要:最近的研究表明,基于监督学习的多语语音识别已经取得了很大的成果。随着音频和文本数据预训练方法的发展,迫切需要从无监督的多语言模型中转移知识,以便于识别,尤其是在数据有限的多种语言中。我们的工作研究了两种模式使用两种预训练模型的有效性:wav2vec 2.0用于音频,MBART50用于文本,以及自适应权重技术,以大幅提高包含CommonVoice和Europarl的公共数据集的识别质量。总的来说,我们注意到与纯监督学习相比,有44%的改进,更重要的是,每种技术在不同的语言中都提供了不同的强化。我们还探索了通过稍微增加对架构的深度或相对关注度来获得最佳模型的其他可能性。
摘要:Multilingual speech recognition with supervised learning has achieved great results as reflected in recent research. With the development of pretraining methods on audio and text data, it is imperative to transfer the knowledge from unsupervised multilingual models to facilitate recognition, especially in many languages with limited data. Our work investigated the effectiveness of using two pretrained models for two modalities: wav2vec 2.0 for audio and MBART50 for text, together with the adaptive weight techniques to massively improve the recognition quality on the public datasets containing CommonVoice and Europarl. Overall, we noticed an 44% improvement over purely supervised learning, and more importantly, each technique provides a different reinforcement in different languages. We also explore other possibilities to potentially obtain the best model by slightly adding either depth or relative attention to the architecture.


机器翻译,仅供参考