今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理4篇。
cs.SD语音
【1】 Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands
标题:Kaggle大赛:粤语视听语音识别车内指令
链接:https://arxiv.org/abs/2207.02663
作者:Wenliang Dai,Samuel Cahyawijaya,Tiezheng Yu,Elham J Barezi,Pascale Fung摘要:随着深度学习和智能车辆的兴起,智能助手已成为车内必不可少的组件,以方便驾驶并提供额外功能。车内智能助理应能够处理一般命令以及与汽车相关的命令,并执行相应的操作,从而简化驾驶并提高安全性。然而,在这一研究领域,大多数数据集是以主要语言编写的,如英语和汉语。低资源语言存在巨大的数据稀缺问题,阻碍了更广泛社区的研究和应用开发。因此,有更多的基准来提高对低资源语言的认识和激励研究是至关重要的。为了缓解这个问题,我们收集了一个新的数据集,即广东话车内视听语音识别(CI-AVSR),用于使用视频和音频数据的广东话车内语音识别。同时,我们提出了车内命令的粤语视听语音识别,作为社区应对车内场景下低资源语音识别的新挑战。摘要:With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. However, in this research field, most datasets are in major languages, such as English and Chinese. There is a huge data scarcity issue for low-resource languages, hindering the development of research and applications for broader communities. Therefore, it is crucial to have more benchmarks to raise awareness and motivate the research in low-resource languages. To mitigate this problem, we collect a new dataset, namely Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR), for in-car speech recognition in the Cantonese language with video and audio data. Together with it, we propose Cantonese Audio-Visual Speech Recognition for In-car Commands as a new challenge for the community to tackle low-resource speech recognition under in-car scenarios.
【2】 Compute Cost Amortized Transformer for Streaming ASR
标题:用于流ASR的计算成本摊销Transformer
链接:https://arxiv.org/abs/2207.02393
作者:Yi Xie,Jonathan Macoskey,Martin Radfar,Feng-Ju Chang,Brian King,Ariya Rastrow,Athanasios Mouchtaris,Grant P. Strimel摘要:我们提出了一种基于变换器的流式端到端自动语音识别(ASR)架构,该架构通过计算成本分摊实现了高效的神经推理。我们的架构在推理时动态创建稀疏计算路径,从而在整个解码过程中选择性地使用计算资源,在对准确性影响最小的情况下显著减少计算量。完全可微体系结构是端到端训练的,伴随着在帧级操作的轻量级仲裁机制,以对每个输入做出动态决策,同时使用可调损失函数来根据预测性能调整整体计算水平。我们报告了使用计算摊销Transformer-传感器(T-T)模型对LibriSpeech数据进行实验的经验结果。我们的最佳模型可以实现60%的计算成本降低,而相对文字错误率(WER)仅增加3%。摘要:We present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Our architecture creates sparse computation pathways dynamically at inference time, resulting in selective use of compute resources throughout decoding, enabling significant reductions in compute with minimal impact on accuracy. The fully differentiable architecture is trained end-to-end with an accompanying lightweight arbitrator mechanism operating at the frame-level to make dynamic decisions on each input while a tunable loss function is used to regularize the overall level of compute against predictive performance. We report empirical results from experiments using the compute amortized Transformer-Transducer (T-T) model conducted on LibriSpeech data. Our best model can achieve a 60% compute cost reduction with only a 3% relative word error rate (WER) increase.
【3】 Ultra-Low-Bitrate Speech Coding with Pretrained Transformers
标题:基于预训练变换的超低码率语音编码
链接:https://arxiv.org/abs/2207.02262
作者:Ali Siahkoohi,Michael Chinen,Tom Denton,W. Bastiaan Kleijn,Jan Skoglund备注:Proceedings of INTERSPEECH 2022摘要:语音编码有助于在低带宽网络上以最小失真传输语音。最近,基于神经网络的语音编解码器与传统方法相比在质量上有了显著改善。虽然新一代编解码器能够合成高保真语音,但其使用循环层或卷积层往往限制了其有效接收场,从而阻碍了其有效压缩语音。我们建议通过使用预训练Transformer来进一步降低神经语音编解码器的比特率,该Transformer能够利用输入信号中由于其感应偏置而产生的长程依赖性。因此,我们将预训练的变换器与卷积编码器串联使用,卷积编码器通过量化器和生成对抗网络解码器进行端到端训练。我们的数值实验表明,使用Transformer语音嵌入来补充神经语音编解码器的卷积编码器,可以得到比特率为600美元的语音编解码器,在相同比特率下训练时,其合成语音质量优于原始神经语音编解码器。主观人类评估表明,生成的编解码器的质量与以三到四倍速率运行的传统编解码器相当或更好。摘要:Speech coding facilitates the transmission of speech over low-bandwidth networks with minimal distortion. Neural-network based speech codecs have recently demonstrated significant improvements in quality over traditional approaches. While this new generation of codecs is capable of synthesizing high-fidelity speech, their use of recurrent or convolutional layers often restricts their effective receptive fields, which prevents them from compressing speech efficiently. We propose to further reduce the bitrate of neural speech codecs through the use of pretrained Transformers, capable of exploiting long-range dependencies in the input signal due to their inductive bias. As such, we use a pretrained Transformer in tandem with a convolutional encoder, which is trained end-to-end with a quantizer and a generative adversarial net decoder. Our numerical experiments show that supplementing the convolutional encoder of a neural speech codec with Transformer speech embeddings yields a speech codec with a bitrate of $600\,\mathrm{bps}$ that outperforms the original neural speech codec in synthesized speech quality when trained at the same bitrate. Subjective human evaluations suggest that the quality of the resulting codec is comparable or better than that of conventional codecs operating at three to four times the rate.
【4】 Improving Streaming End-to-End ASR on Transformer-based Causal Models with Encoder States Revision Strategies
标题:采用编码器状态修正策略改进基于Transformer因果模型的流端到端ASR
链接:https://arxiv.org/abs/2207.02495
作者:Zehan Li,Haoran Miao,Keqi Deng,Gaofeng Cheng,Sanli Tian,Ta Li,Yonghong Yan备注:Accepted by Interspeech 2022摘要:在流式自动语音识别(ASR)中,通常需要在性能和延迟之间进行权衡。传统方法,如前瞻和基于块的方法,通常需要来自未来帧的信息来提高识别精度,即使计算速度足够快,也不可避免地会产生延迟。在没有任何未来帧的情况下进行计算的因果模型可以避免这种延迟,但其性能明显低于传统方法。在本文中,我们提出了相应的修正策略来改进因果模型。首先,我们引入了一种实时编码器状态修正策略来修改以前的状态。编码器前向计算在收到数据后开始,并在几帧后修改之前的编码器状态,无需等待任何正确的上下文。此外,设计了一种连接时序分类尖峰位置对齐解码算法,以减少修正策略带来的时间开销。实验都是在Librispeech数据集上进行的。在基于连接时序分类的wav2vec2.0模型上进行微调,我们的最佳方法可以在测试清洁集/其他集上实现3.7/9.2 WER,这也与基于块的方法和知识提取方法相竞争。摘要:There is often a trade-off between performance and latency in streaming automatic speech recognition (ASR). Traditional methods such as look-ahead and chunk-based methods, usually require information from future frames to advance recognition accuracy, which incurs inevitable latency even if the computation is fast enough. A causal model that computes without any future frames can avoid this latency, but its performance is significantly worse than traditional methods. In this paper, we propose corresponding revision strategies to improve the causal model. Firstly, we introduce a real-time encoder states revision strategy to modify previous states. Encoder forward computation starts once the data is received and revises the previous encoder states after several frames, which is no need to wait for any right context. Furthermore, a CTC spike position alignment decoding algorithm is designed to reduce time costs brought by the revision strategy. Experiments are all conducted on Librispeech datasets. Fine-tuning on the CTC-based wav2vec2.0 model, our best method can achieve 3.7/9.2 WERs on test-clean/other sets, which is also competitive with the chunk-based methods and the knowledge distillation methods.
【1】 Improving Streaming End-to-End ASR on Transformer-based Causal Models with Encoder States Revision Strategies
标题:采用编码器状态修正策略改进基于Transformer因果模型的流端到端ASR
链接:https://arxiv.org/abs/2207.02495
* 与cs.SD语音【4】为同一篇
作者:Zehan Li,Haoran Miao,Keqi Deng,Gaofeng Cheng,Sanli Tian,Ta Li,Yonghong Yan备注:Accepted by Interspeech 2022摘要:在流式自动语音识别(ASR)中,通常需要在性能和延迟之间进行权衡。传统方法,如前瞻和基于块的方法,通常需要来自未来帧的信息来提高识别精度,即使计算速度足够快,也不可避免地会产生延迟。在没有任何未来帧的情况下进行计算的因果模型可以避免这种延迟,但其性能明显低于传统方法。在本文中,我们提出了相应的修正策略来改进因果模型。首先,我们引入了一种实时编码器状态修正策略来修改以前的状态。编码器前向计算在收到数据后开始,并在几帧后修改之前的编码器状态,无需等待任何正确的上下文。此外,设计了一种连接时序分类尖峰位置对齐解码算法,以减少修正策略带来的时间开销。实验都是在Librispeech数据集上进行的。在基于连接时序分类的wav2vec2.0模型上进行微调,我们的最佳方法可以在测试清洁集/其他集上实现3.7/9.2 WER,这也与基于块的方法和知识提取方法相竞争。摘要:There is often a trade-off between performance and latency in streaming automatic speech recognition (ASR). Traditional methods such as look-ahead and chunk-based methods, usually require information from future frames to advance recognition accuracy, which incurs inevitable latency even if the computation is fast enough. A causal model that computes without any future frames can avoid this latency, but its performance is significantly worse than traditional methods. In this paper, we propose corresponding revision strategies to improve the causal model. Firstly, we introduce a real-time encoder states revision strategy to modify previous states. Encoder forward computation starts once the data is received and revises the previous encoder states after several frames, which is no need to wait for any right context. Furthermore, a CTC spike position alignment decoding algorithm is designed to reduce time costs brought by the revision strategy. Experiments are all conducted on Librispeech datasets. Fine-tuning on the CTC-based wav2vec2.0 model, our best method can achieve 3.7/9.2 WERs on test-clean/other sets, which is also competitive with the chunk-based methods and the knowledge distillation methods.
【2】 Kaggle Competition: Cantonese Audio-Visual Speech Recognition for In-car Commands
标题:Kaggle大赛:粤语视听语音识别车内指令
链接:https://arxiv.org/abs/2207.02663
* 与cs.SD语音【1】为同一篇
作者:Wenliang Dai,Samuel Cahyawijaya,Tiezheng Yu,Elham J Barezi,Pascale Fung摘要:随着深度学习和智能车辆的兴起,智能助手已成为车内必不可少的组件,以方便驾驶并提供额外功能。车内智能助理应能够处理一般命令以及与汽车相关的命令,并执行相应的操作,从而简化驾驶并提高安全性。然而,在这一研究领域,大多数数据集是以主要语言编写的,如英语和汉语。低资源语言存在巨大的数据稀缺问题,阻碍了更广泛社区的研究和应用开发。因此,有更多的基准来提高对低资源语言的认识和激励研究是至关重要的。为了缓解这个问题,我们收集了一个新的数据集,即广东话车内视听语音识别(CI-AVSR),用于使用视频和音频数据的广东话车内语音识别。同时,我们提出了车内命令的粤语视听语音识别,作为社区应对车内场景下低资源语音识别的新挑战。摘要:With the rise of deep learning and intelligent vehicles, the smart assistant has become an essential in-car component to facilitate driving and provide extra functionalities. In-car smart assistants should be able to process general as well as car-related commands and perform corresponding actions, which eases driving and improves safety. However, in this research field, most datasets are in major languages, such as English and Chinese. There is a huge data scarcity issue for low-resource languages, hindering the development of research and applications for broader communities. Therefore, it is crucial to have more benchmarks to raise awareness and motivate the research in low-resource languages. To mitigate this problem, we collect a new dataset, namely Cantonese In-car Audio-Visual Speech Recognition (CI-AVSR), for in-car speech recognition in the Cantonese language with video and audio data. Together with it, we propose Cantonese Audio-Visual Speech Recognition for In-car Commands as a new challenge for the community to tackle low-resource speech recognition under in-car scenarios.
【3】 Compute Cost Amortized Transformer for Streaming ASR
标题:用于流ASR的计算成本摊销Transformer
链接:https://arxiv.org/abs/2207.02393
* 与cs.SD语音【2】为同一篇
作者:Yi Xie,Jonathan Macoskey,Martin Radfar,Feng-Ju Chang,Brian King,Ariya Rastrow,Athanasios Mouchtaris,Grant P. Strimel摘要:我们提出了一种基于变换器的流式端到端自动语音识别(ASR)架构,该架构通过计算成本分摊实现了高效的神经推理。我们的架构在推理时动态创建稀疏计算路径,从而在整个解码过程中选择性地使用计算资源,在对准确性影响最小的情况下显著减少计算量。完全可微体系结构是端到端训练的,伴随着在帧级操作的轻量级仲裁机制,以对每个输入做出动态决策,同时使用可调损失函数来根据预测性能调整整体计算水平。我们报告了使用计算摊销Transformer-传感器(T-T)模型对LibriSpeech数据进行实验的经验结果。我们的最佳模型可以实现60%的计算成本降低,而相对文字错误率(WER)仅增加3%。摘要:We present a streaming, Transformer-based end-to-end automatic speech recognition (ASR) architecture which achieves efficient neural inference through compute cost amortization. Our architecture creates sparse computation pathways dynamically at inference time, resulting in selective use of compute resources throughout decoding, enabling significant reductions in compute with minimal impact on accuracy. The fully differentiable architecture is trained end-to-end with an accompanying lightweight arbitrator mechanism operating at the frame-level to make dynamic decisions on each input while a tunable loss function is used to regularize the overall level of compute against predictive performance. We report empirical results from experiments using the compute amortized Transformer-Transducer (T-T) model conducted on LibriSpeech data. Our best model can achieve a 60% compute cost reduction with only a 3% relative word error rate (WER) increase.
【4】 Ultra-Low-Bitrate Speech Coding with Pretrained Transformers
标题:基于预训练变换的超低码率语音编码
链接:https://arxiv.org/abs/2207.02262
* 与cs.SD语音【3】为同一篇
作者:Ali Siahkoohi,Michael Chinen,Tom Denton,W. Bastiaan Kleijn,Jan Skoglund备注:Proceedings of INTERSPEECH 2022摘要:语音编码有助于在低带宽网络上以最小失真传输语音。最近,基于神经网络的语音编解码器与传统方法相比在质量上有了显著改善。虽然新一代编解码器能够合成高保真语音,但其使用循环层或卷积层往往限制了其有效接收场,从而阻碍了其有效压缩语音。我们建议通过使用预训练Transformer来进一步降低神经语音编解码器的比特率,该Transformer能够利用输入信号中由于其感应偏置而产生的长程依赖性。因此,我们将预训练的变换器与卷积编码器串联使用,卷积编码器通过量化器和生成对抗网络解码器进行端到端训练。我们的数值实验表明,使用Transformer语音嵌入来补充神经语音编解码器的卷积编码器,可以得到比特率为600美元的语音编解码器,在相同比特率下训练时,其合成语音质量优于原始神经语音编解码器。主观人类评估表明,生成的编解码器的质量与以三到四倍速率运行的传统编解码器相当或更好。摘要:Speech coding facilitates the transmission of speech over low-bandwidth networks with minimal distortion. Neural-network based speech codecs have recently demonstrated significant improvements in quality over traditional approaches. While this new generation of codecs is capable of synthesizing high-fidelity speech, their use of recurrent or convolutional layers often restricts their effective receptive fields, which prevents them from compressing speech efficiently. We propose to further reduce the bitrate of neural speech codecs through the use of pretrained Transformers, capable of exploiting long-range dependencies in the input signal due to their inductive bias. As such, we use a pretrained Transformer in tandem with a convolutional encoder, which is trained end-to-end with a quantizer and a generative adversarial net decoder. Our numerical experiments show that supplementing the convolutional encoder of a neural speech codec with Transformer speech embeddings yields a speech codec with a bitrate of $600\,\mathrm{bps}$ that outperforms the original neural speech codec in synthesized speech quality when trained at the same bitrate. Subjective human evaluations suggest that the quality of the resulting codec is comparable or better than that of conventional codecs operating at three to four times the rate.
机器翻译,仅供参考