今天跟大家分享一篇语音相关的论文合集:cs.SD语音10篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily
cs.SD语音
【1】 Personalized Acoustic Echo Cancellation for Full-duplex Communications
标题:用于全双工通信的个性化声学回波消除
链接:https://arxiv.org/abs/2205.15195
作者:Shimin Zhang,Ziteng Wang,Yukai Ju,Yihui Fu,Yueyue Na,Qiang Fu,Lei Xie
备注:submitted to INTERSPEECH 22
摘要:深度神经网络(DNN)在声回波抵消(AEC)方面显示出了良好的效果。但基于DNN的AEC模型允许所有近端扬声器(包括干扰语音)通过。根据最近关于个性化语音增强的研究,我们研究了在全双工通信中,背景噪声和干扰扬声器可能与声回波共存的情况下,个性化声回波消除(PAEC)的可行性。具体而言,我们首先提出了一种称为门控时间卷积神经网络(GTCNN)的新型主干神经网络,它在性能上优于最先进的AEC模型。进一步采用d向量等说话人嵌入作为辅助信息,引导GTCNN聚焦目标说话人。PAEC中的一个特例是通话双方的语音片段都被登记。实验结果表明,无论是近端扬声器还是远端扬声器的辅助信息都可以改善基于DNN的AEC性能。然而,在有限维扬声器嵌入的利用方面仍有很大的改进空间。
摘要:Deep neural networks (DNNs) have shown promising results for acoustic echo cancellation (AEC). But the DNN-based AEC models let through all near-end speakers including the interfering speech. In light of recent studies on personalized speech enhancement, we investigate the feasibility of personalized acoustic echo cancellation (PAEC) in this paper for full-duplex communications, where background noise and interfering speakers may coexist with acoustic echoes. Specifically, we first propose a novel backbone neural network termed as gated temporal convolutional neural network (GTCNN) that outperforms state-of-the-art AEC models in performance. Speaker embeddings like d-vectors are further adopted as auxiliary information to guide the GTCNN to focus on the target speaker. A special case in PAEC is that speech snippets of both parties on the call are enrolled. Experimental results show that auxiliary information from either the near-end speaker or the far-end speaker can improve the DNN-based AEC performance. Nevertheless, there is still much room for improvement in the utilization of the finite-dimensional speaker embeddings.


【2】 Play it by Ear: Learning Skills amidst Occlusion through Audio-Visual  Imitation Learning

标题:随机应变:通过视听模仿学习在遮挡中学习技能

链接:https://arxiv.org/abs/2205.14850

作者:Maximilian Du,Olivia Y. Lee,Suraj Nair,Chelsea Finn
备注:None
摘要:人类能够完成一系列具有挑战性的操作任务,这些任务需要对视觉、触觉和声音等模式进行联合推理。此外,许多这样的任务被部分观察到;例如,从背包中取出笔记本会导致视觉障碍,需要对音频或触觉信息的历史进行推理。虽然在机器人上捕捉强大的触觉感可能成本高昂,但机器人手爪附近或手爪上的麦克风是获取接触事件音频反馈的一种廉价而简单的方法,在没有视觉的情况下,它可能是一个非常有价值的感知数据源。基于声音缓解视觉遮挡的潜力,我们的目标是从视觉和音频输入中学习一组具有挑战性的部分观察操作任务。我们提出的系统通过结合从少量远程操作演示中的离线模仿学习和使用人工干预的在线微调来学习这些任务。在一组模拟任务中,我们发现我们的系统受益于使用音频,通过使用在线干预,我们能够将离线模仿学习的成功率提高约20%。最后,我们发现,我们的系统可以在Franka Emika熊猫机器人上完成一组具有挑战性的、部分观察到的任务,例如从袋子中提取钥匙,成功率为70%,比不使用音频的策略高出50%。
摘要:Humans are capable of completing a range of challenging manipulation tasks that require reasoning jointly over modalities such as vision, touch, and sound. Moreover, many such tasks are partially-observed; for example, taking a notebook out of a backpack will lead to visual occlusion and require reasoning over the history of audio or tactile information. While robust tactile sensing can be costly to capture on robots, microphones near or on a robot's gripper are a cheap and easy way to acquire audio feedback of contact events, which can be a surprisingly valuable data source for perception in the absence of vision. Motivated by the potential for sound to mitigate visual occlusion, we aim to learn a set of challenging partially-observed manipulation tasks from visual and audio inputs. Our proposed system learns these tasks by combining offline imitation learning from a modest number of tele-operated demonstrations and online finetuning using human provided interventions. In a set of simulated tasks, we find that our system benefits from using audio, and that by using online interventions we are able to improve the success rate of offline imitation learning by ~20%. Finally, we find that our system can complete a set of challenging, partially-observed tasks on a Franka Emika Panda robot, like extracting keys from a bag, with a 70% success rate, 50% higher than a policy that does not use audio.


【3】 Modeling Beats and Downbeats with a Time-Frequency Transformer

标题:用时频Transformer模拟心跳和心跳

链接:https://arxiv.org/abs/2205.14701

作者:Yun-Ning Hung,Ju-Chiang Wang,Xuchen Song,Wei-Tsung Lu,Minz Won
备注:This paper is accepted for publication at ICASSP 2022
摘要:Transformer是一种成功的深度神经网络(DNN)体系结构,它不仅在自然语言处理方面,而且在音乐信息检索(MIR)方面都显示了其多功能性。在本文中,我们提出了一种新的基于Transformer的方法来处理拍频和下拍跟踪。这种方法使用SpecTNT(Transformer中的频谱时间转换器),这是转换器的一种变体,它对音乐音频的时频输入的频谱和时间维度进行建模。SpecTNT模型使用一组块,其中每个块由两个级别的Transformer编码器组成。较低级别(或光谱)编码器处理光谱特征,使模型能够关注每个帧的谐波分量。由于向下拍指示条形边界,并且通常伴随着谐波变化,因此此步骤可能有助于向下拍建模。上层(或时间)编码器聚集有用的局部光谱信息,以关注拍/下拍位置。我们还提出了一种将SpecTNT与最先进的模型时间卷积网络(TCN)相结合的体系结构,以进一步提高性能。大量的实验表明,我们的方法在拍频跟踪方面可以显著优于TCN,同时在拍频跟踪方面保持可比的结果。
摘要:Transformer is a successful deep neural network (DNN) architecture that has shown its versatility not only in natural language processing but also in music information retrieval (MIR). In this paper, we present a novel Transformer-based approach to tackle beat and downbeat tracking. This approach employs SpecTNT (Spectral-Temporal Transformer in Transformer), a variant of Transformer that models both spectral and temporal dimensions of a time-frequency input of music audio. A SpecTNT model uses a stack of blocks, where each consists of two levels of Transformer encoders. The lower-level (or spectral) encoder handles the spectral features and enables the model to pay attention to harmonic components of each frame. Since downbeats indicate bar boundaries and are often accompanied by harmonic changes, this step may help downbeat modeling. The upper-level (or temporal) encoder aggregates useful local spectral information to pay attention to beat/downbeat positions. We also propose an architecture that combines SpecTNT with a state-of-the-art model, Temporal Convolutional Networks (TCN), to further improve the performance. Extensive experiments demonstrate that our approach can significantly outperform TCN in downbeat tracking while maintaining comparable result in beat tracking.


【4】 Speaker Identification using Speech Recognition

标题:基于语音识别的说话人识别

链接:https://arxiv.org/abs/2205.14649

作者:Syeda Rabia Arshad,Syed Mujtaba Haider,Abdul Basit Mughal
备注:3 pages
摘要:随着电话通话、视频会议和语音信息的增加,全球范围内的音频数据日益增多。本研究提供了一种基于人声生物特征(如音调、幅度、频率等)的音频文件说话人识别机制。我们提出了一种无监督学习模型,该模型可以在有限的数据集上学习语音表示。本研究使用Librispeech数据集,我们能够实现1.8的单词错误率。
摘要:The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the human voice biometric features like pitch, amplitude, frequency etc. We proposed an unsupervised learning model where the model can learn speech representation with limited dataset. Librispeech dataset was used in this research and we were able to achieve word error rate of 1.8.


【5】 SuperVoice: Text-Independent Speaker Verification Using Ultrasound  Energy in Human Speech

标题:SuperVoice:利用语音中的超声能量进行与文本无关的说话人确认

链接:https://arxiv.org/abs/2205.14496

作者:Hanqing Guo,Qiben Yan,Nikolay Ivanov,Ying Zhu,Li Xiao,Eric J. Hunter
摘要:语音激活系统集成到各种桌面、移动和物联网(IoT)设备中。然而,语音欺骗攻击,如模拟和重播攻击,其中恶意攻击者合成受害者的语音或简单重播受害者的语音,带来了越来越多的安全问题。现有的说话人验证技术通过从语音命令的音频频率范围中提取的光谱特征来区分单个说话人。然而,它们通常具有高错误率和/或长延迟。在本文中,我们通过研究人类语音在超声频段的独特特性,探索了人类语音研究的一个新方向。我们的研究表明,20到48 kHz的高频超声成分(例如语音摩擦)可以显著提高说话人验证的安全性和准确性。我们提出了一个说话人验证系统SUPERVOICE,该系统使用双流DNN结构和特征融合机制来生成不同的说话人模型。为了测试该系统,我们创建了一个语音数据集,其中包含来自127名参与者的12小时音频(8950个语音样本)。此外,我们还创建了第二个欺骗语音数据集来评估其安全性。为了在受控录音和真实应用之间取得平衡,录音由8种不同的录音设备从两个安静的房间收集,其中包括7部智能手机和一个超声波麦克风。我们的评估表明,SUPERVOICE在说话人验证任务中实现了0.58%的等错误率,测试传入话语只需120 ms,优于所有现有的说话人验证系统。此外,在91 ms的处理时间内,SUPERVOICE在检测由5个不同扬声器发起的重播攻击时实现了0%的同等错误率。
摘要:Voice-activated systems are integrated into a variety of desktop, mobile, and Internet-of-Things (IoT) devices. However, voice spoofing attacks, such as impersonation and replay attacks, in which malicious attackers synthesize the voice of a victim or simply replay it, have brought growing security concerns. Existing speaker verification techniques distinguish individual speakers via the spectrographic features extracted from an audible frequency range of voice commands. However, they often have high error rates and/or long delays. In this paper, we explore a new direction of human voice research by scrutinizing the unique characteristics of human speech at the ultrasound frequency band. Our research indicates that the high-frequency ultrasound components (e.g. speech fricatives) from 20 to 48 kHz can significantly enhance the security and accuracy of speaker verification. We propose a speaker verification system, SUPERVOICE that uses a two-stream DNN architecture with a feature fusion mechanism to generate distinctive speaker models. To test the system, we create a speech dataset with 12 hours of audio (8,950 voice samples) from 127 participants. In addition, we create a second spoofed voice dataset to evaluate its security. In order to balance between controlled recordings and real-world applications, the audio recordings are collected from two quiet rooms by 8 different recording devices, including 7 smartphones and an ultrasound microphone. Our evaluation shows that SUPERVOICE achieves 0.58% equal error rate in the speaker verification task, it only takes 120 ms for testing an incoming utterance, outperforming all existing speaker verification systems. Moreover, within 91 ms processing time, SUPERVOICE achieves 0% equal error rate in detecting replay attacks launched by 5 different loudspeakers.


【6】 Feature Pyramid Attention based Residual Neural Network for  Environmental Sound Classification

标题:基于特征金字塔注意力的残差神经网络环境声分类

链接:https://arxiv.org/abs/2205.14411

作者:Liguang Zhou,Yuhongze Zhou,Xiaonan Qi,Junjie Hu,Tin Lun Lam,Yangsheng Xu
摘要:由于声音信号中存在非结构化的时空关系,环境声音分类是一个具有挑战性的问题。近年来,许多研究集中于从卷积神经网络中提取特征,而忽视了语音信号语义相关框架的学习。为此,我们提出了一个端到端的框架,即特征金字塔注意网络(FPAM),重点是抽象ESC的语义相关特征。我们首先通过主干网络提取预处理后的声音波形谱图的特征图。然后,为了构建声谱图的多尺度层次特征,我们通过聚合多尺度层的特征图来构建声谱图的特征金字塔表示,其中语义相关帧的时间帧和空间位置通过FPAM进行定位。具体而言,多个特征首先由尺寸对齐模块处理。然后,附加金字塔空间注意模块(PSA)以使用空间注意模块(SAM)在空间上定位重要频率区域。最后,通过金字塔通道注意(PCA)对处理后的特征图进行细化,以定位重要的时间帧。为了证明所提出的FPAM的有效性,已经提出了在光谱图上显示注意图的方法。可视化结果表明,FPAM可以更关注语义相关区域,而忽略噪声。在两个广泛使用的ESC数据集:ESC-50和ESC-10数据集上验证了所提出方法的有效性。实验结果表明,FPAM的性能与最先进的方法相当。与基线方法相比,FPAM实现了显著的性能提高。
摘要:Environmental sound classification (ESC) is a challenging problem due to the unstructured spatial-temporal relations that exist in the sound signals. Recently, many studies have focused on abstracting features from convolutional neural networks while the learning of semantically relevant frames of sound signals has been overlooked. To this end, we present an end-to-end framework, namely feature pyramid attention network (FPAM), focusing on abstracting the semantically relevant features for ESC. We first extract the feature maps of the preprocessed spectrogram of the sound waveform by a backbone network. Then, to build multi-scale hierarchical features of sound spectrograms, we construct a feature pyramid representation of the sound spectrograms by aggregating the feature maps from multi-scale layers, where the temporal frames and spatial locations of semantically relevant frames are localized by FPAM. Specifically, the multiple features are first processed by a dimension alignment module. Afterward, the pyramid spatial attention module (PSA) is attached to localize the important frequency regions spatially with a spatial attention module (SAM). Last, the processed feature maps are refined by a pyramid channel attention (PCA) to localize the important temporal frames. To justify the effectiveness of the proposed FPAM, visualization of attention maps on the spectrograms has been presented. The visualization results show that FPAM can focus more on the semantic relevant regions while neglecting the noises. The effectiveness of the proposed methods is validated on two widely used ESC datasets: the ESC-50 and ESC-10 datasets. The experimental results show that the FPAM yields comparable performance to state-of-the-art methods. A substantial performance increase has been achieved by FPAM compared with the baseline methods.


【7】 Speech Augmentation Based Unsupervised Learning for Keyword Spotting

标题:基于语音增强的无监督学习关键词检测

链接:https://arxiv.org/abs/2205.14329

作者:Jian Luo,Jianzong Wang,Ning Cheng,Haobin Tang,Jing Xiao
备注:accepted by WCCI 2022
摘要:在本文中,我们研究了一种基于语音增强的无监督学习方法,用于关键词发现(KWS)任务。KWS是一个有用的语音应用程序,但也严重依赖于标记的数据。我们设计了一个CNN注意力架构来执行KWS任务。CNN层侧重于局部声学特征,而注意层则为长期依赖性建模。为了提高KWS模型的鲁棒性,我们还提出了一种无监督学习方法。无监督损失基于原始和增强语音特征之间的相似性以及音频重建信息。在无监督学习中探索了两种语音增强方法:速度和强度。在Google Speech Commands V2数据集上的实验表明,我们的CNN注意模型具有竞争性的结果。此外,基于增广的无监督学习可以进一步提高KWS任务的分类精度。在我们的实验中,使用基于增强的无监督学习,我们的KWS模型比其他无监督方法(如CPC、APC和MPC)取得了更好的性能。
摘要:In this paper, we investigated a speech augmentation based unsupervised learning approach for keyword spotting (KWS) task. KWS is a useful speech application, yet also heavily depends on the labeled data. We designed a CNN-Attention architecture to conduct the KWS task. CNN layers focus on the local acoustic features, and attention layers model the long-time dependency. To improve the robustness of KWS model, we also proposed an unsupervised learning method. The unsupervised loss is based on the similarity between the original and augmented speech features, as well as the audio reconstructing information. Two speech augmentation methods are explored in the unsupervised learning: speed and intensity. The experiments on Google Speech Commands V2 Dataset demonstrated that our CNN-Attention model has competitive results. Moreover, the augmentation based unsupervised learning could further improve the classification accuracy of KWS task. In our experiments, with augmentation based unsupervised learning, our KWS model achieves better performance than other unsupervised methods, such as CPC, APC, and MPC.


【8】 Adaptive Activation Network For Low Resource Multilingual Speech  Recognition

标题:低资源多语种语音识别的自适应激活网络

链接:https://arxiv.org/abs/2205.14326

作者:Jian Luo,Jianzong Wang,Ning Cheng,Zhenpeng Zheng,Jing Xiao
备注:accepted by WCCI 2022
摘要:低资源自动语音识别(ASR)是一项有用但棘手的任务,因为深入学习ASR模型通常需要大量的训练数据。现有的模型大多通过在大型源语言上进行预训练,然后转移到资源较低的目标语言上,建立了瓶颈(BN)层。在这项工作中,我们在ASR模型的上层引入了一个自适应激活网络,并将不同的激活函数应用于不同的语言。我们还提出了两种方法来训练该模型:(1)跨语言学习,取代从源语言到目标语言的激活函数;(2)多语言学习,联合训练每种语言的连接主义时间分类(CTC)损失和不同语言的相关性。我们在IARPA Babel数据集上的实验表明,我们的方法优于从头开始的训练和传统的基于瓶颈特征的方法。此外,将跨语言学习和多语言学习相结合可以进一步提高多语言语音识别的性能。
摘要:Low resource automatic speech recognition (ASR) is a useful but thorny task, since deep learning ASR models usually need huge amounts of training data. The existing models mostly established a bottleneck (BN) layer by pre-training on a large source language, and transferring to the low resource target language. In this work, we introduced an adaptive activation network to the upper layers of ASR model, and applied different activation functions to different languages. We also proposed two approaches to train the model: (1) cross-lingual learning, replacing the activation function from source language to target language, (2) multilingual learning, jointly training the Connectionist Temporal Classification (CTC) loss of each language and the relevance of different languages. Our experiments on IARPA Babel datasets demonstrated that our approaches outperform the from-scratch training and traditional bottleneck feature based methods. In addition, combining the cross-lingual learning and multilingual learning together could further improve the performance of multilingual speech recognition.


【9】 BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for  Binaural Audio Synthesis

标题:BinauralGrad:双耳音频合成的两阶段条件扩散概率模型

链接:https://arxiv.org/abs/2205.14807

作者:Yichong Leng,Zehua Chen,Junliang Guo,Haohe Liu,Jiawei Chen,Xu Tan,Danilo Mandic,Lei He,Xiang-Yang Li,Tao Qin,Sheng Zhao,Tie-Yan Liu
备注:Demo page: this https URL
摘要:双耳音频在构建身临其境的增强和虚拟现实中起着重要作用。由于从真实世界录制双耳音频的成本很高,因此从单声道音频合成双耳音频引起了越来越多的关注。这一合成过程不仅涉及单声道音频的基本物理扭曲,还涉及房间混响和头/耳相关过滤,但在传统数字信号处理中很难精确模拟。在本文中,我们从不同的角度阐述了合成过程,将双耳音频分解为左声道和右声道共享的公共部分,以及每个声道中不同的特定部分。因此,我们提出了BinauralGrad,这是一种新的两阶段框架,配备了扩散模型来分别合成它们。具体地说,在第一阶段中,双耳音频的公共信息由以单声道音频为条件的单声道扩散模型生成,在此基础上,双耳音频由第二阶段中的双声道扩散模型生成。将这种新的两阶段合成观点与高级生成模型(即扩散模型)相结合,提出的双耳雷达能够生成准确且高保真的双耳音频样本。实验结果表明,在基准数据集上,BinauralGrad在对象和主题评估指标方面都大大优于现有基线(Wave L2:0.128 vs.0.157,MOS:3.80 vs.3.61)。生成的音频样本可在线获取。
摘要:Binaural audio plays a significant role in constructing immersive augmented and virtual realities. As it is expensive to record binaural audio from the real world, synthesizing them from mono audio has attracted increasing attention. This synthesis process involves not only the basic physical warping of the mono audio, but also room reverberations and head/ear related filtrations, which, however, are difficult to accurately simulate in traditional digital signal processing. In this paper, we formulate the synthesis process from a different perspective by decomposing the binaural audio into a common part that shared by the left and right channels as well as a specific part that differs in each channel. Accordingly, we propose BinauralGrad, a novel two-stage framework equipped with diffusion models to synthesize them respectively. Specifically, in the first stage, the common information of the binaural audio is generated with a single-channel diffusion model conditioned on the mono audio, based on which the binaural audio is generated by a two-channel diffusion model in the second stage. Combining this novel perspective of two-stage synthesis with advanced generative models (i.e., the diffusion models),the proposed BinauralGrad is able to generate accurate and high-fidelity binaural audio samples. Experiment results show that on a benchmark dataset, BinauralGrad outperforms the existing baselines by a large margin in terms of both object and subject evaluation metrics (Wave L2: 0.128 vs. 0.157, MOS: 3.80 vs. 3.61). The generated audio samples are available online.


【10】 To catch a chorus, verse, intro, or anything else: Analyzing a song with  structural functions

标题:领会合唱、诗句、序曲或其他:分析一首具有结构功能的歌曲

链接:https://arxiv.org/abs/2205.14700

作者:Ju-Chiang Wang,Yun-Ning Hung,Jordan B. L. Smith
备注:This manuscript is accepted by ICASSP 2022
摘要:传统的音乐结构分析算法旨在将歌曲划分为多个片段,并使用抽象标签(例如,“a”、“B”和“C”)对其进行分组。然而,明确识别每个片段的功能(例如,“韵文”或“合唱”)很少尝试,但有许多应用。我们引入了一个多任务深度学习框架,通过估计“经度”、“欢唱度”等作为时间函数,直接从音频中为这些结构语义标签建模。我们提出了一个7级分类法(即intro、verse、chorus、bridge、outro、instrumental和silence),并提供了规则来整合来自四个不同数据集的注释。我们还建议使用一种基于光谱时间变换器的模型,称为SpecTNT,该模型可以通过额外的连接主义时间定位(CTL)损失进行训练。在使用四个公共数据集的交叉数据集评估中,我们证明了SpecTNT模型和CTL丢失的有效性,并总体上获得了很好的结果:所提出的系统在检测合唱和边界方面分别优于最先进的合唱检测和边界检测方法。
摘要:Conventional music structure analysis algorithms aim to divide a song into segments and to group them with abstract labels (e.g., 'A', 'B', and 'C'). However, explicitly identifying the function of each segment (e.g., 'verse' or 'chorus') is rarely attempted, but has many applications. We introduce a multi-task deep learning framework to model these structural semantic labels directly from audio by estimating "verseness," "chorusness," and so forth, as a function of time. We propose a 7-class taxonomy (i.e., intro, verse, chorus, bridge, outro, instrumental, and silence) and provide rules to consolidate annotations from four disparate datasets. We also propose to use a spectral-temporal Transformer-based model, called SpecTNT, which can be trained with an additional connectionist temporal localization (CTL) loss. In cross-dataset evaluations using four public datasets, we demonstrate the effectiveness of the SpecTNT model and CTL loss, and obtain strong results overall: the proposed system outperforms state-of-the-art chorus-detection and boundary-detection methods at detecting choruses and boundaries, respectively.


eess.AS音频处理

【1】 BinauralGrad: A Two-Stage Conditional Diffusion Probabilistic Model for  Binaural Audio Synthesis

标题:BinauralGrad:双耳音频合成的两阶段条件扩散概率模型

链接:https://arxiv.org/abs/2205.14807

作者:Yichong Leng,Zehua Chen,Junliang Guo,Haohe Liu,Jiawei Chen,Xu Tan,Danilo Mandic,Lei He,Xiang-Yang Li,Tao Qin,Sheng Zhao,Tie-Yan Liu
备注:Demo page: this https URL
摘要:双耳音频在构建身临其境的增强和虚拟现实中起着重要作用。由于从真实世界录制双耳音频的成本很高,因此从单声道音频合成双耳音频引起了越来越多的关注。这一合成过程不仅涉及单声道音频的基本物理扭曲,还涉及房间混响和头/耳相关过滤,但在传统数字信号处理中很难精确模拟。在本文中,我们从不同的角度阐述了合成过程,将双耳音频分解为左声道和右声道共享的公共部分,以及每个声道中不同的特定部分。因此,我们提出了BinauralGrad,这是一种新的两阶段框架,配备了扩散模型来分别合成它们。具体地说,在第一阶段中,双耳音频的公共信息由以单声道音频为条件的单声道扩散模型生成,在此基础上,双耳音频由第二阶段中的双声道扩散模型生成。将这种新的两阶段合成观点与高级生成模型(即扩散模型)相结合,提出的双耳雷达能够生成准确且高保真的双耳音频样本。实验结果表明,在基准数据集上,BinauralGrad在对象和主题评估指标方面都大大优于现有基线(Wave L2:0.128 vs.0.157,MOS:3.80 vs.3.61)。生成的音频样本可在线获取。
摘要:Binaural audio plays a significant role in constructing immersive augmented and virtual realities. As it is expensive to record binaural audio from the real world, synthesizing them from mono audio has attracted increasing attention. This synthesis process involves not only the basic physical warping of the mono audio, but also room reverberations and head/ear related filtrations, which, however, are difficult to accurately simulate in traditional digital signal processing. In this paper, we formulate the synthesis process from a different perspective by decomposing the binaural audio into a common part that shared by the left and right channels as well as a specific part that differs in each channel. Accordingly, we propose BinauralGrad, a novel two-stage framework equipped with diffusion models to synthesize them respectively. Specifically, in the first stage, the common information of the binaural audio is generated with a single-channel diffusion model conditioned on the mono audio, based on which the binaural audio is generated by a two-channel diffusion model in the second stage. Combining this novel perspective of two-stage synthesis with advanced generative models (i.e., the diffusion models),the proposed BinauralGrad is able to generate accurate and high-fidelity binaural audio samples. Experiment results show that on a benchmark dataset, BinauralGrad outperforms the existing baselines by a large margin in terms of both object and subject evaluation metrics (Wave L2: 0.128 vs. 0.157, MOS: 3.80 vs. 3.61). The generated audio samples are available online.


【2】 To catch a chorus, verse, intro, or anything else: Analyzing a song with  structural functions

标题:领会合唱、诗句、序曲或其他:分析一首具有结构功能的歌曲

链接:https://arxiv.org/abs/2205.14700

作者:Ju-Chiang Wang,Yun-Ning Hung,Jordan B. L. Smith
备注:This manuscript is accepted by ICASSP 2022
摘要:传统的音乐结构分析算法旨在将歌曲划分为多个片段,并使用抽象标签(例如,“a”、“B”和“C”)对其进行分组。然而,明确识别每个片段的功能(例如,“韵文”或“合唱”)很少尝试,但有许多应用。我们引入了一个多任务深度学习框架,通过估计“经度”、“欢唱度”等作为时间函数,直接从音频中为这些结构语义标签建模。我们提出了一个7级分类法(即intro、verse、chorus、bridge、outro、instrumental和silence),并提供了规则来整合来自四个不同数据集的注释。我们还建议使用一种基于光谱时间变换器的模型,称为SpecTNT,该模型可以通过额外的连接主义时间定位(CTL)损失进行训练。在使用四个公共数据集的交叉数据集评估中,我们证明了SpecTNT模型和CTL丢失的有效性,并总体上获得了很好的结果:所提出的系统在检测合唱和边界方面分别优于最先进的合唱检测和边界检测方法。
摘要:Conventional music structure analysis algorithms aim to divide a song into segments and to group them with abstract labels (e.g., 'A', 'B', and 'C'). However, explicitly identifying the function of each segment (e.g., 'verse' or 'chorus') is rarely attempted, but has many applications. We introduce a multi-task deep learning framework to model these structural semantic labels directly from audio by estimating "verseness," "chorusness," and so forth, as a function of time. We propose a 7-class taxonomy (i.e., intro, verse, chorus, bridge, outro, instrumental, and silence) and provide rules to consolidate annotations from four disparate datasets. We also propose to use a spectral-temporal Transformer-based model, called SpecTNT, which can be trained with an additional connectionist temporal localization (CTL) loss. In cross-dataset evaluations using four public datasets, we demonstrate the effectiveness of the SpecTNT model and CTL loss, and obtain strong results overall: the proposed system outperforms state-of-the-art chorus-detection and boundary-detection methods at detecting choruses and boundaries, respectively.


【3】 Deep Representation Decomposition for Rate-Invariant Speaker  Verification

标题:用于速率不变说话人确认的深度表示分解

链接:https://arxiv.org/abs/2205.14294

作者:Fuchuan Tong,Siqi Zheng,Haodong Zhou,Xingjia Xie,Qingyang Hong,Lin Li
备注:Accepted by Odyssey 2022
摘要:虽然说话人深度嵌入在说话人验证方面取得了很好的性能,但在说话人风格可变性的情况下,这一优势将会减弱。在实际的说话人确认系统中,经常会出现语速不匹配现象,这实际上可能会降低系统性能。为了减少由语速引起的类内差异,我们提出了一种基于对抗学习的深度表示分解方法来学习语速不变的说话人嵌入。具体来说,通过多任务训练,我们采用注意块将原始嵌入分解为身份相关组件和速率相关组件。此外,为了减少两个分解分量之间的潜在关系,我们进一步提出了一个余弦映射块来反向训练参数,以最小化两个分解分量之间的余弦相似性。因此,与身份相关的特征对语速变得鲁棒,然后用于验证。在VoxCeleb1数据和HI-MIA数据上进行了实验,证明了该方法的有效性。
摘要:While promising performance for speaker verification has been achieved by deep speaker embeddings, the advantage would reduce in the case of speaking-style variability. Speaking rate mismatch is often observed in practical speaker verification systems, which may actually degrade the system performance. To reduce intra-class discrepancy caused by speaking rate, we propose a deep representation decomposition approach with adversarial learning to learn speaking rate-invariant speaker embeddings. Specifically, adopting an attention block, we decompose the original embedding into an identity-related component and a rate-related component through multi-task training. Additionally, to reduce the latent relationship between the two decomposed components, we further propose a cosine mapping block to train the parameters adversarially to minimize the cosine similarity between the two decomposed components. As a result, identity-related features become robust to speaking rate and then are used for verification. Experiments are conducted on VoxCeleb1 data and HI-MIA data to demonstrate the effectiveness of our proposed approach.


【4】 Personalized Acoustic Echo Cancellation for Full-duplex Communications

标题:用于全双工通信的个性化声学回波消除

链接:https://arxiv.org/abs/2205.15195

作者:Shimin Zhang,Ziteng Wang,Yukai Ju,Yihui Fu,Yueyue Na,Qiang Fu,Lei Xie
备注:submitted to INTERSPEECH 22
摘要:深度神经网络(DNN)在声回波抵消(AEC)方面显示出了良好的效果。但基于DNN的AEC模型允许所有近端扬声器(包括干扰语音)通过。根据最近关于个性化语音增强的研究,我们研究了在全双工通信中,背景噪声和干扰扬声器可能与声回波共存的情况下,个性化声回波消除(PAEC)的可行性。具体而言,我们首先提出了一种称为门控时间卷积神经网络(GTCNN)的新型主干神经网络,它在性能上优于最先进的AEC模型。进一步采用d向量等说话人嵌入作为辅助信息,引导GTCNN聚焦目标说话人。PAEC中的一个特例是通话双方的语音片段都被登记。实验结果表明,无论是近端扬声器还是远端扬声器的辅助信息都可以改善基于DNN的AEC性能。然而,在有限维扬声器嵌入的利用方面仍有很大的改进空间。
摘要:Deep neural networks (DNNs) have shown promising results for acoustic echo cancellation (AEC). But the DNN-based AEC models let through all near-end speakers including the interfering speech. In light of recent studies on personalized speech enhancement, we investigate the feasibility of personalized acoustic echo cancellation (PAEC) in this paper for full-duplex communications, where background noise and interfering speakers may coexist with acoustic echoes. Specifically, we first propose a novel backbone neural network termed as gated temporal convolutional neural network (GTCNN) that outperforms state-of-the-art AEC models in performance. Speaker embeddings like d-vectors are further adopted as auxiliary information to guide the GTCNN to focus on the target speaker. A special case in PAEC is that speech snippets of both parties on the call are enrolled. Experimental results show that auxiliary information from either the near-end speaker or the far-end speaker can improve the DNN-based AEC performance. Nevertheless, there is still much room for improvement in the utilization of the finite-dimensional speaker embeddings.


【5】 Play it by Ear: Learning Skills amidst Occlusion through Audio-Visual  Imitation Learning

标题:随机应变:通过视听模仿学习在遮挡中学习技能

链接:https://arxiv.org/abs/2205.14850

作者:Maximilian Du,Olivia Y. Lee,Suraj Nair,Chelsea Finn
备注:None
摘要:人类能够完成一系列具有挑战性的操作任务,这些任务需要对视觉、触觉和声音等模式进行联合推理。此外,许多这样的任务被部分观察到;例如,从背包中取出笔记本会导致视觉障碍,需要对音频或触觉信息的历史进行推理。虽然在机器人上捕捉强大的触觉感可能成本高昂,但机器人手爪附近或手爪上的麦克风是获取接触事件音频反馈的一种廉价而简单的方法,在没有视觉的情况下,它可能是一个非常有价值的感知数据源。基于声音缓解视觉遮挡的潜力,我们的目标是从视觉和音频输入中学习一组具有挑战性的部分观察操作任务。我们提出的系统通过结合从少量远程操作演示中的离线模仿学习和使用人工干预的在线微调来学习这些任务。在一组模拟任务中,我们发现我们的系统受益于使用音频,通过使用在线干预,我们能够将离线模仿学习的成功率提高约20%。最后,我们发现,我们的系统可以在Franka Emika熊猫机器人上完成一组具有挑战性的、部分观察到的任务,例如从袋子中提取钥匙,成功率为70%,比不使用音频的策略高出50%。
摘要:Humans are capable of completing a range of challenging manipulation tasks that require reasoning jointly over modalities such as vision, touch, and sound. Moreover, many such tasks are partially-observed; for example, taking a notebook out of a backpack will lead to visual occlusion and require reasoning over the history of audio or tactile information. While robust tactile sensing can be costly to capture on robots, microphones near or on a robot's gripper are a cheap and easy way to acquire audio feedback of contact events, which can be a surprisingly valuable data source for perception in the absence of vision. Motivated by the potential for sound to mitigate visual occlusion, we aim to learn a set of challenging partially-observed manipulation tasks from visual and audio inputs. Our proposed system learns these tasks by combining offline imitation learning from a modest number of tele-operated demonstrations and online finetuning using human provided interventions. In a set of simulated tasks, we find that our system benefits from using audio, and that by using online interventions we are able to improve the success rate of offline imitation learning by ~20%. Finally, we find that our system can complete a set of challenging, partially-observed tasks on a Franka Emika Panda robot, like extracting keys from a bag, with a 70% success rate, 50% higher than a policy that does not use audio.


【6】 Modeling Beats and Downbeats with a Time-Frequency Transformer

标题:用时频Transformer模拟心跳和心跳

链接:https://arxiv.org/abs/2205.14701

作者:Yun-Ning Hung,Ju-Chiang Wang,Xuchen Song,Wei-Tsung Lu,Minz Won
备注:This paper is accepted for publication at ICASSP 2022
摘要:Transformer是一种成功的深度神经网络(DNN)体系结构,它不仅在自然语言处理方面,而且在音乐信息检索(MIR)方面都显示了其多功能性。在本文中,我们提出了一种新的基于Transformer的方法来处理拍频和下拍跟踪。这种方法使用SpecTNT(Transformer中的频谱时间转换器),这是转换器的一种变体,它对音乐音频的时频输入的频谱和时间维度进行建模。SpecTNT模型使用一组块,其中每个块由两个级别的Transformer编码器组成。较低级别(或光谱)编码器处理光谱特征,使模型能够关注每个帧的谐波分量。由于向下拍指示条形边界,并且通常伴随着谐波变化,因此此步骤可能有助于向下拍建模。上层(或时间)编码器聚集有用的局部光谱信息,以关注拍/下拍位置。我们还提出了一种将SpecTNT与最先进的模型时间卷积网络(TCN)相结合的体系结构,以进一步提高性能。大量的实验表明,我们的方法在拍频跟踪方面可以显著优于TCN,同时在拍频跟踪方面保持可比的结果。
摘要:Transformer is a successful deep neural network (DNN) architecture that has shown its versatility not only in natural language processing but also in music information retrieval (MIR). In this paper, we present a novel Transformer-based approach to tackle beat and downbeat tracking. This approach employs SpecTNT (Spectral-Temporal Transformer in Transformer), a variant of Transformer that models both spectral and temporal dimensions of a time-frequency input of music audio. A SpecTNT model uses a stack of blocks, where each consists of two levels of Transformer encoders. The lower-level (or spectral) encoder handles the spectral features and enables the model to pay attention to harmonic components of each frame. Since downbeats indicate bar boundaries and are often accompanied by harmonic changes, this step may help downbeat modeling. The upper-level (or temporal) encoder aggregates useful local spectral information to pay attention to beat/downbeat positions. We also propose an architecture that combines SpecTNT with a state-of-the-art model, Temporal Convolutional Networks (TCN), to further improve the performance. Extensive experiments demonstrate that our approach can significantly outperform TCN in downbeat tracking while maintaining comparable result in beat tracking.


【7】 Speaker Identification using Speech Recognition

标题:基于语音识别的说话人识别

链接:https://arxiv.org/abs/2205.14649

作者:Syeda Rabia Arshad,Syed Mujtaba Haider,Abdul Basit Mughal
备注:3 pages
摘要:随着电话通话、视频会议和语音信息的增加,全球范围内的音频数据日益增多。本研究提供了一种基于人声生物特征(如音调、幅度、频率等)的音频文件说话人识别机制。我们提出了一种无监督学习模型,该模型可以在有限的数据集上学习语音表示。本研究使用Librispeech数据集,我们能够实现1.8的单词错误率。
摘要:The audio data is increasing day by day throughout the globe with the increase of telephonic conversations, video conferences and voice messages. This research provides a mechanism for identifying a speaker in an audio file, based on the human voice biometric features like pitch, amplitude, frequency etc. We proposed an unsupervised learning model where the model can learn speech representation with limited dataset. Librispeech dataset was used in this research and we were able to achieve word error rate of 1.8.


【8】 SuperVoice: Text-Independent Speaker Verification Using Ultrasound  Energy in Human Speech

标题:SuperVoice:利用语音中的超声能量进行与文本无关的说话人确认

链接:https://arxiv.org/abs/2205.14496

作者:Hanqing Guo,Qiben Yan,Nikolay Ivanov,Ying Zhu,Li Xiao,Eric J. Hunter
摘要:语音激活系统集成到各种桌面、移动和物联网(IoT)设备中。然而,语音欺骗攻击,如模拟和重播攻击,其中恶意攻击者合成受害者的语音或简单重播受害者的语音,带来了越来越多的安全问题。现有的说话人验证技术通过从语音命令的音频频率范围中提取的光谱特征来区分单个说话人。然而,它们通常具有高错误率和/或长延迟。在本文中,我们通过研究人类语音在超声频段的独特特性,探索了人类语音研究的一个新方向。我们的研究表明,20到48 kHz的高频超声成分(例如语音摩擦)可以显著提高说话人验证的安全性和准确性。我们提出了一个说话人验证系统SUPERVOICE,该系统使用双流DNN结构和特征融合机制来生成不同的说话人模型。为了测试该系统,我们创建了一个语音数据集,其中包含来自127名参与者的12小时音频(8950个语音样本)。此外,我们还创建了第二个欺骗语音数据集来评估其安全性。为了在受控录音和真实应用之间取得平衡,录音由8种不同的录音设备从两个安静的房间收集,其中包括7部智能手机和一个超声波麦克风。我们的评估表明,SUPERVOICE在说话人验证任务中实现了0.58%的等错误率,测试传入话语只需120 ms,优于所有现有的说话人验证系统。此外,在91 ms的处理时间内,SUPERVOICE在检测由5个不同扬声器发起的重播攻击时实现了0%的同等错误率。
摘要:Voice-activated systems are integrated into a variety of desktop, mobile, and Internet-of-Things (IoT) devices. However, voice spoofing attacks, such as impersonation and replay attacks, in which malicious attackers synthesize the voice of a victim or simply replay it, have brought growing security concerns. Existing speaker verification techniques distinguish individual speakers via the spectrographic features extracted from an audible frequency range of voice commands. However, they often have high error rates and/or long delays. In this paper, we explore a new direction of human voice research by scrutinizing the unique characteristics of human speech at the ultrasound frequency band. Our research indicates that the high-frequency ultrasound components (e.g. speech fricatives) from 20 to 48 kHz can significantly enhance the security and accuracy of speaker verification. We propose a speaker verification system, SUPERVOICE that uses a two-stream DNN architecture with a feature fusion mechanism to generate distinctive speaker models. To test the system, we create a speech dataset with 12 hours of audio (8,950 voice samples) from 127 participants. In addition, we create a second spoofed voice dataset to evaluate its security. In order to balance between controlled recordings and real-world applications, the audio recordings are collected from two quiet rooms by 8 different recording devices, including 7 smartphones and an ultrasound microphone. Our evaluation shows that SUPERVOICE achieves 0.58% equal error rate in the speaker verification task, it only takes 120 ms for testing an incoming utterance, outperforming all existing speaker verification systems. Moreover, within 91 ms processing time, SUPERVOICE achieves 0% equal error rate in detecting replay attacks launched by 5 different loudspeakers.


【9】 Feature Pyramid Attention based Residual Neural Network for  Environmental Sound Classification

标题:基于特征金字塔注意力的残差神经网络环境声分类

链接:https://arxiv.org/abs/2205.14411

作者:Liguang Zhou,Yuhongze Zhou,Xiaonan Qi,Junjie Hu,Tin Lun Lam,Yangsheng Xu
摘要:由于声音信号中存在非结构化的时空关系,环境声音分类是一个具有挑战性的问题。近年来,许多研究集中于从卷积神经网络中提取特征,而忽视了语音信号语义相关框架的学习。为此,我们提出了一个端到端的框架,即特征金字塔注意网络(FPAM),重点是抽象ESC的语义相关特征。我们首先通过主干网络提取预处理后的声音波形谱图的特征图。然后,为了构建声谱图的多尺度层次特征,我们通过聚合多尺度层的特征图来构建声谱图的特征金字塔表示,其中语义相关帧的时间帧和空间位置通过FPAM进行定位。具体而言,多个特征首先由尺寸对齐模块处理。然后,附加金字塔空间注意模块(PSA)以使用空间注意模块(SAM)在空间上定位重要频率区域。最后,通过金字塔通道注意(PCA)对处理后的特征图进行细化,以定位重要的时间帧。为了证明所提出的FPAM的有效性,已经提出了在光谱图上显示注意图的方法。可视化结果表明,FPAM可以更关注语义相关区域,而忽略噪声。在两个广泛使用的ESC数据集:ESC-50和ESC-10数据集上验证了所提出方法的有效性。实验结果表明,FPAM的性能与最先进的方法相当。与基线方法相比,FPAM实现了显著的性能提高。
摘要:Environmental sound classification (ESC) is a challenging problem due to the unstructured spatial-temporal relations that exist in the sound signals. Recently, many studies have focused on abstracting features from convolutional neural networks while the learning of semantically relevant frames of sound signals has been overlooked. To this end, we present an end-to-end framework, namely feature pyramid attention network (FPAM), focusing on abstracting the semantically relevant features for ESC. We first extract the feature maps of the preprocessed spectrogram of the sound waveform by a backbone network. Then, to build multi-scale hierarchical features of sound spectrograms, we construct a feature pyramid representation of the sound spectrograms by aggregating the feature maps from multi-scale layers, where the temporal frames and spatial locations of semantically relevant frames are localized by FPAM. Specifically, the multiple features are first processed by a dimension alignment module. Afterward, the pyramid spatial attention module (PSA) is attached to localize the important frequency regions spatially with a spatial attention module (SAM). Last, the processed feature maps are refined by a pyramid channel attention (PCA) to localize the important temporal frames. To justify the effectiveness of the proposed FPAM, visualization of attention maps on the spectrograms has been presented. The visualization results show that FPAM can focus more on the semantic relevant regions while neglecting the noises. The effectiveness of the proposed methods is validated on two widely used ESC datasets: the ESC-50 and ESC-10 datasets. The experimental results show that the FPAM yields comparable performance to state-of-the-art methods. A substantial performance increase has been achieved by FPAM compared with the baseline methods.


【10】 Speech Augmentation Based Unsupervised Learning for Keyword Spotting

标题:基于语音增强的无监督学习关键词检测

链接:https://arxiv.org/abs/2205.14329

作者:Jian Luo,Jianzong Wang,Ning Cheng,Haobin Tang,Jing Xiao
备注:accepted by WCCI 2022
摘要:在本文中,我们研究了一种基于语音增强的无监督学习方法,用于关键词发现(KWS)任务。KWS是一个有用的语音应用程序,但也严重依赖于标记的数据。我们设计了一个CNN注意力架构来执行KWS任务。CNN层侧重于局部声学特征,而注意层则为长期依赖性建模。为了提高KWS模型的鲁棒性,我们还提出了一种无监督学习方法。无监督损失基于原始和增强语音特征之间的相似性以及音频重建信息。在无监督学习中探索了两种语音增强方法:速度和强度。在Google Speech Commands V2数据集上的实验表明,我们的CNN注意模型具有竞争性的结果。此外,基于增广的无监督学习可以进一步提高KWS任务的分类精度。在我们的实验中,使用基于增强的无监督学习,我们的KWS模型比其他无监督方法(如CPC、APC和MPC)取得了更好的性能。
摘要:In this paper, we investigated a speech augmentation based unsupervised learning approach for keyword spotting (KWS) task. KWS is a useful speech application, yet also heavily depends on the labeled data. We designed a CNN-Attention architecture to conduct the KWS task. CNN layers focus on the local acoustic features, and attention layers model the long-time dependency. To improve the robustness of KWS model, we also proposed an unsupervised learning method. The unsupervised loss is based on the similarity between the original and augmented speech features, as well as the audio reconstructing information. Two speech augmentation methods are explored in the unsupervised learning: speed and intensity. The experiments on Google Speech Commands V2 Dataset demonstrated that our CNN-Attention model has competitive results. Moreover, the augmentation based unsupervised learning could further improve the classification accuracy of KWS task. In our experiments, with augmentation based unsupervised learning, our KWS model achieves better performance than other unsupervised methods, such as CPC, APC, and MPC.


【11】 Adaptive Activation Network For Low Resource Multilingual Speech  Recognition

标题:低资源多语种语音识别的自适应激活网络

链接:https://arxiv.org/abs/2205.14326

作者:Jian Luo,Jianzong Wang,Ning Cheng,Zhenpeng Zheng,Jing Xiao
备注:accepted by WCCI 2022
摘要:低资源自动语音识别(ASR)是一项有用但棘手的任务,因为深入学习ASR模型通常需要大量的训练数据。现有的模型大多通过在大型源语言上进行预训练,然后转移到资源较低的目标语言上,建立了瓶颈(BN)层。在这项工作中,我们在ASR模型的上层引入了一个自适应激活网络,并将不同的激活函数应用于不同的语言。我们还提出了两种方法来训练该模型:(1)跨语言学习,取代从源语言到目标语言的激活函数;(2)多语言学习,联合训练每种语言的连接主义时间分类(CTC)损失和不同语言的相关性。我们在IARPA Babel数据集上的实验表明,我们的方法优于从头开始的训练和传统的基于瓶颈特征的方法。此外,将跨语言学习和多语言学习相结合可以进一步提高多语言语音识别的性能。
摘要:Low resource automatic speech recognition (ASR) is a useful but thorny task, since deep learning ASR models usually need huge amounts of training data. The existing models mostly established a bottleneck (BN) layer by pre-training on a large source language, and transferring to the low resource target language. In this work, we introduced an adaptive activation network to the upper layers of ASR model, and applied different activation functions to different languages. We also proposed two approaches to train the model: (1) cross-lingual learning, replacing the activation function from source language to target language, (2) multilingual learning, jointly training the Connectionist Temporal Classification (CTC) loss of each language and the relevance of different languages. Our experiments on IARPA Babel datasets demonstrated that our approaches outperform the from-scratch training and traditional bottleneck feature based methods. In addition, combining the cross-lingual learning and multilingual learning together could further improve the performance of multilingual speech recognition.


【12】 Is Lip Region-of-Interest Sufficient for Lipreading?

标题:嘴唇感兴趣区域足以进行唇语阅读吗?

链接:https://arxiv.org/abs/2205.14295

作者:Jing-Xuan Zhang,Gen-Shun Wan,Jia Pan
备注:preprint
摘要:唇部感兴趣区域(ROI)通常用于唇读任务中的视觉输入。很少有研究将整个人脸作为视觉输入,因为人脸的除唇部分通常被认为是多余的,与视觉语音识别无关。然而,人脸包含比嘴唇更详细的信息,如说话人的头部姿势、情绪、身份等。我们认为,如果训练一个使用整个人脸的强大特征提取程序,这些信息可能有助于视觉语音识别。在这项工作中,我们建议采用全脸唇读和自我监督学习。实验采用视听多模式自监督学习框架avhubert。我们的实验结果表明,与使用lip作为视觉输入的基线方法相比,采用整个人脸的唇读任务的相对单词错误率(WER)降低了16%。在没有自我监督预训练的情况下,在有限训练数据(30小时)的情况下,人脸输入的模型比嘴唇输入的模型获得了更高的WER,而在使用大量训练数据(433小时)的情况下,WER略低。

摘要:Lip region-of-interest (ROI) is conventionally used for visual input in the lipreading task. Few works have adopted the entire face as visual input because lip-excluded parts of the face are usually considered to be redundant and irrelevant to visual speech recognition. However, faces contain much more detailed information than lips, such as speakers' head pose, emotion, identity etc. We argue that such information might benefit visual speech recognition if a powerful feature extractor employing the entire face is trained. In this work, we propose to adopt the entire face for lipreading with self-supervised learning. AV-HuBERT, an audio-visual multi-modal self-supervised learning framework, was adopted in our experiments. Our experimental results showed that adopting the entire face achieved 16% relative word error rate (WER) reduction on the lipreading task, compared with the baseline method using lip as visual input. Without self-supervised pretraining, the model with face input achieved a higher WER than that using lip input in the case of limited training data (30 hours), while a slightly lower WER when using large amount of training data (433 hours).