今天跟大家分享一篇语音相关的论文合集:cs.SD语音17篇,eess.AS音频处理17篇。
【1】 Audio-Visual Segmentation
标题:视听分割
链接:https://arxiv.org/abs/2207.05042
作者:Jinxing Zhou,Jianyuan Wang,Jiayi Zhang,Weixuan Sun,Jing Zhang,Stan Birchfield,Dan Guo,Lingpeng Kong,Meng Wang,Yiran Zhong备注:Accepted to ECCV 2022; Jinxing Zhou and Jianyuan Wang contributed equally; Meng Wang and Yiran Zhong are corresponding authors; Code is available at this https URL摘要:我们提出了一个新的问题,称为视听分割(AVS),其目标是输出在图像帧时产生声音的对象的像素级地图。为了促进这项研究,我们构建了第一个视听分割基准(AVSBench),为音频视频中的发声对象提供像素级注释。该基准研究了两种设置:1)单声源半监督视听分割和2)多声源全监督视听分割。为了解决AVS问题,我们提出了一种新方法,该方法使用时间像素级视听交互模块注入音频语义,作为视觉分割过程的指导。我们还设计了一种正则化损失,以鼓励在训练期间进行视听映射。在AVSBench上进行的定量和定性实验将我们的方法与相关任务中的几种现有方法进行了比较,表明该方法有望在音频和像素视觉语义之间建立桥梁。代码位于https://github.com/OpenNLPLab/AVSBench.摘要:We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark (AVSBench), providing pixel-wise annotations for the sounding objects in audible videos. Two settings are studied with this benchmark: 1) semi-supervised audio-visual segmentation with a single sound source and 2) fully-supervised audio-visual segmentation with multiple sound sources. To deal with the AVS problem, we propose a novel method that uses a temporal pixel-wise audio-visual interaction module to inject audio semantics as guidance for the visual segmentation process. We also design a regularization loss to encourage the audio-visual mapping during training. Quantitative and qualitative experiments on the AVSBench compare our approach to several existing methods from related tasks, demonstrating that the proposed method is promising for building a bridge between the audio and pixel-wise visual semantics. Code is available at https://github.com/OpenNLPLab/AVSBench.
【2】 Speaker Anonymization with Phonetic Intermediate Representations
标题:基于语音中间表示的说话人匿名
链接:https://arxiv.org/abs/2207.04834
作者:Sarina Meyer,Florian Lux,Pavel Denisov,Julia Koch,Pascal Tilli,Ngoc Thang Vu备注:Accepted at Interspeech 2022摘要:在这项工作中,我们提出了一种说话人匿名管道,该管道利用高质量的自动语音识别和合成系统来生成以语音转录和匿名说话人嵌入为条件的语音。使用手机作为中间表示可以确保几乎完全消除输入中的说话人身份信息,同时尽可能保留原始语音内容。我们在LibriSpeech和VCTK语料库上的实验结果揭示了两个关键发现:1)虽然自动语音识别产生不完美的转录,但我们的神经语音合成系统可以处理此类错误,使我们的系统可行且鲁棒;2)将来自不同资源的说话人嵌入相结合是有益的,并且它们的适当归一化至关重要。总的来说,我们的最终最佳系统在对懒惰的知情攻击者的隐私鲁棒性方面显著优于2020年语音隐私挑战赛中提供的基线,同时保持匿名语音的高可懂度和自然度。摘要:In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker embeddings. Using phones as the intermediate representation ensures near complete elimination of speaker identity information from the input while preserving the original phonetic content as much as possible. Our experimental results on LibriSpeech and VCTK corpora reveal two key findings: 1) although automatic speech recognition produces imperfect transcriptions, our neural speech synthesis system can handle such errors, making our system feasible and robust, and 2) combining speaker embeddings from different resources is beneficial and their appropriate normalization is crucial. Overall, our final best system outperforms significantly the baselines provided in the Voice Privacy Challenge 2020 in terms of privacy robustness against a lazy-informed attacker while maintaining high intelligibility and naturalness of the anonymized speech.
【3】 Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition
标题:Wav2vec 2.0和BERT的多层融合多模式情感识别
链接:https://arxiv.org/abs/2207.04697
作者:Zihan Zhao,Yanfeng Wang,Yu Wang备注:Accepted to INTERSPEECH 2022摘要:近年来,多模态情感识别的研究和应用越来越广泛。然而,多模式情感识别面临着缺乏数据的挑战。为了解决这个问题,我们建议使用转移学习,它利用最先进的预训练模型,包括wav2vec 2.0和BERT来完成这项任务。探索了多层次融合方法,包括基于协同注意的早期融合和基于两种嵌入训练模型的晚期融合。此外,提出了一种多粒度框架,该框架不仅提取帧级语音嵌入,还提取段级嵌入,包括电话、音节和单词级语音嵌入,以进一步提高性能。通过将基于共同注意力的早期融合模型和晚期融合模型与多粒度特征提取框架相结合,我们在IEMOCAP数据集上获得了优于最佳基线方法1.3%未加权精度(UA)的结果。摘要:The research and applications of multimodal emotion recognition have become increasingly popular recently. However, multimodal emotion recognition faces the challenge of lack of data. To solve this problem, we propose to use transfer learning which leverages state-of-the-art pre-trained models including wav2vec 2.0 and BERT for this task. Multi-level fusion approaches including coattention-based early fusion and late fusion with the models trained on both embeddings are explored. Also, a multi-granularity framework which extracts not only frame-level speech embeddings but also segment-level embeddings including phone, syllable and word-level speech embeddings is proposed to further boost the performance. By combining our coattention-based early fusion model and late fusion model with the multi-granularity feature extraction framework, we obtain result that outperforms best baseline approaches by 1.3% unweighted accuracy (UA) on the IEMOCAP dataset.
【4】 The HCCL System for the NIST SRE21
标题:NIST SRE21的HCCL系统
链接:https://arxiv.org/abs/2207.04676
作者:Zhuo Li,Runqiu Xiao,Hangting Chen,Zhenduo Zhao,Zihan Zhang,Wenchao Wang备注:accepted by interspeech 2022摘要:本文描述了HCCL团队为NIST 2021说话人识别评估(NIST SRE21)开发的系统。我们首先探索了各种最先进的说话人嵌入提取器,并结合一种新的圆损耗来获得有区别的深度说话人嵌入。考虑到跨通道和跨语言说话人识别是SRE21的关键挑战,我们介绍了几种技术来减少跨域失配。具体来说,将编解码器和语音增强直接应用于原始语音,以消除编解码器和环境噪声不匹配。我们将直接作用于语音以消除相对显式失配的方法统称为数据自适应方法。实验表明,数据自适应方法比我们的基线提高了15%。此外,在说话人嵌入上部署了一些流行的后端域自适应算法,以缓解由隐式失配引起的说话人性能下降。分数校准是我们在SRE21中的一个重大失败。原因是,参数过多的分数校准容易导致过度拟合问题。摘要:This paper describes the systems developed by the HCCL team for the NIST 2021 speaker recognition evaluation (NIST SRE21).We first explore various state-of-the-art speaker embedding extractors combined with a novel circle loss to obtain discriminative deep speaker embeddings. Considering that cross-channel and cross-linguistic speaker recognition are the key challenges of SRE21, we introduce several techniques to reduce the cross-domain mismatch. Specifically, Codec and speech enhancement are directly applied to the raw speech to eliminate the codecs and the environment noise mismatch. We denote the methods that work directly on speech to eliminate the relatively explicit mismatches collectively as data adaptation methods. Experiments show that data adaption methods achieve 15\% improvements over our baseline. Furthermore, some popular back-ends domain adaptation algorithms are deployed on speaker embeddings to alleviate speaker performance degradation caused by the implicit mismatch. Score calibration is a major failure for us in SRE21. The reason is that score calibration with too many parameters easily lead to overfitting problems.
【5】 Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text data
标题:基于非配对文本数据的TTS和ASR半监督联合训练说话人一致性损失及逐步优化
链接:https://arxiv.org/abs/2207.04659
作者:Naoki Makishima,Satoshi Suzuki,Atsushi Ando,Ryo Masumura备注:Accepted to INTERSPEECH 2022摘要:在本文中,我们研究了文本到语音(TTS)和自动语音识别(ASR)的半监督联合训练,其中有少量配对数据和大量未配对文本数据可用。传统的研究形成了一个称为TTS-ASR管道的循环,其中多峰值TTS模型从具有参考语音的文本中合成语音,ASR模型从合成语音中重建文本,然后两个模型都以循环一致性损失进行训练。然而,合成语音不能反映参考语音的说话人特征,训练后合成语音变得过于容易ASR模型识别。这不仅降低了TTS模型的质量,而且限制了ASR模型的改进。为了解决这个问题,我们提出了改进基于循环一致性的训练,减少说话人一致性损失和逐步优化。说话人一致性损失使合成语音的说话人特征更接近参考语音的说话人特征。在分步优化中,我们首先在训练两个模型之前冻结TTS模型的参数,以避免TTS模型对ASR模型的过度适应。实验结果证明了该方法的有效性。摘要:In this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available. Conventional studies form a cycle called the TTS-ASR pipeline, where the multispeaker TTS model synthesizes speech from text with a reference speech and the ASR model reconstructs the text from the synthesized speech, after which both models are trained with a cycle-consistency loss. However, the synthesized speech does not reflect the speaker characteristics of the reference speech and the synthesized speech becomes overly easy for the ASR model to recognize after training. This not only decreases the TTS model quality but also limits the ASR model improvement. To solve this problem, we propose improving the cycleconsistency-based training with a speaker consistency loss and step-wise optimization. The speaker consistency loss brings the speaker characteristics of the synthesized speech closer to that of the reference speech. In the step-wise optimization, we first freeze the parameter of the TTS model before both models are trained to avoid over-adaptation of the TTS model to the ASR model. Experimental results demonstrate the efficacy of the proposed method.
【6】 DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
标题:DelightfulTTS 2:采用对抗性矢量量化自动编码器的端到端语音合成
链接:https://arxiv.org/abs/2207.04646
作者:Yanqing Liu,Ruiqing Xue,Lei He,Xu Tan,Sheng Zhao备注:To appear in Interspeech 2022摘要:当前的文本到语音(TTS)系统通常利用级联声学模型和声码器管道,以mel频谱图作为中间表示,这有两个局限性:1)声学模型和声码器单独训练,而不是联合优化,这会产生级联错误;2) 中间语音表示(例如,mel频谱图)是预先设计的,并且丢失了次优的相位信息。为了解决这些问题,本文开发了DelightfulTTS 2,这是一种新的端到端语音合成系统,具有自动学习的语音表示和联合优化的声学模型和声码器。具体来说,1)我们提出了一种新的基于矢量量化自动编码器和对抗训练的编解码器网络(VQ-GAN),用于提取中间帧级语音表示(而不是传统的mel谱图表示)并重建语音波形;2) 我们联合优化了声学模型(基于TTS)和声码器(VQ-GAN的解码器),并在声学模型上添加了辅助损耗,以预测中间语音表示。实验表明,与DelightfulTTS相比,DelightfulTTS 2实现了+0.14的CMOS增益,更多的方法分析进一步验证了所开发系统的有效性。摘要:Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from two limitations: 1) the acoustic model and vocoder are separately trained instead of jointly optimized, which incurs cascaded errors; 2) the intermediate speech representations (e.g., mel-spectrogram) are pre-designed and lose phase information, which are sub-optimal. To solve these problems, in this paper, we develop DelightfulTTS 2, a new end-to-end speech synthesis system with automatically learned speech representations and jointly optimized acoustic model and vocoder. Specifically, 1) we propose a new codec network based on vector-quantized auto-encoders with adversarial training (VQ-GAN) to extract intermediate frame-level speech representations (instead of traditional representations like mel-spectrograms) and reconstruct speech waveform; 2) we jointly optimize the acoustic model (based on DelightfulTTS) and the vocoder (the decoder of VQ-GAN), with an auxiliary loss on the acoustic model to predict intermediate speech representations. Experiments show that DelightfulTTS 2 achieves a CMOS gain +0.14 over DelightfulTTS, and more method analyses further verify the effectiveness of the developed system.
【7】 Towards Proper Contrastive Self-supervised Learning Strategies For Music Audio Representation
标题:音乐音频表征中恰当的对比性自我监督学习策略
链接:https://arxiv.org/abs/2207.04471
作者:Jeong Choi,Seongwon Jang,Hyunsouk Cho,Sehee Chung备注:2022 IEEE International Conference on Multimedia and Expo (ICME)摘要:自监督学习的共同研究目标是提取任意下游任务将受益的一般表示。在这项工作中,我们研究了从不同的对比自监督学习方案中学习到的音乐音频表示,并在涉及不同音乐感知水平的各种音乐信息检索(MIR)任务中实证评估了嵌入向量。我们对结果进行分析,以探讨针对不同MIR任务的对比学习策略的正确方向。我们表明,这些表征总体上传达了关于音乐听觉特征的全面信息,尽管每种自我监督策略在某些信息方面都有其自身的有效性。摘要:The common research goal of self-supervised learning is to extract a general representation which an arbitrary downstream task would benefit from. In this work, we investigate music audio representation learned from different contrastive self-supervised learning schemes and empirically evaluate the embedded vectors on various music information retrieval (MIR) tasks where different levels of the music perception are concerned. We analyze the results to discuss the proper direction of contrastive learning strategies for different MIR tasks. We show that these representations convey a comprehensive information about the auditory characteristics of music in general, although each of the self-supervision strategies has its own effectiveness in certain aspect of information.
【8】 Joint Analysis of Acoustic Scenes and Sound Events with Weakly labeled Data
标题:基于弱标签数据的声场景与声事件联合分析
链接:https://arxiv.org/abs/2207.04357
作者:Shunsuke Tsubaki,Keisuke Imoto,Nobutaka Ono备注:Accepted to IWAENC2022摘要:考虑到声场景和声音事件彼此密切相关,在以前的一些论文中,提出了利用基于多任务学习(MTL)的神经网络对声场景和声音事件进行联合分析。在传统方法中,将强监督方案应用于MTL模型中的声音事件检测,这需要在模型训练中对声音事件进行强标记;然而,注释强事件标签相当耗时。因此,在本文中,我们提出了一种基于MTL框架的声音场景和声音事件联合分析方法,该框架具有声音事件的弱标签。特别是,在该方法中,我们引入了用于声音事件检测弱监督训练的多实例学习方案,并评估了四个池函数,即最大池、平均池、指数软最大池和注意力池。使用部分TUT声学场景2016/2017和TUT声音事件2016/2017数据集获得的实验结果表明,在场景分类和事件检测性能方面,基于MTL的弱标签方法优于传统的基于单任务的弱标签场景分类和事件检测模型。摘要:Considering that acoustic scenes and sound events are closely related to each other, in some previous papers, a joint analysis of acoustic scenes and sound events utilizing multitask learning (MTL)-based neural networks was proposed. In conventional methods, a strongly supervised scheme is applied to sound event detection in MTL models, which requires strong labels of sound events in model training; however, annotating strong event labels is quite time-consuming. In this paper, we thus propose a method for the joint analysis of acoustic scenes and sound events based on the MTL framework with weak labels of sound events. In particular, in the proposed method, we introduce the multiple-instance learning scheme for weakly supervised training of sound event detection and evaluate four pooling functions, namely, max pooling, average pooling, exponential softmax pooling, and attention pooling. Experimental results obtained using parts of the TUT Acoustic Scenes 2016/2017 and TUT Sound Events 2016/2017 datasets show that the proposed MTL-based method with weak labels outperforms the conventional single-task-based scene classification and event detection models with weak labels in terms of both the scene classification and event detection performances.
【9】 A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
标题:基于自监督语音表示的语音转换方法的比较研究
链接:https://arxiv.org/abs/2207.04356
作者:Wen-Chin Huang,Shu-Wen Yang,Tomoki Hayashi,Tomoki Toda备注:Accepted to IEEE Journal of Selected Topics in Signal Processing. arXiv admin note: substantial text overlap with arXiv:2110.06280摘要:我们对基于自监督语音表示(S3R)的语音转换(VC)进行了大规模比较研究。在识别合成VC中,S3R具有替代昂贵的监督表示(如语音后验图(PPG))的潜力,因此具有吸引力,这是最先进的VC系统通常采用的。使用我们之前开发的开源VC软件S3PRL-VC,我们使用语音转换挑战2020(VCC2020)数据集,在三种VC设置下提供了一系列深入的客观和主观分析:语内/跨语言任意对一(A2O)和任意对任意(A2A)VC。我们从多个方面研究了基于S3R的VC,包括模型类型、多语言性和监督。我们还研究了k-均值聚类后离散化过程的效果,并展示了它在A2A设置中的改进。最后,通过与最先进的VC系统的比较,证明了基于S3R的VC的竞争力,并揭示了可能的改进方向。摘要:We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attractive owing to their potential to replace expensive supervised representations such as phonetic posteriorgrams (PPGs), which are commonly adopted by state-of-the-art VC systems. Using S3PRL-VC, an open-source VC software we previously developed, we provide a series of in-depth objective and subjective analyses under three VC settings: intra-/cross-lingual any-to-one (A2O) and any-to-any (A2A) VC, using the voice conversion challenge 2020 (VCC2020) dataset. We investigated S3R-based VC in various aspects, including model type, multilinguality, and supervision. We also studied the effect of a post-discretization process with k-means clustering and showed how it improves in the A2A setting. Finally, the comparison with state-of-the-art VC systems demonstrates the competitiveness of S3R-based VC and also sheds light on the possible improving directions.
【10】 Dual-path Attention is All You Need for Audio-Visual Speech Extraction
标题:双路径注意力是音视频语音提取所需要的全部
链接:https://arxiv.org/abs/2207.04213
作者:Zhongweiyang Xu,Xulin Fan,Mark Hasegawa-Johnson摘要:视听目标语音提取旨在通过观察嘴唇运动从噪声混合物中提取特定说话人的语音,结合时域语音分离模型和视觉特征提取器(CNN)取得了重大进展。融合音频和视频信息的一个问题是它们具有不同的时间分辨率。目前大多数研究都是沿着时间维度对视觉特征进行采样,以便音频和视频特征能够及时对齐。然而,我们认为唇部运动应该主要包含长期或电话层面的信息。基于这一假设,我们提出了一种新的融合视听特征的方法。我们观察到,对于DPRN,块间维度的时间分辨率可能非常接近视频帧的时间分辨率。与{sepformer}类似,DPRN中的LSTM被块内和块间自注意力所取代,但在该算法中,块间注意力将视觉特征合并为一个额外的特征流。这可以防止视觉线索的上采样,从而实现更高效的视听融合。结果表明,与其他基于时域的视听融合模型相比,我们取得了更好的结果。摘要:Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature extractors (CNN). One problem of fusing audio and video information is that they have different time resolutions. Most current research upsamples the visual features along the time dimension so that audio and video features are able to align in time. However, we believe that lip movement should mostly contain long-term, or phone-level information. Based on this assumption, we propose a new way to fuse audio-visual features. We observe that for DPRNN \cite{dprnn}, the interchunk dimension's time resolution could be very close to the time resolution of video frames. Like \cite{sepformer}, the LSTM in DPRNN is replaced by intra-chunk and inter-chunk self-attention, but in the proposed algorithm, inter-chunk attention incorporates the visual features as an additional feature stream. This prevents the upsampling of visual cues, resulting in more efficient audio-visual fusion. The result shows we achieve superior results compared with other time-domain based audio-visual fusion models.
【11】 Learning to Separate Voices by Spatial Regions
标题:学习按空间区域分离声音
链接:https://arxiv.org/abs/2207.04203
作者:Zhongweiyang Xu,Romit Roy Choudhury备注:Accepted to ICML 2020. For associated audio samples, see this https URL摘要:我们考虑双耳应用中的音频-语音分离问题,例如耳机和助听器。虽然今天的神经网络表现非常好(用两个麦克风分离$4+$个源),但它们假设已知或固定的最大源数K。此外,今天的模型是以有监督的方式训练的,使用从通用源、环境和人头形状合成的训练数据。本文旨在放松这两个约束,但代价是对问题定义进行轻微修改。我们观察到,当接收到的混合信号包含太多信源时,按区域将其分离仍然有帮助,即,将信号混合信号与用户头部周围的每个锥形扇区隔离。这需要学习每个区域的细粒度空间特性,包括人的头部施加的信号失真。我们提出了一种两阶段自监督框架,其中对来自耳机的无意听到的声音进行预处理,以提取相对干净的个性化信号,然后用于训练区域分离模型。结果显示了良好的性能,强调了个性化相对于一般监督方法的重要性。(音频样本可在我们的项目网站上获得:https://uiuc-earable-computing.github.io/binaural/.我们相信这一结果可以帮助在选择性听力、噪声消除和音频增强现实中的实际应用。摘要:We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed maximum number of sources, K. Moreover, today's models are trained in a supervised manner, using training data synthesized from generic sources, environments, and human head shapes. This paper intends to relax both these constraints at the expense of a slight alteration in the problem definition. We observe that, when a received mixture contains too many sources, it is still helpful to separate them by region, i.e., isolating signal mixtures from each conical sector around the user's head. This requires learning the fine-grained spatial properties of each region, including the signal distortions imposed by a person's head. We propose a two-stage self-supervised framework in which overheard voices from earphones are pre-processed to extract relatively clean personalized signals, which are then used to train a region-wise separation model. Results show promising performance, underscoring the importance of personalization over a generic supervised approach. (audio samples available at our project website: https://uiuc-earable-computing.github.io/binaural/. We believe this result could help real-world applications in selective hearing, noise cancellation, and audio augmented reality.
【12】 Automated Audio Captioning and Language-Based Audio Retrieval
标题:自动音频字幕和基于语言的音频检索
链接:https://arxiv.org/abs/2207.04156
作者:Clive Gomes,Hyejin Park,Patrick Kollman,Yi Song备注:DCASE 2022 Competition (Task 6)摘要:该项目参与了DCASE 2022竞赛(任务6),该竞赛有两个子任务:(1)自动音频字幕和(2)基于语言的音频检索。第一个子任务涉及生成音频样本的文本描述,而第二个子任务的目标是在固定数据集中找到与给定描述匹配的音频样本。对于这两个子任务,都使用了Clotho数据集。这些模型在BLEU1、BLEU2、BLEU3、ROUGEL、METEOR、苹果酒、SPICE和SPIDEr的音频字幕评分以及R1、R5、R10和mARP10的音频检索评分上进行了评估。我们进行了一些实验,修改了这些任务的基线模型。我们的自动音频字幕的最终架构接近基线性能,而基于语言的音频检索模型已经超过了它的对应模型。摘要:This project involved participation in the DCASE 2022 Competition (Task 6) which had two subtasks: (1) Automated Audio Captioning and (2) Language-Based Audio Retrieval. The first subtask involved the generation of a textual description for audio samples, while the goal of the second was to find audio samples within a fixed dataset that match a given description. For both subtasks, the Clotho dataset was used. The models were evaluated on BLEU1, BLEU2, BLEU3, ROUGEL, METEOR, CIDEr, SPICE, and SPIDEr scores for audio captioning and R1, R5, R10 and mARP10 scores for audio retrieval. We have conducted a handful of experiments that modify the baseline models for these tasks. Our final architecture for Automated Audio Captioning is close to the baseline performance, while our model for Language-Based Audio Retrieval has surpassed its counterpart.
【13】 pMCT: Patched Multi-Condition Training for Robust Speech Recognition
标题:PMCT:用于稳健语音识别的补丁多条件训练
链接:https://arxiv.org/abs/2207.04949
作者:Pablo Peso Parada,Agnieszka Dobrowolska,Karthikeyan Saravanan,Mete Ozay备注:Accepted at Interspeech 2022摘要:我们提出了一种新的用于鲁棒自动语音识别(ASR)的补丁多条件训练(pMCT)方法。pMCT通过混合{it patches}从干净和扭曲的语音中提取的相同语音,采用多条件音频修改和修补(MAMP)。使用贴片修正信号的训练提高了模型在噪声混响场景中的鲁棒性。在LibriSpeech数据集上对我们提出的pMCT进行了评估,结果表明,与使用vanilla多条件训练(MCT)相比,pMCT有了改进。为了分析鲁棒ASR,我们在语音数据集上采用了pMCT,语音数据集是一个使用LibriSpeech中的语音创建的噪声混响数据集。在分析中,与MCT相比,pMCT实现了23.1%的相对功率降低。摘要:We propose a novel Patched Multi-Condition Training (pMCT) method for robust Automatic Speech Recognition (ASR). pMCT employs Multi-condition Audio Modification and Patching (MAMP) via mixing {\it patches} of the same utterance extracted from clean and distorted speech. Training using patch-modified signals improves robustness of models in noisy reverberant scenarios. Our proposed pMCT is evaluated on the LibriSpeech dataset showing improvement over using vanilla Multi-Condition Training (MCT). For analyses on robust ASR, we employed pMCT on the VOiCES dataset which is a noisy reverberant dataset created using utterances from LibriSpeech. In the analyses, pMCT achieves 23.1% relative WER reduction compared to the MCT.
【14】 Multi-Frequency Information Enhanced Channel Attention Module for Speaker Representation Learning
标题:用于说话人表征学习的多频信息增强通道注意模块
链接:https://arxiv.org/abs/2207.04540
作者:Mufan Sang,John H. L. Hansen备注:Accepted to Interspeech 2022摘要:最近,注意力机制已成功应用于基于神经网络的说话人验证系统。将压缩和激励块合并到卷积神经网络中取得了显著的性能。然而,它使用全局平均池(GAP)来简单地沿时间和频率维度平均特征,这无法在特征映射中保留足够的说话人信息。在本研究中,我们证明了GAP是时频域离散余弦变换(DCT)的特例,在数学上仅使用频率分解中的最低频率分量。为了增强说话人信息提取能力,我们提出利用多频率信息,设计两个新颖有效的注意模块,即单频单通道(SFSC)注意模块和多频单通道(MFSC)注意模块。基于离散余弦变换,该注意力模块可以有效地从多个频率分量中捕获更多的说话人信息。我们在VoxCeleb数据集上进行了全面的实验,并在第一个48-UTD法医语料库上进行了探索性评估。实验结果表明,我们提出的SFSC和MFSC注意模块可以有效地生成更具辨别力的说话人表示,在不添加额外网络参数的情况下,其性能优于ResNet34 SE和ECAPA-TDNN系统,EER分别降低了20.9%和20.2%。摘要:Recently, attention mechanisms have been applied successfully in neural network-based speaker verification systems. Incorporating the Squeeze-and-Excitation block into convolutional neural networks has achieved remarkable performance. However, it uses global average pooling (GAP) to simply average the features along time and frequency dimensions, which is incapable of preserving sufficient speaker information in the feature maps. In this study, we show that GAP is a special case of a discrete cosine transform (DCT) on time-frequency domain mathematically using only the lowest frequency component in frequency decomposition. To strengthen the speaker information extraction ability, we propose to utilize multi-frequency information and design two novel and effective attention modules, called Single-Frequency Single-Channel (SFSC) attention module and Multi-Frequency Single-Channel (MFSC) attention module. The proposed attention modules can effectively capture more speaker information from multiple frequency components on the basis of DCT. We conduct comprehensive experiments on the VoxCeleb datasets and a probe evaluation on the 1st 48-UTD forensic corpus. Experimental results demonstrate that our proposed SFSC and MFSC attention modules can efficiently generate more discriminative speaker representations and outperform ResNet34-SE and ECAPA-TDNN systems with relative 20.9% and 20.2% reduction in EER, without adding extra network parameters.
【15】 Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder
标题:基于注意力的语音识别中的中间层输出正则化
链接:https://arxiv.org/abs/2207.04177
作者:Jicheng Zhang,Yizhou Peng,Haihua Xu,Yi He,Eng Siong Chng,Hao Huang备注:5 pages. Submitted to INTERSPEECH 2022摘要:通过编码器端的多任务训练实现中间层输出(ILO)正则化已被证明是在广泛的端到端ASR框架上产生改进结果的有效方法。在本文中,我们提出了一种新的方法来进行不同的国际劳工组织正规化训练。我们没有使用需要更多训练开销的传统多任务方法,而是直接将中间层输出作为解码器的输入,也就是说,我们的解码器不仅接受最终编码器层的输出作为输入,还将编码器ILO的输出作为训练期间的输入。使用该方法,由于编码器和解码器同时“正则化”,与基于ILO的连接时序分类方法以及未使用该方法的原始基于注意力的建模方法相比,网络训练更充分,一致性更好。摘要:Intermediate layer output (ILO) regularization by means of multitask training on encoder side has been shown to be an effective approach to yielding improved results on a wide range of end-to-end ASR frameworks. In this paper, we propose a novel method to do ILO regularized training differently. Instead of using conventional multitask methods that entail more training overhead, we directly make the intermediate layer output as input to the decoder, that is, our decoder not only accepts the output of the final encoder layer as input, it also takes the output of the encoder ILO as input during training. With the proposed method, as both encoder and decoder are simultaneously "regularized", the network is more sufficiently trained, consistently leading to improved results, over the ILO-based CTC method, as well as over the original attention-based modeling method without the proposed method employed.
【16】 Internal Language Model Estimation based Language Model Fusion for Cross-Domain Code-Switching Speech Recognition
标题:基于内部语言模型估计的跨域码变语音识别语言模型融合
链接:https://arxiv.org/abs/2207.04176
作者:Yizhou Peng,Yufei Liu,Jicheng Zhang,Haihua Xu,Yi He,Hao Huang,Eng Siong Chng备注:5 pages. Submitted to INTERSPEECH 2022摘要:在域内和跨域语音识别任务中,基于内部语言模型估计(ILME)的语言模型(LM)融合比传统的浅层融合显著改善了识别结果。在本文中,我们尝试将ILME方法应用于跨域码切换语音识别(CSSR)工作。具体来说,我们的好奇心来自几个方面。首先,我们想知道基于ILME的LM融合对于域内和跨域CSSR任务的有效性。我们在合并或不合并两个代码交换域的情况下验证了这一点。更重要的是,我们通过合并两个单语数据集来训练端到端(E2E)语音识别模型,并观察所提出的基于ILME的LM融合对CSSR的有效性。在来自东南亚和另一个中国大陆CS数据集的SEAM上的实验结果证明了所提出的基于ILME的LM融合方法的有效性。摘要:Internal Language Model Estimation (ILME) based language model (LM) fusion has been shown significantly improved recognition results over conventional shallow fusion in both intra-domain and cross-domain speech recognition tasks. In this paper, we attempt to apply our ILME method to cross-domain code-switching speech recognition (CSSR) work. Specifically, our curiosity comes from several aspects. First, we are curious about how effective the ILME-based LM fusion is for both intra-domain and cross-domain CSSR tasks. We verify this with or without merging two code-switching domains. More importantly, we train an end-to-end (E2E) speech recognition model by means of merging two monolingual data sets and observe the efficacy of the proposed ILME-based LM fusion for CSSR. Experimental results on SEAME that is from Southeast Asian and another Chinese Mainland CS data set demonstrate the effectiveness of the proposed ILME-based LM fusion method.
【17】 Graph-based Multi-View Fusion and Local Adaptation: Mitigating Within-Household Confusability for Speaker Identification
标题:基于图的多视点融合和局部自适应:减少说话人辨认的家庭内混淆
链接:https://arxiv.org/abs/2207.04081
作者:Long Chen,Yixiong Meng,Venkatesh Ravichandran,Andreas Stolcke备注:To appear in Interspeech 2022. arXiv admin note: text overlap with arXiv:2106.08207摘要:家庭场景中的说话人识别(SID)(例如,对于智能说话人)是一个重要但具有挑战性的问题,因为标记(注册)话语数量有限,声音容易混淆,人口统计不平衡。传统的说话人识别系统是从说话人的大量随机样本中概括出来的,导致来自特定群体的家庭的识别效果不佳,或者表现出高度的易混淆性。在这项工作中,我们提出了一种基于图的半监督学习方法,通过局部自适应图归一化和多视图图的多信号融合来提高家庭级SID的准确性和鲁棒性。与其他关于家庭SID、公平性和信号融合的工作不同,这项工作侧重于说话人标签推理(评分),并提供了一种简单的解决方案,以实现家庭特定的自适应和多信号融合,而无需调整嵌入或训练融合网络。在VoxCeleb数据集上的实验表明,我们的方法持续提高了不同客户群和混淆程度的家庭的性能。摘要:Speaker identification (SID) in the household scenario (e.g., for smart speakers) is an important but challenging problem due to limited number of labeled (enrollment) utterances, confusable voices, and demographic imbalances. Conventional speaker recognition systems generalize from a large random sample of speakers, causing the recognition to underperform for households drawn from specific cohorts or otherwise exhibiting high confusability. In this work, we propose a graph-based semi-supervised learning approach to improve household-level SID accuracy and robustness with locally adapted graph normalization and multi-signal fusion with multi-view graphs. Unlike other work on household SID, fairness, and signal fusion, this work focuses on speaker label inference (scoring) and provides a simple solution to realize household-specific adaptation and multi-signal fusion without tuning the embeddings or training a fusion network. Experiments on the VoxCeleb dataset demonstrate that our approach consistently improves the performance across households with different customer cohorts and degrees of confusability.
【1】 pMCT: Patched Multi-Condition Training for Robust Speech Recognition
标题:PMCT:用于稳健语音识别的补丁多条件训练
链接:https://arxiv.org/abs/2207.04949
* 与cs.SD语音【13】为同一篇
作者:Pablo Peso Parada,Agnieszka Dobrowolska,Karthikeyan Saravanan,Mete Ozay备注:Accepted at Interspeech 2022摘要:我们提出了一种新的用于鲁棒自动语音识别(ASR)的补丁多条件训练(pMCT)方法。pMCT通过混合{it patches}从干净和扭曲的语音中提取的相同语音,采用多条件音频修改和修补(MAMP)。使用贴片修正信号的训练提高了模型在噪声混响场景中的鲁棒性。在LibriSpeech数据集上对我们提出的pMCT进行了评估,结果表明,与使用vanilla多条件训练(MCT)相比,pMCT有了改进。为了分析鲁棒ASR,我们在语音数据集上采用了pMCT,语音数据集是一个使用LibriSpeech中的语音创建的噪声混响数据集。在分析中,与MCT相比,pMCT实现了23.1%的相对功率降低。摘要:We propose a novel Patched Multi-Condition Training (pMCT) method for robust Automatic Speech Recognition (ASR). pMCT employs Multi-condition Audio Modification and Patching (MAMP) via mixing {\it patches} of the same utterance extracted from clean and distorted speech. Training using patch-modified signals improves robustness of models in noisy reverberant scenarios. Our proposed pMCT is evaluated on the LibriSpeech dataset showing improvement over using vanilla Multi-Condition Training (MCT). For analyses on robust ASR, we employed pMCT on the VOiCES dataset which is a noisy reverberant dataset created using utterances from LibriSpeech. In the analyses, pMCT achieves 23.1% relative WER reduction compared to the MCT.
【2】 Multi-Frequency Information Enhanced Channel Attention Module for Speaker Representation Learning
标题:用于说话人表征学习的多频信息增强通道注意模块
链接:https://arxiv.org/abs/2207.04540
* 与cs.SD语音【14】为同一篇
作者:Mufan Sang,John H. L. Hansen备注:Accepted to Interspeech 2022摘要:最近,注意力机制已成功应用于基于神经网络的说话人验证系统。将压缩和激励块合并到卷积神经网络中取得了显著的性能。然而,它使用全局平均池(GAP)来简单地沿时间和频率维度平均特征,这无法在特征映射中保留足够的说话人信息。在本研究中,我们证明了GAP是时频域离散余弦变换(DCT)的特例,在数学上仅使用频率分解中的最低频率分量。为了增强说话人信息提取能力,我们提出利用多频率信息,设计两个新颖有效的注意模块,即单频单通道(SFSC)注意模块和多频单通道(MFSC)注意模块。基于离散余弦变换,该注意力模块可以有效地从多个频率分量中捕获更多的说话人信息。我们在VoxCeleb数据集上进行了全面的实验,并在第一个48-UTD法医语料库上进行了探索性评估。实验结果表明,我们提出的SFSC和MFSC注意模块可以有效地生成更具辨别力的说话人表示,在不添加额外网络参数的情况下,其性能优于ResNet34 SE和ECAPA-TDNN系统,EER分别降低了20.9%和20.2%。摘要:Recently, attention mechanisms have been applied successfully in neural network-based speaker verification systems. Incorporating the Squeeze-and-Excitation block into convolutional neural networks has achieved remarkable performance. However, it uses global average pooling (GAP) to simply average the features along time and frequency dimensions, which is incapable of preserving sufficient speaker information in the feature maps. In this study, we show that GAP is a special case of a discrete cosine transform (DCT) on time-frequency domain mathematically using only the lowest frequency component in frequency decomposition. To strengthen the speaker information extraction ability, we propose to utilize multi-frequency information and design two novel and effective attention modules, called Single-Frequency Single-Channel (SFSC) attention module and Multi-Frequency Single-Channel (MFSC) attention module. The proposed attention modules can effectively capture more speaker information from multiple frequency components on the basis of DCT. We conduct comprehensive experiments on the VoxCeleb datasets and a probe evaluation on the 1st 48-UTD forensic corpus. Experimental results demonstrate that our proposed SFSC and MFSC attention modules can efficiently generate more discriminative speaker representations and outperform ResNet34-SE and ECAPA-TDNN systems with relative 20.9% and 20.2% reduction in EER, without adding extra network parameters.
【3】 Intermediate-layer output Regularization for Attention-based Speech Recognition with Shared Decoder
标题:基于注意力的语音识别中的中间层输出正则化
链接:https://arxiv.org/abs/2207.04177
* 与cs.SD语音【15】为同一篇
作者:Jicheng Zhang,Yizhou Peng,Haihua Xu,Yi He,Eng Siong Chng,Hao Huang备注:5 pages. Submitted to INTERSPEECH 2022摘要:通过编码器端的多任务训练实现中间层输出(ILO)正则化已被证明是在广泛的端到端ASR框架上产生改进结果的有效方法。在本文中,我们提出了一种新的方法来进行不同的国际劳工组织正规化训练。我们没有使用需要更多训练开销的传统多任务方法,而是直接将中间层输出作为解码器的输入,也就是说,我们的解码器不仅接受最终编码器层的输出作为输入,还将编码器ILO的输出作为训练期间的输入。使用该方法,由于编码器和解码器同时“正则化”,与基于ILO的连接时序分类方法以及未使用该方法的原始基于注意力的建模方法相比,网络训练更充分,一致性更好。摘要:Intermediate layer output (ILO) regularization by means of multitask training on encoder side has been shown to be an effective approach to yielding improved results on a wide range of end-to-end ASR frameworks. In this paper, we propose a novel method to do ILO regularized training differently. Instead of using conventional multitask methods that entail more training overhead, we directly make the intermediate layer output as input to the decoder, that is, our decoder not only accepts the output of the final encoder layer as input, it also takes the output of the encoder ILO as input during training. With the proposed method, as both encoder and decoder are simultaneously "regularized", the network is more sufficiently trained, consistently leading to improved results, over the ILO-based CTC method, as well as over the original attention-based modeling method without the proposed method employed.
【4】 Internal Language Model Estimation based Language Model Fusion for Cross-Domain Code-Switching Speech Recognition
标题:基于内部语言模型估计的跨域码变语音识别语言模型融合
链接:https://arxiv.org/abs/2207.04176
* 与cs.SD语音【16】为同一篇
作者:Yizhou Peng,Yufei Liu,Jicheng Zhang,Haihua Xu,Yi He,Hao Huang,Eng Siong Chng备注:5 pages. Submitted to INTERSPEECH 2022摘要:在域内和跨域语音识别任务中,基于内部语言模型估计(ILME)的语言模型(LM)融合比传统的浅层融合显著改善了识别结果。在本文中,我们尝试将ILME方法应用于跨域码切换语音识别(CSSR)工作。具体来说,我们的好奇心来自几个方面。首先,我们想知道基于ILME的LM融合对于域内和跨域CSSR任务的有效性。我们在合并或不合并两个代码交换域的情况下验证了这一点。更重要的是,我们通过合并两个单语数据集来训练端到端(E2E)语音识别模型,并观察所提出的基于ILME的LM融合对CSSR的有效性。在来自东南亚和另一个中国大陆CS数据集的SEAM上的实验结果证明了所提出的基于ILME的LM融合方法的有效性。摘要:Internal Language Model Estimation (ILME) based language model (LM) fusion has been shown significantly improved recognition results over conventional shallow fusion in both intra-domain and cross-domain speech recognition tasks. In this paper, we attempt to apply our ILME method to cross-domain code-switching speech recognition (CSSR) work. Specifically, our curiosity comes from several aspects. First, we are curious about how effective the ILME-based LM fusion is for both intra-domain and cross-domain CSSR tasks. We verify this with or without merging two code-switching domains. More importantly, we train an end-to-end (E2E) speech recognition model by means of merging two monolingual data sets and observe the efficacy of the proposed ILME-based LM fusion for CSSR. Experimental results on SEAME that is from Southeast Asian and another Chinese Mainland CS data set demonstrate the effectiveness of the proposed ILME-based LM fusion method.
【5】 Graph-based Multi-View Fusion and Local Adaptation: Mitigating Within-Household Confusability for Speaker Identification
标题:基于图的多视点融合和局部自适应:减少说话人辨认的家庭内混淆
链接:https://arxiv.org/abs/2207.04081
* 与cs.SD语音【17】为同一篇
作者:Long Chen,Yixiong Meng,Venkatesh Ravichandran,Andreas Stolcke备注:To appear in Interspeech 2022. arXiv admin note: text overlap with arXiv:2106.08207摘要:家庭场景中的说话人识别(SID)(例如,对于智能说话人)是一个重要但具有挑战性的问题,因为标记(注册)话语数量有限,声音容易混淆,人口统计不平衡。传统的说话人识别系统是从说话人的大量随机样本中概括出来的,导致来自特定群体的家庭的识别效果不佳,或者表现出高度的易混淆性。在这项工作中,我们提出了一种基于图的半监督学习方法,通过局部自适应图归一化和多视图图的多信号融合来提高家庭级SID的准确性和鲁棒性。与其他关于家庭SID、公平性和信号融合的工作不同,这项工作侧重于说话人标签推理(评分),并提供了一种简单的解决方案,以实现家庭特定的自适应和多信号融合,而无需调整嵌入或训练融合网络。在VoxCeleb数据集上的实验表明,我们的方法持续提高了不同客户群和混淆程度的家庭的性能。摘要:Speaker identification (SID) in the household scenario (e.g., for smart speakers) is an important but challenging problem due to limited number of labeled (enrollment) utterances, confusable voices, and demographic imbalances. Conventional speaker recognition systems generalize from a large random sample of speakers, causing the recognition to underperform for households drawn from specific cohorts or otherwise exhibiting high confusability. In this work, we propose a graph-based semi-supervised learning approach to improve household-level SID accuracy and robustness with locally adapted graph normalization and multi-signal fusion with multi-view graphs. Unlike other work on household SID, fairness, and signal fusion, this work focuses on speaker label inference (scoring) and provides a simple solution to realize household-specific adaptation and multi-signal fusion without tuning the embeddings or training a fusion network. Experiments on the VoxCeleb dataset demonstrate that our approach consistently improves the performance across households with different customer cohorts and degrees of confusability.
【6】 Audio-Visual Segmentation
标题:视听分割
链接:https://arxiv.org/abs/2207.05042
* 与cs.SD语音【1】为同一篇
作者:Jinxing Zhou,Jianyuan Wang,Jiayi Zhang,Weixuan Sun,Jing Zhang,Stan Birchfield,Dan Guo,Lingpeng Kong,Meng Wang,Yiran Zhong备注:Accepted to ECCV 2022; Jinxing Zhou and Jianyuan Wang contributed equally; Meng Wang and Yiran Zhong are corresponding authors; Code is available at this https URL摘要:我们提出了一个新的问题,称为视听分割(AVS),其目标是输出在图像帧时产生声音的对象的像素级地图。为了促进这项研究,我们构建了第一个视听分割基准(AVSBench),为音频视频中的发声对象提供像素级注释。该基准研究了两种设置:1)单声源半监督视听分割和2)多声源全监督视听分割。为了解决AVS问题,我们提出了一种新方法,该方法使用时间像素级视听交互模块注入音频语义,作为视觉分割过程的指导。我们还设计了一种正则化损失,以鼓励在训练期间进行视听映射。在AVSBench上进行的定量和定性实验将我们的方法与相关任务中的几种现有方法进行了比较,表明该方法有望在音频和像素视觉语义之间建立桥梁。代码位于https://github.com/OpenNLPLab/AVSBench.摘要:We propose to explore a new problem called audio-visual segmentation (AVS), in which the goal is to output a pixel-level map of the object(s) that produce sound at the time of the image frame. To facilitate this research, we construct the first audio-visual segmentation benchmark (AVSBench), providing pixel-wise annotations for the sounding objects in audible videos. Two settings are studied with this benchmark: 1) semi-supervised audio-visual segmentation with a single sound source and 2) fully-supervised audio-visual segmentation with multiple sound sources. To deal with the AVS problem, we propose a novel method that uses a temporal pixel-wise audio-visual interaction module to inject audio semantics as guidance for the visual segmentation process. We also design a regularization loss to encourage the audio-visual mapping during training. Quantitative and qualitative experiments on the AVSBench compare our approach to several existing methods from related tasks, demonstrating that the proposed method is promising for building a bridge between the audio and pixel-wise visual semantics. Code is available at https://github.com/OpenNLPLab/AVSBench.
【7】 Speaker Anonymization with Phonetic Intermediate Representations
标题:基于语音中间表示的说话人匿名
链接:https://arxiv.org/abs/2207.04834
* 与cs.SD语音【2】为同一篇
作者:Sarina Meyer,Florian Lux,Pavel Denisov,Julia Koch,Pascal Tilli,Ngoc Thang Vu备注:Accepted at Interspeech 2022摘要:在这项工作中,我们提出了一种说话人匿名管道,该管道利用高质量的自动语音识别和合成系统来生成以语音转录和匿名说话人嵌入为条件的语音。使用手机作为中间表示可以确保几乎完全消除输入中的说话人身份信息,同时尽可能保留原始语音内容。我们在LibriSpeech和VCTK语料库上的实验结果揭示了两个关键发现:1)虽然自动语音识别产生不完美的转录,但我们的神经语音合成系统可以处理此类错误,使我们的系统可行且鲁棒;2)将来自不同资源的说话人嵌入相结合是有益的,并且它们的适当归一化至关重要。总的来说,我们的最终最佳系统在对懒惰的知情攻击者的隐私鲁棒性方面显著优于2020年语音隐私挑战赛中提供的基线,同时保持匿名语音的高可懂度和自然度。摘要:In this work, we propose a speaker anonymization pipeline that leverages high quality automatic speech recognition and synthesis systems to generate speech conditioned on phonetic transcriptions and anonymized speaker embeddings. Using phones as the intermediate representation ensures near complete elimination of speaker identity information from the input while preserving the original phonetic content as much as possible. Our experimental results on LibriSpeech and VCTK corpora reveal two key findings: 1) although automatic speech recognition produces imperfect transcriptions, our neural speech synthesis system can handle such errors, making our system feasible and robust, and 2) combining speaker embeddings from different resources is beneficial and their appropriate normalization is crucial. Overall, our final best system outperforms significantly the baselines provided in the Voice Privacy Challenge 2020 in terms of privacy robustness against a lazy-informed attacker while maintaining high intelligibility and naturalness of the anonymized speech.
【8】 Multi-level Fusion of Wav2vec 2.0 and BERT for Multimodal Emotion Recognition
标题:Wav2vec 2.0和BERT的多层融合多模式情感识别
链接:https://arxiv.org/abs/2207.04697
* 与cs.SD语音【3】为同一篇
作者:Zihan Zhao,Yanfeng Wang,Yu Wang备注:Accepted to INTERSPEECH 2022摘要:近年来,多模态情感识别的研究和应用越来越广泛。然而,多模式情感识别面临着缺乏数据的挑战。为了解决这个问题,我们建议使用转移学习,它利用最先进的预训练模型,包括wav2vec 2.0和BERT来完成这项任务。探索了多层次融合方法,包括基于协同注意的早期融合和基于两种嵌入训练模型的晚期融合。此外,提出了一种多粒度框架,该框架不仅提取帧级语音嵌入,还提取段级嵌入,包括电话、音节和单词级语音嵌入,以进一步提高性能。通过将基于共同注意力的早期融合模型和晚期融合模型与多粒度特征提取框架相结合,我们在IEMOCAP数据集上获得了优于最佳基线方法1.3%未加权精度(UA)的结果。摘要:The research and applications of multimodal emotion recognition have become increasingly popular recently. However, multimodal emotion recognition faces the challenge of lack of data. To solve this problem, we propose to use transfer learning which leverages state-of-the-art pre-trained models including wav2vec 2.0 and BERT for this task. Multi-level fusion approaches including coattention-based early fusion and late fusion with the models trained on both embeddings are explored. Also, a multi-granularity framework which extracts not only frame-level speech embeddings but also segment-level embeddings including phone, syllable and word-level speech embeddings is proposed to further boost the performance. By combining our coattention-based early fusion model and late fusion model with the multi-granularity feature extraction framework, we obtain result that outperforms best baseline approaches by 1.3% unweighted accuracy (UA) on the IEMOCAP dataset.
【9】 The HCCL System for the NIST SRE21
标题:NIST SRE21的HCCL系统
链接:https://arxiv.org/abs/2207.04676
* 与cs.SD语音【4】为同一篇
作者:Zhuo Li,Runqiu Xiao,Hangting Chen,Zhenduo Zhao,Zihan Zhang,Wenchao Wang备注:accepted by interspeech 2022摘要:本文描述了HCCL团队为NIST 2021说话人识别评估(NIST SRE21)开发的系统。我们首先探索了各种最先进的说话人嵌入提取器,并结合一种新的圆损耗来获得有区别的深度说话人嵌入。考虑到跨通道和跨语言说话人识别是SRE21的关键挑战,我们介绍了几种技术来减少跨域失配。具体来说,将编解码器和语音增强直接应用于原始语音,以消除编解码器和环境噪声不匹配。我们将直接作用于语音以消除相对显式失配的方法统称为数据自适应方法。实验表明,数据自适应方法比我们的基线提高了15%。此外,在说话人嵌入上部署了一些流行的后端域自适应算法,以缓解由隐式失配引起的说话人性能下降。分数校准是我们在SRE21中的一个重大失败。原因是,参数过多的分数校准容易导致过度拟合问题。摘要:This paper describes the systems developed by the HCCL team for the NIST 2021 speaker recognition evaluation (NIST SRE21).We first explore various state-of-the-art speaker embedding extractors combined with a novel circle loss to obtain discriminative deep speaker embeddings. Considering that cross-channel and cross-linguistic speaker recognition are the key challenges of SRE21, we introduce several techniques to reduce the cross-domain mismatch. Specifically, Codec and speech enhancement are directly applied to the raw speech to eliminate the codecs and the environment noise mismatch. We denote the methods that work directly on speech to eliminate the relatively explicit mismatches collectively as data adaptation methods. Experiments show that data adaption methods achieve 15\% improvements over our baseline. Furthermore, some popular back-ends domain adaptation algorithms are deployed on speaker embeddings to alleviate speaker performance degradation caused by the implicit mismatch. Score calibration is a major failure for us in SRE21. The reason is that score calibration with too many parameters easily lead to overfitting problems.
【10】 Speaker consistency loss and step-wise optimization for semi-supervised joint training of TTS and ASR using unpaired text data
标题:基于非配对文本数据的TTS和ASR半监督联合训练说话人一致性损失及逐步优化
链接:https://arxiv.org/abs/2207.04659
* 与cs.SD语音【5】为同一篇
作者:Naoki Makishima,Satoshi Suzuki,Atsushi Ando,Ryo Masumura备注:Accepted to INTERSPEECH 2022摘要:在本文中,我们研究了文本到语音(TTS)和自动语音识别(ASR)的半监督联合训练,其中有少量配对数据和大量未配对文本数据可用。传统的研究形成了一个称为TTS-ASR管道的循环,其中多峰值TTS模型从具有参考语音的文本中合成语音,ASR模型从合成语音中重建文本,然后两个模型都以循环一致性损失进行训练。然而,合成语音不能反映参考语音的说话人特征,训练后合成语音变得过于容易ASR模型识别。这不仅降低了TTS模型的质量,而且限制了ASR模型的改进。为了解决这个问题,我们提出了改进基于循环一致性的训练,减少说话人一致性损失和逐步优化。说话人一致性损失使合成语音的说话人特征更接近参考语音的说话人特征。在分步优化中,我们首先在训练两个模型之前冻结TTS模型的参数,以避免TTS模型对ASR模型的过度适应。实验结果证明了该方法的有效性。摘要:In this paper, we investigate the semi-supervised joint training of text to speech (TTS) and automatic speech recognition (ASR), where a small amount of paired data and a large amount of unpaired text data are available. Conventional studies form a cycle called the TTS-ASR pipeline, where the multispeaker TTS model synthesizes speech from text with a reference speech and the ASR model reconstructs the text from the synthesized speech, after which both models are trained with a cycle-consistency loss. However, the synthesized speech does not reflect the speaker characteristics of the reference speech and the synthesized speech becomes overly easy for the ASR model to recognize after training. This not only decreases the TTS model quality but also limits the ASR model improvement. To solve this problem, we propose improving the cycleconsistency-based training with a speaker consistency loss and step-wise optimization. The speaker consistency loss brings the speaker characteristics of the synthesized speech closer to that of the reference speech. In the step-wise optimization, we first freeze the parameter of the TTS model before both models are trained to avoid over-adaptation of the TTS model to the ASR model. Experimental results demonstrate the efficacy of the proposed method.
【11】 DelightfulTTS 2: End-to-End Speech Synthesis with Adversarial Vector-Quantized Auto-Encoders
标题:DelightfulTTS 2:采用对抗性矢量量化自动编码器的端到端语音合成
链接:https://arxiv.org/abs/2207.04646
* 与cs.SD语音【6】为同一篇
作者:Yanqing Liu,Ruiqing Xue,Lei He,Xu Tan,Sheng Zhao备注:To appear in Interspeech 2022摘要:当前的文本到语音(TTS)系统通常利用级联声学模型和声码器管道,以mel频谱图作为中间表示,这有两个局限性:1)声学模型和声码器单独训练,而不是联合优化,这会产生级联错误;2) 中间语音表示(例如,mel频谱图)是预先设计的,并且丢失了次优的相位信息。为了解决这些问题,本文开发了DelightfulTTS 2,这是一种新的端到端语音合成系统,具有自动学习的语音表示和联合优化的声学模型和声码器。具体来说,1)我们提出了一种新的基于矢量量化自动编码器和对抗训练的编解码器网络(VQ-GAN),用于提取中间帧级语音表示(而不是传统的mel谱图表示)并重建语音波形;2) 我们联合优化了声学模型(基于TTS)和声码器(VQ-GAN的解码器),并在声学模型上添加了辅助损耗,以预测中间语音表示。实验表明,与DelightfulTTS相比,DelightfulTTS 2实现了+0.14的CMOS增益,更多的方法分析进一步验证了所开发系统的有效性。摘要:Current text to speech (TTS) systems usually leverage a cascaded acoustic model and vocoder pipeline with mel-spectrograms as the intermediate representations, which suffer from two limitations: 1) the acoustic model and vocoder are separately trained instead of jointly optimized, which incurs cascaded errors; 2) the intermediate speech representations (e.g., mel-spectrogram) are pre-designed and lose phase information, which are sub-optimal. To solve these problems, in this paper, we develop DelightfulTTS 2, a new end-to-end speech synthesis system with automatically learned speech representations and jointly optimized acoustic model and vocoder. Specifically, 1) we propose a new codec network based on vector-quantized auto-encoders with adversarial training (VQ-GAN) to extract intermediate frame-level speech representations (instead of traditional representations like mel-spectrograms) and reconstruct speech waveform; 2) we jointly optimize the acoustic model (based on DelightfulTTS) and the vocoder (the decoder of VQ-GAN), with an auxiliary loss on the acoustic model to predict intermediate speech representations. Experiments show that DelightfulTTS 2 achieves a CMOS gain +0.14 over DelightfulTTS, and more method analyses further verify the effectiveness of the developed system.
【12】 Towards Proper Contrastive Self-supervised Learning Strategies For Music Audio Representation
标题:音乐音频表征中恰当的对比性自我监督学习策略
链接:https://arxiv.org/abs/2207.04471
* 与cs.SD语音【7】为同一篇
作者:Jeong Choi,Seongwon Jang,Hyunsouk Cho,Sehee Chung备注:2022 IEEE International Conference on Multimedia and Expo (ICME)摘要:自监督学习的共同研究目标是提取任意下游任务将受益的一般表示。在这项工作中,我们研究了从不同的对比自监督学习方案中学习到的音乐音频表示,并在涉及不同音乐感知水平的各种音乐信息检索(MIR)任务中实证评估了嵌入向量。我们对结果进行分析,以探讨针对不同MIR任务的对比学习策略的正确方向。我们表明,这些表征总体上传达了关于音乐听觉特征的全面信息,尽管每种自我监督策略在某些信息方面都有其自身的有效性。摘要:The common research goal of self-supervised learning is to extract a general representation which an arbitrary downstream task would benefit from. In this work, we investigate music audio representation learned from different contrastive self-supervised learning schemes and empirically evaluate the embedded vectors on various music information retrieval (MIR) tasks where different levels of the music perception are concerned. We analyze the results to discuss the proper direction of contrastive learning strategies for different MIR tasks. We show that these representations convey a comprehensive information about the auditory characteristics of music in general, although each of the self-supervision strategies has its own effectiveness in certain aspect of information.
【13】 Joint Analysis of Acoustic Scenes and Sound Events with Weakly labeled Data
标题:基于弱标签数据的声场景与声事件联合分析
链接:https://arxiv.org/abs/2207.04357
* 与cs.SD语音【8】为同一篇
作者:Shunsuke Tsubaki,Keisuke Imoto,Nobutaka Ono备注:Accepted to IWAENC2022摘要:考虑到声场景和声音事件彼此密切相关,在以前的一些论文中,提出了利用基于多任务学习(MTL)的神经网络对声场景和声音事件进行联合分析。在传统方法中,将强监督方案应用于MTL模型中的声音事件检测,这需要在模型训练中对声音事件进行强标记;然而,注释强事件标签相当耗时。因此,在本文中,我们提出了一种基于MTL框架的声音场景和声音事件联合分析方法,该框架具有声音事件的弱标签。特别是,在该方法中,我们引入了用于声音事件检测弱监督训练的多实例学习方案,并评估了四个池函数,即最大池、平均池、指数软最大池和注意力池。使用部分TUT声学场景2016/2017和TUT声音事件2016/2017数据集获得的实验结果表明,在场景分类和事件检测性能方面,基于MTL的弱标签方法优于传统的基于单任务的弱标签场景分类和事件检测模型。摘要:Considering that acoustic scenes and sound events are closely related to each other, in some previous papers, a joint analysis of acoustic scenes and sound events utilizing multitask learning (MTL)-based neural networks was proposed. In conventional methods, a strongly supervised scheme is applied to sound event detection in MTL models, which requires strong labels of sound events in model training; however, annotating strong event labels is quite time-consuming. In this paper, we thus propose a method for the joint analysis of acoustic scenes and sound events based on the MTL framework with weak labels of sound events. In particular, in the proposed method, we introduce the multiple-instance learning scheme for weakly supervised training of sound event detection and evaluate four pooling functions, namely, max pooling, average pooling, exponential softmax pooling, and attention pooling. Experimental results obtained using parts of the TUT Acoustic Scenes 2016/2017 and TUT Sound Events 2016/2017 datasets show that the proposed MTL-based method with weak labels outperforms the conventional single-task-based scene classification and event detection models with weak labels in terms of both the scene classification and event detection performances.
【14】 A Comparative Study of Self-supervised Speech Representation Based Voice Conversion
标题:基于自监督语音表示的语音转换方法的比较研究
链接:https://arxiv.org/abs/2207.04356
* 与cs.SD语音【9】为同一篇
作者:Wen-Chin Huang,Shu-Wen Yang,Tomoki Hayashi,Tomoki Toda备注:Accepted to IEEE Journal of Selected Topics in Signal Processing. arXiv admin note: substantial text overlap with arXiv:2110.06280摘要:我们对基于自监督语音表示(S3R)的语音转换(VC)进行了大规模比较研究。在识别合成VC中,S3R具有替代昂贵的监督表示(如语音后验图(PPG))的潜力,因此具有吸引力,这是最先进的VC系统通常采用的。使用我们之前开发的开源VC软件S3PRL-VC,我们使用语音转换挑战2020(VCC2020)数据集,在三种VC设置下提供了一系列深入的客观和主观分析:语内/跨语言任意对一(A2O)和任意对任意(A2A)VC。我们从多个方面研究了基于S3R的VC,包括模型类型、多语言性和监督。我们还研究了k-均值聚类后离散化过程的效果,并展示了它在A2A设置中的改进。最后,通过与最先进的VC系统的比较,证明了基于S3R的VC的竞争力,并揭示了可能的改进方向。摘要:We present a large-scale comparative study of self-supervised speech representation (S3R)-based voice conversion (VC). In the context of recognition-synthesis VC, S3Rs are attractive owing to their potential to replace expensive supervised representations such as phonetic posteriorgrams (PPGs), which are commonly adopted by state-of-the-art VC systems. Using S3PRL-VC, an open-source VC software we previously developed, we provide a series of in-depth objective and subjective analyses under three VC settings: intra-/cross-lingual any-to-one (A2O) and any-to-any (A2A) VC, using the voice conversion challenge 2020 (VCC2020) dataset. We investigated S3R-based VC in various aspects, including model type, multilinguality, and supervision. We also studied the effect of a post-discretization process with k-means clustering and showed how it improves in the A2A setting. Finally, the comparison with state-of-the-art VC systems demonstrates the competitiveness of S3R-based VC and also sheds light on the possible improving directions.
【15】 Dual-path Attention is All You Need for Audio-Visual Speech Extraction
标题:双路径注意力是音视频语音提取所需要的全部
链接:https://arxiv.org/abs/2207.04213
* 与cs.SD语音【10】为同一篇
作者:Zhongweiyang Xu,Xulin Fan,Mark Hasegawa-Johnson摘要:视听目标语音提取旨在通过观察嘴唇运动从噪声混合物中提取特定说话人的语音,结合时域语音分离模型和视觉特征提取器(CNN)取得了重大进展。融合音频和视频信息的一个问题是它们具有不同的时间分辨率。目前大多数研究都是沿着时间维度对视觉特征进行采样,以便音频和视频特征能够及时对齐。然而,我们认为唇部运动应该主要包含长期或电话层面的信息。基于这一假设,我们提出了一种新的融合视听特征的方法。我们观察到,对于DPRN,块间维度的时间分辨率可能非常接近视频帧的时间分辨率。与{sepformer}类似,DPRN中的LSTM被块内和块间自注意力所取代,但在该算法中,块间注意力将视觉特征合并为一个额外的特征流。这可以防止视觉线索的上采样,从而实现更高效的视听融合。结果表明,与其他基于时域的视听融合模型相比,我们取得了更好的结果。摘要:Audio-visual target speech extraction, which aims to extract a certain speaker's speech from the noisy mixture by looking at lip movements, has made significant progress combining time-domain speech separation models and visual feature extractors (CNN). One problem of fusing audio and video information is that they have different time resolutions. Most current research upsamples the visual features along the time dimension so that audio and video features are able to align in time. However, we believe that lip movement should mostly contain long-term, or phone-level information. Based on this assumption, we propose a new way to fuse audio-visual features. We observe that for DPRNN \cite{dprnn}, the interchunk dimension's time resolution could be very close to the time resolution of video frames. Like \cite{sepformer}, the LSTM in DPRNN is replaced by intra-chunk and inter-chunk self-attention, but in the proposed algorithm, inter-chunk attention incorporates the visual features as an additional feature stream. This prevents the upsampling of visual cues, resulting in more efficient audio-visual fusion. The result shows we achieve superior results compared with other time-domain based audio-visual fusion models.
【16】 Learning to Separate Voices by Spatial Regions
标题:学习按空间区域分离声音
链接:https://arxiv.org/abs/2207.04203
* 与cs.SD语音【11】为同一篇
作者:Zhongweiyang Xu,Romit Roy Choudhury备注:Accepted to ICML 2020. For associated audio samples, see this https URL摘要:我们考虑双耳应用中的音频-语音分离问题,例如耳机和助听器。虽然今天的神经网络表现非常好(用两个麦克风分离$4+$个源),但它们假设已知或固定的最大源数K。此外,今天的模型是以有监督的方式训练的,使用从通用源、环境和人头形状合成的训练数据。本文旨在放松这两个约束,但代价是对问题定义进行轻微修改。我们观察到,当接收到的混合信号包含太多信源时,按区域将其分离仍然有帮助,即,将信号混合信号与用户头部周围的每个锥形扇区隔离。这需要学习每个区域的细粒度空间特性,包括人的头部施加的信号失真。我们提出了一种两阶段自监督框架,其中对来自耳机的无意听到的声音进行预处理,以提取相对干净的个性化信号,然后用于训练区域分离模型。结果显示了良好的性能,强调了个性化相对于一般监督方法的重要性。(音频样本可在我们的项目网站上获得:https://uiuc-earable-computing.github.io/binaural/.我们相信这一结果可以帮助在选择性听力、噪声消除和音频增强现实中的实际应用。摘要:We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed maximum number of sources, K. Moreover, today's models are trained in a supervised manner, using training data synthesized from generic sources, environments, and human head shapes. This paper intends to relax both these constraints at the expense of a slight alteration in the problem definition. We observe that, when a received mixture contains too many sources, it is still helpful to separate them by region, i.e., isolating signal mixtures from each conical sector around the user's head. This requires learning the fine-grained spatial properties of each region, including the signal distortions imposed by a person's head. We propose a two-stage self-supervised framework in which overheard voices from earphones are pre-processed to extract relatively clean personalized signals, which are then used to train a region-wise separation model. Results show promising performance, underscoring the importance of personalization over a generic supervised approach. (audio samples available at our project website: https://uiuc-earable-computing.github.io/binaural/. We believe this result could help real-world applications in selective hearing, noise cancellation, and audio augmented reality.
【17】 Automated Audio Captioning and Language-Based Audio Retrieval
标题:自动音频字幕和基于语言的音频检索
链接:https://arxiv.org/abs/2207.04156
* 与cs.SD语音【12】为同一篇
作者:Clive Gomes,Hyejin Park,Patrick Kollman,Yi Song备注:DCASE 2022 Competition (Task 6)摘要:该项目参与了DCASE 2022竞赛(任务6),该竞赛有两个子任务:(1)自动音频字幕和(2)基于语言的音频检索。第一个子任务涉及生成音频样本的文本描述,而第二个子任务的目标是在固定数据集中找到与给定描述匹配的音频样本。对于这两个子任务,都使用了Clotho数据集。这些模型在BLEU1、BLEU2、BLEU3、ROUGEL、METEOR、苹果酒、SPICE和SPIDEr的音频字幕评分以及R1、R5、R10和mARP10的音频检索评分上进行了评估。我们进行了一些实验,修改了这些任务的基线模型。我们的自动音频字幕的最终架构接近基线性能,而基于语言的音频检索模型已经超过了它的对应模型。摘要:This project involved participation in the DCASE 2022 Competition (Task 6) which had two subtasks: (1) Automated Audio Captioning and (2) Language-Based Audio Retrieval. The first subtask involved the generation of a textual description for audio samples, while the goal of the second was to find audio samples within a fixed dataset that match a given description. For both subtasks, the Clotho dataset was used. The models were evaluated on BLEU1, BLEU2, BLEU3, ROUGEL, METEOR, CIDEr, SPICE, and SPIDEr scores for audio captioning and R1, R5, R10 and mARP10 scores for audio retrieval. We have conducted a handful of experiments that modify the baseline models for these tasks. Our final architecture for Automated Audio Captioning is close to the baseline performance, while our model for Language-Based Audio Retrieval has surpassed its counterpart.
机器翻译,仅供参考