今天跟大家分享一篇语音相关的论文合集:cs.SD语音10篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily


cs.SD语音

【1】 An Experimental Study on Private Aggregation of Teacher Ensemble  Learning for End-to-End Speech Recognition

标题:端到端语音识别中教师群体学习的私人聚集实验研究

链接:https://arxiv.org/abs/2210.05614

作者:Chao-Han Huck Yang,I-Fan Chen,Andreas Stolcke,Sabato Marco Siniscalchi,Chin-Hui Lee
机构:Georgia Institute of Technology, USA,  Amazon Alexa AI, USA and ,NTNU, Norway
备注:5 pages. Accepted to IEEE SLT 2022. A first version draft was finished in Aug 2021
摘要:差分隐私(Differential Privacy,DP)是一种通过对隐私数据施加噪声失真来保护用于训练深度模型的用户信息的数据保护方法。这样的噪声扰动经常导致自动语音识别(ASR)中的严重性能降级,以便满足隐私预算。当处理由小的$\varepsilon$值控制的噪声影响时,教师集合的私有聚合(PATE)利用集合概率来提高ASR准确性。在这项工作中,我们将PATE学习扩展到处理动态模式,即语音,并对ASR进行了首次实验研究,以避免声学数据泄漏。我们在开源的LibriSpeech和TIMIT语料库上评估了三个端到端的深度模型,包括LAS、混合注意力/CTC和RNN转换器。PATE学习增强的ASR模型优于基准DP-SGD机制,尤其是在严格的DP预算下,使用LibriSpeech评估RNN换能器模型时,相对误字率降低了26.2% ~ 27.5%。本文还介绍了另一种基于公共语料库预训练的保留动态预测的ASR方案。
摘要:Differential privacy (DP) is one data protection avenue to safeguard user information used for training deep models by imposing noisy distortion on privacy data. Such a noise perturbation often results in a severe performance degradation in automatic speech recognition (ASR) in order to meet a privacy budget $\varepsilon$. Private aggregation of teacher ensemble (PATE) utilizes ensemble probabilities to improve ASR accuracy when dealing with the noise effects controlled by small values of $\varepsilon$. In this work, we extend PATE learning to work with dynamic patterns, namely speech, and perform one very first experimental study on ASR to avoid acoustic data leakage. We evaluate three end-to-end deep models, including LAS, hybrid attention/CTC, and RNN transducer, on the open-source LibriSpeech and TIMIT corpora. PATE learning-enhanced ASR models outperform the benchmark DP-SGD mechanisms, especially under strict DP budgets, giving relative word error rate reductions between 26.2% and 27.5% for RNN transducer model evaluated with LibriSpeech. We also introduce another DP-preserving ASR solution with public speech corpus pre-training.


【2】 On the Use of Semantically-Aligned Speech Representations for Spoken  Language Understanding

标题:论语义对齐言语表征在口语理解中的应用

链接:https://arxiv.org/abs/2210.05291

作者:Gaëlle Laperrière,Valentin Pelloin,Mickaël Rouvier,Themos Stafylakis,Yannick Estève
机构:LIA - Avignon Universit´e, France,  LIUM - Le Mans Universit´e, France,  Omilia - Conversational Intelligence, Greece
备注:Accepted in IEEE SLT 2022. This work was performed using HPC resources from GENCI/IDRIS (grant 2022 AD011012565) and received funding from the EU H2020 research and innovation programme under the Marie Sklodowska-Curie ESPERANTO project (grant agreement No 101007666), through the SELMA project (grant No 957017) and from the French ANR through the AISSPER project (ANR-19-CE23-0004)
摘要:在这篇文章中,我们研究了语义对齐的语音表示在端到端口语理解中的应用。我们采用了最近引入的SAMU-XLSR模型,该模型被设计为生成单个嵌入,该嵌入在话语级别捕获语义,跨不同语言在语义上对齐。该模型结合了声学帧级语音表示学习模型(XLS-R)和语言不可知BERT语句嵌入(LaBSE)模型。结果表明,在端到端的SLU框架下,使用SAMU-XLSR模型代替初始的XLS-R模型可以显著提高系统的性能。最后,我们展示了使用该模型在SLU中实现语言可移植性的好处。
摘要:In this paper we examine the use of semantically-aligned speech representations for end-to-end spoken language understanding (SLU). We employ the recently-introduced SAMU-XLSR model, which is designed to generate a single embedding that captures the semantics at the utterance level, semantically aligned across different languages. This model combines the acoustic frame-level speech representation learning model (XLS-R) with the Language Agnostic BERT Sentence Embedding (LaBSE) model. We show that the use of the SAMU-XLSR model instead of the initial XLS-R model improves significantly the performance in the framework of end-to-end SLU. Finally, we present the benefits of using this model towards language portability in SLU.


【3】 GAN You Hear Me? Reclaiming Unconditional Speech Synthesis from  Diffusion Models

标题:你能听到我说话吗?从扩散模型中回收无条件语音合成

链接:https://arxiv.org/abs/2210.05271

作者:Matthew Baas,Herman Kamper
机构:MediaLab, Electrical & Electronic Engineering, Stellenbosch University, South Africa
备注:6 pages, 2 figures, 2 tables. Accepted at IEEE SLT 2022
摘要:本文提出了一种新的生成式对抗网络(GAN),即音频风格GAN(ASGAN)。如在图像合成模型的StyleGAN系列中,ASGAN将采样噪声映射到解纠缠的潜在向量,该潜在向量然后被映射到音频特征序列,使得在每一层抑制信号混叠。为了成功训练ASGAN,我们引入了一些新技术,包括对自适应鉴别器增强的修改,以概率性地跳过鉴别器更新。ASGAN在Google Speech Commands数据集的无条件语音合成方面取得了最先进的成果。它也比性能最好的扩散模型快得多。通过一个鼓励解开纠缠的设计,ASGAN能够执行语音转换和语音编辑,而不需要显式的培训。ASGAN证明了GANs仍然具有与扩散模型的高度竞争力。编码、型号、样品:https://github.com/RF5/simple-asgan/。
摘要:We propose AudioStyleGAN (ASGAN), a new generative adversarial network (GAN) for unconditional speech synthesis. As in the StyleGAN family of image synthesis models, ASGAN maps sampled noise to a disentangled latent vector which is then mapped to a sequence of audio features so that signal aliasing is suppressed at every layer. To successfully train ASGAN, we introduce a number of new techniques, including a modification to adaptive discriminator augmentation to probabilistically skip discriminator updates. ASGAN achieves state-of-the-art results in unconditional speech synthesis on the Google Speech Commands dataset. It is also substantially faster than the top-performing diffusion models. Through a design that encourages disentanglement, ASGAN is able to perform voice conversion and speech editing without being explicitly trained to do so. ASGAN demonstrates that GANs are still highly competitive with diffusion models. Code, models, samples: https://github.com/RF5/simple-asgan/.


【4】 MFCCA:Multi-Frame Cross-Channel attention for multi-speaker ASR in  Multi-party meeting scenario

标题:MFCCA:多方会议场景下多说话人ASR的多帧跨信道注意力

链接:https://arxiv.org/abs/2210.05265

作者:Fan Yu,Shiliang Zhang,Pengcheng Guo,Yuhao Liang,Zhihao Du,Yuxiao Lin,Lei Xie
机构:Northwestern Polytechnical University, Xi’an, China, College of Computer Science and Technology, Zhejiang University, Hangzhou, China
备注:Accepted by SLT 2022
摘要:近年来,更好地利用来自麦克风阵列的多通道信号的跨通道注意力在多方会议场景中显示出有希望的结果。跨通道的注意力集中在学习不同通道序列之间的全局相关性或在每个时间步长有效地利用细粒度的通道信息。考虑到麦克风阵列接收声音的延迟,提出了一种多帧跨通道注意力模型,该模型对相邻帧之间的跨通道信息进行建模,以利用帧级和通道级知识的互补性。此外,还提出了一种多层卷积机制来融合多通道输出,并提出了一种通道掩蔽策略来解决训练和推理之间的通道数失配问题.在真实语料库AliMeeting上的实验结果表明,该模型在Eval集和Test集上的CER分别比单通道模型减少了31.7%和37.0%.此外,在具有可比性的模型参数和训练数据的情况下,我们提出的模型在AliMeeting语料库上取得了新的SOTA性能,与最近举行的ICASSP 2022 M2MeT挑战赛(多通道多说话人ASR挑战赛)中排名靠前的系统相比。
摘要:Recently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel attention focuses on either learning global correlations between sequences of different channels or exploiting fine-grained channel-wise information effectively at each time step. Considering the delay of microphone array receiving sound, we propose a multi-frame cross-channel attention, which models cross-channel information between adjacent frames to exploit the complementarity of both frame-wise and channel-wise knowledge. Besides, we also propose a multi-layer convolutional mechanism to fuse the multi-channel output and a channel masking strategy to combat the channel number mismatch problem between training and inference. Experiments on the AliMeeting, a real-world corpus, reveal that our proposed model outperforms single-channel model by 31.7\% and 37.0\% CER reduction on Eval and Test sets. Moreover, with comparable model parameters and training data, our proposed model achieves a new SOTA performance on the AliMeeting corpus, as compared with the top ranking systems in the ICASSP2022 M2MeT challenge, a recently held multi-channel multi-speaker ASR challenge.


【5】 Deep Spectro-temporal Artifacts for Detecting Synthesized Speech

标题:用于检测合成语音的深谱-时间伪影

链接:https://arxiv.org/abs/2210.05254

作者:Xiaohui Liu,Meng Liu,Lin Zhang,Linjuan Zhang,Chang Zeng,Kai Li,Nan Li,Kong Aik Lee,Longbiao Wang,Jianwu Dang
机构:Tianjin University, Tianjin, China, National Institute of Informatics, Tokyo, Japan, Taiyuan University of Technology, Taiyuan, China, Japan Advanced Institute of Science, and Technology, Nomi, Ishikawa, Japan, Institute for Infocomm Research, A★STAR, Singapore
备注:7 pages, 1 figures, Accecpted by Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia
摘要:音频深度合成检测(ADD)挑战赛已举办,以检测生成的类人语音。利用我们提交的系统,本文提供了对音轨1(低质量虚假音频检测)和音轨2(部分虚假音频检测)的总体评估。本文利用原始时域信号、频谱特征以及深度嵌入特征检测频谱-时域伪影。为了解决跟踪1,低质量数据增强、通过微调的域自适应以及各种互补特征信息融合在我们的系统中被集合。通过可视化方法分析了不同特征子系统的聚类特性,说明了贪婪融合策略的有效性。对于轨迹2,使用自监督学习结构来检测帧转换和平滑,以捕获时域中的PF攻击操纵。我们在第一和第二赛道分别排名第四和第五。
摘要:The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of track 1 (Low-quality Fake Audio Detection) and track 2 (Partially Fake Audio Detection). In this paper, spectro-temporal artifacts were detected using raw temporal signals, spectral features, as well as deep embedding features. To address track 1, low-quality data augmentation, domain adaptation via finetuning, and various complementary feature information fusion were aggregated in our system. Furthermore, we analyzed the clustering characteristics of subsystems with different features by visualization method and explained the effectiveness of our proposed greedy fusion strategy. As for track 2, frame transition and smoothing were detected using self-supervised learning structure to capture the manipulation of PF attacks in the time domain. We ranked 4th and 5th in track 1 and track 2, respectively.


【6】 CTC Alignments Improve Autoregressive Translation

标题:CTC对齐改进自回归翻译

链接:https://arxiv.org/abs/2210.05200

作者:Brian Yan,Siddharth Dalmia,Yosuke Higuchi,Graham Neubig,Florian Metze,Alan W Black,Shinji Watanabe
机构:Language Technologies Institute, Carnegie Mellon University, USA, Department of Communications and Computer Engineering, Waseda University, Japan, Human Language Technology Center of Excellence, Johns Hopkins University, USA
摘要:连接主义时态分类(CTC)是一种广泛用于自动语音识别(ASR)的方法,其执行条件独立单调对齐。然而,对于翻译,由于任务的上下文和非单调性,CTC表现出明显的局限性,因此在翻译质量方面落后于注意解码器方法。在本研究中,我们认为,如果将CTC应用于CTC/注意力联合框架中,CTC的核心特性可以弥补纯注意力模型在训练和解码过程中的几个关键弱点,那么CTC实际上对翻译是有意义的。为了验证这一猜想,我们修改了最初为ASR提出的混合CTC/注意力模型,使其支持文本到文本的翻译(MT)和语音到文本的翻译(ST)。我们提出的CTC/注意联合模型在六个基准翻译任务中的表现优于纯注意基线。
摘要:Connectionist Temporal Classification (CTC) is a widely used approach for automatic speech recognition (ASR) that performs conditionally independent monotonic alignment. However for translation, CTC exhibits clear limitations due to the contextual and non-monotonic nature of the task and thus lags behind attentional decoder approaches in terms of translation quality. In this work, we argue that CTC does in fact make sense for translation if applied in a joint CTC/attention framework wherein CTC's core properties can counteract several key weaknesses of pure-attention models during training and decoding. To validate this conjecture, we modify the Hybrid CTC/Attention model originally proposed for ASR to support text-to-text translation (MT) and speech-to-text translation (ST). Our proposed joint CTC/attention models outperform pure-attention baselines across six benchmark translation tasks.


【7】 DiffRoll: Diffusion-based Generative Music Transcription with  Unsupervised Pretraining Capability

标题:DiffRoll:具有无监督预训练能力的基于扩散的生成性音乐转录

链接:https://arxiv.org/abs/2210.05148

作者:Kin Wai Cheuk,Ryosuke Sawata,Toshimitsu Uesaka,Naoki Murata,Naoya Takahashi,Shusuke Takahashi,Dorien Herremans,Yuki Mitsufuji
机构:Singapore University of Technology and Design, Singapore, Agency for Science, Technology and Research, Singapore, Sony Group Corporation, Tokyo, Japan
摘要:本文提出了一种新的生成式方法DiffRoll来解决音乐自动转录问题。我们不把AMT看作是一个区别性任务,在这个任务中,模型被训练成把声谱图转换成钢琴卷,我们把它看作是一个条件生成任务,在这个任务中,我们训练我们的模型,从声谱图条件下的纯高斯噪声中生成看起来逼真的钢琴卷。这种新的AMT配方使DiffRoll能够转录、生成甚至修复音乐。由于无分类器的性质,DiffRoll也能够在只有钢琴卷可用的不成对数据集上进行训练。我们的实验表明,DiffRoll比它的区别性对应物高17.9个百分点(ppt)。并且我们的消融研究还表明它比类似的现有方法的性能好3.70ppt。
摘要:In this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our model to generate realistic looking piano rolls from pure Gaussian noise conditioned on spectrograms. This new AMT formulation enables DiffRoll to transcribe, generate and even inpaint music. Due to the classifier-free nature, DiffRoll is also able to be trained on unpaired datasets where only piano rolls are available. Our experiments show that DiffRoll outperforms its discriminative counterpart by 17.9 percentage points (ppt.) and our ablation studies also indicate that it outperforms similar existing methods by 3.70 ppt.


【8】 The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge  2022

标题:VoxCeleb说话人识别挑战赛2022的DKU-Tencent系统

链接:https://arxiv.org/abs/2210.05092

作者:Xiaoyi Qin,Na Li,Yuke Lin,Yiwei Ding,Chao Weng,Dan Su,Ming Li
机构:Data Science Research Center, Duke Kunshan University, Kunshan, China,  Tencent AI Lab, Shenzhen, China
摘要:本文是DKU-Tencent系统对VoxCeleb说话人识别挑战赛2022(VoxSRC 22)的系统描述。在此挑战中,我们将重点关注轨道1和轨道3。对于track 1,采用多个骨干网络来提取帧级特征。由于track 1侧重于跨年龄情景,因此我们采用跨年龄试验并进行QMF来校准评分。基于量值的质量度量实现了较大的改进。对于track 3这一半监督域自适应任务,采用伪标签法进行域自适应。考虑到聚类过程中的噪声标签,将ArcFace替换为子中心ArcFace。最终提交的任务1中实现了0.107 mDCF,任务3中实现了7.135% EER。
摘要:This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For track1, multiple backbone networks are adopted to extract frame-level features. Since track1 focus on the cross-age scenarios, we adopt the cross-age trials and perform QMF to calibrate score. The magnitude-based quality measures achieve a large improvement. For track3, the semi-supervised domain adaptation task, the pseudo label method is adopted to make domain adaptation. Considering the noise labels in clustering, the ArcFace is replaced by Sub-center ArcFace. The final submission achieves 0.107 mDCF in task1 and 7.135% EER in task3.


【9】 ConchShell: A Generative Adversarial Networks that Turns Pictures into  Piano Music

标题:ConchShell:一个将图片转化为钢琴音乐的生成性对抗性网络

链接:https://arxiv.org/abs/2210.05076

作者:Wanpeng Fan,Yuanzhi Su,Yuxin Huang
机构:† School of Computer Science and Cyber Engineering, Guangzhou University, Guangzhou, China, ⋆ Law School, Guangzhou University, Guangzhou, China
备注:5 pages
摘要:我们提出了ConchShell,一个多模态生成对抗框架,它将图片作为网络的输入,并生成与图片上下文匹配的钢琴音乐样本。受I3 D的启发,提出了一种新的图像特征表示方法:时间卷积神经网络(TCNN),其用于在时间维度上伪造图像的特征。虽然我们的图像数据只包括六个类别,但我们提出的框架将具有创新性和商业意义。该项目将为3D游戏配音、短视频配乐、元实境背景音乐的实时生成等工作提供技术思路。我们还发布了一个新的数据集--海滩-海洋-钢琴数据集(BOPD)1,其中包含3,000多幅图像和1,500多首钢琴作品。该数据集将支持多模式图像到音乐的研究。
摘要:We present ConchShell, a multi-modal generative adversarial framework that takes pictures as input to the network and generates piano music samples that match the picture context. Inspired by I3D, we introduce a novel image feature representation method: time-convolutional neural network (TCNN), which is used to forge features for images in the temporal dimension. Although our image data consists of only six categories, our proposed framework will be innovative and commercially meaningful. The project will provide technical ideas for work such as 3D game voice overs, short-video soundtracks, and real-time generation of metaverse background music.We have also released a new dataset, the Beach-Ocean-Piano Dataset (BOPD) 1, which contains more than 3,000 images and more than 1,500 piano pieces. This dataset will support multimodal image-to-music research.


【10】 Automated Audio Captioning via Fusion of Low- and High- Dimensional  Features

标题:融合低维和高维特征的自动音频字幕

链接:https://arxiv.org/abs/2210.05037

作者:Jianyuan Sun,Xubo Liu,Xinhao Mei,Mark D. Plumbley,Volkan Kilic,Wenwu Wang
机构:Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey, UK, Department of Electrical and Electronics Engineering, Izmir Katip Celebi University, Turkey, College of Computer Science and Technology, Qingdao University, China
摘要:自动音频字幕(AAC)旨在使用简单的句子描述音频剪辑的内容。现有的AAC方法是基于编码器-解码器架构开发的,该架构的成功归因于使用称为PANN的预训练CNN 10作为编码器来学习丰富的音频表示。AAC是一项极具挑战性的任务,因为它的高维人才空间涉及到各种场景的音频。现有方法仅使用PANN的高维表示作为解码器的输入。然而,低维表示可以保留与可以忽略的高维表示一样多的音频信息。另外,虽然高维方法可以通过从现有音频字幕中学习来预测音频字幕,但是其缺乏鲁棒性和效率。针对这些问题,提出了一种融合低维和高维特征的AAC框架。本文提出了一种新的AAC编解码器框架--低高维特征融合(LHDFF)模型。此外,在LHDFF中,通过融合中间卷积层输出的低维特征和最终层输出的高维特征,提出了一种新的PANs编码器--残差PANs(RPANs)。为了充分挖掘低维融合特征、高维融合特征和高维特征各自的信息,提出了双变换解码器结构来并行生成字幕。特别地,提出了一种概率融合方法,该方法可以集中两种Transformer解码器各自的优点,确保系统的整体性能得到改善。实验结果表明,LHDFF在Clotho和AudioCaps数据集上的性能优于其他模型
摘要:Automated audio captioning (AAC) aims to describe the content of an audio clip using simple sentences. Existing AAC methods are developed based on an encoder-decoder architecture that success is attributed to the use of a pre-trained CNN10 called PANNs as the encoder to learn rich audio representations. AAC is a highly challenging task due to its high-dimensional talent space involves audio of various scenarios. Existing methods only use the high-dimensional representation of the PANNs as the input of the decoder. However, the low-dimension representation may retain as much audio information as the high-dimensional representation may be neglected. In addition, although the high-dimensional approach may predict the audio captions by learning from existing audio captions, which lacks robustness and efficiency. To deal with these challenges, a fusion model which integrates low- and high-dimensional features AAC framework is proposed. In this paper, a new encoder-decoder framework is proposed called the Low- and High-Dimensional Feature Fusion (LHDFF) model for AAC. Moreover, in LHDFF, a new PANNs encoder is proposed called Residual PANNs (RPANNs) by fusing the low-dimensional feature from the intermediate convolution layer output and the high-dimensional feature from the final layer output of PANNs. To fully explore the information of the low- and high-dimensional fusion feature and high-dimensional feature respectively, we proposed dual transformer decoder structures to generate the captions in parallel. Especially, a probabilistic fusion approach is proposed that can ensure the overall performance of the system is improved by concentrating on the respective advantages of the two transformer decoders. Experimental results show that LHDFF achieves the best performance on the Clotho and AudioCaps datasets compared with other existing models


eess.AS音频处理

【1】 An Experimental Study on Private Aggregation of Teacher Ensemble  Learning for End-to-End Speech Recognition

标题:端到端语音识别中教师群体学习的私人聚集实验研究

链接:https://arxiv.org/abs/2210.05614

* 与cs.SD语音【1】为同一篇

作者:Chao-Han Huck Yang,I-Fan Chen,Andreas Stolcke,Sabato Marco Siniscalchi,Chin-Hui Lee
机构:Georgia Institute of Technology, USA,  Amazon Alexa AI, USA and ,NTNU, Norway
备注:5 pages. Accepted to IEEE SLT 2022. A first version draft was finished in Aug 2021
摘要:差分隐私(Differential Privacy,DP)是一种通过对隐私数据施加噪声失真来保护用于训练深度模型的用户信息的数据保护方法。这样的噪声扰动经常导致自动语音识别(ASR)中的严重性能降级,以便满足隐私预算。当处理由小的$\varepsilon$值控制的噪声影响时,教师集合的私有聚合(PATE)利用集合概率来提高ASR准确性。在这项工作中,我们将PATE学习扩展到处理动态模式,即语音,并对ASR进行了首次实验研究,以避免声学数据泄漏。我们在开源的LibriSpeech和TIMIT语料库上评估了三个端到端的深度模型,包括LAS、混合注意力/CTC和RNN转换器。PATE学习增强的ASR模型优于基准DP-SGD机制,尤其是在严格的DP预算下,使用LibriSpeech评估RNN换能器模型时,相对误字率降低了26.2% ~ 27.5%。本文还介绍了另一种基于公共语料库预训练的保留动态预测的ASR方案。
摘要:Differential privacy (DP) is one data protection avenue to safeguard user information used for training deep models by imposing noisy distortion on privacy data. Such a noise perturbation often results in a severe performance degradation in automatic speech recognition (ASR) in order to meet a privacy budget $\varepsilon$. Private aggregation of teacher ensemble (PATE) utilizes ensemble probabilities to improve ASR accuracy when dealing with the noise effects controlled by small values of $\varepsilon$. In this work, we extend PATE learning to work with dynamic patterns, namely speech, and perform one very first experimental study on ASR to avoid acoustic data leakage. We evaluate three end-to-end deep models, including LAS, hybrid attention/CTC, and RNN transducer, on the open-source LibriSpeech and TIMIT corpora. PATE learning-enhanced ASR models outperform the benchmark DP-SGD mechanisms, especially under strict DP budgets, giving relative word error rate reductions between 26.2% and 27.5% for RNN transducer model evaluated with LibriSpeech. We also introduce another DP-preserving ASR solution with public speech corpus pre-training.


【2】 On the Use of Semantically-Aligned Speech Representations for Spoken  Language Understanding

标题:论语义对齐言语表征在口语理解中的应用

链接:https://arxiv.org/abs/2210.05291

* 与cs.SD语音【2】为同一篇

作者:Gaëlle Laperrière,Valentin Pelloin,Mickaël Rouvier,Themos Stafylakis,Yannick Estève
机构:LIA - Avignon Universit´e, France,  LIUM - Le Mans Universit´e, France,  Omilia - Conversational Intelligence, Greece
备注:Accepted in IEEE SLT 2022. This work was performed using HPC resources from GENCI/IDRIS (grant 2022 AD011012565) and received funding from the EU H2020 research and innovation programme under the Marie Sklodowska-Curie ESPERANTO project (grant agreement No 101007666), through the SELMA project (grant No 957017) and from the French ANR through the AISSPER project (ANR-19-CE23-0004)
摘要:在这篇文章中,我们研究了语义对齐的语音表示在端到端口语理解中的应用。我们采用了最近引入的SAMU-XLSR模型,该模型被设计为生成单个嵌入,该嵌入在话语级别捕获语义,跨不同语言在语义上对齐。该模型结合了声学帧级语音表示学习模型(XLS-R)和语言不可知BERT语句嵌入(LaBSE)模型。结果表明,在端到端的SLU框架下,使用SAMU-XLSR模型代替初始的XLS-R模型可以显著提高系统的性能。最后,我们展示了使用该模型在SLU中实现语言可移植性的好处。
摘要:In this paper we examine the use of semantically-aligned speech representations for end-to-end spoken language understanding (SLU). We employ the recently-introduced SAMU-XLSR model, which is designed to generate a single embedding that captures the semantics at the utterance level, semantically aligned across different languages. This model combines the acoustic frame-level speech representation learning model (XLS-R) with the Language Agnostic BERT Sentence Embedding (LaBSE) model. We show that the use of the SAMU-XLSR model instead of the initial XLS-R model improves significantly the performance in the framework of end-to-end SLU. Finally, we present the benefits of using this model towards language portability in SLU.


【3】 GAN You Hear Me? Reclaiming Unconditional Speech Synthesis from  Diffusion Models

标题:你能听到我说话吗?从扩散模型中回收无条件语音合成

链接:https://arxiv.org/abs/2210.05271

* 与cs.SD语音【3】为同一篇

作者:Matthew Baas,Herman Kamper
机构:MediaLab, Electrical & Electronic Engineering, Stellenbosch University, South Africa
备注:6 pages, 2 figures, 2 tables. Accepted at IEEE SLT 2022
摘要:本文提出了一种新的生成式对抗网络(GAN),即音频风格GAN(ASGAN)。如在图像合成模型的StyleGAN系列中,ASGAN将采样噪声映射到解纠缠的潜在向量,该潜在向量然后被映射到音频特征序列,使得在每一层抑制信号混叠。为了成功训练ASGAN,我们引入了一些新技术,包括对自适应鉴别器增强的修改,以概率性地跳过鉴别器更新。ASGAN在Google Speech Commands数据集的无条件语音合成方面取得了最先进的成果。它也比性能最好的扩散模型快得多。通过一个鼓励解开纠缠的设计,ASGAN能够执行语音转换和语音编辑,而不需要显式的培训。ASGAN证明了GANs仍然具有与扩散模型的高度竞争力。编码、型号、样品:https://github.com/RF5/simple-asgan/。
摘要:We propose AudioStyleGAN (ASGAN), a new generative adversarial network (GAN) for unconditional speech synthesis. As in the StyleGAN family of image synthesis models, ASGAN maps sampled noise to a disentangled latent vector which is then mapped to a sequence of audio features so that signal aliasing is suppressed at every layer. To successfully train ASGAN, we introduce a number of new techniques, including a modification to adaptive discriminator augmentation to probabilistically skip discriminator updates. ASGAN achieves state-of-the-art results in unconditional speech synthesis on the Google Speech Commands dataset. It is also substantially faster than the top-performing diffusion models. Through a design that encourages disentanglement, ASGAN is able to perform voice conversion and speech editing without being explicitly trained to do so. ASGAN demonstrates that GANs are still highly competitive with diffusion models. Code, models, samples: https://github.com/RF5/simple-asgan/.


【4】 MFCCA:Multi-Frame Cross-Channel attention for multi-speaker ASR in  Multi-party meeting scenario

标题:MFCCA:多方会议场景下多说话人ASR的多帧跨信道注意力

链接:https://arxiv.org/abs/2210.05265

* 与cs.SD语音【4】为同一篇

作者:Fan Yu,Shiliang Zhang,Pengcheng Guo,Yuhao Liang,Zhihao Du,Yuxiao Lin,Lei Xie
机构:Northwestern Polytechnical University, Xi’an, China, College of Computer Science and Technology, Zhejiang University, Hangzhou, China
备注:Accepted by SLT 2022
摘要:近年来,更好地利用来自麦克风阵列的多通道信号的跨通道注意力在多方会议场景中显示出有希望的结果。跨通道的注意力集中在学习不同通道序列之间的全局相关性或在每个时间步长有效地利用细粒度的通道信息。考虑到麦克风阵列接收声音的延迟,提出了一种多帧跨通道注意力模型,该模型对相邻帧之间的跨通道信息进行建模,以利用帧级和通道级知识的互补性。此外,还提出了一种多层卷积机制来融合多通道输出,并提出了一种通道掩蔽策略来解决训练和推理之间的通道数失配问题.在真实语料库AliMeeting上的实验结果表明,该模型在Eval集和Test集上的CER分别比单通道模型减少了31.7%和37.0%.此外,在具有可比性的模型参数和训练数据的情况下,我们提出的模型在AliMeeting语料库上取得了新的SOTA性能,与最近举行的ICASSP 2022 M2MeT挑战赛(多通道多说话人ASR挑战赛)中排名靠前的系统相比。
摘要:Recently cross-channel attention, which better leverages multi-channel signals from microphone array, has shown promising results in the multi-party meeting scenario. Cross-channel attention focuses on either learning global correlations between sequences of different channels or exploiting fine-grained channel-wise information effectively at each time step. Considering the delay of microphone array receiving sound, we propose a multi-frame cross-channel attention, which models cross-channel information between adjacent frames to exploit the complementarity of both frame-wise and channel-wise knowledge. Besides, we also propose a multi-layer convolutional mechanism to fuse the multi-channel output and a channel masking strategy to combat the channel number mismatch problem between training and inference. Experiments on the AliMeeting, a real-world corpus, reveal that our proposed model outperforms single-channel model by 31.7\% and 37.0\% CER reduction on Eval and Test sets. Moreover, with comparable model parameters and training data, our proposed model achieves a new SOTA performance on the AliMeeting corpus, as compared with the top ranking systems in the ICASSP2022 M2MeT challenge, a recently held multi-channel multi-speaker ASR challenge.


【5】 Deep Spectro-temporal Artifacts for Detecting Synthesized Speech

标题:用于检测合成语音的深谱-时间伪影

链接:https://arxiv.org/abs/2210.05254

* 与cs.SD语音【5】为同一篇

作者:Xiaohui Liu,Meng Liu,Lin Zhang,Linjuan Zhang,Chang Zeng,Kai Li,Nan Li,Kong Aik Lee,Longbiao Wang,Jianwu Dang
机构:Tianjin University, Tianjin, China, National Institute of Informatics, Tokyo, Japan, Taiyuan University of Technology, Taiyuan, China, Japan Advanced Institute of Science, and Technology, Nomi, Ishikawa, Japan, Institute for Infocomm Research, A★STAR, Singapore
备注:7 pages, 1 figures, Accecpted by Proceedings of the 1st International Workshop on Deepfake Detection for Audio Multimedia
摘要:音频深度合成检测(ADD)挑战赛已举办,以检测生成的类人语音。利用我们提交的系统,本文提供了对音轨1(低质量虚假音频检测)和音轨2(部分虚假音频检测)的总体评估。本文利用原始时域信号、频谱特征以及深度嵌入特征检测频谱-时域伪影。为了解决跟踪1,低质量数据增强、通过微调的域自适应以及各种互补特征信息融合在我们的系统中被集合。通过可视化方法分析了不同特征子系统的聚类特性,说明了贪婪融合策略的有效性。对于轨迹2,使用自监督学习结构来检测帧转换和平滑,以捕获时域中的PF攻击操纵。我们在第一和第二赛道分别排名第四和第五。
摘要:The Audio Deep Synthesis Detection (ADD) Challenge has been held to detect generated human-like speech. With our submitted system, this paper provides an overall assessment of track 1 (Low-quality Fake Audio Detection) and track 2 (Partially Fake Audio Detection). In this paper, spectro-temporal artifacts were detected using raw temporal signals, spectral features, as well as deep embedding features. To address track 1, low-quality data augmentation, domain adaptation via finetuning, and various complementary feature information fusion were aggregated in our system. Furthermore, we analyzed the clustering characteristics of subsystems with different features by visualization method and explained the effectiveness of our proposed greedy fusion strategy. As for track 2, frame transition and smoothing were detected using self-supervised learning structure to capture the manipulation of PF attacks in the time domain. We ranked 4th and 5th in track 1 and track 2, respectively.


【6】 CTC Alignments Improve Autoregressive Translation

标题:CTC对齐改进自回归翻译

链接:https://arxiv.org/abs/2210.05200

* 与cs.SD语音【6】为同一篇

作者:Brian Yan,Siddharth Dalmia,Yosuke Higuchi,Graham Neubig,Florian Metze,Alan W Black,Shinji Watanabe
机构:Language Technologies Institute, Carnegie Mellon University, USA, Department of Communications and Computer Engineering, Waseda University, Japan, Human Language Technology Center of Excellence, Johns Hopkins University, USA
摘要:连接主义时态分类(CTC)是一种广泛用于自动语音识别(ASR)的方法,其执行条件独立单调对齐。然而,对于翻译,由于任务的上下文和非单调性,CTC表现出明显的局限性,因此在翻译质量方面落后于注意解码器方法。在本研究中,我们认为,如果将CTC应用于CTC/注意力联合框架中,CTC的核心特性可以弥补纯注意力模型在训练和解码过程中的几个关键弱点,那么CTC实际上对翻译是有意义的。为了验证这一猜想,我们修改了最初为ASR提出的混合CTC/注意力模型,使其支持文本到文本的翻译(MT)和语音到文本的翻译(ST)。我们提出的CTC/注意联合模型在六个基准翻译任务中的表现优于纯注意基线。
摘要:Connectionist Temporal Classification (CTC) is a widely used approach for automatic speech recognition (ASR) that performs conditionally independent monotonic alignment. However for translation, CTC exhibits clear limitations due to the contextual and non-monotonic nature of the task and thus lags behind attentional decoder approaches in terms of translation quality. In this work, we argue that CTC does in fact make sense for translation if applied in a joint CTC/attention framework wherein CTC's core properties can counteract several key weaknesses of pure-attention models during training and decoding. To validate this conjecture, we modify the Hybrid CTC/Attention model originally proposed for ASR to support text-to-text translation (MT) and speech-to-text translation (ST). Our proposed joint CTC/attention models outperform pure-attention baselines across six benchmark translation tasks.


【7】 DiffRoll: Diffusion-based Generative Music Transcription with  Unsupervised Pretraining Capability

标题:DiffRoll:具有无监督预训练能力的基于扩散的生成性音乐转录

链接:https://arxiv.org/abs/2210.05148

* 与cs.SD语音【7】为同一篇

作者:Kin Wai Cheuk,Ryosuke Sawata,Toshimitsu Uesaka,Naoki Murata,Naoya Takahashi,Shusuke Takahashi,Dorien Herremans,Yuki Mitsufuji
机构:Singapore University of Technology and Design, Singapore, Agency for Science, Technology and Research, Singapore, Sony Group Corporation, Tokyo, Japan
摘要:本文提出了一种新的生成式方法DiffRoll来解决音乐自动转录问题。我们不把AMT看作是一个区别性任务,在这个任务中,模型被训练成把声谱图转换成钢琴卷,我们把它看作是一个条件生成任务,在这个任务中,我们训练我们的模型,从声谱图条件下的纯高斯噪声中生成看起来逼真的钢琴卷。这种新的AMT配方使DiffRoll能够转录、生成甚至修复音乐。由于无分类器的性质,DiffRoll也能够在只有钢琴卷可用的不成对数据集上进行训练。我们的实验表明,DiffRoll比它的区别性对应物高17.9个百分点(ppt)。并且我们的消融研究还表明它比类似的现有方法的性能好3.70ppt。
摘要:In this paper we propose a novel generative approach, DiffRoll, to tackle automatic music transcription (AMT). Instead of treating AMT as a discriminative task in which the model is trained to convert spectrograms into piano rolls, we think of it as a conditional generative task where we train our model to generate realistic looking piano rolls from pure Gaussian noise conditioned on spectrograms. This new AMT formulation enables DiffRoll to transcribe, generate and even inpaint music. Due to the classifier-free nature, DiffRoll is also able to be trained on unpaired datasets where only piano rolls are available. Our experiments show that DiffRoll outperforms its discriminative counterpart by 17.9 percentage points (ppt.) and our ablation studies also indicate that it outperforms similar existing methods by 3.70 ppt.


【8】 The DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge  2022

标题:VoxCeleb说话人识别挑战赛2022的DKU-Tencent系统

链接:https://arxiv.org/abs/2210.05092

* 与cs.SD语音【8】为同一篇

作者:Xiaoyi Qin,Na Li,Yuke Lin,Yiwei Ding,Chao Weng,Dan Su,Ming Li
机构:Data Science Research Center, Duke Kunshan University, Kunshan, China,  Tencent AI Lab, Shenzhen, China
摘要:本文是DKU-Tencent系统对VoxCeleb说话人识别挑战赛2022(VoxSRC 22)的系统描述。在此挑战中,我们将重点关注轨道1和轨道3。对于track 1,采用多个骨干网络来提取帧级特征。由于track 1侧重于跨年龄情景,因此我们采用跨年龄试验并进行QMF来校准评分。基于量值的质量度量实现了较大的改进。对于track 3这一半监督域自适应任务,采用伪标签法进行域自适应。考虑到聚类过程中的噪声标签,将ArcFace替换为子中心ArcFace。最终提交的任务1中实现了0.107 mDCF,任务3中实现了7.135% EER。
摘要:This paper is the system description of the DKU-Tencent System for the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC22). In this challenge, we focus on track1 and track3. For track1, multiple backbone networks are adopted to extract frame-level features. Since track1 focus on the cross-age scenarios, we adopt the cross-age trials and perform QMF to calibrate score. The magnitude-based quality measures achieve a large improvement. For track3, the semi-supervised domain adaptation task, the pseudo label method is adopted to make domain adaptation. Considering the noise labels in clustering, the ArcFace is replaced by Sub-center ArcFace. The final submission achieves 0.107 mDCF in task1 and 7.135% EER in task3.


【9】 ConchShell: A Generative Adversarial Networks that Turns Pictures into  Piano Music

标题:ConchShell:一个将图片转化为钢琴音乐的生成性对抗性网络

链接:https://arxiv.org/abs/2210.05076

* 与cs.SD语音【9】为同一篇

作者:Wanpeng Fan,Yuanzhi Su,Yuxin Huang
机构:† School of Computer Science and Cyber Engineering, Guangzhou University, Guangzhou, China, ⋆ Law School, Guangzhou University, Guangzhou, China
备注:5 pages
摘要:我们提出了ConchShell,一个多模态生成对抗框架,它将图片作为网络的输入,并生成与图片上下文匹配的钢琴音乐样本。受I3 D的启发,提出了一种新的图像特征表示方法:时间卷积神经网络(TCNN),其用于在时间维度上伪造图像的特征。虽然我们的图像数据只包括六个类别,但我们提出的框架将具有创新性和商业意义。该项目将为3D游戏配音、短视频配乐、元实境背景音乐的实时生成等工作提供技术思路。我们还发布了一个新的数据集--海滩-海洋-钢琴数据集(BOPD)1,其中包含3,000多幅图像和1,500多首钢琴作品。该数据集将支持多模式图像到音乐的研究。
摘要:We present ConchShell, a multi-modal generative adversarial framework that takes pictures as input to the network and generates piano music samples that match the picture context. Inspired by I3D, we introduce a novel image feature representation method: time-convolutional neural network (TCNN), which is used to forge features for images in the temporal dimension. Although our image data consists of only six categories, our proposed framework will be innovative and commercially meaningful. The project will provide technical ideas for work such as 3D game voice overs, short-video soundtracks, and real-time generation of metaverse background music.We have also released a new dataset, the Beach-Ocean-Piano Dataset (BOPD) 1, which contains more than 3,000 images and more than 1,500 piano pieces. This dataset will support multimodal image-to-music research.


【10】 Automated Audio Captioning via Fusion of Low- and High- Dimensional  Features

标题:融合低维和高维特征的自动音频字幕

链接:https://arxiv.org/abs/2210.05037

* 与cs.SD语音【10】为同一篇

作者:Jianyuan Sun,Xubo Liu,Xinhao Mei,Mark D. Plumbley,Volkan Kilic,Wenwu Wang
机构:Centre for Vision, Speech and Signal Processing (CVSSP), University of Surrey, UK, Department of Electrical and Electronics Engineering, Izmir Katip Celebi University, Turkey, College of Computer Science and Technology, Qingdao University, China
摘要:自动音频字幕(AAC)旨在使用简单的句子描述音频剪辑的内容。现有的AAC方法是基于编码器-解码器架构开发的,该架构的成功归因于使用称为PANN的预训练CNN 10作为编码器来学习丰富的音频表示。AAC是一项极具挑战性的任务,因为它的高维人才空间涉及到各种场景的音频。现有方法仅使用PANN的高维表示作为解码器的输入。然而,低维表示可以保留与可以忽略的高维表示一样多的音频信息。另外,虽然高维方法可以通过从现有音频字幕中学习来预测音频字幕,但是其缺乏鲁棒性和效率。针对这些问题,提出了一种融合低维和高维特征的AAC框架。本文提出了一种新的AAC编解码器框架--低高维特征融合(LHDFF)模型。此外,在LHDFF中,通过融合中间卷积层输出的低维特征和最终层输出的高维特征,提出了一种新的PANs编码器--残差PANs(RPANs)。为了充分挖掘低维融合特征、高维融合特征和高维特征各自的信息,提出了双变换解码器结构来并行生成字幕。特别地,提出了一种概率融合方法,该方法可以集中两种Transformer解码器各自的优点,确保系统的整体性能得到改善。实验结果表明,LHDFF在Clotho和AudioCaps数据集上的性能优于其他模型
摘要:Automated audio captioning (AAC) aims to describe the content of an audio clip using simple sentences. Existing AAC methods are developed based on an encoder-decoder architecture that success is attributed to the use of a pre-trained CNN10 called PANNs as the encoder to learn rich audio representations. AAC is a highly challenging task due to its high-dimensional talent space involves audio of various scenarios. Existing methods only use the high-dimensional representation of the PANNs as the input of the decoder. However, the low-dimension representation may retain as much audio information as the high-dimensional representation may be neglected. In addition, although the high-dimensional approach may predict the audio captions by learning from existing audio captions, which lacks robustness and efficiency. To deal with these challenges, a fusion model which integrates low- and high-dimensional features AAC framework is proposed. In this paper, a new encoder-decoder framework is proposed called the Low- and High-Dimensional Feature Fusion (LHDFF) model for AAC. Moreover, in LHDFF, a new PANNs encoder is proposed called Residual PANNs (RPANNs) by fusing the low-dimensional feature from the intermediate convolution layer output and the high-dimensional feature from the final layer output of PANNs. To fully explore the information of the low- and high-dimensional fusion feature and high-dimensional feature respectively, we proposed dual transformer decoder structures to generate the captions in parallel. Especially, a probabilistic fusion approach is proposed that can ensure the overall performance of the system is improved by concentrating on the respective advantages of the two transformer decoders. Experimental results show that LHDFF achieves the best performance on the Clotho and AudioCaps datasets compared with other existing models


机器翻译,仅供参考