【1】 Position Prediction as an Effective Pretraining Strategy
标题:职位预测是一种有效的职前训练策略
链接:https://arxiv.org/abs/2207.07611
作者:Shuangfei Zhai,Navdeep Jaitly,Jason Ramapuram,Dan Busbridge,Tatiana Likhomanenko,Joseph Yitan Cheng,Walter Talbott,Chen Huang,Hanlin Goh,Joshua Susskind摘要:Transformer由于其强大的表征能力,在自然语言处理、计算机视觉和语音识别等广泛应用中越来越受欢迎。然而,有效利用这种表征能力需要大量数据和/或强正则化,以缓解过度拟合。最近,基于掩码自动编码器的自监督预训练策略解锁了Transformer的功率,该策略依赖于直接或对比地从未掩码内容重构掩码输入。这种预训练策略已用于非线性规划中的BERT模型、语音中的Wav2Vec模型,以及最近在视觉中的MAE模型中,迫使模型使用自动编码相关目标来了解输入不同部分内容之间的关系。在本文中,我们提出了一种新的、但令人惊讶的简单的内容重建替代方法,即根据内容预测位置,而不提供位置信息。这样做需要Transformer仅从输入的内容来理解输入的不同部分之间的位置关系。这相当于一种有效的实现,其中借口任务是每个输入令牌的所有可能位置之间的分类问题。我们在视觉和语音基准上进行了实验,我们的方法比强监督训练基线带来了改进,与现代无监督/自监督预训练方法相当。我们的方法还使未经位置嵌入训练的Transformer优于使用完整位置信息训练的Transformer。摘要:Transformers have gained increasing popularity in a wide range of applications, including Natural Language Processing (NLP), Computer Vision and Speech Recognition, because of their powerful representational capacity. However, harnessing this representational capacity effectively requires a large amount of data, strong regularization, or both, to mitigate overfitting. Recently, the power of the Transformer has been unlocked by self-supervised pretraining strategies based on masked autoencoders which rely on reconstructing masked inputs, directly, or contrastively from unmasked content. This pretraining strategy which has been used in BERT models in NLP, Wav2Vec models in Speech and, recently, in MAE models in Vision, forces the model to learn about relationships between the content in different parts of the input using autoencoding related objectives. In this paper, we propose a novel, but surprisingly simple alternative to content reconstruction~-- that of predicting locations from content, without providing positional information for it. Doing so requires the Transformer to understand the positional relationships between different parts of the input, from their content alone. This amounts to an efficient implementation where the pretext task is a classification problem among all possible positions for each input token. We experiment on both Vision and Speech benchmarks, where our approach brings improvements over strong supervised training baselines and is comparable to modern unsupervised/self-supervised pretraining methods. Our method also enables Transformers trained without position embeddings to outperform ones trained with full position information.
【2】 Low-bit Shift Network for End-to-End Spoken Language Understanding
标题:用于端到端口语理解的低位移位网络
链接:https://arxiv.org/abs/2207.07497
作者:Anderson R. Avila,Khalil Bibi,Rui Heng Yang,Xinlin Li,Chao Xing,Xiao Chen备注:Accepted at INTERSPEECH 2022摘要:深度神经网络(DNN)在多个领域取得了令人瞩目的成功。多年来,这些模型的准确性随着更深层和更复杂架构的增加而提高。因此,最先进的解决方案通常计算成本高昂,因此不适合部署在边缘计算平台上。为了减轻推断卷积神经网络(CNN)的高计算、内存和功率要求,我们提出使用两个量化的功率,将连续参数量化为两个值的低比特功率。这通过消除昂贵的乘法运算和使用低比特权重来降低计算复杂度。我们采用ResNet作为解决方案的构建块,并在口语理解任务中对所提出的模型进行了评估。实验结果表明,移位神经网络结构的性能得到改善,我们的低比特量化在测试集上达到98.76%,其性能与全精度对应和最先进的解决方案相当。摘要:Deep neural networks (DNN) have achieved impressive success in multiple domains. Over the years, the accuracy of these models has increased with the proliferation of deeper and more complex architectures. Thus, state-of-the-art solutions are often computationally expensive, which makes them unfit to be deployed on edge computing platforms. In order to mitigate the high computation, memory, and power requirements of inferring convolutional neural networks (CNNs), we propose the use of power-of-two quantization, which quantizes continuous parameters into low-bit power-of-two values. This reduces computational complexity by removing expensive multiplication operations and with the use of low-bit weights. ResNet is adopted as the building block of our solution and the proposed model is evaluated on a spoken language understanding (SLU) task. Experimental results show improved performance for shift neural network architectures, with our low-bit quantization achieving 98.76 \% on the test set which is comparable performance to its full-precision counterpart and state-of-the-art solutions.
【3】 Continual Learning For On-Device Environmental Sound Classification
标题:基于持续学习的设备环境声分类方法
链接:https://arxiv.org/abs/2207.07429
作者:Yang Xiao,Xubo Liu,James King,Arshdeep Sing,Eng Siong Chng,Mark D. Plumbley,Wenwu Wang备注:The first two authors contributed equally, 5 pages one figure, submitted to DCASE2022 Workshop摘要:考虑到计算资源(例如,模型大小、运行内存)的限制,在没有灾难性遗忘的情况下连续学习新类对于设备上的环境声音分类是一个具有挑战性的问题。为了解决这个问题,我们提出了一种简单有效的连续学习方法。我们的方法通过测量每个样本的分类不确定性来选择历史数据进行训练。具体来说,我们通过观察数据的分类概率如何随着添加到分类器嵌入中的并行扰动而波动来测量不确定性。这样,与向原始数据添加扰动相比,计算成本可以显著降低。在DCASE 2019任务1和ESC-50数据集上的实验结果表明,我们提出的方法在分类精度和计算效率方面优于基线连续学习方法,表明我们的方法可以有效地增量学习新类,而不会出现设备上环境声音分类的灾难性遗忘问题。摘要:Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation resources (e.g., model size, running memory). To address this issue, we propose a simple and efficient continual learning method. Our method selects the historical data for the training by measuring the per-sample classification uncertainty. Specifically, we measure the uncertainty by observing how the classification probability of data fluctuates against the parallel perturbations added to the classifier embedding. In this way, the computation cost can be significantly reduced compared with adding perturbation to the raw data. Experimental results on the DCASE 2019 Task 1 and ESC-50 dataset show that our proposed method outperforms baseline continual learning methods on classification accuracy and computational efficiency, indicating our method can efficiently and incrementally learn new classes without the catastrophic forgetting problem for on-device environmental sound classification.
【4】 PodcastMix: A dataset for separating music and speech in podcasts
标题:Podcast Mix:用于在播客中分离音乐和语音的数据集
链接:https://arxiv.org/abs/2207.07403
作者:Nicolás Schmidt,Jordi Pons,Marius Miron备注:In proceedings of INTERSPEECH2022. Project webpage: this http URL摘要:我们介绍了PodcastMix,这是一个数据集,用于将播客中的背景音乐和前景语音分离。我们的目标是定义一个适用于训练和评估(深度学习)信源分离模型的基准。为此,我们发布了一个基于编程生成的播客的大型和多样化的训练数据集。然而,当前(深度学习)模型可能会引发泛化问题,特别是在对合成数据进行训练时。为了解决潜在的泛化问题,我们发布了一个基于真实播客的评估集,并为此设计了客观和主观测试。通过对真实播客的实验,我们发现当前(深度学习)模型可能存在泛化问题。然而,这些可以胜任,例如,我们的最佳基线分离语音,平均意见分数为3.84(评级“整体分离质量”从1到5)。数据集和基线可以在线访问。摘要:We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and evaluating (deep learning) source separation models. To that end, we release a large and diverse training dataset based on programatically generated podcasts. However, current (deep learning) models can incur into generalization issues, specially when trained on synthetic data. To target potential generalization issues, we release an evaluation set based on real podcasts for which we design objective and subjective tests. Out of our experiments with real podcasts, we find that current (deep learning) models may have generalization issues. Yet, these can perform competently, e.g., our best baseline separates speech with a mean opinion score of 3.84 (rating "overall separation quality" from 1 to 5). The dataset and baselines are accessible online.
【5】 Audio-guided Album Cover Art Generation with Genetic Algorithms
标题:基于遗传算法的音频制导专辑封面艺术生成
链接:https://arxiv.org/abs/2207.07162
作者:James Marien,Sam Leroux,Bart Dhoedt,Cedric De Boom备注:8 pages, 6 figures, 4 tables摘要:Spotify每天发布超过60000首歌曲,争夺听众注意力的竞争非常激烈。在这方面,不能低估吸引人的封面艺术的重要性,因为它与歌曲的性格和艺术家的身份息息相关,并且仍然是引导人们发现音乐的最重要途径之一。然而,封面艺术的设计是一个高度创造性、漫长且有时昂贵的过程,这可能会让人望而却步,尤其是对于非专业艺术家而言。因此,我们提出了一种新的深度学习框架来生成由音频特征引导的封面艺术。受VGAN-CLIP的启发,我们的方法非常灵活,因为单个组件可以轻松更换,无需任何再训练。本文概述了我们模型的架构细节,并讨论了由此产生的优化挑战。更具体地说,我们将利用遗传算法来克服糟糕的局部极小值和对抗性示例。我们发现,我们的框架可以为大多数体裁生成合适的封面艺术,并且视觉特征能够适应音频特征的变化。鉴于这些结果,我们相信我们的框架为扩展和更高级的音频引导视频生成任务的应用铺平了道路。摘要:Over 60,000 songs are released on Spotify every day, and the competition for the listener's attention is immense. In that regard, the importance of captivating and inviting cover art cannot be underestimated, because it is deeply entangled with a song's character and the artist's identity, and remains one of the most important gateways to lead people to discover music. However, designing cover art is a highly creative, lengthy and sometimes expensive process that can be daunting, especially for non-professional artists. For this reason, we propose a novel deep-learning framework to generate cover art guided by audio features. Inspired by VQGAN-CLIP, our approach is highly flexible because individual components can easily be replaced without the need for any retraining. This paper outlines the architectural details of our models and discusses the optimization challenges that emerge from them. More specifically, we will exploit genetic algorithms to overcome bad local minima and adversarial examples. We find that our framework can generate suitable cover art for most genres, and that the visual features adapt themselves to audio feature changes. Given these results, we believe that our framework paves the road for extensions and more advanced applications in audio-guided visual generation tasks.
【6】 The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification Challenge
标题:面向2022年欺骗感知说话人确认挑战赛的DKU-OPPO系统
链接:https://arxiv.org/abs/2207.07510
作者:Xingming Wang,Xiaoyi Qin,Yikang Wang,Yunfei Xu,Ming Li备注:Accepted by Interspeech2022摘要:本文介绍了我们的DKU-OPPO系统,用于2022年防欺骗说话人验证(SASV)挑战。首先,我们将联合任务分为说话人验证(SV)和欺骗对抗(CM),这两个任务分别进行了优化。对于ASV系统,采用了四种最先进的方法。对于CM系统,我们在挑战基线的基础上提出了两种方法来进一步提高性能,即嵌入随机抽样增强(ERSA)和一类混淆损失(OCCL)。其次,我们还探讨了SV嵌入是否有助于提高CM系统的性能。我们观察到,在域不匹配的Voxceleb2数据集上,现有CM系统的性能急剧下降。第三,我们比较了不同的融合策略,包括并行分数融合和顺序级联系统。与1.71%的SASV-EER基线相比,我们提交的级联系统在挑战官方评估集上获得了0.21%的SASV-EER。摘要:This paper describes our DKU-OPPO system for the 2022 Spoofing-Aware Speaker Verification (SASV) Challenge. First, we split the joint task into speaker verification (SV) and spoofing countermeasure (CM), these two tasks which are optimized separately. For ASV systems, four state-of-the-art methods are employed. For CM systems, we propose two methods on top of the challenge baseline to further improve the performance, namely Embedding Random Sampling Augmentation (ERSA) and One-Class Confusion Loss(OCCL). Second, we also explore whether SV embedding could help improve CM system performance. We observe a dramatic performance degradation of existing CM systems on the domain-mismatched Voxceleb2 dataset. Third, we compare different fusion strategies, including parallel score fusion and sequential cascaded systems. Compared to the 1.71% SASV-EER baseline, our submitted cascaded system obtains a 0.21% SASV-EER on the challenge official evaluation set.
【7】 PoLyScribers: Joint Training of Vocal Extractor and Lyrics Transcriber for Polyphonic Music
标题:PoLyScribers:复调音乐声乐抽取员和歌词转录员的联合训练
链接:https://arxiv.org/abs/2207.07336
作者:Xiaoxue Gao,Chitralekha Gupta,Haizhou Li备注:14 pages, TALSP submission摘要:复调音乐的歌词转录具有挑战性,因为背景音乐会影响歌词的可懂度。通常,歌词转录可以通过两步管道执行,即歌唱人声提取前端,然后是歌词转录解码器后端,其中前端和后端分别进行训练。这样的两步流水线同时存在语音提取不完善和前端与后端不匹配的问题。在这项工作中,我们提出了一种新的端到端联合训练框架,我们称之为多分类器,以联合优化用于复调音乐中歌词转录的人声提取器前端和歌词转录器后端。实验结果表明,与现有的公开测试数据集方法相比,我们提出的联合训练模型实现了实质性的改进。摘要:Lyrics transcription of polyphonic music is challenging as the background music affects lyrics intelligibility. Typically, lyrics transcription can be performed by a two step pipeline, i.e. singing vocal extraction frontend, followed by a lyrics transcriber decoder backend, where the frontend and backend are trained separately. Such a two step pipeline suffers from both imperfect vocal extraction and mismatch between frontend and backend. In this work, we propose novel end-to-end joint-training framework, that we call PoLyScribers, to jointly optimize the vocal extractor front-end and lyrics transcriber backend for lyrics transcription in polyphonic music. The experimental results show that our proposed joint-training model achieves substantial improvements over the existing approaches on publicly available test datasets.
【8】 MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound Sources
标题:MIMO-DOAnet:声源数目未知的多通道输入多输出DOA网络
链接:https://arxiv.org/abs/2207.07307
作者:Haoran Yin,Meng Ge,Yanjie Fu,Gaoyan Zhang,Longbiao Wang,Lei Zhang,Lin Qiu,Jianwu Dang备注:Accepted by Interspeech 2022摘要:最近基于神经网络的波达方向(DoA)估计算法在未知声源数的情况下表现良好。这些算法通常通过将多通道音频输入映射到单个输出(即所有源的整体空间伪频谱(SPS))来实现,即MISO。然而,这种MISO算法强烈依赖于经验阈值设置和角度假设,即声源之间的角度大于固定角度。为了解决这些局限性,我们提出了一种新的多通道输入多输出DoA网络,称为MIMO-DoAnet。与一般MISO算法不同,MIMO DoAnet借助信息空间协方差矩阵预测每个声源的SPS编码。通过这样做,检测声源数量的阈值任务变得更容易检测每个输出中是否有声源,并且声源之间的严重交互在推理阶段消失。实验结果表明,与MISO基线系统相比,MIMO DoAnet在3、4信源场景中的F1分数分别提高了18.6%和13.3%,34.4%和20.2%。结果还表明,MIMO-DoAnet有效地解决了阈值设置问题和角度假设问题。摘要:Recent neural network based Direction of Arrival (DoA) estimation algorithms have performed well on unknown number of sound sources scenarios. These algorithms are usually achieved by mapping the multi-channel audio input to the single output (i.e. overall spatial pseudo-spectrum (SPS) of all sources), that is called MISO. However, such MISO algorithms strongly depend on empirical threshold setting and the angle assumption that the angles between the sound sources are greater than a fixed angle. To address these limitations, we propose a novel multi-channel input and multiple outputs DoA network called MIMO-DoAnet. Unlike the general MISO algorithms, MIMO-DoAnet predicts the SPS coding of each sound source with the help of the informative spatial covariance matrix. By doing so, the threshold task of detecting the number of sound sources becomes an easier task of detecting whether there is a sound source in each output, and the serious interaction between sound sources disappears during inference stage. Experimental results show that MIMO-DoAnet achieves relative 18.6% and absolute 13.3%, relative 34.4% and absolute 20.2% F1 score improvement compared with the MISO baseline system in 3, 4 sources scenes. The results also demonstrate MIMO-DoAnet alleviates the threshold setting problem and solves the angle assumption problem effectively.
【9】 Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments
标题:真实噪声环境下基于增强现实耳机的方向感知自适应在线神经语音增强
链接:https://arxiv.org/abs/2207.07296
作者:Kouhei Sekiguchi,Aditya Arie Nugraha,Yicheng Du,Yoshiaki Bando,Mathieu Fontaine,Kazuyoshi Yoshii摘要:本文描述了增强现实(AR)耳机在线语音增强的实用响应和性能感知开发,该耳机可帮助用户理解在真实嘈杂回声环境(例如鸡尾酒会)中进行的对话。可以使用一种称为快速多通道非负矩阵分解(FastMNMF)的最先进的盲源分离方法,由于其无监督性质,该方法在各种环境中都能很好地工作。然而,其高昂的计算成本阻碍了其在实时处理中的应用。相反,使用深度神经网络(DNN)估计语音和噪声的空间信息的监督波束形成方法很容易适应实时处理,但在不匹配的情况下性能会急剧下降。鉴于这种互补特性,我们提出了一种基于DNN波束形成和快速MNMF引导自适应的双过程鲁棒在线语音增强方法。FastMNMF(后端)以小批量方式执行,噪声和增强语音对与原始并行训练数据一起使用,以计算允许的间隔使用反向传播更新方向感知DNN(前端)。该方法与称为加权预测误差(WPE)的盲去混响方法一起使用,用于转录说话人的噪声混响语音,该混响语音可以从视频中检测到,或通过用户的手势或眼睛注视来选择,以流式方式,并用AR技术在空间上显示转录。我们的实验表明,运行时自适应只需要12分钟的观察时间,文字错误率就提高了10多个点。摘要:This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source separation method called fast multichannel nonnegative matrix factorization (FastMNMF) that works well in various environments thanks to its unsupervised nature. Its heavy computational cost, however, prevents its application to real-time processing. In contrast, a supervised beamforming method that uses a deep neural network (DNN) for estimating spatial information of speech and noise readily fits real-time processing, but suffers from drastic performance degradation in mismatched conditions. Given such complementary characteristics, we propose a dual-process robust online speech enhancement method based on DNN-based beamforming with FastMNMF-guided adaptation. FastMNMF (back end) is performed in a mini-batch style and the noisy and enhanced speech pairs are used together with the original parallel training data for updating the direction-aware DNN (front end) with backpropagation at a computationally-allowable interval. This method is used with a blind dereverberation method called weighted prediction error (WPE) for transcribing the noisy reverberant speech of a speaker, which can be detected from video or selected by a user's hand gesture or eye gaze, in a streaming manner and spatially showing the transcriptions with an AR technique. Our experiment showed that the word error rate was improved by more than 10 points with the run-time adaptation using only twelve minutes of observation.
【10】 Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational Environments
标题:多方对话环境下神经语音增强与识别的方向感知联合自适应
链接:https://arxiv.org/abs/2207.07273
作者:Yicheng Du,Aditya Arie Nugraha,Kouhei Sekiguchi,Yoshiaki Bando,Mathieu Fontaine,Kazuyoshi Yoshii摘要:本文描述了增强现实耳机的噪声语音识别,该耳机有助于在真实的多方对话环境中进行语音通信。在模拟环境中积极研究的一种主要方法是基于在监督方式下训练的深度神经网络(DNN)顺序执行语音增强和自动语音识别(ASR)。然而,在我们的任务中,由于训练和测试条件与用户头部运动之间的不匹配,这种预训练系统无法工作。为了仅增强目标说话人的语音,我们使用基于DNN的语音掩码估计器的波束形成,该估计器可以自适应地提取对应于头部特定方向的语音分量。我们提出了一种半监督自适应方法,使用具有地面真实转录的干净语音信号和具有高度置信估计转录的噪声语音信号,在运行时联合更新掩码估计器和ASR模型。使用最先进的远程语音识别系统进行的对比实验表明,该方法显著提高了自动语音识别的性能。摘要:This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simulated environments is to sequentially perform speech enhancement and automatic speech recognition (ASR) based on deep neural networks (DNNs) trained in a supervised manner. In our task, however, such a pretrained system fails to work due to the mismatch between the training and test conditions and the head movements of the user. To enhance only the utterances of a target speaker, we use beamforming based on a DNN-based speech mask estimator that can adaptively extract the speech components corresponding to a head-relative particular direction. We propose a semi-supervised adaptation method that jointly updates the mask estimator and the ASR model at run-time using clean speech signals with ground-truth transcriptions and noisy speech signals with highly-confident estimated transcriptions. Comparative experiments using the state-of-the-art distant speech recognition system show that the proposed method significantly improves the ASR performance.
【1】 The DKU-OPPO System for the 2022 Spoofing-Aware Speaker Verification Challenge
标题:面向2022年欺骗感知说话人确认挑战赛的DKU-OPPO系统
链接:https://arxiv.org/abs/2207.07510
* 与cs.SD语音【6】为同一篇
作者:Xingming Wang,Xiaoyi Qin,Yikang Wang,Yunfei Xu,Ming Li备注:Accepted by Interspeech2022摘要:本文介绍了我们的DKU-OPPO系统,用于2022年防欺骗说话人验证(SASV)挑战。首先,我们将联合任务分为说话人验证(SV)和欺骗对抗(CM),这两个任务分别进行了优化。对于ASV系统,采用了四种最先进的方法。对于CM系统,我们在挑战基线的基础上提出了两种方法来进一步提高性能,即嵌入随机抽样增强(ERSA)和一类混淆损失(OCCL)。其次,我们还探讨了SV嵌入是否有助于提高CM系统的性能。我们观察到,在域不匹配的Voxceleb2数据集上,现有CM系统的性能急剧下降。第三,我们比较了不同的融合策略,包括并行分数融合和顺序级联系统。与1.71%的SASV-EER基线相比,我们提交的级联系统在挑战官方评估集上获得了0.21%的SASV-EER。摘要:This paper describes our DKU-OPPO system for the 2022 Spoofing-Aware Speaker Verification (SASV) Challenge. First, we split the joint task into speaker verification (SV) and spoofing countermeasure (CM), these two tasks which are optimized separately. For ASV systems, four state-of-the-art methods are employed. For CM systems, we propose two methods on top of the challenge baseline to further improve the performance, namely Embedding Random Sampling Augmentation (ERSA) and One-Class Confusion Loss(OCCL). Second, we also explore whether SV embedding could help improve CM system performance. We observe a dramatic performance degradation of existing CM systems on the domain-mismatched Voxceleb2 dataset. Third, we compare different fusion strategies, including parallel score fusion and sequential cascaded systems. Compared to the 1.71% SASV-EER baseline, our submitted cascaded system obtains a 0.21% SASV-EER on the challenge official evaluation set.
【2】 PoLyScribers: Joint Training of Vocal Extractor and Lyrics Transcriber for Polyphonic Music
标题:PoLyScribers:复调音乐声乐抽取员和歌词转录员的联合训练
链接:https://arxiv.org/abs/2207.07336
* 与cs.SD语音【7】为同一篇
作者:Xiaoxue Gao,Chitralekha Gupta,Haizhou Li备注:14 pages, TALSP submission摘要:复调音乐的歌词转录具有挑战性,因为背景音乐会影响歌词的可懂度。通常,歌词转录可以通过两步管道执行,即歌唱人声提取前端,然后是歌词转录解码器后端,其中前端和后端分别进行训练。这样的两步流水线同时存在语音提取不完善和前端与后端不匹配的问题。在这项工作中,我们提出了一种新的端到端联合训练框架,我们称之为多分类器,以联合优化用于复调音乐中歌词转录的人声提取器前端和歌词转录器后端。实验结果表明,与现有的公开测试数据集方法相比,我们提出的联合训练模型实现了实质性的改进。摘要:Lyrics transcription of polyphonic music is challenging as the background music affects lyrics intelligibility. Typically, lyrics transcription can be performed by a two step pipeline, i.e. singing vocal extraction frontend, followed by a lyrics transcriber decoder backend, where the frontend and backend are trained separately. Such a two step pipeline suffers from both imperfect vocal extraction and mismatch between frontend and backend. In this work, we propose novel end-to-end joint-training framework, that we call PoLyScribers, to jointly optimize the vocal extractor front-end and lyrics transcriber backend for lyrics transcription in polyphonic music. The experimental results show that our proposed joint-training model achieves substantial improvements over the existing approaches on publicly available test datasets.
【3】 MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound Sources
标题:MIMO-DOAnet:声源数目未知的多通道输入多输出DOA网络
链接:https://arxiv.org/abs/2207.07307
* 与cs.SD语音【8】为同一篇
作者:Haoran Yin,Meng Ge,Yanjie Fu,Gaoyan Zhang,Longbiao Wang,Lei Zhang,Lin Qiu,Jianwu Dang备注:Accepted by Interspeech 2022摘要:最近基于神经网络的波达方向(DoA)估计算法在未知声源数的情况下表现良好。这些算法通常通过将多通道音频输入映射到单个输出(即所有源的整体空间伪频谱(SPS))来实现,即MISO。然而,这种MISO算法强烈依赖于经验阈值设置和角度假设,即声源之间的角度大于固定角度。为了解决这些局限性,我们提出了一种新的多通道输入多输出DoA网络,称为MIMO-DoAnet。与一般MISO算法不同,MIMO DoAnet借助信息空间协方差矩阵预测每个声源的SPS编码。通过这样做,检测声源数量的阈值任务变得更容易检测每个输出中是否有声源,并且声源之间的严重交互在推理阶段消失。实验结果表明,与MISO基线系统相比,MIMO DoAnet在3、4信源场景中的F1分数分别提高了18.6%和13.3%,34.4%和20.2%。结果还表明,MIMO-DoAnet有效地解决了阈值设置问题和角度假设问题。摘要:Recent neural network based Direction of Arrival (DoA) estimation algorithms have performed well on unknown number of sound sources scenarios. These algorithms are usually achieved by mapping the multi-channel audio input to the single output (i.e. overall spatial pseudo-spectrum (SPS) of all sources), that is called MISO. However, such MISO algorithms strongly depend on empirical threshold setting and the angle assumption that the angles between the sound sources are greater than a fixed angle. To address these limitations, we propose a novel multi-channel input and multiple outputs DoA network called MIMO-DoAnet. Unlike the general MISO algorithms, MIMO-DoAnet predicts the SPS coding of each sound source with the help of the informative spatial covariance matrix. By doing so, the threshold task of detecting the number of sound sources becomes an easier task of detecting whether there is a sound source in each output, and the serious interaction between sound sources disappears during inference stage. Experimental results show that MIMO-DoAnet achieves relative 18.6% and absolute 13.3%, relative 34.4% and absolute 20.2% F1 score improvement compared with the MISO baseline system in 3, 4 sources scenes. The results also demonstrate MIMO-DoAnet alleviates the threshold setting problem and solves the angle assumption problem effectively.
【4】 Direction-Aware Adaptive Online Neural Speech Enhancement with an Augmented Reality Headset in Real Noisy Conversational Environments
标题:真实噪声环境下基于增强现实耳机的方向感知自适应在线神经语音增强
链接:https://arxiv.org/abs/2207.07296
* 与cs.SD语音【9】为同一篇
作者:Kouhei Sekiguchi,Aditya Arie Nugraha,Yicheng Du,Yoshiaki Bando,Mathieu Fontaine,Kazuyoshi Yoshii摘要:本文描述了增强现实(AR)耳机在线语音增强的实用响应和性能感知开发,该耳机可帮助用户理解在真实嘈杂回声环境(例如鸡尾酒会)中进行的对话。可以使用一种称为快速多通道非负矩阵分解(FastMNMF)的最先进的盲源分离方法,由于其无监督性质,该方法在各种环境中都能很好地工作。然而,其高昂的计算成本阻碍了其在实时处理中的应用。相反,使用深度神经网络(DNN)估计语音和噪声的空间信息的监督波束形成方法很容易适应实时处理,但在不匹配的情况下性能会急剧下降。鉴于这种互补特性,我们提出了一种基于DNN波束形成和快速MNMF引导自适应的双过程鲁棒在线语音增强方法。FastMNMF(后端)以小批量方式执行,噪声和增强语音对与原始并行训练数据一起使用,以计算允许的间隔使用反向传播更新方向感知DNN(前端)。该方法与称为加权预测误差(WPE)的盲去混响方法一起使用,用于转录说话人的噪声混响语音,该混响语音可以从视频中检测到,或通过用户的手势或眼睛注视来选择,以流式方式,并用AR技术在空间上显示转录。我们的实验表明,运行时自适应只需要12分钟的观察时间,文字错误率就提高了10多个点。摘要:This paper describes the practical response- and performance-aware development of online speech enhancement for an augmented reality (AR) headset that helps a user understand conversations made in real noisy echoic environments (e.g., cocktail party). One may use a state-of-the-art blind source separation method called fast multichannel nonnegative matrix factorization (FastMNMF) that works well in various environments thanks to its unsupervised nature. Its heavy computational cost, however, prevents its application to real-time processing. In contrast, a supervised beamforming method that uses a deep neural network (DNN) for estimating spatial information of speech and noise readily fits real-time processing, but suffers from drastic performance degradation in mismatched conditions. Given such complementary characteristics, we propose a dual-process robust online speech enhancement method based on DNN-based beamforming with FastMNMF-guided adaptation. FastMNMF (back end) is performed in a mini-batch style and the noisy and enhanced speech pairs are used together with the original parallel training data for updating the direction-aware DNN (front end) with backpropagation at a computationally-allowable interval. This method is used with a blind dereverberation method called weighted prediction error (WPE) for transcribing the noisy reverberant speech of a speaker, which can be detected from video or selected by a user's hand gesture or eye gaze, in a streaming manner and spatially showing the transcriptions with an AR technique. Our experiment showed that the word error rate was improved by more than 10 points with the run-time adaptation using only twelve minutes of observation.
【5】 Direction-Aware Joint Adaptation of Neural Speech Enhancement and Recognition in Real Multiparty Conversational Environments
标题:多方对话环境下神经语音增强与识别的方向感知联合自适应
链接:https://arxiv.org/abs/2207.07273
* 与cs.SD语音【10】为同一篇
作者:Yicheng Du,Aditya Arie Nugraha,Kouhei Sekiguchi,Yoshiaki Bando,Mathieu Fontaine,Kazuyoshi Yoshii摘要:本文描述了增强现实耳机的噪声语音识别,该耳机有助于在真实的多方对话环境中进行语音通信。在模拟环境中积极研究的一种主要方法是基于在监督方式下训练的深度神经网络(DNN)顺序执行语音增强和自动语音识别(ASR)。然而,在我们的任务中,由于训练和测试条件与用户头部运动之间的不匹配,这种预训练系统无法工作。为了仅增强目标说话人的语音,我们使用基于DNN的语音掩码估计器的波束形成,该估计器可以自适应地提取对应于头部特定方向的语音分量。我们提出了一种半监督自适应方法,使用具有地面真实转录的干净语音信号和具有高度置信估计转录的噪声语音信号,在运行时联合更新掩码估计器和ASR模型。使用最先进的远程语音识别系统进行的对比实验表明,该方法显著提高了自动语音识别的性能。摘要:This paper describes noisy speech recognition for an augmented reality headset that helps verbal communication within real multiparty conversational environments. A major approach that has actively been studied in simulated environments is to sequentially perform speech enhancement and automatic speech recognition (ASR) based on deep neural networks (DNNs) trained in a supervised manner. In our task, however, such a pretrained system fails to work due to the mismatch between the training and test conditions and the head movements of the user. To enhance only the utterances of a target speaker, we use beamforming based on a DNN-based speech mask estimator that can adaptively extract the speech components corresponding to a head-relative particular direction. We propose a semi-supervised adaptation method that jointly updates the mask estimator and the ASR model at run-time using clean speech signals with ground-truth transcriptions and noisy speech signals with highly-confident estimated transcriptions. Comparative experiments using the state-of-the-art distant speech recognition system show that the proposed method significantly improves the ASR performance.
【6】 Position Prediction as an Effective Pretraining Strategy
标题:职位预测是一种有效的职前训练策略
链接:https://arxiv.org/abs/2207.07611
* 与cs.SD语音【1】为同一篇
作者:Shuangfei Zhai,Navdeep Jaitly,Jason Ramapuram,Dan Busbridge,Tatiana Likhomanenko,Joseph Yitan Cheng,Walter Talbott,Chen Huang,Hanlin Goh,Joshua Susskind摘要:Transformer由于其强大的表征能力,在自然语言处理、计算机视觉和语音识别等广泛应用中越来越受欢迎。然而,有效利用这种表征能力需要大量数据和/或强正则化,以缓解过度拟合。最近,基于掩码自动编码器的自监督预训练策略解锁了Transformer的功率,该策略依赖于直接或对比地从未掩码内容重构掩码输入。这种预训练策略已用于非线性规划中的BERT模型、语音中的Wav2Vec模型,以及最近在视觉中的MAE模型中,迫使模型使用自动编码相关目标来了解输入不同部分内容之间的关系。在本文中,我们提出了一种新的、但令人惊讶的简单的内容重建替代方法,即根据内容预测位置,而不提供位置信息。这样做需要Transformer仅从输入的内容来理解输入的不同部分之间的位置关系。这相当于一种有效的实现,其中借口任务是每个输入令牌的所有可能位置之间的分类问题。我们在视觉和语音基准上进行了实验,我们的方法比强监督训练基线带来了改进,与现代无监督/自监督预训练方法相当。我们的方法还使未经位置嵌入训练的Transformer优于使用完整位置信息训练的Transformer。摘要:Transformers have gained increasing popularity in a wide range of applications, including Natural Language Processing (NLP), Computer Vision and Speech Recognition, because of their powerful representational capacity. However, harnessing this representational capacity effectively requires a large amount of data, strong regularization, or both, to mitigate overfitting. Recently, the power of the Transformer has been unlocked by self-supervised pretraining strategies based on masked autoencoders which rely on reconstructing masked inputs, directly, or contrastively from unmasked content. This pretraining strategy which has been used in BERT models in NLP, Wav2Vec models in Speech and, recently, in MAE models in Vision, forces the model to learn about relationships between the content in different parts of the input using autoencoding related objectives. In this paper, we propose a novel, but surprisingly simple alternative to content reconstruction~-- that of predicting locations from content, without providing positional information for it. Doing so requires the Transformer to understand the positional relationships between different parts of the input, from their content alone. This amounts to an efficient implementation where the pretext task is a classification problem among all possible positions for each input token. We experiment on both Vision and Speech benchmarks, where our approach brings improvements over strong supervised training baselines and is comparable to modern unsupervised/self-supervised pretraining methods. Our method also enables Transformers trained without position embeddings to outperform ones trained with full position information.
【7】 Low-bit Shift Network for End-to-End Spoken Language Understanding
标题:用于端到端口语理解的低位移位网络
链接:https://arxiv.org/abs/2207.07497
* 与cs.SD语音【2】为同一篇
作者:Anderson R. Avila,Khalil Bibi,Rui Heng Yang,Xinlin Li,Chao Xing,Xiao Chen备注:Accepted at INTERSPEECH 2022摘要:深度神经网络(DNN)在多个领域取得了令人瞩目的成功。多年来,这些模型的准确性随着更深层和更复杂架构的增加而提高。因此,最先进的解决方案通常计算成本高昂,因此不适合部署在边缘计算平台上。为了减轻推断卷积神经网络(CNN)的高计算、内存和功率要求,我们提出使用两个量化的功率,将连续参数量化为两个值的低比特功率。这通过消除昂贵的乘法运算和使用低比特权重来降低计算复杂度。我们采用ResNet作为解决方案的构建块,并在口语理解任务中对所提出的模型进行了评估。实验结果表明,移位神经网络结构的性能得到改善,我们的低比特量化在测试集上达到98.76%,其性能与全精度对应和最先进的解决方案相当。摘要:Deep neural networks (DNN) have achieved impressive success in multiple domains. Over the years, the accuracy of these models has increased with the proliferation of deeper and more complex architectures. Thus, state-of-the-art solutions are often computationally expensive, which makes them unfit to be deployed on edge computing platforms. In order to mitigate the high computation, memory, and power requirements of inferring convolutional neural networks (CNNs), we propose the use of power-of-two quantization, which quantizes continuous parameters into low-bit power-of-two values. This reduces computational complexity by removing expensive multiplication operations and with the use of low-bit weights. ResNet is adopted as the building block of our solution and the proposed model is evaluated on a spoken language understanding (SLU) task. Experimental results show improved performance for shift neural network architectures, with our low-bit quantization achieving 98.76 \% on the test set which is comparable performance to its full-precision counterpart and state-of-the-art solutions.
【8】 Continual Learning For On-Device Environmental Sound Classification
标题:基于持续学习的设备环境声分类方法
链接:https://arxiv.org/abs/2207.07429
* 与cs.SD语音【3】为同一篇
作者:Yang Xiao,Xubo Liu,James King,Arshdeep Sing,Eng Siong Chng,Mark D. Plumbley,Wenwu Wang备注:The first two authors contributed equally, 5 pages one figure, submitted to DCASE2022 Workshop摘要:考虑到计算资源(例如,模型大小、运行内存)的限制,在没有灾难性遗忘的情况下连续学习新类对于设备上的环境声音分类是一个具有挑战性的问题。为了解决这个问题,我们提出了一种简单有效的连续学习方法。我们的方法通过测量每个样本的分类不确定性来选择历史数据进行训练。具体来说,我们通过观察数据的分类概率如何随着添加到分类器嵌入中的并行扰动而波动来测量不确定性。这样,与向原始数据添加扰动相比,计算成本可以显著降低。在DCASE 2019任务1和ESC-50数据集上的实验结果表明,我们提出的方法在分类精度和计算效率方面优于基线连续学习方法,表明我们的方法可以有效地增量学习新类,而不会出现设备上环境声音分类的灾难性遗忘问题。摘要:Continuously learning new classes without catastrophic forgetting is a challenging problem for on-device environmental sound classification given the restrictions on computation resources (e.g., model size, running memory). To address this issue, we propose a simple and efficient continual learning method. Our method selects the historical data for the training by measuring the per-sample classification uncertainty. Specifically, we measure the uncertainty by observing how the classification probability of data fluctuates against the parallel perturbations added to the classifier embedding. In this way, the computation cost can be significantly reduced compared with adding perturbation to the raw data. Experimental results on the DCASE 2019 Task 1 and ESC-50 dataset show that our proposed method outperforms baseline continual learning methods on classification accuracy and computational efficiency, indicating our method can efficiently and incrementally learn new classes without the catastrophic forgetting problem for on-device environmental sound classification.
【9】 PodcastMix: A dataset for separating music and speech in podcasts
标题:Podcast Mix:用于在播客中分离音乐和语音的数据集
链接:https://arxiv.org/abs/2207.07403
* 与cs.SD语音【4】为同一篇
作者:Nicolás Schmidt,Jordi Pons,Marius Miron备注:In proceedings of INTERSPEECH2022. Project webpage: this http URL摘要:我们介绍了PodcastMix,这是一个数据集,用于将播客中的背景音乐和前景语音分离。我们的目标是定义一个适用于训练和评估(深度学习)信源分离模型的基准。为此,我们发布了一个基于编程生成的播客的大型和多样化的训练数据集。然而,当前(深度学习)模型可能会引发泛化问题,特别是在对合成数据进行训练时。为了解决潜在的泛化问题,我们发布了一个基于真实播客的评估集,并为此设计了客观和主观测试。通过对真实播客的实验,我们发现当前(深度学习)模型可能存在泛化问题。然而,这些可以胜任,例如,我们的最佳基线分离语音,平均意见分数为3.84(评级“整体分离质量”从1到5)。数据集和基线可以在线访问。摘要:We introduce PodcastMix, a dataset formalizing the task of separating background music and foreground speech in podcasts. We aim at defining a benchmark suitable for training and evaluating (deep learning) source separation models. To that end, we release a large and diverse training dataset based on programatically generated podcasts. However, current (deep learning) models can incur into generalization issues, specially when trained on synthetic data. To target potential generalization issues, we release an evaluation set based on real podcasts for which we design objective and subjective tests. Out of our experiments with real podcasts, we find that current (deep learning) models may have generalization issues. Yet, these can perform competently, e.g., our best baseline separates speech with a mean opinion score of 3.84 (rating "overall separation quality" from 1 to 5). The dataset and baselines are accessible online.
【10】 Audio-guided Album Cover Art Generation with Genetic Algorithms
标题:基于遗传算法的音频制导专辑封面艺术生成
链接:https://arxiv.org/abs/2207.07162
* 与cs.SD语音【5】为同一篇
作者:James Marien,Sam Leroux,Bart Dhoedt,Cedric De Boom备注:8 pages, 6 figures, 4 tables摘要:Spotify每天发布超过60000首歌曲,争夺听众注意力的竞争非常激烈。在这方面,不能低估吸引人的封面艺术的重要性,因为它与歌曲的性格和艺术家的身份息息相关,并且仍然是引导人们发现音乐的最重要途径之一。然而,封面艺术的设计是一个高度创造性、漫长且有时昂贵的过程,这可能会让人望而却步,尤其是对于非专业艺术家而言。因此,我们提出了一种新的深度学习框架来生成由音频特征引导的封面艺术。受VGAN-CLIP的启发,我们的方法非常灵活,因为单个组件可以轻松更换,无需任何再训练。本文概述了我们模型的架构细节,并讨论了由此产生的优化挑战。更具体地说,我们将利用遗传算法来克服糟糕的局部极小值和对抗性示例。我们发现,我们的框架可以为大多数体裁生成合适的封面艺术,并且视觉特征能够适应音频特征的变化。鉴于这些结果,我们相信我们的框架为扩展和更高级的音频引导视频生成任务的应用铺平了道路。摘要:Over 60,000 songs are released on Spotify every day, and the competition for the listener's attention is immense. In that regard, the importance of captivating and inviting cover art cannot be underestimated, because it is deeply entangled with a song's character and the artist's identity, and remains one of the most important gateways to lead people to discover music. However, designing cover art is a highly creative, lengthy and sometimes expensive process that can be daunting, especially for non-professional artists. For this reason, we propose a novel deep-learning framework to generate cover art guided by audio features. Inspired by VQGAN-CLIP, our approach is highly flexible because individual components can easily be replaced without the need for any retraining. This paper outlines the architectural details of our models and discusses the optimization challenges that emerge from them. More specifically, we will exploit genetic algorithms to overcome bad local minima and adversarial examples. We find that our framework can generate suitable cover art for most genres, and that the visual features adapt themselves to audio feature changes. Given these results, we believe that our framework paves the road for extensions and more advanced applications in audio-guided visual generation tasks.
机器翻译,仅供参考