今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理8篇。

cs.SD语音

【1】 LCSM: A Lightweight Complex Spectral Mapping Framework for Stereophonic  Acoustic Echo Cancellation

标题:LCSM:一种用于立体声回声消除的轻量级复谱映射框架

链接:https://arxiv.org/abs/2208.07277

作者:Chenggang Zhang,Jinjiang Liu,Xueliang Zhang
机构:Department of Computer Science, Inner Mongolia University, China
备注:Accepted to Interspeech 2022
摘要:传统的自适应算法在处理立体声回波抵消时会面临非唯一性问题本文首先提出了一种高效的多输入多输出(multi-input-multi-output,SAEC)算法(MIMO)方案来一次从所有麦克风信号中滤除回声。然后,我们采用一种轻量级的复杂频谱映射框架,(LCSM)方法实现端到端SAEC,无需对扬声器信号进行去相关预处理,利用现场卷积和通道空间建模来保证近端SAEC得准确性.实验结果表明,LCSM方法的泛化性能明显优于已有的方法,而且所提出的因果框架仅包含55万个参数,大大少于同类基于深度学习的方法,这对于资源有限的设备是重要的。
摘要:The traditional adaptive algorithms will face the non-uniqueness problem when dealing with stereophonic acoustic echo cancellation (SAEC). In this paper, we first propose an efficient multi-input and multi-output (MIMO) scheme based on deep learning to filter out echoes from all microphone signals at once. Then, we employ a lightweight complex spectral mapping framework (LCSM) for end-to-end SAEC without decorrelation preprocessing to the loudspeaker signals. Inplace convolution and channel-wise spatial modeling are utilized to ensure the near-end signal information is preserved. Finally, a cross-domain loss function is designed for better generalization capability. Experiments are evaluated on a variety of untrained conditions and results demonstrate that the LCSM significantly outperforms previous methods. Moreover, the proposed causal framework only has 0.55 million parameters, much less than the similar deep learning-based methods, which is important for the resource-limited devices.


【2】 Towards Parametric Speech Synthesis Using Gaussian-Markov Model of  Spectral Envelope and Wavelet-Based Decomposition of F0

标题:基于Gaussian-Markov谱包络模型和F0小波分解的参数语音合成

链接:https://arxiv.org/abs/2208.07122

作者:Mohammed Salah Al-Radhi,Tamás Gábor Csapó,Csaba Zainkó,Géza Németh
机构:Department of Telecommunications and Media Informatics, Budapest University of Technology and Economics, Budapest, Hungary
备注:accepted at EUSIPCO2022
摘要:基于神经网络的文语转换技术显著提高了合成语音的质量。(例如,Tacotron2、FastSpeech、FastPitch)通常从文本生成Mel谱图,然后使用声码器合成语音(例如,WaveNet、WaveGlow、HiFiGAN)。与传统的参数方法相比(如STRAIGHT和WORLD),基于神经网络的端到端模型存在推理速度慢,合成语音鲁棒性差和可控性差的缺点,本文提出了一种新的更新声码器,该模型训练简单,波形生成容易,采用高斯—马尔可夫模型对频谱包络进行鲁棒学习,利用基于小波的统计信号处理对F0特征进行表征和分解,既能保持良好的频谱包络,又能实现自然语音的高可控性。实验结果表明,本文提出的声码器比传统的STRAIGHT声码器具有更好的重构语音自然度,略优于WaveNet,并且比WaveRNN稍差。
摘要:Neural network-based Text-to-Speech has significantly improved the quality of synthesized speech. Prominent methods (e.g., Tacotron2, FastSpeech, FastPitch) usually generate Mel-spectrogram from text and then synthesize speech using vocoder (e.g., WaveNet, WaveGlow, HiFiGAN). Compared with traditional parametric approaches (e.g., STRAIGHT and WORLD), neural vocoder based end-to-end models suffer from slow inference speed, and the synthesized speech is usually not robust and lack of controllability. In this work, we propose a novel updated vocoder, which is a simple signal model to train and easy to generate waveforms. We use the Gaussian-Markov model toward robust learning of spectral envelope and wavelet-based statistical signal processing to characterize and decompose F0 features. It can retain the fine spectral envelope and achieve high controllability of natural speech. The experimental results demonstrate that our proposed vocoder achieves better naturalness of reconstructed speech than the conventional STRAIGHT vocoder, slightly better than WaveNet, and somewhat worse than the WaveRNN.


【3】 Analysis of impact of emotions on target speech extraction and speech  separation

标题:情感对目标语音提取和语音分离的影响分析

链接:https://arxiv.org/abs/2208.07091

作者:Ján Švec,Kateřina Žmolíková,Martin Kocour,Marc Delcroix,Tsubasa Ochiai,Ladislav Mošner,Jan Černocký
机构:Ladislav Moˇsner, Jan “Honza” ˇCernock´y, Brno University of Technology, IT,I Centre of Excellence, NTT Corporation, Japan
备注:Accepted to IWAENC 2022
摘要:近年来,盲语音分离的性能(BSS)和目标语音提取(TSE)取得了很大的进展.然而,大多数工作集中在相对良好控制的条件下,例如使用朗读语音.在更真实的情况下,性能可能会下降.造成这种下降的因素之一可能是内在的说话者可变性,如情感,这在真实语音中是常见的.本文在分析了语音的可变性的基础上,提出了一种基于语音可变性的语音识别方法,我们研究了情感对TSE和BSS的影响,并结合LibriSpeech和Ryerson的情感语音和歌曲视听数据库,建立了一个新的情感混合测试数据集,用于TSE和BSS的评价通过对照实验,我们可以分析不同的情绪对BSS和TSE性能的影响,我们观察到BSS对情绪相对鲁棒,而TSE,该方法需要识别和提取目标说话人的语音,对情感更加敏感,在对比说话人确认实验中,我们发现识别目标说话人在处理情感语音时尤其困难,我们概述了可能提高BSS和TSE系统对情感语音的鲁棒性的潜在未来方向。
摘要:Recently, the performance of blind speech separation (BSS) and target speech extraction (TSE) has greatly progressed. Most works, however, focus on relatively well-controlled conditions using, e.g., read speech. The performance may degrade in more realistic situations. One of the factors causing such degradation may be intrinsic speaker variability, such as emotions, occurring commonly in realistic speech. In this paper, we investigate the influence of emotions on TSE and BSS. We create a new test dataset of emotional mixtures for the evaluation of TSE and BSS. This dataset combines LibriSpeech and Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). Through controlled experiments, we can analyze the impact of different emotions on the performance of BSS and TSE. We observe that BSS is relatively robust to emotions, while TSE, which requires identifying and extracting the speech of a target speaker, is much more sensitive to emotions. On comparative speaker verification experiments we show that identifying the target speaker may be particularly challenging when dealing with emotional speech. Using our findings, we outline potential future directions that could improve the robustness of BSS and TSE systems toward emotional speech.


【4】 Models of Music Cognition and Composition

标题:音乐认知与创作模式

链接:https://arxiv.org/abs/2208.06878

作者:Abhimanyu Sethia,Aayush
机构:Indian Institute of Technology, Kanpur
备注:TLDR: literature review of models of music cognition and composition
摘要:与大多数认知研究一样,音乐认知也是一个跨学科领域,它试图应用认知科学的方法(神经学的,计算的和实验的)来理解音乐创作的感知和过程。本文,我们首先提出为什么音乐与认知科学家相关,并概述音乐认知的计算建模方法。然后,我们回顾了关于音乐感知的各种模型的文献,包括非计算模型,计算非认知模型和计算认知模型。最后,我们回顾了关于创造性行为建模和能够创作音乐的计算机系统的文献。由于大量来自音乐理论的专业术语被使用,我们在最后附上了相关术语及其定义的列表。
摘要:Much like most of cognition research, music cognition is an interdisciplinary field, which attempts to apply methods of cognitive science (neurological, computational and experimental) to understand the perception and process of composition of music. In this paper, we first motivate why music is relevant to cognitive scientists and give an overview of the approaches to computational modelling of music cognition. We then review literature on the various models of music perception, including non-computational models, computational non-cognitive models and computational cognitive models. Lastly, we review literature on modelling the creative behaviour and on computer systems capable of composing music. Since a lot of technical terms from music theory have been used, we have appended a list of relevant terms and their definitions at the end.


【5】 Differentiable WORLD Synthesizer-based Neural Vocoder With Application  To End-To-End Audio Style Transfer

标题:基于可微分WORLD合成器的神经声码器及其在端到端音频风格传输中的应用

链接:https://arxiv.org/abs/2208.07282

作者:Shahan Nercessian
机构:Zotope, Inc.
备注:11 pages, 2 figures
摘要:在本文中,我们提出了一种可区分WORLD合成器,并演示了其在端到端音频风格传输任务中的使用,例如(歌唱)语音转换和DDSP音色转换任务。因此,我们的基线可微分合成器没有模型参数,我们可以通过添加轻量级的黑色-一种可选择的可微分方法考虑直接提取源激发光谱,这可以改善自然度,尽管是针对较窄类别的风格转印应用。由我们的方法使用的声学特征参数化具有附加的益处,即其自然地解开音调和音色信息,使得它们可以被分别建模。此外,由于存在一种从单声道音频源估计这些声学特征的鲁棒装置,所以它允许将参数损失项添加到端到端目标函数,这可以帮助收敛和/或进一步稳定(对抗性)训练。
摘要:In this paper, we propose a differentiable WORLD synthesizer and demonstrate its use in end-to-end audio style transfer tasks such as (singing) voice conversion and the DDSP timbre transfer task. Accordingly, our baseline differentiable synthesizer has no model parameters, yet it yields adequate synthesis quality. We can extend the baseline synthesizer by appending lightweight black-box postnets which apply further processing to the baseline output in order to improve fidelity. An alternative differentiable approach considers extraction of the source excitation spectrum directly, which can improve naturalness albeit for a narrower class of style transfer applications. The acoustic feature parameterization used by our approaches has the added benefit that it naturally disentangles pitch and timbral information so that they can be modeled separately. Moreover, as there exists a robust means of estimating these acoustic features from monophonic audio sources, it allows for parameter loss terms to be added to an end-to-end objective function, which can help convergence and/or further stabilize (adversarial) training.


【6】 How Should We Evaluate Synthesized Environmental Sounds

标题:如何评价合成环境声

链接:https://arxiv.org/abs/2208.07679

作者:Yuki Okamoto,Keisuke Imoto,Shinnosuke Takamichi,Takahiro Fukumori,Yoichi Yamashita
机构:∗ Ritsumeikan University, Japan, † Doshisha University, Japan, ‡ The University of Tokyo, Japan备注:Submitted APSIPA ASC 2022
摘要:环境声的合成方法有多种,但对合成后的环境声如何进行评价却没有讨论,传统的评价方法只进行主观评价或客观评价,对哪种评价方法不明确,本文提出了一种新的评价方法———环境声综合评价法,我们研究如何评价合成环境声音。我们还提出了一种主观评价方法来评价合成的声音是否适当地代表了输入到环境中的信息在实验中,我们将所提出的评价方法与传统的评价方法进行了比较,结果表明主观评价的结果往往与客观评价的结果不同,从这些结果中我们得出结论,不仅要进行客观评价,而且要进行主观评价。
摘要:Although several methods of environmental sound synthesis have been proposed, there has been no discussion on how synthesized environmental sounds should be evaluated. Only either subjective or objective evaluations have been conducted in conventional evaluations, and it is not clear what type of evaluation should be carried out. In this paper, we investigate how to evaluate synthesized environmental sounds. We also propose a subjective evaluation methodology to evaluate whether the synthesized sound appropriately represents the information input to the environmental sound synthesis system. In our experiments, we compare the proposed and conventional evaluation methods and show that the results of subjective evaluations tended to differ from those of objective evaluations. From these results, we conclude that it is necessary to conduct not only objective evaluation but also subjective evaluation.


【7】 Uconv-Conformer: High Reduction of Input Sequence Length for End-to-End  Speech Recognition

标题:Uconv-符合者:用于端到端语音识别的输入序列长度的高度缩减

链接:https://arxiv.org/abs/2208.07657

作者:Andrei Andrusenko,Rauf Nasretdinov,Aleksei Romanenko
机构:ITMO University, St. Petersburg, Russia, STC-innovations Ltd, St. Petersburg, Russia
备注:5 pages, 1 figure
摘要:本文在标准Conformer模型的基础上提出了一种新的Uconv-Conformer结构,该结构将输入序列长度一致地减少了16倍,为了解决在时间维数显著降低的情况下的收敛问题,我们使用类似于U-Net架构的上采样模块,以确保正确的CTC丢失计算和稳定的网络训练。Uconv-Conformer架构似乎不仅在训练和推理方面更快,而且与基线Conformer相比显示出更好的WER。我们的最佳Uconv-Conformer模型显示出40.3%的历元训练时间减少,47.8%,和GPU上的推理加速分别降低了23.5%和23.5%,Librispeech test_clean和test_other上的相对WER分别降低了7.3%和9.2%。
摘要:Optimization of modern ASR architectures is among the highest priority tasks since it saves many computational resources for model training and inference. The work proposes a new Uconv-Conformer architecture based on the standard Conformer model that consistently reduces the input sequence length by 16 times, which results in speeding up the work of the intermediate layers. To solve the convergence problem with such a significant reduction of the time dimension, we use upsampling blocks similar to the U-Net architecture to ensure the correct CTC loss calculation and stabilize network training. The Uconv-Conformer architecture appears to be not only faster in terms of training and inference but also shows better WER compared to the baseline Conformer. Our best Uconv-Conformer model showed 40.3% epoch training time reduction, 47.8%, and 23.5% inference acceleration on the CPU and GPU, respectively. Relative WER on Librispeech test_clean and test_other decreased by 7.3% and 9.2%.


【8】 C3-DINO: Joint Contrastive and Non-contrastive Self-Supervised Learning  for Speaker Verification

标题:C3-DINO:基于对比和非对比自监督学习的说话人确认

链接:https://arxiv.org/abs/2208.07446

作者:Chunlei Zhang,Dong Yu
备注:Accepted to IEEE Journal of Selected Topics in Signal Processing
摘要:自监督学习(SSL)是近年来语音处理领域的一个研究热点,近年来的研究表明,对比学习能够以自监督的方式学习具有区分性的说话人嵌入,然而,基于对比自监督学习的说话人嵌入算法在语音处理领域的应用还不多见.(CSSL)假定从锚实例的视图和其它实例的任何视图生成的对都是负的,这一问题被称为“类—冲突”问题,一直是阻碍基于CSSL的说话人确认的主要问题之一(SV)系统性能的提高,研究表明无负样本的SSL框架在学习说话人或图像表示方面表现良好,本文首先分析了CSSL系统中假阴性对得影响,然后提出了一种多级类碰撞校正算法,该算法可以有效地提高类碰撞校正得性能.在CSSL模型的基础上,进一步提出了一种无负样本的SSL目标,该目标是一个基于CSSL的说话人嵌入系统(即DINO)来微调说话人嵌入网络。(C3-DINO)在Voxceleb1测试集上用简单的余弦距离评分方法获得2.5%的EER,性能优于以前的SOTA SSL系统在Voxceleb2训练集上进行说话人聚类和伪标记,LDA/CDS后端应用在C3-DINO扬声器嵌入上能够进一步将EER推到2.2%。在Voxceleb基准测试和我们内部数据集上的综合实验结果表明了我们所提方法的有效性,SSL SV与有监督的SSL SV之间的性能差距进一步缩小.
摘要:Self-supervised learning (SSL) has drawn an increased attention in the field of speech processing. Recent studies have demonstrated that contrastive learning is able to learn discriminative speaker embeddings in a self-supervised manner. However, base contrastive self-supervised learning (CSSL) assumes that the pairs generated from a view of anchor instance and any view of other instances are all negative, which introduces many false negative pairs in constructing the loss function. The problem is referred as $class$-$collision$, which remains as one major issue that impedes the CSSL based speaker verification (SV) systems from achieving better performances. In the meanwhile, studies reveal that negative sample free SSL frameworks perform well in learning speaker or image representations. In this study, we investigate SSL techniques that lead to an improved SV performance. We first analyse the impact of false negative pairs in the CSSL systems. Then, a multi-stage Class-Collision Correction (C3) method is proposed, which leads to the state-of-the-art CSSL based speaker embedding system. On the basis of the pretrained CSSL model, we further propose to employ a negative sample free SSL objective (i.e., DINO) to fine-tune the speaker embedding network. The resulting speaker embedding system (C3-DINO) achieves 2.5% EER with a simple Cosine Distance Scoring method on Voxceleb1 test set, which outperforms the previous SOTA SSL system (4.86%) by a significant +45% relative improvement. With speaker clustering and pseudo labeling on Voxceleb2 training set, a LDA/CDS back-end applying on the C3-DINO speaker embeddings is able to further push the EER to 2.2%. Comprehensive experimental investigations of the Voxceleb benchmarks and our internal dataset demonstrate the effectiveness of our proposed methods, and the performance gap between the SSL SV and the supervised counterpart narrows further.


eess.AS音频处理

【1】 Differentiable WORLD Synthesizer-based Neural Vocoder With Application  To End-To-End Audio Style Transfer

标题:基于可微分WORLD合成器的神经声码器及其在端到端音频风格传输中的应用

链接:https://arxiv.org/abs/2208.07282

* 与cs.SD语音【5】为同一篇

作者:Shahan Nercessian
机构:Zotope, Inc.
备注:11 pages, 2 figures
摘要:在本文中,我们提出了一种可区分WORLD合成器,并演示了其在端到端音频风格传输任务中的使用,例如(歌唱)语音转换和DDSP音色转换任务。因此,我们的基线可微分合成器没有模型参数,我们可以通过添加轻量级的黑色-一种可选择的可微分方法考虑直接提取源激发光谱,这可以改善自然度,尽管是针对较窄类别的风格转印应用。由我们的方法使用的声学特征参数化具有附加的益处,即其自然地解开音调和音色信息,使得它们可以被分别建模。此外,由于存在一种从单声道音频源估计这些声学特征的鲁棒装置,所以它允许将参数损失项添加到端到端目标函数,这可以帮助收敛和/或进一步稳定(对抗性)训练。
摘要:In this paper, we propose a differentiable WORLD synthesizer and demonstrate its use in end-to-end audio style transfer tasks such as (singing) voice conversion and the DDSP timbre transfer task. Accordingly, our baseline differentiable synthesizer has no model parameters, yet it yields adequate synthesis quality. We can extend the baseline synthesizer by appending lightweight black-box postnets which apply further processing to the baseline output in order to improve fidelity. An alternative differentiable approach considers extraction of the source excitation spectrum directly, which can improve naturalness albeit for a narrower class of style transfer applications. The acoustic feature parameterization used by our approaches has the added benefit that it naturally disentangles pitch and timbral information so that they can be modeled separately. Moreover, as there exists a robust means of estimating these acoustic features from monophonic audio sources, it allows for parameter loss terms to be added to an end-to-end objective function, which can help convergence and/or further stabilize (adversarial) training.


【2】 LCSM: A Lightweight Complex Spectral Mapping Framework for Stereophonic  Acoustic Echo Cancellation

标题:LCSM:一种用于立体声回声消除的轻量级复谱映射框架

链接:https://arxiv.org/abs/2208.07277

* 与cs.SD语音【1】为同一篇

作者:Chenggang Zhang,Jinjiang Liu,Xueliang Zhang
机构:Department of Computer Science, Inner Mongolia University, China
备注:Accepted to Interspeech 2022
摘要:传统的自适应算法在处理立体声回波抵消时会面临非唯一性问题本文首先提出了一种高效的多输入多输出(multi-input-multi-output,SAEC)算法(MIMO)方案来一次从所有麦克风信号中滤除回声。然后,我们采用一种轻量级的复杂频谱映射框架,(LCSM)方法实现端到端SAEC,无需对扬声器信号进行去相关预处理,利用现场卷积和通道空间建模来保证近端SAEC得准确性.实验结果表明,LCSM方法的泛化性能明显优于已有的方法,而且所提出的因果框架仅包含55万个参数,大大少于同类基于深度学习的方法,这对于资源有限的设备是重要的。
摘要:The traditional adaptive algorithms will face the non-uniqueness problem when dealing with stereophonic acoustic echo cancellation (SAEC). In this paper, we first propose an efficient multi-input and multi-output (MIMO) scheme based on deep learning to filter out echoes from all microphone signals at once. Then, we employ a lightweight complex spectral mapping framework (LCSM) for end-to-end SAEC without decorrelation preprocessing to the loudspeaker signals. Inplace convolution and channel-wise spatial modeling are utilized to ensure the near-end signal information is preserved. Finally, a cross-domain loss function is designed for better generalization capability. Experiments are evaluated on a variety of untrained conditions and results demonstrate that the LCSM significantly outperforms previous methods. Moreover, the proposed causal framework only has 0.55 million parameters, much less than the similar deep learning-based methods, which is important for the resource-limited devices.


【3】 Towards Parametric Speech Synthesis Using Gaussian-Markov Model of  Spectral Envelope and Wavelet-Based Decomposition of F0

标题:基于Gaussian-Markov谱包络模型和F0小波分解的参数语音合成

链接:https://arxiv.org/abs/2208.07122

* 与cs.SD语音【2】为同一篇

作者:Mohammed Salah Al-Radhi,Tamás Gábor Csapó,Csaba Zainkó,Géza Németh
机构:Department of Telecommunications and Media Informatics, Budapest University of Technology and Economics, Budapest, Hungary
备注:accepted at EUSIPCO2022
摘要:基于神经网络的文语转换技术显著提高了合成语音的质量。(例如,Tacotron2、FastSpeech、FastPitch)通常从文本生成Mel谱图,然后使用声码器合成语音(例如,WaveNet、WaveGlow、HiFiGAN)。与传统的参数方法相比(如STRAIGHT和WORLD),基于神经网络的端到端模型存在推理速度慢,合成语音鲁棒性差和可控性差的缺点,本文提出了一种新的更新声码器,该模型训练简单,波形生成容易,采用高斯—马尔可夫模型对频谱包络进行鲁棒学习,利用基于小波的统计信号处理对F0特征进行表征和分解,既能保持良好的频谱包络,又能实现自然语音的高可控性。实验结果表明,本文提出的声码器比传统的STRAIGHT声码器具有更好的重构语音自然度,略优于WaveNet,并且比WaveRNN稍差。
摘要:Neural network-based Text-to-Speech has significantly improved the quality of synthesized speech. Prominent methods (e.g., Tacotron2, FastSpeech, FastPitch) usually generate Mel-spectrogram from text and then synthesize speech using vocoder (e.g., WaveNet, WaveGlow, HiFiGAN). Compared with traditional parametric approaches (e.g., STRAIGHT and WORLD), neural vocoder based end-to-end models suffer from slow inference speed, and the synthesized speech is usually not robust and lack of controllability. In this work, we propose a novel updated vocoder, which is a simple signal model to train and easy to generate waveforms. We use the Gaussian-Markov model toward robust learning of spectral envelope and wavelet-based statistical signal processing to characterize and decompose F0 features. It can retain the fine spectral envelope and achieve high controllability of natural speech. The experimental results demonstrate that our proposed vocoder achieves better naturalness of reconstructed speech than the conventional STRAIGHT vocoder, slightly better than WaveNet, and somewhat worse than the WaveRNN.


【4】 Analysis of impact of emotions on target speech extraction and speech  separation

标题:情感对目标语音提取和语音分离的影响分析

链接:https://arxiv.org/abs/2208.07091

* 与cs.SD语音【3】为同一篇

作者:Ján Švec,Kateřina Žmolíková,Martin Kocour,Marc Delcroix,Tsubasa Ochiai,Ladislav Mošner,Jan Černocký
机构:Ladislav Moˇsner, Jan “Honza” ˇCernock´y, Brno University of Technology, IT,I Centre of Excellence, NTT Corporation, Japan
备注:Accepted to IWAENC 2022
摘要:近年来,盲语音分离的性能(BSS)和目标语音提取(TSE)取得了很大的进展.然而,大多数工作集中在相对良好控制的条件下,例如使用朗读语音.在更真实的情况下,性能可能会下降.造成这种下降的因素之一可能是内在的说话者可变性,如情感,这在真实语音中是常见的.本文在分析了语音的可变性的基础上,提出了一种基于语音可变性的语音识别方法,我们研究了情感对TSE和BSS的影响,并结合LibriSpeech和Ryerson的情感语音和歌曲视听数据库,建立了一个新的情感混合测试数据集,用于TSE和BSS的评价通过对照实验,我们可以分析不同的情绪对BSS和TSE性能的影响,我们观察到BSS对情绪相对鲁棒,而TSE,该方法需要识别和提取目标说话人的语音,对情感更加敏感,在对比说话人确认实验中,我们发现识别目标说话人在处理情感语音时尤其困难,我们概述了可能提高BSS和TSE系统对情感语音的鲁棒性的潜在未来方向。
摘要:Recently, the performance of blind speech separation (BSS) and target speech extraction (TSE) has greatly progressed. Most works, however, focus on relatively well-controlled conditions using, e.g., read speech. The performance may degrade in more realistic situations. One of the factors causing such degradation may be intrinsic speaker variability, such as emotions, occurring commonly in realistic speech. In this paper, we investigate the influence of emotions on TSE and BSS. We create a new test dataset of emotional mixtures for the evaluation of TSE and BSS. This dataset combines LibriSpeech and Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS). Through controlled experiments, we can analyze the impact of different emotions on the performance of BSS and TSE. We observe that BSS is relatively robust to emotions, while TSE, which requires identifying and extracting the speech of a target speaker, is much more sensitive to emotions. On comparative speaker verification experiments we show that identifying the target speaker may be particularly challenging when dealing with emotional speech. Using our findings, we outline potential future directions that could improve the robustness of BSS and TSE systems toward emotional speech.


【5】 Models of Music Cognition and Composition

标题:音乐认知与创作模式

链接:https://arxiv.org/abs/2208.06878

* 与cs.SD语音【4】为同一篇

作者:Abhimanyu Sethia,Aayush
机构:Indian Institute of Technology, Kanpur
备注:TLDR: literature review of models of music cognition and composition
摘要:与大多数认知研究一样,音乐认知也是一个跨学科领域,它试图应用认知科学的方法(神经学的,计算的和实验的)来理解音乐创作的感知和过程。本文,我们首先提出为什么音乐与认知科学家相关,并概述音乐认知的计算建模方法。然后,我们回顾了关于音乐感知的各种模型的文献,包括非计算模型,计算非认知模型和计算认知模型。最后,我们回顾了关于创造性行为建模和能够创作音乐的计算机系统的文献。由于大量来自音乐理论的专业术语被使用,我们在最后附上了相关术语及其定义的列表。
摘要:Much like most of cognition research, music cognition is an interdisciplinary field, which attempts to apply methods of cognitive science (neurological, computational and experimental) to understand the perception and process of composition of music. In this paper, we first motivate why music is relevant to cognitive scientists and give an overview of the approaches to computational modelling of music cognition. We then review literature on the various models of music perception, including non-computational models, computational non-cognitive models and computational cognitive models. Lastly, we review literature on modelling the creative behaviour and on computer systems capable of composing music. Since a lot of technical terms from music theory have been used, we have appended a list of relevant terms and their definitions at the end.


【6】 Uconv-Conformer: High Reduction of Input Sequence Length for End-to-End  Speech Recognition

标题:Uconv-符合者:用于端到端语音识别的输入序列长度的高度缩减

链接:https://arxiv.org/abs/2208.07657

* 与cs.SD语音【7】为同一篇

作者:Andrei Andrusenko,Rauf Nasretdinov,Aleksei Romanenko
机构:ITMO University, St. Petersburg, Russia, STC-innovations Ltd, St. Petersburg, Russia
备注:5 pages, 1 figure
摘要:本文在标准Conformer模型的基础上提出了一种新的Uconv-Conformer结构,该结构将输入序列长度一致地减少了16倍,为了解决在时间维数显著降低的情况下的收敛问题,我们使用类似于U-Net架构的上采样模块,以确保正确的CTC丢失计算和稳定的网络训练。Uconv-Conformer架构似乎不仅在训练和推理方面更快,而且与基线Conformer相比显示出更好的WER。我们的最佳Uconv-Conformer模型显示出40.3%的历元训练时间减少,47.8%,和GPU上的推理加速分别降低了23.5%和23.5%,Librispeech test_clean和test_other上的相对WER分别降低了7.3%和9.2%。
摘要:Optimization of modern ASR architectures is among the highest priority tasks since it saves many computational resources for model training and inference. The work proposes a new Uconv-Conformer architecture based on the standard Conformer model that consistently reduces the input sequence length by 16 times, which results in speeding up the work of the intermediate layers. To solve the convergence problem with such a significant reduction of the time dimension, we use upsampling blocks similar to the U-Net architecture to ensure the correct CTC loss calculation and stabilize network training. The Uconv-Conformer architecture appears to be not only faster in terms of training and inference but also shows better WER compared to the baseline Conformer. Our best Uconv-Conformer model showed 40.3% epoch training time reduction, 47.8%, and 23.5% inference acceleration on the CPU and GPU, respectively. Relative WER on Librispeech test_clean and test_other decreased by 7.3% and 9.2%.


【7】 C3-DINO: Joint Contrastive and Non-contrastive Self-Supervised Learning  for Speaker Verification

标题:C3-DINO:基于对比和非对比自监督学习的说话人确认

链接:https://arxiv.org/abs/2208.07446

* 与cs.SD语音【8】为同一篇

作者:Chunlei Zhang,Dong Yu
备注:Accepted to IEEE Journal of Selected Topics in Signal Processing
摘要:自监督学习(SSL)是近年来语音处理领域的一个研究热点,近年来的研究表明,对比学习能够以自监督的方式学习具有区分性的说话人嵌入,然而,基于对比自监督学习的说话人嵌入算法在语音处理领域的应用还不多见.(CSSL)假定从锚实例的视图和其它实例的任何视图生成的对都是负的,这一问题被称为“类—冲突”问题,一直是阻碍基于CSSL的说话人确认的主要问题之一(SV)系统性能的提高,研究表明无负样本的SSL框架在学习说话人或图像表示方面表现良好,本文首先分析了CSSL系统中假阴性对得影响,然后提出了一种多级类碰撞校正算法,该算法可以有效地提高类碰撞校正得性能.在CSSL模型的基础上,进一步提出了一种无负样本的SSL目标,该目标是一个基于CSSL的说话人嵌入系统(即DINO)来微调说话人嵌入网络。(C3-DINO)在Voxceleb1测试集上用简单的余弦距离评分方法获得2.5%的EER,性能优于以前的SOTA SSL系统在Voxceleb2训练集上进行说话人聚类和伪标记,LDA/CDS后端应用在C3-DINO扬声器嵌入上能够进一步将EER推到2.2%。在Voxceleb基准测试和我们内部数据集上的综合实验结果表明了我们所提方法的有效性,SSL SV与有监督的SSL SV之间的性能差距进一步缩小.
摘要:Self-supervised learning (SSL) has drawn an increased attention in the field of speech processing. Recent studies have demonstrated that contrastive learning is able to learn discriminative speaker embeddings in a self-supervised manner. However, base contrastive self-supervised learning (CSSL) assumes that the pairs generated from a view of anchor instance and any view of other instances are all negative, which introduces many false negative pairs in constructing the loss function. The problem is referred as $class$-$collision$, which remains as one major issue that impedes the CSSL based speaker verification (SV) systems from achieving better performances. In the meanwhile, studies reveal that negative sample free SSL frameworks perform well in learning speaker or image representations. In this study, we investigate SSL techniques that lead to an improved SV performance. We first analyse the impact of false negative pairs in the CSSL systems. Then, a multi-stage Class-Collision Correction (C3) method is proposed, which leads to the state-of-the-art CSSL based speaker embedding system. On the basis of the pretrained CSSL model, we further propose to employ a negative sample free SSL objective (i.e., DINO) to fine-tune the speaker embedding network. The resulting speaker embedding system (C3-DINO) achieves 2.5% EER with a simple Cosine Distance Scoring method on Voxceleb1 test set, which outperforms the previous SOTA SSL system (4.86%) by a significant +45% relative improvement. With speaker clustering and pseudo labeling on Voxceleb2 training set, a LDA/CDS back-end applying on the C3-DINO speaker embeddings is able to further push the EER to 2.2%. Comprehensive experimental investigations of the Voxceleb benchmarks and our internal dataset demonstrate the effectiveness of our proposed methods, and the performance gap between the SSL SV and the supervised counterpart narrows further.


【8】 How Should We Evaluate Synthesized Environmental Sounds

标题:如何评价合成环境声

链接:https://arxiv.org/abs/2208.07679

* 与cs.SD语音【6】为同一篇

作者:Yuki Okamoto,Keisuke Imoto,Shinnosuke Takamichi,Takahiro Fukumori,Yoichi Yamashita
机构:∗ Ritsumeikan University, Japan, † Doshisha University, Japan, ‡ The University of Tokyo, Japan备注:Submitted APSIPA ASC 2022
摘要:环境声的合成方法有多种,但对合成后的环境声如何进行评价却没有讨论,传统的评价方法只进行主观评价或客观评价,对哪种评价方法不明确,本文提出了一种新的评价方法———环境声综合评价法,我们研究如何评价合成环境声音。我们还提出了一种主观评价方法来评价合成的声音是否适当地代表了输入到环境中的信息在实验中,我们将所提出的评价方法与传统的评价方法进行了比较,结果表明主观评价的结果往往与客观评价的结果不同,从这些结果中我们得出结论,不仅要进行客观评价,而且要进行主观评价。
摘要:Although several methods of environmental sound synthesis have been proposed, there has been no discussion on how synthesized environmental sounds should be evaluated. Only either subjective or objective evaluations have been conducted in conventional evaluations, and it is not clear what type of evaluation should be carried out. In this paper, we investigate how to evaluate synthesized environmental sounds. We also propose a subjective evaluation methodology to evaluate whether the synthesized sound appropriately represents the information input to the environmental sound synthesis system. In our experiments, we compare the proposed and conventional evaluation methods and show that the results of subjective evaluations tended to differ from those of objective evaluations. From these results, we conclude that it is necessary to conduct not only objective evaluation but also subjective evaluation.


机器翻译,仅供参考