今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily


cs.SD语音

【1】 On Compressing Sequences for Self-Supervised Speech Models

标题:关于自监督语音模型的序列压缩

链接:https://arxiv.org/abs/2210.07189

作者:Yen Meng,Hsuan-Jui Chen,Jiatong Shi,Shinji Watanabe,Paola Garcia,Hung-yi Lee,Hao Tang
机构:National Taiwan University,Carnegie Mellon University,Johns Hopkins University, The University of Edinburgh
备注:Accepted to IEEE SLT 2022
摘要:随着自监督模型变得越来越大,压缩自监督模型变得越来越有必要。虽然先前的方法主要集中于压缩模型大小,但是缩短序列在降低计算成本方面也是有效的。本文研究了自监督学习中沿时间轴的定长和变长子采样。我们探讨了单个下游任务对输入帧速率的敏感性。在训练自监督模型的同时进行子采样,不仅可以在一定的帧速率下提高下游任务的整体性能,而且可以显著提高推理速度。可变长度子采样在低帧速率下表现得特别好。此外,如果我们可以访问语音边界,我们发现对于低至10Hz的平均帧速率,性能没有下降。
摘要:Compressing self-supervised models has become increasingly necessary, as self-supervised models become larger. While previous approaches have primarily focused on compressing the model size, shortening sequences is also effective in reducing the computational cost. In this work, we study fixed-length and variable-length subsampling along the time axis in self-supervised learning. We explore how individual downstream tasks are sensitive to input frame rates. Subsampling while training self-supervised models not only improves the overall performance on downstream tasks under certain frame rates, but also brings significant speed-up in inference. Variable-length subsampling performs particularly well under low frame rates. In addition, if we have access to phonetic boundaries, we find no degradation in performance for an average frame rate as low as 10 Hz.


【2】 Sparse in Space and Time: Audio-visual Synchronisation with Trainable  Selectors

标题:时空稀疏:与可训练选择器的视听同步

链接:https://arxiv.org/abs/2210.07055

作者:Vladimir Iashin,Weidi Xie,Esa Rahtu,Andrew Zisserman
机构:Computing Sciences, Tampere University, Tampere, Finland,  Coop. Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, China,  Visual Geometry Group, Department of Engineering Science, University of Oxford, Oxford, UK
备注:Accepted as a spotlight presentation for the BMVC 2022. Code: this https URL Project page: this https URL
摘要:本文的目标是“在野外”的一般视频的视听同步。对于这样的视频,可以被利用用于同步提示的事件可以在空间上很小,并且可以在许多秒长的视频剪辑期间仅偶尔发生,即同步信号在空间和时间上是稀疏的。这与同步会说话的人的视频的情况形成了对比,在这种情况下,视听通信在时间和空间上都是密集的。  我们做出了四项贡献:(i)为了处理稀疏同步信号所需的较长时间序列,我们设计了一种多模态变换器模型,该模型采用“选择器”将长音频和视频流提取成小序列,然后使用这些小序列来预测流之间的时间偏移。(ii)我们识别可能由用于音频和视频的压缩编解码器引起的伪像,并且可以由训练中的视听模型使用以人工解决同步任务。(iii)我们整理了一个时间和空间同步信号稀疏的数据集;以及(iv)在密集和稀疏数据集上定量和定性地显示了所提出的模型的有效性。  项目页面:v-iashin.github.io/SparseSync
摘要:The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spatially small and may occur only infrequently during a many seconds-long video clip, i.e. the synchronisation signal is 'sparse in space and time'. This contrasts with the case of synchronising videos of talking heads, where audio-visual correspondence is dense in both time and space.  We make four contributions: (i) in order to handle longer temporal sequences required for sparse synchronisation signals, we design a multi-modal transformer model that employs 'selectors' to distil the long audio and visual streams into small sequences that are then used to predict the temporal offset between streams. (ii) We identify artefacts that can arise from the compression codecs used for audio and video and can be used by audio-visual models in training to artificially solve the synchronisation task. (iii) We curate a dataset with only sparse in time and space synchronisation signals; and (iv) the effectiveness of the proposed model is shown on both dense and sparse datasets quantitatively and qualitatively.  Project page: v-iashin.github.io/SparseSync


【3】 Anonymizing Speech with Generative Adversarial Networks to Preserve  Speaker Privacy

标题:利用产生式对抗网络保护说话人隐私的语音匿名

链接:https://arxiv.org/abs/2210.07002

作者:Sarina Meyer,Pascal Tilli,Pavel Denisov,Florian Lux,Julia Koch,Ngoc Thang Vu
机构:Institute for Natural Language Processing (IMS), University of Stuttgart, Germany
备注:IEEE Spoken Language Technology Workshop 2022
摘要:为了保护语音数据的隐私性,说话人匿名化通过改变语音记录中的声音来隐藏说话人的身份。这通常伴随着在保护个人和数据对下游应用程序的可用性之间的隐私-效用权衡。在这方面的挑战之一是创造不存在的声音听起来尽可能自然。  在本文中,我们提出以Wasserstein距离为代价函数,使用生成式对抗网络来产生说话人嵌入,以解决这个问题。通过将这些人工嵌入到语音到文本到语音的流水线中,我们在隐私和实用性方面优于先前的方法。根据标准的客观度量和人工评估,我们的方法生成原始记录的可理解的、内容保持但隐私保护的版本。
摘要:In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes with a privacy-utility trade-off between protection of individuals and usability of the data for downstream applications. One of the challenges in this context is to create non-existent voices that sound as natural as possible.  In this work, we propose to tackle this issue by generating speaker embeddings using a generative adversarial network with Wasserstein distance as cost function. By incorporating these artificial embeddings into a speech-to-text-to-speech pipeline, we outperform previous approaches in terms of privacy and utility. According to standard objective metrics and human evaluation, our approach generates intelligible and content-preserving yet privacy-protecting versions of the original recordings.


【4】 Multilingual Zero Resource Speech Recognition Base on Self-Supervise  Pre-Trained Acoustic Models

标题:基于自监督预训练声学模型的多语种零资源语音识别

链接:https://arxiv.org/abs/2210.06936

作者:Haoyu Wang,Wei-Qiang Zhang,Hongbin Suo,Yulong Wan
机构:† Beijing National Research Center for Information Science and Technology, Department of Electronic Engineering, Tsinghua University, Beijing , China, ‡ Data & AI Engineering System, OPPO, Beijing , China
备注:accepted by ISCSLP 2022
摘要:标记的音频数据不足以为世界上的大多数语言建立令人满意的语音识别系统。已经有一些零资源方法试图在没有目标语言的标记音频数据的情况下执行音素或词级语音识别,但是这些方法的错误率通常太高而不能应用于现实世界场景。近年来,人们发现,自监督预训练模型的表示能力在零资源音素识别中非常有益。就我们而言,本文是第一次尝试将预训练模型的使用扩展到词级零资源语音识别。这是通过微调关于IPA音素转录的预先训练的模型和用关于额外文本训练的语言模型解码来完成的。在Wav2vec2.0和HuBERT模型上的实验结果表明,该方法对部分语种的分词错误率小于20%,对8种语种的平均分词错误率为33.77%。
摘要:Labeled audio data is insufficient to build satisfying speech recognition systems for most of the languages in the world. There have been some zero-resource methods trying to perform phoneme or word-level speech recognition without labeled audio data of the target language, but the error rate of these methods is usually too high to be applied in real-world scenarios. Recently, the representation ability of self-supervise pre-trained models has been found to be extremely beneficial in zero-resource phoneme recognition. As far as we are concerned, this paper is the first attempt to extend the use of pre-trained models into word-level zero-resource speech recognition. This is done by fine-tuning the pre-trained models on IPA phoneme transcriptions and decoding with a language model trained on extra texts. Experiments on Wav2vec 2.0 and HuBERT models show that this method can achieve less than 20% word error rate on some languages, and the average error rate on 8 languages is 33.77%.


【5】 Inner speech recognition through electroencephalographic signals

标题:基于脑电信号的内部语音识别

链接:https://arxiv.org/abs/2210.06472

作者:Francesca Gasparini,Elisa Cazzaniga,Aurora Saibene
机构:Saibene,[,−,−,−,],  University of Milano-Bicocca, Viale Sarca , Milano, Italy,  NeuroMI, Milan Center for Neuroscience, Piazza dell’Ateneo Nuovo
备注:Submitted to the Italian Workshop on Artificial Intelligence for Human Machine Interaction (AIxHMI 2022), December 02, 2022, Udine, Italy
摘要:本文主要研究从脑电信号出发的内部语音识别。内部言语识别是指人在纯粹意义上思考的内化过程,通常与自己内心“声音”的听觉意象相关联。将EEG解码成文本应该理解为对有限数量的单词(命令)或音素(组成单词的声音单位)的存在进行分类。语音相关的BCI提供了有效的语音通信策略,用于通过从脑信号解释的语音命令来控制设备,通过恢复与他们的环境的通信来改善失去说话能力的人的生活质量。分析了两个公开的内部语音数据集。利用这些数据,研究并实现了一些分类模型,从基本方法(如支持向量机)到集成方法(如极限梯度提升分类器),再到神经网络(如长短时记忆(LSTM)和双向长短时记忆(BiLSTM))的使用。利用LSTM和BiLSTM模型,获得了与现有技术中存在的结果一致或优于现有技术中存在的结果的结果,所述LSTM和BiLSTM模型通常未在内部语音识别的文献中使用。
摘要:This work focuses on inner speech recognition starting from EEG signals. Inner speech recognition is defined as the internalized process in which the person thinks in pure meanings, generally associated with an auditory imagery of own inner "voice". The decoding of the EEG into text should be understood as the classification of a limited number of words (commands) or the presence of phonemes (units of sound that make up words). Speech-related BCIs provide effective vocal communication strategies for controlling devices through speech commands interpreted from brain signals, improving the quality of life of people who have lost the capability to speak, by restoring communication with their environment. Two public inner speech datasets are analysed. Using this data, some classification models are studied and implemented starting from basic methods such as Support Vector Machines, to ensemble methods such as the eXtreme Gradient Boosting classifier up to the use of neural networks such as Long Short Term Memory (LSTM) and Bidirectional Long Short Term Memory (BiLSTM). With the LSTM and BiLSTM models, generally not used in the literature of inner speech recognition, results in line with or superior to those present in the stateof-the-art are obtained.


【6】 Deepfake Detection System for the ADD Challenge Track 3.2 Based on Score  Fusion

标题:基于分数融合的ADD挑战赛3.2赛道深伪检测系统

链接:https://arxiv.org/abs/2210.06818

作者:Yuxiang Zhang,Jingze Lu,Xingming Wang,Zhuo Li,Runqiu Xiao,Wenchao Wang,Ming Li,Pengyuan Zhang
机构:Institute of Acoustics, Chinese, Academy of Sciences, Haidian District, Beijing, China, University of Chinese Academy of, Shijingshan District, Beijing, China, Data Science Research Center, Duke, Kunshan University, Kunshan, Jiangsu Province, China
备注:Accepted by ACM Multimedia 2022 Workshop: First International Workshop on Deepfake Detection for Audio Multimedia
摘要:本文描述了提交给音频深度合成检测(ADD)挑战赛轨道3.2的deepfake音频检测系统,并给出了分数融合的分析。所提出的系统是几个基于轻卷积神经网络(LCNN)的模型的分数级融合。各种前端被用作输入特征,包括低频短时傅里叶变换和常数Q变换。由于复杂的噪声和丰富的合成算法,直接利用训练集很难获得理想的性能。在线数据增强方法有效地提高了虚假音频检测系统的鲁棒性。特别地,通过分数分布的可视化以及与另一数据集上的分数分布的比较,探索了分数融合改进较差的原因。模型对训练集的过拟合导致得分的极值和得分分布的低相关性,这使得得分融合困难。与部分伪造音频检测系统的融合进一步提高了系统性能。轨道3.2上的提交获得了11.04%的加权等误率(WEER),这是挑战中表现最好的系统之一。
摘要:This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light convolutional neural network (LCNN) based models. Various front-ends are used as input features, including low-frequency short-time Fourier transform and Constant Q transform. Due to the complex noise and rich synthesis algorithms, it is difficult to obtain the desired performance using the training set directly. Online data augmentation methods effectively improve the robustness of fake audio detection systems. In particular, the reasons for the poor improvement of score fusion are explored through visualization of the score distributions and comparison with score distribution on another dataset. The overfitting of the model to the training set leads to extreme values of the scores and low correlation of the score distributions, which makes score fusion difficult. Fusion with partially fake audio detection system improves system performance further. The submission on track 3.2 obtained the weighted equal error rate (WEER) of 11.04\%, which is one of the best performing systems in the challenge.


【7】 An Analysis Method for Metric-Level Switching in Beat Tracking

标题:节拍跟踪中度量级切换的一种分析方法

链接:https://arxiv.org/abs/2210.06817

作者:Ching-Yu Chiu,Meinard Müller,Matthew E. P. Davies,Alvin Wen-Yu Su,Yi-Hsuan Yang
机构:Davies is with the Department of Informatics Engineering, Centre for Informatics and Systems of the University of Coimbra,  Universityof Coimbra
备注:Accepted to IEEE Signal Processing Letters (Oct. 2022)
摘要:对于富有表现力的音乐,节奏可能随时间而改变,这对通过自动模型跟踪节拍提出了挑战。模型可以首先轻敲到正确的节奏,但是随后可能无法适应节奏变化,或者在几个不正确但是感觉上似乎合理的节奏之间切换(例如,半或双倍速度)。现有的用于心跳跟踪的评估度量不反映这样的行为,因为它们通常假设参考心跳和估计心跳之间的固定关系。本文提出了一种新的性能分析方法--标注覆盖率(ACR),该方法考虑了节拍跟踪器的各种可能的度量级切换行为。其思想是针对每两个连续的参考搏动导出所有度量水平的修改的参考搏动的序列,并且将修改的参考搏动的每个序列与估计搏动的子序列进行比较。我们通过在三个不同类型的数据集上的实验,展示了ACR在与现有度量一起使用时的有用性,并讨论了将获得的新见解。
摘要:For expressive music, the tempo may change over time, posing challenges to tracking the beats by an automatic model. The model may first tap to the correct tempo, but then may fail to adapt to a tempo change, or switch between several incorrect but perceptually plausible ones (e.g., half- or double-tempo). Existing evaluation metrics for beat tracking do not reflect such behaviors, as they typically assume a fixed relationship between the reference beats and estimated beats. In this paper, we propose a new performance analysis method, called annotation coverage ratio (ACR), that accounts for a variety of possible metric-level switching behaviors of beat trackers. The idea is to derive sequences of modified reference beats of all metrical levels for every two consecutive reference beats, and compare every sequence of modified reference beats to the subsequences of estimated beats. We show via experiments on three datasets of different genres the usefulness of ACR when utilized alongside existing metrics, and discuss the new insights to be gained.


eess.AS音频处理

【1】 Deepfake Detection System for the ADD Challenge Track 3.2 Based on Score  Fusion

标题:基于分数融合的ADD挑战赛3.2赛道深伪检测系统

链接:https://arxiv.org/abs/2210.06818

* 与cs.SD语音【6】为同一篇

作者:Yuxiang Zhang,Jingze Lu,Xingming Wang,Zhuo Li,Runqiu Xiao,Wenchao Wang,Ming Li,Pengyuan Zhang
机构:Institute of Acoustics, Chinese, Academy of Sciences, Haidian District, Beijing, China, University of Chinese Academy of, Shijingshan District, Beijing, China, Data Science Research Center, Duke, Kunshan University, Kunshan, Jiangsu Province, China
备注:Accepted by ACM Multimedia 2022 Workshop: First International Workshop on Deepfake Detection for Audio Multimedia
摘要:本文描述了提交给音频深度合成检测(ADD)挑战赛轨道3.2的deepfake音频检测系统,并给出了分数融合的分析。所提出的系统是几个基于轻卷积神经网络(LCNN)的模型的分数级融合。各种前端被用作输入特征,包括低频短时傅里叶变换和常数Q变换。由于复杂的噪声和丰富的合成算法,直接利用训练集很难获得理想的性能。在线数据增强方法有效地提高了虚假音频检测系统的鲁棒性。特别地,通过分数分布的可视化以及与另一数据集上的分数分布的比较,探索了分数融合改进较差的原因。模型对训练集的过拟合导致得分的极值和得分分布的低相关性,这使得得分融合困难。与部分伪造音频检测系统的融合进一步提高了系统性能。轨道3.2上的提交获得了11.04%的加权等误率(WEER),这是挑战中表现最好的系统之一。
摘要:This paper describes the deepfake audio detection system submitted to the Audio Deep Synthesis Detection (ADD) Challenge Track 3.2 and gives an analysis of score fusion. The proposed system is a score-level fusion of several light convolutional neural network (LCNN) based models. Various front-ends are used as input features, including low-frequency short-time Fourier transform and Constant Q transform. Due to the complex noise and rich synthesis algorithms, it is difficult to obtain the desired performance using the training set directly. Online data augmentation methods effectively improve the robustness of fake audio detection systems. In particular, the reasons for the poor improvement of score fusion are explored through visualization of the score distributions and comparison with score distribution on another dataset. The overfitting of the model to the training set leads to extreme values of the scores and low correlation of the score distributions, which makes score fusion difficult. Fusion with partially fake audio detection system improves system performance further. The submission on track 3.2 obtained the weighted equal error rate (WEER) of 11.04\%, which is one of the best performing systems in the challenge.


【2】 An Analysis Method for Metric-Level Switching in Beat Tracking

标题:节拍跟踪中度量级切换的一种分析方法

链接:https://arxiv.org/abs/2210.06817

* 与cs.SD语音【7】为同一篇

作者:Ching-Yu Chiu,Meinard Müller,Matthew E. P. Davies,Alvin Wen-Yu Su,Yi-Hsuan Yang
机构:Davies is with the Department of Informatics Engineering, Centre for Informatics and Systems of the University of Coimbra,  Universityof Coimbra
备注:Accepted to IEEE Signal Processing Letters (Oct. 2022)
摘要:对于富有表现力的音乐,节奏可能随时间而改变,这对通过自动模型跟踪节拍提出了挑战。模型可以首先轻敲到正确的节奏,但是随后可能无法适应节奏变化,或者在几个不正确但是感觉上似乎合理的节奏之间切换(例如,半或双倍速度)。现有的用于心跳跟踪的评估度量不反映这样的行为,因为它们通常假设参考心跳和估计心跳之间的固定关系。本文提出了一种新的性能分析方法--标注覆盖率(ACR),该方法考虑了节拍跟踪器的各种可能的度量级切换行为。其思想是针对每两个连续的参考搏动导出所有度量水平的修改的参考搏动的序列,并且将修改的参考搏动的每个序列与估计搏动的子序列进行比较。我们通过在三个不同类型的数据集上的实验,展示了ACR在与现有度量一起使用时的有用性,并讨论了将获得的新见解。
摘要:For expressive music, the tempo may change over time, posing challenges to tracking the beats by an automatic model. The model may first tap to the correct tempo, but then may fail to adapt to a tempo change, or switch between several incorrect but perceptually plausible ones (e.g., half- or double-tempo). Existing evaluation metrics for beat tracking do not reflect such behaviors, as they typically assume a fixed relationship between the reference beats and estimated beats. In this paper, we propose a new performance analysis method, called annotation coverage ratio (ACR), that accounts for a variety of possible metric-level switching behaviors of beat trackers. The idea is to derive sequences of modified reference beats of all metrical levels for every two consecutive reference beats, and compare every sequence of modified reference beats to the subsequences of estimated beats. We show via experiments on three datasets of different genres the usefulness of ACR when utilized alongside existing metrics, and discuss the new insights to be gained.


【3】 On Compressing Sequences for Self-Supervised Speech Models

标题:关于自监督语音模型的序列压缩

链接:https://arxiv.org/abs/2210.07189

* 与cs.SD语音【1】为同一篇

作者:Yen Meng,Hsuan-Jui Chen,Jiatong Shi,Shinji Watanabe,Paola Garcia,Hung-yi Lee,Hao Tang
机构:National Taiwan University,Carnegie Mellon University,Johns Hopkins University, The University of Edinburgh
备注:Accepted to IEEE SLT 2022
摘要:随着自监督模型变得越来越大,压缩自监督模型变得越来越有必要。虽然先前的方法主要集中于压缩模型大小,但是缩短序列在降低计算成本方面也是有效的。本文研究了自监督学习中沿时间轴的定长和变长子采样。我们探讨了单个下游任务对输入帧速率的敏感性。在训练自监督模型的同时进行子采样,不仅可以在一定的帧速率下提高下游任务的整体性能,而且可以显著提高推理速度。可变长度子采样在低帧速率下表现得特别好。此外,如果我们可以访问语音边界,我们发现对于低至10Hz的平均帧速率,性能没有下降。
摘要:Compressing self-supervised models has become increasingly necessary, as self-supervised models become larger. While previous approaches have primarily focused on compressing the model size, shortening sequences is also effective in reducing the computational cost. In this work, we study fixed-length and variable-length subsampling along the time axis in self-supervised learning. We explore how individual downstream tasks are sensitive to input frame rates. Subsampling while training self-supervised models not only improves the overall performance on downstream tasks under certain frame rates, but also brings significant speed-up in inference. Variable-length subsampling performs particularly well under low frame rates. In addition, if we have access to phonetic boundaries, we find no degradation in performance for an average frame rate as low as 10 Hz.


【4】 On the Utility of Self-supervised Models for Prosody-related Tasks

标题:自我监督模型在韵律相关任务中的应用研究

链接:https://arxiv.org/abs/2210.07185

作者:Guan-Ting Lin,Chi-Luen Feng,Wei-Ping Huang,Yuan Tseng,Tzu-Han Lin,Chen-An Li,Hung-yi Lee,Nigel G. Ward
机构:National Taiwan University, Taiwan, University of Texas at El Paso, USA
备注:Accepted to IEEE SLT 2022
摘要:来自语音数据的自监督学习(SSL)已经产生了在许多任务中实现了显著性能的模型,并且已知这些模型隐含地表示潜在地存在于语音信号中的信息的许多方面。然而,关于这些模型对韵律相关任务的适用性或它们编码韵律信息的程度,相对知之甚少。本文提出了一个新的评价框架SUPERB-韵律,它由三个与韵律相关的下游任务和两个伪任务组成。我们发现,15个SSL模型中有13个模型在所有韵律相关任务上的表现都优于基线。我们还在两个伪任务上表现出良好的性能:韵律重建和未来韵律预测。我们进一步分析了SSL模型的分层贡献。总体而言,我们得出结论,SSL语音模型是非常有效的韵律相关的任务。
摘要:Self-Supervised Learning (SSL) from speech data has produced models that have achieved remarkable performance in many tasks, and that are known to implicitly represent many aspects of information latently present in speech signals. However, relatively little is known about the suitability of such models for prosody-related tasks or the extent to which they encode prosodic information. We present a new evaluation framework, SUPERB-prosody, consisting of three prosody-related downstream tasks and two pseudo tasks. We find that 13 of the 15 SSL models outperformed the baseline on all the prosody-related tasks. We also show good performance on two pseudo tasks: prosody reconstruction and future prosody prediction. We further analyze the layerwise contributions of the SSL models. Overall we conclude that SSL speech models are highly effective for prosody-related tasks.


【5】 Sparse in Space and Time: Audio-visual Synchronisation with Trainable  Selectors

标题:时空稀疏:与可训练选择器的视听同步

链接:https://arxiv.org/abs/2210.07055

* 与cs.SD语音【2】为同一篇

作者:Vladimir Iashin,Weidi Xie,Esa Rahtu,Andrew Zisserman
机构:Computing Sciences, Tampere University, Tampere, Finland,  Coop. Medianet Innovation Center, Shanghai Jiao Tong University, Shanghai, China,  Visual Geometry Group, Department of Engineering Science, University of Oxford, Oxford, UK
备注:Accepted as a spotlight presentation for the BMVC 2022. Code: this https URL Project page: this https URL
摘要:本文的目标是“在野外”的一般视频的视听同步。对于这样的视频,可以被利用用于同步提示的事件可以在空间上很小,并且可以在许多秒长的视频剪辑期间仅偶尔发生,即同步信号在空间和时间上是稀疏的。这与同步会说话的人的视频的情况形成了对比,在这种情况下,视听通信在时间和空间上都是密集的。  我们做出了四项贡献:(i)为了处理稀疏同步信号所需的较长时间序列,我们设计了一种多模态变换器模型,该模型采用“选择器”将长音频和视频流提取成小序列,然后使用这些小序列来预测流之间的时间偏移。(ii)我们识别可能由用于音频和视频的压缩编解码器引起的伪像,并且可以由训练中的视听模型使用以人工解决同步任务。(iii)我们整理了一个时间和空间同步信号稀疏的数据集;以及(iv)在密集和稀疏数据集上定量和定性地显示了所提出的模型的有效性。  项目页面:v-iashin.github.io/SparseSync
摘要:The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spatially small and may occur only infrequently during a many seconds-long video clip, i.e. the synchronisation signal is 'sparse in space and time'. This contrasts with the case of synchronising videos of talking heads, where audio-visual correspondence is dense in both time and space.  We make four contributions: (i) in order to handle longer temporal sequences required for sparse synchronisation signals, we design a multi-modal transformer model that employs 'selectors' to distil the long audio and visual streams into small sequences that are then used to predict the temporal offset between streams. (ii) We identify artefacts that can arise from the compression codecs used for audio and video and can be used by audio-visual models in training to artificially solve the synchronisation task. (iii) We curate a dataset with only sparse in time and space synchronisation signals; and (iv) the effectiveness of the proposed model is shown on both dense and sparse datasets quantitatively and qualitatively.  Project page: v-iashin.github.io/SparseSync


【6】 Anonymizing Speech with Generative Adversarial Networks to Preserve  Speaker Privacy

标题:利用产生式对抗网络保护说话人隐私的语音匿名

链接:https://arxiv.org/abs/2210.07002

* 与cs.SD语音【3】为同一篇

作者:Sarina Meyer,Pascal Tilli,Pavel Denisov,Florian Lux,Julia Koch,Ngoc Thang Vu
机构:Institute for Natural Language Processing (IMS), University of Stuttgart, Germany
备注:IEEE Spoken Language Technology Workshop 2022
摘要:为了保护语音数据的隐私性,说话人匿名化通过改变语音记录中的声音来隐藏说话人的身份。这通常伴随着在保护个人和数据对下游应用程序的可用性之间的隐私-效用权衡。在这方面的挑战之一是创造不存在的声音听起来尽可能自然。  在本文中,我们提出以Wasserstein距离为代价函数,使用生成式对抗网络来产生说话人嵌入,以解决这个问题。通过将这些人工嵌入到语音到文本到语音的流水线中,我们在隐私和实用性方面优于先前的方法。根据标准的客观度量和人工评估,我们的方法生成原始记录的可理解的、内容保持但隐私保护的版本。
摘要:In order to protect the privacy of speech data, speaker anonymization aims for hiding the identity of a speaker by changing the voice in speech recordings. This typically comes with a privacy-utility trade-off between protection of individuals and usability of the data for downstream applications. One of the challenges in this context is to create non-existent voices that sound as natural as possible.  In this work, we propose to tackle this issue by generating speaker embeddings using a generative adversarial network with Wasserstein distance as cost function. By incorporating these artificial embeddings into a speech-to-text-to-speech pipeline, we outperform previous approaches in terms of privacy and utility. According to standard objective metrics and human evaluation, our approach generates intelligible and content-preserving yet privacy-protecting versions of the original recordings.


【7】 Multilingual Zero Resource Speech Recognition Base on Self-Supervise  Pre-Trained Acoustic Models

标题:基于自监督预训练声学模型的多语种零资源语音识别

链接:https://arxiv.org/abs/2210.06936

* 与cs.SD语音【4】为同一篇

作者:Haoyu Wang,Wei-Qiang Zhang,Hongbin Suo,Yulong Wan
机构:† Beijing National Research Center for Information Science and Technology, Department of Electronic Engineering, Tsinghua University, Beijing , China, ‡ Data & AI Engineering System, OPPO, Beijing , China
备注:accepted by ISCSLP 2022
摘要:标记的音频数据不足以为世界上的大多数语言建立令人满意的语音识别系统。已经有一些零资源方法试图在没有目标语言的标记音频数据的情况下执行音素或词级语音识别,但是这些方法的错误率通常太高而不能应用于现实世界场景。近年来,人们发现,自监督预训练模型的表示能力在零资源音素识别中非常有益。就我们而言,本文是第一次尝试将预训练模型的使用扩展到词级零资源语音识别。这是通过微调关于IPA音素转录的预先训练的模型和用关于额外文本训练的语言模型解码来完成的。在Wav2vec2.0和HuBERT模型上的实验结果表明,该方法对部分语种的分词错误率小于20%,对8种语种的平均分词错误率为33.77%。
摘要:Labeled audio data is insufficient to build satisfying speech recognition systems for most of the languages in the world. There have been some zero-resource methods trying to perform phoneme or word-level speech recognition without labeled audio data of the target language, but the error rate of these methods is usually too high to be applied in real-world scenarios. Recently, the representation ability of self-supervise pre-trained models has been found to be extremely beneficial in zero-resource phoneme recognition. As far as we are concerned, this paper is the first attempt to extend the use of pre-trained models into word-level zero-resource speech recognition. This is done by fine-tuning the pre-trained models on IPA phoneme transcriptions and decoding with a language model trained on extra texts. Experiments on Wav2vec 2.0 and HuBERT models show that this method can achieve less than 20% word error rate on some languages, and the average error rate on 8 languages is 33.77%.


【8】 Inner speech recognition through electroencephalographic signals

标题:基于脑电信号的内部语音识别

链接:https://arxiv.org/abs/2210.06472

* 与cs.SD语音【5】为同一篇

作者:Francesca Gasparini,Elisa Cazzaniga,Aurora Saibene
机构:Saibene,[,−,−,−,],  University of Milano-Bicocca, Viale Sarca , Milano, Italy,  NeuroMI, Milan Center for Neuroscience, Piazza dell’Ateneo Nuovo
备注:Submitted to the Italian Workshop on Artificial Intelligence for Human Machine Interaction (AIxHMI 2022), December 02, 2022, Udine, Italy
摘要:本文主要研究从脑电信号出发的内部语音识别。内部言语识别是指人在纯粹意义上思考的内化过程,通常与自己内心“声音”的听觉意象相关联。将EEG解码成文本应该理解为对有限数量的单词(命令)或音素(组成单词的声音单位)的存在进行分类。语音相关的BCI提供了有效的语音通信策略,用于通过从脑信号解释的语音命令来控制设备,通过恢复与他们的环境的通信来改善失去说话能力的人的生活质量。分析了两个公开的内部语音数据集。利用这些数据,研究并实现了一些分类模型,从基本方法(如支持向量机)到集成方法(如极限梯度提升分类器),再到神经网络(如长短时记忆(LSTM)和双向长短时记忆(BiLSTM))的使用。利用LSTM和BiLSTM模型,获得了与现有技术中存在的结果一致或优于现有技术中存在的结果的结果,所述LSTM和BiLSTM模型通常未在内部语音识别的文献中使用。
摘要:This work focuses on inner speech recognition starting from EEG signals. Inner speech recognition is defined as the internalized process in which the person thinks in pure meanings, generally associated with an auditory imagery of own inner "voice". The decoding of the EEG into text should be understood as the classification of a limited number of words (commands) or the presence of phonemes (units of sound that make up words). Speech-related BCIs provide effective vocal communication strategies for controlling devices through speech commands interpreted from brain signals, improving the quality of life of people who have lost the capability to speak, by restoring communication with their environment. Two public inner speech datasets are analysed. Using this data, some classification models are studied and implemented starting from basic methods such as Support Vector Machines, to ensemble methods such as the eXtreme Gradient Boosting classifier up to the use of neural networks such as Long Short Term Memory (LSTM) and Bidirectional Long Short Term Memory (BiLSTM). With the LSTM and BiLSTM models, generally not used in the literature of inner speech recognition, results in line with or superior to those present in the stateof-the-art are obtained.


机器翻译,仅供参考