今天跟大家分享一篇语音相关的论文合集:cs.SD语音13篇,eess.AS音频处理14篇。
【1】 Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models
标题:唇听:基于词汇模型的跨通道知识提炼的混合意义理解唇语
链接:https://arxiv.org/abs/2207.05692
作者:Hadeel Mabrouk,Omar Abugabal,Nourhan Sakr,Hesham M. Eraqi备注:arXiv admin note: text overlap with arXiv:2108.03543摘要:在这项工作中,我们提出了一种将语音识别能力从音频语音识别系统转移到视觉语音识别器的技术,我们的目标是在唇读模型训练期间利用音频数据。语音和视听系统在语音识别领域取得了令人瞩目的进展。然而,由于某些音素的视觉模糊性,视觉语音识别系统仍有许多有待探索的地方。为此,鉴于音频模型的不稳定性,开发视觉语音识别模型至关重要。这项工作的主要贡献是i)通过将序列级和框架级知识提取(KD)集成到其系统中,在最新的基于单词的唇读模型的基础上进行构建;ii)在训练视觉模型期间利用音频数据,这是以前基于文字的工作中没有使用过的专长;iii)在帧级KD中提出高斯形状平均,作为一种有效的技术,帮助模型在序列模型编码器处提取知识。这项工作提出了一种新颖且具有竞争力的唇读架构,因为我们证明了性能的显著改善,在LRW数据集上设置了一个新的基准,相当于88.64%。摘要:In this work, we propose a technique to transfer speech recognition capabilities from audio speech recognition systems to visual speech recognizers, where our goal is to utilize audio data during lipreading model training. Impressive progress in the domain of speech recognition has been exhibited by audio and audio-visual systems. Nevertheless, there is still much to be explored with regards to visual speech recognition systems due to the visual ambiguity of some phonemes. To this end, the development of visual speech recognition models is crucial given the instability of audio models. The main contributions of this work are i) building on recent state-of-the-art word-based lipreading models by integrating sequence-level and frame-level Knowledge Distillation (KD) to their systems; ii) leveraging audio data during training visual models, a feat which has not been utilized in prior word-based work; iii) proposing the Gaussian-shaped averaging in frame-level KD, as an efficient technique that aids the model in distilling knowledge at the sequence model encoder. This work proposes a novel and competitive architecture for lip-reading, as we demonstrate a noticeable improvement in performance, setting a new benchmark equals to 88.64% on the LRW dataset.
【2】 ReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody Relationships
标题:ReLyMe:通过融合歌词和旋律的关系来改进歌词到旋律的生成
链接:https://arxiv.org/abs/2207.05688
作者:Chen Zhang,Luchin Chang,Songruoyao Wu,Xu Tan,Tao Qin,Tie-Yan Liu,Kejun Zhang备注:Accepted by ACMMM 2022, oral摘要:歌词到旋律生成是最重要的自动音乐创作任务之一,它根据给定的歌词生成旋律。随着深度学习的快速发展,以前的工作都是用端到端的神经网络模型来解决这一问题。然而,深度学习模型无法很好地捕捉歌词和旋律之间严格但微妙的关系,这损害了歌词和生成的旋律之间的和谐。在本文中,我们提出了ReLyMe,一种从音乐理论中结合歌词和旋律之间关系的方法,以确保歌词和旋律之间的和谐。具体来说,我们首先介绍了歌词和旋律在音调、节奏和结构关系方面应遵循的几个原则。然后,通过在解码过程中添加相应的约束,将这些原则集成到神经网络歌词到旋律模型中,以改善歌词和旋律之间的和谐。我们使用一系列客观和主观指标来评估生成的旋律。在英文和中文歌曲数据集上的实验都表明了ReLyMe的有效性,证明了将音乐领域的抒情-旋律关系纳入神经抒情-旋律生成的优越性。摘要:Lyric-to-melody generation, which generates melody according to given lyrics, is one of the most important automatic music composition tasks. With the rapid development of deep learning, previous works address this task with end-to-end neural network models. However, deep learning models cannot well capture the strict but subtle relationships between lyrics and melodies, which compromises the harmony between lyrics and generated melodies. In this paper, we propose ReLyMe, a method that incorporates Relationships between Lyrics and Melodies from music theory to ensure the harmony between lyrics and melodies. Specifically, we first introduce several principles that lyrics and melodies should follow in terms of tone, rhythm, and structure relationships. These principles are then integrated into neural network lyric-to-melody models by adding corresponding constraints during the decoding process to improve the harmony between lyrics and melodies. We use a series of objective and subjective metrics to evaluate the generated melodies. Experiments on both English and Chinese song datasets show the effectiveness of ReLyMe, demonstrating the superiority of incorporating lyric-melody relationships from the music domain into neural lyric-to-melody generation.
【3】 The Contribution of Lyrics and Acoustics to Collaborative Understanding of Mood
标题:歌词和声学对情绪协作性理解的贡献
链接:https://arxiv.org/abs/2207.05680
作者:Shahrzad Naseri,Sravana Reddy,Joana Correia,Jussi Karlgren,Rosie Jones摘要:在这项工作中,我们通过数据驱动的分析来研究歌词和情绪之间的关联。我们的数据集由近100万首歌曲组成,歌曲情绪关联来自Spotify流媒体平台上的用户播放列表。我们利用基于transformers的最先进的自然语言处理模型来学习歌词和情绪之间的关联。我们发现,在Zero-Shot设置下,基于预训练transformer的语言模型(即开箱即用,无需对数据进行进一步训练)对于捕捉歌曲语气关联非常有效。此外,我们还说明了对歌曲情绪关联的训练可以得到一个高度准确的模型,该模型可以预测未看到歌曲的这些关联。此外,通过比较使用歌词和使用声学特征的模型的预测,我们观察到,与声学相比,歌词在情绪预测中的相对重要性取决于特定的情绪。最后,我们通过注释任务验证模型是否捕捉到与人类相同的歌词和声学信息,在注释任务中,我们根据歌词和声学获得人类对情绪歌曲相关性的判断。摘要:In this work, we study the association between song lyrics and mood through a data-driven analysis. Our data set consists of nearly one million songs, with song-mood associations derived from user playlists on the Spotify streaming platform. We take advantage of state-of-the-art natural language processing models based on transformers to learn the association between the lyrics and moods. We find that a pretrained transformer-based language model in a zero-shot setting -- i.e., out of the box with no further training on our data -- is powerful for capturing song-mood associations. Moreover, we illustrate that training on song-mood associations results in a highly accurate model that predicts these associations for unseen songs. Furthermore, by comparing the prediction of a model using lyrics with one using acoustic features, we observe that the relative importance of lyrics for mood prediction in comparison with acoustics depends on the specific mood. Finally, we verify if the models are capturing the same information about lyrics and acoustics as humans through an annotation task where we obtain human judgments of mood-song relevance based on lyrics and acoustics.
【4】 EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use
标题:EfficientLEAF:一个使用有问题的更快的可学习音频前端
链接:https://arxiv.org/abs/2207.05508
作者:Jan Schlüter,Gerald Gutenbrunner备注:Accepted at EUSIPCO 2022. Code at this https URL摘要:在音频分类中,参数较少的可微听觉滤波器组覆盖了硬编码频谱图和原始音频之间的中间地带。LEAF(arXiv:2101.08596)是一种基于Gabor的滤波器组,与每通道能量归一化(PCEN)相结合,已显示出有希望的结果,但计算成本较高。对于卷积核大小和步长不均匀的情况,通过用更好的并行操作替换PCEN,我们可以更有效地获得类似的结果。在六个音频分类任务的实验中,我们的前端以3%的成本匹配LEAF的准确性,但两者都未能始终优于固定的mel滤波器组。对可学习音频前端的探索尚未解决。摘要:In audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio. LEAF (arXiv:2101.08596), a Gabor-based filterbank combined with Per-Channel Energy Normalization (PCEN), has shown promising results, but is computationally expensive. With inhomogeneous convolution kernel sizes and strides, and by replacing PCEN with better parallelizable operations, we can reach similar results more efficiently. In experiments on six audio classification tasks, our frontend matches the accuracy of LEAF at 3% of the cost, but both fail to consistently outperform a fixed mel filterbank. The quest for learnable audio frontends is not solved.
【5】 A Generative deep learning approach for shape recognition of arbitrary objects from phaseless acoustic scattering data
标题:无相位声散射数据中任意目标形状识别的生成式深度学习方法
链接:https://arxiv.org/abs/2207.05433
作者:W. W. Ahmed,M. Farhat,P. -Y. Chen,X. Zhang,Y. Wu摘要:我们提出并演示了一种生成式深度学习方法,用于根据任意物体的声散射特性进行形状识别。该策略利用深度神经网络学习二维声目标的潜在空间与远场散射振幅之间的映射。将神经网络设计为对抗式自动编码器,并通过无监督学习进行训练,以确定声学对象的潜在空间。物体的重要结构特征嵌入到低维潜在空间中,该空间支持形状生成器的建模,并加速逆设计过程中的学习。所提出的逆设计使用具有编码器和解码器类似结构的变分推理方法,其中解码器由两个预训练神经网络组成,即生成器和正向模型。数据驱动框架找到了不适定逆散射问题的精确解,其中非唯一解空间被多频无相位远场模式克服。这种逆方法是一种强大的设计工具,不需要复杂的分析计算,为实际实现、任意形状潜艇或大型鱼类的自动识别以及其他水下应用开辟了新途径。摘要:We propose and demonstrate a generative deep learning approach for the shape recognition of an arbitrary object from its acoustic scattering properties. The strategy exploits deep neural networks to learn the mapping between the latent space of a two-dimensional acoustic object and the far-field scattering amplitudes. A neural network is designed as an Adversarial autoencoder and trained via unsupervised learning to determine the latent space of the acoustic object. Important structural features of the object are embedded in lower-dimensional latent space which supports the modeling of a shape generator and accelerates the learning in the inverse design process.The proposed inverse design uses the variational inference approach with encoder and decoder-like architecture where the decoder is composed of two pretrained neural networks, the generator and the forward model. The data-driven framework finds an accurate solution to the ill-posed inverse scattering problem, where non-unique solution space is overcome by the multifrequency phaseless far-field patterns. This inverse method is a powerful design tool that does not require complex analytical calculation and opens up new avenues for practical realization, automatic recognition of arbitrary shaped submarines or large fish, and other underwater applications.
【6】 Western Mediterranean wetlands bird species classification: evaluating small-footprint deep learning approaches on a new annotated dataset
标题:西地中海湿地鸟类分类:在一个新的注解数据集上评估小规模深度学习方法
链接:https://arxiv.org/abs/2207.05393
作者:Juan Gómez-Gómez,Ester Vidaña-Vila,Xavier Sevillano备注:17 pages, 8 figures, 3 tables摘要:部署一个运行在无线声学传感器网络上的专家系统,该网络由生物声学监测设备组成,这些设备可以从鸟类的声音中识别鸟类物种,这将使许多具有生态价值的任务实现自动化,包括分析鸟类种群组成或检测环境相关区域的濒危物种。由于人工智能的最新进展,赋予这些设备准确的音频分类能力是可能的,其中深度学习技术是最优秀的。然而,使生物声学设备价格合理的一个关键问题是使用可嵌入资源和电池受限硬件平台的小占地面积深度神经网络。因此,这项工作对两种重型和大型足迹深度神经网络(VGG16和ResNet50)和一种轻型替代品MobileNetV2进行了关键的对比分析。我们的实验结果表明,MobileNet V2的F1平均分数比ResNet50低5%(0.789比0.834),性能优于VGG16,足迹大小比VGG16小近40倍。此外,为了比较模型,我们创建并公布了西地中海湿地鸟类数据集,包括201.6分钟和5795段自然公园艾瓜莫尔斯德兰波德20种特有鸟类的音频摘录。摘要:The deployment of an expert system running over a wireless acoustic sensors network made up of bioacoustic monitoring devices that recognise bird species from their sounds would enable the automation of many tasks of ecological value, including the analysis of bird population composition or the detection of endangered species in areas of environmental interest. Endowing these devices with accurate audio classification capabilities is possible thanks to the latest advances in artificial intelligence, among which deep learning techniques excel. However, a key issue to make bioacoustic devices affordable is the use of small footprint deep neural networks that can be embedded in resource and battery constrained hardware platforms. For this reason, this work presents a critical comparative analysis between two heavy and large footprint deep neural networks (VGG16 and ResNet50) and a lightweight alternative, MobileNetV2. Our experimental results reveal that MobileNetV2 achieves an average F1-score less than a 5\% lower than ResNet50 (0.789 vs. 0.834), performing better than VGG16 with a footprint size nearly 40 times smaller. Moreover, to compare the models, we have created and made public the Western Mediterranean Wetland Birds dataset, consisting of 201.6 minutes and 5,795 audio excerpts of 20 endemic bird species of the Aiguamolls de l'Empord\`a Natural Park.
【7】 Multitask Learning from Augmented Auxiliary Data for Improving Speech Emotion Recognition
标题:基于增强辅助数据的多任务学习改进语音情感识别
链接:https://arxiv.org/abs/2207.05298
作者:Siddique Latif,Rajib Rana,Sara Khalifa,Raja Jurdak,Björn W. Schuller备注:Under review IEEE Transactions on Affective Computing摘要:尽管语音情感识别(SER)最近取得了进展,但最先进的系统缺乏跨不同条件的通用性。泛化较差的一个关键潜在原因是情感数据集的缺乏,这是设计鲁棒机器学习(ML)模型的一个重要障碍。SER最近的工作重点是利用多任务学习(MTL)方法通过学习共享表示来提高泛化。然而,大多数研究提出了MTL解决方案,要求辅助任务具有元标签,这限制了SER系统的训练。本文提出了一个MTL框架(MTL-AUG),该框架从增强数据中学习广义表示。我们使用增强类型分类和无监督重建作为辅助任务,允许在增强数据上训练SER系统,而不需要辅助任务的任何元标签。MTL-AUG的半监督性质允许利用丰富的未标记数据来进一步提高SER的性能。我们在以下环境中对提出的框架进行了综合评估:(1)语料库内,(2)跨语料库和跨语言,(3)噪声语音,(4)和对抗性攻击。我们使用广泛使用的IEMOCAP、MSP-IMPROV和EMODB数据集进行的评估表明,与现有最先进的方法相比,结果有所改善。摘要:Despite the recent progress in speech emotion recognition (SER), state-of-the-art systems lack generalisation across different conditions. A key underlying reason for poor generalisation is the scarcity of emotion datasets, which is a significant roadblock to designing robust machine learning (ML) models. Recent works in SER focus on utilising multitask learning (MTL) methods to improve generalisation by learning shared representations. However, most of these studies propose MTL solutions with the requirement of meta labels for auxiliary tasks, which limits the training of SER systems. This paper proposes an MTL framework (MTL-AUG) that learns generalised representations from augmented data. We utilise augmentation-type classification and unsupervised reconstruction as auxiliary tasks, which allow training SER systems on augmented data without requiring any meta labels for auxiliary tasks. The semi-supervised nature of MTL-AUG allows for the exploitation of the abundant unlabelled data to further boost the performance of SER. We comprehensively evaluate the proposed framework in the following settings: (1) within corpus, (2) cross-corpus and cross-language, (3) noisy speech, (4) and adversarial attacks. Our evaluations using the widely used IEMOCAP, MSP-IMPROV, and EMODB datasets show improved results compared to existing state-of-the-art methods.
【8】 Indoor optical fiber eavesdropping approach and its avoidance
标题:室内光纤窃听方法及其避免
链接:https://arxiv.org/abs/2207.05267
作者:Haiqing Hao,Zhongwang Pang,Guan Wang,Bo Wang备注:8 pages, 4 figures, submitted to Optics Express摘要:光纤网络已成为全球的基础设施。除了在通信中的基本功能外,其感知能力也越来越受到重视。在本文中,我们讨论了家用光纤用于窃听的风险,并在实验室演示了其性能。在家用光学调制解调器前使用3米长的尾纤,可以用激光干涉仪窃听正常人类语音,并在1.1公里外恢复。定量分析了检测距离极限和系统噪声。我们还提供了一些实用的方法来防止通过家用光纤进行窃听。摘要:The optical fiber network has become a worldwide infrastructure. In addition to the basic functions in telecommunication, its sensing ability has attracted more and more attention. In this paper, we discuss the risk of household fiber being used for eavesdropping and demonstrate its performance in the lab. Using a 3-meter tail fiber in front of the household optical modem, voices of normal human speech can be eavesdropped by a laser interferometer and recovered 1.1 km away. The detection distance limit and system noise are analyzed quantitatively. We also give some practical ways to prevent eavesdropping through household fiber.
【9】 Online Continual Learning of End-to-End Speech Recognition Models
标题:端到端语音识别模型的在线持续学习
链接:https://arxiv.org/abs/2207.05071
作者:Muqiao Yang,Ian Lane,Shinji Watanabe备注:Accepted at InterSpeech 2022摘要:持续学习也称为终身学习,目的是在新数据可用时不断学习。虽然先前对自动语音识别中的连续学习的研究集中在跨多个不同语音识别任务的模型自适应上,但在本文中,我们提出了一种用于单个任务的自动语音识别的实验设置。特别关注同一任务的额外训练数据随时间增量可用的情况,我们证明了使用在线梯度情节记忆(GEM)方法对端到端语音识别模型执行增量模型更新的有效性。此外,我们表明,通过在线连续学习和选择性采样策略,我们可以保持与从头开始重新训练模型类似的准确性,同时需要显著降低计算成本。我们还使用自监督学习(SSL)功能验证了我们的方法。摘要:Continual Learning, also known as Lifelong Learning, aims to continually learn from new data as it becomes available. While prior research on continual learning in automatic speech recognition has focused on the adaptation of models across multiple different speech recognition tasks, in this paper we propose an experimental setting for \textit{online continual learning} for automatic speech recognition of a single task. Specifically focusing on the case where additional training data for the same task becomes available incrementally over time, we demonstrate the effectiveness of performing incremental model updates to end-to-end speech recognition models with an online Gradient Episodic Memory (GEM) method. Moreover, we show that with online continual learning and a selective sampling strategy, we can maintain an accuracy that is similar to retraining a model from scratch while requiring significantly lower computation costs. We have also verified our method with self-supervised learning (SSL) features.
【10】 PoeticTTS -- Controllable Poetry Reading for Literary Studies
标题:PoeticTTS--文学研究的可控诗歌阅读
链接:https://arxiv.org/abs/2207.05549
作者:Julia Koch,Florian Lux,Nadja Schauffler,Toni Bernhart,Felix Dieterle,Jonas Kuhn,Sandra Richter,Gabriel Viehhauser,Ngoc Thang Vu备注:Accepted to Interspeech 2022摘要:由于诗歌语音固有的特定语调模式,诗歌语音合成具有挑战性。在这项工作中,我们提出了一种方法来合成几乎像人类一样自然的诗歌,以使文学学者能够系统地研究关于文本、口语实现和听者对诗歌感知之间相互作用的假设。为了满足文学研究的这些特殊要求,我们通过从人类参考背诵中克隆韵律值来重新合成诗歌,然后利用细粒度韵律控制在人在回路环境中操纵合成语音,以改变背诵的特定现象。我们发现,微调我们的诗歌TTS模型在很大程度上捕捉了诗歌的语调模式,这有利于韵律的克隆和操纵,并验证了我们的方法在客观评价和人类研究中的成功。摘要:Speech synthesis for poetry is challenging due to specific intonation patterns inherent to poetic speech. In this work, we propose an approach to synthesise poems with almost human like naturalness in order to enable literary scholars to systematically examine hypotheses on the interplay between text, spoken realisation, and the listener's perception of poems. To meet these special requirements for literary studies, we resynthesise poems by cloning prosodic values from a human reference recitation, and afterwards make use of fine-grained prosody control to manipulate the synthetic speech in a human-in-the-loop setting to alter the recitation w.r.t. specific phenomena. We find that finetuning our TTS model on poetry captures poetic intonation patterns to a large extent which is beneficial for prosody cloning and manipulation and verify the success of our approach both in an objective evaluation as well as in human studies.
【11】 Statistics of the interaural parameters for dichotic tones in diotic noise ($N_0 S_ψ$)
标题:双耳噪声中双耳音耳间参数的统计($N_0 S_ψ$)
链接:https://arxiv.org/abs/2207.05541
作者:Jörg Encke,Mathias Dietz摘要:在双耳听觉领域中,通常使用由耳间相移音调组成的刺激(通常称为$N\u 0 S\u psi$)。由于将二分音噪声与二分音混合,这类刺激包含耳间相位差和电平差的随机波动。本研究报告了两个双耳差异的联合概率密度函数,作为音调振幅和双耳相位的函数。此外,推导了双耳相位差和刺激瞬时功率的第二个联合概率密度函数。摘要:Stimuli consisting of an interaurally phase-shifted tone in diotic noise -- often referred to as $N_0 S_\psi$ -- are commonly used in the field of binaural hearing. As a consequence of mixing diotic noise with a dichotic tone, this type of stimulus contains random fluctuations in both interaural phase- and level-difference. This study reports the joint probability density functions of the two interaural differences as a function of amplitude and interaural phase of the tone. Furthermore, a second joint probability density function for interaural phase differences and the instantaneous power of the stimulus is derived.
【12】 Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
标题:基于信息最大化和对比学习的标签有效自监督说话人确认
链接:https://arxiv.org/abs/2207.05506
作者:Théo Lepage,Réda Dehak摘要:最先进的说话人验证系统本质上依赖于某种人类监督,因为它们是在大量标记数据上训练的。然而,手动注释话语速度慢、成本高,并且无法扩展到今天可用的数据量。在本研究中,我们通过直接从原始音频中学习表示来探索用于说话人验证的自监督学习。目标是产生鲁棒的说话人嵌入,具有较小的说话人内方差和较大的说话人间方差。我们的方法基于最近的信息最大化学习框架和密集的数据增强预处理步骤。我们评估了这些方法在没有对比样本的情况下工作的能力,然后表明它们在与对比损失相结合时取得了更好的性能。此外,我们进行的实验表明,与现有技术相比,我们的方法达到了有竞争力的结果,并且在使用少量标记数据进行微调时,与监督基线相比,可以获得更好的性能。摘要:State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to the amount of data available today. In this study, we explore self-supervised learning for speaker verification by learning representations directly from raw audio. The objective is to produce robust speaker embeddings that have small intra-speaker and large inter-speaker variance. Our approach is based on recent information maximization learning frameworks and an intensive data augmentation pre-processing step. We evaluate the ability of these methods to work without contrastive samples before showing that they achieve better performance when combined with a contrastive loss. Furthermore, we conduct experiments to show that our method reaches competitive results compared to existing techniques and can get better performances compared to a supervised baseline when fine-tuned with a small portion of labeled data.
【13】 End-to-end speech recognition modeling from de-identified data
标题:基于未识别数据的端到端语音识别建模
链接:https://arxiv.org/abs/2207.05469
作者:Martin Flechl,Shou-Chun Yin,Junho Park,Peter Skala备注:Accepted to INTERSPEECH 2022摘要:用于自动语音识别建模的数据去识别是保护隐私的关键组件,尤其是在医疗领域。然而,简单地从端到端模型训练数据中删除所有个人识别信息(PII)会导致显著的性能下降,尤其是在识别相似类别的名称、日期、位置和单词方面。我们提出并评估了部分恢复该损失的两步方法。首先,识别PII,并用相同类别的随机词序列替换每个出现。然后,通过文本到语音或将从语料库中提取的匹配音频片段拼接在一起,生成相应的音频。这些人工音频/标签对,以及来自没有PII的原始数据的说话人转向,用于训练模型。我们评估了该方法在医疗对话内部数据上的性能,并观察到在总体文字错误率方面几乎完全恢复了性能下降,同时仍然保持了很强的日志化性能。我们的主要重点是提高PII相关词识别的召回率和准确性。根据PII类别,使用我们提出的方法可以恢复50%-90%的性能下降。摘要:De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end model training data leads to a significant performance degradation in particular for the recognition of names, dates, locations, and words from similar categories. We propose and evaluate a two-step method for partially recovering this loss. First, PII is identified, and each occurrence is replaced with a random word sequence of the same category. Then, corresponding audio is produced via text-to-speech or by splicing together matching audio fragments extracted from the corpus. These artificial audio/label pairs, together with speaker turns from the original data without PII, are used to train models. We evaluate the performance of this method on in-house data of medical conversations and observe a recovery of almost the entire performance degradation in the general word error rate while still maintaining a strong diarization performance. Our main focus is the improvement of recall and precision in the recognition of PII-related words. Depending on the PII category, between $50\% - 90\%$ of the performance degradation can be recovered using our proposed method.
【1】 PoeticTTS -- Controllable Poetry Reading for Literary Studies
标题:PoeticTTS--文学研究的可控诗歌阅读
链接:https://arxiv.org/abs/2207.05549
* 与cs.SD语音【10】为同一篇
作者:Julia Koch,Florian Lux,Nadja Schauffler,Toni Bernhart,Felix Dieterle,Jonas Kuhn,Sandra Richter,Gabriel Viehhauser,Ngoc Thang Vu备注:Accepted to Interspeech 2022摘要:由于诗歌语音固有的特定语调模式,诗歌语音合成具有挑战性。在这项工作中,我们提出了一种方法来合成几乎像人类一样自然的诗歌,以使文学学者能够系统地研究关于文本、口语实现和听者对诗歌感知之间相互作用的假设。为了满足文学研究的这些特殊要求,我们通过从人类参考背诵中克隆韵律值来重新合成诗歌,然后利用细粒度韵律控制在人在回路环境中操纵合成语音,以改变背诵的特定现象。我们发现,微调我们的诗歌TTS模型在很大程度上捕捉了诗歌的语调模式,这有利于韵律的克隆和操纵,并验证了我们的方法在客观评价和人类研究中的成功。摘要:Speech synthesis for poetry is challenging due to specific intonation patterns inherent to poetic speech. In this work, we propose an approach to synthesise poems with almost human like naturalness in order to enable literary scholars to systematically examine hypotheses on the interplay between text, spoken realisation, and the listener's perception of poems. To meet these special requirements for literary studies, we resynthesise poems by cloning prosodic values from a human reference recitation, and afterwards make use of fine-grained prosody control to manipulate the synthetic speech in a human-in-the-loop setting to alter the recitation w.r.t. specific phenomena. We find that finetuning our TTS model on poetry captures poetic intonation patterns to a large extent which is beneficial for prosody cloning and manipulation and verify the success of our approach both in an objective evaluation as well as in human studies.
【2】 Statistics of the interaural parameters for dichotic tones in diotic noise ($N_0 S_ψ$)
标题:双耳噪声中双耳音耳间参数的统计($N_0 S_ψ$)
链接:https://arxiv.org/abs/2207.05541
* 与cs.SD语音【11】为同一篇
作者:Jörg Encke,Mathias Dietz摘要:在双耳听觉领域中,通常使用由耳间相移音调组成的刺激(通常称为$N\u 0 S\u psi$)。由于将二分音噪声与二分音混合,这类刺激包含耳间相位差和电平差的随机波动。本研究报告了两个双耳差异的联合概率密度函数,作为音调振幅和双耳相位的函数。此外,推导了双耳相位差和刺激瞬时功率的第二个联合概率密度函数。摘要:Stimuli consisting of an interaurally phase-shifted tone in diotic noise -- often referred to as $N_0 S_\psi$ -- are commonly used in the field of binaural hearing. As a consequence of mixing diotic noise with a dichotic tone, this type of stimulus contains random fluctuations in both interaural phase- and level-difference. This study reports the joint probability density functions of the two interaural differences as a function of amplitude and interaural phase of the tone. Furthermore, a second joint probability density function for interaural phase differences and the instantaneous power of the stimulus is derived.
【3】 Label-Efficient Self-Supervised Speaker Verification With Information Maximization and Contrastive Learning
标题:基于信息最大化和对比学习的标签有效自监督说话人确认
链接:https://arxiv.org/abs/2207.05506
* 与cs.SD语音【12】为同一篇
作者:Théo Lepage,Réda Dehak摘要:最先进的说话人验证系统本质上依赖于某种人类监督,因为它们是在大量标记数据上训练的。然而,手动注释话语速度慢、成本高,并且无法扩展到今天可用的数据量。在本研究中,我们通过直接从原始音频中学习表示来探索用于说话人验证的自监督学习。目标是产生鲁棒的说话人嵌入,具有较小的说话人内方差和较大的说话人间方差。我们的方法基于最近的信息最大化学习框架和密集的数据增强预处理步骤。我们评估了这些方法在没有对比样本的情况下工作的能力,然后表明它们在与对比损失相结合时取得了更好的性能。此外,我们进行的实验表明,与现有技术相比,我们的方法达到了有竞争力的结果,并且在使用少量标记数据进行微调时,与监督基线相比,可以获得更好的性能。摘要:State-of-the-art speaker verification systems are inherently dependent on some kind of human supervision as they are trained on massive amounts of labeled data. However, manually annotating utterances is slow, expensive and not scalable to the amount of data available today. In this study, we explore self-supervised learning for speaker verification by learning representations directly from raw audio. The objective is to produce robust speaker embeddings that have small intra-speaker and large inter-speaker variance. Our approach is based on recent information maximization learning frameworks and an intensive data augmentation pre-processing step. We evaluate the ability of these methods to work without contrastive samples before showing that they achieve better performance when combined with a contrastive loss. Furthermore, we conduct experiments to show that our method reaches competitive results compared to existing techniques and can get better performances compared to a supervised baseline when fine-tuned with a small portion of labeled data.
【4】 End-to-end speech recognition modeling from de-identified data
标题:基于未识别数据的端到端语音识别建模
链接:https://arxiv.org/abs/2207.05469
* 与cs.SD语音【13】为同一篇
作者:Martin Flechl,Shou-Chun Yin,Junho Park,Peter Skala备注:Accepted to INTERSPEECH 2022摘要:用于自动语音识别建模的数据去识别是保护隐私的关键组件,尤其是在医疗领域。然而,简单地从端到端模型训练数据中删除所有个人识别信息(PII)会导致显著的性能下降,尤其是在识别相似类别的名称、日期、位置和单词方面。我们提出并评估了部分恢复该损失的两步方法。首先,识别PII,并用相同类别的随机词序列替换每个出现。然后,通过文本到语音或将从语料库中提取的匹配音频片段拼接在一起,生成相应的音频。这些人工音频/标签对,以及来自没有PII的原始数据的说话人转向,用于训练模型。我们评估了该方法在医疗对话内部数据上的性能,并观察到在总体文字错误率方面几乎完全恢复了性能下降,同时仍然保持了很强的日志化性能。我们的主要重点是提高PII相关词识别的召回率和准确性。根据PII类别,使用我们提出的方法可以恢复50%-90%的性能下降。摘要:De-identification of data used for automatic speech recognition modeling is a critical component in protecting privacy, especially in the medical domain. However, simply removing all personally identifiable information (PII) from end-to-end model training data leads to a significant performance degradation in particular for the recognition of names, dates, locations, and words from similar categories. We propose and evaluate a two-step method for partially recovering this loss. First, PII is identified, and each occurrence is replaced with a random word sequence of the same category. Then, corresponding audio is produced via text-to-speech or by splicing together matching audio fragments extracted from the corpus. These artificial audio/label pairs, together with speaker turns from the original data without PII, are used to train models. We evaluate the performance of this method on in-house data of medical conversations and observe a recovery of almost the entire performance degradation in the general word error rate while still maintaining a strong diarization performance. Our main focus is the improvement of recall and precision in the recognition of PII-related words. Depending on the PII category, between $50\% - 90\%$ of the performance degradation can be recovered using our proposed method.
【5】 Lip-Listening: Mixing Senses to Understand Lips using Cross Modality Knowledge Distillation for Word-Based Models
标题:唇听:基于词汇模型的跨通道知识提炼的混合意义理解唇语
链接:https://arxiv.org/abs/2207.05692
* 与cs.SD语音【1】为同一篇
作者:Hadeel Mabrouk,Omar Abugabal,Nourhan Sakr,Hesham M. Eraqi备注:arXiv admin note: text overlap with arXiv:2108.03543摘要:在这项工作中,我们提出了一种将语音识别能力从音频语音识别系统转移到视觉语音识别器的技术,我们的目标是在唇读模型训练期间利用音频数据。语音和视听系统在语音识别领域取得了令人瞩目的进展。然而,由于某些音素的视觉模糊性,视觉语音识别系统仍有许多有待探索的地方。为此,鉴于音频模型的不稳定性,开发视觉语音识别模型至关重要。这项工作的主要贡献是i)通过将序列级和框架级知识提取(KD)集成到其系统中,在最新的基于单词的唇读模型的基础上进行构建;ii)在训练视觉模型期间利用音频数据,这是以前基于文字的工作中没有使用过的专长;iii)在帧级KD中提出高斯形状平均,作为一种有效的技术,帮助模型在序列模型编码器处提取知识。这项工作提出了一种新颖且具有竞争力的唇读架构,因为我们证明了性能的显著改善,在LRW数据集上设置了一个新的基准,相当于88.64%。摘要:In this work, we propose a technique to transfer speech recognition capabilities from audio speech recognition systems to visual speech recognizers, where our goal is to utilize audio data during lipreading model training. Impressive progress in the domain of speech recognition has been exhibited by audio and audio-visual systems. Nevertheless, there is still much to be explored with regards to visual speech recognition systems due to the visual ambiguity of some phonemes. To this end, the development of visual speech recognition models is crucial given the instability of audio models. The main contributions of this work are i) building on recent state-of-the-art word-based lipreading models by integrating sequence-level and frame-level Knowledge Distillation (KD) to their systems; ii) leveraging audio data during training visual models, a feat which has not been utilized in prior word-based work; iii) proposing the Gaussian-shaped averaging in frame-level KD, as an efficient technique that aids the model in distilling knowledge at the sequence model encoder. This work proposes a novel and competitive architecture for lip-reading, as we demonstrate a noticeable improvement in performance, setting a new benchmark equals to 88.64% on the LRW dataset.
【6】 The MuSe 2022 Multimodal Sentiment Analysis Challenge: Humor, Emotional Reactions, and Stress
标题:缪斯2022多通道情绪分析挑战:幽默、情绪反应和压力
链接:https://arxiv.org/abs/2207.05691
作者:Lukas Christ,Shahin Amiriparian,Alice Baird,Panagiotis Tzirakis,Alexander Kathan,Niklas Müller,Lukas Stappen,Eva-Maria Meßner,Andreas König,Alan Cowen,Erik Cambria,Björn W. Schuller备注:Preliminary baseline paper for the 3rd Multimodal Sentiment Analysis Challenge (MuSe) 2022, a full-day workshop at ACM Multimedia 2022摘要:2022年多模式情绪分析挑战赛(MuSe)致力于多模式情绪和情绪识别。对于今年的挑战,我们提供了三个数据集:(i)Passau自发性足球教练幽默(Passau SFCH)数据集,其中包含德国足球教练的视听记录,标记为幽默的存在;(ii)休谟反应数据集,其中个体对情绪刺激的反应已根据七种情绪表达强度进行注释,以及(iii)Ulm Trier社会压力测试(Ulm TSST)数据集,包括标记有压力倾向的人的连续情绪值(觉醒和效价)的视听数据。使用引入的数据集,MuSe 2022 2022解决了三个当代情感计算问题:在幽默检测子挑战(MuSe幽默)中,必须识别自发幽默;在情绪反应子挑战(缪斯反应)中,必须预测七种细粒度的“野生”情绪;在情绪压力子挑战(MuSe压力)中,持续预测压力情绪值。这项挑战旨在吸引不同的研究群体,鼓励他们学科的融合。MuSe 2022主要针对视听情感识别、健康信息学和符号情感分析社区。本文描述了数据集以及从中提取的特征集。使用带有LSTM单元的递归神经网络在每个子挑战的测试分区上设置竞争基线结果。我们报告了的曲线下面积(AUC)。8480表示缪斯幽默。2801平均值(来自7类)皮尔逊相关系数的MuSe反应,以及。4931协和相关系数(CCC)和。缪斯应激中的配价和唤醒分别为4761。摘要:The Multimodal Sentiment Analysis Challenge (MuSe) 2022 is dedicated to multimodal sentiment and emotion recognition. For this year's challenge, we feature three datasets: (i) the Passau Spontaneous Football Coach Humor (Passau-SFCH) dataset that contains audio-visual recordings of German football coaches, labelled for the presence of humour; (ii) the Hume-Reaction dataset in which reactions of individuals to emotional stimuli have been annotated with respect to seven emotional expression intensities, and (iii) the Ulm-Trier Social Stress Test (Ulm-TSST) dataset comprising of audio-visual data labelled with continuous emotion values (arousal and valence) of people in stressful dispositions. Using the introduced datasets, MuSe 2022 2022 addresses three contemporary affective computing problems: in the Humor Detection Sub-Challenge (MuSe-Humor), spontaneous humour has to be recognised; in the Emotional Reactions Sub-Challenge (MuSe-Reaction), seven fine-grained `in-the-wild' emotions have to be predicted; and in the Emotional Stress Sub-Challenge (MuSe-Stress), a continuous prediction of stressed emotion values is featured. The challenge is designed to attract different research communities, encouraging a fusion of their disciplines. Mainly, MuSe 2022 targets the communities of audio-visual emotion recognition, health informatics, and symbolic sentiment analysis. This baseline paper describes the datasets as well as the feature sets extracted from them. A recurrent neural network with LSTM cells is used to set competitive baseline results on the test partitions for each sub-challenge. We report an Area Under the Curve (AUC) of .8480 for MuSe-Humor; .2801 mean (from 7-classes) Pearson's Correlations Coefficient for MuSe-Reaction, as well as .4931 Concordance Correlation Coefficient (CCC) and .4761 for valence and arousal in MuSe-Stress, respectively.
【7】 ReLyMe: Improving Lyric-to-Melody Generation by Incorporating Lyric-Melody Relationships
标题:ReLyMe:通过融合歌词和旋律的关系来改进歌词到旋律的生成
链接:https://arxiv.org/abs/2207.05688
* 与cs.SD语音【2】为同一篇
作者:Chen Zhang,Luchin Chang,Songruoyao Wu,Xu Tan,Tao Qin,Tie-Yan Liu,Kejun Zhang备注:Accepted by ACMMM 2022, oral摘要:歌词到旋律生成是最重要的自动音乐创作任务之一,它根据给定的歌词生成旋律。随着深度学习的快速发展,以前的工作都是用端到端的神经网络模型来解决这一问题。然而,深度学习模型无法很好地捕捉歌词和旋律之间严格但微妙的关系,这损害了歌词和生成的旋律之间的和谐。在本文中,我们提出了ReLyMe,一种从音乐理论中结合歌词和旋律之间关系的方法,以确保歌词和旋律之间的和谐。具体来说,我们首先介绍了歌词和旋律在音调、节奏和结构关系方面应遵循的几个原则。然后,通过在解码过程中添加相应的约束,将这些原则集成到神经网络歌词到旋律模型中,以改善歌词和旋律之间的和谐。我们使用一系列客观和主观指标来评估生成的旋律。在英文和中文歌曲数据集上的实验都表明了ReLyMe的有效性,证明了将音乐领域的抒情-旋律关系纳入神经抒情-旋律生成的优越性。摘要:Lyric-to-melody generation, which generates melody according to given lyrics, is one of the most important automatic music composition tasks. With the rapid development of deep learning, previous works address this task with end-to-end neural network models. However, deep learning models cannot well capture the strict but subtle relationships between lyrics and melodies, which compromises the harmony between lyrics and generated melodies. In this paper, we propose ReLyMe, a method that incorporates Relationships between Lyrics and Melodies from music theory to ensure the harmony between lyrics and melodies. Specifically, we first introduce several principles that lyrics and melodies should follow in terms of tone, rhythm, and structure relationships. These principles are then integrated into neural network lyric-to-melody models by adding corresponding constraints during the decoding process to improve the harmony between lyrics and melodies. We use a series of objective and subjective metrics to evaluate the generated melodies. Experiments on both English and Chinese song datasets show the effectiveness of ReLyMe, demonstrating the superiority of incorporating lyric-melody relationships from the music domain into neural lyric-to-melody generation.
【8】 The Contribution of Lyrics and Acoustics to Collaborative Understanding of Mood
标题:歌词和声学对情绪协作性理解的贡献
链接:https://arxiv.org/abs/2207.05680
* 与cs.SD语音【3】为同一篇
作者:Shahrzad Naseri,Sravana Reddy,Joana Correia,Jussi Karlgren,Rosie Jones摘要:在这项工作中,我们通过数据驱动的分析来研究歌词和情绪之间的关联。我们的数据集由近100万首歌曲组成,歌曲情绪关联来自Spotify流媒体平台上的用户播放列表。我们利用基于transformers的最先进的自然语言处理模型来学习歌词和情绪之间的关联。我们发现,在Zero-Shot设置下,基于预训练transformer的语言模型(即开箱即用,无需对数据进行进一步训练)对于捕捉歌曲语气关联非常有效。此外,我们还说明了对歌曲情绪关联的训练可以得到一个高度准确的模型,该模型可以预测未看到歌曲的这些关联。此外,通过比较使用歌词和使用声学特征的模型的预测,我们观察到,与声学相比,歌词在情绪预测中的相对重要性取决于特定的情绪。最后,我们通过注释任务验证模型是否捕捉到与人类相同的歌词和声学信息,在注释任务中,我们根据歌词和声学获得人类对情绪歌曲相关性的判断。摘要:In this work, we study the association between song lyrics and mood through a data-driven analysis. Our data set consists of nearly one million songs, with song-mood associations derived from user playlists on the Spotify streaming platform. We take advantage of state-of-the-art natural language processing models based on transformers to learn the association between the lyrics and moods. We find that a pretrained transformer-based language model in a zero-shot setting -- i.e., out of the box with no further training on our data -- is powerful for capturing song-mood associations. Moreover, we illustrate that training on song-mood associations results in a highly accurate model that predicts these associations for unseen songs. Furthermore, by comparing the prediction of a model using lyrics with one using acoustic features, we observe that the relative importance of lyrics for mood prediction in comparison with acoustics depends on the specific mood. Finally, we verify if the models are capturing the same information about lyrics and acoustics as humans through an annotation task where we obtain human judgments of mood-song relevance based on lyrics and acoustics.
【9】 EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use
标题:EfficientLEAF:一个使用有问题的更快的可学习音频前端
链接:https://arxiv.org/abs/2207.05508
* 与cs.SD语音【4】为同一篇
作者:Jan Schlüter,Gerald Gutenbrunner备注:Accepted at EUSIPCO 2022. Code at this https URL摘要:在音频分类中,参数较少的可微听觉滤波器组覆盖了硬编码频谱图和原始音频之间的中间地带。LEAF(arXiv:2101.08596)是一种基于Gabor的滤波器组,与每通道能量归一化(PCEN)相结合,已显示出有希望的结果,但计算成本较高。对于卷积核大小和步长不均匀的情况,通过用更好的并行操作替换PCEN,我们可以更有效地获得类似的结果。在六个音频分类任务的实验中,我们的前端以3%的成本匹配LEAF的准确性,但两者都未能始终优于固定的mel滤波器组。对可学习音频前端的探索尚未解决。摘要:In audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio. LEAF (arXiv:2101.08596), a Gabor-based filterbank combined with Per-Channel Energy Normalization (PCEN), has shown promising results, but is computationally expensive. With inhomogeneous convolution kernel sizes and strides, and by replacing PCEN with better parallelizable operations, we can reach similar results more efficiently. In experiments on six audio classification tasks, our frontend matches the accuracy of LEAF at 3% of the cost, but both fail to consistently outperform a fixed mel filterbank. The quest for learnable audio frontends is not solved.
【10】 A Generative deep learning approach for shape recognition of arbitrary objects from phaseless acoustic scattering data
标题:无相位声散射数据中任意目标形状识别的生成式深度学习方法
链接:https://arxiv.org/abs/2207.05433
* 与cs.SD语音【5】为同一篇
作者:W. W. Ahmed,M. Farhat,P. -Y. Chen,X. Zhang,Y. Wu摘要:我们提出并演示了一种生成式深度学习方法,用于根据任意物体的声散射特性进行形状识别。该策略利用深度神经网络学习二维声目标的潜在空间与远场散射振幅之间的映射。将神经网络设计为对抗式自动编码器,并通过无监督学习进行训练,以确定声学对象的潜在空间。物体的重要结构特征嵌入到低维潜在空间中,该空间支持形状生成器的建模,并加速逆设计过程中的学习。所提出的逆设计使用具有编码器和解码器类似结构的变分推理方法,其中解码器由两个预训练神经网络组成,即生成器和正向模型。数据驱动框架找到了不适定逆散射问题的精确解,其中非唯一解空间被多频无相位远场模式克服。这种逆方法是一种强大的设计工具,不需要复杂的分析计算,为实际实现、任意形状潜艇或大型鱼类的自动识别以及其他水下应用开辟了新途径。摘要:We propose and demonstrate a generative deep learning approach for the shape recognition of an arbitrary object from its acoustic scattering properties. The strategy exploits deep neural networks to learn the mapping between the latent space of a two-dimensional acoustic object and the far-field scattering amplitudes. A neural network is designed as an Adversarial autoencoder and trained via unsupervised learning to determine the latent space of the acoustic object. Important structural features of the object are embedded in lower-dimensional latent space which supports the modeling of a shape generator and accelerates the learning in the inverse design process.The proposed inverse design uses the variational inference approach with encoder and decoder-like architecture where the decoder is composed of two pretrained neural networks, the generator and the forward model. The data-driven framework finds an accurate solution to the ill-posed inverse scattering problem, where non-unique solution space is overcome by the multifrequency phaseless far-field patterns. This inverse method is a powerful design tool that does not require complex analytical calculation and opens up new avenues for practical realization, automatic recognition of arbitrary shaped submarines or large fish, and other underwater applications.
【11】 Western Mediterranean wetlands bird species classification: evaluating small-footprint deep learning approaches on a new annotated dataset
标题:西地中海湿地鸟类分类:在一个新的注解数据集上评估小规模深度学习方法
链接:https://arxiv.org/abs/2207.05393
* 与cs.SD语音【6】为同一篇
作者:Juan Gómez-Gómez,Ester Vidaña-Vila,Xavier Sevillano备注:17 pages, 8 figures, 3 tables摘要:部署一个运行在无线声学传感器网络上的专家系统,该网络由生物声学监测设备组成,这些设备可以从鸟类的声音中识别鸟类物种,这将使许多具有生态价值的任务实现自动化,包括分析鸟类种群组成或检测环境相关区域的濒危物种。由于人工智能的最新进展,赋予这些设备准确的音频分类能力是可能的,其中深度学习技术是最优秀的。然而,使生物声学设备价格合理的一个关键问题是使用可嵌入资源和电池受限硬件平台的小占地面积深度神经网络。因此,这项工作对两种重型和大型足迹深度神经网络(VGG16和ResNet50)和一种轻型替代品MobileNetV2进行了关键的对比分析。我们的实验结果表明,MobileNet V2的F1平均分数比ResNet50低5%(0.789比0.834),性能优于VGG16,足迹大小比VGG16小近40倍。此外,为了比较模型,我们创建并公布了西地中海湿地鸟类数据集,包括201.6分钟和5795段自然公园艾瓜莫尔斯德兰波德20种特有鸟类的音频摘录。摘要:The deployment of an expert system running over a wireless acoustic sensors network made up of bioacoustic monitoring devices that recognise bird species from their sounds would enable the automation of many tasks of ecological value, including the analysis of bird population composition or the detection of endangered species in areas of environmental interest. Endowing these devices with accurate audio classification capabilities is possible thanks to the latest advances in artificial intelligence, among which deep learning techniques excel. However, a key issue to make bioacoustic devices affordable is the use of small footprint deep neural networks that can be embedded in resource and battery constrained hardware platforms. For this reason, this work presents a critical comparative analysis between two heavy and large footprint deep neural networks (VGG16 and ResNet50) and a lightweight alternative, MobileNetV2. Our experimental results reveal that MobileNetV2 achieves an average F1-score less than a 5\% lower than ResNet50 (0.789 vs. 0.834), performing better than VGG16 with a footprint size nearly 40 times smaller. Moreover, to compare the models, we have created and made public the Western Mediterranean Wetland Birds dataset, consisting of 201.6 minutes and 5,795 audio excerpts of 20 endemic bird species of the Aiguamolls de l'Empord\`a Natural Park.
【12】 Multitask Learning from Augmented Auxiliary Data for Improving Speech Emotion Recognition
标题:基于增强辅助数据的多任务学习改进语音情感识别
链接:https://arxiv.org/abs/2207.05298
* 与cs.SD语音【7】为同一篇
作者:Siddique Latif,Rajib Rana,Sara Khalifa,Raja Jurdak,Björn W. Schuller备注:Under review IEEE Transactions on Affective Computing摘要:尽管语音情感识别(SER)最近取得了进展,但最先进的系统缺乏跨不同条件的通用性。泛化较差的一个关键潜在原因是情感数据集的缺乏,这是设计鲁棒机器学习(ML)模型的一个重要障碍。SER最近的工作重点是利用多任务学习(MTL)方法通过学习共享表示来提高泛化。然而,大多数研究提出了MTL解决方案,要求辅助任务具有元标签,这限制了SER系统的训练。本文提出了一个MTL框架(MTL-AUG),该框架从增强数据中学习广义表示。我们使用增强类型分类和无监督重建作为辅助任务,允许在增强数据上训练SER系统,而不需要辅助任务的任何元标签。MTL-AUG的半监督性质允许利用丰富的未标记数据来进一步提高SER的性能。我们在以下环境中对提出的框架进行了综合评估:(1)语料库内,(2)跨语料库和跨语言,(3)噪声语音,(4)和对抗性攻击。我们使用广泛使用的IEMOCAP、MSP-IMPROV和EMODB数据集进行的评估表明,与现有最先进的方法相比,结果有所改善。摘要:Despite the recent progress in speech emotion recognition (SER), state-of-the-art systems lack generalisation across different conditions. A key underlying reason for poor generalisation is the scarcity of emotion datasets, which is a significant roadblock to designing robust machine learning (ML) models. Recent works in SER focus on utilising multitask learning (MTL) methods to improve generalisation by learning shared representations. However, most of these studies propose MTL solutions with the requirement of meta labels for auxiliary tasks, which limits the training of SER systems. This paper proposes an MTL framework (MTL-AUG) that learns generalised representations from augmented data. We utilise augmentation-type classification and unsupervised reconstruction as auxiliary tasks, which allow training SER systems on augmented data without requiring any meta labels for auxiliary tasks. The semi-supervised nature of MTL-AUG allows for the exploitation of the abundant unlabelled data to further boost the performance of SER. We comprehensively evaluate the proposed framework in the following settings: (1) within corpus, (2) cross-corpus and cross-language, (3) noisy speech, (4) and adversarial attacks. Our evaluations using the widely used IEMOCAP, MSP-IMPROV, and EMODB datasets show improved results compared to existing state-of-the-art methods.
【13】 Indoor optical fiber eavesdropping approach and its avoidance
标题:室内光纤窃听方法及其避免
链接:https://arxiv.org/abs/2207.05267
* 与cs.SD语音【8】为同一篇
作者:Haiqing Hao,Zhongwang Pang,Guan Wang,Bo Wang备注:8 pages, 4 figures, submitted to Optics Express摘要:光纤网络已成为全球的基础设施。除了在通信中的基本功能外,其感知能力也越来越受到重视。在本文中,我们讨论了家用光纤用于窃听的风险,并在实验室演示了其性能。在家用光学调制解调器前使用3米长的尾纤,可以用激光干涉仪窃听正常人类语音,并在1.1公里外恢复。定量分析了检测距离极限和系统噪声。我们还提供了一些实用的方法来防止通过家用光纤进行窃听。摘要:The optical fiber network has become a worldwide infrastructure. In addition to the basic functions in telecommunication, its sensing ability has attracted more and more attention. In this paper, we discuss the risk of household fiber being used for eavesdropping and demonstrate its performance in the lab. Using a 3-meter tail fiber in front of the household optical modem, voices of normal human speech can be eavesdropped by a laser interferometer and recovered 1.1 km away. The detection distance limit and system noise are analyzed quantitatively. We also give some practical ways to prevent eavesdropping through household fiber.
【14】 Online Continual Learning of End-to-End Speech Recognition Models
标题:端到端语音识别模型的在线持续学习
链接:https://arxiv.org/abs/2207.05071
* 与cs.SD语音【9】为同一篇
作者:Muqiao Yang,Ian Lane,Shinji Watanabe备注:Accepted at InterSpeech 2022摘要:持续学习也称为终身学习,目的是在新数据可用时不断学习。虽然先前对自动语音识别中的连续学习的研究集中在跨多个不同语音识别任务的模型自适应上,但在本文中,我们提出了一种用于单个任务的自动语音识别的实验设置。特别关注同一任务的额外训练数据随时间增量可用的情况,我们证明了使用在线梯度情节记忆(GEM)方法对端到端语音识别模型执行增量模型更新的有效性。此外,我们表明,通过在线连续学习和选择性采样策略,我们可以保持与从头开始重新训练模型类似的准确性,同时需要显著降低计算成本。我们还使用自监督学习(SSL)功能验证了我们的方法。摘要:Continual Learning, also known as Lifelong Learning, aims to continually learn from new data as it becomes available. While prior research on continual learning in automatic speech recognition has focused on the adaptation of models across multiple different speech recognition tasks, in this paper we propose an experimental setting for \textit{online continual learning} for automatic speech recognition of a single task. Specifically focusing on the case where additional training data for the same task becomes available incrementally over time, we demonstrate the effectiveness of performing incremental model updates to end-to-end speech recognition models with an online Gradient Episodic Memory (GEM) method. Moreover, we show that with online continual learning and a selective sampling strategy, we can maintain an accuracy that is similar to retraining a model from scratch while requiring significantly lower computation costs. We have also verified our method with self-supervised learning (SSL) features.
机器翻译,仅供参考