今天跟大家分享一篇语音相关的论文合集:cs.SD语音13篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily
cs.SD语音

【1】 Canonical Cortical Graph Neural Networks and its Application for Speech  Enhancement in Future Audio-Visual Hearing Aids

标题:规范皮层图神经网络及其在未来视听助听器语音增强中的应用

链接:https://arxiv.org/abs/2206.02671

作者:Leandro A. Passos,João Paulo Papa,Ahsan Adeel
机构:CMI Lab, School of Engineering and Informatics, University of Wolverhampton, Wolverhampton, United Kingdom,  Department of Computing, S˜ao Paulo State University, Bauru, Brazil,  deepCI.org ,, Parkside Terrace, Edinburgh, United Kingdom
摘要:尽管机器学习算法最近取得了成功,但在考虑需要不同来源之间交互的更复杂任务时,大多数模型仍面临一些缺点,例如多模式输入数据和逻辑时序。另一方面,生物大脑在这个意义上是高度敏锐的,经过数百万年的进化,它能够自动管理和整合这样的信息流。在这种背景下,本文从大脑皮层回路的最新发现中得到启发,提出了一种生物学上更合理的自我监督机器学习方法,该方法将使用层内调制的多模态信息与典型相关分析(CCA)相结合,以及一种跟踪时间数据的记忆机制,所谓的典型皮层图神经网络。考虑到更好的干净音频重建和能量效率,该方法的表现优于最近的最新研究结果,这可以通过减少和抑制神经元放电率分布来描述,表明该模型是未来视听助听器中语音增强的合适方法。
摘要:Despite the recent success of machine learning algorithms, most of these models still face several drawbacks when considering more complex tasks requiring interaction between different sources, such as multimodal input data and logical time sequence. On the other hand, the biological brain is highly sharpened in this sense, empowered to automatically manage and integrate such a stream of information through millions of years of evolution. In this context, this paper finds inspiration from recent discoveries on cortical circuits in the brain to propose a more biologically plausible self-supervised machine learning approach that combines multimodal information using intra-layer modulations together with canonical correlation analysis (CCA), as well as a memory mechanism to keep track of temporal data, the so-called Canonical Cortical Graph Neural networks. The approach outperformed recent state-of-the-art results considering both better clean audio reconstruction and energy efficiency, described by a reduced and smother neuron firing rate distribution, suggesting the model as a suitable approach for speech enhancement in future audio-visual hearing aid devices.


【2】 Tagged-MRI2Audio with Attention Guided Heterogeneous Translator

标题:带有注意力引导的异类翻译器的标记-MRI2音频

链接:https://arxiv.org/abs/2206.02284

作者:Xiaofeng Liu,Fangxu Xing,Jerry L. Prince,Jiachen Zhuo,Maureen Stone,Georges El Fakhri,Jonghye Woo
机构:Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA,  Johns Hopkins University, Baltimore, MD, USA,  University of Maryland, Baltimore, MD, USA
备注:MICCAI 2022 (early accept)
摘要:理解标记MRI和可理解语音中舌和口咽肌变形之间的潜在关系,对于推进言语运动控制理论和言语相关疾病的治疗具有重要作用。然而,由于其不同的表现形式,两种模式之间的直接映射——即二维(中矢状切片)加上时间标记的MRI序列及其相应的一维波形——并不简单。相反,我们求助于二维谱图作为中间表示,其中包含基音和共振,从中开发端到端的深度学习框架,以将标记的MRI序列转换为数据集大小有限的相应音频波形。我们的框架基于一种全新的完全卷积不对称翻译器,并在自我剩余注意策略的指导下,专门利用语音中的运动肌肉结构。此外,我们利用具有相同话语的样本的成对关联,并采用潜在空间表示解纠缠策略。此外,我们将对抗训练方法与生成对抗网络相结合,以提高生成光谱图的真实性。我们的实验结果显示,我们的框架能够从标记的MRI序列中生成清晰的音频波形,超过了其他方法。实验结果共有63个标记的MRI序列以及语音声学。
摘要:Understanding the underlying relationship between tongue and oropharyngeal muscle deformation seen in tagged-MRI and intelligible speech plays an important role in advancing speech motor control theories and treatment of speech related-disorders. Because of their heterogeneous representations, however, direct mapping between the two modalities -- i.e., two-dimensional (mid-sagittal slice) plus time tagged-MRI sequence and its corresponding one-dimensional waveform -- is not straightforward. Instead, we resort to two-dimensional spectrograms as an intermediate representation, which contains both pitch and resonance, from which to develop an end-to-end deep learning framework to translate from a sequence of tagged-MRI to its corresponding audio waveform with limited dataset size. Our framework is based on a novel fully convolutional asymmetry translator with guidance of a self residual attention strategy to specifically exploit the moving muscular structures during speech. In addition, we leverage a pairwise correlation of the samples with the same utterances with a latent space representation disentanglement strategy. Furthermore, we incorporate an adversarial training approach with generative adversarial networks to offer improved realism on our generated spectrograms. Our experimental results, carried out with a total of 63 tagged-MRI sequences alongside speech acoustics, showed that our framework enabled the generation of clear audio waveforms from a sequence of tagged-MRI, surpassing competing methods.


【3】 Zero-Shot Voice Conditioning for Denoising Diffusion TTS Models

标题:用于扩散TTS模型去噪的零发声条件

链接:https://arxiv.org/abs/2206.02246

作者:Alon Levkovitch,Eliya Nachmani,Lior Wolf
机构:Tel Aviv University, Facebook AI Research
摘要:我们提出了一种新的方法来调节经过预训练的去噪扩散语音模型,以在训练过程中看不到的新人物的声音中生成语音。该方法需要从目标人员处获得一个短样本(约3秒),并在推理时引导生成,无需任何训练步骤。该方法的核心是一个采样过程,该过程将去噪模型的估计与新说话人样本的低通版本相结合。客观和主观评估表明,我们的采样方法可以生成与目标说话人频率相似的声音,准确度与最先进的方法相当,并且无需训练。
摘要:We present a novel way of conditioning a pretrained denoising diffusion speech model to produce speech in the voice of a novel person unseen during training. The method requires a short (~3 seconds) sample from the target person, and generation is steered at inference time, without any training steps. At the heart of the method lies a sampling process that combines the estimation of the denoising model with a low-pass version of the new speaker's sample. The objective and subjective evaluations show that our sampling method can generate a voice similar to that of the target speaker in terms of frequency, with an accuracy comparable to state-of-the-art methods, and without training.


【4】 Variable-rate hierarchical CPC leads to acoustic unit discovery in  speech

标题:可变速率分层CPC导致语音中的声学单元发现

链接:https://arxiv.org/abs/2206.02211

作者:Santiago Cuervo,Adrian Łańcucki,Ricard Marxer,Paweł Rychlikowski,Jan Chorowski
机构:University of Wrocław, Poland, Adrian Ła´ncucki, NVIDIA, Poland, Université de Toulon, France, NavAlgo, France
备注:Submitted to 36th Conference on Neural Information Processing Systems (NeurIPS 2022)
摘要:深度学习的成功来自于它通过学习由低级表示定义的高级表示来捕获数据层次结构的能力。在本文中,我们通过应用多层次对比预测编码(CPC)来探索语音层次表示的自监督学习。我们观察到,简单地堆叠两个CPC模型并不会比单级架构产生显著的改进。受语音通常被描述为时间上分布不均匀的离散单元序列这一事实的启发,我们提出了一个模型,其中低级CPC模块的输出是非均匀降采样的,以直接最小化高级CPC模块的损失。后者还通过聚焦负采样和预测目标量化来实现连续高层表示的相异性,从而在其表示中实现可分性和离散性先验。根据下游语音识别任务的测量结果,考虑语音信号的结构改进了单级CPC特征,增强了学习表示的分离,同时产生了与电话边界非常相似的有意义的信号分割。
摘要:The success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learning of hierarchical representations of speech by applying multiple levels of Contrastive Predictive Coding (CPC). We observe that simply stacking two CPC models does not yield significant improvements over single-level architectures. Inspired by the fact that speech is often described as a sequence of discrete units unevenly distributed in time, we propose a model in which the output of a low-level CPC module is non-uniformly downsampled to directly minimize the loss of a high-level CPC module. The latter is designed to also enforce a prior of separability and discreteness in its representations by enforcing dissimilarity of successive high-level representations through focused negative sampling, and by quantization of the prediction targets. Accounting for the structure of the speech signal improves upon single-level CPC features and enhances the disentanglement of the learned representations, as measured by downstream speech recognition tasks, while resulting in a meaningful segmentation of the signal that closely resembles phone boundaries.


【5】 M2FNet: Multi-modal Fusion Network for Emotion Recognition in  Conversation

标题:M2FNet:面向会话情感识别的多模式融合网络

链接:https://arxiv.org/abs/2206.02187

作者:Vishal Chudasama,Purbayan Kar,Ashish Gudmalwar,Nirmesh Shah,Pankaj Wasnik,Naoyuki Onoe
机构:Media Analysis Group, Sony Research India, Bangalore, India
备注:Accepted for publication in the 5th Multimodal Learning and Applications (MULA) Workshop at CVPR 2022
摘要:对话中的情感识别(ERC)对于发展富有同情心的人机交互至关重要。在对话视频中,情感可以以多种形式呈现,即音频、视频和文字记录。然而,由于这些模式的固有特点,多模式ERC一直被认为是一项具有挑战性的任务。现有的ERC研究主要集中在讨论中使用文本信息,而忽略了其他两种方式。我们期望通过采用多模态方法可以提高情绪识别的准确性。因此,在本研究中,我们提出了一种多模态融合网络(M2FNet),从视觉、音频和文本模态中提取情感相关特征。它采用了一种基于多头注意的融合机制,将输入数据中情感丰富的潜在表示结合起来。我们引入了一种新的特征抽取器来从音频和视频模态中提取潜在特征。该特征提取器使用一种新的基于边缘的自适应三重损失函数进行训练,以从音频和视频数据中学习情感相关的特征。在ERC领域,现有的方法在一个基准数据集上表现良好,但在其他数据集上表现不佳。我们的结果表明,在著名的MELD和IEMOCAP数据集上,所提出的M2FNet体系结构在加权平均F1得分方面优于所有其他方法,并在ERC中创造了新的最先进性能。
摘要:Emotion Recognition in Conversations (ERC) is crucial in developing sympathetic human-machine interaction. In conversational videos, emotion can be present in multiple modalities, i.e., audio, video, and transcript. However, due to the inherent characteristics of these modalities, multi-modal ERC has always been considered a challenging undertaking. Existing ERC research focuses mainly on using text information in a discussion, ignoring the other two modalities. We anticipate that emotion recognition accuracy can be improved by employing a multi-modal approach. Thus, in this study, we propose a Multi-modal Fusion Network (M2FNet) that extracts emotion-relevant features from visual, audio, and text modality. It employs a multi-head attention-based fusion mechanism to combine emotion-rich latent representations of the input data. We introduce a new feature extractor to extract latent features from the audio and visual modality. The proposed feature extractor is trained with a novel adaptive margin-based triplet loss function to learn emotion-relevant features from the audio and visual data. In the domain of ERC, the existing methods perform well on one benchmark dataset but not on others. Our results show that the proposed M2FNet architecture outperforms all other methods in terms of weighted average F1 score on well-known MELD and IEMOCAP datasets and sets a new state-of-the-art performance in ERC.


【6】 Learning Speaker-specific Lip-to-Speech Generation

标题:学习特定于说话人的唇语转换生成

链接:https://arxiv.org/abs/2206.02050

作者:Munender Varshney,Ravindra Yadav,Vinay P. Namboodiri,Rajesh M Hegde
机构:† Electrical department, Indian institute of Technology Kanpur, India, ‡ Computer Science and Engineering department, Indian institute of Technology Kanpur, India, §University of Bath, UK
备注:Accepted at ICPR 2022
摘要:众所周知,对于普通人来说,理解嘴唇运动并从中推断出语音是很困难的。准确的唇读任务从说话者的各种线索及其上下文或环境设置中得到帮助。每个说话人都有不同的口音和说话风格,这可以从他们的视觉和言语特征中推断出来。这项工作的目的是在一个不受限制的大词汇量中,了解单个说话人的语音和嘴唇运动序列之间的相关性/映射。我们将帧序列建模为自动编码器设置中Transformer之前的帧序列,并学习利用音频和视频的时间特性的联合嵌入。我们使用深度度量学习学习时间同步,它引导解码器生成与输入嘴唇运动同步的语音。因此,预测后验可以提供说话人说话风格的生成语音。我们在网格和Lip2Wav化学讲座数据集上对我们的模型进行了约束,以评估在无约束自然环境中通过嘴唇运动生成单说话人自然语音的任务。使用各种定性和定量指标以及人工评估进行的广泛评估也表明,我们的方法在几乎所有评估指标上都优于Lip2Wav化学数据集(无约束环境下的大词汇量),并且略优于最先进的网格数据集。
摘要:Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every speaker has a different accent and speaking style, which can be inferred from their visual and speech features. This work aims to understand the correlation/mapping between speech and the sequence of lip movement of individual speakers in an unconstrained and large vocabulary. We model the frame sequence as a prior to the transformer in an auto-encoder setting and learned a joint embedding that exploits temporal properties of both audio and video. We learn temporal synchronization using deep metric learning, which guides the decoder to generate speech in sync with input lip movements. The predictive posterior thus gives us the generated speech in speaker speaking style. We have trained our model on the Grid and Lip2Wav Chemistry lecture dataset to evaluate single speaker natural speech generation tasks from lip movement in an unconstrained natural setting. Extensive evaluation using various qualitative and quantitative metrics with human evaluation also shows that our method outperforms the Lip2Wav Chemistry dataset(large vocabulary in an unconstrained setting) by a good margin across almost all evaluation metrics and marginally outperforms the state-of-the-art on GRID dataset.


【7】 Continuous-Time Analog Filters for Audio Edge Intelligence: Review and  Analysis on Design Techniques

标题:用于音频边缘智能的连续时间模拟滤波器:设计技术回顾与分析

链接:https://arxiv.org/abs/2206.02639

作者:Kwantae Kim,Shih-Chii Liu
机构:The authors are with the Institute of Neuroinformatics,  University of Z¨urichand ETH Z¨urich
备注:16 pages, 17 figures
摘要:硅耳蜗设计体现了生物耳蜗的功能。它们的用途已被探索用于人工耳蜗应用,最近又被探索用于边缘音频设备,这些设备需要支持始终在线操作。由于其严格的电源限制带来了一些设计挑战,IC设计师不得不寻找使用低备用电源的解决方案。一种很有希望的生物启发方法是将硅耳蜗的连续时间模拟滤波器通道与一个小内存占用深度神经网络相结合,该神经网络在边缘任务(如关键字定位)上进行训练,从而允许将所有块嵌入IC中。本文从硅耳蜗的原始双四阶滤波器电路开始,回顾了当前边缘音频设备中用作特征提取器的模拟滤波器电路。我们的分析从将基本双二次滤波器解释为双积分环路拓扑开始,并回顾了二阶低通和带通滤波器的设计进展,从基于OTA的结构到基于源极跟随器的结构。我们还推导和分析了小信号传递函数,并讨论了这些滤波器的性能方面。这些不同滤波器配置的分析可应用于其他应用领域,例如采用前端带通滤波器的生物医学设备。
摘要:Silicon cochlea designs capture the functionality of the biological cochlea. Their use has been explored for cochlea prosthesis applications and more recently in edge audio devices which are required to support always-on operation. As their stringent power constraints pose several design challenges, IC designers are forced to look for solutions that use low standby power. One promising bio-inspired approach is to combine the continuous-time analog filter channels of the silicon cochlea with a small memory footprint deep neural network that is trained on edge tasks such as keyword spotting, thereby allowing all blocks to be embedded in an IC. This paper reviews the analog filter circuits used as feature extractors for current edge audio devices, starting with the original biquad filter circuits proposed for the silicon cochlea. Our analysis starts from the interpretation of a basic biquad filter as a two-integrator-loop topology and reviews the progression in the design of second-order low-pass and band-pass filters ranging from OTA-based to source-follower-based architectures. We also derive and analyze the small-signal transfer function and discuss performance aspects of these filters. The analysis of these different filter configurations can be applied to other application domains such as biomedical devices which employ a front-end bandpass filter.


【8】 UTTS: Unsupervised TTS with Conditional Disentangled Sequential  Variational Auto-encoder

标题:UTTS:条件解缠序列变分自动编码器的无监督TTS

链接:https://arxiv.org/abs/2206.02512

作者:Jiachen Lian,Chunlei Zhang,Gopala Krishna Anumanchipalli,Dong Yu
机构:UC Berkeley, EECS, CA, Tencent AI Lab, Bellevue, WA, Gopala K. Anumanchipalli
摘要:在本文中,我们提出了一种新的无监督文本到语音(UTTS)框架,该框架不需要文本-音频对来进行TTS声学建模(AM)。UTTS是从非纠缠语音表征学习的角度开发的多说话人语音合成器。该框架为TTS推理提供了说话人时长模型、音色特征(身份)和内容的灵活选择。我们利用自监督语音表示学习以及语音合成前端技术的最新进展进行系统开发。具体地说,我们利用一个词典将输入文本映射到音素序列,然后将其扩展到框架级的强制对齐(FA),并使用一个与说话人相关的持续时间模型。然后,我们开发了一个对齐映射模块,将FA转换为无监督对齐(UA)。最后,作为自我监督TTS AM的条件分离顺序变分自动编码器(C-DSVAE)将预测的UA和目标说话人嵌入生成mel谱图,并最终使用神经声码器将其转换为波形。我们展示了我们的方法如何在不使用成对TTS语料库的情况下实现语音合成。实验表明,UTTS能够合成出自然度高、可懂度高的语音,并能进行客观评价。
摘要:In this paper, we propose a novel unsupervised text-to-speech (UTTS) framework which does not require text-audio pairs for the TTS acoustic modeling (AM). UTTS is a multi-speaker speech synthesizer developed from the perspective of disentangled speech representation learning. The framework offers a flexible choice of a speaker's duration model, timbre feature (identity) and content for TTS inference. We leverage recent advancements in self-supervised speech representation learning as well as speech synthesis front-end techniques for the system development. Specifically, we utilize a lexicon to map input text to the phoneme sequence, which is expanded to the frame-level forced alignment (FA) with a speaker-dependent duration model. Then, we develop an alignment mapping module that converts the FA to the unsupervised alignment (UA). Finally, a Conditional Disentangled Sequential Variational Auto-encoder (C-DSVAE), serving as the self-supervised TTS AM, takes the predicted UA and a target speaker embedding to generate the mel spectrogram, which is ultimately converted to waveform with a neural vocoder. We show how our method enables speech synthesis without using a paired TTS corpus. Experiments demonstrate that UTTS can synthesize speech of high naturalness and intelligibility measured by human and objective evaluations.


【9】 Online Neural Diarization of Unlimited Numbers of Speakers

标题:无限说话人的在线神经网络二值化

链接:https://arxiv.org/abs/2206.02432

作者:Shota Horiguchi,Shinji Watanabe,Paola Garcia,Yuki Takashima,Yohei Kawaguchi
机构:Watanabe is with Carnegie Mellon University,  Garc´ıa is with Johns Hopkins University
摘要:本文描述了一种对无限数量的说话人进行离线和在线说话人日记的方法。端到端神经二值化(EEND)通过将说话人二值化描述为多标签分类问题,实现了重叠感知说话人二值化。通过引入基于说话人的吸引子,它还扩展到了灵活数量的说话人。然而,基于吸引子的EEND的说话人的输出数量在经验上是有限的;它不能处理推理过程中出现的说话人数量高于训练过程中出现的说话人数量的情况,因为它的说话人计数是在完全监督的方式下训练的。我们的方法EEND-GLA通过在基于吸引子的EEND中引入无监督聚类来解决这个问题。该方法首先将输入音频划分为短块,然后对每个块进行基于吸引子的二值化,最后根据局部计算的吸引子之间的相似度对每个块的结果进行聚类。虽然输出扬声器的数量限制在每个块内,但整个输入估计的扬声器总数可能高于限制。为了以在线方式使用EEND-GLA,我们的方法还扩展了说话人跟踪缓冲区,该缓冲区最初是为了支持传统EEND的在线推理。我们引入了一种分块缓冲区更新,使说话人跟踪缓冲区与EEND-GLA兼容。最后,为了改进在线日志化,我们的方法改进了缓冲区更新方法,并重新讨论了EEND的可变块大小训练。实验结果表明,EEND-GLA可以在离线和在线推理中对未知数量的说话人进行说话人日记。
摘要:A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a multi-label classification problem. It has also been extended for a flexible number of speakers by introducing speaker-wise attractors. However, the output number of speakers of attractor-based EEND is empirically capped; it cannot deal with cases where the number of speakers appearing during inference is higher than that during training because its speaker counting is trained in a fully supervised manner. Our method, EEND-GLA, solves this problem by introducing unsupervised clustering into attractor-based EEND. In the method, the input audio is first divided into short blocks, then attractor-based diarization is performed for each block, and finally the results of each blocks are clustered on the basis of the similarity between locally-calculated attractors. While the number of output speakers is limited within each block, the total number of speakers estimated for the entire input can be higher than the limitation. To use EEND-GLA in an online manner, our method also extends the speaker-tracing buffer, which was originally proposed to enable online inference of conventional EEND. We introduces a block-wise buffer update to make the speaker-tracing buffer compatible with EEND-GLA. Finally, to improve online diarization, our method improves the buffer update method and revisits the variable chunk-size training of EEND. The experimental results demonstrate that EEND-GLA can perform speaker diarization of an unseen number of speakers in both offline and online inferences.


【10】 Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for  Text-to-Speech

标题:DICT-TTS:学习利用先验词典知识进行语音转换的发音

链接:https://arxiv.org/abs/2206.02147

作者:Ziyue Jiang,Su Zhe,Zhou Zhao,Qian Yang,Yi Ren,Jinglin Liu,Zhenhui Ye
机构:Zhejiang University, Zhe Su∗
摘要:多音字消歧旨在从自然文本序列中获取准确的语音知识,以实现可靠的文语转换(TTS)系统。然而,以前的方法需要大量带注释的训练数据和语言专家的额外努力,因此很难将高质量的神经TTS系统扩展到域外的日常对话和世界各地的无数语言。本文从简洁新颖的角度解决了多音字的消歧问题:我们提出了Dict-TTS,这是一种语义感知的生成性文语转换模型,具有在线网站词典(自然语言中现有的先验信息)。具体而言,我们设计了语义到发音注意(S2PA)模块,将输入文本序列与词典中的先验语义之间的语义模式进行匹配,并获得相应的发音;S2PA模块可以使用端到端TTS模型轻松训练,无需任何带注释的音素标签。在三种语言上的实验结果表明,我们的模型在发音准确性方面优于几种强基线模型,并改进了TTS系统的韵律建模。对不同语言编码器的进一步广泛分析表明,Dict-TTS中的每种设计都是有效的。音频示例位于\url{https://dicttts.github.io/DictTTS-Demo/}.
摘要:Polyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems. However, previous approaches require substantial annotated training data and additional efforts from language experts, making it difficult to extend high-quality neural TTS systems to out-of-domain daily conversations and countless languages worldwide. This paper tackles the polyphone disambiguation problem from a concise and novel perspective: we propose Dict-TTS, a semantic-aware generative text-to-speech model with an online website dictionary (the existing prior information in the natural language). Specifically, we design a semantics-to-pronunciation attention (S2PA) module to match the semantic patterns between the input text sequence and the prior semantics in the dictionary and obtain the corresponding pronunciations; The S2PA module can be easily trained with the end-to-end TTS model without any annotated phoneme labels. Experimental results in three languages show that our model outperforms several strong baseline models in terms of pronunciation accuracy and improves the prosody modeling of TTS systems. Further extensive analyses with different linguistic encoders demonstrate that each design in Dict-TTS is effective. Audio samples are available at \url{https://dicttts.github.io/DictTTS-Demo/}.


【11】 Geometrically-Motivated Primary-Ambient Decomposition With  Center-Channel Extraction

标题:中心通道提取的几何激励原生环境分解

链接:https://arxiv.org/abs/2206.02125

作者:Jouni Paulus,Matteo Torcoli
机构:Fraunhofer Institute for Integrated Circuits IIS, and, International Audio Laboratories Erlangen†, Erlangen, Germany
备注:accepted into EUSIPCO 2022
摘要:提出了一种几何激励的初级环境分解方法,并在上混合应用中进行了评估。该方法包括两个步骤,并提供了特别直观的解释。第一步包括对输入立体声场景应用信号自适应旋转,将主要声源转换为旋转场景的中心。第二步应用基于简单信号模型的中心信道提取方法,并在均方误差意义下进行优化。通过使用估计的环境分量来评估性能,以实现从真实立体声信号开始的环绕声。在报告的听力测试中,参与者被要求调整音频场景包络,并找到最让他们满意的音频设置。该方法实现上混音的可能性得到了广泛的应用,与原来的立体声混音相比,用户满意度显著提高。
摘要:A geometrically-motivated method for primary-ambient decomposition is proposed and evaluated in an up-mixing application. The method consists of two steps, accommodating a particularly intuitive explanation. The first step consists of signal-adaptive rotations applied on the input stereo scene, which translate the primary sound sources into the center of the rotated scene. The second step applies a center-channel extraction method, based on a simple signal model and optimal in the mean-squared-error sense. The performance is evaluated by using the estimated ambient component to enable surround sound starting from real-world stereo signals. The participants in the reported listening test are asked to adjust the audio scene envelopment and find the audio settings that pleases them the most. The possibility for up-mixing enabled by the proposed method is used extensively, and the user satisfaction is significantly increased compared to the original stereo mix.


【12】 Sampling Frequency Independent Dialogue Separation标

题:采样频率无关的对话分离

链接:https://arxiv.org/abs/2206.02124

作者:Jouni Paulus,Matteo Torcoli
机构:Fraunhofer Institute for Integrated Circuits IIS, and, International Audio Laboratories Erlangen†, Erlangen, Germany
备注:accepted into EUSIPCO 2022
摘要:在某些用于音频源分离的DNN中,相关模型参数与用于训练的音频采样频率无关。考虑到对话分离的应用,这显示了两种DNN体系结构:U网络和完全卷积模型。模型使用8 kHz的音频采样进行训练。学习到的参数被传输到模型中,用于处理48 kHz的音频。将分离的音频源与使用48 kHz版本的相同训练数据训练的相同模型架构产生的音频源进行比较。听力测试和计算测量结果表明,用8 kHz或48 kHz训练的模型之间没有显著的知觉差异。学习参数的这种可转移性允许更快且计算成本更低的训练。它还支持以低于现有应用程序所需的采样频率使用可用的训练数据集,或使用具有多个采样频率的数据采集。
摘要:In some DNNs for audio source separation, the relevant model parameters are independent of the sampling frequency of the audio used for training. Considering the application of dialogue separation, this is shown for two DNN architectures: a U-Net and a fully-convolutional model. The models are trained with audio sampled at 8 kHz. The learned parameters are transferred to models for processing audio at 48 kHz. The separated audio sources are compared with the ones produced by the same model architectures trained with 48 kHz versions of the same training data. A listening test and computational measures show that there is no significant perceptual difference between the models trained with 8 kHz or with 48 kHz. This transferability of the learned parameters allows for a faster and computationally less costly training. It also enables using training datasets available at a lower sampling frequency than the one needed by the application at hand, or using data collections with multiple sampling frequencies.


【13】 STARSS22: A dataset of spatial recordings of real scenes with  spatiotemporal annotations of sound events

标题:STARSS22:具有声音事件时空标注的真实场景空间记录的数据集

链接:https://arxiv.org/abs/2206.01948

作者:Archontis Politis,Kazuki Shimada,Parthasaarathy Sudarsanam,Sharath Adavanne,Daniel Krause,Yuichiro Koyama,Naoya Takahashi,Shusuke Takahashi,Yuki Mitsufuji,Tuomas Virtanen
机构:Audio Research Group, Tampere University, Tampere, Finland,  Sony Group Corporation, Tokyo, Japan
摘要:本报告介绍了Sony TAu Realistic Spatial Soundscapes 2022(STARS22)数据集,用于声音事件定位和检测,包括在两个不同地点的不同内部收集的真实场景的空间记录。数据集由一个高分辨率球形麦克风阵列捕获,并以两种4通道格式提供,一阶环境声学和四面体麦克风阵列。数据集中属于13个目标声音类别的声音事件通过人工注释和光学跟踪的组合在时间和空间上进行注释。该数据集是DCASE2022声音事件定位和检测挑战任务3的开发和评估数据集,与之前基于合成空间化声音场景记录的迭代相比,该数据集为该任务带来了重大的新挑战。详细介绍了数据集规范,包括记录和注释过程、目标类及其存在,以及开发和评估拆分的详细信息。此外,报告还介绍了挑战中伴随数据集的基线系统,重点介绍了与先前迭代基线的差异;即,引入多ACCDOA表示来处理同一类事件的多个同时发生,并支持麦克风阵列格式的其他改进输入功能。基线测试结果表明,通过适当的训练策略,可以在真实的声场记录上获得合理的检测和定位性能。数据集在中可用https://zenodo.org/record/6387880.
摘要:This report presents the Sony-TAu Realistic Spatial Soundscapes 2022 (STARS22) dataset for sound event localization and detection, comprised of spatial recordings of real scenes collected in various interiors of two different sites. The dataset is captured with a high resolution spherical microphone array and delivered in two 4-channel formats, first-order Ambisonics and tetrahedral microphone array. Sound events in the dataset belonging to 13 target sound classes are annotated both temporally and spatially through a combination of human annotation and optical tracking. The dataset serves as the development and evaluation dataset for the Task 3 of the DCASE2022 Challenge on Sound Event Localization and Detection and introduces significant new challenges for the task compared to the previous iterations, which were based on synthetic spatialized sound scene recordings. Dataset specifications are detailed including recording and annotation process, target classes and their presence, and details on the development and evaluation splits. Additionally, the report presents the baseline system that accompanies the dataset in the challenge with emphasis on the differences with the baseline of the previous iterations; namely, introduction of the multi-ACCDOA representation to handle multiple simultaneous occurences of events of the same class, and support for additional improved input features for the microphone array format. Results of the baseline indicate that with a suitable training strategy a reasonable detection and localization performance can be achieved on real sound scene recordings. The dataset is available in https://zenodo.org/record/6387880.


eess.AS音频处理

【1】 Continuous-Time Analog Filters for Audio Edge Intelligence: Review and  Analysis on Design Techniques

标题:用于音频边缘智能的连续时间模拟滤波器:设计技术回顾与分析

链接:https://arxiv.org/abs/2206.02639

作者:Kwantae Kim,Shih-Chii Liu
机构:The authors are with the Institute of Neuroinformatics,  University of Z¨urichand ETH Z¨urich
备注:16 pages, 17 figures
摘要:硅耳蜗设计体现了生物耳蜗的功能。它们的用途已被探索用于人工耳蜗应用,最近又被探索用于边缘音频设备,这些设备需要支持始终在线操作。由于其严格的电源限制带来了一些设计挑战,IC设计师不得不寻找使用低备用电源的解决方案。一种很有希望的生物启发方法是将硅耳蜗的连续时间模拟滤波器通道与一个小内存占用深度神经网络相结合,该神经网络在边缘任务(如关键字定位)上进行训练,从而允许将所有块嵌入IC中。本文从硅耳蜗的原始双四阶滤波器电路开始,回顾了当前边缘音频设备中用作特征提取器的模拟滤波器电路。我们的分析从将基本双二次滤波器解释为双积分环路拓扑开始,并回顾了二阶低通和带通滤波器的设计进展,从基于OTA的结构到基于源极跟随器的结构。我们还推导和分析了小信号传递函数,并讨论了这些滤波器的性能方面。这些不同滤波器配置的分析可应用于其他应用领域,例如采用前端带通滤波器的生物医学设备。
摘要:Silicon cochlea designs capture the functionality of the biological cochlea. Their use has been explored for cochlea prosthesis applications and more recently in edge audio devices which are required to support always-on operation. As their stringent power constraints pose several design challenges, IC designers are forced to look for solutions that use low standby power. One promising bio-inspired approach is to combine the continuous-time analog filter channels of the silicon cochlea with a small memory footprint deep neural network that is trained on edge tasks such as keyword spotting, thereby allowing all blocks to be embedded in an IC. This paper reviews the analog filter circuits used as feature extractors for current edge audio devices, starting with the original biquad filter circuits proposed for the silicon cochlea. Our analysis starts from the interpretation of a basic biquad filter as a two-integrator-loop topology and reviews the progression in the design of second-order low-pass and band-pass filters ranging from OTA-based to source-follower-based architectures. We also derive and analyze the small-signal transfer function and discuss performance aspects of these filters. The analysis of these different filter configurations can be applied to other application domains such as biomedical devices which employ a front-end bandpass filter.


【2】 UTTS: Unsupervised TTS with Conditional Disentangled Sequential  Variational Auto-encoder

标题:UTTS:条件解缠序列变分自动编码器的无监督TTS

链接:https://arxiv.org/abs/2206.02512

作者:Jiachen Lian,Chunlei Zhang,Gopala Krishna Anumanchipalli,Dong Yu
机构:UC Berkeley, EECS, CA, Tencent AI Lab, Bellevue, WA, Gopala K. Anumanchipalli
摘要:在本文中,我们提出了一种新的无监督文本到语音(UTTS)框架,该框架不需要文本-音频对来进行TTS声学建模(AM)。UTTS是从非纠缠语音表征学习的角度开发的多说话人语音合成器。该框架为TTS推理提供了说话人时长模型、音色特征(身份)和内容的灵活选择。我们利用自监督语音表示学习以及语音合成前端技术的最新进展进行系统开发。具体地说,我们利用一个词典将输入文本映射到音素序列,然后将其扩展到框架级的强制对齐(FA),并使用一个与说话人相关的持续时间模型。然后,我们开发了一个对齐映射模块,将FA转换为无监督对齐(UA)。最后,作为自我监督TTS AM的条件分离顺序变分自动编码器(C-DSVAE)将预测的UA和目标说话人嵌入生成mel谱图,并最终使用神经声码器将其转换为波形。我们展示了我们的方法如何在不使用成对TTS语料库的情况下实现语音合成。实验表明,UTTS能够合成出自然度高、可懂度高的语音,并能进行客观评价。
摘要:In this paper, we propose a novel unsupervised text-to-speech (UTTS) framework which does not require text-audio pairs for the TTS acoustic modeling (AM). UTTS is a multi-speaker speech synthesizer developed from the perspective of disentangled speech representation learning. The framework offers a flexible choice of a speaker's duration model, timbre feature (identity) and content for TTS inference. We leverage recent advancements in self-supervised speech representation learning as well as speech synthesis front-end techniques for the system development. Specifically, we utilize a lexicon to map input text to the phoneme sequence, which is expanded to the frame-level forced alignment (FA) with a speaker-dependent duration model. Then, we develop an alignment mapping module that converts the FA to the unsupervised alignment (UA). Finally, a Conditional Disentangled Sequential Variational Auto-encoder (C-DSVAE), serving as the self-supervised TTS AM, takes the predicted UA and a target speaker embedding to generate the mel spectrogram, which is ultimately converted to waveform with a neural vocoder. We show how our method enables speech synthesis without using a paired TTS corpus. Experiments demonstrate that UTTS can synthesize speech of high naturalness and intelligibility measured by human and objective evaluations.


【3】 Online Neural Diarization of Unlimited Numbers of Speakers

标题:无限说话人的在线神经网络二值化

链接:https://arxiv.org/abs/2206.02432

作者:Shota Horiguchi,Shinji Watanabe,Paola Garcia,Yuki Takashima,Yohei Kawaguchi
机构:Watanabe is with Carnegie Mellon University,  Garc´ıa is with Johns Hopkins University
摘要:本文描述了一种对无限数量的说话人进行离线和在线说话人日记的方法。端到端神经二值化(EEND)通过将说话人二值化描述为多标签分类问题,实现了重叠感知说话人二值化。通过引入基于说话人的吸引子,它还扩展到了灵活数量的说话人。然而,基于吸引子的EEND的说话人的输出数量在经验上是有限的;它不能处理推理过程中出现的说话人数量高于训练过程中出现的说话人数量的情况,因为它的说话人计数是在完全监督的方式下训练的。我们的方法EEND-GLA通过在基于吸引子的EEND中引入无监督聚类来解决这个问题。该方法首先将输入音频划分为短块,然后对每个块进行基于吸引子的二值化,最后根据局部计算的吸引子之间的相似度对每个块的结果进行聚类。虽然输出扬声器的数量限制在每个块内,但整个输入估计的扬声器总数可能高于限制。为了以在线方式使用EEND-GLA,我们的方法还扩展了说话人跟踪缓冲区,该缓冲区最初是为了支持传统EEND的在线推理。我们引入了一种分块缓冲区更新,使说话人跟踪缓冲区与EEND-GLA兼容。最后,为了改进在线日志化,我们的方法改进了缓冲区更新方法,并重新讨论了EEND的可变块大小训练。实验结果表明,EEND-GLA可以在离线和在线推理中对未知数量的说话人进行说话人日记。
摘要:A method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a multi-label classification problem. It has also been extended for a flexible number of speakers by introducing speaker-wise attractors. However, the output number of speakers of attractor-based EEND is empirically capped; it cannot deal with cases where the number of speakers appearing during inference is higher than that during training because its speaker counting is trained in a fully supervised manner. Our method, EEND-GLA, solves this problem by introducing unsupervised clustering into attractor-based EEND. In the method, the input audio is first divided into short blocks, then attractor-based diarization is performed for each block, and finally the results of each blocks are clustered on the basis of the similarity between locally-calculated attractors. While the number of output speakers is limited within each block, the total number of speakers estimated for the entire input can be higher than the limitation. To use EEND-GLA in an online manner, our method also extends the speaker-tracing buffer, which was originally proposed to enable online inference of conventional EEND. We introduces a block-wise buffer update to make the speaker-tracing buffer compatible with EEND-GLA. Finally, to improve online diarization, our method improves the buffer update method and revisits the variable chunk-size training of EEND. The experimental results demonstrate that EEND-GLA can perform speaker diarization of an unseen number of speakers in both offline and online inferences.


【4】 Dict-TTS: Learning to Pronounce with Prior Dictionary Knowledge for  Text-to-Speech

标题:DICT-TTS:学习利用先验词典知识进行语音转换的发音

链接:https://arxiv.org/abs/2206.02147

作者:Ziyue Jiang,Su Zhe,Zhou Zhao,Qian Yang,Yi Ren,Jinglin Liu,Zhenhui Ye
机构:Zhejiang University, Zhe Su∗
摘要:多音字消歧旨在从自然文本序列中获取准确的语音知识,以实现可靠的文语转换(TTS)系统。然而,以前的方法需要大量带注释的训练数据和语言专家的额外努力,因此很难将高质量的神经TTS系统扩展到域外的日常对话和世界各地的无数语言。本文从简洁新颖的角度解决了多音字的消歧问题:我们提出了Dict-TTS,这是一种语义感知的生成性文语转换模型,具有在线网站词典(自然语言中现有的先验信息)。具体而言,我们设计了语义到发音注意(S2PA)模块,将输入文本序列与词典中的先验语义之间的语义模式进行匹配,并获得相应的发音;S2PA模块可以使用端到端TTS模型轻松训练,无需任何带注释的音素标签。在三种语言上的实验结果表明,我们的模型在发音准确性方面优于几种强基线模型,并改进了TTS系统的韵律建模。对不同语言编码器的进一步广泛分析表明,Dict-TTS中的每种设计都是有效的。音频示例位于\url{https://dicttts.github.io/DictTTS-Demo/}.
摘要:Polyphone disambiguation aims to capture accurate pronunciation knowledge from natural text sequences for reliable Text-to-speech (TTS) systems. However, previous approaches require substantial annotated training data and additional efforts from language experts, making it difficult to extend high-quality neural TTS systems to out-of-domain daily conversations and countless languages worldwide. This paper tackles the polyphone disambiguation problem from a concise and novel perspective: we propose Dict-TTS, a semantic-aware generative text-to-speech model with an online website dictionary (the existing prior information in the natural language). Specifically, we design a semantics-to-pronunciation attention (S2PA) module to match the semantic patterns between the input text sequence and the prior semantics in the dictionary and obtain the corresponding pronunciations; The S2PA module can be easily trained with the end-to-end TTS model without any annotated phoneme labels. Experimental results in three languages show that our model outperforms several strong baseline models in terms of pronunciation accuracy and improves the prosody modeling of TTS systems. Further extensive analyses with different linguistic encoders demonstrate that each design in Dict-TTS is effective. Audio samples are available at \url{https://dicttts.github.io/DictTTS-Demo/}.


【5】 Geometrically-Motivated Primary-Ambient Decomposition With  Center-Channel Extraction

标题:中心通道提取的几何激励原生环境分解

链接:https://arxiv.org/abs/2206.02125

作者:Jouni Paulus,Matteo Torcoli
机构:Fraunhofer Institute for Integrated Circuits IIS, and, International Audio Laboratories Erlangen†, Erlangen, Germany
备注:accepted into EUSIPCO 2022
摘要:提出了一种几何激励的初级环境分解方法,并在上混合应用中进行了评估。该方法包括两个步骤,并提供了特别直观的解释。第一步包括对输入立体声场景应用信号自适应旋转,将主要声源转换为旋转场景的中心。第二步应用基于简单信号模型的中心信道提取方法,并在均方误差意义下进行优化。通过使用估计的环境分量来评估性能,以实现从真实立体声信号开始的环绕声。在报告的听力测试中,参与者被要求调整音频场景包络,并找到最让他们满意的音频设置。该方法实现上混音的可能性得到了广泛的应用,与原来的立体声混音相比,用户满意度显著提高。
摘要:A geometrically-motivated method for primary-ambient decomposition is proposed and evaluated in an up-mixing application. The method consists of two steps, accommodating a particularly intuitive explanation. The first step consists of signal-adaptive rotations applied on the input stereo scene, which translate the primary sound sources into the center of the rotated scene. The second step applies a center-channel extraction method, based on a simple signal model and optimal in the mean-squared-error sense. The performance is evaluated by using the estimated ambient component to enable surround sound starting from real-world stereo signals. The participants in the reported listening test are asked to adjust the audio scene envelopment and find the audio settings that pleases them the most. The possibility for up-mixing enabled by the proposed method is used extensively, and the user satisfaction is significantly increased compared to the original stereo mix.


【6】 Sampling Frequency Independent Dialogue Separation

标题:采样频率无关的对话分离

链接:https://arxiv.org/abs/2206.02124

作者:Jouni Paulus,Matteo Torcoli
机构:Fraunhofer Institute for Integrated Circuits IIS, and, International Audio Laboratories Erlangen†, Erlangen, Germany
备注:accepted into EUSIPCO 2022
摘要:在某些用于音频源分离的DNN中,相关模型参数与用于训练的音频采样频率无关。考虑到对话分离的应用,这显示了两种DNN体系结构:U网络和完全卷积模型。模型使用8 kHz的音频采样进行训练。学习到的参数被传输到模型中,用于处理48 kHz的音频。将分离的音频源与使用48 kHz版本的相同训练数据训练的相同模型架构产生的音频源进行比较。听力测试和计算测量结果表明,用8 kHz或48 kHz训练的模型之间没有显著的知觉差异。学习参数的这种可转移性允许更快且计算成本更低的训练。它还支持以低于现有应用程序所需的采样频率使用可用的训练数据集,或使用具有多个采样频率的数据采集。
摘要:In some DNNs for audio source separation, the relevant model parameters are independent of the sampling frequency of the audio used for training. Considering the application of dialogue separation, this is shown for two DNN architectures: a U-Net and a fully-convolutional model. The models are trained with audio sampled at 8 kHz. The learned parameters are transferred to models for processing audio at 48 kHz. The separated audio sources are compared with the ones produced by the same model architectures trained with 48 kHz versions of the same training data. A listening test and computational measures show that there is no significant perceptual difference between the models trained with 8 kHz or with 48 kHz. This transferability of the learned parameters allows for a faster and computationally less costly training. It also enables using training datasets available at a lower sampling frequency than the one needed by the application at hand, or using data collections with multiple sampling frequencies.


【7】 STARSS22: A dataset of spatial recordings of real scenes with  spatiotemporal annotations of sound events

标题:STARSS22:具有声音事件时空标注的真实场景空间记录的数据集

链接:https://arxiv.org/abs/2206.01948

作者:Archontis Politis,Kazuki Shimada,Parthasaarathy Sudarsanam,Sharath Adavanne,Daniel Krause,Yuichiro Koyama,Naoya Takahashi,Shusuke Takahashi,Yuki Mitsufuji,Tuomas Virtanen
机构:Audio Research Group, Tampere University, Tampere, Finland,  Sony Group Corporation, Tokyo, Japan
摘要:本报告介绍了Sony TAu Realistic Spatial Soundscapes 2022(STARS22)数据集,用于声音事件定位和检测,包括在两个不同地点的不同内部收集的真实场景的空间记录。数据集由一个高分辨率球形麦克风阵列捕获,并以两种4通道格式提供,一阶环境声学和四面体麦克风阵列。数据集中属于13个目标声音类别的声音事件通过人工注释和光学跟踪的组合在时间和空间上进行注释。该数据集是DCASE2022声音事件定位和检测挑战任务3的开发和评估数据集,与之前基于合成空间化声音场景记录的迭代相比,该数据集为该任务带来了重大的新挑战。详细介绍了数据集规范,包括记录和注释过程、目标类及其存在,以及开发和评估拆分的详细信息。此外,报告还介绍了挑战中伴随数据集的基线系统,重点介绍了与先前迭代基线的差异;即,引入多ACCDOA表示来处理同一类事件的多个同时发生,并支持麦克风阵列格式的其他改进输入功能。基线测试结果表明,通过适当的训练策略,可以在真实的声场记录上获得合理的检测和定位性能。数据集在中可用https://zenodo.org/record/6387880.
摘要:This report presents the Sony-TAu Realistic Spatial Soundscapes 2022 (STARS22) dataset for sound event localization and detection, comprised of spatial recordings of real scenes collected in various interiors of two different sites. The dataset is captured with a high resolution spherical microphone array and delivered in two 4-channel formats, first-order Ambisonics and tetrahedral microphone array. Sound events in the dataset belonging to 13 target sound classes are annotated both temporally and spatially through a combination of human annotation and optical tracking. The dataset serves as the development and evaluation dataset for the Task 3 of the DCASE2022 Challenge on Sound Event Localization and Detection and introduces significant new challenges for the task compared to the previous iterations, which were based on synthetic spatialized sound scene recordings. Dataset specifications are detailed including recording and annotation process, target classes and their presence, and details on the development and evaluation splits. Additionally, the report presents the baseline system that accompanies the dataset in the challenge with emphasis on the differences with the baseline of the previous iterations; namely, introduction of the multi-ACCDOA representation to handle multiple simultaneous occurences of events of the same class, and support for additional improved input features for the microphone array format. Results of the baseline indicate that with a suitable training strategy a reasonable detection and localization performance can be achieved on real sound scene recordings. The dataset is available in https://zenodo.org/record/6387880.


【8】 Canonical Cortical Graph Neural Networks and its Application for Speech  Enhancement in Future Audio-Visual Hearing Aids

标题:规范皮层图神经网络及其在未来视听助听器语音增强中的应用

链接:https://arxiv.org/abs/2206.02671

作者:Leandro A. Passos,João Paulo Papa,Ahsan Adeel
机构:CMI Lab, School of Engineering and Informatics, University of Wolverhampton, Wolverhampton, United Kingdom,  Department of Computing, S˜ao Paulo State University, Bauru, Brazil,  deepCI.org ,, Parkside Terrace, Edinburgh, United Kingdom
摘要:尽管机器学习算法最近取得了成功,但在考虑需要不同来源之间交互的更复杂任务时,大多数模型仍面临一些缺点,例如多模式输入数据和逻辑时序。另一方面,生物大脑在这个意义上是高度敏锐的,经过数百万年的进化,它能够自动管理和整合这样的信息流。在这种背景下,本文从大脑皮层回路的最新发现中得到启发,提出了一种生物学上更合理的自我监督机器学习方法,该方法将使用层内调制的多模态信息与典型相关分析(CCA)相结合,以及一种跟踪时间数据的记忆机制,所谓的典型皮层图神经网络。考虑到更好的干净音频重建和能量效率,该方法的表现优于最近的最新研究结果,这可以通过减少和抑制神经元放电率分布来描述,表明该模型是未来视听助听器中语音增强的合适方法。
摘要:Despite the recent success of machine learning algorithms, most of these models still face several drawbacks when considering more complex tasks requiring interaction between different sources, such as multimodal input data and logical time sequence. On the other hand, the biological brain is highly sharpened in this sense, empowered to automatically manage and integrate such a stream of information through millions of years of evolution. In this context, this paper finds inspiration from recent discoveries on cortical circuits in the brain to propose a more biologically plausible self-supervised machine learning approach that combines multimodal information using intra-layer modulations together with canonical correlation analysis (CCA), as well as a memory mechanism to keep track of temporal data, the so-called Canonical Cortical Graph Neural networks. The approach outperformed recent state-of-the-art results considering both better clean audio reconstruction and energy efficiency, described by a reduced and smother neuron firing rate distribution, suggesting the model as a suitable approach for speech enhancement in future audio-visual hearing aid devices.


【9】 Tagged-MRI2Audio with Attention Guided Heterogeneous Translator

标题:带有注意力引导的异类翻译器的标记-MRI2音频

链接:https://arxiv.org/abs/2206.02284

作者:Xiaofeng Liu,Fangxu Xing,Jerry L. Prince,Jiachen Zhuo,Maureen Stone,Georges El Fakhri,Jonghye Woo
机构:Massachusetts General Hospital and Harvard Medical School, Boston, MA, USA,  Johns Hopkins University, Baltimore, MD, USA,  University of Maryland, Baltimore, MD, USA
备注:MICCAI 2022 (early accept)
摘要:理解标记MRI和可理解语音中舌和口咽肌变形之间的潜在关系,对于推进言语运动控制理论和言语相关疾病的治疗具有重要作用。然而,由于其不同的表现形式,两种模式之间的直接映射——即二维(中矢状切片)加上时间标记的MRI序列及其相应的一维波形——并不简单。相反,我们求助于二维谱图作为中间表示,其中包含基音和共振,从中开发端到端的深度学习框架,以将标记的MRI序列转换为数据集大小有限的相应音频波形。我们的框架基于一种全新的完全卷积不对称翻译器,并在自我剩余注意策略的指导下,专门利用语音中的运动肌肉结构。此外,我们利用具有相同话语的样本的成对关联,并采用潜在空间表示解纠缠策略。此外,我们将对抗训练方法与生成对抗网络相结合,以提高生成光谱图的真实性。我们的实验结果显示,我们的框架能够从标记的MRI序列中生成清晰的音频波形,超过了其他方法。实验结果共有63个标记的MRI序列以及语音声学。
摘要:Understanding the underlying relationship between tongue and oropharyngeal muscle deformation seen in tagged-MRI and intelligible speech plays an important role in advancing speech motor control theories and treatment of speech related-disorders. Because of their heterogeneous representations, however, direct mapping between the two modalities -- i.e., two-dimensional (mid-sagittal slice) plus time tagged-MRI sequence and its corresponding one-dimensional waveform -- is not straightforward. Instead, we resort to two-dimensional spectrograms as an intermediate representation, which contains both pitch and resonance, from which to develop an end-to-end deep learning framework to translate from a sequence of tagged-MRI to its corresponding audio waveform with limited dataset size. Our framework is based on a novel fully convolutional asymmetry translator with guidance of a self residual attention strategy to specifically exploit the moving muscular structures during speech. In addition, we leverage a pairwise correlation of the samples with the same utterances with a latent space representation disentanglement strategy. Furthermore, we incorporate an adversarial training approach with generative adversarial networks to offer improved realism on our generated spectrograms. Our experimental results, carried out with a total of 63 tagged-MRI sequences alongside speech acoustics, showed that our framework enabled the generation of clear audio waveforms from a sequence of tagged-MRI, surpassing competing methods.


【10】 Zero-Shot Voice Conditioning for Denoising Diffusion TTS Models

标题:用于扩散TTS模型去噪的零发声条件

链接:https://arxiv.org/abs/2206.02246

作者:Alon Levkovitch,Eliya Nachmani,Lior Wolf
机构:Tel Aviv University, Facebook AI Research
摘要:我们提出了一种新的方法来调节经过预训练的去噪扩散语音模型,以在训练过程中看不到的新人物的声音中生成语音。该方法需要从目标人员处获得一个短样本(约3秒),并在推理时引导生成,无需任何训练步骤。该方法的核心是一个采样过程,该过程将去噪模型的估计与新说话人样本的低通版本相结合。客观和主观评估表明,我们的采样方法可以生成与目标说话人频率相似的声音,准确度与最先进的方法相当,并且无需训练。
摘要:We present a novel way of conditioning a pretrained denoising diffusion speech model to produce speech in the voice of a novel person unseen during training. The method requires a short (~3 seconds) sample from the target person, and generation is steered at inference time, without any training steps. At the heart of the method lies a sampling process that combines the estimation of the denoising model with a low-pass version of the new speaker's sample. The objective and subjective evaluations show that our sampling method can generate a voice similar to that of the target speaker in terms of frequency, with an accuracy comparable to state-of-the-art methods, and without training.


【11】 Variable-rate hierarchical CPC leads to acoustic unit discovery in  speech

标题:可变速率分层CPC导致语音中的声学单元发现

链接:https://arxiv.org/abs/2206.02211

作者:Santiago Cuervo,Adrian Łańcucki,Ricard Marxer,Paweł Rychlikowski,Jan Chorowski
机构:University of Wrocław, Poland, Adrian Ła´ncucki, NVIDIA, Poland, Université de Toulon, France, NavAlgo, France
备注:Submitted to 36th Conference on Neural Information Processing Systems (NeurIPS 2022)
摘要:深度学习的成功来自于它通过学习由低级表示定义的高级表示来捕获数据层次结构的能力。在本文中,我们通过应用多层次对比预测编码(CPC)来探索语音层次表示的自监督学习。我们观察到,简单地堆叠两个CPC模型并不会比单级架构产生显著的改进。受语音通常被描述为时间上分布不均匀的离散单元序列这一事实的启发,我们提出了一个模型,其中低级CPC模块的输出是非均匀降采样的,以直接最小化高级CPC模块的损失。后者还通过聚焦负采样和预测目标量化来实现连续高层表示的相异性,从而在其表示中实现可分性和离散性先验。根据下游语音识别任务的测量结果,考虑语音信号的结构改进了单级CPC特征,增强了学习表示的分离,同时产生了与电话边界非常相似的有意义的信号分割。
摘要:The success of deep learning comes from its ability to capture the hierarchical structure of data by learning high-level representations defined in terms of low-level ones. In this paper we explore self-supervised learning of hierarchical representations of speech by applying multiple levels of Contrastive Predictive Coding (CPC). We observe that simply stacking two CPC models does not yield significant improvements over single-level architectures. Inspired by the fact that speech is often described as a sequence of discrete units unevenly distributed in time, we propose a model in which the output of a low-level CPC module is non-uniformly downsampled to directly minimize the loss of a high-level CPC module. The latter is designed to also enforce a prior of separability and discreteness in its representations by enforcing dissimilarity of successive high-level representations through focused negative sampling, and by quantization of the prediction targets. Accounting for the structure of the speech signal improves upon single-level CPC features and enhances the disentanglement of the learned representations, as measured by downstream speech recognition tasks, while resulting in a meaningful segmentation of the signal that closely resembles phone boundaries.


【12】 M2FNet: Multi-modal Fusion Network for Emotion Recognition in  Conversation

标题:M2FNet:面向会话情感识别的多模式融合网络

链接:https://arxiv.org/abs/2206.02187

作者:Vishal Chudasama,Purbayan Kar,Ashish Gudmalwar,Nirmesh Shah,Pankaj Wasnik,Naoyuki Onoe
机构:Media Analysis Group, Sony Research India, Bangalore, India
备注:Accepted for publication in the 5th Multimodal Learning and Applications (MULA) Workshop at CVPR 2022
摘要:对话中的情感识别(ERC)对于发展富有同情心的人机交互至关重要。在对话视频中,情感可以以多种形式呈现,即音频、视频和文字记录。然而,由于这些模式的固有特点,多模式ERC一直被认为是一项具有挑战性的任务。现有的ERC研究主要集中在讨论中使用文本信息,而忽略了其他两种方式。我们期望通过采用多模态方法可以提高情绪识别的准确性。因此,在本研究中,我们提出了一种多模态融合网络(M2FNet),从视觉、音频和文本模态中提取情感相关特征。它采用了一种基于多头注意的融合机制,将输入数据中情感丰富的潜在表示结合起来。我们引入了一种新的特征抽取器来从音频和视频模态中提取潜在特征。该特征提取器使用一种新的基于边缘的自适应三重损失函数进行训练,以从音频和视频数据中学习情感相关的特征。在ERC领域,现有的方法在一个基准数据集上表现良好,但在其他数据集上表现不佳。我们的结果表明,在著名的MELD和IEMOCAP数据集上,所提出的M2FNet体系结构在加权平均F1得分方面优于所有其他方法,并在ERC中创造了新的最先进性能。
摘要:Emotion Recognition in Conversations (ERC) is crucial in developing sympathetic human-machine interaction. In conversational videos, emotion can be present in multiple modalities, i.e., audio, video, and transcript. However, due to the inherent characteristics of these modalities, multi-modal ERC has always been considered a challenging undertaking. Existing ERC research focuses mainly on using text information in a discussion, ignoring the other two modalities. We anticipate that emotion recognition accuracy can be improved by employing a multi-modal approach. Thus, in this study, we propose a Multi-modal Fusion Network (M2FNet) that extracts emotion-relevant features from visual, audio, and text modality. It employs a multi-head attention-based fusion mechanism to combine emotion-rich latent representations of the input data. We introduce a new feature extractor to extract latent features from the audio and visual modality. The proposed feature extractor is trained with a novel adaptive margin-based triplet loss function to learn emotion-relevant features from the audio and visual data. In the domain of ERC, the existing methods perform well on one benchmark dataset but not on others. Our results show that the proposed M2FNet architecture outperforms all other methods in terms of weighted average F1 score on well-known MELD and IEMOCAP datasets and sets a new state-of-the-art performance in ERC.


【13】 Learning Speaker-specific Lip-to-Speech Generation

标题:学习特定于说话人的唇语转换生成

链接:https://arxiv.org/abs/2206.02050

作者:Munender Varshney,Ravindra Yadav,Vinay P. Namboodiri,Rajesh M Hegde
机构:† Electrical department, Indian institute of Technology Kanpur, India, ‡ Computer Science and Engineering department, Indian institute of Technology Kanpur, India, §University of Bath, UK
备注:Accepted at ICPR 2022
摘要:众所周知,对于普通人来说,理解嘴唇运动并从中推断出语音是很困难的。准确的唇读任务从说话者的各种线索及其上下文或环境设置中得到帮助。每个说话人都有不同的口音和说话风格,这可以从他们的视觉和言语特征中推断出来。这项工作的目的是在一个不受限制的大词汇量中,了解单个说话人的语音和嘴唇运动序列之间的相关性/映射。我们将帧序列建模为自动编码器设置中Transformer之前的帧序列,并学习利用音频和视频的时间特性的联合嵌入。我们使用深度度量学习学习时间同步,它引导解码器生成与输入嘴唇运动同步的语音。因此,预测后验可以提供说话人说话风格的生成语音。我们在网格和Lip2Wav化学讲座数据集上对我们的模型进行了约束,以评估在无约束自然环境中通过嘴唇运动生成单说话人自然语音的任务。使用各种定性和定量指标以及人工评估进行的广泛评估也表明,我们的方法在几乎所有评估指标上都优于Lip2Wav化学数据集(无约束环境下的大词汇量),并且略优于最先进的网格数据集。
摘要:Understanding the lip movement and inferring the speech from it is notoriously difficult for the common person. The task of accurate lip-reading gets help from various cues of the speaker and its contextual or environmental setting. Every speaker has a different accent and speaking style, which can be inferred from their visual and speech features. This work aims to understand the correlation/mapping between speech and the sequence of lip movement of individual speakers in an unconstrained and large vocabulary. We model the frame sequence as a prior to the transformer in an auto-encoder setting and learned a joint embedding that exploits temporal properties of both audio and video. We learn temporal synchronization using deep metric learning, which guides the decoder to generate speech in sync with input lip movements. The predictive posterior thus gives us the generated speech in speaker speaking style. We have trained our model on the Grid and Lip2Wav Chemistry lecture dataset to evaluate single speaker natural speech generation tasks from lip movement in an unconstrained natural setting. Extensive evaluation using various qualitative and quantitative metrics with human evaluation also shows that our method outperforms the Lip2Wav Chemistry dataset(large vocabulary in an unconstrained setting) by a good margin across almost all evaluation metrics and marginally outperforms the state-of-the-art on GRID dataset.


机器翻译,仅供参考