今天跟大家分享一篇语音相关的论文合集:cs.SD语音10篇,eess.AS音频处理14篇。
【1】 Towards Cross-speaker Reading Style Transfer on Audiobook Dataset
标题:基于有声读物数据集的跨语者阅读风格迁移研究
链接:https://arxiv.org/abs/2208.05359
作者:Xiang Li,Changhe Song,Xianhao Wei,Zhiyong Wu,Jia Jia,Helen Meng机构:Tsinghua-CUHK Joint Research Center for Media Sciences, Technologies and Systems, Shenzhen International Graduate School, Tsinghua University, Shenzhen, China, Department of Systems Engineering and Engineering Management备注:5 pages, 3 figures, accepted to INTERSPEECH 2022摘要:跨说话人风格迁移的目的是从给定的参考语音中提取出可以在任意目标说话人的音色中再现的风格.现有的跨说话人风格迁移方法已经探索了利用话语级风格标记通过全局或局部尺度风格表示来执行风格迁移.然而,有声读物数据集通常同时具有局部韵律和全局风格的特征,本文提出了一种基于语块的跨语者风格模型,该模型能够同时描述有声读物中的语类和局部韵律特征,并且能够在语块层次上对不同语者的风格进行有效的转换,通过用所提出的可切换对抗分类器来解开说话者音色和风格,实验结果表明,该模型能够将特定的阅读风格迁移到新的目标语者。在局部韵律和全局体裁类型预测器的支持下,进一步展示了该方法在多说话人有声读物生成中的潜力。摘要:Cross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers. Existing methods on this topic have explored utilizing utterance-level style labels to perform style transfer via either global or local scale style representations. However, audiobook datasets are typically characterized by both the local prosody and global genre, and are rarely accompanied by utterance-level style labels. Thus, properly transferring the reading style across different speakers remains a challenging task. This paper aims to introduce a chunk-wise multi-scale cross-speaker style model to capture both the global genre and the local prosody in audiobook speeches. Moreover, by disentangling speaker timbre and style with the proposed switchable adversarial classifiers, the extracted reading style is made adaptable to the timbre of different speakers. Experiment results confirm that the model manages to transfer a given reading style to new target speakers. With the support of local prosody and global genre type predictor, the potentiality of the proposed method in multi-speaker audiobook generation is further revealed.
【2】 Controlling Perceived Emotion in Symbolic Music Generation with Monte Carlo Tree Search
标题:基于蒙特卡罗树搜索的符号音乐生成中的情感控制
链接:https://arxiv.org/abs/2208.05162
作者:Lucas N. Ferreira,Lili Mou,Jim Whitehead,Levi H. S. Lelis机构:Alberta Machine Intelligence Institute (Amii), University of Alberta, Computational Media Department, University of California, Santa Cruz备注:Accepted for publication at the 18th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE-22)摘要:本文提出了一种基于蒙特卡罗树搜索的符号音乐生成中情感控制的新方法.我们使用蒙特卡罗树搜索作为解码机制来引导语言模型学习到的概率分布朝向给定的情感.在解码过程的每一步,我们使用树的预测值置信上限-(PUCT)搜索使由情感分类器和鉴别器给出的情感和质量的平均值最大化的序列,我们使用语言模型作为PUCT的策略,使用情感分类器和鉴别器的组合作为PUCT的值函数,我们从搜索过程中创建的节点访问分布中进行采样。我们评估生成的样本的质量,我们还进行了一项用户研究,以评估人类受试者如何感知所生成样本的质量和情感。我们将PUCT与随机双目标波束搜索进行了比较结果表明,PUCT在音乐质量和情感的几乎所有指标上都优于SBBS和CS。摘要:This paper presents a new approach for controlling emotion in symbolic music generation with Monte Carlo Tree Search. We use Monte Carlo Tree Search as a decoding mechanism to steer the probability distribution learned by a language model towards a given emotion. At every step of the decoding process, we use Predictor Upper Confidence for Trees (PUCT) to search for sequences that maximize the average values of emotion and quality as given by an emotion classifier and a discriminator, respectively. We use a language model as PUCT's policy and a combination of the emotion classifier and the discriminator as its value function. To decode the next token in a piece of music, we sample from the distribution of node visits created during the search. We evaluate the quality of the generated samples with respect to human-composed pieces using a set of objective metrics computed directly from the generated samples. We also perform a user study to evaluate how human subjects perceive the generated samples' quality and emotion. We compare PUCT against Stochastic Bi-Objective Beam Search (SBBS) and Conditional Sampling (CS). Results suggest that PUCT outperforms SBBS and CS in almost all metrics of music quality and emotion.
【3】 Subjective Evaluation of Deep Neural Network Based Speech Enhancement Systems in Real-World Conditions
标题:基于深度神经网络的语音增强系统在真实环境中的主观评价
链接:https://arxiv.org/abs/2208.05057
作者:Gaurav Naithani,Kirsi Pietilä,Riitta Niemistö,Erkki Paajanen,Tero Takala,Tuomas Virtanen机构:Audio Research Group, Tampere University, Tampere, Finland, Huawei Tampere Wireless Headset Audio Lab, Tampere, Finland备注:Accepted for publication in IEEE MMSP 2022摘要:将两个低延迟深度神经网络(DNN)的主观评估结果与基于传统维纳滤波器的噪声抑制器的成熟版本进行比较。目标用例是真实世界单通道语音增强应用,例如,通信。包括由加性平稳和非平稳噪声类型组成的真实世界记录。评估分为四个结果:结果表明,与传统的维纳滤波器基线相比,DNN在所有条件下都改善了噪声抑制,而没有显著降低语音质量和噪声透明度,同时保持了比基线更好的语音可懂度。摘要:Subjective evaluation results for two low-latency deep neural networks (DNN) are compared to a matured version of a traditional Wiener-filter based noise suppressor. The target use-case is real-world single-channel speech enhancement applications, e.g., communications. Real-world recordings consisting of additive stationary and non-stationary noise types are included. The evaluation is divided into four outcomes: speech quality, noise transparency, speech intelligibility or listening effort, and noise level w.r.t. speech. It is shown that DNNs improve noise suppression in all conditions in comparison to the traditional Wiener-filter baseline without major degradation in speech quality and noise transparency while maintaining speech intelligibility better than the baseline.
【4】 Generative Data Augmentation Guided by Triplet Loss for Speech Emotion Recognition
标题:语音情感识别中基于三重音缺失的生成式数据增强
链接:https://arxiv.org/abs/2208.04994
作者:Shijun Wang,Hamed Hemati,Jón Guðnason,Damian Borth机构:University of St. Gallen, Switzerland, Reykjavik University, Iceland备注:Published in INTERSPEECH 2022摘要:语音情感识别(SER)是人机交互的关键,但由于两个主要障碍,SER仍然是一个挑战性的问题:数据缺乏和不平衡。许多用于SER的数据集实质上是不平衡的,其中一类(最常见的是中性)的数据话语比其他类的数据话语频繁得多。此外,对于许多现有的口语来说,只有很少的数据资源可用。为了解决这些问题,我们开发了一种由三元组网络指导的基于GAN的增强模型。以改善训练数据不均衡和不足时的误码率性能。1)在高度不平衡的数据集上,我们的增强策略显著提高了SER性能(与基线相比提高了8%的召回率)。2)此外,在一个跨语言的基准测试中,我们用足够多的源语言话语而很少的目标语言话语(在我们的实验中大约为50个)训练模型,我们的增强策略对所有三种目标语言的SER性能都带来了好处。摘要:Speech Emotion Recognition (SER) is crucial for human-computer interaction but still remains a challenging problem because of two major obstacles: data scarcity and imbalance. Many datasets for SER are substantially imbalanced, where data utterances of one class (most often Neutral) are much more frequent than those of other classes. Furthermore, only a few data resources are available for many existing spoken languages. To address these problems, we exploit a GAN-based augmentation model guided by a triplet network, to improve SER performance given imbalanced and insufficient training data. We conduct experiments and demonstrate: 1) With a highly imbalanced dataset, our augmentation strategy significantly improves the SER performance (+8% recall score compared with the baseline). 2) Moreover, in a cross-lingual benchmark, where we train a model with enough source language utterances but very few target language utterances (around 50 in our experiments), our augmentation strategy brings benefits for the SER performance of all three target languages.
【5】 Mathematical Foundations of Complex Tonality
标题:复调性的数学基础
链接:https://arxiv.org/abs/2208.04974
作者:Jeffrey R. Boland,Lane P. Hughston机构:Syndikat LLC, South Santa Fe Avenue, Los Angeles, California , USA, Department of Computing, Goldsmiths University of London, New Cross, London SE,NW, UK摘要:平均律,其中半音以$2^{1/12}的非理性比例调音:1$,最好被看作是一个可行的折衷方案,牺牲了纯粹性以获得灵活性。公正的语调,其中间隔由$2$,$3$和$5$的幂的乘积给出,更自然,但灵活性有限。我们提出了一个新的方案,其中高斯整数的比率形成了一个抽象的音调系统的基础。三音,在公正的音律上如此有问题,由$45:32美元、64美元、45美元、36美元25美元,或25:18$,没有一个是令人满意的,在我们的方案中由复数比率$1 + \rm{i}:1$。由$\tfrac{9}{8}$和$\tfrac{10}{9}$的音程给出的大全音和小全音可以分别因式分解为复数半音的乘积,从而得到大复数半音$\tfrac{3}{4}(1 + \rm{i})$和一个小复半音$\tfrac{1}{3}(3 + \rm{i})$。完全三度,由区间$\tfrac{5}{4}$给出,分解为复数全音的乘积$\tfrac{1}{2}(1 + 2){i}及其复共轭。用这些补充音调来增强,基于高斯素数的幂的乘积的复区间的结果方案非常自然地导致了大尺度和小尺度的完整系统的构造在所有键中。摘要:Equal temperament, in which semitones are tuned in the irrational ratio of $2^{1/12} : 1$, is best seen as a serviceable compromise, sacrificing purity for flexibility. Just intonation, in which intervals given by products of powers of $2$, $3$, and $5$, is more natural, but of limited flexibility. We propose a new scheme in which ratios of Gaussian integers form the basis of an abstract tonal system. The tritone, so problematic in just temperament, given ambiguously by $45:32$, $64:45$, $36:25$, or $25:18$, none satisfactory, is in our scheme represented by the complex ratio $1 + \rm{i} : 1$. The major and minor whole tones, given by intervals of $\tfrac{9}{8}$ and $\tfrac{10}{9}$, can each be factorized into products of complex semitones, giving us a major complex semitone $\tfrac{3}{4}(1 + \rm{i})$ and a minor complex semitone $\tfrac{1}{3}(3 + \rm{i})$. The perfect third, given by the interval $\tfrac{5}{4}$, factorizes into the product of a complex whole tone $\tfrac{1}{2}(1 + 2\rm{i})$ and its complex conjugate. Augmented with these supplementary tones, the resulting scheme of complex intervals based on products of powers of Gaussian primes leads very naturally to the construction of a complete system of major and minor scales in all keys.
【6】 Preserving the beamforming effect for spatial cue-based pseudo-binaural dereverberation of a single source
标题:保持波束形成效果的基于空间线索的单源伪双耳去混响
链接:https://arxiv.org/abs/2208.05184
作者:Sania Gul,Muhammad Salman Khan,Syed Waqar Shah机构:Department of Electrical Engineering, University of Engineering and Technology Peshawar, Pakistan., Department of Electrical Engineering, College of Engineering, Qatar University, Doha, Qatar.备注:25 pages, 7 figures摘要:混响在封闭空间中是不可避免的,它不仅降低了听力受损者和非母语听力正常者的可懂度,而且也降低了机器听音的性能.本文提出了一种新的双耳去混响方法,使用直接路径信号和混响的耳间线索的差异。两个波束形成器,以耳间距离隔开,由这些混响产生的耳间线索和由直接路径信号产生的耳间线索充当两类数据集,用于U-Net的培训在其训练之后,波束形成器被移除,并且训练的U网络连同最大似然估计一起(MLE)算法用于区分直接路径线索和混响线索,当系统暴露于混响语音信号的耳间谱图时,我们提出的模型在倒谱距离方面优于经典的信号处理去混响模型的加权预测误差(CEP),频率加权分段信噪比(FWSEGSNR)和信号与混响调制能量比(SRMR)分别提高了1.4个百分点、8dB和0.6dB。与基于深度学习的去混响模型相比,该模型在与FWSEGSNR相当的情况下,CEP性能提高了1.3个百分点,所提出的模型在相对类似的不可见声学条件下以及在其训练位置附近的位置处也维持其性能。摘要:Reverberations are unavoidable in enclosures, resulting in reduced intelligibility for hearing impaired and non native listeners and even for the normal hearing listeners in noisy circumstances. It also degrades the performance of machine listening applications. In this paper, we propose a novel approach of binaural dereverberation of a single speech source, using the differences in the interaural cues of the direct path signal and the reverberations. Two beamformers, spaced at an interaural distance, are used to extract the reverberations from the reverberant speech. The interaural cues generated by these reverberations and those generated by the direct path signal act as a two class dataset, used for the training of U-Net (a deep convolutional neural network). After its training, the beamformers are removed and the trained U-Net along with the maximum likelihood estimation (MLE) algorithm is used to discriminate between the direct path cues from the reverberation cues, when the system is exposed to the interaural spectrogram of the reverberant speech signal. Our proposed model has outperformed the classical signal processing dereverberation model weighted prediction error in terms of cepstral distance (CEP), frequency weighted segmental signal to noise ratio (FWSEGSNR) and signal to reverberation modulation energy ratio (SRMR) by 1.4 points, 8 dB and 0.6dB. It has achieved better performance than the deep learning based dereverberation model by gaining 1.3 points improvement in CEP with comparable FWSEGSNR, using training dataset which is almost 8 times smaller than required for that model. The proposed model also sustained its performance under relatively similar unseen acoustic conditions and at positions in the vicinity of its training position.
【7】 ROC: A New Paradigm for Lyric-to-Melody Generation
标题:ROC:歌词到旋律生成的新范式
链接:https://arxiv.org/abs/2208.05697
作者:Ang Lv,Xu Tan,Tao Qin,Tie-Yan Liu,Rui Yan机构:Renmin University of China, Beijing, China, Microsoft Research Asia摘要:Lyric-to-melody generation is an important task in songwriting, and is also quite challenging due to its distinctive characteristics: the generated melodies should not only follow good musical patterns, but also align with features in lyrics such as rhythms and structures. These characteristics cannot be well handled by neural generation models that learn lyric-to-melody mapping in an end-to-end way, due to several issues: (1) lack of aligned lyric-melody training data to sufficiently learn lyric-melody feature alignment; (2) lack of controllability in generation to explicitly guarantee the lyric-melody feature alignment. In this paper, we propose ROC, a new paradigm for lyric-to-melody generation that addresses the above issues through a generation-retrieval pipeline. Specifically, our paradigm has two stages: (1) creation stage, where a huge amount of music pieces are generated by a neural-based melody language model and indexed in a database through several key features (e.g., chords, tonality, rhythm, and structural information including chorus or verse); (2) re-creation stage, where melodies are recreated by retrieving music pieces from the database according to the key features from lyrics and concatenating best music pieces based on composition guidelines and melody language model scores. Our ROC paradigm has several advantages: (1) It only needs unpaired melody data to train melody language model, instead of paired lyric-melody data in previous models. (2) It achieves good lyric-melody feature alignment in lyric-to-melody generation. Experiments on English and Chinese datasets demonstrate that ROC outperforms previous neural based lyric-to-melody generation models on both objective and subjective metrics.
【8】 Symbolic Music Loop Generation with Neural Discrete Representations
标题:用神经离散表示法生成符号音乐环路
链接:https://arxiv.org/abs/2208.05605
作者:Sangjun Han,Hyeongrae Ihm,Moontae Lee,Woohyung Lim机构:LG AI Research, University of Illinois at Chicago摘要:Since most of music has repetitive structures from motifs to phrases, repeating musical ideas can be a basic operation for music composition. The basic block that we focus on is conceptualized as loops which are essential ingredients of music. Furthermore, meaningful note patterns can be formed in a finite space, so it is sufficient to represent them with combinations of discrete symbols as done in other domains. In this work, we propose symbolic music loop generation via learning discrete representations. We first extract loops from MIDI datasets using a loop detector and then learn an autoregressive model trained by discrete latent codes of the extracted loops. We show that our model outperforms well-known music generative models in terms of both fidelity and diversity, evaluating on random space. Our code and supplementary materials are available at https://github.com/sjhan91/Loop_VQVAE_Official.
【9】 Speech Enhancement and Dereverberation with Diffusion-based Generative Models
标题:基于扩散的生成模型的语音增强和去混响
链接:https://arxiv.org/abs/2208.05830
作者:Julius Richter,Simon Welker,Jean-Marie Lemercier,Bunlong Lay,Timo Gerkmann机构:Student Member, IEEE, Jean-Marie, Lemercier, Senior Member, IEEE备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible摘要:Recently, diffusion-based generative models have been introduced to the task of speech enhancement. The corruption of clean speech is modeled as a fixed forward process in which increasing amounts of noise are gradually added. By learning to reverse this process in an iterative fashion conditioned on the noisy input, clean speech is generated. We build upon our previous work and derive the training task within the formalism of stochastic differential equations. We present a detailed theoretical review of the underlying score matching objective and explore different sampler configurations for solving the reverse process at test time. By using a sophisticated network architecture from natural image generation literature, we significantly improve performance compared to our previous publication. We also show that we can compete with recent discriminative models and achieve better generalization when evaluating on a different corpus than used for training. We complement the evaluation results with a subjective listening test, in which our proposed method is rated best. Furthermore, we show that the proposed method achieves remarkable state-of-the-art performance in single-channel speech dereverberation. Our code and audio examples are available online, see https://uhh.de/inf-sp-sgmse
【10】 Comparison and Analysis of New Curriculum Criteria for End-to-End ASR
标题:端到端ASR新课程标准的比较分析
链接:https://arxiv.org/abs/2208.05782
作者:Georgios Karakasidis,Tamás Grósz,Mikko Kurimo机构:Department of Signal Processing and Acoustics, Aalto University, Finland备注:5 pages, 2 figures, in Proceedings Interspeech 2022摘要:It is common knowledge that the quantity and quality of the training data play a significant role in the creation of a good machine learning model. In this paper, we take it one step further and demonstrate that the way the training examples are arranged is also of crucial importance. Curriculum Learning is built on the observation that organized and structured assimilation of knowledge has the ability to enable faster training and better comprehension. When humans learn to speak, they first try to utter basic phones and then gradually move towards more complex structures such as words and sentences. This methodology is known as Curriculum Learning, and we employ it in the context of Automatic Speech Recognition. We hypothesize that end-to-end models can achieve better performance when provided with an organized training set consisting of examples that exhibit an increasing level of difficulty (i.e. a curriculum). To impose structure on the training set and to define the notion of an easy example, we explored multiple scoring functions that either use feedback from an external neural network or incorporate feedback from the model itself. Empirical results show that with different curriculums we can balance the training times and the network's performance.
【1】 Non-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech
标题:基于非对比自监督学习的语音信息提取
链接:https://arxiv.org/abs/2208.05445
作者:Jaejin Cho,Jes'us Villalba,Laureano Moro-Velazquez,Najim Dehak机构:an increasing number of studies have started topropose new techniques to exploit unlabeled data during theAll authors are associated with the Department of Electrical and ComputerEngineering, Johns Hopkins University备注:EARLY ACCESS of IEEE JSTSP Special Issue on Self-Supervised Learning for Speech and Audio Processing摘要:在最近的研究中,自监督预训练模型在迁移学习中的表现往往优于有监督预训练模型,特别是,话语级语音表示的自监督学习(SSL)可用于要求话语中一致属性的区别性表示的语音应用:说话者、语言、情感和年龄。现有的帧级自监督语音表示,例如wav2vec,可以用作具有池的话语级表示,但是模型通常很大。也有SSL技术来学习话语级表示。最成功的一种是对比方法,其需要负采样:本文提出了一种非对比的自监督方法来学习话语级嵌入,将计算机视觉中的DINO(DIstillation with NO labels)算法应用到语音中,与对比方法不同的是,DINO不需要负采样,我们将DINO与以监督方式训练的x向量进行比较,当转移到下游任务(说话人确认、语音情感识别(SER)和阿尔茨海默病检测)时,我们研究了迁移学习过程中的几个方面的影响,如微调过程的步骤划分、组块长度、扩充等,首先调整最后一个仿射层,然后调整整个网络,这一点超过了一次微调。使用较短的块长度,虽然它们会生成更多的不同输入,但并不一定会提高性能,这意味着每个应用程序需要至少具有特定长度的语音段才能获得更好的性能。增强在SER中很有帮助。摘要:In recent studies, self-supervised pre-trained models tend to outperform supervised pre-trained models in transfer learning. In particular, self-supervised learning (SSL) of utterance-level speech representation can be used in speech applications that require discriminative representation of consistent attributes within an utterance: speaker, language, emotion, and age. Existing frame-level self-supervised speech representation, e.g., wav2vec, can be used as utterance-level representation with pooling, but the models are usually large. There are also SSL techniques to learn utterance-level representation. One of the most successful is a contrastive method, which requires negative sampling: selecting alternative samples to contrast with the current sample (anchor). However, this does not ensure that all the negative samples belong to classes different from the anchor class without labels. This paper applies a non-contrastive self-supervised method to learn utterance-level embeddings. We adapted DIstillation with NO labels (DINO) from computer vision to speech. Unlike contrastive methods, DINO does not require negative sampling. We compared DINO to x-vector trained in a supervised manner. When transferred to down-stream tasks (speaker verification, speech emotion recognition (SER), and Alzheimer's disease detection), DINO outperformed x-vector. We studied the influence of several aspects during transfer learning such as dividing the fine-tuning process into steps, chunk lengths, or augmentation. During fine-tuning, tuning the last affine layers first and then the whole network surpassed fine-tuning all at once. Using shorter chunk lengths, although they generate more diverse inputs, did not necessarily improve performance, implying speech segments at least with a specific length are required for better performance per application. Augmentation was helpful in SER.
【2】 Non-Contrastive Self-Supervised Learning of Utterance-Level Speech Representations
标题:话语级语音表征的非对比自监督学习
链接:https://arxiv.org/abs/2208.05413
作者:Jaejin Cho,Raghavendra Pappagari,Piotr Żelasko,Laureano Moro-Velazquez,Jesús Villalba,Najim Dehak机构:Center for Language and Speech Processing, Johns Hopkins University, Baltimore, MD, USA, Human Language Technology Center of Excellence, Johns Hopkins University, Baltimore, MD备注:Accepted at Interspeech 2022摘要:考虑到大量的未标记语音数据和高标记成本,无监督学习方法对于更好的系统开发是必不可少的。最成功的方法之一是对比自监督方法,其需要负采样:对备选样本进行采样以与当前样本进行对比本文在一个无标注的语音语料库上应用一种非对比的自监督学习方法来学习话语级嵌入,我们使用了无标注的DIstillation方法来学习话语级嵌入,并在此基础上提出了一种基于非对比的自监督学习方法来学习话语级嵌入的算法(DINO),并将其应用于语音领域.与对比方法不同,DINO方法不需要负采样.在说话人确认和情感识别中对这些嵌入方法进行了评估.在说话人确认中,具有余弦评分的无监督DINO嵌入在VoxCeleb1测试试验中提供了4.38%的EER。这在EER方面比最好的对比自监督方法高出40%。迭代伪标记训练流水线,不需要说话者标记,在情感识别方面,DINO嵌入算法在IEMOCAP、Crema-D和MSP-Podcast上的micro-f1得分分别为60.87,79.21和56.98%,实验结果表明DINO嵌入算法对不同语音应用具有通用性。摘要:Considering the abundance of unlabeled speech data and the high labeling costs, unsupervised learning methods can be essential for better system development. One of the most successful methods is contrastive self-supervised methods, which require negative sampling: sampling alternative samples to contrast with the current sample (anchor). However, it is hard to ensure if all the negative samples belong to classes different from the anchor class without labels. This paper applies a non-contrastive self-supervised learning method on an unlabeled speech corpus to learn utterance-level embeddings. We used DIstillation with NO labels (DINO), proposed in computer vision, and adapted it to the speech domain. Unlike the contrastive methods, DINO does not require negative sampling. These embeddings were evaluated on speaker verification and emotion recognition. In speaker verification, the unsupervised DINO embedding with cosine scoring provided 4.38% EER on the VoxCeleb1 test trial. This outperforms the best contrastive self-supervised method by 40% relative in EER. An iterative pseudo-labeling training pipeline, not requiring speaker labels, further improved the EER to 1.89%. In emotion recognition, the DINO embedding performed 60.87, 79.21, and 56.98% in micro-f1 score on IEMOCAP, Crema-D, and MSP-Podcast, respectively. The results imply the generality of the DINO embedding to different speech applications.
【3】 Preserving the beamforming effect for spatial cue-based pseudo-binaural dereverberation of a single source
标题:保持波束形成效果的基于空间线索的单源伪双耳去混响
链接:https://arxiv.org/abs/2208.05184
* 与cs.SD语音【6】为同一篇
作者:Sania Gul,Muhammad Salman Khan,Syed Waqar Shah机构:Department of Electrical Engineering, University of Engineering and Technology Peshawar, Pakistan., Department of Electrical Engineering, College of Engineering, Qatar University, Doha, Qatar.备注:25 pages, 7 figures摘要:混响在封闭空间中是不可避免的,它不仅降低了听力受损者和非母语听力正常者的可懂度,而且也降低了机器听音的性能.本文提出了一种新的双耳去混响方法,使用直接路径信号和混响的耳间线索的差异。两个波束形成器,以耳间距离隔开,由这些混响产生的耳间线索和由直接路径信号产生的耳间线索充当两类数据集,用于U-Net的培训在其训练之后,波束形成器被移除,并且训练的U网络连同最大似然估计一起(MLE)算法用于区分直接路径线索和混响线索,当系统暴露于混响语音信号的耳间谱图时,我们提出的模型在倒谱距离方面优于经典的信号处理去混响模型的加权预测误差(CEP),频率加权分段信噪比(FWSEGSNR)和信号与混响调制能量比(SRMR)分别提高了1.4个百分点、8dB和0.6dB。与基于深度学习的去混响模型相比,该模型在与FWSEGSNR相当的情况下,CEP性能提高了1.3个百分点,所提出的模型在相对类似的不可见声学条件下以及在其训练位置附近的位置处也维持其性能。摘要:Reverberations are unavoidable in enclosures, resulting in reduced intelligibility for hearing impaired and non native listeners and even for the normal hearing listeners in noisy circumstances. It also degrades the performance of machine listening applications. In this paper, we propose a novel approach of binaural dereverberation of a single speech source, using the differences in the interaural cues of the direct path signal and the reverberations. Two beamformers, spaced at an interaural distance, are used to extract the reverberations from the reverberant speech. The interaural cues generated by these reverberations and those generated by the direct path signal act as a two class dataset, used for the training of U-Net (a deep convolutional neural network). After its training, the beamformers are removed and the trained U-Net along with the maximum likelihood estimation (MLE) algorithm is used to discriminate between the direct path cues from the reverberation cues, when the system is exposed to the interaural spectrogram of the reverberant speech signal. Our proposed model has outperformed the classical signal processing dereverberation model weighted prediction error in terms of cepstral distance (CEP), frequency weighted segmental signal to noise ratio (FWSEGSNR) and signal to reverberation modulation energy ratio (SRMR) by 1.4 points, 8 dB and 0.6dB. It has achieved better performance than the deep learning based dereverberation model by gaining 1.3 points improvement in CEP with comparable FWSEGSNR, using training dataset which is almost 8 times smaller than required for that model. The proposed model also sustained its performance under relatively similar unseen acoustic conditions and at positions in the vicinity of its training position.
【4】 Improving Hypernasality Estimation with Automatic Speech Recognition in Cleft Palate Speech
标题:自动语音识别技术在腭裂语音中的应用
链接:https://arxiv.org/abs/2208.05122
作者:Kaitao Song,Teng Wan,Bixia Wang,Huiqiang Jiang,Luna Qiu,Jiahang Xu,Liping Jiang,Qun Lou,Yuqing Yang,Dongsheng Li,Xudong Wang,Lili Qiu机构:Microsoft Research, Department of Oral and Craniomaxillofacial Surgery, Shanghai Ninth People’s Hospital, Shanghai Jiao Tong University School of Medicine备注:Accepted by InterSpeech 2022摘要:高鼻音是人类语音产生过程中的一种异常共振现象,尤其是在腭裂等颅面畸形患者中,在临床应用中,高鼻音的估计对于腭裂的诊断至关重要,其结果决定了后续的手术和额外的语音治疗,因此,设计一种自动鼻音过度评估方法将有助于语言病理学家做出精确的诊断。针对现有的高鼻音估计方法只能在低资源的腭裂数据集上利用统计或神经网络特征进行声学分析的不足,提出了一种基于自动语音识别模型的高鼻音估计方法,我们首先在自动语音识别中预训练编码器—解码器框架,(ASR)目标,然后在腭裂数据集上对ASR编码器进行微调,以实现高鼻音估计。我们的高鼻音估计模型具有ASR模型的优点:1)与低资源的腭裂语音数据集相比,ASR任务通常包含大规模的一般域语音数据,这使得模型泛化能力更好;在两个腭裂数据集上的实验结果表明,该方法与已有的方法相比具有更好的性能。摘要:Hypernasality is an abnormal resonance in human speech production, especially in patients with craniofacial anomalies such as cleft palate. In clinical application, hypernasality estimation is crucial in cleft palate diagnosis, as its results determine the subsequent surgery and additional speech therapy. Therefore, designing an automatic hypernasality assessment method will facilitate speech-language pathologists to make precise diagnoses. Existing methods for hypernasality estimation only conduct acoustic analysis based on low-resource cleft palate dataset, by using statistical or neural network-based features. In this paper, we propose a novel approach that uses automatic speech recognition model to improve hypernasality estimation. Specifically, we first pre-train an encoder-decoder framework in an automatic speech recognition (ASR) objective by using speech-to-text dataset, and then fine-tune ASR encoder on the cleft palate dataset for hypernasality estimation. Benefiting from such design, our model for hypernasality estimation can enjoy the advantages of ASR model: 1) compared with low-resource cleft palate dataset, the ASR task usually includes large-scale speech data in the general domain, which enables better model generalization; 2) the text annotations in ASR dataset guide model to extract better acoustic features. Experimental results on two cleft palate datasets demonstrate that our method achieves superior performance compared with previous approaches.
【5】 Towards Cross-speaker Reading Style Transfer on Audiobook Dataset
标题:基于有声读物数据集的跨语者阅读风格迁移研究
链接:https://arxiv.org/abs/2208.05359
* 与cs.SD语音【1】为同一篇
作者:Xiang Li,Changhe Song,Xianhao Wei,Zhiyong Wu,Jia Jia,Helen Meng机构:Tsinghua-CUHK Joint Research Center for Media Sciences, Technologies and Systems, Shenzhen International Graduate School, Tsinghua University, Shenzhen, China, Department of Systems Engineering and Engineering Management备注:5 pages, 3 figures, accepted to INTERSPEECH 2022摘要:跨说话人风格迁移的目的是从给定的参考语音中提取出可以在任意目标说话人的音色中再现的风格.现有的跨说话人风格迁移方法已经探索了利用话语级风格标记通过全局或局部尺度风格表示来执行风格迁移.然而,有声读物数据集通常同时具有局部韵律和全局风格的特征,本文提出了一种基于语块的跨语者风格模型,该模型能够同时描述有声读物中的语类和局部韵律特征,并且能够在语块层次上对不同语者的风格进行有效的转换,通过用所提出的可切换对抗分类器来解开说话者音色和风格,实验结果表明,该模型能够将特定的阅读风格迁移到新的目标语者。在局部韵律和全局体裁类型预测器的支持下,进一步展示了该方法在多说话人有声读物生成中的潜力。摘要:Cross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers. Existing methods on this topic have explored utilizing utterance-level style labels to perform style transfer via either global or local scale style representations. However, audiobook datasets are typically characterized by both the local prosody and global genre, and are rarely accompanied by utterance-level style labels. Thus, properly transferring the reading style across different speakers remains a challenging task. This paper aims to introduce a chunk-wise multi-scale cross-speaker style model to capture both the global genre and the local prosody in audiobook speeches. Moreover, by disentangling speaker timbre and style with the proposed switchable adversarial classifiers, the extracted reading style is made adaptable to the timbre of different speakers. Experiment results confirm that the model manages to transfer a given reading style to new target speakers. With the support of local prosody and global genre type predictor, the potentiality of the proposed method in multi-speaker audiobook generation is further revealed.
【6】 Controlling Perceived Emotion in Symbolic Music Generation with Monte Carlo Tree Search
标题:基于蒙特卡罗树搜索的符号音乐生成中的情感控制
链接:https://arxiv.org/abs/2208.05162
* 与cs.SD语音【2】为同一篇
作者:Lucas N. Ferreira,Lili Mou,Jim Whitehead,Levi H. S. Lelis机构:Alberta Machine Intelligence Institute (Amii), University of Alberta, Computational Media Department, University of California, Santa Cruz备注:Accepted for publication at the 18th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE-22)摘要:本文提出了一种基于蒙特卡罗树搜索的符号音乐生成中情感控制的新方法.我们使用蒙特卡罗树搜索作为解码机制来引导语言模型学习到的概率分布朝向给定的情感.在解码过程的每一步,我们使用树的预测值置信上限-(PUCT)搜索使由情感分类器和鉴别器给出的情感和质量的平均值最大化的序列,我们使用语言模型作为PUCT的策略,使用情感分类器和鉴别器的组合作为PUCT的值函数,我们从搜索过程中创建的节点访问分布中进行采样。我们评估生成的样本的质量,我们还进行了一项用户研究,以评估人类受试者如何感知所生成样本的质量和情感。我们将PUCT与随机双目标波束搜索进行了比较结果表明,PUCT在音乐质量和情感的几乎所有指标上都优于SBBS和CS。摘要:This paper presents a new approach for controlling emotion in symbolic music generation with Monte Carlo Tree Search. We use Monte Carlo Tree Search as a decoding mechanism to steer the probability distribution learned by a language model towards a given emotion. At every step of the decoding process, we use Predictor Upper Confidence for Trees (PUCT) to search for sequences that maximize the average values of emotion and quality as given by an emotion classifier and a discriminator, respectively. We use a language model as PUCT's policy and a combination of the emotion classifier and the discriminator as its value function. To decode the next token in a piece of music, we sample from the distribution of node visits created during the search. We evaluate the quality of the generated samples with respect to human-composed pieces using a set of objective metrics computed directly from the generated samples. We also perform a user study to evaluate how human subjects perceive the generated samples' quality and emotion. We compare PUCT against Stochastic Bi-Objective Beam Search (SBBS) and Conditional Sampling (CS). Results suggest that PUCT outperforms SBBS and CS in almost all metrics of music quality and emotion.
【7】 Subjective Evaluation of Deep Neural Network Based Speech Enhancement Systems in Real-World Conditions
标题:基于深度神经网络的语音增强系统在真实环境中的主观评价
链接:https://arxiv.org/abs/2208.05057
* 与cs.SD语音【3】为同一篇
作者:Gaurav Naithani,Kirsi Pietilä,Riitta Niemistö,Erkki Paajanen,Tero Takala,Tuomas Virtanen机构:Audio Research Group, Tampere University, Tampere, Finland, Huawei Tampere Wireless Headset Audio Lab, Tampere, Finland备注:Accepted for publication in IEEE MMSP 2022摘要:Subjective evaluation results for two low-latency deep neural networks (DNN) are compared to a matured version of a traditional Wiener-filter based noise suppressor. The target use-case is real-world single-channel speech enhancement applications, e.g., communications. Real-world recordings consisting of additive stationary and non-stationary noise types are included. The evaluation is divided into four outcomes: speech quality, noise transparency, speech intelligibility or listening effort, and noise level w.r.t. speech. It is shown that DNNs improve noise suppression in all conditions in comparison to the traditional Wiener-filter baseline without major degradation in speech quality and noise transparency while maintaining speech intelligibility better than the baseline.
【8】 Generative Data Augmentation Guided by Triplet Loss for Speech Emotion Recognition
标题:语音情感识别中基于三重音缺失的生成式数据增强
链接:https://arxiv.org/abs/2208.04994
* 与cs.SD语音【4】为同一篇
作者:Shijun Wang,Hamed Hemati,Jón Guðnason,Damian Borth机构:University of St. Gallen, Switzerland, Reykjavik University, Iceland备注:Published in INTERSPEECH 2022摘要:Speech Emotion Recognition (SER) is crucial for human-computer interaction but still remains a challenging problem because of two major obstacles: data scarcity and imbalance. Many datasets for SER are substantially imbalanced, where data utterances of one class (most often Neutral) are much more frequent than those of other classes. Furthermore, only a few data resources are available for many existing spoken languages. To address these problems, we exploit a GAN-based augmentation model guided by a triplet network, to improve SER performance given imbalanced and insufficient training data. We conduct experiments and demonstrate: 1) With a highly imbalanced dataset, our augmentation strategy significantly improves the SER performance (+8% recall score compared with the baseline). 2) Moreover, in a cross-lingual benchmark, where we train a model with enough source language utterances but very few target language utterances (around 50 in our experiments), our augmentation strategy brings benefits for the SER performance of all three target languages.
【9】 Mathematical Foundations of Complex Tonality
标题:复调性的数学基础
链接:https://arxiv.org/abs/2208.04974
* 与cs.SD语音【5】为同一篇
作者:Jeffrey R. Boland,Lane P. Hughston机构:Syndikat LLC, South Santa Fe Avenue, Los Angeles, California , USA, Department of Computing, Goldsmiths University of London, New Cross, London SE,NW, UK摘要:平均律,其中半音以$2^{1/12}的非理性比例调音:1$,最好被看作是一个可行的折衷方案,牺牲了纯粹性以获得灵活性。公正的语调,其中间隔由$2$,$3$和$5$的幂的乘积给出,更自然,但灵活性有限。我们提出了一个新的方案,其中高斯整数的比率形成了一个抽象的音调系统的基础。三音,在公正的音律上如此有问题,由$45:32美元、64美元、45美元、36美元25美元,或25:18$,没有一个是令人满意的,在我们的方案中由复数比率$1 + \rm{i}:1$。由$\tfrac{9}{8}$和$\tfrac{10}{9}$的音程给出的大全音和小全音可以分别因式分解为复数半音的乘积,从而得到大复数半音$\tfrac{3}{4}(1 + \rm{i})$和一个小复半音$\tfrac{1}{3}(3 + \rm{i})$。完全三度,由区间$\tfrac{5}{4}$给出,分解为复数全音的乘积$\tfrac{1}{2}(1 + 2){i}及其复共轭。用这些补充音调来增强,基于高斯素数的幂的乘积的复区间的结果方案非常自然地导致了大尺度和小尺度的完整系统的构造在所有键中。摘要:Equal temperament, in which semitones are tuned in the irrational ratio of $2^{1/12} : 1$, is best seen as a serviceable compromise, sacrificing purity for flexibility. Just intonation, in which intervals given by products of powers of $2$, $3$, and $5$, is more natural, but of limited flexibility. We propose a new scheme in which ratios of Gaussian integers form the basis of an abstract tonal system. The tritone, so problematic in just temperament, given ambiguously by $45:32$, $64:45$, $36:25$, or $25:18$, none satisfactory, is in our scheme represented by the complex ratio $1 + \rm{i} : 1$. The major and minor whole tones, given by intervals of $\tfrac{9}{8}$ and $\tfrac{10}{9}$, can each be factorized into products of complex semitones, giving us a major complex semitone $\tfrac{3}{4}(1 + \rm{i})$ and a minor complex semitone $\tfrac{1}{3}(3 + \rm{i})$. The perfect third, given by the interval $\tfrac{5}{4}$, factorizes into the product of a complex whole tone $\tfrac{1}{2}(1 + 2\rm{i})$ and its complex conjugate. Augmented with these supplementary tones, the resulting scheme of complex intervals based on products of powers of Gaussian primes leads very naturally to the construction of a complete system of major and minor scales in all keys.
【10】 Speech Enhancement and Dereverberation with Diffusion-based Generative Models
标题:基于扩散的生成模型的语音增强和去混响
链接:https://arxiv.org/abs/2208.05830
* 与cs.SD语音【9】为同一篇
作者:Julius Richter,Simon Welker,Jean-Marie Lemercier,Bunlong Lay,Timo Gerkmann机构:Student Member, IEEE, Jean-Marie, Lemercier, Senior Member, IEEE备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible摘要:Recently, diffusion-based generative models have been introduced to the task of speech enhancement. The corruption of clean speech is modeled as a fixed forward process in which increasing amounts of noise are gradually added. By learning to reverse this process in an iterative fashion conditioned on the noisy input, clean speech is generated. We build upon our previous work and derive the training task within the formalism of stochastic differential equations. We present a detailed theoretical review of the underlying score matching objective and explore different sampler configurations for solving the reverse process at test time. By using a sophisticated network architecture from natural image generation literature, we significantly improve performance compared to our previous publication. We also show that we can compete with recent discriminative models and achieve better generalization when evaluating on a different corpus than used for training. We complement the evaluation results with a subjective listening test, in which our proposed method is rated best. Furthermore, we show that the proposed method achieves remarkable state-of-the-art performance in single-channel speech dereverberation. Our code and audio examples are available online, see https://uhh.de/inf-sp-sgmse
【11】 Comparison and Analysis of New Curriculum Criteria for End-to-End ASR
标题:端到端ASR新课程标准的比较分析
链接:https://arxiv.org/abs/2208.05782
* 与cs.SD语音【10】为同一篇
作者:Georgios Karakasidis,Tamás Grósz,Mikko Kurimo机构:Department of Signal Processing and Acoustics, Aalto University, Finland备注:5 pages, 2 figures, in Proceedings Interspeech 2022摘要:It is common knowledge that the quantity and quality of the training data play a significant role in the creation of a good machine learning model. In this paper, we take it one step further and demonstrate that the way the training examples are arranged is also of crucial importance. Curriculum Learning is built on the observation that organized and structured assimilation of knowledge has the ability to enable faster training and better comprehension. When humans learn to speak, they first try to utter basic phones and then gradually move towards more complex structures such as words and sentences. This methodology is known as Curriculum Learning, and we employ it in the context of Automatic Speech Recognition. We hypothesize that end-to-end models can achieve better performance when provided with an organized training set consisting of examples that exhibit an increasing level of difficulty (i.e. a curriculum). To impose structure on the training set and to define the notion of an easy example, we explored multiple scoring functions that either use feedback from an external neural network or incorporate feedback from the model itself. Empirical results show that with different curriculums we can balance the training times and the network's performance.
【12】 Chewing Detection from Commercial Smart-glasses
标题:商用智能眼镜的咀嚼检测
链接:https://arxiv.org/abs/2208.05735
作者:Vasileios Papapanagiotou,Anastasia Liapi,Anastasios Delopoulos机构:Multimedia Understanding Group, Electrical and Computer Engineering, Dpt., Aristotle University of Thessaloniki, Thessaloniki, Greece摘要:Automatic dietary monitoring has progressed significantly during the last years, offering a variety of solutions, both in terms of sensors and algorithms as well as in terms of what aspect or parameters of eating behavior are measured and monitored. Automatic detection of eating based on chewing sounds has been studied extensively, however, it requires a microphone to be mounted on the subject's head for capturing the relevant sounds. In this work, we evaluate the feasibility of using an off-the-shelf commercial device, the Razer Anzu smart-glasses, for automatic chewing detection. The smart-glasses are equipped with stereo speakers and microphones that communicate with smart-phones via Bluetooth. The microphone placement is not optimal for capturing chewing sounds, however, we find that it does not significantly affect the detection effectiveness. We apply an algorithm from literature with some adjustments on a challenging dataset that we have collected in house. Leave-one-subject-out experiments yield promising results, with an F1-score of 0.96 for the best case of duration-based evaluation of eating time.
【13】 ROC: A New Paradigm for Lyric-to-Melody Generation
标题:ROC:歌词到旋律生成的新范式
链接:https://arxiv.org/abs/2208.05697
* 与cs.SD语音【7】为同一篇
作者:Ang Lv,Xu Tan,Tao Qin,Tie-Yan Liu,Rui Yan机构:Renmin University of China, Beijing, China, Microsoft Research Asia摘要:Lyric-to-melody generation is an important task in songwriting, and is also quite challenging due to its distinctive characteristics: the generated melodies should not only follow good musical patterns, but also align with features in lyrics such as rhythms and structures. These characteristics cannot be well handled by neural generation models that learn lyric-to-melody mapping in an end-to-end way, due to several issues: (1) lack of aligned lyric-melody training data to sufficiently learn lyric-melody feature alignment; (2) lack of controllability in generation to explicitly guarantee the lyric-melody feature alignment. In this paper, we propose ROC, a new paradigm for lyric-to-melody generation that addresses the above issues through a generation-retrieval pipeline. Specifically, our paradigm has two stages: (1) creation stage, where a huge amount of music pieces are generated by a neural-based melody language model and indexed in a database through several key features (e.g., chords, tonality, rhythm, and structural information including chorus or verse); (2) re-creation stage, where melodies are recreated by retrieving music pieces from the database according to the key features from lyrics and concatenating best music pieces based on composition guidelines and melody language model scores. Our ROC paradigm has several advantages: (1) It only needs unpaired melody data to train melody language model, instead of paired lyric-melody data in previous models. (2) It achieves good lyric-melody feature alignment in lyric-to-melody generation. Experiments on English and Chinese datasets demonstrate that ROC outperforms previous neural based lyric-to-melody generation models on both objective and subjective metrics.
【14】 Symbolic Music Loop Generation with Neural Discrete Representations
标题:用神经离散表示法生成符号音乐环路
链接:https://arxiv.org/abs/2208.05605
* 与cs.SD语音【8】为同一篇
作者:Sangjun Han,Hyeongrae Ihm,Moontae Lee,Woohyung Lim机构:LG AI Research, University of Illinois at Chicago摘要:Since most of music has repetitive structures from motifs to phrases, repeating musical ideas can be a basic operation for music composition. The basic block that we focus on is conceptualized as loops which are essential ingredients of music. Furthermore, meaningful note patterns can be formed in a finite space, so it is sufficient to represent them with combinations of discrete symbols as done in other domains. In this work, we propose symbolic music loop generation via learning discrete representations. We first extract loops from MIDI datasets using a loop detector and then learn an autoregressive model trained by discrete latent codes of the extracted loops. We show that our model outperforms well-known music generative models in terms of both fidelity and diversity, evaluating on random space. Our code and supplementary materials are available at https://github.com/sjhan91/Loop_VQVAE_Official.
机器翻译,仅供参考