今天跟大家分享一篇语音相关的论文合集:cs.SD语音13篇,eess.AS音频处理17篇。

cs.SD语音

【1】 Volume-Independent Music Matching by Frequency Spectrum Comparison

标题:基于频谱比较法的音量无关音乐匹配

链接:https://arxiv.org/abs/2206.15426

作者:Anthony Lee
摘要:我经常听到一首音乐,不知道它叫什么名字。事实上,有一些应用程序,例如Shazam应用程序,可以提供音乐匹配。然而,这些应用程序的局限性在于,如果不是同一段录音,就无法识别同一位音乐家演奏的同一首乐曲。沙扎姆识别的是它的录音,而不是音乐。这是因为沙扎姆匹配的是音量的变化,而不是声音的频率。这项研究试图以人类理解音乐的方式来匹配音乐:通过音乐的频谱,而不是音量变化。基本上,这个想法是预先计算数据库中所有音乐的频谱,然后提取未知片段,尝试将其频谱与数据库中每个音乐的每个片段相匹配。我通过将窗口滑动0.1秒,将未知片段的频谱与我们的数据库进行匹配,并通过取绝对值、归一化音频、减去归一化数组和绝对差之和来计算误差。误差最小的段被视为匹配的候选段。事实证明,匹配性能取决于音乐的复杂性。匹配简单的音乐,如单音符片段,是成功的。然而,更复杂的作品,如肖邦民谣4,没有成功,也就是说,该算法无法在数据库中的任何音乐中产生低误差值。我怀疑这与注释过多有关:高次谐波中的失配增加了大量错误,从而淹没了计算。
摘要:Often, I hear a piece of music and wonder what the name of the piece is. Indeed, there are applications such as Shazam app that provides music matching. However, the limitations of those apps are that the same piece performed by the same musician cannot be identified if it is not the same recording. Shazam identifies the recording of it, not the music. This is because Shazam matches the variation in volume, not the frequencies of the sound. This research attempts to match music the way humans understand it: by the frequency spectrum of music, not the volume variation. Essentially, the idea is to precompute the frequency spectrums of all the music in the database, then take the unknown piece and try to match its frequency spectrum against every segment of every music in the database. I did it by matching the frequency spectrum of the unknown piece to our database by sliding the window by 0.1 seconds and calculating the error by taking Absolute value, normalizing the audio, subtracting the normalized arrays, and taking the sum of absolute differences. The segment that shows the least error is considered the candidate for the match. The matching performance proved to be dependent on the complexity of the music. Matching simple music, such as single note pieces, was successful. However, more complex pieces, such as Chopins Ballade 4, were not successful, that is, the algorithm could not produce low error values in any of the music in the database. I suspect that it has to do with having too many notes: mismatches in the higher harmonics added up to a significant amount of errors, which swamps the calculations.


【2】 Implicit Neural Spatial Filtering for Multichannel Source Separation in  the Waveform Domain

标题:隐式神经空间滤波在波形域中的多道源分离

链接:https://arxiv.org/abs/2206.15423

作者:Dejan Markovic,Alexandre Defossez,Alexander Richard
备注:Interspeech 2022
摘要:我们提出了一种单级随机波形到波形多通道模型,该模型可以根据移动声源在动态声学场景中的广泛空间位置来分离移动声源。我们将场景分为两个空间区域,分别包含目标和干扰声源。该模型经过端到端的训练并隐式执行空间处理,没有任何基于传统处理或使用手工制作的空间特征的组件。我们在真实数据集上对所提出的模型进行了评估,结果表明,该模型与oracle波束形成器以及最先进的单通道增强网络的性能相匹配。
摘要:We present a single-stage casual waveform-to-waveform multichannel model that can separate moving sound sources based on their broad spatial locations in a dynamic acoustic scene. We divide the scene into two spatial regions containing, respectively, the target and the interfering sound sources. The model is trained end-to-end and performs spatial processing implicitly, without any components based on traditional processing or use of hand-crafted spatial features. We evaluate the proposed model on a real-world dataset and show that the model matches the performance of an oracle beamformer followed by a state-of-the-art single-channel enhancement network.


【3】 Sonification as a Reliable Alternative to Conventional Visual Surgical  Navigation

标题:可听化是传统视觉外科导航的可靠替代方法

链接:https://arxiv.org/abs/2206.15291

作者:Sasan Matinfar,Mehrdad Salehi,Daniel Suter,Matthias Seibold,Navid Navab,Shervin Dehghani,Florian Wanivenhaus,Philipp Fürnstahl,Mazda Farshad,Nassir Navab
备注:19 pages, 7 figures
摘要:尽管图像引导手术辅助系统在准确性方面具有无可否认的优势,但此类系统尚未完全满足外科医生在可用性、时间效率以及将其集成到手术流程中方面的需求或期望。另一方面,感知研究表明,通过涉及不同感觉模式的多模式反馈呈现独立但因果相关的信息可以提高任务绩效。本文研究了一种计算机辅助手术导航的替代方法,介绍了一种用于导航椎弓根螺钉放置的新型超声方法,并讨论了基于多传感器反馈的高级解决方案。该方法包括一种基于调频(FM)合成的四自由度对准任务的新型超声解。我们比较了所提出的超声方法与视觉导航的结果准确性和执行时间,视觉导航目前被认为是最先进的。我们进行了一项模拟研究,其中17名外科医生在拟议的基于超声的方法或传统视觉导航方法的指导下在腰椎中执行椎弓根螺钉放置任务。结果表明,该方法与现有技术一样精确,同时减少了外科医生在任务执行过程中对视觉导航显示的需要,而不是对手术工具和目标解剖的自然关注。
摘要:Despite the undeniable advantages of image-guided surgical assistance systems in terms of accuracy, such systems have not yet fully met surgeons' needs or expectations regarding usability, time efficiency, and their integration into the surgical workflow. On the other hand, perceptual studies have shown that presenting independent but causally correlated information via multimodal feedback involving different sensory modalities can improve task performance. This article investigates an alternative method for computer-assisted surgical navigation, introduces a novel sonification methodology for navigated pedicle screw placement, and discusses advanced solutions based on multisensory feedback. The proposed method comprises a novel sonification solution for alignment tasks in four degrees of freedom based on frequency modulation (FM) synthesis. We compared the resulting accuracy and execution time of the proposed sonification method with visual navigation, which is currently considered the state of the art. We conducted a phantom study in which 17 surgeons executed the pedicle screw placement task in the lumbar spine, guided by either the proposed sonification-based or the traditional visual navigation method. The results demonstrated that the proposed method is as accurate as the state of the art while decreasing the surgeon's need to focus on visual navigation displays instead of the natural focus on surgical tools and targeted anatomy during task execution.


【4】 R-MelNet: Reduced Mel-Spectral Modeling for Neural TTS

标题:R-MelNet:神经TTS的简化Mel谱建模

链接:https://arxiv.org/abs/2206.15276

作者:Kyle Kastner,Aaron Courville
摘要:本文介绍了R-MelNet,这是一种两部分自回归结构,前端基于MelNet的第一层,后端是用于神经文本语音合成的WaveRNN风格的音频解码器。该模型将字符和音素的混合序列作为输入,并带有可选的音频启动序列,生成低分辨率的mel频谱特征,由WaveRNN解码器插值并用于生成音频波形。再加上半精度训练,R-MelNet在单个商品GPU(NVIDIA 2080Ti)上使用的GPU内存不足11GB。我们详细介绍了稳定半精度训练的一些关键实现细节,包括近似的、数值稳定的物流注意力混合。使用随机、多样本每步推理方案,生成的模型生成高度变化的音频,同时允许基于文本和音频的控件修改输出波形。在单说话人TTS数据集上训练的R-MelNet系统的定性和定量评估证明了我们方法的有效性。
摘要:This paper introduces R-MelNet, a two-part autoregressive architecture with a frontend based on the first tier of MelNet and a backend WaveRNN-style audio decoder for neural text-to-speech synthesis. Taking as input a mixed sequence of characters and phonemes, with an optional audio priming sequence, this model produces low-resolution mel-spectral features which are interpolated and used by a WaveRNN decoder to produce an audio waveform. Coupled with half precision training, R-MelNet uses under 11 gigabytes of GPU memory on a single commodity GPU (NVIDIA 2080Ti). We detail a number of critical implementation details for stable half precision training, including an approximate, numerically stable mixture of logistics attention. Using a stochastic, multi-sample per step inference scheme, the resulting model generates highly varied audio, while enabling text and audio based controls to modify output waveforms. Qualitative and quantitative evaluations of an R-MelNet system trained on a single speaker TTS dataset demonstrate the effectiveness of our approach.


【5】 libACA, pyACA, and ACA-Code: Audio Content Analysis in 3 Languages

标题:LibACA、pyACA和ACA-Code:3种语言的音频内容分析

链接:https://arxiv.org/abs/2206.15219

作者:Alexander Lerch
备注:Preprint submitted to "Software Impacts"
摘要:libACA、pyACA和ACA代码这三个包为使用三种不同语言(C++、Python和Matlab)分析音乐音频信号的基本方法和算法提供了参考实现。这三个软件包涵盖了相同的算法,例如低电平音频特征提取、基频估计,以及和弦识别、音乐关键点检测和开始检测的简单方法。此外,还提供了在音频内容分析中有用的更通用算法的it实现,如动态时间扭曲和维特比算法。因此,这三个软件包为实现音频分析算法的学生和工程师提供了实用的跨语言和跨平台参考,并支持以实现为中心的音频内容分析和音乐信息检索算法学习。
摘要:The three packages libACA, pyACA, and ACA-Code provide reference implementations for basic approaches and algorithms for the analysis of musical audio signals in three different languages: C++, Python, and Matlab. All three packages cover the same algorithms, such as extraction of low level audio features, fundamental frequency estimation, as well as simple approaches to chord recognition, musical key detection, and onset detection. In addition, it implementations of more generic algorithms useful in audio content analysis such as dynamic time warping and the Viterbi algorithm are provided. The three packages thus provide a practical cross-language and cross-platform reference to students and engineers implementing audio analysis algorithms and enable implementation-focused learning of algorithms for audio content analysis and music information retrieval.


【6】 An Evaluation of Three-Stage Voice Conversion Framework for Noisy and  Reverberant Conditions

标题:三级语音转换框架在噪声和混响条件下的评价

链接:https://arxiv.org/abs/2206.15155

作者:Yeonjong Choi,Chao Xie,Tomoki Toda
备注:Accepted to INTERSPEECH 2022
摘要:本文提出了一种能够同时处理加性噪声和混响的新语音转换框架,并对其性能进行了评估。已有一些VC研究侧重于语音数据受到背景噪声和混响干扰的真实环境。为了处理没有干净目标数据集的更实际的情况,一种可能的方法是零炮VC,但与使用足够数量的目标语音数据的VC相比,其性能往往会下降。为了利用大量噪声混响目标语音数据,我们提出了一种三阶段VC框架,基于使用预训练去噪模型的去噪过程、使用去冗余模型的去冗余过程和使用基于变分自动编码器的非并行VC模型的VC过程。实验结果表明,1)噪声和混响会导致VC性能显著下降,2)该方法缓解了噪声和混响带来的不利影响,显著优于在噪声混响语音数据上直接训练的基线,3)去噪和去冗余带来的潜在退化仍然会对VC性能造成明显的不利影响。
摘要:This paper presents a new voice conversion (VC) framework capable of dealing with both additive noise and reverberation, and its performance evaluation. There have been studied some VC researches focusing on real-world circumstances where speech data are interfered with background noise and reverberation. To deal with more practical conditions where no clean target dataset is available, one possible approach is zero-shot VC, but its performance tends to degrade compared with VC using sufficient amount of target speech data. To leverage large amount of noisy-reverberant target speech data, we propose a three-stage VC framework based on denoising process using a pretrained denoising model, dereverberation process using a dereverberation model, and VC process using a nonparallel VC model based on a variational autoencoder. The experimental results show that 1) noise and reverberation additively cause significant VC performance degradation, 2) the proposed method alleviates the adverse effects caused by both noise and reverberation, and significantly outperforms the baseline directly trained on the noisy-reverberant speech data, and 3) the potential degradation introduced by the denoising and dereverberation still causes noticeable adverse effects on VC performance.


【7】 Language Model-Based Emotion Prediction Methods for Emotional Speech  Synthesis Systems

标题:基于语言模型的情感语音合成系统情感预测方法

链接:https://arxiv.org/abs/2206.15067

作者:Hyun-Wook Yoon,Ohsung Kwon,Hoyeon Lee,Ryuichi Yamamoto,Eunwoo Song,Jae-Min Kim,Min-Jae Hwang
备注:Accepted in INTERSPEECH2022
摘要:本文提出了一种基于预训练语言模型(LM)的情感预测方法的有效情感文本到语音(TTS)系统。与需要手动定义情感类等辅助输入的传统系统不同,我们的系统直接从输入文本中估计情感相关属性。具体来说,我们利用生成预训练变换器(GPT)-3分别联合预测情感类别及其表示情感粗糙和精细属性的强度。然后,将这些属性组合在情感嵌入空间中,并用作TTS模型的条件特征,以生成输出语音信号。因此,该系统只能从文本中产生情感语音,而无需任何辅助输入。此外,由于GPT-3能够捕捉连续句子之间的情感语境,因此该方法可以有效地处理情感语音的段落级生成。
摘要:This paper proposes an effective emotional text-to-speech (TTS) system with a pre-trained language model (LM)-based emotion prediction method. Unlike conventional systems that require auxiliary inputs such as manually defined emotion classes, our system directly estimates emotion-related attributes from the input text. Specifically, we utilize generative pre-trained transformer (GPT)-3 to jointly predict both an emotion class and its strength in representing emotions coarse and fine properties, respectively. Then, these attributes are combined in the emotional embedding space and used as conditional features of the TTS model for generating output speech signals. Consequently, the proposed system can produce emotional speech only from text without any auxiliary inputs. Furthermore, because the GPT-3 enables to capture emotional context among the consecutive sentences, the proposed method can effectively handle the paragraph-level generation of emotional speech.


【8】 FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised  Learning Features in Robust End-to-end Speech Recognition

标题:无畏:稳健端到端语音识别中集成自监督学习特征的特征细化损失

链接:https://arxiv.org/abs/2206.15056

作者:Szu-Jui Chen,Jiamin Xie,John H. L. Hansen
备注:Accepted for Interspeech 2022
摘要:自监督学习表示(SSLR)为许多领域的下游任务带来了强大的特征。最近,一些SSLR在自动语音识别(ASR)基准语料库上显示了有希望的结果。然而,以前的研究仅表明孤立SSLR的性能作为ASR模型的输入特征。在本研究中,我们提议在端到端(E2E)ASR模型中使用各种融合方法来研究不同SSLR组合的有效性。此外,我们将显示这些提取的SSLR之间存在相关性。因此,我们进一步提出了用于去相关的特征细化损失,以有效地组合输入特征集。为了评估,我们表明,对于《华尔街日报》和无畏步骤挑战(FSC)语料库,拟议的“无畏学习特征”比没有拟议特征细化损失的系统表现更好。
摘要:Self-supervised learning representations (SSLR) have resulted in robust features for downstream tasks in many fields. Recently, several SSLRs have shown promising results on automatic speech recognition (ASR) benchmark corpora. However, previous studies have only shown performance for solitary SSLRs as an input feature for ASR models. In this study, we propose to investigate the effectiveness of diverse SSLR combinations using various fusion methods within end-to-end (E2E) ASR models. In addition, we will show there are correlations between these extracted SSLRs. As such, we further propose a feature refinement loss for decorrelation to efficiently combine the set of input features. For evaluation, we show that the proposed 'FeaRLESS learning features' perform better than systems without the proposed feature refinement loss for both the WSJ and Fearless Steps Challenge (FSC) corpora.


【9】 Interpretable Melody Generation from Lyrics with Discrete-Valued  Adversarial Training

标题:基于离散值对抗性训练的歌词可解释旋律生成

链接:https://arxiv.org/abs/2206.15027

作者:Wei Duan,Zhe Zhang,Yi Yu,Keizo Oyama
备注:3 pages, 3 figures
摘要:在人工智能和音乐领域,从歌词中生成旋律是一项有趣但富有挑战性的任务。然而,难以保持输入歌词和生成旋律之间的一致性限制了以往作品的生成质量。在我们的提案中,我们展示了我们提出的可解释歌词到旋律生成系统,该系统可以与用户交互,以了解生成过程并重新创建所需的歌曲。为了提高匹配歌词的旋律生成的可靠性,利用互信息来增强歌词和生成的旋律之间的一致性。利用Gumbel Softmax解决了生成对抗网络生成离散音乐属性的不可微性问题。此外,利用生成器输出的预测概率来推荐音乐属性。与我们的歌词到旋律生成系统交互,用户可以收听生成的人工智能歌曲,并通过从推荐的音乐属性中选择来重新创建一首新歌。
摘要:Generating melody from lyrics is an interesting yet challenging task in the area of artificial intelligence and music. However, the difficulty of keeping the consistency between input lyrics and generated melody limits the generation quality of previous works. In our proposal, we demonstrate our proposed interpretable lyrics-to-melody generation system which can interact with users to understand the generation process and recreate the desired songs. To improve the reliability of melody generation that matches lyrics, mutual information is exploited to strengthen the consistency between lyrics and generated melodies. Gumbel-Softmax is exploited to solve the non-differentiability problem of generating discrete music attributes by Generative Adversarial Networks (GANs). Moreover, the predicted probabilities output by the generator is utilized to recommend music attributes. Interacting with our lyrics-to-melody generation system, users can listen to the generated AI song as well as recreate a new song by selecting from recommended music attributes.


【10】 Acoustic Room Compensation Using Local PCA-based Room Average PSD  Estimation

标题:基于局部主成分分析的声学房间平均PSD估计

链接:https://arxiv.org/abs/2206.15356

作者:Wenyu Jin,Patrick McPherson,Chris Pike,Adib Mehrabi
备注:5 pages, 7 figures, accepted to IWAENC 2022
摘要:声学房间补偿技术已被广泛研究,该技术允许声音再现系统抵消由于过度房间共振引起的声音场景的不希望的改变。据报道,人们作出了广泛的努力来扩大房间均衡有效的区域,并对比房间传递函数在空间中的变化。扬声器调谐技术“Trueplay”允许用户基于房间的空间平均功率谱密度(PSD)在扩展的收听区域上补偿不希望的房间效果,通常在用户在房间内走动时使用便携式设备上的麦克风测量。在这项工作中,我们提出了一种新的系统,该系统利用扬声器回波路径自响应的测量,使用基于局部主成分分析的方法预测房间平均功率谱密度。实验结果证实了所提出的估计方法的有效性,这进一步导致了一种房间补偿滤波器设计,与具有地面真实房间平均功率谱密度的参考系统相比,该设计实现了良好的声音相似性,同时优于未利用所提估计器的其他系统。
摘要:Acoustic room compensation techniques, which allow a sound reproduction system to counteract undesired alteration to the sound scene due to excessive room resonances, have been widely studied. Extensive efforts have been reported to enlarge the region over which room equalization is effective and to contrast variations of room transfer functions in space. A speaker-tuning technology "Trueplay" allows users to compensate for undesired room effects over an extended listening area based on a spatially averaged power spectral density (PSD) of the room, which is conventionally measured using microphones on portable devices when users move around the room. In this work, we propose a novel system that leverages the measurement of the speaker echo path self-response to predict the room average PSD using a local PCA based approach. Experimental results confirm the effectiveness of the proposed estimation method, which further leads to a room compensation filter design that achieves a good sound similarity compared to the reference system with the ground-truth room average PSD while outperforming other systems that do not leverage the proposed estimator.


【11】 TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis  using ranking support vector machine with variational autoencoder

标题:TTS-by-TTS 2:基于变分自动编码器的排序支持向量机的神经语音合成数据选择性增强

链接:https://arxiv.org/abs/2206.14984

作者:Eunwoo Song,Ryuichi Yamamoto,Ohsung Kwon,Chan-Ho Song,Min-Jae Hwang,Suhyeon Oh,Hyun-Wook Yoon,Jin-Seob Kim,Jae-Min Kim
备注:Accepted to the conference of INTERSPEECH 2022
摘要:合成语音质量的最新进展使我们能够使用合成语料库训练文本到语音(TTS)系统。然而,仅仅增加合成数据量并不总是有利于提高训练效率。我们在本研究中的目的是有选择地选择有利于训练过程的合成数据。在该方法中,我们首先采用一种变分自动编码器,其后验分布用于提取代表记录语料库和合成语料库之间声学相似性的潜在特征。通过使用这些学习的特征,我们训练了一个排序支持向量机(RankSVM),该机以有效地对二进制类之间的相对属性进行排序而闻名。通过将记录语音和合成语音设置为两个相反的类,RankSVM用于确定合成语音在声学上与记录数据的相似程度。然后,从大规模合成语料库中选择分布接近记录数据的合成TTS数据。通过使用这些数据重新训练TTS模型,可以显著提高合成质量。客观和主观评价结果表明,该方法优于传统方法。
摘要:Recent advances in synthetic speech quality have enabled us to train text-to-speech (TTS) systems by using synthetic corpora. However, merely increasing the amount of synthetic data is not always advantageous for improving training efficiency. Our aim in this study is to selectively choose synthetic data that are beneficial to the training process. In the proposed method, we first adopt a variational autoencoder whose posterior distribution is utilized to extract latent features representing acoustic similarity between the recorded and synthetic corpora. By using those learned features, we then train a ranking support vector machine (RankSVM) that is well known for effectively ranking relative attributes among binary classes. By setting the recorded and synthetic ones as two opposite classes, RankSVM is used to determine how the synthesized speech is acoustically similar to the recorded data. Then, synthetic TTS data, whose distribution is close to the recorded data, are selected from large-scale synthetic corpora. By using these data for retraining the TTS model, the synthetic quality can be significantly improved. Objective and subjective evaluation results show the superiority of the proposed method over the conventional methods.


【12】 Improving Visual Speech Enhancement Network by Learning Audio-visual  Affinity with Multi-head Attention

标题:基于多头注意的视听亲和力学习改进视觉语音增强网络

链接:https://arxiv.org/abs/2206.14964

作者:Xinmeng Xu,Yang Wang,Jie Jia,Binbin Chen,Dejun Li
备注:Accepted by Interspeech 2022. arXiv admin note: substantial text overlap with arXiv:2101.06268摘要:视听语音增强系统被认为是分离和增强目标说话人语音的一种很有前景的解决方案。典型的方法侧重于通过基于卷积神经网络的简单编解码结构预测纯净语音频谱,这些方法a)不足以充分利用数据,b)无法有效平衡视听特征。该模型通过a)在编码阶段逐层融合音频和视频特征,并将融合的音频和视频特征提供给每个相应的解码器层,更重要的是,b)引入两阶段多头交叉注意力(MHCA)机制来推断视听语音增强,以平衡融合的视听特征并消除不相关的特征。本文提出了注意力视听多层特征融合模型,其中MHCA单元应用于解码器每一层的特征映射。该模型证明了网络相对于最先进模型的优越性能。
摘要:Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based encoder-decoder architecture, and these methods a) are not adequate to use data fully, b) are unable to effectively balance audio-visual features. The proposed model alleviates these drawbacks by a) applying a model that fuses audio and visual features layer by layer in encoding phase, and that feeds fused audio-visual features to each corresponding decoder layer, and more importantly, b) introducing a 2-stage multi-head cross attention (MHCA) mechanism to infer audio-visual speech enhancement for balancing the fused audio-visual features and eliminating irrelevant features. This paper proposes attentional audio-visual multi-layer feature fusion model, in which MHCA units are applied to feature mapping at every layer of decoder. The proposed model demonstrates the superior performance of the network against the state-of-the-art models.


【13】 GLD-Net: Improving Monaural Speech Enhancement by Learning Global and  Local Dependency Features with GLD Block

标题:GLD-Net:通过GLD块学习全局和局部依赖特征来改进单声道语音增强

链接:https://arxiv.org/abs/2206.14962

作者:Xinmeng Xu,Yang Wang,Jie Jia,Binbin Chen,Jianjun Hao
备注:Accepted by Interspeech 2022
摘要:对于单耳语音增强,语境信息对于准确的语音估计非常重要。然而,常用的卷积神经网络(CNN)在捕捉时间上下文方面很弱,因为它们一次只构建处理一个局部邻域的块。为了解决这个问题,我们从人类听觉感知中引入了一种两阶段可训练的推理机制,称为全局局部依赖(GLD)块。GLD块从噪声频谱图中捕获全局和局部水平的时频箱的长期依赖性,以帮助检测语音部分、噪声部分和整个噪声输入之间的相关性。此外,我们还开发了一个单耳语音增强网络GLD网,该网络采用编解码结构,由语音对象分支、干扰分支和全局噪声分支组成。在全局级和局部级提取的语音特征在每个分支中有效地进行推理和聚合。我们在WSJ0和需求数据集上将提出的GLD网络与现有的最先进方法进行了比较。结果表明,GLD-Net在PESQ和STOI方面优于最先进的方法。
摘要:For monaural speech enhancement, contextual information is important for accurate speech estimation. However, commonly used convolution neural networks (CNNs) are weak in capturing temporal contexts since they only build blocks that process one local neighborhood at a time. To address this problem, we learn from human auditory perception to introduce a two-stage trainable reasoning mechanism, referred as global-local dependency (GLD) block. GLD blocks capture long-term dependency of time-frequency bins both in global level and local level from the noisy spectrogram to help detecting correlations among speech part, noise part, and whole noisy input. What is more, we conduct a monaural speech enhancement network called GLD-Net, which adopts encoder-decoder architecture and consists of speech object branch, interference branch, and global noisy branch. The extracted speech feature at global-level and local-level are efficiently reasoned and aggregated in each of the branches. We compare the proposed GLD-Net with existing state-of-art methods on WSJ0 and DEMAND dataset. The results show that GLD-Net outperforms the state-of-the-art methods in terms of PESQ and STOI.
eess.AS音频处理

【1】 Challenges and Opportunities in Multi-device Speech Processing

标题:多设备语音处理面临的挑战和机遇

链接:https://arxiv.org/abs/2206.15432

作者:Gregory Ciccarelli,Jarred Barber,Arun Nair,Israel Cohen,Tao Zhang
备注:Accepted for INTERSPEECH 2022
摘要:我们回顾了多设备家庭环境中自动语音识别、关键词识别、设备仲裁、语音增强和源定位的当前解决方案和技术挑战,为INTERSPEECH 2022特别会议“多智能设备信号处理和机器学习的挑战和机遇”提供背景。我们还确定了支持这些研究领域所需的数据集。基于以上回顾和我们在多设备领域的研究经验,我们对未来的发展进行了展望。
摘要:We review current solutions and technical challenges for automatic speech recognition, keyword spotting, device arbitration, speech enhancement, and source localization in multidevice home environments to provide context for the INTERSPEECH 2022 special session, "Challenges and opportunities for signal processing and machine learning for multiple smart devices". We also identify the datasets needed to support these research areas. Based on the review and our research experience in the multi-device domain, we conclude with an outlook on the future evolution


【2】 Few-Shot Cross-Lingual TTS Using Transferable Phoneme Embedding

标题:利用可转移音素嵌入的Few-Shot跨语言TTS

链接:https://arxiv.org/abs/2206.15427

作者:Wei-Ping Huang,Po-Chun Chen,Sung-Feng Huang,Hung-yi Lee
备注:Submitted to Interspeech 2022
摘要:本文研究了一种可转移音素嵌入框架,旨在处理Few-Shot设置下的跨语言文本到语音(TTS)问题。当涉及到Few-Shot学习时,转移学习是一种常见的方法,因为在Few-Shot训练数据上从头开始训练必然会过拟合。然而,我们发现,在极少的镜头设置下,原始的迁移学习方法无法适应看不见的语言,在这种情况下,提供的数据不到8分钟。我们通过提出一个由基于音素的TTS模型和码本模块组成的框架来处理这个问题,该框架将来自不同语言的音素投影到学习的潜在空间中。此外,通过利用音素级平均自监督学习特征,我们有效地提高了合成语音的质量。实验表明,使用我们的框架,当适应一种看不见的语言时,使用4个话语(约30秒的数据)就足以合成可理解的语音。
摘要:This paper studies a transferable phoneme embedding framework that aims to deal with the cross-lingual text-to-speech (TTS) problem under the few-shot setting. Transfer learning is a common approach when it comes to few-shot learning since training from scratch on few-shot training data is bound to overfit. Still, we find that the naive transfer learning approach fails to adapt to unseen languages under extremely few-shot settings, where less than 8 minutes of data is provided. We deal with the problem by proposing a framework that consists of a phoneme-based TTS model and a codebook module to project phonemes from different languages into a learned latent space. Furthermore, by utilizing phoneme-level averaged self-supervised learned features, we effectively improve the quality of synthesized speeches. Experiments show that using 4 utterances, which is about 30 seconds of data, is enough to synthesize intelligible speech when adapting to an unseen language using our framework.


【3】 Sub-8-Bit Quantization Aware Training for 8-Bit Neural Network  Accelerator with On-Device Speech Recognition

标题:具有设备上语音识别的8位神经网络加速器的亚8位量化感知训练

链接:https://arxiv.org/abs/2206.15408

作者:Kai Zhen,Hieu Duy Nguyen,Raviteja Chinta,Nathan Susanj,Athanasios Mouchtaris,Tariq Afzal,Ariya Rastrow
备注:Accepted for publication in INTERSPEECH 2022
摘要:我们提出了一种新的用于8位神经网络加速器的亚8位量化感知训练(S8BKAT)方案。我们的方法受Lloyd-Max压缩理论的启发,在训练期间对可行的计算开销进行了实际调整。利用从32位基线导出的量化质心,我们使用多区域绝对余弦(MRACos)正则化器来增加训练损失,该正则化器将权重聚集到其最近的质心,有效地充当伪压缩器。此外,引入了一个周期调用的硬压缩器,通过模拟运行时模型权重量化来提高收敛速度。我们使用递归神经网络传感器(RNN-T)架构将S8BKAT应用于语音识别任务。使用S8BQAT,我们能够增加模型参数大小,相对减少文字错误率4-16%,同时仍将延迟提高5%。
摘要:We present a novel sub-8-bit quantization-aware training (S8BQAT) scheme for 8-bit neural network accelerators. Our method is inspired from Lloyd-Max compression theory with practical adaptations for a feasible computational overhead during training. With the quantization centroids derived from a 32-bit baseline, we augment training loss with a Multi-Regional Absolute Cosine (MRACos) regularizer that aggregates weights towards their nearest centroid, effectively acting as a pseudo compressor. Additionally, a periodically invoked hard compressor is introduced to improve the convergence rate by emulating runtime model weight quantization. We apply S8BQAT on speech recognition tasks using Recurrent Neural NetworkTransducer (RNN-T) architecture. With S8BQAT, we are able to increase the model parameter size to reduce the word error rate by 4-16% relatively, while still improving latency by 5%.


【4】 Learning Audio-Text Agreement for Open-vocabulary Keyword Spotting

标题:用于开放词汇关键词识别的学习音文协议

链接:https://arxiv.org/abs/2206.15400

作者:Hyeon-Kyeong Shin,Hyewon Han,Doyeon Kim,Soo-Whan Chung,Hong-Goo Kang
备注:Accepted to Interspeech 2022
摘要:在本文中,我们提出了一种新的端到端用户定义的关键字识别方法,该方法利用语音和文本序列之间的语言对应模式。与以前需要语音关键字注册的方法不同,我们的方法将输入查询与注册的文本关键字序列进行比较。为了将音频和文本表示放在一个公共的潜在空间中,我们采用了一种基于注意力的跨模式匹配方法,该方法以端到端的方式进行训练,具有单调匹配损失和关键字分类损失。我们还利用声学嵌入网络的去噪损耗来提高噪声环境中的鲁棒性。此外,我们还介绍了LibriPhrase数据集,这是一种基于LibriSpeech的新的短短语数据集,用于有效地训练关键词识别模型。与其他单模式和跨模式基线相比,我们提出的方法在各种评估集上取得了有竞争力的结果。
摘要:In this paper, we propose a novel end-to-end user-defined keyword spotting method that utilizes linguistically corresponding patterns between speech and text sequences. Unlike previous approaches requiring speech keyword enrollment, our method compares input queries with an enrolled text keyword sequence. To place the audio and text representations within a common latent space, we adopt an attention-based cross-modal matching approach that is trained in an end-to-end manner with monotonic matching loss and keyword classification loss. We also utilize a de-noising loss for the acoustic embedding network to improve robustness in noisy environments. Additionally, we introduce the LibriPhrase dataset, a new short-phrase dataset based on LibriSpeech for efficiently training keyword spotting models. Our proposed method achieves competitive results on various evaluation sets compared to other single-modal and cross-modal baselines.


【5】 Acoustic Room Compensation Using Local PCA-based Room Average PSD  Estimation

标题:基于局部主成分分析的声学房间平均PSD估计

链接:https://arxiv.org/abs/2206.15356

* 与cs.SD语音【10】为同一篇

作者:Wenyu Jin,Patrick McPherson,Chris Pike,Adib Mehrabi
备注:5 pages, 7 figures, accepted to IWAENC 2022
摘要:声学房间补偿技术已被广泛研究,该技术允许声音再现系统抵消由于过度房间共振引起的声音场景的不希望的改变。据报道,人们作出了广泛的努力来扩大房间均衡有效的区域,并对比房间传递函数在空间中的变化。扬声器调谐技术“Trueplay”允许用户基于房间的空间平均功率谱密度(PSD)在扩展的收听区域上补偿不希望的房间效果,通常在用户在房间内走动时使用便携式设备上的麦克风测量。在这项工作中,我们提出了一种新的系统,该系统利用扬声器回波路径自响应的测量,使用基于局部主成分分析的方法预测房间平均功率谱密度。实验结果证实了所提出的估计方法的有效性,这进一步导致了一种房间补偿滤波器设计,与具有地面真实房间平均功率谱密度的参考系统相比,该设计实现了良好的声音相似性,同时优于未利用所提估计器的其他系统。
摘要:Acoustic room compensation techniques, which allow a sound reproduction system to counteract undesired alteration to the sound scene due to excessive room resonances, have been widely studied. Extensive efforts have been reported to enlarge the region over which room equalization is effective and to contrast variations of room transfer functions in space. A speaker-tuning technology "Trueplay" allows users to compensate for undesired room effects over an extended listening area based on a spatially averaged power spectral density (PSD) of the room, which is conventionally measured using microphones on portable devices when users move around the room. In this work, we propose a novel system that leverages the measurement of the speaker echo path self-response to predict the room average PSD using a local PCA based approach. Experimental results confirm the effectiveness of the proposed estimation method, which further leads to a room compensation filter design that achieves a good sound similarity compared to the reference system with the ground-truth room average PSD while outperforming other systems that do not leverage the proposed estimator.


【6】 TTS-by-TTS 2: Data-selective augmentation for neural speech synthesis  using ranking support vector machine with variational autoencoder

标题:TTS-by-TTS 2:基于变分自动编码器的排序支持向量机的神经语音合成数据选择性增强

链接:https://arxiv.org/abs/2206.14984

* 与cs.SD语音【11】为同一篇

作者:Eunwoo Song,Ryuichi Yamamoto,Ohsung Kwon,Chan-Ho Song,Min-Jae Hwang,Suhyeon Oh,Hyun-Wook Yoon,Jin-Seob Kim,Jae-Min Kim
备注:Accepted to the conference of INTERSPEECH 2022
摘要:合成语音质量的最新进展使我们能够使用合成语料库训练文本到语音(TTS)系统。然而,仅仅增加合成数据量并不总是有利于提高训练效率。我们在本研究中的目的是有选择地选择有利于训练过程的合成数据。在该方法中,我们首先采用一种变分自动编码器,其后验分布用于提取代表记录语料库和合成语料库之间声学相似性的潜在特征。通过使用这些学习的特征,我们训练了一个排序支持向量机(RankSVM),该机以有效地对二进制类之间的相对属性进行排序而闻名。通过将记录语音和合成语音设置为两个相反的类,RankSVM用于确定合成语音在声学上与记录数据的相似程度。然后,从大规模合成语料库中选择分布接近记录数据的合成TTS数据。通过使用这些数据重新训练TTS模型,可以显著提高合成质量。客观和主观评价结果表明,该方法优于传统方法。
摘要:Recent advances in synthetic speech quality have enabled us to train text-to-speech (TTS) systems by using synthetic corpora. However, merely increasing the amount of synthetic data is not always advantageous for improving training efficiency. Our aim in this study is to selectively choose synthetic data that are beneficial to the training process. In the proposed method, we first adopt a variational autoencoder whose posterior distribution is utilized to extract latent features representing acoustic similarity between the recorded and synthetic corpora. By using those learned features, we then train a ranking support vector machine (RankSVM) that is well known for effectively ranking relative attributes among binary classes. By setting the recorded and synthetic ones as two opposite classes, RankSVM is used to determine how the synthesized speech is acoustically similar to the recorded data. Then, synthetic TTS data, whose distribution is close to the recorded data, are selected from large-scale synthetic corpora. By using these data for retraining the TTS model, the synthetic quality can be significantly improved. Objective and subjective evaluation results show the superiority of the proposed method over the conventional methods.


【7】 Improving Visual Speech Enhancement Network by Learning Audio-visual  Affinity with Multi-head Attention

标题:基于多头注意的视听亲和力学习改进视觉语音增强网络

链接:https://arxiv.org/abs/2206.14964

* 与cs.SD语音【12】为同一篇

作者:Xinmeng Xu,Yang Wang,Jie Jia,Binbin Chen,Dejun Li
备注:Accepted by Interspeech 2022. arXiv admin note: substantial text overlap with arXiv:2101.06268摘要:视听语音增强系统被认为是分离和增强目标说话人语音的一种很有前景的解决方案。典型的方法侧重于通过基于卷积神经网络的简单编解码结构预测纯净语音频谱,这些方法a)不足以充分利用数据,b)无法有效平衡视听特征。该模型通过a)在编码阶段逐层融合音频和视频特征,并将融合的音频和视频特征提供给每个相应的解码器层,更重要的是,b)引入两阶段多头交叉注意力(MHCA)机制来推断视听语音增强,以平衡融合的视听特征并消除不相关的特征。本文提出了注意力视听多层特征融合模型,其中MHCA单元应用于解码器每一层的特征映射。该模型证明了网络相对于最先进模型的优越性能。
摘要:Audio-visual speech enhancement system is regarded as one of promising solutions for isolating and enhancing speech of desired speaker. Typical methods focus on predicting clean speech spectrum via a naive convolution neural network based encoder-decoder architecture, and these methods a) are not adequate to use data fully, b) are unable to effectively balance audio-visual features. The proposed model alleviates these drawbacks by a) applying a model that fuses audio and visual features layer by layer in encoding phase, and that feeds fused audio-visual features to each corresponding decoder layer, and more importantly, b) introducing a 2-stage multi-head cross attention (MHCA) mechanism to infer audio-visual speech enhancement for balancing the fused audio-visual features and eliminating irrelevant features. This paper proposes attentional audio-visual multi-layer feature fusion model, in which MHCA units are applied to feature mapping at every layer of decoder. The proposed model demonstrates the superior performance of the network against the state-of-the-art models.


【8】 GLD-Net: Improving Monaural Speech Enhancement by Learning Global and  Local Dependency Features with GLD Block

标题:GLD-Net:通过GLD块学习全局和局部依赖特征来改进单声道语音增强

链接:https://arxiv.org/abs/2206.14962

* 与cs.SD语音【13】为同一篇

作者:Xinmeng Xu,Yang Wang,Jie Jia,Binbin Chen,Jianjun Hao
备注:Accepted by Interspeech 2022
摘要:对于单耳语音增强,语境信息对于准确的语音估计非常重要。然而,常用的卷积神经网络(CNN)在捕捉时间上下文方面很弱,因为它们一次只构建处理一个局部邻域的块。为了解决这个问题,我们从人类听觉感知中引入了一种两阶段可训练的推理机制,称为全局局部依赖(GLD)块。GLD块从噪声频谱图中捕获全局和局部水平的时频箱的长期依赖性,以帮助检测语音部分、噪声部分和整个噪声输入之间的相关性。此外,我们还开发了一个单耳语音增强网络GLD网,该网络采用编解码结构,由语音对象分支、干扰分支和全局噪声分支组成。在全局级和局部级提取的语音特征在每个分支中有效地进行推理和聚合。我们在WSJ0和需求数据集上将提出的GLD网络与现有的最先进方法进行了比较。结果表明,GLD-Net在PESQ和STOI方面优于最先进的方法。
摘要:For monaural speech enhancement, contextual information is important for accurate speech estimation. However, commonly used convolution neural networks (CNNs) are weak in capturing temporal contexts since they only build blocks that process one local neighborhood at a time. To address this problem, we learn from human auditory perception to introduce a two-stage trainable reasoning mechanism, referred as global-local dependency (GLD) block. GLD blocks capture long-term dependency of time-frequency bins both in global level and local level from the noisy spectrogram to help detecting correlations among speech part, noise part, and whole noisy input. What is more, we conduct a monaural speech enhancement network called GLD-Net, which adopts encoder-decoder architecture and consists of speech object branch, interference branch, and global noisy branch. The extracted speech feature at global-level and local-level are efficiently reasoned and aggregated in each of the branches. We compare the proposed GLD-Net with existing state-of-art methods on WSJ0 and DEMAND dataset. The results show that GLD-Net outperforms the state-of-the-art methods in terms of PESQ and STOI.


【9】 iEmoTTS: Toward Robust Cross-Speaker Emotion Transfer and Control for  Speech Synthesis based on Disentanglement between Prosody and Timbre

标题:IEmoTTS:基于韵律和音色分离的语音合成中稳健的交叉说话人情感传递与控制

链接:https://arxiv.org/abs/2206.14866

作者:Guangyan Zhang,Ying Qin,Wenjie Zhang,Jialun Wu,Mei Li,Yutao Gai,Feijun Jiang,Tan Lee
备注:Submitted to IEEE Transactions on Audio, Speech, and Language Processing
摘要:在人机交互的许多应用中,需要具有生成具有特定类型情感的语音的能力。当目标说话人的带有情感标签的语音无法用于模型训练时,跨说话人情感转移是生成情感语音的常用方法。本文提出了一种新的跨说话人情感传递系统iEmoTTS。该系统由情感编码器、韵律预测器和音色编码器组成。情感编码器从输入语音的mel频谱图中提取情感类型的身份以及相应的情感强度。情感强度由输入话语携带该情感的后验概率来衡量。韵律预测器用于为情感传递提供韵律特征。木材编码器为系统提供与木材相关的信息。与其他许多侧重于解开说话人和语音风格因素的研究不同,iEmoTTS旨在通过韵律和音色之间的解开来实现跨说话人的情感传递。韵律是情感相关语音特征的主要载体,音色是说话人识别的基本特征。零触发情感传递,即在模型训练中看不到目标说话人的语音,也通过iEmoTTS实现。进行了大量的主观评价实验。结果表明,与最近提出的其他跨说话人情感转移系统相比,iEmoTTS是有效的。结果表明,iEmoTTS可以产生指定情绪类型和可控情绪强度的语音。通过适当的信息瓶颈容量,iEmoTTS能够有效地将情感信息传递给新的说话人。音频样本可公开获取\脚注{https://patrick-g-zhang.github.io/iemotts/}.
摘要:The capability of generating speech with specific type of emotion is desired for many applications of human-computer interaction. Cross-speaker emotion transfer is a common approach to generating emotional speech when speech with emotion labels from target speakers is not available for model training. This paper presents a novel cross-speaker emotion transfer system, named iEmoTTS. The system is composed of an emotion encoder, a prosody predictor, and a timbre encoder. The emotion encoder extracts the identity of emotion type as well as the respective emotion intensity from the mel-spectrogram of input speech. The emotion intensity is measured by the posterior probability that the input utterance carries that emotion. The prosody predictor is used to provide prosodic features for emotion transfer. The timber encoder provides timbre-related information for the system. Unlike many other studies which focus on disentangling speaker and style factors of speech, the iEmoTTS is designed to achieve cross-speaker emotion transfer via disentanglement between prosody and timbre. Prosody is considered as the main carrier of emotion-related speech characteristics and timbre accounts for the essential characteristics for speaker identification. Zero-shot emotion transfer, meaning that speech of target speakers are not seen in model training, is also realized with iEmoTTS. Extensive experiments of subjective evaluation have been carried out. The results demonstrate the effectiveness of iEmoTTS as compared with other recently proposed systems of cross-speaker emotion transfer. It is shown that iEmoTTS can produce speech with designated emotion type and controllable emotion intensity. With appropriate information bottleneck capacity, iEmoTTS is able to effectively transfer emotion information to a new speaker. Audio samples are publicly available\footnote{https://patrick-g-zhang.github.io/iemotts/}.


【10】 Volume-Independent Music Matching by Frequency Spectrum Comparison

标题:基于频谱比较法的音量无关音乐匹配

链接:https://arxiv.org/abs/2206.15426

* 与cs.SD语音【1】为同一篇

作者:Anthony Lee
摘要:我经常听到一首音乐,不知道它叫什么名字。事实上,有一些应用程序,例如Shazam应用程序,可以提供音乐匹配。然而,这些应用程序的局限性在于,如果不是同一段录音,就无法识别同一位音乐家演奏的同一首乐曲。沙扎姆识别的是它的录音,而不是音乐。这是因为沙扎姆匹配的是音量的变化,而不是声音的频率。这项研究试图以人类理解音乐的方式来匹配音乐:通过音乐的频谱,而不是音量变化。基本上,这个想法是预先计算数据库中所有音乐的频谱,然后提取未知片段,尝试将其频谱与数据库中每个音乐的每个片段相匹配。我通过将窗口滑动0.1秒,将未知片段的频谱与我们的数据库进行匹配,并通过取绝对值、归一化音频、减去归一化数组和绝对差之和来计算误差。误差最小的段被视为匹配的候选段。事实证明,匹配性能取决于音乐的复杂性。匹配简单的音乐,如单音符片段,是成功的。然而,更复杂的作品,如肖邦民谣4,没有成功,也就是说,该算法无法在数据库中的任何音乐中产生低误差值。我怀疑这与注释过多有关:高次谐波中的失配增加了大量错误,从而淹没了计算。
摘要:Often, I hear a piece of music and wonder what the name of the piece is. Indeed, there are applications such as Shazam app that provides music matching. However, the limitations of those apps are that the same piece performed by the same musician cannot be identified if it is not the same recording. Shazam identifies the recording of it, not the music. This is because Shazam matches the variation in volume, not the frequencies of the sound. This research attempts to match music the way humans understand it: by the frequency spectrum of music, not the volume variation. Essentially, the idea is to precompute the frequency spectrums of all the music in the database, then take the unknown piece and try to match its frequency spectrum against every segment of every music in the database. I did it by matching the frequency spectrum of the unknown piece to our database by sliding the window by 0.1 seconds and calculating the error by taking Absolute value, normalizing the audio, subtracting the normalized arrays, and taking the sum of absolute differences. The segment that shows the least error is considered the candidate for the match. The matching performance proved to be dependent on the complexity of the music. Matching simple music, such as single note pieces, was successful. However, more complex pieces, such as Chopins Ballade 4, were not successful, that is, the algorithm could not produce low error values in any of the music in the database. I suspect that it has to do with having too many notes: mismatches in the higher harmonics added up to a significant amount of errors, which swamps the calculations.


【11】 Implicit Neural Spatial Filtering for Multichannel Source Separation in  the Waveform Domain

标题:隐式神经空间滤波在波形域中的多道源分离

链接:https://arxiv.org/abs/2206.15423

* 与cs.SD语音【2】为同一篇

作者:Dejan Markovic,Alexandre Defossez,Alexander Richard
备注:Interspeech 2022
摘要:我们提出了一种单级随机波形到波形多通道模型,该模型可以根据移动声源在动态声学场景中的广泛空间位置来分离移动声源。我们将场景分为两个空间区域,分别包含目标和干扰声源。该模型经过端到端的训练并隐式执行空间处理,没有任何基于传统处理或使用手工制作的空间特征的组件。我们在真实数据集上对所提出的模型进行了评估,结果表明,该模型与oracle波束形成器以及最先进的单通道增强网络的性能相匹配。
摘要:We present a single-stage casual waveform-to-waveform multichannel model that can separate moving sound sources based on their broad spatial locations in a dynamic acoustic scene. We divide the scene into two spatial regions containing, respectively, the target and the interfering sound sources. The model is trained end-to-end and performs spatial processing implicitly, without any components based on traditional processing or use of hand-crafted spatial features. We evaluate the proposed model on a real-world dataset and show that the model matches the performance of an oracle beamformer followed by a state-of-the-art single-channel enhancement network.


【12】 Sonification as a Reliable Alternative to Conventional Visual Surgical  Navigation

标题:可听化是传统视觉外科导航的可靠替代方法

链接:https://arxiv.org/abs/2206.15291

* 与cs.SD语音【3】为同一篇

作者:Sasan Matinfar,Mehrdad Salehi,Daniel Suter,Matthias Seibold,Navid Navab,Shervin Dehghani,Florian Wanivenhaus,Philipp Fürnstahl,Mazda Farshad,Nassir Navab
备注:19 pages, 7 figures
摘要:尽管图像引导手术辅助系统在准确性方面具有无可否认的优势,但此类系统尚未完全满足外科医生在可用性、时间效率以及将其集成到手术流程中方面的需求或期望。另一方面,感知研究表明,通过涉及不同感觉模式的多模式反馈呈现独立但因果相关的信息可以提高任务绩效。本文研究了一种计算机辅助手术导航的替代方法,介绍了一种用于导航椎弓根螺钉放置的新型超声方法,并讨论了基于多传感器反馈的高级解决方案。该方法包括一种基于调频(FM)合成的四自由度对准任务的新型超声解。我们比较了所提出的超声方法与视觉导航的结果准确性和执行时间,视觉导航目前被认为是最先进的。我们进行了一项模拟研究,其中17名外科医生在拟议的基于超声的方法或传统视觉导航方法的指导下在腰椎中执行椎弓根螺钉放置任务。结果表明,该方法与现有技术一样精确,同时减少了外科医生在任务执行过程中对视觉导航显示的需要,而不是对手术工具和目标解剖的自然关注。
摘要:Despite the undeniable advantages of image-guided surgical assistance systems in terms of accuracy, such systems have not yet fully met surgeons' needs or expectations regarding usability, time efficiency, and their integration into the surgical workflow. On the other hand, perceptual studies have shown that presenting independent but causally correlated information via multimodal feedback involving different sensory modalities can improve task performance. This article investigates an alternative method for computer-assisted surgical navigation, introduces a novel sonification methodology for navigated pedicle screw placement, and discusses advanced solutions based on multisensory feedback. The proposed method comprises a novel sonification solution for alignment tasks in four degrees of freedom based on frequency modulation (FM) synthesis. We compared the resulting accuracy and execution time of the proposed sonification method with visual navigation, which is currently considered the state of the art. We conducted a phantom study in which 17 surgeons executed the pedicle screw placement task in the lumbar spine, guided by either the proposed sonification-based or the traditional visual navigation method. The results demonstrated that the proposed method is as accurate as the state of the art while decreasing the surgeon's need to focus on visual navigation displays instead of the natural focus on surgical tools and targeted anatomy during task execution.


【13】 R-MelNet: Reduced Mel-Spectral Modeling for Neural TTS

标题:R-MelNet:神经TTS的简化Mel谱建模

链接:https://arxiv.org/abs/2206.15276

* 与cs.SD语音【4】为同一篇

作者:Kyle Kastner,Aaron Courville
摘要:本文介绍了R-MelNet,这是一种两部分自回归结构,前端基于MelNet的第一层,后端是用于神经文本语音合成的WaveRNN风格的音频解码器。该模型将字符和音素的混合序列作为输入,并带有可选的音频启动序列,生成低分辨率的mel频谱特征,由WaveRNN解码器插值并用于生成音频波形。再加上半精度训练,R-MelNet在单个商品GPU(NVIDIA 2080Ti)上使用的GPU内存不足11GB。我们详细介绍了稳定半精度训练的一些关键实现细节,包括近似的、数值稳定的物流注意力混合。使用随机、多样本每步推理方案,生成的模型生成高度变化的音频,同时允许基于文本和音频的控件修改输出波形。在单说话人TTS数据集上训练的R-MelNet系统的定性和定量评估证明了我们方法的有效性。
摘要:This paper introduces R-MelNet, a two-part autoregressive architecture with a frontend based on the first tier of MelNet and a backend WaveRNN-style audio decoder for neural text-to-speech synthesis. Taking as input a mixed sequence of characters and phonemes, with an optional audio priming sequence, this model produces low-resolution mel-spectral features which are interpolated and used by a WaveRNN decoder to produce an audio waveform. Coupled with half precision training, R-MelNet uses under 11 gigabytes of GPU memory on a single commodity GPU (NVIDIA 2080Ti). We detail a number of critical implementation details for stable half precision training, including an approximate, numerically stable mixture of logistics attention. Using a stochastic, multi-sample per step inference scheme, the resulting model generates highly varied audio, while enabling text and audio based controls to modify output waveforms. Qualitative and quantitative evaluations of an R-MelNet system trained on a single speaker TTS dataset demonstrate the effectiveness of our approach.


【14】 libACA, pyACA, and ACA-Code: Audio Content Analysis in 3 Languages

标题:LibACA、pyACA和ACA-Code:3种语言的音频内容分析

链接:https://arxiv.org/abs/2206.15219

* 与cs.SD语音【5】为同一篇

作者:Alexander Lerch
备注:Preprint submitted to "Software Impacts"
摘要:libACA、pyACA和ACA代码这三个包为使用三种不同语言(C++、Python和Matlab)分析音乐音频信号的基本方法和算法提供了参考实现。这三个软件包涵盖了相同的算法,例如低电平音频特征提取、基频估计,以及和弦识别、音乐关键点检测和开始检测的简单方法。此外,还提供了在音频内容分析中有用的更通用算法的it实现,如动态时间扭曲和维特比算法。因此,这三个软件包为实现音频分析算法的学生和工程师提供了实用的跨语言和跨平台参考,并支持以实现为中心的音频内容分析和音乐信息检索算法学习。
摘要:The three packages libACA, pyACA, and ACA-Code provide reference implementations for basic approaches and algorithms for the analysis of musical audio signals in three different languages: C++, Python, and Matlab. All three packages cover the same algorithms, such as extraction of low level audio features, fundamental frequency estimation, as well as simple approaches to chord recognition, musical key detection, and onset detection. In addition, it implementations of more generic algorithms useful in audio content analysis such as dynamic time warping and the Viterbi algorithm are provided. The three packages thus provide a practical cross-language and cross-platform reference to students and engineers implementing audio analysis algorithms and enable implementation-focused learning of algorithms for audio content analysis and music information retrieval.


【15】 An Evaluation of Three-Stage Voice Conversion Framework for Noisy and  Reverberant Conditions

标题:三级语音转换框架在噪声和混响条件下的评价

链接:https://arxiv.org/abs/2206.15155

* 与cs.SD语音【6】为同一篇

作者:Yeonjong Choi,Chao Xie,Tomoki Toda
备注:Accepted to INTERSPEECH 2022
摘要:本文提出了一种能够同时处理加性噪声和混响的新语音转换框架,并对其性能进行了评估。已有一些VC研究侧重于语音数据受到背景噪声和混响干扰的真实环境。为了处理没有干净目标数据集的更实际的情况,一种可能的方法是零炮VC,但与使用足够数量的目标语音数据的VC相比,其性能往往会下降。为了利用大量噪声混响目标语音数据,我们提出了一种三阶段VC框架,基于使用预训练去噪模型的去噪过程、使用去冗余模型的去冗余过程和使用基于变分自动编码器的非并行VC模型的VC过程。实验结果表明,1)噪声和混响会导致VC性能显著下降,2)该方法缓解了噪声和混响带来的不利影响,显著优于在噪声混响语音数据上直接训练的基线,3)去噪和去冗余带来的潜在退化仍然会对VC性能造成明显的不利影响。
摘要:This paper presents a new voice conversion (VC) framework capable of dealing with both additive noise and reverberation, and its performance evaluation. There have been studied some VC researches focusing on real-world circumstances where speech data are interfered with background noise and reverberation. To deal with more practical conditions where no clean target dataset is available, one possible approach is zero-shot VC, but its performance tends to degrade compared with VC using sufficient amount of target speech data. To leverage large amount of noisy-reverberant target speech data, we propose a three-stage VC framework based on denoising process using a pretrained denoising model, dereverberation process using a dereverberation model, and VC process using a nonparallel VC model based on a variational autoencoder. The experimental results show that 1) noise and reverberation additively cause significant VC performance degradation, 2) the proposed method alleviates the adverse effects caused by both noise and reverberation, and significantly outperforms the baseline directly trained on the noisy-reverberant speech data, and 3) the potential degradation introduced by the denoising and dereverberation still causes noticeable adverse effects on VC performance.


【16】 Language Model-Based Emotion Prediction Methods for Emotional Speech  Synthesis Systems

标题:基于语言模型的情感语音合成系统情感预测方法

链接:https://arxiv.org/abs/2206.15067

* 与cs.SD语音【7】为同一篇

作者:Hyun-Wook Yoon,Ohsung Kwon,Hoyeon Lee,Ryuichi Yamamoto,Eunwoo Song,Jae-Min Kim,Min-Jae Hwang
备注:Accepted in INTERSPEECH2022
摘要:本文提出了一种基于预训练语言模型(LM)的情感预测方法的有效情感文本到语音(TTS)系统。与需要手动定义情感类等辅助输入的传统系统不同,我们的系统直接从输入文本中估计情感相关属性。具体来说,我们利用生成预训练变换器(GPT)-3分别联合预测情感类别及其表示情感粗糙和精细属性的强度。然后,将这些属性组合在情感嵌入空间中,并用作TTS模型的条件特征,以生成输出语音信号。因此,该系统只能从文本中产生情感语音,而无需任何辅助输入。此外,由于GPT-3能够捕捉连续句子之间的情感语境,因此该方法可以有效地处理情感语音的段落级生成。
摘要:This paper proposes an effective emotional text-to-speech (TTS) system with a pre-trained language model (LM)-based emotion prediction method. Unlike conventional systems that require auxiliary inputs such as manually defined emotion classes, our system directly estimates emotion-related attributes from the input text. Specifically, we utilize generative pre-trained transformer (GPT)-3 to jointly predict both an emotion class and its strength in representing emotions coarse and fine properties, respectively. Then, these attributes are combined in the emotional embedding space and used as conditional features of the TTS model for generating output speech signals. Consequently, the proposed system can produce emotional speech only from text without any auxiliary inputs. Furthermore, because the GPT-3 enables to capture emotional context among the consecutive sentences, the proposed method can effectively handle the paragraph-level generation of emotional speech.


【17】 FeaRLESS: Feature Refinement Loss for Ensembling Self-Supervised  Learning Features in Robust End-to-end Speech Recognition

标题:无畏:稳健端到端语音识别中集成自监督学习特征的特征细化损失

链接:https://arxiv.org/abs/2206.15056

* 与cs.SD语音【8】为同一篇

作者:Szu-Jui Chen,Jiamin Xie,John H. L. Hansen
备注:Accepted for Interspeech 2022
摘要:自监督学习表示(SSLR)为许多领域的下游任务带来了强大的特征。最近,一些SSLR在自动语音识别(ASR)基准语料库上显示了有希望的结果。然而,以前的研究仅表明孤立SSLR的性能作为ASR模型的输入特征。在本研究中,我们提议在端到端(E2E)ASR模型中使用各种融合方法来研究不同SSLR组合的有效性。此外,我们将显示这些提取的SSLR之间存在相关性。因此,我们进一步提出了用于去相关的特征细化损失,以有效地组合输入特征集。为了评估,我们表明,对于《华尔街日报》和无畏步骤挑战(FSC)语料库,拟议的“无畏学习特征”比没有拟议特征细化损失的系统表现更好。
摘要:Self-supervised learning representations (SSLR) have resulted in robust features for downstream tasks in many fields. Recently, several SSLRs have shown promising results on automatic speech recognition (ASR) benchmark corpora. However, previous studies have only shown performance for solitary SSLRs as an input feature for ASR models. In this study, we propose to investigate the effectiveness of diverse SSLR combinations using various fusion methods within end-to-end (E2E) ASR models. In addition, we will show there are correlations between these extracted SSLRs. As such, we further propose a feature refinement loss for decorrelation to efficiently combine the set of input features. For evaluation, we show that the proposed 'FeaRLESS learning features' perform better than systems without the proposed feature refinement loss for both the WSJ and Fearless Steps Challenge (FSC) corpora.


机器翻译,仅供参考