今天跟大家分享一篇语音相关的论文合集:cs.SD语音11篇,eess.AS音频处理17篇。本文经arXiv每日学术速递授权转载
【1】 Robust Time Series Denoising with Learnable Wavelet Packet Transform标题:基于可学习小波包变换的时间序列稳健降噪
链接:https://arxiv.org/abs/2206.06126
作者:Gaetan Frusque,Olga Fink备注:15 pages, 13 figures, 8 tables摘要:在许多应用中,信号去噪通常是后续分析或学习任务之前的第一个预处理步骤。在本文中,我们建议应用一种受信号处理启发的深度学习去噪模型,这是一种可学习的小波包变换。该算法具有显著的学习能力,可解释参数少,初始化直观。我们提出了一种参数的后学习修改,以适应不同噪声水平的去噪。我们在两个案例研究中评估了所提方法的性能,并将其与其他最先进的方法进行了比较,包括小波schrinkage去噪、卷积神经网络、自动编码器和U-net深度模型。第一个案例研究基于设计的函数,这些函数通常用于研究算法的去噪特性。第二个案例研究是音频背景移除任务。我们展示了该算法如何与信号处理方法的通用性和深度学习方法的学习能力相关联。特别是,我们评估了在用于训练的课堂内外对结构化噪声信号所获得的去噪性能。除了在训练课内外的信号去噪方面具有良好的性能外,我们的方法在添加不同的噪声级、噪声类型和伪影时表现出特别的鲁棒性。摘要:In many applications, signal denoising is often the first pre-processing step before any subsequent analysis or learning task. In this paper, we propose to apply a deep learning denoising model inspired by a signal processing, a learnable version of wavelet packet transform. The proposed algorithm has signficant learning capabilities with few interpretable parameters and has an intuitive initialisation. We propose a post-learning modification of the parameters to adapt the denoising to different noise levels. We evaluate the performance of the proposed methodology on two case studies and compare it to other state of the art approaches, including wavelet schrinkage denoising, convolutional neural network, autoencoder and U-net deep models. The first case study is based on designed functions that have typically been used to study denoising properties of the algorithms. The second case study is an audio background removal task. We demonstrate how the proposed algorithm relates to the universality of signal processing methods and the learning capabilities of deep learning approaches. In particular, we evaluate the obtained denoising performances on structured noisy signals inside and outside the classes used for training. In addition to having good performance in denoising signals inside and outside to the training class, our method shows to be particularly robust when different noise levels, noise types and artifacts are added.
【2】 Optimizing musical chord inversions using the cartesian coordinate system
标题:利用笛卡尔坐标系优化音乐和弦反转
链接:https://arxiv.org/abs/2206.06117
摘要:在古典音乐和任何当代音乐流派中,演奏所用的音调元素或音符都是相同的。一首曲子中给定实例的无数和弦的可能性使演奏总体上非常复杂和先进。这个理论听起来很琐碎,但应用程序有很多选择,每种选择都会导致无可争议的不同结果,其特点是科学和音乐原则。和弦及其重要性不言而喻。和弦是一起演奏的一串音符。就科学家而言,这是一组音调频率一起回响,产生辅音/不谐音的声音。众所周知,和弦的音符可以重新排列,以产生同一和弦的各种voicings(1),这使作曲家/演奏者能够选择最理想的一个来表达他们想要表达的情感。尽管有许多可能性,但认为只有一种适合音调运动特定情况的声音是科学的。在这项研究中,我们试图通过将和弦视为三维笛卡尔坐标系中的点来寻找最佳发声,并进一步加深对音乐理论中数学的基本理解。摘要:In classical music and in any genre of contemporary music, the tonal elements or notes used for playing are the same. The numerous possibilities of chords for a given instance in a piece make the playing, in general, very intricate, and advanced. The theory sounds quite trivial, yet the application has vast options, each leading to inarguably different outcomes, characterized by scientific and musical principles. Chords and their importance are self-explanatory. A chord is a bunch of notes played together. As far as scientists are concerned, it is a set of tonal frequencies ringing together resulting in a consonant/dissonant sound. It is well-known that the notes of a chord can be rearranged to come up with various voicings (1) of the same chord which enables a composer/player to choose the most optimal one to convey the emotion they wish to convey. Though there are numerous possibilities, it is scientific to think that there is just one appropriate voicing for a particular situation of tonal movements. In this study, we attempt to find the optimal voicings by considering chords to be points in a 3-dimensional cartesian coordinate system and further the fundamental understanding of mathematics in music theory.
【3】 Low-complexity deep learning frameworks for acoustic scene classification
标题:用于声学场景分类的低复杂度深度学习框架
链接:https://arxiv.org/abs/2206.06057
作者:Lam Pham,Dat Ngo,Anahid Jalali,Alexander Schindler摘要:在这份报告中,我们提出了低复杂度的声学场景分类深度学习框架(ASC)。提出的框架可分为四个主要步骤:前端谱图提取、在线数据增强、后端分类和预测概率的后期融合。特别是,我们最初将音频记录转换为Mel、Gammatone和CQT光谱图。接下来,使用随机裁剪、Specaugment和Mixup等数据增强方法生成增强的光谱图,然后再将其输入基于深度学习的分类器。最后,为了获得最佳性能,我们融合了从三个单独分类器中获得的概率,这些分类器使用三种类型的谱图进行独立训练。我们在DCASE 2022 Task 1开发数据集上进行的实验充分满足了低复杂性的要求,实现了60.1%的最佳分类精度,将DCASE基线提高了17.2%。摘要:In this report, we presents low-complexity deep learning frameworks for acoustic scene classification (ASC). The proposed frameworks can be separated into four main steps: Front-end spectrogram extraction, online data augmentation, back-end classification, and late fusion of predicted probabilities. In particular, we initially transform audio recordings into Mel, Gammatone, and CQT spectrograms. Next, data augmentation methods of Random Cropping, Specaugment, and Mixup are then applied to generate augmented spectrograms before being fed into deep learning based classifiers. Finally, to achieve the best performance, we fuse probabilities which obtained from three individual classifiers, which are independently-trained with three type of spectrograms. Our experiments conducted on DCASE 2022 Task 1 Development dataset have fullfiled the requirement of low-complexity and achieved the best classification accuracy of 60.1%, improving DCASE baseline by 17.2%.
【4】 Improvement of Serial Approach to Anomalous Sound Detection by Incorporating Two Binary Cross-Entropies for Outlier Exposure
标题:融合两个二进制交叉熵的异常声序列检测方法的改进
链接:https://arxiv.org/abs/2206.05929
作者:Ibuki Kuroyanagi,Tomoki Hayashi,Kazuya Takeda,Tomoki Toda备注:5 pages, 3 figures, 3 tables, EUSIPCO 2022摘要:异常声音检测系统必须仅使用正常音频数据检测未知、非典型声音。传统方法使用串行方法,将异常值暴露(OE)与inlier建模(IM)相结合,前者将正常和伪异常数据分类并获得嵌入,后者对嵌入的概率分布进行建模。虽然串行方法由于OE的强大特征提取和IM的鲁棒性而显示出很高的性能,但OE仍然存在一个问题,即当正常数据和伪异常数据太相似或太不同时,OE不能很好地工作。为了明确区分这些数据,该方法在训练OE时使用两个二进制交叉熵的多任务学习。第一种是一种损失,它对目标机器发出的声音进行分类,该产品用于处理正常数据和伪异常数据过于相似的情况。第二种是识别声音是否从目标机器发出的损失,它处理正常数据和伪异常数据相差太大的情况。我们使用DCASE 2021任务2数据集进行了实验。我们提出的单模型方法在AUC方面比组合多个模型的排名靠前的方法要好2.1%。摘要:Anomalous sound detection systems must detect unknown, atypical sounds using only normal audio data. Conventional methods use the serial method, a combination of outlier exposure (OE), which classifies normal and pseudo-anomalous data and obtains embedding, and inlier modeling (IM), which models the probability distribution of the embedding. Although the serial method shows high performance due to the powerful feature extraction of OE and the robustness of IM, OE still has a problem that doesn't work well when the normal and pseudo-anomalous data are too similar or too different. To explicitly distinguish these data, the proposed method uses multi-task learning of two binary cross-entropies when training OE. The first is a loss that classifies the sound of the target machine to which product it is emitted from, which deals with the case where the normal data and the pseudo-anomalous data are too similar. The second is a loss that identifies whether the sound is emitted from the target machine or not, which deals with the case where the normal data and the pseudo-anomalous data are too different. We perform our experiments with DCASE 2021 Task~2 dataset. Our proposed single-model method outperforms the top-ranked method, which combines multiple models, by 2.1% in AUC.
【5】 Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques
标题:描述和讨论DCASE 2022挑战任务2:应用领域泛化技术进行机器状态监测的无监督异常声音检测
链接:https://arxiv.org/abs/2206.05876
作者:Kota Dohi,Keisuke Imoto,Noboru Harada,Daisuke Niizumi,Yuma Koizumi,Tomoya Nishida,Harsh Purohit,Takashi Endo,Masaaki Yamamoto,Yohei Kawaguchi备注:arXiv admin note: substantial text overlap with arXiv:2106.04492摘要:我们介绍了声学场景和事件检测与分类(DCASE)2022挑战任务2的任务描述:“应用领域泛化技术进行机器状态监测的无监督异常声音检测(ASD)”。领域转移是ASD系统应用中的一个关键问题。由于域移动会改变数据的声学特性,因此在源域中训练的模型对于目标域的性能较差。在DCASE 2021挑战任务2中,我们组织了一个ASD任务来处理域转移。在这项任务中,假设域移动的发生是已知的。然而,在实践中,可能不会给出每个样本的域,域移动可能会隐式发生。在2022年的任务2中,我们重点关注领域泛化技术,该技术可以检测异常,而不管领域发生了什么变化。具体来说,测试数据中没有给出每个样本的域,所有域只允许一个阈值。我们将在挑战提交截止日期后添加挑战结果和对提交内容的分析。摘要:We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2022 Challenge Task 2: "Unsupervised anomalous sound detection (ASD) for machine condition monitoring applying domain generalization techniques". Domain shifts are a critical problem for the application of ASD systems. Because domain shifts can change the acoustic characteristics of data, a model trained in a source domain performs poorly for a target domain. In DCASE 2021 Challenge Task 2, we organized an ASD task for handling domain shifts. In this task, it was assumed that the occurrences of domain shifts are known. However, in practice, the domain of each sample may not be given, and the domain shifts can occur implicitly. In 2022 Task 2, we focus on domain generalization techniques that detects anomalies regardless of the domain shifts. Specifically, the domain of each sample is not given in the test data and only one threshold is allowed for all domains. We will add challenge results and analysis of the submissions after the challenge submission deadline.
【6】 Multi-instrument Music Synthesis with Spectrogram Diffusion
标题:基于谱图扩散的多乐器音乐合成
链接:https://arxiv.org/abs/2206.05408
作者:Curtis Hawthorne,Ian Simon,Adam Roberts,Neil Zeghidour,Josh Gardner,Ethan Manilow,Jesse Engel摘要:理想的音乐合成器应具有交互性和表现力,能够为乐器和音符的任意组合实时生成高保真音频。最近的神经合成器在特定领域的模型(只提供特定乐器的详细控制)和原始波形模型(可以对所有音乐进行训练,但控制最少且生成速度较慢)之间进行了权衡。在这项工作中,我们将重点放在神经合成器的中间地带,该合成器可以通过MIDI序列和任意组合的乐器实时生成音频。这使得可以使用单一模型在广泛的转录数据集上进行训练,进而提供对广泛仪器的组成和仪器的注释级控制。我们使用一个简单的两阶段过程:MIDI到频谱图,使用编码器-解码器转换器,然后频谱图到音频,使用生成对抗网络(GAN)频谱图转换器。我们比较了将解码器作为自回归模型和去噪扩散概率模型(DDPM)进行训练,发现DDPM方法在定性和通过音频重建和Fr距离度量衡量方面都优于传统方法。考虑到这种方法的交互性和通用性,我们发现这是实现交互式和表达性神经合成的第一步,可用于任意组合的乐器和音符。摘要:An ideal music synthesizer should be both interactive and expressive, generating high-fidelity audio in realtime for arbitrary combinations of instruments and notes. Recent neural synthesizers have exhibited a tradeoff between domain-specific models that offer detailed control of only specific instruments, or raw waveform models that can train on all of music but with minimal control and slow generation. In this work, we focus on a middle ground of neural synthesizers that can generate audio from MIDI sequences with arbitrary combinations of instruments in realtime. This enables training on a wide range of transcription datasets with a single model, which in turn offers note-level control of composition and instrumentation across a wide range of instruments. We use a simple two-stage process: MIDI to spectrograms with an encoder-decoder Transformer, then spectrograms to audio with a generative adversarial network (GAN) spectrogram inverter. We compare training the decoder as an autoregressive model and as a Denoising Diffusion Probabilistic Model (DDPM) and find that the DDPM approach is superior both qualitatively and as measured by audio reconstruction and Fr\'echet distance metrics. Given the interactivity and generality of this approach, we find this to be a promising first step towards interactive and expressive neural synthesis for arbitrary combinations of instruments and notes.
【7】 AHD ConvNet for Speech Emotion Classification
标题:用于语音情感分类的AND转换网
链接:https://arxiv.org/abs/2206.05286
作者:Asfand Ali,Danial Nasir,Mohammad Hassan Jawad摘要:人工智能领域的成就被用于推动计算和智能机器的制造,以方便人类和改善用户体验。情绪对人们来说是最基本的,它影响思维和日常练习,如通信、学习和指导。语音情感识别是这方面的一个研究领域,在这项工作中,我们提出了一种新的mel谱图学习方法,其中我们的模型使用数据点从流行的CREMA-D数据集中给定的wav形式语音注释中学习情感。我们的模型使用对数mel谱图作为特征,mel数=64。与用于解决情感语音识别问题的其他方法相比,它所需的训练时间更少。摘要:Accomplishments in the field of artificial intelligence are utilized in the advancement of computing and making of intelligent machines for facilitating mankind and improving user experience. Emotions are rudimentary for people, affecting thinking and ordinary exercises like correspondence, learning and direction. Speech emotion recognition is domain of interest in this regard and in this work, we propose a novel mel spectrogram learning approach in which our model uses the datapoints to learn emotions from the given wav form voice notes in the popular CREMA-D dataset. Our model uses log mel-spectrogram as feature with number of mels = 64. It took less training time compared to other approaches used to address the problem of emotion speech recognition.
【8】 Automated Evaluation of Standardized Dementia Screening Tests
标题:痴呆症标准化筛查测验的自动化评价
链接:https://arxiv.org/abs/2206.06208
作者:Franziska Braun,Markus Förstel,Bastian Oppermann,Andreas Erzigkeit,Thomas Hillemacher,Hartmut Lehfeld,Korbinian Riedhammer备注:Submitted to Interspeech 2022. arXiv admin note: text overlap with arXiv:2206.05018摘要:对于痴呆症筛查和监测,标准化测试在临床常规中起着关键作用,因为它们旨在通过测量各种认知任务的表现来最大限度地减少主观性。在本文中,我们报告了一项研究,该研究包括一项半标准化的病史记录,然后进行两项标准化的神经心理学测试,即SKT和CERAD-NB。这些测试包括命名对象、学习单词列表等基本任务,还包括MMSE等广泛使用的工具。大多数任务都是口头完成的,因此应该适合基于成绩单的自动评分。对于第一批30名患者,我们分析了专家手动评估和基于手动和自动转录的自动评估之间的相关性。对于SKT和CERAD-NB,我们使用手动转录本观察到高度到完美的相关性;对于某些相关性较低的任务,自动评分比人工参考更严格,因为它仅限于音频。使用自动转录,相关性会像预期的那样下降,并与识别精度相关;然而,我们仍然观察到高达0.98(SKT)和0.85(CERAD-NB)的高度相关性。我们表明,使用词汇替代有助于减少识别错误,并随后改善与专家得分的相关性。摘要:For dementia screening and monitoring, standardized tests play a key role in clinical routine since they aim at minimizing subjectivity by measuring performance on a variety of cognitive tasks. In this paper, we report on a study that consists of a semi-standardized history taking followed by two standardized neuropsychological tests, namely the SKT and the CERAD-NB. The tests include basic tasks such as naming objects, learning word lists, but also widely used tools such as the MMSE. Most of the tasks are performed verbally and should thus be suitable for automated scoring based on transcripts. For the first batch of 30 patients, we analyze the correlation between expert manual evaluations and automatic evaluations based on manual and automatic transcriptions. For both SKT and CERAD-NB, we observe high to perfect correlations using manual transcripts; for certain tasks with lower correlation, the automatic scoring is stricter than the human reference since it is limited to the audio. Using automatic transcriptions, correlations drop as expected and are related to recognition accuracy; however, we still observe high correlations of up to 0.98 (SKT) and 0.85 (CERAD-NB). We show that using word alternatives helps to mitigate recognition errors and subsequently improves correlation with expert scores.
【9】 Toward Zero Oracle Word Error Rate on the Switchboard Benchmark
标题:在Switchboard基准测试中走向零Oracle字错误率
链接:https://arxiv.org/abs/2206.06192
作者:Arlo Faria,Adam Janin,Korbinian Riedhammer,Sidhi Adkoli备注:Submitted to Interspeech 2022摘要:“交换机基准测试”是自动语音识别(ASR)研究中非常著名的测试集,为声称人类水平转录准确性的系统建立了记录设置性能。这项工作强调了该评估中鲜为人知的实际考虑因素,通过更正参考记录和偏离官方评分方法,证明了单词错误率(WER)的重大改进。在这个更详细和可复制的方案中,即使是商业ASR系统也可以得分低于5%的WER,并且研究系统的既定记录降低到2.3%。提出了一种替代的转录精度指标,该指标不惩罚删除,并且似乎对人类与机器的性能更具区分性。虽然商用ASR系统仍低于这一阈值,但研究表明,一个研究系统明显超过了商用人类语音识别的准确性。这项工作还探讨了如何使用标准化评分工具,通过在一系列备选方案中选择最佳方案来计算oracle WER。将短语替代表示与话语级N-最佳列表和单词级数据结构进行比较;使用密集格和添加词汇表外的单词,这实现了0.18%的oracle WER。摘要:The "Switchboard benchmark" is a very well-known test set in automatic speech recognition (ASR) research, establishing record-setting performance for systems that claim human-level transcription accuracy. This work highlights lesser-known practical considerations of this evaluation, demonstrating major improvements in word error rate (WER) by correcting the reference transcriptions and deviating from the official scoring methodology. In this more detailed and reproducible scheme, even commercial ASR systems can score below 5\% WER and the established record for a research system is lowered to 2.3%. An alternative metric of transcript precision is proposed, which does not penalize deletions and appears to be more discriminating for human vs. machine performance. While commercial ASR systems are still below this threshold, a research system is shown to clearly surpass the accuracy of commercial human speech recognition. This work also explores using standardized scoring tools to compute oracle WER by selecting the best among a list of alternatives. A phrase alternatives representation is compared to utterance-level N-best lists and word-level data structures; using dense lattices and adding out-of-vocabulary words, this achieves an oracle WER of 0.18%.
【10】 Signal-informed DNN-based DOA Estimation combining an External Microphone and GCC-PHAT Features
标题:基于信号通知离散神经网络结合外置麦克风和GCC特征的波达方向估计
链接:https://arxiv.org/abs/2206.05606
作者:Ulrik Kowalk,Simon Doclo,Joerg Bitzer摘要:为了在多说话人环境中利用麦克风阵列估计目标说话人的到达方向(DOA),本文提出了一种利用目标说话人外部麦克风可用性的信号通知方法。该方法将二进制掩码应用于卷积神经网络的GCC-PHAT输入特征,其中二进制掩码基于外部麦克风信号的功率分布计算。对最多有四个干扰说话人的混响场景的实验结果表明,信号通知掩蔽提高了定位精度,而不需要任何关于干扰说话人的知识。摘要:Aiming at estimating the direction of arrival (DOA) of a desired speaker in a multi-talker environment using a microphone array, in this paper we propose a signal-informed method exploiting the availability of an external microphone attached to the desired speaker. The proposed method applies a binary mask to the GCC-PHAT input features of a convolutional neural network, where the binary mask is computed based on the power distribution of the external microphone signal. Experimental results for a reverberant scenario with up to four interfering speakers demonstrate that the signal-informed masking improves the localization accuracy, without requiring any knowledge about the interfering speakers.
【11】 Svadhyaya system for the Second Diagnosing COVID-19 using Acoustics Challenge 2021
标题:声学挑战赛2021第二次诊断新冠肺炎的Svadhyaya系统
链接:https://arxiv.org/abs/2206.05462
作者:Deepak Mittal,Amir H. Poorjam,Debottam Dutta,Debarpan Bhattacharya,Zemin Yu,Sriram Ganapathy,Maneesh Singh摘要:本报告描述了在第二次DiCOVA挑战中,使用三种不同的声学模式,即语音、呼吸和咳嗽,检测2019冠状病毒疾病阳性的系统。所提出的系统基于4种不同方法的组合,每种方法都更多地关注问题的一个方面,在呼吸、咳嗽和语音轨迹中分别达到86.41、77.60和84.55的盲测试AUC,在这三个轨迹的融合中达到85.37的AUC。摘要:This report describes the system used for detecting COVID-19 positives using three different acoustic modalities, namely speech, breathing, and cough in the second DiCOVA challenge. The proposed system is based on the combination of 4 different approaches, each focusing more on one aspect of the problem, and reaches the blind test AUCs of 86.41, 77.60, and 84.55, in the breathing, cough, and speech tracks, respectively, and the AUC of 85.37 in the fusion of these three tracks.
【1】 Realistic Gramophone Noise Synthesis using a Diffusion Model标题:基于扩散模型的真实留声机噪声合成
链接:https://arxiv.org/abs/2206.06259
作者:Eloi Moliner,Vesa Välimäki备注:submitted to DAFx 20in22摘要:本文介绍了一种新的数据驱动策略,用于合成留声机噪声纹理。采用扩散概率模型产生高度真实的准周期噪声。提出的模型旨在生成长度等于一个圆盘旋转的样本,但也提出了一种在旋转之间生成合理周期变化的方法。引导方法也被用作调节方法,其中通过手动调谐信号处理生成的音频信号通过反向扩散进行细化,以显示更真实的声音。这种方法已经在一次主观听力测试中进行了评估,在该测试中,参与者往往无法识别合成信号和真实信号。用最好的无条件方法产生的合成噪声在统计上与真实噪声记录无法区分。这项工作显示了扩散模型在高逼真度音频合成任务中的潜力。摘要:This paper introduces a novel data-driven strategy for synthesizing gramophone noise textures. A diffusion probabilistic model is applied to generate highly realistic quasiperiodic noises. The proposed model is designed to generate samples of length equal to one disk revolution, but a method to generate plausible periodic variations between revolutions is also proposed. A guided approach is also applied as a conditioning method, where an audio signal generated with manually-tuned signal processing is refined via reverse diffusion to appear more realistically sounding. The method has been evaluated in a subjective listening test, in which the participants were often unable to recognize the synthesized signals from the real ones. The synthetic noises produced with the best proposed unconditional method are statistically indistinguishable from real noise recordings. This work shows the potential of diffusion models for highly realistic audio synthesis tasks.
【2】 Automated Evaluation of Standardized Dementia Screening Tests
标题:痴呆症标准化筛查测验的自动化评价
链接:https://arxiv.org/abs/2206.06208
作者:Franziska Braun,Markus Förstel,Bastian Oppermann,Andreas Erzigkeit,Thomas Hillemacher,Hartmut Lehfeld,Korbinian Riedhammer备注:Submitted to Interspeech 2022. arXiv admin note: text overlap with arXiv:2206.05018摘要:对于痴呆症筛查和监测,标准化测试在临床常规中起着关键作用,因为它们旨在通过测量各种认知任务的表现来最大限度地减少主观性。在本文中,我们报告了一项研究,该研究包括一项半标准化的病史记录,然后进行两项标准化的神经心理学测试,即SKT和CERAD-NB。这些测试包括命名对象、学习单词列表等基本任务,还包括MMSE等广泛使用的工具。大多数任务都是口头完成的,因此应该适合基于成绩单的自动评分。对于第一批30名患者,我们分析了专家手动评估和基于手动和自动转录的自动评估之间的相关性。对于SKT和CERAD-NB,我们使用手动转录本观察到高度到完美的相关性;对于某些相关性较低的任务,自动评分比人工参考更严格,因为它仅限于音频。使用自动转录,相关性会像预期的那样下降,并与识别精度相关;然而,我们仍然观察到高达0.98(SKT)和0.85(CERAD-NB)的高度相关性。我们表明,使用词汇替代有助于减少识别错误,并随后改善与专家得分的相关性。摘要:For dementia screening and monitoring, standardized tests play a key role in clinical routine since they aim at minimizing subjectivity by measuring performance on a variety of cognitive tasks. In this paper, we report on a study that consists of a semi-standardized history taking followed by two standardized neuropsychological tests, namely the SKT and the CERAD-NB. The tests include basic tasks such as naming objects, learning word lists, but also widely used tools such as the MMSE. Most of the tasks are performed verbally and should thus be suitable for automated scoring based on transcripts. For the first batch of 30 patients, we analyze the correlation between expert manual evaluations and automatic evaluations based on manual and automatic transcriptions. For both SKT and CERAD-NB, we observe high to perfect correlations using manual transcripts; for certain tasks with lower correlation, the automatic scoring is stricter than the human reference since it is limited to the audio. Using automatic transcriptions, correlations drop as expected and are related to recognition accuracy; however, we still observe high correlations of up to 0.98 (SKT) and 0.85 (CERAD-NB). We show that using word alternatives helps to mitigate recognition errors and subsequently improves correlation with expert scores.
【3】 Toward Zero Oracle Word Error Rate on the Switchboard Benchmark
标题:在Switchboard基准测试中走向零Oracle字错误率
链接:https://arxiv.org/abs/2206.06192
作者:Arlo Faria,Adam Janin,Korbinian Riedhammer,Sidhi Adkoli备注:Submitted to Interspeech 2022摘要:“交换机基准测试”是自动语音识别(ASR)研究中非常著名的测试集,为声称人类水平转录准确性的系统建立了记录设置性能。这项工作强调了该评估中鲜为人知的实际考虑因素,通过更正参考记录和偏离官方评分方法,证明了单词错误率(WER)的重大改进。在这个更详细和可复制的方案中,即使是商业ASR系统也可以得分低于5%的WER,并且研究系统的既定记录降低到2.3%。提出了一种替代的转录精度指标,该指标不惩罚删除,并且似乎对人类与机器的性能更具区分性。虽然商用ASR系统仍低于这一阈值,但研究表明,一个研究系统明显超过了商用人类语音识别的准确性。这项工作还探讨了如何使用标准化评分工具,通过在一系列备选方案中选择最佳方案来计算oracle WER。将短语替代表示与话语级N-最佳列表和单词级数据结构进行比较;使用密集格和添加词汇表外的单词,这实现了0.18%的oracle WER。摘要:The "Switchboard benchmark" is a very well-known test set in automatic speech recognition (ASR) research, establishing record-setting performance for systems that claim human-level transcription accuracy. This work highlights lesser-known practical considerations of this evaluation, demonstrating major improvements in word error rate (WER) by correcting the reference transcriptions and deviating from the official scoring methodology. In this more detailed and reproducible scheme, even commercial ASR systems can score below 5\% WER and the established record for a research system is lowered to 2.3%. An alternative metric of transcript precision is proposed, which does not penalize deletions and appears to be more discriminating for human vs. machine performance. While commercial ASR systems are still below this threshold, a research system is shown to clearly surpass the accuracy of commercial human speech recognition. This work also explores using standardized scoring tools to compute oracle WER by selecting the best among a list of alternatives. A phrase alternatives representation is compared to utterance-level N-best lists and word-level data structures; using dense lattices and adding out-of-vocabulary words, this achieves an oracle WER of 0.18%.
【4】 AmbiSep: Ambisonic-to-Ambisonic Reverberant Speech Separation Using Transformer Networks
标题:AmbiSep:基于Transformer网络的混响语音分离
链接:https://arxiv.org/abs/2206.06184
作者:Adrian Herzog,Srikanth Raj Chetupalli,Emanuël A. P. Habets备注:Preprint submitted to IWAENC 2022 (this https URL)摘要:考虑一个包含多个混响语音信号的多通道环境音录音。从混音中盲取对应于单个语音源的混响环境音信号是一项具有挑战性的任务,因为它需要估计每个源的多个信号通道。在这项工作中,我们提出了一种基于深度神经网络的平面波域掩蔽方法AmbiSep来解决这一问题。掩蔽网络在三路径处理配置中使用学习的特征表示和变换器。我们在空间化WSJ0-2mix数据集上对所提出的网络结构进行了训练和评估,结果表明,该方法在盲测试集上实现了17.7 dB的多通道尺度不变信噪比改善,同时保留了分离声音的空间特征。摘要:Consider a multichannel Ambisonic recording containing a mixture of several reverberant speech signals. Retreiving the reverberant Ambisonic signals corresponding to the individual speech sources blindly from the mixture is a challenging task as it requires to estimate multiple signal channels for each source. In this work, we propose AmbiSep, a deep neural network-based plane-wave domain masking approach to solve this task. The masking network uses learned feature representations and transformers in a triple-path processing configuration. We train and evaluate the proposed network architecture on a spatialized WSJ0-2mix dataset, and show that the method achieves a multichannel scale-invariant signal-to-distortion ratio improvement of 17.7 dB on the blind test set, while preserving the spatial characteristics of the separated sounds.
【5】 DCASE 2022 Challenge Task 6B: Language-Based Audio Retrieval
标题:DCASE 2022挑战任务6B:基于语言的音频检索
链接:https://arxiv.org/abs/2206.06108
作者:Huang Xie,Samuel Lipping,Tuomas Virtanen摘要:在本报告中,我们介绍了DCASE 2022挑战任务6:基于语言的音频检索子任务的子任务B的任务设置和基线系统。对于这个子任务,Clotho v2数据集被用作开发数据集,另外一个数据集由1000个音频字幕对组成,作为评估数据集。我们使用开发数据集对基线系统进行训练,并在评估数据集上对其进行评估,以提供此子任务的一些初始结果。摘要:In this report, we introduce the task setup and the baseline system for the sub-task B of the DCASE 2022 Challenge Task 6: language-based audio retrieval subtask. For this subtask, the Clotho v2 dataset is utilized as the development dataset, and an additional dataset consisting of 1,000 audio-caption pairs as the evaluation dataset. We train the baseline system with the development dataset, and evaluate it on the evaluation dataset to provide some initial results for this subtask.
【6】 Signal-informed DNN-based DOA Estimation combining an External Microphone and GCC-PHAT Features
标题:基于信号通知离散神经网络结合外置麦克风和GCC特征的波达方向估计
链接:https://arxiv.org/abs/2206.05606
作者:Ulrik Kowalk,Simon Doclo,Joerg Bitzer摘要:为了在多说话人环境中利用麦克风阵列估计目标说话人的到达方向(DOA),本文提出了一种利用目标说话人外部麦克风可用性的信号通知方法。该方法将二进制掩码应用于卷积神经网络的GCC-PHAT输入特征,其中二进制掩码基于外部麦克风信号的功率分布计算。对最多有四个干扰说话人的混响场景的实验结果表明,信号通知掩蔽提高了定位精度,而不需要任何关于干扰说话人的知识。摘要:Aiming at estimating the direction of arrival (DOA) of a desired speaker in a multi-talker environment using a microphone array, in this paper we propose a signal-informed method exploiting the availability of an external microphone attached to the desired speaker. The proposed method applies a binary mask to the GCC-PHAT input features of a convolutional neural network, where the binary mask is computed based on the power distribution of the external microphone signal. Experimental results for a reverberant scenario with up to four interfering speakers demonstrate that the signal-informed masking improves the localization accuracy, without requiring any knowledge about the interfering speakers.
【7】 Svadhyaya system for the Second Diagnosing COVID-19 using Acoustics Challenge 2021
标题:声学挑战赛2021第二次诊断新冠肺炎的Svadhyaya系统
链接:https://arxiv.org/abs/2206.05462
作者:Deepak Mittal,Amir H. Poorjam,Debottam Dutta,Debarpan Bhattacharya,Zemin Yu,Sriram Ganapathy,Maneesh Singh摘要:本报告描述了在第二次DiCOVA挑战中,使用三种不同的声学模式,即语音、呼吸和咳嗽,检测2019冠状病毒疾病阳性的系统。所提出的系统基于4种不同方法的组合,每种方法都更多地关注问题的一个方面,在呼吸、咳嗽和语音轨迹中分别达到86.41、77.60和84.55的盲测试AUC,在这三个轨迹的融合中达到85.37的AUC。摘要:This report describes the system used for detecting COVID-19 positives using three different acoustic modalities, namely speech, breathing, and cough in the second DiCOVA challenge. The proposed system is based on the combination of 4 different approaches, each focusing more on one aspect of the problem, and reaches the blind test AUCs of 86.41, 77.60, and 84.55, in the breathing, cough, and speech tracks, respectively, and the AUC of 85.37 in the fusion of these three tracks.
【8】 Robust Time Series Denoising with Learnable Wavelet Packet Transform
标题:基于可学习小波包变换的时间序列稳健降噪
链接:https://arxiv.org/abs/2206.06126
作者:Gaetan Frusque,Olga Fink备注:15 pages, 13 figures, 8 tables摘要:在许多应用中,信号去噪通常是后续分析或学习任务之前的第一个预处理步骤。在本文中,我们建议应用一种受信号处理启发的深度学习去噪模型,这是一种可学习的小波包变换。该算法具有显著的学习能力,可解释参数少,初始化直观。我们提出了一种参数的后学习修改,以适应不同噪声水平的去噪。我们在两个案例研究中评估了所提方法的性能,并将其与其他最先进的方法进行了比较,包括小波schrinkage去噪、卷积神经网络、自动编码器和U-net深度模型。第一个案例研究基于设计的函数,这些函数通常用于研究算法的去噪特性。第二个案例研究是音频背景移除任务。我们展示了该算法如何与信号处理方法的通用性和深度学习方法的学习能力相关联。特别是,我们评估了在用于训练的课堂内外对结构化噪声信号所获得的去噪性能。除了在训练课内外的信号去噪方面具有良好的性能外,我们的方法在添加不同的噪声级、噪声类型和伪影时表现出特别的鲁棒性。摘要:In many applications, signal denoising is often the first pre-processing step before any subsequent analysis or learning task. In this paper, we propose to apply a deep learning denoising model inspired by a signal processing, a learnable version of wavelet packet transform. The proposed algorithm has signficant learning capabilities with few interpretable parameters and has an intuitive initialisation. We propose a post-learning modification of the parameters to adapt the denoising to different noise levels. We evaluate the performance of the proposed methodology on two case studies and compare it to other state of the art approaches, including wavelet schrinkage denoising, convolutional neural network, autoencoder and U-net deep models. The first case study is based on designed functions that have typically been used to study denoising properties of the algorithms. The second case study is an audio background removal task. We demonstrate how the proposed algorithm relates to the universality of signal processing methods and the learning capabilities of deep learning approaches. In particular, we evaluate the obtained denoising performances on structured noisy signals inside and outside the classes used for training. In addition to having good performance in denoising signals inside and outside to the training class, our method shows to be particularly robust when different noise levels, noise types and artifacts are added.
【9】 Optimizing musical chord inversions using the cartesian coordinate system
标题:利用笛卡尔坐标系优化音乐和弦反转
链接:https://arxiv.org/abs/2206.06117
摘要:在古典音乐和任何当代音乐流派中,演奏所用的音调元素或音符都是相同的。一首曲子中给定实例的无数和弦的可能性使演奏总体上非常复杂和先进。这个理论听起来很琐碎,但应用程序有很多选择,每种选择都会导致无可争议的不同结果,其特点是科学和音乐原则。和弦及其重要性不言而喻。和弦是一起演奏的一串音符。就科学家而言,这是一组音调频率一起回响,产生辅音/不谐音的声音。众所周知,和弦的音符可以重新排列,以产生同一和弦的各种voicings(1),这使作曲家/演奏者能够选择最理想的一个来表达他们想要表达的情感。尽管有许多可能性,但认为只有一种适合音调运动特定情况的声音是科学的。在这项研究中,我们试图通过将和弦视为三维笛卡尔坐标系中的点来寻找最佳发声,并进一步加深对音乐理论中数学的基本理解。摘要:In classical music and in any genre of contemporary music, the tonal elements or notes used for playing are the same. The numerous possibilities of chords for a given instance in a piece make the playing, in general, very intricate, and advanced. The theory sounds quite trivial, yet the application has vast options, each leading to inarguably different outcomes, characterized by scientific and musical principles. Chords and their importance are self-explanatory. A chord is a bunch of notes played together. As far as scientists are concerned, it is a set of tonal frequencies ringing together resulting in a consonant/dissonant sound. It is well-known that the notes of a chord can be rearranged to come up with various voicings (1) of the same chord which enables a composer/player to choose the most optimal one to convey the emotion they wish to convey. Though there are numerous possibilities, it is scientific to think that there is just one appropriate voicing for a particular situation of tonal movements. In this study, we attempt to find the optimal voicings by considering chords to be points in a 3-dimensional cartesian coordinate system and further the fundamental understanding of mathematics in music theory.
【10】 Low-complexity deep learning frameworks for acoustic scene classification
标题:用于声学场景分类的低复杂度深度学习框架
链接:https://arxiv.org/abs/2206.06057
作者:Lam Pham,Dat Ngo,Anahid Jalali,Alexander Schindler摘要:在这份报告中,我们提出了低复杂度的声学场景分类深度学习框架(ASC)。提出的框架可分为四个主要步骤:前端谱图提取、在线数据增强、后端分类和预测概率的后期融合。特别是,我们最初将音频记录转换为Mel、Gammatone和CQT光谱图。接下来,使用随机裁剪、Specaugment和Mixup等数据增强方法生成增强的光谱图,然后再将其输入基于深度学习的分类器。最后,为了获得最佳性能,我们融合了从三个单独分类器中获得的概率,这些分类器使用三种类型的谱图进行独立训练。我们在DCASE 2022 Task 1开发数据集上进行的实验充分满足了低复杂性的要求,实现了60.1%的最佳分类精度,将DCASE基线提高了17.2%。摘要:In this report, we presents low-complexity deep learning frameworks for acoustic scene classification (ASC). The proposed frameworks can be separated into four main steps: Front-end spectrogram extraction, online data augmentation, back-end classification, and late fusion of predicted probabilities. In particular, we initially transform audio recordings into Mel, Gammatone, and CQT spectrograms. Next, data augmentation methods of Random Cropping, Specaugment, and Mixup are then applied to generate augmented spectrograms before being fed into deep learning based classifiers. Finally, to achieve the best performance, we fuse probabilities which obtained from three individual classifiers, which are independently-trained with three type of spectrograms. Our experiments conducted on DCASE 2022 Task 1 Development dataset have fullfiled the requirement of low-complexity and achieved the best classification accuracy of 60.1%, improving DCASE baseline by 17.2%.
【11】 Improvement of Serial Approach to Anomalous Sound Detection by Incorporating Two Binary Cross-Entropies for Outlier Exposure
标题:融合两个二进制交叉熵的异常声序列检测方法的改进
链接:https://arxiv.org/abs/2206.05929
作者:Ibuki Kuroyanagi,Tomoki Hayashi,Kazuya Takeda,Tomoki Toda备注:5 pages, 3 figures, 3 tables, EUSIPCO 2022摘要:异常声音检测系统必须仅使用正常音频数据检测未知、非典型声音。传统方法使用串行方法,将异常值暴露(OE)与inlier建模(IM)相结合,前者将正常和伪异常数据分类并获得嵌入,后者对嵌入的概率分布进行建模。虽然串行方法由于OE的强大特征提取和IM的鲁棒性而显示出很高的性能,但OE仍然存在一个问题,即当正常数据和伪异常数据太相似或太不同时,OE不能很好地工作。为了明确区分这些数据,该方法在训练OE时使用两个二进制交叉熵的多任务学习。第一种是一种损失,它对目标机器发出的声音进行分类,该产品用于处理正常数据和伪异常数据过于相似的情况。第二种是识别声音是否从目标机器发出的损失,它处理正常数据和伪异常数据相差太大的情况。我们使用DCASE 2021任务2数据集进行了实验。我们提出的单模型方法在AUC方面比组合多个模型的排名靠前的方法要好2.1%。摘要:Anomalous sound detection systems must detect unknown, atypical sounds using only normal audio data. Conventional methods use the serial method, a combination of outlier exposure (OE), which classifies normal and pseudo-anomalous data and obtains embedding, and inlier modeling (IM), which models the probability distribution of the embedding. Although the serial method shows high performance due to the powerful feature extraction of OE and the robustness of IM, OE still has a problem that doesn't work well when the normal and pseudo-anomalous data are too similar or too different. To explicitly distinguish these data, the proposed method uses multi-task learning of two binary cross-entropies when training OE. The first is a loss that classifies the sound of the target machine to which product it is emitted from, which deals with the case where the normal data and the pseudo-anomalous data are too similar. The second is a loss that identifies whether the sound is emitted from the target machine or not, which deals with the case where the normal data and the pseudo-anomalous data are too different. We perform our experiments with DCASE 2021 Task~2 dataset. Our proposed single-model method outperforms the top-ranked method, which combines multiple models, by 2.1% in AUC.
【12】 Description and Discussion on DCASE 2022 Challenge Task 2: Unsupervised Anomalous Sound Detection for Machine Condition Monitoring Applying Domain Generalization Techniques
标题:描述和讨论DCASE 2022挑战任务2:应用领域泛化技术进行机器状态监测的无监督异常声音检测
链接:https://arxiv.org/abs/2206.05876
作者:Kota Dohi,Keisuke Imoto,Noboru Harada,Daisuke Niizumi,Yuma Koizumi,Tomoya Nishida,Harsh Purohit,Takashi Endo,Masaaki Yamamoto,Yohei Kawaguchi备注:arXiv admin note: substantial text overlap with arXiv:2106.04492摘要:我们介绍了声学场景和事件检测与分类(DCASE)2022挑战任务2的任务描述:“应用领域泛化技术进行机器状态监测的无监督异常声音检测(ASD)”。领域转移是ASD系统应用中的一个关键问题。由于域移动会改变数据的声学特性,因此在源域中训练的模型对于目标域的性能较差。在DCASE 2021挑战任务2中,我们组织了一个ASD任务来处理域转移。在这项任务中,假设域移动的发生是已知的。然而,在实践中,可能不会给出每个样本的域,域移动可能会隐式发生。在2022年的任务2中,我们重点关注领域泛化技术,该技术可以检测异常,而不管领域发生了什么变化。具体来说,测试数据中没有给出每个样本的域,所有域只允许一个阈值。我们将在挑战提交截止日期后添加挑战结果和对提交内容的分析。摘要:We present the task description of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2022 Challenge Task 2: "Unsupervised anomalous sound detection (ASD) for machine condition monitoring applying domain generalization techniques". Domain shifts are a critical problem for the application of ASD systems. Because domain shifts can change the acoustic characteristics of data, a model trained in a source domain performs poorly for a target domain. In DCASE 2021 Challenge Task 2, we organized an ASD task for handling domain shifts. In this task, it was assumed that the occurrences of domain shifts are known. However, in practice, the domain of each sample may not be given, and the domain shifts can occur implicitly. In 2022 Task 2, we focus on domain generalization techniques that detects anomalies regardless of the domain shifts. Specifically, the domain of each sample is not given in the test data and only one threshold is allowed for all domains. We will add challenge results and analysis of the submissions after the challenge submission deadline.
【13】 The YiTrans End-to-End Speech Translation System for IWSLT 2022 Offline Shared Task
标题:面向IWSLT 2022离线共享任务的宜传端到端语音翻译系统
链接:https://arxiv.org/abs/2206.05777
作者:Ziqiang Zhang,Junyi Ao,Shujie Liu,Furu Wei,Jinyu Li摘要:本文描述了我们为IWSLT 2022离线任务提交的端到端YiTrans语音翻译系统,该系统可将英语音频翻译为德语、汉语和日语。YiTrans系统构建在大规模预训练编码器-解码器模型上。更具体地说,我们首先设计了一个多阶段的预训练策略,用大量的标记和未标记数据构建一个多模态模型。然后,我们为下游语音翻译任务微调模型的相应组件。此外,我们还努力提高性能,如数据过滤、数据增强、语音分割、模型集成等。实验结果表明,我们的YiTrans系统在三个翻译方向上都比强基线有了显著的改进,并且在tst2021英德版上比去年的最佳端到端系统实现了+5.2 BLEU改进。在自动评估指标方面,我们的最终提交资料在英德和英中端到端系统中排名第一。我们公开了我们的代码和模型。摘要:This paper describes the submission of our end-to-end YiTrans speech translation system for the IWSLT 2022 offline task, which translates from English audio to German, Chinese, and Japanese. The YiTrans system is built on large-scale pre-trained encoder-decoder models. More specifically, we first design a multi-stage pre-training strategy to build a multi-modality model with a large amount of labeled and unlabeled data. We then fine-tune the corresponding components of the model for the downstream speech translation tasks. Moreover, we make various efforts to improve performance, such as data filtering, data augmentation, speech segmentation, model ensemble, and so on. Experimental results show that our YiTrans system obtains a significant improvement than the strong baseline on three translation directions, and it achieves +5.2 BLEU improvements over last year's optimal end-to-end system on tst2021 English-German. Our final submissions rank first on English-German and English-Chinese end-to-end systems in terms of the automatic evaluation metric. We make our code and models publicly available.
【14】 Investigation of Ensemble features of Self-Supervised Pretrained Models for Automatic Speech Recognition
标题:自动语音识别中自监督预训练模型的集成特征研究
链接:https://arxiv.org/abs/2206.05518
作者:A Arunkumar,Vrunda N Sukhadia,S. Umesh备注:4 pages , 2 figures,submitted to interspeech 2022摘要:基于自监督学习(SSL)的模型可以生成强大的表示,可以用来提高下游语音任务的性能。有几种最先进的SSL模型可用,并且每种模型都优化了不同的损耗,这使得它们的功能可以互补。本文提出使用这种SSL表示和模型的集成,利用各种预训练模型提取的特征的互补性。我们假设这会导致更丰富的特征表示,并显示ASR下游任务的结果。为此,我们使用了三个在ASR任务上显示出优异结果的SSL模型,即HuBERT、Wav2vec2.0和WaveLM。我们探索了为ASR任务微调的模型集合,以及使用从下游ASR任务的预训练模型中获得的嵌入来实现的特征集合。我们使用Librispeech(100h)和WSJ数据集为下游任务改进了单个模型和预先训练的特征的性能。摘要:Self-supervised learning (SSL) based models have been shown to generate powerful representations that can be used to improve the performance of downstream speech tasks. Several state-of-the-art SSL models are available, and each of these models optimizes a different loss which gives rise to the possibility of their features being complementary. This paper proposes using an ensemble of such SSL representations and models, which exploits the complementary nature of the features extracted by the various pretrained models. We hypothesize that this results in a richer feature representation and shows results for the ASR downstream task. To this end, we use three SSL models that have shown excellent results on ASR tasks, namely HuBERT, Wav2vec2.0, and WaveLM. We explore the ensemble of models fine-tuned for the ASR task and the ensemble of features using the embeddings obtained from the pre-trained models for a downstream ASR task. We get improved performance over individual models and pre-trained features using Librispeech(100h) and WSJ dataset for the downstream tasks.
【15】 Hierarchical Conditional Variational Autoencoder Based Acoustic Anomaly Detection
标题:基于分层条件变分自动编码器的声学异常检测
链接:https://arxiv.org/abs/2206.05460
作者:Harsh Purohit,Takashi Endo,Masaaki Yamamoto,Yohei Kawaguchi摘要:本文旨在开发一种基于声信号的无监督机器自动监测异常检测方法。现有的方法,如深度自动编码器(DAE)、变分自动编码器(VAE)、条件变分自动编码器(CVAE)等,在潜在空间中的表示能力有限,因此异常检测性能较差。必须为每种不同类型的机器训练不同的模型,以准确执行异常检测任务。为了解决这个问题,我们提出了一种新的方法,称为分层条件变分自动编码器(HCVAE)。该方法利用现有的工业设施分类层次知识对潜在空间表示进行细化。这些知识也有助于改进模型的异常检测性能。通过使用适当的条件,我们证明了单个HCVAE模型对不同类型机器的泛化能力。此外,为了证明所提方法的实用性,(i)我们在不同领域评估了HCVAE模型,(ii)我们检查了部分层次知识的影响。我们的结果表明,HCVAE方法验证了这两个点,并且在AUC得分指标上,它在异常检测任务上的表现比基线系统高出15%。摘要:This paper aims to develop an acoustic signal-based unsupervised anomaly detection method for automatic machine monitoring. Existing approaches such as deep autoencoder (DAE), variational autoencoder (VAE), conditional variational autoencoder (CVAE) etc. have limited representation capabilities in the latent space and, hence, poor anomaly detection performance. Different models have to be trained for each different kind of machines to accurately perform the anomaly detection task. To solve this issue, we propose a new method named as hierarchical conditional variational autoencoder (HCVAE). This method utilizes available taxonomic hierarchical knowledge about industrial facility to refine the latent space representation. This knowledge helps model to improve the anomaly detection performance as well. We demonstrated the generalization capability of a single HCVAE model for different types of machines by using appropriate conditions. Additionally, to show the practicability of the proposed approach, (i) we evaluated HCVAE model on different domain and (ii) we checked the effect of partial hierarchical knowledge. Our results show that HCVAE method validates both of these points, and it outperforms the baseline system on anomaly detection task by utmost 15 % on the AUC score metric.
【16】 Multi-instrument Music Synthesis with Spectrogram Diffusion
标题:基于谱图扩散的多乐器音乐合成
链接:https://arxiv.org/abs/2206.05408
作者:Curtis Hawthorne,Ian Simon,Adam Roberts,Neil Zeghidour,Josh Gardner,Ethan Manilow,Jesse Engel摘要:理想的音乐合成器应具有交互性和表现力,能够为乐器和音符的任意组合实时生成高保真音频。最近的神经合成器在特定领域的模型(只提供特定乐器的详细控制)和原始波形模型(可以对所有音乐进行训练,但控制最少且生成速度较慢)之间进行了权衡。在这项工作中,我们将重点放在神经合成器的中间地带,该合成器可以通过MIDI序列和任意组合的乐器实时生成音频。这使得可以使用单一模型在广泛的转录数据集上进行训练,进而提供对广泛仪器的组成和仪器的注释级控制。我们使用一个简单的两阶段过程:MIDI到频谱图,使用编码器-解码器转换器,然后频谱图到音频,使用生成对抗网络(GAN)频谱图转换器。我们比较了将解码器作为自回归模型和去噪扩散概率模型(DDPM)进行训练,发现DDPM方法在定性和通过音频重建和Fr距离度量衡量方面都优于传统方法。考虑到这种方法的交互性和通用性,我们发现这是实现交互式和表达性神经合成的第一步,可用于任意组合的乐器和音符。摘要:An ideal music synthesizer should be both interactive and expressive, generating high-fidelity audio in realtime for arbitrary combinations of instruments and notes. Recent neural synthesizers have exhibited a tradeoff between domain-specific models that offer detailed control of only specific instruments, or raw waveform models that can train on all of music but with minimal control and slow generation. In this work, we focus on a middle ground of neural synthesizers that can generate audio from MIDI sequences with arbitrary combinations of instruments in realtime. This enables training on a wide range of transcription datasets with a single model, which in turn offers note-level control of composition and instrumentation across a wide range of instruments. We use a simple two-stage process: MIDI to spectrograms with an encoder-decoder Transformer, then spectrograms to audio with a generative adversarial network (GAN) spectrogram inverter. We compare training the decoder as an autoregressive model and as a Denoising Diffusion Probabilistic Model (DDPM) and find that the DDPM approach is superior both qualitatively and as measured by audio reconstruction and Fr\'echet distance metrics. Given the interactivity and generality of this approach, we find this to be a promising first step towards interactive and expressive neural synthesis for arbitrary combinations of instruments and notes.
【17】 AHD ConvNet for Speech Emotion Classification
标题:用于语音情感分类的AND转换网
链接:https://arxiv.org/abs/2206.05286
作者:Asfand Ali,Danial Nasir,Mohammad Hassan Jawad摘要:人工智能领域的成就被用于推动计算和智能机器的制造,以方便人类和改善用户体验。情绪对人们来说是最基本的,它影响思维和日常练习,如通信、学习和指导。语音情感识别是这方面的一个研究领域,在这项工作中,我们提出了一种新的mel谱图学习方法,其中我们的模型使用数据点从流行的CREMA-D数据集中给定的wav形式语音注释中学习情感。我们的模型使用对数mel谱图作为特征,mel数=64。与用于解决情感语音识别问题的其他方法相比,它所需的训练时间更少。摘要:Accomplishments in the field of artificial intelligence are utilized in the advancement of computing and making of intelligent machines for facilitating mankind and improving user experience. Emotions are rudimentary for people, affecting thinking and ordinary exercises like correspondence, learning and direction. Speech emotion recognition is domain of interest in this regard and in this work, we propose a novel mel spectrogram learning approach in which our model uses the datapoints to learn emotions from the given wav form voice notes in the popular CREMA-D dataset. Our model uses log mel-spectrogram as feature with number of mels = 64. It took less training time compared to other approaches used to address the problem of emotion speech recognition.
机器翻译,仅供参考