今日论文合集:cs.SD语音20篇,eess.AS音频处理21篇。

本文经arXiv每日学术速递授权转载

cs.SD语音

【1】DITTO: Diffusion Inference-Time T-Optimization for Music Generation
标题:同上:音乐生成的扩散推理时间T-优化
链接:https://arxiv.org/abs/2401.12179
作者:Zachary Novack,Julian McAuley,Taylor Berg-Kirkpatrick,Nicholas J. Bryan
摘要:我们提出了扩散推理时间T优化(DITTO),这是一种通用框架,用于通过优化初始噪声潜伏期来控制预训练的文本到音乐扩散模型。我们的方法可用于通过任何可微特征匹配损失进行优化,以实现目标(风格化)输出,并利用梯度检查点提高内存效率。我们展示了一个令人惊讶的广泛的应用程序的音乐生成,包括修复,outpainting,循环以及强度,旋律和音乐结构控制-所有没有微调的基础模型。当我们将我们的方法与相关的训练,指导和基于优化的方法进行比较时,我们发现DITTO在几乎所有任务上都达到了最先进的性能,包括在可控性,音频质量和计算效率方面优于可比方法,从而为高质量,灵活,无训练的扩散模型控制打开了大门。可以在https://DITTO-Music.github.io/web/上找到合理的例子。
摘要:We propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose frame-work for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents. Our method can be used to optimize through any differentiable feature matching loss to achieve a target (stylized) output and leverages gradient checkpointing for memory efficiency. We demonstrate a surprisingly wide-range of applications for music generation including inpainting, outpainting, and looping as well as intensity, melody, and musical structure control - all without ever fine-tuning the underlying model. When we compare our approach against related training, guidance, and optimization-based methods, we find DITTO achieves state-of-the-art performance on nearly all tasks, including outperforming comparable approaches on controllability, audio quality, and computational efficiency, thus opening the door for high-quality, flexible, training-free control of diffusion models. Sound examples can be found at https://DITTO-Music.github.io/web/.


【2】 Resource-constrained stereo singing voice cancellation
标题:资源受限的立体声演唱语音消除
链接:https://arxiv.org/abs/2401.12068
作者:Clara Borrelli,James Rae,Dogac Basaran,Matt McVicar,Mehrez Souden,Matthias Mauch
摘要:我们研究的问题,立体声歌唱的声音消除,音乐源分离的子任务,其目标是从立体声混合估计乐器背景。我们将探索如何从一个小型、高效的实时语音分离模型开始,实现与大型最先进的源分离网络相似的性能。这种模型在内存和计算有限并且歌声处理必须以有限的前瞻运行时是有用的。在实践中,这是通过调整现有的单声道模型来处理立体声输入来实现的。通过调整模型参数和扩展训练集来提高质量。此外,我们强调立体声模型带来的好处,通过引入一个新的度量,检测通道之间的衰减不一致。我们的方法使用客观的离线指标和大规模的MUSHRA试验进行评估,证实了我们的技术在严格的听力测试中的有效性。
摘要:We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art source separation networks starting from a small, efficient model for real-time speech separation. Such a model is useful when memory and compute are limited and singing voice processing has to run with limited look-ahead. In practice, this is realised by adapting an existing mono model to handle stereo input. Improvements in quality are obtained by tuning model parameters and expanding the training set. Moreover, we highlight the benefits a stereo model brings by introducing a new metric which detects attenuation inconsistencies between channels. Our approach is evaluated using objective offline metrics and a large-scale MUSHRA trial, confirming the effectiveness of our techniques in stringent listening tests.


【3】 NEUROSEC: FPGA-Based Neuromorphic Audio Security
标题:NeurOSEC:基于FPGA的神经形态音频安全
链接:https://arxiv.org/abs/2401.12055
作者:Murat Isik,Hiruna Vishwamith,Yusuf Sur,Kayode Inadagbo,I. Can Dikmen
备注:Audio processing, FPGA, Hardware Security, Neuromorphic Computing
摘要:受人脑复杂性和功能性的启发,神经形态系统由于其在广泛应用中无与伦比的潜力而引起了学术和工业界的关注。虽然它们的能力预示着创新,但必须强调的是,这些计算范式与传统的计算范式类似,并非不受安全威胁的影响。虽然用于图像和视频处理的神经形态方法的探索已经被严格地追求,但神经形态音频处理的领域仍然处于早期阶段。我们的研究结果突出了我们的基于FPGA的神经形态系统的鲁棒性和精度。具体来说,我们的系统在期望信号和背景噪声之间实现了值得称赞的平衡,有效的尖峰速率编码,以及对FGSM和PGD等对抗性攻击的无与伦比的弹性。我们的框架的一个突出特点是其检测率为94%,与其他方法相比,强调了其在5.39 dB范围内识别和缓解威胁的能力,这是一个值得称赞的SNR比。此外,神经形态计算和硬件安全服务于关键任务和隐私保护应用中的许多传感器领域。
摘要:Neuromorphic systems, inspired by the complexity and functionality of the human brain, have gained interest in academic and industrial attention due to their unparalleled potential across a wide range of applications. While their capabilities herald innovation, it is imperative to underscore that these computational paradigms, analogous to their traditional counterparts, are not impervious to security threats. Although the exploration of neuromorphic methodologies for image and video processing has been rigorously pursued, the realm of neuromorphic audio processing remains in its early stages. Our results highlight the robustness and precision of our FPGA-based neuromorphic system. Specifically, our system showcases a commendable balance between desired signal and background noise, efficient spike rate encoding, and unparalleled resilience against adversarial attacks such as FGSM and PGD. A standout feature of our framework is its detection rate of 94%, which, when compared to other methodologies, underscores its greater capability in identifying and mitigating threats within 5.39 dB, a commendable SNR ratio. Furthermore, neuromorphic computing and hardware security serve many sensor domains in mission-critical and privacy-preserving applications.


【4】 Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
标题:看、听、认:角色感知的视听字幕
链接:https://arxiv.org/abs/2401.12039
作者:Bruno Korbar,Jaesung Huh,Andrew Zisserman
备注:Accepted for publication in ICASSP 2024
摘要:本文的目标是自动字符感知字幕生成。给定一个视频和最少量的元数据,我们提出了一个视听方法,生成一个完整的对话记录,精确的语音时间戳,和字符说话识别。其核心思想是首先使用视听线索为每个字符选择一组高精度的音频样本,然后使用这些样本对所有语音段进行说话人身份分类。值得注意的是,该方法不需要面部检测或跟踪。我们评估了各种电视情景喜剧,包括宋飞,Fraiser和灌木丛的方法。我们设想这个系统是有用的自动生成字幕,以提高现代流媒体服务上可用的大量视频的可访问性。项目页面:\url{https://www.robots.ox.ac.uk/juvengg/research/look-juven-juvenise/}
摘要:The goal of this paper is automatic character-aware subtitle generation. Given a video and a minimal amount of metadata, we propose an audio-visual method that generates a full transcript of the dialogue, with precise speech timestamps, and the character speaking identified. The key idea is to first use audio-visual cues to select a set of high-precision audio exemplars for each character, and then use these exemplars to classify all speech segments by speaker identity. Notably, the method does not require face detection or tracking. We evaluate the method over a variety of TV sitcoms, including Seinfeld, Fraiser and Scrubs. We envision this system being useful for the automatic generation of subtitles to improve the accessibility of the vast amount of videos available on modern streaming services. Project page : \url{https://www.robots.ox.ac.uk/~vgg/research/look-listen-recognise/}


【5】 Lightweight Protection for Privacy in Offloaded Speech Understanding
标题:卸载语音理解中的轻量级隐私保护
链接:https://arxiv.org/abs/2401.11983
作者:Dongqi Cai,Shangguang Wang,Zeling Zhang,Felix Xiaozhu Lin,Mengwei Xu
备注:under review摘要:语音是移动嵌入式设备的常见输入方法,但基于云的语音识别系统会带来隐私风险。基于解纠缠的编码器被设计为通过从语音信号中过滤敏感信息来保护用户隐私,不幸的是需要大量的存储器和计算资源,这限制了它们在功能较弱的设备中的使用。为了克服这一点,我们介绍了一种新的系统,XXX,优化了这样的设备。XXX是建立在这样一种见解之上的:语音理解主要依赖于理解整个话语的长期依赖性,而隐私问题通常与短期细节有关。因此,XXX专注于选择性地掩蔽这些短期元素,保持长期语音理解的质量。XXX的核心是一个创新的差分掩码生成器,以可解释学习为基础,对掩码过程进行微调。我们在STM32H7微控制器上测试了XXX,评估了其在各种潜在攻击场景中的性能。结果表明,XXX保持了与现有编码器相当的语音理解准确性和隐私性,但在效率上有了显着提高,处理速度提高了53.3倍,内存占用减少了134.1倍。
摘要:Speech is a common input method for mobile embedded devices, but cloud-based speech recognition systems pose privacy risks. Disentanglement-based encoders, designed to safeguard user privacy by filtering sensitive information from speech signals, unfortunately require substantial memory and computational resources, which limits their use in less powerful devices. To overcome this, we introduce a novel system, XXX, optimized for such devices. XXX is built on the insight that speech understanding primarily relies on understanding the entire utterance's long-term dependencies, while privacy concerns are often linked to short-term details. Therefore, XXX focuses on selectively masking these short-term elements, preserving the quality of long-term speech understanding. The core of XXX is an innovative differential mask generator, grounded in interpretable learning, which fine-tunes the masking process. We tested XXX on the STM32H7 microcontroller, assessing its performance in various potential attack scenarios. The results show that XXX maintains speech understanding accuracy and privacy at levels comparable to existing encoders, but with a significant improvement in efficiency, achieving up to 53.3$\times$ faster processing and a 134.1$\times$ smaller memory footprint.


【6】 Keep Decoding Parallel with Effective Knowledge Distillation from  Language Models to End-to-end Speech Recognisers
标题:保持解码与从语言模型到端到端语音识别器的有效知识提取的并行性
链接:https://arxiv.org/abs/2401.11700
作者:Michael Hentschel,Yuta Nishikawa,Tatsuya Komatsu,Yusuke Fujita
备注:Accepted at ICASSP 2024
摘要:本研究提出一种新的方法,知识蒸馏(KD)从一个BERT教师模型的自动语音识别(ASR)模型使用中间层。为了验证教师的知识,我们使用了一个注意力解码器,它从BERT的令牌概率中学习。我们的方法表明,语言模型(LM)的信息可以更有效地提取到一个ASR模型使用的中间层和最终层。通过将中间层作为蒸馏目标,我们可以更有效地将知识蒸馏到网络的较低层。使用我们的方法,我们实现了更好的识别精度比浅融合的外部LM,使我们能够保持快速并行解码。LibriSpeech数据集上的实验证明了我们的方法在增强贪婪解码与连接主义时间分类(CTC)的有效性。
摘要:This study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil the teacher's knowledge, we use an attention decoder that learns from BERT's token probabilities. Our method shows that language model (LM) information can be more effectively distilled into an ASR model using both the intermediate layers and the final layer. By using the intermediate layers as distillation target, we can more effectively distil LM knowledge into the lower network layers. Using our method, we achieve better recognition accuracy than with shallow fusion of an external LM, allowing us to maintain fast parallel decoding. Experiments on the LibriSpeech dataset demonstrate the effectiveness of our approach in enhancing greedy decoding with connectionist temporal classification (CTC).


【7】 Word-Level ASR Quality Estimation for Efficient Corpus Sampling and  Post-Editing through Analyzing Attentions of a Reference-Free Metric
标题:基于无参考度量的词级ASR质量估计
链接:https://arxiv.org/abs/2401.11268
作者:Golara Javadi,Kamer Ali Yuksel,Yunsu Kim,Thiago Castro Ferreira,Mohamed Al-Badrashiny
摘要:在自动语音识别(ASR)领域,对模型的追求不仅具有高准确性,而且在决策过程中提供透明度是至关重要的。质量估计(QE)指标的潜力,介绍和评估作为一种新的工具,以提高可解释的人工智能(XAI)在ASR系统。通过实验和分析,NoRefER(无参考错误率)指标的能力进行了探索,在识别字级错误,以帮助后编辑在细化ASR假设。调查还扩展到在语料库建设过程中的NoRefER的效用,证明了其有效性,在增强数据集有见地的注释。的诊断方面的NoRefER检查,揭示其提供有价值的见解模型的行为和决策模式的能力。事实证明,这有利于在编辑后工作流程中优先考虑假设和微调ASR模型。研究结果表明,NoRefER不仅是一个错误检测工具,而且是一个提高ASR系统透明度、效率和有效性的综合框架。为了确保结果的重现性,本研究的所有源代码都是公开的。
摘要:In the realm of automatic speech recognition (ASR), the quest for models that not only perform with high accuracy but also offer transparency in their decision-making processes is crucial. The potential of quality estimation (QE) metrics is introduced and evaluated as a novel tool to enhance explainable artificial intelligence (XAI) in ASR systems. Through experiments and analyses, the capabilities of the NoRefER (No Reference Error Rate) metric are explored in identifying word-level errors to aid post-editors in refining ASR hypotheses. The investigation also extends to the utility of NoRefER in the corpus-building process, demonstrating its effectiveness in augmenting datasets with insightful annotations. The diagnostic aspects of NoRefER are examined, revealing its ability to provide valuable insights into model behaviors and decision patterns. This has proven beneficial for prioritizing hypotheses in post-editing workflows and fine-tuning ASR models. The findings suggest that NoRefER is not merely a tool for error detection but also a comprehensive framework for enhancing ASR systems' transparency, efficiency, and effectiveness. To ensure the reproducibility of the results, all source codes of this study are made publicly available.


【8】 Projected Belief Networks With Discriminative Alignment for Acoustic  Event Classification: Rivaling State of the Art CNNs
标题:用于声事件分类的判别对齐投影信念网络:与现有CNN相媲美
链接:https://arxiv.org/abs/2401.11199
作者:Paul M. Baggenstoss,Kevin Wilkinghoff,Felix Govaers,Frank Kurth
备注:15 Pages. Submitted to IEEE-TNNLS
摘要:投影信念网络(PBN)是一种基于前馈神经网络(FFNN)的具有易处理似然函数的生成式随机网络。生成函数通过FFNN进行“备份”操作。PBN是两个网络合二为一,一个是在前向方向上运行的FFNN,另一个是在后向方向上运行的生成网络。两个网络基于相同的参数集共存,具有自己的成本函数,并且可以单独或联合训练。因此,PBN有可能拥有最好的品质的歧视和生成分类器。为了实现这种潜力,在每个类上训练一个单独的PBN,最大化给定类的生成似然函数,同时最小化FFNN对“所有其他类”的区分成本。这种技术被称为判别对齐(PBN-DA),将似然函数的轮廓与决策边界对齐,并大大提高了分类性能,可与最先进的判别网络相媲美。该方法可以使用隐马尔可夫模型(HMM)作为PBN的组件(称为PBN-DA-HMM)来进一步改进。本文提供了PBN、PBN-DA和PBN-DA-HMM的综合处理。此外,两个新的分类实验的结果提供。第一个实验使用空气声学事件,第二个实验使用由海洋哺乳动物叫声组成的水声数据。在这两个实验中,PBN-DA-HMM获得了与现有技术CNN相当或更好的性能,并且在与CNN组合时获得了两倍的误差减少。
摘要:The projected belief network (PBN) is a generative stochastic network with tractable likelihood function based on a feed-forward neural network (FFNN). The generative function operates by "backing up" through the FFNN. The PBN is two networks in one, a FFNN that operates in the forward direction, and a generative network that operates in the backward direction. Both networks co-exist based on the same parameter set, have their own cost functions, and can be separately or jointly trained. The PBN therefore has the potential to possess the best qualities of both discriminative and generative classifiers. To realize this potential, a separate PBN is trained on each class, maximizing the generative likelihood function for the given class, while minimizing the discriminative cost for the FFNN against "all other classes". This technique, called discriminative alignment (PBN-DA), aligns the contours of the likelihood function to the decision boundaries and attains vastly improved classification performance, rivaling that of state of the art discriminative networks. The method may be further improved using a hidden Markov model (HMM) as a component of the PBN, called PBN-DA-HMM. This paper provides a comprehensive treatment of PBN, PBN-DA, and PBN-DA-HMM. In addition, the results of two new classification experiments are provided. The first experiment uses air-acoustic events, and the second uses underwater acoustic data consisting of marine mammal calls. In both experiments, PBN-DA-HMM attains comparable or better performance as a state of the art CNN, and attain a factor of two error reduction when combined with the CNN.

【9】 Generalizing Speaker Verification for Spoof Awareness in the Embedding  Space
标题:嵌入空间中基于欺骗感知的说话人确认泛化
链接:https://arxiv.org/abs/2401.11156
作者:Xuechen Liu,Md Sahidullah,Kong Aik Lee,Tomi Kinnunen
备注:To appear in IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:现在众所周知,自动说话人验证(ASV)系统可以使用各种类型的对手进行欺骗。对抗ASV系统的通常方法是开发一个单独的欺骗对策(CM)模块,将语音输入分类为真实的或欺骗的话语。然而,这样的设计在认证阶段需要额外的计算和利用工作。另一种策略涉及一个单一的单片ASV系统,旨在处理零努力冒名顶替者(非目标)和欺骗攻击。这种欺骗感知ASV系统有可能提供更强的保护和更经济的计算。为此,我们建议推广独立的ASV(G-SASV)来对抗欺骗攻击,其中我们利用来自CM的有限训练数据来增强嵌入空间中的简单后端,而无需在测试(认证)阶段涉及单独的CM模块。我们提出了一种基于深度神经网络的新颖而简单的后端分类器,并在训练阶段通过域自适应和欺骗嵌入的多任务集成进行研究。在ASVspoof 2019逻辑访问数据集上进行了实验,我们在联合(真实和欺骗)和欺骗条件下分别将统计ASV后端的性能提高了36.2%和49.8%。
摘要:It is now well-known that automatic speaker verification (ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofing countermeasure (CM) module to classify speech input either as a bonafide, or a spoofed utterance. Nevertheless, such a design requires additional computation and utilization efforts at the authentication stage. An alternative strategy involves a single monolithic ASV system designed to handle both zero-effort imposter (non-targets) and spoofing attacks. Such spoof-aware ASV systems have the potential to provide stronger protections and more economic computations. To this end, we propose to generalize the standalone ASV (G-SASV) against spoofing attacks, where we leverage limited training data from CM to enhance a simple backend in the embedding space, without the involvement of a separate CM module during the test (authentication) phase. We propose a novel yet simple backend classifier based on deep neural networks and conduct the study via domain adaptation and multi-task integration of spoof embeddings at the training stage. Experiments are conducted on the ASVspoof 2019 logical access dataset, where we improve the performance of statistical ASV backends on the joint (bonafide and spoofed) and spoofed conditions by a maximum of 36.2% and 49.8% in terms of equal error rates, respectively.

【10】 Gaussian Adaptive Attention is All You Need: Robust Contextual  Representations Across Multiple Modalities
标题:高斯自适应注意力是您所需要的全部:跨多个通道的强健上下文表示
链接:https://arxiv.org/abs/2401.11143
作者:Georgios Ioannides,Aman Chadha,Aaron Elkins
摘要:我们提出了多头高斯自适应注意力机制(GAAM),一种新的概率注意力框架,和高斯自适应Transformer(GAT),旨在增强跨多个模态的信息聚合,包括语音,文本和视觉。GAAM将可学习的均值和方差集成到其注意力机制中,在多头框架中实现,使其能够集体建模任何概率分布,以动态重新校准特征重要性。这种方法表现出显着的改进,特别是对于高度非平稳的数据,通过识别特征空间内的关键元素,在模型性能方面超过了最先进的注意力技术(准确度高达约+20%)。GAAM与基于点积的注意力模型的兼容性和相对较低的参数数量显示了其适应性和潜力,以提高现有的注意力框架。从经验上讲,GAAM在各种任务中表现出卓越的适应性和有效性,包括语音中的情感识别,图像分类和文本分类,从而建立了处理多模态数据的鲁棒性和多功能性。此外,我们还引入了重要性因子(IF),这是一种新的基于学习的度量,可以增强使用基于GAAM的方法训练的模型的可解释性。总的来说,GAAM代表了跨多种模式开发更好的性能和更可解释的注意力模型的进步。
摘要:We propose the Multi-Head Gaussian Adaptive Attention Mechanism (GAAM), a novel probabilistic attention framework, and the Gaussian Adaptive Transformer (GAT), designed to enhance information aggregation across multiple modalities, including Speech, Text and Vision. GAAM integrates learnable mean and variance into its attention mechanism, implemented in a Multi-Headed framework enabling it to collectively model any Probability Distribution for dynamic recalibration of feature significance. This method demonstrates significant improvements, especially with highly non-stationary data, surpassing the state-of-the-art attention techniques in model performance (up to approximately +20% in accuracy) by identifying key elements within the feature space. GAAM's compatibility with dot-product-based attention models and relatively low number of parameters showcases its adaptability and potential to boost existing attention frameworks. Empirically, GAAM exhibits superior adaptability and efficacy across a diverse range of tasks, including emotion recognition in speech, image classification, and text classification, thereby establishing its robustness and versatility in handling multi-modal data. Furthermore, we introduce the Importance Factor (IF), a new learning-based metric that enhances the explainability of models trained with GAAM-based methods. Overall, GAAM represents an advancement towards development of better performing and more explainable attention models across multiple modalities.


【11】 ASM: Audio Spectrogram Mixer
标题:ASM:音频频谱混合器
链接:https://arxiv.org/abs/2401.11102
作者:Qingfeng Ji,Jicun Zhang,Yuxin Wang
摘要:Transformer结构最近在深度学习领域表现出了出色的技能,显著提高了各种领域模型的准确性。研究人员已经开始质疑这种复杂的网络结构是否真的有必要,以及由于其复杂的网络拓扑结构和高推理成本,是否可以在降低推理成本的情况下获得同样出色的结果。为了证明Mixer在三个数据集Speech Commands、UrbanSound 8 k和CASIA中文情感语料库上的有效性,本文将Mixer的精简版本应用于音频分类任务,并与基于Transformer-based Audio Spectrogram Transformer(AST)模型进行了对比实验。此外,本文还对GeLU、Mish、Swish和Mixer-C等几种激活函数在Mixer中的应用进行了对比实验。此外,本文还通过对比实验,比较了Mixer中GeLU、Mish、Swish和Mix-C等激活函数的使用情况。此外,还指出了AST模型的一些缺陷,并对本文提出的模型进行了改进。总之,在这项研究中提出了一个称为音频频谱图混合器的模型,这是第一个使用混合器进行音频分类的模型,并研究了该模型未来的改进方向。
摘要:Transformer structures have demonstrated outstanding skills in the deep learning space recently, significantly increasing the accuracy of models across a variety of domains. Researchers have started to question whether such a sophisticated network structure is actually necessary and whether equally outstanding results can be reached with reduced inference cost due to its complicated network topology and high inference cost. In order to prove the Mixer's efficacy on three datasets Speech Commands, UrbanSound8k, and CASIA Chinese Sentiment Corpus this paper applies amore condensed version of the Mixer to an audio classification task and conducts comparative experiments with the Transformer-based Audio Spectrogram Transformer (AST)model. In addition, this paper conducts comparative experiments on the application of several activation functions in Mixer, namely GeLU, Mish, Swish and Acon-C. Further-more, the use of various activation functions in Mixer, including GeLU, Mish, Swish, and Acon-C, is compared in this research through comparison experiments. Additionally, some AST model flaws are highlighted, and the model suggested in this study is improved as a result. In conclusion, a model called the Audio Spectrogram Mixer, which is the first model for audio classification with Mixer, is suggested in this study and the model's future directions for improvement are examined.

【12】 Sound Unblending: Exploring Sound Manipulations for Accessible  Mixed-Reality Awareness
标题:声音分离:探索声音操作以获得可接近的混合现实意识
链接:https://arxiv.org/abs/2401.11095
作者:Ruei-Che Chang,Chia-Sheng Hung,Bing-Yu Chen,Dhruv Jain,Anhong Guo
摘要:混合现实(MR)音景将真实世界的声音与来自听力设备的虚拟音频融合在一起,呈现出难以辨别和区分的复杂听觉信息。这对于盲人或视障人士来说尤其具有挑战性,他们在日常生活中依赖声音和描述。为了了解复杂的音频信息是如何被消费的,我们分析了盲人社区中的在线论坛帖子,确定了普遍的挑战,需求和期望的解决方案。我们综合了这些结果,并提出了用于提高MR声音意识的声音分解,其中包括六种声音操作:Ambience Builder,Feature Shifter,Earcon Generator,Prioritizer,Spatializer和Stylizer。为了评估声音分解的有效性,我们在三个模拟MR场景中与18名盲人参与者进行了一项用户研究,参与者在复杂的音景中识别出特定的声音。我们发现,声音不混合增加MR声音意识和最小化认知负荷。最后,我们开发了三个真实世界的示例应用程序来演示声音分解的实用性。
摘要:Mixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually impaired individuals, who rely on sounds and descriptions in their everyday lives. To understand how complex audio information is consumed, we analyzed online forum posts within the blind community, identifying prevailing challenges, needs, and desired solutions. We synthesized the results and proposed Sound Unblending for increasing MR sound awareness, which includes six sound manipulations: Ambience Builder, Feature Shifter, Earcon Generator, Prioritizer, Spatializer, and Stylizer. To evaluate the effectiveness of sound unblending, we conducted a user study with 18 blind participants across three simulated MR scenarios, where participants identified specific sounds within intricate soundscapes. We found that sound unblending increased MR sound awareness and minimized cognitive load. Finally, we developed three real-world example applications to demonstrate the practicality of sound unblending.


【13】 Consistency Based Unsupervised Self-training For ASR Personalisation
标题:基于一致性的ASR个性化无监督自我训练
链接:https://arxiv.org/abs/2401.12085
作者:Jisi Zhang,Vandana Rajan,Haaris Mehmood,David Tuckey,Pablo Peso Parada,Md Asif Jalal,Karthikeyan Saravanan,Gil Ho Lee,Jungin Lee,Seokyeong Jung
备注:Accepted for IEEE ASRU 2023
摘要:在大量人群的语音数据上训练的设备上自动语音识别(ASR)模型可能会在训练期间看不见的个人中表现不佳。这是由于用户数据和原始训练数据之间的域移位,不同的用户的说话特性和环境声学条件。ASR个性化是一种旨在利用用户数据来提高模型鲁棒性的解决方案。大多数ASR个性化方法都假设标记的用户数据用于监督。由于有限的数据大小和录制的音频样本质量差,没有任何标记数据的个性化是具有挑战性的。这项工作通过伪标签开发一种新的基于一致性的训练方法来解决无监督的个性化。与预训练模型相比,我们的方法在未标记的训练数据上实现了17.3%的相对字错误率降低(WERR),在保留数据上实现了8.1%的相对字错误率降低(WERR),并且优于当前最先进的方法。
摘要:On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain shift between user data and the original training data, differed by user's speaking characteristics and environmental acoustic conditions. ASR personalisation is a solution that aims to exploit user data to improve model robustness. The majority of ASR personalisation methods assume labelled user data for supervision. Personalisation without any labelled data is challenging due to limited data size and poor quality of recorded audio samples. This work addresses unsupervised personalisation by developing a novel consistency based training method via pseudo-labelling. Our method achieves a relative Word Error Rate Reduction (WERR) of 17.3% on unlabelled training data and 8.1% on held-out data compared to a pre-trained model, and outperforms the current state-of-the art methods.


【14】 Adversarial speech for voice privacy protection from Personalized Speech  generation
标题:针对个性化语音生成的语音隐私保护的对抗性语音
链接:https://arxiv.org/abs/2401.11857
作者:Shihao Chen,Liping Chen,Jie Zhang,KongAik Lee,Zhenhua Ling,Lirong Dai
备注:Accepted by icassp 2024
摘要:个性化语音生成技术的快速发展,包括个性化文本到语音(TTS)和语音转换(VC),对人类收听者区分生成的语音和真实语音提出了挑战,从而迫切需要保护说话者的语音免受恶意滥用。在这方面,我们提出了一种基于对抗性攻击的说话人保护方法。所提出的方法通过最小限度地改变原始语音来扰动语音信号,同时使下游语音生成模型无法准确地生成目标说话者的语音。为了验证,我们采用开源的预训练的YourTTS模型进行语音生成,并在白盒场景中保护目标说话者的语音。自动说话人确认(ASV)的评价进行了对所产生的语音的语音保护能力的评估。我们的实验结果表明,我们成功地扰动扬声器编码器的YourTTS模型使用基于梯度的I-FGSM对抗扰动方法。此外,对抗性扰动在防止YourTTS模型生成目标说话者的语音方面是有效的。音频样本可以在https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS上找到。
摘要:The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicious misuse. In this regard, we propose a speaker protection method based on adversarial attacks. The proposed method perturbs speech signals by minimally altering the original speech while rendering downstream speech generation models unable to accurately generate the voice of the target speaker. For validation, we employ the open-source pre-trained YourTTS model for speech generation and protect the target speaker's speech in the white-box scenario. Automatic speaker verification (ASV) evaluations were carried out on the generated speech as the assessment of the voice protection capability. Our experimental results show that we successfully perturbed the speaker encoder of the YourTTS model using the gradient-based I-FGSM adversarial perturbation method. Furthermore, the adversarial perturbation is effective in preventing the YourTTS model from generating the speech of the target speaker. Audio samples can be found in https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS.


【15】 Intelligibility Enhancement of Acoustic Noisy Speech for Autism Spectrum  Disorder Condition
标题:提高自闭症谱系障碍声学含噪语音的清晰度
链接:https://arxiv.org/abs/2401.11832
作者:M. Pillonetto,A. Queiroz,R. Coelho
备注:5 pages, 3 figues, 2 tables
摘要:本文介绍了一种时域个性化方法(pGTFF0),用于改善自闭症谱系障碍(ASD)患者的噪声语音清晰度。对于这个建议,从语音帧估计的谐波特征被认为是Gammatone听觉滤波器组的中心频率。增益因子被进一步应用于滤波样本的输出。关键目标是模拟为ASD患者量身定制的外部噪声过滤。感知听力测试表明,ASD志愿者达到较低的可懂率比神经(NT)。所提出的解决方案相比,三个竞争的方法,考虑在不同的信号噪声比的四个声学噪声。评估还采用了两个客观指标(ESTOI和PESQ)。实验结果表明,个性化的解决方案优于竞争的方法,在可懂度和质量的改善。
摘要:This work introduces a time domain personalized method (pGTFF0) to achieve intelligibility improvement of noisy speech for Autism Spectrum Disorder (ASD) situation. For this proposal, harmonic features estimated from speech frames are considered as center frequencies of Gammatone auditory filterbanks. A gain factor is further applied to the output of the filtered samples. The key goal is the emulation of an external noise filtering tailored for individuals with ASD. A perceptual listening test demonstrates that ASD volunteers attained lower intelligibility rates than Neurotypical (NT). The proposed solution is compared to three competing approaches considering four acoustic noises at different signal-to-noise ratios. Two objective measures (ESTOI and PESQ) are also adopted for evaluation. The experimental results show that the personalized solution outperformed the competing approaches in terms of intelligibility and quality improvement.


【16】 Harmonic Detection from Noisy Speech with Auditory Frame Gain for  Intelligibility Enhancement标题:利用听觉帧增益提高清晰度的噪声语音中的谐波检测
链接:https://arxiv.org/abs/2401.11829
作者:A. Queiroz,R. Coelho
备注:9 pages, 6 figures, 4 tables
摘要:本文介绍了一种新的方法(HDAG -谐波检测听觉增益)的语音清晰度增强在嘈杂的情况下。在该方案中,采用一系列的选择性Gammachirp滤波器来强调语音的谐波成分,减少声学噪声的掩蔽效应。利用HHT-DFT技术对基频进行了估计。根据FSFFE低/高音调分离检测和调整以低精度估计的谐波图案。滤波器组的中心频率被定义为考虑最适合覆盖与可懂度最相关的区域的第三倍频程子带。在信号重构之前,通过由FSFFE分类调节的增益因子来放大伽马线性滤波分量。建议HDAG解决方案和三个基线技术进行检查,考虑六个背景噪声与四个信噪比。采用三种客观的方法来评价语音清晰度和语音质量。几个实验进行了证明,所提出的方案实现了更好的语音清晰度的改善相比,竞争的方法。一个感性的听力测试进一步考虑和佐证的客观结果。
摘要:This paper introduces a novel (HDAG - Harmonic Detection for Auditory Gain) method for speech intelligibility enhancement in noisy scenarios. In the proposed scheme, a series of selective Gammachirp filters are adopted to emphasize the harmonic components of speech reducing the masking effects of acoustic noises. The fundamental frequency are estimated by the HHT-Amp technique. Harmonic patterns estimated with low accuracy are detected and adjusted according the FSFFE low/high pitch separation. The central frequencies of the filterbank are defined considering the third octave subbands which are best suited to cover the regions most relevant to intelligibility. Before signal reconstruction, the gammachirp filtered components are amplified by gain factors regulated by FSFFE classification. The proposed HDAG solution and three baseline techniques are examined considering six background noises with four signal-to-noise ratios. Three objective measures are adopted for the evaluation of speech intelligibility and quality. Several experiments are conducted to demonstrate that the proposed scheme achieves better speech intelligibility improvement when compared to the competing approaches. A perceptual listening test is further considered and corroborates with the objective results.


【17】 Advancing Accessibility: Voice Cloning and Speech Synthesis for  Individuals with Speech Disorders
标题:提高可获得性:语音障碍患者的语音克隆和语音合成
链接:https://arxiv.org/abs/2401.11771
作者:Vinotha R,Hepsiba D,L. D. Vijay Anand,Deepak John Reji
摘要:神经文本到语音(TTS)合成是一种功能强大的技术,可以使用神经网络生成语音。TTS合成最显著的特点之一是它能够产生不同说话人的语音。本文介绍了语音克隆和语音合成https://pypi.org/project/voice-cloning/,这是一个开源的python软件包,用于帮助语音障碍者更有效地交流,以及寻求将语音克隆或语音合成功能集成到他们的项目中的专业人士。该软件包旨在生成听起来像个人自然声音的合成语音,但它不会取代自然的人类声音。该系统的结构包括说话人确认系统、合成器、声码器和降噪。说话人确认系统在不同的说话人集合上进行训练,以实现最佳的泛化性能,而不依赖于transmittance。合成器使用从文本生成Mel频谱图的音频和transmittance两者以及将生成的Mel频谱图转换成相应的音频信号的声码器来训练。然后通过降噪算法处理音频信号以消除不需要的噪声并增强语音清晰度。然后使用主观和客观评价,如平均意见分数(MOS),总音高误差(GPE),和频谱失真(SD)的合成语音的性能进行评估。该模型可以通过包括随机选择的说话者特征来创建不同声音的语音。
摘要:Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This paper introduces voice cloning and speech synthesis https://pypi.org/project/voice-cloning/ an open-source python package for helping speech disorders to communicate more effectively as well as for professionals seeking to integrate voice cloning or speech synthesis capabilities into their projects. This package aims to generate synthetic speech that sounds like the natural voice of an individual, but it does not replace the natural human voice. The architecture of the system comprises a speaker verification system, a synthesizer, a vocoder, and noise reduction. Speaker verification system trained on a varied set of speakers to achieve optimal generalization performance without relying on transcriptions. Synthesizer is trained using both audio and transcriptions that generate Mel spectrogram from a text and vocoder which converts the generated Mel Spectrogram into corresponding audio signal. Then the audio signal is processed by a noise reduction algorithm to eliminate unwanted noise and enhance speech clarity. The performance of synthesized speech from seen and unseen speakers are then evaluated using subjective and objective evaluation such as Mean Opinion Score (MOS), Gross Pitch Error (GPE), and Spectral distortion (SD). The model can create speech in distinct voices by including speaker characteristics that are chosen randomly.

【18】 Streaming Bilingual End-to-End ASR model using Attention over Multiple  Softmax
标题:基于多软最大关注的流媒体双语端到端ASR模型
链接:https://arxiv.org/abs/2401.11645
作者:Aditya Patil,Vikas Joshi,Purvi Agrawal,Rupesh Mehta
备注:None
摘要:即使在多语言建模方面取得了一些进展,使用单个神经模型识别多种语言仍然具有挑战性,而不知道输入语言,并且大多数多语言模型假设输入语言的可用性。在这项工作中,我们提出了一种新的双语端到端(E2E)建模方法,其中单个神经模型可以识别两种语言,并支持语言之间的切换,而无需用户输入任何语言。该模型具有共享的编码器和预测网络,以及通过自注意机制组合的特定于语言的联合网络。当语言特定的后验被组合时,它在所有输出符号上产生单个后验概率,从而实现单波束搜索解码,并且还允许在语言之间动态切换。该方法优于传统的双语基线与13.3%,8.23%和1.3%的字错误率相对减少印地语,英语和代码混合测试集,分别。
摘要:Even with several advancements in multilingual modeling, it is challenging to recognize multiple languages using a single neural model, without knowing the input language and most multilingual models assume the availability of the input language. In this work, we propose a novel bilingual end-to-end (E2E) modeling approach, where a single neural model can recognize both languages and also support switching between the languages, without any language input from the user. The proposed model has shared encoder and prediction networks, with language-specific joint networks that are combined via a self-attention mechanism. As the language-specific posteriors are combined, it produces a single posterior probability over all the output symbols, enabling a single beam search decoding and also allowing dynamic switching between the languages. The proposed approach outperforms the conventional bilingual baseline with 13.3%, 8.23% and 1.3% word error rate relative reduction on Hindi, English and code-mixed test sets, respectively.

【19】 StreamVoice: Streamable Context-Aware Language Modeling for Real-time  Zero-Shot Voice Conversion
标题:StreamVoice:用于实时零发声转换的可流上下文感知语言建模
链接:https://arxiv.org/abs/2401.11053
作者:Zhichao Wang,Yuanzhe Chen,Xinsheng Wang,Zhuo Chen,Lei Xie,Yuping Wang,Yuxuan Wang
摘要:最近的语言模型(LM)进步展示了令人印象深刻的zero-shot语音转换(VC)性能。然而,现有的基于LM的VC模型通常应用从源语义到声学特征的离线转换,要求完整的源语音,并限制其部署到实时应用。在本文中,我们介绍了StreamVoice,一种新的流LM为基础的模型zero-shot VC,方便实时转换任意扬声器提示和源语音。具体而言,为了实现流传输能力,StreamVoice采用具有时间独立声学预测器的完全因果上下文感知LM,同时在自回归的每个时间步长交替处理语义和声学特征,这消除了对完整源语音的依赖。为了解决流媒体处理中不完整上下文可能带来的性能下降问题,我们通过两种策略增强LM的上下文感知能力:1)教师引导的上下文预见,使用教师模型在训练过程中总结当前和未来的语义上下文,以指导模型对缺失上下文的预测;(2)语义掩蔽策略,促进对先前已破坏的语义和声学输入的预测,增强上下文学习能力。值得注意的是,StreamVoice是第一个基于LM的流媒体zero-shot VC模型,没有任何未来展望。实验结果表明,StreamVoice的流转换能力,同时保持zero-shot性能相比,非流VC系统。
摘要:Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance. However, existing LM-based VC models usually apply offline conversion from source semantics to acoustic features, demanding the complete source speech, and limiting their deployment to real-time applications. In this paper, we introduce StreamVoice, a novel streaming LM-based model for zero-shot VC, facilitating real-time conversion given arbitrary speaker prompts and source speech. Specifically, to enable streaming capability, StreamVoice employs a fully causal context-aware LM with a temporal-independent acoustic predictor, while alternately processing semantic and acoustic features at each time step of autoregression which eliminates the dependence on complete source speech. To address the potential performance degradation from the incomplete context in streaming processing, we enhance the context-awareness of the LM through two strategies: 1) teacher-guided context foresight, using a teacher model to summarize the present and future semantic context during training to guide the model's forecasting for missing context; 2) semantic masking strategy, promoting acoustic prediction from preceding corrupted semantic and acoustic input, enhancing context-learning ability. Notably, StreamVoice is the first LM-based streaming zero-shot VC model without any future look-ahead. Experimental results demonstrate StreamVoice's streaming conversion capability while maintaining zero-shot performance comparable to non-streaming VC systems.

【20】 Revealing Emotional Clusters in Speaker Embeddings: A Contrastive  Learning Strategy for Speech Emotion Recognition
标题:在说话人嵌入中揭示情感簇:一种语音情感识别的对比学习策略
链接:https://arxiv.org/abs/2401.11017
作者:Ismail Rasim Ulgen,Zongyang Du,Carlos Busso,Berrak Sisman
备注:Accepted to ICASSP 2024
摘要:说话人信息中含有丰富的情感信息,是提高语音情感识别能力的有效途径,特别是在有限的标注数据下。传统上,人们一直认为情感信息是间接嵌入在说话人嵌入,导致他们的利用不足。我们的研究揭示了一个直接的和有用的联系之间的情感和国家的最先进的扬声器嵌入的形式内扬声器集群。通过进行彻底的聚类分析,我们证明了情感信息可以很容易地从说话人嵌入提取。为了利用这些信息,我们引入了一种新的对比预训练方法,应用于语音情感识别的情感未标记的数据。所提出的方法涉及的积极和消极的例子的基础上的说话人嵌入的说话人内集群的采样。所提出的策略,它利用了广泛的情绪未标记的数据,导致SER性能的显着改善,无论是作为一个独立的预训练任务或集成到一个多任务预训练设置。
摘要:Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion information is indirectly embedded within speaker embeddings, leading to their under-utilization. Our study reveals a direct and useful link between emotion and state-of-the-art speaker embeddings in the form of intra-speaker clusters. By conducting a thorough clustering analysis, we demonstrate that emotion information can be readily extracted from speaker embeddings. In order to leverage this information, we introduce a novel contrastive pretraining approach applied to emotion-unlabeled data for speech emotion recognition. The proposed approach involves the sampling of positive and the negative examples based on the intra-speaker clusters of speaker embeddings. The proposed strategy, which leverages extensive emotion-unlabeled data, leads to a significant improvement in SER performance, whether employed as a standalone pretraining task or integrated into a multi-task pretraining setting.

eess.AS音频处理
【1】 ScoreDec: A Phase-preserving High-Fidelity Audio Codec with A  Generalized Score-based Diffusion Post-filter
标题:ScoreDec:一种基于广义分数扩散后置滤波的保相高保真音频编解码器
链接:https://arxiv.org/abs/2401.12160
作者:Yi-Chiao Wu,Dejan Marković,Steven Krenn,Israel D. Gebru,Alexander Richard
备注:5 pages, 3 figures, 2 tables. Proc. ICASSP, 2024
摘要:虽然最近主流的波形域端到端(E2E)神经音频编解码器以非常低的比特率实现了令人印象深刻的编码音频质量,但编码音频和自然音频之间的质量差距仍然很大。由于直接相位建模的困难,这些E2E神经编解码器通常需要生成对抗网络(GAN)训练。然而,这种对抗性学习阻碍了这些编解码器保留原始相位信息。为了以合理的比特率实现人类水平的自然度,保留原始相位,并摆脱棘手和不透明的GAN训练,我们在复频谱域中开发了一个基于分数的扩散后滤波器(SPF),并将我们以前的AudioDec与SPF结合起来,提出了ScoreDec,它可以只使用频谱和分数匹配损失进行训练。客观和主观实验结果均表明,在24kbps码率下,ScoreDec对48kHz全频带语音进行编解码时,语音的自然度和相位信息都得到了很好的保留。
摘要:Although recent mainstream waveform-domain end-to-end (E2E) neural audio codecs achieve impressive coded audio quality with a very low bitrate, the quality gap between the coded and natural audio is still significant. A generative adversarial network (GAN) training is usually required for these E2E neural codecs because of the difficulty of direct phase modeling. However, such adversarial learning hinders these codecs from preserving the original phase information. To achieve human-level naturalness with a reasonable bitrate, preserve the original phase, and get rid of the tricky and opaque GAN training, we develop a score-based diffusion post-filter (SPF) in the complex spectral domain and combine our previous AudioDec with the SPF to propose ScoreDec, which can be trained using only spectral and score-matching losses. Both the objective and subjective experimental results show that ScoreDec with a 24~kbps bitrate encodes and decodes full-band 48~kHz speech with human-level naturalness and well-preserved phase information.


【2】 Consistency Based Unsupervised Self-training For ASR Personalisation
标题:基于一致性的ASR个性化无监督自我训练
链接:https://arxiv.org/abs/2401.12085
作者:Jisi Zhang,Vandana Rajan,Haaris Mehmood,David Tuckey,Pablo Peso Parada,Md Asif Jalal,Karthikeyan Saravanan,Gil Ho Lee,Jungin Lee,Seokyeong Jung
备注:Accepted for IEEE ASRU 2023
摘要:在大量人群的语音数据上训练的设备上自动语音识别(ASR)模型可能会在训练期间看不见的个人中表现不佳。这是由于用户数据和原始训练数据之间的域移位,不同的用户的说话特性和环境声学条件。ASR个性化是一种旨在利用用户数据来提高模型鲁棒性的解决方案。大多数ASR个性化方法都假设标记的用户数据用于监督。由于有限的数据大小和录制的音频样本质量差,没有任何标记数据的个性化是具有挑战性的。这项工作通过伪标签开发一种新的基于一致性的训练方法来解决无监督的个性化。与预训练模型相比,我们的方法在未标记的训练数据上实现了17.3%的相对字错误率降低(WERR),在保留数据上实现了8.1%的相对字错误率降低(WERR),并且优于当前最先进的方法。
摘要:On-device Automatic Speech Recognition (ASR) models trained on speech data of a large population might underperform for individuals unseen during training. This is due to a domain shift between user data and the original training data, differed by user's speaking characteristics and environmental acoustic conditions. ASR personalisation is a solution that aims to exploit user data to improve model robustness. The majority of ASR personalisation methods assume labelled user data for supervision. Personalisation without any labelled data is challenging due to limited data size and poor quality of recorded audio samples. This work addresses unsupervised personalisation by developing a novel consistency based training method via pseudo-labelling. Our method achieves a relative Word Error Rate Reduction (WERR) of 17.3% on unlabelled training data and 8.1% on held-out data compared to a pre-trained model, and outperforms the current state-of-the art methods.


【3】 Adversarial speech for voice privacy protection from Personalized Speech  generation
标题:对抗性语音用于个性化语音生成中的语音隐私保护
链接:https://arxiv.org/abs/2401.11857
作者:Shihao Chen,Liping Chen,Jie Zhang,KongAik Lee,Zhenhua Ling,Lirong Dai
备注:Accepted by icassp 2024
摘要:个性化语音生成技术的快速发展,包括个性化文本到语音(TTS)和语音转换(VC),对人类收听者区分生成的语音和真实语音提出了挑战,从而迫切需要保护说话者的语音免受恶意滥用。在这方面,我们提出了一种基于对抗性攻击的说话人保护方法。所提出的方法通过最小限度地改变原始语音来扰动语音信号,同时使下游语音生成模型无法准确地生成目标说话者的语音。为了验证,我们采用开源的预训练的YourTTS模型进行语音生成,并在白盒场景中保护目标说话者的语音。自动说话人确认(ASV)的评价进行了对所产生的语音的语音保护能力的评估。我们的实验结果表明,我们成功地扰动扬声器编码器的YourTTS模型使用基于梯度的I-FGSM对抗扰动方法。此外,对抗性扰动在防止YourTTS模型生成目标说话者的语音方面是有效的。音频样本可以在https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS上找到。
摘要:The rapid progress in personalized speech generation technology, including personalized text-to-speech (TTS) and voice conversion (VC), poses a challenge in distinguishing between generated and real speech for human listeners, resulting in an urgent demand in protecting speakers' voices from malicious misuse. In this regard, we propose a speaker protection method based on adversarial attacks. The proposed method perturbs speech signals by minimally altering the original speech while rendering downstream speech generation models unable to accurately generate the voice of the target speaker. For validation, we employ the open-source pre-trained YourTTS model for speech generation and protect the target speaker's speech in the white-box scenario. Automatic speaker verification (ASV) evaluations were carried out on the generated speech as the assessment of the voice protection capability. Our experimental results show that we successfully perturbed the speaker encoder of the YourTTS model using the gradient-based I-FGSM adversarial perturbation method. Furthermore, the adversarial perturbation is effective in preventing the YourTTS model from generating the speech of the target speaker. Audio samples can be found in https://voiceprivacy.github.io/Adeversarial-Speech-with-YourTTS.

【4】 Intelligibility Enhancement of Acoustic Noisy Speech for Autism Spectrum  Disorder Condition
标题:提高自闭症谱系障碍声学含噪语音的清晰度
链接:https://arxiv.org/abs/2401.11832
作者:M. Pillonetto,A. Queiroz,R. Coelho
备注:5 pages, 3 figues, 2 tables
摘要:本文介绍了一种时域个性化方法(pGTFF0),用于改善自闭症谱系障碍(ASD)患者的噪声语音清晰度。对于这个建议,从语音帧估计的谐波特征被认为是Gammatone听觉滤波器组的中心频率。增益因子被进一步应用于滤波样本的输出。关键目标是模拟为ASD患者量身定制的外部噪声过滤。感知听力测试表明,ASD志愿者达到较低的可懂率比神经(NT)。所提出的解决方案相比,三个竞争的方法,考虑在不同的信号噪声比的四个声学噪声。评估还采用了两个客观指标(ESTOI和PESQ)。实验结果表明,个性化的解决方案优于竞争的方法,在可懂度和质量的改善。
摘要:This work introduces a time domain personalized method (pGTFF0) to achieve intelligibility improvement of noisy speech for Autism Spectrum Disorder (ASD) situation. For this proposal, harmonic features estimated from speech frames are considered as center frequencies of Gammatone auditory filterbanks. A gain factor is further applied to the output of the filtered samples. The key goal is the emulation of an external noise filtering tailored for individuals with ASD. A perceptual listening test demonstrates that ASD volunteers attained lower intelligibility rates than Neurotypical (NT). The proposed solution is compared to three competing approaches considering four acoustic noises at different signal-to-noise ratios. Two objective measures (ESTOI and PESQ) are also adopted for evaluation. The experimental results show that the personalized solution outperformed the competing approaches in terms of intelligibility and quality improvement.


【5】 Harmonic Detection from Noisy Speech with Auditory Frame Gain for  Intelligibility Enhancement标题:利用听觉帧增益提高清晰度的噪声语音中的谐波检测
链接:https://arxiv.org/abs/2401.11829
作者:A. Queiroz,R. Coelho
备注:9 pages, 6 figures, 4 tables
摘要:本文介绍了一种新的方法(HDAG -谐波检测听觉增益)的语音清晰度增强在嘈杂的情况下。在该方案中,采用一系列的选择性Gammachirp滤波器来强调语音的谐波成分,减少声学噪声的掩蔽效应。利用HHT-DFT技术对基频进行了估计。根据FSFFE低/高音调分离检测和调整以低精度估计的谐波图案。滤波器组的中心频率被定义为考虑最适合覆盖与可懂度最相关的区域的第三倍频程子带。在信号重构之前,通过由FSFFE分类调节的增益因子来放大伽马线性滤波分量。建议HDAG解决方案和三个基线技术进行检查,考虑六个背景噪声与四个信噪比。采用三种客观的方法来评价语音清晰度和语音质量。几个实验进行了证明,所提出的方案实现了更好的语音清晰度的改善相比,竞争的方法。一个感性的听力测试进一步考虑和佐证的客观结果。
摘要:This paper introduces a novel (HDAG - Harmonic Detection for Auditory Gain) method for speech intelligibility enhancement in noisy scenarios. In the proposed scheme, a series of selective Gammachirp filters are adopted to emphasize the harmonic components of speech reducing the masking effects of acoustic noises. The fundamental frequency are estimated by the HHT-Amp technique. Harmonic patterns estimated with low accuracy are detected and adjusted according the FSFFE low/high pitch separation. The central frequencies of the filterbank are defined considering the third octave subbands which are best suited to cover the regions most relevant to intelligibility. Before signal reconstruction, the gammachirp filtered components are amplified by gain factors regulated by FSFFE classification. The proposed HDAG solution and three baseline techniques are examined considering six background noises with four signal-to-noise ratios. Three objective measures are adopted for the evaluation of speech intelligibility and quality. Several experiments are conducted to demonstrate that the proposed scheme achieves better speech intelligibility improvement when compared to the competing approaches. A perceptual listening test is further considered and corroborates with the objective results.


【6】 Advancing Accessibility: Voice Cloning and Speech Synthesis for  Individuals with Speech Disorders
标题:提高可获得性:语音障碍患者的语音克隆和语音合成
链接:https://arxiv.org/abs/2401.11771
作者:Vinotha R,Hepsiba D,L. D. Vijay Anand,Deepak John Reji
摘要:神经文本到语音(TTS)合成是一种功能强大的技术,可以使用神经网络生成语音。TTS合成最显著的特点之一是它能够产生不同说话人的语音。本文介绍了语音克隆和语音合成https://pypi.org/project/voice-cloning/,这是一个开源的python软件包,用于帮助语音障碍者更有效地交流,以及寻求将语音克隆或语音合成功能集成到他们的项目中的专业人士。该软件包旨在生成听起来像个人自然声音的合成语音,但它不会取代自然的人类声音。该系统的结构包括说话人确认系统、合成器、声码器和降噪。说话人确认系统在不同的说话人集合上进行训练,以实现最佳的泛化性能,而不依赖于transmittance。合成器使用从文本生成Mel频谱图的音频和transmittance两者以及将生成的Mel频谱图转换成相应的音频信号的声码器来训练。然后通过降噪算法处理音频信号以消除不需要的噪声并增强语音清晰度。然后使用主观和客观评价,如平均意见分数(MOS),总音高误差(GPE),和频谱失真(SD)的合成语音的性能进行评估。该模型可以通过包括随机选择的说话者特征来创建不同声音的语音。
摘要:Neural Text-to-speech (TTS) synthesis is a powerful technology that can generate speech using neural networks. One of the most remarkable features of TTS synthesis is its capability to produce speech in the voice of different speakers. This paper introduces voice cloning and speech synthesis https://pypi.org/project/voice-cloning/ an open-source python package for helping speech disorders to communicate more effectively as well as for professionals seeking to integrate voice cloning or speech synthesis capabilities into their projects. This package aims to generate synthetic speech that sounds like the natural voice of an individual, but it does not replace the natural human voice. The architecture of the system comprises a speaker verification system, a synthesizer, a vocoder, and noise reduction. Speaker verification system trained on a varied set of speakers to achieve optimal generalization performance without relying on transcriptions. Synthesizer is trained using both audio and transcriptions that generate Mel spectrogram from a text and vocoder which converts the generated Mel Spectrogram into corresponding audio signal. Then the audio signal is processed by a noise reduction algorithm to eliminate unwanted noise and enhance speech clarity. The performance of synthesized speech from seen and unseen speakers are then evaluated using subjective and objective evaluation such as Mean Opinion Score (MOS), Gross Pitch Error (GPE), and Spectral distortion (SD). The model can create speech in distinct voices by including speaker characteristics that are chosen randomly.

【7】 Streaming Bilingual End-to-End ASR model using Attention over Multiple  Softmax
标题:基于多软最大关注的流媒体双语端到端ASR模型
链接:https://arxiv.org/abs/2401.11645
作者:Aditya Patil,Vikas Joshi,Purvi Agrawal,Rupesh Mehta
备注:None
摘要:即使在多语言建模方面取得了一些进展,使用单个神经模型识别多种语言仍然具有挑战性,而不知道输入语言,并且大多数多语言模型假设输入语言的可用性。在这项工作中,我们提出了一种新的双语端到端(E2E)建模方法,其中单个神经模型可以识别两种语言,并支持语言之间的切换,而无需用户输入任何语言。该模型具有共享的编码器和预测网络,以及通过自注意机制组合的特定于语言的联合网络。当语言特定的后验被组合时,它在所有输出符号上产生单个后验概率,从而实现单波束搜索解码,并且还允许在语言之间动态切换。该方法优于传统的双语基线与13.3%,8.23%和1.3%的字错误率相对减少印地语,英语和代码混合测试集,分别。
摘要:Even with several advancements in multilingual modeling, it is challenging to recognize multiple languages using a single neural model, without knowing the input language and most multilingual models assume the availability of the input language. In this work, we propose a novel bilingual end-to-end (E2E) modeling approach, where a single neural model can recognize both languages and also support switching between the languages, without any language input from the user. The proposed model has shared encoder and prediction networks, with language-specific joint networks that are combined via a self-attention mechanism. As the language-specific posteriors are combined, it produces a single posterior probability over all the output symbols, enabling a single beam search decoding and also allowing dynamic switching between the languages. The proposed approach outperforms the conventional bilingual baseline with 13.3%, 8.23% and 1.3% word error rate relative reduction on Hindi, English and code-mixed test sets, respectively.

【8】 StreamVoice: Streamable Context-Aware Language Modeling for Real-time  Zero-Shot Voice Conversion
标题:StreamVoice:用于实时零发声转换的可流上下文感知语言建模
链接:https://arxiv.org/abs/2401.11053
作者:Zhichao Wang,Yuanzhe Chen,Xinsheng Wang,Zhuo Chen,Lei Xie,Yuping Wang,Yuxuan Wang
摘要:最近的语言模型(LM)进步展示了令人印象深刻的zero-shot语音转换(VC)性能。然而,现有的基于LM的VC模型通常应用从源语义到声学特征的离线转换,要求完整的源语音,并限制其部署到实时应用。在本文中,我们介绍了StreamVoice,一种新的流LM为基础的模型zero-shot VC,方便实时转换任意扬声器提示和源语音。具体而言,为了实现流传输能力,StreamVoice采用具有时间独立声学预测器的完全因果上下文感知LM,同时在自回归的每个时间步长交替处理语义和声学特征,这消除了对完整源语音的依赖。为了解决流媒体处理中不完整上下文可能带来的性能下降问题,我们通过两种策略增强LM的上下文感知能力:1)教师引导的上下文预见,使用教师模型在训练过程中总结当前和未来的语义上下文,以指导模型对缺失上下文的预测;(2)语义掩蔽策略,促进对先前已破坏的语义和声学输入的预测,增强上下文学习能力。值得注意的是,StreamVoice是第一个基于LM的流媒体zero-shot VC模型,没有任何未来展望。实验结果表明,StreamVoice的流转换能力,同时保持zero-shot性能相比,非流VC系统。
摘要:Recent language model (LM) advancements have showcased impressive zero-shot voice conversion (VC) performance. However, existing LM-based VC models usually apply offline conversion from source semantics to acoustic features, demanding the complete source speech, and limiting their deployment to real-time applications. In this paper, we introduce StreamVoice, a novel streaming LM-based model for zero-shot VC, facilitating real-time conversion given arbitrary speaker prompts and source speech. Specifically, to enable streaming capability, StreamVoice employs a fully causal context-aware LM with a temporal-independent acoustic predictor, while alternately processing semantic and acoustic features at each time step of autoregression which eliminates the dependence on complete source speech. To address the potential performance degradation from the incomplete context in streaming processing, we enhance the context-awareness of the LM through two strategies: 1) teacher-guided context foresight, using a teacher model to summarize the present and future semantic context during training to guide the model's forecasting for missing context; 2) semantic masking strategy, promoting acoustic prediction from preceding corrupted semantic and acoustic input, enhancing context-learning ability. Notably, StreamVoice is the first LM-based streaming zero-shot VC model without any future look-ahead. Experimental results demonstrate StreamVoice's streaming conversion capability while maintaining zero-shot performance comparable to non-streaming VC systems.

【9】 Revealing Emotional Clusters in Speaker Embeddings: A Contrastive  Learning Strategy for Speech Emotion Recognition
标题:在说话人嵌入中揭示情感簇:一种语音情感识别的对比学习策略
链接:https://arxiv.org/abs/2401.11017
作者:Ismail Rasim Ulgen,Zongyang Du,Carlos Busso,Berrak Sisman
备注:Accepted to ICASSP 2024
摘要:说话人信息中含有丰富的情感信息,是提高语音情感识别能力的有效途径,特别是在有限的标注数据下。传统上,人们一直认为情感信息是间接嵌入在说话人嵌入,导致他们的利用不足。我们的研究揭示了一个直接的和有用的联系之间的情感和国家的最先进的扬声器嵌入的形式内扬声器集群。通过进行彻底的聚类分析,我们证明了情感信息可以很容易地从说话人嵌入提取。为了利用这些信息,我们引入了一种新的对比预训练方法,应用于语音情感识别的情感未标记的数据。所提出的方法涉及的积极和消极的例子的基础上的说话人嵌入的说话人内集群的采样。所提出的策略,它利用了广泛的情绪未标记的数据,导致SER性能的显着改善,无论是作为一个独立的预训练任务或集成到一个多任务预训练设置。
摘要:Speaker embeddings carry valuable emotion-related information, which makes them a promising resource for enhancing speech emotion recognition (SER), especially with limited labeled data. Traditionally, it has been assumed that emotion information is indirectly embedded within speaker embeddings, leading to their under-utilization. Our study reveals a direct and useful link between emotion and state-of-the-art speaker embeddings in the form of intra-speaker clusters. By conducting a thorough clustering analysis, we demonstrate that emotion information can be readily extracted from speaker embeddings. In order to leverage this information, we introduce a novel contrastive pretraining approach applied to emotion-unlabeled data for speech emotion recognition. The proposed approach involves the sampling of positive and the negative examples based on the intra-speaker clusters of speaker embeddings. The proposed strategy, which leverages extensive emotion-unlabeled data, leads to a significant improvement in SER performance, whether employed as a standalone pretraining task or integrated into a multi-task pretraining setting.


【10】 DITTO: Diffusion Inference-Time T-Optimization for Music Generation
标题:同上:音乐生成的扩散推理时间T-优化
链接:https://arxiv.org/abs/2401.12179
作者:Zachary Novack,Julian McAuley,Taylor Berg-Kirkpatrick,Nicholas J. Bryan
摘要:我们提出了扩散推理时间T优化(DITTO),这是一种通用框架,用于通过优化初始噪声潜伏期来控制预训练的文本到音乐扩散模型。我们的方法可用于通过任何可微特征匹配损失进行优化,以实现目标(风格化)输出,并利用梯度检查点提高内存效率。我们展示了一个令人惊讶的广泛的应用程序的音乐生成,包括修复,outpainting,循环以及强度,旋律和音乐结构控制-所有没有微调的基础模型。当我们将我们的方法与相关的训练,指导和基于优化的方法进行比较时,我们发现DITTO在几乎所有任务上都达到了最先进的性能,包括在可控性,音频质量和计算效率方面优于可比方法,从而为高质量,灵活,无训练的扩散模型控制打开了大门。可以在https://DITTO-Music.github.io/web/上找到合理的例子。
摘要:We propose Diffusion Inference-Time T-Optimization (DITTO), a general-purpose frame-work for controlling pre-trained text-to-music diffusion models at inference-time via optimizing initial noise latents. Our method can be used to optimize through any differentiable feature matching loss to achieve a target (stylized) output and leverages gradient checkpointing for memory efficiency. We demonstrate a surprisingly wide-range of applications for music generation including inpainting, outpainting, and looping as well as intensity, melody, and musical structure control - all without ever fine-tuning the underlying model. When we compare our approach against related training, guidance, and optimization-based methods, we find DITTO achieves state-of-the-art performance on nearly all tasks, including outperforming comparable approaches on controllability, audio quality, and computational efficiency, thus opening the door for high-quality, flexible, training-free control of diffusion models. Sound examples can be found at https://DITTO-Music.github.io/web/.


【11】 Resource-constrained stereo singing voice cancellation
标题:资源受限的立体声演唱语音消除
链接:https://arxiv.org/abs/2401.12068
作者:Clara Borrelli,James Rae,Dogac Basaran,Matt McVicar,Mehrez Souden,Matthias Mauch
摘要:我们研究的问题,立体声歌唱的声音消除,音乐源分离的子任务,其目标是从立体声混合估计乐器背景。我们将探索如何从一个小型、高效的实时语音分离模型开始,实现与大型最先进的源分离网络相似的性能。这种模型在内存和计算有限并且歌声处理必须以有限的前瞻运行时是有用的。在实践中,这是通过调整现有的单声道模型来处理立体声输入来实现的。通过调整模型参数和扩展训练集来提高质量。此外,我们强调立体声模型带来的好处,通过引入一个新的度量,检测通道之间的衰减不一致。我们的方法使用客观的离线指标和大规模的MUSHRA试验进行评估,证实了我们的技术在严格的听力测试中的有效性。
摘要:We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art source separation networks starting from a small, efficient model for real-time speech separation. Such a model is useful when memory and compute are limited and singing voice processing has to run with limited look-ahead. In practice, this is realised by adapting an existing mono model to handle stereo input. Improvements in quality are obtained by tuning model parameters and expanding the training set. Moreover, we highlight the benefits a stereo model brings by introducing a new metric which detects attenuation inconsistencies between channels. Our approach is evaluated using objective offline metrics and a large-scale MUSHRA trial, confirming the effectiveness of our techniques in stringent listening tests.


【12】 NEUROSEC: FPGA-Based Neuromorphic Audio Security
标题:NeurOSEC:基于FPGA的神经形态音频安全
链接:https://arxiv.org/abs/2401.12055
作者:Murat Isik,Hiruna Vishwamith,Yusuf Sur,Kayode Inadagbo,I. Can Dikmen
备注:Audio processing, FPGA, Hardware Security, Neuromorphic Computing
摘要:受人脑复杂性和功能性的启发,神经形态系统由于其在广泛应用中无与伦比的潜力而引起了学术和工业界的关注。虽然它们的能力预示着创新,但必须强调的是,这些计算范式与传统的计算范式类似,并非不受安全威胁的影响。虽然用于图像和视频处理的神经形态方法的探索已经被严格地追求,但神经形态音频处理的领域仍然处于早期阶段。我们的研究结果突出了我们的基于FPGA的神经形态系统的鲁棒性和精度。具体来说,我们的系统在期望信号和背景噪声之间实现了值得称赞的平衡,有效的尖峰速率编码,以及对FGSM和PGD等对抗性攻击的无与伦比的弹性。我们的框架的一个突出特点是其检测率为94%,与其他方法相比,强调了其在5.39 dB范围内识别和缓解威胁的能力,这是一个值得称赞的SNR比。此外,神经形态计算和硬件安全服务于关键任务和隐私保护应用中的许多传感器领域。
摘要:Neuromorphic systems, inspired by the complexity and functionality of the human brain, have gained interest in academic and industrial attention due to their unparalleled potential across a wide range of applications. While their capabilities herald innovation, it is imperative to underscore that these computational paradigms, analogous to their traditional counterparts, are not impervious to security threats. Although the exploration of neuromorphic methodologies for image and video processing has been rigorously pursued, the realm of neuromorphic audio processing remains in its early stages. Our results highlight the robustness and precision of our FPGA-based neuromorphic system. Specifically, our system showcases a commendable balance between desired signal and background noise, efficient spike rate encoding, and unparalleled resilience against adversarial attacks such as FGSM and PGD. A standout feature of our framework is its detection rate of 94%, which, when compared to other methodologies, underscores its greater capability in identifying and mitigating threats within 5.39 dB, a commendable SNR ratio. Furthermore, neuromorphic computing and hardware security serve many sensor domains in mission-critical and privacy-preserving applications.


【13】 Look, Listen and Recognise: Character-Aware Audio-Visual Subtitling
标题:看、听、认:角色感知的视听字幕
链接:https://arxiv.org/abs/2401.12039
作者:Bruno Korbar,Jaesung Huh,Andrew Zisserman
备注:Accepted for publication in ICASSP 2024
摘要:本文的目标是自动字符感知字幕生成。给定一个视频和最少量的元数据,我们提出了一个视听方法,生成一个完整的对话记录,精确的语音时间戳,和字符说话识别。其核心思想是首先使用视听线索为每个字符选择一组高精度的音频样本,然后使用这些样本对所有语音段进行说话人身份分类。值得注意的是,该方法不需要面部检测或跟踪。我们评估了各种电视情景喜剧,包括宋飞,Fraiser和灌木丛的方法。我们设想这个系统是有用的自动生成字幕,以提高现代流媒体服务上可用的大量视频的可访问性。项目页面:\url{https://www.robots.ox.ac.uk/juvengg/research/look-juven-juvenise/}
摘要:The goal of this paper is automatic character-aware subtitle generation. Given a video and a minimal amount of metadata, we propose an audio-visual method that generates a full transcript of the dialogue, with precise speech timestamps, and the character speaking identified. The key idea is to first use audio-visual cues to select a set of high-precision audio exemplars for each character, and then use these exemplars to classify all speech segments by speaker identity. Notably, the method does not require face detection or tracking. We evaluate the method over a variety of TV sitcoms, including Seinfeld, Fraiser and Scrubs. We envision this system being useful for the automatic generation of subtitles to improve the accessibility of the vast amount of videos available on modern streaming services. Project page : \url{https://www.robots.ox.ac.uk/~vgg/research/look-listen-recognise/}

【14】 Lightweight Protection for Privacy in Offloaded Speech Understanding
标题:卸载语音理解中的轻量级隐私保护
链接:https://arxiv.org/abs/2401.11983
作者:Dongqi Cai,Shangguang Wang,Zeling Zhang,Felix Xiaozhu Lin,Mengwei Xu
备注:under review
摘要:语音是移动嵌入式设备的常见输入方法,但基于云的语音识别系统会带来隐私风险。基于解纠缠的编码器被设计为通过从语音信号中过滤敏感信息来保护用户隐私,不幸的是需要大量的存储器和计算资源,这限制了它们在功能较弱的设备中的使用。为了克服这一点,我们介绍了一种新的系统,XXX,优化了这样的设备。XXX是建立在这样一种见解之上的:语音理解主要依赖于理解整个话语的长期依赖性,而隐私问题通常与短期细节有关。因此,XXX专注于选择性地掩蔽这些短期元素,保持长期语音理解的质量。XXX的核心是一个创新的差分掩码生成器,以可解释学习为基础,对掩码过程进行微调。我们在STM32H7微控制器上测试了XXX,评估了其在各种潜在攻击场景中的性能。结果表明,XXX保持了与现有编码器相当的语音理解准确性和隐私性,但在效率上有了显着提高,处理速度提高了53.3倍,内存占用减少了134.1倍。
摘要:Speech is a common input method for mobile embedded devices, but cloud-based speech recognition systems pose privacy risks. Disentanglement-based encoders, designed to safeguard user privacy by filtering sensitive information from speech signals, unfortunately require substantial memory and computational resources, which limits their use in less powerful devices. To overcome this, we introduce a novel system, XXX, optimized for such devices. XXX is built on the insight that speech understanding primarily relies on understanding the entire utterance's long-term dependencies, while privacy concerns are often linked to short-term details. Therefore, XXX focuses on selectively masking these short-term elements, preserving the quality of long-term speech understanding. The core of XXX is an innovative differential mask generator, grounded in interpretable learning, which fine-tunes the masking process. We tested XXX on the STM32H7 microcontroller, assessing its performance in various potential attack scenarios. The results show that XXX maintains speech understanding accuracy and privacy at levels comparable to existing encoders, but with a significant improvement in efficiency, achieving up to 53.3$\times$ faster processing and a 134.1$\times$ smaller memory footprint.

【15】 Keep Decoding Parallel with Effective Knowledge Distillation from  Language Models to End-to-end Speech Recognisers
标题:保持解码与从语言模型到端到端语音识别器的有效知识提取的并行性
链接:https://arxiv.org/abs/2401.11700
作者:Michael Hentschel,Yuta Nishikawa,Tatsuya Komatsu,Yusuke Fujita
备注:Accepted at ICASSP 2024
摘要:本研究提出一种新的方法,知识蒸馏(KD)从一个BERT教师模型的自动语音识别(ASR)模型使用中间层。为了验证教师的知识,我们使用了一个注意力解码器,它从BERT的令牌概率中学习。我们的方法表明,语言模型(LM)的信息可以更有效地提取到一个ASR模型使用的中间层和最终层。通过将中间层作为蒸馏目标,我们可以更有效地将知识蒸馏到网络的较低层。使用我们的方法,我们实现了更好的识别精度比浅融合的外部LM,使我们能够保持快速并行解码。LibriSpeech数据集上的实验证明了我们的方法在增强贪婪解码与连接主义时间分类(CTC)的有效性。
摘要:This study presents a novel approach for knowledge distillation (KD) from a BERT teacher model to an automatic speech recognition (ASR) model using intermediate layers. To distil the teacher's knowledge, we use an attention decoder that learns from BERT's token probabilities. Our method shows that language model (LM) information can be more effectively distilled into an ASR model using both the intermediate layers and the final layer. By using the intermediate layers as distillation target, we can more effectively distil LM knowledge into the lower network layers. Using our method, we achieve better recognition accuracy than with shallow fusion of an external LM, allowing us to maintain fast parallel decoding. Experiments on the LibriSpeech dataset demonstrate the effectiveness of our approach in enhancing greedy decoding with connectionist temporal classification (CTC).


【16】 Word-Level ASR Quality Estimation for Efficient Corpus Sampling and  Post-Editing through Analyzing Attentions of a Reference-Free Metric
标题:基于词级ASR质量评估的语料库采样和后期编辑
链接:https://arxiv.org/abs/2401.11268
作者:Golara Javadi,Kamer Ali Yuksel,Yunsu Kim,Thiago Castro Ferreira,Mohamed Al-Badrashiny
摘要:在自动语音识别(ASR)领域,对模型的追求不仅具有高准确性,而且在决策过程中提供透明度是至关重要的。质量估计(QE)指标的潜力,介绍和评估作为一种新的工具,以提高可解释的人工智能(XAI)在ASR系统。通过实验和分析,NoRefER(无参考错误率)指标的能力进行了探索,在识别字级错误,以帮助后编辑在细化ASR假设。调查还扩展到在语料库建设过程中的NoRefER的效用,证明了其有效性,在增强数据集有见地的注释。的诊断方面的NoRefER检查,揭示其提供有价值的见解模型的行为和决策模式的能力。事实证明,这有利于在编辑后工作流程中优先考虑假设和微调ASR模型。研究结果表明,NoRefER不仅是一个错误检测工具,而且是一个提高ASR系统透明度、效率和有效性的综合框架。为了确保结果的重现性,本研究的所有源代码都是公开的。
摘要:In the realm of automatic speech recognition (ASR), the quest for models that not only perform with high accuracy but also offer transparency in their decision-making processes is crucial. The potential of quality estimation (QE) metrics is introduced and evaluated as a novel tool to enhance explainable artificial intelligence (XAI) in ASR systems. Through experiments and analyses, the capabilities of the NoRefER (No Reference Error Rate) metric are explored in identifying word-level errors to aid post-editors in refining ASR hypotheses. The investigation also extends to the utility of NoRefER in the corpus-building process, demonstrating its effectiveness in augmenting datasets with insightful annotations. The diagnostic aspects of NoRefER are examined, revealing its ability to provide valuable insights into model behaviors and decision patterns. This has proven beneficial for prioritizing hypotheses in post-editing workflows and fine-tuning ASR models. The findings suggest that NoRefER is not merely a tool for error detection but also a comprehensive framework for enhancing ASR systems' transparency, efficiency, and effectiveness. To ensure the reproducibility of the results, all source codes of this study are made publicly available.

【17】 Projected Belief Networks With Discriminative Alignment for Acoustic  Event Classification: Rivaling State of the Art CNNs
标题:用于声事件分类的判别对齐投影信念网络:与现有CNN相媲美
链接:https://arxiv.org/abs/2401.11199
作者:Paul M. Baggenstoss,Kevin Wilkinghoff,Felix Govaers,Frank Kurth
备注:15 Pages. Submitted to IEEE-TNNLS
摘要:投影信念网络(PBN)是一种基于前馈神经网络(FFNN)的具有易处理似然函数的生成式随机网络。生成函数通过FFNN进行“备份”操作。PBN是两个网络合二为一,一个是在前向方向上运行的FFNN,另一个是在后向方向上运行的生成网络。两个网络基于相同的参数集共存,具有自己的成本函数,并且可以单独或联合训练。因此,PBN有可能拥有最好的品质的歧视和生成分类器。为了实现这种潜力,在每个类上训练一个单独的PBN,最大化给定类的生成似然函数,同时最小化FFNN对“所有其他类”的区分成本。这种技术被称为判别对齐(PBN-DA),将似然函数的轮廓与决策边界对齐,并大大提高了分类性能,可与最先进的判别网络相媲美。该方法可以使用隐马尔可夫模型(HMM)作为PBN的组件(称为PBN-DA-HMM)来进一步改进。本文提供了PBN、PBN-DA和PBN-DA-HMM的综合处理。此外,两个新的分类实验的结果提供。第一个实验使用空气声学事件,第二个实验使用由海洋哺乳动物叫声组成的水声数据。在这两个实验中,PBN-DA-HMM获得了与现有技术CNN相当或更好的性能,并且在与CNN组合时获得了两倍的误差减少。
摘要:The projected belief network (PBN) is a generative stochastic network with tractable likelihood function based on a feed-forward neural network (FFNN). The generative function operates by "backing up" through the FFNN. The PBN is two networks in one, a FFNN that operates in the forward direction, and a generative network that operates in the backward direction. Both networks co-exist based on the same parameter set, have their own cost functions, and can be separately or jointly trained. The PBN therefore has the potential to possess the best qualities of both discriminative and generative classifiers. To realize this potential, a separate PBN is trained on each class, maximizing the generative likelihood function for the given class, while minimizing the discriminative cost for the FFNN against "all other classes". This technique, called discriminative alignment (PBN-DA), aligns the contours of the likelihood function to the decision boundaries and attains vastly improved classification performance, rivaling that of state of the art discriminative networks. The method may be further improved using a hidden Markov model (HMM) as a component of the PBN, called PBN-DA-HMM. This paper provides a comprehensive treatment of PBN, PBN-DA, and PBN-DA-HMM. In addition, the results of two new classification experiments are provided. The first experiment uses air-acoustic events, and the second uses underwater acoustic data consisting of marine mammal calls. In both experiments, PBN-DA-HMM attains comparable or better performance as a state of the art CNN, and attain a factor of two error reduction when combined with the CNN.

【18】 Generalizing Speaker Verification for Spoof Awareness in the Embedding  Space
标题:嵌入空间中基于欺骗感知的说话人确认泛化
链接:https://arxiv.org/abs/2401.11156
作者:Xuechen Liu,Md Sahidullah,Kong Aik Lee,Tomi Kinnunen
备注:To appear in IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:现在众所周知,自动说话人验证(ASV)系统可以使用各种类型的对手进行欺骗。对抗ASV系统的通常方法是开发一个单独的欺骗对策(CM)模块,将语音输入分类为真实的或欺骗的话语。然而,这样的设计在认证阶段需要额外的计算和利用工作。另一种策略涉及一个单一的单片ASV系统,旨在处理零努力冒名顶替者(非目标)和欺骗攻击。这种欺骗感知ASV系统有可能提供更强的保护和更经济的计算。为此,我们建议推广独立的ASV(G-SASV)来对抗欺骗攻击,其中我们利用来自CM的有限训练数据来增强嵌入空间中的简单后端,而无需在测试(认证)阶段涉及单独的CM模块。我们提出了一种基于深度神经网络的新颖而简单的后端分类器,并在训练阶段通过域自适应和欺骗嵌入的多任务集成进行研究。在ASVspoof 2019逻辑访问数据集上进行了实验,我们在联合(真实和欺骗)和欺骗条件下分别将统计ASV后端的性能提高了36.2%和49.8%。
摘要:It is now well-known that automatic speaker verification (ASV) systems can be spoofed using various types of adversaries. The usual approach to counteract ASV systems against such attacks is to develop a separate spoofing countermeasure (CM) module to classify speech input either as a bonafide, or a spoofed utterance. Nevertheless, such a design requires additional computation and utilization efforts at the authentication stage. An alternative strategy involves a single monolithic ASV system designed to handle both zero-effort imposter (non-targets) and spoofing attacks. Such spoof-aware ASV systems have the potential to provide stronger protections and more economic computations. To this end, we propose to generalize the standalone ASV (G-SASV) against spoofing attacks, where we leverage limited training data from CM to enhance a simple backend in the embedding space, without the involvement of a separate CM module during the test (authentication) phase. We propose a novel yet simple backend classifier based on deep neural networks and conduct the study via domain adaptation and multi-task integration of spoof embeddings at the training stage. Experiments are conducted on the ASVspoof 2019 logical access dataset, where we improve the performance of statistical ASV backends on the joint (bonafide and spoofed) and spoofed conditions by a maximum of 36.2% and 49.8% in terms of equal error rates, respectively.


【19】 Gaussian Adaptive Attention is All You Need: Robust Contextual  Representations Across Multiple Modalities
标题:高斯自适应注意力是您所需要的全部:跨多个通道的强健上下文表示
链接:https://arxiv.org/abs/2401.11143
作者:Georgios Ioannides,Aman Chadha,Aaron Elkins
摘要:我们提出了多头高斯自适应注意力机制(GAAM),一种新的概率注意力框架,和高斯自适应Transformer(GAT),旨在增强跨多个模态的信息聚合,包括语音,文本和视觉。GAAM将可学习的均值和方差集成到其注意力机制中,在多头框架中实现,使其能够集体建模任何概率分布,以动态重新校准特征重要性。这种方法表现出显着的改进,特别是对于高度非平稳的数据,通过识别特征空间内的关键元素,在模型性能方面超过了最先进的注意力技术(准确度高达约+20%)。GAAM与基于点积的注意力模型的兼容性和相对较低的参数数量显示了其适应性和潜力,以提高现有的注意力框架。从经验上讲,GAAM在各种任务中表现出卓越的适应性和有效性,包括语音中的情感识别,图像分类和文本分类,从而建立了处理多模态数据的鲁棒性和多功能性。此外,我们还引入了重要性因子(IF),这是一种新的基于学习的度量,可以增强使用基于GAAM的方法训练的模型的可解释性。总的来说,GAAM代表了跨多种模式开发更好的性能和更可解释的注意力模型的进步。
摘要:We propose the Multi-Head Gaussian Adaptive Attention Mechanism (GAAM), a novel probabilistic attention framework, and the Gaussian Adaptive Transformer (GAT), designed to enhance information aggregation across multiple modalities, including Speech, Text and Vision. GAAM integrates learnable mean and variance into its attention mechanism, implemented in a Multi-Headed framework enabling it to collectively model any Probability Distribution for dynamic recalibration of feature significance. This method demonstrates significant improvements, especially with highly non-stationary data, surpassing the state-of-the-art attention techniques in model performance (up to approximately +20% in accuracy) by identifying key elements within the feature space. GAAM's compatibility with dot-product-based attention models and relatively low number of parameters showcases its adaptability and potential to boost existing attention frameworks. Empirically, GAAM exhibits superior adaptability and efficacy across a diverse range of tasks, including emotion recognition in speech, image classification, and text classification, thereby establishing its robustness and versatility in handling multi-modal data. Furthermore, we introduce the Importance Factor (IF), a new learning-based metric that enhances the explainability of models trained with GAAM-based methods. Overall, GAAM represents an advancement towards development of better performing and more explainable attention models across multiple modalities.


【20】 ASM: Audio Spectrogram Mixer
标题:ASM:音频频谱混合器
链接:https://arxiv.org/abs/2401.11102
作者:Qingfeng Ji,Jicun Zhang,Yuxin Wang
摘要:Transformer结构最近在深度学习领域表现出了出色的技能,显著提高了各种领域模型的准确性。研究人员已经开始质疑这种复杂的网络结构是否真的有必要,以及由于其复杂的网络拓扑结构和高推理成本,是否可以在降低推理成本的情况下获得同样出色的结果。为了证明Mixer在三个数据集Speech Commands、UrbanSound 8 k和CASIA中文情感语料库上的有效性,本文将Mixer的精简版本应用于音频分类任务,并与基于Transformer-based Audio Spectrogram Transformer(AST)模型进行了对比实验。此外,本文还对GeLU、Mish、Swish和Mixer-C等几种激活函数在Mixer中的应用进行了对比实验。此外,本文还通过对比实验,比较了Mixer中GeLU、Mish、Swish和Mix-C等激活函数的使用情况。此外,还指出了AST模型的一些缺陷,并对本文提出的模型进行了改进。总之,在这项研究中提出了一个称为音频频谱图混合器的模型,这是第一个使用混合器进行音频分类的模型,并研究了该模型未来的改进方向。
摘要:Transformer structures have demonstrated outstanding skills in the deep learning space recently, significantly increasing the accuracy of models across a variety of domains. Researchers have started to question whether such a sophisticated network structure is actually necessary and whether equally outstanding results can be reached with reduced inference cost due to its complicated network topology and high inference cost. In order to prove the Mixer's efficacy on three datasets Speech Commands, UrbanSound8k, and CASIA Chinese Sentiment Corpus this paper applies amore condensed version of the Mixer to an audio classification task and conducts comparative experiments with the Transformer-based Audio Spectrogram Transformer (AST)model. In addition, this paper conducts comparative experiments on the application of several activation functions in Mixer, namely GeLU, Mish, Swish and Acon-C. Further-more, the use of various activation functions in Mixer, including GeLU, Mish, Swish, and Acon-C, is compared in this research through comparison experiments. Additionally, some AST model flaws are highlighted, and the model suggested in this study is improved as a result. In conclusion, a model called the Audio Spectrogram Mixer, which is the first model for audio classification with Mixer, is suggested in this study and the model's future directions for improvement are examined.

【21】 Sound Unblending: Exploring Sound Manipulations for Accessible  Mixed-Reality Awareness
标题:声音分离:探索声音操作以获得可接近的混合现实意识
链接:https://arxiv.org/abs/2401.11095
作者:Ruei-Che Chang,Chia-Sheng Hung,Bing-Yu Chen,Dhruv Jain,Anhong Guo
摘要:混合现实(MR)音景将真实世界的声音与来自听力设备的虚拟音频融合在一起,呈现出难以辨别和区分的复杂听觉信息。这对于盲人或视障人士来说尤其具有挑战性,他们在日常生活中依赖声音和描述。为了了解复杂的音频信息是如何被消费的,我们分析了盲人社区中的在线论坛帖子,确定了普遍的挑战,需求和期望的解决方案。我们综合了这些结果,并提出了用于提高MR声音意识的声音分解,其中包括六种声音操作:Ambience Builder,Feature Shifter,Earcon Generator,Prioritizer,Spatializer和Stylizer。为了评估声音分解的有效性,我们在三个模拟MR场景中与18名盲人参与者进行了一项用户研究,参与者在复杂的音景中识别出特定的声音。我们发现,声音不混合增加MR声音意识和最小化认知负荷。最后,我们开发了三个真实世界的示例应用程序来演示声音分解的实用性。
摘要:Mixed-reality (MR) soundscapes blend real-world sound with virtual audio from hearing devices, presenting intricate auditory information that is hard to discern and differentiate. This is particularly challenging for blind or visually impaired individuals, who rely on sounds and descriptions in their everyday lives. To understand how complex audio information is consumed, we analyzed online forum posts within the blind community, identifying prevailing challenges, needs, and desired solutions. We synthesized the results and proposed Sound Unblending for increasing MR sound awareness, which includes six sound manipulations: Ambience Builder, Feature Shifter, Earcon Generator, Prioritizer, Spatializer, and Stylizer. To evaluate the effectiveness of sound unblending, we conducted a user study with 18 blind participants across three simulated MR scenarios, where participants identified specific sounds within intricate soundscapes. We found that sound unblending increased MR sound awareness and minimized cognitive load. Finally, we developed three real-world example applications to demonstrate the practicality of sound unblending.

机器翻译由腾讯交互翻译提供,仅供参考