【1】Synchformer: Efficient Synchronization from Sparse Cues标题:同步形成器:稀疏线索的高效同步链接:https://arxiv.org/abs/2401.16423作者:Vladimir Iashin,Weidi Xie,Esa Rahtu,Andrew Zisserman备注:Extended version of the ICASSP 24 paper. Project page: this https URL Code: this https URL摘要:我们的目标是视听同步,重点是“在野外”的视频,如YouTube上的视频,其中同步线索可能是稀疏的。我们的贡献包括一个新的视听同步模型,并通过多模态段级对比预训练,从同步建模中提取特征。这种方法在密集和稀疏设置中都实现了最先进的性能。我们还将同步模型训练扩展到AudioSet百万级“野外”数据集,研究证据归因技术的可解释性,并探索同步模型的新功能:视听同步性。摘要:Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that decouples feature extraction from synchronization modelling through multi-modal segment-level contrastive pre-training. This approach achieves state-of-the-art performance in both dense and sparse settings. We also extend synchronization model training to AudioSet a million-scale 'in-the-wild' dataset, investigate evidence attribution techniques for interpretability, and explore a new capability for synchronization models: audio-visual synchronizability. 【2】 Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings标题:连续目标语音提取:增强复杂录音的个性化二值化和提取链接:https://arxiv.org/abs/2401.15993作者:He Zhao,Hangting Chen,Jianwei Yu,Yuehai Wang备注:8 pages, 6 figures摘要:目标说话人提取(TSE)的目的是从输入的混合语音中提取出目标说话人的语音。以前的研究集中在高度重叠的场景。然而,现实世界的应用通常会遇到更复杂的场景,如可变扬声器重叠和目标扬声器缺席。在本文中,我们介绍了一个框架来执行连续TSE(C-TSE),包括目标说话人语音激活检测(TSVAD)和TSE模型。该框架显着提高了TSE性能相似的扬声器和增强个性化,这是缺乏传统的日记方法。详细地说,不同于传统的TSVAD部署细化的日记结果,建议的注意目标说话人语音激活检测(A-TSVAD)直接生成目标说话人的时间戳。通过对级联和并行方法的比较,探讨了A-TSVAD和TSE的不同集成方法。该框架的有效性评估使用一系列的指标,包括日记和增强指标。我们的实验表明,A-TSVAD优于传统的方法,在减少日记错误。此外,以顺序级联方式集成A-TSVAD和TSE进一步提高了提取精度。摘要:Target speaker extraction (TSE) aims to extract the target speaker's voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-world applications usually meet more complex scenarios like variable speaker overlapping and target speaker absence. In this paper, we introduces a framework to perform continuous TSE (C-TSE), comprising a target speaker voice activation detection (TSVAD) and a TSE model. This framework significantly improves TSE performance on similar speakers and enhances personalization, which is lacking in traditional diarization methods. In detail, unlike conventional TSVAD deployed to refine the diarization results, the proposed Attention-target speaker voice activation detection (A-TSVAD) directly generates timestamps of the target speaker. We also explore some different integration methods of A-TSVAD and TSE by comparing the cascaded and parallel methods. The framework's effectiveness is assessed using a range of metrics, including diarization and enhancement metrics. Our experiments demonstrate that A-TSVAD outperforms conventional methods in reducing diarization errors. Furthermore, the integration of A-TSVAD and TSE in a sequential cascaded manner further enhances extraction accuracy. 【3】 Masked Audio Modeling with CLAP and Multi-Objective Learning标题:基于CLAP和多目标学习的掩蔽音频建模链接:https://arxiv.org/abs/2401.15953作者:Yifei Xin,Xiulian Peng,Yan Lu备注:Accepted by Interspeech2023摘要:大多数现有的掩蔽音频建模(MAM)方法通过掩蔽和重建局部频谱图补丁来学习音频表示。然而,重建损失主要占重建的频谱图的信号级质量,并且在提取高级音频语义方面仍然受到限制。在本文中,我们建议通过从掩蔽和非掩蔽区域的对比语言音频预训练(CLAP)表示(MAM-CLAP)中提取跨模态知识,并利用具有监督分类分支(SupMAM)的多目标学习策略,从而为MAM提供更多的语义知识,使其能够有效地从标签中学习全局特征,从而增强MAM的语义建模。实验表明,我们的方法显着提高多个下游任务的性能。此外,通过将我们的MAM-CLAP与SupMAM相结合,我们可以在各种音频和语音分类任务上获得新的最先进的结果,超过以前的自监督学习和监督预训练方法。摘要:Most existing masked audio modeling (MAM) methods learn audio representations by masking and reconstructing local spectrogram patches. However, the reconstruction loss mainly accounts for the signal-level quality of the reconstructed spectrogram and is still limited in extracting high-level audio semantics. In this paper, we propose to enhance the semantic modeling of MAM by distilling cross-modality knowledge from contrastive language-audio pretraining (CLAP) representations for both masked and unmasked regions (MAM-CLAP) and leveraging a multi-objective learning strategy with a supervised classification branch (SupMAM), thereby providing more semantic knowledge for MAM and enabling it to effectively learn global features from labels. Experiments show that our methods significantly improve the performance on multiple downstream tasks. Furthermore, by combining our MAM-CLAP with SupMAM, we can achieve new state-of-the-art results on various audio and speech classification tasks, exceeding previous self-supervised learning and supervised pretraining methods. 【4】 Phoneme-Based Proactive Anti-Eavesdropping with Controlled Recording Privilege标题:基于音素的可控录音权限主动反窃听链接:https://arxiv.org/abs/2401.15704作者:Peng Huang,Yao Wei,Peng Cheng,Zhongjie Ba,Li Lu,Feng Lin,Yang Wang,Kui Ren备注:14 pages, 28 figures; submitted to IEEE TDSC摘要:随着智能设备的普及,人们越来越担心被窃听的问题。为了增强语音隐私,最近的研究利用麦克风的非线性特性,用听不见的超声波干扰录音机。然而,现有的解决方案仅仅依赖于能量掩蔽。它们的简单形式的噪声导致了几个问题,例如高能量要求和容易通过语音增强技术去除。此外,这些解决方案中的大多数不支持授权录制,这限制了它们的使用场景。在本文中,我们设计了一个高效而强大的系统,可以干扰麦克风,同时保留授权录音。具体来说,我们提出了一种新的基于音素的噪声与信息掩蔽的想法,它可以分散机器和人类的注意力,并抵抗去噪技术。此外,我们优化了噪声传输策略,以获得更广泛的覆盖范围,并实现了我们的系统的硬件原型。实验结果表明,在所有测试的语音识别系统下,该系统都能将录音的识别准确率降低到50%以下,这比现有的解决方案要好得多。摘要:The widespread smart devices raise people's concerns of being eavesdropped on. To enhance voice privacy, recent studies exploit the nonlinearity in microphone to jam audio recorders with inaudible ultrasound. However, existing solutions solely rely on energetic masking. Their simple-form noise leads to several problems, such as high energy requirements and being easily removed by speech enhancement techniques. Besides, most of these solutions do not support authorized recording, which restricts their usage scenarios. In this paper, we design an efficient yet robust system that can jam microphones while preserving authorized recording. Specifically, we propose a novel phoneme-based noise with the idea of informational masking, which can distract both machines and humans and is resistant to denoising techniques. Besides, we optimize the noise transmission strategy for broader coverage and implement a hardware prototype of our system. Experimental results show that our system can reduce the recognition accuracy of recordings to below 50\% under all tested speech recognition systems, which is much better than existing solutions.
【5】 Evaluating Echo State Network for Parkinson's Disease Prediction using Voice Features标题:利用语音特征评价回声状态网络预测帕金森病链接:https://arxiv.org/abs/2401.15672作者:Seyedeh Zahra Seyedi Hosseininian,Ahmadreza Tajari,Mohsen Ghalehnoie,Alireza Alfi摘要:帕金森氏病(PD)是一种使人衰弱的神经系统疾病,需要精确和早期诊断,以便有效地护理患者。本研究旨在开发一种诊断模型,能够实现高准确性和最小化假阴性,这是临床实践中的一个关键因素。在有限的训练数据下,采用方差分析的特征选择策略来识别信息量最大的特征。随后,采用了各种机器学习方法,包括回声状态网络(ESN),随机森林,k-最近邻,支持向量分类器,极端梯度提升和决策树,并进行了全面的评估。结果的统计分析突出了ESN的卓越性能,不仅显示了优越的准确性,而且在所有方法中假阴性率最低。同样,统计数据表明,ESN方法在83%的情况下始终保持低于8%的假阴性率。ESN在诊断精度和最小化错误分类之间取得微妙平衡的能力使其成为PD诊断的示例性选择,特别是在数据有限的情况下。这项研究标志着朝着更有效和可靠的PD诊断迈出了重要一步,对改善患者预后和医疗动态具有潜在意义。摘要:Parkinson's disease (PD) is a debilitating neurological disorder that necessitates precise and early diagnosis for effective patient care. This study aims to develop a diagnostic model capable of achieving both high accuracy and minimizing false negatives, a critical factor in clinical practice. Given the limited training data, a feature selection strategy utilizing ANOVA is employed to identify the most informative features. Subsequently, various machine learning methods, including Echo State Networks (ESN), Random Forest, k-nearest Neighbors, Support Vector Classifier, Extreme Gradient Boosting, and Decision Tree, are employed and thoroughly evaluated. The statistical analyses of the results highlight ESN's exceptional performance, showcasing not only superior accuracy but also the lowest false negative rate among all methods. Consistently, statistical data indicates that the ESN method consistently maintains a false negative rate of less than 8% in 83% of cases. ESN's capacity to strike a delicate balance between diagnostic precision and minimizing misclassifications positions it as an exemplary choice for PD diagnosis, especially in scenarios characterized by limited data. This research marks a significant step towards more efficient and reliable PD diagnosis, with potential implications for enhanced patient outcomes and healthcare dynamics.
【6】 MunTTS: A Text-to-Speech System for Mundari标题:MunTTS:一个面向Mundari的文语转换系统链接:https://arxiv.org/abs/2401.15579作者:Varun Gumma,Rishav Hada,Aditya Yadavalli,Pamir Gogoi,Ishani Mondal,Vivek Seshadri,Kalika Bali备注:Accepted to ComputEL-7摘要:我们提出了MunTTS,一个端到端的文本到语音(TTS)系统专门为蒙达里,低资源的印度语言的Austo-Asiatic家庭。我们的工作是通过收集和处理数据来建立一个语音合成系统,从而解决语言技术中代表性不足的语言的差距。我们通过收集大量的Mundari文本和语音数据集并训练端到端语音模型来开始我们的研究。我们还深入研究了用于训练模型的方法,以确保它们在数据限制的情况下仍然有效。我们评估我们的系统与母语和客观的指标,展示其作为一种工具,在数字时代保存和促进蒙达里语言的潜力。摘要:We present MunTTS, an end-to-end text-to-speech (TTS) system specifically for Mundari, a low-resource Indian language of the Austo-Asiatic family. Our work addresses the gap in linguistic technology for underrepresented languages by collecting and processing data to build a speech synthesis system. We begin our study by gathering a substantial dataset of Mundari text and speech and train end-to-end speech models. We also delve into the methods used for training our models, ensuring they are efficient and effective despite the data constraints. We evaluate our system with native speakers and objective metrics, demonstrating its potential as a tool for preserving and promoting the Mundari language in the digital age. 【7】 Byte Pair Encoding Is All You Need For Automatic Bengali Speech Recognition标题:字节对编码是孟加拉自动语音识别所需的全部功能链接:https://arxiv.org/abs/2401.15532作者:Ahnaf Mozib Samin备注:Under-review摘要:字节对编码(BPE)作为一种有效的标记化方法出现,用于解决各种自然语言和语音处理任务中的词汇外(OOV)挑战。最近的研究强调了BPE子词标记化的功效对语言的形态性质的依赖性,特别是在富于曲折形态的语言中,较少的BPE合并足以生成高生产力的标记。受此启发,我们的研究经验确定了孟加拉语的最佳BPE令牌数量,孟加拉语是一种以其形态复杂性而闻名的语言,从而提高了分布外自动语音识别(ASR)性能。实验评估表明,过高数量的BPE令牌可能导致过拟合,而大约500-1000个令牌会导致更好的OOV性能。此外,我们进行了比较分析的BPE与基于字符和基于单字的标记化方法。通过将BPE标记化引入孟加拉语ASR,我们实现了字错误率(WER)的大幅降低,从基于字符的基线系统中的66.44%降低到LB-ASRTD评估集的63.80%,从SHRUTI评估集的46.34%降低到42.80%,这两个都包括分布外数据。摘要:Byte pair encoding (BPE) emerges as an effective tokenization method for tackling the out-of-vocabulary (OOV) challenge in various natural language and speech processing tasks. Recent research highlights the dependency of BPE subword tokenization's efficacy on the morphological nature of the language, particularly in languages rich in inflectional morphology, where fewer BPE merges suffice for generating highly productive tokens. Motivated by this, our study empirically identifies the optimal number of BPE tokens for Bengali, a language known for its morphological complexity, thus enhancing out-of-distribution automatic speech recognition (ASR) performance. Experimental evaluation reveals that an excessively high number of BPE tokens can lead to overfitting, while approximately 500-1000 tokens result in superior OOV performance. Furthermore, we conduct a comparative analysis of BPE with character-based and unigram-based tokenization methods. By introducing BPE tokenization to Bengali ASR, we achieve a substantial reduction in the word error rate (WER) from 66.44% in our character-based baseline system to 63.80% on the LB-ASRTD eval set and from 46.34% to 42.80% on the SHRUTI eval set, both of which include out-of-distribution data.
【8】 Validation of artificial neural networks to model the acoustic behaviour of induction motors标题:人工神经网络在感应电机声学特性建模中的验证链接:https://arxiv.org/abs/2401.15377作者:F. J. Jimenez-Romero,D. Guijo-Rubio,F. R. Lara-Raya,A. Ruiz-Gonzalez,C. Hervas-Martinez摘要:近十年来,感应电动机的声品质一直是研究的热点。特别是,由于其大量的应用,人们暴露于由噪声排放引起的身体和心理不适。因此,有必要尽量减少其对人口的心理影响。以这种方式,这项工作的主要目标是评估使用多任务人工神经网络作为建模技术,同时预测心理声学参数的感应电机。使用几个输入,例如电机功率信号的电幅度和极数,而不是将电动机的噪声与环境噪声分离。提出了两种不同类型的人工神经网络,以等效声压、响度、粗糙度和尖锐度作为输出,对感应电机的声品质进行评价。具体而言,两种不同的拓扑结构被认为是:简单的模型和更复杂的模型。前者更具有可解释性,而后者以隐藏因果关系为代价导致更高的准确性。专注于简单的可解释的模型,产品单元神经网络取得了最好的结果:MSE和SEP。这种产品单元模型的主要好处是它的简单性,因为只有10个输入变量,概述了多任务人工神经网络的有效传输机制,以提取多个任务的共同特征。最后,利用最优积单元神经网络对感应电机的声学质量进行了深入的分析。摘要:In the last decade, the sound quality of electric induction motors is a hot topic in the research field. Specially, due to its high number of applications, the population is exposed to physical and psychological discomfort caused by the noise emission. Therefore, it is necessary to minimise its psychological impact on the population. In this way, the main goal of this work is to evaluate the use of multitask artificial neural networks as a modelling technique for simultaneously predicting psychoacoustic parameters of induction motors. Several inputs are used, such as, the electrical magnitudes of the motor power signal and the number of poles, instead of separating the noise of the electric motor from the environmental noise. Two different kind of artificial neural networks are proposed to evaluate the acoustic quality of induction motors, by using the equivalent sound pressure, the loudness, the roughness and the sharpness as outputs. Concretely, two different topologies have been considered: simple models and more complex models. The former are more interpretable, while the later lead to higher accuracy at the cost of hiding the cause-effect relationship. Focusing on the simple interpretable models, product unit neural networks achieved the best results: for MSE and for SEP. The main benefit of this product unit model is its simplicity, since only 10 inputs variables are used, outlining the effective transfer mechanism of multitask artificial neural networks to extract common features of multiple tasks. Finally, a deep analysis of the acoustic quality of induction motors in done using the best product unit neural networks.
【9】 Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training标题:通过领域对抗性训练学习的具有稳健音乐表示的音乐自动标记链接:https://arxiv.org/abs/2401.15323作者:Haesun Joung,Kyogu Lee备注:5 pages, 3 figures, accepted to ICASSP 2024摘要:音乐自动标记对于增强音乐发现和推荐至关重要。音乐信息检索(MIR)中的现有模型与多媒体内容中的环境和语音声音等现实世界的噪声作斗争。本研究提出了一种受语音相关任务启发的方法,以提高噪声环境下的音乐自动标记性能。该方法将域对抗训练(DAT)集成到音乐域中,从而实现了能够承受噪声的强大音乐表示。与以前的研究不同,这种方法涉及一个额外的预训练阶段的领域分类器,以避免性能下降,在随后的阶段。添加各种合成的嘈杂音乐数据可以提高模型在不同噪声水平下的泛化能力。该架构通过有效利用未标记的嘈杂音乐数据,展示了增强的音乐自动标记性能。补充未标记数据的额外实验进一步提高了模型的性能,强调了其强大的泛化能力和广泛的适用性。摘要:Music auto-tagging is crucial for enhancing music discovery and recommendation. Existing models in Music Information Retrieval (MIR) struggle with real-world noise such as environmental and speech sounds in multimedia content. This study proposes a method inspired by speech-related tasks to enhance music auto-tagging performance in noisy settings. The approach integrates Domain Adversarial Training (DAT) into the music domain, enabling robust music representations that withstand noise. Unlike previous research, this approach involves an additional pretraining phase for the domain classifier, to avoid performance degradation in the subsequent phase. Adding various synthesized noisy music data improves the model's generalization across different noise levels. The proposed architecture demonstrates enhanced performance in music auto-tagging by effectively utilizing unlabeled noisy music data. Additional experiments with supplementary unlabeled data further improves the model's performance, underscoring its robust generalization capabilities and broad applicability. 【10】 AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations标题:AMUSE:群话中说话人情感识别的自适应多模式分析链接:https://arxiv.org/abs/2401.15164作者:Naresh Kumar Devulapally,Sidharth Anand,Sreyasee Das Bhattacharjee,Junsong Yuan,Yu-Ping Chang摘要:在群体对话中分析个体情绪对于开发能够进行自然人机交互的智能代理至关重要。虽然可靠的情感识别技术依赖于不同的模态(文本,音频,视频),这些模态之间的固有异质性和受个体独特行为模式影响的动态跨模态交互,使得情感识别的任务非常具有挑战性。这种困难在群体环境中更加复杂,在群体环境中,情绪及其时间演变不仅受到个体的影响,而且还受到外部环境的影响,如观众的反应和正在进行的对话的背景。为了应对这一挑战,我们提出了一个多模态注意力网络,它通过共同学习其交互式的特定于模式的外围网络和中央网络,在不同的空间抽象层次上捕获跨模态的交互。建议的MAN注入跨模态的注意力,通过其外围键值对内的每一层的模式特定的中央查询网络。然后,使用自适应融合技术将所得到的交叉参与的模式特定描述符组合,该技术使模型能够在实例特定的多模式描述符内集成有区别的和互补的模式特定数据模式。给定由一系列话语表示的对话,所提出的AMuSE模型将空间和时间特征压缩成两个密集的描述符:说话者级别和话语级别。这不仅有助于在大规模公共数据集中提供更好的分类性能(加权F1提高3-5%,准确性提高5-7%),还有助于用户通过其多模态可解释性可视化模块理解模型所做的每个情感预测背后的推理。摘要:Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio, video), the inherent heterogeneity between these modalities and the dynamic cross-modal interactions influenced by an individual's unique behavioral patterns make the task of emotion recognition very challenging. This difficulty is compounded in group settings, where the emotion and its temporal evolution are not only influenced by the individual but also by external contexts like audience reaction and context of the ongoing conversation. To meet this challenge, we propose a Multimodal Attention Network that captures cross-modal interactions at various levels of spatial abstraction by jointly learning its interactive bunch of mode-specific Peripheral and Central networks. The proposed MAN injects cross-modal attention via its Peripheral key-value pairs within each layer of a mode-specific Central query network. The resulting cross-attended mode-specific descriptors are then combined using an Adaptive Fusion technique that enables the model to integrate the discriminative and complementary mode-specific data patterns within an instance-specific multimodal descriptor. Given a dialogue represented by a sequence of utterances, the proposed AMuSE model condenses both spatial and temporal features into two dense descriptors: speaker-level and utterance-level. This helps not only in delivering better classification performance (3-5% improvement in Weighted-F1 and 5-7% improvement in Accuracy) in large-scale public datasets but also helps the users in understanding the reasoning behind each emotion prediction made by the model via its Multimodal Explainability Visualization module.
【11】 Generalisations of Euler's Tonnetz on triangulated surfaces标题:三角曲面上Euler‘s Tonnetz的推广链接:https://arxiv.org/abs/2401.15692作者:Konstanze Rietsch备注:17 pages, 1 table, 12 figures摘要:我们给出了一个定义,我们称之为一个'tonnetz'的三角曲面,概括著名的tonnetz欧拉从1739年。在欧拉的音圈中,平面的规则“A_2 $三角形”的顶点用音符或音高类标记。在我们的推广,我们允许更一般的三角曲面的标号。特别是,边缘标注结果导致了一组丰富的例子。我们构建自然的例子,是有关晶体反射群和生活在三角环。在这些基础上,我们观察到数学朗兰兹对偶和大/小对偶之间的一种奇怪的关系。我们还构建'异国情调的'类型-$A_2$的例子(不同于欧拉的Tonnetz),和一个tonnetz上的球体,编码所有主要的第九和弦。摘要:We give a definition of a what we call a `tonnetz' on a triangulated surface, generalising the famous tonnetz of Euler from 1739. In Euler's tonnetz the vertices of a regular `$A_2$ triangulation' of the plane are labelled with notes, or pitch-classes. In our generalisation we allow much more general labellings of triangulated surfaces. In particular, edge labellings turn out to lead to a rich set of examples. We construct natural examples that are related to crystallographic reflection groups and live on triangulations of tori. Underlying these we observe a curious relationship between mathematical Langlands duality and major/minor duality. We also construct `exotic' type-$A_2$ examples (different from Euler's Tonnetz), and a tonnetz on a sphere that encodes all major ninth chords.
【12】 On Speaker Attribution with SURT标题:论说话人的归因与苏尔特链接:https://arxiv.org/abs/2401.15676作者:Desh Raj,Matthew Wiesner,Matthew Maciejewski,Leibny Paola Garcia-Perera,Daniel Povey,Sanjeev Khudanpur备注:8 pages, 6 figures, 6 tables. Submitted to Odyssey 2024摘要:流分解和识别转换器(SURT)最近已经成为一个流行的框架,连续,流,多说话人语音识别(ASR)。随着体系结构,目标和混合模拟方法的进步,它被证明是一个有效的流媒体方法,为扬声器不可知转录的真实会议。在这项工作中,我们进一步推动这个框架,提出了方法来执行扬声器归因的转录与SURT,为短的混合和长的录音。我们通过向SURT添加辅助扬声器分支来实现这一点,并通过HAT风格的空白因子分解将其标签预测与ASR令牌预测同步。为了确保录音中不同话语组的相对扬声器标签的一致性,我们提出了“扬声器前缀”-在每个块后面添加前一个块中识别的扬声器的高置信度帧,以建立相对顺序。我们进行了广泛的消融实验合成LibriSpeech混合物,以验证我们的设计选择,并证明我们的最终模型对AMI语料库的有效性。摘要:The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in architecture, objectives, and mixture simulation methods, it was demonstrated that SURT can be an efficient streaming method for speaker-agnostic transcription of real meetings. In this work, we push this framework further by proposing methods to perform speaker-attributed transcription with SURT, for both short mixtures and long recordings. We achieve this by adding an auxiliary speaker branch to SURT, and synchronizing its label prediction with ASR token prediction through HAT-style blank factorization. In order to ensure consistency in relative speaker labels across different utterance groups in a recording, we propose "speaker prefixing" -- appending each chunk with high-confidence frames of speakers identified in previous chunks, to establish the relative order. We perform extensive ablation experiments on synthetic LibriSpeech mixtures to validate our design choices, and demonstrate the efficacy of our final model on the AMI corpus. eess.AS音频处理【1】 On Speaker Attribution with SURT标题:论说话人的归因与苏尔特链接:https://arxiv.org/abs/2401.15676作者:Desh Raj,Matthew Wiesner,Matthew Maciejewski,Leibny Paola Garcia-Perera,Daniel Povey,Sanjeev Khudanpur备注:8 pages, 6 figures, 6 tables. Submitted to Odyssey 2024摘要:流分解和识别转换器(SURT)最近已经成为一个流行的框架,连续,流,多说话人语音识别(ASR)。随着体系结构,目标和混合模拟方法的进步,它被证明是一个有效的流媒体方法,为扬声器不可知转录的真实会议。在这项工作中,我们进一步推动这个框架,提出了方法来执行扬声器归因的转录与SURT,为短的混合和长的录音。我们通过向SURT添加辅助扬声器分支来实现这一点,并通过HAT风格的空白因子分解将其标签预测与ASR令牌预测同步。为了确保录音中不同话语组的相对扬声器标签的一致性,我们提出了“扬声器前缀”-在每个块后面添加前一个块中识别的扬声器的高置信度帧,以建立相对顺序。我们进行了广泛的消融实验合成LibriSpeech混合物,以验证我们的设计选择,并证明我们的最终模型对AMI语料库的有效性。摘要:The Streaming Unmixing and Recognition Transducer (SURT) has recently become a popular framework for continuous, streaming, multi-talker speech recognition (ASR). With advances in architecture, objectives, and mixture simulation methods, it was demonstrated that SURT can be an efficient streaming method for speaker-agnostic transcription of real meetings. In this work, we push this framework further by proposing methods to perform speaker-attributed transcription with SURT, for both short mixtures and long recordings. We achieve this by adding an auxiliary speaker branch to SURT, and synchronizing its label prediction with ASR token prediction through HAT-style blank factorization. In order to ensure consistency in relative speaker labels across different utterance groups in a recording, we propose "speaker prefixing" -- appending each chunk with high-confidence frames of speakers identified in previous chunks, to establish the relative order. We perform extensive ablation experiments on synthetic LibriSpeech mixtures to validate our design choices, and demonstrate the efficacy of our final model on the AMI corpus.
【2】 Synchformer: Efficient Synchronization from Sparse Cues标题:同步形成器:稀疏线索的高效同步链接:https://arxiv.org/abs/2401.16423作者:Vladimir Iashin,Weidi Xie,Esa Rahtu,Andrew Zisserman备注:Extended version of the ICASSP 24 paper. Project page: this https URL Code: this https URL摘要:我们的目标是视听同步,重点是“在野外”的视频,如YouTube上的视频,其中同步线索可能是稀疏的。我们的贡献包括一个新的视听同步模型,并通过多模态段级对比预训练,从同步建模中提取特征。这种方法在密集和稀疏设置中都实现了最先进的性能。我们还将同步模型训练扩展到AudioSet百万级“野外”数据集,研究证据归因技术的可解释性,并探索同步模型的新功能:视听同步性。摘要:Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model, and training that decouples feature extraction from synchronization modelling through multi-modal segment-level contrastive pre-training. This approach achieves state-of-the-art performance in both dense and sparse settings. We also extend synchronization model training to AudioSet a million-scale 'in-the-wild' dataset, investigate evidence attribution techniques for interpretability, and explore a new capability for synchronization models: audio-visual synchronizability.
【3】 Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings标题:连续目标语音提取:增强复杂录音的个性化二值化和提取链接:https://arxiv.org/abs/2401.15993作者:He Zhao,Hangting Chen,Jianwei Yu,Yuehai Wang备注:8 pages, 6 figures摘要:目标说话人提取(TSE)的目的是从输入的混合语音中提取出目标说话人的语音。以前的研究集中在高度重叠的场景。然而,现实世界的应用通常会遇到更复杂的场景,如可变扬声器重叠和目标扬声器缺席。在本文中,我们介绍了一个框架来执行连续TSE(C-TSE),包括目标说话人语音激活检测(TSVAD)和TSE模型。该框架显着提高了TSE性能相似的扬声器和增强个性化,这是缺乏传统的日记方法。详细地说,不同于传统的TSVAD部署细化的日记结果,建议的注意目标说话人语音激活检测(A-TSVAD)直接生成目标说话人的时间戳。通过对级联和并行方法的比较,探讨了A-TSVAD和TSE的不同集成方法。该框架的有效性评估使用一系列的指标,包括日记和增强指标。我们的实验表明,A-TSVAD优于传统的方法,在减少日记错误。此外,以顺序级联方式集成A-TSVAD和TSE进一步提高了提取精度。摘要:Target speaker extraction (TSE) aims to extract the target speaker's voice from the input mixture. Previous studies have concentrated on high-overlapping scenarios. However, real-world applications usually meet more complex scenarios like variable speaker overlapping and target speaker absence. In this paper, we introduces a framework to perform continuous TSE (C-TSE), comprising a target speaker voice activation detection (TSVAD) and a TSE model. This framework significantly improves TSE performance on similar speakers and enhances personalization, which is lacking in traditional diarization methods. In detail, unlike conventional TSVAD deployed to refine the diarization results, the proposed Attention-target speaker voice activation detection (A-TSVAD) directly generates timestamps of the target speaker. We also explore some different integration methods of A-TSVAD and TSE by comparing the cascaded and parallel methods. The framework's effectiveness is assessed using a range of metrics, including diarization and enhancement metrics. Our experiments demonstrate that A-TSVAD outperforms conventional methods in reducing diarization errors. Furthermore, the integration of A-TSVAD and TSE in a sequential cascaded manner further enhances extraction accuracy.
【4】 Masked Audio Modeling with CLAP and Multi-Objective Learning标题:基于CLAP和多目标学习的掩蔽音频建模链接:https://arxiv.org/abs/2401.15953作者:Yifei Xin,Xiulian Peng,Yan Lu备注:Accepted by Interspeech2023摘要:大多数现有的掩蔽音频建模(MAM)方法通过掩蔽和重建局部频谱图补丁来学习音频表示。然而,重建损失主要占重建的频谱图的信号级质量,并且在提取高级音频语义方面仍然受到限制。在本文中,我们建议通过从掩蔽和非掩蔽区域的对比语言音频预训练(CLAP)表示(MAM-CLAP)中提取跨模态知识,并利用具有监督分类分支(SupMAM)的多目标学习策略,从而为MAM提供更多的语义知识,使其能够有效地从标签中学习全局特征,从而增强MAM的语义建模。实验表明,我们的方法显着提高多个下游任务的性能。此外,通过将我们的MAM-CLAP与SupMAM相结合,我们可以在各种音频和语音分类任务上获得新的最先进的结果,超过以前的自监督学习和监督预训练方法。摘要:Most existing masked audio modeling (MAM) methods learn audio representations by masking and reconstructing local spectrogram patches. However, the reconstruction loss mainly accounts for the signal-level quality of the reconstructed spectrogram and is still limited in extracting high-level audio semantics. In this paper, we propose to enhance the semantic modeling of MAM by distilling cross-modality knowledge from contrastive language-audio pretraining (CLAP) representations for both masked and unmasked regions (MAM-CLAP) and leveraging a multi-objective learning strategy with a supervised classification branch (SupMAM), thereby providing more semantic knowledge for MAM and enabling it to effectively learn global features from labels. Experiments show that our methods significantly improve the performance on multiple downstream tasks. Furthermore, by combining our MAM-CLAP with SupMAM, we can achieve new state-of-the-art results on various audio and speech classification tasks, exceeding previous self-supervised learning and supervised pretraining methods. 【5】 Phoneme-Based Proactive Anti-Eavesdropping with Controlled Recording Privilege标题:基于音素的可控录音权限主动反窃听链接:https://arxiv.org/abs/2401.15704作者:Peng Huang,Yao Wei,Peng Cheng,Zhongjie Ba,Li Lu,Feng Lin,Yang Wang,Kui Ren备注:14 pages, 28 figures; submitted to IEEE TDSC摘要:随着智能设备的普及,人们越来越担心被窃听的问题。为了增强语音隐私,最近的研究利用麦克风的非线性特性,用听不见的超声波干扰录音机。然而,现有的解决方案仅仅依赖于能量掩蔽。它们的简单形式的噪声导致了几个问题,例如高能量要求和容易通过语音增强技术去除。此外,这些解决方案中的大多数不支持授权录制,这限制了它们的使用场景。在本文中,我们设计了一个高效而强大的系统,可以干扰麦克风,同时保留授权录音。具体来说,我们提出了一种新的基于音素的噪声与信息掩蔽的想法,它可以分散机器和人类的注意力,并抵抗去噪技术。此外,我们优化了噪声传输策略,以获得更广泛的覆盖范围,并实现了我们的系统的硬件原型。实验结果表明,在所有测试的语音识别系统下,该系统都能将录音的识别准确率降低到50%以下,这比现有的解决方案要好得多。摘要:The widespread smart devices raise people's concerns of being eavesdropped on. To enhance voice privacy, recent studies exploit the nonlinearity in microphone to jam audio recorders with inaudible ultrasound. However, existing solutions solely rely on energetic masking. Their simple-form noise leads to several problems, such as high energy requirements and being easily removed by speech enhancement techniques. Besides, most of these solutions do not support authorized recording, which restricts their usage scenarios. In this paper, we design an efficient yet robust system that can jam microphones while preserving authorized recording. Specifically, we propose a novel phoneme-based noise with the idea of informational masking, which can distract both machines and humans and is resistant to denoising techniques. Besides, we optimize the noise transmission strategy for broader coverage and implement a hardware prototype of our system. Experimental results show that our system can reduce the recognition accuracy of recordings to below 50\% under all tested speech recognition systems, which is much better than existing solutions. 【6】 Evaluating Echo State Network for Parkinson's Disease Prediction using Voice Features标题:利用语音特征评价回声状态网络预测帕金森病链接:https://arxiv.org/abs/2401.15672作者:Seyedeh Zahra Seyedi Hosseininian,Ahmadreza Tajari,Mohsen Ghalehnoie,Alireza Alfi摘要:帕金森氏病(PD)是一种使人衰弱的神经系统疾病,需要精确和早期诊断,以便有效地护理患者。本研究旨在开发一种诊断模型,能够实现高准确性和最小化假阴性,这是临床实践中的一个关键因素。在有限的训练数据下,采用方差分析的特征选择策略来识别信息量最大的特征。随后,采用了各种机器学习方法,包括回声状态网络(ESN),随机森林,k-最近邻,支持向量分类器,极端梯度提升和决策树,并进行了全面的评估。结果的统计分析突出了ESN的卓越性能,不仅显示了优越的准确性,而且在所有方法中假阴性率最低。同样,统计数据表明,ESN方法在83%的情况下始终保持低于8%的假阴性率。ESN在诊断精度和最小化错误分类之间取得微妙平衡的能力使其成为PD诊断的示例性选择,特别是在数据有限的情况下。这项研究标志着朝着更有效和可靠的PD诊断迈出了重要一步,对改善患者预后和医疗动态具有潜在意义。摘要:Parkinson's disease (PD) is a debilitating neurological disorder that necessitates precise and early diagnosis for effective patient care. This study aims to develop a diagnostic model capable of achieving both high accuracy and minimizing false negatives, a critical factor in clinical practice. Given the limited training data, a feature selection strategy utilizing ANOVA is employed to identify the most informative features. Subsequently, various machine learning methods, including Echo State Networks (ESN), Random Forest, k-nearest Neighbors, Support Vector Classifier, Extreme Gradient Boosting, and Decision Tree, are employed and thoroughly evaluated. The statistical analyses of the results highlight ESN's exceptional performance, showcasing not only superior accuracy but also the lowest false negative rate among all methods. Consistently, statistical data indicates that the ESN method consistently maintains a false negative rate of less than 8% in 83% of cases. ESN's capacity to strike a delicate balance between diagnostic precision and minimizing misclassifications positions it as an exemplary choice for PD diagnosis, especially in scenarios characterized by limited data. This research marks a significant step towards more efficient and reliable PD diagnosis, with potential implications for enhanced patient outcomes and healthcare dynamics.
【7】 MunTTS: A Text-to-Speech System for Mundari标题:MunTTS:一个面向Mundari的文语转换系统链接:https://arxiv.org/abs/2401.15579作者:Varun Gumma,Rishav Hada,Aditya Yadavalli,Pamir Gogoi,Ishani Mondal,Vivek Seshadri,Kalika Bali备注:Accepted to ComputEL-7摘要:我们提出了MunTTS,一个端到端的文本到语音(TTS)系统专门为蒙达里,低资源的印度语言的Austo-Asiatic家庭。我们的工作是通过收集和处理数据来建立一个语音合成系统,从而解决语言技术中代表性不足的语言的差距。我们通过收集大量的Mundari文本和语音数据集并训练端到端语音模型来开始我们的研究。我们还深入研究了用于训练模型的方法,以确保它们在数据限制的情况下仍然有效。我们评估我们的系统与母语和客观的指标,展示其作为一种工具,在数字时代保存和促进蒙达里语言的潜力。摘要:We present MunTTS, an end-to-end text-to-speech (TTS) system specifically for Mundari, a low-resource Indian language of the Austo-Asiatic family. Our work addresses the gap in linguistic technology for underrepresented languages by collecting and processing data to build a speech synthesis system. We begin our study by gathering a substantial dataset of Mundari text and speech and train end-to-end speech models. We also delve into the methods used for training our models, ensuring they are efficient and effective despite the data constraints. We evaluate our system with native speakers and objective metrics, demonstrating its potential as a tool for preserving and promoting the Mundari language in the digital age. 【8】 Byte Pair Encoding Is All You Need For Automatic Bengali Speech Recognition标题:字节对编码是孟加拉自动语音识别所需的全部功能链接:https://arxiv.org/abs/2401.15532作者:Ahnaf Mozib Samin备注:Under-review摘要:字节对编码(BPE)作为一种有效的标记化方法出现,用于解决各种自然语言和语音处理任务中的词汇外(OOV)挑战。最近的研究强调了BPE子词标记化的功效对语言的形态性质的依赖性,特别是在富于曲折形态的语言中,较少的BPE合并足以生成高生产力的标记。受此启发,我们的研究经验确定了孟加拉语的最佳BPE令牌数量,孟加拉语是一种以其形态复杂性而闻名的语言,从而提高了分布外自动语音识别(ASR)性能。实验评估表明,过高数量的BPE令牌可能导致过拟合,而大约500-1000个令牌会导致更好的OOV性能。此外,我们进行了比较分析的BPE与基于字符和基于单字的标记化方法。通过将BPE标记化引入孟加拉语ASR,我们实现了字错误率(WER)的大幅降低,从基于字符的基线系统中的66.44%降低到LB-ASRTD评估集的63.80%,从SHRUTI评估集的46.34%降低到42.80%,这两个都包括分布外数据。摘要:Byte pair encoding (BPE) emerges as an effective tokenization method for tackling the out-of-vocabulary (OOV) challenge in various natural language and speech processing tasks. Recent research highlights the dependency of BPE subword tokenization's efficacy on the morphological nature of the language, particularly in languages rich in inflectional morphology, where fewer BPE merges suffice for generating highly productive tokens. Motivated by this, our study empirically identifies the optimal number of BPE tokens for Bengali, a language known for its morphological complexity, thus enhancing out-of-distribution automatic speech recognition (ASR) performance. Experimental evaluation reveals that an excessively high number of BPE tokens can lead to overfitting, while approximately 500-1000 tokens result in superior OOV performance. Furthermore, we conduct a comparative analysis of BPE with character-based and unigram-based tokenization methods. By introducing BPE tokenization to Bengali ASR, we achieve a substantial reduction in the word error rate (WER) from 66.44% in our character-based baseline system to 63.80% on the LB-ASRTD eval set and from 46.34% to 42.80% on the SHRUTI eval set, both of which include out-of-distribution data. 【9】 Validation of artificial neural networks to model the acoustic behaviour of induction motors标题:人工神经网络在感应电机声学特性建模中的验证链接:https://arxiv.org/abs/2401.15377作者:F. J. Jimenez-Romero,D. Guijo-Rubio,F. R. Lara-Raya,A. Ruiz-Gonzalez,C. Hervas-Martinez摘要:近十年来,感应电动机的声品质一直是研究的热点。特别是,由于其大量的应用,人们暴露于由噪声排放引起的身体和心理不适。因此,有必要尽量减少其对人口的心理影响。以这种方式,这项工作的主要目标是评估使用多任务人工神经网络作为建模技术,同时预测心理声学参数的感应电机。使用几个输入,例如电机功率信号的电幅度和极数,而不是将电动机的噪声与环境噪声分离。提出了两种不同类型的人工神经网络,以等效声压、响度、粗糙度和尖锐度作为输出,对感应电机的声品质进行评价。具体而言,两种不同的拓扑结构被认为是:简单的模型和更复杂的模型。前者更具有可解释性,而后者以隐藏因果关系为代价导致更高的准确性。专注于简单的可解释的模型,产品单元神经网络取得了最好的结果:MSE和SEP。这种产品单元模型的主要好处是它的简单性,因为只有10个输入变量,概述了多任务人工神经网络的有效传输机制,以提取多个任务的共同特征。最后,利用最优积单元神经网络对感应电机的声学质量进行了深入的分析。摘要:In the last decade, the sound quality of electric induction motors is a hot topic in the research field. Specially, due to its high number of applications, the population is exposed to physical and psychological discomfort caused by the noise emission. Therefore, it is necessary to minimise its psychological impact on the population. In this way, the main goal of this work is to evaluate the use of multitask artificial neural networks as a modelling technique for simultaneously predicting psychoacoustic parameters of induction motors. Several inputs are used, such as, the electrical magnitudes of the motor power signal and the number of poles, instead of separating the noise of the electric motor from the environmental noise. Two different kind of artificial neural networks are proposed to evaluate the acoustic quality of induction motors, by using the equivalent sound pressure, the loudness, the roughness and the sharpness as outputs. Concretely, two different topologies have been considered: simple models and more complex models. The former are more interpretable, while the later lead to higher accuracy at the cost of hiding the cause-effect relationship. Focusing on the simple interpretable models, product unit neural networks achieved the best results: for MSE and for SEP. The main benefit of this product unit model is its simplicity, since only 10 inputs variables are used, outlining the effective transfer mechanism of multitask artificial neural networks to extract common features of multiple tasks. Finally, a deep analysis of the acoustic quality of induction motors in done using the best product unit neural networks. 【10】 Music Auto-Tagging with Robust Music Representation Learned via Domain Adversarial Training标题:通过领域对抗性训练学习的具有稳健音乐表示的音乐自动标记链接:https://arxiv.org/abs/2401.15323作者:Haesun Joung,Kyogu Lee备注:5 pages, 3 figures, accepted to ICASSP 2024摘要:音乐自动标记对于增强音乐发现和推荐至关重要。音乐信息检索(MIR)中的现有模型与多媒体内容中的环境和语音声音等现实世界的噪声作斗争。本研究提出了一种受语音相关任务启发的方法,以提高噪声环境下的音乐自动标记性能。该方法将域对抗训练(DAT)集成到音乐域中,从而实现了能够承受噪声的强大音乐表示。与以前的研究不同,这种方法涉及一个额外的预训练阶段的领域分类器,以避免性能下降,在随后的阶段。添加各种合成的嘈杂音乐数据可以提高模型在不同噪声水平下的泛化能力。该架构通过有效利用未标记的嘈杂音乐数据,展示了增强的音乐自动标记性能。补充未标记数据的额外实验进一步提高了模型的性能,强调了其强大的泛化能力和广泛的适用性。摘要:Music auto-tagging is crucial for enhancing music discovery and recommendation. Existing models in Music Information Retrieval (MIR) struggle with real-world noise such as environmental and speech sounds in multimedia content. This study proposes a method inspired by speech-related tasks to enhance music auto-tagging performance in noisy settings. The approach integrates Domain Adversarial Training (DAT) into the music domain, enabling robust music representations that withstand noise. Unlike previous research, this approach involves an additional pretraining phase for the domain classifier, to avoid performance degradation in the subsequent phase. Adding various synthesized noisy music data improves the model's generalization across different noise levels. The proposed architecture demonstrates enhanced performance in music auto-tagging by effectively utilizing unlabeled noisy music data. Additional experiments with supplementary unlabeled data further improves the model's performance, underscoring its robust generalization capabilities and broad applicability. 【11】 AMuSE: Adaptive Multimodal Analysis for Speaker Emotion Recognition in Group Conversations标题:AMUSE:群话中说话人情感识别的自适应多模式分析链接:https://arxiv.org/abs/2401.15164作者:Naresh Kumar Devulapally,Sidharth Anand,Sreyasee Das Bhattacharjee,Junsong Yuan,Yu-Ping Chang摘要:在群体对话中分析个体情绪对于开发能够进行自然人机交互的智能代理至关重要。虽然可靠的情感识别技术依赖于不同的模态(文本,音频,视频),这些模态之间的固有异质性和受个体独特行为模式影响的动态跨模态交互,使得情感识别的任务非常具有挑战性。这种困难在群体环境中更加复杂,在群体环境中,情绪及其时间演变不仅受到个体的影响,而且还受到外部环境的影响,如观众的反应和正在进行的对话的背景。为了应对这一挑战,我们提出了一个多模态注意力网络,它通过共同学习其交互式的特定于模式的外围网络和中央网络,在不同的空间抽象层次上捕获跨模态的交互。建议的MAN注入跨模态的注意力,通过其外围键值对内的每一层的模式特定的中央查询网络。然后,使用自适应融合技术将所得到的交叉参与的模式特定描述符组合,该技术使模型能够在实例特定的多模式描述符内集成有区别的和互补的模式特定数据模式。给定由一系列话语表示的对话,所提出的AMuSE模型将空间和时间特征压缩成两个密集的描述符:说话者级别和话语级别。这不仅有助于在大规模公共数据集中提供更好的分类性能(加权F1提高3-5%,准确性提高5-7%),还有助于用户通过其多模态可解释性可视化模块理解模型所做的每个情感预测背后的推理。摘要:Analyzing individual emotions during group conversation is crucial in developing intelligent agents capable of natural human-machine interaction. While reliable emotion recognition techniques depend on different modalities (text, audio, video), the inherent heterogeneity between these modalities and the dynamic cross-modal interactions influenced by an individual's unique behavioral patterns make the task of emotion recognition very challenging. This difficulty is compounded in group settings, where the emotion and its temporal evolution are not only influenced by the individual but also by external contexts like audience reaction and context of the ongoing conversation. To meet this challenge, we propose a Multimodal Attention Network that captures cross-modal interactions at various levels of spatial abstraction by jointly learning its interactive bunch of mode-specific Peripheral and Central networks. The proposed MAN injects cross-modal attention via its Peripheral key-value pairs within each layer of a mode-specific Central query network. The resulting cross-attended mode-specific descriptors are then combined using an Adaptive Fusion technique that enables the model to integrate the discriminative and complementary mode-specific data patterns within an instance-specific multimodal descriptor. Given a dialogue represented by a sequence of utterances, the proposed AMuSE model condenses both spatial and temporal features into two dense descriptors: speaker-level and utterance-level. This helps not only in delivering better classification performance (3-5% improvement in Weighted-F1 and 5-7% improvement in Accuracy) in large-scale public datasets but also helps the users in understanding the reasoning behind each emotion prediction made by the model via its Multimodal Explainability Visualization module.