【1】Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds标题:混合溷合法用于稀有无尾蛙音的多标记分类链接:https://arxiv.org/abs/2403.09598作者:Ilyass Moummad,Nicolas Farrugia,Romain Serizel,Jeremy Froidevaux,Vincent Lostanlen摘要:多标签不平衡分类在机器学习中提出了一个重大挑战,特别是在生物声学中,动物的声音经常同时出现,某些声音比其他声音频率低得多。本文重点介绍了使用数据集AnuraSet对无尾类动物物种声音进行分类的具体情况,该数据集包含类别不平衡和多标签示例。为了解决这些挑战,我们引入了Mixup(Mix 2),这是一个利用混合正则化方法Mixup,Manifold Mixup和MultiMix的框架。实验结果表明,单独使用这些方法可能会导致次优结果;然而,当随机应用时,在每次训练迭代中选择一个,它们证明可以有效地解决上述挑战,特别是对于很少出现的罕见类别。进一步的分析表明,Mix 2还擅长在不同级别的类同现中对声音进行分类。摘要:Multi-label imbalanced classification poses a significant challenge in machine learning, particularly evident in bioacoustics where animal sounds often co-occur, and certain sounds are much less frequent than others. This paper focuses on the specific case of classifying anuran species sounds using the dataset AnuraSet, that contains both class imbalance and multi-label examples. To address these challenges, we introduce Mixture of Mixups (Mix2), a framework that leverages mixing regularization methods Mixup, Manifold Mixup, and MultiMix. Experimental results show that these methods, individually, may lead to suboptimal results; however, when applied randomly, with one selected at each training iteration, they prove effective in addressing the mentioned challenges, particularly for rare classes with few occurrences. Further analysis reveals that Mix2 is also proficient in classifying sounds across various levels of class co-occurrences. 【2】 uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures标题:uaMix—MAE:使用无监督音频混合的预训练音频变换器的有效调谐链接:https://arxiv.org/abs/2403.09579作者:Afrina Tabassum,Dung Tran,Trung Dang,Ismini Lourentzou,Kazuhito Koishida备注:5 pages, 6 figures, 4 tables. To appear in ICASSP'2024摘要:Masked Autoencoders(MAE)从未标记数据中学习丰富的低级表示,但需要大量的标记数据才能有效地适应下游任务。相反,实例区分(ID)强调高级语义,提供了一个潜在的解决方案,以减轻在MAE的注释要求。虽然结合这两种方法可以解决具有有限标记数据的下游任务,但天真地将ID集成到MAE中会导致延长的训练时间和高计算成本。为了应对这一挑战,我们引入了uaMix-MAE,这是一种利用无监督音频混合的高效ID调谐策略。利用对比调优,uaMix-MAE对齐预训练的MAE的表示,从而促进对特定于任务的语义的有效适应。为了用少量未标记数据优化模型,我们提出了一种音频混合技术,该技术在输入和虚拟标签空间中操纵音频样本。在低/Few-Shot设置中的实验表明,在使用有限的未标记数据(如AudioSet-20 K)进行调整时,\modelname在各种基准测试中实现了4-6%的准确性提高。代码可在https://github.com/PLAN-Lab/uamix-MAE上获得摘要:Masked Autoencoders (MAEs) learn rich low-level representations from unlabeled data but require substantial labeled data to effectively adapt to downstream tasks. Conversely, Instance Discrimination (ID) emphasizes high-level semantics, offering a potential solution to alleviate annotation requirements in MAEs. Although combining these two approaches can address downstream tasks with limited labeled data, naively integrating ID into MAEs leads to extended training times and high computational costs. To address this challenge, we introduce uaMix-MAE, an efficient ID tuning strategy that leverages unsupervised audio mixtures. Utilizing contrastive tuning, uaMix-MAE aligns the representations of pretrained MAEs, thereby facilitating effective adaptation to task-specific semantics. To optimize the model with small amounts of unlabeled data, we propose an audio mixing technique that manipulates audio samples in both input and virtual label spaces. Experiments in low/few-shot settings demonstrate that \modelname achieves 4-6% accuracy improvements over various benchmarks when tuned with limited unlabeled data, such as AudioSet-20K. Code is available at https://github.com/PLAN-Lab/uamix-MAE
【3】 The Neural-SRP method for positional sound source localization标题:位置声源定位的神经SRP方法链接:https://arxiv.org/abs/2403.09455作者:Eric Grinstein,Toon van Waterschoot,Mike Brookes,Patrick A. Naylor备注:Presented at Asilomar Conference on Signals, Systems, and Computers摘要:转向响应功率(SRP)是一种广泛应用于麦克风阵列声源定位的方法,在许多实际场景中表现出令人满意的定位性能。然而,其性能在高混响环境下会降低。尽管之前已经提出了深度神经网络(DNN)来克服这一限制,但大多数都是针对具有固定空间坐标的特定数量的麦克风进行训练的。这限制了它们在无线声传感器网络中经常观察到的场景中的实际应用,其中每个应用程序都具有自组织麦克风拓扑结构。我们提出了Neural-SRP,这是一种DNN,它将SRP的灵活性与DNN的性能增益相结合。我们使用模拟数据和迁移学习来训练我们的网络,并在记录和模拟数据上评估我们的方法。结果验证了Neural-SRP的本地化性能显着优于基线。摘要:Steered Response Power (SRP) is a widely used method for the task of sound source localization using microphone arrays, showing satisfactory localization performance on many practical scenarios. However, its performance is diminished under highly reverberant environments. Although Deep Neural Networks (DNNs) have been previously proposed to overcome this limitation, most are trained for a specific number of microphones with fixed spatial coordinates. This restricts their practical application on scenarios frequently observed in wireless acoustic sensor networks, where each application has an ad-hoc microphone topology. We propose Neural-SRP, a DNN which combines the flexibility of SRP with the performance gains of DNNs. We train our network using simulated data and transfer learning, and evaluate our approach on recorded and simulated data. Results verify that Neural-SRP's localization performance significantly outperforms the baselines. 【4】 M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment标题:M & M:整合视听线索的认知负荷评估模型链接:https://arxiv.org/abs/2403.09451作者:Long Nguyen-Phuoc,Renald Gaboriau,Dimitri Delacroix,Laurent Navarro备注:None摘要:本文介绍了M&M模型,一种新的多模态多任务学习框架,应用于AVCAffe数据集的认知负荷评估(CLA)。M&M通过双通道架构独特地集成了视听提示,具有用于音频和视频输入的专用流。一个关键的创新在于它的跨通道多头注意机制,融合了同步多任务的不同通道。另一个值得注意的特点是该模型的三个专门分支,每个分支都针对特定的认知负荷标签,从而实现细致入微的特定任务分析。虽然它表现出适度的性能相比,AVCAffe的单任务基线,M\&M展示了一个有前途的框架,综合多模态处理。这项工作为未来多模态多任务学习系统的增强铺平了道路,强调了复杂任务处理的不同数据类型的融合。摘要:This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling.
【5】 LM2D: Lyrics- and Music-Driven Dance Synthesis标题:LM2D:歌词和音乐驱动的舞蹈合成链接:https://arxiv.org/abs/2403.09407作者:Wenjie Yin,Xuejiao Zhao,Yi Yu,Hang Yin,Danica Kragic,Mårten Björkman摘要:舞蹈通常涉及专业的编舞,复杂的动作遵循音乐节奏,也可以受到抒情内容的影响。歌词与听觉维度的结合丰富了基调,使动作生成更符合语义。然而,现有的舞蹈合成方法倾向于对仅以音频信号为条件的运动进行建模。在这项工作中,我们做了两个贡献,以弥合这一差距。首先,我们提出了LM2D,一种新的概率架构,它采用了多模态扩散模型与一致性蒸馏,旨在创建舞蹈条件的音乐和歌词在一个扩散生成步骤。其次,我们介绍了第一个三维舞蹈运动数据集,包括音乐和歌词,获得姿态估计技术。我们使用客观指标和人类评估(包括舞者和编舞)来评估我们的模型。结果表明,LM2D能够产生与歌词和音乐相匹配的逼真和多样化的舞蹈。视频摘要可以在https://youtu.be/4XCgvYookvA上访问。摘要:Dance typically involves professional choreography with complex movements that follow a musical rhythm and can also be influenced by lyrical content. The integration of lyrics in addition to the auditory dimension, enriches the foundational tone and makes motion generation more amenable to its semantic meanings. However, existing dance synthesis methods tend to model motions only conditioned on audio signals. In this work, we make two contributions to bridge this gap. First, we propose LM2D, a novel probabilistic architecture that incorporates a multimodal diffusion model with consistency distillation, designed to create dance conditioned on both music and lyrics in one diffusion generation step. Second, we introduce the first 3D dance-motion dataset that encompasses both music and lyrics, obtained with pose estimation technologies. We evaluate our model against music-only baseline models with objective metrics and human evaluations, including dancers and choreographers. The results demonstrate LM2D is able to produce realistic and diverse dance matching both lyrics and music. A video summary can be accessed at: https://youtu.be/4XCgvYookvA.
【6】 A Practical Guide to Spectrogram Analysis for Audio Signal Processing标题:音频信号处理用频谱分析实用指南链接:https://arxiv.org/abs/2403.09321作者:Zulfidin Khodzhaev摘要:本文对谱图进行了概述,并给出了谱图在信号处理中的实际应用。为了进行分析,以441000 Hz和96000 Hz的采样率记录手指咬合。分析了分段数对功率谱密度(PSD)和谱图的影响,并进行了可视化。摘要:The paper summarizes spectrogram and gives practical application of spectrogram in signal processing. For analysis, finger-snapping is recorded with a sampling rate of 441000 Hz and 96000 Hz. The effects of the number of segments on the Power Spectral Density (PSD) and spectrogram are analyzed and visualized.
【7】 More than words: Advancements and challenges in speech recognition for singing标题:不仅仅是文字:歌唱语音识别的进步和挑战链接:https://arxiv.org/abs/2403.09298作者:Anna Kruspe备注:Conference on Electronic Speech Signal Processing (ESSV) 2024, Keynote摘要:本文讨论了唱歌语音识别的挑战和进步,这是一个与标准语音识别截然不同的领域。歌唱包含独特的挑战,包括广泛的音高变化,不同的声乐风格和背景音乐干扰。我们探索的关键领域,如音素识别,在歌曲中的语言识别,关键字发现,和完整的歌词转录。我将描述我自己在这些任务上进行研究时的一些经验,就像他们开始获得牵引力一样,但也将展示深度学习和大规模数据集的最新发展如何推动这一领域的进步。我的目标是阐明将语音识别应用于唱歌的复杂性,评估当前的能力,并概述未来的研究方向。摘要:This paper addresses the challenges and advancements in speech recognition for singing, a domain distinctly different from standard speech recognition. Singing encompasses unique challenges, including extensive pitch variations, diverse vocal styles, and background music interference. We explore key areas such as phoneme recognition, language identification in songs, keyword spotting, and full lyrics transcription. I will describe some of my own experiences when performing research on these tasks just as they were starting to gain traction, but will also show how recent developments in deep learning and large-scale datasets have propelled progress in this field. My goal is to illuminate the complexities of applying speech recognition to singing, evaluate current capabilities, and outline future research directions.
【8】 An AI-Driven Approach to Wind Turbine Bearing Fault Diagnosis from Acoustic Signals标题:人工智能驱动的风力机轴承故障声信号诊断方法链接:https://arxiv.org/abs/2403.09030作者:Zhao Wang,Xiaomeng Li,Na Li,Longlong Shu摘要:本研究旨在开发一种深度学习模型,用于从声学信号中对风力涡轮发电机的轴承故障进行分类。通过使用来自五种预定义故障类型的音频数据进行训练和验证,成功构建和训练了卷积LSTM模型。为了创建数据集,收集原始音频信号数据并在帧中处理以捕获时域和频域信息。该模型在训练样本上表现出了出色的准确性,并在验证过程中表现出良好的泛化能力,表明其泛化能力的熟练程度。在测试样本上,该模型取得了显著的分类性能,总体准确率超过99.5%,正常状态下的假阳性率小于1%。研究结果为风力发电机组轴承故障的诊断和维护提供了必要的支持,具有提高风力发电机组可靠性和效率的潜力。摘要:This study aimed to develop a deep learning model for the classification of bearing faults in wind turbine generators from acoustic signals. A convolutional LSTM model was successfully constructed and trained by using audio data from five predefined fault types for both training and validation. To create the dataset, raw audio signal data was collected and processed in frames to capture time and frequency domain information. The model exhibited outstanding accuracy on training samples and demonstrated excellent generalization ability during validation, indicating its proficiency of generalization capability. On the test samples, the model achieved remarkable classification performance, with an overall accuracy exceeding 99.5%, and a false positive rate of less than 1% for normal status. The findings of this study provide essential support for the diagnosis and maintenance of bearing faults in wind turbine generators, with the potential to enhance the reliability and efficiency of wind power generation.
eess.AS音频处理【1】 WavCraft: Audio Editing and Generation with Natural Language Prompts标题:WavCraft:使用自然语言脚本进行音频编辑和生成链接:https://arxiv.org/abs/2403.09527作者:Jinhua Liang,Huan Zhang,Haohe Liu,Yin Cao,Qiuqiang Kong,Xubo Liu,Wenwu Wang,Mark D. Plumbley,Huy Phan,Emmanouil Benetos摘要:我们介绍了WavCraft,这是一个集体系统,它利用大型语言模型(LLM)来连接各种特定于任务的模型,以进行音频内容创建和编辑。具体来说,WavCraft以自然语言描述原始声音材料的内容,并根据音频描述和用户的请求提示LLM。WavCraft利用LLM的上下文学习能力将用户的指令分解为几个任务,并与音频专家模块协作处理每个任务。通过任务分解以及一组特定于任务的模型,WavCraft遵循输入指令创建或编辑具有更多细节和原理的音频内容,方便用户控制。此外,WavCraft能够通过对话交互与用户合作,甚至在没有明确用户命令的情况下生成音频内容。实验结果表明,WavCraft算法比现有方法具有更好的性能,尤其是在调整音频片段局部区域时。此外,WavCraft可以遵循复杂的指令来编辑,甚至在输入录音的顶部创建音频内容,为音频制作者提供更广泛的应用。我们的实现和演示可以在https://github.com/JinhuaLiang/WavCraft上找到。摘要:We introduce WavCraft, a collective system that leverages large language models (LLMs) to connect diverse task-specific models for audio content creation and editing. Specifically, WavCraft describes the content of raw sound materials in natural language and prompts the LLM conditioned on audio descriptions and users' requests. WavCraft leverages the in-context learning ability of the LLM to decomposes users' instructions into several tasks and tackle each task collaboratively with audio expert modules. Through task decomposition along with a set of task-specific models, WavCraft follows the input instruction to create or edit audio content with more details and rationales, facilitating users' control. In addition, WavCraft is able to cooperate with users via dialogue interaction and even produce the audio content without explicit user commands. Experiments demonstrate that WavCraft yields a better performance than existing methods, especially when adjusting the local regions of audio clips. Moreover, WavCraft can follow complex instructions to edit and even create audio content on the top of input recordings, facilitating audio producers in a broader range of applications. Our implementation and demos are available at https://github.com/JinhuaLiang/WavCraft. 【2】 Physics-Informed Neural Network for Volumetric Sound field Reconstruction of Speech Signals标题:语音信号体积声场重建的物理信息神经网络链接:https://arxiv.org/abs/2403.09524作者:Marco Olivieri,Xenofon Karakonstantis,Mirco Pezzoli,Fabio Antonacci,Augusto Sarti,Efren Fernandez-Grande备注:Submitted to EURASIP Journal on Audio, Speech, and Music Processing摘要:声学信号处理的最新发展已经看到了深度学习方法的整合,以及基于经典波扩展的方法的持续突出,特别是在声场重建中。物理信息神经网络(PINN)已经成为一种新的框架,弥合了数据驱动和基于模型的技术之间的差距,以解决由偏微分方程控制的物理现象。本文介绍了一种基于PINN的任意体积声场恢复方法。该网络结合了波动方程,在时域中对信号重建施加正则化。这种方法使网络能够学习声音传播的基本物理学,并允许基于有限的一组观察结果对声场进行完整的表征。所提出的方法的有效性进行了验证,通过实验涉及语音信号在现实世界的环境中,考虑不同数量的可用测量。此外,对现有文献中最先进的频域和时域重建方法进行了比较分析,突出了各种测量配置的精度提高。摘要:Recent developments in acoustic signal processing have seen the integration of deep learning methodologies, alongside the continued prominence of classical wave expansion-based approaches, particularly in sound field reconstruction. Physics-Informed Neural Networks (PINNs) have emerged as a novel framework, bridging the gap between data-driven and model-based techniques for addressing physical phenomena governed by partial differential equations. This paper introduces a PINN-based approach for the recovery of arbitrary volumetric acoustic fields. The network incorporates the wave equation to impose a regularization on signal reconstruction in the time domain. This methodology enables the network to learn the underlying physics of sound propagation and allows for the complete characterization of the sound field based on a limited set of observations. The proposed method's efficacy is validated through experiments involving speech signals in a real-world environment, considering varying numbers of available measurements. Moreover, a comparative analysis is undertaken against state-of-the-art frequency-domain and time-domain reconstruction methods from existing literature, highlighting the increased accuracy across the various measurement configurations. 【3】 Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds标题:用于稀有无尾两栖动物音多标签分类的混合算法链接:https://arxiv.org/abs/2403.09598作者:Ilyass Moummad,Nicolas Farrugia,Romain Serizel,Jeremy Froidevaux,Vincent Lostanlen摘要:多标签不平衡分类在机器学习中提出了一个重大挑战,特别是在生物声学中,动物的声音经常同时出现,某些声音比其他声音频率低得多。本文重点介绍了使用数据集AnuraSet对无尾类动物物种声音进行分类的具体情况,该数据集包含类别不平衡和多标签示例。为了解决这些挑战,我们引入了Mixup(Mix 2),这是一个利用混合正则化方法Mixup,Manifold Mixup和MultiMix的框架。实验结果表明,单独使用这些方法可能会导致次优结果;然而,当随机应用时,在每次训练迭代中选择一个,它们证明可以有效地解决上述挑战,特别是对于很少出现的罕见类别。进一步的分析表明,Mix 2还擅长在不同级别的类同现中对声音进行分类。摘要:Multi-label imbalanced classification poses a significant challenge in machine learning, particularly evident in bioacoustics where animal sounds often co-occur, and certain sounds are much less frequent than others. This paper focuses on the specific case of classifying anuran species sounds using the dataset AnuraSet, that contains both class imbalance and multi-label examples. To address these challenges, we introduce Mixture of Mixups (Mix2), a framework that leverages mixing regularization methods Mixup, Manifold Mixup, and MultiMix. Experimental results show that these methods, individually, may lead to suboptimal results; however, when applied randomly, with one selected at each training iteration, they prove effective in addressing the mentioned challenges, particularly for rare classes with few occurrences. Further analysis reveals that Mix2 is also proficient in classifying sounds across various levels of class co-occurrences.
【4】 uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures标题:uaMix—MAE:使用无监督音频混合的预训练音频变换器的有效调谐链接:https://arxiv.org/abs/2403.09579作者:Afrina Tabassum,Dung Tran,Trung Dang,Ismini Lourentzou,Kazuhito Koishida备注:5 pages, 6 figures, 4 tables. To appear in ICASSP'2024摘要:Masked Autoencoders(MAE)从未标记数据中学习丰富的低级表示,但需要大量的标记数据才能有效地适应下游任务。相反,实例区分(ID)强调高级语义,提供了一个潜在的解决方案,以减轻在MAE的注释要求。虽然结合这两种方法可以解决具有有限标记数据的下游任务,但天真地将ID集成到MAE中会导致延长的训练时间和高计算成本。为了应对这一挑战,我们引入了uaMix-MAE,这是一种利用无监督音频混合的高效ID调谐策略。利用对比调优,uaMix-MAE对齐预训练的MAE的表示,从而促进对特定于任务的语义的有效适应。为了用少量未标记数据优化模型,我们提出了一种音频混合技术,该技术在输入和虚拟标签空间中操纵音频样本。在低/Few-Shot设置中的实验表明,在使用有限的未标记数据(如AudioSet-20 K)进行调整时,\modelname在各种基准测试中实现了4-6%的准确性提高。代码可在https://github.com/PLAN-Lab/uamix-MAE上获得摘要:Masked Autoencoders (MAEs) learn rich low-level representations from unlabeled data but require substantial labeled data to effectively adapt to downstream tasks. Conversely, Instance Discrimination (ID) emphasizes high-level semantics, offering a potential solution to alleviate annotation requirements in MAEs. Although combining these two approaches can address downstream tasks with limited labeled data, naively integrating ID into MAEs leads to extended training times and high computational costs. To address this challenge, we introduce uaMix-MAE, an efficient ID tuning strategy that leverages unsupervised audio mixtures. Utilizing contrastive tuning, uaMix-MAE aligns the representations of pretrained MAEs, thereby facilitating effective adaptation to task-specific semantics. To optimize the model with small amounts of unlabeled data, we propose an audio mixing technique that manipulates audio samples in both input and virtual label spaces. Experiments in low/few-shot settings demonstrate that \modelname achieves 4-6% accuracy improvements over various benchmarks when tuned with limited unlabeled data, such as AudioSet-20K. Code is available at https://github.com/PLAN-Lab/uamix-MAE 【5】 The Neural-SRP method for positional sound source localization标题:定位声源定位的Neural—SRP方法链接:https://arxiv.org/abs/2403.09455作者:Eric Grinstein,Toon van Waterschoot,Mike Brookes,Patrick A. Naylor备注:Presented at Asilomar Conference on Signals, Systems, and Computers摘要:转向响应功率(SRP)是一种广泛应用于麦克风阵列声源定位的方法,在许多实际场景中表现出令人满意的定位性能。然而,其性能在高混响环境下会降低。尽管之前已经提出了深度神经网络(DNN)来克服这一限制,但大多数都是针对具有固定空间坐标的特定数量的麦克风进行训练的。这限制了它们在无线声传感器网络中经常观察到的场景中的实际应用,其中每个应用程序都具有自组织麦克风拓扑结构。我们提出了Neural-SRP,这是一种DNN,它将SRP的灵活性与DNN的性能增益相结合。我们使用模拟数据和迁移学习来训练我们的网络,并在记录和模拟数据上评估我们的方法。结果验证了Neural-SRP的本地化性能显着优于基线。摘要:Steered Response Power (SRP) is a widely used method for the task of sound source localization using microphone arrays, showing satisfactory localization performance on many practical scenarios. However, its performance is diminished under highly reverberant environments. Although Deep Neural Networks (DNNs) have been previously proposed to overcome this limitation, most are trained for a specific number of microphones with fixed spatial coordinates. This restricts their practical application on scenarios frequently observed in wireless acoustic sensor networks, where each application has an ad-hoc microphone topology. We propose Neural-SRP, a DNN which combines the flexibility of SRP with the performance gains of DNNs. We train our network using simulated data and transfer learning, and evaluate our approach on recorded and simulated data. Results verify that Neural-SRP's localization performance significantly outperforms the baselines.
【6】 M&M: Multimodal-Multitask Model Integrating Audiovisual Cues in Cognitive Load Assessment标题:M & M:整合视听线索的认知负荷评估模型链接:https://arxiv.org/abs/2403.09451作者:Long Nguyen-Phuoc,Renald Gaboriau,Dimitri Delacroix,Laurent Navarro备注:None摘要:本文介绍了M&M模型,一种新的多模态多任务学习框架,应用于AVCAffe数据集的认知负荷评估(CLA)。M&M通过双通道架构独特地集成了视听提示,具有用于音频和视频输入的专用流。一个关键的创新在于它的跨通道多头注意机制,融合了同步多任务的不同通道。另一个值得注意的特点是该模型的三个专门分支,每个分支都针对特定的认知负荷标签,从而实现细致入微的特定任务分析。虽然它表现出适度的性能相比,AVCAffe的单任务基线,M\&M展示了一个有前途的框架,综合多模态处理。这项工作为未来多模态多任务学习系统的增强铺平了道路,强调了复杂任务处理的不同数据类型的融合。摘要:This paper introduces the M&M model, a novel multimodal-multitask learning framework, applied to the AVCAffe dataset for cognitive load assessment (CLA). M&M uniquely integrates audiovisual cues through a dual-pathway architecture, featuring specialized streams for audio and video inputs. A key innovation lies in its cross-modality multihead attention mechanism, fusing the different modalities for synchronized multitasking. Another notable feature is the model's three specialized branches, each tailored to a specific cognitive load label, enabling nuanced, task-specific analysis. While it shows modest performance compared to the AVCAffe's single-task baseline, M\&M demonstrates a promising framework for integrated multimodal processing. This work paves the way for future enhancements in multimodal-multitask learning systems, emphasizing the fusion of diverse data types for complex task handling.
【7】 LM2D: Lyrics- and Music-Driven Dance Synthesis标题:LM2D:歌词和音乐驱动的舞蹈合成链接:https://arxiv.org/abs/2403.09407作者:Wenjie Yin,Xuejiao Zhao,Yi Yu,Hang Yin,Danica Kragic,Mårten Björkman摘要:舞蹈通常涉及专业的编舞,复杂的动作遵循音乐节奏,也可以受到抒情内容的影响。歌词与听觉维度的结合丰富了基调,使动作生成更符合语义。然而,现有的舞蹈合成方法倾向于对仅以音频信号为条件的运动进行建模。在这项工作中,我们做了两个贡献,以弥合这一差距。首先,我们提出了LM2D,一种新的概率架构,它采用了多模态扩散模型与一致性蒸馏,旨在创建舞蹈条件的音乐和歌词在一个扩散生成步骤。其次,我们介绍了第一个三维舞蹈运动数据集,包括音乐和歌词,获得姿态估计技术。我们使用客观指标和人类评估(包括舞者和编舞)来评估我们的模型。结果表明,LM2D能够产生与歌词和音乐相匹配的逼真和多样化的舞蹈。视频摘要可以在https://youtu.be/4XCgvYookvA上访问。摘要:Dance typically involves professional choreography with complex movements that follow a musical rhythm and can also be influenced by lyrical content. The integration of lyrics in addition to the auditory dimension, enriches the foundational tone and makes motion generation more amenable to its semantic meanings. However, existing dance synthesis methods tend to model motions only conditioned on audio signals. In this work, we make two contributions to bridge this gap. First, we propose LM2D, a novel probabilistic architecture that incorporates a multimodal diffusion model with consistency distillation, designed to create dance conditioned on both music and lyrics in one diffusion generation step. Second, we introduce the first 3D dance-motion dataset that encompasses both music and lyrics, obtained with pose estimation technologies. We evaluate our model against music-only baseline models with objective metrics and human evaluations, including dancers and choreographers. The results demonstrate LM2D is able to produce realistic and diverse dance matching both lyrics and music. A video summary can be accessed at: https://youtu.be/4XCgvYookvA. 【8】 A Practical Guide to Spectrogram Analysis for Audio Signal Processing标题:音频信号处理的频谱图分析实用指南链接:https://arxiv.org/abs/2403.09321作者:Zulfidin Khodzhaev摘要:本文对谱图进行了概述,并给出了谱图在信号处理中的实际应用。为了进行分析,以441000 Hz和96000 Hz的采样率记录手指咬合。分析了分段数对功率谱密度(PSD)和谱图的影响,并进行了可视化。摘要:The paper summarizes spectrogram and gives practical application of spectrogram in signal processing. For analysis, finger-snapping is recorded with a sampling rate of 441000 Hz and 96000 Hz. The effects of the number of segments on the Power Spectral Density (PSD) and spectrogram are analyzed and visualized. 【9】 More than words: Advancements and challenges in speech recognition for singing标题:超越文字:歌唱语音识别的进展和挑战链接:https://arxiv.org/abs/2403.09298作者:Anna Kruspe备注:Conference on Electronic Speech Signal Processing (ESSV) 2024, Keynote摘要:本文讨论了唱歌语音识别的挑战和进步,这是一个与标准语音识别截然不同的领域。歌唱包含独特的挑战,包括广泛的音高变化,不同的声乐风格和背景音乐干扰。我们探索的关键领域,如音素识别,在歌曲中的语言识别,关键字发现,和完整的歌词转录。我将描述我自己在这些任务上进行研究时的一些经验,就像他们开始获得牵引力一样,但也将展示深度学习和大规模数据集的最新发展如何推动这一领域的进步。我的目标是阐明将语音识别应用于唱歌的复杂性,评估当前的能力,并概述未来的研究方向。摘要:This paper addresses the challenges and advancements in speech recognition for singing, a domain distinctly different from standard speech recognition. Singing encompasses unique challenges, including extensive pitch variations, diverse vocal styles, and background music interference. We explore key areas such as phoneme recognition, language identification in songs, keyword spotting, and full lyrics transcription. I will describe some of my own experiences when performing research on these tasks just as they were starting to gain traction, but will also show how recent developments in deep learning and large-scale datasets have propelled progress in this field. My goal is to illuminate the complexities of applying speech recognition to singing, evaluate current capabilities, and outline future research directions.
【10】 An AI-Driven Approach to Wind Turbine Bearing Fault Diagnosis from Acoustic Signals标题:人工智能驱动的风力机轴承声信号故障诊断方法链接:https://arxiv.org/abs/2403.09030作者:Zhao Wang,Xiaomeng Li,Na Li,Longlong Shu摘要:本研究旨在开发一种深度学习模型,用于从声学信号中对风力涡轮发电机的轴承故障进行分类。通过使用来自五种预定义故障类型的音频数据进行训练和验证,成功构建和训练了卷积LSTM模型。为了创建数据集,收集原始音频信号数据并在帧中处理以捕获时域和频域信息。该模型在训练样本上表现出了出色的准确性,并在验证过程中表现出良好的泛化能力,表明其泛化能力的熟练程度。在测试样本上,该模型取得了显著的分类性能,总体准确率超过99.5%,正常状态下的假阳性率小于1%。研究结果为风力发电机组轴承故障的诊断和维护提供了必要的支持,具有提高风力发电机组可靠性和效率的潜力。摘要:This study aimed to develop a deep learning model for the classification of bearing faults in wind turbine generators from acoustic signals. A convolutional LSTM model was successfully constructed and trained by using audio data from five predefined fault types for both training and validation. To create the dataset, raw audio signal data was collected and processed in frames to capture time and frequency domain information. The model exhibited outstanding accuracy on training samples and demonstrated excellent generalization ability during validation, indicating its proficiency of generalization capability. On the test samples, the model achieved remarkable classification performance, with an overall accuracy exceeding 99.5%, and a false positive rate of less than 1% for normal status. The findings of this study provide essential support for the diagnosis and maintenance of bearing faults in wind turbine generators, with the potential to enhance the reliability and efficiency of wind power generation.