今日论文合集:cs.SD语音14篇,eess.AS音频处理14篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Emotion-Driven Melody Harmonization via Melodic Variation and Functional Representation
标题: 通过旋律变奏和功能性表现实现音乐驱动的旋律协调
作者:Jingyue Huang,Yi-Hsuan Yang
备注:This work is the initial version of the ISMIR 2024 paper EMO-Disentanger
链接:点击下载PDF文件
摘要:旋律和声是指在同一旋律中产生不同的和声,以传达所期望的情感。以往的研究发现,由于旋律本身的限制和现有音乐表现形式的局限性,很难通过用不同的和弦来协调同一旋律来改变人们对铅片的情感效价。在本文中,我们提出了一种新的功能表示的象征音乐。这种新的方法考虑到音乐键,认识到他们在塑造音乐的情感特征,通过大小调调性的重要作用。它还允许相对于键的旋律变化,并解决了数据稀缺的问题,以实现更好的情感建模。一个Transformer被用来协调可调旋律,允许以基于规则或基于模型的方式确定调。实验结果证实了我们的新表示在生成关键意识的和声的有效性,客观和主观的评价肯定了我们的方法,以传达特定的效价多才多艺的旋律的潜力。摘要:Emotion-driven melody harmonization aims to generate diverse harmonies for a single melody to convey desired emotions. Previous research found it hard to alter the perceived emotional valence of lead sheets only by harmonizing the same melody with different chords, which may be attributed to the constraints imposed by the melody itself and the limitation of existing music representation. In this paper, we propose a novel functional representation for symbolic music. This new method takes musical keys into account, recognizing their significant role in shaping music's emotional character through major-minor tonality. It also allows for melodic variation with respect to keys and addresses the problem of data scarcity for better emotion modeling. A Transformer is employed to harmonize key-adaptable melodies, allowing for keys determined in rule-based or model-based manner. Experimental results confirm the effectiveness of our new representation in generating key-aware harmonies, with objective and subjective evaluations affirming the potential of our approach to convey specific valence for versatile melody.

【2】 Enhancing Anti-spoofing Countermeasures Robustness through Joint Optimization and Transfer Learning
标题: 通过联合优化和迁移学习增强反欺骗对策的鲁棒性
作者:Yikang Wang,Xingming Wang,Hiromitsu Nishizaki,Ming Li
备注:29 pages, 4 figures, Journal Papers
链接:点击下载PDF文件
摘要:目前合成语音检测的研究主要集中在无噪声语音的未知欺骗方法的检测系统的推广。然而,反欺骗对抗(CM)系统的性能往往不工作,以及在更具挑战性的场景,如那些涉及噪声和混响。针对增强CM系统鲁棒性的问题,提出了一种基于迁移学习的语音增强前端联合优化(TL-SEJ)方法,研究了该方法在提高抗噪声和混响鲁棒性方面的有效性。我们通过一系列的比较和烧蚀实验来评估所提出的方法的性能。实验结果表明,在不同信噪比测试条件下,与基线相比,所提出的TL-SEJ方法将识别准确率提高了2.7%至15.8%。与传统的数据增强方法相比,我们的系统在各种噪声条件下实现了从0.7%到5.8%的精度提高,在不同的RT 60混响场景下实现了从1.7%到2.8%的精度提高。实验结果表明,该方法有效地提高了系统在噪声和混响条件下的鲁棒性。摘要:Current research in synthesized speech detection primarily focuses on the generalization of detection systems to unknown spoofing methods of noise-free speech. However, the performance of anti-spoofing countermeasures (CM) system is often don't work as well in more challenging scenarios, such as those involving noise and reverberation. To address the problem of enhancing the robustness of CM systems, we propose a transfer learning-based speech enhancement front-end joint optimization (TL-SEJ) method, investigating its effectiveness in improving robustness against noise and reverberation. We evaluated the proposed method's performance through a series of comparative and ablation experiments. The experimental results show that, across different signal-to-noise ratio test conditions, the proposed TL-SEJ method improves recognition accuracy by 2.7% to 15.8% compared to the baseline. Compared to conventional data augmentation methods, our system achieves an accuracy improvement ranging from 0.7% to 5.8% in various noisy conditions and from 1.7% to 2.8% under different RT60 reverberation scenarios. These experiments demonstrate that the proposed method effectively enhances system robustness in noisy and reverberant conditions.

【3】 Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
标题: 通过具有结构嵌入的大型语言模型生成实用且可复制的符号音乐
作者:Seungyeon Rhyu,Kichang Yang,Sungjun Cho,Jaehyeon Kim,Kyogu Lee,Moontae Lee
备注:9 pages, 6 figures, 4 tables
链接:点击下载PDF文件
摘要:音乐生成为大型语言模型带来了具有挑战性的复杂性。音乐的象征性结构通常包括垂直和谐和水平对位,促使各种改编和增强大型Transformers。然而,现有的作品有三个主要缺点:1)它们的标记化需要特定于域的注释,例如在原始数据中通常缺失的小节和节拍; 2)在没有特定于域的注释的情况下,很难检查增强标记嵌入方法的纯粹影响;以及3)克服上述缺点的现有作品,例如MuseNet,缺乏可重复性。为了解决这些限制,我们开发了一个受MuseNet启发的基于MIDI的音乐生成框架,实证研究了两个不依赖于特定于领域的注释的结构嵌入。我们提供各种指标和见解,可以指导合适的编码部署。我们还验证了多个嵌入配置可以选择性地提高某些音乐方面。通过HuggingFace提供开源实现,我们的研究结果揭示了利用大型语言模型来实现实用和可再现的音乐生成。摘要:Music generation introduces challenging complexities to large language models. Symbolic structures of music often include vertical harmonization as well as horizontal counterpoint, urging various adaptations and enhancements for large-scale Transformers. However, existing works share three major drawbacks: 1) their tokenization requires domain-specific annotations, such as bars and beats, that are typically missing in raw MIDI data; 2) the pure impact of enhancing token embedding methods is hardly examined without domain-specific annotations; and 3) existing works to overcome the aforementioned drawbacks, such as MuseNet, lack reproducibility. To tackle such limitations, we develop a MIDI-based music generation framework inspired by MuseNet, empirically studying two structural embeddings that do not rely on domain-specific annotations. We provide various metrics and insights that can guide suitable encoding to deploy. We also verify that multiple embedding configurations can selectively boost certain musical aspects. By providing open-source implementations via HuggingFace, our findings shed light on leveraging large language models toward practical and reproducible music generation.

【4】 Wavespace: A Highly Explorable Wavetable Generator
标题: Wavespace:一个高度探索的波表生成器
作者:Hazounne Lee,Kihong Kim,Sungho Lee,Kyogu Lee
链接:点击下载PDF文件
摘要:波表合成通过内插一系列波形(称为波表)来生成音调的准周期波形。由于利用潜在表示的生成模型提供了各种方法在音乐应用中的波形生成,最近也出现了可逆结构的波表生成的研究。虽然它们是有前途的,它仍然是具有挑战性的,以产生详细的控制,在潜在的表示内解开因素的波表。作为回应,我们提出了Wavespace,一个新的框架,使用户具有增强的参数控制的波表生成。我们的模型允许用户将预定义的条件应用于输出波表。我们采用变分自动编码器并将其潜在空间完全分解为不同的波形风格。我们还条件的生成器与辅助音色和形态描述符。这样,用户可以通过独立地操作每个潜在子空间和描述符参数来创建唯一的波表。我们的框架对于实际使用来说足够高效;我们制作了振荡器插件原型,作为Wavespace在数字音频工作空间(DW)中实时集成的概念验证。摘要:Wavetable synthesis generates quasi-periodic waveforms of musical tones by interpolating a list of waveforms called wavetable. As generative models that utilize latent representations offer various methods in waveform generation for musical applications, studies in wavetable generation with invertible architecture have also arisen recently. While they are promising, it is still challenging to generate wavetables with detailed controls in disentangling factors within the latent representation. In response, we present Wavespace, a novel framework for wavetable generation that empowers users with enhanced parameter controls. Our model allows users to apply pre-defined conditions to the output wavetables. We employ a variational autoencoder and completely factorize its latent space to different waveform styles. We also condition the generator with auxiliary timbral and morphological descriptors. This way, users can create unique wavetables by independently manipulating each latent subspace and descriptor parameters. Our framework is efficient enough for practical use; we prototyped an oscillator plug-in as a proof of concept for real-time integration of Wavespace within digital audio workspaces (DAWs).

【5】 Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
标题: 分析和缩小音乐信息检索中综合到真实的传输差距:鼓自动转录任务
作者:Mickaël Zehren,Marco Alunno,Paolo Bientinesi
备注:21 pages, 4 figures
链接:点击下载PDF文件
摘要:自动鼓转录是音乐信息检索中提取和分析音乐曲目节奏的关键工具,但它受到可用于训练的数据集大小的限制。一种用于增加数据量的流行方法是通过从用虚拟乐器呈现的乐谱合成生成它们。这种方法可以产生几乎无限数量的轨迹,但经验证据表明,在先前创建的合成数据集上训练的模型不能很好地转换为真实轨迹。在这项工作中,除了增加数据量外,我们还确定和评估了三种策略,从业者可以使用这些策略来提高生成数据的真实性,从而缩小合成数据与真实数据之间的差距。为了探索它们的有效性,我们使用它们来构建一个新的合成数据集,然后测量模型的性能如何扩展,特别是当增加不同数据集的训练轨道数量时,它将停滞在什么值。通过这样做,我们能够证明上述策略有助于使我们的数据集成为我们评估的合成数据集中具有最真实数据分布和最低合成到真实传输差距的数据集。最后,我们强调了鼓转录中无限数据训练的局限性,并展示了如何克服这些局限性。摘要:Automatic drum transcription is a critical tool in Music Information Retrieval for extracting and analyzing the rhythm of a music track, but it is limited by the size of the datasets available for training. A popular method used to increase the amount of data is by generating them synthetically from music scores rendered with virtual instruments. This method can produce a virtually infinite quantity of tracks, but empirical evidence shows that models trained on previously created synthetic datasets do not transfer well to real tracks. In this work, besides increasing the amount of data, we identify and evaluate three more strategies that practitioners can use to improve the realism of the generated data and, thus, narrow the synthetic-to-real transfer gap. To explore their efficacy, we used them to build a new synthetic dataset and then we measured how the performance of a model scales and, specifically, at what value it will stagnate when increasing the number of training tracks for different datasets. By doing this, we were able to prove that the aforementioned strategies contribute to make our dataset the one with the most realistic data distribution and the lowest synthetic-to-real transfer gap among the synthetic datasets we evaluated. We conclude by highlighting the limits of training with infinite data in drum transcription and we show how they can be overcome.

【6】 Navigating the United States Legislative Landscape on Voice Privacy: Existing Laws, Proposed Bills, Protection for Children, and Synthetic Data for AI
标题: 了解美国语音隐私立法格局:现有法律、拟议法案、儿童保护和人工智能合成数据
作者:Satwik Dutta,John H. L. Hansen
备注:5 pages, 2 figures, accepted at the Interspeech SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:隐私是包括美国在内的全球政策制定者的热门话题。人工智能的不断发展以及对滥用个人数据的担忧促使政策制定者起草关于可信人工智能和公民隐私保护的立法。本文介绍了美国国会隐私立法的现状,并概述了如何将语音数据视为立法定义的一部分。本文还审查了对儿童的额外隐私保护。本文对美国50个州的已颁布和拟议的隐私法进行了全面审查,并对语音数据进行了考虑,包括处理儿童数据的指导方针。作为实际人类数据的开创性替代方案,道德生成的合成数据可以灵活地保持人工智能创新的进展。鉴于决策者在人工智能立法中对合成数据的考虑相对较新,与隐私法相比,本文回顾了合成数据的监管考虑。摘要:Privacy is a hot topic for policymakers across the globe, including the United States. Evolving advances in AI and emerging concerns about the misuse of personal data have pushed policymakers to draft legislation on trustworthy AI and privacy protection for its citizens. This paper presents the state of the privacy legislation at the U.S. Congress and outlines how voice data is considered as part of the legislation definition. This paper also reviews additional privacy protection for children. This paper presents a holistic review of enacted and proposed privacy laws, and consideration for voice data, including guidelines for processing children's data, in those laws across the fifty U.S. states. As a groundbreaking alternative to actual human data, ethically generated synthetic data allows much flexibility to keep AI innovation in progress. Given the consideration of synthetic data in AI legislation by policymakers to be relatively new, as compared to that of privacy laws, this paper reviews regulatory considerations for synthetic data.

【7】 Towards Robust Few-shot Class Incremental Learning in Audio Classification using Contrastive Representation
标题: 使用对比表示在音频分类中实现鲁棒的Few-Shot类增量学习
作者:Riyansha Singh,Parinita Nema,Vinod K Kurmi
链接:点击下载PDF文件
摘要:在机器学习应用中,渐进式数据进入很常见,特别是在音频处理中,增量学习对于实时分析至关重要。Few-Shot类增量学习解决了有限传入数据带来的挑战。现有方法通常集成额外的可训练组件或依赖于固定的嵌入提取器在基础会话上进行训练后,以减轻与灾难性遗忘和模型过拟合危险相关的担忧。然而,在基本会话训练期间单独使用交叉熵损失对于音频数据是次优的。为了解决这个问题,我们建议结合监督对比学习来细化表示空间,增强区分能力并导致更好的泛化,因为它有助于在到达时无缝集成增量类。在具有100个类的NSynth和LibriSpeech数据集以及具有50个和10个类的ESC数据集上的实验结果证明了最先进的性能。摘要:In machine learning applications, gradual data ingress is common, especially in audio processing where incremental learning is vital for real-time analytics. Few-shot class-incremental learning addresses challenges arising from limited incoming data. Existing methods often integrate additional trainable components or rely on a fixed embedding extractor post-training on base sessions to mitigate concerns related to catastrophic forgetting and the dangers of model overfitting. However, using cross-entropy loss alone during base session training is suboptimal for audio data. To address this, we propose incorporating supervised contrastive learning to refine the representation space, enhancing discriminative power and leading to better generalization since it facilitates seamless integration of incremental classes, upon arrival. Experimental results on NSynth and LibriSpeech datasets with 100 classes, as well as ESC dataset with 50 and 10 classes, demonstrate state-of-the-art performance.

【8】 RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
标题: RAVD:视觉线索缺失的多说话者场景中的鲁棒视听语音分离
作者:Tianrui Pan,Jie Liu,Bohan Wang,Jie Tang,Gangshan Wu
链接:点击下载PDF文件
摘要:虽然现有的视听语音分离(AVSS)方法主要集中在两个扬声器分离的视听融合策略,他们表现出严重的性能下降,在多个扬声器分离的情况下。通常,AVSS方法采用引导视频来从给定的音频混合物中顺序地隔离各个扬声器,从而导致在分离的语音的各个片段中出现明显的缺失和噪声部分。在这项研究中,我们提出了一个同时多扬声器分离框架,可以促进单一过程中的多个扬声器的并发分离。我们引入发言者明智的互动,建立扬声器之间的区别和相关性。在VoxCeleb2和LRS3数据集上的实验结果表明,我们的方法分别在分离具有2个,3个,4个和5个扬声器的混合物时达到了最先进的性能。此外,我们的模型可以利用具有完整视听信息的扬声器来减轻其他视觉缺陷的扬声器,从而增强其对丢失视觉线索的弹性。我们还进行实验,其中特定扬声器的视觉信息完全缺失或视觉帧部分缺失。结果表明,我们的模型始终优于其他模型,在涉及2,3,4和5个扬声器的所有设置中表现出最小的性能下降。摘要:While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios. Typically, AVSS methods employ guiding videos to sequentially isolate individual speakers from the given audio mixture, resulting in notable missing and noisy parts across various segments of the separated speech. In this study, we propose a simultaneous multi-speaker separation framework that can facilitate the concurrent separation of multiple speakers within a singular process. We introduce speaker-wise interactions to establish distinctions and correlations among speakers. Experimental results on the VoxCeleb2 and LRS3 datasets demonstrate that our method achieves state-of-the-art performance in separating mixtures with 2, 3, 4, and 5 speakers, respectively. Additionally, our model can utilize speakers with complete audio-visual information to mitigate other visual-deficient speakers, thereby enhancing its resilience to missing visual cues. We also conduct experiments where visual information for specific speakers is entirely absent or visual frames are partially missing. The results demonstrate that our model consistently outperforms others, exhibiting the smallest performance drop across all settings involving 2, 3, 4, and 5 speakers.

【9】 Implementation and Applications of WakeWords Integrated with Speaker Recognition: A Case Study
标题: 唤醒词与说话人识别集成的实现与应用:案例研究
作者:Alexandre Costa Ferro Filho,Elisa Ayumi Masasi de Oliveira,Iago Alves Brito,Pedro Martins Bittencourt
链接:点击下载PDF文件
摘要:本文探讨了人工智能技术在音频和语音处理中的应用,重点是唤醒词和说话人识别在嵌入式系统中的安全访问的集成。随着亚马逊Alexa等声控设备的日益普及,确保安全和用户特定的交互变得至关重要。我们的研究旨在通过利用唤醒词进行初始激活和说话人识别来验证用户权限,从而增强这些系统的安全框架。通过整合这些人工智能驱动的方法,我们提出了一个强大的解决方案,将系统使用限制在授权个人,从而降低未经授权的访问风险。本研究深入研究了唤醒词检测和说话人识别的算法和技术,评估了它们在现实应用中的有效性,并讨论了它们在各种嵌入式系统中实现的潜力,强调了安全性和用户便利性。研究结果强调了使用这些人工智能技术来创建安全,用户友好的语音激活系统的可行性和优势。摘要:This paper explores the application of artificial intelligence techniques in audio and voice processing, focusing on the integration of wake words and speaker recognition for secure access in embedded systems. With the growing prevalence of voice-activated devices such as Amazon Alexa, ensuring secure and user-specific interactions has become paramount. Our study aims to enhance the security framework of these systems by leveraging wake words for initial activation and speaker recognition to validate user permissions. By incorporating these AI-driven methodologies, we propose a robust solution that restricts system usage to authorized individuals, thereby mitigating unauthorized access risks. This research delves into the algorithms and technologies underpinning wake word detection and speaker recognition, evaluates their effectiveness in real-world applications, and discusses the potential for their implementation in various embedded systems, emphasizing security and user convenience. The findings underscore the feasibility and advantages of employing these AI techniques to create secure, user-friendly voice-activated systems.

【10】 Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
标题: 用于小足迹噪音口语关键词发现的频率和渠道注意力网络
作者:Yuanxi Lin,Yuriy Evgenyevich Gapanyuk
备注:Submitted to the APSIPA ASC 2024
链接:点击下载PDF文件
摘要:在本文中,我们的目标是提高鲁棒性的关键字定位(KWS)系统在嘈杂的环境中,同时保持一个小的内存占用。我们提出了一种新的卷积神经网络(CNN)称为FCA-Net,它结合了基于混合器单元的特征交互和基于二维卷积的注意力模块。首先,我们介绍和比较轻量级的注意力方法,以提高CNN中的噪声鲁棒性。然后,我们提出了一个注意力模块,它创建细粒度的注意力权重来捕获通道和频率特定的信息,提高了模型处理噪声条件的能力。通过将基于混合器单元的特征交互与注意力模块相结合,我们提高了性能。此外,我们还使用了一种基于神经网络的多条件训练策略。我们的实验表明,我们的系统优于目前最先进的解决方案,在嘈杂的环境中的小足迹KWS,使其可靠的现实世界中使用。摘要:In this paper, we aim to improve the robustness of Keyword Spotting (KWS) systems in noisy environments while keeping a small memory footprint. We propose a new convolutional neural network (CNN) called FCA-Net, which combines mixer unit-based feature interaction with a two-dimensional convolution-based attention module. First, we introduce and compare lightweight attention methods to enhance noise robustness in CNN. Then, we propose an attention module that creates fine-grained attention weights to capture channel and frequency-specific information, boosting the model's ability to handle noisy conditions. By combining the mixer unit-based feature interaction with the attention module, we enhance performance. Additionally, we use a curriculum-based multi-condition training strategy. Our experiments show that our system outperforms current state-of-the-art solutions for small-footprint KWS in noisy environments, making it reliable for real-world use.

【11】 UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
标题: UNQA:音频、图像、视频和视听内容的统一无参考质量评估
作者:Yuqin Cao,Xiongkuo Min,Yixuan Gao,Wei Sun,Weisi Lin,Guangtao Zhai
链接:点击下载PDF文件
摘要:随着多媒体数据在互联网上的大量传播,多媒体数据的质量评估(QA)成为数字媒体应用的关键。由于多媒体数据包括音频、图像、视频和视听(A V)内容等多种形式,因此研究人员开发了一系列质量保证方法来评估不同形式数据的质量。虽然他们专注于解决单一模态的QA问题,但仍然缺少一个可以处理多种模态的不同媒体的统一QA模型,而后者可以更好地模拟人类的感知行为,并且具有更广泛的应用。在本文中,我们提出了统一的无参考质量评估模型(UNQA)的音频,图像,视频和A V内容,它试图训练一个单一的QA模型在不同的媒体形式。为了解决不同的QA数据库之间的质量尺度不一致的问题,我们开发了一种多模态策略,联合训练UNQA多个QA数据库。基于输入模态,UNQA选择性地提取空间特征、运动特征和音频特征,并通过四个相应的模态回归模块计算最终质量分数。与现有的QA方法相比,UNQA具有两个优点:1)多模态训练策略使QA模型学习更通用和鲁棒的质量感知特征表示,与最先进的QA方法相比,UNQA的性能优越。2)UNQA减少了评估不同模式的多媒体数据所需的模型数量。并且易于部署到实际应用中。摘要:As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications.

【12】 ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement
标题: ctPuLSE:近距离交谈和基于伪标签的远场语音增强
作者:Zhong-Qiu Wang
备注:in submission
链接:点击下载PDF文件
摘要:目前用于神经语音增强的主要方法是通过对模拟的远场噪声-混响语音对(即,混合物)和干净的讲话。然而,经过训练的模型通常对真实记录的混合物表现出有限的泛化能力。为了解决这个问题,本文直接在真实混合物上研究训练增强模型。然而,挑战这种方法的一个主要困难是,由于干净的语音的真实混合是不可用的,缺乏一个很好的监督真实的混合。在这种情况下,假设由近距离说话和远场混合的真实记录对组成的训练集是可用的,我们提出通过近距离说话语音增强来解决这个困难,其中首先在模拟混合物上训练增强模型以增强真实记录的近距离说话混合物,然后可以将估计的近距离说话语音用作监督(即,伪标签),用于直接在成对的真实记录的远场混合物上训练远场语音增强模型。我们将所提出的系统命名为$ textit{ctPuLSE}$。在CHiME-4数据集上的评估结果表明,ctPuLSE可以获得高质量的伪标签,并产生对真实数据具有较强泛化能力的远场语音增强模型。摘要:The current dominant approach for neural speech enhancement is via purely-supervised deep learning on simulated pairs of far-field noisy-reverberant speech (i.e., mixtures) and clean speech. The trained models, however, often exhibit limited generalizability to real-recorded mixtures. To deal with this, this paper investigates training enhancement models directly on real mixtures. However, a major difficulty challenging this approach is that, since the clean speech of real mixtures is unavailable, there lacks a good supervision for real mixtures. In this context, assuming that a training set consisting of real-recorded pairs of close-talk and far-field mixtures is available, we propose to address this difficulty via close-talk speech enhancement, where an enhancement model is first trained on simulated mixtures to enhance real-recorded close-talk mixtures and the estimated close-talk speech can then be utilized as a supervision (i.e., pseudo-label) for training far-field speech enhancement models directly on the paired real-recorded far-field mixtures. We name the proposed system $ textit{ctPuLSE}$. Evaluation results on the CHiME-4 dataset show that ctPuLSE can derive high-quality pseudo-labels and yield far-field speech enhancement models with strong generalizability to real data.

【13】 ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
标题: ASGIR:音频频谱图Transformer引导的鸟类分类和信息检索
作者:Yashwardhan Chaudhuri,Paridhi Mundra,Arnesh Batra,Orchid Chetia Phukan,Arun Balaji Buduru
备注:Accepted to INTERSPEECH'24
链接:点击下载PDF文件
摘要:鸟类鸣叫声的识别和解释是鸟类学研究和生态保护工作的关键,因为它们在了解鸟类行为,进行栖息地评估和判断生态健康方面具有重要意义。本文提出了一个音频频谱引导的分类框架称为ASGIR,以提高鸟类的声音识别和信息检索。我们的工作是伴随着一个简单易用,两步信息检索系统,使用地理位置和鸟的声音,本地化和检索相关的鸟类信息,通过抓取维基百科页面信息的识别鸟类。ASGIR在来自欧洲国家的51类Xeno-Canto数据集鸟类声音的随机子集上提供了实质性的性能,在F1,精度和灵敏度指标上的平均性能为100%。我们的代码如下:https: github.com MainSample1234 AS-GIR。摘要:Recognition and interpretation of bird vocalizations are pivotal in ornithological research and ecological conservation efforts due to their significance in understanding avian behaviour, performing habitat assessment and judging ecological health. This paper presents an audio spectrogram-guided classification framework called ASGIR for improved bird sound recognition and information retrieval. Our work is accompanied by a simple-to-use, two-step information retrieval system that uses geographical location and bird sounds to localize and retrieve relevant bird information by scraping Wikipedia page information of recognized birds. ASGIR offers a substantial performance on a random subset of 51 classes of Xeno-Canto dataset Bird sounds from European countries with a median of 100 % performance on F1, Precision and Sensitivity metrics. Our code is available as follows: https: github.com MainSample1234 AS-GIR .

【14】 VoxMed: One-Step Respiratory Disease Classifier using Digital Stethoscope Sounds
标题: VoxMed:使用数字听诊器声音的一步呼吸道疾病分类器
作者:Paridhi Mundra,Manik Sharma,Yashwardhan Chaudhuri,Orchid Chetia Phukan,Arun Balaji Buduru
备注:Accepted to INTERSPEECH'24
链接:点击下载PDF文件
摘要:随着呼吸道疾病变得越来越普遍,快速准确地检测它们以改善患者护理至关重要。需要改进的诊断方法,用于立即进行医学评估,以获得最佳的患者结果。本文介绍了VoxMed,一个UI辅助的一步分类器,使用数字听诊器记录来诊断呼吸系统疾病。它采用音频频谱图Transformer(AST)进行特征提取,并采用基于1-D CNN的架构对呼吸系统疾病进行分类,为专业人员提供有关患者呼吸系统健康的信息。我们使用ICBHI数据集,其中包括从希腊和葡萄牙患者收集的听诊器记录,对呼吸系统疾病进行分类。GitHub存储库:https: github.com Sample-User131001 VoxMed摘要:As respiratory illnesses become more common, it is crucial to quickly and accurately detect them to improve patient care. There is a need for improved diagnostic methods for immediate medical assessments for optimal patient outcomes. This paper introduces VoxMed, a UI-assisted one-step classifier that uses digital stethoscope recordings to diagnose respiratory diseases. It employs an Audio Spectrogram Transformer(AST) for feature extraction and a 1-D CNN-based architecture to classify respiratory diseases, offering professionals information regarding their patients respiratory health in seconds. We use the ICBHI dataset, which includes stethoscope recordings collected from patients in Greece and Portugal, to classify respiratory diseases. GitHub repository: https: github.com Sample-User131001 VoxMed


eess.AS音频处理
【1】 Blind Acoustic Parameter Estimation Through Task-Agnostic Embeddings Using Latent Approximations
标题: 使用潜在逼近通过任务不可知嵌入进行盲声学参数估计
作者:Philipp Götz,Cagdas Tuna,Andreas Brendel,Andreas Walther,Emanuël A. P. Habets
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
摘要:本文提出了一种从单通道混响语音中盲估计声学参数的方法。该方法分为三个阶段。在第一阶段中,训练变分自动编码器以提取表示为梅尔频谱图的声学脉冲响应的潜在表示。在第二阶段,一个单独的语音编码器进行训练,估计低维表示混响语音的短片段。最后,将预训练的语音编码器与小回归模型相结合,并在两个参数回归任务上进行评估。实验表明,所提出的方法优于完全端到端训练的基线模型。摘要:We present a method for blind acoustic parameter estimation from single-channel reverberant speech. The method is structured into three stages. In the first stage, a variational auto-encoder is trained to extract latent representations of acoustic impulse responses represented as mel-spectrograms. In the second stage, a separate speech encoder is trained to estimate low-dimensional representations from short segments of reverberant speech. Finally, the pre-trained speech encoder is combined with a small regression model and evaluated on two parameter regression tasks. Experimentally, the proposed method is shown to outperform a fully end-to-end trained baseline model.

【2】 Frequency & Channel Attention Network for Small Footprint Noisy Spoken Keyword Spotting
标题: 用于小足迹噪音口语关键词发现的频率和渠道注意力网络
作者:Yuanxi Lin,Yuriy Evgenyevich Gapanyuk
备注:Submitted to the APSIPA ASC 2024
链接:点击下载PDF文件
摘要:在本文中,我们的目标是提高鲁棒性的关键字定位(KWS)系统在嘈杂的环境中,同时保持一个小的内存占用。我们提出了一种新的卷积神经网络(CNN)称为FCA-Net,它结合了基于混合器单元的特征交互和基于二维卷积的注意力模块。首先,我们介绍和比较轻量级的注意力方法,以提高CNN中的噪声鲁棒性。然后,我们提出了一个注意力模块,它创建细粒度的注意力权重来捕获通道和频率特定的信息,提高了模型处理噪声条件的能力。通过将基于混合器单元的特征交互与注意力模块相结合,我们提高了性能。此外,我们还使用了一种基于神经网络的多条件训练策略。我们的实验表明,我们的系统优于目前最先进的解决方案,在嘈杂的环境中的小足迹KWS,使其可靠的现实世界中使用。摘要:In this paper, we aim to improve the robustness of Keyword Spotting (KWS) systems in noisy environments while keeping a small memory footprint. We propose a new convolutional neural network (CNN) called FCA-Net, which combines mixer unit-based feature interaction with a two-dimensional convolution-based attention module. First, we introduce and compare lightweight attention methods to enhance noise robustness in CNN. Then, we propose an attention module that creates fine-grained attention weights to capture channel and frequency-specific information, boosting the model's ability to handle noisy conditions. By combining the mixer unit-based feature interaction with the attention module, we enhance performance. Additionally, we use a curriculum-based multi-condition training strategy. Our experiments show that our system outperforms current state-of-the-art solutions for small-footprint KWS in noisy environments, making it reliable for real-world use.

【3】 UNQA: Unified No-Reference Quality Assessment for Audio, Image, Video, and Audio-Visual Content
标题: UNQA:音频、图像、视频和视听内容的统一无参考质量评估
作者:Yuqin Cao,Xiongkuo Min,Yixuan Gao,Wei Sun,Weisi Lin,Guangtao Zhai
链接:点击下载PDF文件
摘要:随着多媒体数据在互联网上的大量传播,多媒体数据的质量评估(QA)成为数字媒体应用的关键。由于多媒体数据包括音频、图像、视频和视听(A V)内容等多种形式,因此研究人员开发了一系列质量保证方法来评估不同形式数据的质量。虽然他们专注于解决单一模态的QA问题,但仍然缺少一个可以处理多种模态的不同媒体的统一QA模型,而后者可以更好地模拟人类的感知行为,并且具有更广泛的应用。在本文中,我们提出了统一的无参考质量评估模型(UNQA)的音频,图像,视频和A V内容,它试图训练一个单一的QA模型在不同的媒体形式。为了解决不同的QA数据库之间的质量尺度不一致的问题,我们开发了一种多模态策略,联合训练UNQA多个QA数据库。基于输入模态,UNQA选择性地提取空间特征、运动特征和音频特征,并通过四个相应的模态回归模块计算最终质量分数。与现有的QA方法相比,UNQA具有两个优点:1)多模态训练策略使QA模型学习更通用和鲁棒的质量感知特征表示,与最先进的QA方法相比,UNQA的性能优越。2)UNQA减少了评估不同模式的多媒体数据所需的模型数量。并且易于部署到实际应用中。摘要:As multimedia data flourishes on the Internet, quality assessment (QA) of multimedia data becomes paramount for digital media applications. Since multimedia data includes multiple modalities including audio, image, video, and audio-visual (A V) content, researchers have developed a range of QA methods to evaluate the quality of different modality data. While they exclusively focus on addressing the single modality QA issues, a unified QA model that can handle diverse media across multiple modalities is still missing, whereas the latter can better resemble human perception behaviour and also have a wider range of applications. In this paper, we propose the Unified No-reference Quality Assessment model (UNQA) for audio, image, video, and A V content, which tries to train a single QA model across different media modalities. To tackle the issue of inconsistent quality scales among different QA databases, we develop a multi-modality strategy to jointly train UNQA on multiple QA databases. Based on the input modality, UNQA selectively extracts the spatial features, motion features, and audio features, and calculates a final quality score via the four corresponding modality regression modules. Compared with existing QA methods, UNQA has two advantages: 1) the multi-modality training strategy makes the QA model learn more general and robust quality-aware feature representation as evidenced by the superior performance of UNQA compared to state-of-the-art QA methods. 2) UNQA reduces the number of models required to assess multimedia data across different modalities. and is friendly to deploy to practical applications.

【4】 ctPuLSE: Close-Talk, and Pseudo-Label Based Far-Field, Speech Enhancement
标题: ctPuLSE:近距离交谈和基于伪标签的远场语音增强
作者:Zhong-Qiu Wang
备注:in submission
链接:点击下载PDF文件
摘要:目前用于神经语音增强的主要方法是通过对模拟的远场噪声-混响语音对(即,混合物)和干净的讲话。然而,经过训练的模型通常对真实记录的混合物表现出有限的泛化能力。为了解决这个问题,本文直接在真实混合物上研究训练增强模型。然而,挑战这种方法的一个主要困难是,由于干净的语音的真实混合是不可用的,缺乏一个很好的监督真实的混合。在这种情况下,假设由近距离说话和远场混合的真实记录对组成的训练集是可用的,我们提出通过近距离说话语音增强来解决这个困难,其中首先在模拟混合物上训练增强模型以增强真实记录的近距离说话混合物,然后可以将估计的近距离说话语音用作监督(即,伪标签),用于直接在成对的真实记录的远场混合物上训练远场语音增强模型。我们将所提出的系统命名为$ textit{ctPuLSE}$。在CHiME-4数据集上的评估结果表明,ctPuLSE可以获得高质量的伪标签,并产生对真实数据具有较强泛化能力的远场语音增强模型。摘要:The current dominant approach for neural speech enhancement is via purely-supervised deep learning on simulated pairs of far-field noisy-reverberant speech (i.e., mixtures) and clean speech. The trained models, however, often exhibit limited generalizability to real-recorded mixtures. To deal with this, this paper investigates training enhancement models directly on real mixtures. However, a major difficulty challenging this approach is that, since the clean speech of real mixtures is unavailable, there lacks a good supervision for real mixtures. In this context, assuming that a training set consisting of real-recorded pairs of close-talk and far-field mixtures is available, we propose to address this difficulty via close-talk speech enhancement, where an enhancement model is first trained on simulated mixtures to enhance real-recorded close-talk mixtures and the estimated close-talk speech can then be utilized as a supervision (i.e., pseudo-label) for training far-field speech enhancement models directly on the paired real-recorded far-field mixtures. We name the proposed system $ textit{ctPuLSE}$. Evaluation results on the CHiME-4 dataset show that ctPuLSE can derive high-quality pseudo-labels and yield far-field speech enhancement models with strong generalizability to real data.

【5】 Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
标题: 基于数据驱动分层修剪的语音识别动态编码器大小
作者:Jingjing Xu,Wei Zhou,Zijian Yang,Eugen Beck,Ralf Schlueter
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在不同的硬件和 或应用程序约束(如内存和延迟)下部署ASR系统通常需要不同大小的模型。为了避免不同大小的单个模型的冗余训练和优化工作,我们提出了动态编码器大小方法,该方法从头开始在一个超网中联合训练多个性能模型。这些不同大小的节点是从超网中逐层修剪的,因此可以享受完全的参数共享。通过将基于分数的剪枝与超网训练相结合,我们提出了两种新的方法,Simple-Top-k和Iterative-Zero-Out,以数据驱动的方式自动选择性能最好的搜索引擎,避免了资源密集型搜索工作。我们使用CTC在Librispeech和TED-LIUM-v2语料库上进行的实验表明,我们的方法可以实现与每个大小类别的单独训练模型相同的性能。此外,我们的方法始终为全尺寸超网带来小的性能改进。摘要:Varying-size models are often required to deploy ASR systems under different hardware and or application constraints such as memory and latency. To avoid redundant training and optimization efforts for individual models of different sizes, we present the dynamic encoder size approach, which jointly trains multiple performant models within one supernet from scratch. These subnets of various sizes are layer-wise pruned from the supernet, and thus, enjoy full parameter sharing. By combining score-based pruning with supernet training, we propose two novel methods, Simple-Top-k and Iterative-Zero-Out, to automatically select the best-performing subnets in a data-driven manner, avoiding resource-intensive search efforts. Our experiments using CTC on both Librispeech and TED-LIUM-v2 corpora show that our methods can achieve on-par performance as individually trained models of each size category. Also, our approach consistently brings small performance improvements for the full-size supernet.

【6】 ASGIR: Audio Spectrogram Transformer Guided Classification And Information Retrieval For Birds
标题: ASGIR:音频频谱图Transformer引导的鸟类分类和信息检索
作者:Yashwardhan Chaudhuri,Paridhi Mundra,Arnesh Batra,Orchid Chetia Phukan,Arun Balaji Buduru
备注:Accepted to INTERSPEECH'24
链接:点击下载PDF文件
摘要:鸟类鸣叫声的识别和解释是鸟类学研究和生态保护工作的关键,因为它们在了解鸟类行为,进行栖息地评估和判断生态健康方面具有重要意义。本文提出了一个音频频谱引导的分类框架称为ASGIR,以提高鸟类的声音识别和信息检索。我们的工作是伴随着一个简单易用,两步信息检索系统,使用地理位置和鸟的声音,本地化和检索相关的鸟类信息,通过抓取维基百科页面信息的识别鸟类。ASGIR在来自欧洲国家的51类Xeno-Canto数据集鸟类声音的随机子集上提供了实质性的性能,在F1,精度和灵敏度指标上的平均性能为100%。我们的代码如下:https: github.com MainSample1234 AS-GIR。摘要:Recognition and interpretation of bird vocalizations are pivotal in ornithological research and ecological conservation efforts due to their significance in understanding avian behaviour, performing habitat assessment and judging ecological health. This paper presents an audio spectrogram-guided classification framework called ASGIR for improved bird sound recognition and information retrieval. Our work is accompanied by a simple-to-use, two-step information retrieval system that uses geographical location and bird sounds to localize and retrieve relevant bird information by scraping Wikipedia page information of recognized birds. ASGIR offers a substantial performance on a random subset of 51 classes of Xeno-Canto dataset Bird sounds from European countries with a median of 100 % performance on F1, Precision and Sensitivity metrics. Our code is available as follows: https: github.com MainSample1234 AS-GIR .

【7】 VoxMed: One-Step Respiratory Disease Classifier using Digital Stethoscope Sounds
标题: VoxMed:使用数字听诊器声音的一步呼吸道疾病分类器
作者:Paridhi Mundra,Manik Sharma,Yashwardhan Chaudhuri,Orchid Chetia Phukan,Arun Balaji Buduru
备注:Accepted to INTERSPEECH'24
链接:点击下载PDF文件
摘要:随着呼吸道疾病变得越来越普遍,快速准确地检测它们以改善患者护理至关重要。需要改进的诊断方法,用于立即进行医学评估,以获得最佳的患者结果。本文介绍了VoxMed,一个UI辅助的一步分类器,使用数字听诊器记录来诊断呼吸系统疾病。它采用音频频谱图Transformer(AST)进行特征提取,并采用基于1-D CNN的架构对呼吸系统疾病进行分类,为专业人员提供有关患者呼吸健康的信息。我们使用ICBHI数据集,其中包括从希腊和葡萄牙患者收集的听诊器记录,对呼吸系统疾病进行分类。GitHub存储库:https: github.com Sample-User131001 VoxMed摘要:As respiratory illnesses become more common, it is crucial to quickly and accurately detect them to improve patient care. There is a need for improved diagnostic methods for immediate medical assessments for optimal patient outcomes. This paper introduces VoxMed, a UI-assisted one-step classifier that uses digital stethoscope recordings to diagnose respiratory diseases. It employs an Audio Spectrogram Transformer(AST) for feature extraction and a 1-D CNN-based architecture to classify respiratory diseases, offering professionals information regarding their patients respiratory health in seconds. We use the ICBHI dataset, which includes stethoscope recordings collected from patients in Greece and Portugal, to classify respiratory diseases. GitHub repository: https: github.com Sample-User131001 VoxMed

【8】 Emotion-Driven Melody Harmonization via Melodic Variation and Functional Representation
标题: 通过旋律变奏和功能性表现实现音乐驱动的旋律协调
作者:Jingyue Huang,Yi-Hsuan Yang
备注:This work is the initial version of the ISMIR 2024 paper EMO-Disentanger
链接:点击下载PDF文件
摘要:旋律和声是指在同一旋律中产生不同的和声,以传达所期望的情感。以往的研究发现,由于旋律本身的限制和现有音乐表现形式的局限性,很难通过用不同的和弦来协调同一旋律来改变人们对铅片的情感效价。在本文中,我们提出了一种新的功能表示的象征音乐。这种新的方法考虑到音乐键,认识到他们在塑造音乐的情感特征,通过大小调调性的重要作用。它还允许相对于键的旋律变化,并解决了数据稀缺的问题,以实现更好的情感建模。一个Transformer被用来协调可调旋律,允许以基于规则或基于模型的方式确定调。实验结果证实了我们的新表示在生成关键意识的和声的有效性,客观和主观的评价肯定了我们的方法,以传达特定的效价多才多艺的旋律的潜力。摘要:Emotion-driven melody harmonization aims to generate diverse harmonies for a single melody to convey desired emotions. Previous research found it hard to alter the perceived emotional valence of lead sheets only by harmonizing the same melody with different chords, which may be attributed to the constraints imposed by the melody itself and the limitation of existing music representation. In this paper, we propose a novel functional representation for symbolic music. This new method takes musical keys into account, recognizing their significant role in shaping music's emotional character through major-minor tonality. It also allows for melodic variation with respect to keys and addresses the problem of data scarcity for better emotion modeling. A Transformer is employed to harmonize key-adaptable melodies, allowing for keys determined in rule-based or model-based manner. Experimental results confirm the effectiveness of our new representation in generating key-aware harmonies, with objective and subjective evaluations affirming the potential of our approach to convey specific valence for versatile melody.

【9】 Enhancing Anti-spoofing Countermeasures Robustness through Joint Optimization and Transfer Learning
标题: 通过联合优化和迁移学习增强反欺骗对策的鲁棒性
作者:Yikang Wang,Xingming Wang,Hiromitsu Nishizaki,Ming Li
备注:29 pages, 4 figures, Journal Papers
链接:点击下载PDF文件
摘要:目前合成语音检测的研究主要集中在无噪声语音的未知欺骗方法的检测系统的推广。然而,反欺骗对抗(CM)系统的性能往往不工作,以及在更具挑战性的场景,如那些涉及噪声和混响。针对增强CM系统鲁棒性的问题,提出了一种基于迁移学习的语音增强前端联合优化(TL-SEJ)方法,研究了该方法在提高抗噪声和混响鲁棒性方面的有效性。我们通过一系列的比较和烧蚀实验来评估所提出的方法的性能。实验结果表明,在不同信噪比测试条件下,与基线相比,所提出的TL-SEJ方法将识别准确率提高了2.7%至15.8%。与传统的数据增强方法相比,我们的系统在各种噪声条件下实现了从0.7%到5.8%的精度提高,在不同的RT 60混响场景下实现了从1.7%到2.8%的精度提高。实验结果表明,该方法有效地提高了系统在噪声和混响条件下的鲁棒性。摘要:Current research in synthesized speech detection primarily focuses on the generalization of detection systems to unknown spoofing methods of noise-free speech. However, the performance of anti-spoofing countermeasures (CM) system is often don't work as well in more challenging scenarios, such as those involving noise and reverberation. To address the problem of enhancing the robustness of CM systems, we propose a transfer learning-based speech enhancement front-end joint optimization (TL-SEJ) method, investigating its effectiveness in improving robustness against noise and reverberation. We evaluated the proposed method's performance through a series of comparative and ablation experiments. The experimental results show that, across different signal-to-noise ratio test conditions, the proposed TL-SEJ method improves recognition accuracy by 2.7% to 15.8% compared to the baseline. Compared to conventional data augmentation methods, our system achieves an accuracy improvement ranging from 0.7% to 5.8% in various noisy conditions and from 1.7% to 2.8% under different RT60 reverberation scenarios. These experiments demonstrate that the proposed method effectively enhances system robustness in noisy and reverberant conditions.

【10】 Practical and Reproducible Symbolic Music Generation by Large Language Models with Structural Embeddings
标题: 通过具有结构嵌入的大型语言模型生成实用且可复制的符号音乐
作者:Seungyeon Rhyu,Kichang Yang,Sungjun Cho,Jaehyeon Kim,Kyogu Lee,Moontae Lee
备注:9 pages, 6 figures, 4 tables
链接:点击下载PDF文件
摘要:音乐生成为大型语言模型带来了具有挑战性的复杂性。音乐的象征性结构通常包括垂直和谐和水平对位,促使各种改编和增强大型Transformers。然而,现有的作品有三个主要缺点:1)它们的标记化需要特定于域的注释,例如在原始数据中通常缺失的小节和节拍; 2)在没有特定于域的注释的情况下,很难检查增强标记嵌入方法的纯粹影响;以及3)克服上述缺点的现有作品,例如MuseNet,缺乏可重复性。为了解决这样的限制,我们开发了一个基于MIDI的音乐生成框架的灵感来自MuseNet,实证研究两个结构嵌入,不依赖于特定领域的注释。我们提供各种指标和见解,可以指导合适的编码部署。我们还验证了多个嵌入配置可以选择性地提高某些音乐方面。通过HuggingFace提供开源实现,我们的研究结果揭示了利用大型语言模型来实现实用和可再现的音乐生成。摘要:Music generation introduces challenging complexities to large language models. Symbolic structures of music often include vertical harmonization as well as horizontal counterpoint, urging various adaptations and enhancements for large-scale Transformers. However, existing works share three major drawbacks: 1) their tokenization requires domain-specific annotations, such as bars and beats, that are typically missing in raw MIDI data; 2) the pure impact of enhancing token embedding methods is hardly examined without domain-specific annotations; and 3) existing works to overcome the aforementioned drawbacks, such as MuseNet, lack reproducibility. To tackle such limitations, we develop a MIDI-based music generation framework inspired by MuseNet, empirically studying two structural embeddings that do not rely on domain-specific annotations. We provide various metrics and insights that can guide suitable encoding to deploy. We also verify that multiple embedding configurations can selectively boost certain musical aspects. By providing open-source implementations via HuggingFace, our findings shed light on leveraging large language models toward practical and reproducible music generation.

【11】 Wavespace: A Highly Explorable Wavetable Generator
标题: Wavespace:一个高度探索的波表生成器
作者:Hazounne Lee,Kihong Kim,Sungho Lee,Kyogu Lee
链接:点击下载PDF文件
摘要:波表合成通过内插一系列波形(称为波表)来生成音调的准周期波形。由于利用潜在表示的生成模型提供了各种方法在音乐应用中的波形生成,最近也出现了可逆结构的波表生成的研究。虽然它们是有前途的,它仍然是具有挑战性的,以产生详细的控制,在潜在的表示内解开因素的波表。作为回应,我们提出了Wavespace,一个新的框架,使用户具有增强的参数控制的波表生成。我们的模型允许用户将预定义的条件应用于输出波表。我们采用了变分自动编码器和完全因式分解其潜在的空间,以不同的波形风格。我们还条件的生成器与辅助音色和形态描述符。这样,用户可以通过独立地操作每个潜在子空间和描述符参数来创建唯一的波表。我们的框架对于实际使用来说足够高效;我们制作了振荡器插件原型,作为Wavespace在数字音频工作空间(DW)中实时集成的概念验证。摘要:Wavetable synthesis generates quasi-periodic waveforms of musical tones by interpolating a list of waveforms called wavetable. As generative models that utilize latent representations offer various methods in waveform generation for musical applications, studies in wavetable generation with invertible architecture have also arisen recently. While they are promising, it is still challenging to generate wavetables with detailed controls in disentangling factors within the latent representation. In response, we present Wavespace, a novel framework for wavetable generation that empowers users with enhanced parameter controls. Our model allows users to apply pre-defined conditions to the output wavetables. We employ a variational autoencoder and completely factorize its latent space to different waveform styles. We also condition the generator with auxiliary timbral and morphological descriptors. This way, users can create unique wavetables by independently manipulating each latent subspace and descriptor parameters. Our framework is efficient enough for practical use; we prototyped an oscillator plug-in as a proof of concept for real-time integration of Wavespace within digital audio workspaces (DAWs).

【12】 Analyzing and reducing the synthetic-to-real transfer gap in Music Information Retrieval: the task of automatic drum transcription
标题: 分析和缩小音乐信息检索中综合到真实的传输差距:鼓自动转录任务
作者:Mickaël Zehren,Marco Alunno,Paolo Bientinesi
备注:21 pages, 4 figures
链接:点击下载PDF文件
摘要:自动鼓转录是音乐信息检索中提取和分析音乐曲目节奏的关键工具,但它受到可用于训练的数据集大小的限制。一种用于增加数据量的流行方法是通过从用虚拟乐器呈现的乐谱合成生成它们。这种方法可以产生几乎无限数量的轨迹,但经验证据表明,在先前创建的合成数据集上训练的模型不能很好地转换为真实轨迹。在这项工作中,除了增加数据量外,我们还确定和评估了三种策略,从业者可以使用这些策略来提高生成数据的真实性,从而缩小合成数据与真实数据之间的差距。为了探索它们的有效性,我们使用它们来构建一个新的合成数据集,然后测量模型的性能如何扩展,特别是当增加不同数据集的训练轨道数量时,它将停滞在什么值。通过这样做,我们能够证明上述策略有助于使我们的数据集成为我们评估的合成数据集中具有最真实数据分布和最低合成到真实传输差距的数据集。最后,我们强调了鼓转录中无限数据训练的局限性,并展示了如何克服这些局限性。摘要:Automatic drum transcription is a critical tool in Music Information Retrieval for extracting and analyzing the rhythm of a music track, but it is limited by the size of the datasets available for training. A popular method used to increase the amount of data is by generating them synthetically from music scores rendered with virtual instruments. This method can produce a virtually infinite quantity of tracks, but empirical evidence shows that models trained on previously created synthetic datasets do not transfer well to real tracks. In this work, besides increasing the amount of data, we identify and evaluate three more strategies that practitioners can use to improve the realism of the generated data and, thus, narrow the synthetic-to-real transfer gap. To explore their efficacy, we used them to build a new synthetic dataset and then we measured how the performance of a model scales and, specifically, at what value it will stagnate when increasing the number of training tracks for different datasets. By doing this, we were able to prove that the aforementioned strategies contribute to make our dataset the one with the most realistic data distribution and the lowest synthetic-to-real transfer gap among the synthetic datasets we evaluated. We conclude by highlighting the limits of training with infinite data in drum transcription and we show how they can be overcome.

【13】 Navigating the United States Legislative Landscape on Voice Privacy: Existing Laws, Proposed Bills, Protection for Children, and Synthetic Data for AI
标题: 了解美国语音隐私立法格局:现有法律、拟议法案、儿童保护和人工智能合成数据
作者:Satwik Dutta,John H. L. Hansen
备注:5 pages, 2 figures, accepted at the Interspeech SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:隐私是包括美国在内的全球政策制定者的热门话题。人工智能的不断发展以及对滥用个人数据的担忧促使政策制定者起草关于可信人工智能和公民隐私保护的立法。本文介绍了美国国会隐私立法的现状,并概述了如何将语音数据视为立法定义的一部分。本文还审查了对儿童的额外隐私保护。本文对美国50个州的已颁布和拟议的隐私法进行了全面审查,并对语音数据进行了考虑,包括处理儿童数据的指导方针。作为实际人类数据的开创性替代方案,道德生成的合成数据可以灵活地保持人工智能创新的进展。鉴于决策者在人工智能立法中对合成数据的考虑相对较新,与隐私法相比,本文回顾了合成数据的监管考虑。摘要:Privacy is a hot topic for policymakers across the globe, including the United States. Evolving advances in AI and emerging concerns about the misuse of personal data have pushed policymakers to draft legislation on trustworthy AI and privacy protection for its citizens. This paper presents the state of the privacy legislation at the U.S. Congress and outlines how voice data is considered as part of the legislation definition. This paper also reviews additional privacy protection for children. This paper presents a holistic review of enacted and proposed privacy laws, and consideration for voice data, including guidelines for processing children's data, in those laws across the fifty U.S. states. As a groundbreaking alternative to actual human data, ethically generated synthetic data allows much flexibility to keep AI innovation in progress. Given the consideration of synthetic data in AI legislation by policymakers to be relatively new, as compared to that of privacy laws, this paper reviews regulatory considerations for synthetic data.

【14】 RAVSS: Robust Audio-Visual Speech Separation in Multi-Speaker Scenarios with Missing Visual Cues
标题: RAVD:视觉线索缺失的多说话者场景中的鲁棒视听语音分离
作者:Tianrui Pan,Jie Liu,Bohan Wang,Jie Tang,Gangshan Wu
链接:点击下载PDF文件
摘要:虽然现有的视听语音分离(AVSS)方法主要集中在两个扬声器分离的视听融合策略,他们表现出严重的性能下降,在多个扬声器分离的情况下。通常,AVSS方法采用引导视频来从给定的音频混合物中顺序地隔离各个扬声器,从而导致在分离的语音的各个片段中出现明显的缺失和噪声部分。在这项研究中,我们提出了一个同时多扬声器分离框架,可以促进单一过程中的多个扬声器的并发分离。我们引入发言者明智的互动,建立扬声器之间的区别和相关性。在VoxCeleb2和LRS3数据集上的实验结果表明,我们的方法分别在分离具有2个,3个,4个和5个扬声器的混合物时达到了最先进的性能。此外,我们的模型可以利用具有完整视听信息的扬声器来减轻其他视觉缺陷扬声器的影响,从而增强其对丢失视觉线索的恢复能力。我们还进行实验,其中特定扬声器的视觉信息完全缺失或视觉帧部分缺失。结果表明,我们的模型始终优于其他模型,在涉及2,3,4和5个扬声器的所有设置中表现出最小的性能下降。摘要:While existing Audio-Visual Speech Separation (AVSS) methods primarily concentrate on the audio-visual fusion strategy for two-speaker separation, they demonstrate a severe performance drop in the multi-speaker separation scenarios. Typically, AVSS methods employ guiding videos to sequentially isolate individual speakers from the given audio mixture, resulting in notable missing and noisy parts across various segments of the separated speech. In this study, we propose a simultaneous multi-speaker separation framework that can facilitate the concurrent separation of multiple speakers within a singular process. We introduce speaker-wise interactions to establish distinctions and correlations among speakers. Experimental results on the VoxCeleb2 and LRS3 datasets demonstrate that our method achieves state-of-the-art performance in separating mixtures with 2, 3, 4, and 5 speakers, respectively. Additionally, our model can utilize speakers with complete audio-visual information to mitigate other visual-deficient speakers, thereby enhancing its resilience to missing visual cues. We also conduct experiments where visual information for specific speakers is entirely absent or visual frames are partially missing. The results demonstrate that our model consistently outperforms others, exhibiting the smallest performance drop across all settings involving 2, 3, 4, and 5 speakers.


机器翻译,仅供参考