本文经arXiv每日学术速递授权转载
标题: RapVerse:文本中的连贯人声和全身动作
作者:Jiaben Chen,Xin Yan,Yihang Chen,Siyuan Cen,Qinwei Ma,Haoyu Zhen,Kaizhi Qian,Lie Lu,Chuang Gan
备注:Project website: this https URL
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了一个具有挑战性的任务,同时生成3D整体身体运动和直接从文本歌词输入唱歌,超越现有的作品,通常解决这两种模式的隔离。为了便于实现这一点,我们首先收集RapVerse数据集,这是一个包含同步说唱人声、歌词和高质量3D整体人体网格的大型数据集。使用RapVerse数据集,我们研究了跨语言,音频和运动的自回归多模态Transformers的缩放程度,可以增强声音和全身人体运动的连贯性和真实性。对于模态统一,采用矢量量化变分自编码器将全身运动序列编码为离散运动令牌,而利用声乐到单元模型来获得保留内容、韵律信息和歌手身份的量化音频令牌。通过联合执行Transformer建模,这三种形式在一个统一的方式,我们的框架,确保了一个无缝的和现实的融合人声和人体运动。大量的实验表明,我们的统一生成框架不仅产生连贯和逼真的歌声,以及直接从文本输入的人类运动,但也竞争对手的性能专门的单模态生成系统,建立新的基准联合声乐运动生成。该项目的页面可在https: vis-www.cs.umass.edu RapVerse上进行研究。摘要:In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in isolation. To facilitate this, we first collect the RapVerse dataset, a large dataset containing synchronous rapping vocals, lyrics, and high-quality 3D holistic body meshes. With the RapVerse dataset, we investigate the extent to which scaling autoregressive multimodal transformers across language, audio, and motion can enhance the coherent and realistic generation of vocals and whole-body human motions. For modality unification, a vector-quantized variational autoencoder is employed to encode whole-body motion sequences into discrete motion tokens, while a vocal-to-unit model is leveraged to obtain quantized audio tokens preserving content, prosodic information, and singer identity. By jointly performing transformer modeling on these three modalities in a unified way, our framework ensures a seamless and realistic blend of vocals and human motions. Extensive experiments demonstrate that our unified generation framework not only produces coherent and realistic singing vocals alongside human motions directly from textual inputs but also rivals the performance of specialized single-modality generation systems, establishing new benchmarks for joint vocal-motion generation. The project page is available for research purposes at https: vis-www.cs.umass.edu RapVerse.
【2】 DITTO-2: Distilled Diffusion Inference-Time T-Optimization for Music Generation
标题: DITTO-2:音乐生成的蒸馏扩散推理时间T优化
作者:Zachary Novack,Julian McAuley,Taylor Berg-Kirkpatrick,Nicholas Bryan
链接:点击下载PDF文件
摘要:可控的音乐生成方法对于以人为中心的基于AI的音乐创作至关重要,但目前受到速度,质量和控制设计权衡的限制。特别是扩散推断时间T优化(DITTO),提供了最先进的结果,但比实时慢10倍以上,限制了实际使用。我们提出了蒸馏扩散推理时间T -优化(或DITTO-2),这是一种新方法,可以加速基于推理时间优化的控制,并为各种应用程序(如音乐修复,outpainting,强度,旋律和音乐结构控制)解锁比实时更快的生成。我们的方法通过以下方式工作:(1)通过有效的、修改的一致性或一致性轨迹蒸馏过程来蒸馏用于快速采样的预训练扩散模型;(2)使用我们的蒸馏模型执行推理时间优化,其中一步采样作为有效的替代优化任务;以及(3)运行最终的多步采样生成(解码)使用我们估计的噪声潜伏期以获得最佳质量、快速、可控的生成。通过全面的评估,我们发现我们的方法不仅加快了10- 20倍的生成速度,而且同时提高了控制依从性和生成质量。此外,我们将我们的方法应用于最大化文本坚持(CLAP得分)的新应用,并表明我们可以将无文本输入的无条件扩散模型转换为产生最先进的文本控制的模型。可以在https: ditto-music.github.io ditto2 上找到合理的例子。摘要:Controllable music generation methods are critical for human-centered AI-based music creation, but are currently limited by speed, quality, and control design trade-offs. Diffusion Inference-Time T-optimization (DITTO), in particular, offers state-of-the-art results, but is over 10x slower than real-time, limiting practical use. We propose Distilled Diffusion Inference-Time T -Optimization (or DITTO-2), a new method to speed up inference-time optimization-based control and unlock faster-than-real-time generation for a wide-variety of applications such as music inpainting, outpainting, intensity, melody, and musical structure control. Our method works by (1) distilling a pre-trained diffusion model for fast sampling via an efficient, modified consistency or consistency trajectory distillation process (2) performing inference-time optimization using our distilled model with one-step sampling as an efficient surrogate optimization task and (3) running a final multi-step sampling generation (decoding) using our estimated noise latents for best-quality, fast, controllable generation. Through thorough evaluation, we find our method not only speeds up generation over 10-20x, but simultaneously improves control adherence and generation quality all at once. Furthermore, we apply our approach to a new application of maximizing text adherence (CLAP score) and show we can convert an unconditional diffusion model without text inputs into a model that yields state-of-the-art text control. Sound examples can be found at https: ditto-music.github.io ditto2 .
【3】 Iterative Feature Boosting for Explainable Speech Emotion Recognition
标题: 可解释语音情感识别的迭代特征增强
作者:Alaa Nfissi,Wassim Bouachir,Nizar Bouguila,Brian Mishara
Journal-ref:2023 International Conference on Machine Learning and Applications (ICMLA), Jacksonville, FL, USA, 2023, pp. 543-549
链接:点击下载PDF文件
摘要:在语音情感识别(SER)中,使用预定义的特征而不考虑其实际重要性可能会导致高维数据集,包括冗余和不相关的信息。因此,高维学习通常会导致模型精度降低,同时增加计算复杂度。我们的工作强调了仔细考虑和分析功能,以建立有效的SER系统的重要性。我们提出了一种新的监督SER方法的基础上,一个有效的特征工程方法。我们特别注意结果的可解释性,以评估特征相关性并细化特征集。这是通过特征评估循环迭代执行的,使用Shapley值来增强特征选择并提高整体框架性能。因此,我们的方法允许平衡模型性能和透明度之间的利益。该方法在TESS数据集上的情感识别中优于人类水平性能(HLP)和最先进的机器学习方法。摘要:In speech emotion recognition (SER), using predefined features without considering their practical importance may lead to high dimensional datasets, including redundant and irrelevant information. Consequently, high-dimensional learning often results in decreasing model accuracy while increasing computational complexity. Our work underlines the importance of carefully considering and analyzing features in order to build efficient SER systems. We present a new supervised SER method based on an efficient feature engineering approach. We pay particular attention to the explainability of results to evaluate feature relevance and refine feature sets. This is performed iteratively through feature evaluation loop, using Shapley values to boost feature selection and improve overall framework performance. Our approach allows thus to balance the benefits between model performance and transparency. The proposed method outperforms human-level performance (HLP) and state-of-the-art machine learning methods in emotion recognition on the TESS dataset.
【4】 Fill in the Gap! Combining Self-supervised Representation Learning with Neural Audio Synthesis for Speech Inpainting
标题: 填补空白!将自我监督表示学习与神经音频合成相结合用于语音修复
作者:Ihab Asaad,Maxime Jacquelin,Olivier Perrotin,Laurent Girin,Thomas Hueber
链接:点击下载PDF文件
摘要:大多数语音自监督学习(SSL)模型都是用一个借口任务来训练的,该任务包括预测输入信号中缺失的部分,无论是未来的片段(因果预测)还是输入中任何地方被屏蔽的片段(非因果预测)。然后可以将学习的语音表示有效地传送到下游任务(例如,自动语音或说话人识别)。在本研究中,我们研究了使用语音SSL模型进行语音修复,即从其周围上下文重建语音信号的缺失部分,即,完成与所述借口任务非常相似的下游任务。为此,我们将SSL编码器(即HuBERT)与神经声码器(即HiFiGAN)相结合,扮演解码器的角色。特别是,我们提出了两种解决方案来匹配HuBERT输出与HiFiGAN输入,通过冻结一个并微调另一个,反之亦然。在单扬声器和多扬声器设置中评估了两种方法的性能,用于知情和盲修复配置(即,掩模的位置分别是已知的或未知的),具有不同的客观度量和感知评估。性能表明,如果两种解决方案都允许正确重建高达200 ms(在某些情况下甚至是400 ms)的信号部分,那么微调SSL编码器在单扬声器设置情况下提供更准确的信号重建,而在处理多扬声器数据时,冻结它(并训练神经声码器)是一种更好的策略。摘要:Most speech self-supervised learning (SSL) models are trained with a pretext task which consists in predicting missing parts of the input signal, either future segments (causal prediction) or segments masked anywhere within the input (non-causal prediction). Learned speech representations can then be efficiently transferred to downstream tasks (e.g., automatic speech or speaker recognition). In the present study, we investigate the use of a speech SSL model for speech inpainting, that is reconstructing a missing portion of a speech signal from its surrounding context, i.e., fulfilling a downstream task that is very similar to the pretext task. To that purpose, we combine an SSL encoder, namely HuBERT, with a neural vocoder, namely HiFiGAN, playing the role of a decoder. In particular, we propose two solutions to match the HuBERT output with the HiFiGAN input, by freezing one and fine-tuning the other, and vice versa. Performance of both approaches was assessed in single- and multi-speaker settings, for both informed and blind inpainting configurations (i.e., the position of the mask is known or unknown, respectively), with different objective metrics and a perceptual evaluation. Performances show that if both solutions allow to correctly reconstruct signal portions up to the size of 200ms (and even 400ms in some cases), fine-tuning the SSL encoder provides a more accurate signal reconstruction in the single-speaker setting case, while freezing it (and training the neural vocoder instead) is a better strategy when dealing with multi-speaker data.
【5】 Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
标题: 歌唱声音的光谱映射:U-Net辅助的人声分割
作者:Adam Sorrenti
链接:点击下载PDF文件
摘要:从音轨中分离出人声成分是音频信号处理中的一个长期挑战。这项研究解决了从音乐声谱图的声乐成分的明显分离。我们采用短时傅立叶变换(STFT)提取音频波到详细的频率-时间频谱图,利用基准MUSDB 18数据集的音乐分离。随后,我们实现了一个UNet神经网络分割的声谱图图像,旨在描绘和提取准确的歌声成分。我们使用基于U-Net的模型在音频源分离方面取得了显著的成果。频率轴归一化与最小 最大缩放和平均绝对误差(MAE)损失函数的组合实现了7.1 dB的最高源失真比(SDR),表明在分离期间保持原始信号质量的高精度。该设置还记录了令人印象深刻的源干扰比(SIR)和源干扰比(SAR)分数,分别为25.2 dB和7.2 dB。这些值显著优于其他配置,特别是使用基于分位数的归一化或均方误差(MSE)损失函数的配置。我们的源代码、模型权重和演示材料可以在项目的GitHub存储库中找到:https: github.com mbrotos SoundSeg摘要:Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform (STFT) to extract audio waves into detailed frequency-time spectrograms, utilizing the benchmark MUSDB18 dataset for music separation. Subsequently, we implement a UNet neural network to segment the spectrogram image, aiming to delineate and extract singing voice components accurately. We achieved noteworthy results in audio source separation using of our U-Net-based models. The combination of frequency-axis normalization with Min Max scaling and the Mean Absolute Error (MAE) loss function achieved the highest Source-to-Distortion Ratio (SDR) of 7.1 dB, indicating a high level of accuracy in preserving the quality of the original signal during separation. This setup also recorded impressive Source-to-Interference Ratio (SIR) and Source-to-Artifact Ratio (SAR) scores of 25.2 dB and 7.2 dB, respectively. These values significantly outperformed other configurations, particularly those using Quantile-based normalization or a Mean Squared Error (MSE) loss function. Our source code, model weights, and demo material can be found at the project's GitHub repository: https: github.com mbrotos SoundSeg
【6】 Explainable Attribute-Based Speaker Verification
标题: 可解释的基于属性的说话人验证
作者:Xiaoliang Wu,Chau Luu,Peter Bell,Ajitha Rajan
链接:点击下载PDF文件
摘要:本文提出了一种完全可解释的方法说话人确认(SV),一项任务,从根本上依赖于个人说话人的特征。在当前SV系统中不透明地使用说话者属性引起了信任问题。针对这一点,我们提出了一个基于属性的解释SV系统,通过比较个人属性,如性别,国籍和年龄自动提取的语音记录来识别扬声器。我们相信这种方法更符合人类推理,比传统方法更容易理解。在Voxceleb1测试集上进行评估,我们的系统的最佳性能与使用所有正确属性时建立的地面实况相当,证明了其有效性。虽然与不可解释的方法相比,我们的方法牺牲了一些性能,但我们相信它使我们更接近透明,可解释的AI的目标,并为未来通过属性扩展进行增强奠定了基础。摘要:This paper proposes a fully explainable approach to speaker verification (SV), a task that fundamentally relies on individual speaker characteristics. The opaque use of speaker attributes in current SV systems raises concerns of trust. Addressing this, we propose an attribute-based explainable SV system that identifies speakers by comparing personal attributes such as gender, nationality, and age extracted automatically from voice recordings. We believe this approach better aligns with human reasoning, making it more understandable than traditional methods. Evaluated on the Voxceleb1 test set, the best performance of our system is comparable with the ground truth established when using all correct attributes, proving its efficacy. Whilst our approach sacrifices some performance compared to non-explainable methods, we believe that it moves us closer to the goal of transparent, interpretable AI and lays the groundwork for future enhancements through attribute expansion.
【7】 Deep Learning for Assessment of Oral Reading Fluency
标题: 深度学习评估口语阅读流利度
作者:Mithilesh Vaidya,Binaya Kumar Sahoo,Preeti Rao
链接:点击下载PDF文件
摘要:阅读流畅性评估是扫盲方案的一个重要组成部分,有助于指导和监测早期教育干预措施。鉴于教师进行的练习的资源密集型的性质,自动工具,可以操作的录音口语阅读的发展是有吸引力的,作为一个客观的和高度可扩展的解决方案。准确性、速度和表达性等多个复杂的方面构成了人们对阅读流畅性的判断。在这项工作中,我们研究了由人类专家标记的故事文本的儿童音频记录的训练数据集上的端到端建模。预训练的wav2vec2.0模型被采用,因为它有可能减轻来自有限数量的标记数据的挑战。我们报告了一些系统的相关措施的变化的性能,也探讨了学习嵌入已知的重要的阅读流畅性的感知词汇和声学韵律特征。摘要:Reading fluency assessment is a critical component of literacy programmes, serving to guide and monitor early education interventions. Given the resource intensive nature of the exercise when conducted by teachers, the development of automatic tools that can operate on audio recordings of oral reading is attractive as an objective and highly scalable solution. Multiple complex aspects such as accuracy, rate and expressiveness underlie human judgements of reading fluency. In this work, we investigate end-to-end modeling on a training dataset of children's audio recordings of story texts labeled by human experts. The pre-trained wav2vec2.0 model is adopted due its potential to alleviate the challenges from the limited amount of labeled data. We report the performance of a number of system variations on the relevant measures, and also probe the learned embeddings for lexical and acoustic-prosodic features known to be important to the perception of reading fluency.
【8】 Luganda Speech Intent Recognition for IoT Applications
标题: 适用于物联网应用的Luganda语音意图识别
作者:Andrew Katumba,Sudi Murindanyi,John Trevor Kasule,Elvis Mugume
备注:Presented as a conference paper at ICLR 2024AfricaNLP
链接:点击下载PDF文件
摘要:物联网(IoT)技术的出现引起了人们对语音控制智能家居的极大兴趣。虽然许多语音控制的智能家居系统旨在理解和支持英语等广泛使用的语言,但像Luganda这样的低资源语言的使用者可能需要更多的支持。该研究项目旨在为物联网应用开发一个Luganda语音意图分类系统,以将本地语言集成到智能家居环境中。该项目使用Raspberry Pi、Wio Terminal和ESP32节点等硬件组件作为微控制器。Raspberry Pi处理Luganda语音命令,Wio终端是显示设备,ESP 32节点控制物联网设备。这项工作的最终目标是使用Luganda实现语音控制,这是通过部署在Raspberry Pi上的自然语言处理(NLP)模型实现的。NLP模型利用Mel频率倒谱系数(MFCC)作为声学特征,并利用卷积神经网络(Conv2D)架构进行语音意图分类。为此目的,我们策划了一个Luganda语音命令数据集,并且已经开源。这项工作通过整合Luganda语音命令解决了物联网应用中的本地化挑战和语言多样性,使用户能够在不精通英语的情况下与智能家居设备进行交互,特别是在当地语言占主导地位的地区。摘要:The advent of Internet of Things (IoT) technology has generated massive interest in voice-controlled smart homes. While many voice-controlled smart home systems are designed to understand and support widely spoken languages like English, speakers of low-resource languages like Luganda may need more support. This research project aimed to develop a Luganda speech intent classification system for IoT applications to integrate local languages into smart home environments. The project uses hardware components such as Raspberry Pi, Wio Terminal, and ESP32 nodes as microcontrollers. The Raspberry Pi processes Luganda voice commands, the Wio Terminal is a display device, and the ESP32 nodes control the IoT devices. The ultimate objective of this work was to enable voice control using Luganda, which was accomplished through a natural language processing (NLP) model deployed on the Raspberry Pi. The NLP model utilized Mel Frequency Cepstral Coefficients (MFCCs) as acoustic features and a Convolutional Neural Network (Conv2D) architecture for speech intent classification. A dataset of Luganda voice commands was curated for this purpose and this has been made open-source. This work addresses the localization challenges and linguistic diversity in IoT applications by incorporating Luganda voice commands, enabling users to interact with smart home devices without English proficiency, especially in regions where local languages are predominant.
【9】 Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
标题: Sonos语音控制偏见评估数据集:语音助手人口统计偏见评估方法
作者:Chloé Sekkat,Fanny Leroy,Salima Mdhaffar,Blake Perry Smith,Yannick Estève,Joseph Dureau,Alice Coucke
链接:点击下载PDF文件
摘要:最近的研究表明,语音助手并不是对每个人都表现得一样好,但对语音技术的人口统计学鲁棒性的研究仍然很少。这主要是由于具有受控人口统计标签的大型数据集的稀缺性。本文介绍了Sonos语音控制偏差评估数据集,这是一个开放的数据集,由音乐领域的北美英语语音助理请求组成(1,038个扬声器,166小时,170k音频样本,9,040个唯一标记的成绩单),具有受控的人口统计多样性(性别,年龄,方言地区和种族)。我们还发布了一种统计人口统计偏见评估方法,在单变量和多变量水平上,针对此特定用例量身定制,并利用口语理解指标而不是转录准确性,我们认为这是用户体验的更好代表。为了证明该数据集和统计方法检测人口统计偏差的能力,我们考虑了一对最先进的自动语音识别和口语理解模型。结果显示,在不同年龄,方言地区和种族的表现有统计学显着差异。多变量测试对于揭示方言地区、性别和年龄之间的混合效应至关重要。摘要:Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled demographic tags. This paper introduces the Sonos Voice Control Bias Assessment Dataset, an open dataset composed of voice assistant requests for North American English in the music domain (1,038 speakers, 166 hours, 170k audio samples, with 9,040 unique labelled transcripts) with a controlled demographic diversity (gender, age, dialectal region and ethnicity). We also release a statistical demographic bias assessment methodology, at the univariate and multivariate levels, tailored to this specific use case and leveraging spoken language understanding metrics rather than transcription accuracy, which we believe is a better proxy for user experience. To demonstrate the capabilities of this dataset and statistical method to detect demographic bias, we consider a pair of state-of-the-art Automatic Speech Recognition and Spoken Language Understanding models. Results show statistically significant differences in performance across age, dialectal region and ethnicity. Multivariate tests are crucial to shed light on mixed effects between dialectal region, gender and age.
【10】 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem
标题: 奥德赛情感识别挑战任务1第一名解决方案:解决班级失衡问题
作者:Mingjie Chen,Hezhao Zhang,Yuanchao Li,Jiachen Luo,Wen Wu,Ziyang Ma,Peter Bell,Catherine Lai,Joshua Reiss,Lin Wang,Philip C. Woodland,Xie Chen,Huy Phan,Thomas Hain
链接:点击下载PDF文件
摘要:语音情感识别是一个具有挑战性的分类任务,自然的情感语音,特别是当情感类型的分布是不平衡的训练和测试数据。在这种情况下,模型更难学习分离少数类,导致这些少数类有时被忽略或经常被错误分类。以前的工作利用类加权损失进行训练,但问题仍然存在,因为它有时会导致过拟合的小类或下拟合的主要类。本文介绍了一个多站点团队为参加奥德赛2024情感识别挑战赛轨道1开发的系统。挑战数据具有上述属性,因此所提出的系统旨在通过在应用类别加权损失时在优化中引入焦点损失来解决这些问题。具体地,焦点损失进一步由基于先验的类权重加权。实验结果表明,结合这两种方法带来了更好的整体性能,牺牲主要类的性能。该系统还采用多数投票策略来组合7个模型的集合的输出。这些模型使用不同的声学特征和损失函数进行独立训练,目的是为不同的数据提供不同的属性。因此,这些模型在主要类别和次要类别上表现出不同的性能偏好。集成系统的输出在挑战中获得了最好的表现,在68个提交的作品中排名前1。它也优于我们集合中的所有单个模型。在Odyssey 2024情感识别挑战任务-1数据上,该系统获得了35.69%的宏观F1分数和37.32%的准确率。摘要:Speech emotion recognition is a challenging classification task with natural emotional speech, especially when the distribution of emotion types is imbalanced in the training and test data. In this case, it is more difficult for a model to learn to separate minority classes, resulting in those sometimes being ignored or frequently misclassified. Previous work has utilised class weighted loss for training, but problems remain as it sometimes causes over-fitting for minor classes or under-fitting for major classes. This paper presents the system developed by a multi-site team for the participation in the Odyssey 2024 Emotion Recognition Challenge Track-1. The challenge data has the aforementioned properties and therefore the presented systems aimed to tackle these issues, by introducing focal loss in optimisation when applying class weighted loss. Specifically, the focal loss is further weighted by prior-based class weights. Experimental results show that combining these two approaches brings better overall performance, by sacrificing performance on major classes. The system further employs a majority voting strategy to combine the outputs of an ensemble of 7 models. The models are trained independently, using different acoustic features and loss functions - with the aim to have different properties for different data. Hence these models show different performance preferences on major classes and minor classes. The ensemble system output obtained the best performance in the challenge, ranking top-1 among 68 submissions. It also outperformed all single models in our set. On the Odyssey 2024 Emotion Recognition Challenge Task-1 data the system obtained a Macro-F1 score of 35.69% and an accuracy of 37.32%.
【11】 Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
标题: 用于具有未配对数据的音频域传输的高斯流桥
作者:Eloi Moliner,Sebastian Braun,Hannes Gamper
备注:Submitted to IWAENC 2024
链接:点击下载PDF文件
摘要:音频域传输是修改音频信号以匹配不同域的特性,同时保留原始内容的过程。本文探讨了高斯流桥,一种新兴的方法在生成建模,这个问题的潜力。所提出的框架解决了通过一系列的两个确定性的概率流的实现跨不同的音频信号分布的传输问题。所提出的框架便于操纵的目标分布属性,通过一个连续的控制变量,它定义了目标域的某个方面。值得注意的是,这种方法不依赖于成对的样本进行训练。为了解决在保持语音内容一致性方面所面临的挑战,我们建议采用一种训练策略,该策略结合了数据样本和噪声的基于块的小批量最优传输耦合。将我们的无监督方法与已建立的基线进行比较,我们发现在混响和失真操作任务中具有竞争力的性能。尽管存在局限性,但本研究中获得的有趣结果强调了进一步探索的潜力。摘要:Audio domain transfer is the process of modifying audio signals to match characteristics of a different domain, while retaining the original content. This paper investigates the potential of Gaussian Flow Bridges, an emerging approach in generative modeling, for this problem. The presented framework addresses the transport problem across different distributions of audio signals through the implementation of a series of two deterministic probability flows. The proposed framework facilitates manipulation of the target distribution properties through a continuous control variable, which defines a certain aspect of the target domain. Notably, this approach does not rely on paired examples for training. To address identified challenges on maintaining the speech content consistent, we recommend a training strategy that incorporates chunk-based minibatch Optimal Transport couplings of data samples and noise. Comparing our unsupervised method with established baselines, we find competitive performance in tasks of reverberation and distortion manipulation. Despite encoutering limitations, the intriguing results obtained in this study underscore potential for further exploration.
eess.AS音频处理
【1】 1st Place Solution to Odyssey Emotion Recognition Challenge Task1: Tackling Class Imbalance Problem标题: 奥德赛情感识别挑战任务1第一名解决方案:解决班级失衡问题
作者:Mingjie Chen,Hezhao Zhang,Yuanchao Li,Jiachen Luo,Wen Wu,Ziyang Ma,Peter Bell,Catherine Lai,Joshua Reiss,Lin Wang,Philip C. Woodland,Xie Chen,Huy Phan,Thomas Hain
链接:点击下载PDF文件
摘要:语音情感识别是一个具有挑战性的分类任务,自然的情感语音,特别是当情感类型的分布是不平衡的训练和测试数据。在这种情况下,模型更难学习分离少数类,导致这些少数类有时被忽略或经常被错误分类。以前的工作利用类加权损失进行训练,但问题仍然存在,因为它有时会导致过拟合的小类或下拟合的主要类。本文介绍了一个多站点团队为参加奥德赛2024情感识别挑战赛轨道1开发的系统。挑战数据具有上述属性,因此所提出的系统旨在通过在应用类别加权损失时在优化中引入焦点损失来解决这些问题。具体地,焦点损失进一步由基于先验的类权重加权。实验结果表明,结合这两种方法带来了更好的整体性能,牺牲主要类的性能。该系统还采用多数投票策略来组合7个模型的集合的输出。这些模型使用不同的声学特征和损失函数进行独立训练,目的是为不同的数据提供不同的属性。因此,这些模型在主要类别和次要类别上表现出不同的性能偏好。集成系统的输出在挑战中获得了最好的表现,在68个提交的作品中排名前1。它也优于我们集合中的所有单个模型。在Odyssey 2024情感识别挑战任务-1数据上,该系统获得了35.69%的宏观F1分数和37.32%的准确率。摘要:Speech emotion recognition is a challenging classification task with natural emotional speech, especially when the distribution of emotion types is imbalanced in the training and test data. In this case, it is more difficult for a model to learn to separate minority classes, resulting in those sometimes being ignored or frequently misclassified. Previous work has utilised class weighted loss for training, but problems remain as it sometimes causes over-fitting for minor classes or under-fitting for major classes. This paper presents the system developed by a multi-site team for the participation in the Odyssey 2024 Emotion Recognition Challenge Track-1. The challenge data has the aforementioned properties and therefore the presented systems aimed to tackle these issues, by introducing focal loss in optimisation when applying class weighted loss. Specifically, the focal loss is further weighted by prior-based class weights. Experimental results show that combining these two approaches brings better overall performance, by sacrificing performance on major classes. The system further employs a majority voting strategy to combine the outputs of an ensemble of 7 models. The models are trained independently, using different acoustic features and loss functions - with the aim to have different properties for different data. Hence these models show different performance preferences on major classes and minor classes. The ensemble system output obtained the best performance in the challenge, ranking top-1 among 68 submissions. It also outperformed all single models in our set. On the Odyssey 2024 Emotion Recognition Challenge Task-1 data the system obtained a Macro-F1 score of 35.69% and an accuracy of 37.32%.
【2】 Gaussian Flow Bridges for Audio Domain Transfer with Unpaired Data
标题: 用于具有未配对数据的音频域传输的高斯流桥
作者:Eloi Moliner,Sebastian Braun,Hannes Gamper
备注:Submitted to IWAENC 2024
链接:点击下载PDF文件
摘要:音频域传输是修改音频信号以匹配不同域的特性,同时保留原始内容的过程。本文探讨了高斯流桥,一种新兴的方法在生成建模,这个问题的潜力。所提出的框架解决了通过一系列的两个确定性的概率流的实现跨不同的音频信号分布的传输问题。所提出的框架便于操纵的目标分布属性,通过一个连续的控制变量,它定义了目标域的某个方面。值得注意的是,这种方法不依赖于成对的样本进行训练。为了解决在保持语音内容一致性方面所面临的挑战,我们建议采用一种训练策略,该策略结合了数据样本和噪声的基于块的小批量最优传输耦合。将我们的无监督方法与已建立的基线进行比较,我们发现在混响和失真操作任务中具有竞争力的性能。尽管存在局限性,但本研究中获得的有趣结果强调了进一步探索的潜力。摘要:Audio domain transfer is the process of modifying audio signals to match characteristics of a different domain, while retaining the original content. This paper investigates the potential of Gaussian Flow Bridges, an emerging approach in generative modeling, for this problem. The presented framework addresses the transport problem across different distributions of audio signals through the implementation of a series of two deterministic probability flows. The proposed framework facilitates manipulation of the target distribution properties through a continuous control variable, which defines a certain aspect of the target domain. Notably, this approach does not rely on paired examples for training. To address identified challenges on maintaining the speech content consistent, we recommend a training strategy that incorporates chunk-based minibatch Optimal Transport couplings of data samples and noise. Comparing our unsupervised method with established baselines, we find competitive performance in tasks of reverberation and distortion manipulation. Despite encoutering limitations, the intriguing results obtained in this study underscore potential for further exploration.
【3】 RapVerse: Coherent Vocals and Whole-Body Motions Generations from Text
标题: RapVerse:文本中的连贯人声和全身动作
作者:Jiaben Chen,Xin Yan,Yihang Chen,Siyuan Cen,Qinwei Ma,Haoyu Zhen,Kaizhi Qian,Lie Lu,Chuang Gan
备注:Project website: this https URL
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了一个具有挑战性的任务,同时生成3D整体身体运动和直接从文本歌词输入唱歌,超越现有的作品,通常解决这两种模式的隔离。为了便于实现这一点,我们首先收集RapVerse数据集,这是一个包含同步说唱人声、歌词和高质量3D整体人体网格的大型数据集。使用RapVerse数据集,我们研究了跨语言,音频和运动的自回归多模态Transformers的缩放程度,可以增强声音和全身人体运动的连贯性和真实性。对于模态统一,采用矢量量化变分自编码器将全身运动序列编码为离散运动令牌,而利用声乐到单元模型来获得保留内容、韵律信息和歌手身份的量化音频令牌。通过联合执行Transformer建模,这三种形式在一个统一的方式,我们的框架,确保了一个无缝的和现实的融合人声和人体运动。大量的实验表明,我们的统一生成框架不仅产生连贯和逼真的歌声,以及直接从文本输入的人类运动,但也竞争对手的性能专门的单模态生成系统,建立新的基准联合声乐运动生成。该项目的页面可在https: vis-www.cs.umass.edu RapVerse上进行研究。摘要:In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in isolation. To facilitate this, we first collect the RapVerse dataset, a large dataset containing synchronous rapping vocals, lyrics, and high-quality 3D holistic body meshes. With the RapVerse dataset, we investigate the extent to which scaling autoregressive multimodal transformers across language, audio, and motion can enhance the coherent and realistic generation of vocals and whole-body human motions. For modality unification, a vector-quantized variational autoencoder is employed to encode whole-body motion sequences into discrete motion tokens, while a vocal-to-unit model is leveraged to obtain quantized audio tokens preserving content, prosodic information, and singer identity. By jointly performing transformer modeling on these three modalities in a unified way, our framework ensures a seamless and realistic blend of vocals and human motions. Extensive experiments demonstrate that our unified generation framework not only produces coherent and realistic singing vocals alongside human motions directly from textual inputs but also rivals the performance of specialized single-modality generation systems, establishing new benchmarks for joint vocal-motion generation. The project page is available for research purposes at https: vis-www.cs.umass.edu RapVerse.
【4】 Iterative Feature Boosting for Explainable Speech Emotion Recognition
标题: 可解释语音情感识别的迭代特征增强
作者:Alaa Nfissi,Wassim Bouachir,Nizar Bouguila,Brian Mishara
Journal-ref:2023 International Conference on Machine Learning and Applications (ICMLA), Jacksonville, FL, USA, 2023, pp. 543-549
链接:点击下载PDF文件
摘要:在语音情感识别(SER)中,使用预定义的特征而不考虑其实际重要性可能会导致高维数据集,包括冗余和不相关的信息。因此,高维学习通常会导致模型精度降低,同时增加计算复杂度。我们的工作强调了仔细考虑和分析功能,以建立有效的SER系统的重要性。我们提出了一种新的监督SER方法的基础上,一个有效的特征工程方法。我们特别注意结果的可解释性,以评估特征相关性并细化特征集。这是通过特征评估循环迭代执行的,使用Shapley值来增强特征选择并提高整体框架性能。因此,我们的方法允许平衡模型性能和透明度之间的利益。该方法在TESS数据集上的情感识别中优于人类水平性能(HLP)和最先进的机器学习方法。摘要:In speech emotion recognition (SER), using predefined features without considering their practical importance may lead to high dimensional datasets, including redundant and irrelevant information. Consequently, high-dimensional learning often results in decreasing model accuracy while increasing computational complexity. Our work underlines the importance of carefully considering and analyzing features in order to build efficient SER systems. We present a new supervised SER method based on an efficient feature engineering approach. We pay particular attention to the explainability of results to evaluate feature relevance and refine feature sets. This is performed iteratively through feature evaluation loop, using Shapley values to boost feature selection and improve overall framework performance. Our approach allows thus to balance the benefits between model performance and transparency. The proposed method outperforms human-level performance (HLP) and state-of-the-art machine learning methods in emotion recognition on the TESS dataset.
【5】 Fill in the Gap! Combining Self-supervised Representation Learning with Neural Audio Synthesis for Speech Inpainting
标题: 填补空白!将自我监督表示学习与神经音频合成相结合用于语音修复
作者:Ihab Asaad,Maxime Jacquelin,Olivier Perrotin,Laurent Girin,Thomas Hueber
链接:点击下载PDF文件
摘要:大多数语音自监督学习(SSL)模型都是用一个借口任务来训练的,该任务包括预测输入信号中缺失的部分,无论是未来的片段(因果预测)还是输入中任何地方被屏蔽的片段(非因果预测)。然后可以将学习的语音表示有效地传送到下游任务(例如,自动语音或说话人识别)。在本研究中,我们研究了使用语音SSL模型进行语音修复,即从其周围上下文重建语音信号的缺失部分,即,完成与所述借口任务非常相似的下游任务。为此,我们将SSL编码器(即HuBERT)与神经声码器(即HiFiGAN)相结合,扮演解码器的角色。特别是,我们提出了两种解决方案来匹配HuBERT输出与HiFiGAN输入,通过冻结一个并微调另一个,反之亦然。在单扬声器和多扬声器设置中评估了两种方法的性能,用于知情和盲修复配置(即,掩模的位置分别是已知的或未知的),具有不同的客观度量和感知评估。性能表明,如果两种解决方案都允许正确重建高达200 ms(在某些情况下甚至是400 ms)的信号部分,那么微调SSL编码器在单扬声器设置情况下提供更准确的信号重建,而在处理多扬声器数据时,冻结它(并训练神经声码器)是一种更好的策略。摘要:Most speech self-supervised learning (SSL) models are trained with a pretext task which consists in predicting missing parts of the input signal, either future segments (causal prediction) or segments masked anywhere within the input (non-causal prediction). Learned speech representations can then be efficiently transferred to downstream tasks (e.g., automatic speech or speaker recognition). In the present study, we investigate the use of a speech SSL model for speech inpainting, that is reconstructing a missing portion of a speech signal from its surrounding context, i.e., fulfilling a downstream task that is very similar to the pretext task. To that purpose, we combine an SSL encoder, namely HuBERT, with a neural vocoder, namely HiFiGAN, playing the role of a decoder. In particular, we propose two solutions to match the HuBERT output with the HiFiGAN input, by freezing one and fine-tuning the other, and vice versa. Performance of both approaches was assessed in single- and multi-speaker settings, for both informed and blind inpainting configurations (i.e., the position of the mask is known or unknown, respectively), with different objective metrics and a perceptual evaluation. Performances show that if both solutions allow to correctly reconstruct signal portions up to the size of 200ms (and even 400ms in some cases), fine-tuning the SSL encoder provides a more accurate signal reconstruction in the single-speaker setting case, while freezing it (and training the neural vocoder instead) is a better strategy when dealing with multi-speaker data.
【6】 Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
标题: 歌唱声音的光谱映射:U-Net辅助的人声分割
作者:Adam Sorrenti
链接:点击下载PDF文件
摘要:从音轨中分离出人声成分是音频信号处理中的一个长期挑战。这项研究解决了从音乐声谱图的声乐成分的明显分离。我们采用短时傅立叶变换(STFT)提取音频波到详细的频率-时间频谱图,利用基准MUSDB 18数据集的音乐分离。随后,我们实现了一个UNet神经网络分割的声谱图图像,旨在描绘和提取准确的歌声成分。我们使用基于U-Net的模型在音频源分离方面取得了显著的成果。频率轴归一化与最小 最大缩放和平均绝对误差(MAE)损失函数的组合实现了7.1 dB的最高源失真比(SDR),表明在分离期间保持原始信号质量的高精度。该设置还记录了令人印象深刻的源干扰比(SIR)和源干扰比(SAR)分数,分别为25.2 dB和7.2 dB。这些值显著优于其他配置,特别是使用基于分位数的归一化或均方误差(MSE)损失函数的配置。我们的源代码、模型权重和演示材料可以在项目的GitHub存储库中找到:https: github.com mbrotos SoundSeg摘要:Separating vocal elements from musical tracks is a longstanding challenge in audio signal processing. This study tackles the distinct separation of vocal components from musical spectrograms. We employ the Short Time Fourier Transform (STFT) to extract audio waves into detailed frequency-time spectrograms, utilizing the benchmark MUSDB18 dataset for music separation. Subsequently, we implement a UNet neural network to segment the spectrogram image, aiming to delineate and extract singing voice components accurately. We achieved noteworthy results in audio source separation using of our U-Net-based models. The combination of frequency-axis normalization with Min Max scaling and the Mean Absolute Error (MAE) loss function achieved the highest Source-to-Distortion Ratio (SDR) of 7.1 dB, indicating a high level of accuracy in preserving the quality of the original signal during separation. This setup also recorded impressive Source-to-Interference Ratio (SIR) and Source-to-Artifact Ratio (SAR) scores of 25.2 dB and 7.2 dB, respectively. These values significantly outperformed other configurations, particularly those using Quantile-based normalization or a Mean Squared Error (MSE) loss function. Our source code, model weights, and demo material can be found at the project's GitHub repository: https: github.com mbrotos SoundSeg
【7】 Explainable Attribute-Based Speaker Verification
标题: 可解释的基于属性的说话人验证
作者:Xiaoliang Wu,Chau Luu,Peter Bell,Ajitha Rajan
链接:点击下载PDF文件
摘要:本文提出了一种完全可解释的方法说话人确认(SV),一项任务,从根本上依赖于个人说话人的特征。在当前SV系统中不透明地使用说话者属性引起了信任问题。针对这一点,我们提出了一个基于属性的解释SV系统,通过比较个人属性,如性别,国籍和年龄自动提取的语音记录来识别扬声器。我们相信这种方法更符合人类推理,比传统方法更容易理解。在Voxceleb1测试集上进行评估,我们的系统的最佳性能与使用所有正确属性时建立的地面实况相当,证明了其有效性。虽然与不可解释的方法相比,我们的方法牺牲了一些性能,但我们相信它使我们更接近透明,可解释的AI的目标,并为未来通过属性扩展进行增强奠定了基础。摘要:This paper proposes a fully explainable approach to speaker verification (SV), a task that fundamentally relies on individual speaker characteristics. The opaque use of speaker attributes in current SV systems raises concerns of trust. Addressing this, we propose an attribute-based explainable SV system that identifies speakers by comparing personal attributes such as gender, nationality, and age extracted automatically from voice recordings. We believe this approach better aligns with human reasoning, making it more understandable than traditional methods. Evaluated on the Voxceleb1 test set, the best performance of our system is comparable with the ground truth established when using all correct attributes, proving its efficacy. Whilst our approach sacrifices some performance compared to non-explainable methods, we believe that it moves us closer to the goal of transparent, interpretable AI and lays the groundwork for future enhancements through attribute expansion.
【8】 Deep Learning for Assessment of Oral Reading Fluency
标题: 深度学习评估口语阅读流利度
作者:Mithilesh Vaidya,Binaya Kumar Sahoo,Preeti Rao
链接:点击下载PDF文件
摘要:阅读流畅性评估是扫盲方案的一个重要组成部分,有助于指导和监测早期教育干预措施。鉴于教师进行的练习的资源密集型的性质,自动工具,可以操作的录音口语阅读的发展是有吸引力的,作为一个客观的和高度可扩展的解决方案。准确性、速度和表达性等多个复杂的方面构成了人们对阅读流畅性的判断。在这项工作中,我们研究了由人类专家标记的故事文本的儿童音频记录的训练数据集上的端到端建模。预训练的wav2vec2.0模型被采用,因为它有可能减轻来自有限数量的标记数据的挑战。我们报告了一些系统的相关措施的变化的性能,并探讨了学习嵌入已知的重要的阅读流畅性的感知词汇和声学韵律特征。摘要:Reading fluency assessment is a critical component of literacy programmes, serving to guide and monitor early education interventions. Given the resource intensive nature of the exercise when conducted by teachers, the development of automatic tools that can operate on audio recordings of oral reading is attractive as an objective and highly scalable solution. Multiple complex aspects such as accuracy, rate and expressiveness underlie human judgements of reading fluency. In this work, we investigate end-to-end modeling on a training dataset of children's audio recordings of story texts labeled by human experts. The pre-trained wav2vec2.0 model is adopted due its potential to alleviate the challenges from the limited amount of labeled data. We report the performance of a number of system variations on the relevant measures, and also probe the learned embeddings for lexical and acoustic-prosodic features known to be important to the perception of reading fluency.
【9】 Luganda Speech Intent Recognition for IoT Applications
标题: 适用于物联网应用的Luganda语音意图识别
作者:Andrew Katumba,Sudi Murindanyi,John Trevor Kasule,Elvis Mugume
备注:Presented as a conference paper at ICLR 2024AfricaNLP
链接:点击下载PDF文件
摘要:物联网(IoT)技术的出现引起了人们对语音控制智能家居的极大兴趣。虽然许多语音控制的智能家居系统旨在理解和支持英语等广泛使用的语言,但像Luganda这样的低资源语言的使用者可能需要更多的支持。该研究项目旨在为物联网应用开发一个Luganda语音意图分类系统,以将本地语言集成到智能家居环境中。该项目使用Raspberry Pi、Wio Terminal和ESP32节点等硬件组件作为微控制器。Raspberry Pi处理Luganda语音命令,Wio终端是显示设备,ESP 32节点控制物联网设备。这项工作的最终目标是使用Luganda实现语音控制,这是通过部署在Raspberry Pi上的自然语言处理(NLP)模型实现的。NLP模型利用Mel频率倒谱系数(MFCC)作为声学特征,并利用卷积神经网络(Conv2D)架构进行语音意图分类。为此目的,我们策划了一个Luganda语音命令数据集,并且已经开源。这项工作通过整合Luganda语音命令解决了物联网应用中的本地化挑战和语言多样性,使用户能够在不精通英语的情况下与智能家居设备进行交互,特别是在当地语言占主导地位的地区。摘要:The advent of Internet of Things (IoT) technology has generated massive interest in voice-controlled smart homes. While many voice-controlled smart home systems are designed to understand and support widely spoken languages like English, speakers of low-resource languages like Luganda may need more support. This research project aimed to develop a Luganda speech intent classification system for IoT applications to integrate local languages into smart home environments. The project uses hardware components such as Raspberry Pi, Wio Terminal, and ESP32 nodes as microcontrollers. The Raspberry Pi processes Luganda voice commands, the Wio Terminal is a display device, and the ESP32 nodes control the IoT devices. The ultimate objective of this work was to enable voice control using Luganda, which was accomplished through a natural language processing (NLP) model deployed on the Raspberry Pi. The NLP model utilized Mel Frequency Cepstral Coefficients (MFCCs) as acoustic features and a Convolutional Neural Network (Conv2D) architecture for speech intent classification. A dataset of Luganda voice commands was curated for this purpose and this has been made open-source. This work addresses the localization challenges and linguistic diversity in IoT applications by incorporating Luganda voice commands, enabling users to interact with smart home devices without English proficiency, especially in regions where local languages are predominant.
【10】 Sonos Voice Control Bias Assessment Dataset: A Methodology for Demographic Bias Assessment in Voice Assistants
标题: Sonos语音控制偏见评估数据集:语音助手人口统计偏见评估方法
作者:Chloé Sekkat,Fanny Leroy,Salima Mdhaffar,Blake Perry Smith,Yannick Estève,Joseph Dureau,Alice Coucke
链接:点击下载PDF文件
摘要:最近的研究表明,语音助手并不是对每个人都表现得一样好,但对语音技术的人口统计学鲁棒性的研究仍然很少。这主要是由于具有受控人口统计标签的大型数据集的稀缺性。本文介绍了Sonos语音控制偏差评估数据集,这是一个开放的数据集,由音乐领域的北美英语语音助理请求组成(1,038个扬声器,166小时,170k音频样本,9,040个唯一标记的成绩单),具有受控的人口统计多样性(性别,年龄,方言地区和种族)。我们还发布了一种统计人口统计偏见评估方法,在单变量和多变量水平上,针对此特定用例量身定制,并利用口语理解指标而不是转录准确性,我们认为这是用户体验的更好代表。为了证明该数据集和统计方法检测人口统计偏差的能力,我们考虑了一对最先进的自动语音识别和口语理解模型。结果显示,在不同年龄,方言地区和种族的表现有统计学显着差异。多变量测试对于揭示方言地区、性别和年龄之间的混合效应至关重要。摘要:Recent works demonstrate that voice assistants do not perform equally well for everyone, but research on demographic robustness of speech technologies is still scarce. This is mainly due to the rarity of large datasets with controlled demographic tags. This paper introduces the Sonos Voice Control Bias Assessment Dataset, an open dataset composed of voice assistant requests for North American English in the music domain (1,038 speakers, 166 hours, 170k audio samples, with 9,040 unique labelled transcripts) with a controlled demographic diversity (gender, age, dialectal region and ethnicity). We also release a statistical demographic bias assessment methodology, at the univariate and multivariate levels, tailored to this specific use case and leveraging spoken language understanding metrics rather than transcription accuracy, which we believe is a better proxy for user experience. To demonstrate the capabilities of this dataset and statistical method to detect demographic bias, we consider a pair of state-of-the-art Automatic Speech Recognition and Spoken Language Understanding models. Results show statistically significant differences in performance across age, dialectal region and ethnicity. Multivariate tests are crucial to shed light on mixed effects between dialectal region, gender and age.
机器翻译,仅供参考
