今日论文合集:cs.SD语音4篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载

cs.SD语音
【1】 Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
标题: 研究用于基于语音语言模型的语音生成的神经音频编解码器
作者:Jiaqi Li,Dongmei Wang,Xiaofei Wang,Yao Qian,Long Zhou,Shujie Liu,Midia Yousefi,Canrun Li,Chung-Hsien Tsai,Zhen Xiao,Yanqing Liu,Junkun Chen,Sheng Zhao,Jinyu Li,Zhizheng Wu,Michael Zeng
备注:Accepted by SLT-2024
链接:点击下载PDF文件
摘要:神经音频编解码器令牌充当基于语音语言模型(SLM)的语音生成的基本构建块。然而,有没有系统的了解如何编解码器系统影响的SLM的语音生成性能。在这项工作中,我们研究了SLM语音生成框架内的编解码器令牌,为有效的编解码器设计提供见解。我们在相同的数据集和损失函数上重新训练现有的高性能神经编解码器模型,以比较它们在统一设置中的性能。我们将编解码器令牌集成到两个SLM系统中:基于掩码的并行语音生成系统和基于自回归(AR)加非自回归(NAR)模型的系统。我们的研究结果表明,在编解码器系统中更好的语音重建并不能保证改善SLM中的语音生成。高质量的编解码器是SLM中自然语音产生的关键,而语音可懂度更多地取决于量化机制。摘要:Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism.

【2】 Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
标题: 在语音情感识别中寻找有效的预处理方法和具有高效通道注意力的基于CNN的架构
作者:Byunggun Kim,Younghun Kwon
链接:点击下载PDF文件
摘要:语音情感识别(SER)是利用计算机模型对语音中的人类情感进行分类。最近,随着深度学习技术的适应,SER的性能稳步提高。然而,与许多使用语音数据的领域不同,SER模型中用于训练的数据是不足的。这会导致神经网络训练的过拟合,导致性能下降。事实上,成功的情感识别需要有效的预处理方法和有效使用权重参数数量的模型结构。在本研究中,我们建议使用八个具有不同频率-时间分辨率的数据集版本来寻找有效的情感语音预处理方法。我们提出了一个6层卷积神经网络(CNN)模型与有效的通道注意力(ECA)追求一个有效的模型结构。特别是,良好定位的ECA块可以改善信道特征表示,只有几个参数。在交互式情绪二元运动捕捉(IEMOCAP)数据集上,提高情感语音预处理的频率分辨率可以提高情感识别性能。此外,深度卷积层之后的ECA可以有效地增加信道特征表示。因此,最好的结果(79.37UA 79.68WA),可以得到超过以前的SER模型的性能。此外,为了弥补情感语音数据的不足,我们尝试了多种预处理数据方法,这些方法可以增强用一个样本中所有不同设置预处理的可训练数据。在实验中,我们可以达到最高的结果(80.28UA 80.46WA)。摘要:Speech emotion recognition (SER) classifies human emotions in speech with a computer model. Recently, performance in SER has steadily increased as deep learning techniques have adapted. However, unlike many domains that use speech data, data for training in the SER model is insufficient. This causes overfitting of training of the neural network, resulting in performance degradation. In fact, successful emotion recognition requires an effective preprocessing method and a model structure that efficiently uses the number of weight parameters. In this study, we propose using eight dataset versions with different frequency-time resolutions to search for an effective emotional speech preprocessing method. We propose a 6-layer convolutional neural network (CNN) model with efficient channel attention (ECA) to pursue an efficient model structure. In particular, the well-positioned ECA blocks can improve channel feature representation with only a few parameters. With the interactive emotional dyadic motion capture (IEMOCAP) dataset, increasing the frequency resolution in preprocessing emotional speech can improve emotion recognition performance. Also, ECA after the deep convolution layer can effectively increase channel feature representation. Consequently, the best result (79.37UA 79.68WA) can be obtained, exceeding the performance of previous SER models. Furthermore, to compensate for the lack of emotional speech data, we experiment with multiple preprocessing data methods that augment trainable data preprocessed with all different settings from one sample. In the experiment, we can achieve the highest result (80.28UA 80.46WA).

【3】 MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
标题: MetaCGM:动态原声转换,实现具有环境感知和个性化的连续多场景体验
作者:Haoxuan Liu,Zihao Wang,Haorong Hong,Youwei Feng,Jiaxin Yu,Han Diao,Yunfei Xu,Kejun Zhang
链接:点击下载PDF文件
摘要:本文介绍了MetaBGM,这是一个开创性的框架,用于生成适应动态场景和实时用户交互的背景音乐。我们将多场景定义为环境上下文的变化,例如游戏设置或电影场景中的过渡。为了解决将后端数据转换为音频生成模型的音乐描述文本的挑战,MetaBGM采用了一种新颖的两阶段生成方法,将连续的场景和用户状态数据转换为这些文本,然后将其输入到音频生成模型中进行实时配乐创建。实验结果表明,MetaBGM有效地生成上下文相关的和动态的背景音乐的交互式应用程序。摘要:This paper introduces MetaBGM, a groundbreaking framework for generating background music that adapts to dynamic scenes and real-time user interactions. We define multi-scene as variations in environmental contexts, such as transitions in game settings or movie scenes. To tackle the challenge of converting backend data into music description texts for audio generation models, MetaBGM employs a novel two-stage generation approach that transforms continuous scene and user state data into these texts, which are then fed into an audio generation model for real-time soundtrack creation. Experimental results demonstrate that MetaBGM effectively generates contextually relevant and dynamic background music for interactive applications.

【4】 Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence
标题: 巴西葡萄牙语空间化噪音句子(LiSN-S)听力测试的开发:演示软件、言语刺激和句子等效
作者:Bruno S. Masiero,Leticia R. Borges,Harvey Dillon,Maria Francisca Colella-Santos
链接:点击下载PDF文件
摘要:空间化噪声句子听力(LiSN-S)是一种评估听觉空间处理的测试,目前仅在英语中可用。它在耳机下产生三维听觉环境,并使用简单的重复响应协议来确定在各种条件下竞争语音中呈现的句子的语音接收阈值(SRT)。为了开发巴西葡萄牙语的LiSN-S测试,有必要准备一个由专业配音演员录制的语音数据库,并设计演示软件。这些句子被呈现给35名成年人(年龄在19岁到40岁之间)和24名儿童(年龄在8岁到10岁之间),他们都有正常的听力,通过音调和言语测听和鼓室压测试来验证,并且在学校表现良好。我们使用描述单词错误率与呈现水平的逻辑曲线,为每个句子进行拟合,以选择一组120个句子进行测试。此外,所有选定的句子都进行了幅度调整,以获得相同的可懂度。巴西葡萄牙语的LiSN-S框架已准备好进行规范性数据分析。在其结论之后,我们相信它将有助于诊断和康复巴西儿童与噪音环境中的听力困难有关的投诉摘要:The Listening in Spatialized Noise Sentences (LiSN-S) is a test to evaluate auditory spatial processing currently only available in the English language. It produces a three-dimensional auditory environment under headphones and uses a simple repetition response protocol to determine speech reception thresholds (SRTs) for sentences presented in competing speech under various conditions. In order to develop the LiSN-S test in Brazilian Portuguese, it was necessary to prepare a speech database recorded by professional voice actresses and to devise presentation software. These sentences were presented to 35 adults (aged between 19 and 40 years) and 24 children (aged between 8 and 10 years), all with normal hearing-verified through tone and speech audiometry and tympanometry-and good performance at school. We used a logistic curve describing word error rate versus presentation level, fitted for each sentence, to select a set of 120 sentences for the test. Furthermore, all selected sentences were adjusted in amplitude for equal intelligibility. The framework of LiSN-S in Brazilian Portuguese is ready for normative data analysis. After its conclusion, we believe it will contribute to diagnosing and rehabilitating Brazilian children with complaints related to hearing difficulties in noisy environments


eess.AS音频处理
【1】 NPU-NTU System for Voice Privacy 2024 Challenge
标题: NPU-NTU语音隐私系统2024挑战
作者:Jixun Yao,Nikita Kuzmin,Qing Wang,Pengcheng Guo,Ziqian Ning,Dake Guo,Kong Aik Lee,Eng-Siong Chng,Lei Xie
备注:System description for VPC 2024
链接:点击下载PDF文件
摘要:说话人匿名是一种有效的隐私保护方案,它在隐藏说话人身份的同时保留了原始语音的语言内容和非语言信息。为建立公平的基准,并促进说话者匿名化系统的比较,语音隐私挑战赛(VPC)于2020年及2022年举行,并计划于2024年举行新一届。在本文中,我们描述了我们提出的VPC 2024扬声器匿名化系统。我们的系统采用了一个解纠缠的神经编解码器架构和串行解纠缠策略,逐步解开全球扬声器的身份和时变的语言内容和非语言信息。我们介绍了多种蒸馏方法来解开语言内容,说话人身份和情感。这些方法包括语义提取、监督说话人提取和帧级情感提取。基于这些蒸馏,我们匿名的原始扬声器身份使用一组候选扬声器身份和随机生成的扬声器身份的加权和。我们的系统在VPC 2024中实现了隐私保护和情感保护的最佳平衡。摘要:Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair benchmark and facilitate comparison of speaker anonymization systems, the VoicePrivacy Challenge (VPC) was held in 2020 and 2022, with a new edition planned for 2024. In this paper, we describe our proposed speaker anonymization system for VPC 2024. Our system employs a disentangled neural codec architecture and a serial disentanglement strategy to gradually disentangle the global speaker identity and time-variant linguistic content and paralinguistic information. We introduce multiple distillation methods to disentangle linguistic content, speaker identity, and emotion. These methods include semantic distillation, supervised speaker distillation, and frame-level emotion distillation. Based on these distillations, we anonymize the original speaker identity using a weighted sum of a set of candidate speaker identities and a randomly generated speaker identity. Our system achieves the best trade-off of privacy protection and emotion preservation in VPC 2024.

【2】 Low Complexity Own Voice Reconstruction for Hearables with an In-ear Microphone
标题: 使用入耳式麦克风进行低复杂度的可听语音重建
作者:Mattes Ohlenbusch,Christian Rollwage,Simon Doclo
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:配备有一个或多个麦克风的可听设备通常用于语音通信。在这里,我们考虑的情况下,一个hearable是用来捕捉用户自己的声音在嘈杂的环境。在这种情况下,自己的声音重建(OVR)是必不可少的,以提高记录的嘈杂自己的声音信号的质量和可懂度。在以前的工作中,我们开发了一个基于深度学习的OVR系统,旨在通过使用具有自身语音传输特征的音素相关模型的数据增强来减少用于训练的特定于设备的录音量。鉴于有限的计算资源上的听觉,在本文中,我们提出了低复杂度的变体的OVR系统的基础上的FT-JNF架构,并调查所需的量的设备特定的记录有效的数据增强和微调。仿真结果表明,所提出的OVR系统大大提高了语音质量,即使在低复杂度和有限的设备特定的录音量的约束下。摘要:Hearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user's own voice in a noisy environment. In this scenario, own voice reconstruction (OVR) is essential for enhancing the quality and intelligibility of the recorded noisy own voice signals. In previous work, we developed a deep learning-based OVR system, aiming to reduce the amount of device-specific recordings for training by using data augmentation with phoneme-dependent models of own voice transfer characteristics. Given the limited computational resources available on hearables, in this paper we propose low-complexity variants of an OVR system based on the FT-JNF architecture and investigate the required amount of device-specific recordings for effective data augmentation and fine-tuning. Simulation results show that the proposed OVR system considerably improves speech quality, even under constraints of low complexity and a limited amount of device-specific recordings.

【3】 Development of the Listening in Spatialized Noise-Sentences (LiSN-S) Test in Brazilian Portuguese: Presentation Software, Speech Stimuli, and Sentence Equivalence
标题: 巴西葡萄牙语空间化噪音句子(LiSN-S)听力测试的开发:演示软件、言语刺激和句子等效
作者:Bruno S. Masiero,Leticia R. Borges,Harvey Dillon,Maria Francisca Colella-Santos
链接:点击下载PDF文件
摘要:空间化噪声句子听力(LiSN-S)是一种评估听觉空间处理的测试,目前仅在英语中可用。它在耳机下产生三维听觉环境,并使用简单的重复响应协议来确定在各种条件下竞争语音中呈现的句子的语音接收阈值(SRT)。为了开发巴西葡萄牙语的LiSN-S测试,有必要准备一个由专业配音演员录制的语音数据库,并设计演示软件。这些句子被呈现给35名成年人(年龄在19岁到40岁之间)和24名儿童(年龄在8岁到10岁之间),他们都有正常的听力,通过音调和言语测听和鼓室压测试来验证,并且在学校表现良好。我们使用描述单词错误率与呈现水平的逻辑曲线,为每个句子进行拟合,以选择一组120个句子进行测试。此外,所有选定的句子都进行了幅度调整,以获得相同的可懂度。巴西葡萄牙语的LiSN-S框架已准备好进行规范性数据分析。在其结论之后,我们相信它将有助于诊断和康复巴西儿童与噪音环境中的听力困难有关的投诉摘要:The Listening in Spatialized Noise Sentences (LiSN-S) is a test to evaluate auditory spatial processing currently only available in the English language. It produces a three-dimensional auditory environment under headphones and uses a simple repetition response protocol to determine speech reception thresholds (SRTs) for sentences presented in competing speech under various conditions. In order to develop the LiSN-S test in Brazilian Portuguese, it was necessary to prepare a speech database recorded by professional voice actresses and to devise presentation software. These sentences were presented to 35 adults (aged between 19 and 40 years) and 24 children (aged between 8 and 10 years), all with normal hearing-verified through tone and speech audiometry and tympanometry-and good performance at school. We used a logistic curve describing word error rate versus presentation level, fitted for each sentence, to select a set of 120 sentences for the test. Furthermore, all selected sentences were adjusted in amplitude for equal intelligibility. The framework of LiSN-S in Brazilian Portuguese is ready for normative data analysis. After its conclusion, we believe it will contribute to diagnosing and rehabilitating Brazilian children with complaints related to hearing difficulties in noisy environments

【4】 Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
标题: 研究用于基于语音语言模型的语音生成的神经音频编解码器
作者:Jiaqi Li,Dongmei Wang,Xiaofei Wang,Yao Qian,Long Zhou,Shujie Liu,Midia Yousefi,Canrun Li,Chung-Hsien Tsai,Zhen Xiao,Yanqing Liu,Junkun Chen,Sheng Zhao,Jinyu Li,Zhizheng Wu,Michael Zeng
备注:Accepted by SLT-2024
链接:点击下载PDF文件
摘要:神经音频编解码器令牌充当基于语音语言模型(SLM)的语音生成的基本构建块。然而,有没有系统的了解如何编解码器系统影响的SLM的语音生成性能。在这项工作中,我们研究了SLM语音生成框架内的编解码器令牌,为有效的编解码器设计提供见解。我们在相同的数据集和损失函数上重新训练现有的高性能神经编解码器模型,以比较它们在统一设置中的性能。我们将编解码器令牌集成到两个SLM系统中:基于掩码的并行语音生成系统和基于自回归(AR)加非自回归(NAR)模型的系统。我们的研究结果表明,在编解码器系统中更好的语音重建并不能保证改善SLM中的语音生成。高质量的编解码器是SLM中自然语音产生的关键,而语音可懂度更多地取决于量化机制。摘要:Neural audio codec tokens serve as the fundamental building blocks for speech language model (SLM)-based speech generation. However, there is no systematic understanding on how the codec system affects the speech generation performance of the SLM. In this work, we examine codec tokens within SLM framework for speech generation to provide insights for effective codec design. We retrain existing high-performing neural codec models on the same data set and loss functions to compare their performance in a uniform setting. We integrate codec tokens into two SLM systems: masked-based parallel speech generation system and an auto-regressive (AR) plus non-auto-regressive (NAR) model-based system. Our findings indicate that better speech reconstruction in codec systems does not guarantee improved speech generation in SLM. A high-quality codec decoder is crucial for natural speech production in SLM, while speech intelligibility depends more on quantization mechanism.

【5】 Searching for Effective Preprocessing Method and CNN-based Architecture with Efficient Channel Attention on Speech Emotion Recognition
标题: 在语音情感识别中寻找有效的预处理方法和具有高效通道注意力的基于CNN的架构
作者:Byunggun Kim,Younghun Kwon
链接:点击下载PDF文件
摘要:语音情感识别(SER)是利用计算机模型对语音中的人类情感进行分类。最近,随着深度学习技术的适应,SER的性能稳步提高。然而,与许多使用语音数据的领域不同,SER模型中用于训练的数据是不足的。这会导致神经网络训练的过拟合,导致性能下降。事实上,成功的情感识别需要有效的预处理方法和有效使用权重参数数量的模型结构。在这项研究中,我们提出使用八个数据集版本具有不同的频率-时间分辨率,以寻找一个有效的情感语音预处理方法。我们提出了一个6层卷积神经网络(CNN)模型与有效的通道注意力(ECA)追求一个有效的模型结构。特别是,位置良好的ECA块只需几个参数就可以改善通道特征表示。在交互式情绪二元运动捕捉(IEMOCAP)数据集上,提高情感语音预处理的频率分辨率可以提高情感识别性能。此外,深度卷积层之后的ECA可以有效地增加信道特征表示。因此,最好的结果(79.37UA 79.68WA),可以得到超过以前的SER模型的性能。此外,为了弥补情感语音数据的不足,我们尝试了多种预处理数据方法,这些方法可以增强用一个样本中所有不同设置预处理的可训练数据。在实验中,我们可以获得最高的结果(80.28UA 80.46WA)。摘要:Speech emotion recognition (SER) classifies human emotions in speech with a computer model. Recently, performance in SER has steadily increased as deep learning techniques have adapted. However, unlike many domains that use speech data, data for training in the SER model is insufficient. This causes overfitting of training of the neural network, resulting in performance degradation. In fact, successful emotion recognition requires an effective preprocessing method and a model structure that efficiently uses the number of weight parameters. In this study, we propose using eight dataset versions with different frequency-time resolutions to search for an effective emotional speech preprocessing method. We propose a 6-layer convolutional neural network (CNN) model with efficient channel attention (ECA) to pursue an efficient model structure. In particular, the well-positioned ECA blocks can improve channel feature representation with only a few parameters. With the interactive emotional dyadic motion capture (IEMOCAP) dataset, increasing the frequency resolution in preprocessing emotional speech can improve emotion recognition performance. Also, ECA after the deep convolution layer can effectively increase channel feature representation. Consequently, the best result (79.37UA 79.68WA) can be obtained, exceeding the performance of previous SER models. Furthermore, to compensate for the lack of emotional speech data, we experiment with multiple preprocessing data methods that augment trainable data preprocessed with all different settings from one sample. In the experiment, we can achieve the highest result (80.28UA 80.46WA).

【6】 MetaBGM: Dynamic Soundtrack Transformation For Continuous Multi-Scene Experiences With Ambient Awareness And Personalization
标题: MetaCGM:动态原声转换,实现具有环境感知和个性化的连续多场景体验
作者:Haoxuan Liu,Zihao Wang,Haorong Hong,Youwei Feng,Jiaxin Yu,Han Diao,Yunfei Xu,Kejun Zhang
链接:点击下载PDF文件
摘要:本文介绍了MetaBGM,这是一个开创性的框架,用于生成适应动态场景和实时用户交互的背景音乐。我们将多场景定义为环境上下文的变化,例如游戏设置或电影场景中的过渡。为了解决将后端数据转换为音频生成模型的音乐描述文本的挑战,MetaBGM采用了一种新颖的两阶段生成方法,将连续的场景和用户状态数据转换为这些文本,然后将其输入到音频生成模型中进行实时配乐创建。实验结果表明,MetaBGM有效地生成上下文相关的和动态的背景音乐的交互式应用程序。摘要:This paper introduces MetaBGM, a groundbreaking framework for generating background music that adapts to dynamic scenes and real-time user interactions. We define multi-scene as variations in environmental contexts, such as transitions in game settings or movie scenes. To tackle the challenge of converting backend data into music description texts for audio generation models, MetaBGM employs a novel two-stage generation approach that transforms continuous scene and user state data into these texts, which are then fed into an audio generation model for real-time soundtrack creation. Experimental results demonstrate that MetaBGM effectively generates contextually relevant and dynamic background music for interactive applications.


机器翻译,仅供参考