今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Learning Expressive Disentangled Speech Representations with Soft Speech  Units and Adversarial Style Augmentation
标题:使用软语音单元和对抗风格增强学习表达性解纠缠语音表示
链接:https://arxiv.org/abs/2405.00603
作者:Yimin Deng,Jianzong Wang,Xulong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:语音转换是指在保留语音内容信息的前提下,对源语音进行语音特征转换的任务。如今,自监督表示学习模型越来越多地用于内容提取。然而,在这些表示中,大量的隐藏说话人信息导致音色泄漏,而隐藏单元的韵律信息缺乏使用。为了解决这些问题,我们提出了一个新的框架表达的语音转换称为“SAVC”的基础上软语音单元从休伯特软。以软语音单元为输入,设计了一个属性编码器,分别提取语音的内容特征和韵律特征。具体来说,我们首先引入统计扰动施加对抗风格增强,以消除扬声器信息。然后,韵律隐式建模的软语音单元与知识蒸馏。实验结果表明,转换后语音的可懂度和自然度均优于前人的工作。
摘要:Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in these representations, a lot of hidden speaker information leads to timbre leakage while the prosodic information of hidden units lacks use. To address these issues, we propose a novel framework for expressive voice conversion called "SAVC" based on soft speech units from HuBert-soft. Taking soft speech units as input, we design an attribute encoder to extract content and prosody features respectively. Specifically, we first introduce statistic perturbation imposed by adversarial style augmentation to eliminate speaker information. Then the prosody is implicitly modeled on soft speech units with knowledge distillation. Experiment results show that the intelligibility and naturalness of converted speech outperform previous work.

【2】 Visual and audio scene classification for detecting discrepancies in  video: a baseline method and experimental protocol
标题:用于检测视频差异的视觉和音频场景分类:基线方法和实验协议
链接:https://arxiv.org/abs/2405.00384
作者:Konstantinos Apostolidis,Jakob Abesser,Luca Cuccovillo,Vasileios Mezaris
备注:Accepted for publication, 3rd ACM Int. Workshop on Multimedia AI against Disinformation (MAD'24) at ACM ICMR'24, June 10, 2024, Phuket, Thailand. This is the "accepted version"
摘要:本文提出了一种基线方法和实验协议的一个特定的内容验证问题:检测多媒体内容中的音频和视频模态之间的差异。我们首先设计和优化的视听场景分类器,与现有的分类基线,使用这两种方式进行比较。然后,通过将该分类器分别应用于音频和视觉模态,我们可以检测它们之间的场景类不一致。为了便于进一步的研究,并提供一个共同的评估平台,我们介绍了一个实验协议和基准数据集模拟这种不一致。我们的方法在场景分类和视听差异检测方面取得了最先进的结果,突出了其在内容验证应用中的潜力。
摘要:This paper presents a baseline approach and an experimental protocol for a specific content verification problem: detecting discrepancies between the audio and video modalities in multimedia content. We first design and optimize an audio-visual scene classifier, to compare with existing classification baselines that use both modalities. Then, by applying this classifier separately to the audio and the visual modality, we can detect scene-class inconsistencies between them. To facilitate further research and provide a common evaluation platform, we introduce an experimental protocol and a benchmark dataset simulating such inconsistencies. Our approach achieves state-of-the-art results in scene classification and promising outcomes in audio-visual discrepancies detection, highlighting its potential in content verification applications.

【3】 Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data  Manipulation
标题:基于距离采样的Paraphraser利用ChatGPT进行文本数据操作
链接:https://arxiv.org/abs/2405.00367
作者:Yoori Oh,Yoseob Han,Kyogu Lee
备注:Accepted at SIGIR 2024 short paper track
摘要:音频—语言检索的研究越来越受到人们的关注,其目标是建立音频和文本模态之间的相关性。然而,大多数音频文本配对数据集往往缺乏丰富的表达的文本数据相比,音频样本。音频文本数据集面临的一个重大挑战是,尽管音频样本不同,但仍存在相似或相同的字幕。因此,在多对一的映射条件下,音频文本数据集导致检索任务的性能较差。在本文中,我们提出了一种新的方法来解决音频语言检索任务中的数据不平衡问题。为了克服这一局限性,我们引入了一种方法,该方法采用基于距离采样的释义器,利用ChatGPT,利用距离函数来生成可控分布的操纵文本数据。对于具有相同上下文的一组句子,距离用于计算任意两个句子的操纵程度,并且ChatGPT的Few-Shot提示使用具有由Jaccard相似度定义的相似距离的文本聚类来执行。因此,ChatGPT在应用于具有文本聚类的Few-Shot提示时,可以基于距离来调整所操作的文本的多样性。所提出的方法显着提高性能的音频文本检索,优于传统的文本增强技术。
摘要:There has been growing interest in audio-language retrieval research, where the objective is to establish the correlation between audio and text modalities. However, most audio-text paired datasets often lack rich expression of the text data compared to the audio samples. One of the significant challenges facing audio-text datasets is the presence of similar or identical captions despite different audio samples. Therefore, under many-to-one mapping conditions, audio-text datasets lead to poor performance of retrieval tasks. In this paper, we propose a novel approach to tackle the data imbalance problem in audio-language retrieval task. To overcome the limitation, we introduce a method that employs a distance sampling-based paraphraser leveraging ChatGPT, utilizing distance function to generate a controllable distribution of manipulated text data. For a set of sentences with the same context, the distance is used to calculate a degree of manipulation for any two sentences, and ChatGPT's few-shot prompting is performed using a text cluster with a similar distance defined by the Jaccard similarity. Therefore, ChatGPT, when applied to few-shot prompting with text clusters, can adjust the diversity of the manipulated text based on the distance. The proposed approach is shown to significantly enhance performance in audio-text retrieval, outperforming conventional text augmentation techniques.

【4】 Active Learning with Task Adaptation Pre-training for Speech Emotion  Recognition
标题:语音情感识别的任务适应预训练的主动学习
链接:https://arxiv.org/abs/2405.00307
作者:Dongyuan Li,Ying Zhang,Yusong Wang,Funakoshi Kataro,Manabu Okumura
备注:Accepted by Journal of Natural Language Processing. arXiv admin note: text overlap with arXiv:2310.00283
摘要:语音情感识别(SER)因其在人机交互、虚拟助手、心理健康辅助等领域的广泛应用而受到越来越多的关注。然而,现有的SER方法往往忽略了预训练语音识别任务和下游SER任务之间的信息差距,导致次优性能。此外,目前的方法需要大量的时间对每个特定的语音数据集进行微调,如IEMOCAP,这限制了它们在具有大规模噪声数据的真实世界场景中的有效性。为了解决这些问题,我们提出了一个基于主动学习(AL)的SER微调框架,称为\textsc {After},它利用任务适应预训练(TAPT)和AL方法来提高性能和效率。具体来说,我们首先使用TAPT来最小化预训练语音识别任务和下游语音情感识别任务之间的信息差距。然后,AL方法被用来迭代地选择一个子集的最具信息性和多样性的样本进行微调,从而减少时间消耗。实验表明,该方法只使用了20%的样本,准确率提高了8.45%,时间消耗减少了79%. \textsc {After}和消融研究的额外扩展进一步证实了其有效性和对各种真实场景的适用性。我们的源代码可以在Github上获得,以便重复使用。(https://github.com/Clearloveyuan/AFTER).
摘要:Speech emotion recognition (SER) has garnered increasing attention due to its wide range of applications in various fields, including human-machine interaction, virtual assistants, and mental health assistance. However, existing SER methods often overlook the information gap between the pre-training speech recognition task and the downstream SER task, resulting in sub-optimal performance. Moreover, current methods require much time for fine-tuning on each specific speech dataset, such as IEMOCAP, which limits their effectiveness in real-world scenarios with large-scale noisy data. To address these issues, we propose an active learning (AL)-based fine-tuning framework for SER, called \textsc{After}, that leverages task adaptation pre-training (TAPT) and AL methods to enhance performance and efficiency. Specifically, we first use TAPT to minimize the information gap between the pre-training speech recognition task and the downstream speech emotion recognition task. Then, AL methods are employed to iteratively select a subset of the most informative and diverse samples for fine-tuning, thereby reducing time consumption. Experiments demonstrate that our proposed method \textsc{After}, using only 20\% of samples, improves accuracy by 8.45\% and reduces time consumption by 79\%. The additional extension of \textsc{After} and ablation studies further confirm its effectiveness and applicability to various real-world scenarios. Our source code is available on Github for reproducibility. (https://github.com/Clearloveyuan/AFTER).


【5】 Who is Authentic Speaker
标题:谁是真正的演讲者
链接:https://arxiv.org/abs/2405.00248
作者:Qiang Huang
摘要:使用深度学习技术的语音转换(VC)现在可以生成高质量的一对多语音,因此已被用于娱乐和医疗保健等实际应用领域。然而,当操纵的声音被用于欺骗目的时,声音转换可能会造成潜在的社会问题。此外,由于源说话人的声学特性发生了很大的变化,因此从转换后的语音中发现谁是真正的说话人是一个很大的挑战。在本文中,我们试图探讨的可行性,确定真实的发言人从转换的声音。这项研究是在这样一个假设下进行的:即使当源说话人的声音转换成不同的目标声音时,源说话人的某些信息仍然存在。因此,我们的实验是面向识别源扬声器的转换后的声音,这是通过使用FragmentVC从源和目标扬声器的随机配对的话语产生的。为了提高对转换语音的鲁棒性,我们的识别模型是通过在深度神经网络中使用局部聚合描述符(VLAD)的分层向量来构建的。本文主要从两个方面对真实说话人识别系统进行了测试,包括转换语音质量的影响和VLAD的变化。在这项工作中使用的数据集是VCTK语料库,其中源和目标扬声器随机配对。转换后的话语上得到的结果显示出有前途的性能,在识别真实的扬声器从转换的声音。
摘要:Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice conversion can pose potential social issues when manipulated voices are employed for deceptive purposes. Moreover, it is a big challenge to find who are real speakers from the converted voices as the acoustic characteristics of source speakers are changed greatly. In this paper we attempt to explore the feasibility of identifying authentic speakers from converted voices. This study is conducted with the assumption that certain information from the source speakers persists, even when their voices undergo conversion into different target voices. Therefore our experiments are geared towards recognising the source speakers given the converted voices, which are generated by using FragmentVC on the randomly paired utterances from source and target speakers. To improve the robustness against converted voices, our recognition model is constructed by using hierarchical vector of locally aggregated descriptors (VLAD) in deep neural networks. The authentic speaker recognition system is mainly tested in two aspects, including the impact of quality of converted voices and the variations of VLAD. The dataset used in this work is VCTK corpus, where source and target speakers are randomly paired. The results obtained on the converted utterances show promising performances in recognising authentic speakers from converted voices.


【6】 SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General  Sound
标题:SemantiCodec:用于通用声音的超低比特率语义音频编解码器
链接:https://arxiv.org/abs/2405.00233
作者:Haohe Liu,Xuenan Xu,Yi Yuan,Mengyue Wu,Wenwu Wang,Mark D. Plumbley
备注:Demo and code: this https URL
摘要:大型语言模型(LLM)通过将音频转换为离散令牌的音频编解码器大大提高了音频处理,从而能够将语言建模技术应用于音频数据。然而,传统的编解码器通常以高比特率或在语音等狭窄的领域内运行,并且缺乏有效语言建模所需的语义线索。为了应对这些挑战,我们引入了SemantiCodec,这是一种新颖的编解码器,旨在将音频压缩为每秒不到100个令牌,用于各种音频类型,包括语音,通用音频和音乐,而不会影响质量。SemantiCodec具有双编码器架构:使用自监督AudioMAE的语义编码器,使用k均值聚类对大量音频数据进行离散化,以及捕获剩余细节的声学编码器。语义和声学编码器输出用于经由基于扩散模型的解码器重构音频。SemantiCodec有三种变体,令牌速率分别为每秒25、50和100,支持0.31 kbps至1.43 kbps的超低比特率范围。实验结果表明,SemantiCodec的重建质量显着优于国家的最先进的描述编解码器。我们的研究结果还表明,SemantiCodec包含显着丰富的语义信息比所有评估的音频编解码器,即使在显着较低的比特率。我们的代码和演示可以在https://haoheliu.github.io/SemantiCodec/上找到。
摘要:Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelling techniques to audio data. However, traditional codecs often operate at high bitrates or within narrow domains such as speech and lack the semantic clues required for efficient language modelling. Addressing these challenges, we introduce SemantiCodec, a novel codec designed to compress audio into fewer than a hundred tokens per second across diverse audio types, including speech, general audio, and music, without compromising quality. SemantiCodec features a dual-encoder architecture: a semantic encoder using a self-supervised AudioMAE, discretized using k-means clustering on extensive audio data, and an acoustic encoder to capture the remaining details. The semantic and acoustic encoder outputs are used to reconstruct audio via a diffusion-model-based decoder. SemantiCodec is presented in three variants with token rates of 25, 50, and 100 per second, supporting a range of ultra-low bit rates between 0.31 kbps and 1.43 kbps. Experimental results demonstrate that SemantiCodec significantly outperforms the state-of-the-art Descript codec on reconstruction quality. Our results also suggest that SemantiCodec contains significantly richer semantic information than all evaluated audio codecs, even at significantly lower bitrates. Our code and demos are available at https://haoheliu.github.io/SemantiCodec/.

eess.AS音频处理
【1】 Learning Expressive Disentangled Speech Representations with Soft Speech  Units and Adversarial Style Augmentation
标题:使用软语音单元和对抗风格增强学习表达性解纠缠语音表示
链接:https://arxiv.org/abs/2405.00603
作者:Yimin Deng,Jianzong Wang,Xulong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:语音转换是指在保留语音内容信息的前提下,对源语音进行语音特征转换的任务。如今,自监督表示学习模型越来越多地用于内容提取。然而,在这些表示中,大量的隐藏说话人信息导致音色泄漏,而隐藏单元的韵律信息缺乏使用。为了解决这些问题,我们提出了一个新的框架表达的语音转换称为“SAVC”的基础上软语音单元从休伯特软。以软语音单元为输入,设计了一个属性编码器,分别提取语音的内容特征和韵律特征。具体来说,我们首先引入统计扰动施加对抗风格增强,以消除扬声器信息。然后,韵律隐式建模的软语音单元与知识蒸馏。实验结果表明,转换后语音的可懂度和自然度均优于前人的工作。
摘要:Voice conversion is the task to transform voice characteristics of source speech while preserving content information. Nowadays, self-supervised representation learning models are increasingly utilized in content extraction. However, in these representations, a lot of hidden speaker information leads to timbre leakage while the prosodic information of hidden units lacks use. To address these issues, we propose a novel framework for expressive voice conversion called "SAVC" based on soft speech units from HuBert-soft. Taking soft speech units as input, we design an attribute encoder to extract content and prosody features respectively. Specifically, we first introduce statistic perturbation imposed by adversarial style augmentation to eliminate speaker information. Then the prosody is implicitly modeled on soft speech units with knowledge distillation. Experiment results show that the intelligibility and naturalness of converted speech outperform previous work.


【2】 Visual and audio scene classification for detecting discrepancies in  video: a baseline method and experimental protocol
标题:用于检测视频差异的视觉和音频场景分类:基线方法和实验协议
链接:https://arxiv.org/abs/2405.00384
作者:Konstantinos Apostolidis,Jakob Abesser,Luca Cuccovillo,Vasileios Mezaris
备注:Accepted for publication, 3rd ACM Int. Workshop on Multimedia AI against Disinformation (MAD'24) at ACM ICMR'24, June 10, 2024, Phuket, Thailand. This is the "accepted version"
摘要:本文提出了一种基线方法和实验协议的一个特定的内容验证问题:检测多媒体内容中的音频和视频模态之间的差异。我们首先设计和优化的视听场景分类器,与现有的分类基线,使用这两种方式进行比较。然后,通过将该分类器分别应用于音频和视觉模态,我们可以检测它们之间的场景类不一致。为了便于进一步的研究,并提供一个共同的评估平台,我们介绍了一个实验协议和基准数据集模拟这种不一致。我们的方法在场景分类和视听差异检测方面取得了最先进的结果,突出了其在内容验证应用中的潜力。
摘要:This paper presents a baseline approach and an experimental protocol for a specific content verification problem: detecting discrepancies between the audio and video modalities in multimedia content. We first design and optimize an audio-visual scene classifier, to compare with existing classification baselines that use both modalities. Then, by applying this classifier separately to the audio and the visual modality, we can detect scene-class inconsistencies between them. To facilitate further research and provide a common evaluation platform, we introduce an experimental protocol and a benchmark dataset simulating such inconsistencies. Our approach achieves state-of-the-art results in scene classification and promising outcomes in audio-visual discrepancies detection, highlighting its potential in content verification applications.

【3】 Distance Sampling-based Paraphraser Leveraging ChatGPT for Text Data  Manipulation
标题:基于距离采样的Paraphraser利用ChatGPT进行文本数据操作
链接:https://arxiv.org/abs/2405.00367
作者:Yoori Oh,Yoseob Han,Kyogu Lee
备注:Accepted at SIGIR 2024 short paper track
摘要:音频-语言检索的研究越来越受到人们的关注,其目标是建立音频和文本模态之间的相关性。然而,大多数音频文本配对数据集往往缺乏丰富的表达的文本数据相比,音频样本。音频文本数据集面临的一个重大挑战是,尽管音频样本不同,但仍存在相似或相同的字幕。因此,在多对一的映射条件下,音频文本数据集导致检索任务的性能较差。在本文中,我们提出了一种新的方法来解决音频语言检索任务中的数据不平衡问题。为了克服这一局限性,我们引入了一种方法,该方法采用基于距离采样的释义器,利用ChatGPT,利用距离函数来生成可控分布的操纵文本数据。对于具有相同上下文的一组句子,距离用于计算任意两个句子的操纵程度,并且ChatGPT的Few-Shot提示使用具有由Jaccard相似度定义的相似距离的文本聚类来执行。因此,ChatGPT在应用于具有文本聚类的Few-Shot提示时,可以基于距离来调整所操作的文本的多样性。所提出的方法显着提高性能的音频文本检索,优于传统的文本增强技术。
摘要:There has been growing interest in audio-language retrieval research, where the objective is to establish the correlation between audio and text modalities. However, most audio-text paired datasets often lack rich expression of the text data compared to the audio samples. One of the significant challenges facing audio-text datasets is the presence of similar or identical captions despite different audio samples. Therefore, under many-to-one mapping conditions, audio-text datasets lead to poor performance of retrieval tasks. In this paper, we propose a novel approach to tackle the data imbalance problem in audio-language retrieval task. To overcome the limitation, we introduce a method that employs a distance sampling-based paraphraser leveraging ChatGPT, utilizing distance function to generate a controllable distribution of manipulated text data. For a set of sentences with the same context, the distance is used to calculate a degree of manipulation for any two sentences, and ChatGPT's few-shot prompting is performed using a text cluster with a similar distance defined by the Jaccard similarity. Therefore, ChatGPT, when applied to few-shot prompting with text clusters, can adjust the diversity of the manipulated text based on the distance. The proposed approach is shown to significantly enhance performance in audio-text retrieval, outperforming conventional text augmentation techniques.

【4】 Active Learning with Task Adaptation Pre-training for Speech Emotion  Recognition
标题:语音情感识别的任务适应预训练的主动学习
链接:https://arxiv.org/abs/2405.00307
作者:Dongyuan Li,Ying Zhang,Yusong Wang,Funakoshi Kataro,Manabu Okumura
备注:Accepted by Journal of Natural Language Processing. arXiv admin note: text overlap with arXiv:2310.00283
摘要:语音情感识别(SER)因其在人机交互、虚拟助手、心理健康辅助等领域的广泛应用而受到越来越多的关注。然而,现有的SER方法往往忽略了预训练语音识别任务和下游SER任务之间的信息差距,导致次优性能。此外,目前的方法需要大量的时间对每个特定的语音数据集进行微调,如IEMOCAP,这限制了它们在具有大规模噪声数据的真实世界场景中的有效性。为了解决这些问题,我们提出了一个基于主动学习(AL)的SER微调框架,称为\textsc{After},它利用任务适应预训练(TAPT)和AL方法来提高性能和效率。具体来说,我们首先使用TAPT来最小化预训练语音识别任务和下游语音情感识别任务之间的信息差距。然后,AL方法被用来迭代地选择一个子集的最具信息性和多样性的样本进行微调,从而减少时间消耗。实验表明,该方法只使用了20%的样本,准确率提高了8.45%,时间消耗减少了79%. \textsc{After}和消融研究的额外扩展进一步证实了其有效性和对各种真实场景的适用性。我们的源代码可以在Github上获得,以便重复使用。(https://github.com/Clearloveyuan/AFTER).
摘要:Speech emotion recognition (SER) has garnered increasing attention due to its wide range of applications in various fields, including human-machine interaction, virtual assistants, and mental health assistance. However, existing SER methods often overlook the information gap between the pre-training speech recognition task and the downstream SER task, resulting in sub-optimal performance. Moreover, current methods require much time for fine-tuning on each specific speech dataset, such as IEMOCAP, which limits their effectiveness in real-world scenarios with large-scale noisy data. To address these issues, we propose an active learning (AL)-based fine-tuning framework for SER, called \textsc{After}, that leverages task adaptation pre-training (TAPT) and AL methods to enhance performance and efficiency. Specifically, we first use TAPT to minimize the information gap between the pre-training speech recognition task and the downstream speech emotion recognition task. Then, AL methods are employed to iteratively select a subset of the most informative and diverse samples for fine-tuning, thereby reducing time consumption. Experiments demonstrate that our proposed method \textsc{After}, using only 20\% of samples, improves accuracy by 8.45\% and reduces time consumption by 79\%. The additional extension of \textsc{After} and ablation studies further confirm its effectiveness and applicability to various real-world scenarios. Our source code is available on Github for reproducibility. (https://github.com/Clearloveyuan/AFTER).


【5】 Who is Authentic Speaker
标题:谁是真正的演讲者
链接:https://arxiv.org/abs/2405.00248
作者:Qiang Huang
摘要:使用深度学习技术的语音转换(VC)现在可以生成高质量的一对多语音,因此已被用于娱乐和医疗保健等实际应用领域。然而,当操纵的声音被用于欺骗目的时,声音转换可能会造成潜在的社会问题。此外,由于源说话人的声学特性发生了很大的变化,因此从转换后的语音中发现谁是真正的说话人是一个很大的挑战。在本文中,我们试图探讨的可行性,确定真实的发言人从转换的声音。这项研究是在这样一个假设下进行的:即使当源说话人的声音转换成不同的目标声音时,源说话人的某些信息仍然存在。因此,我们的实验是面向识别源扬声器的转换后的声音,这是通过使用FragmentVC从源和目标扬声器的随机配对的话语产生的。为了提高对转换语音的鲁棒性,我们的识别模型是通过在深度神经网络中使用局部聚合描述符(VLAD)的分层向量来构建的。本文主要从两个方面对真实说话人识别系统进行了测试,包括转换语音质量的影响和VLAD的变化。在这项工作中使用的数据集是VCTK语料库,其中源和目标扬声器随机配对。转换后的话语上得到的结果显示出有前途的性能,在识别真实的扬声器从转换的声音。
摘要:Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice conversion can pose potential social issues when manipulated voices are employed for deceptive purposes. Moreover, it is a big challenge to find who are real speakers from the converted voices as the acoustic characteristics of source speakers are changed greatly. In this paper we attempt to explore the feasibility of identifying authentic speakers from converted voices. This study is conducted with the assumption that certain information from the source speakers persists, even when their voices undergo conversion into different target voices. Therefore our experiments are geared towards recognising the source speakers given the converted voices, which are generated by using FragmentVC on the randomly paired utterances from source and target speakers. To improve the robustness against converted voices, our recognition model is constructed by using hierarchical vector of locally aggregated descriptors (VLAD) in deep neural networks. The authentic speaker recognition system is mainly tested in two aspects, including the impact of quality of converted voices and the variations of VLAD. The dataset used in this work is VCTK corpus, where source and target speakers are randomly paired. The results obtained on the converted utterances show promising performances in recognising authentic speakers from converted voices.

【6】 SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General  Sound
标题:SemantiCodec:用于通用声音的超低比特率语义音频编解码器
链接:https://arxiv.org/abs/2405.00233
作者:Haohe Liu,Xuenan Xu,Yi Yuan,Mengyue Wu,Wenwu Wang,Mark D. Plumbley
备注:Demo and code: this https URL
摘要:大型语言模型(LLM)通过将音频转换为离散令牌的音频编解码器大大提高了音频处理,从而能够将语言建模技术应用于音频数据。然而,传统的编解码器通常以高比特率或在语音等狭窄的领域内运行,并且缺乏有效语言建模所需的语义线索。为了应对这些挑战,我们引入了SemantiCodec,这是一种新颖的编解码器,旨在将音频压缩为每秒不到100个令牌,用于各种音频类型,包括语音,通用音频和音乐,而不会影响质量。SemantiCodec具有双编码器架构:使用自监督AudioMAE的语义编码器,使用k均值聚类对大量音频数据进行离散化,以及捕获剩余细节的声学编码器。语义和声学编码器输出用于经由基于扩散模型的解码器重构音频。SemantiCodec有三种变体,令牌速率分别为每秒25、50和100,支持0.31 kbps至1.43 kbps的超低比特率范围。实验结果表明,SemantiCodec的重建质量显着优于国家的最先进的描述编解码器。我们的研究结果还表明,SemantiCodec包含显着丰富的语义信息比所有评估的音频编解码器,即使在显着较低的比特率。我们的代码和演示可以在https://haoheliu.github.io/SemantiCodec/上找到。
摘要:Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modelling techniques to audio data. However, traditional codecs often operate at high bitrates or within narrow domains such as speech and lack the semantic clues required for efficient language modelling. Addressing these challenges, we introduce SemantiCodec, a novel codec designed to compress audio into fewer than a hundred tokens per second across diverse audio types, including speech, general audio, and music, without compromising quality. SemantiCodec features a dual-encoder architecture: a semantic encoder using a self-supervised AudioMAE, discretized using k-means clustering on extensive audio data, and an acoustic encoder to capture the remaining details. The semantic and acoustic encoder outputs are used to reconstruct audio via a diffusion-model-based decoder. SemantiCodec is presented in three variants with token rates of 25, 50, and 100 per second, supporting a range of ultra-low bit rates between 0.31 kbps and 1.43 kbps. Experimental results demonstrate that SemantiCodec significantly outperforms the state-of-the-art Descript codec on reconstruction quality. Our results also suggest that SemantiCodec contains significantly richer semantic information than all evaluated audio codecs, even at significantly lower bitrates. Our code and demos are available at https://haoheliu.github.io/SemantiCodec/.

机器翻译由腾讯交互翻译提供,仅供参考