今日论文合集:cs.SD语音8篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Fake it to make it: Using synthetic data to remedy the data shortage in  joint multimodal speech-and-gesture synthesis

标题:伪造它来制造它:使用合成数据来弥补联合多模式语音和手势合成中的数据短缺

链接:https://arxiv.org/abs/2404.19622

作者:Shivam Mehta,Anna Deichler,Jim O'Regan,Birger Moëll,Jonas Beskow,Gustav Eje Henter,Simon Alexanderson

备注:13+1 pages, 2 figures, accepted at the Human Motion Generation workshop (HuMoGen) at CVPR 2024

摘要:虽然参与面对面交谈的人同时进行口头和非口头通信,但用于从文本联合和统一合成语音音频和共语音3D手势运动的方法是一个新的新兴领域。这些技术为更人性化、高效、富有表现力和强大的合成通信带来了巨大的希望,但目前由于缺乏适当的大数据集而受到阻碍,因为现有方法是在来自所有组成模态的并行数据上训练的。受学生—教师方法的启发,我们提出了一个简单的解决数据短缺的方法,通过简单地合成额外的培训材料。具体来说,我们使用在大型数据集上训练的单峰合成模型来创建多峰(但合成)并行训练数据,然后在该材料上预训练联合合成模型。此外,我们提出了一个新的合成架构,增加了更好的和更可控的韵律建模的国家的最先进的方法在该领域。我们的研究结果证实,对大量合成数据进行预训练可以提高多模态模型合成的语音和运动的质量,并且在对合成数据进行预训练时,所提出的架构可以带来进一步的好处。有关输出示例,请参见www.example.com。

摘要:Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speech 3D gesture motion from text are a new and emerging field. These technologies hold great promise for more human-like, efficient, expressive, and robust synthetic communication, but are currently held back by the lack of suitably large datasets, as existing methods are trained on parallel data from all constituent modalities. Inspired by student-teacher methods, we propose a straightforward solution to the data shortage, by simply synthesising additional training material. Specifically, we use unimodal synthesis models trained on large datasets to create multimodal (but synthetic) parallel training data, and then pre-train a joint synthesis model on that material. In addition, we propose a new synthesis architecture that adds better and more controllable prosody modelling to the state-of-the-art method in the field. Our results confirm that pre-training on large amounts of synthetic data improves the quality of both the speech and the motion synthesised by the multimodal model, with the proposed architecture yielding further benefits when pre-trained on the synthetic data. See https://shivammehta25.github.io/MAGI/ for example output.


【2】 SemiPL: A Semi-supervised Method for Event Sound Source Localization
标题:SemIPL:一种半监督的事件声音源定位方法
链接:https://arxiv.org/abs/2404.19615
作者:Yue Li,Baiqiao Yin,Jinfu Liu,Jiajun Wen,Jiaying Lin,Mengyuan Liu
摘要:近年来,事件声源定位在各个领域得到了广泛的应用。最近的作品通常依赖于对比学习框架显示令人印象深刻的性能。然而,所有的工作都是基于相对简单的大型数据集。在许多应用中,理解和分析人类行为(人的动作和交互)、声音和混沌事件中的声音也至关重要,例如,人群管理和紧急响应服务。本文将现有的模型应用到一个更复杂的数据集上,探讨了参数对模型的影响,并提出了一种半监督的改进方法SemiPL。随着数据量的增加和标签质量的影响,自监督学习将是不可阻挡的趋势。实验表明,参数调整将积极影响现有的模型。特别是,与提供的结果相比,SSPL在混沌世界中实现了12.2% cIoU和0.56% AUC的改善。代码可在以下网址获得:www.example.com
摘要:In recent years, Event Sound Source Localization has been widely applied in various fields. Recent works typically relying on the contrastive learning framework show impressive performance. However, all work is based on large relatively simple datasets. It's also crucial to understand and analyze human behaviors (actions and interactions of people), voices, and sounds in chaotic events in many applications, e.g., crowd management, and emergency response services. In this paper, we apply the existing model to a more complex dataset, explore the influence of parameters on the model, and propose a semi-supervised improvement method SemiPL. With the increase in data quantity and the influence of label quality, self-supervised learning will be an unstoppable trend. The experiment shows that the parameter adjustment will positively affect the existing model. In particular, SSPL achieved an improvement of 12.2% cIoU and 0.56% AUC in Chaotic World compared to the results provided. The code is available at: https://github.com/ly245422/SSPL


【3】 ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized  Transformers
标题:ESM:使用跨尺度残留量量化变换器的高效语音编码
链接:https://arxiv.org/abs/2404.19441
作者:Yuzhe Gu,Enmao Diao
备注:Preprint
摘要:现有的神经音频编解码器通常为了音频质量而牺牲计算复杂度。它们主要在卷积块上构建特征变换层,卷积块本身并不适合捕获音频信号的局部冗余。作为补偿,无论是对抗损失从一个编码器或大量的模型参数,需要改善编解码器。为此,我们提出了高效的语音编解码器(ESC),一个轻量级的参数高效的编解码器跨尺度残差矢量量化和Transformers。我们的模型利用镜像分层窗口注意Transformer块,并执行从粗到细的特征表示逐步解码。为了提高码本利用率,我们设计了一个学习范式,涉及一个预训练阶段,以协助编解码器的训练。大量的实验结果表明,ESC能够以较低的复杂度获得较高的音频质量,是替代现有编解码器的一种有前景的方案。
摘要:Existing neural audio codecs usually sacrifice computational complexity for audio quality. They build the feature transformation layers mainly on convolutional blocks, which are not inherently appropriate for capturing local redundancies of audio signals. As compensation, either adversarial losses from a discriminator or a large number of model parameters are required to improve the codec. To that end, we propose Efficient Speech Codec (ESC), a lightweight parameter-efficient codec laid on cross-scale residual vector quantization and transformers. Our model leverages mirrored hierarchical window-attention transformer blocks and performs step-wise decoding from coarse-to-fine feature representations. To enhance codebook utilization, we design a learning paradigm that involves a pre-training stage to assist with codec training. Extensive results show that ESC can achieve high audio quality with much lower complexity, which is a prospective alternative in place of existing codecs.

【4】 EfficientASR: Speech Recognition Network Compression via Attention  Redundancy and Chunk-Level FFN Optimization
标题:EfficientASB:通过注意力冗余和块级FFN优化的语音识别网络压缩
链接:https://arxiv.org/abs/2404.19214
作者:Jianzong Wang,Ziqi Liang,Xulong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:近年来,Transformer网络在语音识别任务中表现出了卓越的性能。然而,由于高计算和存储资源需求,它们的部署带来了挑战。为了解决这个问题,本文提出了一个轻量级的模型称为EfficientASR,旨在提高Transformer模型的通用性。EfficientASR采用两个主要模块:共享剩余多头注意力(SRMHA)和块级前馈网络(CFFN)。SRMHA模块有效地减少了网络中的冗余计算,而CFFN模块捕获空间知识并减少了参数的数量。EfficientASR模型的有效性在两个公共数据集上进行了验证,即Aishell—1和HKUST。实验结果表明,与基线Transformer网络相比,参数减少了36%,Aishell—1和HKUST数据集上的字符错误率(CER)分别提高了0.3%和0.2%。
摘要:In recent years, Transformer networks have shown remarkable performance in speech recognition tasks. However, their deployment poses challenges due to high computational and storage resource requirements. To address this issue, a lightweight model called EfficientASR is proposed in this paper, aiming to enhance the versatility of Transformer models. EfficientASR employs two primary modules: Shared Residual Multi-Head Attention (SRMHA) and Chunk-Level Feedforward Networks (CFFN). The SRMHA module effectively reduces redundant computations in the network, while the CFFN module captures spatial knowledge and reduces the number of parameters. The effectiveness of the EfficientASR model is validated on two public datasets, namely Aishell-1 and HKUST. Experimental results demonstrate a 36% reduction in parameters compared to the baseline Transformer network, along with improvements of 0.3% and 0.2% in Character Error Rate (CER) on the Aishell-1 and HKUST datasets, respectively.


【5】 EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with  IFUB Estimator and Joint Text-Guided Consistent Learning
标题:EAD-VC:利用IFLU估计器和联合文本引导一致学习增强语音转换的语音自动解纠缠
链接:https://arxiv.org/abs/2404.19212
作者:Ziqi Liang,Jianzong Wang,Xulong Zhang,Yong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:利用无监督学习将语音分解为内容、节奏、音高和音色进行语音转换已经成为一个热门的研究课题。现有的工作一般考虑到解开语音成分通过人为制造的瓶颈特征,不能实现足够的信息解开,而音高和节奏仍然可能混合在一起。在解缠过程中存在信息重叠的风险,这导致语音自然度降低。为了克服这些限制,我们提出了一个两阶段的模型来解开语音表示在一个自我监督的方式没有人为的瓶颈设计,它使用的互信息(MI)与设计的上限估计(IFUB)分离语音组件之间的重叠信息。此外,我们设计了一个联合文本引导一致(TGC)模块,以指导语音内容的提取和消除音色泄漏问题。实验表明,我们的模型可以实现更好的性能比基线,关于解纠缠的有效性,语音自然度,和相似性。音频样本可以在https://largeaudiomodel.com/eadvc上找到。
摘要:Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into account disentangling speech components through human-crafted bottleneck features which can not achieve sufficient information disentangling, while pitch and rhythm may still be mixed together. There is a risk of information overlap in the disentangling process which results in less speech naturalness. To overcome such limits, we propose a two-stage model to disentangle speech representations in a self-supervised manner without a human-crafted bottleneck design, which uses the Mutual Information (MI) with the designed upper bound estimator (IFUB) to separate overlapping information between speech components. Moreover, we design a Joint Text-Guided Consistent (TGC) module to guide the extraction of speech content and eliminate timbre leakage issues. Experiments show that our model can achieve a better performance than the baseline, regarding disentanglement effectiveness, speech naturalness, and similarity. Audio samples can be found at https://largeaudiomodel.com/eadvc.


【6】 CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness  Condition
标题:指挥者:用音调和表现力来美化歌声
链接:https://arxiv.org/abs/2404.19187
作者:Jianzong Wang,Pengcheng Li,Xulong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:歌唱嗓音矫正是一项在人们日常生活中具有应用价值的新课题,其目的是在不改变原有音色和内容的前提下,矫正歌唱声音的音高,提高歌唱声音的表现力。现有的方法依赖于成对的数据或仅集中于音高的校正。但是,专业歌曲和业余歌曲很难从同一个人那里获得,而且歌唱声音的调整不仅仅是音高的调整,还包括情感和节奏等方面。由于我们提出了一个快速和高保真的歌唱语音识别系统称为ConTuner,一个扩散模型结合修改的条件,以产生美化的梅尔频谱图,其中修改的条件是由优化的音高和表现力。对于基音周期校正,我们建立了基音周期、频谱包络到基音周期的映射关系。为了使业余歌唱更有表现力,我们提出了在潜在空间中的表现力增强器,将业余声乐音调转换为专业声乐音调。ConTuner对中文和英文歌曲都达到了令人满意的美化效果。消融实验表明,ConTuner中的表现力增强器和基于生成器的加速方法是有效的。
摘要:Singing voice beautifying is a novel task that has application value in people's daily life, aiming to correct the pitch of the singing voice and improve the expressiveness without changing the original timbre and content. Existing methods rely on paired data or only concentrate on the correction of pitch. However, professional songs and amateur songs from the same person are hard to obtain, and singing voice beautifying doesn't only contain pitch correction but other aspects like emotion and rhythm. Since we propose a fast and high-fidelity singing voice beautifying system called ConTuner, a diffusion model combined with the modified condition to generate the beautified Mel-spectrogram, where the modified condition is composed of optimized pitch and expressiveness. For pitch correction, we establish a mapping relationship from MIDI, spectrum envelope to pitch. To make amateur singing more expressive, we propose the expressiveness enhancer in the latent space to convert amateur vocal tone to professional. ConTuner achieves a satisfactory beautification effect on both Mandarin and English songs. Ablation study demonstrates that the expressiveness enhancer and generator-based accelerate method in ConTuner are effective.


【7】 Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
标题:鲁棒纯解码器文本到语音的注意力限制推理
链接:https://arxiv.org/abs/2404.19723
作者:Hankun Wang,Chenpeng Du,Yiwei Guo,Shuai Wang,Xie Chen,Kai Yu
摘要:最近流行的仅解码器的文本到语音模型以其生成自然声音的语音的能力而闻名。然而,这样的模型有时会遭受单词跳过和重复,由于缺乏明确的单调对齐约束。在本文中,我们注意到从注意力地图,一些特定的注意力头的解码器只模型表示语音和文本之间的对齐。我们称这些头部的注意力地图为对齐浮现注意力地图(AEAM)。基于这一发现,我们提出了一种新的推理方法,而不改变训练过程,命名为注意力约束推理(ACI),以促进单调合成。它首先使用注意力扫描算法识别AEAM,然后在AEAM上应用约束掩码。我们的实验结果表明,解码器只TTS模型VALL—E的合成语音的误码率率降低了20.5%,而自然度和说话人相似度相当。
摘要:Recent popular decoder-only text-to-speech models are known for their ability of generating natural-sounding speech. However, such models sometimes suffer from word skipping and repeating due to the lack of explicit monotonic alignment constraints. In this paper, we notice from the attention maps that some particular attention heads of the decoder-only model indicate the alignments between speech and text. We call the attention maps of those heads Alignment-Emerged Attention Maps (AEAMs). Based on this discovery, we propose a novel inference method without altering the training process, named Attention-Constrained Inference (ACI), to facilitate monotonic synthesis. It first identifies AEAMs using the Attention Sweeping algorithm and then applies constraining masks on AEAMs. Our experimental results on decoder-only TTS model VALL-E show that the WER of synthesized speech is reduced by up to 20.5% relatively with ACI while the naturalness and speaker similarity are comparable.

【8】 Deep low-latency joint speech transmission and enhancement over a  gaussian channel
标题:高斯通道上的深度低延迟联合语音传输和增强
链接:https://arxiv.org/abs/2404.19375
作者:Mohammad Bokaei,Jesper Jensen,Simon Doclo,Jan Østergaard
摘要:确保听力辅助设备在低延迟场景中的可理解语音通信在语音增强、编码和传输方面提出了重大挑战。在本文中,我们提出了利用深度神经网络(DNN)进行低延迟联合语音传输和增强的新解决方案。我们的方法集成了两种最先进的DNN架构,用于低延迟语音增强和低延迟模拟基于源通道的联合传输,创建了一个组合的低延迟系统,并以端到端的方法联合训练这两个系统。由于增强系统的计算需求,当解码器中没有高计算能力时,例如听力辅助设备,该顺序是合适的。所提出的系统能够配置总延迟,即使在延迟低至3 ms时也能实现高性能,这通常是具有挑战性的。仿真结果提供了令人信服的证据,联合增强和传输系统是优于一个简单的级联系统在不同的设置,包括各种无线信道条件下,lavidity,和背景噪声的情况。
摘要:Ensuring intelligible speech communication for hearing assistive devices in low-latency scenarios presents significant challenges in terms of speech enhancement, coding and transmission. In this paper, we propose novel solutions for low-latency joint speech transmission and enhancement, leveraging deep neural networks (DNNs). Our approach integrates two state-of-the-art DNN architectures for low-latency speech enhancement and low-latency analog joint source-channel-based transmission, creating a combined low-latency system and jointly training both systems in an end-to-end approach. Due to the computational demands of the enhancement system, this order is suitable when high computational power is unavailable in the decoder, like hearing assistive devices. The proposed system enables the configuration of total latency, achieving high performance even at latencies as low as 3 ms, which is typically challenging to attain. The simulation results provide compelling evidence that a joint enhancement and transmission system is superior to a simple concatenation system in diverse settings, encompassing various wireless channel conditions, latencies, and background noise scenarios.

eess.AS音频处理
【1】 Attention-Constrained Inference for Robust Decoder-Only Text-to-Speech
标题:鲁棒纯解码器文本到语音的注意力限制推理
链接:https://arxiv.org/abs/2404.19723
作者:Hankun Wang,Chenpeng Du,Yiwei Guo,Shuai Wang,Xie Chen,Kai Yu
摘要:最近流行的仅解码器的文本到语音模型以其生成自然声音的语音的能力而闻名。然而,这样的模型有时会遭受单词跳过和重复,由于缺乏明确的单调对齐约束。在本文中,我们注意到从注意力地图,一些特定的注意力头的解码器只模型表示语音和文本之间的对齐。我们称这些头部的注意力地图为对齐浮现注意力地图(AEAM)。基于这一发现,我们提出了一种新的推理方法,而不改变训练过程,命名为注意力约束推理(ACI),以促进单调合成。它首先使用注意力扫描算法识别AEAM,然后在AEAM上应用约束掩码。我们的实验结果表明,解码器只TTS模型VALL-E的合成语音的误码率率降低了20.5%,而自然度和说话人相似度相当。
摘要:Recent popular decoder-only text-to-speech models are known for their ability of generating natural-sounding speech. However, such models sometimes suffer from word skipping and repeating due to the lack of explicit monotonic alignment constraints. In this paper, we notice from the attention maps that some particular attention heads of the decoder-only model indicate the alignments between speech and text. We call the attention maps of those heads Alignment-Emerged Attention Maps (AEAMs). Based on this discovery, we propose a novel inference method without altering the training process, named Attention-Constrained Inference (ACI), to facilitate monotonic synthesis. It first identifies AEAMs using the Attention Sweeping algorithm and then applies constraining masks on AEAMs. Our experimental results on decoder-only TTS model VALL-E show that the WER of synthesized speech is reduced by up to 20.5% relatively with ACI while the naturalness and speaker similarity are comparable.

【2】 Deep low-latency joint speech transmission and enhancement over a  gaussian channel
标题:高斯通道上的深度低延迟联合语音传输和增强
链接:https://arxiv.org/abs/2404.19375
作者:Mohammad Bokaei,Jesper Jensen,Simon Doclo,Jan Østergaard
摘要:确保听力辅助设备在低延迟场景中的可理解语音通信在语音增强、编码和传输方面提出了重大挑战。在本文中,我们提出了利用深度神经网络(DNN)进行低延迟联合语音传输和增强的新解决方案。我们的方法集成了两种最先进的DNN架构,用于低延迟语音增强和低延迟模拟基于源通道的联合传输,创建了一个组合的低延迟系统,并以端到端的方法联合训练这两个系统。由于增强系统的计算需求,当解码器中没有高计算能力时,例如听力辅助设备,该顺序是合适的。所提出的系统能够配置总延迟,即使在延迟低至3 ms时也能实现高性能,这通常是具有挑战性的。仿真结果提供了令人信服的证据,联合增强和传输系统是优于一个简单的级联系统在不同的设置,包括各种无线信道条件下,lavidity,和背景噪声的情况。
摘要:Ensuring intelligible speech communication for hearing assistive devices in low-latency scenarios presents significant challenges in terms of speech enhancement, coding and transmission. In this paper, we propose novel solutions for low-latency joint speech transmission and enhancement, leveraging deep neural networks (DNNs). Our approach integrates two state-of-the-art DNN architectures for low-latency speech enhancement and low-latency analog joint source-channel-based transmission, creating a combined low-latency system and jointly training both systems in an end-to-end approach. Due to the computational demands of the enhancement system, this order is suitable when high computational power is unavailable in the decoder, like hearing assistive devices. The proposed system enables the configuration of total latency, achieving high performance even at latencies as low as 3 ms, which is typically challenging to attain. The simulation results provide compelling evidence that a joint enhancement and transmission system is superior to a simple concatenation system in diverse settings, encompassing various wireless channel conditions, latencies, and background noise scenarios.


【3】 Fake it to make it: Using synthetic data to remedy the data shortage in  joint multimodal speech-and-gesture synthesis
标题:伪造它来制造它:使用合成数据来弥补联合多模式语音和手势合成中的数据短缺
链接:https://arxiv.org/abs/2404.19622
作者:Shivam Mehta,Anna Deichler,Jim O'Regan,Birger Moëll,Jonas Beskow,Gustav Eje Henter,Simon Alexanderson
备注:13+1 pages, 2 figures, accepted at the Human Motion Generation workshop (HuMoGen) at CVPR 2024
摘要:虽然参与面对面交谈的人同时进行口头和非口头通信,但用于从文本联合和统一合成语音音频和共语音3D手势运动的方法是一个新的新兴领域。这些技术为更人性化、高效、富有表现力和强大的合成通信带来了巨大的希望,但目前由于缺乏适当的大数据集而受到阻碍,因为现有方法是在来自所有组成模态的并行数据上训练的。受学生-教师方法的启发,我们提出了一个简单的解决数据短缺的方法,通过简单地合成额外的培训材料。具体来说,我们使用在大型数据集上训练的单峰合成模型来创建多峰(但合成)并行训练数据,然后在该材料上预训练联合合成模型。此外,我们提出了一个新的合成架构,增加了更好的和更可控的韵律建模的国家的最先进的方法在该领域。我们的研究结果证实,对大量合成数据进行预训练可以提高多模态模型合成的语音和运动的质量,并且在对合成数据进行预训练时,所提出的架构可以带来进一步的好处。有关输出示例,请参见https://shivammehta25.github.io/MAGI/。
摘要:Although humans engaged in face-to-face conversation simultaneously communicate both verbally and non-verbally, methods for joint and unified synthesis of speech audio and co-speech 3D gesture motion from text are a new and emerging field. These technologies hold great promise for more human-like, efficient, expressive, and robust synthetic communication, but are currently held back by the lack of suitably large datasets, as existing methods are trained on parallel data from all constituent modalities. Inspired by student-teacher methods, we propose a straightforward solution to the data shortage, by simply synthesising additional training material. Specifically, we use unimodal synthesis models trained on large datasets to create multimodal (but synthetic) parallel training data, and then pre-train a joint synthesis model on that material. In addition, we propose a new synthesis architecture that adds better and more controllable prosody modelling to the state-of-the-art method in the field. Our results confirm that pre-training on large amounts of synthetic data improves the quality of both the speech and the motion synthesised by the multimodal model, with the proposed architecture yielding further benefits when pre-trained on the synthetic data. See https://shivammehta25.github.io/MAGI/ for example output.


【4】 SemiPL: A Semi-supervised Method for Event Sound Source Localization
标题:SemIPL:一种半监督的事件声音源定位方法
链接:https://arxiv.org/abs/2404.19615
作者:Yue Li,Baiqiao Yin,Jinfu Liu,Jiajun Wen,Jiaying Lin,Mengyuan Liu
摘要:近年来,事件声源定位在各个领域得到了广泛的应用。最近的作品通常依赖于对比学习框架显示令人印象深刻的性能。然而,所有的工作都是基于相对简单的大型数据集。在许多应用中,理解和分析人类行为(人的动作和交互)、声音和混沌事件中的声音也至关重要,例如,人群管理和紧急响应服务。本文将现有的模型应用到一个更复杂的数据集上,探讨了参数对模型的影响,并提出了一种半监督的改进方法SemiPL。随着数据量的增加和标签质量的影响,自监督学习将是不可阻挡的趋势。实验表明,参数调整将积极影响现有的模型。特别是,与提供的结果相比,SSPL在混沌世界中实现了12.2% cIoU和0.56% AUC的改善。代码可在以下网址获得:www.example.com
摘要:In recent years, Event Sound Source Localization has been widely applied in various fields. Recent works typically relying on the contrastive learning framework show impressive performance. However, all work is based on large relatively simple datasets. It's also crucial to understand and analyze human behaviors (actions and interactions of people), voices, and sounds in chaotic events in many applications, e.g., crowd management, and emergency response services. In this paper, we apply the existing model to a more complex dataset, explore the influence of parameters on the model, and propose a semi-supervised improvement method SemiPL. With the increase in data quantity and the influence of label quality, self-supervised learning will be an unstoppable trend. The experiment shows that the parameter adjustment will positively affect the existing model. In particular, SSPL achieved an improvement of 12.2% cIoU and 0.56% AUC in Chaotic World compared to the results provided. The code is available at: https://github.com/ly245422/SSPL

【5】 ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized  Transformers
标题:ESM:使用跨尺度残留量量化变换器的高效语音编码
链接:https://arxiv.org/abs/2404.19441
作者:Yuzhe Gu,Enmao Diao
备注:Preprint
摘要:现有的神经音频编解码器通常为了音频质量而牺牲计算复杂度。它们主要在卷积块上构建特征变换层,卷积块本身并不适合捕获音频信号的局部冗余。作为补偿,无论是对抗损失从一个编码器或大量的模型参数,需要改善编解码器。为此,我们提出了高效的语音编解码器(ESC),一个轻量级的参数高效的编解码器跨尺度残差矢量量化和Transformers。我们的模型利用镜像分层窗口注意Transformer块,并执行从粗到细的特征表示逐步解码。为了提高码本利用率,我们设计了一个学习范式,涉及一个预训练阶段,以协助编解码器的训练。大量的实验结果表明,ESC能够以较低的复杂度获得较高的音频质量,是替代现有编解码器的一种有前景的方案。
摘要:Existing neural audio codecs usually sacrifice computational complexity for audio quality. They build the feature transformation layers mainly on convolutional blocks, which are not inherently appropriate for capturing local redundancies of audio signals. As compensation, either adversarial losses from a discriminator or a large number of model parameters are required to improve the codec. To that end, we propose Efficient Speech Codec (ESC), a lightweight parameter-efficient codec laid on cross-scale residual vector quantization and transformers. Our model leverages mirrored hierarchical window-attention transformer blocks and performs step-wise decoding from coarse-to-fine feature representations. To enhance codebook utilization, we design a learning paradigm that involves a pre-training stage to assist with codec training. Extensive results show that ESC can achieve high audio quality with much lower complexity, which is a prospective alternative in place of existing codecs.

【6】 EfficientASR: Speech Recognition Network Compression via Attention  Redundancy and Chunk-Level FFN Optimization
标题:EfficientASB:通过注意力冗余和块级FFN优化的语音识别网络压缩
链接:https://arxiv.org/abs/2404.19214
作者:Jianzong Wang,Ziqi Liang,Xulong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:近年来,Transformer网络在语音识别任务中表现出了卓越的性能。然而,由于高计算和存储资源需求,它们的部署带来了挑战。为了解决这个问题,本文提出了一个轻量级的模型称为EfficientASR,旨在提高Transformer模型的通用性。EfficientASR采用两个主要模块:共享剩余多头注意力(SRMHA)和块级前馈网络(CFFN)。SRMHA模块有效地减少了网络中的冗余计算,而CFFN模块捕获空间知识并减少了参数的数量。EfficientASR模型的有效性在两个公共数据集上进行了验证,即Aishell-1和HKUST。实验结果表明,与基线Transformer网络相比,参数减少了36%,Aishell-1和HKUST数据集上的字符错误率(CER)分别提高了0.3%和0.2%。
摘要:In recent years, Transformer networks have shown remarkable performance in speech recognition tasks. However, their deployment poses challenges due to high computational and storage resource requirements. To address this issue, a lightweight model called EfficientASR is proposed in this paper, aiming to enhance the versatility of Transformer models. EfficientASR employs two primary modules: Shared Residual Multi-Head Attention (SRMHA) and Chunk-Level Feedforward Networks (CFFN). The SRMHA module effectively reduces redundant computations in the network, while the CFFN module captures spatial knowledge and reduces the number of parameters. The effectiveness of the EfficientASR model is validated on two public datasets, namely Aishell-1 and HKUST. Experimental results demonstrate a 36% reduction in parameters compared to the baseline Transformer network, along with improvements of 0.3% and 0.2% in Character Error Rate (CER) on the Aishell-1 and HKUST datasets, respectively.

【7】 EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with  IFUB Estimator and Joint Text-Guided Consistent Learning
标题:EAD-VC:利用IFLU估计器和联合文本引导一致学习增强语音转换的语音自动解纠缠
链接:https://arxiv.org/abs/2404.19212
作者:Ziqi Liang,Jianzong Wang,Xulong Zhang,Yong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:利用无监督学习将语音分解为内容、节奏、音高和音色进行语音转换已经成为一个热门的研究课题。现有的工作一般考虑到解开语音成分通过人为制造的瓶颈特征,不能实现足够的信息解开,而音高和节奏仍然可能混合在一起。在解缠过程中存在信息重叠的风险,这导致语音自然度降低。为了克服这些限制,我们提出了一个两阶段的模型来解开语音表示在一个自我监督的方式没有人为的瓶颈设计,它使用的互信息(MI)与设计的上限估计(IFUB)分离语音组件之间的重叠信息。此外,我们设计了一个联合文本引导一致(TGC)模块,以指导语音内容的提取和消除音色泄漏问题。实验表明,我们的模型可以实现更好的性能比基线,关于解纠缠的有效性,语音自然度,和相似性。音频样本可以在https://largeaudiomodel.com/eadvc上找到。
摘要:Using unsupervised learning to disentangle speech into content, rhythm, pitch, and timbre for voice conversion has become a hot research topic. Existing works generally take into account disentangling speech components through human-crafted bottleneck features which can not achieve sufficient information disentangling, while pitch and rhythm may still be mixed together. There is a risk of information overlap in the disentangling process which results in less speech naturalness. To overcome such limits, we propose a two-stage model to disentangle speech representations in a self-supervised manner without a human-crafted bottleneck design, which uses the Mutual Information (MI) with the designed upper bound estimator (IFUB) to separate overlapping information between speech components. Moreover, we design a Joint Text-Guided Consistent (TGC) module to guide the extraction of speech content and eliminate timbre leakage issues. Experiments show that our model can achieve a better performance than the baseline, regarding disentanglement effectiveness, speech naturalness, and similarity. Audio samples can be found at https://largeaudiomodel.com/eadvc.


【8】 CONTUNER: Singing Voice Beautifying with Pitch and Expressiveness  Condition
标题:指挥者:用音调和表现力来美化歌声
链接:https://arxiv.org/abs/2404.19187
作者:Jianzong Wang,Pengcheng Li,Xulong Zhang,Ning Cheng,Jing Xiao
备注:Accepted by the 2024 International Joint Conference on Neural Networks (IJCNN 2024)
摘要:歌唱嗓音矫正是一项在人们日常生活中具有应用价值的新课题,其目的是在不改变原有音色和内容的前提下,矫正歌唱声音的音高,提高歌唱声音的表现力。现有的方法依赖于成对的数据或仅集中于音高的校正。但是,专业歌曲和业余歌曲很难从同一个人那里获得,而且歌唱声音的调整不仅仅是音高的调整,还包括情感和节奏等方面。由于我们提出了一个快速和高保真的歌唱语音识别系统称为ConTuner,一个扩散模型结合修改的条件,以产生美化的梅尔频谱图,其中修改的条件是由优化的音高和表现力。对于基音周期校正,我们建立了基音周期、频谱包络到基音周期的映射关系。为了使业余歌唱更有表现力,我们提出了在潜在空间中的表现力增强器,将业余声乐音调转换为专业声乐音调。ConTuner对中文和英文歌曲都达到了令人满意的美化效果。消融实验表明,ConTuner中的表现力增强器和基于生成器的加速方法是有效的。
摘要:Singing voice beautifying is a novel task that has application value in people's daily life, aiming to correct the pitch of the singing voice and improve the expressiveness without changing the original timbre and content. Existing methods rely on paired data or only concentrate on the correction of pitch. However, professional songs and amateur songs from the same person are hard to obtain, and singing voice beautifying doesn't only contain pitch correction but other aspects like emotion and rhythm. Since we propose a fast and high-fidelity singing voice beautifying system called ConTuner, a diffusion model combined with the modified condition to generate the beautified Mel-spectrogram, where the modified condition is composed of optimized pitch and expressiveness. For pitch correction, we establish a mapping relationship from MIDI, spectrum envelope to pitch. To make amateur singing more expressive, we propose the expressiveness enhancer in the latent space to convert amateur vocal tone to professional. ConTuner achieves a satisfactory beautification effect on both Mandarin and English songs. Ablation study demonstrates that the expressiveness enhancer and generator-based accelerate method in ConTuner are effective.


机器翻译由腾讯交互翻译提供,仅供参考