今日论文合集:cs.SD语音7篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Dynamic Cross Attention for Audio-Visual Person Verification
标题:视听人员验证中的动态交叉注意
链接:https://arxiv.org/abs/2403.04661
作者:R. Gnana Praveen,Jahangir Alam
备注:Accepted to FG2024
摘要:虽然个人或身份验证主要是使用个人模式,如面部和声音,视听融合最近显示出巨大的潜力,优于单峰的方法。视听形式往往被认为是一种强有力的互补关系,这在有效的视听融合中起着至关重要的作用。然而,它们可能并不总是强烈地相互补充,它们也可能表现出弱的互补关系,导致视听特征表征不佳。在本文中,我们提出了一个动态交叉注意力(DCA)模型,可以动态地选择交叉出席或无人值守功能的基础上飞的强或弱的互补关系,分别跨音频和视觉模态。特别是,一个条件门控层的设计,以评估的贡献的交叉注意机制,并选择交叉参加的功能,只有当他们表现出很强的互补关系,否则无人值守的功能。在Voxceleb1数据集上进行了大量的实验,以证明所提出的模型的鲁棒性。结果表明,该模型一致地提高了交叉注意的多个变量的性能,同时优于最先进的方法。
摘要:Although person or identity verification has been predominantly explored using individual modalities such as face and voice, audio-visual fusion has recently shown immense potential to outperform unimodal approaches. Audio and visual modalities are often expected to pose strong complementary relationships, which plays a crucial role in effective audio-visual fusion. However, they may not always strongly complement each other, they may also exhibit weak complementary relationships, resulting in poor audio-visual feature representations. In this paper, we propose a Dynamic Cross-Attention (DCA) model that can dynamically select the cross-attended or unattended features on the fly based on the strong or weak complementary relationships, respectively, across audio and visual modalities. In particular, a conditional gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit strong complementary relationships, otherwise unattended features. Extensive experiments are conducted on the Voxceleb1 dataset to demonstrate the robustness of the proposed model. Results indicate that the proposed model consistently improves the performance on multiple variants of cross-attention while outperforming the state-of-the-art methods.


【2】 Audio-Visual Person Verification based on Recursive Fusion of Joint  Cross-Attention
标题:基于联合交叉注意递归融合的视听人身份验证
链接:https://arxiv.org/abs/2403.04654
作者:R. Gnana Praveen,Jahangir Alam
备注:Accepted to FG2024
摘要:由于人脸和声音彼此之间有着密切的联系,因此使用视听融合的个人或身份验证最近受到了很多关注。传统的视听融合方法依赖于分数级或早期特征级融合技术。虽然现有的方法显示出对单峰系统的改进,但视听融合用于人员验证的潜力尚未得到充分利用。在本文中,我们已经调查了有效地捕获跨音频和视觉模态的模态内和模态间关系的前景,这可以在显着提高单峰系统的融合性能方面发挥至关重要的作用。特别是,我们引入了一个递归融合的联合交叉注意模型,其中一个联合视听特征表示采用在交叉注意框架中的递归方式逐步完善的特征表示,可以有效地捕捉内和模态间的关系。为了进一步增强视听特征表示,我们还探索了BLSTM来改进视听特征表示的时间建模。在Voxceleb1数据集上进行了大量的实验,以评估所提出的模型。结果表明,该模型显示出良好的融合性能的改善,熟练地捕捉跨音频和视觉模态的内部和跨模态的关系。
摘要:Person or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion rely on score-level or early feature-level fusion techniques. Though existing approaches showed improvement over unimodal systems, the potential of audio-visual fusion for person verification is not fully exploited. In this paper, we have investigated the prospect of effectively capturing both the intra- and inter-modal relationships across audio and visual modalities, which can play a crucial role in significantly improving the fusion performance over unimodal systems. In particular, we introduce a recursive fusion of a joint cross-attentional model, where a joint audio-visual feature representation is employed in the cross-attention framework in a recursive fashion to progressively refine the feature representations that can efficiently capture the intra-and inter-modal relationships. To further enhance the audio-visual feature representations, we have also explored BLSTMs to improve the temporal modeling of audio-visual feature representations. Extensive experiments are conducted on the Voxceleb1 dataset to evaluate the proposed model. Results indicate that the proposed model shows promising improvement in fusion performance by adeptly capturing the intra-and inter-modal relationships across audio and visual modalities.


【3】 A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
标题:一种使用单事件声音的详细音文数据模拟流水线
链接:https://arxiv.org/abs/2403.04594
作者:Xuenan Xu,Xiaohang Xu,Zeyu Xie,Pingyue Zhang,Mengyue Wu,Kai Yu
摘要:近年来,有声文本跨模态学习越来越受到人们的关注。然而,大多数现有的音频文本数据集只包含简单的声音事件的描述。与分类标签相比,这种描述的优势明显有限。在本文中,我们首先分析了详细的信息,人类描述的音频可能包含以外的声音事件标签。基于分析,我们提出了一个自动管道,用于策划具有丰富细节的音频文本对。利用声音可以在时域中混合和连接的属性,我们在四个方面控制细节:时间关系,响度,扬声器身份和出现次数,在模拟音频混合。相应的细节通过大型语言模型转换为字幕。从而获得在文本描述中具有丰富细节的音频文本对。我们用少量的模拟数据验证了我们的管道的有效性,证明了模拟数据使模型能够学习详细的音频字幕。
摘要:Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such descriptions are significantly limited. In this paper, we first analyze the detailed information that human descriptions of audio may contain beyond sound event labels. Based on the analysis, we propose an automatic pipeline for curating audio-text pairs with rich details. Leveraging the property that sounds can be mixed and concatenated in the time domain, we control details in four aspects: temporal relationship, loudness, speaker identity, and occurrence number, in simulating audio mixtures. Corresponding details are transformed into captions by large language models. Audio-text pairs with rich details in text descriptions are thereby obtained. We validate the effectiveness of our pipeline with a small amount of simulated data, demonstrating that the simulated data enables models to learn detailed audio captioning.

【4】 A Study of Dropout-Induced Modality Bias on Robustness to Missing Video  Frames for Audio-Visual Speech Recognition
标题:丢失诱导的通道偏差对视听语音识别中丢失视频帧的稳健性研究
链接:https://arxiv.org/abs/2403.04245
作者:Yusheng Dai,Hang Chen,Jun Du,Ruoyu Wang,Shihao Chen,Jiefeng Ma,Haotian Wang,Chin-Hui Lee备注:the paper is accepted by CVPR2024
摘要:高级视听语音识别(AVSR)系统已被观察到对丢失的视频帧敏感,表现甚至比单模态模型更差。虽然将dropout技术应用于视频模态增强了对丢失帧的鲁棒性,但当处理完整的数据输入时,它同时导致性能损失。本文从情态偏误的角度对这一现象进行了研究,揭示了由于丢失而导致的对音频的过度情态偏误是造成这一现象的根本原因。此外,我们提出了模态偏差假设(MBH),系统地描述了模态偏差和鲁棒性之间的关系,对丢失的模态系统。基于这些发现,我们提出了一种新的多模态分布近似与知识蒸馏(MDA-KD)框架,以减少对音频模态的过度依赖,并同时保持性能和鲁棒性。最后,为了解决一个完全缺失的模态,我们采用适配器来动态切换决策策略。我们提出的方法的有效性进行了评估和验证,通过一系列的综合实验使用MISP 2021和MISP 2022数据集。我们的代码可在https://github.com/dalision/ModalBiasAVSR上获得
摘要:Advanced Audio-Visual Speech Recognition (AVSR) systems have been observed to be sensitive to missing video frames, performing even worse than single-modality models. While applying the dropout technique to the video modality enhances robustness to missing frames, it simultaneously results in a performance loss when dealing with complete data input. In this paper, we investigate this contrasting phenomenon from the perspective of modality bias and reveal that an excessive modality bias on the audio caused by dropout is the underlying reason. Moreover, we present the Modality Bias Hypothesis (MBH) to systematically describe the relationship between modality bias and robustness against missing modality in multimodal systems. Building on these findings, we propose a novel Multimodal Distribution Approximation with Knowledge Distillation (MDA-KD) framework to reduce over-reliance on the audio modality and to maintain performance and robustness simultaneously. Finally, to address an entirely missing modality, we adopt adapters to dynamically switch decision strategies. The effectiveness of our proposed approach is evaluated and validated through a series of comprehensive experiments using the MISP2021 and MISP2022 datasets. Our code is available at https://github.com/dalision/ModalBiasAVSR

【5】 Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
标题:语音到语音机器翻译中重音转移的尝试
链接:https://arxiv.org/abs/2403.04178
作者:Sai Akarsh,Vamshi Raghusimha,Anindita Mondal,Anil Vuppala
摘要:印度教育部门的语言多样性构成了重大挑战,阻碍了包容性。尽管通过在线教育内容实现了知识的民主化,但英语作为互联网通用语言的主导地位限制了可访问性,强调了翻译成印度语言的关键需求。尽管现有的语音到语音机器翻译(SSMT)技术,但这些系统中缺乏语调,翻译单调,导致观众失去兴趣,脱离内容。为了解决这个问题,我们的论文介绍了一个数据集与印度英语的压力注释,也是一个文本到语音(TTS)架构能够将压力合成语音。该数据集用于训练重音检测模型,然后在SSMT系统中用于检测源语音中的重音并将其转换为目标语言语音。TTS架构基于FastPitch,可以根据给定的重音单词修改方差。我们提出了一个印度英语到印地语SSMT系统,可以转移压力,旨在提高教育内容的整体质量和参与度。
摘要:The language diversity in India's education sector poses a significant challenge, hindering inclusivity. Despite the democratization of knowledge through online educational content, the dominance of English, as the internet's lingua franca, limits accessibility, emphasizing the crucial need for translation into Indian languages. Despite existing Speech-to-Speech Machine Translation (SSMT) technologies, the lack of intonation in these systems gives monotonous translations, leading to a loss of audience interest and disengagement from the content. To address this, our paper introduces a dataset with stress annotations in Indian English and also a Text-to-Speech (TTS) architecture capable of incorporating stress into synthesized speech. This dataset is used for training a stress detection model, which is then used in the SSMT system for detecting stress in the source speech and transferring it into the target language speech. The TTS architecture is based on FastPitch and can modify the variances based on stressed words given. We present an Indian English-to-Hindi SSMT system that can transfer stress and aim to enhance the overall quality and engagement of educational content.


【6】 Multi-Level Attention Aggregation for Language-Agnostic Speaker  Replication
标题:多层次注意力聚合的语音不可知说话人复制算法
链接:https://arxiv.org/abs/2403.04111
作者:Yejin Jeon,Gary Geunbae Lee
备注:Accepted to EACL Main 2024
摘要:本文探讨了语言不可知论的扬声器复制,一种新颖的努力,旨在复制扬声器的声音,无论他们说的语言的任务。为此,我们引入了一个多层次的注意力聚合方法,系统地探测和放大各种扬声器的特定属性的分层方式。通过严格的评估,在广泛的场景,包括看到的和看不见的扬声器在看到的和看不见的语言交谈,我们建立我们提出的模型是能够实现大量的扬声器相似性,并能够推广到域外(OOD)的情况下。
摘要:This paper explores the task of language-agnostic speaker replication, a novel endeavor that seeks to replicate a speaker's voice irrespective of the language they are speaking. Towards this end, we introduce a multi-level attention aggregation approach that systematically probes and amplifies various speaker-specific attributes in a hierarchical manner. Through rigorous evaluations across a wide range of scenarios including seen and unseen speakers conversing in seen and unseen lingua, we establish that our proposed model is able to achieve substantial speaker similarity, and is able to generalize to out-of-domain (OOD) cases.


【7】 On the Use of Autoregressive Methods for Audio Inpainting
标题:自回归方法在音频修复中的应用
链接:https://arxiv.org/abs/2403.04433
作者:Ondřej Mokrý,Pavel Rajmic
摘要:本文提出了一种流行的音频修复方法的自回归建模的基础上,即基于外推和杨森方法的评估。还提出了一种新的变体的Janssen方法适用于间隙修复。特别流行的方法之间的主要区别指出,并提出了一个中等规模的计算实验。结果表明,AR模型估计量的选择和新的间隙明智的杨森方法的适用性的重要性。
摘要:The paper presents an evaluation of popular audio inpainting methods based on autoregressive modelling, namely, extrapolation-based and Janssen methods. A novel variant of the Janssen method suitable for gap inpainting is also proposed. The main differences between the particular popular approaches are pointed out, and a mid-scale computational experiment is presented. The results demonstrate the importance of the choice of the AR model estimator and the suitability of the new gap-wise Janssen method.

eess.AS音频处理
【1】 Speech Emotion Recognition Via CNN-Transforemr and Multidimensional  Attention Mechanism
标题:基于CNN-Transformmr和多维注意机制的语音情感识别
链接:https://arxiv.org/abs/2403.04743
作者:Xiaoyu Tang,Yixin Lin,Ting Dang,Yuanfang Zhang,Jintao Cheng
摘要:语音情感识别(SER)在人机交互中至关重要。主流方法利用卷积神经网络或递归神经网络来从语音信息中学习语音片段的局部能量特征表示,但难以捕获诸如语音中的能量持续时间之类的全局信息。有些使用Transformers来捕获全局信息,但在参数计数和性能方面还有改进的余地。此外,现有的注意力机制集中在空间或通道维度,阻碍了重要的时间信息的语音学习。本文提出了一种基于CNN-Transformer和多维注意力机制的语音情感识别网络,用于对语音中不同粒度级别的局部和全局信息进行建模,并捕捉语音信号中的时间、空间和通道依赖性。具体地,CNN块的堆栈专用于从时间-频率角度捕获语音中的本地信息。此外,时间-通道-空间的注意力机制用于增强跨三个维度的特征。此外,我们使用具有可分离卷积的大型卷积核和轻量级Transformer模块来建模特征序列的局部和全局依赖性。我们在IEMOCAP和EMO-DB数据集上评估了所提出的方法,并表明我们的方法比最先进的方法显着提高了性能。我们的代码可在https://github.com/SCNU-RISLAB/CNN-Transforemr-and-Multidimensional-Attention-Mechanism上获得
摘要:Speech Emotion Recognition (SER) is crucial in human-machine interactions. Mainstream approaches utilize Convolutional Neural Networks or Recurrent Neural Networks to learn local energy feature representations of speech segments from speech information, but struggle with capturing global information such as the duration of energy in speech. Some use Transformers to capture global information, but there is room for improvement in terms of parameter count and performance. Furthermore, existing attention mechanisms focus on spatial or channel dimensions, hindering learning of important temporal information in speech. In this paper, to model local and global information at different levels of granularity in speech and capture temporal, spatial and channel dependencies in speech signals, we propose a Speech Emotion Recognition network based on CNN-Transformer and multi-dimensional attention mechanisms. Specifically, a stack of CNN blocks is dedicated to capturing local information in speech from a time-frequency perspective. In addition, a time-channel-space attention mechanism is used to enhance features across three dimensions. Moreover, we model local and global dependencies of feature sequences using large convolutional kernels with depthwise separable convolutions and lightweight Transformer modules. We evaluate the proposed method on IEMOCAP and Emo-DB datasets and show our approach significantly improves the performance over the state-of-the-art methods. Our code is available on https://github.com/SCNU-RISLAB/CNN-Transforemr-and-Multidimensional-Attention-Mechanism


【2】 On the Use of Autoregressive Methods for Audio Inpainting
标题:自回归方法在音频修复中的应用
链接:https://arxiv.org/abs/2403.04433
作者:Ondřej Mokrý,Pavel Rajmic
摘要:本文提出了一种流行的音频修复方法的自回归建模的基础上,即基于外推和杨森方法的评估。还提出了一种新的变体的Janssen方法适用于间隙修复。特别流行的方法之间的主要区别指出,并提出了一个中等规模的计算实验。结果表明,AR模型估计量的选择和新的间隙明智的杨森方法的适用性的重要性。
摘要:The paper presents an evaluation of popular audio inpainting methods based on autoregressive modelling, namely, extrapolation-based and Janssen methods. A novel variant of the Janssen method suitable for gap inpainting is also proposed. The main differences between the particular popular approaches are pointed out, and a mid-scale computational experiment is presented. The results demonstrate the importance of the choice of the AR model estimator and the suitability of the new gap-wise Janssen method.


【3】 Dynamic Cross Attention for Audio-Visual Person Verification
标题:视听人员验证中的动态交叉注意
链接:https://arxiv.org/abs/2403.04661
作者:R. Gnana Praveen,Jahangir Alam
备注:Accepted to FG2024
摘要:虽然个人或身份验证主要是使用个人模式,如面部和声音,视听融合最近显示出巨大的潜力,优于单峰的方法。视听形式往往被认为是一种强有力的互补关系,这在有效的视听融合中起着至关重要的作用。然而,它们可能并不总是强烈地相互补充,它们也可能表现出弱的互补关系,导致视听特征表征不佳。在本文中,我们提出了一个动态交叉注意力(DCA)模型,可以动态地选择交叉出席或无人值守功能的基础上飞的强或弱的互补关系,分别跨音频和视觉模态。特别是,一个条件门控层的设计,以评估的贡献的交叉注意机制,并选择交叉参加的功能,只有当他们表现出很强的互补关系,否则无人值守的功能。在Voxceleb1数据集上进行了大量的实验,以证明所提出的模型的鲁棒性。结果表明,该模型一致地提高了交叉注意的多个变量的性能,同时优于最先进的方法。
摘要:Although person or identity verification has been predominantly explored using individual modalities such as face and voice, audio-visual fusion has recently shown immense potential to outperform unimodal approaches. Audio and visual modalities are often expected to pose strong complementary relationships, which plays a crucial role in effective audio-visual fusion. However, they may not always strongly complement each other, they may also exhibit weak complementary relationships, resulting in poor audio-visual feature representations. In this paper, we propose a Dynamic Cross-Attention (DCA) model that can dynamically select the cross-attended or unattended features on the fly based on the strong or weak complementary relationships, respectively, across audio and visual modalities. In particular, a conditional gating layer is designed to evaluate the contribution of the cross-attention mechanism and choose cross-attended features only when they exhibit strong complementary relationships, otherwise unattended features. Extensive experiments are conducted on the Voxceleb1 dataset to demonstrate the robustness of the proposed model. Results indicate that the proposed model consistently improves the performance on multiple variants of cross-attention while outperforming the state-of-the-art methods.


【4】 Audio-Visual Person Verification based on Recursive Fusion of Joint  Cross-Attention
标题:基于联合交叉注意递归融合的视听人身份验证
链接:https://arxiv.org/abs/2403.04654
作者:R. Gnana Praveen,Jahangir Alam
备注:Accepted to FG2024
摘要:由于人脸和声音彼此之间有着密切的联系,因此使用视听融合的个人或身份验证最近受到了很多关注。传统的视听融合方法依赖于分数级或早期特征级融合技术。虽然现有的方法显示出对单峰系统的改进,但视听融合用于人员验证的潜力尚未得到充分利用。在本文中,我们已经调查了有效地捕获跨音频和视觉模态的模态内和模态间关系的前景,这可以在显着提高单峰系统的融合性能方面发挥至关重要的作用。特别是,我们引入了一个递归融合的联合交叉注意模型,其中一个联合视听特征表示采用在交叉注意框架中的递归方式逐步完善的特征表示,可以有效地捕捉内和模态间的关系。为了进一步增强视听特征表示,我们还探索了BLSTM来改进视听特征表示的时间建模。在Voxceleb1数据集上进行了大量的实验,以评估所提出的模型。结果表明,该模型显示出良好的融合性能的改善,熟练地捕捉跨音频和视觉模态的内部和跨模态的关系。
摘要:Person or identity verification has been recently gaining a lot of attention using audio-visual fusion as faces and voices share close associations with each other. Conventional approaches based on audio-visual fusion rely on score-level or early feature-level fusion techniques. Though existing approaches showed improvement over unimodal systems, the potential of audio-visual fusion for person verification is not fully exploited. In this paper, we have investigated the prospect of effectively capturing both the intra- and inter-modal relationships across audio and visual modalities, which can play a crucial role in significantly improving the fusion performance over unimodal systems. In particular, we introduce a recursive fusion of a joint cross-attentional model, where a joint audio-visual feature representation is employed in the cross-attention framework in a recursive fashion to progressively refine the feature representations that can efficiently capture the intra-and inter-modal relationships. To further enhance the audio-visual feature representations, we have also explored BLSTMs to improve the temporal modeling of audio-visual feature representations. Extensive experiments are conducted on the Voxceleb1 dataset to evaluate the proposed model. Results indicate that the proposed model shows promising improvement in fusion performance by adeptly capturing the intra-and inter-modal relationships across audio and visual modalities.

【5】 A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
标题:一种使用单事件声音的详细音文数据模拟流水线
链接:https://arxiv.org/abs/2403.04594
作者:Xuenan Xu,Xiaohang Xu,Zeyu Xie,Pingyue Zhang,Mengyue Wu,Kai Yu
摘要:近年来,有声文本跨模态学习越来越受到人们的关注。然而,大多数现有的音频文本数据集只包含简单的声音事件的描述。与分类标签相比,这种描述的优势明显有限。在本文中,我们首先分析了详细的信息,人类描述的音频可能包含以外的声音事件标签。基于分析,我们提出了一个自动管道,用于策划具有丰富细节的音频文本对。利用声音可以在时域中混合和连接的属性,我们在四个方面控制细节:时间关系,响度,扬声器身份和出现次数,在模拟音频混合。相应的细节通过大型语言模型转换为字幕。从而获得在文本描述中具有丰富细节的音频文本对。我们用少量的模拟数据验证了我们的管道的有效性,证明了模拟数据使模型能够学习详细的音频字幕。
摘要:Recently, there has been an increasing focus on audio-text cross-modal learning. However, most of the existing audio-text datasets contain only simple descriptions of sound events. Compared with classification labels, the advantages of such descriptions are significantly limited. In this paper, we first analyze the detailed information that human descriptions of audio may contain beyond sound event labels. Based on the analysis, we propose an automatic pipeline for curating audio-text pairs with rich details. Leveraging the property that sounds can be mixed and concatenated in the time domain, we control details in four aspects: temporal relationship, loudness, speaker identity, and occurrence number, in simulating audio mixtures. Corresponding details are transformed into captions by large language models. Audio-text pairs with rich details in text descriptions are thereby obtained. We validate the effectiveness of our pipeline with a small amount of simulated data, demonstrating that the simulated data enables models to learn detailed audio captioning.

【6】 A Study of Dropout-Induced Modality Bias on Robustness to Missing Video  Frames for Audio-Visual Speech Recognition
标题:丢失诱导的通道偏差对视听语音识别中丢失视频帧的稳健性研究
链接:https://arxiv.org/abs/2403.04245
作者:Yusheng Dai,Hang Chen,Jun Du,Ruoyu Wang,Shihao Chen,Jiefeng Ma,Haotian Wang,Chin-Hui Lee备注:the paper is accepted by CVPR2024
摘要:高级视听语音识别(AVSR)系统已被观察到对丢失的视频帧敏感,表现甚至比单模态模型更差。虽然将dropout技术应用于视频模态增强了对丢失帧的鲁棒性,但当处理完整的数据输入时,它同时导致性能损失。本文从情态偏误的角度对这一现象进行了研究,揭示了由于丢失而导致的对音频的过度情态偏误是造成这一现象的根本原因。此外,我们提出了模态偏差假设(MBH),系统地描述了模态偏差和鲁棒性之间的关系,对丢失的模态系统。基于这些发现,我们提出了一种新的多模态分布近似与知识蒸馏(MDA-KD)框架,以减少对音频模态的过度依赖,并同时保持性能和鲁棒性。最后,为了解决一个完全缺失的模态,我们采用适配器来动态切换决策策略。我们提出的方法的有效性进行了评估和验证,通过一系列的综合实验使用MISP 2021和MISP 2022数据集。我们的代码可在https://github.com/dalision/ModalBiasAVSR上获得
摘要:Advanced Audio-Visual Speech Recognition (AVSR) systems have been observed to be sensitive to missing video frames, performing even worse than single-modality models. While applying the dropout technique to the video modality enhances robustness to missing frames, it simultaneously results in a performance loss when dealing with complete data input. In this paper, we investigate this contrasting phenomenon from the perspective of modality bias and reveal that an excessive modality bias on the audio caused by dropout is the underlying reason. Moreover, we present the Modality Bias Hypothesis (MBH) to systematically describe the relationship between modality bias and robustness against missing modality in multimodal systems. Building on these findings, we propose a novel Multimodal Distribution Approximation with Knowledge Distillation (MDA-KD) framework to reduce over-reliance on the audio modality and to maintain performance and robustness simultaneously. Finally, to address an entirely missing modality, we adopt adapters to dynamically switch decision strategies. The effectiveness of our proposed approach is evaluated and validated through a series of comprehensive experiments using the MISP2021 and MISP2022 datasets. Our code is available at https://github.com/dalision/ModalBiasAVSR

【7】 Attempt Towards Stress Transfer in Speech-to-Speech Machine Translation
标题:语音到语音机器翻译中重音转移的尝试
链接:https://arxiv.org/abs/2403.04178
作者:Sai Akarsh,Vamshi Raghusimha,Anindita Mondal,Anil Vuppala
摘要:印度教育部门的语言多样性构成了重大挑战,阻碍了包容性。尽管通过在线教育内容实现了知识的民主化,但英语作为互联网通用语言的主导地位限制了可访问性,强调了翻译成印度语言的关键需求。尽管现有的语音到语音机器翻译(SSMT)技术,但这些系统中缺乏语调,翻译单调,导致观众失去兴趣,脱离内容。为了解决这个问题,我们的论文介绍了一个数据集与印度英语的压力注释,也是一个文本到语音(TTS)架构能够将压力合成语音。该数据集用于训练重音检测模型,然后在SSMT系统中用于检测源语音中的重音并将其转换为目标语言语音。TTS架构基于FastPitch,可以根据给定的重音单词修改方差。我们提出了一个印度英语到印地语SSMT系统,可以转移压力,旨在提高教育内容的整体质量和参与度。
摘要:The language diversity in India's education sector poses a significant challenge, hindering inclusivity. Despite the democratization of knowledge through online educational content, the dominance of English, as the internet's lingua franca, limits accessibility, emphasizing the crucial need for translation into Indian languages. Despite existing Speech-to-Speech Machine Translation (SSMT) technologies, the lack of intonation in these systems gives monotonous translations, leading to a loss of audience interest and disengagement from the content. To address this, our paper introduces a dataset with stress annotations in Indian English and also a Text-to-Speech (TTS) architecture capable of incorporating stress into synthesized speech. This dataset is used for training a stress detection model, which is then used in the SSMT system for detecting stress in the source speech and transferring it into the target language speech. The TTS architecture is based on FastPitch and can modify the variances based on stressed words given. We present an Indian English-to-Hindi SSMT system that can transfer stress and aim to enhance the overall quality and engagement of educational content.


【8】 Multi-Level Attention Aggregation for Language-Agnostic Speaker  Replication
标题:语言无关说话人复制的多层次注意力聚合
链接:https://arxiv.org/abs/2403.04111
作者:Yejin Jeon,Gary Geunbae Lee
备注:Accepted to EACL Main 2024
摘要:本文探讨了语言不可知论的扬声器复制,一种新颖的努力,旨在复制扬声器的声音,无论他们说的语言的任务。为此,我们引入了一个多层次的注意力聚合方法,系统地探测和放大各种扬声器的特定属性的分层方式。通过严格的评估,在广泛的场景,包括看到的和看不见的扬声器在看到的和看不见的语言交谈,我们建立我们提出的模型是能够实现大量的扬声器相似性,并能够推广到域外(OOD)的情况下。
摘要:This paper explores the task of language-agnostic speaker replication, a novel endeavor that seeks to replicate a speaker's voice irrespective of the language they are speaking. Towards this end, we introduce a multi-level attention aggregation approach that systematically probes and amplifies various speaker-specific attributes in a hierarchical manner. Through rigorous evaluations across a wide range of scenarios including seen and unseen speakers conversing in seen and unseen lingua, we establish that our proposed model is able to achieve substantial speaker similarity, and is able to generalize to out-of-domain (OOD) cases.


机器翻译由腾讯交互翻译提供,仅供参考