今日论文合集:cs.SD语音7篇,eess.AS音频处理9篇。本文经arXiv每日学术速递授权转载
【1】Siamese Vision Transformers are ScalableAudio-visual Learners
标题:Siamese Vision Transformers是可扩展的视听学习者
链接:https://arxiv.org/abs/2403.19638
作者:Yan-Bo Lin,Gedas Bertasius
摘要:传统的视听方法依赖于独立的音频和视频骨干,这是昂贵的,不可扩展的。在这项工作中,我们研究使用视听连体网络(AVSiam)进行高效和可扩展的视听预训练。我们的框架使用一个共享的Vision Transformer主干来处理音频和视频输入,提高了其参数效率,减少了GPU内存占用,并允许我们将我们的方法扩展到更大的数据集和模型大小。我们使用具有多比率随机掩蔽方案的对比视听匹配目标来预训练我们的模型,这使我们的模型能够处理更大的视听实例批次,有助于对比学习。与以前的视听方法不同,我们的方法可以鲁棒地处理音频,视频和视听输入与一个单一的共享ViT骨干。此外,尽管两种模式都使用了共享的主干,AVSiam在视听分类和检索方面取得了比AudioSet和VGGSound更好的结果。我们的代码可在https://github.com/GenjiB/AVSiam上获得
摘要:Traditional audio-visual methods rely on independent audio and visual backbones, which is costly and not scalable. In this work, we investigate using an audio-visual siamese network (AVSiam) for efficient and scalable audio-visual pretraining. Our framework uses a single shared vision transformer backbone to process audio and visual inputs, improving its parameter efficiency, reducing the GPU memory footprint, and allowing us to scale our method to larger datasets and model sizes. We pretrain our model using a contrastive audio-visual matching objective with a multi-ratio random masking scheme, which enables our model to process larger audio-visual instance batches, helpful for contrastive learning. Unlike prior audio-visual methods, our method can robustly handle audio, visual, and audio-visual inputs with a single shared ViT backbone. Furthermore, despite using the shared backbone for both modalities, AVSiam achieves competitive or even better results than prior methods on AudioSet and VGGSound for audio-visual classification and retrieval. Our code is available at https://github.com/GenjiB/AVSiam
【2】 Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2标题:非对称和试验相关建模:LIA对SdSV挑战任务2的贡献作者:Pierre-Michel Bousquet,Mickael Rouvier备注:LIA system description for the Short Duration Speaker Verification (SdSv) challenge 2020 Task 2摘要:SdSv挑战任务2提供了一个机会,以评估现代文本无关的说话人确认系统的效率和鲁棒性。但这也使得有可能测试新的方法,能够考虑到这一挑战的主要问题(期限、语言.)。本文介绍了我们实验室在说话人识别领域的贡献。这些贡献突出了除了短持续时间和语言之外的另外两个挑战:招募和测试数据之间的不匹配以及评估试验数据集子集之间的不匹配。实验表明,所提出的方法的相关性和效率的SdSv评估,并可能在许多现实生活中的应用感兴趣。摘要:The SdSv challenge Task 2 provided an opportunity to assess efficiency and robustness of modern text-independent speaker verification systems. But it also made it possible to test new approaches, capable of taking into account the main issues of this challenge (duration, language, ...). This paper describes the contributions of our laboratory to the speaker recognition field. These contributions highlight two other challenges in addition to short-duration and language: the mismatch between enrollment and test data and the one between subsets of the evaluation trial dataset. The proposed approaches experimentally show their relevance and efficiency on the SdSv evaluation, and could be of interest in many real-life applications.
【3】 Phonetic Segmentation of the UCLA Phonetics Lab Archive标题:加州大学洛杉矶分校语音实验室档案中的语音分割作者:Eleanor Chodroff,Blaž Pažon,Annie Baker,Steven Moran备注:Accepted at LREC-COLING 2024摘要:语音技术和比较语言学的研究依赖于获得多样化和可访问的语音数据。加州大学洛杉矶分校语音学实验室档案是最早的多语种语音语料库之一,具有314种语言的长格式音频记录和语音音译(Ladefoged et al.,2009年)。最近,这些语言中的95种与单词级语音音译时间对齐(Li等人,2021年)。在这里,我们提出了VoxAngeles,一个语料库的审计语音音译和语音水平的调整加州大学洛杉矶分校语音实验室档案,它使用95种语言的CMU重新发布作为我们的出发点。VoxAngeles还包括来自原始UCLA语料库的单词和音素级分割,以及单词和音素持续时间,元音共振峰和元音f0的语音测量。该语料库增强了原始数据的可用性,特别是对于定量语音类型学,如通过元音内在f0的案例研究所示。我们还讨论了跨语言语音学的一般研究和教学的VoxAngeles语料库的效用,以及低资源和多语言语音技术。VoxAngeles可以在CC-BY-NC 4.0许可下免费下载和使用。摘要:Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio recordings and phonetic transcriptions for 314 languages (Ladefoged et al., 2009). Recently, 95 of these languages were time-aligned with word-level phonetic transcriptions (Li et al., 2021). Here we present VoxAngeles, a corpus of audited phonetic transcriptions and phone-level alignments of the UCLA Phonetics Lab Archive, which uses the 95-language CMU re-release as our starting point. VoxAngeles also includes word- and phone-level segmentations from the original UCLA corpus, as well as phonetic measurements of word and phone durations, vowel formants, and vowel f0. This corpus enhances the usability of the original data, particularly for quantitative phonetic typology, as demonstrated through a case study of vowel intrinsic f0. We also discuss the utility of the VoxAngeles corpus for general research and pedagogy in crosslinguistic phonetics, as well as for low-resource and multilingual speech technologies. VoxAngeles is free to download and use under a CC-BY-NC 4.0 license.
【4】 A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews作者:Mamadou Dia,Ghazaleh Khodabandelou,Alice Othmani摘要:创伤后应激障碍(PTSD)是一种精神障碍,可以在目睹或经历极端创伤事件后发展。PTSD可以影响任何人,无论种族或文化。据估计,每11个人中就有一个人在一生中会经历PTSD。临床医生管理的PTSD量表(CAPS)和平民PTSD检查表(PCL-C)访谈是诊断PTSD的金标准。这些问卷可能会被受试者的回答所欺骗。这项工作提出了一种基于深度学习的方法,该方法在临床访谈期间使用音频记录实现了最先进的PTSD检测性能。我们的方法是基于从临床访谈的音频记录中提取的MFCC低级特征,然后使用随机Transformer进行深度高级学习。我们提出的方法在eDAIC数据集上实现了最先进的性能,RMSE为2.92,这要归功于随机深度、随机深度学习层和随机激活函数。摘要:Post-traumatic stress disorder (PTSD) is a mental disorder that can be developed after witnessing or experiencing extremely traumatic events. PTSD can affect anyone, regardless of ethnicity, or culture. An estimated one in every eleven people will experience PTSD during their lifetime. The Clinician-Administered PTSD Scale (CAPS) and the PTSD Check List for Civilians (PCL-C) interviews are gold standards in the diagnosis of PTSD. These questionnaires can be fooled by the subject's responses. This work proposes a deep learning-based approach that achieves state-of-the-art performances for PTSD detection using audio recordings during clinical interviews. Our approach is based on MFCC low-level features extracted from audio recordings of clinical interviews, followed by deep high-level learning using a Stochastic Transformer. Our proposed approach achieves state-of-the-art performances with an RMSE of 2.92 on the eDAIC dataset thanks to the stochastic depth, stochastic deep learning layers, and stochastic activation function.【5】 Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition作者:Siyuan Shen,Yu Gao,Feng Liu,Hanyang Wang,Aimin Zhou备注:Accepted by 49th IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024)摘要:语音情感识别的主流范式是识别整个话语的单个情感标签。这一系列的工作忽略了情感动态在细时间粒度和大多数未能充分利用语音信号的语言信息显式。在本文中,我们提出了情感神经传感器的细粒度语音情感识别与自动语音识别(ASR)联合训练。我们首先扩展了典型的神经传感器与情感联合网络,构建情感格细粒度的SER,然后我们提出了对齐格上的格最大池,以方便区分情感和非情感帧。为了使细粒度的SER适应传感器推理方式,我们进一步将ASR的特殊符号blank作为底层情感指示器,得到了因子化的情感神经传感器。对于典型的话语级SER,我们的ENT模型优于国家的最先进的方法IEMOCAP在低字错误率。在IEMOCAP和最新的语音情感日志数据集ZED上的实验也证明了细粒度情感建模的优越性。我们的代码可在https://github.com/ECNU-Cross-Innovation-Lab/ENT上获得。摘要:The mainstream paradigm of speech emotion recognition (SER) is identifying the single emotion label of the entire utterance. This line of works neglect the emotion dynamics at fine temporal granularity and mostly fail to leverage linguistic information of speech signal explicitly. In this paper, we propose Emotion Neural Transducer for fine-grained speech emotion recognition with automatic speech recognition (ASR) joint training. We first extend typical neural transducer with emotion joint network to construct emotion lattice for fine-grained SER. Then we propose lattice max pooling on the alignment lattice to facilitate distinguishing emotional and non-emotional frames. To adapt fine-grained SER to transducer inference manner, we further make blank, the special symbol of ASR, serve as underlying emotion indicator as well, yielding Factorized Emotion Neural Transducer. For typical utterance-level SER, our ENT models outperform state-of-the-art methods on IEMOCAP in low word error rate. Experiments on IEMOCAP and the latest speech emotion diarization dataset ZED also demonstrate the superiority of fine-grained emotion modeling. Our code is available at https://github.com/ECNU-Cross-Innovation-Lab/ENT.
【6】 Robust Active Speaker Detection in Noisy Environments作者:Siva Sai Nagender Vasireddy,Chenxu Zhang,Xiaohu Guo,Yapeng Tian摘要:针对噪声环境下的主动说话人检测问题,提出了一种鲁棒的主动说话人检测方法。现有的ASD方法利用音频和视觉模态,但周围环境中的非语音声音会对性能产生负面影响。为了克服这一点,我们提出了一个新的框架,利用视听语音分离作为指导,学习无噪声的音频功能。这些功能然后在ASD模型中使用,并且这两个任务在端到端框架中联合优化。我们提出的框架减轻了残留噪声和音频质量降低的问题,可以发生在一个天真的级联两阶段的框架,直接使用分离的语音ASD,并使这两个任务同时优化。为了进一步增强音频特征的鲁棒性和处理固有的语音噪声,我们提出了一种动态加权损失的方法来训练语音分离器。我们还收集了一个真实世界的噪音音频数据集,以方便调查。实验表明,非语音音频噪声显着影响ASD模型,我们提出的方法提高了ASD在噪声环境中的性能。该框架是通用的,可以应用于不同的ASD方法,以提高其鲁棒性。我们的代码、模型和数据将被发布。摘要:This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds in the surrounding environment can negatively impact performance. To overcome this, we propose a novel framework that utilizes audio-visual speech separation as guidance to learn noise-free audio features. These features are then utilized in an ASD model, and both tasks are jointly optimized in an end-to-end framework. Our proposed framework mitigates residual noise and audio quality reduction issues that can occur in a naive cascaded two-stage framework that directly uses separated speech for ASD, and enables the two tasks to be optimized simultaneously. To further enhance the robustness of the audio features and handle inherent speech noises, we propose a dynamic weighted loss approach to train the speech separator. We also collected a real-world noise audio dataset to facilitate investigations. Experiments demonstrate that non-speech audio noises significantly impact ASD models, and our proposed approach improves ASD performance in noisy environments. The framework is general and can be applied to different ASD approaches to improve their robustness. Our code, models, and data will be released.【7】 JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition标题:JEP—KD:基于联合嵌入预测结构的视觉语音识别知识提取作者:Chang Sun,Hong Yang,Bo Qin摘要:视觉语音识别(VSR)任务通常被认为具有比自动语音识别(ASR)更低的理论性能上限,这是由于视觉传达语义信息的固有限制。为了缓解这一挑战,本文介绍了一种使用联合嵌入预测架构(JEPA)的高级知识蒸馏方法,称为JEP-KD,旨在更有效地利用模型训练期间的音频特征。JEP-KD的核心是在嵌入层中包含一个生成网络,这增强了视频编码器的语义特征提取能力,并使其与来自预训练ASR模型编码器的音频特征更接近。这种方法旨在逐步缩小VSR和ASR之间的性能差距。此外,还为JEP-KD框架建立了一个全面的多模式、多阶段训练方案,增强了训练过程的鲁棒性和有效性。实验结果表明,JEP-KD显著提高了VSR模型的性能,并在不同的VSR平台上展示了通用性,表明其在其他多模态任务中的更广泛的应用潜力。摘要:Visual Speech Recognition (VSR) tasks are generally recognized to have a lower theoretical performance ceiling than Automatic Speech Recognition (ASR), owing to the inherent limitations of conveying semantic information visually. To mitigate this challenge, this paper introduces an advanced knowledge distillation approach using a Joint-Embedding Predictive Architecture (JEPA), named JEP-KD, designed to more effectively utilize audio features during model training. Central to JEP-KD is the inclusion of a generative network within the embedding layer, which enhances the video encoder's capacity for semantic feature extraction and brings it into closer alignment with the audio features from a pre-trained ASR model's encoder. This approach aims to progressively reduce the performance gap between VSR and ASR. Moreover, a comprehensive multimodal, multistage training regimen for the JEP-KD framework is established, bolstering the robustness and efficacy of the training process. Experiment results demonstrate that JEP-KD significantly improves the performance of VSR models and demonstrates versatility across different VSR platforms, indicating its potential for broader application within other multimodal tasks.【1】 Blind Identification of Binaural Room Impulse Responses from Smart Glasses作者:Thomas Deppisch,Nils Meyer-Kahlen,Sebastià V. Amengual Garí摘要:智能眼镜越来越被认为是增强现实的关键媒介,它提供了一个免提平台,集成了麦克风和非耳朵遮挡扬声器,可以将虚拟声源无缝地混合到现实世界的声学场景中。为了令人信服地集成虚拟声源,虚拟声源的室内声学渲染必须与真实世界的声学效果相匹配。然而,关于用户的声学环境的信息通常不可用。这项工作使用一副智能眼镜中的麦克风阵列,从现实世界环境中的几秒钟语音中盲识别双耳房间脉冲响应(BRIR)。所提出的方法使用去混响和波束成形来生成伪参考信号,该伪参考信号由多通道维纳滤波器用于估计房间脉冲响应,然后将房间脉冲响应转换为BRIR。多通道房间脉冲响应可以用来估计房间声学参数,这被证明是优于基线算法在混响时间和直接混响能量比的估计。来自收听实验的结果进一步指示所估计的BRIR通常比来自具有类似几何形状的其它房间的所测量的BRIR在感知上更令人信服地再现真实世界房间声学。摘要:Smart glasses are increasingly recognized as a key medium for augmented reality, offering a hands-free platform with integrated microphones and non-ear-occluding loudspeakers to seamlessly mix virtual sound sources into the real-world acoustic scene. To convincingly integrate virtual sound sources, the room acoustic rendering of the virtual sources must match the real-world acoustics. Information about a user's acoustic environment however is typically not available. This work uses a microphone array in a pair of smart glasses to blindly identify binaural room impulse responses (BRIRs) from a few seconds of speech in the real-world environment. The proposed method uses dereverberation and beamforming to generate a pseudo reference signal that is used by a multichannel Wiener filter to estimate room impulse responses which are then converted to BRIRs. The multichannel room impulse responses can be used to estimate room acoustic parameters which is shown to outperform baseline algorithms in the estimation of reverberation time and direct-to-reverberant energy ratio. Results from a listening experiment further indicate that the estimated BRIRs often reproduce the real-world room acoustics perceptually more convincingly than measured BRIRs from other rooms with similar geometry.
【2】 LV-CTC: Non-autoregressive ASR with CTC and latent variable models标题:LV—CTC:非自回归ASR与CTC和潜在变量模型作者:Yuya Fujita,Shinji Watanabe,Xuankai Chang,Takashi Maekaku摘要:用于自动语音识别(ASR)的非自回归(NAR)模型旨在通过简化传统模型的自回归(AR)生成过程来实现高精度和快速推理。连接主义时态分类(CTC)是NAR ASR模型的关键技术之一。在本文中,我们提出了一个新的模型结合CTC和一个潜在的变量模型,这是一个国家的最先进的模型在神经机器翻译研究领域。介绍了一种新的用于汽车自动识别的神经网络结构和公式。在所提出的模型中,CTC对齐被假设为依赖于潜在变量,这些潜在变量被期望捕获令牌之间的依赖关系。在100小时的Libri语音语料集上的实验结果表明,基于CTC的NAR模型的识别准确率最高。在TED-LIUM 2语料库上,实现了最佳识别精度,包括具有更快推理速度的AR E2 E模型。摘要:Non-autoregressive (NAR) models for automatic speech recognition (ASR) aim to achieve high accuracy and fast inference by simplifying the autoregressive (AR) generation process of conventional models. Connectionist temporal classification (CTC) is one of the key techniques used in NAR ASR models. In this paper, we propose a new model combining CTC and a latent variable model, which is one of the state-of-the-art models in the neural machine translation research field. A new neural network architecture and formulation specialized for ASR application are introduced. In the proposed model, CTC alignment is assumed to be dependent on the latent variables that are expected to capture dependencies between tokens. Experimental results on a 100 hours subset of Librispeech corpus showed the best recognition accuracy among CTC-based NAR models. On the TED-LIUM2 corpus, the best recognition accuracy is achieved including AR E2E models with faster inference speed.【3】 Siamese Vision Transformers are Scalable Audio-visual Learners标题:Siamese Vision Transformers是可扩展的视听学习者作者:Yan-Bo Lin,Gedas Bertasius摘要:传统的视听方法依赖于独立的音频和视频骨干,这是昂贵的,不可扩展的。在这项工作中,我们研究使用视听连体网络(AVSiam)进行高效和可扩展的视听预训练。我们的框架使用一个共享的Vision Transformer主干来处理音频和视频输入,提高了其参数效率,减少了GPU内存占用,并允许我们将我们的方法扩展到更大的数据集和模型大小。我们使用具有多比率随机掩蔽方案的对比视听匹配目标来预训练我们的模型,这使我们的模型能够处理更大的视听实例批次,有助于对比学习。与以前的视听方法不同,我们的方法可以鲁棒地处理音频,视频和视听输入与一个单一的共享ViT骨干。此外,尽管两种模式都使用了共享的主干,AVSiam在视听分类和检索方面取得了比AudioSet和VGGSound更好的结果。我们的代码可在https://github.com/GenjiB/AVSiam上获得摘要:Traditional audio-visual methods rely on independent audio and visual backbones, which is costly and not scalable. In this work, we investigate using an audio-visual siamese network (AVSiam) for efficient and scalable audio-visual pretraining. Our framework uses a single shared vision transformer backbone to process audio and visual inputs, improving its parameter efficiency, reducing the GPU memory footprint, and allowing us to scale our method to larger datasets and model sizes. We pretrain our model using a contrastive audio-visual matching objective with a multi-ratio random masking scheme, which enables our model to process larger audio-visual instance batches, helpful for contrastive learning. Unlike prior audio-visual methods, our method can robustly handle audio, visual, and audio-visual inputs with a single shared ViT backbone. Furthermore, despite using the shared backbone for both modalities, AVSiam achieves competitive or even better results than prior methods on AudioSet and VGGSound for audio-visual classification and retrieval. Our code is available at https://github.com/GenjiB/AVSiam【4】 Asymmetric and trial-dependent modeling: the contribution of LIA to SdSV Challenge Task 2标题:非对称和试验相关建模:LIA对SdSV挑战任务2的贡献作者:Pierre-Michel Bousquet,Mickael Rouvier备注:LIA system description for the Short Duration Speaker Verification (SdSv) challenge 2020 Task 2摘要:SdSv挑战任务2提供了一个机会,以评估现代文本无关的说话人确认系统的效率和鲁棒性。但这也使得有可能测试新的方法,能够考虑到这一挑战的主要问题(期限、语言.)。本文介绍了我们实验室在说话人识别领域的贡献。这些贡献突出了除了短持续时间和语言之外的另外两个挑战:招募和测试数据之间的不匹配以及评估试验数据集子集之间的不匹配。实验表明,所提出的方法的相关性和效率的SdSv评估,并可能在许多现实生活中的应用感兴趣。摘要:The SdSv challenge Task 2 provided an opportunity to assess efficiency and robustness of modern text-independent speaker verification systems. But it also made it possible to test new approaches, capable of taking into account the main issues of this challenge (duration, language, ...). This paper describes the contributions of our laboratory to the speaker recognition field. These contributions highlight two other challenges in addition to short-duration and language: the mismatch between enrollment and test data and the one between subsets of the evaluation trial dataset. The proposed approaches experimentally show their relevance and efficiency on the SdSv evaluation, and could be of interest in many real-life applications.【5】 Phonetic Segmentation of the UCLA Phonetics Lab Archive标题:加州大学洛杉矶分校语音实验室档案中的语音分割作者:Eleanor Chodroff,Blaž Pažon,Annie Baker,Steven Moran备注:Accepted at LREC-COLING 2024摘要:语音技术和比较语言学的研究依赖于获得多样化和可访问的语音数据。加州大学洛杉矶分校语音学实验室档案是最早的多语种语音语料库之一,具有314种语言的长格式音频记录和语音音译(Ladefoged et al.,2009年)。最近,这些语言中的95种与单词级语音音译时间对齐(Li等人,2021年)。在这里,我们提出了VoxAngeles,一个语料库的审计语音音译和语音水平的调整加州大学洛杉矶分校语音实验室档案,它使用95种语言的CMU重新发布作为我们的出发点。VoxAngeles还包括来自原始UCLA语料库的单词和音素级分割,以及单词和音素持续时间,元音共振峰和元音f0的语音测量。该语料库增强了原始数据的可用性,特别是对于定量语音类型学,如通过元音内在f0的案例研究所示。我们还讨论了跨语言语音学的一般研究和教学的VoxAngeles语料库的效用,以及低资源和多语言语音技术。VoxAngeles可以在CC-BY-NC 4.0许可下免费下载和使用。摘要:Research in speech technologies and comparative linguistics depends on access to diverse and accessible speech data. The UCLA Phonetics Lab Archive is one of the earliest multilingual speech corpora, with long-form audio recordings and phonetic transcriptions for 314 languages (Ladefoged et al., 2009). Recently, 95 of these languages were time-aligned with word-level phonetic transcriptions (Li et al., 2021). Here we present VoxAngeles, a corpus of audited phonetic transcriptions and phone-level alignments of the UCLA Phonetics Lab Archive, which uses the 95-language CMU re-release as our starting point. VoxAngeles also includes word- and phone-level segmentations from the original UCLA corpus, as well as phonetic measurements of word and phone durations, vowel formants, and vowel f0. This corpus enhances the usability of the original data, particularly for quantitative phonetic typology, as demonstrated through a case study of vowel intrinsic f0. We also discuss the utility of the VoxAngeles corpus for general research and pedagogy in crosslinguistic phonetics, as well as for low-resource and multilingual speech technologies. VoxAngeles is free to download and use under a CC-BY-NC 4.0 license.
【6】 A Novel Stochastic Transformer-based Approach for Post-Traumatic Stress Disorder Detection using Audio Recording of Clinical Interviews作者:Mamadou Dia,Ghazaleh Khodabandelou,Alice Othmani摘要:创伤后应激障碍(PTSD)是一种精神障碍,可以在目睹或经历极端创伤事件后发展。PTSD可以影响任何人,无论种族或文化。据估计,每11个人中就有一个人在一生中会经历PTSD。临床医生管理的PTSD量表(CAPS)和平民PTSD检查表(PCL-C)访谈是诊断PTSD的金标准。这些问卷可能会被受试者的回答所欺骗。这项工作提出了一种基于深度学习的方法,该方法在临床访谈期间使用音频记录实现了最先进的PTSD检测性能。我们的方法是基于从临床访谈的音频记录中提取的MFCC低级特征,然后使用随机Transformer进行深度高级学习。我们提出的方法在eDAIC数据集上实现了最先进的性能,RMSE为2.92,这要归功于随机深度、随机深度学习层和随机激活函数。摘要:Post-traumatic stress disorder (PTSD) is a mental disorder that can be developed after witnessing or experiencing extremely traumatic events. PTSD can affect anyone, regardless of ethnicity, or culture. An estimated one in every eleven people will experience PTSD during their lifetime. The Clinician-Administered PTSD Scale (CAPS) and the PTSD Check List for Civilians (PCL-C) interviews are gold standards in the diagnosis of PTSD. These questionnaires can be fooled by the subject's responses. This work proposes a deep learning-based approach that achieves state-of-the-art performances for PTSD detection using audio recordings during clinical interviews. Our approach is based on MFCC low-level features extracted from audio recordings of clinical interviews, followed by deep high-level learning using a Stochastic Transformer. Our proposed approach achieves state-of-the-art performances with an RMSE of 2.92 on the eDAIC dataset thanks to the stochastic depth, stochastic deep learning layers, and stochastic activation function.【7】 Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition作者:Siyuan Shen,Yu Gao,Feng Liu,Hanyang Wang,Aimin Zhou备注:Accepted by 49th IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2024)摘要:语音情感识别的主流范式是识别整个话语的单个情感标签。这一系列的工作忽略了情感动态在细时间粒度和大多数未能充分利用语音信号的语言信息显式。在本文中,我们提出了情感神经传感器的细粒度语音情感识别与自动语音识别(ASR)联合训练。我们首先扩展了典型的神经传感器与情感联合网络,构建情感格细粒度的SER,然后我们提出了对齐格上的格最大池,以方便区分情感和非情感帧。为了使细粒度的SER适应传感器推理方式,我们进一步将ASR的特殊符号blank作为底层情感指示器,得到了因子化的情感神经传感器。对于典型的话语级SER,我们的ENT模型优于国家的最先进的方法IEMOCAP在低字错误率。在IEMOCAP和最新的语音情感日志数据集ZED上的实验也证明了细粒度情感建模的优越性。我们的代码可在https://github.com/ECNU-Cross-Innovation-Lab/ENT上获得。摘要:The mainstream paradigm of speech emotion recognition (SER) is identifying the single emotion label of the entire utterance. This line of works neglect the emotion dynamics at fine temporal granularity and mostly fail to leverage linguistic information of speech signal explicitly. In this paper, we propose Emotion Neural Transducer for fine-grained speech emotion recognition with automatic speech recognition (ASR) joint training. We first extend typical neural transducer with emotion joint network to construct emotion lattice for fine-grained SER. Then we propose lattice max pooling on the alignment lattice to facilitate distinguishing emotional and non-emotional frames. To adapt fine-grained SER to transducer inference manner, we further make blank, the special symbol of ASR, serve as underlying emotion indicator as well, yielding Factorized Emotion Neural Transducer. For typical utterance-level SER, our ENT models outperform state-of-the-art methods on IEMOCAP in low word error rate. Experiments on IEMOCAP and the latest speech emotion diarization dataset ZED also demonstrate the superiority of fine-grained emotion modeling. Our code is available at https://github.com/ECNU-Cross-Innovation-Lab/ENT.
【8】 Robust Active Speaker Detection in Noisy Environments作者:Siva Sai Nagender Vasireddy,Chenxu Zhang,Xiaohu Guo,Yapeng Tian摘要:针对噪声环境下的主动说话人检测问题,提出了一种鲁棒的主动说话人检测方法。现有的ASD方法利用音频和视觉模态,但周围环境中的非语音声音会对性能产生负面影响。为了克服这一点,我们提出了一个新的框架,利用视听语音分离作为指导,学习无噪声的音频功能。这些功能然后在ASD模型中使用,并且这两个任务在端到端框架中联合优化。我们提出的框架减轻了残留噪声和音频质量降低的问题,可以发生在一个天真的级联两阶段的框架,直接使用分离的语音ASD,并使这两个任务同时优化。为了进一步增强音频特征的鲁棒性和处理固有的语音噪声,我们提出了一种动态加权损失的方法来训练语音分离器。我们还收集了一个真实世界的噪音音频数据集,以方便调查。实验表明,非语音音频噪声显着影响ASD模型,我们提出的方法提高了ASD在噪声环境中的性能。该框架是通用的,可以应用于不同的ASD方法,以提高其鲁棒性。我们的代码、模型和数据将被发布。摘要:This paper addresses the issue of active speaker detection (ASD) in noisy environments and formulates a robust active speaker detection (rASD) problem. Existing ASD approaches leverage both audio and visual modalities, but non-speech sounds in the surrounding environment can negatively impact performance. To overcome this, we propose a novel framework that utilizes audio-visual speech separation as guidance to learn noise-free audio features. These features are then utilized in an ASD model, and both tasks are jointly optimized in an end-to-end framework. Our proposed framework mitigates residual noise and audio quality reduction issues that can occur in a naive cascaded two-stage framework that directly uses separated speech for ASD, and enables the two tasks to be optimized simultaneously. To further enhance the robustness of the audio features and handle inherent speech noises, we propose a dynamic weighted loss approach to train the speech separator. We also collected a real-world noise audio dataset to facilitate investigations. Experiments demonstrate that non-speech audio noises significantly impact ASD models, and our proposed approach improves ASD performance in noisy environments. The framework is general and can be applied to different ASD approaches to improve their robustness. Our code, models, and data will be released.【9】 JEP-KD: Joint-Embedding Predictive Architecture Based Knowledge Distillation for Visual Speech Recognition标题:JEP—KD:基于联合嵌入预测结构的视觉语音识别知识提取作者:Chang Sun,Hong Yang,Bo Qin摘要:视觉语音识别(VSR)任务通常被认为具有比自动语音识别(ASR)更低的理论性能上限,这是由于视觉传达语义信息的固有限制。为了缓解这一挑战,本文介绍了一种使用联合嵌入预测架构(JEPA)的高级知识蒸馏方法,称为JEP-KD,旨在更有效地利用模型训练期间的音频特征。JEP-KD的核心是在嵌入层中包含一个生成网络,这增强了视频编码器的语义特征提取能力,并使其与来自预训练ASR模型编码器的音频特征更接近。这种方法旨在逐步缩小VSR和ASR之间的性能差距。此外,还为JEP-KD框架建立了一个全面的多模式、多阶段训练方案,增强了训练过程的鲁棒性和有效性。实验结果表明,JEP-KD显著提高了VSR模型的性能,并在不同的VSR平台上展示了通用性,表明其在其他多模态任务中的更广泛的应用潜力。摘要:Visual Speech Recognition (VSR) tasks are generally recognized to have a lower theoretical performance ceiling than Automatic Speech Recognition (ASR), owing to the inherent limitations of conveying semantic information visually. To mitigate this challenge, this paper introduces an advanced knowledge distillation approach using a Joint-Embedding Predictive Architecture (JEPA), named JEP-KD, designed to more effectively utilize audio features during model training. Central to JEP-KD is the inclusion of a generative network within the embedding layer, which enhances the video encoder's capacity for semantic feature extraction and brings it into closer alignment with the audio features from a pre-trained ASR model's encoder. This approach aims to progressively reduce the performance gap between VSR and ASR. Moreover, a comprehensive multimodal, multistage training regimen for the JEP-KD framework is established, bolstering the robustness and efficacy of the training process. Experiment results demonstrate that JEP-KD significantly improves the performance of VSR models and demonstrates versatility across different VSR platforms, indicating its potential for broader application within other multimodal tasks.