今天跟大家分享一篇语音相关的论文合集:cs.SD语音7篇,eess.AS音频处理7篇。本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
【1】 Contrastive Audio-Language Learning for Music
标题:音乐视听语言对比学习
链接:https://arxiv.org/abs/2208.12208
作者:Ilaria Manco,Emmanouil Benetos,Elio Quinton,György Fazekas
机构:George Fazekas, School of EECS, Queen Mary University of London, London, U.K, Music & Audio Machine Learning Lab, Universal Music Group, London, U.K.
备注:Accepted to ISMIR 2022
摘要:作为人类已知的最直观的界面之一,自然语言具有调解许多涉及人机交互的任务的潜力,特别是在以应用为中心的领域,如音乐信息检索。在这项工作中,我们探索跨通道学习,试图在音乐领域的音频和语言的桥梁。为此,我们提出了MusCALL,一个音乐对比音频语言学习的框架。我们的方法由一个双编码器架构组成,它学习音乐音频和描述性句子对之间的对齐,产生多模态嵌入,可用于文本到音频和音频到文本的开箱即用检索。由于这个属性,MusCALL可以被转移到几乎任何可以被转换为基于文本的检索的任务。我们的实验表明,我们的方法在检索与文本描述匹配的音频以及与音频查询匹配的文本方面的性能明显优于基线。我们还证明了我们的模型的多模态对齐能力可以成功地扩展到两个公共数据集上的体裁分类和自动标注的zero-shot传输场景。
摘要:As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information Retrieval. In this work, we explore cross-modal learning in an attempt to bridge audio and language in the music domain. To this end, we propose MusCALL, a framework for Music Contrastive Audio-Language Learning. Our approach consists of a dual-encoder architecture that learns the alignment between pairs of music audio and descriptive sentences, producing multimodal embeddings that can be used for text-to-audio and audio-to-text retrieval out-of-the-box. Thanks to this property, MusCALL can be transferred to virtually any task that can be cast as text-based retrieval. Our experiments show that our method performs significantly better than the baselines at retrieving audio that matches a textual description and, conversely, text that matches an audio query. We also demonstrate that the multimodal alignment capability of our model can be successfully extended to the zero-shot transfer scenario for genre classification and auto-tagging on two public datasets.
【2】 The ReprGesture entry to the GENEA Challenge 2022
标题:2022年Genea挑战赛的ReprGesture条目
链接:https://arxiv.org/abs/2208.12133
作者:Sicheng Yang,Zhiyong Wu,Minglei Li,Mengchen Zhao,Jiuxin Lin,Liyang Chen,Weihong Bao机构:Tsinghua University, China and The Chinese University of Hong Kong备注:8 pages, 4 figures, ICMI 2022摘要:本文介绍了2022年体现代理非语言行为的生成和评估(GENEA)挑战赛的ReprGesture参赛作品。GENEA挑战赛提供经过处理的数据集,并执行众包评估,以比较不同手势生成系统的性能。本文研究了一种基于多模态表示学习的手势自动生成系统。我们对音频使用WavLM特征,对文本使用FastText特征,对手势使用位置和旋转矩阵特征。每个模态被投影到两个不同的子空间:形态不变和形态特定。在训练过程中,使用基于梯度反转层的对抗性分类器和模态重构解码器,学习模态间不变量的共性,并捕获模态特定表示的特征。手势解码器使用与音频中的节奏相关的所有表示和特征来生成适当的手势。我们的代码、预先培训的模型和演示可在www.example.com上获得https://github.com/YoungSeng/ReprGesture。摘要:This paper describes the ReprGesture entry to the Generation and Evaluation of Non-verbal Behaviour for Embodied Agents (GENEA) challenge 2022. The GENEA challenge provides the processed datasets and performs crowdsourced evaluations to compare the performance of different gesture generation systems. In this paper, we explore an automatic gesture generation system based on multimodal representation learning. We use WavLM features for audio, FastText features for text and position and rotation matrix features for gesture. Each modality is projected to two distinct subspaces: modality-invariant and modality-specific. To learn inter-modality-invariant commonalities and capture the characters of modality-specific representations, gradient reversal layer based adversarial classifier and modality reconstruction decoders are used during training. The gesture decoder generates proper gestures using all representations and features related to the rhythm in the audio. Our code, pre-trained models and demo are available at https://github.com/YoungSeng/ReprGesture.
【3】 A Study on Broadcast Networks for Music Genre Classification
标题:音乐流派分类的广播网研究
链接:https://arxiv.org/abs/2208.12086
作者:Ahmed Heakl,Abdelrahman Abdelgawad,Victor Parque机构:Department of Computer Science, Egypt-Japan University of Science and Technology, Alexandria, Egypt, Department of Mechatronics and Robotics, Egypt-Japan University of Science and Technology, Alexandria, Egypt备注:accepted for oral presentation at the World Congress on Computational Intelligence (WCCI 2022) - International Joint Conference on Neural Networks (IJCNN 2022)摘要:随着人们对音乐流媒体/推荐服务需求的不断增长以及音乐信息检索框架的不断发展,音乐流派分类(Music Genre Classification,MGC)已经引起了社会各界的广泛关注。然而,已知基于卷积的方法缺乏有效地编码和定位时间特征的能力。本文研究了基于广播的神经网络,旨在提高其在小参数集(约180k)下的局部化和泛化能力,并研究了12种广播网络,讨论了块结构、池方法、激活函数、归一化机制、标记平滑、信道相关性、LSTM块包含和初始方案的变体对网络性能的影响。我们使用GTZAN、Extended Ballroom、HOMBURG和Free Music Archive(FMA)等相关数据集进行的计算实验表明,该方法在音乐流派分类中具有最高的分类精度.我们的方法提供了见解和潜力,使紧凑和通用的广播网络的音乐和音频分类。摘要:Due to the increased demand for music streaming/recommender services and the recent developments of music information retrieval frameworks, Music Genre Classification (MGC) has attracted the community's attention. However, convolutional-based approaches are known to lack the ability to efficiently encode and localize temporal features. In this paper, we study the broadcast-based neural networks aiming to improve the localization and generalizability under a small set of parameters (about 180k) and investigate twelve variants of broadcast networks discussing the effect of block configuration, pooling method, activation function, normalization mechanism, label smoothing, channel interdependency, LSTM block inclusion, and variants of inception schemes. Our computational experiments using relevant datasets such as GTZAN, Extended Ballroom, HOMBURG, and Free Music Archive (FMA) show state-of-the-art classification accuracies in Music Genre Classification. Our approach offers insights and the potential to enable compact and generalizable broadcast networks for music and audio classification.
【4】 Digital Audio Tampering Detection Based on ENF Spatio-temporal Features Representation Learning
标题:基于神经网络时空特征表示学习的数字音频篡改检测
链接:https://arxiv.org/abs/2208.11920
作者:Chunyan Zeng,Shuai Kong,Zhifeng Wang,Xiangkui Wan,Yunfan Chen机构:Hubei Key Laboratory for High-efficiency Utilization of Solar Energy and Operation, Control of Energy Storage System, Hubei University of Technology, Nanli Road , Department of Digital Media Technology, Central China Normal University, Luoyu Road摘要:基于电网频率的数字音频篡改检测方法大多只利用电网频率的静态空间信息,忽略了电网频率在时间序列上的变化,限制了电网频率特征的表达能力,降低了篡改检测的准确性。提出了一种基于ENF时空特征表示学习的数字音频篡改检测方法。利用CNN和BiLSTM构建并行时空网络模型,深度提取ENF空间特征信息和ENF时间特征信息,增强特征表示能力,提高篡改检测准确率。为了提取ENF信号的时空特征,首先利用数字音频高精度离散傅里叶变换分析提取ENF信号的相位序列。通过自适应帧移位将非等相位序列划分为帧,以获得相同大小的特征矩阵来表示ENF的空间特征。同时,基于ENF时间变化信息将相位序列划分为帧,以表示ENF的时间特征。然后分别利用CNN和BiLSTM进一步提取深层时空特征,并利用注意力机制自适应地为深层时空特征分配权重,得到具有更强表示能力的时空特征.最后,利用深度神经网络判断音频是否被篡改。实验结果表明,-7.12在新西班牙语Carioca公共数据库上,该方法与现有方法相比,准确率提高了2.12% www.example.com %.摘要:Most digital audio tampering detection methods based on electrical network frequency (ENF) only utilize the static spatial information of ENF, ignoring the variation of ENF in time series, which limit the ability of ENF feature representation and reduce the accuracy of tampering detection. This paper proposes a new method for digital audio tampering detection based on ENF spatio-temporal features representation learning. A parallel spatio-temporal network model is constructed using CNN and BiLSTM, which deeply extracts ENF spatial feature information and ENF temporal feature information to enhance the feature representation capability to improve the tampering detection accuracy. In order to extract the spatial and temporal features of the ENF, this paper firstly uses digital audio high-precision Discrete Fourier Transform analysis to extract the phase sequences of the ENF. The unequal phase series is divided into frames by adaptive frame shifting to obtain feature matrices of the same size to represent the spatial features of the ENF. At the same time, the phase sequences are divided into frames based on ENF time changes information to represent the temporal features of the ENF. Then deep spatial and temporal features are further extracted using CNN and BiLSTM respectively, and an attention mechanism is used to adaptively assign weights to the deep spatial and temporal features to obtain spatio-temporal features with stronger representation capability. Finally, the deep neural network is used to determine whether the audio has been tampered with. The experimental results show that the proposed method improves the accuracy by 2.12%-7.12% compared with state-of-the-art methods under the public database Carioca, New Spanish.
【5】 Interpretable Multimodal Emotion Recognition using Hybrid Fusion of Speech and Image Data
标题:基于语音和图像数据混合融合的可解释多通道情感识别
链接:https://arxiv.org/abs/2208.11868
作者:Puneet Kumar,Sarthak Malik,Balasubramanian Raman机构:Computer Science & Engineering Department, Indian Institute of Technology, Roorkee, India, Electrical Engineering Department, Indian Institute of Technology, Roorkee, India, Article history:, Affective Computing, Multimodal Anal-备注:arXiv admin note: text overlap with arXiv:2208.11450摘要:提出了一种基于混合融合的多模态情感识别系统,该系统将语音和相应图像所描述的情感进行离散分类。一种新的可解释性技术已经被开发出来,以识别重要的语音和图像特征,从而预测特定的情感类别。申报系统的架构已通过大量消融研究确定。它融合语音和图像特征,然后将语音、图像和中间融合输出进行组合。所提出的可解释性技术结合了分治方法来计算表示每个语音和图像特征的重要性的形状值。我们还构建了一个大规模的数据集(IIT-R SIER数据集),它由语音话语、对应的图像和类别标签组成,即:“愤怒”、"快乐“、"憎恨”和“悲伤”。该系统对情感识别的准确率达到了83.29%。所提出的系统的增强的性能表明利用来自多个模态的互补信息进行情绪识别的重要性。摘要:This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been developed to identify the important speech & image features leading to the prediction of particular emotion classes. The proposed system's architecture has been determined through intensive ablation studies. It fuses the speech & image features and then combines speech, image, and intermediate fusion outputs. The proposed interpretability technique incorporates the divide & conquer approach to compute shapely values denoting each speech & image feature's importance. We have also constructed a large-scale dataset (IIT-R SIER dataset), consisting of speech utterances, corresponding images, and class labels, i.e., 'anger,' 'happy,' 'hate,' and 'sad.' The proposed system has achieved 83.29% accuracy for emotion recognition. The enhanced performance of the proposed system advocates the importance of utilizing complementary information from multiple modalities for emotion recognition.
【6】 IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian languages
标题:INDICSUPERB:一种面向印度语言的语音处理通用性能基准
链接:https://arxiv.org/abs/2208.11761
作者:Tahir Javed,Kaushal Santosh Bhogale,Abhigyan Raman,Anoop Kunchukuttan,Pratyush Kumar,Mitesh M. Khapra机构:Indian Institute of Technology, Madras, AI,Bharat, Microsoft摘要:人工智能研究的一个基石是创建和采用标准化的训练和测试数据集,以标记最先进模型的进展。一个特别成功的例子是用于训练和评估英语的自然语言理解(NLU)模型的GLUE数据集。基于自监督BERT的语言模型的大量研究围绕着GLUE中NLU任务的性能改进。为了评估其他语言中的语言模型,创建了几个特定语言的GLUE数据集。言语语言理解(SLU)领域也遵循了类似的轨迹。诸如wav2vec2之类的大型自监督模型的成功使得能够创建具有相对容易访问未标记数据的语音模型。然后,可以在SLU任务(如SUPERB基准测试)上评估这些模型。在本文中,我们通过发布IndicSUPERB基准测试将其扩展到印度语。具体来说,我们做出了以下三点贡献。(i)我们从印度203个地区的1,218个贡献者那里收集了包含1,684小时的12种印度语言的标记语音数据的Kathbath。(ii)使用Kathbath,我们创建了6个语音任务的基准:自动语音识别、说话人验证、说话人识别(单声道/多声道)、语言识别、示例查询和12种语言的关键字定位。(iii)在已发布的基准测试中,我们使用常用的基准FBANK来训练和评估不同的自监督模型。实验结果表明,在大多数任务中,语言识别模型的准确率都高于基线模型,其中在语言识别任务中,语言识别模型的准确率与基线模型的准确率差距高达76%。然而,对于说话人识别,在大数据集上训练的自监督模型显示出优势。我们希望IndicSUPERB能为印度语言的语音理解模型的发展做出贡献。摘要:A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. The large body of research around self-supervised BERT-based language models revolved around performance improvements on NLU tasks in GLUE. To evaluate language models in other languages, several language-specific GLUE datasets were created. The area of speech language understanding (SLU) has followed a similar trajectory. The success of large self-supervised models such as wav2vec2 enable creation of speech models with relatively easy to access unlabelled data. These models can then be evaluated on SLU tasks, such as the SUPERB benchmark. In this work, we extend this to Indic languages by releasing the IndicSUPERB benchmark. Specifically, we make the following three contributions. (i) We collect Kathbath containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India. (ii) Using Kathbath, we create benchmarks across 6 speech tasks: Automatic Speech Recognition, Speaker Verification, Speaker Identification (mono/multi), Language Identification, Query By Example, and Keyword Spotting for 12 languages. (iii) On the released benchmarks, we train and evaluate different self-supervised models alongside a commonly used baseline FBANK. We show that language-specific fine-tuned models are more accurate than baseline on most of the tasks, including a large gap of 76\% for the Language Identification task. However, for speaker identification, self-supervised models trained on large datasets demonstrate an advantage. We hope IndicSUPERB contributes to the progress of developing speech language understanding models for Indian languages.
【7】 Low-Level Physiological Implications of End-to-End Learning of Speech Recognition
标题:语音识别端到端学习的低水平生理学意义
链接:https://arxiv.org/abs/2208.11700
作者:Louise Coppieters de Gibson,Philip N. Garner机构:Idiap Research Institute, Martigny, Switzerland, Ecole polytechnique f´ed´erale de Lausanne (EPFL), Switzerland备注:Submitted to INTERSPEECH 2022摘要:从机器学习的观点来看,当前的语音识别体系结构表现得非常好,因此用户交互也是如此。这表明它们很好地模仿了人类的生物系统。我们研究的推论是否可以反过来提供洞察生物系统;特别是听询机制。使用SincNet,我们确认端到端系统确实学习了众所周知的滤波器组结构。然而,我们也表明,较宽的带宽滤波器在学习的结构中是重要的。虽然初始化窄带和宽带滤波器都可以获得一些好处,但生理限制表明,这种滤波器出现在中脑而不是耳蜗中。我们表明,标准的机器学习架构必须修改,以允许这一过程被神经地模拟。摘要:Current speech recognition architectures perform very well from the point of view of machine learning, hence user interaction. This suggests that they are emulating the human biological system well. We investigate whether the inference can be inverted to provide insights into that biological system; in particular the hearing mechanism. Using SincNet, we confirm that end-to-end systems do learn well known filterbank structures. However, we also show that wider band-width filters are important in the learned structure. Whilst some benefits can be gained by initialising both narrow and wide-band filters, physiological constraints suggest that such filters arise in mid-brain rather than the cochlea. We show that standard machine learning architectures must be modified to allow this process to be emulated neurally.
【1】 Contrastive Audio-Language Learning for Music
标题:音乐视听语言对比学习
链接:https://arxiv.org/abs/2208.12208
* 与cs.SD语音【1】为同一篇
作者:Ilaria Manco,Emmanouil Benetos,Elio Quinton,György Fazekas机构:George Fazekas, School of EECS, Queen Mary University of London, London, U.K, Music & Audio Machine Learning Lab, Universal Music Group, London, U.K.备注:Accepted to ISMIR 2022摘要:作为人类已知的最直观的界面之一,自然语言具有调解许多涉及人机交互的任务的潜力,特别是在以应用为中心的领域,如音乐信息检索。在这项工作中,我们探索跨通道学习,试图在音乐领域的音频和语言的桥梁。为此,我们提出了MusCALL,一个音乐对比音频语言学习的框架。我们的方法由一个双编码器架构组成,它学习音乐音频和描述性句子对之间的对齐,产生多模态嵌入,可用于文本到音频和音频到文本的开箱即用检索。由于这个属性,MusCALL可以被转移到几乎任何可以被转换为基于文本的检索的任务。我们的实验表明,我们的方法在检索与文本描述匹配的音频以及与音频查询匹配的文本方面的性能明显优于基线。我们还证明了我们的模型的多模态对齐能力可以成功地扩展到两个公共数据集上的体裁分类和自动标注的zero-shot传输场景。摘要:As one of the most intuitive interfaces known to humans, natural language has the potential to mediate many tasks that involve human-computer interaction, especially in application-focused fields like Music Information Retrieval. In this work, we explore cross-modal learning in an attempt to bridge audio and language in the music domain. To this end, we propose MusCALL, a framework for Music Contrastive Audio-Language Learning. Our approach consists of a dual-encoder architecture that learns the alignment between pairs of music audio and descriptive sentences, producing multimodal embeddings that can be used for text-to-audio and audio-to-text retrieval out-of-the-box. Thanks to this property, MusCALL can be transferred to virtually any task that can be cast as text-based retrieval. Our experiments show that our method performs significantly better than the baselines at retrieving audio that matches a textual description and, conversely, text that matches an audio query. We also demonstrate that the multimodal alignment capability of our model can be successfully extended to the zero-shot transfer scenario for genre classification and auto-tagging on two public datasets.
【2】 The ReprGesture entry to the GENEA Challenge 2022
标题:2022年Genea挑战赛的ReprGesture条目
链接:https://arxiv.org/abs/2208.12133
* 与cs.SD语音【2】为同一篇
作者:Sicheng Yang,Zhiyong Wu,Minglei Li,Mengchen Zhao,Jiuxin Lin,Liyang Chen,Weihong Bao机构:Tsinghua University, China and The Chinese University of Hong Kong备注:8 pages, 4 figures, ICMI 2022摘要:本文介绍了2022年体现代理非语言行为的生成和评估(GENEA)挑战赛的ReprGesture参赛作品。GENEA挑战赛提供经过处理的数据集,并执行众包评估,以比较不同手势生成系统的性能。本文研究了一种基于多模态表示学习的手势自动生成系统。我们对音频使用WavLM特征,对文本使用FastText特征,对手势使用位置和旋转矩阵特征。每个模态被投影到两个不同的子空间:形态不变和形态特定。在训练过程中,使用基于梯度反转层的对抗性分类器和模态重构解码器,学习模态间不变量的共性,并捕获模态特定表示的特征。手势解码器使用与音频中的节奏相关的所有表示和特征来生成适当的手势。我们的代码、预先培训的模型和演示可在www.example.com上获得https://github.com/YoungSeng/ReprGesture。摘要:This paper describes the ReprGesture entry to the Generation and Evaluation of Non-verbal Behaviour for Embodied Agents (GENEA) challenge 2022. The GENEA challenge provides the processed datasets and performs crowdsourced evaluations to compare the performance of different gesture generation systems. In this paper, we explore an automatic gesture generation system based on multimodal representation learning. We use WavLM features for audio, FastText features for text and position and rotation matrix features for gesture. Each modality is projected to two distinct subspaces: modality-invariant and modality-specific. To learn inter-modality-invariant commonalities and capture the characters of modality-specific representations, gradient reversal layer based adversarial classifier and modality reconstruction decoders are used during training. The gesture decoder generates proper gestures using all representations and features related to the rhythm in the audio. Our code, pre-trained models and demo are available at https://github.com/YoungSeng/ReprGesture.
【3】 A Study on Broadcast Networks for Music Genre Classification
标题:音乐流派分类的广播网研究
链接:https://arxiv.org/abs/2208.12086
* 与cs.SD语音【3】为同一篇
作者:Ahmed Heakl,Abdelrahman Abdelgawad,Victor Parque机构:Department of Computer Science, Egypt-Japan University of Science and Technology, Alexandria, Egypt, Department of Mechatronics and Robotics, Egypt-Japan University of Science and Technology, Alexandria, Egypt备注:accepted for oral presentation at the World Congress on Computational Intelligence (WCCI 2022) - International Joint Conference on Neural Networks (IJCNN 2022)摘要:随着人们对音乐流媒体/推荐服务需求的不断增长以及音乐信息检索框架的不断发展,音乐流派分类(Music Genre Classification,MGC)已经引起了社会各界的广泛关注。然而,已知基于卷积的方法缺乏有效地编码和定位时间特征的能力。本文研究了基于广播的神经网络,旨在提高其在小参数集(约180k)下的局部化和泛化能力,并研究了12种广播网络,讨论了块结构、池方法、激活函数、归一化机制、标记平滑、信道相关性、LSTM块包含和初始方案的变体对网络性能的影响。我们使用GTZAN、Extended Ballroom、HOMBURG和Free Music Archive(FMA)等相关数据集进行的计算实验表明,该方法在音乐流派分类中具有最高的分类精度.我们的方法提供了见解和潜力,使紧凑和通用的广播网络的音乐和音频分类。摘要:Due to the increased demand for music streaming/recommender services and the recent developments of music information retrieval frameworks, Music Genre Classification (MGC) has attracted the community's attention. However, convolutional-based approaches are known to lack the ability to efficiently encode and localize temporal features. In this paper, we study the broadcast-based neural networks aiming to improve the localization and generalizability under a small set of parameters (about 180k) and investigate twelve variants of broadcast networks discussing the effect of block configuration, pooling method, activation function, normalization mechanism, label smoothing, channel interdependency, LSTM block inclusion, and variants of inception schemes. Our computational experiments using relevant datasets such as GTZAN, Extended Ballroom, HOMBURG, and Free Music Archive (FMA) show state-of-the-art classification accuracies in Music Genre Classification. Our approach offers insights and the potential to enable compact and generalizable broadcast networks for music and audio classification.
【4】 Digital Audio Tampering Detection Based on ENF Spatio-temporal Features Representation Learning
标题:基于神经网络时空特征表示学习的数字音频篡改检测
链接:https://arxiv.org/abs/2208.11920
* 与cs.SD语音【4】为同一篇
作者:Chunyan Zeng,Shuai Kong,Zhifeng Wang,Xiangkui Wan,Yunfan Chen机构:Hubei Key Laboratory for High-efficiency Utilization of Solar Energy and Operation, Control of Energy Storage System, Hubei University of Technology, Nanli Road , Department of Digital Media Technology, Central China Normal University, Luoyu Road摘要:基于电网频率的数字音频篡改检测方法大多只利用电网频率的静态空间信息,忽略了电网频率在时间序列上的变化,限制了电网频率特征的表达能力,降低了篡改检测的准确性。提出了一种基于ENF时空特征表示学习的数字音频篡改检测方法。利用CNN和BiLSTM构建并行时空网络模型,深度提取ENF空间特征信息和ENF时间特征信息,增强特征表示能力,提高篡改检测准确率。为了提取ENF信号的时空特征,首先利用数字音频高精度离散傅里叶变换分析提取ENF信号的相位序列。通过自适应帧移位将非等相位序列划分为帧,以获得相同大小的特征矩阵来表示ENF的空间特征。同时,基于ENF时间变化信息将相位序列划分为帧,以表示ENF的时间特征。然后分别利用CNN和BiLSTM进一步提取深层时空特征,并利用注意力机制自适应地为深层时空特征分配权重,得到具有更强表示能力的时空特征.最后,利用深度神经网络判断音频是否被篡改。实验结果表明,-7.12在新西班牙语Carioca公共数据库上,该方法与现有方法相比,准确率提高了2.12% www.example.com %.摘要:Most digital audio tampering detection methods based on electrical network frequency (ENF) only utilize the static spatial information of ENF, ignoring the variation of ENF in time series, which limit the ability of ENF feature representation and reduce the accuracy of tampering detection. This paper proposes a new method for digital audio tampering detection based on ENF spatio-temporal features representation learning. A parallel spatio-temporal network model is constructed using CNN and BiLSTM, which deeply extracts ENF spatial feature information and ENF temporal feature information to enhance the feature representation capability to improve the tampering detection accuracy. In order to extract the spatial and temporal features of the ENF, this paper firstly uses digital audio high-precision Discrete Fourier Transform analysis to extract the phase sequences of the ENF. The unequal phase series is divided into frames by adaptive frame shifting to obtain feature matrices of the same size to represent the spatial features of the ENF. At the same time, the phase sequences are divided into frames based on ENF time changes information to represent the temporal features of the ENF. Then deep spatial and temporal features are further extracted using CNN and BiLSTM respectively, and an attention mechanism is used to adaptively assign weights to the deep spatial and temporal features to obtain spatio-temporal features with stronger representation capability. Finally, the deep neural network is used to determine whether the audio has been tampered with. The experimental results show that the proposed method improves the accuracy by 2.12%-7.12% compared with state-of-the-art methods under the public database Carioca, New Spanish.
【5】 Interpretable Multimodal Emotion Recognition using Hybrid Fusion of Speech and Image Data
标题:基于语音和图像数据混合融合的可解释多通道情感识别
链接:https://arxiv.org/abs/2208.11868
* 与cs.SD语音【5】为同一篇
作者:Puneet Kumar,Sarthak Malik,Balasubramanian Raman机构:Computer Science & Engineering Department, Indian Institute of Technology, Roorkee, India, Electrical Engineering Department, Indian Institute of Technology, Roorkee, India, Article history:, Affective Computing, Multimodal Anal-备注:arXiv admin note: text overlap with arXiv:2208.11450摘要:提出了一种基于混合融合的多模态情感识别系统,该系统将语音和相应图像所描述的情感进行离散分类。一种新的可解释性技术已经被开发出来,以识别重要的语音和图像特征,从而预测特定的情感类别。申报系统的架构已通过大量消融研究确定。它融合语音和图像特征,然后将语音、图像和中间融合输出进行组合。所提出的可解释性技术结合了分治方法来计算表示每个语音和图像特征的重要性的形状值。我们还构建了一个大规模的数据集(IIT-R SIER数据集),它由语音话语、对应的图像和类别标签组成,即:“愤怒”、"快乐“、"憎恨”和“悲伤”。该系统对情感识别的准确率达到了83.29%。所提出的系统的增强的性能表明利用来自多个模态的互补信息进行情绪识别的重要性。摘要:This paper proposes a multimodal emotion recognition system based on hybrid fusion that classifies the emotions depicted by speech utterances and corresponding images into discrete classes. A new interpretability technique has been developed to identify the important speech & image features leading to the prediction of particular emotion classes. The proposed system's architecture has been determined through intensive ablation studies. It fuses the speech & image features and then combines speech, image, and intermediate fusion outputs. The proposed interpretability technique incorporates the divide & conquer approach to compute shapely values denoting each speech & image feature's importance. We have also constructed a large-scale dataset (IIT-R SIER dataset), consisting of speech utterances, corresponding images, and class labels, i.e., 'anger,' 'happy,' 'hate,' and 'sad.' The proposed system has achieved 83.29% accuracy for emotion recognition. The enhanced performance of the proposed system advocates the importance of utilizing complementary information from multiple modalities for emotion recognition.
【6】 IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian languages
标题:INDICSUPERB:一种面向印度语言的语音处理通用性能基准
链接:https://arxiv.org/abs/2208.11761
* 与cs.SD语音【6】为同一篇
作者:Tahir Javed,Kaushal Santosh Bhogale,Abhigyan Raman,Anoop Kunchukuttan,Pratyush Kumar,Mitesh M. Khapra机构:Indian Institute of Technology, Madras, AI,Bharat, Microsoft摘要:人工智能研究的一个基石是创建和采用标准化的训练和测试数据集,以标记最先进模型的进展。一个特别成功的例子是用于训练和评估英语的自然语言理解(NLU)模型的GLUE数据集。基于自监督BERT的语言模型的大量研究围绕着GLUE中NLU任务的性能改进。为了评估其他语言中的语言模型,创建了几个特定语言的GLUE数据集。言语语言理解(SLU)领域也遵循了类似的轨迹。诸如wav2vec2之类的大型自监督模型的成功使得能够创建具有相对容易访问未标记数据的语音模型。然后,可以在SLU任务(如SUPERB基准测试)上评估这些模型。在本文中,我们通过发布IndicSUPERB基准测试将其扩展到印度语。具体来说,我们做出了以下三点贡献。(i)我们从印度203个地区的1,218个贡献者那里收集了包含1,684小时的12种印度语言的标记语音数据的Kathbath。(ii)使用Kathbath,我们创建了6个语音任务的基准:自动语音识别、说话人验证、说话人识别(单声道/多声道)、语言识别、示例查询和12种语言的关键字定位。(iii)在已发布的基准测试中,我们使用常用的基准FBANK来训练和评估不同的自监督模型。实验结果表明,在大多数任务中,语言识别模型的准确率都高于基线模型,其中在语言识别任务中,语言识别模型的准确率与基线模型的准确率差距高达76%。然而,对于说话人识别,在大数据集上训练的自监督模型显示出优势。我们希望IndicSUPERB能为印度语言的语音理解模型的发展做出贡献。摘要:A cornerstone in AI research has been the creation and adoption of standardized training and test datasets to earmark the progress of state-of-the-art models. A particularly successful example is the GLUE dataset for training and evaluating Natural Language Understanding (NLU) models for English. The large body of research around self-supervised BERT-based language models revolved around performance improvements on NLU tasks in GLUE. To evaluate language models in other languages, several language-specific GLUE datasets were created. The area of speech language understanding (SLU) has followed a similar trajectory. The success of large self-supervised models such as wav2vec2 enable creation of speech models with relatively easy to access unlabelled data. These models can then be evaluated on SLU tasks, such as the SUPERB benchmark. In this work, we extend this to Indic languages by releasing the IndicSUPERB benchmark. Specifically, we make the following three contributions. (i) We collect Kathbath containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India. (ii) Using Kathbath, we create benchmarks across 6 speech tasks: Automatic Speech Recognition, Speaker Verification, Speaker Identification (mono/multi), Language Identification, Query By Example, and Keyword Spotting for 12 languages. (iii) On the released benchmarks, we train and evaluate different self-supervised models alongside a commonly used baseline FBANK. We show that language-specific fine-tuned models are more accurate than baseline on most of the tasks, including a large gap of 76\% for the Language Identification task. However, for speaker identification, self-supervised models trained on large datasets demonstrate an advantage. We hope IndicSUPERB contributes to the progress of developing speech language understanding models for Indian languages.
【7】 Low-Level Physiological Implications of End-to-End Learning of Speech Recognition
标题:语音识别端到端学习的低水平生理学意义
链接:https://arxiv.org/abs/2208.11700
* 与cs.SD语音【7】为同一篇
作者:Louise Coppieters de Gibson,Philip N. Garner机构:Idiap Research Institute, Martigny, Switzerland, Ecole polytechnique f´ed´erale de Lausanne (EPFL), Switzerland备注:Submitted to INTERSPEECH 2022摘要:从机器学习的观点来看,当前的语音识别体系结构表现得非常好,因此用户交互也是如此。这表明它们很好地模仿了人类的生物系统。我们研究的推论是否可以反过来提供洞察生物系统;特别是听询机制。使用SincNet,我们确认端到端系统确实学习了众所周知的滤波器组结构。然而,我们也表明,较宽的带宽滤波器在学习的结构中是重要的。虽然初始化窄带和宽带滤波器都可以获得一些好处,但生理限制表明,这种滤波器出现在中脑而不是耳蜗中。我们表明,标准的机器学习架构必须修改,以允许这一过程被神经地模拟。摘要:Current speech recognition architectures perform very well from the point of view of machine learning, hence user interaction. This suggests that they are emulating the human biological system well. We investigate whether the inference can be inverted to provide insights into that biological system; in particular the hearing mechanism. Using SincNet, we confirm that end-to-end systems do learn well known filterbank structures. However, we also show that wider band-width filters are important in the learned structure. Whilst some benefits can be gained by initialising both narrow and wide-band filters, physiological constraints suggest that such filters arise in mid-brain rather than the cochlea. We show that standard machine learning architectures must be modified to allow this process to be emulated neurally.
机器翻译,仅供参考