cs.SD语音,共计7篇,eess.AS音频处理,共计8篇


1.cs.SD语音:

【1】 The AI Mechanic: Acoustic Vehicle Characterization Neural Networks

标题:人工智能机械:声学车辆特征神经网络

链接:https://arxiv.org/abs/2205.09667

作者:Adam M. Terwilliger,Joshua E. Siegel
机构:Department of Computer Science and Engineering, Michigan State University, East Lansing, MI备注:34 pages, 12 figures, 28 tables
摘要:在一个日益依赖公路运输的世界里,了解车辆至关重要。我们介绍了AI Mechanical,一种声学车辆特征深度学习系统,作为一种综合方法,它使用从移动设备捕获的声音来增强非专家用户对车辆及其状况的透明度和理解。我们开发并实现了用于车辆理解的新型级联体系结构,我们将其定义为顺序、条件、多级网络,用于处理原始音频以提取高粒度的见解。为了展示级联结构的可行性,我们构建了一个多任务卷积神经网络,用于预测和级联车辆属性以增强故障检测。我们在一个合成数据集上对这些模型进行训练和测试,该数据集反映了40多小时的增强音频,并在属性(燃油类型、发动机配置、气缸数和吸气类型)上实现了超过92%的验证集准确性。我们的级联架构在失火故障预测方面还实现了93.6%的验证和86.8%的测试集准确性,与原始基线和平行基线相比,改进幅度分别为16.4%/7.8%和4.2%/1.5%。我们探索了以声学特征、数据增强、特征融合和数据可靠性为重点的实验研究。最后,我们讨论了这项工作的更广泛含义、未来方向和应用领域。
摘要:In a world increasingly dependent on road-based transportation, it is essential to understand vehicles. We introduce the AI mechanic, an acoustic vehicle characterization deep learning system, as an integrated approach using sound captured from mobile devices to enhance transparency and understanding of vehicles and their condition for non-expert users. We develop and implement novel cascading architectures for vehicle understanding, which we define as sequential, conditional, multi-level networks that process raw audio to extract highly-granular insights. To showcase the viability of cascading architectures, we build a multi-task convolutional neural network that predicts and cascades vehicle attributes to enhance fault detection. We train and test these models on a synthesized dataset reflecting more than 40 hours of augmented audio and achieve >92% validation set accuracy on attributes (fuel type, engine configuration, cylinder count and aspiration type). Our cascading architecture additionally achieved 93.6% validation and 86.8% test set accuracy on misfire fault prediction, demonstrating margins of 16.4% / 7.8% and 4.2% / 1.5% improvement over na\"ive and parallel baselines. We explore experimental studies focused on acoustic features, data augmentation, feature fusion, and data reliability. Finally, we conclude with a discussion of broader implications, future directions, and application areas for this work.


【2】 Automatic Spoken Language Identification using a Time-Delay Neural Network

标题:基于时延神经网络的口语自动识别

链接:https://arxiv.org/abs/2205.09564

作者:Benjamin Kepecs,Homayoon Beigi
机构:Recognition Technologies, Inc. and Columbia University
备注:6 pages, 6 figures, Technical Report Recognition Technologies, Inc
摘要:闭集口语识别的任务是从一组已知语言中识别录制的音频片段中所说的语言。在这项研究中,建立了一个语言识别系统,并对其进行了训练,以仅基于记录的语音来区分阿拉伯语、西班牙语、法语和土耳其语。使用预先存在的多语言数据集,在Tedlium TDNN模型的基础上训练一系列声学模型,以执行自动语音识别。该系统提供了一个定制的多语种语言模型和一个专门的语音词典,语音词典中的语言名称预先添加到手机上。训练后的模型用于生成电话比对,以测试所有四种语言的数据,并根据选择话语中最常见语言前置词的投票方案预测语言。准确度是通过将预测语言与已知语言进行比较来衡量的,确定西班牙语和阿拉伯语的识别率非常高,而土耳其语和法语的识别率稍低。
摘要:Closed-set spoken language identification is the task of recognizing the language being spoken in a recorded audio clip from a set of known languages. In this study, a language identification system was built and trained to distinguish between Arabic, Spanish, French, and Turkish based on nothing more than recorded speech. A pre-existing multilingual dataset was used to train a series of acoustic models based on the Tedlium TDNN model to perform automatic speech recognition. The system was provided with a custom multilingual language model and a specialized pronunciation lexicon with language names prepended to phones. The trained model was used to generate phone alignments to test data from all four languages, and languages were predicted based on a voting scheme choosing the most common language prepend in an utterance. Accuracy was measured by comparing predicted languages to known languages, and was determined to be very high in identifying Spanish and Arabic, and somewhat lower in identifying Turkish and French.


【3】 Insights on Neural Representations for End-to-End Speech Recognition

标题:端到端语音识别中神经表示的研究

链接:https://arxiv.org/abs/2205.09456

作者:Anna Ollerenshaw,Md Asif Jalal,Thomas Hain
机构:Speech and Hearing Research Group, University of Sheffield, UK
备注:None
摘要:端到端自动语音识别(ASR)模型旨在学习通用的语音表示。然而,只有有限的工具可用于理解模型体系结构中的内部功能和层次依赖关系的影响。理解分层表示之间的相关性,深入了解神经表示与性能之间的关系至关重要。之前使用相关分析技术对网络相似性进行的研究尚未针对端到端ASR模型进行探索。本文采用典型相关分析(CCA)和中心核对齐(CKA)方法,通过CNN、LSTM和基于Transformer的方法对训练过程中各层之间的内部动力学进行了分析和探索。研究发现,随着层深度的增加,CNN层内的神经表征表现出层次相关性,但这主要限于神经表征相关性更密切的情况。在LSTM体系结构中未观察到这种行为,但在整个训练过程中观察到自下而上的模式,而Transformer编码器层随着神经深度的增加表现出不规则的系数相关性。总之,这些结果为神经结构对语音识别性能的作用提供了新的见解。更具体地说,这些技术可以用作建立性能更好的语音识别模型的指标。
摘要:End-to-end automatic speech recognition (ASR) models aim to learn a generalised speech representation. However, there are limited tools available to understand the internal functions and the effect of hierarchical dependencies within the model architecture. It is crucial to understand the correlations between the layer-wise representations, to derive insights on the relationship between neural representations and performance. Previous investigations of network similarities using correlation analysis techniques have not been explored for End-to-End ASR models. This paper analyses and explores the internal dynamics between layers during training with CNN, LSTM and Transformer based approaches using Canonical correlation analysis (CCA) and centered kernel alignment (CKA) for the experiments. It was found that neural representations within CNN layers exhibit hierarchical correlation dependencies as layer depth increases but this is mostly limited to cases where neural representation correlates more closely. This behaviour is not observed in LSTM architecture, however there is a bottom-up pattern observed across the training process, while Transformer encoder layers exhibit irregular coefficiency correlation as neural depth increases. Altogether, these results provide new insights into the role that neural architectures have upon speech recognition performance. More specifically, these techniques can be used as indicators to build better performing speech recognition models.


【4】 MESH2IR: Neural Acoustic Impulse Response Generator for Complex 3D Scenes

标题:MESH2IR:面向复杂3D场景的神经声学脉冲响应产生器

链接:https://arxiv.org/abs/2205.09248

作者:Anton Ratnarajah,Zhenyu Tang,Rohith Chandrashekar Aralikatti,Dinesh Manocha
机构:University of Maryland, College Park, USA
备注:More results and source code is available at this https URL摘要:我们提出了一种基于网格的神经网络(MESH2IR)来为使用网格表示的室内3D场景生成声脉冲响应(IRs)。IRs用于在交互式应用程序和音频处理中创建高质量的声音体验。我们的方法可以处理任意拓扑(2K-3M三角形)的输入三角形网格。我们提出了一种新的训练技术来训练MESH2IR使用能量衰减救济,并强调其优点。我们还表明,使用我们提出的技术在IRs预处理上训练MESH2IR可以显著提高IR生成的准确性。我们通过使用图卷积网络将三维场景网格转换为潜在空间来减少网格空间中的非线性。我们的MESH2IR比CPU上的几何声学算法快200多倍,在NVIDIA GeForce RTX 2080 Ti GPU上,可以为给定的室内3D场景每秒生成10000多个IRs。声学度量用于描述声学环境。我们表明,从我们的MESH2IR预测的IRs的声学指标与地面真实值相匹配,误差小于10%。我们还强调了MESH2IR在语音去冗余和语音分离等音频和语音处理应用中的优势。据我们所知,我们是第一个基于神经网络的方法,可以从给定的3D场景网格实时预测IRs。
摘要:We propose a mesh-based neural network (MESH2IR) to generate acoustic impulse responses (IRs) for indoor 3D scenes represented using a mesh. The IRs are used to create a high-quality sound experience in interactive applications and audio processing. Our method can handle input triangular meshes with arbitrary topologies (2K - 3M triangles). We present a novel training technique to train MESH2IR using energy decay relief and highlight its benefits. We also show that training MESH2IR on IRs preprocessed using our proposed technique significantly improves the accuracy of IR generation. We reduce the non-linearity in the mesh space by transforming 3D scene meshes to latent space using a graph convolution network. Our MESH2IR is more than 200 times faster than a geometric acoustic algorithm on a CPU and can generate more than 10,000 IRs per second on an NVIDIA GeForce RTX 2080 Ti GPU for a given furnished indoor 3D scene. The acoustic metrics are used to characterize the acoustic environment. We show that the acoustic metrics of the IRs predicted from our MESH2IR match the ground truth with less than 10% error. We also highlight the benefits of MESH2IR on audio and speech processing applications such as speech dereverberation and speech separation. To the best of our knowledge, ours is the first neural-network-based approach to predict IRs from a given 3D scene mesh in real-time.


【5】 Neural network for multi-exponential sound energy decay analysis

标题:用于多指数声能衰减分析的神经网络

链接:https://arxiv.org/abs/2205.09644

作者:Georg Götz,Ricardo Falcón Pérez,Sebastian J. Schlecht,Ville Pulkki
机构:Aalto Acoustics Lab, Department of Signal Processing and Acoustics, Aalto University, P.O. Box, Aalto, Finland
备注:The following article has been submitted to the Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this http URL
摘要:声能衰减函数(EDF)的一个已建立模型是多个指数和一个噪声项的叠加。本文提出了一种基于神经网络的EDF模型参数估计方法。该网络接受合成EDF的训练,并在不同声学环境下进行的超过20000次EDF测量的两个大型数据集上进行评估。评估表明,所提出的神经网络结构能够从大量实测EDF数据集中稳健地估计模型参数,同时具有轻量级和计算效率高的特点。提出的神经网络的实现是公开的。
摘要:An established model for sound energy decay functions (EDFs) is the superposition of multiple exponentials and a noise term. This work proposes a neural-network-based approach for estimating the model parameters from EDFs. The network is trained on synthetic EDFs and evaluated on two large datasets of over 20000 EDF measurements conducted in various acoustic environments. The evaluation shows that the proposed neural network architecture robustly estimates the model parameters from large datasets of measured EDFs, while being lightweight and computationally efficient. An implementation of the proposed neural network is publicly available.


【6】 Bias Analysis of Spatial Coherence-Based RTF Vector Estimation for Acoustic Sensor Networks in a Diffuse Sound Field

标题:扩散声场中基于空间相干的声传感器网络RTF矢量估计的偏差分析

链接:https://arxiv.org/abs/2205.09401

作者:Wiebke Middelberg,Simon Doclo
机构:University of Oldenburg, Department of Medical Physics and Acoustics, and Cluster of Excellence Hearing,all, Oldenburg, Germany
摘要:在许多多麦克风算法中,需要估计所需扬声器的相对传递函数(RTF)。最近,提出了一种计算效率高的RTF矢量估计方法,该方法假设局部麦克风阵列和多个外部麦克风之间的噪声分量的空间相干性(SC)较低。该方法以优化输出信噪比(SNR)为目标,将多个RTF向量估计线性组合,其中复数权重使用广义特征值分解(GEVD)计算。在本文中,我们对基于SC的多外置话筒RTF矢量估计方法进行了理论偏差分析。假设噪声场的某个模型,我们推导了权重的解析表达式,表明基于模型的最优权重是实值的,并且仅取决于外部麦克风的输入信噪比。真实记录的仿真结果表明,基于GEVD的权重与基于模型的权重具有良好的一致性。然而,结果也表明,在实践中,会出现基于模型的权重无法解释的估计误差。
摘要:In many multi-microphone algorithms, an estimate of the relative transfer functions (RTFs) of the desired speaker is required. Recently, a computationally efficient RTF vector estimation method was proposed for acoustic sensor networks, assuming that the spatial coherence (SC) of the noise component between a local microphone array and multiple external microphones is low. Aiming at optimizing the output signal-to-noise ratio (SNR), this method linearly combines multiple RTF vector estimates, where the complex-valued weights are computed using a generalized eigenvalue decomposition (GEVD). In this paper, we perform a theoretical bias analysis for the SC-based RTF vector estimation method with multiple external microphones. Assuming a certain model for the noise field, we derive an analytical expression for the weights, showing that the optimal model-based weights are real-valued and only depend on the input SNR in the external microphones. Simulations with real-world recordings show a good accordance of the GEVD-based and the model-based weights. Nevertheless, the results also indicate that in practice, estimation errors occur which the model-based weights cannot account for.


【7】 Macedonian Speech Synthesis for Assistive Technology Applications

标题:用于辅助技术应用的马其顿语语音合成

链接:https://arxiv.org/abs/2205.09198

作者:Bojan Sofronievski,Elena Velovska,Martin Velichkovski,Violeta Argirova,Tea Veljkovikj,Risto Chavdarov,Stefan Janev,Kristijan Lazarev,Toni Bachvarovski,Zoran Ivanovski,Dimitar Tashkovski,Branislav Gerazov
机构:FEEIT, Ss Cyril and Methodius University in Skopje, RN Macedonia, Melon Inc., Skopje, RN Macedonia, Association for Assistive Technologies “Open the Windows”, Skopje, RN Macedonia
备注:5 pages, 2 figures, EUSIPCO conference 2022
摘要:随着语音设备和服务的发展,语音技术变得越来越普遍。在增强型和替代性通信工具中使用语音合成,有助于将有语音障碍的个人纳入其中,使他们能够使用语音与周围环境进行通信。尽管世界上大多数口语语言都有许多语音合成系统,但对较小语言的提供仍然有限。我们提出并比较了三个使用参数化和深度学习技术构建的模型,用于在新记录的语料库上训练马其顿语。我们的目标是为增强和替代通信和辅助技术(如通信板和屏幕阅读器)部署低资源优势。听力测试结果表明,与更先进的深度学习模型相比,参数化语音合成的性能相当。由于它还需要较少的资源,并提供全语音速率和基音控制,因此它是为该应用场景构建马其顿TTS系统的首选。
摘要:Speech technology is becoming ever more ubiquitous with the advance of speech enabled devices and services. The use of speech synthesis in Augmentative and Alternative Communication tools, has facilitated inclusion of individuals with speech impediments allowing them to communicate with their surroundings using speech. Although there are numerous speech synthesis systems for the most spoken world languages, there is still a limited offer for smaller languages. We propose and compare three models built using parametric and deep learning techniques for Macedonian trained on a newly recorded corpus. We target low-resource edge deployment for Augmentative and Alternative Communication and assistive technologies, such as communication boards and screen readers. The listening test results show that parametric speech synthesis is as performant compared to the more advanced deep learning models. Since it also requires less resources, and offers full speech rate and pitch control, it is the preferred choice for building a Macedonian TTS system for this application scenario.


2.eess.AS音频处理:

【1】 Bi-LSTM Scoring Based Similarity Measurement with Agglomerative Hierarchical Clustering (AHC) for Speaker Diarization

标题:基于BI-LSTM评分和凝聚层次聚类(AHC)的说话人二值化相似性度量

链接:https://arxiv.org/abs/2205.09709

作者:Siddharth S. Nijhawan,Homayoon Beigi
机构:Electrical Engineering Dept., Columbia University, Recognition Technologies, Inc. and Columbia University
备注:8 pages, 3 figures, 2 tables, 1 algorithm, Technical Report: Recognition Technologies, Inc
摘要:不同场景中的大多数语音信号都无法与仅包含一个扬声器的定义良好的音频段一起使用。两个说话人之间的典型对话由声音重叠、相互打断或在多个句子之间暂停的部分组成。最近,日记技术的发展利用基于神经网络的方法改进了说话人日记系统的多个子系统,包括提取分段嵌入特征和检测对话过程中说话人的变化。然而,为了通过聚类识别说话人,模型依赖于PLDA等方法在从给定会话音频中提取的两个片段之间生成相似性度量。由于这些算法忽略了会话的时间结构,它们往往会实现更高的日记错误率(DER),从而导致说话人和变化识别方面的误判。因此,为了独立地和顺序地比较两段语音的相似性,我们提出了一种双向长短时记忆网络来估计相似度矩阵中的元素。一旦相似矩阵生成,凝聚层次聚类(AHC)被应用于基于阈值的说话人分段进一步识别。为了评估性能,使用了二值化错误率(DER%)指标。与传统的基于PLDA的相似性度量机制相比,该模型在ICSI会议语料库的音频样本测试集上的DER较低,为34.80%,而基于PLDA的相似性度量机制的DER为39.90%。
摘要:Majority of speech signals across different scenarios are never available with well-defined audio segments containing only a single speaker. A typical conversation between two speakers consists of segments where their voices overlap, interrupt each other or halt their speech in between multiple sentences. Recent advancements in diarization technology leverage neural network-based approaches to improvise multiple subsystems of speaker diarization system comprising of extracting segment-wise embedding features and detecting changes in the speaker during conversation. However, to identify speaker through clustering, models depend on methodologies like PLDA to generate similarity measure between two extracted segments from a given conversational audio. Since these algorithms ignore the temporal structure of conversations, they tend to achieve a higher Diarization Error Rate (DER), thus leading to misdetections both in terms of speaker and change identification. Therefore, to compare similarity of two speech segments both independently and sequentially, we propose a Bi-directional Long Short-term Memory network for estimating the elements present in the similarity matrix. Once the similarity matrix is generated, Agglomerative Hierarchical Clustering (AHC) is applied to further identify speaker segments based on thresholding. To evaluate the performance, Diarization Error Rate (DER%) metric is used. The proposed model achieves a low DER of 34.80% on a test set of audio samples derived from ICSI Meeting Corpus as compared to traditional PLDA based similarity measurement mechanism which achieved a DER of 39.90%.


【2】 Neural network for multi-exponential sound energy decay analysis

标题:用于多指数声能衰减分析的神经网络

链接:https://arxiv.org/abs/2205.09644

作者:Georg Götz,Ricardo Falcón Pérez,Sebastian J. Schlecht,Ville Pulkki
机构:Aalto Acoustics Lab, Department of Signal Processing and Acoustics, Aalto University, P.O. Box, Aalto, Finland
备注:The following article has been submitted to the Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this http URL
摘要:声能衰减函数(EDF)的一个已建立模型是多个指数和一个噪声项的叠加。本文提出了一种基于神经网络的EDF模型参数估计方法。该网络接受合成EDF的训练,并在不同声学环境下进行的超过20000次EDF测量的两个大型数据集上进行评估。评估表明,所提出的神经网络结构能够从大量实测EDF数据集中稳健地估计模型参数,同时具有轻量级和计算效率高的特点。提出的神经网络的实现是公开的。
摘要:An established model for sound energy decay functions (EDFs) is the superposition of multiple exponentials and a noise term. This work proposes a neural-network-based approach for estimating the model parameters from EDFs. The network is trained on synthetic EDFs and evaluated on two large datasets of over 20000 EDF measurements conducted in various acoustic environments. The evaluation shows that the proposed neural network architecture robustly estimates the model parameters from large datasets of measured EDFs, while being lightweight and computationally efficient. An implementation of the proposed neural network is publicly available.


【3】 Bias Analysis of Spatial Coherence-Based RTF Vector Estimation for Acoustic Sensor Networks in a Diffuse Sound Field

标题:扩散声场中基于空间相干的声传感器网络RTF矢量估计的偏差分析

链接:https://arxiv.org/abs/2205.09401

作者:Wiebke Middelberg,Simon Doclo
机构:University of Oldenburg, Department of Medical Physics and Acoustics, and Cluster of Excellence Hearing,all, Oldenburg, Germany
摘要:在许多多麦克风算法中,需要估计所需扬声器的相对传递函数(RTF)。最近,提出了一种计算效率高的RTF矢量估计方法,该方法假设局部麦克风阵列和多个外部麦克风之间的噪声分量的空间相干性(SC)较低。该方法以优化输出信噪比(SNR)为目标,将多个RTF向量估计线性组合,其中复数权重使用广义特征值分解(GEVD)计算。在本文中,我们对基于SC的多外置话筒RTF矢量估计方法进行了理论偏差分析。假设噪声场的某个模型,我们推导了权重的解析表达式,表明基于模型的最优权重是实值的,并且仅取决于外部麦克风的输入信噪比。真实记录的仿真结果表明,基于GEVD的权重与基于模型的权重具有良好的一致性。然而,结果也表明,在实践中,会出现基于模型的权重无法解释的估计误差。
摘要:In many multi-microphone algorithms, an estimate of the relative transfer functions (RTFs) of the desired speaker is required. Recently, a computationally efficient RTF vector estimation method was proposed for acoustic sensor networks, assuming that the spatial coherence (SC) of the noise component between a local microphone array and multiple external microphones is low. Aiming at optimizing the output signal-to-noise ratio (SNR), this method linearly combines multiple RTF vector estimates, where the complex-valued weights are computed using a generalized eigenvalue decomposition (GEVD). In this paper, we perform a theoretical bias analysis for the SC-based RTF vector estimation method with multiple external microphones. Assuming a certain model for the noise field, we derive an analytical expression for the weights, showing that the optimal model-based weights are real-valued and only depend on the input SNR in the external microphones. Simulations with real-world recordings show a good accordance of the GEVD-based and the model-based weights. Nevertheless, the results also indicate that in practice, estimation errors occur which the model-based weights cannot account for.


【4】 Macedonian Speech Synthesis for Assistive Technology Applications

标题:用于辅助技术应用的马其顿语语音合成

链接:https://arxiv.org/abs/2205.09198

作者:Bojan Sofronievski,Elena Velovska,Martin Velichkovski,Violeta Argirova,Tea Veljkovikj,Risto Chavdarov,Stefan Janev,Kristijan Lazarev,Toni Bachvarovski,Zoran Ivanovski,Dimitar Tashkovski,Branislav Gerazov
机构:FEEIT, Ss Cyril and Methodius University in Skopje, RN Macedonia, Melon Inc., Skopje, RN Macedonia, Association for Assistive Technologies “Open the Windows”, Skopje, RN Macedonia
备注:5 pages, 2 figures, EUSIPCO conference 2022
摘要:随着语音设备和服务的发展,语音技术变得越来越普遍。在增强型和替代性通信工具中使用语音合成,有助于将有语音障碍的个人纳入其中,使他们能够使用语音与周围环境进行通信。尽管世界上大多数口语语言都有许多语音合成系统,但对较小语言的提供仍然有限。我们提出并比较了三个使用参数化和深度学习技术构建的模型,用于在新记录的语料库上训练马其顿语。我们的目标是为增强和替代通信和辅助技术(如通信板和屏幕阅读器)部署低资源优势。听力测试结果表明,与更先进的深度学习模型相比,参数化语音合成的性能相当。由于它还需要较少的资源,并提供全语音速率和基音控制,因此它是为该应用场景构建马其顿TTS系统的首选。
摘要:Speech technology is becoming ever more ubiquitous with the advance of speech enabled devices and services. The use of speech synthesis in Augmentative and Alternative Communication tools, has facilitated inclusion of individuals with speech impediments allowing them to communicate with their surroundings using speech. Although there are numerous speech synthesis systems for the most spoken world languages, there is still a limited offer for smaller languages. We propose and compare three models built using parametric and deep learning techniques for Macedonian trained on a newly recorded corpus. We target low-resource edge deployment for Augmentative and Alternative Communication and assistive technologies, such as communication boards and screen readers. The listening test results show that parametric speech synthesis is as performant compared to the more advanced deep learning models. Since it also requires less resources, and offers full speech rate and pitch control, it is the preferred choice for building a Macedonian TTS system for this application scenario.


【5】 The AI Mechanic: Acoustic Vehicle Characterization Neural Networks

标题:人工智能机械:声学车辆特征神经网络

链接:https://arxiv.org/abs/2205.09667

作者:Adam M. Terwilliger,Joshua E. Siegel
机构:Department of Computer Science and Engineering, Michigan State University, East Lansing, MI
备注:34 pages, 12 figures, 28 tables
摘要:在一个日益依赖公路运输的世界里,了解车辆至关重要。我们介绍了AI Mechanical,一种声学车辆特征深度学习系统,作为一种综合方法,它使用从移动设备捕获的声音来增强非专家用户对车辆及其状况的透明度和理解。我们开发并实现了用于车辆理解的新型级联体系结构,我们将其定义为顺序、条件、多级网络,用于处理原始音频以提取高粒度的见解。为了展示级联结构的可行性,我们构建了一个多任务卷积神经网络,用于预测和级联车辆属性以增强故障检测。我们在一个合成数据集上对这些模型进行训练和测试,该数据集反映了40多小时的增强音频,并在属性(燃油类型、发动机配置、气缸数和吸气类型)上实现了超过92%的验证集准确性。我们的级联架构在失火故障预测方面还实现了93.6%的验证和86.8%的测试集准确性,与原始基线和平行基线相比,改进幅度分别为16.4%/7.8%和4.2%/1.5%。我们探索了以声学特征、数据增强、特征融合和数据可靠性为重点的实验研究。最后,我们讨论了这项工作的更广泛含义、未来方向和应用领域。
摘要:In a world increasingly dependent on road-based transportation, it is essential to understand vehicles. We introduce the AI mechanic, an acoustic vehicle characterization deep learning system, as an integrated approach using sound captured from mobile devices to enhance transparency and understanding of vehicles and their condition for non-expert users. We develop and implement novel cascading architectures for vehicle understanding, which we define as sequential, conditional, multi-level networks that process raw audio to extract highly-granular insights. To showcase the viability of cascading architectures, we build a multi-task convolutional neural network that predicts and cascades vehicle attributes to enhance fault detection. We train and test these models on a synthesized dataset reflecting more than 40 hours of augmented audio and achieve >92% validation set accuracy on attributes (fuel type, engine configuration, cylinder count and aspiration type). Our cascading architecture additionally achieved 93.6% validation and 86.8% test set accuracy on misfire fault prediction, demonstrating margins of 16.4% / 7.8% and 4.2% / 1.5% improvement over na\"ive and parallel baselines. We explore experimental studies focused on acoustic features, data augmentation, feature fusion, and data reliability. Finally, we conclude with a discussion of broader implications, future directions, and application areas for this work.


【6】 Automatic Spoken Language Identification using a Time-Delay Neural Network

标题:基于时延神经网络的口语自动识别

链接:https://arxiv.org/abs/2205.09564

作者:Benjamin Kepecs,Homayoon Beigi
机构:Recognition Technologies, Inc. and Columbia University
备注:6 pages, 6 figures, Technical Report Recognition Technologies, Inc
摘要:闭集口语识别的任务是从一组已知语言中识别录制的音频片段中所说的语言。在这项研究中,建立了一个语言识别系统,并对其进行了训练,以仅基于记录的语音来区分阿拉伯语、西班牙语、法语和土耳其语。使用预先存在的多语言数据集,在Tedlium TDNN模型的基础上训练一系列声学模型,以执行自动语音识别。该系统提供了一个定制的多语种语言模型和一个专门的语音词典,语音词典中的语言名称预先添加到手机上。训练后的模型用于生成电话比对,以测试所有四种语言的数据,并根据选择话语中最常见语言前置词的投票方案预测语言。准确度是通过将预测语言与已知语言进行比较来衡量的,确定西班牙语和阿拉伯语的识别率非常高,而土耳其语和法语的识别率稍低。
摘要:Closed-set spoken language identification is the task of recognizing the language being spoken in a recorded audio clip from a set of known languages. In this study, a language identification system was built and trained to distinguish between Arabic, Spanish, French, and Turkish based on nothing more than recorded speech. A pre-existing multilingual dataset was used to train a series of acoustic models based on the Tedlium TDNN model to perform automatic speech recognition. The system was provided with a custom multilingual language model and a specialized pronunciation lexicon with language names prepended to phones. The trained model was used to generate phone alignments to test data from all four languages, and languages were predicted based on a voting scheme choosing the most common language prepend in an utterance. Accuracy was measured by comparing predicted languages to known languages, and was determined to be very high in identifying Spanish and Arabic, and somewhat lower in identifying Turkish and French.


【7】 Insights on Neural Representations for End-to-End Speech Recognition

标题:端到端语音识别中神经表示的研究

链接:https://arxiv.org/abs/2205.09456

作者:Anna Ollerenshaw,Md Asif Jalal,Thomas Hain
机构:Speech and Hearing Research Group, University of Sheffield, UK
备注:None
摘要:端到端自动语音识别(ASR)模型旨在学习通用的语音表示。然而,只有有限的工具可用于理解模型体系结构中的内部功能和层次依赖关系的影响。理解分层表示之间的相关性,深入了解神经表示与性能之间的关系至关重要。之前使用相关分析技术对网络相似性进行的研究尚未针对端到端ASR模型进行探索。本文采用典型相关分析(CCA)和中心核对齐(CKA)方法,通过CNN、LSTM和基于Transformer的方法对训练过程中各层之间的内部动力学进行了分析和探索。研究发现,随着层深度的增加,CNN层内的神经表征表现出层次相关性,但这主要限于神经表征相关性更密切的情况。在LSTM体系结构中未观察到这种行为,但在整个训练过程中观察到自下而上的模式,而Transformer编码器层随着神经深度的增加表现出不规则的系数相关性。总之,这些结果为神经结构对语音识别性能的作用提供了新的见解。更具体地说,这些技术可以用作建立性能更好的语音识别模型的指标。
摘要:End-to-end automatic speech recognition (ASR) models aim to learn a generalised speech representation. However, there are limited tools available to understand the internal functions and the effect of hierarchical dependencies within the model architecture. It is crucial to understand the correlations between the layer-wise representations, to derive insights on the relationship between neural representations and performance. Previous investigations of network similarities using correlation analysis techniques have not been explored for End-to-End ASR models. This paper analyses and explores the internal dynamics between layers during training with CNN, LSTM and Transformer based approaches using Canonical correlation analysis (CCA) and centered kernel alignment (CKA) for the experiments. It was found that neural representations within CNN layers exhibit hierarchical correlation dependencies as layer depth increases but this is mostly limited to cases where neural representation correlates more closely. This behaviour is not observed in LSTM architecture, however there is a bottom-up pattern observed across the training process, while Transformer encoder layers exhibit irregular coefficiency correlation as neural depth increases. Altogether, these results provide new insights into the role that neural architectures have upon speech recognition performance. More specifically, these techniques can be used as indicators to build better performing speech recognition models.


【8】 MESH2IR: Neural Acoustic Impulse Response Generator for Complex 3D Scenes

标题:MESH2IR:面向复杂3D场景的神经声学脉冲响应产生器

链接:https://arxiv.org/abs/2205.09248

作者:Anton Ratnarajah,Zhenyu Tang,Rohith Chandrashekar Aralikatti,Dinesh Manocha
机构:University of Maryland, College Park, USA
备注:More results and source code is available at this https URL
摘要:我们提出了一种基于网格的神经网络(MESH2IR)来为使用网格表示的室内3D场景生成声脉冲响应(IRs)。IRs用于在交互式应用程序和音频处理中创建高质量的声音体验。我们的方法可以处理任意拓扑(2K-3M三角形)的输入三角形网格。我们提出了一种新的训练技术来训练MESH2IR使用能量衰减救济,并强调其优点。我们还表明,使用我们提出的技术在IRs预处理上训练MESH2IR可以显著提高IR生成的准确性。我们通过使用图卷积网络将三维场景网格转换为潜在空间来减少网格空间中的非线性。我们的MESH2IR比CPU上的几何声学算法快200多倍,在NVIDIA GeForce RTX 2080 Ti GPU上,可以为给定的室内3D场景每秒生成10000多个IRs。声学度量用于描述声学环境。我们表明,从我们的MESH2IR预测的IRs的声学指标与地面真实值相匹配,误差小于10%。我们还强调了MESH2IR在语音去冗余和语音分离等音频和语音处理应用中的优势。据我们所知,我们是第一个基于神经网络的方法,可以从给定的3D场景网格实时预测IRs。
摘要:We propose a mesh-based neural network (MESH2IR) to generate acoustic impulse responses (IRs) for indoor 3D scenes represented using a mesh. The IRs are used to create a high-quality sound experience in interactive applications and audio processing. Our method can handle input triangular meshes with arbitrary topologies (2K - 3M triangles). We present a novel training technique to train MESH2IR using energy decay relief and highlight its benefits. We also show that training MESH2IR on IRs preprocessed using our proposed technique significantly improves the accuracy of IR generation. We reduce the non-linearity in the mesh space by transforming 3D scene meshes to latent space using a graph convolution network. Our MESH2IR is more than 200 times faster than a geometric acoustic algorithm on a CPU and can generate more than 10,000 IRs per second on an NVIDIA GeForce RTX 2080 Ti GPU for a given furnished indoor 3D scene. The acoustic metrics are used to characterize the acoustic environment. We show that the acoustic metrics of the IRs predicted from our MESH2IR match the ground truth with less than 10% error. We also highlight the benefits of MESH2IR on audio and speech processing applications such as speech dereverberation and speech separation. To the best of our knowledge, ours is the first neural-network-based approach to predict IRs from a given 3D scene mesh in real-time.


机器翻译,仅供参考