今日论文合集:cs.SD语音11篇,eess.AS音频处理14篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Multilingual Turn-taking Prediction Using Voice Activity Projection
标题:基于语音活动投影的多语种话轮转换预测
链接:https://arxiv.org/abs/2403.06487
作者:Koji Inoue,Bing'er Jiang,Erik Ekstedt,Tatsuya Kawahara,Gabriel Skantze
备注:This paper has been accepted for presentation at The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) and represents the author's version of the work
摘要:本文研究了语音活动投影(VAP),一个预测轮对模型的口语对话,对多语种数据,包括英语,汉语和日语的应用。VAP模型持续预测二元对话中参与者即将到来的语音活动,利用交叉注意Transformer来捕获参与者之间的动态交互。结果表明,在一种语言上训练的单语VAP模型在应用于其他语言时不能做出很好的预测。然而,在所有三种语言上训练的多语言模型显示出与所有语言的单语言模型相当的预测性能。进一步的分析表明,多语言模型已经学会识别输入信号的语言。我们还分析了对音高的敏感性,这是一种被认为对话轮转换很重要的韵律线索。最后,我们比较了两种不同的音频编码器,即在英语上预训练的对比预测编码(CPC)和基于多语言wav2vec 2.0(MMS)的最新模型。
摘要:This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS).


【2】 Cosine Scoring with Uncertainty for Neural Speaker Embedding
标题:神经说话人嵌入的不确定余弦评分
链接:https://arxiv.org/abs/2403.06404
作者:Qiongqiong Wang,Kong Aik Lee
备注:None
摘要:说话人表示中的不确定性建模旨在学习语音话语中存在的可变性。虽然传统的余弦评分是计算效率和普遍的说话人识别,它缺乏处理不确定性的能力。为了解决这一问题,本文提出了一种方法,用于估计不确定性的扬声器嵌入前端和传播到余弦评分后端。在VoxCeleb和SITW数据集上进行的实验证实了该方法在处理嵌入估计引起的不确定性方面的有效性。与传统的余弦相似性相比,它实现了EER和minDCF平均降低8.5%和9.8%的改善。它在实践中也是计算效率高的。
摘要:Uncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to handle uncertainty. To address this challenge, this paper proposes an approach for estimating uncertainty at the speaker embedding front-end and propagating it to the cosine scoring back-end. Experiments conducted on the VoxCeleb and SITW datasets confirmed the efficacy of the proposed method in handling uncertainty arising from embedding estimation. It achieved improvement with 8.5% and 9.8% average reductions in EER and minDCF compared to the conventional cosine similarity. It is also computationally efficient in practice.


【3】 Towards Decoupling Frontend Enhancement and Backend Recognition in  Monaural Robust ASR标题:单声道鲁棒ASR的前端增强与后端识别解耦研究
链接:https://arxiv.org/abs/2403.06387
作者:Yufeng Yang,Ashutosh Pandey,DeLiang Wang
备注:Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2210.13318
摘要:研究表明,语音增强(SE)算法可以提高含噪语音的可懂度。然而,单声道SE还没有被建立为一个有效的前端自动语音识别(ASR)在嘈杂的条件下相比,一个ASR模型直接训练嘈杂的语音。SE和ASR之间的鸿沟阻碍了强大的ASR系统的发展,特别是SE近年来取得了重大进展。本文的重点是消除这种鸿沟与ARN(注意递归网络)时域和CrossNet时频域增强模型。所提出的系统完全解耦前端增强和后端ASR只在干净的语音上训练。WSJ、CHiME-2、LibriSpeech和CHiME-4语料库上的结果表明,ARN和CrossNet增强语音在嘈杂和混响环境中都能改善ASR结果,并很好地推广到真实的声学场景。所提出的系统优于直接在损坏的语音上训练的基线。此外,该算法将CHiME-2上的最佳字错误率(WER)降低了28.4美元,WER为5.57美元,在没有CHiME-4训练的情况下,单通道CHiME-4模拟/真实测试数据的WER为3.32/4.44美元。
摘要:It has been shown that the intelligibility of noisy speech can be improved by speech enhancement (SE) algorithms. However, monaural SE has not been established as an effective frontend for automatic speech recognition (ASR) in noisy conditions compared to an ASR model trained on noisy speech directly. The divide between SE and ASR impedes the progress of robust ASR systems, especially as SE has made major advances in recent years. This paper focuses on eliminating this divide with an ARN (attentive recurrent network) time-domain and a CrossNet time-frequency domain enhancement models. The proposed systems fully decouple frontend enhancement and backend ASR trained only on clean speech. Results on the WSJ, CHiME-2, LibriSpeech, and CHiME-4 corpora demonstrate that ARN and CrossNet enhanced speech both translate to improved ASR results in noisy and reverberant environments, and generalize well to real acoustic scenarios. The proposed system outperforms the baselines trained on corrupted speech directly. Furthermore, it cuts the previous best word error rate (WER) on CHiME-2 by $28.4\%$ relatively with a $5.57\%$ WER, and achieves $3.32/4.44\%$ WER on single-channel CHiME-4 simulated/real test data without training on CHiME-4.

【4】 SCORE: Self-supervised Correspondence Fine-tuning for Improved Content  Representations
标题:得分:自我监督通信微调以改进内容表示
链接:https://arxiv.org/abs/2403.06260
作者:Amit Meghanani,Thomas Hain
备注:Accepted at ICASSP 2024
摘要:有越来越多的兴趣在成本效益的自我监督微调(SSFT)的自我监督学习(SSL)为基础的语音模型,以获得特定于任务的表示。这些特定于任务的表示通过对标记数据进行微调来用于各种下游任务的鲁棒性能。这项工作提出了一种具有成本效益的SSFT方法命名为自监督对应(SCORE)微调,以适应SSL语音表示内容相关的任务。该方法使用对应训练策略,旨在从扰动语音和原始语音中学习相似的表示。通常使用的内容相关的任务(ASR)的数据增强技术,以获得扰动的语音。SCORE微调的HuBERT在SUPERB基准测试中的表现优于普通的HuBERT,在单个GPU上进行自动语音识别,音素识别和按示例查询任务的微调仅需几个小时(< 5小时),相对提高分别为1.09%,3.58%和12.65%。SCORE提供了与最近提出的SSFT方法SPIN竞争的结果,与SPIN相比,仅使用1/3的处理语音。
摘要:There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used for robust performance on various downstream tasks by fine-tuning on the labelled data. This work presents a cost-effective SSFT method named Self-supervised Correspondence (SCORE) fine-tuning to adapt the SSL speech representations for content-related tasks. The proposed method uses a correspondence training strategy, aiming to learn similar representations from perturbed speech and original speech. Commonly used data augmentation techniques for content-related tasks (ASR) are applied to obtain perturbed speech. SCORE fine-tuned HuBERT outperforms the vanilla HuBERT on SUPERB benchmark with only a few hours of fine-tuning (< 5 hrs) on a single GPU for automatic speech recognition, phoneme recognition, and query-by-example tasks, with relative improvements of 1.09%, 3.58%, and 12.65%, respectively. SCORE provides competitive results with the recently proposed SSFT method SPIN, using only 1/3 of the processed speech compared to SPIN.

【5】 HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot  Text-to-Speech with Model and Data Scaling
标题:HAM-TTS:基于模型和数据缩放的基于令牌的零发声文语合成的分层声学建模
链接:https://arxiv.org/abs/2403.05989
作者:Chunhui Wang,Chang Zeng,Bowen Zhang,Ziyang Ma,Yefan Zhu,Zifeng Cai,Jian Zhao,Zhonglin Jiang,Yong Chen
摘要:基于令牌的文本到语音(TTS)模型已经成为生成自然和逼真语音的一种有前途的途径,但它们面临着发音准确性低、说话风格和音色不一致以及对多样化训练数据的大量需求。作为回应,我们引入了一种新的分层声学建模方法,辅以量身定制的数据增强策略,并结合真实数据和合成数据对其进行训练,将数据大小扩展到65万小时,从而获得具有0.8B参数的zero-shot TTS模型。具体来说,我们的方法采用了一个潜在的变量序列包含补充声学信息的基础上,完善的自监督学习(SSL)离散单元到TTS模型的预测。这显著地减轻了合成语音中的发音错误和风格突变。在训练过程中,我们策略性地替换和复制数据片段,以提高音色的一致性。此外,利用预先训练的Few-Shot语音转换模型来生成具有相同内容但不同音色的过多语音。这促进了话语级一对多映射的显式学习,丰富了语音多样性,并确保了音色的一致性。对比实验(演示页面:https://anonymous.4open.science/w/ham-tts/)证明了我们的模型在发音精度和保持说话风格以及音色连续性方面优于VALL-E。
摘要:Token-based text-to-speech (TTS) models have emerged as a promising avenue for generating natural and realistic speech, yet they grapple with low pronunciation accuracy, speaking style and timbre inconsistency, and a substantial need for diverse training data. In response, we introduce a novel hierarchical acoustic modeling approach complemented by a tailored data augmentation strategy and train it on the combination of real and synthetic data, scaling the data size up to 650k hours, leading to the zero-shot TTS model with 0.8B parameters. Specifically, our method incorporates a latent variable sequence containing supplementary acoustic information based on refined self-supervised learning (SSL) discrete units into the TTS model by a predictor. This significantly mitigates pronunciation errors and style mutations in synthesized speech. During training, we strategically replace and duplicate segments of the data to enhance timbre uniformity. Moreover, a pretrained few-shot voice conversion model is utilized to generate a plethora of voices with identical content yet varied timbres. This facilitates the explicit learning of utterance-level one-to-many mappings, enriching speech diversity and also ensuring consistency in timbre. Comparative experiments (Demo page: https://anonymous.4open.science/w/ham-tts/)demonstrate our model's superiority over VALL-E in pronunciation precision and maintaining speaking style, as well as timbre continuity.

【6】 Enhancing Expressiveness in Dance Generation via Integrating Frequency  and Music Style Information
标题:通过整合频率和音乐风格信息增强舞蹈生成的表现力
链接:https://arxiv.org/abs/2403.05834
作者:Qiaochu Huang,Xu He,Boshi Tang,Haolin Zhuang,Liyang Chen,Shuochen Gao,Zhiyong Wu,Haozhi Huang,Helen Meng
摘要:舞蹈生成作为人体运动生成的一个分支,越来越受到人们的关注。近年来,一些作品试图从不同的角度来增强舞蹈的表现力,包括体裁匹配、节拍对齐和舞蹈动态。但是,由于缺乏对上述三个因素的综合考虑,这种改进是相当有限的。在本文中,我们提出了ExpressiveBailando,一种新颖的舞蹈生成方法,旨在生成富有表现力的舞蹈,同时考虑到所有三个因素。具体来说,我们通过将频率信息纳入VQ-VAE来缓解速度均匀化的问题,从而改善舞蹈动态。此外,我们通过预训练的音乐模型提取流派和节拍相关的特征来整合音乐风格信息,从而实现其他两个因素的改进。大量的实验结果表明,我们提出的方法可以生成具有高表现力的舞蹈,并优于现有的方法在定性和定量。
摘要:Dance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre matching, beat alignment, and dance dynamics, from certain aspects. However, the enhancement is quite limited as they lack comprehensive consideration of the aforementioned three factors. In this paper, we propose ExpressiveBailando, a novel dance generation method designed to generate expressive dances, concurrently taking all three factors into account. Specifically, we mitigate the issue of speed homogenization by incorporating frequency information into VQ-VAE, thus improving dance dynamics. Additionally, we integrate music style information by extracting genre- and beat-related features with a pre-trained music model, hence achieving improvements in the other two factors. Extensive experimental results demonstrate that our proposed method can generate dances with high expressiveness and outperforms existing methods both qualitatively and quantitatively.

【7】 An Audio-textual Diffusion Model For Converting Speech Signals Into  Ultrasound Tongue Imaging Data
标题:一种将语音信号转换为超声舌象数据的音频-文本扩散模型
链接:https://arxiv.org/abs/2403.05820
作者:Yudong Yang,Rongfeng Su,Xiaokang Liu,Nan Yan,Lan Wang
备注:ICASSP2024 Accept
摘要:声学发音反转(AAI)是将音频转换为发音器运动,例如超声舌成像(UTI)数据。现有的AAI方法的一个问题是仅使用个性化的声学信息来导出舌头运动的一般模式,因此生成的UTI数据的质量是有限的。为了解决这个问题,本文提出了一个音频文本扩散模型的UTI数据生成任务。在该模型中,与舌头运动细节相关的个体固有声学特征使用wav 2 vec 2.0编码,而与舌头运动的普遍性相关的ASR transmittance使用BERT编码。然后通过使用扩散模块生成UTI数据。实验结果表明,该扩散模型可以生成高质量的UTI数据,具有清晰的舌头轮廓,这是至关重要的语言分析和临床评估。该项目可在网站\footnote{https:yangyudong2020.github.io/wav2uti/
摘要:Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the general patterns of tongue motions, and thus the quality of generated UTI data is limited. To address this issue, this paper proposes an audio-textual diffusion model for the UTI data generation task. In this model, the inherent acoustic characteristics of individuals related to the tongue motion details are encoded by using wav2vec 2.0, while the ASR transcriptions related to the universality of tongue motions are encoded by using BERT. UTI data are then generated by using a diffusion module. Experimental results showed that the proposed diffusion model could generate high-quality UTI data with clear tongue contour that is crucial for the linguistic analysis and clinical assessment. The project can be found on the website\footnote{https://yangyudong2020.github.io/wav2uti/

【8】 sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection  with Spiking Neural Networks
标题:SVAD:一种基于尖峰神经网络的健壮、低功耗、轻量级语音活动检测
链接:https://arxiv.org/abs/2403.05772
作者:Qu Yang,Qianhui Liu,Nan Li,Meng Ge,Zeyang Song,Haizhou Li
备注:Accepted by ICASSP 2024
摘要:语音应用被期望在噪声条件下是低功耗和鲁棒的。有效的语音活动检测(VAD)前端降低了计算需求。尖峰神经网络(SNN)被认为是生物学上合理的和节能的。然而,基于SNN的VAD尚未实现噪声鲁棒性,并且通常需要大型模型来实现高性能。本文介绍了一种新的SNN为基础的VAD模型,简称为sVAD,它具有一个听觉编码器与SNN为基础的注意力机制。特别地,它通过SincNet和1D卷积提供有效的听觉特征表示,并通过注意机制提高噪声鲁棒性。该分类器利用尖峰递归神经网络(sRNN)来利用时间语音信息。实验结果表明,我们的sVAD实现了显着的噪声鲁棒性,同时保持低功耗和小的足迹,使其成为一个有前途的解决方案,为现实世界的VAD应用。
摘要:Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achieve noise robustness and often require large models for high performance. This paper introduces a novel SNN-based VAD model, referred to as sVAD, which features an auditory encoder with an SNN-based attention mechanism. Particularly, it provides effective auditory feature representation through SincNet and 1D convolution, and improves noise robustness with attention mechanisms. The classifier utilizes Spiking Recurrent Neural Networks (sRNN) to exploit temporal speech information. Experimental results demonstrate that our sVAD achieves remarkable noise robustness and meanwhile maintains low power consumption and a small footprint, making it a promising solution for real-world VAD applications.

【9】 A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
标题:一种基于LLM增强识别的无声语音跨模式识别方法
链接:https://arxiv.org/abs/2403.05583
作者:Tyler Benster,Guy Wilson,Reshef Elisha,Francis R Willett,Shaul Druckmann
摘要:无声语音接口(SSI)为无声的口头交流提供了脑机接口的非侵入性替代方案。我们介绍多模态口面神经音频(MONA),一个系统,利用跨模态对齐通过新的损失函数-交叉对比度(crossCon)和监督时间对比度(supTcon)-训练一个多模态模型与共享的潜在表示。这种架构允许使用像LibriSpeech这样的纯音频数据集来改进无声语音识别。此外,我们引入了大型语言模型(LLM)综合评分调整(LISA),显着提高了识别精度。总之,MONA LISA将Gaddy(2020)开放词汇表上无声语音基准数据集中最先进的单词错误率(WER)从28.8%降低到12.2%。对于声音EMG记录,我们的方法将最先进的WER从23.3%提高到3.7%。在Brain-to-Text 2024竞赛中,LISA表现最好,将顶级WER从9.8%提高到8.9%。据我们所知,这项工作代表了第一个实例,在开放词汇表上的非侵入性无声语音识别已经清除了15%的WER阈值,表明SSI可以成为自动语音识别(ASR)的可行替代方案。我们的工作不仅缩小了无声语音和有声语音之间的性能差距,而且还为人机交互开辟了新的可能性,展示了在嘈杂和数据有限的情况下跨模态方法的潜力。
摘要:Silent Speech Interfaces (SSIs) offer a noninvasive alternative to brain-computer interfaces for soundless verbal communication. We introduce Multimodal Orofacial Neural Audio (MONA), a system that leverages cross-modal alignment through novel loss functions--cross-contrast (crossCon) and supervised temporal contrast (supTcon)--to train a multimodal model with a shared latent representation. This architecture enables the use of audio-only datasets like LibriSpeech to improve silent speech recognition. Additionally, our introduction of Large Language Model (LLM) Integrated Scoring Adjustment (LISA) significantly improves recognition accuracy. Together, MONA LISA reduces the state-of-the-art word error rate (WER) from 28.8% to 12.2% in the Gaddy (2020) benchmark dataset for silent speech on an open vocabulary. For vocal EMG recordings, our method improves the state-of-the-art from 23.3% to 3.7% WER. In the Brain-to-Text 2024 competition, LISA performs best, improving the top WER from 9.8% to 8.9%. To the best of our knowledge, this work represents the first instance where noninvasive silent speech recognition on an open vocabulary has cleared the threshold of 15% WER, demonstrating that SSIs can be a viable alternative to automatic speech recognition (ASR). Our work not only narrows the performance gap between silent and vocalized speech but also opens new possibilities in human-computer interaction, demonstrating the potential of cross-modal approaches in noisy and data-limited regimes.


【10】 SonoTraceLab - A Raytracing-Based Acoustic Modelling System for  Simulating Echolocation Behavior of Bats
标题:SonoTraceLab--基于光线跟踪的蝙蝠回声定位行为声学模拟系统
链接:https://arxiv.org/abs/2403.06847
作者:Wouter Jansen,Jan Steckel
摘要:回声定位是许多蝙蝠物种的主要感知方式,它们在复杂和非结构化的环境中表现出复杂的能力来执行大量任务。了解这种特殊的感觉运动相互作用是构建更强大和性能更高的人造声纳传感器的关键方面。为了更好地理解潜在的感知机制,重要的是要深入了解蝙蝠感知到的反射信号的性质。虽然声穿透实验是更好地理解这些信号的性质的重要方式,但它们既耗时又信息量大。在本文中,我们提出了SonoTraceLab,一个开源软件包,用于模拟技术以及生物声纳系统在复杂的场景。使用模拟方法可以大大增加对生物回声定位系统性质的了解,同时减少执行它们的时间和材料复杂性。
摘要:Echolocation is the prime sensing modality for many species of bats, who show the intricate ability to perform a plethora of tasks in complex and unstructured environments. Understanding this exceptional feat of sensorimotor interaction is a key aspect into building more robust and performant man-made sonar sensors. In order to better understand the underlying perception mechanisms it is important to get a good insight into the nature of the reflected signals that the bat perceives. While ensonification experiments are in important way to better understand the nature of these signals, they are as time-consuming to perform as they are informative. In this paper we present SonoTraceLab, an open-source software package for simulating both technical as well as biological sonar systems in complex scenes. Using simulation approaches can drastically increase insights into the nature of biological echolocation systems, while reducing the time- and material complexity of performing them.

【11】 Asynchronous Microphone Array Calibration using Hybrid TDOA Information
标题:利用混合时差信息的异步麦克风阵列校准
链接:https://arxiv.org/abs/2403.05791
作者:Chengjie Zhang,Jiang Wang,You-Fu Li,He Kong
摘要:异步麦克风阵列校准是大多数试听机器人应用的先决条件。在实践中,校准需要同时估计麦克风位置、时间偏移、时钟漂移率和声音事件位置。现有的方法提出了基于图的同时定位和映射(Graph-SLAM)利用共同的TDOA,两个麦克风之间的到达时间差(TDOA-M)和里程测量,但是,它严重依赖于初始值。本文提出了一种新的TDOA--相邻声事件到达时间差(TDOA-S),将其与TDOA-M相结合,称为混合TDOA,并加入测距测量,构建Graph-SLAM,采用高斯-牛顿(GN)法求解。TDOA-S简单有效,因为它消除了时间偏移而不产生新的变量。仿真和实际实验结果表明,该方法不受麦克风数目的影响,对初始值不敏感,在各种TDOA噪声下具有较好的校准精度和稳定性。仿真结果表明,该方法对麦克风参数具有较低的Cram\'er-Rao下界(CRLB),说明了该方法的优越性.摘要:Asynchronous Microphone array calibration is a prerequisite for most audition robot applications. In practice, the calibration requires estimating microphone positions, time offsets, clock drift rates, and sound event locations simultaneously. The existing method proposed Graph-based Simultaneous Localisation and Mapping (Graph-SLAM) utilizing common TDOA, time difference of arrival between two microphones (TDOA-M), and odometry measurement, however, it heavily depends on the initial value. In this paper, we propose a novel TDOA, time difference of arrival between adjacent sound events (TDOA-S), combine it with TDOA-M, called hybrid TDOA, and add odometry measurement to construct Graph-SLAM and use the Gauss-Newton (GN) method to solve. TDOA-S is simple and efficient because it eliminates time offset without generating new variables. Simulation and real-world experiment results consistently show that our method is independent of microphone number, insensitive to initial values, and has better calibration accuracy and stability under various TDOA noises. In addition, the simulation result demonstrates that our method has a lower Cram\'er-Rao lower bound (CRLB) for microphone parameters, which explains the advantages of my method.


eess.AS音频处理
【1】 Concurrent Speaker Detection: A multi-microphone Transformer-Based  Approach
标题:并发说话人检测:一种基于多麦克风Transformer的方法
链接:https://arxiv.org/abs/2403.06856
作者:Amit Eliav,Sharon Gannot
备注:5 pages, 6 tables, 2 figures
摘要:我们提出了一种深度学习方法,使用修改后的Transformer模型的并发说话人检测(CSD)的任务。我们的模型被设计为处理多麦克风数据,但也可以在单麦克风的情况下工作。该方法可以将音频片段分类为三个类别之一:1)没有语音活动(仅噪声),2)仅单个说话者是活动的,以及3)多于一个说话者是活动的。我们将成本敏感(CS)的损失和信心校准的培训过程。该方法使用三个真实世界的数据库进行评估:AMI,AliMeeting和CHiME 5,证明了对现有方法的改进。
摘要:We present a deep-learning approach for the task of Concurrent Speaker Detection (CSD) using a modified transformer model. Our model is designed to handle multi- microphone data but can also work in the single-microphone case. The method can classify audio segments into one of three classes: 1) no speech activity (noise only), 2) only a single speaker is active, and 3) more than one speaker is active. We incorporate a Cost-Sensitive (CS) loss and a confidence calibration to the training procedure. The approach is evaluated using three real- world databases: AMI, AliMeeting, and CHiME 5, demonstrating an improvement over existing approaches.


【2】 SonoTraceLab - A Raytracing-Based Acoustic Modelling System for  Simulating Echolocation Behavior of Bats
标题:SonoTraceLab--基于光线跟踪的蝙蝠回声定位行为声学模拟系统
链接:https://arxiv.org/abs/2403.06847
作者:Wouter Jansen,Jan Steckel
摘要:回声定位是许多蝙蝠物种的主要感知方式,它们在复杂和非结构化的环境中表现出复杂的能力来执行大量任务。了解这种特殊的感觉运动相互作用是构建更强大和性能更高的人造声纳传感器的关键方面。为了更好地理解潜在的感知机制,重要的是要深入了解蝙蝠感知到的反射信号的性质。虽然声穿透实验是更好地理解这些信号的性质的重要方式,但它们既耗时又信息量大。在本文中,我们提出了SonoTraceLab,一个开源软件包,用于模拟技术以及生物声纳系统在复杂的场景。使用模拟方法可以大大增加对生物回声定位系统性质的了解,同时减少执行它们的时间和材料复杂性。
摘要:Echolocation is the prime sensing modality for many species of bats, who show the intricate ability to perform a plethora of tasks in complex and unstructured environments. Understanding this exceptional feat of sensorimotor interaction is a key aspect into building more robust and performant man-made sonar sensors. In order to better understand the underlying perception mechanisms it is important to get a good insight into the nature of the reflected signals that the bat perceives. While ensonification experiments are in important way to better understand the nature of these signals, they are as time-consuming to perform as they are informative. In this paper we present SonoTraceLab, an open-source software package for simulating both technical as well as biological sonar systems in complex scenes. Using simulation approaches can drastically increase insights into the nature of biological echolocation systems, while reducing the time- and material complexity of performing them.

【3】 Aligning Speech to Languages to Enhance Code-switching Speech  Recognition
标题:将语音与语言对齐以增强代码切换语音识别
链接:https://arxiv.org/abs/2403.05887
作者:Hexin Liu,Xiangyu Zhang,Leibny Paola Garcia,Andy W. H. Khong,Eng Siong Chng,Shinji Watanabe
备注:Manuscript submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:语码转换(CS)是指语音信号中语言的转换,并导致自动语音识别(ASR)的语言混乱。为了解决语言混乱,我们提出了语言对齐损失,使用从ASR解码器学习的伪语言标签执行帧级语言识别。这消除了对帧级语言注释的需要。为了进一步解决双语场景中语言建模的复杂令牌替代方案,我们建议通过生成纠错方法采用大型语言模型。一个语言暗示,结合语言信息(来自建议的语言对齐损失和解码的假设),以指导大型语言模型的提示。在SEAME数据集和ASRU 2019年普通话-英语代码切换语音识别挑战赛的数据上对所提出的方法进行了评估。与基线模型相比,所提出的语言对齐损失的并入证明了更高的CS-ASR性能,两个数据集上的参数数量仅增加了可忽略不计。这项工作还强调了语言对齐损失在训练过程中平衡主要语言占主导地位的双语数据的有效性,与基线模型相比,ASRU数据集相对改善了8.6%。使用大型语言模型进行的性能评估显示了语言提示的优势,在ASRU和SEAME数据集的测试集上分别实现了14.1%和5.5%的相对改进。
摘要:Code-switching (CS) refers to the switching of languages within a speech signal and results in language confusion for automatic speech recognition (ASR). To address language confusion, we propose the language alignment loss that performs frame-level language identification using pseudo language labels learned from the ASR decoder. This eliminates the need for frame-level language annotations. To further tackle the complex token alternatives for language modeling in bilingual scenarios, we propose to employ large language models via a generative error correction method. A linguistic hint that incorporates language information (derived from the proposed language alignment loss and decoded hypotheses) is introduced to guide the prompting of large language models. The proposed methods are evaluated on the SEAME dataset and data from the ASRU 2019 Mandarin-English code-switching speech recognition challenge. The incorporation of the proposed language alignment loss demonstrates a higher CS-ASR performance with only a negligible increase in the number of parameters on both datasets compared to the baseline model. This work also highlights the efficacy of language alignment loss in balancing primary-language-dominant bilingual data during training, with an 8.6% relative improvement on the ASRU dataset compared to the baseline model. Performance evaluation using large language models reveals the advantage of the linguistic hint by achieving 14.1% and 5.5% relative improvement on test sets of the ASRU and SEAME datasets, respectively.

【4】 Asynchronous Microphone Array Calibration using Hybrid TDOA Information
标题:利用混合时差信息的异步麦克风阵列校准
链接:https://arxiv.org/abs/2403.05791
作者:Chengjie Zhang,Jiang Wang,You-Fu Li,He Kong
摘要:异步麦克风阵列校准是大多数试听机器人应用的先决条件。在实践中,校准需要同时估计麦克风位置、时间偏移、时钟漂移率和声音事件位置。现有的方法提出了基于图的同时定位和映射(Graph-SLAM)利用共同的TDOA,两个麦克风之间的到达时间差(TDOA-M)和里程测量,但是,它严重依赖于初始值。本文提出了一种新的TDOA--相邻声事件到达时间差(TDOA-S),将其与TDOA-M相结合,称为混合TDOA,并加入测距测量,构建Graph-SLAM,采用高斯-牛顿(GN)法求解。TDOA-S简单有效,因为它消除了时间偏移而不产生新的变量。仿真和实际实验结果表明,该方法不受麦克风数目的影响,对初始值不敏感,在各种TDOA噪声下具有较好的校准精度和稳定性。仿真结果表明,该方法对麦克风参数具有较低的Cram\'er-Rao下界(CRLB),说明了该方法的优越性.
摘要:Asynchronous Microphone array calibration is a prerequisite for most audition robot applications. In practice, the calibration requires estimating microphone positions, time offsets, clock drift rates, and sound event locations simultaneously. The existing method proposed Graph-based Simultaneous Localisation and Mapping (Graph-SLAM) utilizing common TDOA, time difference of arrival between two microphones (TDOA-M), and odometry measurement, however, it heavily depends on the initial value. In this paper, we propose a novel TDOA, time difference of arrival between adjacent sound events (TDOA-S), combine it with TDOA-M, called hybrid TDOA, and add odometry measurement to construct Graph-SLAM and use the Gauss-Newton (GN) method to solve. TDOA-S is simple and efficient because it eliminates time offset without generating new variables. Simulation and real-world experiment results consistently show that our method is independent of microphone number, insensitive to initial values, and has better calibration accuracy and stability under various TDOA noises. In addition, the simulation result demonstrates that our method has a lower Cram\'er-Rao lower bound (CRLB) for microphone parameters, which explains the advantages of my method.


【5】 Multilingual Turn-taking Prediction Using Voice Activity Projection
标题:基于语音活动投影的多语种话轮转换预测
链接:https://arxiv.org/abs/2403.06487
作者:Koji Inoue,Bing'er Jiang,Erik Ekstedt,Tatsuya Kawahara,Gabriel Skantze
备注:This paper has been accepted for presentation at The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) and represents the author's version of the work
摘要:本文研究了语音活动投影(VAP),一个预测轮对模型的口语对话,对多语种数据,包括英语,汉语和日语的应用。VAP模型持续预测二元对话中参与者即将到来的语音活动,利用交叉注意Transformer来捕获参与者之间的动态交互。结果表明,在一种语言上训练的单语VAP模型在应用于其他语言时不能做出很好的预测。然而,在所有三种语言上训练的多语言模型显示出与所有语言的单语言模型相当的预测性能。进一步的分析表明,多语言模型已经学会识别输入信号的语言。我们还分析了对音高的敏感性,这是一种被认为对话轮转换很重要的韵律线索。最后,我们比较了两种不同的音频编码器,即在英语上预训练的对比预测编码(CPC)和基于多语言wav2vec 2.0(MMS)的最新模型。
摘要:This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS).


【6】 Cosine Scoring with Uncertainty for Neural Speaker Embedding
标题:神经说话人嵌入的不确定余弦评分
链接:https://arxiv.org/abs/2403.06404
作者:Qiongqiong Wang,Kong Aik Lee
备注:None
摘要:说话人表示中的不确定性建模旨在学习语音话语中存在的可变性。虽然传统的余弦评分是计算效率和普遍的说话人识别,它缺乏处理不确定性的能力。为了解决这一问题,本文提出了一种方法,用于估计不确定性的扬声器嵌入前端和传播到余弦评分后端。在VoxCeleb和SITW数据集上进行的实验证实了该方法在处理嵌入估计引起的不确定性方面的有效性。与传统的余弦相似性相比,它实现了EER和minDCF平均降低8.5%和9.8%的改善。它在实践中也是计算效率高的。
摘要:Uncertainty modeling in speaker representation aims to learn the variability present in speech utterances. While the conventional cosine-scoring is computationally efficient and prevalent in speaker recognition, it lacks the capability to handle uncertainty. To address this challenge, this paper proposes an approach for estimating uncertainty at the speaker embedding front-end and propagating it to the cosine scoring back-end. Experiments conducted on the VoxCeleb and SITW datasets confirmed the efficacy of the proposed method in handling uncertainty arising from embedding estimation. It achieved improvement with 8.5% and 9.8% average reductions in EER and minDCF compared to the conventional cosine similarity. It is also computationally efficient in practice.

【7】 Towards Decoupling Frontend Enhancement and Backend Recognition in  Monaural Robust ASR
标题:单声道鲁棒ASR的前端增强与后端识别解耦研究
链接:https://arxiv.org/abs/2403.06387
作者:Yufeng Yang,Ashutosh Pandey,DeLiang Wang
备注:Submitted to IEEE/ACM Transactions on Audio, Speech and Language Processing. arXiv admin note: text overlap with arXiv:2210.13318
摘要:研究表明,语音增强(SE)算法可以提高含噪语音的可懂度。然而,单声道SE还没有被建立为一个有效的前端自动语音识别(ASR)在嘈杂的条件下相比,一个ASR模型直接训练嘈杂的语音。SE和ASR之间的鸿沟阻碍了强大的ASR系统的发展,特别是SE近年来取得了重大进展。本文的重点是消除这种鸿沟与ARN(注意递归网络)时域和CrossNet时频域增强模型。所提出的系统完全解耦前端增强和后端ASR只在干净的语音上训练。WSJ、CHiME-2、LibriSpeech和CHiME-4语料库上的结果表明,ARN和CrossNet增强语音在嘈杂和混响环境中都能改善ASR结果,并很好地推广到真实的声学场景。所提出的系统优于直接在损坏的语音上训练的基线。此外,该算法将CHiME-2上的最佳字错误率(WER)降低了28.4美元,WER为5.57美元,在没有CHiME-4训练的情况下,单通道CHiME-4模拟/真实测试数据的WER为3.32/4.44美元。
摘要:It has been shown that the intelligibility of noisy speech can be improved by speech enhancement (SE) algorithms. However, monaural SE has not been established as an effective frontend for automatic speech recognition (ASR) in noisy conditions compared to an ASR model trained on noisy speech directly. The divide between SE and ASR impedes the progress of robust ASR systems, especially as SE has made major advances in recent years. This paper focuses on eliminating this divide with an ARN (attentive recurrent network) time-domain and a CrossNet time-frequency domain enhancement models. The proposed systems fully decouple frontend enhancement and backend ASR trained only on clean speech. Results on the WSJ, CHiME-2, LibriSpeech, and CHiME-4 corpora demonstrate that ARN and CrossNet enhanced speech both translate to improved ASR results in noisy and reverberant environments, and generalize well to real acoustic scenarios. The proposed system outperforms the baselines trained on corrupted speech directly. Furthermore, it cuts the previous best word error rate (WER) on CHiME-2 by $28.4\%$ relatively with a $5.57\%$ WER, and achieves $3.32/4.44\%$ WER on single-channel CHiME-4 simulated/real test data without training on CHiME-4.


【8】 SCORE: Self-supervised Correspondence Fine-tuning for Improved Content  Representations
标题:得分:自我监督通信微调以改进内容表示
链接:https://arxiv.org/abs/2403.06260
作者:Amit Meghanani,Thomas Hain
备注:Accepted at ICASSP 2024
摘要:有越来越多的兴趣在成本效益的自我监督微调(SSFT)的自我监督学习(SSL)为基础的语音模型,以获得特定于任务的表示。这些特定于任务的表示通过对标记数据进行微调来用于各种下游任务的鲁棒性能。这项工作提出了一种具有成本效益的SSFT方法命名为自监督对应(SCORE)微调,以适应SSL语音表示内容相关的任务。该方法使用对应训练策略,旨在从扰动语音和原始语音中学习相似的表示。通常使用的内容相关的任务(ASR)的数据增强技术,以获得扰动的语音。SCORE微调的HuBERT在SUPERB基准测试中的表现优于普通的HuBERT,在单个GPU上进行自动语音识别,音素识别和按示例查询任务的微调仅需几个小时(< 5小时),相对提高分别为1.09%,3.58%和12.65%。SCORE提供了与最近提出的SSFT方法SPIN竞争的结果,与SPIN相比,仅使用1/3的处理语音。
摘要:There is a growing interest in cost-effective self-supervised fine-tuning (SSFT) of self-supervised learning (SSL)-based speech models to obtain task-specific representations. These task-specific representations are used for robust performance on various downstream tasks by fine-tuning on the labelled data. This work presents a cost-effective SSFT method named Self-supervised Correspondence (SCORE) fine-tuning to adapt the SSL speech representations for content-related tasks. The proposed method uses a correspondence training strategy, aiming to learn similar representations from perturbed speech and original speech. Commonly used data augmentation techniques for content-related tasks (ASR) are applied to obtain perturbed speech. SCORE fine-tuned HuBERT outperforms the vanilla HuBERT on SUPERB benchmark with only a few hours of fine-tuning (< 5 hrs) on a single GPU for automatic speech recognition, phoneme recognition, and query-by-example tasks, with relative improvements of 1.09%, 3.58%, and 12.65%, respectively. SCORE provides competitive results with the recently proposed SSFT method SPIN, using only 1/3 of the processed speech compared to SPIN.


【9】 Automatic design optimization of preference-based subjective evaluation  with online learning in crowdsourcing environment
标题:众包环境下基于偏好的在线学习主观评价自动优化设计
链接:https://arxiv.org/abs/2403.06100
作者:Yusuke Yasuda,Tomoki Toda
摘要:基于偏好的主观评价是可靠评价生成媒体的关键方法。然而,其庞大的配对组合使其无法应用于使用众包的大规模评估。为了解决这个问题,我们提出了一种自动优化方法,基于偏好的主观评价对组合的选择和分配的评价卷与在线学习在众包环境中。我们使用基于偏好的在线学习方法的排序算法的基础上,以最小的样本量来确定总的顺序的评价目标。我们的在线学习算法支持在众包所需的固定预算条件下的并行和异步执行。我们的实验基于偏好的主观评价的合成语音表明,我们的方法成功地优化了测试,减少对组合从351到83和分配最佳的评价卷为每对从30到663不等,而不影响评价精度和浪费预算分配。
摘要:A preference-based subjective evaluation is a key method for evaluating generative media reliably. However, its huge combinations of pairs prohibit it from being applied to large-scale evaluation using crowdsourcing. To address this issue, we propose an automatic optimization method for preference-based subjective evaluation in terms of pair combination selections and allocation of evaluation volumes with online learning in a crowdsourcing environment. We use a preference-based online learning method based on a sorting algorithm to identify the total order of evaluation targets with minimum sample volumes. Our online learning algorithm supports parallel and asynchronous execution under fixed-budget conditions required for crowdsourcing. Our experiment on preference-based subjective evaluation of synthetic speech shows that our method successfully optimizes the test by reducing pair combinations from 351 to 83 and allocating optimal evaluation volumes for each pair ranging from 30 to 663 without compromising evaluation accuracies and wasting budget allocations.

【10】 HAM-TTS: Hierarchical Acoustic Modeling for Token-Based Zero-Shot  Text-to-Speech with Model and Data Scaling
标题:HAM-TTS:基于模型和数据缩放的基于令牌的零发声文语合成的分层声学建模
链接:https://arxiv.org/abs/2403.05989
作者:Chunhui Wang,Chang Zeng,Bowen Zhang,Ziyang Ma,Yefan Zhu,Zifeng Cai,Jian Zhao,Zhonglin Jiang,Yong Chen
摘要:基于令牌的文本到语音(TTS)模型已经成为生成自然和逼真语音的一种有前途的途径,但它们面临着发音准确性低、说话风格和音色不一致以及对多样化训练数据的大量需求。作为回应,我们引入了一种新的分层声学建模方法,辅以量身定制的数据增强策略,并结合真实数据和合成数据对其进行训练,将数据大小扩展到65万小时,从而获得具有0.8B参数的zero-shot TTS模型。具体来说,我们的方法采用了一个潜在的变量序列包含补充声学信息的基础上,完善的自监督学习(SSL)离散单元到TTS模型的预测。这显著地减轻了合成语音中的发音错误和风格突变。在训练过程中,我们策略性地替换和复制数据片段,以提高音色的一致性。此外,利用预先训练的Few-Shot语音转换模型来生成具有相同内容但不同音色的过多语音。这促进了话语级一对多映射的显式学习,丰富了语音多样性,并确保了音色的一致性。对比实验(演示页面:https://anonymous.4open.science/w/ham-tts/)证明了我们的模型在发音精度和保持说话风格以及音色连续性方面优于VALL-E。
摘要:Token-based text-to-speech (TTS) models have emerged as a promising avenue for generating natural and realistic speech, yet they grapple with low pronunciation accuracy, speaking style and timbre inconsistency, and a substantial need for diverse training data. In response, we introduce a novel hierarchical acoustic modeling approach complemented by a tailored data augmentation strategy and train it on the combination of real and synthetic data, scaling the data size up to 650k hours, leading to the zero-shot TTS model with 0.8B parameters. Specifically, our method incorporates a latent variable sequence containing supplementary acoustic information based on refined self-supervised learning (SSL) discrete units into the TTS model by a predictor. This significantly mitigates pronunciation errors and style mutations in synthesized speech. During training, we strategically replace and duplicate segments of the data to enhance timbre uniformity. Moreover, a pretrained few-shot voice conversion model is utilized to generate a plethora of voices with identical content yet varied timbres. This facilitates the explicit learning of utterance-level one-to-many mappings, enriching speech diversity and also ensuring consistency in timbre. Comparative experiments (Demo page: https://anonymous.4open.science/w/ham-tts/)demonstrate our model's superiority over VALL-E in pronunciation precision and maintaining speaking style, as well as timbre continuity.

【11】 Enhancing Expressiveness in Dance Generation via Integrating Frequency  and Music Style Information
标题:通过整合频率和音乐风格信息增强舞蹈生成的表现力
链接:https://arxiv.org/abs/2403.05834
作者:Qiaochu Huang,Xu He,Boshi Tang,Haolin Zhuang,Liyang Chen,Shuochen Gao,Zhiyong Wu,Haozhi Huang,Helen Meng
摘要:舞蹈生成作为人体运动生成的一个分支,越来越受到人们的关注。近年来,一些作品试图从不同的角度来增强舞蹈的表现力,包括体裁匹配、节拍对齐和舞蹈动态。但是,由于缺乏对上述三个因素的综合考虑,这种改进是相当有限的。在本文中,我们提出了ExpressiveBailando,一种新颖的舞蹈生成方法,旨在生成富有表现力的舞蹈,同时考虑到所有三个因素。具体来说,我们通过将频率信息纳入VQ-VAE来缓解速度均匀化的问题,从而改善舞蹈动态。此外,我们通过预训练的音乐模型提取流派和节拍相关的特征来整合音乐风格信息,从而实现其他两个因素的改进。大量的实验结果表明,我们提出的方法可以生成具有高表现力的舞蹈,并优于现有的方法在定性和定量。
摘要:Dance generation, as a branch of human motion generation, has attracted increasing attention. Recently, a few works attempt to enhance dance expressiveness, which includes genre matching, beat alignment, and dance dynamics, from certain aspects. However, the enhancement is quite limited as they lack comprehensive consideration of the aforementioned three factors. In this paper, we propose ExpressiveBailando, a novel dance generation method designed to generate expressive dances, concurrently taking all three factors into account. Specifically, we mitigate the issue of speed homogenization by incorporating frequency information into VQ-VAE, thus improving dance dynamics. Additionally, we integrate music style information by extracting genre- and beat-related features with a pre-trained music model, hence achieving improvements in the other two factors. Extensive experimental results demonstrate that our proposed method can generate dances with high expressiveness and outperforms existing methods both qualitatively and quantitatively.

【12】 An Audio-textual Diffusion Model For Converting Speech Signals Into  Ultrasound Tongue Imaging Data
标题:一种将语音信号转换为超声舌象数据的音频-文本扩散模型
链接:https://arxiv.org/abs/2403.05820
作者:Yudong Yang,Rongfeng Su,Xiaokang Liu,Nan Yan,Lan Wang
备注:ICASSP2024 Accept
摘要:声学发音反转(AAI)是将音频转换为发音器运动,例如超声舌成像(UTI)数据。现有的AAI方法的一个问题是仅使用个性化的声学信息来导出舌头运动的一般模式,因此生成的UTI数据的质量是有限的。为了解决这个问题,本文提出了一个音频文本扩散模型的UTI数据生成任务。在该模型中,与舌头运动细节相关的个体固有声学特征使用wav 2 vec 2.0编码,而与舌头运动的普遍性相关的ASR transmittance使用BERT编码。然后通过使用扩散模块生成UTI数据。实验结果表明,该扩散模型可以生成高质量的UTI数据,具有清晰的舌头轮廓,这是至关重要的语言分析和临床评估。该项目可在网站\footnote{https:yangyudong2020.github.io/wav2uti/
摘要:Acoustic-to-articulatory inversion (AAI) is to convert audio into articulator movements, such as ultrasound tongue imaging (UTI) data. An issue of existing AAI methods is only using the personalized acoustic information to derive the general patterns of tongue motions, and thus the quality of generated UTI data is limited. To address this issue, this paper proposes an audio-textual diffusion model for the UTI data generation task. In this model, the inherent acoustic characteristics of individuals related to the tongue motion details are encoded by using wav2vec 2.0, while the ASR transcriptions related to the universality of tongue motions are encoded by using BERT. UTI data are then generated by using a diffusion module. Experimental results showed that the proposed diffusion model could generate high-quality UTI data with clear tongue contour that is crucial for the linguistic analysis and clinical assessment. The project can be found on the website\footnote{https://yangyudong2020.github.io/wav2uti/


【13】 sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection  with Spiking Neural Networks
标题:SVAD:一种基于尖峰神经网络的健壮、低功耗、轻量级语音活动检测
链接:https://arxiv.org/abs/2403.05772
作者:Qu Yang,Qianhui Liu,Nan Li,Meng Ge,Zeyang Song,Haizhou Li
备注:Accepted by ICASSP 2024
摘要:语音应用被期望在噪声条件下是低功耗和鲁棒的。有效的语音活动检测(VAD)前端降低了计算需求。尖峰神经网络(SNN)被认为是生物学上合理的和节能的。然而,基于SNN的VAD尚未实现噪声鲁棒性,并且通常需要大型模型来实现高性能。本文介绍了一种新的SNN为基础的VAD模型,简称为sVAD,它具有一个听觉编码器与SNN为基础的注意力机制。特别地,它通过SincNet和1D卷积提供有效的听觉特征表示,并通过注意机制提高噪声鲁棒性。该分类器利用尖峰递归神经网络(sRNN)来利用时间语音信息。实验结果表明,我们的sVAD实现了显着的噪声鲁棒性,同时保持低功耗和小的足迹,使其成为一个有前途的解决方案,为现实世界的VAD应用。
摘要:Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achieve noise robustness and often require large models for high performance. This paper introduces a novel SNN-based VAD model, referred to as sVAD, which features an auditory encoder with an SNN-based attention mechanism. Particularly, it provides effective auditory feature representation through SincNet and 1D convolution, and improves noise robustness with attention mechanisms. The classifier utilizes Spiking Recurrent Neural Networks (sRNN) to exploit temporal speech information. Experimental results demonstrate that our sVAD achieves remarkable noise robustness and meanwhile maintains low power consumption and a small footprint, making it a promising solution for real-world VAD applications.

【14】 A Cross-Modal Approach to Silent Speech with LLM-Enhanced Recognition
标题:一种基于LLM增强识别的无声语音跨模式识别方法
链接:https://arxiv.org/abs/2403.05583
作者:Tyler Benster,Guy Wilson,Reshef Elisha,Francis R Willett,Shaul Druckmann
摘要:无声语音接口(SSI)为无声的口头交流提供了脑机接口的非侵入性替代方案。我们介绍多模态口面神经音频(MONA),一个系统,利用跨模态对齐通过新的损失函数-交叉对比度(crossCon)和监督时间对比度(supTcon)-训练一个多模态模型与共享的潜在表示。这种架构允许使用像LibriSpeech这样的纯音频数据集来改进无声语音识别。此外,我们引入了大型语言模型(LLM)综合评分调整(LISA),显着提高了识别精度。总之,MONA LISA将Gaddy(2020)开放词汇表上无声语音基准数据集中最先进的单词错误率(WER)从28.8%降低到12.2%。对于声音EMG记录,我们的方法将最先进的WER从23.3%提高到3.7%。在Brain-to-Text 2024竞赛中,LISA表现最好,将顶级WER从9.8%提高到8.9%。据我们所知,这项工作代表了第一个实例,在开放词汇表上的非侵入性无声语音识别已经清除了15%的WER阈值,表明SSI可以成为自动语音识别(ASR)的可行替代方案。我们的工作不仅缩小了无声语音和有声语音之间的性能差距,而且还为人机交互开辟了新的可能性,展示了在嘈杂和数据有限的情况下跨模态方法的潜力。
摘要:Silent Speech Interfaces (SSIs) offer a noninvasive alternative to brain-computer interfaces for soundless verbal communication. We introduce Multimodal Orofacial Neural Audio (MONA), a system that leverages cross-modal alignment through novel loss functions--cross-contrast (crossCon) and supervised temporal contrast (supTcon)--to train a multimodal model with a shared latent representation. This architecture enables the use of audio-only datasets like LibriSpeech to improve silent speech recognition. Additionally, our introduction of Large Language Model (LLM) Integrated Scoring Adjustment (LISA) significantly improves recognition accuracy. Together, MONA LISA reduces the state-of-the-art word error rate (WER) from 28.8% to 12.2% in the Gaddy (2020) benchmark dataset for silent speech on an open vocabulary. For vocal EMG recordings, our method improves the state-of-the-art from 23.3% to 3.7% WER. In the Brain-to-Text 2024 competition, LISA performs best, improving the top WER from 9.8% to 8.9%. To the best of our knowledge, this work represents the first instance where noninvasive silent speech recognition on an open vocabulary has cleared the threshold of 15% WER, demonstrating that SSIs can be a viable alternative to automatic speech recognition (ASR). Our work not only narrows the performance gap between silent and vocalized speech but also opens new possibilities in human-computer interaction, demonstrating the potential of cross-modal approaches in noisy and data-limited regimes.


机器翻译由腾讯交互翻译提供,仅供参考