今日论文合集:cs.SD语音10篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing  Using Discrete Speech Unit Challenge

标题:UTDUSS:UTokyo—SaruLab Interspeech 2024语音处理系统使用离散语音单元挑战赛
链接:https://arxiv.org/abs/2403.13720
作者:Wataru Nakata,Kazuki Yamauchi,Dong Yang,Hiroaki Hyodo,Yuki Saito
备注:5 pages, 3 figures
摘要:我们提出UTDUSS,UTokyo-SaruLab系统提交给Interspeech 2024使用离散语音单元挑战的语音处理。挑战的重点是使用从大型语音语料库中学习的离散语音单元来完成某些任务。我们将UTDUSS系统提交给两个文本到语音的轨道:声码器和声学+声码器。我们的系统采用了神经音频编解码器(NAC)只在语音语料库上进行预训练,这使得学习的编解码器代表了高保真语音重建所需的丰富声学特征。对于声学+声码器音轨,我们训练了一个基于Transformer编码器-解码器的声学模型,该编码器-解码器从文本输入预测预训练的NAC标记。我们描述了构建这些模型的策略,例如数据选择,下采样和超参数调整。我们的系统分别在声码器和声学+声码器音轨中排名第二和第一。
摘要:We present UTDUSS, the UTokyo-SaruLab system submitted to Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge. The challenge focuses on using discrete speech unit learned from large speech corpora for some tasks. We submitted our UTDUSS system to two text-to-speech tracks: Vocoder and Acoustic+Vocoder. Our system incorporates neural audio codec (NAC) pre-trained on only speech corpora, which makes the learned codec represent rich acoustic features that are necessary for high-fidelity speech reconstruction. For the acoustic+vocoder track, we trained an acoustic model based on Transformer encoder-decoder that predicted the pre-trained NAC tokens from text input. We describe our strategies to build these models, such as data selection, downsampling, and hyper-parameter tuning. Our system ranked in second and first for the Vocoder and Acoustic+Vocoder tracks, respectively.


【2】 Recursive Cross-Modal Attention for Multimodal Fusion in Dimensional  Emotion Recognition
标题:多维情感识别中多模态融合的递归跨模态注意力
链接:https://arxiv.org/abs/2403.13659
作者:R. Gnana Praveen,Jahangir Alam
备注:arXiv admin note: substantial text overlap with arXiv:2209.09068; text overlap with arXiv:2203.14779 by other authors
摘要:多模态情感识别最近得到了很多关注,因为它可以利用多种模态(如音频,视觉和文本)的多样和互补关系。大多数最先进的多模态融合方法依赖于循环网络或传统的注意力机制,这些机制不能有效地利用模态的互补性。在本文中,我们专注于三维情感识别的基础上融合的面部,声音和文本形式提取的视频。具体来说,我们提出了一个递归的跨模态注意(RCMA),以有效地捕捉互补的关系,在一个递归的方式跨模态。该模型能够有效地捕捉跨通道的关系,通过计算跨单个通道的交叉注意力权重和其他两个通道的联合表示。为了进一步改善模态间关系,所获得的个体模态的关注特征再次作为输入被馈送到跨模态关注,以细化个体模态的特征表示。除此之外,我们还使用了时间卷积网络(TCN)来捕获各个模态的时间建模(模态内关系)。通过以递归方式部署TCN以及跨模态注意,我们能够有效地捕获跨音频,视觉和文本模态的模态内和模态间关系。在AffWild2数据集的验证集视频上的实验结果表明,我们提出的融合模型能够在2024年野外情感行为分析(ABAW6)竞赛的第六次挑战中实现显著的改进。
摘要:Multi-modal emotion recognition has recently gained a lot of attention since it can leverage diverse and complementary relationships over multiple modalities, such as audio, visual, and text. Most state-of-the-art methods for multimodal fusion rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complementary nature of the modalities. In this paper, we focus on dimensional emotion recognition based on the fusion of facial, vocal, and text modalities extracted from videos. Specifically, we propose a recursive cross-modal attention (RCMA) to effectively capture the complementary relationships across the modalities in a recursive fashion. The proposed model is able to effectively capture the inter-modal relationships by computing the cross-attention weights across the individual modalities and the joint representation of the other two modalities. To further improve the inter-modal relationships, the obtained attended features of the individual modalities are again fed as input to the cross-modal attention to refine the feature representations of the individual modalities. In addition to that, we have used Temporal convolution networks (TCNs) to capture the temporal modeling (intra-modal relationships) of the individual modalities. By deploying the TCNs as well cross-modal attention in a recursive fashion, we are able to effectively capture both intra- and inter-modal relationships across the audio, visual, and text modalities. Experimental results on validation-set videos from the AffWild2 dataset indicate that our proposed fusion model is able to achieve significant improvement over the baseline for the sixth challenge of Affective Behavior Analysis in-the-Wild 2024 (ABAW6) competition.


【3】 Advanced Long-Content Speech Recognition With Factorized Neural  Transducer
标题:基于因式分解神经传感器的高级长内容语音识别
链接:https://arxiv.org/abs/2403.13423
作者:Xun Gong,Yu Wu,Jinyu Li,Shujie Liu,Rui Zhao,Xie Chen,Yanmin Qian
备注:None
摘要:在本文中,我们提出了两种新的方法,将长内容信息集成到基于非流(简称为LongFNT)和流(简称为SLongFNT)场景的因子分解神经换能器(FNT)架构中。我们首先调查是否长含量transmantine可以改善香草构象转换器(C-T)模型。我们的实验表明,香草C-T模型没有表现出改善的性能时,利用长内容transmittance,可能是由于预测网络的C-T模型不作为一个纯语言模型。相反,FNT显示其潜力,利用长内容的信息,在这里我们提出了LongFNT模型,并探讨文本(LongFNT-Text)和语音(LongFNT-Speech)的长内容信息的影响。所提出的LongFNT-Text和LongFNT-Speech模型进一步相互补充以实现更好的性能,转录历史对模型更有价值。在LibriSpeech和GigaSpeech语料库上评估了LongFNT方法的有效性,并分别获得了相对19%和12%的单词错误率降低。此外,我们扩展的LongFNT模型的流媒体场景,这是命名为SLongFNT,包括SLongFNT-Text和SLongFNT-Speech方法,利用长内容的文本和语音信息。实验结果表明,与FNT基线相比,所提出的SLongFNT模型在LibriSpeech和GigaSpeech上分别实现了相对26%和17%的WER降低,同时保持了良好的延迟。总的来说,我们提出的LongFNT和SLongFNT突出了考虑长内容语音和转录知识对改善非流和流语音识别系统的重要性。
摘要:In this paper, we propose two novel approaches, which integrate long-content information into the factorized neural transducer (FNT) based architecture in both non-streaming (referred to as LongFNT ) and streaming (referred to as SLongFNT ) scenarios. We first investigate whether long-content transcriptions can improve the vanilla conformer transducer (C-T) models. Our experiments indicate that the vanilla C-T models do not exhibit improved performance when utilizing long-content transcriptions, possibly due to the predictor network of C-T models not functioning as a pure language model. Instead, FNT shows its potential in utilizing long-content information, where we propose the LongFNT model and explore the impact of long-content information in both text (LongFNT-Text) and speech (LongFNT-Speech). The proposed LongFNT-Text and LongFNT-Speech models further complement each other to achieve better performance, with transcription history proving more valuable to the model. The effectiveness of our LongFNT approach is evaluated on LibriSpeech and GigaSpeech corpora, and obtains relative 19% and 12% word error rate reduction, respectively. Furthermore, we extend the LongFNT model to the streaming scenario, which is named SLongFNT , consisting of SLongFNT-Text and SLongFNT-Speech approaches to utilize long-content text and speech information. Experiments show that the proposed SLongFNT model achieves relative 26% and 17% WER reduction on LibriSpeech and GigaSpeech respectively while keeping a good latency, compared to the FNT baseline. Overall, our proposed LongFNT and SLongFNT highlight the significance of considering long-content speech and transcription knowledge for improving both non-streaming and streaming speech recognition systems.


【4】 Building speech corpus with diverse voice characteristics for its  prompt-based representation
标题:基于语音特征的语音库的构建
链接:https://arxiv.org/abs/2403.13353
作者:Aya Watanabe,Shinnosuke Takamichi,Yuki Saito,Wataru Nakata,Detai Xin,Hiroshi Saruwatari
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing. arXiv admin note: text overlap with arXiv:2309.13509
摘要:在文本到语音合成中,控制语音特征的能力对于各种应用至关重要。通过利用蓬勃发展的基于文本的生成技术,应该可以增强对语音特征的细微控制。虽然以前的研究已经探索了基于语音特征的操纵,但大多数研究都使用预先录制的语音,这限制了可用语音特征的多样性。因此,我们的目标是通过创建一个新的语料库,并开发一个模型,在文本到语音合成的语音特征的基于人工智能的操作,以解决这一差距,促进更广泛的语音特征。具体来说,我们提出了一种方法来建立一个相当大的语料库配对语音特征描述与相应的语音样本。这涉及自动从互联网收集与语音相关的语音数据,确保其质量,并使用众包手动注释。我们用日语数据实现了该方法,并对结果进行了分析,以验证其有效性。随后,我们提出了一种基于对比学习方法的模型的构建方法,从语音特征描述中检索语音。我们训练的模型,不仅使用保守的对比学习,但也特征预测学习,预测语音特征对应的定量语音特征。我们通过实验来评估模型的性能与我们上面构建的语料库。
摘要:In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice characteristics. While previous research has explored the prompt-based manipulation of voice characteristics, most studies have used pre-recorded speech, which limits the diversity of voice characteristics available. Thus, we aim to address this gap by creating a novel corpus and developing a model for prompt-based manipulation of voice characteristics in text-to-speech synthesis, facilitating a broader range of voice characteristics. Specifically, we propose a method to build a sizable corpus pairing voice characteristics descriptions with corresponding speech samples. This involves automatically gathering voice-related speech data from the Internet, ensuring its quality, and manually annotating it using crowdsourcing. We implement this method with Japanese language data and analyze the results to validate its effectiveness. Subsequently, we propose a construction method of the model to retrieve speech from voice characteristics descriptions based on a contrastive learning method. We train the model using not only conservative contrastive learning but also feature prediction learning to predict quantitative speech features corresponding to voice characteristics. We evaluate the model performance via experiments with the corpus we constructed above.


【5】 Onset and offset weighted loss function for sound event detection
标题:用于声音事件检测的起始和偏移加权损失函数
链接:https://arxiv.org/abs/2403.13254
作者:Tao Song
摘要:在典型的声音事件检测(SED)系统中,在帧级检测声音事件的存在,并且将检测到相同事件的连续帧组合为一个声音事件。中值滤波器被应用作为后处理步骤,以尽可能地去除检测误差。然而,发生在声音事件的起始和偏移周围的检测误差超出了中值滤波器的能力。为了解决这个问题,本文提出了一种起始和偏移加权二进制交叉熵(OWBCE)损失函数,它训练DNN模型在(a)起始和偏移周围的帧上更加鲁棒。在DCASE 2022任务4的背景下进行实验。结果表明,OWBCE优于BCE时,考虑不同的模型。对于基本CRNN,OWBCE可以实现事件F1的6.43%,PSDS 1的1.96%和PSDS 2的2.43%的相对改善。
摘要:In a typical sound event detection (SED) system, the existence of a sound event is detected at a frame level, and consecutive frames with the same event detected are combined as one sound event. The median filter is applied as a post-processing step to remove detection errors as much as possible. However, detection errors occurring around the onset and offset of a sound event are beyond the capacity of the median filter. To address this issue, an onset and offset weighted binary cross-entropy (OWBCE) loss function is proposed in this paper, which trains the DNN model to be more robust on frames around (a) onsets and offsets. Experiments are carried out in the context of DCASE 2022 task 4. Results show that OWBCE outperforms BCE when different models are considered. For a basic CRNN, relative improvements of 6.43% in event-F1, 1.96% in PSDS1, and 2.43% in PSDS2 can be achieved by OWBCE.


【6】 Frequency-aware convolution for sound event detection
标题:用于声音事件检测的频率感知卷积
链接:https://arxiv.org/abs/2403.13252
作者:Tao Song
摘要:在声音事件检测(SED)中,卷积神经网络(CNN)被广泛用于从输入频谱图中提取时频模式。然而,由CNN提取的特征可能对时频模式沿频率轴的偏移不敏感。为了解决这个问题,已经提出了频率动态卷积(FDY),其将不同的内核应用于不同的频率分量。与vannila CNN相比,FDY需要多几倍的参数。在本文中,一个更有效的解决方案命名为频率感知卷积(FAC)。在FAC中,频率位置信息被编码在矢量中并被添加到输入频谱图。为了匹配输入的幅度,编码矢量自适应缩放和通道无关。在DCASE 2022任务4的背景下进行了实验,结果表明,FAC可以实现与FDY的性能相当,只有515个额外的参数,而FDY需要802万个额外的参数。消融研究表明,自适应缩放的编码矢量和信道无关的是至关重要的FAC的性能。
摘要:In sound event detection (SED), convolution neural networks (CNNs) are widely used to extract time-frequency patterns from the input spectrogram. However, features extracted by CNN can be insensitive to the shift of time-frequency patterns along the frequency axis. To address this issue, frequency dynamic convolution (FDY) has been proposed, which applies different kernels to different frequency components. Compared to the vannila CNN, FDY requires several times more parameters. In this paper, a more efficient solution named frequency-aware convolution (FAC) is proposed. In FAC, frequency-positional information is encoded in a vector and added to the input spectrogram. To match the amplitude of input, the encoding vector is scaled adaptively and channel-independently. Experiments are carried out in the context of DCASE 2022 task 4, and the results demonstrate that FAC can achieve comparable performance to that of FDY with only 515 additional parameters, while FDY requires 8.02 million additional parameters. The ablation study shows that scaling the encoding vector adaptively and channel-independently is critical to the performance of FAC.


【7】 Listenable Maps for Audio Classifiers
标题:音频分类器的可听映射
链接:https://arxiv.org/abs/2403.13086
作者:Francesco Paissan,Mirco Ravanelli,Cem Subakan
摘要:尽管深度学习模型在不同任务中的表现令人印象深刻,但其复杂性给解释带来了挑战。这一挑战对于音频信号尤其明显,其中传达解释变得固有地困难。为了解决这个问题,我们引入了可听地图音频分类器(L-MAC),一个事后解释方法,产生忠实和可解释的解释。L-MAC利用预训练分类器之上的解码器来生成突出输入音频的相关部分的二进制掩码。我们训练具有特殊损失的解码器,该特殊损失最大化分类器对音频的掩蔽部分的决策的置信度,同时最小化掩蔽部分的模型输出的概率。对域内和域外数据的定量评价表明,L-MAC始终比几种基于梯度和掩蔽的方法产生更忠实的解释。此外,用户研究证实,平均而言,用户更喜欢所提出的技术产生的解释。
摘要:Despite the impressive performance of deep learning models across diverse tasks, their complexity poses challenges for interpretation. This challenge is particularly evident for audio signals, where conveying interpretations becomes inherently difficult. To address this issue, we introduce Listenable Maps for Audio Classifiers (L-MAC), a posthoc interpretation method that generates faithful and listenable interpretations. L-MAC utilizes a decoder on top of a pretrained classifier to generate binary masks that highlight relevant portions of the input audio. We train the decoder with a special loss that maximizes the confidence of the classifier decision on the masked-in portion of the audio while minimizing the probability of model output for the masked-out portion. Quantitative evaluations on both in-domain and out-of-domain data demonstrate that L-MAC consistently produces more faithful interpretations than several gradient and masking-based methodologies. Furthermore, a user study confirms that, on average, users prefer the interpretations generated by the proposed technique.

【8】 Vibration Sensitivity of one-port and two-port MEMS microphones
标题:单端口和双端口MEMS麦克风的振动灵敏度
链接:https://arxiv.org/abs/2403.13643
作者:Francis Doyon-D'Amour,Carly Stalder,Timothy Hodges,Michel Stephan,Lixiue Wu,Triantafillos Koukoulas,Stephane Leahy,Raphael St-Gelais
备注:8 pages, 14 figures
摘要:具有两个声学端口的微机电系统(MEMS)麦克风(MEMS)目前受到相当大的关注,与传统的单端口架构相比,其有望实现更高的方向灵敏度。然而,在双端口麦克风中测量压差通常需要比单端口麦克风更软的传感元件,因此可能更容易受到外部振动的干扰。在这里,我们推导出一个通用的麦克风振动灵敏度的表达式,我们实验证明其有效性的几个新兴的双端口麦克风技术。我们还在单端口麦克风上进行振动测量,从而提供单端口和双端口传感方法之间的一站式直接比较。我们发现,双端口MEMS器件的声学参考振动灵敏度,以每外部加速度测量的声压为单位(即,帕斯卡每克),不取决于敏感元件的刚度,也不取决于其固有频率。我们还表明,这种振动的灵敏度在两个端口的麦克风与频率成反比,而不是在一个端口的麦克风观察到的频率无关的行为。这是证实了几种类型的麦克风封装实验。
摘要:Micro-electro-mechanical system (MEMS) microphones (mics) with two acoustic ports are currently receiving considerable interest, with the promise of achieving higher directional sensitivity compared to traditional one-port architectures. However, measuring pressure differences in two-port microphones typically commands sensing elements that are softer than in one-port mics, and are therefore presumably more prone to interference from external vibration. Here we derive a universal expression for microphone sensitivity to vibration and we experimentally demonstrate its validity for several emerging two-port microphone technologies. We also perform vibration measurements on a one-port mic, thus providing a one-stop direct comparison between one-port and two-port sensing approaches. We find that the acoustically-referred vibration sensitivity of two-port MEMS mics, in units of measured acoustic pressure per external acceleration (i.e., Pascals per g), does not depend on the sensing element stiffness nor on its natural frequency. We also show that this vibration sensitivity in two-port mics is inversely proportional to frequency as opposed to the frequency independent behavior observed in one-port mics. This is confirmed experimentally for several types of microphone packages.

【9】 KunquDB: An Attempt for Speaker Verification in the Chinese Opera  Scenario
标题:昆曲DB:戏曲场景中说话人确认的尝试
链接:https://arxiv.org/abs/2403.13356
作者:Huali Zhou,Yuke Lin,Dong Liu,Ming Li
摘要:本研究旨在克服资料的局限性,促进中国戏曲音乐和语言领域的研究。我们介绍KunquDB,一个相对大规模的,注释良好的视听数据集,包括339个扬声器和128小时的内容。昆曲数据库源于《昆曲艺术宝典》,由台词精心构建,提供明确的注释,包括人物名称、说话人名称、性别信息、发声方式分类,并附有初步的文本翻译。KunquDB为以角色为中心的声学研究和语音相关研究的进步提供了多功能的基础,包括自动说话人验证(ASV)。除了丰富歌剧研究,该数据集还弥合了艺术表达和技术创新之间的差距。本研究首先探讨戏曲中的主动语态,并针对戏曲中的两种不同发声方式:舞台发声(ST)与演唱(S),建构了四个测试实验。实现域自适应方法有效地减轻了这些发声方式变化引起的域失配,同时作为基准仍有进一步改进的空间。
摘要:This work aims to promote Chinese opera research in both musical and speech domains, with a primary focus on overcoming the data limitations. We introduce KunquDB, a relatively large-scale, well-annotated audio-visual dataset comprising 339 speakers and 128 hours of content. Originating from the Kunqu Opera Art Canon (Kunqu yishu dadian), KunquDB is meticulously structured by dialogue lines, providing explicit annotations including character names, speaker names, gender information, vocal manner classifications, and accompanied by preliminary text transcriptions. KunquDB provides a versatile foundation for role-centric acoustic studies and advancements in speech-related research, including Automatic Speaker Verification (ASV). Beyond enriching opera research, this dataset bridges the gap between artistic expression and technological innovation. Pioneering the exploration of ASV in Chinese opera, we construct four test trials considering two distinct vocal manners in opera voices: stage speech (ST) and singing (S). Implementing domain adaptation methods effectively mitigates domain mismatches induced by these vocal manner variations while there is still room for further improvement as a benchmark.


【10】 TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration  Transducer
标题:TDT—KWS:利用令牌和持续时间传感器快速准确地识别关键词
链接:https://arxiv.org/abs/2403.13332
作者:Yu Xi,Hao Li,Baochen Yang,Haoyu Li,Hainan Xu,Kai Yu
备注:Accepted by ICASSP2024
摘要:设计一个高效的关键字识别(KWS)系统,在资源受限的边缘设备上提供卓越的性能一直是一个备受关注的主题。现有的KWS搜索算法通常遵循帧同步方法,其中在每个帧处重复地做出搜索决策,尽管大多数帧是关键字不相关的。在本文中,我们提出了TDT-KWS,它利用令牌和持续时间传感器(TDT)的KWS任务。我们还提出了一种新的KWS特定于任务的解码算法,基于传感器的模型,它支持高效的帧异步关键字搜索流语音场景。通过对公共Hey Snips和自构建的LibriKWS-20数据集进行评估,我们提出的KWS解码算法比传统的ASR解码算法产生更准确的结果。此外,TDT-KWS实现了与RNN-T和传统TDT-ASR系统相当或更好的唤醒字检测性能,同时实现了显著的推理速度提升。此外,实验表明,TDT-KWS是更强大的噪声环境相比,RNN-T KWS。
摘要:Designing an efficient keyword spotting (KWS) system that delivers exceptional performance on resource-constrained edge devices has long been a subject of significant attention. Existing KWS search algorithms typically follow a frame-synchronous approach, where search decisions are made repeatedly at each frame despite the fact that most frames are keyword-irrelevant. In this paper, we propose TDT-KWS, which leverages token-and-duration Transducers (TDT) for KWS tasks. We also propose a novel KWS task-specific decoding algorithm for Transducer-based models, which supports highly effective frame-asynchronous keyword search in streaming speech scenarios. With evaluations conducted on both the public Hey Snips and self-constructed LibriKWS-20 datasets, our proposed KWS-decoding algorithm produces more accurate results than conventional ASR decoding algorithms. Additionally, TDT-KWS achieves on-par or better wake word detection performance than both RNN-T and traditional TDT-ASR systems while achieving significant inference speed-up. Furthermore, experiments show that TDT-KWS is more robust to noisy environments compared to RNN-T KWS.


eess.AS音频处理
【1】 Vibration Sensitivity of one-port and two-port MEMS microphones
标题:单端口和双端口MEMS麦克风的振动灵敏度
链接:https://arxiv.org/abs/2403.13643
作者:Francis Doyon-D'Amour,Carly Stalder,Timothy Hodges,Michel Stephan,Lixiue Wu,Triantafillos Koukoulas,Stephane Leahy,Raphael St-Gelais
备注:8 pages, 14 figures
摘要:具有两个声学端口的微机电系统(MEMS)麦克风(MEMS)目前受到相当大的关注,与传统的单端口架构相比,其有望实现更高的方向灵敏度。然而,在双端口麦克风中测量压差通常需要比单端口麦克风更软的传感元件,因此可能更容易受到外部振动的干扰。在这里,我们推导出一个通用的麦克风振动灵敏度的表达式,我们实验证明其有效性的几个新兴的双端口麦克风技术。我们还在单端口麦克风上进行振动测量,从而提供单端口和双端口传感方法之间的一站式直接比较。我们发现,双端口MEMS器件的声学参考振动灵敏度,以每外部加速度测量的声压为单位(即,帕斯卡每克),不取决于敏感元件的刚度,也不取决于其固有频率。我们还表明,这种振动的灵敏度在两个端口的麦克风与频率成反比,而不是在一个端口的麦克风观察到的频率无关的行为。这是证实了几种类型的麦克风封装实验。
摘要:Micro-electro-mechanical system (MEMS) microphones (mics) with two acoustic ports are currently receiving considerable interest, with the promise of achieving higher directional sensitivity compared to traditional one-port architectures. However, measuring pressure differences in two-port microphones typically commands sensing elements that are softer than in one-port mics, and are therefore presumably more prone to interference from external vibration. Here we derive a universal expression for microphone sensitivity to vibration and we experimentally demonstrate its validity for several emerging two-port microphone technologies. We also perform vibration measurements on a one-port mic, thus providing a one-stop direct comparison between one-port and two-port sensing approaches. We find that the acoustically-referred vibration sensitivity of two-port MEMS mics, in units of measured acoustic pressure per external acceleration (i.e., Pascals per g), does not depend on the sensing element stiffness nor on its natural frequency. We also show that this vibration sensitivity in two-port mics is inversely proportional to frequency as opposed to the frequency independent behavior observed in one-port mics. This is confirmed experimentally for several types of microphone packages.

【2】 BanglaNum -- A Public Dataset for Bengali Digit Recognition from Speech
标题:孟加拉语——一个面向孟加拉语语音数字识别的公共数据集
链接:https://arxiv.org/abs/2403.13465
作者:Mir Sayeed Mohammad,Azizul Zahid,Md Asif Iqbal
摘要:自动语音识别(ASR)将人类的声音转换为易于理解和分类的文本或单词。虽然孟加拉语是世界上使用最广泛的语言之一,但对孟加拉语ASR的研究很少,特别是对孟加拉口音的孟加拉语。在这项研究中,从大学生的口语数字(0-9)的音频记录被用来创建孟加拉语语音数字数据集,可用于训练人工神经网络的语音为基础的数字输入系统。本文还使用频谱图比较了几种卷积神经网络(CNN)的孟加拉数字识别准确率,并表明在我们的数据集上使用参数有效的模型(如SqueezeNet)可以实现98.23%的测试准确率。
摘要:Automatic speech recognition (ASR) converts the human voice into readily understandable and categorized text or words. Although Bengali is one of the most widely spoken languages in the world, there have been very few studies on Bengali ASR, particularly on Bangladeshi-accented Bengali. In this study, audio recordings of spoken digits (0-9) from university students were used to create a Bengali speech digits dataset that may be employed to train artificial neural networks for voice-based digital input systems. This paper also compares the Bengali digit recognition accuracy of several Convolutional Neural Networks (CNNs) using spectrograms and shows that a test accuracy of 98.23% is achievable using parameter-efficient models such as SqueezeNet on our dataset.

【3】 KunquDB: An Attempt for Speaker Verification in the Chinese Opera  Scenario
标题:昆曲DB:戏曲场景中说话人确认的尝试
链接:https://arxiv.org/abs/2403.13356
作者:Huali Zhou,Yuke Lin,Dong Liu,Ming Li
摘要:本研究旨在克服资料的局限性,促进中国戏曲音乐和语言领域的研究。我们介绍KunquDB,一个相对大规模的,注释良好的视听数据集,包括339个扬声器和128小时的内容。昆曲数据库源于《昆曲艺术宝典》,由台词精心构建,提供明确的注释,包括人物名称、说话人名称、性别信息、发声方式分类,并附有初步的文本翻译。KunquDB为以角色为中心的声学研究和语音相关研究的进步提供了多功能的基础,包括自动说话人验证(ASV)。除了丰富歌剧研究,该数据集还弥合了艺术表达和技术创新之间的差距。本研究首先探讨戏曲中的主动语态,并针对戏曲中的两种不同发声方式:舞台发声(ST)与演唱(S),建构了四个测试实验。实现域自适应方法有效地减轻了这些发声方式变化引起的域失配,同时作为基准仍有进一步改进的空间。
摘要:This work aims to promote Chinese opera research in both musical and speech domains, with a primary focus on overcoming the data limitations. We introduce KunquDB, a relatively large-scale, well-annotated audio-visual dataset comprising 339 speakers and 128 hours of content. Originating from the Kunqu Opera Art Canon (Kunqu yishu dadian), KunquDB is meticulously structured by dialogue lines, providing explicit annotations including character names, speaker names, gender information, vocal manner classifications, and accompanied by preliminary text transcriptions. KunquDB provides a versatile foundation for role-centric acoustic studies and advancements in speech-related research, including Automatic Speaker Verification (ASV). Beyond enriching opera research, this dataset bridges the gap between artistic expression and technological innovation. Pioneering the exploration of ASV in Chinese opera, we construct four test trials considering two distinct vocal manners in opera voices: stage speech (ST) and singing (S). Implementing domain adaptation methods effectively mitigates domain mismatches induced by these vocal manner variations while there is still room for further improvement as a benchmark.

【4】 TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration  Transducer
标题:TDT—KWS:利用令牌和持续时间传感器快速准确地识别关键词
链接:https://arxiv.org/abs/2403.13332
作者:Yu Xi,Hao Li,Baochen Yang,Haoyu Li,Hainan Xu,Kai Yu
备注:Accepted by ICASSP2024
摘要:设计一个高效的关键字识别(KWS)系统,在资源受限的边缘设备上提供卓越的性能一直是一个备受关注的主题。现有的KWS搜索算法通常遵循帧同步方法,其中在每个帧处重复地做出搜索决策,尽管大多数帧是关键字不相关的。在本文中,我们提出了TDT-KWS,它利用令牌和持续时间传感器(TDT)的KWS任务。我们还提出了一种新的KWS特定于任务的解码算法,基于传感器的模型,它支持高效的帧异步关键字搜索流语音场景。通过对公共Hey Snips和自构建的LibriKWS-20数据集进行评估,我们提出的KWS解码算法比传统的ASR解码算法产生更准确的结果。此外,TDT-KWS实现了与RNN-T和传统TDT-ASR系统相当或更好的唤醒字检测性能,同时实现了显著的推理速度提升。此外,实验表明,TDT-KWS是更强大的噪声环境相比,RNN-T KWS。
摘要:Designing an efficient keyword spotting (KWS) system that delivers exceptional performance on resource-constrained edge devices has long been a subject of significant attention. Existing KWS search algorithms typically follow a frame-synchronous approach, where search decisions are made repeatedly at each frame despite the fact that most frames are keyword-irrelevant. In this paper, we propose TDT-KWS, which leverages token-and-duration Transducers (TDT) for KWS tasks. We also propose a novel KWS task-specific decoding algorithm for Transducer-based models, which supports highly effective frame-asynchronous keyword search in streaming speech scenarios. With evaluations conducted on both the public Hey Snips and self-constructed LibriKWS-20 datasets, our proposed KWS-decoding algorithm produces more accurate results than conventional ASR decoding algorithms. Additionally, TDT-KWS achieves on-par or better wake word detection performance than both RNN-T and traditional TDT-ASR systems while achieving significant inference speed-up. Furthermore, experiments show that TDT-KWS is more robust to noisy environments compared to RNN-T KWS.

【5】 UTDUSS: UTokyo-SaruLab System for Interspeech2024 Speech Processing  Using Discrete Speech Unit Challenge
标题:UTDUSS:UTokyo—SaruLab Interspeech 2024语音处理系统使用离散语音单元挑战赛
链接:https://arxiv.org/abs/2403.13720
作者:Wataru Nakata,Kazuki Yamauchi,Dong Yang,Hiroaki Hyodo,Yuki Saito备注:5 pages, 3 figures摘要:我们提出UTDUSS,UTokyo-SaruLab系统提交给Interspeech 2024使用离散语音单元挑战的语音处理。挑战的重点是使用从大型语音语料库中学习的离散语音单元来完成某些任务。我们将UTDUSS系统提交给两个文本到语音的轨道:声码器和声学+声码器。我们的系统采用了神经音频编解码器(NAC)只在语音语料库上进行预训练,这使得学习的编解码器代表了高保真语音重建所需的丰富声学特征。对于声学+声码器音轨,我们训练了一个基于Transformer编码器-解码器的声学模型,该编码器-解码器从文本输入预测预训练的NAC标记。我们描述了构建这些模型的策略,例如数据选择,下采样和超参数调整。我们的系统分别在声码器和声学+声码器音轨中排名第二和第一。
摘要:We present UTDUSS, the UTokyo-SaruLab system submitted to Interspeech2024 Speech Processing Using Discrete Speech Unit Challenge. The challenge focuses on using discrete speech unit learned from large speech corpora for some tasks. We submitted our UTDUSS system to two text-to-speech tracks: Vocoder and Acoustic+Vocoder. Our system incorporates neural audio codec (NAC) pre-trained on only speech corpora, which makes the learned codec represent rich acoustic features that are necessary for high-fidelity speech reconstruction. For the acoustic+vocoder track, we trained an acoustic model based on Transformer encoder-decoder that predicted the pre-trained NAC tokens from text input. We describe our strategies to build these models, such as data selection, downsampling, and hyper-parameter tuning. Our system ranked in second and first for the Vocoder and Acoustic+Vocoder tracks, respectively.

【6】 Recursive Cross-Modal Attention for Multimodal Fusion in Dimensional  Emotion Recognition
标题:多维情感识别中多模态融合的递归跨模态注意力
链接:https://arxiv.org/abs/2403.13659
作者:R. Gnana Praveen,Jahangir Alam
备注:arXiv admin note: substantial text overlap with arXiv:2209.09068; text overlap with arXiv:2203.14779 by other authors
摘要:多模态情感识别最近得到了很多关注,因为它可以利用多种模态(如音频,视觉和文本)的多样和互补关系。大多数最先进的多模态融合方法依赖于循环网络或传统的注意力机制,这些机制不能有效地利用模态的互补性。在本文中,我们专注于三维情感识别的基础上融合的面部,声音和文本形式提取的视频。具体来说,我们提出了一个递归的跨模态注意(RCMA),以有效地捕捉互补的关系,在一个递归的方式跨模态。该模型能够有效地捕捉跨通道的关系,通过计算跨单个通道的交叉注意力权重和其他两个通道的联合表示。为了进一步改善模态间关系,所获得的个体模态的关注特征再次作为输入被馈送到跨模态关注,以细化个体模态的特征表示。除此之外,我们还使用了时间卷积网络(TCN)来捕获各个模态的时间建模(模态内关系)。通过以递归方式部署TCN以及跨模态注意,我们能够有效地捕获跨音频,视觉和文本模态的模态内和模态间关系。在AffWild2数据集的验证集视频上的实验结果表明,我们提出的融合模型能够在2024年野外情感行为分析(ABAW6)竞赛的第六次挑战中实现显著的改进。
摘要:Multi-modal emotion recognition has recently gained a lot of attention since it can leverage diverse and complementary relationships over multiple modalities, such as audio, visual, and text. Most state-of-the-art methods for multimodal fusion rely on recurrent networks or conventional attention mechanisms that do not effectively leverage the complementary nature of the modalities. In this paper, we focus on dimensional emotion recognition based on the fusion of facial, vocal, and text modalities extracted from videos. Specifically, we propose a recursive cross-modal attention (RCMA) to effectively capture the complementary relationships across the modalities in a recursive fashion. The proposed model is able to effectively capture the inter-modal relationships by computing the cross-attention weights across the individual modalities and the joint representation of the other two modalities. To further improve the inter-modal relationships, the obtained attended features of the individual modalities are again fed as input to the cross-modal attention to refine the feature representations of the individual modalities. In addition to that, we have used Temporal convolution networks (TCNs) to capture the temporal modeling (intra-modal relationships) of the individual modalities. By deploying the TCNs as well cross-modal attention in a recursive fashion, we are able to effectively capture both intra- and inter-modal relationships across the audio, visual, and text modalities. Experimental results on validation-set videos from the AffWild2 dataset indicate that our proposed fusion model is able to achieve significant improvement over the baseline for the sixth challenge of Affective Behavior Analysis in-the-Wild 2024 (ABAW6) competition.


【7】 Advanced Long-Content Speech Recognition With Factorized Neural  Transducer
标题:基于因子分解神经传感器的高级长内容语音识别
链接:https://arxiv.org/abs/2403.13423
作者:Xun Gong,Yu Wu,Jinyu Li,Shujie Liu,Rui Zhao,Xie Chen,Yanmin Qian
备注:None
摘要:在本文中,我们提出了两种新的方法,将长内容信息集成到基于非流(简称为LongFNT)和流(简称为SLongFNT)场景的因子分解神经换能器(FNT)架构中。我们首先调查是否长含量transmantine可以改善香草构象转换器(C-T)模型。我们的实验表明,香草C-T模型没有表现出改善的性能时,利用长内容transmittance,可能是由于预测网络的C-T模型不作为一个纯语言模型。相反,FNT显示其潜力,利用长内容的信息,在这里我们提出了LongFNT模型,并探讨文本(LongFNT-Text)和语音(LongFNT-Speech)的长内容信息的影响。所提出的LongFNT-Text和LongFNT-Speech模型进一步相互补充以实现更好的性能,转录历史对模型更有价值。在LibriSpeech和GigaSpeech语料库上评估了LongFNT方法的有效性,并分别获得了相对19%和12%的单词错误率降低。此外,我们扩展的LongFNT模型的流媒体场景,这是命名为SLongFNT,包括SLongFNT-Text和SLongFNT-Speech方法,利用长内容的文本和语音信息。实验结果表明,与FNT基线相比,所提出的SLongFNT模型在LibriSpeech和GigaSpeech上分别实现了相对26%和17%的WER降低,同时保持了良好的延迟。总的来说,我们提出的LongFNT和SLongFNT突出了考虑长内容语音和转录知识对改善非流和流语音识别系统的重要性。
摘要:In this paper, we propose two novel approaches, which integrate long-content information into the factorized neural transducer (FNT) based architecture in both non-streaming (referred to as LongFNT ) and streaming (referred to as SLongFNT ) scenarios. We first investigate whether long-content transcriptions can improve the vanilla conformer transducer (C-T) models. Our experiments indicate that the vanilla C-T models do not exhibit improved performance when utilizing long-content transcriptions, possibly due to the predictor network of C-T models not functioning as a pure language model. Instead, FNT shows its potential in utilizing long-content information, where we propose the LongFNT model and explore the impact of long-content information in both text (LongFNT-Text) and speech (LongFNT-Speech). The proposed LongFNT-Text and LongFNT-Speech models further complement each other to achieve better performance, with transcription history proving more valuable to the model. The effectiveness of our LongFNT approach is evaluated on LibriSpeech and GigaSpeech corpora, and obtains relative 19% and 12% word error rate reduction, respectively. Furthermore, we extend the LongFNT model to the streaming scenario, which is named SLongFNT , consisting of SLongFNT-Text and SLongFNT-Speech approaches to utilize long-content text and speech information. Experiments show that the proposed SLongFNT model achieves relative 26% and 17% WER reduction on LibriSpeech and GigaSpeech respectively while keeping a good latency, compared to the FNT baseline. Overall, our proposed LongFNT and SLongFNT highlight the significance of considering long-content speech and transcription knowledge for improving both non-streaming and streaming speech recognition systems.


【8】 Building speech corpus with diverse voice characteristics for its  prompt-based representation
标题:基于语音特征的语音库的构建
链接:https://arxiv.org/abs/2403.13353
作者:Aya Watanabe,Shinnosuke Takamichi,Yuki Saito,Wataru Nakata,Detai Xin,Hiroshi Saruwatari
备注:Submitted to IEEE/ACM Transactions on Audio, Speech, and Language Processing. arXiv admin note: text overlap with arXiv:2309.13509
摘要:在文本到语音合成中,控制语音特征的能力对于各种应用至关重要。通过利用蓬勃发展的基于文本的生成技术,应该可以增强对语音特征的细微控制。虽然以前的研究已经探索了基于语音特征的操纵,但大多数研究都使用预先录制的语音,这限制了可用语音特征的多样性。因此,我们的目标是通过创建一个新的语料库,并开发一个模型,在文本到语音合成的语音特征的基于人工智能的操作,以解决这一差距,促进更广泛的语音特征。具体来说,我们提出了一种方法来建立一个相当大的语料库配对语音特征描述与相应的语音样本。这涉及自动从互联网收集与语音相关的语音数据,确保其质量,并使用众包手动注释。我们用日语数据实现了该方法,并对结果进行了分析,以验证其有效性。随后,我们提出了一种基于对比学习方法的模型的构建方法,从语音特征描述中检索语音。我们训练的模型,不仅使用保守的对比学习,但也特征预测学习,预测语音特征对应的定量语音特征。我们通过实验来评估模型的性能与我们上面构建的语料库。
摘要:In text-to-speech synthesis, the ability to control voice characteristics is vital for various applications. By leveraging thriving text prompt-based generation techniques, it should be possible to enhance the nuanced control of voice characteristics. While previous research has explored the prompt-based manipulation of voice characteristics, most studies have used pre-recorded speech, which limits the diversity of voice characteristics available. Thus, we aim to address this gap by creating a novel corpus and developing a model for prompt-based manipulation of voice characteristics in text-to-speech synthesis, facilitating a broader range of voice characteristics. Specifically, we propose a method to build a sizable corpus pairing voice characteristics descriptions with corresponding speech samples. This involves automatically gathering voice-related speech data from the Internet, ensuring its quality, and manually annotating it using crowdsourcing. We implement this method with Japanese language data and analyze the results to validate its effectiveness. Subsequently, we propose a construction method of the model to retrieve speech from voice characteristics descriptions based on a contrastive learning method. We train the model using not only conservative contrastive learning but also feature prediction learning to predict quantitative speech features corresponding to voice characteristics. We evaluate the model performance via experiments with the corpus we constructed above.


【9】 Onset and offset weighted loss function for sound event detection
标题:用于声音事件检测的起始和偏移加权损失函数
链接:https://arxiv.org/abs/2403.13254
作者:Tao Song
摘要:在典型的声音事件检测(SED)系统中,在帧级检测声音事件的存在,并且将检测到相同事件的连续帧组合为一个声音事件。中值滤波器被应用作为后处理步骤,以尽可能地去除检测误差。然而,发生在声音事件的起始和偏移周围的检测误差超出了中值滤波器的能力。为了解决这个问题,本文提出了一种起始和偏移加权二进制交叉熵(OWBCE)损失函数,它训练DNN模型在(a)起始和偏移周围的帧上更加鲁棒。在DCASE 2022任务4的背景下进行实验。结果表明,OWBCE优于BCE时,考虑不同的模型。对于基本CRNN,OWBCE可以实现事件F1的6.43%,PSDS 1的1.96%和PSDS 2的2.43%的相对改善。
摘要:In a typical sound event detection (SED) system, the existence of a sound event is detected at a frame level, and consecutive frames with the same event detected are combined as one sound event. The median filter is applied as a post-processing step to remove detection errors as much as possible. However, detection errors occurring around the onset and offset of a sound event are beyond the capacity of the median filter. To address this issue, an onset and offset weighted binary cross-entropy (OWBCE) loss function is proposed in this paper, which trains the DNN model to be more robust on frames around (a) onsets and offsets. Experiments are carried out in the context of DCASE 2022 task 4. Results show that OWBCE outperforms BCE when different models are considered. For a basic CRNN, relative improvements of 6.43% in event-F1, 1.96% in PSDS1, and 2.43% in PSDS2 can be achieved by OWBCE.


【10】 Document Author Classification Using Parsed Language Structure
标题:基于解析语言结构的文档作者分类
链接:https://arxiv.org/abs/2403.13253
作者:Todd K Moon,Jacob H. Gunther
备注:None
摘要:多年来,人们一直对基于文本的统计特性来检测文本的作者身份感兴趣,例如通过使用非上下文单词的出现率。在以前的工作中,这些技术已经被用来,例如,确定所有的联邦党人论文的作者。这些方法在更现代的时代可能有助于检测假冒或人工智能作者。统计自然语言分析器的进步引入了使用语法结构来检测作者的可能性。在本文中,我们探索了一种新的可能性,使用统计自然语言解析器提取的语法结构信息检测作者。本文提供了一个概念的证明,测试作者分类的基础上的一组“证明文本”的语法结构,联邦党人文件和Sanditon已作为测试案例,在以前的作者身份检测研究。研究了从统计自然语言解析器中提取的几个特征:任何级别的某个深度的所有子树;解析树中某个深度的有根子树、词性和按级别的词性。发现将特征投影到较低维空间中是有帮助的。对这些文档的统计实验表明,来自统计分析器的信息实际上可以帮助区分作者。
摘要:Over the years there has been ongoing interest in detecting authorship of a text based on statistical properties of the text, such as by using occurrence rates of noncontextual words. In previous work, these techniques have been used, for example, to determine authorship of all of \emph{The Federalist Papers}. Such methods may be useful in more modern times to detect fake or AI authorship. Progress in statistical natural language parsers introduces the possibility of using grammatical structure to detect authorship. In this paper we explore a new possibility for detecting authorship using grammatical structural information extracted using a statistical natural language parser. This paper provides a proof of concept, testing author classification based on grammatical structure on a set of "proof texts," The Federalist Papers and Sanditon which have been as test cases in previous authorship detection studies. Several features extracted from the statistical natural language parser were explored: all subtrees of some depth from any level; rooted subtrees of some depth, part of speech, and part of speech by level in the parse tree. It was found to be helpful to project the features into a lower dimensional space. Statistical experiments on these documents demonstrate that information from a statistical parser can, in fact, assist in distinguishing authors.


【11】 Frequency-aware convolution for sound event detection
标题:用于声音事件检测的频率感知卷积
链接:https://arxiv.org/abs/2403.13252
作者:Tao Song
摘要:在声音事件检测(SED)中,卷积神经网络(CNN)被广泛用于从输入频谱图中提取时频模式。然而,由CNN提取的特征可能对时频模式沿频率轴的偏移不敏感。为了解决这个问题,已经提出了频率动态卷积(FDY),其将不同的内核应用于不同的频率分量。与vannila CNN相比,FDY需要多几倍的参数。在本文中,一个更有效的解决方案命名为频率感知卷积(FAC)。在FAC中,频率位置信息被编码在矢量中并被添加到输入频谱图。为了匹配输入的幅度,编码矢量自适应缩放和通道无关。在DCASE 2022任务4的背景下进行了实验,结果表明,FAC可以实现与FDY的性能相当,只有515个额外的参数,而FDY需要802万个额外的参数。消融研究表明,自适应缩放的编码矢量和信道无关的是至关重要的FAC的性能。
摘要:In sound event detection (SED), convolution neural networks (CNNs) are widely used to extract time-frequency patterns from the input spectrogram. However, features extracted by CNN can be insensitive to the shift of time-frequency patterns along the frequency axis. To address this issue, frequency dynamic convolution (FDY) has been proposed, which applies different kernels to different frequency components. Compared to the vannila CNN, FDY requires several times more parameters. In this paper, a more efficient solution named frequency-aware convolution (FAC) is proposed. In FAC, frequency-positional information is encoded in a vector and added to the input spectrogram. To match the amplitude of input, the encoding vector is scaled adaptively and channel-independently. Experiments are carried out in the context of DCASE 2022 task 4, and the results demonstrate that FAC can achieve comparable performance to that of FDY with only 515 additional parameters, while FDY requires 8.02 million additional parameters. The ablation study shows that scaling the encoding vector adaptively and channel-independently is critical to the performance of FAC.

【12】 Listenable Maps for Audio Classifiers
标题:音频分类器的可监听映射
链接:https://arxiv.org/abs/2403.13086
作者:Francesco Paissan,Mirco Ravanelli,Cem Subakan
摘要:尽管深度学习模型在不同任务中的表现令人印象深刻,但其复杂性给解释带来了挑战。这一挑战对于音频信号尤其明显,其中传达解释变得固有地困难。为了解决这个问题,我们引入了可听地图音频分类器(L-MAC),一个事后解释方法,产生忠实和可解释的解释。L-MAC利用预训练分类器之上的解码器来生成突出输入音频的相关部分的二进制掩码。我们训练具有特殊损失的解码器,该特殊损失最大化分类器对音频的掩蔽部分的决策的置信度,同时最小化掩蔽部分的模型输出的概率。对域内和域外数据的定量评价表明,L-MAC始终比几种基于梯度和掩蔽的方法产生更忠实的解释。此外,用户研究证实,平均而言,用户更喜欢所提出的技术产生的解释。
摘要:Despite the impressive performance of deep learning models across diverse tasks, their complexity poses challenges for interpretation. This challenge is particularly evident for audio signals, where conveying interpretations becomes inherently difficult. To address this issue, we introduce Listenable Maps for Audio Classifiers (L-MAC), a posthoc interpretation method that generates faithful and listenable interpretations. L-MAC utilizes a decoder on top of a pretrained classifier to generate binary masks that highlight relevant portions of the input audio. We train the decoder with a special loss that maximizes the confidence of the classifier decision on the masked-in portion of the audio while minimizing the probability of model output for the masked-out portion. Quantitative evaluations on both in-domain and out-of-domain data demonstrate that L-MAC consistently produces more faithful interpretations than several gradient and masking-based methodologies. Furthermore, a user study confirms that, on average, users prefer the interpretations generated by the proposed technique.

机器翻译由腾讯交互翻译提供,仅供参考