今天跟大家分享一篇语音相关的论文合集:cs.SD语音6篇,eess.AS音频处理7篇。
cs.SD语音
【1】 EEG2Mel: Reconstructing Sound from Brain Responses to Music
标题: EEG2Mel:从大脑对音乐的反应中重建声音
链接:https://arxiv.org/abs/2207.13845
作者:Adolfo G. Ramirez-Aristizabal,Chris Kello备注:5 figures, 2 tables, listening examples and code provided摘要:通过对被试在记录脑电信号的同时所呈现的歌曲名称和图像类别的分类,从听觉和视觉刺激的脑反应中提取信息已经取得了成功。以重构听觉刺激的形式提取信息也取得了一定的成功,但在这里我们改进了以前的方法,通过重构音乐刺激,使其足够好地被独立地感知和识别。此外,针对EEG记录的每个相应的一秒窗口,在时间对准的音乐刺激谱上训练深度学习模型,与现有的研究相比,这大大减少了所需的特征提取步骤。NMED-Tempo和NMED-Hindi数据集被用来训练和验证卷积神经网络测试了原始电压对功率谱输入和线性对mel谱图输出的功效,并将所有输入和输出转换为2D图像,通过训练分类器评估重建谱图的质量,mel谱图的准确率为81%,线性谱图的准确率为72(10%的机会准确率)。最后,在两种选择的匹配—样本任务中,听觉音乐刺激的重建被听者以85%的成功率(50%的机会)辨别。摘要:Information retrieval from brain responses to auditory and visual stimuli has shown success through classification of song names and image classes presented to participants while recording EEG signals. Information retrieval in the form of reconstructing auditory stimuli has also shown some success, but here we improve on previous methods by reconstructing music stimuli well enough to be perceived and identified independently. Furthermore, deep learning models were trained on time-aligned music stimuli spectrum for each corresponding one-second window of EEG recording, which greatly reduces feature extraction steps needed when compared to prior studies. The NMED-Tempo and NMED-Hindi datasets of participants passively listening to full length songs were used to train and validate Convolutional Neural Network (CNN) regressors. The efficacy of raw voltage versus power spectrum inputs and linear versus mel spectrogram outputs were tested, and all inputs and outputs were converted into 2D images. The quality of reconstructed spectrograms was assessed by training classifiers which showed 81% accuracy for mel-spectrograms and 72% for linear spectrograms (10% chance accuracy). Lastly, reconstructions of auditory music stimuli were discriminated by listeners at an 85% success rate (50% chance) in a two-alternative match-to-sample task.
【2】 Deep Learning-Based Acoustic Mosquito Detection in Noisy Conditions Using Trainable Kernels and Augmentations
标题: 噪声环境下基于深度学习的可训练核和增强声学蚊子检测
链接:https://arxiv.org/abs/2207.13843
作者:Devesh Khandelwal,Sean Campos,Shwetha Nagaraj,Fred Nugen,Alberto Todeschini摘要:本文提出了一种新的方法,我们展示了一种独特的方法,通过将预处理技术融合到深度学习模型中来提高音频机器学习方法的有效性。我们的解决方案通过训练优化超参数,而不是代价高昂的随机搜索,来加快训练和推理性能,从而从音频信号中构建可靠的蚊子检测器。本文中的实验和结果是ACM 2022挑战赛MOS C提交的一部分。我们的结果在未公布的测试集上比已公布的基线高出212%。我们相信这是最好的真实测试集之一。建立在嘈杂条件下提供可靠蚊子检测的鲁棒生物声学系统的世界实例。摘要:In this paper, we demonstrate a unique recipe to enhance the effectiveness of audio machine learning approaches by fusing pre-processing techniques into a deep learning model. Our solution accelerates training and inference performance by optimizing hyper-parameters through training instead of costly random searches to build a reliable mosquito detector from audio signals. The experiments and the results presented here are part of the MOS C submission of the ACM 2022 challenge. Our results outperform the published baseline by 212% on the unpublished test set. We believe that this is one of the best real-world examples of building a robust bio-acoustic system that provides reliable mosquito detection in noisy conditions.
【3】 SoundChoice: Grapheme-to-Phoneme Models with Semantic Disambiguation
标题: 声音选择:具有语义消歧功能的字形—音素模型
链接:https://arxiv.org/abs/2207.13703
作者:Artem Ploujnikov,Mirco Ravanelli备注:5 pages, submitted to INTERSPEECH 2022摘要:端到端语音合成模型直接将输入字符转换为音频表示尽管它们的性能令人印象深刻,但是这样的模型难以消除相同拼写的单词的发音的歧义。提出了一种基于SoundChoice(G2P)模型的语音合成方法,提出了一种新的G2P体系结构,该体系结构处理整个句子而不是在词的层次上进行操作,该体系结构利用了加权的同形异义词损失(可改善歧义消除),利用课程学习(逐渐从单词级切换到句子级G2P),并集成了来自BERT的单词嵌入此外,该模型继承了语音识别领域的最佳实践,包括多任务学习和连接时态分类(CTC)和嵌入式语言模型束搜索,使用LibriSpeech和维基百科的数据,SoundChoice在整句转录上实现了2.65%的音素错误率(PER)。索引术语字形到音素,语音合成,文本到语音,语音学,发音,消歧。摘要:End-to-end speech synthesis models directly convert the input characters into an audio representation (e.g., spectrograms). Despite their impressive performance, such models have difficulty disambiguating the pronunciations of identically spelled words. To mitigate this issue, a separate Grapheme-to-Phoneme (G2P) model can be employed to convert the characters into phonemes before synthesizing the audio. This paper proposes SoundChoice, a novel G2P architecture that processes entire sentences rather than operating at the word level. The proposed architecture takes advantage of a weighted homograph loss (that improves disambiguation), exploits curriculum learning (that gradually switches from word-level to sentence-level G2P), and integrates word embeddings from BERT (for further performance improvement). Moreover, the model inherits the best practices in speech recognition, including multi-task learning with Connectionist Temporal Classification (CTC) and beam search with an embedded language model. As a result, SoundChoice achieves a Phoneme Error Rate (PER) of 2.65% on whole-sentence transcription using data from LibriSpeech and Wikipedia. Index Terms grapheme-to-phoneme, speech synthesis, text-tospeech, phonetics, pronunciation, disambiguation.
【4】 Extending RNN-T-based speech recognition systems with emotion and language classification
标题: 基于情感和语言分类的RNN-T语音识别系统的扩展
链接:https://arxiv.org/abs/2207.13965
作者:Zvi Kons,Hagai Aronowitz,Edmilson Morais,Matheus Damasceno,Hong-Kwang Kuo,Samuel Thomas,George Saon备注:Accepted for publication in Interspeech 2022摘要:语音转录,情感识别,和语言识别通常被认为是三个不同的任务,每一个任务都需要一个不同的模型,具有不同的结构和训练过程,我们提出使用一个递归神经网络传感器基于RNN-T的语音到文本(STT)系统作为可用于情感识别和语言识别以及用于语音识别的公共组件。我们的工作通过最小的改动扩展了用于情感分类的STT系统,并在IEMOCAP和MELD数据集上显示了成功的结果。此外,我们证明了通过向RNN-T模块添加一个轻量级组件,它也可以用于语言识别。在我们的评估中,新的分类器在NIST-LRE-07数据集上显示了最先进的准确性。摘要:Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different architecture and training process. We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition. Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets. In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification. In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset.
【5】 A Unifying View on Blind Source Separation of Convolutive Mixtures based on Independent Component Analysis
标题: 基于独立分量分析的卷积混合信号盲分离的统一观点
链接:https://arxiv.org/abs/2207.13934
作者:Andreas Brendel,Thomas Haubner,Walter Kellermann摘要:在许多日常生活场景中,在封闭空间中记录的声源只能在有其他干扰源的情况下才能被观察到,因此,卷积盲源分离(BSS)是音频信号处理中的一个核心问题,基于独立分量分析(ICA)的方法(国际化学品协会)因为它们只需要很少的和弱的假设,并且允许关于原始源信号和声传播的盲目性当前使用的大多数算法属于以下三个家族中的一个:频域独立分量分析(FD-ICA),独立向量分析(IVA)和卷积混合物的三N独立分量分析ICA、FD-ICA和IVA之间的关系由于它们的构造而变得明显,但与TRINICON的关系还没有很好地建立起来,本文通过提供这些算法的共同构建块和它们的差异的深入处理来填补这一空白,从而为所有考虑的算法提供公共框架。摘要:In many daily-life scenarios, acoustic sources recorded in an enclosure can only be observed with other interfering sources. Hence, convolutive Blind Source Separation (BSS) is a central problem in audio signal processing. Methods based on Independent Component Analysis (ICA) are especially important in this field as they require only few and weak assumptions and allow for blindness regarding the original source signals and the acoustic propagation path. Most of the currently used algorithms belong to one of the following three families: Frequency Domain ICA (FD-ICA), Independent Vector Analysis (IVA), and TRIple-N Independent component analysis for CONvolutive mixtures (TRINICON). While the relation between ICA, FD-ICA and IVA becomes apparent due to their construction, the relation to TRINICON is not well established yet. This paper fills this gap by providing an in-depth treatment of the common building blocks of these algorithms and their differences, and thus provides a common framework for all considered algorithms.
【6】 Utterance-by-utterance overlap-aware neural diarization with Graph-PIT
标题: 基于Graph-PIT的逐话语重叠感知神经元日记化
链接:https://arxiv.org/abs/2207.13888
作者:Keisuke Kinoshita,Thilo von Neumann,Marc Delcroix,Christoph Boeddeker,Reinhold Haeb-Umbach备注:Accepted to Interspeech 2022 (5 pages, 1 figure)摘要:最近的说话人日记化研究表明,端到端神经日记化的整合这种方法首先将观察到的信号划分为固定长度的段,然后基于EEND模块执行{\i段级}局部二值化,并通过聚类的方法将分割结果合并为一个最终的全局Diarization结果.由于现有的EEND不能处理大量的说话人,我们认为这种涉及分段的方法具有几个问题;例如,它不可避免地面临这样一个困境:较大的分段大小既增加了可用于提高性能的上下文,又增加了本地EEND模块要处理的说话者数量.为了解决这个问题,本文提出了一种新的框架,但它仍然可以处理包含许多说话者和大量重叠语音的挑战性数据。所提出的方法可以采用整个会议来进行推断并且执行根据说话者对话语活动进行聚类的{\itutterance-by-utterance}日记化。为此,我们利用最近提出的一种称为Graph-PIT的神经网络训练方案来进行神经源分离。Like数据和CALLHOME数据的实验结果表明了该方法的优越性。摘要:Recent speaker diarization studies showed that integration of end-to-end neural diarization (EEND) and clustering-based diarization is a promising approach for achieving state-of-the-art performance on various tasks. Such an approach first divides an observed signal into fixed-length segments, then performs {\it segment-level} local diarization based on an EEND module, and merges the segment-level results via clustering to form a final global diarization result. The segmentation is done to limit the number of speakers in each segment since the current EEND cannot handle a large number of speakers. In this paper, we argue that such an approach involving the segmentation has several issues; for example, it inevitably faces a dilemma that larger segment sizes increase both the context available for enhancing the performance and the number of speakers for the local EEND module to handle. To resolve such a problem, this paper proposes a novel framework that performs diarization without segmentation. However, it can still handle challenging data containing many speakers and a significant amount of overlapping speech. The proposed method can take an entire meeting for inference and perform {\it utterance-by-utterance} diarization that clusters utterance activities in terms of speakers. To this end, we leverage a neural network training scheme called Graph-PIT proposed recently for neural source separation. Experiments with simulated active-meeting-like data and CALLHOME data show the superiority of the proposed approach over the conventional methods.
【1】 Dialogue Enhancement and Listening Effort in Broadcast Audio: A Multimodal Evaluation
标题: 广播音频中的对话增强和收听努力:多模态评价
链接:https://arxiv.org/abs/2207.14240
作者:Matteo Torcoli,Thomas Robotham,Emanuël A. P. Habets备注:Paper accepted to 14th International Conference on Quality of Multimedia Experience (QoMEX), Lippstadt, Germany, 2022摘要:对话增强(DE)在广播中起着至关重要的作用,使得前景语音和背景音乐之间的相对水平和效果能够个性化。DE已经被证明改善了体验质量、可理解性和自我报告的倾听努力从听力学研究中已知的LE的生理指标是瞳孔大小。瞳孔大小与眼视功能的关系通常使用人工语句和广播内容中没有遇到的背景噪声进行研究,本工作以包括瞳孔大小的多模态方式评估眼视功能对眼视功能的影响(由VR耳机跟踪)和来自电视的真实世界音频摘录。在理想的收听条件下,28名听力正常的参与者听了30段音频片段,这些音频片段以随机顺序呈现,并通过改变前景和背景音频之间的相对电平的条件进行处理。这些条件之一采用了最近提出的源分离系统来衰减给定原始混合作为唯一输入的背景。在收听每个摘录之后,被试被要求复述所听到的句子并自我报告LE。分析平均瞳孔扩张和峰值瞳孔扩张,并与自我报告和单词回忆率进行比较。多模式评估显示LE随着背景水平的降低而降低的趋势一致。DE,也当由源分离启用时,显著减小瞳孔大小和自报告LE。这突出了用户端个性化功能的优势。摘要:Dialogue enhancement (DE) plays a vital role in broadcasting, enabling the personalization of the relative level between foreground speech and background music and effects. DE has been shown to improve the quality of experience, intelligibility, and self-reported listening effort (LE). A physiological indicator of LE known from audiology studies is pupil size. The relation between pupil size and LE is typically studied using artificial sentences and background noises not encountered in broadcast content. This work evaluates the effect of DE on LE in a multimodal manner that includes pupil size (tracked by a VR headset) and real-world audio excerpts from TV. Under ideal listening conditions, 28 normal-hearing participants listened to 30 audio excerpts presented in random order and processed by conditions varying the relative level between foreground and background audio. One of these conditions employed a recently proposed source separation system to attenuate the background given the original mixture as the sole input. After listening to each excerpt, subjects were asked to repeat the heard sentence and self-report the LE. Mean pupil dilation and peak pupil dilation were analyzed and compared with the self-report and the word recall rate. The multimodal evaluation shows a consistent trend of decreasing LE along with decreasing background level. DE, also when enabled by source separation, significantly reduces the pupil size as well as the self-reported LE. This highlights the benefit of personalization functionalities at the user's end.
【2】 Extending RNN-T-based speech recognition systems with emotion and language classification
标题: 基于情感和语言分类的RNN-T语音识别系统的扩展
链接:https://arxiv.org/abs/2207.13965
* 与cs.SD语音【4】为同一篇
作者:Zvi Kons,Hagai Aronowitz,Edmilson Morais,Matheus Damasceno,Hong-Kwang Kuo,Samuel Thomas,George Saon备注:Accepted for publication in Interspeech 2022摘要:语音转录,情感识别,和语言识别通常被认为是三个不同的任务,每一个任务都需要一个不同的模型,具有不同的结构和训练过程,我们提出使用一个递归神经网络传感器基于RNN-T的语音到文本(STT)系统作为可用于情感识别和语言识别以及用于语音识别的公共组件。我们的工作通过最小的改动扩展了用于情感分类的STT系统,并在IEMOCAP和MELD数据集上显示了成功的结果。此外,我们证明了通过向RNN-T模块添加一个轻量级组件,它也可以用于语言识别。在我们的评估中,新的分类器在NIST-LRE-07数据集上显示了最先进的准确性。摘要:Speech transcription, emotion recognition, and language identification are usually considered to be three different tasks. Each one requires a different model with a different architecture and training process. We propose using a recurrent neural network transducer (RNN-T)-based speech-to-text (STT) system as a common component that can be used for emotion recognition and language identification as well as for speech recognition. Our work extends the STT system for emotion classification through minimal changes, and shows successful results on the IEMOCAP and MELD datasets. In addition, we demonstrate that by adding a lightweight component to the RNN-T module, it can also be used for language identification. In our evaluations, this new classifier demonstrates state-of-the-art accuracy for the NIST-LRE-07 dataset.
【3】 A Unifying View on Blind Source Separation of Convolutive Mixtures based on Independent Component Analysis
标题: 基于独立分量分析的卷积混合信号盲分离的统一观点
链接:https://arxiv.org/abs/2207.13934
* 与cs.SD语音【5】为同一篇
作者:Andreas Brendel,Thomas Haubner,Walter Kellermann摘要:在许多日常生活场景中,在封闭空间中记录的声源只能在有其他干扰源的情况下才能被观察到,因此,卷积盲源分离(BSS)是音频信号处理中的一个核心问题,基于独立分量分析(ICA)的方法(国际化学品协会)因为它们只需要很少的和弱的假设,并且允许关于原始源信号和声传播的盲目性当前使用的大多数算法属于以下三个家族中的一个:频域独立分量分析(FD-ICA),独立向量分析(IVA)和卷积混合物的三N独立分量分析ICA、FD-ICA和IVA之间的关系由于它们的构造而变得明显,但与TRINICON的关系还没有很好地建立起来,本文通过提供这些算法的共同构建块和它们的差异的深入处理来填补这一空白,从而为所有考虑的算法提供公共框架。摘要:In many daily-life scenarios, acoustic sources recorded in an enclosure can only be observed with other interfering sources. Hence, convolutive Blind Source Separation (BSS) is a central problem in audio signal processing. Methods based on Independent Component Analysis (ICA) are especially important in this field as they require only few and weak assumptions and allow for blindness regarding the original source signals and the acoustic propagation path. Most of the currently used algorithms belong to one of the following three families: Frequency Domain ICA (FD-ICA), Independent Vector Analysis (IVA), and TRIple-N Independent component analysis for CONvolutive mixtures (TRINICON). While the relation between ICA, FD-ICA and IVA becomes apparent due to their construction, the relation to TRINICON is not well established yet. This paper fills this gap by providing an in-depth treatment of the common building blocks of these algorithms and their differences, and thus provides a common framework for all considered algorithms.
【4】 Utterance-by-utterance overlap-aware neural diarization with Graph-PIT
标题: 基于Graph-PIT的逐话语重叠感知神经元日记化
链接:https://arxiv.org/abs/2207.13888
* 与cs.SD语音【6】为同一篇
作者:Keisuke Kinoshita,Thilo von Neumann,Marc Delcroix,Christoph Boeddeker,Reinhold Haeb-Umbach备注:Accepted to Interspeech 2022 (5 pages, 1 figure)摘要:最近的说话人日记化研究表明,端到端神经日记化的整合这种方法首先将观察到的信号划分为固定长度的段,然后基于EEND模块执行{\i段级}局部二值化,并通过聚类的方法将分割结果合并为一个最终的全局Diarization结果.由于现有的EEND不能处理大量的说话人,我们认为这种涉及分段的方法具有几个问题;例如,它不可避免地面临这样一个困境:较大的分段大小既增加了可用于提高性能的上下文,又增加了本地EEND模块要处理的说话者数量.为了解决这个问题,本文提出了一种新的框架,但它仍然可以处理包含许多说话者和大量重叠语音的挑战性数据。所提出的方法可以采用整个会议来进行推断并且执行根据说话者对话语活动进行聚类的{\itutterance-by-utterance}日记化。为此,我们利用最近提出的一种称为Graph-PIT的神经网络训练方案来进行神经源分离。Like数据和CALLHOME数据的实验结果表明了该方法的优越性。摘要:Recent speaker diarization studies showed that integration of end-to-end neural diarization (EEND) and clustering-based diarization is a promising approach for achieving state-of-the-art performance on various tasks. Such an approach first divides an observed signal into fixed-length segments, then performs {\it segment-level} local diarization based on an EEND module, and merges the segment-level results via clustering to form a final global diarization result. The segmentation is done to limit the number of speakers in each segment since the current EEND cannot handle a large number of speakers. In this paper, we argue that such an approach involving the segmentation has several issues; for example, it inevitably faces a dilemma that larger segment sizes increase both the context available for enhancing the performance and the number of speakers for the local EEND module to handle. To resolve such a problem, this paper proposes a novel framework that performs diarization without segmentation. However, it can still handle challenging data containing many speakers and a significant amount of overlapping speech. The proposed method can take an entire meeting for inference and perform {\it utterance-by-utterance} diarization that clusters utterance activities in terms of speakers. To this end, we leverage a neural network training scheme called Graph-PIT proposed recently for neural source separation. Experiments with simulated active-meeting-like data and CALLHOME data show the superiority of the proposed approach over the conventional methods.
【5】 EEG2Mel: Reconstructing Sound from Brain Responses to Music
标题: EEG2Mel:从大脑对音乐的反应中重建声音
链接:https://arxiv.org/abs/2207.13845
* 与cs.SD语音【1】为同一篇
作者:Adolfo G. Ramirez-Aristizabal,Chris Kello备注:5 figures, 2 tables, listening examples and code provided摘要:通过对被试在记录脑电信号的同时所呈现的歌曲名称和图像类别的分类,从听觉和视觉刺激的脑反应中提取信息已经取得了成功。以重构听觉刺激的形式提取信息也取得了一定的成功,但在这里我们改进了以前的方法,通过重构音乐刺激,使其足够好地被独立地感知和识别。此外,针对EEG记录的每个相应的一秒窗口,在时间对准的音乐刺激谱上训练深度学习模型,与现有的研究相比,这大大减少了所需的特征提取步骤。NMED-Tempo和NMED-Hindi数据集被用来训练和验证卷积神经网络测试了原始电压对功率谱输入和线性对mel谱图输出的功效,并将所有输入和输出转换为2D图像,通过训练分类器评估重建谱图的质量,mel谱图的准确率为81%,线性谱图的准确率为72(10%的机会准确率)。最后,在两种选择的匹配—样本任务中,听觉音乐刺激的重建被听者以85%的成功率(50%的机会)辨别。摘要:Information retrieval from brain responses to auditory and visual stimuli has shown success through classification of song names and image classes presented to participants while recording EEG signals. Information retrieval in the form of reconstructing auditory stimuli has also shown some success, but here we improve on previous methods by reconstructing music stimuli well enough to be perceived and identified independently. Furthermore, deep learning models were trained on time-aligned music stimuli spectrum for each corresponding one-second window of EEG recording, which greatly reduces feature extraction steps needed when compared to prior studies. The NMED-Tempo and NMED-Hindi datasets of participants passively listening to full length songs were used to train and validate Convolutional Neural Network (CNN) regressors. The efficacy of raw voltage versus power spectrum inputs and linear versus mel spectrogram outputs were tested, and all inputs and outputs were converted into 2D images. The quality of reconstructed spectrograms was assessed by training classifiers which showed 81% accuracy for mel-spectrograms and 72% for linear spectrograms (10% chance accuracy). Lastly, reconstructions of auditory music stimuli were discriminated by listeners at an 85% success rate (50% chance) in a two-alternative match-to-sample task.
【6】 Deep Learning-Based Acoustic Mosquito Detection in Noisy Conditions Using Trainable Kernels and Augmentations
标题: 噪声环境下基于深度学习的可训练核和增强声学蚊子检测
链接:https://arxiv.org/abs/2207.13843
* 与cs.SD语音【2】为同一篇
作者:Devesh Khandelwal,Sean Campos,Shwetha Nagaraj,Fred Nugen,Alberto Todeschini摘要:本文提出了一种新的方法,我们展示了一种独特的方法,通过将预处理技术融合到深度学习模型中来提高音频机器学习方法的有效性。我们的解决方案通过训练优化超参数,而不是代价高昂的随机搜索,来加快训练和推理性能,从而从音频信号中构建可靠的蚊子检测器。本文中的实验和结果是ACM 2022挑战赛MOS C提交的一部分。我们的结果在未公布的测试集上比已公布的基线高出212%。我们相信这是最好的真实测试集之一。建立在嘈杂条件下提供可靠蚊子检测的鲁棒生物声学系统的世界实例。摘要:In this paper, we demonstrate a unique recipe to enhance the effectiveness of audio machine learning approaches by fusing pre-processing techniques into a deep learning model. Our solution accelerates training and inference performance by optimizing hyper-parameters through training instead of costly random searches to build a reliable mosquito detector from audio signals. The experiments and the results presented here are part of the MOS C submission of the ACM 2022 challenge. Our results outperform the published baseline by 212% on the unpublished test set. We believe that this is one of the best real-world examples of building a robust bio-acoustic system that provides reliable mosquito detection in noisy conditions.
【7】 SoundChoice: Grapheme-to-Phoneme Models with Semantic Disambiguation
标题: 声音选择:具有语义消歧功能的字形—音素模型
链接:https://arxiv.org/abs/2207.13703
* 与cs.SD语音【3】为同一篇
作者:Artem Ploujnikov,Mirco Ravanelli备注:5 pages, submitted to INTERSPEECH 2022摘要:端到端语音合成模型直接将输入字符转换为音频表示尽管它们的性能令人印象深刻,但是这样的模型难以消除相同拼写的单词的发音的歧义。提出了一种基于SoundChoice(G2P)模型的语音合成方法,提出了一种新的G2P体系结构,该体系结构处理整个句子而不是在词的层次上进行操作,该体系结构利用了加权的同形异义词损失(可改善歧义消除),利用课程学习(逐渐从单词级切换到句子级G2P),并集成了来自BERT的单词嵌入此外,该模型继承了语音识别领域的最佳实践,包括多任务学习和连接时态分类(CTC)和嵌入式语言模型束搜索,使用LibriSpeech和维基百科的数据,SoundChoice在整句转录上实现了2.65%的音素错误率(PER)。索引术语字形到音素,语音合成,文本到语音,语音学,发音,消歧。摘要:End-to-end speech synthesis models directly convert the input characters into an audio representation (e.g., spectrograms). Despite their impressive performance, such models have difficulty disambiguating the pronunciations of identically spelled words. To mitigate this issue, a separate Grapheme-to-Phoneme (G2P) model can be employed to convert the characters into phonemes before synthesizing the audio. This paper proposes SoundChoice, a novel G2P architecture that processes entire sentences rather than operating at the word level. The proposed architecture takes advantage of a weighted homograph loss (that improves disambiguation), exploits curriculum learning (that gradually switches from word-level to sentence-level G2P), and integrates word embeddings from BERT (for further performance improvement). Moreover, the model inherits the best practices in speech recognition, including multi-task learning with Connectionist Temporal Classification (CTC) and beam search with an embedded language model. As a result, SoundChoice achieves a Phoneme Error Rate (PER) of 2.65% on whole-sentence transcription using data from LibriSpeech and Wikipedia. Index Terms grapheme-to-phoneme, speech synthesis, text-tospeech, phonetics, pronunciation, disambiguation.
机器翻译,仅供参考