今日论文合集:cs.SD语音5篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
标题: 希腊播客文集:具有弱监督数据的低资源语言的竞争语音模型
作者:Georgios Paraskevopoulos,Chara Tsoukala,Athanasios Katsamanis,Vassilis Katsouros
备注:To be presented at Interspeech 2024
链接:点击下载PDF文件
摘要:为数字表示有限的语言开发语音技术带来了重大挑战,主要是由于缺乏可用数据。在大型数据密集型模型时代,这个问题更加严重。最近的研究强调了利用薄弱的监督来扩大现有数据库的潜力。在这项研究中,我们编译了一个800小时的现代希腊语语料库,从播客和使用Whisper large-v3生成银transmits。该语料库被用来微调我们的模型,旨在评估这种方法在提高ASR性能的有效性。我们的分析涵盖了16个不同的播客领域,同时对现代希腊语的既定数据集进行了评估。研究结果表明WER的持续改善,与数据量和模型大小的增加相关。我们的研究证实,组装大型,弱监督语料库作为一个具有成本效益的战略,推进语音技术在资源不足的语言。摘要:The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models. Recent research has underscored the potential of leveraging weak supervision to augment the pool of available data. In this study, we compile an 800-hour corpus of Modern Greek from podcasts and employ Whisper large-v3 to generate silver transcriptions. This corpus is utilized to fine-tune our models, aiming to assess the efficacy of this approach in enhancing ASR performance. Our analysis spans 16 distinct podcast domains, alongside evaluations on established datasets for Modern Greek. The findings indicate consistent WER improvements, correlating with increases in both data volume and model size. Our study confirms that assembling large, weakly supervised corpora serves as a cost-effective strategy for advancing speech technologies in under-resourced languages.

【2】 Machine Learning Techniques in Automatic Music Transcription: A Systematic Survey
标题: 自动音乐转录中的机器学习技术:系统调查
作者:Fatemeh Jamshidi,Gary Pike,Amit Das,Richard Chapman
链接:点击下载PDF文件
摘要:在音乐信息检索(MIR)领域,自动音乐转录(AMT)成为一个核心挑战,旨在将音频信号转换为符号符号符号,如音符或乐谱。这一系统的审查突出了AMT在音乐信号分析中的关键作用,强调其重要性,由于复杂和重叠的频谱结构的音乐和声。通过对AMT中使用的现有机器学习技术的深入研究,我们探索了当前模型和方法的进展和限制。尽管取得了显着的进步,但AMT系统尚未达到人类专家的准确性,这主要是由于音乐和声的复杂性以及对细微差别的解释的需求。这篇评论批判性地评估了全自动和半自动AMT系统,强调了最小用户干预的重要性,并研究了迄今为止提出的各种方法。通过解决现有技术的局限性,并提出改进的途径,我们的目标是引导未来的研究走向全自动AMT系统能够准确,有效地将复杂的音频信号转化为精确的符号表示。这项研究不仅综合了最新的进展,还为克服AMT中现有的挑战提供了路线图,为旨在缩小当前系统与人类水平转录准确性之间差距的研究人员提供了有价值的见解。摘要:In the domain of Music Information Retrieval (MIR), Automatic Music Transcription (AMT) emerges as a central challenge, aiming to convert audio signals into symbolic notations like musical notes or sheet music. This systematic review accentuates the pivotal role of AMT in music signal analysis, emphasizing its importance due to the intricate and overlapping spectral structure of musical harmonies. Through a thorough examination of existing machine learning techniques utilized in AMT, we explore the progress and constraints of current models and methodologies. Despite notable advancements, AMT systems have yet to match the accuracy of human experts, largely due to the complexities of musical harmonies and the need for nuanced interpretation. This review critically evaluates both fully automatic and semi-automatic AMT systems, emphasizing the importance of minimal user intervention and examining various methodologies proposed to date. By addressing the limitations of prior techniques and suggesting avenues for improvement, our objective is to steer future research towards fully automated AMT systems capable of accurately and efficiently translating intricate audio signals into precise symbolic representations. This study not only synthesizes the latest advancements but also lays out a road-map for overcoming existing challenges in AMT, providing valuable insights for researchers aiming to narrow the gap between current systems and human-level transcription accuracy.

【3】 Speech Emotion Recognition under Resource Constraints with Data Distillation
标题: 资源约束下的数据蒸馏语音情感识别
作者:Yi Chang,Zhao Ren,Zhonghao Zhao,Thanh Tam Nguyen,Kun Qian,Tanja Schultz,Björn W. Schuller
链接:点击下载PDF文件
摘要:语音情感识别在人机交互中起着至关重要的作用。物联网(IoT)中边缘设备的出现,由于内存和计算资源的限制,给构建复杂的深度学习模型带来了挑战。此外,情感语音数据通常包含隐私信息,这引起了人们对SER模型部署过程中隐私泄露的担忧。为了应对这些挑战,我们提出了一个数据蒸馏框架,以促进使用合成的,较小的和蒸馏的数据集在物联网应用中有效开发SER模型。我们的实验表明,蒸馏数据集可以有效地利用训练SER模型与固定的初始化,实现性能相比,开发使用原始的完整的情感语音数据集。摘要:Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset.

【4】 GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
标题: Globe:具有全球口音的高质量英语数据库,适用于Zero-Shot说话者自适应文本到语音
作者:Wenbin Wang,Yang Song,Sanjay Jha
备注:Interspeech 2024, 4 pages, 3 figures
链接:点击下载PDF文件
摘要:本文介绍了GLOBE,一个高质量的英语语料库与世界各地的口音,专门设计来解决目前的zero-shot扬声器自适应文本到语音(TTS)系统,表现出较差的泛化能力,适应说话人的口音的局限性。与常用的英语语料库(如LibriTTS和VCTK)相比,GLOBE的独特之处在于它包含了来自23,519名发言者的话语,涵盖了全球164种口音,以及这些发言者的详细元数据。与原始语料库相比,即,Common Voice,GLOBE通过严格的过滤和增强过程显著提高了语音数据的质量,同时还填充了所有缺失的说话者元数据。最终策划的GLOBE语料库包括535小时的语音数据,采样率为24 kHz。我们的基准测试结果表明,在GLOBE语料库上训练的说话人自适应TTS模型可以合成出比在其他流行语料库上训练的语音具有更好的说话人相似性和可比的自然度的语音。我们将在接受后公开发布GLOBE。GLOBE数据集可在https: globecorpus.github.io 上获得。摘要:This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the limitations of current zero-shot speaker adaptive Text-to-Speech (TTS) systems that exhibit poor generalizability in adapting to speakers with accents. Compared to commonly used English corpora, such as LibriTTS and VCTK, GLOBE is unique in its inclusion of utterances from 23,519 speakers and covers 164 accents worldwide, along with detailed metadata for these speakers. Compared to its original corpus, i.e., Common Voice, GLOBE significantly improves the quality of the speech data through rigorous filtering and enhancement processes, while also populating all missing speaker metadata. The final curated GLOBE corpus includes 535 hours of speech data at a 24 kHz sampling rate. Our benchmark results indicate that the speaker adaptive TTS model trained on the GLOBE corpus can synthesize speech with better speaker similarity and comparable naturalness than that trained on other popular corpora. We will release GLOBE publicly after acceptance. The GLOBE dataset is available at https: globecorpus.github.io .

【5】 Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
标题: 使用RNNT丢失的语音前置调整以改善LLM预测
作者:Murali Karthick Baskar,Andrew Rosenberg,Bhuvana Ramabhadran,Neeraj Gaur,Zhong Meng
链接:点击下载PDF文件
摘要:在本文中,我们专注于解决应用LLM到ASR时所面临的约束。最近的作品利用前缀LM型模型,直接将语音作为前缀应用于ASR的LLM。我们已经发现,优化语音前缀导致更好的ASR性能,并建议应用RNNT损失来执行语音前缀调整。这是一种简单的方法,不会增加模型复杂性或改变推理管道。我们还提出了基于语言的软提示,以进一步改善冻结LLM。对10种印度语言的实时测试集的实证分析表明,我们提出的语音前缀调整产生的改进与冻结和微调LLM。我们的平均10个指数的识别结果表明,建议的前缀调整与RNNT损失的结果在WER的12%的相对改善超过基线与微调LLM。我们提出的方法与冻结LLM导致一个31%的相对改善基本软提示前缀LM。摘要:In this paper, we focus on addressing the constraints faced when applying LLMs to ASR. Recent works utilize prefixLM-type models, which directly apply speech as a prefix to LLMs for ASR. We have found that optimizing speech prefixes leads to better ASR performance and propose applying RNNT loss to perform speech prefix-tuning. This is a simple approach and does not increase the model complexity or alter the inference pipeline. We also propose language-based soft prompting to further improve with frozen LLMs. Empirical analysis on realtime testset from 10 Indic languages demonstrate that our proposed speech prefix-tuning yields improvements with both frozen and fine-tuned LLMs. Our recognition results on an average of 10 Indics show that the proposed prefix-tuning with RNNT loss results in a 12 % relative improvement in WER over the baseline with a fine-tuned LLM. Our proposed approches with the frozen LLM leads to a 31 % relative improvement over basic soft-prompting prefixLM.


eess.AS音频处理
【1】 Prompting Whisper for QA-driven Zero-shot End-to-end Spoken Language Understanding
标题: 嵌入Whisper,用于QA驱动的Zero-Shot端到端口语理解
作者:Mohan Li,Simon Keizer,Rama Doddipatla
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:Zero-shot口语理解(SLU)使系统能够在新领域中理解用户话语,而无需事先接触训练数据。最近的研究通常依赖于大型语言模型(LLM),导致过多的足迹和复杂性。本文提出了使用Whisper,一个独立的语音处理模型,为zero-shot端到端(E2 E)SLU。为了处理看不见的语义标签,SLU任务被集成到一个问答(QA)框架中,该框架提示Whisper解码器进行语义推理。该系统通过前缀调整进行有效训练,优化最小参数集,而不是整个Whisper模型。我们表明,所提出的系统实现了40.7%的绝对增益为插槽填充(SLU-F1)SLURP相比,最近推出的zero-shot基准。此外,它执行的Whisper-GPT-2模块化系统在语料库内和跨语料库的评估设置,但相对减少34.8%的模型参数。摘要:Zero-shot spoken language understanding (SLU) enables systems to comprehend user utterances in new domains without prior exposure to training data. Recent studies often rely on large language models (LLMs), leading to excessive footprints and complexity. This paper proposes the use of Whisper, a standalone speech processing model, for zero-shot end-to-end (E2E) SLU. To handle unseen semantic labels, SLU tasks are integrated into a question-answering (QA) framework, which prompts the Whisper decoder for semantics deduction. The system is efficiently trained with prefix-tuning, optimising a minimal set of parameters rather than the entire Whisper model. We show that the proposed system achieves a 40.7% absolute gain for slot filling (SLU-F1) on SLURP compared to a recently introduced zero-shot benchmark. Furthermore, it performs comparably to a Whisper-GPT-2 modular system under both in-corpus and cross-corpus evaluation settings, but with a relative 34.8% reduction in model parameters.

【2】 Exploring Audio-Visual Information Fusion for Sound Event Localization and Detection In Low-Resource Realistic Scenarios
标题: 探索视听信息融合在低资源现实场景中进行声音事件定位和检测
作者:Ya Jiang,Qing Wang,Jun Du,Maocheng Hu,Pengfei Hu,Zeyan Liu,Shi Cheng,Zhaoxu Nian,Yuxuan Dong,Mingqi Cai,Xin Fang,Chin-Hui Lee
备注:accepted by icme2024
链接:点击下载PDF文件
摘要:本研究提出一种视听资讯融合的方法,以在低资源的情况下进行声音事件定位与侦测。我们的目标是通过跨模态学习和多模态融合来利用音频和视频模态信息。首先,我们提出了一个跨模态的师生学习(TSL)框架,将信息从一个只有音频的教师模型,训练了丰富的音频数据集与多种数据增强技术,视听学生模型训练,只有有限的一组多模态数据。接下来,我们提出了一个两阶段的视听融合策略,包括早期的功能融合和后期的视频引导决策融合,以利用音频和视频模态之间的协同作用。最后,我们介绍了一种创新的视频像素交换(VPS)技术扩展的音频通道交换(ACS)方法的视听联合增强。对声学场景和事件的检测和分类(DCASE)2023挑战数据集的评估结果表明,SELD性能有了显着改善。此外,我们提交给DCASE 2023挑战赛的SELD任务排名第一,有效地将所提出的技术集成到模型集成中。摘要:This study presents an audio-visual information fusion approach to sound event localization and detection (SELD) in low-resource scenarios. We aim at utilizing audio and video modality information through cross-modal learning and multi-modal fusion. First, we propose a cross-modal teacher-student learning (TSL) framework to transfer information from an audio-only teacher model, trained on a rich collection of audio data with multiple data augmentation techniques, to an audio-visual student model trained with only a limited set of multi-modal data. Next, we propose a two-stage audio-visual fusion strategy, consisting of an early feature fusion and a late video-guided decision fusion to exploit synergies between audio and video modalities. Finally, we introduce an innovative video pixel swapping (VPS) technique to extend an audio channel swapping (ACS) method to an audio-visual joint augmentation. Evaluation results on the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Challenge data set demonstrate significant improvements in SELD performances. Furthermore, our submission to the SELD task of the DCASE 2023 Challenge ranks first place by effectively integrating the proposed techniques into a model ensemble.

【3】 DExter: Learning and Controlling Performance Expression with Diffusion Models
标题: DExter:使用扩散模型学习和控制绩效表达
作者:Huan Zhang,Shreyan Chowdhury,Carlos Eduardo Cancino-Chacón,Jinhua Liang,Simon Dixon,Gerhard Widmer
备注:in submission to appsci special session
链接:点击下载PDF文件
摘要:在追求发展表现力的音乐表演模型,使用人工智能,本文介绍了DExter,一种新的方法,利用扩散概率模型来渲染西方古典钢琴演奏。在这种方法中,性能参数表示在一个连续的表达空间和扩散模型进行训练,以预测这些连续的参数,同时被限制在乐谱。此外,DExter还能够通过共同调节分数和感知特征表示来生成由感知有意义的特征引导的解释(表演的表达变化)。因此,我们发现,我们的模型是有用的学习表达性能,产生感知导向性能,并转移性能风格。我们通过定量和定性分析来评估模型,重点关注有关发音和清晰度等维度的特定性能指标,以及通过听力测试将生成的性能与不同的人类解释进行比较。结果表明,DExter是能够捕捉的表达参数的时变相关性,并比较现有的渲染模型在主观评价的评级。通过代理模型预测不同转向性能的感知特性,验证了DExter的感知特征条件生成和传输能力。摘要:In the pursuit of developing expressive music performance models using artificial intelligence, this paper introduces DExter, a new approach leveraging diffusion probabilistic models to render Western classical piano performances. In this approach, performance parameters are represented in a continuous expression space and a diffusion model is trained to predict these continuous parameters while being conditioned on the musical score. Furthermore, DExter also enables the generation of interpretations (expressive variations of a performance) guided by perceptually meaningful features by conditioning jointly on score and perceptual feature representations. Consequently, we find that our model is useful for learning expressive performance, generating perceptually steered performances, and transferring performance styles. We assess the model through quantitative and qualitative analyses, focusing on specific performance metrics regarding dimensions like asynchrony and articulation, as well as through listening tests comparing generated performances with different human interpretations. Results show that DExter is able to capture the time-varying correlation of the expressive parameters, and compares well to existing rendering models in subjectively evaluated ratings. The perceptual-feature-conditioned generation and transferring capabilities of DExter are verified by a proxy model predicting perceptual characteristics of differently steered performances.

【4】 Voice Disorder Analysis: a Transformer-based Approach
标题: 语音障碍分析:基于转换器的方法
作者:Alkis Koudounas,Gabriele Ciravegna,Marco Fantini,Giovanni Succo,Erika Crosetti,Tania Cerquitelli,Elena Baralis
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音障碍是显著影响患者生活质量的病理。然而,由于病理语音数据的短缺和用于诊断的记录类型的多样性,这些病理的非侵入性自动诊断仍然是探索不足。本文提出了一种新的解决方案,采用Transformers直接工作在原始语音信号和解决数据不足,通过合成数据生成和数据增强。此外,我们同时考虑许多记录类型,例如句子阅读和持续元音发射,通过采用混合专家集成来对齐不同数据类型的预测。在公共和私人数据集上获得的实验结果表明,我们的解决方案在疾病检测和分类任务中的有效性,并大大改善了现有的方法。摘要:Voice disorders are pathologies significantly affecting patient quality of life. However, non-invasive automated diagnosis of these pathologies is still under-explored, due to both a shortage of pathological voice data, and diversity of the recording types used for the diagnosis. This paper proposes a novel solution that adopts transformers directly working on raw voice signals and addresses data shortage through synthetic data generation and data augmentation. Further, we consider many recording types at the same time, such as sentence reading and sustained vowel emission, by employing a Mixture of Expert ensemble to align the predictions on different data types. The experimental results, obtained on both public and private datasets, show the effectiveness of our solution in the disorder detection and classification tasks and largely improve over existing approaches.

【5】 Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis
标题: 手术室智能语音助手:手术工作流程分析的多模式模型
作者:Kubilay Can Demir,Belen Lojo Rodriguez,Tobias Weise,Andreas Maier,Seung Hee Yang
备注:5 Pages, Interspeech 2024
链接:点击下载PDF文件
摘要:为了开发智能语音助手并将其与术中决策支持框架无缝集成,准确高效的手术阶段识别是先决条件。在这项研究中,我们提出了一个多模态框架的基础上门控多模态单元(GMU)和多阶段时间卷积网络(MS-TCN),以识别端口导管放置操作的手术阶段。我们的方法合并语音和图像模型,并在不同的手术阶段分别使用它们。基于对28个操作的评估,我们报告了92.65 $ pm $3.52%的逐帧准确度和92.30 $ pm $3.82%的F1分数。我们的研究结果表明,这两个指标比以前的工作约10%的改善,并验证了手术阶段识别任务的多模态数据集成的有效性。我们进一步研究的贡献,个别数据通道比较单模态模型与多模态模型。摘要:To develop intelligent speech assistants and integrate them seamlessly with intra-operative decision-support frameworks, accurate and efficient surgical phase recognition is a prerequisite. In this study, we propose a multimodal framework based on Gated Multimodal Units (GMU) and Multi-Stage Temporal Convolutional Networks (MS-TCN) to recognize surgical phases of port-catheter placement operations. Our method merges speech and image models and uses them separately in different surgical phases. Based on the evaluation of 28 operations, we report a frame-wise accuracy of 92.65 $ pm$ 3.52% and an F1-score of 92.30 $ pm$ 3.82%. Our results show approximately 10% improvement in both metrics over previous work and validate the effectiveness of integrating multimodal data for the surgical phase recognition task. We further investigate the contribution of individual data channels by comparing mono-modal models with multimodal models.

【6】 The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
标题: 希腊播客文集:具有弱监督数据的低资源语言的竞争语音模型
作者:Georgios Paraskevopoulos,Chara Tsoukala,Athanasios Katsamanis,Vassilis Katsouros
备注:To be presented at Interspeech 2024
链接:点击下载PDF文件
摘要:为数字表示有限的语言开发语音技术带来了重大挑战,主要是由于缺乏可用数据。在大型数据密集型模型时代,这个问题更加严重。最近的研究强调了利用薄弱的监督来扩大现有数据库的潜力。在这项研究中,我们编译了一个800小时的现代希腊语语料库,从播客和使用Whisper large-v3生成银transmits。该语料库被用来微调我们的模型,旨在评估这种方法在提高ASR性能的有效性。我们的分析涵盖了16个不同的播客领域,同时对现代希腊语的既定数据集进行了评估。研究结果表明WER的持续改善,与数据量和模型大小的增加相关。我们的研究证实,组装大型,弱监督语料库作为一个具有成本效益的战略,推进语音技术在资源不足的语言。摘要:The development of speech technologies for languages with limited digital representation poses significant challenges, primarily due to the scarcity of available data. This issue is exacerbated in the era of large, data-intensive models. Recent research has underscored the potential of leveraging weak supervision to augment the pool of available data. In this study, we compile an 800-hour corpus of Modern Greek from podcasts and employ Whisper large-v3 to generate silver transcriptions. This corpus is utilized to fine-tune our models, aiming to assess the efficacy of this approach in enhancing ASR performance. Our analysis spans 16 distinct podcast domains, alongside evaluations on established datasets for Modern Greek. The findings indicate consistent WER improvements, correlating with increases in both data volume and model size. Our study confirms that assembling large, weakly supervised corpora serves as a cost-effective strategy for advancing speech technologies in under-resourced languages.

【7】 Machine Learning Techniques in Automatic Music Transcription: A Systematic Survey
标题: 自动音乐转录中的机器学习技术:系统调查
作者:Fatemeh Jamshidi,Gary Pike,Amit Das,Richard Chapman
链接:点击下载PDF文件
摘要:在音乐信息检索(MIR)领域,自动音乐转录(AMT)成为一个核心挑战,旨在将音频信号转换为符号,如音符或乐谱。这一系统的审查突出了AMT在音乐信号分析中的关键作用,强调其重要性,由于复杂和重叠的频谱结构的音乐和声。通过对AMT中使用的现有机器学习技术的深入研究,我们探索了当前模型和方法的进展和限制。尽管取得了显着的进步,但AMT系统尚未达到人类专家的准确性,这主要是由于音乐和声的复杂性以及对细微差别的解释的需求。这篇评论批判性地评估了全自动和半自动AMT系统,强调了最小用户干预的重要性,并研究了迄今为止提出的各种方法。通过解决现有技术的局限性,并提出改进的途径,我们的目标是引导未来的研究走向全自动AMT系统能够准确,有效地将复杂的音频信号转化为精确的符号表示。这项研究不仅综合了最新的进展,还为克服AMT中现有的挑战提供了路线图,为旨在缩小当前系统与人类水平转录准确性之间差距的研究人员提供了有价值的见解。摘要:In the domain of Music Information Retrieval (MIR), Automatic Music Transcription (AMT) emerges as a central challenge, aiming to convert audio signals into symbolic notations like musical notes or sheet music. This systematic review accentuates the pivotal role of AMT in music signal analysis, emphasizing its importance due to the intricate and overlapping spectral structure of musical harmonies. Through a thorough examination of existing machine learning techniques utilized in AMT, we explore the progress and constraints of current models and methodologies. Despite notable advancements, AMT systems have yet to match the accuracy of human experts, largely due to the complexities of musical harmonies and the need for nuanced interpretation. This review critically evaluates both fully automatic and semi-automatic AMT systems, emphasizing the importance of minimal user intervention and examining various methodologies proposed to date. By addressing the limitations of prior techniques and suggesting avenues for improvement, our objective is to steer future research towards fully automated AMT systems capable of accurately and efficiently translating intricate audio signals into precise symbolic representations. This study not only synthesizes the latest advancements but also lays out a road-map for overcoming existing challenges in AMT, providing valuable insights for researchers aiming to narrow the gap between current systems and human-level transcription accuracy.

【8】 Speech Emotion Recognition under Resource Constraints with Data Distillation
标题: 资源约束下的数据蒸馏语音情感识别
作者:Yi Chang,Zhao Ren,Zhonghao Zhao,Thanh Tam Nguyen,Kun Qian,Tanja Schultz,Björn W. Schuller
链接:点击下载PDF文件
摘要:语音情感识别在人机交互中起着至关重要的作用。物联网(IoT)中边缘设备的出现,由于内存和计算资源的限制,给构建复杂的深度学习模型带来了挑战。此外,情感语音数据通常包含隐私信息,这引起了人们对SER模型部署过程中隐私泄露的担忧。为了应对这些挑战,我们提出了一个数据蒸馏框架,以促进使用合成的,较小的和蒸馏的数据集在物联网应用中有效开发SER模型。我们的实验表明,蒸馏数据集可以有效地利用训练SER模型与固定的初始化,实现性能相比,开发使用原始的完整的情感语音数据集。摘要:Speech emotion recognition (SER) plays a crucial role in human-computer interaction. The emergence of edge devices in the Internet of Things (IoT) presents challenges in constructing intricate deep learning models due to constraints in memory and computational resources. Moreover, emotional speech data often contains private information, raising concerns about privacy leakage during the deployment of SER models. To address these challenges, we propose a data distillation framework to facilitate efficient development of SER models in IoT applications using a synthesised, smaller, and distilled dataset. Our experiments demonstrate that the distilled dataset can be effectively utilised to train SER models with fixed initialisation, achieving performances comparable to those developed using the original full emotional speech dataset.

【9】 InterBiasing: Boost Unseen Word Recognition through Biasing Intermediate Predictions
标题: 互偏:通过偏置中间预测来提高不可见的单词识别
作者:Yu Nakagome,Michael Hentschel
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:尽管最近在端到端语音识别方法方面取得了进展,但它们的输出偏向于训练数据的词汇,导致对未知术语或专有名词的识别不准确。为了提高识别精度,对于一个给定的一组这样的条款,我们提出了一个自适应参数的基础上自适应CTC的方法。我们的方法通过用正确的标签替换其中间CTC预测,然后将其传递到后续层,来提高误识别目标关键字的识别准确性。首先,我们使用文本到语音和识别模型为关键字列表创建正确的标签和识别错误实例对。我们使用这些对来替换标签的中间预测误差。调节标签上的编码器的后续层,可以从声学上评估目标关键字。在日语中进行的实验表明,我们的方法成功地提高了未知词的F1分数。摘要:Despite recent advances in end-to-end speech recognition methods, their output is biased to the training data's vocabulary, resulting in inaccurate recognition of unknown terms or proper nouns. To improve the recognition accuracy for a given set of such terms, we propose an adaptation parameter-free approach based on Self-conditioned CTC. Our method improves the recognition accuracy of misrecognized target keywords by substituting their intermediate CTC predictions with corrected labels, which are then passed on to the subsequent layers. First, we create pairs of correct labels and recognition error instances for a keyword list using Text-to-Speech and a recognition model. We use these pairs to replace intermediate prediction errors by the labels. Conditioning the subsequent layers of the encoder on the labels, it is possible to acoustically evaluate the target keywords. Experiments conducted in Japanese demonstrated that our method successfully improved the F1 score for unknown words.

【10】 GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
标题: Globe:具有全球口音的高质量英语数据库,适用于Zero-Shot说话者自适应文本到语音
作者:Wenbin Wang,Yang Song,Sanjay Jha
备注:Interspeech 2024, 4 pages, 3 figures
链接:点击下载PDF文件
摘要:本文介绍了GLOBE,一个高质量的英语语料库与世界各地的口音,专门设计来解决目前的zero-shot扬声器自适应文本到语音(TTS)系统,表现出较差的泛化能力,适应说话人的口音的局限性。与常用的英语语料库(如LibriTTS和VCTK)相比,GLOBE的独特之处在于它包含了来自23,519名发言者的话语,涵盖了全球164种口音,以及这些发言者的详细元数据。与原始语料库相比,即,Common Voice,GLOBE通过严格的过滤和增强过程显著提高了语音数据的质量,同时还填充了所有缺失的说话者元数据。最终策划的GLOBE语料库包括535小时的语音数据,采样率为24 kHz。我们的基准测试结果表明,在GLOBE语料库上训练的说话人自适应TTS模型可以合成出比在其他流行语料库上训练的语音具有更好的说话人相似性和可比的自然度的语音。我们将在接受后公开发布GLOBE。GLOBE数据集可在https: globecorpus.github.io 上获得。摘要:This paper introduces GLOBE, a high-quality English corpus with worldwide accents, specifically designed to address the limitations of current zero-shot speaker adaptive Text-to-Speech (TTS) systems that exhibit poor generalizability in adapting to speakers with accents. Compared to commonly used English corpora, such as LibriTTS and VCTK, GLOBE is unique in its inclusion of utterances from 23,519 speakers and covers 164 accents worldwide, along with detailed metadata for these speakers. Compared to its original corpus, i.e., Common Voice, GLOBE significantly improves the quality of the speech data through rigorous filtering and enhancement processes, while also populating all missing speaker metadata. The final curated GLOBE corpus includes 535 hours of speech data at a 24 kHz sampling rate. Our benchmark results indicate that the speaker adaptive TTS model trained on the GLOBE corpus can synthesize speech with better speaker similarity and comparable naturalness than that trained on other popular corpora. We will release GLOBE publicly after acceptance. The GLOBE dataset is available at https: globecorpus.github.io .

【11】 Speech Prefix-Tuning with RNNT Loss for Improving LLM Predictions
标题: 使用RNNT丢失的语音前置调整以改善LLM预测
作者:Murali Karthick Baskar,Andrew Rosenberg,Bhuvana Ramabhadran,Neeraj Gaur,Zhong Meng
链接:点击下载PDF文件
摘要:在本文中,我们专注于解决应用LLM到ASR时所面临的约束。最近的作品利用前缀LM型模型,直接将语音作为前缀应用于ASR的LLM。我们已经发现,优化语音前缀导致更好的ASR性能,并建议应用RNNT损失来执行语音前缀调整。这是一种简单的方法,不会增加模型复杂性或改变推理管道。我们还提出了基于语言的软提示,以进一步改善冻结LLM。对10种印度语言的实时测试集的实证分析表明,我们提出的语音前缀调整产生的改进与冻结和微调LLM。我们的平均10个指数的识别结果表明,建议的前缀调整与RNNT损失的结果在WER的12%的相对改善超过基线与微调LLM。我们提出的方法与冻结LLM导致一个31%的相对改善基本软提示前缀LM。摘要:In this paper, we focus on addressing the constraints faced when applying LLMs to ASR. Recent works utilize prefixLM-type models, which directly apply speech as a prefix to LLMs for ASR. We have found that optimizing speech prefixes leads to better ASR performance and propose applying RNNT loss to perform speech prefix-tuning. This is a simple approach and does not increase the model complexity or alter the inference pipeline. We also propose language-based soft prompting to further improve with frozen LLMs. Empirical analysis on realtime testset from 10 Indic languages demonstrate that our proposed speech prefix-tuning yields improvements with both frozen and fine-tuned LLMs. Our recognition results on an average of 10 Indics show that the proposed prefix-tuning with RNNT loss results in a 12 % relative improvement in WER over the baseline with a fine-tuned LLM. Our proposed approches with the frozen LLM leads to a 31 % relative improvement over basic soft-prompting prefixLM.

【12】 A Contrastive Learning Approach to Mitigate Bias in Speech Models
标题: 减轻语音模型中偏见的对比学习方法
作者:Alkis Koudounas,Flavio Giobergia,Eliana Pastor,Elena Baralis
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:语音模型可能会受到不同人群亚组中性能不平衡的影响,这引发了对这些群体之间公平待遇的担忧。以前的尝试,以减轻不公平的要么集中在用户定义的子组,可能会忽略其他受影响的子组,或不显着改善内部表示在子组级别。本文提出了第一次采用对比学习,以减轻语音模型偏差表现不佳的子组。我们采用了一种三级学习技术,引导模型关注对比损失的不同范围,即,任务、子组和子组内的错误。在两个口语理解数据集和两种语言上的实验表明,我们的方法改进了内部子组表示,从而减少了模型偏差并提高了性能。摘要:Speech models may be affected by performance imbalance in different population subgroups, raising concerns about fair treatment across these groups. Prior attempts to mitigate unfairness either focus on user-defined subgroups, potentially overlooking other affected subgroups, or do not explicitly improve the internal representation at the subgroup level. This paper proposes the first adoption of contrastive learning to mitigate speech model bias in underperforming subgroups. We employ a three-level learning technique that guides the model in focusing on different scopes for the contrastive loss, i.e., task, subgroup, and the errors within subgroups. The experiments on two spoken language understanding datasets and two languages demonstrate that our approach improves internal subgroup representations, thus reducing model bias and enhancing performance.


机器翻译,仅供参考