今日论文合集:cs.SD语音9篇,eess.AS音频处理16篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Distortion Recovery: A Two-Stage Method for Guitar Effect Removal
标题: 失真恢复:消除吉他效应的两阶段方法
作者:Ying-Shuo Lee,Yueh-Po Peng,Jui-Te Wu,Ming Cheng,Li Su,Yi-Hsuan Yang
备注:DAFx 2024
链接:点击下载PDF文件
摘要:从电吉他录音中删除音频效果可以更轻松地进行后期制作和声音编辑。音频失真恢复模型不仅提高了吉他声音的清晰度,还为混音和母带的创造性调整开辟了新的机会。虽然在创建此类模型方面取得了进展,但以前的努力主要集中在合成失真上,这些失真可能过于简单,无法准确捕捉现实世界记录中的复杂性。 在本文中,我们通过使用商业级音频效果VST插件渲染的吉他录音数据集来解决这个问题。此外,我们介绍了一种新的两阶段的音频失真恢复方法。其思想是首先在第一阶段中在Mel频谱图域中处理音频信号,然后在第二阶段中使用神经声码器从处理后的Mel频谱图生成原始的吉他声音。我们报告了一组实验,通过主观和客观的评价指标,证明了我们的方法对现有方法的有效性。摘要:Removing audio effects from electric guitar recordings makes it easier for post-production and sound editing. An audio distortion recovery model not only improves the clarity of the guitar sounds but also opens up new opportunities for creative adjustments in mixing and mastering. While progress have been made in creating such models, previous efforts have largely focused on synthetic distortions that may be too simplistic to accurately capture the complexities seen in real-world recordings. In this paper, we tackle the task by using a dataset of guitar recordings rendered with commercial-grade audio effect VST plugins. Moreover, we introduce a novel two-stage methodology for audio distortion recovery. The idea is to firstly process the audio signal in the Mel-spectrogram domain in the first stage, and then use a neural vocoder to generate the pristine original guitar sound from the processed Mel-spectrogram in the second stage. We report a set of experiments demonstrating the effectiveness of our approach over existing methods, through both subjective and objective evaluation metrics.

【2】 Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
标题: 音频提示适配器:通过轻量级微调释放文本到音乐的音乐编辑能力
作者:Fang-Duo Tsai,Shih-Lun Wu,Haven Kim,Bo-Yu Chen,Hao-Chung Cheng,Yi-Hsuan Yang
备注:Accepted by the 25th International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:文本到音乐模型允许用户使用文本命令生成近乎真实的音乐音频。然而,编辑音乐音频仍然具有挑战性,这是由于在维持简单用户界面的同时对音频执行细粒度更改的冲突需求。为了应对这一挑战,我们提出了音频提示适配器(或AP适配器),这是对预训练文本到音乐模型的轻量级补充。我们利用AudioMAE从输入音频中提取特征,并构建基于注意力的适配器将这些特征馈送到AudioLDM2的内部层,AudioLDM2是一个基于扩散的文本到音乐模型。通过22M可训练参数,AP适配器使用户能够利用全局(例如,流派和音色)和本地(例如,旋律)方面,使用原始音频和短文本作为输入。通过客观和主观的研究,我们评估AP适配器的三个任务:音色转移,体裁转移,和伴奏生成。此外,我们证明了它的有效性域外的音频包含看不见的仪器在训练过程中。摘要:Text-to-music models allow users to generate nearly realistic musical audio with textual commands. However, editing music audios remains challenging due to the conflicting desiderata of performing fine-grained alterations on the audio while maintaining a simple user interface. To address this challenge, we propose Audio Prompt Adapter (or AP-Adapter), a lightweight addition to pretrained text-to-music models. We utilize AudioMAE to extract features from the input audio, and construct attention-based adapters to feedthese features into the internal layers of AudioLDM2, a diffusion-based text-to-music model. With 22M trainable parameters, AP-Adapter empowers users to harness both global (e.g., genre and timbre) and local (e.g., melody) aspects of music, using the original audio and a short text as inputs. Through objective and subjective studies, we evaluate AP-Adapter on three tasks: timbre transfer, genre transfer, and accompaniment generation. Additionally, we demonstrate its effectiveness on out-of-domain audios containing unseen instruments during training.

【3】 Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
标题: 音频时间伪造检测和定位的从粗到细的提案细化框架
作者:Junyan Wu,Wei Lu,Xiangyang Luo,Rui Yang,Qian Wang,Xiaochun Cao
备注:9pages, 3figures. This paper has been accepted for ACM MM 2024
链接:点击下载PDF文件
摘要:最近,一种新形式的音频部分伪造对其取证提出了挑战,需要先进的对策来检测长时间音频内的细微伪造操作。然而,现有的对策仍然服务于分类目的,并且不能对部分伪造段的开始和结束时间戳执行有意义的分析。为了解决这一挑战,我们引入了一种新的粗到细的建议细化框架(CFPRF),它结合了帧级检测网络(FDN)和建议细化网络(PRN)的音频时间伪造检测和定位。具体而言,FDN的目的是挖掘真实和虚假帧之间的信息不一致线索,以获得有利于粗略指示伪造区域的区别性特征。PRN负责预测置信度分数和回归偏移,以细化从FDN导出的粗粒度建议。为了学习鲁棒的判别特征,我们设计了一个差异感知特征学习(DAFL)模块,该模块由对比表示学习指导,以放大由微小操作引起的不同帧之间的敏感差异。我们进一步设计了一个边界感知的特征增强(BAFE)模块来捕获多个过渡边界的上下文信息,并通过交叉注意机制引导边界信息和时间特征之间的交互。大量的实验表明,我们的CFPRF在各种数据集上实现了最先进的性能,包括LAV-DF,ASVS 2019 PS和HAD。摘要:Recently, a novel form of audio partial forgery has posed challenges to its forensics, requiring advanced countermeasures to detect subtle forgery manipulations within long-duration audio. However, existing countermeasures still serve a classification purpose and fail to perform meaningful analysis of the start and end timestamps of partial forgery segments. To address this challenge, we introduce a novel coarse-to-fine proposal refinement framework (CFPRF) that incorporates a frame-level detection network (FDN) and a proposal refinement network (PRN) for audio temporal forgery detection and localization. Specifically, the FDN aims to mine informative inconsistency cues between real and fake frames to obtain discriminative features that are beneficial for roughly indicating forgery regions. The PRN is responsible for predicting confidence scores and regression offsets to refine the coarse-grained proposals derived from the FDN. To learn robust discriminative features, we devise a difference-aware feature learning (DAFL) module guided by contrastive representation learning to enlarge the sensitive differences between different frames induced by minor manipulations. We further design a boundary-aware feature enhancement (BAFE) module to capture the contextual information of multiple transition boundaries and guide the interaction between boundary information and temporal features via a cross-attention mechanism. Extensive experiments show that our CFPRF achieves state-of-the-art performance on various datasets, including LAV-DF, ASVS2019PS, and HAD.

【4】 On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
标题: 关于绒猴叫声分析的语音和音频基础模型的实用性
作者:Eklavya Sarkar,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024 Satellite event (VIHAR 2024)
链接:点击下载PDF文件
摘要:绒猴在它们的叫声中编码重要信息,并作为神经生物学家了解人类声音交流进化起源的替代模型。传统上使用基于信号处理的特征进行分析,最近的方法利用在人类语音上预训练的自监督模型进行特征提取,利用它们独立于声学域学习信号内在结构的能力。然而,在多类分类、带宽和预训练域方面,这些基础模型的效用对于绒猴呼叫分析仍然不清楚。本研究评估了来自语音和一般音频域的特征表示,在4,8和16 kHz的预训练带宽的绒猴呼叫类型和呼叫者分类任务。结果表明,具有更高带宽的模型可以提高性能,并且对语音或一般音频进行预训练会产生相当的结果,在频谱基线上有所改善。摘要:Marmoset monkeys encode vital information in their calls and serve as a surrogate model for neuro-biologists to understand the evolutionary origins of human vocal communication. Traditionally analyzed with signal processing-based features, recent approaches have utilized self-supervised models pre-trained on human speech for feature extraction, capitalizing on their ability to learn a signal's intrinsic structure independently of its acoustic domain. However, the utility of such foundation models remains unclear for marmoset call analysis in terms of multi-class classification, bandwidth, and pre-training domain. This study assesses feature representations derived from speech and general audio domains, across pre-training bandwidths of 4, 8, and 16 kHz for marmoset call-type and caller classification tasks. Results show that models with higher bandwidth improve performance, and pre-training on speech or general audio yields comparable results, improving over a spectral baseline.

【5】 Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
标题: 基于LLM的ASB后错误纠正的进化提示设计
作者:Rithik Sachdev,Zhong-Qiu Wang,Chao-Han Huck Yang
备注:in submission
链接:点击下载PDF文件
摘要:基于现代大型语言模型(LLM)的优势,生成纠错(GEC)已经成为一种有前途的范例,可以提高现代自动语音识别(ASR)系统的性能。一种代表性的方法是利用上下文学习来提示LLM,使得LLM可以基于精心设计的提示和由ASR系统产生的$N$最佳假设列表来生成更好的假设。然而,目前尚不清楚现有的提示是否是最有效的ASR后纠错的任务。在这种情况下,本文首先探讨替代提示,以确定一个初始的有效提示,然后提出采用进化提示优化算法来完善初始提示。在2024 $ GenSEC挑战的任务$1$的CHiME-4子集上的评估结果显示了所提出的算法的有效性和潜力。摘要:Building upon the strength of modern large language models (LLMs), generative error correction (GEC) has emerged as a promising paradigm that can elevate the performance of modern automatic speech recognition (ASR) systems. One representative approach is to leverage in-context learning to prompt LLMs so that a better hypothesis can be generated by the LLMs based on a carefully-designed prompt and an $N$-best list of hypotheses produced by ASR systems. However, it is yet unknown whether the existing prompts are the most effective ones for the task of post-ASR error correction. In this context, this paper first explores alternative prompts to identify an initial set of effective prompts, and then proposes to employ an evolutionary prompt optimization algorithm to refine the initial prompts. Evaluations results on the CHiME-4 subset of the Task $1$ of the SLT $2024$ GenSEC challenge show the effectiveness and potential of the proposed algorithms.

【6】 Multimodal Input Aids a Bayesian Model of Phonetic Learning
标题: 多模式输入辅助语音学习的Bayesian模型
作者:Sophia Zhi,Roger P. Levy,Stephan C. Meylan
备注:12 pages, 5 figures
链接:点击下载PDF文件
摘要:典型的儿童语言学习者面临的许多任务之一是学习区分构成母语单词的独特声音。在这里,我们调查是否多模态信息-特别是成人语音加上视频帧的扬声器的脸-有利于语音学习的计算模型。我们介绍了一种方法,用于创建高质量的合成视频扬声器的脸为现有的音频语料库。我们的学习模型,在视听输入上训练和测试时,与仅在音频输入上训练和测试的模型相比,在音素辨别电池上实现了高达8.1%的相对改善。当两者都在纯音频数据上进行测试时,它的表现也优于音频模型高达3.9%,这表明视觉信息有助于获得声学区别。视觉信息在嘈杂的音频环境中特别有益,其中视听模型相对于无噪声环境关闭了噪声中音频模型的辨别性能损失的67%。这些结果表明,视觉信息有利于一个理想的学习者,并说明了一些方式,儿童可能能够利用视觉线索时,学习区分语音。摘要:One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal information--specifically adult speech coupled with video frames of speakers' faces--benefits a computational model of phonetic learning. We introduce a method for creating high-quality synthetic videos of speakers' faces for an existing audio corpus. Our learning model, when both trained and tested on audiovisual inputs, achieves up to a 8.1% relative improvement on a phoneme discrimination battery compared to a model trained and tested on audio-only input. It also outperforms the audio model by up to 3.9% when both are tested on audio-only data, suggesting that visual information facilitates the acquisition of acoustic distinctions. Visual information is especially beneficial in noisy audio environments, where an audiovisual model closes 67% of the loss in discrimination performance of the audio model in noise relative to a non-noisy environment. These results demonstrate that visual information benefits an ideal learner and illustrate some of the ways that children might be able to leverage visual cues when learning to discriminate speech sounds.

【7】 Automatic Equalization for Individual Instrument Tracks Using Convolutional Neural Networks
标题: 使用卷积神经网络对单个乐器轨迹进行自动均衡
作者:Florian Mockenhaupt,Joscha Simon Rieber,Shahan Nercessian
备注:8 pages, 9 figures. Accepted to the 27th International Conference on Digital Audio Effects (DAFx24)
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,用于自动均衡的个别乐器的轨道。我们的方法首先识别源记录中存在的仪器,以选择其相应的理想光谱作为目标。接下来,计算记录和目标之间的光谱差异,并且相应地,均衡器匹配模型用于预测参数均衡器的设置。为此,我们建立在一个可微参数均衡器匹配神经网络,展示了相对于以前建立的最先进的改进。与过去的方法不同,我们展示了我们的系统如何自然地允许在我们的匹配模型的训练过程中利用真实世界的音频数据,有效地生成适当的生产训练目标在自动化的方式镜像条件在推理时间。因此,我们说明了如何微调我们的匹配模型在这样的例子大大提高了参数均衡器匹配性能在现实世界中的情况下,减少平均绝对误差24%,相对于方法只依赖于随机参数采样技术作为一个自我监督的学习策略。我们进行听力测试,并证明我们提出的自动均衡解决方案主观上增强了常见乐器类型录音的音调特征。摘要:We propose a novel approach for the automatic equalization of individual musical instrument tracks. Our method begins by identifying the instrument present within a source recording in order to choose its corresponding ideal spectrum as a target. Next, the spectral difference between the recording and the target is calculated, and accordingly, an equalizer matching model is used to predict settings for a parametric equalizer. To this end, we build upon a differentiable parametric equalizer matching neural network, demonstrating improvements relative to previously established state-of-the-art. Unlike past approaches, we show how our system naturally allows real-world audio data to be leveraged during the training of our matching model, effectively generating suitably produced training targets in an automated manner mirroring conditions at inference time. Consequently, we illustrate how fine-tuning our matching model on such examples considerably improves parametric equalizer matching performance in real-world scenarios, decreasing mean absolute error by 24% relative to methods relying solely on random parameter sampling techniques as a self-supervised learning strategy. We perform listening tests, and demonstrate that our proposed automatic equalization solution subjectively enhances the tonal characteristics for recordings of common instrument types.

【8】 Synthesizer Sound Matching Using Audio Spectrogram Transformers
标题: 使用音频频谱图变换器的合成器声音匹配
作者:Fred Bruford,Frederik Blang,Shahan Nercessian
备注:4 pages, 1 figure. Accepted to the 27th International Conference on Digital Audio Effects (DAFx24)
链接:点击下载PDF文件
摘要:用于合成器声音匹配的系统,其自动设置合成器的参数以模拟输入声音,具有使合成器编程的过程对于新手和有经验的音乐家来说更快和更容易的潜力,同时还提供与合成器交互的新手段。考虑到市场上各种各样的合成器,以及它们中的许多合成器的复杂性,特别需要以关于底层合成架构的最少知识或先验假设来工作的通用声音匹配系统。考虑到这一点,我们引入了一个基于音频谱图Transformer的合成器声音匹配模型。我们通过在一个由流行的Massive合成器随机生成的样本组成的大型合成数据集上进行训练,证明了该模型的可行性。我们表明,该模型可以重建从一组16个参数生成的样本的参数,突出了其相对于多层感知器和卷积神经网络基线的保真度提高。我们还提供了音频示例,演示了域外模型在模拟人声模仿以及来自其他合成器和乐器的声音时的性能。摘要:Systems for synthesizer sound matching, which automatically set the parameters of a synthesizer to emulate an input sound, have the potential to make the process of synthesizer programming faster and easier for novice and experienced musicians alike, whilst also affording new means of interaction with synthesizers. Considering the enormous variety of synthesizers in the marketplace, and the complexity of many of them, general-purpose sound matching systems that function with minimal knowledge or prior assumptions about the underlying synthesis architecture are particularly desirable. With this in mind, we introduce a synthesizer sound matching model based on the Audio Spectrogram Transformer. We demonstrate the viability of this model by training on a large synthetic dataset of randomly generated samples from the popular Massive synthesizer. We show that this model can reconstruct parameters of samples generated from a set of 16 parameters, highlighting its improved fidelity relative to multi-layer perceptron and convolutional neural network baselines. We also provide audio examples demonstrating the out-of-domain model performance in emulating vocal imitations, and sounds from other synthesizers and musical instruments.

【9】 The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
标题: CHiME-8针对可概括和阵列不可知的远程自动语音识别和规模化的DASB挑战
作者:Samuele Cornell,Taejin Park,Steve Huang,Christoph Boeddeker,Xuankai Chang,Matthew Maciejewski,Matthew Wiesner,Paola Garcia,Shinji Watanabe
链接:点击下载PDF文件
摘要:本文介绍了CHiME-8 DASR挑战赛,该挑战赛是从上一版CHiME-7 DASR(C7 DASR)和过去的CHiME-6挑战赛中进行的。它侧重于联合多通道远程语音识别(DASR)和日记与一个或多个,可能异构,设备。主要目标是促进对会议转录方法的研究,这些方法可以概括任意数量的发言者,不同的设置(正式与非正式对话),会议持续时间,各种各样的声学场景和不同的录音配置。C7 DASR的新颖之处包括:i)添加了NOTSOFAR-1,这是一个额外的办公室 公司会议场景,ii)手动更正的Mixer 6开发集,iii)允许使用大语言模型的新曲目(LLM)iv)评审团奖励机制,鼓励参与者探索更实用和创新的解决方案。为了降低参与者的进入门槛,我们提供了一个独立的工具包,用于下载和准备这些数据集,以及执行文本规范化和评分。此外,今年我们还提供了两个基线系统,一个直接继承自C7 DASR,基于ESPnet,另一个基于NeMo开发,基于去年C7 DASR的NeMo团队提交。基线系统的结果表明,增加了NOTSOFAR-1的情况下显着增加了任务的难度,由于其大量的发言者和非常短的持续时间。摘要:This paper presents the CHiME-8 DASR challenge which carries on from the previous edition CHiME-7 DASR (C7DASR) and the past CHiME-6 challenge. It focuses on joint multi-channel distant speech recognition (DASR) and diarization with one or more, possibly heterogeneous, devices. The main goal is to spur research towards meeting transcription approaches that can generalize across arbitrary number of speakers, diverse settings (formal vs. informal conversations), meeting duration, wide-variety of acoustic scenarios and different recording configurations. Novelties with respect to C7DASR include: i) the addition of NOTSOFAR-1, an additional office corporate meeting scenario, ii) a manually corrected Mixer 6 development set, iii) a new track in which we allow the use of large-language models (LLM) iv) a jury award mechanism to encourage participants to explore also more practical and innovative solutions. To lower the entry barrier for participants, we provide a standalone toolkit for downloading and preparing such datasets as well as performing text normalization and scoring their submissions. Furthermore, this year we also provide two baseline systems, one directly inherited from C7DASR and based on ESPnet and another one developed on NeMo and based on NeMo team submission in last year C7DASR. Baseline system results suggest that the addition of the NOTSOFAR-1 scenario significantly increases the task's difficulty due to its high number of speakers and very short duration.


eess.AS音频处理
【1】 Automatic Equalization for Individual Instrument Tracks Using Convolutional Neural Networks
标题: 使用卷积神经网络对单个乐器轨迹进行自动均衡
作者:Florian Mockenhaupt,Joscha Simon Rieber,Shahan Nercessian
备注:8 pages, 9 figures. Accepted to the 27th International Conference on Digital Audio Effects (DAFx24)
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,用于自动均衡的个别乐器的轨道。我们的方法首先识别源记录中存在的仪器,以选择其相应的理想光谱作为目标。接下来,计算记录和目标之间的光谱差异,并且相应地,均衡器匹配模型用于预测参数均衡器的设置。为此,我们建立在一个可微参数均衡器匹配神经网络,展示了相对于以前建立的最先进的改进。与过去的方法不同,我们展示了我们的系统如何自然地允许在我们的匹配模型的训练过程中利用真实世界的音频数据,有效地生成适当的生产训练目标在自动化的方式镜像条件在推理时间。因此,我们说明了如何微调我们的匹配模型在这样的例子大大提高了参数均衡器匹配性能在现实世界中的情况下,减少平均绝对误差24%,相对于方法只依赖于随机参数采样技术作为一个自我监督的学习策略。我们进行听力测试,并证明我们提出的自动均衡解决方案主观上增强了常见乐器类型录音的音调特征。摘要:We propose a novel approach for the automatic equalization of individual musical instrument tracks. Our method begins by identifying the instrument present within a source recording in order to choose its corresponding ideal spectrum as a target. Next, the spectral difference between the recording and the target is calculated, and accordingly, an equalizer matching model is used to predict settings for a parametric equalizer. To this end, we build upon a differentiable parametric equalizer matching neural network, demonstrating improvements relative to previously established state-of-the-art. Unlike past approaches, we show how our system naturally allows real-world audio data to be leveraged during the training of our matching model, effectively generating suitably produced training targets in an automated manner mirroring conditions at inference time. Consequently, we illustrate how fine-tuning our matching model on such examples considerably improves parametric equalizer matching performance in real-world scenarios, decreasing mean absolute error by 24% relative to methods relying solely on random parameter sampling techniques as a self-supervised learning strategy. We perform listening tests, and demonstrate that our proposed automatic equalization solution subjectively enhances the tonal characteristics for recordings of common instrument types.

【2】 Synthesizer Sound Matching Using Audio Spectrogram Transformers
标题: 使用音频频谱图变换器的合成器声音匹配
作者:Fred Bruford,Frederik Blang,Shahan Nercessian
备注:4 pages, 1 figure. Accepted to the 27th International Conference on Digital Audio Effects (DAFx24)
链接:点击下载PDF文件
摘要:用于合成器声音匹配的系统,其自动设置合成器的参数以模拟输入声音,具有使合成器编程的过程对于新手和有经验的音乐家来说更快和更容易的潜力,同时还提供与合成器交互的新手段。考虑到市场上各种各样的合成器,以及它们中的许多合成器的复杂性,特别需要以关于底层合成架构的最少知识或先验假设来工作的通用声音匹配系统。考虑到这一点,我们引入了一个基于音频谱图Transformer的合成器声音匹配模型。我们通过在一个由流行的Massive合成器随机生成的样本组成的大型合成数据集上进行训练,证明了该模型的可行性。我们表明,该模型可以重建从一组16个参数生成的样本的参数,突出了其相对于多层感知器和卷积神经网络基线的保真度提高。我们还提供了音频示例,演示了域外模型在模拟人声模仿以及来自其他合成器和乐器的声音时的性能。摘要:Systems for synthesizer sound matching, which automatically set the parameters of a synthesizer to emulate an input sound, have the potential to make the process of synthesizer programming faster and easier for novice and experienced musicians alike, whilst also affording new means of interaction with synthesizers. Considering the enormous variety of synthesizers in the marketplace, and the complexity of many of them, general-purpose sound matching systems that function with minimal knowledge or prior assumptions about the underlying synthesis architecture are particularly desirable. With this in mind, we introduce a synthesizer sound matching model based on the Audio Spectrogram Transformer. We demonstrate the viability of this model by training on a large synthetic dataset of randomly generated samples from the popular Massive synthesizer. We show that this model can reconstruct parameters of samples generated from a set of 16 parameters, highlighting its improved fidelity relative to multi-layer perceptron and convolutional neural network baselines. We also provide audio examples demonstrating the out-of-domain model performance in emulating vocal imitations, and sounds from other synthesizers and musical instruments.

【3】 The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
标题: CHiME-8针对可概括和阵列不可知的远程自动语音识别和规模化的DASB挑战
作者:Samuele Cornell,Taejin Park,Steve Huang,Christoph Boeddeker,Xuankai Chang,Matthew Maciejewski,Matthew Wiesner,Paola Garcia,Shinji Watanabe
链接:点击下载PDF文件
摘要:本文介绍了CHiME-8 DASR挑战赛,该挑战赛是从上一版CHiME-7 DASR(C7 DASR)和过去的CHiME-6挑战赛中进行的。它侧重于联合多通道远程语音识别(DASR)和日记与一个或多个,可能异构,设备。主要目标是促进对会议转录方法的研究,这些方法可以概括任意数量的发言者,不同的设置(正式与非正式对话),会议持续时间,各种各样的声学场景和不同的录音配置。C7 DASR的创新之处包括:i)增加了NOTSOFAR-1,这是一个额外的办公室 企业会议场景,ii)一个手动修正的Mixer 6开发集,iii)一个新的轨道,我们允许使用大语言模型(LLM)iv)一个陪审团奖励机制,以鼓励参与者探索更实用和创新的解决方案。为了降低参与者的进入门槛,我们提供了一个独立的工具包,用于下载和准备这些数据集,以及执行文本规范化和评分。此外,今年我们还提供了两个基线系统,一个直接继承自C7 DASR,基于ESPnet,另一个基于NeMo开发,基于去年C7 DASR的NeMo团队提交。基线系统的结果表明,增加了NOTSOFAR-1的情况下显着增加了任务的难度,由于其大量的发言者和非常短的持续时间。摘要:This paper presents the CHiME-8 DASR challenge which carries on from the previous edition CHiME-7 DASR (C7DASR) and the past CHiME-6 challenge. It focuses on joint multi-channel distant speech recognition (DASR) and diarization with one or more, possibly heterogeneous, devices. The main goal is to spur research towards meeting transcription approaches that can generalize across arbitrary number of speakers, diverse settings (formal vs. informal conversations), meeting duration, wide-variety of acoustic scenarios and different recording configurations. Novelties with respect to C7DASR include: i) the addition of NOTSOFAR-1, an additional office corporate meeting scenario, ii) a manually corrected Mixer 6 development set, iii) a new track in which we allow the use of large-language models (LLM) iv) a jury award mechanism to encourage participants to explore also more practical and innovative solutions. To lower the entry barrier for participants, we provide a standalone toolkit for downloading and preparing such datasets as well as performing text normalization and scoring their submissions. Furthermore, this year we also provide two baseline systems, one directly inherited from C7DASR and based on ESPnet and another one developed on NeMo and based on NeMo team submission in last year C7DASR. Baseline system results suggest that the addition of the NOTSOFAR-1 scenario significantly increases the task's difficulty due to its high number of speakers and very short duration.

【4】 Schrödinger Bridge for Generative Speech Enhancement
标题: 用于生成性语音增强的薛定格桥
作者:Ante Jukić,Roman Korostik,Jagadeesh Balam,Boris Ginsburg
链接:点击下载PDF文件
摘要:提出了一种基于Schr odinger桥(SB)的生成式语音增强模型.该模型采用一个易于处理的SB来制定一个干净的语音分布和观察到的噪声语音分布之间的数据到数据的过程。该模型使用数据预测损失进行训练,旨在恢复复值干净的语音系数,并使用辅助时域损失来改善模型的训练。在两个不同的语音增强任务:语音去噪和语音去混响的建议SB为基础的模型的有效性进行评估。实验结果表明,所提出的基于SB的模型在语音质量度量和ASR性能方面优于基于扩散的模型,例如,导致与最佳基线模型相比,去噪的相对字错误率降低20%,去混响的相对字错误率降低6%。所提出的模型还表明,提高了效率,实现了更好的质量比基线相同数量的采样步骤,并降低了计算成本。摘要:This paper proposes a generative speech enhancement model based on Schr "odinger bridge (SB). The proposed model is employing a tractable SB to formulate a data-to-data process between the clean speech distribution and the observed noisy speech distribution. The model is trained with a data prediction loss, aiming to recover the complex-valued clean speech coefficients, and an auxiliary time-domain loss is used to improve training of the model. The effectiveness of the proposed SB-based model is evaluated in two different speech enhancement tasks: speech denoising and speech dereverberation. The experimental results demonstrate that the proposed SB-based outperforms diffusion-based models in terms of speech quality metrics and ASR performance, e.g., resulting in relative word error rate reduction of 20% for denoising and 6% for dereverberation compared to the best baseline model. The proposed model also demonstrates improved efficiency, achieving better quality than the baselines for the same number of sampling steps and with a reduced computational cost.

【5】 Towards scalable efficient on-device ASR with transfer learning
标题: 通过迁移学习实现可扩展、高效的设备上ASB
作者:Laxmi Pandey,Ke Li,Jinxi Guo,Debjyoti Paul,Arthur Guo,Jay Mahadeokar,Xuedong Zhang
链接:点击下载PDF文件
摘要:迁移学习的多语言预训练显著提高了低资源单语ASR模型的鲁棒性。本研究系统地研究了三个主要方面:(a)在初始训练或微调期间迁移学习对模型性能的影响,(b)跨数据集域和语言的迁移学习的影响,以及(c)与非稀有词相比,对稀有词识别的影响。我们的发现表明,RNT-loss预训练,然后通过最小单词错误率(MinWER)损失进行单语微调,可以持续降低意大利语和法语等语言的单词错误率(WER)。与MLS和内部数据集的单语基线相比,WER减少(WERR)达到36.2%和42.8%。域外预训练比域内预训练的WERR高28%。罕见和非罕见单词都受益,罕见单词在域外预训练中表现出更大的改善,而非罕见单词在域内预训练中表现出更大的改善。摘要:Multilingual pretraining for transfer learning significantly boosts the robustness of low-resource monolingual ASR models. This study systematically investigates three main aspects: (a) the impact of transfer learning on model performance during initial training or fine-tuning, (b) the influence of transfer learning across dataset domains and languages, and (c) the effect on rare-word recognition compared to non-rare words. Our finding suggests that RNNT-loss pretraining, followed by monolingual fine-tuning with Minimum Word Error Rate (MinWER) loss, consistently reduces Word Error Rates (WER) across languages like Italian and French. WER Reductions (WERR) reach 36.2% and 42.8% compared to monolingual baselines for MLS and in-house datasets. Out-of-domain pretraining leads to 28% higher WERR than in-domain pretraining. Both rare and non-rare words benefit, with rare words showing greater improvements with out-of-domain pretraining, and non-rare words with in-domain pretraining.

【6】 Distortion Recovery: A Two-Stage Method for Guitar Effect Removal
标题: 失真恢复:消除吉他效应的两阶段方法
作者:Ying-Shuo Lee,Yueh-Po Peng,Jui-Te Wu,Ming Cheng,Li Su,Yi-Hsuan Yang
备注:DAFx 2024
链接:点击下载PDF文件
摘要:从电吉他录音中删除音频效果可以更轻松地进行后期制作和声音编辑。音频失真恢复模型不仅提高了吉他声音的清晰度,还为混音和母带的创造性调整开辟了新的机会。虽然在创建此类模型方面取得了进展,但以前的努力主要集中在合成失真上,这些失真可能过于简单,无法准确捕捉现实世界记录中的复杂性。 在本文中,我们通过使用商业级音频效果VST插件渲染的吉他录音数据集来解决这个问题。此外,我们介绍了一种新的两阶段的音频失真恢复方法。其思想是首先在第一阶段中在Mel频谱图域中处理音频信号,然后在第二阶段中使用神经声码器从处理后的Mel频谱图生成原始的吉他声音。我们报告了一组实验,通过主观和客观的评价指标,证明了我们的方法对现有方法的有效性。摘要:Removing audio effects from electric guitar recordings makes it easier for post-production and sound editing. An audio distortion recovery model not only improves the clarity of the guitar sounds but also opens up new opportunities for creative adjustments in mixing and mastering. While progress have been made in creating such models, previous efforts have largely focused on synthetic distortions that may be too simplistic to accurately capture the complexities seen in real-world recordings. In this paper, we tackle the task by using a dataset of guitar recordings rendered with commercial-grade audio effect VST plugins. Moreover, we introduce a novel two-stage methodology for audio distortion recovery. The idea is to firstly process the audio signal in the Mel-spectrogram domain in the first stage, and then use a neural vocoder to generate the pristine original guitar sound from the processed Mel-spectrogram in the second stage. We report a set of experiments demonstrating the effectiveness of our approach over existing methods, through both subjective and objective evaluation metrics.

【7】 Audio Prompt Adapter: Unleashing Music Editing Abilities for Text-to-Music with Lightweight Finetuning
标题: 音频提示适配器:通过轻量级微调释放文本到音乐的音乐编辑能力
作者:Fang-Duo Tsai,Shih-Lun Wu,Haven Kim,Bo-Yu Chen,Hao-Chung Cheng,Yi-Hsuan Yang
备注:Accepted by the 25th International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:文本到音乐模型允许用户使用文本命令生成近乎真实的音乐音频。然而,编辑音乐音频仍然具有挑战性,这是由于在维持简单用户界面的同时对音频执行细粒度更改的冲突需求。为了应对这一挑战,我们提出了音频提示适配器(或AP适配器),这是对预训练文本到音乐模型的轻量级补充。我们利用AudioMAE从输入音频中提取特征,并构建基于注意力的适配器将这些特征输入到AudioLDM 2(一种基于扩散的文本到音乐模型)的内部层中。通过22M可训练参数,AP适配器使用户能够利用全局(例如,流派和音色)和本地(例如,旋律)方面,使用原始音频和短文本作为输入。通过客观和主观的研究,我们评估AP适配器的三个任务:音色转移,体裁转移,和伴奏生成。此外,我们证明了它的有效性域外的音频包含看不见的仪器在训练过程中。摘要:Text-to-music models allow users to generate nearly realistic musical audio with textual commands. However, editing music audios remains challenging due to the conflicting desiderata of performing fine-grained alterations on the audio while maintaining a simple user interface. To address this challenge, we propose Audio Prompt Adapter (or AP-Adapter), a lightweight addition to pretrained text-to-music models. We utilize AudioMAE to extract features from the input audio, and construct attention-based adapters to feedthese features into the internal layers of AudioLDM2, a diffusion-based text-to-music model. With 22M trainable parameters, AP-Adapter empowers users to harness both global (e.g., genre and timbre) and local (e.g., melody) aspects of music, using the original audio and a short text as inputs. Through objective and subjective studies, we evaluate AP-Adapter on three tasks: timbre transfer, genre transfer, and accompaniment generation. Additionally, we demonstrate its effectiveness on out-of-domain audios containing unseen instruments during training.

【8】 Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and Localization
标题: 音频时间伪造检测和定位的从粗到细的提案细化框架
作者:Junyan Wu,Wei Lu,Xiangyang Luo,Rui Yang,Qian Wang,Xiaochun Cao
备注:9pages, 3figures. This paper has been accepted for ACM MM 2024
链接:点击下载PDF文件
摘要:最近,一种新形式的音频部分伪造对其取证提出了挑战,需要先进的对策来检测长时间音频内的细微伪造操作。然而,现有的对策仍然服务于分类目的,并且不能对部分伪造段的开始和结束时间戳执行有意义的分析。为了解决这一挑战,我们引入了一种新的粗到细的建议细化框架(CFPRF),它结合了帧级检测网络(FDN)和建议细化网络(PRN)的音频时间伪造检测和定位。具体而言,FDN的目的是挖掘真实和虚假帧之间的信息不一致线索,以获得有利于粗略指示伪造区域的区别性特征。PRN负责预测置信度分数和回归偏移,以细化从FDN导出的粗粒度建议。为了学习鲁棒的判别特征,我们设计了一个差异感知特征学习(DAFL)模块,该模块由对比表示学习指导,以放大由微小操作引起的不同帧之间的敏感差异。我们进一步设计了一个边界感知的特征增强(BAFE)模块来捕获多个过渡边界的上下文信息,并通过交叉注意机制引导边界信息和时间特征之间的交互。大量的实验表明,我们的CFPRF在各种数据集上实现了最先进的性能,包括LAV-DF,ASVS 2019 PS和HAD。摘要:Recently, a novel form of audio partial forgery has posed challenges to its forensics, requiring advanced countermeasures to detect subtle forgery manipulations within long-duration audio. However, existing countermeasures still serve a classification purpose and fail to perform meaningful analysis of the start and end timestamps of partial forgery segments. To address this challenge, we introduce a novel coarse-to-fine proposal refinement framework (CFPRF) that incorporates a frame-level detection network (FDN) and a proposal refinement network (PRN) for audio temporal forgery detection and localization. Specifically, the FDN aims to mine informative inconsistency cues between real and fake frames to obtain discriminative features that are beneficial for roughly indicating forgery regions. The PRN is responsible for predicting confidence scores and regression offsets to refine the coarse-grained proposals derived from the FDN. To learn robust discriminative features, we devise a difference-aware feature learning (DAFL) module guided by contrastive representation learning to enlarge the sensitive differences between different frames induced by minor manipulations. We further design a boundary-aware feature enhancement (BAFE) module to capture the contextual information of multiple transition boundaries and guide the interaction between boundary information and temporal features via a cross-attention mechanism. Extensive experiments show that our CFPRF achieves state-of-the-art performance on various datasets, including LAV-DF, ASVS2019PS, and HAD.

【9】 On the Utility of Speech and Audio Foundation Models for Marmoset Call Analysis
标题: 关于绒猴叫声分析的语音和音频基础模型的实用性
作者:Eklavya Sarkar,Mathew Magimai. -Doss
备注:Accepted at Interspeech 2024 Satellite event (VIHAR 2024)
链接:点击下载PDF文件
摘要:绒猴在它们的叫声中编码重要信息,并作为神经生物学家了解人类声音交流进化起源的替代模型。传统上使用基于信号处理的特征进行分析,最近的方法利用在人类语音上预训练的自监督模型进行特征提取,利用它们独立于声学域学习信号内在结构的能力。然而,在多类分类、带宽和预训练域方面,这些基础模型的效用对于绒猴呼叫分析仍然不清楚。本研究评估了来自语音和一般音频域的特征表示,在4,8和16 kHz的预训练带宽的绒猴呼叫类型和呼叫者分类任务。结果表明,具有更高带宽的模型可以提高性能,并且对语音或一般音频进行预训练会产生相当的结果,在频谱基线上有所改善。摘要:Marmoset monkeys encode vital information in their calls and serve as a surrogate model for neuro-biologists to understand the evolutionary origins of human vocal communication. Traditionally analyzed with signal processing-based features, recent approaches have utilized self-supervised models pre-trained on human speech for feature extraction, capitalizing on their ability to learn a signal's intrinsic structure independently of its acoustic domain. However, the utility of such foundation models remains unclear for marmoset call analysis in terms of multi-class classification, bandwidth, and pre-training domain. This study assesses feature representations derived from speech and general audio domains, across pre-training bandwidths of 4, 8, and 16 kHz for marmoset call-type and caller classification tasks. Results show that models with higher bandwidth improve performance, and pre-training on speech or general audio yields comparable results, improving over a spectral baseline.

【10】 Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
标题: 基于LLM的ASB后错误纠正的进化提示设计
作者:Rithik Sachdev,Zhong-Qiu Wang,Chao-Han Huck Yang
备注:in submission
链接:点击下载PDF文件
摘要:基于现代大型语言模型(LLM)的优势,生成纠错(GEC)已经成为一种有前途的范例,可以提高现代自动语音识别(ASR)系统的性能。一种代表性的方法是利用上下文学习来提示LLM,使得LLM可以基于精心设计的提示和由ASR系统产生的$N$最佳假设列表来生成更好的假设。然而,目前尚不清楚现有的提示是否是最有效的ASR后纠错的任务。在这种情况下,本文首先探讨替代提示,以确定一个初始的有效提示,然后提出采用进化提示优化算法来完善初始提示。在2024 $ GenSEC挑战的任务$1$的CHiME-4子集上的评估结果显示了所提出的算法的有效性和潜力。摘要:Building upon the strength of modern large language models (LLMs), generative error correction (GEC) has emerged as a promising paradigm that can elevate the performance of modern automatic speech recognition (ASR) systems. One representative approach is to leverage in-context learning to prompt LLMs so that a better hypothesis can be generated by the LLMs based on a carefully-designed prompt and an $N$-best list of hypotheses produced by ASR systems. However, it is yet unknown whether the existing prompts are the most effective ones for the task of post-ASR error correction. In this context, this paper first explores alternative prompts to identify an initial set of effective prompts, and then proposes to employ an evolutionary prompt optimization algorithm to refine the initial prompts. Evaluations results on the CHiME-4 subset of the Task $1$ of the SLT $2024$ GenSEC challenge show the effectiveness and potential of the proposed algorithms.

【11】 Multimodal Input Aids a Bayesian Model of Phonetic Learning
标题: 多模式输入辅助语音学习的Bayesian模型
作者:Sophia Zhi,Roger P. Levy,Stephan C. Meylan
备注:12 pages, 5 figures
链接:点击下载PDF文件
摘要:典型的儿童语言学习者面临的许多任务之一是学习区分构成母语单词的独特声音。在这里,我们调查是否多模态信息-特别是成人语音加上视频帧的扬声器的脸-有利于语音学习的计算模型。我们介绍了一种方法,用于创建高质量的合成视频扬声器的脸为现有的音频语料库。我们的学习模型,在视听输入上训练和测试时,与仅在音频输入上训练和测试的模型相比,在音素辨别电池上实现了高达8.1%的相对改善。当两者都在纯音频数据上进行测试时,它的表现也优于音频模型高达3.9%,这表明视觉信息有助于获得声学区别。视觉信息在嘈杂的音频环境中特别有益,其中视听模型相对于无噪声环境关闭了噪声中音频模型的辨别性能损失的67%。这些结果表明,视觉信息有利于一个理想的学习者,并说明了一些方式,儿童可能能够利用视觉线索时,学习区分语音。摘要:One of the many tasks facing the typically-developing child language learner is learning to discriminate between the distinctive sounds that make up words in their native language. Here we investigate whether multimodal information--specifically adult speech coupled with video frames of speakers' faces--benefits a computational model of phonetic learning. We introduce a method for creating high-quality synthetic videos of speakers' faces for an existing audio corpus. Our learning model, when both trained and tested on audiovisual inputs, achieves up to a 8.1% relative improvement on a phoneme discrimination battery compared to a model trained and tested on audio-only input. It also outperforms the audio model by up to 3.9% when both are tested on audio-only data, suggesting that visual information facilitates the acquisition of acoustic distinctions. Visual information is especially beneficial in noisy audio environments, where an audiovisual model closes 67% of the loss in discrimination performance of the audio model in noise relative to a non-noisy environment. These results demonstrate that visual information benefits an ideal learner and illustrate some of the ways that children might be able to leverage visual cues when learning to discriminate speech sounds.

【12】 Reading Miscue Detection in Primary School through Automatic Speech Recognition
标题: 基于自动语音识别的小学阅读错误检测
作者:Lingyun Gao,Cristian Tejedor-Garcia,Helmer Strik,Catia Cucchiarini
备注:Proc. INTERSPEECH 2024, 1-5 September 2024. Kos Island, Greece
链接:点击下载PDF文件
摘要:自动阅读诊断系统可以使教师更有效地对阅读练习进行评分,也可以使学生更容易地获得反馈。然而,有有限的研究自动语音识别(ASR)的儿童语音以外的其他语言,并有限的研究,基于ASR的阅读诊断系统。本研究调查了最先进(SOTA)预训练的ASR模型如何有效地识别荷兰本土儿童的语音并设法检测阅读错误。我们发现,Hubert Large对荷兰语语音进行了微调,实现了SOTA音素级别的儿童语音识别(PER为23.1%),而Whisper(Faster Whisper Large-v2)实现了SOTA单词级别的性能(WER为9.8%)。我们的研究结果表明,Wav 2 Vec 2 Large和Whisper是两个最好的阅读错误检测ASR模型。具体来说,Wav 2 Vec 2 Large的查全率最高,为0.83,而Whisper的查准率最高,为0.52,F1得分为0.52。摘要:Automatic reading diagnosis systems can benefit both teachers for more efficient scoring of reading exercises and students for accessing reading exercises with feedback more easily. However, there are limited studies on Automatic Speech Recognition (ASR) for child speech in languages other than English, and limited research on ASR-based reading diagnosis systems. This study investigates how efficiently state-of-the-art (SOTA) pretrained ASR models recognize Dutch native children speech and manage to detect reading miscues. We found that Hubert Large finetuned on Dutch speech achieves SOTA phoneme-level child speech recognition (PER at 23.1 %), while Whisper (Faster Whisper Large-v2) achieves SOTA word-level performance (WER at 9.8 %). Our findings suggest that Wav2Vec2 Large and Whisper are the two best ASR models for reading miscue detection. Specifically, Wav2Vec2 Large shows the highest recall at 0.83, whereas Whisper exhibits the highest precision at 0.52 and an F1 score of 0.52.

【13】 Automatic Speech Recognition of Non-Native Child Speech for Language Learning Applications
标题: 用于语言学习应用的非母语儿童语音自动语音识别
作者:Simone Wills,Yu Bai,Cristian Tejedor-Garcia,Catia Cucchiarini,Helmer Strik
Journal-ref:12th Symposium on Languages, Applications and Technologies (SLATE 2023) (p. 7:1-7:8)
链接:点击下载PDF文件
摘要:语音机器人为支持语言技能的发展提供了新的途径,特别是在第二语言学习的背景下。不过,语音机器人在很大程度上是面向母语为成年人的。我们试图评估两个最先进的ASR系统Wav2Vec2.0和Whisper AI的性能,以期开发一个可以支持儿童学习外语的语音机器人。我们评估了他们的表现,阅读和即兴演讲的本地和非本地荷兰儿童。我们还调查了使用ASR技术来深入了解儿童的发音和流畅性的实用性。结果表明,最近,预训练的ASR变换为基础的模型达到了可接受的性能,从其中可以提取音素发音质量的详细反馈,尽管具有挑战性的性质的儿童和非母语语音。摘要:Voicebots have provided a new avenue for supporting the development of language skills, particularly within the context of second language learning. Voicebots, though, have largely been geared towards native adult speakers. We sought to assess the performance of two state-of-the-art ASR systems, Wav2Vec2.0 and Whisper AI, with a view to developing a voicebot that can support children acquiring a foreign language. We evaluated their performance on read and extemporaneous speech of native and non-native Dutch children. We also investigated the utility of using ASR technology to provide insight into the children's pronunciation and fluency. The results show that recent, pre-trained ASR transformer-based models achieve acceptable performance from which detailed feedback on phoneme pronunciation quality can be extracted, despite the challenging nature of child and non-native speech.

【14】 An ASR-Based Tutor for Learning to Read: How to Optimize Feedback to First Graders
标题: 基于ASB的阅读学习导师:如何优化对一年级学生的反馈
作者:Yu Bai,Cristian Tejedor-Garcia,Ferdy Hubers,Catia Cucchiarini,Helmer Strik
Journal-ref:In: Karpov A., Potapova R. (eds) Speech and Computer. SPECOM 2021. Lecture Notes in Computer Science, vol 12997. Springer, Cham
链接:点击下载PDF文件
摘要:近年来,在阅读实践中使用自动语音识别(ASR)的兴趣一直在增长。在以前的研究中,我们提出了一个基于ASR的荷兰语阅读辅导应用程序,该应用程序的开发是为了向一年级学生学习阅读提供即时反馈。我们看到ASR在阅读过程的这个阶段具有潜力,因为结果表明,学生通过使用该软件在阅读准确性和流畅性方面取得了进步。在目前的研究中,我们使用儿童的语音从现有的语料库(JASMIN)开发两个新的ASR系统,并比较的结果,以前的研究。我们分析正确 不正确的分类的ASR系统使用人类成绩单在单词水平上,通过评价措施,如科恩的Kappa,马修斯相关系数(MCC),精度,召回和F-措施。我们观察到新开发的ASR系统在与基于人类的判断和正确拒绝(CR)的协议方面的改进。ASR系统的准确性因不同的阅读任务和单词类型而异。我们的研究结果表明,在目前的配置,这是很难分类孤立的话。我们讨论这些结果,可能的方法来改善我们的系统和未来的研究途径。摘要:The interest in employing automatic speech recognition (ASR) in applications for reading practice has been growing in recent years. In a previous study, we presented an ASR-based Dutch reading tutor application that was developed to provide instantaneous feedback to first-graders learning to read. We saw that ASR has potential at this stage of the reading process, as the results suggested that pupils made progress in reading accuracy and fluency by using the software. In the current study, we used children's speech from an existing corpus (JASMIN) to develop two new ASR systems, and compared the results to those of the previous study. We analyze correct incorrect classification of the ASR systems using human transcripts at word level, by means of evaluation measures such as Cohen's Kappa, Matthews Correlation Coefficient (MCC), precision, recall and F-measures. We observe improvements for the newly developed ASR systems regarding the agreement with human-based judgment and correct rejection (CR). The accuracy of the ASR systems varies for different reading tasks and word types. Our results suggest that, in the current configuration, it is difficult to classify isolated words. We discuss these results, possible ways to improve our systems and avenues for future research.

【15】 Automatic Assessment of Oral Reading Accuracy for Reading Diagnostics
标题: 阅读诊断中口语阅读准确性的自动评估
作者:Bo Molenaar,Cristian Tejedor-Garcia,Helmer Strik,Catia Cucchiarini
Journal-ref:Proc. INTERSPEECH 2023, pp. 5232-5236, Dublin, Ireland, 20-24, August 2023
链接:点击下载PDF文件
摘要:使用自动语音识别(ASR)的阅读流畅性的自动评估具有很大的潜力,早期发现阅读困难和随后的及时干预。需要精确的评估工具,特别是对于英语以外的语言。在这项研究中,我们评估了六个国家的最先进的ASR为基础的系统自动评估荷兰语口语阅读的准确性,使用Kaldi和耳语。结果显示,我们最成功的系统与人类评估(MCC = .63)达成了实质性的一致。相同的系统达到了最高的相关性之间的强制解码的信心分数和单词的正确性(r = 0.45)。该系统的语言模型(LM)由测试数据的人工拼写转换和阅读提示组成,这表明在LM中包含阅读错误可以提高评估性能。我们讨论了开发自动评估系统的影响,并确定未来研究的可能途径。摘要:Automatic assessment of reading fluency using automatic speech recognition (ASR) holds great potential for early detection of reading difficulties and subsequent timely intervention. Precise assessment tools are required, especially for languages other than English. In this study, we evaluate six state-of-the-art ASR-based systems for automatically assessing Dutch oral reading accuracy using Kaldi and Whisper. Results show our most successful system reached substantial agreement with human evaluations (MCC = .63). The same system reached the highest correlation between forced decoding confidence scores and word correctness (r = .45). This system's language model (LM) consisted of manual orthographic transcriptions and reading prompts of the test data, which shows that including reading errors in the LM improves assessment performance. We discuss the implications for developing automatic assessment systems and identify possible avenues of future research.

【16】 Alzheimer Disease Classification through ASR-based Transcriptions: Exploring the Impact of Punctuation and Pauses
标题: 通过基于SVR的Transit进行阿尔茨海默病分类:探索标点和停顿的影响
作者:Lucía Gómez-Zaragozá,Simone Wills,Cristian Tejedor-Garcia,Javier Marín-Morales,Mariano Alcañiz,Helmer Strik
Journal-ref:Proc. INTERSPEECH 2023, pp. 2403-2407, Dublin, Ireland, 20-24, August 2023
链接:点击下载PDF文件
摘要:阿尔茨海默氏病(AD)是世界上最主要的神经退行性疾病,常导致沟通困难。分析语音可以作为一种诊断工具来识别病情。最近的ADReSS挑战为AD分类提供了数据集,并强调了手动transmittance的实用性。在这项研究中,我们使用了新的国家的最先进的自动语音识别(ASR)模型耳语,以获得transmittance,其中还包括自动标点符号。分类模型分别在手动和ASR成绩单上实现了0.854和0.833的测试准确度分数,并结合了预训练的FastText单词嵌入和递归神经网络。此外,我们还探讨了停顿信息和标点符号在音变中的作用。我们发现,在某些情况下,标点符号只产生轻微的改善,而暂停编码辅助AD分类手动和ASR transmittance在所有的方法研究。摘要:Alzheimer's Disease (AD) is the world's leading neurodegenerative disease, which often results in communication difficulties. Analysing speech can serve as a diagnostic tool for identifying the condition. The recent ADReSS challenge provided a dataset for AD classification and highlighted the utility of manual transcriptions. In this study, we used the new state-of-the-art Automatic Speech Recognition (ASR) model Whisper to obtain the transcriptions, which also include automatic punctuation. The classification models achieved test accuracy scores of 0.854 and 0.833 combining the pretrained FastText word embeddings and recurrent neural networks on manual and ASR transcripts respectively. Additionally, we explored the influence of including pause information and punctuation in the transcriptions. We found that punctuation only yielded minor improvements in some cases, whereas pause encoding aided AD classification for both manual and ASR transcriptions across all approaches investigated.


机器翻译,仅供参考