今日论文合集:cs.SD语音5篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】A Survey of Music Generation in the Context of Interaction

标题:互动语境中的音乐生成研究综述

链接:https://arxiv.org/abs/2402.15294

作者:Ismael Agchar,Ilja Baumann,Franziska Braun,Paula Andrea Perez-Toro,Korbinian Riedhammer,Sebastian Trump,Martin Ullrich

摘要:近年来,机器学习,特别是生成对抗神经网络(GANs)和基于注意力的神经网络(Transformers),已成功用于创作和生成音乐,包括旋律和复调作品。目前的研究主要集中在风格复制(如。产生巴赫风格的赞美诗)或风格转移(例如,古典乐到爵士乐),其基于大量记录或转录的音乐,这反过来也允许相当直接的“性能”评估。然而,这些模型中的大多数都不适合通过实时交互进行人机共同创作,也不清楚如何评估这些模型和由此产生的创作。本文对音乐表示、特征分析、启发式算法、统计和参数建模以及人工和自动评估措施进行了全面的回顾,并讨论了哪些方法和模型似乎最适合现场交互。

摘要:In recent years, machine learning, and in particular generative adversarial neural networks (GANs) and attention-based neural networks (transformers), have been successfully used to compose and generate music, both melodies and polyphonic pieces. Current research focuses foremost on style replication (eg. generating a Bach-style chorale) or style transfer (eg. classical to jazz) based on large amounts of recorded or transcribed music, which in turn also allows for fairly straight-forward "performance" evaluation. However, most of these models are not suitable for human-machine co-creation through live interaction, neither is clear, how such models and resulting creations would be evaluated. This article presents a thorough review of music representation, feature analysis, heuristic algorithms, statistical and parametric modelling, and human and automatic evaluation measures, along with a discussion of which approaches and models seem most suitable for live interaction.


【2】Human Brain Exhibits Distinct Patterns When Listening to Fake Versus  Real Audio: Preliminary Evidence
标题:人类大脑在听假音频和听真音频时表现出不同的模式:初步证据
链接:https://arxiv.org/abs/2402.14982
作者:Mahsa Salehi,Kalin Stefanov,Ehsan Shareghi
备注:9 pages, 4 figures, 3 tables
摘要:在本文中,我们研究了人类大脑活动的变化时,听真实和虚假的音频。我们的初步结果表明,通过最先进的deepfake音频检测算法学习的表示在真实和虚假音频之间没有表现出明显的模式。相比之下,当个体暴露于假音频与真实音频时,通过EEG测量的人类大脑活动显示出不同的模式。这一初步证据为deepfake音频检测等领域的未来研究方向奠定了基础。
摘要:In this paper we study the variations in human brain activity when listening to real and fake audio. Our preliminary results suggest that the representations learned by a state-of-the-art deepfake audio detection algorithm, do not exhibit clear distinct patterns between real and fake audio. In contrast, human brain activity, as measured by EEG, displays distinct patterns when individuals are exposed to fake versus real audio. This preliminary evidence enables future research directions in areas such as deepfake audio detection.


【3】All Thresholds Barred: Direct Estimation of Call Density in Bioacoustic  Data
标题:禁止的所有阈值:直接估计生物声学数据中的呼叫密度
链接:https://arxiv.org/abs/2402.15360
作者:Amanda K. Navine,Tom Denton,Matthew J. Weldy,Patrick J. Hart
备注:14 pages, 6 figures, 3 tables; submitted to Frontiers in Bird Science; Our Hawaiian PAM dataset and classifier scores, as well as annotation information for the three study species, can be found on Zenodo at this https URL The fully annotated Powdermill dataset assembled by Chronister et al. that was used in this study is available at this https URL
摘要:被动声学监测(PAM)研究产生数千小时的音频,可用于监测特定的动物种群,进行广泛的生物多样性调查,检测偷猎者等威胁。用于物种识别的机器学习分类器越来越多地用于处理生物声学调查产生的大量音频,加快分析并提高PAM作为管理工具的实用性。在通常的实践中,阈值被应用于分类器输出分数,并且高于阈值的分数被聚合到检测计数中。阈值的选择产生了有偏差的发声计数,其受到可能在数据集的子集之间变化的假阳性/阴性率的影响。在这项工作中,我们主张直接估计呼叫密度:包含目标发声的检测窗口的比例,无论分类器得分如何。我们的方法针对一个理想的生态估计,并提供了一个更严格的基础,以确定由分布变化所造成的核心问题-当数据分布的定义特征发生变化时-并设计策略来减轻它们。我们提出了一个验证计划,估计调用密度的数据体,并获得,通过贝叶斯推理,概率分布的置信度得分的积极和消极的类。我们使用这些分布来预测站点级别的密度,这可能会受到分布的变化。我们在夏威夷鸟类的真实世界研究中测试了我们提出的方法,并利用现有的完全注释的数据集提供了模拟结果,证明了对呼叫密度和分类器模型质量变化的鲁棒性。
摘要:Passive acoustic monitoring (PAM) studies generate thousands of hours of audio, which may be used to monitor specific animal populations, conduct broad biodiversity surveys, detect threats such as poachers, and more. Machine learning classifiers for species identification are increasingly being used to process the vast amount of audio generated by bioacoustic surveys, expediting analysis and increasing the utility of PAM as a management tool. In common practice, a threshold is applied to classifier output scores, and scores above the threshold are aggregated into a detection count. The choice of threshold produces biased counts of vocalizations, which are subject to false positive/negative rates that may vary across subsets of the dataset. In this work, we advocate for directly estimating call density: The proportion of detection windows containing the target vocalization, regardless of classifier score. Our approach targets a desirable ecological estimator and provides a more rigorous grounding for identifying the core problems caused by distribution shifts -- when the defining characteristics of the data distribution change -- and designing strategies to mitigate them. We propose a validation scheme for estimating call density in a body of data and obtain, through Bayesian reasoning, probability distributions of confidence scores for both the positive and negative classes. We use these distributions to predict site-level densities, which may be subject to distribution shifts. We test our proposed methods on a real-world study of Hawaiian birds and provide simulation results leveraging existing fully annotated datasets, demonstrating robustness to variations in call density and classifier model quality.


【4】High Resolution Guitar Transcription via Domain Adaptation
标题:基于结构域适配的高分辨率吉他转录
链接:https://arxiv.org/abs/2402.15258
作者:Xavier Riley,Drew Edwards,Simon Dixon
备注:Accepted to ICASSP 2024
摘要:由于MAESTRO和MAPS等大型高质量数据集的可用性,自动音乐转录(AMT)已经实现了钢琴的高准确性,但其他乐器尚未提供类似的数据集。然而,在最近的工作中,已经证明将分数与转录模型激活对齐可以为钢琴以外的乐器产生高质量的AMT训练数据。专注于吉他,我们改进了这种方法,使用商业上可用的乐谱-音频对的数据集对乐谱数据进行训练。我们建议使用高分辨率钢琴转录模型来训练新的吉他转录模型。由此产生的模型在zero-shot上下文中获得了GuitarSet上最先进的转录结果,改进了以前发表的方法。
摘要:Automatic music transcription (AMT) has achieved high accuracy for piano due to the availability of large, high-quality datasets such as MAESTRO and MAPS, but comparable datasets are not yet available for other instruments. In recent work, however, it has been demonstrated that aligning scores to transcription model activations can produce high quality AMT training data for instruments other than piano. Focusing on the guitar, we refine this approach to training on score data using a dataset of commercially available score-audio pairs. We propose the use of a high-resolution piano transcription model to train a new guitar transcription model. The resulting model obtains state-of-the-art transcription results on GuitarSet in a zero-shot context, improving on previously published methods.

【5】ChildAugment: Data Augmentation Methods for Zero-Resource Children's  Speaker Verification
标题:ChildAugment:零资源儿童说话人确认的数据扩充方法
链接:https://arxiv.org/abs/2402.15214
作者:Vishwanath Pratap Singh,Md Sahidullah,Tomi Kinnunen
备注:The following article has been accepted by The Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this https URL
摘要:现代自动说话人验证(ASV)系统的准确性,当专门训练成人数据,大幅下降时,适用于儿童的讲话。儿童语音语料库的缺乏阻碍了针对儿童语音的ASV系统的微调。因此,有必要及时探索更有效的方法来重用成人的语音数据。一种有希望的方法是通过儿童特定的数据增强来调整成人和儿童之间的声道参数,这里称为ChildAugment。具体来说,我们修改成人语音的共振峰频率和共振峰带宽,以模仿儿童的语音。利用改进后的谱图训练ECAPA-TDNN(Enhanced Channel Attention,Propagation,and Aggregation in Time Delay Neural Network)识别器。我们将ChildAugment与各种最先进的儿童ASV数据增强技术进行比较。我们还广泛比较了不同的评分方法,包括余弦评分,PLDA(概率线性判别分析)和NPLDA(神经PLDA)。我们还提出了一个低复杂度加权余弦得分极低的资源儿童ASV。我们对CSLU儿童语料库的研究结果表明,ChildAugment有望成为一种简单的声学激励方法,用于改善儿童最先进的基于深度学习的ASV。我们实现了高达12.45%(男孩)和11.96%(女孩)相对于基线的改善。
摘要:The accuracy of modern automatic speaker verification (ASV) systems, when trained exclusively on adult data, drops substantially when applied to children's speech. The scarcity of children's speech corpora hinders fine-tuning ASV systems for children's speech. Hence, there is a timely need to explore more effective ways of reusing adults' speech data. One promising approach is to align vocal-tract parameters between adults and children through children-specific data augmentation, referred here to as ChildAugment. Specifically, we modify the formant frequencies and formant bandwidths of adult speech to emulate children's speech. The modified spectra are used to train ECAPA-TDNN (emphasized channel attention, propagation, and aggregation in time-delay neural network) recognizer for children. We compare ChildAugment against various state-of-the-art data augmentation techniques for children's ASV. We also extensively compare different scoring methods, including cosine scoring, PLDA (probabilistic linear discriminant analysis), and NPLDA (neural PLDA). We also propose a low-complexity weighted cosine score for extremely low-resource children ASV. Our findings on the CSLU kids corpus indicate that ChildAugment holds promise as a simple, acoustics-motivated approach, for improving state-of-the-art deep learning based ASV for children. We achieve up to 12.45% (boys) and 11.96% (girls) relative improvement over the baseline.

eess.AS音频处理
【1】High Resolution Guitar Transcription via Domain Adaptation
标题:基于结构域适配的高分辨率吉他转录
链接:https://arxiv.org/abs/2402.15258
作者:Xavier Riley,Drew Edwards,Simon Dixon
备注:Accepted to ICASSP 2024
摘要:由于MAESTRO和MAPS等大型高质量数据集的可用性,自动音乐转录(AMT)已经实现了钢琴的高准确性,但其他乐器尚未提供类似的数据集。然而,在最近的工作中,已经证明将分数与转录模型激活对齐可以为钢琴以外的乐器产生高质量的AMT训练数据。专注于吉他,我们改进了这种方法,使用商业上可用的乐谱-音频对的数据集对乐谱数据进行训练。我们建议使用高分辨率钢琴转录模型来训练新的吉他转录模型。由此产生的模型在zero-shot上下文中获得了GuitarSet上最先进的转录结果,改进了以前发表的方法。
摘要:Automatic music transcription (AMT) has achieved high accuracy for piano due to the availability of large, high-quality datasets such as MAESTRO and MAPS, but comparable datasets are not yet available for other instruments. In recent work, however, it has been demonstrated that aligning scores to transcription model activations can produce high quality AMT training data for instruments other than piano. Focusing on the guitar, we refine this approach to training on score data using a dataset of commercially available score-audio pairs. We propose the use of a high-resolution piano transcription model to train a new guitar transcription model. The resulting model obtains state-of-the-art transcription results on GuitarSet in a zero-shot context, improving on previously published methods.

【2】ChildAugment: Data Augmentation Methods for Zero-Resource Children's  Speaker Verification
标题:ChildAugment:零资源儿童说话人确认的数据扩充方法
链接:https://arxiv.org/abs/2402.15214
作者:Vishwanath Pratap Singh,Md Sahidullah,Tomi Kinnunen
备注:The following article has been accepted by The Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this https URL
摘要:现代自动说话人验证(ASV)系统的准确性,当专门训练成人数据,大幅下降时,适用于儿童的讲话。儿童语音语料库的缺乏阻碍了针对儿童语音的ASV系统的微调。因此,有必要及时探索更有效的方法来重用成人的语音数据。一种有希望的方法是通过儿童特定的数据增强来调整成人和儿童之间的声道参数,这里称为ChildAugment。具体来说,我们修改成人语音的共振峰频率和共振峰带宽,以模仿儿童的语音。利用改进后的谱图训练ECAPA-TDNN(Enhanced Channel Attention,Propagation,and Aggregation in Time Delay Neural Network)识别器。我们将ChildAugment与各种最先进的儿童ASV数据增强技术进行比较。我们还广泛比较了不同的评分方法,包括余弦评分,PLDA(概率线性判别分析)和NPLDA(神经PLDA)。我们还提出了一个低复杂度加权余弦得分极低的资源儿童ASV。我们对CSLU儿童语料库的研究结果表明,ChildAugment有望成为一种简单的声学激励方法,用于改善儿童最先进的基于深度学习的ASV。我们实现了高达12.45%(男孩)和11.96%(女孩)相对于基线的改善。
摘要:The accuracy of modern automatic speaker verification (ASV) systems, when trained exclusively on adult data, drops substantially when applied to children's speech. The scarcity of children's speech corpora hinders fine-tuning ASV systems for children's speech. Hence, there is a timely need to explore more effective ways of reusing adults' speech data. One promising approach is to align vocal-tract parameters between adults and children through children-specific data augmentation, referred here to as ChildAugment. Specifically, we modify the formant frequencies and formant bandwidths of adult speech to emulate children's speech. The modified spectra are used to train ECAPA-TDNN (emphasized channel attention, propagation, and aggregation in time-delay neural network) recognizer for children. We compare ChildAugment against various state-of-the-art data augmentation techniques for children's ASV. We also extensively compare different scoring methods, including cosine scoring, PLDA (probabilistic linear discriminant analysis), and NPLDA (neural PLDA). We also propose a low-complexity weighted cosine score for extremely low-resource children ASV. Our findings on the CSLU kids corpus indicate that ChildAugment holds promise as a simple, acoustics-motivated approach, for improving state-of-the-art deep learning based ASV for children. We achieve up to 12.45% (boys) and 11.96% (girls) relative improvement over the baseline.


【3】All Thresholds Barred: Direct Estimation of Call Density in Bioacoustic  Data
标题:禁止的所有阈值:直接估计生物声学数据中的呼叫密度
链接:https://arxiv.org/abs/2402.15360
作者:Amanda K. Navine,Tom Denton,Matthew J. Weldy,Patrick J. Hart
备注:14 pages, 6 figures, 3 tables; submitted to Frontiers in Bird Science; Our Hawaiian PAM dataset and classifier scores, as well as annotation information for the three study species, can be found on Zenodo at this https URL The fully annotated Powdermill dataset assembled by Chronister et al. that was used in this study is available at this https URL
摘要:被动声学监测(PAM)研究产生数千小时的音频,可用于监测特定的动物种群,进行广泛的生物多样性调查,检测偷猎者等威胁。用于物种识别的机器学习分类器越来越多地用于处理生物声学调查产生的大量音频,加快分析并提高PAM作为管理工具的实用性。在通常的实践中,阈值被应用于分类器输出分数,并且高于阈值的分数被聚合到检测计数中。阈值的选择产生了有偏差的发声计数,其受到可能在数据集的子集之间变化的假阳性/阴性率的影响。在这项工作中,我们主张直接估计呼叫密度:包含目标发声的检测窗口的比例,无论分类器得分如何。我们的方法针对一个理想的生态估计,并提供了一个更严格的基础,以确定由分布变化所造成的核心问题-当数据分布的定义特征发生变化时-并设计策略来减轻它们。我们提出了一个验证计划,估计调用密度的数据体,并获得,通过贝叶斯推理,概率分布的置信度得分的积极和消极的类。我们使用这些分布来预测站点级别的密度,这可能会受到分布的变化。我们在夏威夷鸟类的真实世界研究中测试了我们提出的方法,并利用现有的完全注释的数据集提供了模拟结果,证明了对呼叫密度和分类器模型质量变化的鲁棒性。
摘要:Passive acoustic monitoring (PAM) studies generate thousands of hours of audio, which may be used to monitor specific animal populations, conduct broad biodiversity surveys, detect threats such as poachers, and more. Machine learning classifiers for species identification are increasingly being used to process the vast amount of audio generated by bioacoustic surveys, expediting analysis and increasing the utility of PAM as a management tool. In common practice, a threshold is applied to classifier output scores, and scores above the threshold are aggregated into a detection count. The choice of threshold produces biased counts of vocalizations, which are subject to false positive/negative rates that may vary across subsets of the dataset. In this work, we advocate for directly estimating call density: The proportion of detection windows containing the target vocalization, regardless of classifier score. Our approach targets a desirable ecological estimator and provides a more rigorous grounding for identifying the core problems caused by distribution shifts -- when the defining characteristics of the data distribution change -- and designing strategies to mitigate them. We propose a validation scheme for estimating call density in a body of data and obtain, through Bayesian reasoning, probability distributions of confidence scores for both the positive and negative classes. We use these distributions to predict site-level densities, which may be subject to distribution shifts. We test our proposed methods on a real-world study of Hawaiian birds and provide simulation results leveraging existing fully annotated datasets, demonstrating robustness to variations in call density and classifier model quality.

【4】A Survey of Music Generation in the Context of Interaction
标题:互动语境中的音乐生成研究综述
链接:https://arxiv.org/abs/2402.15294
作者:Ismael Agchar,Ilja Baumann,Franziska Braun,Paula Andrea Perez-Toro,Korbinian Riedhammer,Sebastian Trump,Martin Ullrich
摘要:近年来,机器学习,特别是生成对抗神经网络(GANs)和基于注意力的神经网络(Transformers),已成功用于创作和生成音乐,包括旋律和复调作品。目前的研究主要集中在风格复制(如。产生巴赫风格的赞美诗)或风格转移(例如,古典乐到爵士乐),其基于大量记录或转录的音乐,这反过来也允许相当直接的“性能”评估。然而,这些模型中的大多数都不适合通过实时交互进行人机共同创作,也不清楚如何评估这些模型和由此产生的创作。本文对音乐表示、特征分析、启发式算法、统计和参数建模以及人工和自动评估措施进行了全面的回顾,并讨论了哪些方法和模型似乎最适合现场交互。
摘要:In recent years, machine learning, and in particular generative adversarial neural networks (GANs) and attention-based neural networks (transformers), have been successfully used to compose and generate music, both melodies and polyphonic pieces. Current research focuses foremost on style replication (eg. generating a Bach-style chorale) or style transfer (eg. classical to jazz) based on large amounts of recorded or transcribed music, which in turn also allows for fairly straight-forward "performance" evaluation. However, most of these models are not suitable for human-machine co-creation through live interaction, neither is clear, how such models and resulting creations would be evaluated. This article presents a thorough review of music representation, feature analysis, heuristic algorithms, statistical and parametric modelling, and human and automatic evaluation measures, along with a discussion of which approaches and models seem most suitable for live interaction.

【5】Where Visual Speech Meets Language: VSP-LLM Framework for Efficient and  Context-Aware Visual Speech Processing
标题:视觉语音与语言相遇的地方:用于高效和上下文感知视觉语音处理的VSP-LLM框架
链接:https://arxiv.org/abs/2402.15151
作者:Jeong Hun Yeo,Seunghee Han,Minsu Kim,Yong Man Ro
摘要:在视觉语音处理中,由于嘴唇运动的模糊性,上下文建模能力是最重要的要求之一。例如,同音异义词,即具有相同的嘴唇运动但产生不同声音的单词,可以通过考虑上下文来区分。在本文中,我们提出了一个新的框架,即视觉语音处理与LLM(VSP-LLM),最大限度地提高上下文建模能力,带来压倒性的LLM的力量。具体而言,VSP-LLM被设计为执行视觉语音识别和翻译的多任务,其中给定的指令控制任务的类型。输入视频映射到LLM的输入潜在空间,通过采用自监督的视觉语音模型。针对输入帧中存在冗余信息这一事实,提出了一种新的去重方法,该方法通过使用视觉语音单元来减少嵌入的视觉特征。通过所提出的去重和低秩适配器(LoRA),可以以计算高效的方式训练VSP-LLM。在翻译数据集中,MuAViC基准测试,我们证明VSP-LLM可以更有效地识别和翻译嘴唇运动,只需15小时的标记数据,而最近的翻译模型训练了433小时的标记数据。
摘要:In visual speech processing, context modeling capability is one of the most important requirements due to the ambiguous nature of lip movements. For example, homophenes, words that share identical lip movements but produce different sounds, can be distinguished by considering the context. In this paper, we propose a novel framework, namely Visual Speech Processing incorporated with LLMs (VSP-LLM), to maximize the context modeling ability by bringing the overwhelming power of LLMs. Specifically, VSP-LLM is designed to perform multi-tasks of visual speech recognition and translation, where the given instructions control the type of task. The input video is mapped to the input latent space of a LLM by employing a self-supervised visual speech model. Focused on the fact that there is redundant information in input frames, we propose a novel deduplication method that reduces the embedded visual features by employing visual speech units. Through the proposed deduplication and Low Rank Adaptors (LoRA), VSP-LLM can be trained in a computationally efficient manner. In the translation dataset, the MuAViC benchmark, we demonstrate that VSP-LLM can more effectively recognize and translate lip movements with just 15 hours of labeled data, compared to the recent translation model trained with 433 hours of labeld data.


【6】Human Brain Exhibits Distinct Patterns When Listening to Fake Versus  Real Audio: Preliminary Evidence
标题:人类大脑在听假音频和听真音频时表现出不同的模式:初步证据
链接:https://arxiv.org/abs/2402.14982
作者:Mahsa Salehi,Kalin Stefanov,Ehsan Shareghi
备注:9 pages, 4 figures, 3 tables
摘要:在本文中,我们研究了人类大脑活动的变化时,听真实和虚假的音频。我们的初步结果表明,通过最先进的deepfake音频检测算法学习的表示在真实和虚假音频之间没有表现出明显的模式。相比之下,当个体暴露于假音频与真实音频时,通过EEG测量的人类大脑活动显示出不同的模式。这一初步证据为deepfake音频检测等领域的未来研究方向奠定了基础。
摘要:In this paper we study the variations in human brain activity when listening to real and fake audio. Our preliminary results suggest that the representations learned by a state-of-the-art deepfake audio detection algorithm, do not exhibit clear distinct patterns between real and fake audio. In contrast, human brain activity, as measured by EEG, displays distinct patterns when individuals are exposed to fake versus real audio. This preliminary evidence enables future research directions in areas such as deepfake audio detection.


机器翻译由腾讯交互翻译提供,仅供参考