今天跟大家分享一篇语音相关的论文合集:cs.SD语音4篇,eess.AS音频处理4篇。本文经arXiv每日学术速递授权转载
【1】 Detecting the Severity of Major Depressive Disorder from Speech: A Novel HARD-Training Methodology
标题:从言语中检测重度抑郁障碍的严重程度:一种新的硬训练方法
链接:https://arxiv.org/abs/2206.01542
作者:Edward L. Campbell,Judith Dineley,Pauline Conde,Faith Matcham,Femke Lamers,Sara Siddi,Laura Docio-Fernandez,Carmen Garcia-Mateo,Nicholas Cummins,the RADAR-CNS Consortium摘要:重度抑郁障碍(MDD)是世界范围内常见的心理健康问题,其社会经济成本较高。因此,MDD的预测和自动检测可以对社会产生巨大影响。语音作为一种非侵入性、易于采集的信号,是一种很有前途的辅助诊断和评估MDD的指标。在这方面,作为远程评估重度抑郁症(RADAR-MDD)研究计划的一部分,收集了语音样本。RADAR-MDD是一项观察性队列研究,从西班牙、英国和荷兰有MDD病史的人群中收集语音和其他数字生物标记物。本文以RADAR-MDD语音语料库为实验框架,在两级抑郁严重度分类范式下,检验了一种具有局部注意机制的序列对序列模型的有效性。此外,还提出了一种新的训练方法&硬训练。它是一种基于为模型训练选择更多模棱两可样本的方法,并受到课程学习范式的启发。我们发现,在所使用的两个语音诱导任务和RADAR-MDD语音语料库的每个采集点中,艰苦训练能够持续提高分类器的性能,平均增长8.6%。有了这种新的方法,我们的序列到序列模型能够有效地检测出MDD的严重性,而不考虑语言。最后,认识到需要更多地了解潜在的算法偏差,我们对每个性别的结果分别进行了额外的分析。摘要:Major Depressive Disorder (MDD) is a common worldwide mental health issue with high associated socioeconomic costs. The prediction and automatic detection of MDD can, therefore, make a huge impact on society. Speech, as a non-invasive, easy to collect signal, is a promising marker to aid the diagnosis and assessment of MDD. In this regard, speech samples were collected as part of the Remote Assessment of Disease and Relapse in Major Depressive Disorder (RADAR-MDD) research programme. RADAR-MDD was an observational cohort study in which speech and other digital biomarkers were collected from a cohort of individuals with a history of MDD in Spain, United Kingdom and the Netherlands. In this paper, the RADAR-MDD speech corpus was taken as an experimental framework to test the efficacy of a Sequence-to-Sequence model with a local attention mechanism in a two-class depression severity classification paradigm. Additionally, a novel training method, HARD-Training, is proposed. It is a methodology based on the selection of more ambiguous samples for the model training, and inspired by the curriculum learning paradigm. HARD-Training was found to consistently improve - with an average increment of 8.6% - the performance of our classifiers for both of two speech elicitation tasks used and each collection site of the RADAR-MDD speech corpus. With this novel methodology, our Sequence-to-Sequence model was able to effectively detect MDD severity regardless of language. Finally, recognising the need for greater awareness of potential algorithmic bias, we conduct an additional analysis of our results separately for each gender.
【2】 Constraining Gaussian processes for physics-informed acoustic emission mapping
标题:物理信息声发射映射的约束高斯过程
链接:https://arxiv.org/abs/2206.01495
作者:Matthew R Jones,Timothy J Rogers,Elizabeth J Cross摘要:结构损伤的自动定位是实现高价值结构预测性或基于状态的维护的一个具有挑战性但至关重要的因素。声发射到达时间图的使用是应对这一挑战的一种很有希望的方法,但由于需要在整个结构上收集密集的人工声发射测量数据集,因此受到了严重的阻碍,导致数据采集过程冗长且往往不切实际。在本文中,我们考虑使用物理信息高斯过程来学习这些映射来缓解这个问题。在该方法中,高斯过程被限制在物理域中,因此与结构的几何和边界条件相关的信息直接嵌入到学习过程中,返回一个模型,该模型可确保所做的任何预测满足边界上物理一致的行为。当训练测量采集受到限制时,会出现许多场景,包括训练数据稀疏的情况,以及对兴趣结构的覆盖范围有限的情况。使用复杂的板状结构作为实验案例研究,我们表明,我们的方法显著减少了数据收集的负担,可以看出,随着训练观测值的减少,结合边界条件知识显著提高了预测精度,尤其是当结构的所有部分都无法获得训练测量值时。摘要:The automated localisation of damage in structures is a challenging but critical ingredient in the path towards predictive or condition-based maintenance of high value structures. The use of acoustic emission time of arrival mapping is a promising approach to this challenge, but is severely hindered by the need to collect a dense set of artificial acoustic emission measurements across the structure, resulting in a lengthy and often impractical data acquisition process. In this paper, we consider the use of physics-informed Gaussian processes for learning these maps to alleviate this problem. In the approach, the Gaussian process is constrained to the physical domain such that information relating to the geometry and boundary conditions of the structure are embedded directly into the learning process, returning a model that guarantees that any predictions made satisfy physically-consistent behaviour at the boundary. A number of scenarios that arise when training measurement acquisition is limited, including where training data are sparse, and also of limited coverage over the structure of interest. Using a complex plate-like structure as an experimental case study, we show that our approach significantly reduces the burden of data collection, where it is seen that incorporation of boundary condition knowledge significantly improves predictive accuracy as training observations are reduced, particularly when training measurements are not available across all parts of the structure.
【3】 The Musical Arrow of Time -- The Role of Temporal Asymmetry in Music and Its Organicist Implications
标题:时间的音乐之箭--时间不对称在音乐中的作用及其有机主义意蕴
链接:https://arxiv.org/abs/2206.01305
摘要:从以演奏者为中心的角度来看,我们经常会遇到两种说法:“音乐流动”和“音乐就像生活”。本论文基于以上两种说法,探索了时间不对称在音乐中的作用(概括“音乐流”)及其与有机主义思想的关系(概括“音乐是逼真的”)。我们关注时间不对称的两个方面。第一个方面涉及我们获取过去和未来知识的截然不同的认知机制。一个特殊的音乐后果如下:重现。过去和未来之间的认知差异塑造了我们对音乐中反复发生的事件的经验和解释。第二个方面涉及时间之箭:强加在时间事件上的明确顺序产生了时间的先验指向性,使得时间不对称且不可逆转。关于热力学的讨论从音乐上告诉我们:时间之箭通过延迟高潮的位置在音乐形式中发挥作用。有机主义作为一个中介话题,与生物体中的生命概念有关。一方面,有机主义通过对生命作为熵减少实体的热力学解释,与科学中的时间不对称有关。另一方面,有机主义是一个音乐固有的话题,这是一种公认的艺术理念,即音乐应被解释为具有意志力的生命力。以有机主义为中介,我们可以更好地理解时间不对称在音乐中的作用。特别是,我们认为音乐形式是一个扩展和细化的过程,类似于有机增长。最后,我们提出了延迟高潮的有机主义解释:将音乐形式视为有机增长的结果,时间之箭转化为对前置结构的偏好,而不是对附加结构的偏好。摘要:Adopting a performer-centric perspective, we frequently encounter two statements: "music flows", and "music is life-like". This dissertation builds on top of the two statements above, resulting in an exploration of the role of temporal asymmetry in music (generalizing "music flows") and its relation to the idea of organicism (generalizing "music is life-like"). We focus on two aspects of temporal asymmetry. The first aspect concerns the vastly different epistemic mechanisms with which we obtain knowledge of the past and the future. A particular musical consequence follows: recurrence. The epistemic difference between the past and the future shapes our experience and interpretation of recurring events in music. The second aspect concerns the arrow of time: the unambiguous ordering imposed on temporal events gives rise to the a priori pointedness of time, rendering time asymmetrical and irreversible. A discussion on thermodynamics informs us musically: the arrow of time effectuates itself in musical forms by delaying the placement of the climax. Organicism serves as a mediating topic, engaging with the concept of life as in organisms. On the one hand, organicism is related to temporal asymmetry in science via a thermodynamical interpretation of life as entropy-reducing entities. On the other hand, organicism is a topic native to music via the universally acknowledged artistic idea that music should be interpreted as a vital force possessing volitional power. With organicism as a mediator, we better understand the role of temporal asymmetry in music. In particular, we view musical form as a process of expansion and elaboration analogous to organic growth. Finally, we present an organicist interpretation of delaying the climax: viewing musical form as the result of organic growth, the arrow of time translates to a preference for prepending structure over appending structure.
【4】 Snow Mountain: Dataset of Audio Recordings of The Bible in Low Resource Languages
标题:雪山:低资源语种的圣经录音数据集
链接:https://arxiv.org/abs/2206.01205
作者:Kavitha Raju,Anjaly V,Ryan Lish,Joel Mathew备注:See dataset at this https URL摘要:自动语音识别(ASR)在现代社会中的应用日益广泛。有许多ASR模型可用于具有大量训练数据的语言,如英语。然而,低资源语言的代表性较差。作为回应,我们创建并发布了一个开放的许可和格式化的数据集,以低资源的北印度语言记录《圣经》。我们建立了多个实验分割,并训练和分析了两个有竞争力的ASR模型,作为使用这些数据进行未来研究的基线。摘要:Automatic Speech Recognition (ASR) has increasing utility in the modern world. There are a many ASR models available for languages with large amounts of training data like English. However, low-resource languages are poorly represented. In response we create and release an open-licensed and formatted dataset of audio recordings of the Bible in low-resource northern Indian languages. We setup multiple experimental splits and train and analyze two competitive ASR models to serve as the baseline for future research using this data.
【1】 Snow Mountain: Dataset of Audio Recordings of The Bible in Low Resource Languages
标题:雪山:低资源语种的圣经录音数据集
链接:https://arxiv.org/abs/2206.01205
作者:Kavitha Raju,Anjaly V,Ryan Lish,Joel Mathew备注:See dataset at this https URL摘要:自动语音识别(ASR)在现代社会中的应用日益广泛。有许多ASR模型可用于具有大量训练数据的语言,如英语。然而,低资源语言的代表性较差。作为回应,我们创建并发布了一个开放的许可和格式化的数据集,以低资源的北印度语言记录《圣经》。我们建立了多个实验分割,并训练和分析了两个有竞争力的ASR模型,作为使用这些数据进行未来研究的基线。摘要:Automatic Speech Recognition (ASR) has increasing utility in the modern world. There are a many ASR models available for languages with large amounts of training data like English. However, low-resource languages are poorly represented. In response we create and release an open-licensed and formatted dataset of audio recordings of the Bible in low-resource northern Indian languages. We setup multiple experimental splits and train and analyze two competitive ASR models to serve as the baseline for future research using this data.
【2】 Detecting the Severity of Major Depressive Disorder from Speech: A Novel HARD-Training Methodology
标题:从言语中检测重度抑郁障碍的严重程度:一种新的硬训练方法
链接:https://arxiv.org/abs/2206.01542
作者:Edward L. Campbell,Judith Dineley,Pauline Conde,Faith Matcham,Femke Lamers,Sara Siddi,Laura Docio-Fernandez,Carmen Garcia-Mateo,Nicholas Cummins,the RADAR-CNS Consortium摘要:重度抑郁障碍(MDD)是世界范围内常见的心理健康问题,其社会经济成本较高。因此,MDD的预测和自动检测可以对社会产生巨大影响。语音作为一种非侵入性、易于采集的信号,是一种很有前途的辅助诊断和评估MDD的指标。在这方面,作为远程评估重度抑郁症(RADAR-MDD)研究计划的一部分,收集了语音样本。RADAR-MDD是一项观察性队列研究,从西班牙、英国和荷兰有MDD病史的人群中收集语音和其他数字生物标记物。本文以RADAR-MDD语音语料库为实验框架,在两级抑郁严重度分类范式下,检验了一种具有局部注意机制的序列对序列模型的有效性。此外,还提出了一种新的训练方法&硬训练。它是一种基于为模型训练选择更多模棱两可样本的方法,并受到课程学习范式的启发。我们发现,在所使用的两个语音诱导任务和RADAR-MDD语音语料库的每个采集点中,艰苦训练能够持续提高分类器的性能,平均增长8.6%。有了这种新的方法,我们的序列到序列模型能够有效地检测出MDD的严重性,而不考虑语言。最后,认识到需要更多地了解潜在的算法偏差,我们对每个性别的结果分别进行了额外的分析。摘要:Major Depressive Disorder (MDD) is a common worldwide mental health issue with high associated socioeconomic costs. The prediction and automatic detection of MDD can, therefore, make a huge impact on society. Speech, as a non-invasive, easy to collect signal, is a promising marker to aid the diagnosis and assessment of MDD. In this regard, speech samples were collected as part of the Remote Assessment of Disease and Relapse in Major Depressive Disorder (RADAR-MDD) research programme. RADAR-MDD was an observational cohort study in which speech and other digital biomarkers were collected from a cohort of individuals with a history of MDD in Spain, United Kingdom and the Netherlands. In this paper, the RADAR-MDD speech corpus was taken as an experimental framework to test the efficacy of a Sequence-to-Sequence model with a local attention mechanism in a two-class depression severity classification paradigm. Additionally, a novel training method, HARD-Training, is proposed. It is a methodology based on the selection of more ambiguous samples for the model training, and inspired by the curriculum learning paradigm. HARD-Training was found to consistently improve - with an average increment of 8.6% - the performance of our classifiers for both of two speech elicitation tasks used and each collection site of the RADAR-MDD speech corpus. With this novel methodology, our Sequence-to-Sequence model was able to effectively detect MDD severity regardless of language. Finally, recognising the need for greater awareness of potential algorithmic bias, we conduct an additional analysis of our results separately for each gender.
【3】 Constraining Gaussian processes for physics-informed acoustic emission mapping
标题:物理信息声发射映射的约束高斯过程
链接:https://arxiv.org/abs/2206.01495
作者:Matthew R Jones,Timothy J Rogers,Elizabeth J Cross摘要:结构损伤的自动定位是实现高价值结构预测性或基于状态的维护的一个具有挑战性但至关重要的因素。声发射到达时间图的使用是应对这一挑战的一种很有希望的方法,但由于需要在整个结构上收集密集的人工声发射测量数据集,因此受到了严重的阻碍,导致数据采集过程冗长且往往不切实际。在本文中,我们考虑使用物理信息高斯过程来学习这些映射来缓解这个问题。在该方法中,高斯过程被限制在物理域中,因此与结构的几何和边界条件相关的信息直接嵌入到学习过程中,返回一个模型,该模型可确保所做的任何预测满足边界上物理一致的行为。当训练测量采集受到限制时,会出现许多场景,包括训练数据稀疏的情况,以及对兴趣结构的覆盖范围有限的情况。使用复杂的板状结构作为实验案例研究,我们表明,我们的方法显著减少了数据收集的负担,可以看出,随着训练观测值的减少,结合边界条件知识显著提高了预测精度,尤其是当结构的所有部分都无法获得训练测量值时。摘要:The automated localisation of damage in structures is a challenging but critical ingredient in the path towards predictive or condition-based maintenance of high value structures. The use of acoustic emission time of arrival mapping is a promising approach to this challenge, but is severely hindered by the need to collect a dense set of artificial acoustic emission measurements across the structure, resulting in a lengthy and often impractical data acquisition process. In this paper, we consider the use of physics-informed Gaussian processes for learning these maps to alleviate this problem. In the approach, the Gaussian process is constrained to the physical domain such that information relating to the geometry and boundary conditions of the structure are embedded directly into the learning process, returning a model that guarantees that any predictions made satisfy physically-consistent behaviour at the boundary. A number of scenarios that arise when training measurement acquisition is limited, including where training data are sparse, and also of limited coverage over the structure of interest. Using a complex plate-like structure as an experimental case study, we show that our approach significantly reduces the burden of data collection, where it is seen that incorporation of boundary condition knowledge significantly improves predictive accuracy as training observations are reduced, particularly when training measurements are not available across all parts of the structure.
【4】 The Musical Arrow of Time -- The Role of Temporal Asymmetry in Music and Its Organicist Implications
标题:时间的音乐之箭--时间不对称在音乐中的作用及其有机主义意蕴
链接:https://arxiv.org/abs/2206.01305
摘要:从以演奏者为中心的角度来看,我们经常会遇到两种说法:“音乐流动”和“音乐就像生活”。本论文基于以上两种说法,探索了时间不对称在音乐中的作用(概括“音乐流”)及其与有机主义思想的关系(概括“音乐是逼真的”)。我们关注时间不对称的两个方面。第一个方面涉及我们获取过去和未来知识的截然不同的认知机制。一个特殊的音乐后果如下:重现。过去和未来之间的认知差异塑造了我们对音乐中反复发生的事件的经验和解释。第二个方面涉及时间之箭:强加在时间事件上的明确顺序产生了时间的先验指向性,使得时间不对称且不可逆转。关于热力学的讨论从音乐上告诉我们:时间之箭通过延迟高潮的位置在音乐形式中发挥作用。有机主义作为一个中介话题,与生物体中的生命概念有关。一方面,有机主义通过对生命作为熵减少实体的热力学解释,与科学中的时间不对称有关。另一方面,有机主义是一个音乐固有的话题,这是一种公认的艺术理念,即音乐应被解释为具有意志力的生命力。以有机主义为中介,我们可以更好地理解时间不对称在音乐中的作用。特别是,我们认为音乐形式是一个扩展和细化的过程,类似于有机增长。最后,我们提出了延迟高潮的有机主义解释:将音乐形式视为有机增长的结果,时间之箭转化为对前置结构的偏好,而不是对附加结构的偏好。摘要:Adopting a performer-centric perspective, we frequently encounter two statements: "music flows", and "music is life-like". This dissertation builds on top of the two statements above, resulting in an exploration of the role of temporal asymmetry in music (generalizing "music flows") and its relation to the idea of organicism (generalizing "music is life-like"). We focus on two aspects of temporal asymmetry. The first aspect concerns the vastly different epistemic mechanisms with which we obtain knowledge of the past and the future. A particular musical consequence follows: recurrence. The epistemic difference between the past and the future shapes our experience and interpretation of recurring events in music. The second aspect concerns the arrow of time: the unambiguous ordering imposed on temporal events gives rise to the a priori pointedness of time, rendering time asymmetrical and irreversible. A discussion on thermodynamics informs us musically: the arrow of time effectuates itself in musical forms by delaying the placement of the climax. Organicism serves as a mediating topic, engaging with the concept of life as in organisms. On the one hand, organicism is related to temporal asymmetry in science via a thermodynamical interpretation of life as entropy-reducing entities. On the other hand, organicism is a topic native to music via the universally acknowledged artistic idea that music should be interpreted as a vital force possessing volitional power. With organicism as a mediator, we better understand the role of temporal asymmetry in music. In particular, we view musical form as a process of expansion and elaboration analogous to organic growth. Finally, we present an organicist interpretation of delaying the climax: viewing musical form as the result of organic growth, the arrow of time translates to a preference for prepending structure over appending structure.
机器翻译,仅供参考