今日论文合集:cs.SD语音3篇,eess.AS音频处理2篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Open vocabulary keyword spotting through transfer learning from speech  synthesis
标题:通过语音合成的迁移学习实现开放词汇关键词识别
链接:https://arxiv.org/abs/2404.03914
作者:Kesavaraj V,Anil Kumar Vuppala
摘要:在开放词汇上下文中识别关键词对于个性化与智能设备的交互至关重要。以前的开放式词汇关键词识别方法依赖于由音频和文本编码器创建的共享嵌入空间。然而,这些方法遭受异构模态表示(即,音频-文本不匹配)。为了解决这个问题,我们提出的框架利用从预训练的文本到语音(TTS)系统获得的知识。这种知识转移允许将对音频投影的感知结合到从文本编码器导出的文本表示中。所提出的方法的性能进行比较,在四个不同的数据集与各种基线方法。我们提出的模型的鲁棒性进行评估,通过评估其性能在不同的单词长度和在一个词汇表(OOV)的情况下。此外,从TTS系统的迁移学习的有效性进行了研究,通过分析其不同的中间表示。实验结果表明,在具有挑战性的LibriPhrase Hard数据集上,所提出的方法在曲线下面积(AUC)和等错误率(EER)方面分别提高了8.22%和12.56%,优于跨模态对应检测器(CMCD)方法。
摘要:Identifying keywords in an open-vocabulary context is crucial for personalizing interactions with smart devices. Previous approaches to open vocabulary keyword spotting dependon a shared embedding space created by audio and text encoders. However, these approaches suffer from heterogeneous modality representations (i.e., audio-text mismatch). To address this issue, our proposed framework leverages knowledge acquired from a pre-trained text-to-speech (TTS) system. This knowledge transfer allows for the incorporation of awareness of audio projections into the text representations derived from the text encoder. The performance of the proposed approach is compared with various baseline methods across four different datasets. The robustness of our proposed model is evaluated by assessing its performance across different word lengths and in an Out-of-Vocabulary (OOV) scenario. Additionally, the effectiveness of transfer learning from the TTS system is investigated by analyzing its different intermediate representations. The experimental results indicate that, in the challenging LibriPhrase Hard dataset, the proposed approach outperformed the cross-modality correspondence detector (CMCD) method by a significant improvement of 8.22% in area under the curve (AUC) and 12.56% in equal error rate (EER).


【2】Multi-Task Learning for Lung sound & Lung disease classification
标题:肺音多任务学习与肺部疾病分类
链接:https://arxiv.org/abs/2404.03908
作者:Suma K V,Deepali Koppad,Preethi Kumar,Neha A Kantikar,Surabhi Ramesh
摘要:近年来,深度学习技术的进步大大提高了医疗诊断的效率和准确性。在这项工作中,提出了一种新的方法,使用多任务学习(MTL)的肺音和肺部疾病的同时分类。我们提出的模型利用MTL和四种不同的深度学习模型,如2D CNN,ResNet 50,MobileNet和Densenet,从肺音记录中提取相关特征。本研究采用了ICBHI 2017呼吸音数据库。MTL for MobileNet模型的肺音分析准确率为74%,肺部疾病分类准确率为91%,优于其他模型。实验结果表明,我们的方法在分类肺音和肺部疾病同时的有效性。  在这项研究中,使用的人口统计数据库中的患者,慢性阻塞性肺疾病的风险水平计算也进行了。对于这个计算,三个机器学习算法,即逻辑回归,SVM和随机森林分类器。在这些ML算法中,随机森林分类器的准确率最高,达到92%。这项工作不仅有助于大大减轻医生诊断病理的负担,而且还有助于有效地与患者沟通可能的原因或结果。
摘要:In recent years, advancements in deep learning techniques have considerably enhanced the efficiency and accuracy of medical diagnostics. In this work, a novel approach using multi-task learning (MTL) for the simultaneous classification of lung sounds and lung diseases is proposed. Our proposed model leverages MTL with four different deep learning models such as 2D CNN, ResNet50, MobileNet and Densenet to extract relevant features from the lung sound recordings. The ICBHI 2017 Respiratory Sound Database was employed in the current study. The MTL for MobileNet model performed better than the other models considered, with an accuracy of74\% for lung sound analysis and 91\% for lung diseases classification. Results of the experimentation demonstrate the efficacy of our approach in classifying both lung sounds and lung diseases concurrently.  In this study,using the demographic data of the patients from the database, risk level computation for Chronic Obstructive Pulmonary Disease is also carried out. For this computation, three machine learning algorithms namely Logistic Regression, SVM and Random Forest classifierswere employed. Among these ML algorithms, the Random Forest classifier had the highest accuracy of 92\%.This work helps in considerably reducing the physician's burden of not just diagnosing the pathology but also effectively communicating to the patient about the possible causes or outcomes.

【3】Holon: a cybernetic interface for bio-semiotics
标题:Holon:生物符号学的控制论界面
链接:https://arxiv.org/abs/2404.03894
作者:Jon McCormack,Elliott Wilson
备注:Paper accepted at ISEA 24, The 29th International Symposium on Electronic Art, Brisbane, Australia, 21-29 June 2024
摘要:本文介绍了一个互动的艺术作品,“Holon”,收集了130个自主的,控制论的有机体,听,使声音与自然环境的合作。这项工作是为在澳大利亚墨尔本的一个列入遗产名录的码头上安装而开发的。介绍了指导工作的概念问题,以及实施的详细技术概述。个体合弄有三种类型,灵感来自动物交流的生物模型:作曲家/发电机,收集器/评论家和破坏者。整体而言,Holon与人类和非人类代理合作整合并占据声学频谱的元素。
摘要:This paper presents an interactive artwork, "Holon", a collection of 130 autonomous, cybernetic organisms that listen and make sound in collaboration with the natural environment. The work was developed for installation on water at a heritage-listed dock in Melbourne, Australia. Conceptual issues informing the work are presented, along with a detailed technical overview of the implementation. Individual holons are of three types, inspired by biological models of animal communication: composer/generators, collector/critics and disruptors. Collectively, Holon integrates and occupies elements of the acoustic spectrum in collaboration with human and non-human agents.


eess.AS音频处理
【1】Open vocabulary keyword spotting through transfer learning from speech  synthesis
标题:通过语音合成的迁移学习实现开放词汇关键词识别
链接:https://arxiv.org/abs/2404.03914
作者:Kesavaraj V,Anil Kumar Vuppala
摘要:在开放词汇上下文中识别关键词对于个性化与智能设备的交互至关重要。以前的开放式词汇关键词识别方法依赖于由音频和文本编码器创建的共享嵌入空间。然而,这些方法遭受异构模态表示(即,音频-文本不匹配)。为了解决这个问题,我们提出的框架利用从预训练的文本到语音(TTS)系统获得的知识。这种知识转移允许将对音频投影的感知结合到从文本编码器导出的文本表示中。所提出的方法的性能进行比较,在四个不同的数据集与各种基线方法。我们提出的模型的鲁棒性进行评估,通过评估其性能在不同的单词长度和在一个词汇表(OOV)的情况下。此外,从TTS系统的迁移学习的有效性进行了研究,通过分析其不同的中间表示。实验结果表明,在具有挑战性的LibriPhrase Hard数据集上,所提出的方法在曲线下面积(AUC)和等错误率(EER)方面分别提高了8.22%和12.56%,优于跨模态对应检测器(CMCD)方法。
摘要:Identifying keywords in an open-vocabulary context is crucial for personalizing interactions with smart devices. Previous approaches to open vocabulary keyword spotting dependon a shared embedding space created by audio and text encoders. However, these approaches suffer from heterogeneous modality representations (i.e., audio-text mismatch). To address this issue, our proposed framework leverages knowledge acquired from a pre-trained text-to-speech (TTS) system. This knowledge transfer allows for the incorporation of awareness of audio projections into the text representations derived from the text encoder. The performance of the proposed approach is compared with various baseline methods across four different datasets. The robustness of our proposed model is evaluated by assessing its performance across different word lengths and in an Out-of-Vocabulary (OOV) scenario. Additionally, the effectiveness of transfer learning from the TTS system is investigated by analyzing its different intermediate representations. The experimental results indicate that, in the challenging LibriPhrase Hard dataset, the proposed approach outperformed the cross-modality correspondence detector (CMCD) method by a significant improvement of 8.22% in area under the curve (AUC) and 12.56% in equal error rate (EER).


【2】Holon: a cybernetic interface for bio-semiotics
标题:Holon:生物符号学的控制论界面
链接:https://arxiv.org/abs/2404.03894
作者:Jon McCormack,Elliott Wilson
备注:Paper accepted at ISEA 24, The 29th International Symposium on Electronic Art, Brisbane, Australia, 21-29 June 2024
摘要:本文介绍了一个互动的艺术作品,“Holon”,收集了130个自主的,控制论的有机体,听,使声音与自然环境的合作。这项工作是为在澳大利亚墨尔本的一个列入遗产名录的码头上安装而开发的。介绍了指导工作的概念问题,以及实施的详细技术概述。个体合弄有三种类型,灵感来自动物交流的生物模型:作曲家/发电机,收集器/评论家和破坏者。整体而言,Holon与人类和非人类代理合作整合并占据声学频谱的元素。
摘要:This paper presents an interactive artwork, "Holon", a collection of 130 autonomous, cybernetic organisms that listen and make sound in collaboration with the natural environment. The work was developed for installation on water at a heritage-listed dock in Melbourne, Australia. Conceptual issues informing the work are presented, along with a detailed technical overview of the implementation. Individual holons are of three types, inspired by biological models of animal communication: composer/generators, collector/critics and disruptors. Collectively, Holon integrates and occupies elements of the acoustic spectrum in collaboration with human and non-human agents.

机器翻译由腾讯交互翻译提供,仅供参考