本文经arXiv每日学术速递授权转载
【1】 Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
标题: 对齐视觉和声音:通过视听对齐进行高级光源定位
作者:Arda Senocak,Hyeonggon Ryu,Junsik Kim,Tae-Hyun Oh,Hanspeter Pfister,Joon Son Chung
备注:Journal Extension of ICCV 2023 paper (arXiV:2309.10724). Code is available at this https URL
链接:点击下载PDF文件
【2】 CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
标题: CogniVoice:用于自发言语轻度认知障碍评估的多模式和多语言融合网络
作者:Jiali Cheng,Mohamed Elgaar,Nidhi Vakil,Hadi Amiri
备注:INTERSPEECH 2024
链接:点击下载PDF文件
【3】 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
标题: 基于语言模型的具有可控自发行为的自发风格文本到语音合成
作者:Weiqin Li,Peiji Yang,Yicheng Zhong,Yixuan Zhou,Zhisheng Wang,Zhiyong Wu,Xixin Wu,Helen Meng
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【4】 Reducing Barriers to the Use of Marginalised Music Genres in AI
标题: 减少人工智能中使用边缘化音乐流派的障碍
作者:Nick Bryan-Kinns,Zijin Li
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
【5】 Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
标题: 通过低成本数据策略提高印度TTC系统的词汇外性能以满足实际应用
作者:Srija Anand,Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【6】 Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
标题: 在损失函数中使用语音基础模型进行助听器语音增强
作者:Robert Sutherland,George Close,Thomas Hain,Stefan Goetze,Jon Barker
备注:Accepted for EUSIPCO 2024
链接:点击下载PDF文件
【7】 Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
标题: 通过弱监督的基于音素的多语言预训练实现Iu Mien语言的低资源语音识别
作者:Lukuan Dong,Donghong Qin,Fengbo Bai,Fanhua Song,Yan Liu,Chen Xu,Zhijian Ou
链接:点击下载PDF文件
【8】 How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
标题: 低频语音音频在野外有多私密?人类和机器对言语可理解性的分析
作者:Ailin Liu,Pepijn Vunderink,Jose Vargas Quiros,Chirag Raman,Hayley Hung
备注:This manuscript has been accepted by Interspeech 2024
链接:点击下载PDF文件
【9】 Underwater Acoustic Signal Denoising Algorithms: A Survey of the State-of-the-art
标题: 水下声学信号去噪算法:最新进展综述
作者:Ruobin Gao,Maohan Liang,Heng Dong,Xuewen Luo,P. N. Suganthan
链接:点击下载PDF文件
【10】 DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
标题: DiveSound:LLM辅助自动分类构建,用于多样化音频生成
作者:Baihan Li,Zeyu Xie,Xuenan Xu,Yiwei Guo,Ming Yan,Ji Zhang,Kai Yu,Mengyue Wu
链接:点击下载PDF文件
【11】 Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems
标题: 受隐私监管的语音翻译系统的预设语音匹配
作者:Daniel Platnick,Bishoy Abdelnour,Eamon Earl,Rahul Kumar,Zahra Rezaei,Thomas Tsangaris,Faraj Lagum
备注:Accepted to the ACL PrivateNLP 2024 Workshop, 7 pages, 2 figures
链接:点击下载PDF文件
【12】 A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
标题: 一种轻量级、高效的标点符号和词性预测模型,用于设备上流媒体ASB
作者:Jian You,Xiangfeng Li
链接:点击下载PDF文件
【13】 Audio-visual Generalized Zero-shot Learning the Easy Way
标题: 视听广义Zero-Shot学习的简单方法
作者:Shentong Mo,Pedro Morgado
链接:点击下载PDF文件
【14】 Modeling and Driving Human Body Soundfields through Acoustic Primitives
标题: 通过声学基元建模和驱动人体音场
作者:Chao Huan,Dejan Markovic,Chenliang Xu,Alexander Richard
备注:ECCV 2024. Project Page: this https URL
链接:点击下载PDF文件
【15】 Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
标题: 预先训练的基础模型表示,以揭示言语中的呼吸模式
作者:Vikramjit Mitra,Anirban Chatterjee,Ke Zhai,Helen Weng,Ayuko Hill,Nicole Hay,Christopher Webb,Jamie Cheng,Erdrin Azemi
备注:8 pages, 6 figures, BioKDD workshop paper
链接:点击下载PDF文件
【16】 Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
标题: 通过关注自动语音识别的声学和置信参考来纠正错误
作者:Yuchun Shu,Bo Hu,Yifeng He,Hao Shi,Longbiao Wang,Jianwu Dang
链接:点击下载PDF文件
【17】 Fade-in Reverberation in Multi-room Environments Using the Common-Slope Model
标题: 使用共坡模型的多房间环境中的渐入混响
作者:Kyung Yun Lee,Nils Meyer-Kahlen,Georg Götz,U. Peter Svensson,Sebastian J. Schlecht,Vesa Välimäki
备注:2024 AES 5th International Conference on Audio for Virtual and Augmented Reality
链接:点击下载PDF文件
【18】 MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
标题: MEDIC:具有解开倒置控制的Zero-Shot音乐编辑
作者:Huadai Liu,Jialei Wang,Rongjie Huang,Yang Liu,Jiayang Xu,Zhou Zhao
链接:点击下载PDF文件
标题: 使用共坡模型的多房间环境中的渐入混响
作者:Kyung Yun Lee,Nils Meyer-Kahlen,Georg Götz,U. Peter Svensson,Sebastian J. Schlecht,Vesa Välimäki
备注:2024 AES 5th International Conference on Audio for Virtual and Augmented Reality
链接:点击下载PDF文件
【2】 MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
标题: MEDIC:具有解开倒置控制的Zero-Shot音乐编辑
作者:Huadai Liu,Jialei Wang,Rongjie Huang,Yang Liu,Jiayang Xu,Zhou Zhao
链接:点击下载PDF文件
【3】 Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
标题: 多迭代多阶段精调Transformer,用于使用异类数据集的声音事件检测
作者:Florian Schmid,Paul Primus,Tobias Morocutti,Jonathan Greif,Gerhard Widmer
备注:Code: this https URL
链接:点击下载PDF文件
【4】 Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
标题: 对齐视觉和声音:通过视听对齐进行高级光源定位
作者:Arda Senocak,Hyeonggon Ryu,Junsik Kim,Tae-Hyun Oh,Hanspeter Pfister,Joon Son Chung
备注:Journal Extension of ICCV 2023 paper (arXiV:2309.10724). Code is available at this https URL
链接:点击下载PDF文件
【5】 CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
标题: CogniVoice:用于自发言语轻度认知障碍评估的多模式和多语言融合网络
作者:Jiali Cheng,Mohamed Elgaar,Nidhi Vakil,Hadi Amiri
备注:INTERSPEECH 2024
链接:点击下载PDF文件
【6】 Accurate Mapping of RNNs on Neuromorphic Hardware with Adaptive Spiking Neurons
标题: 具有自适应尖峰神经元的神经形态硬件上RNN的准确映射
作者:Gauthier Boeshertz,Giacomo Indiveri,Manu Nair,Alpha Renner
备注:5 pages, 3 figures, accepted at ICONS 2024
链接:点击下载PDF文件
【7】 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
标题: 基于语言模型的具有可控自发行为的自发风格文本到语音合成
作者:Weiqin Li,Peiji Yang,Yicheng Zhong,Yixuan Zhou,Zhisheng Wang,Zhiyong Wu,Xixin Wu,Helen Meng
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【8】 Reducing Barriers to the Use of Marginalised Music Genres in AI
标题: 减少人工智能中使用边缘化音乐流派的障碍
作者:Nick Bryan-Kinns,Zijin Li
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
【9】 Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
标题: 通过低成本数据策略提高印度TTC系统的词汇外性能以满足实际应用
作者:Srija Anand,Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【10】 Linear-Complexity Self-Supervised Learning for Speech Processing
标题: 语音处理的线性复杂度自我监督学习
作者:Shucong Zhang,Titouan Parcollet,Rogier van Dalen,Sourav Bhattacharya
备注:Interspeech 2024
链接:点击下载PDF文件
【11】 Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
标题: 在损失函数中使用语音基础模型进行助听器语音增强
作者:Robert Sutherland,George Close,Thomas Hain,Stefan Goetze,Jon Barker
备注:Accepted for EUSIPCO 2024
链接:点击下载PDF文件
【12】 Robust ASR Error Correction with Conservative Data Filtering
标题: 具有保守数据过滤的鲁棒的ASB错误纠正
作者:Takuma Udagawa,Masayuki Suzuki,Masayasu Muraoka,Gakuto Kurata
链接:点击下载PDF文件
【13】 Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
标题: 通过弱监督的基于音素的多语言预训练实现Iu Mien语言的低资源语音识别
作者:Lukuan Dong,Donghong Qin,Fengbo Bai,Fanhua Song,Yan Liu,Chen Xu,Zhijian Ou
链接:点击下载PDF文件
【14】 How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
标题: 低频语音音频在野外有多私密?人类和机器对言语可理解性的分析
作者:Ailin Liu,Pepijn Vunderink,Jose Vargas Quiros,Chirag Raman,Hayley Hung
备注:This manuscript has been accepted by Interspeech 2024
链接:点击下载PDF文件
【15】 Underwater Acoustic Signal Denoising Algorithms: A Survey of the State-of-the-art
标题: 水下声学信号去噪算法:最新进展综述
作者:Ruobin Gao,Maohan Liang,Heng Dong,Xuewen Luo,P. N. Suganthan
链接:点击下载PDF文件
【16】 DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
标题: DiveSound:LLM辅助自动分类构建,用于多样化音频生成
作者:Baihan Li,Zeyu Xie,Xuenan Xu,Yiwei Guo,Ming Yan,Ji Zhang,Kai Yu,Mengyue Wu
链接:点击下载PDF文件
【17】 Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems
标题: 受隐私监管的语音翻译系统的预设语音匹配
作者:Daniel Platnick,Bishoy Abdelnour,Eamon Earl,Rahul Kumar,Zahra Rezaei,Thomas Tsangaris,Faraj Lagum
备注:Accepted to the ACL PrivateNLP 2024 Workshop, 7 pages, 2 figures
链接:点击下载PDF文件
【18】 A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
标题: 一种轻量级、高效的标点符号和词性预测模型,用于设备上流媒体ASB
作者:Jian You,Xiangfeng Li
链接:点击下载PDF文件
【19】 Audio-visual Generalized Zero-shot Learning the Easy Way
标题: 视听广义Zero-Shot学习的简单方法
作者:Shentong Mo,Pedro Morgado
链接:点击下载PDF文件
【20】 Modeling and Driving Human Body Soundfields through Acoustic Primitives
标题: 通过声学基元建模和驱动人体音场
作者:Chao Huan,Dejan Markovic,Chenliang Xu,Alexander Richard
备注:ECCV 2024. Project Page: this https URL
链接:点击下载PDF文件
【21】 Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
标题: 预先训练的基础模型表示,以揭示言语中的呼吸模式
作者:Vikramjit Mitra,Anirban Chatterjee,Ke Zhai,Helen Weng,Ayuko Hill,Nicole Hay,Christopher Webb,Jamie Cheng,Erdrin Azemi
备注:8 pages, 6 figures, BioKDD workshop paper
链接:点击下载PDF文件
【22】 Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
标题: 通过关注自动语音识别的声学和置信参考来纠正错误
作者:Yuchun Shu,Bo Hu,Yifeng He,Hao Shi,Longbiao Wang,Jianwu Dang
链接:点击下载PDF文件
标题: 对齐视觉和声音:通过视听对齐进行高级光源定位
作者:Arda Senocak,Hyeonggon Ryu,Junsik Kim,Tae-Hyun Oh,Hanspeter Pfister,Joon Son Chung
备注:Journal Extension of ICCV 2023 paper (arXiV:2309.10724). Code is available at this https URL
链接:点击下载PDF文件
摘要:近年来,基于学习的声源定位研究主要集中在定位性能方面。然而,以前的工作和现有的基准忽略了一个重要的方面:跨模态的相互作用,这是必不可少的交互式声源定位。跨模态交互对于理解语义匹配或不匹配的视听事件至关重要,例如无声对象或屏幕外的声音。在本文中,我们首先全面研究了现有方法,基准,评估指标和跨模态理解任务的跨模态交互。然后,我们确定了以前的研究的局限性,并提出了一些贡献,以克服这些局限性。首先,我们引入一个用于交互式声源定位的新合成基准。其次,我们引入了新的评估指标来严格评估声源定位方法,重点是准确评估定位性能和跨模态交互能力。第三,我们提出了一个学习框架与跨模态对齐策略,以加强跨模态的互动。最后,我们一起评估交互式声源定位和辅助跨模态检索任务,以彻底评估跨模态交互能力和基准竞争方法。我们的新基准和评估指标揭示了声源定位研究中以前被忽视的问题。我们提出的新方法,增强跨模态对齐,显示出优越的声源定位性能。这项工作提供了迄今为止最全面的声源定位分析,并使用新的标准评估指标对现有和新的基准进行了广泛的竞争方法验证。摘要:Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, which is essential for interactive sound source localization. Cross-modal interaction is vital for understanding semantically matched or mismatched audio-visual events, such as silent objects or off-screen sounds. In this paper, we first comprehensively examine the cross-modal interaction of existing methods, benchmarks, evaluation metrics, and cross-modal understanding tasks. Then, we identify the limitations of previous studies and make several contributions to overcome the limitations. First, we introduce a new synthetic benchmark for interactive sound source localization. Second, we introduce new evaluation metrics to rigorously assess sound source localization methods, focusing on accurately evaluating both localization performance and cross-modal interaction ability. Third, we propose a learning framework with a cross-modal alignment strategy to enhance cross-modal interaction. Lastly, we evaluate both interactive sound source localization and auxiliary cross-modal retrieval tasks together to thoroughly assess cross-modal interaction capabilities and benchmark competing methods. Our new benchmarks and evaluation metrics reveal previously overlooked issues in sound source localization studies. Our proposed novel method, with enhanced cross-modal alignment, shows superior sound source localization performance. This work provides the most comprehensive analysis of sound source localization to date, with extensive validation of competing methods on both existing and new benchmarks using new and standard evaluation metrics.
【2】 CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
标题: CogniVoice:用于自发言语轻度认知障碍评估的多模式和多语言融合网络
作者:Jiali Cheng,Mohamed Elgaar,Nidhi Vakil,Hadi Amiri
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:轻度认知障碍(MCI)是一种医学疾病,其特征是记忆和认知能力明显下降,可能影响个人的日常活动。在本文中,我们介绍CogniVoice,一种新的多语言和多模式的框架来检测MCI和估计的简易精神状态检查(MMSE)分数通过分析语音数据及其文本transanimation。CogniVoice的关键组成部分是一个基于"专家产品“的多模式和多语言综合网络,减少了对捷径解决方案的依赖。使用TAUKADIAL挑战中包含英语和中文的综合数据集,CogniVoice在MCI分类和MMSE回归任务中的表现分别优于最佳基线模型2.8和4.1分,并且可以有效地将不同语言组之间的性能差距缩小0.7分。摘要:Mild Cognitive Impairment (MCI) is a medical condition characterized by noticeable declines in memory and cognitive abilities, potentially affecting individual's daily activities. In this paper, we introduce CogniVoice, a novel multilingual and multimodal framework to detect MCI and estimate Mini-Mental State Examination (MMSE) scores by analyzing speech data and its textual transcriptions. The key component of CogniVoice is an ensemble multimodal and multilingual network based on Product of Experts'' that mitigates reliance on shortcut solutions. Using a comprehensive dataset containing both English and Chinese languages from TAUKADIAL challenge, CogniVoice outperforms the best performing baseline model on MCI classification and MMSE regression tasks by 2.8 and 4.1 points in F1 and RMSE respectively, and can effectively reduce the performance gap across different language groups by 0.7 points in F1.
【3】 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
标题: 基于语言模型的具有可控自发行为的自发风格文本到语音合成
作者:Weiqin Li,Peiji Yang,Yicheng Zhong,Yixuan Zhou,Zhisheng Wang,Zhiyong Wu,Xixin Wu,Helen Meng
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:由于缺乏高质量的数据和模型能力的限制,旨在生成类人语音的自发式语音合成经常遇到挑战。最近的基于语言模型的TTS系统可以在大型,多样化和低质量的语音数据集上进行训练,从而产生高度自然的合成语音。然而,它们受到模拟各种自发行为和捕捉自发语音中的韵律变化的困难的限制。在本文中,我们提出了一种新的自发语音合成系统的语言模型的基础上。我们系统地对各种自发行为进行分类和统一建模。实验结果表明,本文提出的方法在韵律自然度和自发行为自然度方面明显优于传统方法。摘要:Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained on large, diverse, and low-quality speech datasets, resulting in highly natural synthesized speech. However, they are limited by the difficulty of simulating various spontaneous behaviors and capturing prosody variations in spontaneous speech. In this paper, we propose a novel spontaneous speech synthesis system based on language models. We systematically categorize and uniformly model diverse spontaneous behaviors. Moreover, fine-grained prosody modeling is introduced to enhance the model's ability to capture subtle prosody variations in spontaneous speech.Experimental results show that our proposed method significantly outperforms the baseline methods in terms of prosody naturalness and spontaneous behavior naturalness.
【4】 Reducing Barriers to the Use of Marginalised Music Genres in AI
标题: 减少人工智能中使用边缘化音乐流派的障碍
作者:Nick Bryan-Kinns,Zijin Li
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:用于高质量音乐生成的AI系统通常依赖于非常大的音乐数据集来训练AI模型。这就为生成超越主流数据集中所代表的流派(如西方古典音乐或流行音乐)的音乐创造了障碍。我们进行了一项为期4个月的国际研究项目,旨在探索可解释人工智能(XAI)的挑战和机遇,以减少使用人工智能模型的边缘化音乐类型的障碍。确定的XAI机会包括提高AI模型的透明度和控制,解释AI模型的道德和偏见,用小数据集微调大模型以减少偏见,以及解释AI模型的风格转移机会。研究参与者强调,虽然很难使用边缘化音乐和人工智能等小数据集,但这种方法加强了代表性不足的文化的文化代表性,并有助于解决深度学习模型的偏见问题。我们现在正在这个项目的基础上建立一个全球性的国际负责任的人工智能音乐社区,并邀请人们加入我们的网络。摘要:AI systems for high quality music generation typically rely on extremely large musical datasets to train the AI models. This creates barriers to generating music beyond the genres represented in dominant datasets such as Western Classical music or pop music. We undertook a 4 month international research project summarised in this paper to explore the eXplainable AI (XAI) challenges and opportunities associated with reducing barriers to using marginalised genres of music with AI models. XAI opportunities identified included topics of improving transparency and control of AI models, explaining the ethics and bias of AI models, fine tuning large models with small datasets to reduce bias, and explaining style-transfer opportunities with AI models. Participants in the research emphasised that whilst it is hard to work with small datasets such as marginalised music and AI, such approaches strengthen cultural representation of underrepresented cultures and contribute to addressing issues of bias of deep learning models. We are now building on this project to bring together a global International Responsible AI Music community and invite people to join our network.
【5】 Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
标题: 通过低成本数据策略提高印度TTC系统的词汇外性能以满足实际应用
作者:Srija Anand,Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:公开可用的低资源语言(如印地语和泰米尔语)的TTS数据集通常包含10-20小时的数据,导致词汇覆盖率很低。这种限制在下游应用中变得明显,在下游应用中,特定领域的词汇加上与英语的频繁代码混合,导致许多OOV单词。为了突出这个问题,我们创建了一个基准包含OOV字从几个现实世界的应用程序。事实上,最先进的印地语和泰米尔语TTS系统在这个OOV基准上表现不佳,正如可懂度测试所示。为了提高模型的OOV性能,我们提出了一种低成本和经济可行的策略来获得更多的训练数据。具体来说,我们建议使用志愿者而不是高质量的语音艺术家来记录包含训练数据中看不到的字符二元组的单词。我们表明,使用这种廉价的数据,该模型的性能提高OOV的话,而不影响语音质量和域内的性能。摘要:Publicly available TTS datasets for low-resource languages like Hindi and Tamil typically contain 10-20 hours of data, leading to poor vocabulary coverage. This limitation becomes evident in downstream applications where domain-specific vocabulary coupled with frequent code-mixing with English, results in many OOV words. To highlight this problem, we create a benchmark containing OOV words from several real-world applications. Indeed, state-of-the-art Hindi and Tamil TTS systems perform poorly on this OOV benchmark, as indicated by intelligibility tests. To improve the model's OOV performance, we propose a low-effort and economically viable strategy to obtain more training data. Specifically, we propose using volunteers as opposed to high quality voice artists to record words containing character bigrams unseen in the training data. We show that using such inexpensive data, the model's performance improves on OOV words, while not affecting voice quality and in-domain performance.
【6】 Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
标题: 在损失函数中使用语音基础模型进行助听器语音增强
作者:Robert Sutherland,George Close,Thomas Hain,Stefan Goetze,Jon Barker
备注:Accepted for EUSIPCO 2024
链接:点击下载PDF文件
摘要:机器学习技术是助听器语音增强的一个活跃研究领域,其中一个特别关注的是提高嘈杂语音信号的可懂度。最近的工作表明,自监督语音表示模型的特征编码可以有效地捕获语音可懂度。在这项工作中,它表明,清洁和嘈杂的语音的自我监督的语音表示之间的距离更强烈地与人类的可懂度评级比其他基于信号的度量。实验表明,使用该距离作为损失函数的一部分来训练语音增强模型,与使用基于SNR的损失函数相比,可以提高性能,这可以通过HASPI,STOI,PESQ和SI-SNR分数的增加来证明。该方法仅在训练时进行高参数计数模型的推断,这意味着语音增强模型可以保持较小,如助听器所需。摘要:Machine learning techniques are an active area of research for speech enhancement for hearing aids, with one particular focus on improving the intelligibility of a noisy speech signal. Recent work has shown that feature encodings from self-supervised speech representation models can effectively capture speech intelligibility. In this work, it is shown that the distance between self-supervised speech representations of clean and noisy speech correlates more strongly with human intelligibility ratings than other signal-based metrics. Experiments show that training a speech enhancement model using this distance as part of a loss function improves the performance over using an SNR-based loss function, demonstrated by an increase in HASPI, STOI, PESQ and SI-SNR scores. This method takes inference of a high parameter count model only at training time, meaning the speech enhancement model can remain smaller, as is required for hearing aids.
【7】 Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
标题: 通过弱监督的基于音素的多语言预训练实现Iu Mien语言的低资源语音识别
作者:Lukuan Dong,Donghong Qin,Fengbo Bai,Fanhua Song,Yan Liu,Chen Xu,Zhijian Ou
链接:点击下载PDF文件
摘要:主流的自动语音识别(ASR)技术通常需要数百到数千小时的带注释语音数据。低资源ASR的三种方法是基于音素或子词的监督预训练,以及多语言数据的自监督预训练。瑶缅语是我国瑶族的主要民族语言,其资源贫乏,注释语音十分有限。本文以不到10小时的瑶面语为例,研究和比较了三种瑶面语语音识别方法。我们的实验基于最近发布的三种骨干模型,这些模型在CommonVoice数据集(CV-Lang 10)的10种语言上进行了预训练,这对应于低资源ASR的三种方法。结果表明,音素监督比子词监督和自监督能获得更好的结果,从而提供更高的数据效率。特别是Whistle模型,即,通过基于弱监督音素的多语种预训练,获得最具竞争力的结果。摘要:The mainstream automatic speech recognition (ASR) technology usually requires hundreds to thousands of hours of annotated speech data. Three approaches to low-resourced ASR are phoneme or subword based supervised pre-training, and self-supervised pre-training over multilingual data. The Iu Mien language is the main ethnic language of the Yao ethnic group in China and is low-resourced in the sense that the annotated speech is very limited. With less than 10 hours of transcribed Iu Mien language, this paper investigates and compares the three approaches for Iu Mien speech recognition. Our experiments are based on the recently released, three backbone models pretrained over the 10 languages from the CommonVoice dataset (CV-Lang10), which correspond to the three approaches for low-resourced ASR. It is found that phoneme supervision can achieve better results compared to subword supervision and self-supervision, thereby providing higher data-efficiency. Particularly, the Whistle models, i.e., obtained by the weakly-supervised phoneme-based multilingual pre-training, obtain the most competitive results.
【8】 How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
标题: 低频语音音频在野外有多私密?人类和机器对言语可理解性的分析
作者:Ailin Liu,Pepijn Vunderink,Jose Vargas Quiros,Chirag Raman,Hayley Hung
备注:This manuscript has been accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:低频音频已被提出作为一个有前途的隐私保护模式,研究在现实世界中的社会动态。为此,研究人员开发了可穿戴设备,可以以低至1250 Hz的频率记录音频,以减轻对可能包含私人细节的语音的口头内容的自动提取。本文探讨了这一假设的有效性,检查在何种程度上确保言语隐私的低频语音。它包括在各种噪声环境中模拟潜在的隐私攻击。此外,它还探讨了语音活动检测性能(这是理解社交行为的基础)与隐私保护之间的权衡。该评估结合了主观人类可懂度和自动语音识别性能,全面分析了有效的社会行为分析和保护言语隐私之间的微妙平衡。摘要:Low-frequency audio has been proposed as a promising privacy-preserving modality to study social dynamics in real-world settings. To this end, researchers have developed wearable devices that can record audio at frequencies as low as 1250 Hz to mitigate the automatic extraction of the verbal content of speech that may contain private details. This paper investigates the validity of this hypothesis, examining the degree to which low-frequency speech ensures verbal privacy. It includes simulating a potential privacy attack in various noise environments. Further, it explores the trade-off between the performance of voice activity detection, which is fundamental for understanding social behavior, and privacy-preservation. The evaluation incorporates subjective human intelligibility and automatic speech recognition performance, comprehensively analyzing the delicate balance between effective social behavior analysis and preserving verbal privacy.
【9】 Underwater Acoustic Signal Denoising Algorithms: A Survey of the State-of-the-art
标题: 水下声学信号去噪算法:最新进展综述
作者:Ruobin Gao,Maohan Liang,Heng Dong,Xuewen Luo,P. N. Suganthan
链接:点击下载PDF文件
摘要:本文综述了水声信号去噪的最新进展,这是提高水下通信和监测系统可靠性和清晰度的关键领域。尽管该领域取得了重大进展,但水下环境的复杂性带来了独特的挑战,使去噪过程变得复杂。我们首先概述了与水声信号处理相关的基本挑战,包括信号衰减,噪声变化和环境因素的影响。然后,该综述系统地分类和讨论了各种去噪算法,例如传统的,基于分解的和基于学习的技术,突出了它们的应用,优点和局限性。评估指标和实验数据集也进行了审查。本文最后列出了一系列开放性问题和建议,为未来的研究方向,强调需要开发更强大的去噪技术,可以适应动态的水声环境。摘要:This paper comprehensively reviews recent advances in underwater acoustic signal denoising, an area critical for improving the reliability and clarity of underwater communication and monitoring systems. Despite significant progress in the field, the complex nature of underwater environments poses unique challenges that complicate the denoising process. We begin by outlining the fundamental challenges associated with underwater acoustic signal processing, including signal attenuation, noise variability, and the impact of environmental factors. The review then systematically categorizes and discusses various denoising algorithms, such as conventional, decomposition-based, and learning-based techniques, highlighting their applications, advantages, and limitations. Evaluation metrics and experimental datasets are also reviewed. The paper concludes with a list of open questions and recommendations for future research directions, emphasizing the need for developing more robust denoising techniques that can adapt to the dynamic underwater acoustic environment.
【10】 DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
标题: DiveSound:LLM辅助自动分类构建,用于多样化音频生成
作者:Baihan Li,Zeyu Xie,Xuenan Xu,Yiwei Guo,Ming Yan,Ji Zhang,Kai Yu,Mengyue Wu
链接:点击下载PDF文件
摘要:音频生成引起了极大的关注。尽管在音频质量显着提高,现有的模型忽略了多样性评价。这部分是由于缺乏一个系统的健全的类多样性框架和匹配的数据集。为了解决这些问题,我们提出了DiveSound,一个新的框架,用于构建多模态数据集与类内多样化的分类,辅助大型语言模型。由于文本和视觉信息都可以用来指导多样化的生成,DiveSound在数据构建中利用了多模态对比表示。我们的框架是高度自治的,可以很容易地扩展。我们提供了一个文本图像对齐的多样性数据集,其声音事件类标签平均有2.42个子类别。在构建的数据集上进行的文本到音频的实验表明,在视觉信息的引导下,多样性大幅增加。摘要:Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the lack of a systematic sound class diversity framework and a matching dataset. To address these issues, we propose DiveSound, a novel framework for constructing multimodal datasets with in-class diversified taxonomy, assisted by large language models. As both textual and visual information can be utilized to guide diverse generation, DiveSound leverages multimodal contrastive representations in data construction. Our framework is highly autonomous and can be easily scaled up. We provide a textaudio-image aligned diversity dataset whose sound event class tags have an average of 2.42 subcategories. Text-to-audio experiments on the constructed dataset show a substantial increase of diversity with the help of the guidance of visual information.
【11】 Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems
标题: 受隐私监管的语音翻译系统的预设语音匹配
作者:Daniel Platnick,Bishoy Abdelnour,Eamon Earl,Rahul Kumar,Zahra Rezaei,Thomas Tsangaris,Faraj Lagum
备注:Accepted to the ACL PrivateNLP 2024 Workshop, 7 pages, 2 figures
链接:点击下载PDF文件
摘要:近年来,在工业环境中对语音到语音翻译(S2ST)系统的需求已经增加。虽然克隆的S2ST系统已成功商业化,但当被个人滥用时,其分销商将承担责任,当被媒体组织利用时,可能侵犯人格权。这项工作提出了一个规范的S2ST框架称为预设语音匹配(PVM)。PVM通过首先将输入语音与目标语言中类似的先前同意的说话者语音进行匹配来消除S2ST中的跨语言语音克隆。通过这种分离,PVM避免了克隆输入扬声器,确保PVM系统符合法规并降低误用风险。实验结果表明,PVM可以显著提高S2ST系统在多说话人环境下的运行时间和S2ST合成语音的自然度。据我们所知,PVM是第一个明确规定的S2ST框架,利用类似匹配的预设语音进行动态S2ST任务。摘要:In recent years, there has been increased demand for speech-to-speech translation (S2ST) systems in industry settings. Although successfully commercialized, cloning-based S2ST systems expose their distributors to liabilities when misused by individuals and can infringe on personality rights when exploited by media organizations. This work proposes a regulated S2ST framework called Preset-Voice Matching (PVM). PVM removes cross-lingual voice cloning in S2ST by first matching the input voice to a similar prior consenting speaker voice in the target-language. With this separation, PVM avoids cloning the input speaker, ensuring PVM systems comply with regulations and reduce risk of misuse. Our results demonstrate PVM can significantly improve S2ST system run-time in multi-speaker settings and the naturalness of S2ST synthesized speech. To our knowledge, PVM is the first explicitly regulated S2ST framework leveraging similarly-matched preset-voices for dynamic S2ST tasks.
【12】 A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
标题: 一种轻量级、高效的标点符号和词性预测模型,用于设备上流媒体ASB
作者:Jian You,Xiangfeng Li
链接:点击下载PDF文件
摘要:标点符号和词的大小写预测是自动语音识别(ASR)的必要条件。随着设备上端到端流媒体ASR系统的普及,设备上的标点符号和词大小写预测成为必要,但我们发现很少有人讨论这一点。随着Transformer的出现,人们已经为此场景探索了基于Transformer的模型。然而,基于Transformer的模型对于设备上ASR系统来说太大。在本文中,我们提出了一个轻量级和高效的模型,联合预测标点符号和词的大小写实时。该模型基于卷积神经网络(CNN)和双向长短期记忆(BiLSTM)。在IWSLT2011测试集上的实验结果表明,与最好的非Transformer模型相比,该模型在总体F1分数上获得了9%的相对提高.与基于Transformer的模型的代表相比,所提出的模型实现了与代表模型相当的结果,同时仅为其大小的四十分之一,并且在推理时间方面快2.5倍。它适用于设备上流式ASR系统。我们的代码是公开的。摘要:Punctuation and word casing prediction are necessary for automatic speech recognition (ASR). With the popularity of on-device end-to-end streaming ASR systems, the on-device punctuation and word casing prediction become a necessity while we found little discussion on this. With the emergence of Transformer, Transformer based models have been explored for this scenario. However, Transformer based models are too large for on-device ASR systems. In this paper, we propose a light-weight and efficient model that jointly predicts punctuation and word casing in real time. The model is based on Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM). Experimental results on the IWSLT2011 test set show that the proposed model obtains 9% relative improvement compared to the best of non-Transformer models on overall F1-score. Compared to the representative of Transformer based models, the proposed model achieves comparable results to the representative model while being only one-fortieth its size and 2.5 times faster in terms of inference time. It is suitable for on-device streaming ASR systems. Our code is publicly available.
【13】 Audio-visual Generalized Zero-shot Learning the Easy Way
标题: 视听广义Zero-Shot学习的简单方法
作者:Shentong Mo,Pedro Morgado
链接:点击下载PDF文件
摘要:视听广义zero-shot学习是一个快速发展的领域,旨在理解视频中音频和视觉线索之间的复杂关系。总体目标是利用可见类的洞察力来识别以前未见过的实例。先前的方法主要利用同步自动编码器来重建视听属性,这些属性由交叉注意Transformers和投影文本嵌入来通知。然而,这些方法未能有效地捕捉预训练的语言对齐嵌入中固有的跨模态特征和类标签嵌入之间的复杂关系。为了克服这些瓶颈,我们引入了一个简单而有效的框架,用于简单的视听广义Zero-shot学习,名为EZ-AVGZL,将视听嵌入与转换后的文本表示对齐。它利用单一的监督文本视听对比损失来学习视听和文本模态之间的对齐,远离了重建跨模态特征和文本嵌入的传统方法。我们的主要观点是,虽然类名嵌入与基于语言的视听功能很好地结合在一起,但它们并没有提供足够的类分离,从而对zero-shot学习有用。为了解决这个问题,我们的方法利用差分优化将类嵌入转换到一个更具歧视性的空间,同时保留语言表示的语义结构。我们在VGGSound-Gandrel、UCF-Gandrel和ActivityNet-Gandrel基准上进行了广泛的实验。我们的研究结果表明,我们的EZ-AVGZL在视听广义zero-shot学习中达到了最先进的性能。摘要:Audio-visual generalized zero-shot learning is a rapidly advancing domain that seeks to understand the intricate relations between audio and visual cues within videos. The overarching goal is to leverage insights from seen classes to identify instances from previously unseen ones. Prior approaches primarily utilized synchronized auto-encoders to reconstruct audio-visual attributes, which were informed by cross-attention transformers and projected text embeddings. However, these methods fell short of effectively capturing the intricate relationship between cross-modal features and class-label embeddings inherent in pre-trained language-aligned embeddings. To circumvent these bottlenecks, we introduce a simple yet effective framework for Easy Audio-Visual Generalized Zero-shot Learning, named EZ-AVGZL, that aligns audio-visual embeddings with transformed text representations. It utilizes a single supervised text audio-visual contrastive loss to learn an alignment between audio-visual and textual modalities, moving away from the conventional approach of reconstructing cross-modal features and text embeddings. Our key insight is that while class name embeddings are well aligned with language-based audio-visual features, they don't provide sufficient class separation to be useful for zero-shot learning. To address this, our method leverages differential optimization to transform class embeddings into a more discriminative space while preserving the semantic structure of language representations. We conduct extensive experiments on VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL benchmarks. Our results demonstrate that our EZ-AVGZL achieves state-of-the-art performance in audio-visual generalized zero-shot learning.
【14】 Modeling and Driving Human Body Soundfields through Acoustic Primitives
标题: 通过声学基元建模和驱动人体音场
作者:Chao Huan,Dejan Markovic,Chenliang Xu,Alexander Richard
备注:ECCV 2024. Project Page: this https URL
链接:点击下载PDF文件
摘要:虽然真实感3D人体模型的渲染和动画在过去几年中已经成熟并达到了令人印象深刻的质量,但到目前为止,与这种全身模型相关联的空间音频建模在很大程度上被忽视了。在这项工作中,我们提出了一个框架,允许高质量的空间音频生成,能够渲染由人体生成的完整的3D声场,包括语音,脚步,手-身体交互等。给出了一个基本的视听表示的身体的形式的3D身体姿势和音频从头戴式麦克风,我们证明了我们可以在3D空间中的任何一点高效,准确地渲染完整的声学场景。为了实现声音的近场和实时渲染,我们借用了图形神经渲染中的体积基元的想法,并将它们转移到声学域。与以前的方法相比,我们的声学基元可以实现更小数量级的声场表示,并克服了近场渲染中的缺陷。摘要:While rendering and animation of photorealistic 3D human body models have matured and reached an impressive quality over the past years, modeling the spatial audio associated with such full body models has been largely ignored so far. In this work, we present a framework that allows for high-quality spatial audio generation, capable of rendering the full 3D soundfield generated by a human body, including speech, footsteps, hand-body interactions, and others. Given a basic audio-visual representation of the body in form of 3D body pose and audio from a head-mounted microphone, we demonstrate that we can render the full acoustic scene at any point in 3D space efficiently and accurately. To enable near-field and realtime rendering of sound, we borrow the idea of volumetric primitives from graphical neural rendering and transfer them into the acoustic domain. Our acoustic primitives result in an order of magnitude smaller soundfield representations and overcome deficiencies in near-field rendering compared to previous approaches.
【15】 Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
标题: 预先训练的基础模型表示,以揭示言语中的呼吸模式
作者:Vikramjit Mitra,Anirban Chatterjee,Ke Zhai,Helen Weng,Ayuko Hill,Nicole Hay,Christopher Webb,Jamie Cheng,Erdrin Azemi
备注:8 pages, 6 figures, BioKDD workshop paper
链接:点击下载PDF文件
摘要:人类语音产生的过程涉及协调的呼吸动作以引出声学语音信号。通常情况下,当空气从肺中被迫出来并被声道调制时,就会产生语音,其中这种动作被呼吸空气(吸入)的时刻所穿插,以再次重新填充肺部。呼吸率(RR)是一个重要的指标,用于评估个人的整体健康,健身和一般福祉。测量RR(一分钟内呼吸的次数)的现有方法是使用专门的设备或培训进行的。研究表明,机器学习算法可用于使用生物传感器信号作为输入来估计RR。基于语音的RR估计可以提供一种有效的方法来测量生命指标,而不需要任何专门的设备或传感器。这项工作研究了一种基于机器学习的方法来估计RR从语音段从说话的对象到一个近距离说话的麦克风设备。从N=26个个体收集数据,其中通过商业级胸带获得地面真实RR,然后手动校正任何错误。提出了一种卷积长短期记忆网络(Conv-LSTM)来从语音信号中估计呼吸时间序列数据。我们证明,使用从基础模型(如Wav 2 Vec 2)获得的预训练表示,与基线相比,可用于估计具有低均方根误差和高相关系数的呼吸时间序列。模型驱动的时间序列可用于估计$RR$,平均绝对误差(MAE)较低,约为1.6次呼吸 min。摘要:The process of human speech production involves coordinated respiratory action to elicit acoustic speech signals. Typically, speech is produced when air is forced from the lungs and is modulated by the vocal tract, where such actions are interspersed by moments of breathing in air (inhalation) to refill the lungs again. Respiratory rate (RR) is a vital metric that is used to assess the overall health, fitness, and general well-being of an individual. Existing approaches to measure RR (number of breaths one takes in a minute) are performed using specialized equipment or training. Studies have demonstrated that machine learning algorithms can be used to estimate RR using bio-sensor signals as input. Speech-based estimation of RR can offer an effective approach to measure the vital metric without requiring any specialized equipment or sensors. This work investigates a machine learning based approach to estimate RR from speech segments obtained from subjects speaking to a close-talking microphone device. Data were collected from N=26 individuals, where the groundtruth RR was obtained through commercial grade chest-belts and then manually corrected for any errors. A convolutional long-short term memory network (Conv-LSTM) is proposed to estimate respiration time-series data from the speech signal. We demonstrate that the use of pre-trained representations obtained from a foundation model, such as Wav2Vec2, can be used to estimate respiration-time-series with low root-mean-squared error and high correlation coefficient, when compared with the baseline. The model-driven time series can be used to estimate $RR$ with a low mean absolute error (MAE) ~ 1.6 breaths min.
【16】 Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
标题: 通过关注自动语音识别的声学和置信参考来纠正错误
作者:Yuchun Shu,Bo Hu,Yifeng He,Hao Shi,Longbiao Wang,Jianwu Dang
链接:点击下载PDF文件
摘要:语音纠错的目标是准确地发现自动语音识别(ASR)假设中的错误单词并将其恢复为有根据的。本文提出了一种非自回归语音纠错方法。置信度模块测量N个最佳ASR假设中每个词的不确定性,作为找到错误词位置的参考。此外,ASR编码器的声学特征也被用来提供正确的发音参考。使用编辑路径对齐来自ASR的N个最佳候选,以相互确认并恢复一些丢失的字符错误。此外,交叉注意机制融合了错误校正参考和ASR假设之间的信息。实验结果表明,声学参考和置信度参考都有助于纠错。与ASR模型相比,该系统将错误率降低了21%。摘要:Accurately finding the wrong words in the automatic speech recognition (ASR) hypothesis and recovering them well-founded is the goal of speech error correction. In this paper, we propose a non-autoregressive speech error correction method. A Confidence Module measures the uncertainty of each word of the N-best ASR hypotheses as the reference to find the wrong word position. Besides, the acoustic feature from the ASR encoder is also used to provide the correct pronunciation references. N-best candidates from ASR are aligned using the edit path, to confirm each other and recover some missing character errors. Furthermore, the cross-attention mechanism fuses the information between error correction references and the ASR hypothesis. The experimental results show that both the acoustic and confidence references help with error correction. The proposed system reduces the error rate by 21% compared with the ASR model.
【17】 Fade-in Reverberation in Multi-room Environments Using the Common-Slope Model
标题: 使用共坡模型的多房间环境中的渐入混响
作者:Kyung Yun Lee,Nils Meyer-Kahlen,Georg Götz,U. Peter Svensson,Sebastian J. Schlecht,Vesa Välimäki
备注:2024 AES 5th International Conference on Audio for Virtual and Augmented Reality
链接:点击下载PDF文件
摘要:在多房间环境中,由于房间的耦合和不同的源-接收器位置,声音传播的建模是复杂的。一个常见的情况是,当源和接收器在不同的房间,没有一个清晰的视线。对于这样的源-接收器配置,观察到能量的初始增加,称为混响的“淡入”。基于最近的工作,代表不均匀和各向异性混响与共同的衰减时间,这项工作提出了一个扩展的参数模型,使建模的淡入现象。该方法在包络上执行拟合,而不是能量衰减函数,并且允许衰减指数的负幅度。我们在模拟和测量的多房间环境中评估了该方法,在那里我们表明,所提出的方法现在可以对以前的方法无法实现的淡入进行建模。摘要:In multi-room environments, modelling the sound propagation is complex due to the coupling of rooms and diverse source-receiver positions. A common scenario is when the source and the receiver are in different rooms without a clear line of sight. For such source-receiver configurations, an initial increase in energy is observed, referred to as the "fade-in" of reverberation. Based on recent work of representing inhomogeneous and anisotropic reverberation with common decay times, this work proposes an extended parametric model that enables the modelling of the fade-in phenomenon. The method performs fitting on the envelopes, instead of energy decay functions, and allows negative amplitudes of decaying exponentials. We evaluate the method on simulated and measured multi-room environments, where we show that the proposed approach can now model the fade-ins that were unrealisable with the previous method.
【18】 MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
标题: MEDIC:具有解开倒置控制的Zero-Shot音乐编辑
作者:Huadai Liu,Jialei Wang,Rongjie Huang,Yang Liu,Jiayang Xu,Zhou Zhao
链接:点击下载PDF文件
摘要:文本引导的扩散模型催化了音频生成的范式转变,促进了源音频的适应性,以符合特定的文本提示。最近的进展将诸如DDIM反转的反转技术引入到zero-shot编辑,利用预先训练的扩散模型进行音频修改。尽管如此,我们的调查表明,DDIM反转在每个扩散步骤中都存在误差累积,从而削弱了其功效。注意力控制的缺乏阻碍了对音乐的精细操作。为了克服这些局限性,我们引入了 textit{解纠缠反转}技术,该技术旨在将扩散过程分解为三个分支,从而放大它们各自的精确编辑和保存能力。此外,我们提出了 textit{协调注意控制}框架,它统一了相互自我注意和交叉注意与一个额外的谐波分支,以实现所需的组成和结构信息的目标音乐。总的来说,这些创新构成了 textit{解纠缠反转控制(DIC)}框架,在保证结构完整性的同时实现了精确的音乐编辑。为了对音频编辑功效进行基准测试,我们引入了 textit{ZoME-Bench},这是一个全面的音乐编辑基准测试,包含分布在10个不同编辑类别中的1,100个样本,可以促进零拍摄(zero-shot)和基于指令的音乐编辑任务。我们的方法在编辑保真度和基本内容保留方面表现出无与伦比的性能,优于当代最先进的反演技术。摘要:Text-guided diffusion models catalyze a paradigm shift in audio generation, facilitating the adaptability of source audio to conform to specific textual prompts. Recent advancements introduce inversion techniques, like DDIM inversion, to zero-shot editing, exploiting pre-trained diffusion models for audio modification. Nonetheless, our investigation exposes that DDIM inversion suffers from an accumulation of errors across each diffusion step, undermining its efficacy. And the lack of attention control hinders the fine-grained manipulations of music. To counteract these limitations, we introduce the textit{Disentangled Inversion} technique, which is designed to disentangle the diffusion process into triple branches, thereby magnifying their individual capabilities for both precise editing and preservation. Furthermore, we propose the textit{Harmonized Attention Control} framework, which unifies the mutual self-attention and cross-attention with an additional Harmonic Branch to achieve the desired composition and structural information in the target music. Collectively, these innovations comprise the textit{Disentangled Inversion Control (DIC)} framework, enabling accurate music editing whilst safeguarding structural integrity. To benchmark audio editing efficacy, we introduce textit{ZoME-Bench}, a comprehensive music editing benchmark hosting 1,100 samples spread across 10 distinct editing categories, which facilitates both zero-shot and instruction-based music editing tasks. Our method demonstrates unparalleled performance in edit fidelity and essential content preservation, outperforming contemporary state-of-the-art inversion techniques.
eess.AS音频处理
【1】 Fade-in Reverberation in Multi-room Environments Using the Common-Slope Model标题: 使用共坡模型的多房间环境中的渐入混响
作者:Kyung Yun Lee,Nils Meyer-Kahlen,Georg Götz,U. Peter Svensson,Sebastian J. Schlecht,Vesa Välimäki
备注:2024 AES 5th International Conference on Audio for Virtual and Augmented Reality
链接:点击下载PDF文件
摘要:在多房间环境中,由于房间的耦合和不同的源-接收器位置,声音传播的建模是复杂的。一个常见的情况是,当源和接收器在不同的房间,没有一个清晰的视线。对于这样的源-接收器配置,观察到能量的初始增加,称为混响的“淡入”。基于最近的工作,代表不均匀和各向异性混响与共同的衰减时间,这项工作提出了一个扩展的参数模型,使建模的淡入现象。该方法在包络上执行拟合,而不是能量衰减函数,并且允许衰减指数的负幅度。我们在模拟和测量的多房间环境中评估了该方法,在那里我们表明,所提出的方法现在可以对以前的方法无法实现的淡入进行建模。摘要:In multi-room environments, modelling the sound propagation is complex due to the coupling of rooms and diverse source-receiver positions. A common scenario is when the source and the receiver are in different rooms without a clear line of sight. For such source-receiver configurations, an initial increase in energy is observed, referred to as the "fade-in" of reverberation. Based on recent work of representing inhomogeneous and anisotropic reverberation with common decay times, this work proposes an extended parametric model that enables the modelling of the fade-in phenomenon. The method performs fitting on the envelopes, instead of energy decay functions, and allows negative amplitudes of decaying exponentials. We evaluate the method on simulated and measured multi-room environments, where we show that the proposed approach can now model the fade-ins that were unrealisable with the previous method.
【2】 MEDIC: Zero-shot Music Editing with Disentangled Inversion Control
标题: MEDIC:具有解开倒置控制的Zero-Shot音乐编辑
作者:Huadai Liu,Jialei Wang,Rongjie Huang,Yang Liu,Jiayang Xu,Zhou Zhao
链接:点击下载PDF文件
摘要:文本引导的扩散模型催化了音频生成的范式转变,促进了源音频的适应性,以符合特定的文本提示。最近的进展将诸如DDIM反转的反转技术引入到zero-shot编辑,利用预先训练的扩散模型进行音频修改。尽管如此,我们的调查表明,DDIM反转在每个扩散步骤中都存在误差累积,从而削弱了其功效。注意力控制的缺乏阻碍了对音乐的精细操作。为了克服这些局限性,我们引入了 textit{解纠缠反转}技术,该技术旨在将扩散过程分解为三个分支,从而放大它们各自的精确编辑和保存能力。此外,我们提出了 textit{协调注意控制}框架,它统一了相互自我注意和交叉注意与一个额外的谐波分支,以实现所需的组成和结构信息的目标音乐。总的来说,这些创新构成了 textit{解纠缠反转控制(DIC)}框架,在保证结构完整性的同时实现了精确的音乐编辑。为了对音频编辑功效进行基准测试,我们引入了 textit{ZoME-Bench},这是一个全面的音乐编辑基准测试,包含分布在10个不同编辑类别中的1,100个样本,可以促进零拍摄(zero-shot)和基于指令的音乐编辑任务。我们的方法在编辑保真度和基本内容保留方面表现出无与伦比的性能,优于当代最先进的反演技术。摘要:Text-guided diffusion models catalyze a paradigm shift in audio generation, facilitating the adaptability of source audio to conform to specific textual prompts. Recent advancements introduce inversion techniques, like DDIM inversion, to zero-shot editing, exploiting pre-trained diffusion models for audio modification. Nonetheless, our investigation exposes that DDIM inversion suffers from an accumulation of errors across each diffusion step, undermining its efficacy. And the lack of attention control hinders the fine-grained manipulations of music. To counteract these limitations, we introduce the textit{Disentangled Inversion} technique, which is designed to disentangle the diffusion process into triple branches, thereby magnifying their individual capabilities for both precise editing and preservation. Furthermore, we propose the textit{Harmonized Attention Control} framework, which unifies the mutual self-attention and cross-attention with an additional Harmonic Branch to achieve the desired composition and structural information in the target music. Collectively, these innovations comprise the textit{Disentangled Inversion Control (DIC)} framework, enabling accurate music editing whilst safeguarding structural integrity. To benchmark audio editing efficacy, we introduce textit{ZoME-Bench}, a comprehensive music editing benchmark hosting 1,100 samples spread across 10 distinct editing categories, which facilitates both zero-shot and instruction-based music editing tasks. Our method demonstrates unparalleled performance in edit fidelity and essential content preservation, outperforming contemporary state-of-the-art inversion techniques.
【3】 Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
标题: 多迭代多阶段精调Transformer,用于使用异类数据集的声音事件检测
作者:Florian Schmid,Paul Primus,Tobias Morocutti,Jonathan Greif,Gerhard Widmer
备注:Code: this https URL
链接:点击下载PDF文件
摘要:构建有效的声音事件检测系统的一个核心问题是缺乏高质量的,强注释的声音事件数据集。出于这个原因,DCASE 2024挑战的任务4提出了从两个异构数据集中学习,包括用不同的注释粒度和不同的可能事件集标记的音频片段。我们提出了一个多迭代,多阶段的程序微调音频频谱图Transformers联合DESED和MAESTRO真实数据集。第一阶段紧密匹配基线系统设置并训练CRNN模型,同时保持预训练的Transformer模型冻结。在第二阶段,CRNN和Transformer都使用加权自监督损耗进行微调。在第二阶段之后,我们使用微调Transformers的集合来计算训练集中所有音频片段的强伪标签。然后,在第二次迭代中,我们重复两阶段训练过程,并包含基于伪标签的蒸馏损失,在DESED的公共评估集上实现新的单模型,最先进的性能,PSDS1为0.692。基于我们提出的训练过程的单个模型和集成模型在2024年DCASE挑战赛的任务4中排名第一。摘要:A central problem in building effective sound event detection systems is the lack of high-quality, strongly annotated sound event datasets. For this reason, Task 4 of the DCASE 2024 challenge proposes learning from two heterogeneous datasets, including audio clips labeled with varying annotation granularity and with different sets of possible events. We propose a multi-iteration, multi-stage procedure for fine-tuning Audio Spectrogram Transformers on the joint DESED and MAESTRO Real datasets. The first stage closely matches the baseline system setup and trains a CRNN model while keeping the pre-trained transformer model frozen. In the second stage, both CRNN and transformer are fine-tuned using heavily weighted self-supervised losses. After the second stage, we compute strong pseudo-labels for all audio clips in the training set using an ensemble of fine-tuned transformers. Then, in a second iteration, we repeat the two-stage training process and include a distillation loss based on the pseudo-labels, achieving a new single-model, state-of-the-art performance on the public evaluation set of DESED with a PSDS1 of 0.692. A single model and an ensemble, both based on our proposed training procedure, ranked first in Task 4 of the DCASE Challenge 2024.
【4】 Aligning Sight and Sound: Advanced Sound Source Localization Through Audio-Visual Alignment
标题: 对齐视觉和声音:通过视听对齐进行高级光源定位
作者:Arda Senocak,Hyeonggon Ryu,Junsik Kim,Tae-Hyun Oh,Hanspeter Pfister,Joon Son Chung
备注:Journal Extension of ICCV 2023 paper (arXiV:2309.10724). Code is available at this https URL
链接:点击下载PDF文件
摘要:最近关于基于学习的声源定位的研究主要集中在定位性能的角度。然而,之前的工作和现有的基准忽略了一个关键方面:跨模态交互,这对于交互式声源定位至关重要。跨模态交互对于理解语义匹配或不匹配的视听事件(例如无声物体或屏幕外的声音)至关重要。在本文中,我们首先全面研究了现有方法、基准、评估指标和跨模态理解任务的跨模态交互。然后,我们确定了以前的研究的局限性,并提出了一些贡献,以克服这些局限性。首先,我们介绍了一个新的合成基准的交互式声源定位。其次,我们引入了新的评估指标来严格评估声源定位方法,重点是准确评估定位性能和跨模态交互能力。第三,我们提出了一个学习框架与跨模态对齐策略,以加强跨模态的互动。最后,我们一起评估交互式声源定位和辅助跨模态检索任务,以彻底评估跨模态交互能力和基准竞争方法。我们的新基准和评估指标揭示了声源定位研究中以前被忽视的问题。我们提出的新方法,增强跨模态对齐,显示出优越的声源定位性能。这项工作提供了迄今为止最全面的声源定位分析,并使用新的标准评估指标对现有和新的基准进行了广泛的竞争方法验证。摘要:Recent studies on learning-based sound source localization have mainly focused on the localization performance perspective. However, prior work and existing benchmarks overlook a crucial aspect: cross-modal interaction, which is essential for interactive sound source localization. Cross-modal interaction is vital for understanding semantically matched or mismatched audio-visual events, such as silent objects or off-screen sounds. In this paper, we first comprehensively examine the cross-modal interaction of existing methods, benchmarks, evaluation metrics, and cross-modal understanding tasks. Then, we identify the limitations of previous studies and make several contributions to overcome the limitations. First, we introduce a new synthetic benchmark for interactive sound source localization. Second, we introduce new evaluation metrics to rigorously assess sound source localization methods, focusing on accurately evaluating both localization performance and cross-modal interaction ability. Third, we propose a learning framework with a cross-modal alignment strategy to enhance cross-modal interaction. Lastly, we evaluate both interactive sound source localization and auxiliary cross-modal retrieval tasks together to thoroughly assess cross-modal interaction capabilities and benchmark competing methods. Our new benchmarks and evaluation metrics reveal previously overlooked issues in sound source localization studies. Our proposed novel method, with enhanced cross-modal alignment, shows superior sound source localization performance. This work provides the most comprehensive analysis of sound source localization to date, with extensive validation of competing methods on both existing and new benchmarks using new and standard evaluation metrics.
【5】 CogniVoice: Multimodal and Multilingual Fusion Networks for Mild Cognitive Impairment Assessment from Spontaneous Speech
标题: CogniVoice:用于自发言语轻度认知障碍评估的多模式和多语言融合网络
作者:Jiali Cheng,Mohamed Elgaar,Nidhi Vakil,Hadi Amiri
备注:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:轻度认知障碍(MCI)是一种医学疾病,其特征是记忆和认知能力明显下降,可能影响个人的日常活动。在本文中,我们介绍CogniVoice,一种新的多语言和多模式的框架来检测MCI和估计的简易精神状态检查(MMSE)分数通过分析语音数据及其文本transanimation。CogniVoice的关键组成部分是一个基于"专家产品“的多模式和多语言综合网络,减少了对捷径解决方案的依赖。使用TAUKADIAL挑战中包含英语和中文的综合数据集,CogniVoice在MCI分类和MMSE回归任务中的表现分别优于最佳基线模型2.8和4.1分,并且可以有效地将不同语言组之间的性能差距缩小0.7分。摘要:Mild Cognitive Impairment (MCI) is a medical condition characterized by noticeable declines in memory and cognitive abilities, potentially affecting individual's daily activities. In this paper, we introduce CogniVoice, a novel multilingual and multimodal framework to detect MCI and estimate Mini-Mental State Examination (MMSE) scores by analyzing speech data and its textual transcriptions. The key component of CogniVoice is an ensemble multimodal and multilingual network based on Product of Experts'' that mitigates reliance on shortcut solutions. Using a comprehensive dataset containing both English and Chinese languages from TAUKADIAL challenge, CogniVoice outperforms the best performing baseline model on MCI classification and MMSE regression tasks by 2.8 and 4.1 points in F1 and RMSE respectively, and can effectively reduce the performance gap across different language groups by 0.7 points in F1.
【6】 Accurate Mapping of RNNs on Neuromorphic Hardware with Adaptive Spiking Neurons
标题: 具有自适应尖峰神经元的神经形态硬件上RNN的准确映射
作者:Gauthier Boeshertz,Giacomo Indiveri,Manu Nair,Alpha Renner
备注:5 pages, 3 figures, accepted at ICONS 2024
链接:点击下载PDF文件
摘要:由于其并行和稀疏的活动特性,递归神经网络(RNN)非常适合在低功耗神经形态硬件中实现。然而,将基于速率的RNN映射到硬件兼容的尖峰神经网络(SNN)仍然具有挑战性。在这里,我们提出了一个${ Sigma}{ Delta}$-低通RNN(lpRNN):一个RNN架构,采用自适应尖峰神经元模型,使用${ Sigma}{ Delta}$-调制对信号进行编码,并实现精确映射。${ Sigma}{ Delta}$-神经元使用尖峰定时来传递模拟值,并且lpRNN的动态被设置为与处理自然信号(例如语音)的典型时间尺度相匹配。我们的方法集成了速率和时间编码,为RNN到SNN的高效准确转换提供了一个强大的解决方案。我们演示了lpRNN在英特尔神经形态研究芯片Loihi上的实现,使用3位权重在音频基准测试中实现了最先进的分类结果。这些结果要求对基于事件的系统中的递归和自适应进行更深入的研究,这可能会导致对边缘计算应用的深入了解,其中需要高能效的实时推理。摘要:Thanks to their parallel and sparse activity features, recurrent neural networks (RNNs) are well-suited for hardware implementation in low-power neuromorphic hardware. However, mapping rate-based RNNs to hardware-compatible spiking neural networks (SNNs) remains challenging. Here, we present a ${ Sigma}{ Delta}$-low-pass RNN (lpRNN): an RNN architecture employing an adaptive spiking neuron model that encodes signals using ${ Sigma}{ Delta}$-modulation and enables precise mapping. The ${ Sigma}{ Delta}$-neuron communicates analog values using spike timing, and the dynamics of the lpRNN are set to match typical timescales for processing natural signals, such as speech. Our approach integrates rate and temporal coding, offering a robust solution for the efficient and accurate conversion of RNNs to SNNs. We demonstrate the implementation of the lpRNN on Intel's neuromorphic research chip Loihi, achieving state-of-the-art classification results on audio benchmarks using 3-bit weights. These results call for a deeper investigation of recurrency and adaptation in event-based systems, which may lead to insights for edge computing applications where power-efficient real-time inference is required.
【7】 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
标题: 基于语言模型的具有可控自发行为的自发风格文本到语音合成
作者:Weiqin Li,Peiji Yang,Yicheng Zhong,Yixuan Zhou,Zhisheng Wang,Zhiyong Wu,Xixin Wu,Helen Meng
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:由于缺乏高质量的数据和模型能力的限制,旨在生成类人语音的自发式语音合成经常遇到挑战。最近的基于语言模型的TTS系统可以在大型,多样化和低质量的语音数据集上进行训练,从而产生高度自然的合成语音。然而,它们受到模拟各种自发行为和捕捉自发语音中的韵律变化的困难的限制。在本文中,我们提出了一种新的自发语音合成系统的语言模型的基础上。我们系统地对各种自发行为进行分类和统一建模。实验结果表明,本文提出的方法在韵律自然度和自发行为自然度方面明显优于传统方法。摘要:Spontaneous style speech synthesis, which aims to generate human-like speech, often encounters challenges due to the scarcity of high-quality data and limitations in model capabilities. Recent language model-based TTS systems can be trained on large, diverse, and low-quality speech datasets, resulting in highly natural synthesized speech. However, they are limited by the difficulty of simulating various spontaneous behaviors and capturing prosody variations in spontaneous speech. In this paper, we propose a novel spontaneous speech synthesis system based on language models. We systematically categorize and uniformly model diverse spontaneous behaviors. Moreover, fine-grained prosody modeling is introduced to enhance the model's ability to capture subtle prosody variations in spontaneous speech.Experimental results show that our proposed method significantly outperforms the baseline methods in terms of prosody naturalness and spontaneous behavior naturalness.
【8】 Reducing Barriers to the Use of Marginalised Music Genres in AI
标题: 减少人工智能中使用边缘化音乐流派的障碍
作者:Nick Bryan-Kinns,Zijin Li
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:用于高质量音乐生成的AI系统通常依赖于非常大的音乐数据集来训练AI模型。这就为生成超越主流数据集中所代表的流派(如西方古典音乐或流行音乐)的音乐创造了障碍。我们进行了一项为期4个月的国际研究项目,旨在探索可解释人工智能(XAI)的挑战和机遇,以减少使用人工智能模型的边缘化音乐类型的障碍。确定的XAI机会包括提高AI模型的透明度和控制,解释AI模型的道德和偏见,用小数据集微调大模型以减少偏见,以及解释AI模型的风格转移机会。研究参与者强调,虽然很难使用边缘化音乐和人工智能等小数据集,但这种方法加强了代表性不足的文化的文化代表性,并有助于解决深度学习模型的偏见问题。我们现在正在该项目的基础上汇集全球国际负责任人工智能音乐社区,并邀请人们加入我们的网络。摘要:AI systems for high quality music generation typically rely on extremely large musical datasets to train the AI models. This creates barriers to generating music beyond the genres represented in dominant datasets such as Western Classical music or pop music. We undertook a 4 month international research project summarised in this paper to explore the eXplainable AI (XAI) challenges and opportunities associated with reducing barriers to using marginalised genres of music with AI models. XAI opportunities identified included topics of improving transparency and control of AI models, explaining the ethics and bias of AI models, fine tuning large models with small datasets to reduce bias, and explaining style-transfer opportunities with AI models. Participants in the research emphasised that whilst it is hard to work with small datasets such as marginalised music and AI, such approaches strengthen cultural representation of underrepresented cultures and contribute to addressing issues of bias of deep learning models. We are now building on this project to bring together a global International Responsible AI Music community and invite people to join our network.
【9】 Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
标题: 通过低成本数据策略提高印度TTC系统的词汇外性能以满足实际应用
作者:Srija Anand,Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:公开可用的低资源语言(如印地语和泰米尔语)的TTS数据集通常包含10-20小时的数据,导致词汇覆盖率很低。这种限制在下游应用中变得明显,在下游应用中,特定领域的词汇加上与英语的频繁代码混合,导致许多OOV单词。为了突出这个问题,我们创建了一个基准包含OOV字从几个现实世界的应用程序。事实上,最先进的印地语和泰米尔语TTS系统在这个OOV基准上表现不佳,正如可懂度测试所示。为了提高模型的OOV性能,我们提出了一种低成本和经济可行的策略来获得更多的训练数据。具体来说,我们建议使用志愿者而不是高质量的语音艺术家来记录包含训练数据中看不到的字符二元组的单词。我们表明,使用这种廉价的数据,该模型的性能提高OOV的话,而不影响语音质量和域内的性能。摘要:Publicly available TTS datasets for low-resource languages like Hindi and Tamil typically contain 10-20 hours of data, leading to poor vocabulary coverage. This limitation becomes evident in downstream applications where domain-specific vocabulary coupled with frequent code-mixing with English, results in many OOV words. To highlight this problem, we create a benchmark containing OOV words from several real-world applications. Indeed, state-of-the-art Hindi and Tamil TTS systems perform poorly on this OOV benchmark, as indicated by intelligibility tests. To improve the model's OOV performance, we propose a low-effort and economically viable strategy to obtain more training data. Specifically, we propose using volunteers as opposed to high quality voice artists to record words containing character bigrams unseen in the training data. We show that using such inexpensive data, the model's performance improves on OOV words, while not affecting voice quality and in-domain performance.
【10】 Linear-Complexity Self-Supervised Learning for Speech Processing
标题: 语音处理的线性复杂度自我监督学习
作者:Shucong Zhang,Titouan Parcollet,Rogier van Dalen,Sourav Bhattacharya
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:自监督学习(SSL)模型通常需要使用数十个高端GPU进行数周的预训练。这些模型通常具有多头自注意(MHSA)上下文编码器。然而,MHSA在输入长度上占用了二次时间和空间,导致了高的预训练成本。已经提出了MHSA的线性复杂度替代方案。例如,在监督训练中,SummaryMixing模型是第一个在多个语音处理任务中优于MHSA的模型。然而,这些更便宜的替代品尚未被探索用于SSL。本文首次研究了一种线性复杂度的SSL上下文编码器。SummaryMixing在MP3S基准测试的下游任务中具有更好或相当的性能,将wav2vec 2.0模型的预训练时间和峰值VRAM分别减少了18%和23%,从而使155M wav2vec 2.0模型的预训练在一周内完成,使用4个Tesla A100 GPU。代码可在https: github.com SamsungLabs SummaryMixing上获得。摘要:Self-supervised learning (SSL) models usually require weeks of pre-training with dozens of high-end GPUs. These models typically have a multi-headed self-attention (MHSA) context encoder. However, MHSA takes quadratic time and space in the input length, contributing to the high pre-training cost. Linear-complexity alternatives to MHSA have been proposed. For instance, in supervised training, the SummaryMixing model is the first to outperform MHSA across multiple speech processing tasks. However, these cheaper alternatives have not been explored for SSL yet. This paper studies a linear-complexity context encoder for SSL for the first time. With better or equivalent performance for the downstream tasks of the MP3S benchmark, SummaryMixing reduces the pre-training time and peak VRAM of wav2vec 2.0 model by 18% and by 23%, respectively, leading to the pre-training of a 155M wav2vec 2.0 model finished within one week with 4 Tesla A100 GPUs. Code is available at https: github.com SamsungLabs SummaryMixing.
【11】 Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
标题: 在损失函数中使用语音基础模型进行助听器语音增强
作者:Robert Sutherland,George Close,Thomas Hain,Stefan Goetze,Jon Barker
备注:Accepted for EUSIPCO 2024
链接:点击下载PDF文件
摘要:机器学习技术是助听器语音增强的一个活跃研究领域,其中一个特别关注的是提高嘈杂语音信号的可懂度。最近的工作表明,自监督语音表示模型的特征编码可以有效地捕获语音可懂度。在这项工作中,它表明,清洁和嘈杂的语音的自我监督的语音表示之间的距离更强烈地与人类的可懂度评级比其他基于信号的度量。实验表明,使用该距离作为损失函数的一部分来训练语音增强模型,与使用基于SNR的损失函数相比,可以提高性能,这可以通过HASPI,STOI,PESQ和SI-SNR分数的增加来证明。该方法仅在训练时进行高参数计数模型的推断,这意味着语音增强模型可以保持较小,如助听器所需。摘要:Machine learning techniques are an active area of research for speech enhancement for hearing aids, with one particular focus on improving the intelligibility of a noisy speech signal. Recent work has shown that feature encodings from self-supervised speech representation models can effectively capture speech intelligibility. In this work, it is shown that the distance between self-supervised speech representations of clean and noisy speech correlates more strongly with human intelligibility ratings than other signal-based metrics. Experiments show that training a speech enhancement model using this distance as part of a loss function improves the performance over using an SNR-based loss function, demonstrated by an increase in HASPI, STOI, PESQ and SI-SNR scores. This method takes inference of a high parameter count model only at training time, meaning the speech enhancement model can remain smaller, as is required for hearing aids.
【12】 Robust ASR Error Correction with Conservative Data Filtering
标题: 具有保守数据过滤的鲁棒的ASB错误纠正
作者:Takuma Udagawa,Masayuki Suzuki,Masayasu Muraoka,Gakuto Kurata
链接:点击下载PDF文件
摘要:基于大语言模型的纠错技术是一种新兴的提高自动语音识别系统性能的技术。一般来说,EC的训练数据是通过自动配对大量的ASR假设(作为源)及其黄金参考(作为目标)来收集的。然而,这样的对的质量是不能保证的,我们观察到各种类型的噪声,可以使EC模型脆弱,例如,在域外(OOD)设置诱导过校正。在这项工作中,我们提出了EC训练数据应该满足的两个基本标准:即EC目标应该(1)提高源的语言可接受性和(2)从可用的上下文(例如源音素)推断。通过这些标准,我们识别出低质量的EC对,并训练模型在这种情况下不进行任何校正,这个过程我们称为保守数据过滤。在我们的实验中,我们专注于日本ASR使用强一致性CTC作为基线和微调日本LLM EC。通过我们对21个内部基准测试的评估,我们证明了我们的方法可以显着减少过度校正,并在具有挑战性的OOD设置中提高ASR结果的准确性和质量。摘要:Error correction (EC) based on large language models is an emerging technology to enhance the performance of automatic speech recognition (ASR) systems. Generally, training data for EC are collected by automatically pairing a large set of ASR hypotheses (as sources) and their gold references (as targets). However, the quality of such pairs is not guaranteed, and we observed various types of noise which can make the EC models brittle, e.g. inducing overcorrection in out-of-domain (OOD) settings. In this work, we propose two fundamental criteria that EC training data should satisfy: namely, EC targets should (1) improve linguistic acceptability over sources and (2) be inferable from the available context (e.g. source phonemes). Through these criteria, we identify low-quality EC pairs and train the models not to make any correction in such cases, the process we refer to as conservative data filtering. In our experiments, we focus on Japanese ASR using a strong Conformer-CTC as the baseline and finetune Japanese LLMs for EC. Through our evaluation on a suite of 21 internal benchmarks, we demonstrate that our approach can significantly reduce overcorrection and improve both the accuracy and quality of ASR results in the challenging OOD settings.
【13】 Low-Resourced Speech Recognition for Iu Mien Language via Weakly-Supervised Phoneme-based Multilingual Pre-training
标题: 通过弱监督的基于音素的多语言预训练实现Iu Mien语言的低资源语音识别
作者:Lukuan Dong,Donghong Qin,Fengbo Bai,Fanhua Song,Yan Liu,Chen Xu,Zhijian Ou
链接:点击下载PDF文件
摘要:主流的自动语音识别(ASR)技术通常需要数百到数千小时的带注释语音数据。低资源ASR的三种方法是基于音素或子词的监督预训练,以及多语言数据的自监督预训练。瑶缅语是我国瑶族的主要民族语言,其资源贫乏,注释语音十分有限。本文以不到10小时的瑶面语为例,研究和比较了三种瑶面语语音识别方法。我们的实验基于最近发布的三种骨干模型,这些模型在CommonVoice数据集(CV-Lang 10)的10种语言上进行了预训练,这对应于低资源ASR的三种方法。结果表明,音素监督比子词监督和自监督能获得更好的结果,从而提供更高的数据效率。特别是Whistle模型,即,通过基于弱监督音素的多语种预训练,获得最具竞争力的结果。摘要:The mainstream automatic speech recognition (ASR) technology usually requires hundreds to thousands of hours of annotated speech data. Three approaches to low-resourced ASR are phoneme or subword based supervised pre-training, and self-supervised pre-training over multilingual data. The Iu Mien language is the main ethnic language of the Yao ethnic group in China and is low-resourced in the sense that the annotated speech is very limited. With less than 10 hours of transcribed Iu Mien language, this paper investigates and compares the three approaches for Iu Mien speech recognition. Our experiments are based on the recently released, three backbone models pretrained over the 10 languages from the CommonVoice dataset (CV-Lang10), which correspond to the three approaches for low-resourced ASR. It is found that phoneme supervision can achieve better results compared to subword supervision and self-supervision, thereby providing higher data-efficiency. Particularly, the Whistle models, i.e., obtained by the weakly-supervised phoneme-based multilingual pre-training, obtain the most competitive results.
【14】 How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
标题: 低频语音音频在野外有多私密?人类和机器对言语可理解性的分析
作者:Ailin Liu,Pepijn Vunderink,Jose Vargas Quiros,Chirag Raman,Hayley Hung
备注:This manuscript has been accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:低频音频已被提出作为一种有前途的隐私保护模式,以研究在现实世界中的社会动态。为此,研究人员开发了可穿戴设备,可以以低至1250 Hz的频率记录音频,以减轻对可能包含私人细节的语音的口头内容的自动提取。本文探讨了这一假设的有效性,检查在何种程度上确保言语隐私的低频语音。它包括在各种噪声环境中模拟潜在的隐私攻击。此外,它还探讨了语音活动检测性能(这是理解社交行为的基础)与隐私保护之间的权衡。该评估结合了主观的人类可懂度和自动语音识别性能,全面分析了有效的社会行为分析和保护言语隐私之间的微妙平衡。摘要:Low-frequency audio has been proposed as a promising privacy-preserving modality to study social dynamics in real-world settings. To this end, researchers have developed wearable devices that can record audio at frequencies as low as 1250 Hz to mitigate the automatic extraction of the verbal content of speech that may contain private details. This paper investigates the validity of this hypothesis, examining the degree to which low-frequency speech ensures verbal privacy. It includes simulating a potential privacy attack in various noise environments. Further, it explores the trade-off between the performance of voice activity detection, which is fundamental for understanding social behavior, and privacy-preservation. The evaluation incorporates subjective human intelligibility and automatic speech recognition performance, comprehensively analyzing the delicate balance between effective social behavior analysis and preserving verbal privacy.
【15】 Underwater Acoustic Signal Denoising Algorithms: A Survey of the State-of-the-art
标题: 水下声学信号去噪算法:最新进展综述
作者:Ruobin Gao,Maohan Liang,Heng Dong,Xuewen Luo,P. N. Suganthan
链接:点击下载PDF文件
摘要:本文综述了水声信号去噪的最新进展,这是提高水下通信和监测系统可靠性和清晰度的关键领域。尽管该领域取得了重大进展,但水下环境的复杂性带来了独特的挑战,使去噪过程变得复杂。我们首先概述了与水声信号处理相关的基本挑战,包括信号衰减,噪声变化和环境因素的影响。然后,该综述系统地分类和讨论了各种去噪算法,如传统的,基于分解的和基于学习的技术,突出了它们的应用,优点和局限性。评估指标和实验数据集也进行了审查。本文最后列出了一系列开放性问题和建议,为未来的研究方向,强调需要开发更强大的去噪技术,可以适应动态的水声环境。摘要:This paper comprehensively reviews recent advances in underwater acoustic signal denoising, an area critical for improving the reliability and clarity of underwater communication and monitoring systems. Despite significant progress in the field, the complex nature of underwater environments poses unique challenges that complicate the denoising process. We begin by outlining the fundamental challenges associated with underwater acoustic signal processing, including signal attenuation, noise variability, and the impact of environmental factors. The review then systematically categorizes and discusses various denoising algorithms, such as conventional, decomposition-based, and learning-based techniques, highlighting their applications, advantages, and limitations. Evaluation metrics and experimental datasets are also reviewed. The paper concludes with a list of open questions and recommendations for future research directions, emphasizing the need for developing more robust denoising techniques that can adapt to the dynamic underwater acoustic environment.
【16】 DiveSound: LLM-Assisted Automatic Taxonomy Construction for Diverse Audio Generation
标题: DiveSound:LLM辅助自动分类构建,用于多样化音频生成
作者:Baihan Li,Zeyu Xie,Xuenan Xu,Yiwei Guo,Ming Yan,Ji Zhang,Kai Yu,Mengyue Wu
链接:点击下载PDF文件
摘要:音频生成引起了极大的关注。尽管在音频质量显着提高,现有的模型忽略了多样性评价。这部分是由于缺乏一个系统的健全的类多样性框架和匹配的数据集。为了解决这些问题,我们提出了DiveSound,一个新的框架,用于构建多模态数据集与类内多样化的分类,辅助大型语言模型。由于文本和视觉信息都可以用来指导多样化的生成,DiveSound在数据构建中利用了多模态对比表示。我们的框架是高度自治的,可以很容易地扩展。我们提供了一个文本图像对齐的多样性数据集,其声音事件类标签平均有2.42个子类别。在构建的数据集上进行的文本到音频的实验表明,在视觉信息的引导下,多样性大幅增加。摘要:Audio generation has attracted significant attention. Despite remarkable enhancement in audio quality, existing models overlook diversity evaluation. This is partially due to the lack of a systematic sound class diversity framework and a matching dataset. To address these issues, we propose DiveSound, a novel framework for constructing multimodal datasets with in-class diversified taxonomy, assisted by large language models. As both textual and visual information can be utilized to guide diverse generation, DiveSound leverages multimodal contrastive representations in data construction. Our framework is highly autonomous and can be easily scaled up. We provide a textaudio-image aligned diversity dataset whose sound event class tags have an average of 2.42 subcategories. Text-to-audio experiments on the constructed dataset show a substantial increase of diversity with the help of the guidance of visual information.
【17】 Preset-Voice Matching for Privacy Regulated Speech-to-Speech Translation Systems
标题: 受隐私监管的语音翻译系统的预设语音匹配
作者:Daniel Platnick,Bishoy Abdelnour,Eamon Earl,Rahul Kumar,Zahra Rezaei,Thomas Tsangaris,Faraj Lagum
备注:Accepted to the ACL PrivateNLP 2024 Workshop, 7 pages, 2 figures
链接:点击下载PDF文件
摘要:近年来,在工业环境中对语音到语音翻译(S2ST)系统的需求已经增加。虽然克隆的S2ST系统已成功商业化,但当被个人滥用时,其分销商将承担责任,当被媒体组织利用时,可能侵犯人格权。这项工作提出了一个规范的S2ST框架称为预设语音匹配(PVM)。PVM通过首先将输入语音与目标语言中类似的先前同意的说话者语音进行匹配来消除S2ST中的跨语言语音克隆。通过这种分离,PVM避免了克隆输入扬声器,确保PVM系统符合法规并降低误用风险。实验结果表明,PVM可以显著提高S2ST系统在多说话人环境下的运行时间和S2ST合成语音的自然度。据我们所知,PVM是第一个明确规定的S2ST框架,利用类似匹配的预设语音进行动态S2ST任务。摘要:In recent years, there has been increased demand for speech-to-speech translation (S2ST) systems in industry settings. Although successfully commercialized, cloning-based S2ST systems expose their distributors to liabilities when misused by individuals and can infringe on personality rights when exploited by media organizations. This work proposes a regulated S2ST framework called Preset-Voice Matching (PVM). PVM removes cross-lingual voice cloning in S2ST by first matching the input voice to a similar prior consenting speaker voice in the target-language. With this separation, PVM avoids cloning the input speaker, ensuring PVM systems comply with regulations and reduce risk of misuse. Our results demonstrate PVM can significantly improve S2ST system run-time in multi-speaker settings and the naturalness of S2ST synthesized speech. To our knowledge, PVM is the first explicitly regulated S2ST framework leveraging similarly-matched preset-voices for dynamic S2ST tasks.
【18】 A light-weight and efficient punctuation and word casing prediction model for on-device streaming ASR
标题: 一种轻量级、高效的标点符号和词性预测模型,用于设备上流媒体ASB
作者:Jian You,Xiangfeng Li
链接:点击下载PDF文件
摘要:标点符号和词的大小写预测是自动语音识别(ASR)的必要条件。随着设备上端到端流媒体ASR系统的普及,设备上的标点符号和词大小写预测成为必要,但我们发现很少有人讨论这一点。随着Transformer的出现,人们已经为此场景探索了基于Transformer的模型。然而,基于Transformer的模型对于设备上ASR系统来说太大。在本文中,我们提出了一个轻量级和高效的模型,联合预测标点符号和词的大小写实时。该模型基于卷积神经网络(CNN)和双向长短期记忆(BiLSTM)。在IWSLT2011测试集上的实验结果表明,与最好的非Transformer模型相比,该模型在总体F1分数上获得了9%的相对提高.与基于Transformer的模型的代表相比,所提出的模型实现了与代表模型相当的结果,同时仅为其大小的四十分之一,并且在推理时间方面快2.5倍。它适用于设备上流式ASR系统。我们的代码是公开的。摘要:Punctuation and word casing prediction are necessary for automatic speech recognition (ASR). With the popularity of on-device end-to-end streaming ASR systems, the on-device punctuation and word casing prediction become a necessity while we found little discussion on this. With the emergence of Transformer, Transformer based models have been explored for this scenario. However, Transformer based models are too large for on-device ASR systems. In this paper, we propose a light-weight and efficient model that jointly predicts punctuation and word casing in real time. The model is based on Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM). Experimental results on the IWSLT2011 test set show that the proposed model obtains 9% relative improvement compared to the best of non-Transformer models on overall F1-score. Compared to the representative of Transformer based models, the proposed model achieves comparable results to the representative model while being only one-fortieth its size and 2.5 times faster in terms of inference time. It is suitable for on-device streaming ASR systems. Our code is publicly available.
【19】 Audio-visual Generalized Zero-shot Learning the Easy Way
标题: 视听广义Zero-Shot学习的简单方法
作者:Shentong Mo,Pedro Morgado
链接:点击下载PDF文件
摘要:视听广义zero-shot学习是一个快速发展的领域,旨在理解视频中音频和视觉线索之间的复杂关系。首要目标是利用来自可见类的见解来识别来自以前未见过的类的实例。先前的方法主要利用同步自动编码器来重建视听属性,这些属性由交叉注意Transformers和投影文本嵌入来通知。然而,这些方法未能有效地捕捉预训练的语言对齐嵌入中固有的跨模态特征和类标签嵌入之间的复杂关系。为了克服这些瓶颈,我们引入了一个简单而有效的框架,用于简单的视听广义Zero-shot学习,名为EZ-AVGZL,将视听嵌入与转换后的文本表示对齐。它利用单一的监督文本视听对比损失来学习视听和文本模态之间的对齐,远离了重建跨模态特征和文本嵌入的传统方法。我们的主要观点是,虽然类名嵌入与基于语言的视听功能很好地结合在一起,但它们并没有提供足够的类分离,从而对zero-shot学习有用。为了解决这个问题,我们的方法利用差分优化将类嵌入转换到一个更具歧视性的空间,同时保留语言表示的语义结构。我们在VGGSound-Gandrel、UCF-Gandrel和ActivityNet-Gandrel基准上进行了广泛的实验。我们的研究结果表明,我们的EZ-AVGZL在视听广义zero-shot学习中达到了最先进的性能。摘要:Audio-visual generalized zero-shot learning is a rapidly advancing domain that seeks to understand the intricate relations between audio and visual cues within videos. The overarching goal is to leverage insights from seen classes to identify instances from previously unseen ones. Prior approaches primarily utilized synchronized auto-encoders to reconstruct audio-visual attributes, which were informed by cross-attention transformers and projected text embeddings. However, these methods fell short of effectively capturing the intricate relationship between cross-modal features and class-label embeddings inherent in pre-trained language-aligned embeddings. To circumvent these bottlenecks, we introduce a simple yet effective framework for Easy Audio-Visual Generalized Zero-shot Learning, named EZ-AVGZL, that aligns audio-visual embeddings with transformed text representations. It utilizes a single supervised text audio-visual contrastive loss to learn an alignment between audio-visual and textual modalities, moving away from the conventional approach of reconstructing cross-modal features and text embeddings. Our key insight is that while class name embeddings are well aligned with language-based audio-visual features, they don't provide sufficient class separation to be useful for zero-shot learning. To address this, our method leverages differential optimization to transform class embeddings into a more discriminative space while preserving the semantic structure of language representations. We conduct extensive experiments on VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL benchmarks. Our results demonstrate that our EZ-AVGZL achieves state-of-the-art performance in audio-visual generalized zero-shot learning.
【20】 Modeling and Driving Human Body Soundfields through Acoustic Primitives
标题: 通过声学基元建模和驱动人体音场
作者:Chao Huan,Dejan Markovic,Chenliang Xu,Alexander Richard
备注:ECCV 2024. Project Page: this https URL
链接:点击下载PDF文件
摘要:虽然真实感3D人体模型的渲染和动画在过去几年中已经成熟并达到了令人印象深刻的质量,但到目前为止,与这种全身模型相关联的空间音频建模在很大程度上被忽视了。在这项工作中,我们提出了一个框架,允许高质量的空间音频生成,能够渲染由人体生成的完整的3D声场,包括语音,脚步,手-身体交互等。给出了一个基本的视听表示的身体的形式的3D身体姿势和音频从头戴式麦克风,我们证明了我们可以在3D空间中的任何一点高效,准确地渲染完整的声学场景。为了实现声音的近场和实时渲染,我们借用了图形神经渲染中的体积基元的想法,并将它们转移到声学域。与以前的方法相比,我们的声学原语导致了一个数量级的较小的声场表示,并克服了近场渲染的不足。摘要:While rendering and animation of photorealistic 3D human body models have matured and reached an impressive quality over the past years, modeling the spatial audio associated with such full body models has been largely ignored so far. In this work, we present a framework that allows for high-quality spatial audio generation, capable of rendering the full 3D soundfield generated by a human body, including speech, footsteps, hand-body interactions, and others. Given a basic audio-visual representation of the body in form of 3D body pose and audio from a head-mounted microphone, we demonstrate that we can render the full acoustic scene at any point in 3D space efficiently and accurately. To enable near-field and realtime rendering of sound, we borrow the idea of volumetric primitives from graphical neural rendering and transfer them into the acoustic domain. Our acoustic primitives result in an order of magnitude smaller soundfield representations and overcome deficiencies in near-field rendering compared to previous approaches.
【21】 Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
标题: 预先训练的基础模型表示,以揭示言语中的呼吸模式
作者:Vikramjit Mitra,Anirban Chatterjee,Ke Zhai,Helen Weng,Ayuko Hill,Nicole Hay,Christopher Webb,Jamie Cheng,Erdrin Azemi
备注:8 pages, 6 figures, BioKDD workshop paper
链接:点击下载PDF文件
摘要:人类语音产生的过程涉及协调的呼吸动作以引出声学语音信号。通常情况下,当空气从肺中被迫出来并被声道调制时,就会产生语音,其中这种动作被呼吸空气(吸入)的时刻所穿插,以再次重新填充肺部。呼吸率(RR)是一个重要的指标,用于评估个人的整体健康,健身和一般福祉。测量RR(一分钟内呼吸的次数)的现有方法是使用专门的设备或培训进行的。研究表明,机器学习算法可用于使用生物传感器信号作为输入来估计RR。基于语音的RR估计可以提供一种有效的方法来测量生命指标,而不需要任何专门的设备或传感器。这项工作研究了一种基于机器学习的方法来估计RR从语音段从说话的近距离麦克风设备的主题。从N=26个个体收集数据,其中通过商业级胸带获得地面真实RR,然后手动校正任何错误。提出了一种卷积长短期记忆网络(Conv-LSTM)来从语音信号中估计呼吸时间序列数据。我们证明,使用从基础模型(如Wav 2 Vec 2)获得的预训练表示,与基线相比,可用于估计具有低均方根误差和高相关系数的呼吸时间序列。模型驱动的时间序列可用于估计$RR$,平均绝对误差(MAE)较低,约为1.6次呼吸 min。摘要:The process of human speech production involves coordinated respiratory action to elicit acoustic speech signals. Typically, speech is produced when air is forced from the lungs and is modulated by the vocal tract, where such actions are interspersed by moments of breathing in air (inhalation) to refill the lungs again. Respiratory rate (RR) is a vital metric that is used to assess the overall health, fitness, and general well-being of an individual. Existing approaches to measure RR (number of breaths one takes in a minute) are performed using specialized equipment or training. Studies have demonstrated that machine learning algorithms can be used to estimate RR using bio-sensor signals as input. Speech-based estimation of RR can offer an effective approach to measure the vital metric without requiring any specialized equipment or sensors. This work investigates a machine learning based approach to estimate RR from speech segments obtained from subjects speaking to a close-talking microphone device. Data were collected from N=26 individuals, where the groundtruth RR was obtained through commercial grade chest-belts and then manually corrected for any errors. A convolutional long-short term memory network (Conv-LSTM) is proposed to estimate respiration time-series data from the speech signal. We demonstrate that the use of pre-trained representations obtained from a foundation model, such as Wav2Vec2, can be used to estimate respiration-time-series with low root-mean-squared error and high correlation coefficient, when compared with the baseline. The model-driven time series can be used to estimate $RR$ with a low mean absolute error (MAE) ~ 1.6 breaths min.
【22】 Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
标题: 通过关注自动语音识别的声学和置信参考来纠正错误
作者:Yuchun Shu,Bo Hu,Yifeng He,Hao Shi,Longbiao Wang,Jianwu Dang
链接:点击下载PDF文件
摘要:语音纠错的目标是准确地发现自动语音识别(ASR)假设中的错误单词并将其恢复为有根据的。本文提出了一种非自回归语音纠错方法。置信度模块测量N个最佳ASR假设中每个词的不确定性,作为找到错误词位置的参考。此外,ASR编码器的声学特征也被用来提供正确的发音参考。使用编辑路径对齐来自ASR的N个最佳候选,以相互确认并恢复一些丢失的字符错误。此外,交叉注意机制融合了错误校正参考和ASR假设之间的信息。实验结果表明,声学参考和置信度参考都有助于纠错。与ASR模型相比,该系统将错误率降低了21%。摘要:Accurately finding the wrong words in the automatic speech recognition (ASR) hypothesis and recovering them well-founded is the goal of speech error correction. In this paper, we propose a non-autoregressive speech error correction method. A Confidence Module measures the uncertainty of each word of the N-best ASR hypotheses as the reference to find the wrong word position. Besides, the acoustic feature from the ASR encoder is also used to provide the correct pronunciation references. N-best candidates from ASR are aligned using the edit path, to confirm each other and recover some missing character errors. Furthermore, the cross-attention mechanism fuses the information between error correction references and the ASR hypothesis. The experimental results show that both the acoustic and confidence references help with error correction. The proposed system reduces the error rate by 21% compared with the ASR model.
机器翻译,仅供参考
