今日论文合集:cs.SD语音22篇,eess.AS音频处理15篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Few-Shot Keyword Spotting from Mixed Speech
标题: 混合语音中的少数关键词发现
作者:Junming Yuan,Ying Shi,LanTian Li,Dong Wang,Askar Hamdulla
备注:accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:Few-Shot关键词发现(KWS)的目的是检测未知的关键词与有限的训练样本。一种常用的方法是预训练和微调框架。虽然这种方法在干净的条件下有效,但它难以识别混合的关键词-同时检测话语中混合的多个关键词,这在现实世界的应用中至关重要。之前的研究提出了一种混合训练(MT)方法来解决这个问题,然而,它从未在Few-Shot场景中进行过测试。在本文中,我们探讨了使用MT和其他相关方法来解决两个实际挑战的可能性:Few-Shot和混合语音。在LibriSpeech和Google Speech Command语料库上进行的实验表明,MT在预训练阶段或微调阶段都非常有效。此外,将基于SSL的大规模预训练(HuBert)和MT微调相结合,在所有测试条件下都会产生非常强大的结果。摘要:Few-shot keyword spotting (KWS) aims to detect unknown keywords with limited training samples. A commonly used approach is the pre-training and fine-tuning framework. While effective in clean conditions, this approach struggles with mixed keyword spotting -- simultaneously detecting multiple keywords blended in an utterance, which is crucial in real-world applications. Previous research has proposed a Mix-Training (MT) approach to solve the problem, however, it has never been tested in the few-shot scenario. In this paper, we investigate the possibility of using MT and other relevant methods to solve the two practical challenges together: few-shot and mixed speech. Experiments conducted on the LibriSpeech and Google Speech Command corpora demonstrate that MT is highly effective on this task when employed in either the pre-training phase or the fine-tuning phase. Moreover, combining SSL-based large-scale pre-training (HuBert) and MT fine-tuning yields very strong results in all the test conditions.

【2】 MERGE -- A Bimodal Dataset for Static Music Emotion Recognition
标题: 合并--静态音乐情感识别的双峰数据集
作者:Pedro Lima Louro,Hugo Redinho,Ricardo Santos,Ricardo Malheiro,Renato Panda,Rui Pedro Paiva
备注:16 pages, 4 figures, 13 tables, submitted to IEEE Transactions on Affective Computing
链接:点击下载PDF文件
摘要:近年来,音乐情感识别(MER)领域取得了稳步发展,其中包括特征工程、机器学习和深度学习。景观也从以音频为中心的系统转移到结合音频和歌词的双峰合奏。然而,严重缺乏公共和大规模的双峰数据库阻碍了双峰音频歌词系统的发展和改进。本文提出了三个新的音频,歌词和双峰MER研究数据集,统称为MERGE,使用半自动方法创建。为了全面评估拟议的数据集并建立基准测试的基线,我们使用特征工程、机器学习和深度学习方法对每种模式进行了几次实验。此外,我们提出并验证了固定的train-validate-test分裂。所获得的结果证实了所提出的数据集的可行性,使用深度神经网络实现了双峰分类的79.21%F1分数的最佳总体结果。摘要:The Music Emotion Recognition (MER) field has seen steady developments in recent years, with contributions from feature engineering, machine learning, and deep learning. The landscape has also shifted from audio-centric systems to bimodal ensembles that combine audio and lyrics. However, a severe lack of public and sizeable bimodal databases has hampered the development and improvement of bimodal audio-lyrics systems. This article proposes three new audio, lyrics, and bimodal MER research datasets, collectively called MERGE, created using a semi-automatic approach. To comprehensively assess the proposed datasets and establish a baseline for benchmarking, we conducted several experiments for each modality, using feature engineering, machine learning, and deep learning methodologies. In addition, we propose and validate fixed train-validate-test splits. The obtained results confirm the viability of the proposed datasets, achieving the best overall result of 79.21% F1-score for bimodal classification using a deep neural network.

【3】 Cervical Auscultation Machine Learning for Dysphagia Assessment
标题: 用于吞咽困难评估的颈部听诊机器学习
作者:An An Chia,Stacy Lum,Michelle Boo,Rex Tan,Balamurali B T,Jer-Ming Chen
备注:International Conference on Signal Processing and Communications (SPCOM) July 01 - 04, 2024
链接:点击下载PDF文件
摘要:这项研究评估了机器学习的使用,特别是随机森林分类器,以区分正常和病理性吞咽音。采用市售的可穿戴听诊器,我们记录了健康成人和吞咽困难患者的吞咽。分析显示,正常和病理吞咽之间的声学特征,如频谱波峰和过零率的统计学显着差异,而不同的液体和饮食之间没有区别。该系统对吞咽困难表现出相当的灵敏度(平均值± SD:74% ± 8%)和特异性(89% ± 6%)。该模型的总体准确率为83% ± 3%,F1得分为78% ± 5%。这些结果表明,机器学习可以成为非侵入性吞咽困难评估中的一种有价值的工具,尽管注意到了诸如采样率限制以及区分正常和病理声音的灵敏度和特异性的变化等挑战。该研究强调了进一步研究以优化这些技术用于临床的必要性。摘要:This study evaluates the use of machine learning, specifically the Random Forest Classifier, to differentiate normal and pathological swallowing sounds. Employing a commercially available wearable stethoscope, we recorded swallows from both healthy adults and patients with dysphagia. The analysis revealed statistically significant differences in acoustic features, such as spectral crest, and zero-crossing rate between normal and pathological swallows, while no discriminating differences were demonstrated between different fluidand diet consistencies. The system demonstrated fair sensitivity (mean plus or minus SD: 74% plus or minus 8%) and specificity (89% plus or minus 6%) for dysphagic swallows. The model attained an overall accuracy of 83% plus or minus 3%, and F1 score of 78% plus or minus 5%. These results demonstrate that machine learning can be a valuable tool in non-invasive dysphagia assessment, although challenges such as sampling rate limitations and variability in sensitivity and specificity in discriminating between normal and pathological sounds are noted. The study underscores the need for further research to optimize these techniques for clinical use.

【4】 Sequential Contrastive Audio-Visual Learning
标题: 顺序对比视听学习
作者:Ioannis Tsiamas,Santiago Pascual,Chunghsin Yeh,Joan Serrà
链接:点击下载PDF文件
摘要:对比学习已经成为视听表征学习中的一种强大技术,它利用了大量网络视频数据集中视听模态的自然共现来实现重大进步。然而,传统的对比视听学习方法往往依赖于通过时间聚合得到的聚合表示,这忽略了数据的内在顺序性。这种疏忽引起了人们对标准方法捕获和利用序列中细粒度信息的能力的担忧,这些信息对于区分语义相似但不同的示例至关重要。针对这一局限性,我们提出了顺序对比视听学习(SCAV),它使用顺序距离基于非聚合表示空间对比示例。使用VGGSound和Music数据集的检索实验证明了SCAV的有效性,与传统的基于聚合的对比学习和文献中的其他方法相比,SCAV显示出2- 3倍的相对改进。我们还表明,使用SCAV训练的模型在检索所采用的度量方面表现出高度的灵活性,使它们能够在效率-准确性权衡的范围内进行操作,从而可能使它们适用于从小规模到大规模检索的多种场景。摘要:Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in extensive web-scale video datasets to achieve significant advancements. However, conventional contrastive audio-visual learning methodologies often rely on aggregated representations derived through temporal aggregation, which neglects the intrinsic sequential nature of the data. This oversight raises concerns regarding the ability of standard approaches to capture and utilize fine-grained information within sequences, information that is vital for distinguishing between semantically similar yet distinct examples. In response to this limitation, we propose sequential contrastive audio-visual learning (SCAV), which contrasts examples based on their non-aggregated representation space using sequential distances. Retrieval experiments with the VGGSound and Music datasets demonstrate the effectiveness of SCAV, showing 2-3x relative improvements against traditional aggregation-based contrastive learning and other methods from the literature. We also show that models trained with SCAV exhibit a high degree of flexibility regarding the metric employed for retrieval, allowing them to operate on a spectrum of efficiency-accuracy trade-offs, potentially making them applicable in multiple scenarios, from small- to large-scale retrieval.

【5】 MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
标题: MSP-播客BER挑战2024:L ' antenne du Ventoux多模式自我监督学习语音情感识别
作者:Jarod Duret,Mickael Rouvier,Yannick Estève
Journal-ref:Odyssey 2024, Jun 2024, Quebec, France
链接:点击下载PDF文件
摘要:在这项工作中,我们详细介绍了我们提交给2024年版的MSP播客语音情感识别(SER)挑战赛。这个挑战分为两个不同的任务:分类情感识别和情感属性预测。我们将精力集中在任务1上,该任务涉及使用来自MSP-Podcast数据集的数据对八种情绪状态进行分类。我们的方法采用了集成的模型,每个独立训练,然后融合在分数水平使用支持向量机(SVM)分类器。这些模型使用各种策略进行训练,包括跨不同模式的自我监督学习(SSL)微调:单独的语音,单独的文本以及语音和文本相结合的方法。这种联合训练方法旨在增强系统准确分类情绪状态的能力。这种联合训练方法旨在增强系统准确分类情绪状态的能力。因此,该系统在开发集上获得了0.35%的F1-macro。摘要:In this work, we detail our submission to the 2024 edition of the MSP-Podcast Speech Emotion Recognition (SER) Challenge. This challenge is divided into two distinct tasks: Categorical Emotion Recognition and Emotional Attribute Prediction. We concentrated our efforts on Task 1, which involves the categorical classification of eight emotional states using data from the MSP-Podcast dataset. Our approach employs an ensemble of models, each trained independently and then fused at the score level using a Support Vector Machine (SVM) classifier. The models were trained using various strategies, including Self-Supervised Learning (SSL) fine-tuning across different modalities: speech alone, text alone, and a combined speech and text approach. This joint training methodology aims to enhance the system's ability to accurately classify emotional states. This joint training methodology aims to enhance the system's ability to accurately classify emotional states. Thus, the system obtained F1-macro of 0.35 % on development set.

【6】 A Benchmark for Multi-speaker Anonymization
标题: 多扬声器模拟化的基准
作者:Xiaoxiao Miao,Ruijie Tao,Chang Zeng,Xin Wang
链接:点击下载PDF文件
摘要:隐私保护语音保护方法主要是抑制隐私相关的信息,而保留语言的内容,从语言属性。现有的解决方案集中在单个扬声器场景。然而,它们对于现实世界的应用缺乏实用性,即,多扬声器场景。在本文中,我们提出了一个初步的尝试,通过定义的任务和评估协议,提出基准测试的解决方案,并讨论重叠的会话的隐私泄漏提供一个多说话人匿名基准。具体来说,理想的多说话者匿名化应该保留说话者的数量和会话的话轮结构,确保准确的上下文传达,同时保持隐私。为了实现这一点,级联系统使用说话人日记来聚合每个说话人的语音,并使用说话人匿名来隐藏说话人隐私并保留语音内容。此外,我们提出了两个会话级说话人向量匿名化方法,以进一步提高效用。这两种方法的目标都是使每个说话人的原始身份和对应的伪说话人身份不一致,同时保持甚至提高会话中伪说话人之间的可识别性。第一种方法最小化原始会话和匿名会话中说话人对之间的差异相似性,以保持匿名版本中的原始说话人关系。另一种方法最小化匿名说话者之间的聚合相似性,以实现说话者之间的更好区分。在非重叠的模拟和真实数据集上进行的实验证明了所提出的说话人匿名器的多说话人匿名系统的有效性。此外,我们分析了关于隐私泄露的重叠语音,并提供了潜在的解决方案。摘要:Privacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus on single-speaker scenarios. However, they lack practicality for real-world applications, i.e., multi-speaker scenarios. In this paper, we present an initial attempt to provide a multi-speaker anonymization benchmark by defining the task and evaluation protocol, proposing benchmarking solutions, and discussing the privacy leakage of overlapping conversations. Specifically, ideal multi-speaker anonymization should preserve the number of speakers and the turn-taking structure of the conversation, ensuring accurate context conveyance while maintaining privacy. To achieve that, a cascaded system uses speaker diarization to aggregate the speech of each speaker and speaker anonymization to conceal speaker privacy and preserve speech content. Additionally, we propose two conversation-level speaker vector anonymization methods to improve the utility further. Both methods aim to make the original and corresponding pseudo-speaker identities of each speaker unlinkable while preserving or even improving the distinguishability among pseudo-speakers in a conversation. The first method minimizes the differential similarity across speaker pairs in the original and anonymized conversations to maintain original speaker relationships in the anonymized version. The other method minimizes the aggregated similarity across anonymized speakers to achieve better differentiation between speakers. Experiments conducted on both non-overlap simulated and real-world datasets demonstrate the effectiveness of the multi-speaker anonymization system with the proposed speaker anonymizers. Additionally, we analyzed overlapping speech regarding privacy leakage and provide potential solutions.

【7】 Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
标题: 用于ASV欺骗检测的双路径GMM-ResNet和GMM-SENet
作者:Zhenchun Lei,Hui Yan,Changhong Liu,Minglei Ma,Yingen Yang
链接:点击下载PDF文件
摘要:说话人自动确认系统有时容易受到各种欺骗攻击。真假语音的2类高斯混合模型分类器通常被用作欺骗检测的基线。然而,GMM分类器不单独考虑每个高斯分量上的特征帧的分数。此外,GMM独立地累积所有帧上的分数,并且不考虑它们的相关性。我们提出了用于欺骗检测的双路径GMM-ResNet和GMM-SENet模型,其输入是基于分别对真实语音和欺骗语音训练的两个GARCH的高斯概率特征。该模型不仅考虑了GMM分量上的分数分布,而且考虑了相邻帧之间的关系。采用两步训练方案提高系统的鲁棒性。在ASVspoof 2019上的实验表明,与GMM相比,LFCC+GMM-ResNet系统在逻辑访问场景下可以相对降低min-tDCF和EER 76.1%和76.3%,而LFCC+GMM-SENet系统在物理访问场景下可以降低94.4%和95.4%。分数融合后,系统在两种情况下都给出了第二好的结果。摘要:The automatic speaker verification system is sometimes vulnerable to various spoofing attacks. The 2-class Gaussian Mixture Model classifier for genuine and spoofed speech is usually used as the baseline for spoofing detection. However, the GMM classifier does not separately consider the scores of feature frames on each Gaussian component. In addition, the GMM accumulates the scores on all frames independently, and does not consider their correlations. We propose the two-path GMM-ResNet and GMM-SENet models for spoofing detection, whose input is the Gaussian probability features based on two GMMs trained on genuine and spoofed speech respectively. The models consider not only the score distribution on GMM components, but also the relationship between adjacent frames. A two-step training scheme is applied to improve the system robustness. Experiments on the ASVspoof 2019 show that the LFCC+GMM-ResNet system can relatively reduce min-tDCF and EER by 76.1% and 76.3% on logical access scenario compared with the GMM, and the LFCC+GMM-SENet system by 94.4% and 95.4% on physical access scenario. After score fusion, the systems give the second-best results on both scenarios.

【8】 Read, Watch and Scream! Sound Generation from Text and Video
标题: 阅读、观看并尖叫!从文本和视频生成声音
作者:Yujin Jeong,Yunji Kim,Sanghyuk Chun,Jiyoung Lee
备注:Project page: this https URL
链接:点击下载PDF文件
摘要:在强大的扩散模型的帮助下,多模态生成模型已经显示出令人印象深刻的进步。尽管取得了进展,但仅从文本生成声音在确保全面的场景描绘和时间对齐方面带来了挑战。同时,视频到声音的生成限制了为场景内的特定对象优先考虑声音合成的灵活性。为了解决这些挑战,我们提出了一种新的视频和文本到声音的生成方法,称为ReWaS,其中视频作为一个条件控制的文本到音频生成模型。我们的方法估计的结构信息的音频(即,能量)从视频,同时从用户提示接收关键内容线索。我们采用了一个性能良好的文本到声音模型来巩固视频控制,这对于训练具有大量三元组配对(音频-视频-文本)数据的多模态扩散模型更有效。此外,通过分离音频的生成组件,它成为一个更灵活的系统,允许用户根据自己的喜好自由调整能量,周围环境和主要声源。实验结果表明,我们的方法在质量,可控性和训练效率方面表现出优越性。我们的演示可在https: naver-ai.github.io rewas上获得摘要:Multimodal generative models have shown impressive advances with the help of powerful diffusion models. Despite the progress, generating sound solely from text poses challenges in ensuring comprehensive scene depiction and temporal alignment. Meanwhile, video-to-sound generation limits the flexibility to prioritize sound synthesis for specific objects within the scene. To tackle these challenges, we propose a novel video-and-text-to-sound generation method, called ReWaS, where video serves as a conditional control for a text-to-audio generation model. Our method estimates the structural information of audio (namely, energy) from the video while receiving key content cues from a user prompt. We employ a well-performing text-to-sound model to consolidate the video control, which is much more efficient for training multimodal diffusion models with massive triplet-paired (audio-video-text) data. In addition, by separating the generative components of audio, it becomes a more flexible system that allows users to freely adjust the energy, surrounding environment, and primary sound source according to their preferences. Experimental results demonstrate that our method shows superiority in terms of quality, controllability, and training efficiency. Our demo is available at https: naver-ai.github.io rewas

【9】 CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
标题: CosyVoice:基于监督语义令牌的可扩展多语言Zero-Shot文本到语音合成器
作者:Zhihao Du,Qian Chen,Shiliang Zhang,Kai Hu,Heng Lu,Yexin Yang,Hangrui Hu,Siqi Zheng,Yue Gu,Ziyang Ma,Zhijie Yan
备注:work in progress. arXiv admin note: substantial text overlap with arXiv:2407.04051
链接:点击下载PDF文件
摘要:近年来,基于大语言模型(LLM)的文语转换(TTS)以其高自然度和zero-shot容量成为主流。在该范例中,语音信号被离散成令牌序列,该令牌序列由具有文本作为提示的LLM建模,并由基于令牌的声码器重构为波形。显然,语音标记在基于LLM的TTS模型中起着至关重要的作用。目前的语音标记是以无监督的方式学习的,缺乏明确的语义信息和与文本的对齐。在本文中,我们建议表示语音与监督语义令牌,这是来自多语言语音识别模型插入矢量量化到编码器。在此基础上,我们进一步提出了一个可扩展的zero-shot TTS合成器CosyVoice,它由一个用于文本到令牌生成的LLM和一个用于令牌到语音合成的条件流匹配模型组成。实验结果表明,在zero-shot语音克隆中,有监督语义标记在内容一致性和说话人相似性方面明显优于现有的无监督语义标记.此外,我们发现,利用大规模的数据进一步提高合成性能,表明可扩展能力的CosyVoice。据我们所知,这是第一次尝试将有监督的语音令牌纳入TTS模型。摘要:Recent years have witnessed a trend that large language model (LLM) based text-to-speech (TTS) emerges into the mainstream due to their high naturalness and zero-shot capacity. In this paradigm, speech signals are discretized into token sequences, which are modeled by an LLM with text as prompts and reconstructed by a token-based vocoder to waveforms. Obviously, speech tokens play a critical role in LLM-based TTS models. Current speech tokens are learned in an unsupervised manner, which lacks explicit semantic information and alignment to the text. In this paper, we propose to represent speech with supervised semantic tokens, which are derived from a multilingual speech recognition model by inserting vector quantization into the encoder. Based on the tokens, we further propose a scalable zero-shot TTS synthesizer, CosyVoice, which consists of an LLM for text-to-token generation and a conditional flow matching model for token-to-speech synthesis. Experimental results show that supervised semantic tokens significantly outperform existing unsupervised tokens in terms of content consistency and speaker similarity for zero-shot voice cloning. Moreover, we find that utilizing large-scale data further improves the synthesis performance, indicating the scalable capacity of CosyVoice. To the best of our knowledge, this is the first attempt to involve supervised speech tokens into TTS models.

【10】 Research on the Acoustic Emission Source Localization Methodology in Composite Materials based on Artificial Intelligence
标题: 基于人工智能的复合材料声发射源定位方法研究
作者:Jongick Won,Hyuntaik Oh,Jae Sakong
链接:点击下载PDF文件
摘要:本文提出了一种基于人工智能的复合材料声发射源定位方法。选用碳纤维增强塑料作为试件,利用压电器件测量了声发射信号。对测量信号进行小波变换,得到尺度图,作为人工智能模型的训练数据。AESLNet(声发射源定位网络),在这项研究中提出的,由于复合材料的各向异性的卷积层并行构造。声发射源位置坐标的检测是一种回归模型。采用贝叶斯优化方法对网络的超参数进行优化。实验结果表明,该网络可以检测出声发射源的位置,平均误差为3.02mm,分辨率为20mm。摘要:In this study, methodology of acoustic emission source localization in composite materials based on artificial intelligence was presented. Carbon fiber reinforced plastic was selected for specimen, and acoustic emission signal were measured using piezoelectric devices. The measured signal was wavelet-transformed to obtain scalograms, which were used as training data for the artificial intelligence model. AESLNet(acoustic emission source localization network), proposed in this study, was constructed convolutional layers in parallel due to anisotropy of the composited materials. It is regression model to detect the coordinates of acoustic emission source location. Hyper-parameter of network has been optimized by Bayesian optimization. It has been confirmed that network can detect location of acoustic emission source with an average error of 3.02mm and a resolution of 20mm.

【11】 Music Era Recognition Using Supervised Contrastive Learning and Artist Information
标题: 使用监督对比学习和艺术家信息的音乐时代识别
作者:Qiqi He,Xuchen Song,Weituo Hao,Ju-Chiang Wang,Wei-Tsung Lu,Wei Li
链接:点击下载PDF文件
摘要:60年代的流行音乐听起来和90年代的不一样吗?先前的研究表明,在几十年的趋势中,会存在一些与仪器变化和响度增加有关的模式和模式的变化。这表明从诸如音频和艺术家信息的音乐特征感知歌曲的时代是可能的。音乐时代信息可以是用于播放列表生成和推荐的重要特征。然而,在许多情况下,歌曲的发行年份是无法访问的。本文提出了一种新的音乐时代识别任务。我们制定的任务作为一个音乐分类问题,并提出了基于监督对比学习的解决方案。一个基于音频的模型被开发来从音频预测时代。对于艺术家信息可用的情况,我们扩展了基于音频的模型,以获取多模态输入,并开发了一个称为多模态对比(MMC)学习的框架,以增强训练。在百万歌曲数据集上的实验结果表明,基于音频的模型在3年的容错范围内达到了54%的准确率;将艺术家信息与MMC框架相结合进行训练,进一步提高了9%。摘要:Does popular music from the 60s sound different than that of the 90s? Prior study has shown that there would exist some variations of patterns and regularities related to instrumentation changes and growing loudness across multi-decadal trends. This indicates that perceiving the era of a song from musical features such as audio and artist information is possible. Music era information can be an important feature for playlist generation and recommendation. However, the release year of a song can be inaccessible in many circumstances. This paper addresses a novel task of music era recognition. We formulate the task as a music classification problem and propose solutions based on supervised contrastive learning. An audio-based model is developed to predict the era from audio. For the case where the artist information is available, we extend the audio-based model to take multimodal inputs and develop a framework, called MultiModal Contrastive (MMC) learning, to enhance the training. Experimental result on Million Song Dataset demonstrates that the audio-based model achieves 54% in accuracy with a tolerance of 3-years range; incorporating the artist information with the MMC framework for training leads to 9% improvement further.

【12】 A Layer-Anchoring Strategy for Enhancing Cross-Lingual Speech Emotion Recognition
标题: 增强跨语言语音情感识别的分层锚定策略
作者:Shreya G. Upadhyay,Carlos Busso,Chi-Chun Lee
链接:点击下载PDF文件
摘要:跨语言语音情感识别(SER)对于广泛的日常应用是重要的。虽然最近的SER研究在很大程度上依赖于大型的预训练模型进行情绪训练,但现有的研究往往只集中在这些模型的最终Transformer层。但是,由于这些模型的任务特定性质和层次结构,每个Transformer层封装不同级别的信息。利用这种层次结构,我们的研究重点是嵌入在不同层的信息。通过研究不同语言之间的层特征相似性,我们提出了一种新的策略,称为层锚定机制,以促进跨语言SER任务中的情感转移。我们的方法是使用两个不同的语言情感语料库(MSP播客和BIIC播客)进行评估,实现了最好的UAR性能的60.21%的BIIC播客语料库。该分析揭示了对流行的预训练模型行为的有趣见解。摘要:Cross-lingual speech emotion recognition (SER) is important for a wide range of everyday applications. While recent SER research relies heavily on large pretrained models for emotion training, existing studies often concentrate solely on the final transformer layer of these models. However, given the task-specific nature and hierarchical architecture of these models, each transformer layer encapsulates different levels of information. Leveraging this hierarchical structure, our study focuses on the information embedded across different layers. Through an examination of layer feature similarity across different languages, we propose a novel strategy called a layer-anchoring mechanism to facilitate emotion transfer in cross-lingual SER tasks. Our approach is evaluated using two distinct language affective corpora (MSP-Podcast and BIIC-Podcast), achieving a best UAR performance of 60.21% on the BIIC-podcast corpus. The analysis uncovers interesting insights into the behavior of popular pretrained models.

【13】 A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
标题: 使用对比语音-音频预训练的语音查询音频源分离的无参考指标
作者:Feiyang Xiao,Jian Guan,Qiaoxi Zhu,Xubo Liu,Wenbo Wang,Shuhan Qi,Kejia Zhang,Jianyuan Sun,Wenwu Wang
备注:Submitted to DCASE 2024 Workshop
链接:点击下载PDF文件
摘要:语音查询音频源分离(LASS)旨在分离由文本查询引导的音频源,其中基于信号失真比(SDR)的度量通常用于客观地测量分离的音频的质量。然而,基于SDR的度量需要参考信号,这在现实世界的场景中通常难以获得。此外,基于SDR的度量,文本查询的内容信息没有被有效地考虑在LASS。本文介绍了一种使用对比语言音频预训练(CLAP)模块,称为CLAPSCore,它测量分离的音频和文本查询之间的语义相似性的参考无评价指标。与SDR不同,所提出的CLAPScore度量基于文本查询的内容信息来评估分离的音频的质量,而不需要参考信号。实验结果表明,CLAPScore指标提供了一个有效的评估的语义相关性的分离音频的文本查询,相比SDR指标,提供了一种替代的LASS系统的性能评估。摘要:Language-queried audio source separation (LASS) aims to separate an audio source guided by a text query, with the signal-to-distortion ratio (SDR)-based metrics being commonly used to objectively measure the quality of the separated audio. However, the SDR-based metrics require a reference signal, which is often difficult to obtain in real-world scenarios. In addition, with the SDR-based metrics, the content information of the text query is not considered effectively in LASS. This paper introduces a reference-free evaluation metric using a contrastive language-audio pretraining (CLAP) module, termed CLAPScore, which measures the semantic similarity between the separated audio and the text query. Unlike SDR, the proposed CLAPScore metric evaluates the quality of the separated audio based on the content information of the text query, without needing a reference signal. Experimental results show that the CLAPScore metric provides an effective evaluation of the semantic relevance of the separated audio to the text query, as compared to the SDR metric, offering an alternative for the performance evaluation of LASS systems.

【14】 All Neural Low-latency Directional Speech Extraction
标题: 全神经低延迟定向语音提取
作者:Ashutosh Pandey,Sanha Lee,Juan Azcarreta,Daniel Wong,Buye Xu
备注:Accepted for publication at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们介绍了一种新的低延迟方向性语音提取的全神经模型。该模型使用来自预定义空间网格的到达方向(DOA)嵌入,将其转换并融合到基于递归神经网络的语音提取模型中。该过程使模型能够有效地从指定的DOA中提取语音。与以前依赖于手工制作的方向特征的方法不同,所提出的模型使用语音增强损失从头开始训练DOA嵌入,使其适用于低延迟场景。此外,它以高帧速率运行,在每个输入帧中接收DOA,这带来了在高度动态的真实世界场景中快速适应变化场景的能力。我们提供了广泛的评估,以证明该模型的有效性,方向性语音提取,DOA失配的鲁棒性,其能力,以快速适应突然变化的DOA。摘要:We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are transformed and fused into a recurrent neural network based speech extraction model. This process enables the model to effectively extract speech from a specified DOA. Unlike previous methods that relied on hand-crafted directional features, the proposed model trains DOA embeddings from scratch using speech enhancement loss, making it suitable for low-latency scenarios. Additionally, it operates at a high frame rate, taking in DOA with each input frame, which brings in the capability of quickly adapting to changing scene in highly dynamic real-world scenarios. We provide extensive evaluation to demonstrate the model's efficacy in directional speech extraction, robustness to DOA mismatch, and its capability to quickly adapt to abrupt changes in DOA.

【15】 MUSIC-lite: Efficient MUSIC using Approximate Computing: An OFDM Radar Case Study
标题: MUSIC-Lite:使用近似计算的高效音乐:一个CDMA雷达案例研究
作者:Rajat Bhattacharjya,Arnab Sarkar,Biswadip Maity,Nikil Dutt
备注:Paper accepted at ESWEEK-CASES 2024 as a Late Breaking (LB) Result paper. The definitive version of the work will appear in IEEE Embedded Systems Letters
链接:点击下载PDF文件
摘要:多信号分类(MUSIC)是广泛使用的到达方向(DoA) 到达角(AoA)估计算法,其应用于各种应用领域,诸如自动驾驶、医学成像和天文学。然而,MUSIC在计算上是昂贵的,并且在低功率硬件中实现是具有挑战性的,需要在精度、成本和功率之间进行权衡。我们提出了MUSIC-精简,它利用近似计算来生成一个设计空间,探索精度面积功率权衡。这特别适用于正交频分复用(OFDM)雷达用例中MUSIC算法的计算密集型奇异值分解(SVD)组件。MUSIC-lite将近似加法器集成到用于MUSIC硬件实现的迭代CORDIC算法中,产生有趣的精度-面积-功率权衡。我们的实验表明,MUSIC-lite的能力,以节省平均17.25%的片上面积和19.4%的功率与最小的0.14%的错误,有效的MUSIC实现。摘要:Multiple Signal Classification (MUSIC) is a widely used Direction of Arrival (DoA) Angle of Arrival (AoA) estimation algorithm applied to various application domains such as autonomous driving, medical imaging, and astronomy. However, MUSIC is computationally expensive and challenging to implement in low-power hardware, requiring exploration of trade-offs between accuracy, cost, and power. We present MUSIC-lite, which exploits approximate computing to generate a design space exploring accuracy-area-power trade-offs. This is specifically applied to the computationally intensive singular value decomposition (SVD) component of the MUSIC algorithm in an orthogonal frequency-division multiplexing (OFDM) radar use case. MUSIC-lite incorporates approximate adders into the iterative CORDIC algorithm that is used for hardware implementation of MUSIC, generating interesting accuracy-area-power trade-offs. Our experiments demonstrate MUSIC-lite's ability to save an average of 17.25% on-chip area and 19.4% power with a minimal 0.14% error for efficient MUSIC implementations.

【16】 Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
标题: 基于用于对婴儿发声进行集群的基于拓广信号表示的Dirichlet过程混合模型
作者:Guillem Bonafos,Clara Bourot,Pierre Pudlo,Jean-Marc Freyermuth,Laurence Reboul,Samuel Tronçon,Arnaud Rey
链接:点击下载PDF文件
摘要:基于在儿童生命的前12个月每月一次的音频记录,我们提出了一种新的方法来聚类这组发声。我们使用拓扑增强表示的发声,采用两个持久性图为每个发声:一个计算表面上的声谱图和一个Takens的嵌入的发声。一个合成的持久性变量是来自每个图,并添加到MFCC(梅尔频率倒谱系数)。使用这种表示,我们适合一个非参数贝叶斯混合模型与Dirichlet过程之前,模型的组件的数量。这一过程导致了一种新的数据驱动的声乐作品的分类。我们的研究结果揭示了8簇发声的存在,使我们能够比较它们在生命的前12个月的时间分布和声学特征。摘要:Based on audio recordings made once a month during the first 12 months of a child's life, we propose a new method for clustering this set of vocalizations. We use a topologically augmented representation of the vocalizations, employing two persistence diagrams for each vocalization: one computed on the surface of its spectrogram and one on the Takens' embeddings of the vocalization. A synthetic persistent variable is derived for each diagram and added to the MFCCs (Mel-frequency cepstral coefficients). Using this representation, we fit a non-parametric Bayesian mixture model with a Dirichlet process prior to model the number of components. This procedure leads to a novel data-driven categorization of vocal productions. Our findings reveal the presence of 8 clusters of vocalizations, allowing us to compare their temporal distribution and acoustic profiles in the first 12 months of life.

【17】 Automating Urban Soundscape Enhancements with AI: In-situ Assessment of Quality and Restorativeness in Traffic-Exposed Residential Areas
标题: 利用人工智能自动化城市声景增强:对公共设施暴露住宅区的质量和恢复性进行现场评估
作者:Bhan Lam,Zhen-Ting Ong,Kenneth Ooi,Wen-Hui Ong,Trevor Wong,Karn N. Watcharasupat,Vanessa Boey,Irene Lee,Joo Young Hong,Jian Kang,Kar Fye Alvin Lee,Georgios Christopoulos,Woon-Seng Gan
备注:41 pages, 4 figures. Preprint submitted to an Elsevier journal
链接:点击下载PDF文件
摘要:在ISO 12913中正式提出的“声景”方法是向基于感知的城市声音管理的范式转变,旨在减轻噪音污染的重大社会经济成本,以推进联合国可持续发展目标。针对交通暴露的室外住宅场地,我们实现了一个自动掩蔽选择系统(AMSS),利用自然声音来掩蔽(或增强)交通音景。我们采用预先训练的AI模型自动选择最佳掩蔽器并调整其播放水平,以适应周围环境随时间的变化,从而最大限度地提高“愉悦度”,这是ISO 12913中声景质量的感知维度。我们的验证研究涉及($N=68$)居民揭示了一个显着的14.6%的提高“愉快”干预后,与增加的积极性和积极的影响。交通暴露现场的感知增强与较安静的控制现场的感知增强相匹配,低6 dB(A)L_ text{A,eq}$和道路交通噪声优势,肯定了AMSS作为声景干预的有效性,同时简化了劳动密集型的“愉悦度”评估与概率AI预测。摘要:Formalized in ISO 12913, the "soundscape" approach is a paradigmatic shift towards perception-based urban sound management, aiming to alleviate the substantial socioeconomic costs of noise pollution to advance the United Nations Sustainable Development Goals. Focusing on traffic-exposed outdoor residential sites, we implemented an automatic masker selection system (AMSS) utilizing natural sounds to mask (or augment) traffic soundscapes. We employed a pre-trained AI model to automatically select the optimal masker and adjust its playback level, adapting to changes over time in the ambient environment to maximize "Pleasantness", a perceptual dimension of soundscape quality in ISO 12913. Our validation study involving ($N=68$) residents revealed a significant 14.6 % enhancement in "Pleasantness" after intervention, correlating with increased restorativeness and positive affect. Perceptual enhancements at the traffic-exposed site matched those at a quieter control site with 6 dB(A) lower $L_ text{A,eq}$ and road traffic noise dominance, affirming the efficacy of AMSS as a soundscape intervention, while streamlining the labour-intensive assessment of "Pleasantness" with probabilistic AI prediction.

【18】 Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
标题: 平面弦声音物理建模与运动模拟的微振综合
作者:Jin Woo Lee,Jaehyun Park,Min Jun Choi,Kyogu Lee
链接:点击下载PDF文件
摘要:虽然机器学习和计算机听觉中的音乐生成和可微分声音合成已经取得了重大进展,但由物理定律指导的乐器振动模拟尚未得到充分探索。为了解决这一差距,我们引入了一种新的模型来模拟非线性字符串的时空运动,在神经网络框架内集成模态合成和频谱建模。我们的模型利用物理特性和基频作为输入,输出跨时间和空间的弦状态,求解表征非线性弦的偏微分方程。经验评估表明,所提出的架构实现优越的精度相比,现有的基线架构的字符串运动模拟。代码和演示可在线获得。摘要:While significant advancements have been made in music generation and differentiable sound synthesis within machine learning and computer audition, the simulation of instrument vibration guided by physical laws has been underexplored. To address this gap, we introduce a novel model for simulating the spatio-temporal motion of nonlinear strings, integrating modal synthesis and spectral modeling within a neural network framework. Our model leverages physical properties and fundamental frequencies as inputs, outputting string states across time and space that solve the partial differential equation characterizing the nonlinear string. Empirical evaluations demonstrate that the proposed architecture achieves superior accuracy in string motion simulation compared to existing baseline architectures. The code and demo are available online.

【19】 Fine-Grained and Interpretable Neural Speech Editing
标题: 细粒度且可解释的神经语音编辑
作者:Max Morrison,Cameron Churchwell,Nathan Pruyne,Bryan Pardo
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:语音属性$ unicode{x2014}$的细粒度编辑,例如韵律(即,音高、响度和音素持续时间)、发音、说话者身份和共振峰$ unicode{x2014}$可用于微调和修复人类和AI生成的语音记录中的缺陷,以创建播客、电影对话和视频游戏对话。现有的语音合成系统使用纠缠两个或多个这些属性的表示,禁止它们在细粒度的,分离的编辑中使用。在本文中,我们展示了第一个解开和可解释的语音表示具有可比的主观和客观的声码重建精度梅尔频谱。我们的可解释表示与我们提出的数据增强方法相结合,能够训练现有的神经声码器对音高,持续时间,音量,音量的音色相关性,发音,扬声器身份和频谱平衡进行快速,准确和高质量的编辑。摘要:Fine-grained editing of speech attributes$ unicode{x2014}$such as prosody (i.e., the pitch, loudness, and phoneme durations), pronunciation, speaker identity, and formants$ unicode{x2014}$is useful for fine-tuning and fixing imperfections in human and AI-generated speech recordings for creation of podcasts, film dialogue, and video game dialogue. Existing speech synthesis systems use representations that entangle two or more of these attributes, prohibiting their use in fine-grained, disentangled editing. In this paper, we demonstrate the first disentangled and interpretable representation of speech with comparable subjective and objective vocoding reconstruction accuracy to Mel spectrograms. Our interpretable representation, combined with our proposed data augmentation method, enables training an existing neural vocoder to perform fast, accurate, and high-quality editing of pitch, duration, volume, timbral correlates of volume, pronunciation, speaker identity, and spectral balance.

【20】 ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
标题: ASRRL-TTC:用于文本到语音说话人适应的敏捷说话人表示强化学习
作者:Ruibo Fu,Xin Qi,Zhengqi Wen,Jianhua Tao,Tao Wang,Chunyu Qiang,Zhiyong Wang,Yi Lu,Xiaopeng Wang,Shuchen Shi,Yukun Liu,Xuefei Liu,Shuai Zhang
备注:The audio demo is available at this https URL
链接:点击下载PDF文件
摘要:说话人自适应是指在文本到语音转换任务中克隆说话人的语音,由于其在多媒体领域的广泛应用而引起了人们的极大兴趣。尽管最近的进步,现有的方法往往与不足的扬声器表示的准确性和过拟合,特别是在有限的参考语音场景中的斗争。为了解决这些挑战,我们提出了一种敏捷的说话人表示强化学习策略,以提高说话人自适应任务中的说话人相似性。ASRRL是第一个应用强化学习来提高说话人自适应中说话人嵌入建模准确性的工作,解决了解耦语音内容和音色的挑战。我们的方法介绍了两个行动策略,针对不同的参考演讲场景。在单句场景中,采用面向知识的最优例程搜索RL方法,以加快对说话人表示边缘的细化信息的探索和检索。在句子较少的情况下,我们利用动态强化学习方法自适应地融合参考语音,增强了说话人建模的鲁棒性和准确性。为了在目标域中实现最佳结果,提出了一种基于多尺度融合评分机制的奖励模型,该模型评估说话人相似性、语音质量和三维可懂度,确保说话人相似性的改善不会损害语音质量或可懂度。在主流TTS框架下的LibriTTS和VCTK数据集上的实验结果证明了所提出的ASRRL方法的可扩展性和泛化能力。结果表明,ASRRL方法显着优于传统的微调方法,实现更高的说话人相似性和更好的整体语音质量与有限的参考语音。摘要:Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media fields. Despite recent advancements, existing methods often struggle with inadequate speaker representation accuracy and overfitting, particularly in limited reference speeches scenarios. To address these challenges, we propose an Agile Speaker Representation Reinforcement Learning strategy to enhance speaker similarity in speaker adaptation tasks. ASRRL is the first work to apply reinforcement learning to improve the modeling accuracy of speaker embeddings in speaker adaptation, addressing the challenge of decoupling voice content and timbre. Our approach introduces two action strategies tailored to different reference speeches scenarios. In the single-sentence scenario, a knowledge-oriented optimal routine searching RL method is employed to expedite the exploration and retrieval of refinement information on the fringe of speaker representations. In the few-sentence scenario, we utilize a dynamic RL method to adaptively fuse reference speeches, enhancing the robustness and accuracy of speaker modeling. To achieve optimal results in the target domain, a multi-scale fusion scoring mechanism based reward model that evaluates speaker similarity, speech quality, and intelligibility across three dimensions is proposed, ensuring that improvements in speaker similarity do not compromise speech quality or intelligibility. The experimental results on the LibriTTS and VCTK datasets within mainstream TTS frameworks demonstrate the extensibility and generalization capabilities of the proposed ASRRL method. The results indicate that the ASRRL method significantly outperforms traditional fine-tuning approaches, achieving higher speaker similarity and better overall speech quality with limited reference speeches.

【21】 Ternary Spike-based Neuromorphic Signal Processing System
标题: 基于三元尖峰的神经形态信号处理系统
作者:Shuai Wang,Dehao Zhang,Ammar Belatreche,Yichen Xiao,Hongyu Qing,Wenjie We,Malu Zhang,Yang Yang
链接:点击下载PDF文件
摘要:深度神经网络(DNN)已在各种信号处理领域成功实现,从而显著提高了性能。然而,DNN通常需要大量的计算资源,导致显著的经济成本,并对其在资源受限的边缘设备上的部署构成挑战。在这项研究中,我们利用尖峰神经网络(SNN)和量化技术,开发一个节能和轻量级的神经形态信号处理系统。我们的系统的特点是两个主要的创新:阈值自适应编码(TAE)方法和量化的三进制SNN(QT-SNN)。TAE方法可以有效地将时变模拟信号编码成稀疏的三进制尖峰序列,从而减少信号处理的能量和存储器需求。QT-SNN与来自TAE方法的三进制尖峰序列兼容,量化膜电位和突触权重以降低存储器要求,同时保持性能。在两个典型的信号处理任务:语音和脑电识别进行了广泛的实验。结果表明,我们的神经形态信号处理系统实现了最先进的(SOTA)的性能,减少了94%的内存需求。此外,通过理论能耗分析,与其他SNN工程相比,我们的系统节能7.5倍。所提出的系统的效率和功效突出了其作为节能信号处理的有前途的途径的潜力。摘要:Deep Neural Networks (DNNs) have been successfully implemented across various signal processing fields, resulting in significant enhancements in performance. However, DNNs generally require substantial computational resources, leading to significant economic costs and posing challenges for their deployment on resource-constrained edge devices. In this study, we take advantage of spiking neural networks (SNNs) and quantization technologies to develop an energy-efficient and lightweight neuromorphic signal processing system. Our system is characterized by two principal innovations: a threshold-adaptive encoding (TAE) method and a quantized ternary SNN (QT-SNN). The TAE method can efficiently encode time-varying analog signals into sparse ternary spike trains, thereby reducing energy and memory demands for signal processing. QT-SNN, compatible with ternary spike trains from the TAE method, quantifies both membrane potentials and synaptic weights to reduce memory requirements while maintaining performance. Extensive experiments are conducted on two typical signal-processing tasks: speech and electroencephalogram recognition. The results demonstrate that our neuromorphic signal processing system achieves state-of-the-art (SOTA) performance with a 94% reduced memory requirement. Furthermore, through theoretical energy consumption analysis, our system shows 7.5x energy saving compared to other SNN works. The efficiency and efficacy of the proposed system highlight its potential as a promising avenue for energy-efficient signal processing.

【22】 YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
标题: YourMT 3+:具有增强型Transformer架构和跨数据集主干增强的多乐器音乐转录
作者:Sungkyun Chang,Emmanouil Benetos,Holger Kirchhoff,Simon Dixon
备注:Preprint submitted to IEEE MLSP 2024
链接:点击下载PDF文件
摘要:多乐器音乐转录旨在将复调音乐录音转换为分配给每种乐器的乐谱。这项任务对于建模来说是具有挑战性的,因为它需要同时识别多个乐器并转录它们的音高和精确的计时,并且缺乏完全注释的数据增加了训练难度。本文介绍了YourMT 3+,这是一套基于MT3最新语言令牌解码方法的增强型多乐器音乐转录模型。我们通过在时频域中采用分层注意力Transformer和集成专家混合(MoE)来增强其编码器。为了解决数据限制,我们引入了一种新的多通道解码方法来训练不完整注释,并提出了用于数据集混合的茎内和跨茎增强。我们的实验证明了直接的声音转录能力,无需语音分离预处理器。十个公共数据集的基准测试显示了我们的模型与现有转录模型的竞争力或优越性。对流行音乐录音的进一步测试凸显了当前模型的局限性。完全可复制的代码和数据集可在 url{https: github.com mimbres YourMT3}摘要:Multi-instrument music transcription aims to convert polyphonic music recordings into musical scores assigned to each instrument. This task is challenging for modeling as it requires simultaneously identifying multiple instruments and transcribing their pitch and precise timing, and the lack of fully annotated data adds to the training difficulties. This paper introduces YourMT3+, a suite of models for enhanced multi-instrument music transcription based on the recent language token decoding approach of MT3. We strengthen its encoder by adopting a hierarchical attention transformer in the time-frequency domain and integrating a mixture of experts (MoE). To address data limitations, we introduce a new multi-channel decoding method for training with incomplete annotations and propose intra- and cross-stem augmentation for dataset mixing. Our experiments demonstrate direct vocal transcription capabilities, eliminating the need for voice separation pre-processors. Benchmarks across ten public datasets show our models' competitiveness with, or superiority to, existing transcription models. Further testing on pop music recordings highlights the limitations of current models. Fully reproducible code and datasets are available at url{https: github.com mimbres YourMT3}


eess.AS音频处理
【1】 Automating Urban Soundscape Enhancements with AI: In-situ Assessment of Quality and Restorativeness in Traffic-Exposed Residential Areas
标题: 利用人工智能自动化城市声景增强:对公共设施暴露住宅区的质量和恢复性进行现场评估
作者:Bhan Lam,Zhen-Ting Ong,Kenneth Ooi,Wen-Hui Ong,Trevor Wong,Karn N. Watcharasupat,Vanessa Boey,Irene Lee,Joo Young Hong,Jian Kang,Kar Fye Alvin Lee,Georgios Christopoulos,Woon-Seng Gan
备注:41 pages, 4 figures. Preprint submitted to an Elsevier journal
链接:点击下载PDF文件
摘要:在ISO 12913中正式提出的“声景”方法是向基于感知的城市声音管理的范式转变,旨在减轻噪音污染的重大社会经济成本,以推进联合国可持续发展目标。针对交通暴露的室外住宅场地,我们实现了一个自动掩蔽选择系统(AMSS),利用自然声音来掩蔽(或增强)交通音景。我们采用预先训练的AI模型自动选择最佳掩蔽器并调整其播放水平,以适应周围环境随时间的变化,从而最大限度地提高“愉悦度”,这是ISO 12913中声景质量的感知维度。我们的验证研究涉及($N=68$)居民揭示了一个显着的14.6%的提高“愉快”干预后,与增加的积极性和积极的影响。交通暴露现场的感知增强与较安静的控制现场的感知增强相匹配,低6 dB(A)L_ text{A,eq}$和道路交通噪声优势,肯定了AMSS作为声景干预的有效性,同时简化了劳动密集型的“愉悦度”评估与概率AI预测。摘要:Formalized in ISO 12913, the "soundscape" approach is a paradigmatic shift towards perception-based urban sound management, aiming to alleviate the substantial socioeconomic costs of noise pollution to advance the United Nations Sustainable Development Goals. Focusing on traffic-exposed outdoor residential sites, we implemented an automatic masker selection system (AMSS) utilizing natural sounds to mask (or augment) traffic soundscapes. We employed a pre-trained AI model to automatically select the optimal masker and adjust its playback level, adapting to changes over time in the ambient environment to maximize "Pleasantness", a perceptual dimension of soundscape quality in ISO 12913. Our validation study involving ($N=68$) residents revealed a significant 14.6 % enhancement in "Pleasantness" after intervention, correlating with increased restorativeness and positive affect. Perceptual enhancements at the traffic-exposed site matched those at a quieter control site with 6 dB(A) lower $L_ text{A,eq}$ and road traffic noise dominance, affirming the efficacy of AMSS as a soundscape intervention, while streamlining the labour-intensive assessment of "Pleasantness" with probabilistic AI prediction.

【2】 Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
标题: 平面弦声音物理建模与运动模拟的微振综合
作者:Jin Woo Lee,Jaehyun Park,Min Jun Choi,Kyogu Lee
链接:点击下载PDF文件
摘要:虽然机器学习和计算机听觉中的音乐生成和可微分声音合成已经取得了重大进展,但由物理定律指导的乐器振动模拟尚未得到充分探索。为了解决这一差距,我们引入了一种新的模型来模拟非线性字符串的时空运动,在神经网络框架内集成模态合成和频谱建模。我们的模型利用物理特性和基频作为输入,输出跨时间和空间的弦状态,求解表征非线性弦的偏微分方程。经验评估表明,所提出的架构实现优越的精度相比,现有的基线架构的字符串运动模拟。代码和演示可在线获得。摘要:While significant advancements have been made in music generation and differentiable sound synthesis within machine learning and computer audition, the simulation of instrument vibration guided by physical laws has been underexplored. To address this gap, we introduce a novel model for simulating the spatio-temporal motion of nonlinear strings, integrating modal synthesis and spectral modeling within a neural network framework. Our model leverages physical properties and fundamental frequencies as inputs, outputting string states across time and space that solve the partial differential equation characterizing the nonlinear string. Empirical evaluations demonstrate that the proposed architecture achieves superior accuracy in string motion simulation compared to existing baseline architectures. The code and demo are available online.

【3】 Fine-Grained and Interpretable Neural Speech Editing
标题: 细粒度且可解释的神经语音编辑
作者:Max Morrison,Cameron Churchwell,Nathan Pruyne,Bryan Pardo
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:语音属性$ unicode{x2014}$的细粒度编辑,例如韵律(即,音高、响度和音素持续时间)、发音、说话者身份和共振峰$ unicode{x2014}$可用于微调和修复人类和AI生成的语音记录中的缺陷,以创建播客、电影对话和视频游戏对话。现有的语音合成系统使用纠缠两个或多个这些属性的表示,禁止它们在细粒度的,分离的编辑中使用。在本文中,我们展示了第一个解开和可解释的语音表示具有可比的主观和客观的声码重建精度梅尔频谱。我们的可解释表示与我们提出的数据增强方法相结合,能够训练现有的神经声码器对音高,持续时间,音量,音量的音色相关性,发音,扬声器身份和频谱平衡进行快速,准确和高质量的编辑。摘要:Fine-grained editing of speech attributes$ unicode{x2014}$such as prosody (i.e., the pitch, loudness, and phoneme durations), pronunciation, speaker identity, and formants$ unicode{x2014}$is useful for fine-tuning and fixing imperfections in human and AI-generated speech recordings for creation of podcasts, film dialogue, and video game dialogue. Existing speech synthesis systems use representations that entangle two or more of these attributes, prohibiting their use in fine-grained, disentangled editing. In this paper, we demonstrate the first disentangled and interpretable representation of speech with comparable subjective and objective vocoding reconstruction accuracy to Mel spectrograms. Our interpretable representation, combined with our proposed data augmentation method, enables training an existing neural vocoder to perform fast, accurate, and high-quality editing of pitch, duration, volume, timbral correlates of volume, pronunciation, speaker identity, and spectral balance.

【4】 ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
标题: ASRRL-TTC:用于文本到语音说话人适应的敏捷说话人表示强化学习
作者:Ruibo Fu,Xin Qi,Zhengqi Wen,Jianhua Tao,Tao Wang,Chunyu Qiang,Zhiyong Wang,Yi Lu,Xiaopeng Wang,Shuchen Shi,Yukun Liu,Xuefei Liu,Shuai Zhang
备注:The audio demo is available at this https URL
链接:点击下载PDF文件
摘要:None摘要:Speaker adaptation, which involves cloning voices from unseen speakers in the Text-to-Speech task, has garnered significant interest due to its numerous applications in multi-media fields. Despite recent advancements, existing methods often struggle with inadequate speaker representation accuracy and overfitting, particularly in limited reference speeches scenarios. To address these challenges, we propose an Agile Speaker Representation Reinforcement Learning strategy to enhance speaker similarity in speaker adaptation tasks. ASRRL is the first work to apply reinforcement learning to improve the modeling accuracy of speaker embeddings in speaker adaptation, addressing the challenge of decoupling voice content and timbre. Our approach introduces two action strategies tailored to different reference speeches scenarios. In the single-sentence scenario, a knowledge-oriented optimal routine searching RL method is employed to expedite the exploration and retrieval of refinement information on the fringe of speaker representations. In the few-sentence scenario, we utilize a dynamic RL method to adaptively fuse reference speeches, enhancing the robustness and accuracy of speaker modeling. To achieve optimal results in the target domain, a multi-scale fusion scoring mechanism based reward model that evaluates speaker similarity, speech quality, and intelligibility across three dimensions is proposed, ensuring that improvements in speaker similarity do not compromise speech quality or intelligibility. The experimental results on the LibriTTS and VCTK datasets within mainstream TTS frameworks demonstrate the extensibility and generalization capabilities of the proposed ASRRL method. The results indicate that the ASRRL method significantly outperforms traditional fine-tuning approaches, achieving higher speaker similarity and better overall speech quality with limited reference speeches.

【5】 Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
标题: Emilia:用于大规模语音生成的广泛、多语言且多样化的语音数据集
作者:Haorui He,Zengqiang Shang,Chaoren Wang,Xuyuan Li,Yicheng Gu,Hua Hua,Liwei Liu,Chen Yang,Jiaqi Li,Peiyang Shi,Yuancheng Wang,Kai Chen,Pengyuan Zhang,Zhizheng Wu
链接:点击下载PDF文件
摘要:近年来,通过使用大规模训练数据,语音生成模型取得了重大进展。然而,由于缺乏大规模、多样性和自发的语音数据,研究界很难产生高度自发的、类似人类的语音。本文介绍了 textit{Emilia},第一个来自野外语音数据的多语言语音生成数据集,以及Emilia-Pipe,第一个开源预处理管道,旨在将野外语音数据转换为高质量的训练数据,并带有语音生成的注释。Emilia以六种语言超过101,000小时的演讲开始,并以各种演讲风格为特色。为了促进Emilia的规模扩大,开源管道Emilia-Pipe可以在几分钟内处理一小时的原始语音数据,为模型训练做好准备,这使研究社区能够在大规模语音生成研究方面进行合作。实验结果验证了Emilia算法的有效性。演示可在https: emilia-dataset.github.io Emilia-Demo-Page 上获得。摘要:Recently, speech generation models have made significant progress by using large-scale training data. However, the research community struggle to produce highly spontaneous and human-like speech due to the lack of large-scale, diverse, and spontaneous speech data. This paper presents textit{Emilia}, the first multilingual speech generation dataset from in-the-wild speech data, and Emilia-Pipe, the first open-source preprocessing pipeline designed to transform in-the-wild speech data into high-quality training data with annotations for speech generation. Emilia starts with over 101k hours of speech in six languages and features diverse speech with varied speaking styles. To facilitate the scale-up of Emilia, the open-source pipeline Emilia-Pipe can process one hour of raw speech data ready for model training in a few mins, which enables the research community to collaborate on large-scale speech generation research. Experimental results validate the effectiveness of Emilia. Demos are available at: https: emilia-dataset.github.io Emilia-Demo-Page .

【6】 YourMT3+: Multi-instrument Music Transcription with Enhanced Transformer Architectures and Cross-dataset Stem Augmentation
标题: YourMT 3+:具有增强型Transformer架构和跨数据集主干增强的多乐器音乐转录
作者:Sungkyun Chang,Emmanouil Benetos,Holger Kirchhoff,Simon Dixon
备注:Preprint submitted to IEEE MLSP 2024
链接:点击下载PDF文件
摘要:多乐器音乐转录旨在将复调音乐录音转换为分配给每种乐器的乐谱。这项任务对于建模来说是具有挑战性的,因为它需要同时识别多个乐器并转录它们的音高和精确的计时,并且缺乏完全注释的数据增加了训练难度。本文介绍了YourMT 3+,这是一套基于MT3最新语言令牌解码方法的增强型多乐器音乐转录模型。我们通过在时频域中采用分层注意力Transformer和集成专家混合(MoE)来增强其编码器。为了解决数据限制,我们引入了一种新的多通道解码方法来训练不完整注释,并提出了用于数据集混合的茎内和跨茎增强。我们的实验证明了直接的声音转录能力,无需语音分离预处理器。十个公共数据集的基准测试显示了我们的模型与现有转录模型的竞争力或优越性。对流行音乐录音的进一步测试凸显了当前模型的局限性。完全可复制的代码和数据集可在 url{https: github.com mimbres YourMT3}摘要:Multi-instrument music transcription aims to convert polyphonic music recordings into musical scores assigned to each instrument. This task is challenging for modeling as it requires simultaneously identifying multiple instruments and transcribing their pitch and precise timing, and the lack of fully annotated data adds to the training difficulties. This paper introduces YourMT3+, a suite of models for enhanced multi-instrument music transcription based on the recent language token decoding approach of MT3. We strengthen its encoder by adopting a hierarchical attention transformer in the time-frequency domain and integrating a mixture of experts (MoE). To address data limitations, we introduce a new multi-channel decoding method for training with incomplete annotations and propose intra- and cross-stem augmentation for dataset mixing. Our experiments demonstrate direct vocal transcription capabilities, eliminating the need for voice separation pre-processors. Benchmarks across ten public datasets show our models' competitiveness with, or superiority to, existing transcription models. Further testing on pop music recordings highlights the limitations of current models. Fully reproducible code and datasets are available at url{https: github.com mimbres YourMT3}

【7】 Cervical Auscultation Machine Learning for Dysphagia Assessment
标题: 用于吞咽困难评估的颈部听诊机器学习
作者:An An Chia,Stacy Lum,Michelle Boo,Rex Tan,Balamurali B T,Jer-Ming Chen
备注:International Conference on Signal Processing and Communications (SPCOM) July 01 - 04, 2024
链接:点击下载PDF文件
摘要:这项研究评估了机器学习的使用,特别是随机森林分类器,以区分正常和病理性吞咽音。采用市售的可穿戴听诊器,我们记录了健康成人和吞咽困难患者的吞咽。分析显示,正常和病理吞咽之间的声学特征,如频谱波峰和过零率的统计学显着差异,而不同的液体和饮食之间没有区别。该系统对吞咽困难表现出相当的灵敏度(平均值± SD:74% ± 8%)和特异性(89% ± 6%)。该模型的总体准确率为83% ± 3%,F1得分为78% ± 5%。这些结果表明,机器学习可以成为非侵入性吞咽困难评估中的一种有价值的工具,尽管注意到了诸如采样率限制以及区分正常和病理声音的灵敏度和特异性的变化等挑战。该研究强调了进一步研究以优化这些技术用于临床的必要性。摘要:This study evaluates the use of machine learning, specifically the Random Forest Classifier, to differentiate normal and pathological swallowing sounds. Employing a commercially available wearable stethoscope, we recorded swallows from both healthy adults and patients with dysphagia. The analysis revealed statistically significant differences in acoustic features, such as spectral crest, and zero-crossing rate between normal and pathological swallows, while no discriminating differences were demonstrated between different fluidand diet consistencies. The system demonstrated fair sensitivity (mean plus or minus SD: 74% plus or minus 8%) and specificity (89% plus or minus 6%) for dysphagic swallows. The model attained an overall accuracy of 83% plus or minus 3%, and F1 score of 78% plus or minus 5%. These results demonstrate that machine learning can be a valuable tool in non-invasive dysphagia assessment, although challenges such as sampling rate limitations and variability in sensitivity and specificity in discriminating between normal and pathological sounds are noted. The study underscores the need for further research to optimize these techniques for clinical use.

【8】 Sequential Contrastive Audio-Visual Learning
标题: 顺序对比视听学习
作者:Ioannis Tsiamas,Santiago Pascual,Chunghsin Yeh,Joan Serrà
链接:点击下载PDF文件
摘要:对比学习已经成为视听表征学习中的一种强大技术,它利用了大量网络视频数据集中视听模态的自然共现来实现重大进步。然而,传统的对比视听学习方法往往依赖于通过时间聚合得到的聚合表示,这忽略了数据的内在顺序性。这种疏忽引起了人们对标准方法捕获和利用序列中细粒度信息的能力的担忧,这些信息对于区分语义相似但不同的示例至关重要。针对这一局限性,我们提出了顺序对比视听学习(SCAV),它使用顺序距离基于非聚合表示空间对比示例。使用VGGSound和Music数据集的检索实验证明了SCAV的有效性,与传统的基于聚合的对比学习和文献中的其他方法相比,SCAV显示出2- 3倍的相对改进。我们还表明,使用SCAV训练的模型在检索所采用的度量方面表现出高度的灵活性,使它们能够在效率-准确性权衡的范围内进行操作,从而可能使它们适用于从小规模到大规模检索的多种场景。摘要:Contrastive learning has emerged as a powerful technique in audio-visual representation learning, leveraging the natural co-occurrence of audio and visual modalities in extensive web-scale video datasets to achieve significant advancements. However, conventional contrastive audio-visual learning methodologies often rely on aggregated representations derived through temporal aggregation, which neglects the intrinsic sequential nature of the data. This oversight raises concerns regarding the ability of standard approaches to capture and utilize fine-grained information within sequences, information that is vital for distinguishing between semantically similar yet distinct examples. In response to this limitation, we propose sequential contrastive audio-visual learning (SCAV), which contrasts examples based on their non-aggregated representation space using sequential distances. Retrieval experiments with the VGGSound and Music datasets demonstrate the effectiveness of SCAV, showing 2-3x relative improvements against traditional aggregation-based contrastive learning and other methods from the literature. We also show that models trained with SCAV exhibit a high degree of flexibility regarding the metric employed for retrieval, allowing them to operate on a spectrum of efficiency-accuracy trade-offs, potentially making them applicable in multiple scenarios, from small- to large-scale retrieval.

【9】 Dirichlet process mixture model based on topologically augmented signal representation for clustering infant vocalizations
标题: 基于用于对婴儿发声进行集群的基于拓广信号表示的Dirichlet过程混合模型
作者:Guillem Bonafos,Clara Bourot,Pierre Pudlo,Jean-Marc Freyermuth,Laurence Reboul,Samuel Tronçon,Arnaud Rey
链接:点击下载PDF文件
摘要:基于在儿童生命的前12个月每月一次的音频记录,我们提出了一种新的方法来聚类这组发声。我们使用拓扑增强表示的发声,采用两个持久性图为每个发声:一个计算表面上的声谱图和一个Takens的嵌入的发声。一个合成的持久性变量是来自每个图,并添加到MFCC(梅尔频率倒谱系数)。使用这种表示,我们适合一个非参数贝叶斯混合模型与Dirichlet过程之前,模型的组件的数量。这一过程导致了一种新的数据驱动的声乐作品的分类。我们的研究结果揭示了8簇发声的存在,使我们能够比较它们在生命的前12个月的时间分布和声学特征。摘要:Based on audio recordings made once a month during the first 12 months of a child's life, we propose a new method for clustering this set of vocalizations. We use a topologically augmented representation of the vocalizations, employing two persistence diagrams for each vocalization: one computed on the surface of its spectrogram and one on the Takens' embeddings of the vocalization. A synthetic persistent variable is derived for each diagram and added to the MFCCs (Mel-frequency cepstral coefficients). Using this representation, we fit a non-parametric Bayesian mixture model with a Dirichlet process prior to model the number of components. This procedure leads to a novel data-driven categorization of vocal productions. Our findings reveal the presence of 8 clusters of vocalizations, allowing us to compare their temporal distribution and acoustic profiles in the first 12 months of life.

【10】 MSP-Podcast SER Challenge 2024: L'antenne du Ventoux Multimodal Self-Supervised Learning for Speech Emotion Recognition
标题: MSP-播客BER挑战2024:L ' antenne du Ventoux多模式自我监督学习语音情感识别
作者:Jarod Duret,Mickael Rouvier,Yannick Estève
Journal-ref:Odyssey 2024, Jun 2024, Quebec, France
链接:点击下载PDF文件
摘要:在这项工作中,我们详细介绍了我们提交给2024年版的MSP播客语音情感识别(SER)挑战赛。这个挑战分为两个不同的任务:分类情感识别和情感属性预测。我们将精力集中在任务1上,该任务涉及使用来自MSP-Podcast数据集的数据对八种情绪状态进行分类。我们的方法采用了集成的模型,每个独立训练,然后融合在分数水平使用支持向量机(SVM)分类器。这些模型使用各种策略进行训练,包括跨不同模式的自我监督学习(SSL)微调:单独的语音,单独的文本以及语音和文本相结合的方法。这种联合训练方法旨在增强系统准确分类情绪状态的能力。这种联合训练方法旨在增强系统准确分类情绪状态的能力。因此,该系统在开发集上获得了0.35%的F1-macro。摘要:In this work, we detail our submission to the 2024 edition of the MSP-Podcast Speech Emotion Recognition (SER) Challenge. This challenge is divided into two distinct tasks: Categorical Emotion Recognition and Emotional Attribute Prediction. We concentrated our efforts on Task 1, which involves the categorical classification of eight emotional states using data from the MSP-Podcast dataset. Our approach employs an ensemble of models, each trained independently and then fused at the score level using a Support Vector Machine (SVM) classifier. The models were trained using various strategies, including Self-Supervised Learning (SSL) fine-tuning across different modalities: speech alone, text alone, and a combined speech and text approach. This joint training methodology aims to enhance the system's ability to accurately classify emotional states. This joint training methodology aims to enhance the system's ability to accurately classify emotional states. Thus, the system obtained F1-macro of 0.35 % on development set.

【11】 Two-Path GMM-ResNet and GMM-SENet for ASV Spoofing Detection
标题: 用于ASV欺骗检测的双路径GMM-ResNet和GMM-SENet
作者:Zhenchun Lei,Hui Yan,Changhong Liu,Minglei Ma,Yingen Yang
链接:点击下载PDF文件
摘要:说话人自动确认系统有时容易受到各种欺骗攻击。真假语音的2类高斯混合模型分类器通常被用作欺骗检测的基线。然而,GMM分类器不单独考虑每个高斯分量上的特征帧的分数。此外,GMM独立地累积所有帧上的分数,并且不考虑它们的相关性。我们提出了用于欺骗检测的双路径GMM-ResNet和GMM-SENet模型,其输入是基于分别对真实语音和欺骗语音训练的两个GARCH的高斯概率特征。该模型不仅考虑了GMM分量上的分数分布,而且考虑了相邻帧之间的关系。采用两步训练方案提高系统的鲁棒性。在ASVspoof 2019上的实验表明,与GMM相比,LFCC+GMM-ResNet系统在逻辑访问场景下可以相对降低min-tDCF和EER 76.1%和76.3%,而LFCC+GMM-SENet系统在物理访问场景下可以降低94.4%和95.4%。分数融合后,系统在两种情况下都给出了第二好的结果。摘要:The automatic speaker verification system is sometimes vulnerable to various spoofing attacks. The 2-class Gaussian Mixture Model classifier for genuine and spoofed speech is usually used as the baseline for spoofing detection. However, the GMM classifier does not separately consider the scores of feature frames on each Gaussian component. In addition, the GMM accumulates the scores on all frames independently, and does not consider their correlations. We propose the two-path GMM-ResNet and GMM-SENet models for spoofing detection, whose input is the Gaussian probability features based on two GMMs trained on genuine and spoofed speech respectively. The models consider not only the score distribution on GMM components, but also the relationship between adjacent frames. A two-step training scheme is applied to improve the system robustness. Experiments on the ASVspoof 2019 show that the LFCC+GMM-ResNet system can relatively reduce min-tDCF and EER by 76.1% and 76.3% on logical access scenario compared with the GMM, and the LFCC+GMM-SENet system by 94.4% and 95.4% on physical access scenario. After score fusion, the systems give the second-best results on both scenarios.

【12】 Read, Watch and Scream! Sound Generation from Text and Video
标题: 阅读、观看并尖叫!从文本和视频生成声音
作者:Yujin Jeong,Yunji Kim,Sanghyuk Chun,Jiyoung Lee
备注:Project page: this https URL
链接:点击下载PDF文件
摘要:在强大的扩散模型的帮助下,多模态生成模型已经显示出令人印象深刻的进步。尽管取得了进展,但仅从文本生成声音在确保全面的场景描绘和时间对齐方面带来了挑战。同时,视频到声音的生成限制了为场景内的特定对象优先考虑声音合成的灵活性。为了解决这些挑战,我们提出了一种新的视频和文本到声音的生成方法,称为ReWaS,其中视频作为一个条件控制的文本到音频生成模型。我们的方法估计的结构信息的音频(即,能量)从视频,同时从用户提示接收关键内容线索。我们采用了一个性能良好的文本到声音模型来巩固视频控制,这对于训练具有大量三元组配对(音频-视频-文本)数据的多模态扩散模型更有效。此外,通过分离音频的生成组件,它成为一个更灵活的系统,允许用户根据自己的喜好自由调整能量,周围环境和主要声源。实验结果表明,我们的方法在质量,可控性和训练效率方面表现出优越性。我们的演示可在https: naver-ai.github.io rewas上获得摘要:Multimodal generative models have shown impressive advances with the help of powerful diffusion models. Despite the progress, generating sound solely from text poses challenges in ensuring comprehensive scene depiction and temporal alignment. Meanwhile, video-to-sound generation limits the flexibility to prioritize sound synthesis for specific objects within the scene. To tackle these challenges, we propose a novel video-and-text-to-sound generation method, called ReWaS, where video serves as a conditional control for a text-to-audio generation model. Our method estimates the structural information of audio (namely, energy) from the video while receiving key content cues from a user prompt. We employ a well-performing text-to-sound model to consolidate the video control, which is much more efficient for training multimodal diffusion models with massive triplet-paired (audio-video-text) data. In addition, by separating the generative components of audio, it becomes a more flexible system that allows users to freely adjust the energy, surrounding environment, and primary sound source according to their preferences. Experimental results demonstrate that our method shows superiority in terms of quality, controllability, and training efficiency. Our demo is available at https: naver-ai.github.io rewas

【13】 Research on the Acoustic Emission Source Localization Methodology in Composite Materials based on Artificial Intelligence
标题: 基于人工智能的复合材料声发射源定位方法研究
作者:Jongick Won,Hyuntaik Oh,Jae Sakong
链接:点击下载PDF文件
摘要:本文提出了一种基于人工智能的复合材料声发射源定位方法。选用碳纤维增强塑料作为试件,利用压电器件测量了声发射信号。对测量信号进行小波变换,得到尺度图,作为人工智能模型的训练数据。AESLNet(声发射源定位网络),在这项研究中提出的,由于复合材料的各向异性的卷积层并行构造。声发射源位置坐标的检测是一种回归模型。采用贝叶斯优化方法对网络的超参数进行优化。实验结果表明,该网络可以检测出声发射源的位置,平均误差为3.02mm,分辨率为20mm。摘要:In this study, methodology of acoustic emission source localization in composite materials based on artificial intelligence was presented. Carbon fiber reinforced plastic was selected for specimen, and acoustic emission signal were measured using piezoelectric devices. The measured signal was wavelet-transformed to obtain scalograms, which were used as training data for the artificial intelligence model. AESLNet(acoustic emission source localization network), proposed in this study, was constructed convolutional layers in parallel due to anisotropy of the composited materials. It is regression model to detect the coordinates of acoustic emission source location. Hyper-parameter of network has been optimized by Bayesian optimization. It has been confirmed that network can detect location of acoustic emission source with an average error of 3.02mm and a resolution of 20mm.

【14】 A Reference-free Metric for Language-Queried Audio Source Separation using Contrastive Language-Audio Pretraining
标题: 使用对比语音-音频预训练的语音查询音频源分离的无参考指标
作者:Feiyang Xiao,Jian Guan,Qiaoxi Zhu,Xubo Liu,Wenbo Wang,Shuhan Qi,Kejia Zhang,Jianyuan Sun,Wenwu Wang
备注:Submitted to DCASE 2024 Workshop
链接:点击下载PDF文件
摘要:语音查询音频源分离(LASS)旨在分离由文本查询引导的音频源,其中基于信号失真比(SDR)的度量通常用于客观地测量分离的音频的质量。然而,基于SDR的度量需要参考信号,这在现实世界的场景中通常难以获得。此外,基于SDR的度量,文本查询的内容信息没有被有效地考虑在LASS。本文介绍了一种使用对比语言音频预训练(CLAP)模块,称为CLAPSCore,它测量分离的音频和文本查询之间的语义相似性的参考无评价指标。与SDR不同,所提出的CLAPScore度量基于文本查询的内容信息来评估分离的音频的质量,而不需要参考信号。实验结果表明,CLAPScore指标提供了一个有效的评估的语义相关性的分离音频的文本查询,相比SDR指标,提供了一种替代的LASS系统的性能评估。摘要:Language-queried audio source separation (LASS) aims to separate an audio source guided by a text query, with the signal-to-distortion ratio (SDR)-based metrics being commonly used to objectively measure the quality of the separated audio. However, the SDR-based metrics require a reference signal, which is often difficult to obtain in real-world scenarios. In addition, with the SDR-based metrics, the content information of the text query is not considered effectively in LASS. This paper introduces a reference-free evaluation metric using a contrastive language-audio pretraining (CLAP) module, termed CLAPScore, which measures the semantic similarity between the separated audio and the text query. Unlike SDR, the proposed CLAPScore metric evaluates the quality of the separated audio based on the content information of the text query, without needing a reference signal. Experimental results show that the CLAPScore metric provides an effective evaluation of the semantic relevance of the separated audio to the text query, as compared to the SDR metric, offering an alternative for the performance evaluation of LASS systems.

【15】 All Neural Low-latency Directional Speech Extraction
标题: 全神经低延迟定向语音提取
作者:Ashutosh Pandey,Sanha Lee,Juan Azcarreta,Daniel Wong,Buye Xu
备注:Accepted for publication at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们介绍了一种新的低延迟方向性语音提取的全神经模型。该模型使用来自预定义空间网格的到达方向(DOA)嵌入,将其转换并融合到基于递归神经网络的语音提取模型中。该过程使模型能够有效地从指定的DOA中提取语音。与以前依赖于手工制作的方向特征的方法不同,所提出的模型使用语音增强损失从头开始训练DOA嵌入,使其适用于低延迟场景。此外,它以高帧速率运行,在每个输入帧中接收DOA,这带来了在高度动态的真实世界场景中快速适应变化场景的能力。我们提供了广泛的评估,以证明该模型的有效性,方向性语音提取,DOA失配的鲁棒性,其能力,以快速适应突然变化的DOA。摘要:We introduce a novel all neural model for low-latency directional speech extraction. The model uses direction of arrival (DOA) embeddings from a predefined spatial grid, which are transformed and fused into a recurrent neural network based speech extraction model. This process enables the model to effectively extract speech from a specified DOA. Unlike previous methods that relied on hand-crafted directional features, the proposed model trains DOA embeddings from scratch using speech enhancement loss, making it suitable for low-latency scenarios. Additionally, it operates at a high frame rate, taking in DOA with each input frame, which brings in the capability of quickly adapting to changing scene in highly dynamic real-world scenarios. We provide extensive evaluation to demonstrate the model's efficacy in directional speech extraction, robustness to DOA mismatch, and its capability to quickly adapt to abrupt changes in DOA.


机器翻译,仅供参考