今日论文合集:cs.SD语音5篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Emotion-Anchored Contrastive Learning Framework for Emotion Recognition  in Conversation

标题:会话情绪识别的认知锚定对比学习框架

链接:https://arxiv.org/abs/2403.20289

作者:Fangxu Yu,Junjie Guo,Zhen Wu,Xinyu Dai

备注:Accepted by Findings of NAACL 2024

摘要:会话中的情感识别(ERC)涉及检测会话中每个话语背后的潜在情感。有效地生成表示的话语仍然是一个重大的挑战,在这项任务。最近的研究提出了各种模型来解决这个问题,但他们仍然难以区分类似的情绪,如兴奋和幸福。为了缓解这个问题,我们提出了一个预防锚定的对比学习(EACL)框架,可以生成更可区分的话语表示相似的情绪。为了实现这一目标,我们利用标签编码作为锚来指导话语表征的学习,并设计了一个辅助损失,以确保类似情绪的锚的有效分离。此外,提出了一个额外的适应过程,以适应锚作为有效的分类器,以提高分类性能。在广泛的实验中,我们提出的EACL实现了最先进的情感识别性能,并在相似的情感上表现出优异的性能。我们的代码可在https://github.com/Yu-Fangxu/EACL上获得。

摘要:Emotion Recognition in Conversation (ERC) involves detecting the underlying emotion behind each utterance within a conversation. Effectively generating representations for utterances remains a significant challenge in this task. Recent works propose various models to address this issue, but they still struggle with differentiating similar emotions such as excitement and happiness. To alleviate this problem, We propose an Emotion-Anchored Contrastive Learning (EACL) framework that can generate more distinguishable utterance representations for similar emotions. To achieve this, we utilize label encodings as anchors to guide the learning of utterance representations and design an auxiliary loss to ensure the effective separation of anchors for similar emotions. Moreover, an additional adaptation process is proposed to adapt anchors to serve as effective classifiers to improve classification performance. Across extensive experiments, our proposed EACL achieves state-of-the-art emotion recognition performance and exhibits superior performance on similar emotions. Our code is available at https://github.com/Yu-Fangxu/EACL.


【2】 Voice Signal Processing for Machine Learning. The Case of Speaker  Isolation
标题:机器学习语音信号处理扬声器隔离案例
链接:https://arxiv.org/abs/2403.20202
作者:Radan Ganchev
备注:MSc. thesis. for associated source code, see this https URL
摘要:自动语音助理的广泛使用以及其他最新的技术发展增加了对处理音频信号,特别是人类语音的应用程序的需求。语音识别任务通常使用人工智能和机器学习模型来执行。尽管存在端到端模型,但适当地预处理信号可以大大降低任务的复杂性,并允许使用更简单的ML模型和更少的计算资源来解决问题。然而,从事此类任务的机器学习工程师可能没有信号处理的背景,这是一个完全不同的专业领域。这项工作的目的是提供一个简洁的比较分析傅立叶和小波变换,最常用的信号分解方法的音频处理任务。本文还讨论了评价语音可懂度的三种方法,即尺度不变信失真比(SI-SDR)、语音质量感知评价(PESQ)和短时客观可懂度(STOI)。阐述中的详细程度足以让ML工程师在选择、微调和评估特定ML模型的分解方法时做出明智的决策。该博览会包含相关概念的数学定义,并附有直观的非数学解释,以使文本更容易为工程师在信号处理方面没有深入的专业知识。正式的数学定义和定理的证明是故意省略,以保持文字简洁。
摘要:The widespread use of automated voice assistants along with other recent technological developments have increased the demand for applications that process audio signals and human voice in particular. Voice recognition tasks are typically performed using artificial intelligence and machine learning models. Even though end-to-end models exist, properly pre-processing the signal can greatly reduce the complexity of the task and allow it to be solved with a simpler ML model and fewer computational resources. However, ML engineers who work on such tasks might not have a background in signal processing which is an entirely different area of expertise.  The objective of this work is to provide a concise comparative analysis of Fourier and Wavelet transforms that are most commonly used as signal decomposition methods for audio processing tasks. Metrics for evaluating speech intelligibility are also discussed, namely Scale-Invariant Signal-to-Distortion Ratio (SI-SDR), Perceptual Evaluation of Speech Quality (PESQ), and Short-Time Objective Intelligibility (STOI). The level of detail in the exposition is meant to be sufficient for an ML engineer to make informed decisions when choosing, fine-tuning, and evaluating a decomposition method for a specific ML model. The exposition contains mathematical definitions of the relevant concepts accompanied with intuitive non-mathematical explanations in order to make the text more accessible to engineers without deep expertise in signal processing. Formal mathematical definitions and proofs of theorems are intentionally omitted in order to keep the text concise.

【3】 Sound event localization and classification using WASN in Outdoor  Environment
标题:户外环境声事件定位与分类的应用
链接:https://arxiv.org/abs/2403.20130
作者:Dongzhe Zhang,Jianfeng Chen,Jisheng Bai,Mou Wang
摘要:基于深度学习的声音事件定位和分类是无线声学传感器网络中的一个新兴研究领域。然而,当前用于声音事件定位和分类的方法通常依赖于单个麦克风阵列,使得它们容易受到信号衰减和环境噪声的影响,这限制了它们的监测范围。此外,使用多个麦克风阵列的方法通常仅关注源定位,忽略了声音事件分类方面。在本文中,我们提出了一种基于深度学习的方法,该方法采用多个特征和注意力机制来估计声源的位置和类别。我们引入了Soundmap功能来捕获多个频段的空间信息。我们还使用Gammatone滤波器来生成更适合户外环境的声学特征。此外,我们整合注意力机制来学习声学特征内的通道关系和时间依赖性。为了评估我们提出的方法,我们进行了实验,使用不同的噪声水平和监测区域的大小,以及不同的阵列和源位置的模拟数据集。实验结果表明,我们提出的方法优于国家的最先进的方法在两个声音事件分类和声源定位任务。并进一步分析了观测误差产生的原因。
摘要:Deep learning-based sound event localization and classification is an emerging research area within wireless acoustic sensor networks. However, current methods for sound event localization and classification typically rely on a single microphone array, making them susceptible to signal attenuation and environmental noise, which limits their monitoring range. Moreover, methods using multiple microphone arrays often focus solely on source localization, neglecting the aspect of sound event classification. In this paper, we propose a deep learning-based method that employs multiple features and attention mechanisms to estimate the location and class of sound source. We introduce a Soundmap feature to capture spatial information across multiple frequency bands. We also use the Gammatone filter to generate acoustic features more suitable for outdoor environments. Furthermore, we integrate attention mechanisms to learn channel-wise relationships and temporal dependencies within the acoustic features. To evaluate our proposed method, we conduct experiments using simulated datasets with different levels of noise and size of monitoring areas, as well as different arrays and source positions. The experimental results demonstrate the superiority of our proposed method over state-of-the-art methods in both sound event classification and sound source localization tasks. And we provide further analysis to explain the reasons for the observed errors.

【4】 Creating Aesthetic Sonifications on the Web with SIREN
标题:用SIREN在网络上创造审美声音
链接:https://arxiv.org/abs/2403.19763
作者:Tristan Peng,Hongchan Choi,Jonathan Berger
备注:7 pages, 1 figure, 5 listings, submitted to the Web Audio Conference 2024
摘要:SIREN是一个灵活的,可扩展的,可定制的基于Web的通用界面,用于听觉数据显示(发声)。设计为用于发声的数字音频工作站,使用Web Audio API以JavaScript编写的合成器便于直观地将数据映射到听觉参数,用于各种用途。本文探讨了SIREN支持的声音合成技术的广度,并详细介绍了SIREN合成器模块的结构和定义。本文提出了进一步的发展,将增加SIREN的效用。
摘要:SIREN is a flexible, extensible, and customizable web-based general-purpose interface for auditory data display (sonification). Designed as a digital audio workstation for sonification, synthesizers written in JavaScript using the Web Audio API facilitate intuitive mapping of data to auditory parameters for a wide range of purposes.  This paper explores the breadth of sound synthesis techniques supported by SIREN, and details the structure and definition of a SIREN synthesizer module. The paper proposes further development that will increase SIREN's utility.


【5】 Exploring Pathological Speech Quality Assessment with ASR-Powered  Wav2Vec2 in Data-Scarce Context
标题:基于ASR的Wav2Vec2在数据稀缺环境下的病理性语音质量评估中的应用
链接:https://arxiv.org/abs/2403.20184
作者:Tuan Nguyen,Corinne Fredouille,Alain Ghio,Mathieu Balaguer,Virginie Woisard
备注:Accepted at LREC-COLING 2024
摘要:自动语音质量评估作为传统的感知临床评估的替代或支持,引起了越来越多的关注。然而,到目前为止,大多数研究只在简单的任务上获得了良好的结果,如二进制分类,这主要是由于数据稀缺。为了应对这一挑战,目前的工作倾向于将患者的音频文件分割成许多样本来增强数据集。然而,这种方法具有局限性,因为它间接地将整体音频分数与各个片段相关联。本文介绍了一种新的方法,尽管数据稀缺,系统还是在音频级别而不是分段学习。本文提出将预训练的Wav 2 Vec 2架构用于SSL和ASR作为语音评估中的特征提取器。在HNC数据集上进行,与其他方法相比,我们的ASR驱动方法建立了一个新的基线,仅使用95个训练样本,分别获得可懂度和严重程度评分的平均$MSE=0.73$和$MSE=1.15$。结果表明,基于小波2矢量2模型的语音识别效果最好,表明语音识别与语音质量评价之间存在很强的相关性。我们还测量了它的能力可变段持续时间和语音内容,探索影响其决策的因素。
摘要:Automatic speech quality assessment has raised more attention as an alternative or support to traditional perceptual clinical evaluation. However, most research so far only gains good results on simple tasks such as binary classification, largely due to data scarcity. To deal with this challenge, current works tend to segment patients' audio files into many samples to augment the datasets. Nevertheless, this approach has limitations, as it indirectly relates overall audio scores to individual segments. This paper introduces a novel approach where the system learns at the audio level instead of segments despite data scarcity. This paper proposes to use the pre-trained Wav2Vec2 architecture for both SSL, and ASR as feature extractor in speech assessment. Carried out on the HNC dataset, our ASR-driven approach established a new baseline compared with other approaches, obtaining average $MSE=0.73$ and $MSE=1.15$ for the prediction of intelligibility and severity scores respectively, using only 95 training samples. It shows that the ASR based Wav2Vec2 model brings the best results and may indicate a strong correlation between ASR and speech quality assessment. We also measure its ability on variable segment durations and speech content, exploring factors influencing its decision.


eess.AS音频处理
【1】 Exploring Pathological Speech Quality Assessment with ASR-Powered  Wav2Vec2 in Data-Scarce Context
标题:基于ASR的Wav2Vec2在数据稀缺环境下的病理性语音质量评估中的应用
链接:https://arxiv.org/abs/2403.20184
作者:Tuan Nguyen,Corinne Fredouille,Alain Ghio,Mathieu Balaguer,Virginie Woisard
备注:Accepted at LREC-COLING 2024
摘要:自动语音质量评估作为传统的感知临床评估的替代或支持,引起了越来越多的关注。然而,到目前为止,大多数研究只在简单的任务上获得了良好的结果,如二进制分类,这主要是由于数据稀缺。为了应对这一挑战,目前的工作倾向于将患者的音频文件分割成许多样本来增强数据集。然而,这种方法具有局限性,因为它间接地将整体音频分数与各个片段相关联。本文介绍了一种新的方法,尽管数据稀缺,系统还是在音频级别而不是分段学习。本文提出将预训练的Wav 2 Vec 2架构用于SSL和ASR作为语音评估中的特征提取器。在HNC数据集上进行,与其他方法相比,我们的ASR驱动方法建立了一个新的基线,仅使用95个训练样本,分别获得可懂度和严重程度评分的平均$MSE=0.73$和$MSE=1.15$。结果表明,基于小波2矢量2模型的语音识别效果最好,表明语音识别与语音质量评价之间存在很强的相关性。我们还测量了它的能力可变段持续时间和语音内容,探索影响其决策的因素。
摘要:Automatic speech quality assessment has raised more attention as an alternative or support to traditional perceptual clinical evaluation. However, most research so far only gains good results on simple tasks such as binary classification, largely due to data scarcity. To deal with this challenge, current works tend to segment patients' audio files into many samples to augment the datasets. Nevertheless, this approach has limitations, as it indirectly relates overall audio scores to individual segments. This paper introduces a novel approach where the system learns at the audio level instead of segments despite data scarcity. This paper proposes to use the pre-trained Wav2Vec2 architecture for both SSL, and ASR as feature extractor in speech assessment. Carried out on the HNC dataset, our ASR-driven approach established a new baseline compared with other approaches, obtaining average $MSE=0.73$ and $MSE=1.15$ for the prediction of intelligibility and severity scores respectively, using only 95 training samples. It shows that the ASR based Wav2Vec2 model brings the best results and may indicate a strong correlation between ASR and speech quality assessment. We also measure its ability on variable segment durations and speech content, exploring factors influencing its decision.


【2】 Non-Exponential Reverberation Modeling Using Dark Velvet Noise
标题:基于暗天鹅绒噪声的非指数混响建模
链接:https://arxiv.org/abs/2403.20090
作者:Jon Fagerström,Sebastian J. Schelcht,Vesa Välimäki
备注:Accepted for publication in the Journal of Audio Engineering Society
摘要:以前的研究后期混响建模主要集中在指数衰减的房间脉冲响应,而精确建模非指数混响的方法仍然具有挑战性。本文扩展了先前提出的基本暗天鹅绒噪声混响算法,并提出了一个参数化方案建模后期混响与任意时间能量衰减。天鹅绒噪声序列中的每个脉冲被路由到基于加权概率从一组滤波器中选择的单个字典滤波器。概率控制后期混响模型的频谱演变,并通过非负最小二乘优化来优化以拟合目标脉冲响应。以这种方式,目标后期混响脉冲响应的频率相关的能量衰减可以分别与4%和8%的平均和最大T60误差拟合,需要比先前提出的滤波天鹅绒噪声算法少约50%的着色滤波器。此外,扩展的暗丝绒噪声混响算法允许建模的脉冲响应被选通,频率相关的混响时间被修改,并且模型的频谱演化和宽带衰减被解耦。所提出的方法适用于各种声学环境的参数后期混响合成,特别是表现出非指数能量衰减的空间,激励其在音乐音频和虚拟现实中的使用。
摘要:Previous research on late-reverberation modeling has mainly focused on exponentially decaying room impulse responses, whereas methods for accurately modeling non-exponential reverberation remain challenging. This paper extends the previously proposed basic dark-velvet-noise reverberation algorithm and proposes a parametrization scheme for modeling late reverberation with arbitrary temporal energy decay. Each pulse in the velvet-noise sequence is routed to a single dictionary filter that is selected from a set of filters based on weighted probabilities. The probabilities control the spectral evolution of the late-reverberation model and are optimized to fit a target impulse response via non-negative least-squares optimization. In this way, the frequency-dependent energy decay of a target late-reverberation impulse response can be fitted with mean and maximum T60 errors of 4% and 8%, respectively, requiring about 50% less coloration filters than a previously proposed filtered velvet-noise algorithm. Furthermore, the extended dark-velvet-noise reverberation algorithm allows the modeled impulse response to be gated, the frequency-dependent reverberation time to be modified, and the model's spectral evolution and broadband decay to be decoupled. The proposed method is suitable for the parametric late-reverberation synthesis of various acoustic environments, especially spaces that exhibit a non-exponential energy decay, motivating its use in musical audio and virtual reality.

【3】 3D-Speaker-Toolkit: An Open Source Toolkit for Multi-modal Speaker  Verification and Diarization
标题:3D—Speaker—Toolkit:一个开源的多模态说话人验证和拨号工具包
链接:https://arxiv.org/abs/2403.19971
作者:Yafeng Chen,Siqi Zheng,Hui Wang,Luyao Cheng,Tinglong Zhu,Changhe Song,Rongjie Huang,Ziyang Ma,Qian Chen,Shiliang Zhang,Xihao Li
摘要:本文介绍了一个开源的多模态说话人确认和日志化工具包3D-Speaker-Toolkit。它是专为学术研究人员和工业从业人员的需求。3D-Speaker-Toolkit巧妙地利用了声学、语义和视觉数据的综合优势,无缝融合了这些模式,提供了强大的说话人识别功能。声学模块从声学特征中提取说话人嵌入,采用全监督和自监督学习方法。语义模块利用先进的语言模型来理解口语的内容和上下文,从而增强系统通过语言模式区分说话者的能力。最后,视觉模块应用图像处理技术来仔细检查面部特征,这增强了多说话人环境中说话人日记的精度。总的来说,这些模块使3D扬声器工具包在执行扬声器相关任务时达到更高的准确性和可靠性水平,建立了多模态扬声器分析的新基准。3D-Speaker项目还包括一些开源的最先进的模型和一个包含超过10,000个扬声器的大型数据集。该工具包可在https://github.com/alibaba-damo-academy/3D-Speaker上公开查阅。
摘要:This paper introduces 3D-Speaker-Toolkit, an open source toolkit for multi-modal speaker verification and diarization. It is designed for the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the combined strengths of acoustic, semantic, and visual data, seamlessly fusing these modalities to offer robust speaker recognition capabilities. The acoustic module extracts speaker embeddings from acoustic features, employing both fully-supervised and self-supervised learning approaches. The semantic module leverages advanced language models to apprehend the substance and context of spoken language, thereby augmenting the system's proficiency in distinguishing speakers through linguistic patterns. Finally, the visual module applies image processing technologies to scrutinize facial features, which bolsters the precision of speaker diarization in multi-speaker environments. Collectively, these modules empower the 3D-Speaker-Toolkit to attain elevated levels of accuracy and dependability in executing speaker-related tasks, establishing a new benchmark in multi-modal speaker analysis. The 3D-Speaker project also includes a handful of open-sourced state-of-the-art models and a large dataset containing over 10,000 speakers. The toolkit is publicly available at https://github.com/alibaba-damo-academy/3D-Speaker.

【4】 Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of  Large Speech Models
标题:高效多任务适应大型语音模型的分层递归适配器
链接:https://arxiv.org/abs/2403.19709
作者:Tsendsuren Munkhdalai,Youzheng Chen,Khe Chai Sim,Fadi Biadsy,Tara Sainath,Pedro Moreno Mengibar
备注:5 pages, 3 figures, 5 tables
摘要:参数有效的自适应方法已经成为训练大型预训练模型用于下游任务的关键机制。然而,当要适应的下游任务的数量很大时,它们的每任务参数开销被认为仍然很高。我们介绍了一个适配器模块,在大规模的多任务适配场景中有更好的效率。我们的适配器在如何分配适配器参数方面是分层的。该适配器由单个共享控制器网络和多个任务级适配器头组成,以减少每个任务的参数开销,而不会对下游任务造成性能退化。适配器也是循环的,因此整个适配器参数可以在预训练模型的不同层中重复使用。我们的分层递归适配器(HRA)优于以前的适配器为基础的方法,以及完整的模型微调基线在单任务和多任务的自适应设置时,自动语音识别任务进行评估。
摘要:Parameter efficient adaptation methods have become a key mechanism to train large pre-trained models for downstream tasks. However, their per-task parameter overhead is considered still high when the number of downstream tasks to adapt for is large. We introduce an adapter module that has a better efficiency in large scale multi-task adaptation scenario. Our adapter is hierarchical in terms of how the adapter parameters are allocated. The adapter consists of a single shared controller network and multiple task-level adapter heads to reduce the per-task parameter overhead without performance regression on downstream tasks. The adapter is also recurrent so the entire adapter parameters are reused across different layers of the pre-trained model. Our Hierarchical Recurrent Adapter (HRA) outperforms the previous adapter-based approaches as well as full model fine-tuning baseline in both single and multi-task adaptation settings when evaluated on automatic speech recognition tasks.


【5】 Emotion-Anchored Contrastive Learning Framework for Emotion Recognition  in Conversation
标题:会话情绪识别的认知锚定对比学习框架
链接:https://arxiv.org/abs/2403.20289
作者:Fangxu Yu,Junjie Guo,Zhen Wu,Xinyu Dai
备注:Accepted by Findings of NAACL 2024
摘要:会话中的情感识别(ERC)涉及检测会话中每个话语背后的潜在情感。有效地生成表示的话语仍然是一个重大的挑战,在这项任务。最近的研究提出了各种模型来解决这个问题,但他们仍然难以区分类似的情绪,如兴奋和幸福。为了缓解这个问题,我们提出了一个预防锚定的对比学习(EACL)框架,可以生成更可区分的话语表示相似的情绪。为了实现这一目标,我们利用标签编码作为锚来指导话语表征的学习,并设计了一个辅助损失,以确保类似情绪的锚的有效分离。此外,提出了一个额外的适应过程,以适应锚作为有效的分类器,以提高分类性能。在广泛的实验中,我们提出的EACL实现了最先进的情感识别性能,并在相似的情感上表现出优异的性能。我们的代码可在https://github.com/Yu-Fangxu/EACL上获得。
摘要:Emotion Recognition in Conversation (ERC) involves detecting the underlying emotion behind each utterance within a conversation. Effectively generating representations for utterances remains a significant challenge in this task. Recent works propose various models to address this issue, but they still struggle with differentiating similar emotions such as excitement and happiness. To alleviate this problem, We propose an Emotion-Anchored Contrastive Learning (EACL) framework that can generate more distinguishable utterance representations for similar emotions. To achieve this, we utilize label encodings as anchors to guide the learning of utterance representations and design an auxiliary loss to ensure the effective separation of anchors for similar emotions. Moreover, an additional adaptation process is proposed to adapt anchors to serve as effective classifiers to improve classification performance. Across extensive experiments, our proposed EACL achieves state-of-the-art emotion recognition performance and exhibits superior performance on similar emotions. Our code is available at https://github.com/Yu-Fangxu/EACL.


【6】 Voice Signal Processing for Machine Learning. The Case of Speaker  Isolation
标题:机器学习语音信号处理扬声器隔离案例
链接:https://arxiv.org/abs/2403.20202
作者:Radan Ganchev
备注:MSc. thesis. for associated source code, see this https URL
摘要:自动语音助理的广泛使用以及其他最新的技术发展增加了对处理音频信号,特别是人类语音的应用程序的需求。语音识别任务通常使用人工智能和机器学习模型来执行。尽管存在端到端模型,但适当地预处理信号可以大大降低任务的复杂性,并允许使用更简单的ML模型和更少的计算资源来解决问题。然而,从事此类任务的机器学习工程师可能没有信号处理的背景,这是一个完全不同的专业领域。这项工作的目的是提供一个简洁的比较分析傅立叶和小波变换,最常用的信号分解方法的音频处理任务。本文还讨论了评价语音可懂度的三种方法,即尺度不变信失真比(SI-SDR)、语音质量感知评价(PESQ)和短时客观可懂度(STOI)。阐述中的详细程度足以让ML工程师在选择、微调和评估特定ML模型的分解方法时做出明智的决策。该博览会包含相关概念的数学定义,并附有直观的非数学解释,以使文本更容易为工程师在信号处理方面没有深入的专业知识。正式的数学定义和定理的证明是故意省略,以保持文字简洁。
摘要:The widespread use of automated voice assistants along with other recent technological developments have increased the demand for applications that process audio signals and human voice in particular. Voice recognition tasks are typically performed using artificial intelligence and machine learning models. Even though end-to-end models exist, properly pre-processing the signal can greatly reduce the complexity of the task and allow it to be solved with a simpler ML model and fewer computational resources. However, ML engineers who work on such tasks might not have a background in signal processing which is an entirely different area of expertise.  The objective of this work is to provide a concise comparative analysis of Fourier and Wavelet transforms that are most commonly used as signal decomposition methods for audio processing tasks. Metrics for evaluating speech intelligibility are also discussed, namely Scale-Invariant Signal-to-Distortion Ratio (SI-SDR), Perceptual Evaluation of Speech Quality (PESQ), and Short-Time Objective Intelligibility (STOI). The level of detail in the exposition is meant to be sufficient for an ML engineer to make informed decisions when choosing, fine-tuning, and evaluating a decomposition method for a specific ML model. The exposition contains mathematical definitions of the relevant concepts accompanied with intuitive non-mathematical explanations in order to make the text more accessible to engineers without deep expertise in signal processing. Formal mathematical definitions and proofs of theorems are intentionally omitted in order to keep the text concise.


【7】 Sound event localization and classification using WASN in Outdoor  Environment
标题:户外环境声事件定位与分类的应用
链接:https://arxiv.org/abs/2403.20130
作者:Dongzhe Zhang,Jianfeng Chen,Jisheng Bai,Mou Wang
摘要:基于深度学习的声音事件定位和分类是无线声学传感器网络中的一个新兴研究领域。然而,当前用于声音事件定位和分类的方法通常依赖于单个麦克风阵列,使得它们容易受到信号衰减和环境噪声的影响,这限制了它们的监测范围。此外,使用多个麦克风阵列的方法通常仅关注源定位,忽略了声音事件分类方面。在本文中,我们提出了一种基于深度学习的方法,该方法采用多个特征和注意力机制来估计声源的位置和类别。我们引入了Soundmap功能来捕获多个频段的空间信息。我们还使用Gammatone滤波器来生成更适合户外环境的声学特征。此外,我们整合注意力机制来学习声学特征内的通道关系和时间依赖性。为了评估我们提出的方法,我们进行了实验,使用不同的噪声水平和监测区域的大小,以及不同的阵列和源位置的模拟数据集。实验结果表明,我们提出的方法优于国家的最先进的方法在两个声音事件分类和声源定位任务。并进一步分析了观测误差产生的原因。
摘要:Deep learning-based sound event localization and classification is an emerging research area within wireless acoustic sensor networks. However, current methods for sound event localization and classification typically rely on a single microphone array, making them susceptible to signal attenuation and environmental noise, which limits their monitoring range. Moreover, methods using multiple microphone arrays often focus solely on source localization, neglecting the aspect of sound event classification. In this paper, we propose a deep learning-based method that employs multiple features and attention mechanisms to estimate the location and class of sound source. We introduce a Soundmap feature to capture spatial information across multiple frequency bands. We also use the Gammatone filter to generate acoustic features more suitable for outdoor environments. Furthermore, we integrate attention mechanisms to learn channel-wise relationships and temporal dependencies within the acoustic features. To evaluate our proposed method, we conduct experiments using simulated datasets with different levels of noise and size of monitoring areas, as well as different arrays and source positions. The experimental results demonstrate the superiority of our proposed method over state-of-the-art methods in both sound event classification and sound source localization tasks. And we provide further analysis to explain the reasons for the observed errors.

【8】 Creating Aesthetic Sonifications on the Web with SIREN
标题:用SIREN在网络上创造审美声音
链接:https://arxiv.org/abs/2403.19763
作者:Tristan Peng,Hongchan Choi,Jonathan Berger
备注:7 pages, 1 figure, 5 listings, submitted to the Web Audio Conference 2024
摘要:SIREN是一个灵活的,可扩展的,可定制的基于Web的通用界面,用于听觉数据显示(发声)。设计为用于发声的数字音频工作站,使用Web Audio API以JavaScript编写的合成器便于直观地将数据映射到听觉参数,用于各种用途。本文探讨了SIREN支持的声音合成技术的广度,并详细介绍了SIREN合成器模块的结构和定义。本文提出了进一步的发展,将增加SIREN的效用。
摘要:SIREN is a flexible, extensible, and customizable web-based general-purpose interface for auditory data display (sonification). Designed as a digital audio workstation for sonification, synthesizers written in JavaScript using the Web Audio API facilitate intuitive mapping of data to auditory parameters for a wide range of purposes.  This paper explores the breadth of sound synthesis techniques supported by SIREN, and details the structure and definition of a SIREN synthesizer module. The paper proposes further development that will increase SIREN's utility.

机器翻译由腾讯交互翻译提供,仅供参考