今日论文合集:cs.SD语音8篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
标题: PeriodWave:用于高保真波生成的多周期流匹配
作者:Sang-Hoon Lee,Ha-Yeong Choi,Seong-Whan Lee
备注:24 pages, 16 tables, 4 figures
链接:点击下载PDF文件
摘要:最近,通用波形生成任务已被调查条件下的各种分布情况。虽然基于GAN的方法在快速波形生成方面表现出了强大的实力,但它们容易受到训练推理不匹配的情况,例如两阶段文本到语音。与此同时,基于扩散的模型在其他领域显示出强大的生成性能;然而,由于波形生成任务的推理速度较慢,它们一直处于聚光灯下。最重要的是,没有一种生成器架构可以明确地理清高分辨率波形信号的自然周期特征。在本文中,我们提出了PeriodWave,一种新的通用波形生成模型。首先,我们介绍了一个周期感知的流量匹配估计器,可以捕捉波形信号的周期特征时,估计的向量场。此外,我们利用一个多周期估计,避免重叠,以捕捉不同的周期特征的波形信号。虽然增加周期数可以显著提高性能,但这需要更多的计算成本。为了减少这个问题,我们还提出了一个单一的周期条件的通用估计,可以前馈并行周期分批推理。此外,本文还利用离散小波变换对波形信号的频率信息进行离散化去纠缠以实现高频建模,并引入FreeU算法对波形生成过程中的高频噪声进行去噪。实验结果表明,该模型在Mel-声谱图重建和文本到语音转换任务中的性能优于以前的模型。所有源代码都可以在 url{https: github.com sh-lee-prml PeriodWave}上找到。摘要:Recently, universal waveform generation tasks have been investigated conditioned on various out-of-distribution scenarios. Although GAN-based methods have shown their strength in fast waveform generation, they are vulnerable to train-inference mismatch scenarios such as two-stage text-to-speech. Meanwhile, diffusion-based models have shown their powerful generative performance in other domains; however, they stay out of the limelight due to slow inference speed in waveform generation tasks. Above all, there is no generator architecture that can explicitly disentangle the natural periodic features of high-resolution waveform signals. In this paper, we propose PeriodWave, a novel universal waveform generation model. First, we introduce a period-aware flow matching estimator that can capture the periodic features of the waveform signal when estimating the vector fields. Additionally, we utilize a multi-period estimator that avoids overlaps to capture different periodic features of waveform signals. Although increasing the number of periods can improve the performance significantly, this requires more computational costs. To reduce this issue, we also propose a single period-conditional universal estimator that can feed-forward parallel by period-wise batch inference. Additionally, we utilize discrete wavelet transform to losslessly disentangle the frequency information of waveform signals for high-frequency modeling, and introduce FreeU to reduce the high-frequency noise for waveform generation. The experimental results demonstrated that our model outperforms the previous models both in Mel-spectrogram reconstruction and text-to-speech tasks. All source code will be available at url{https: github.com sh-lee-prml PeriodWave}.

【2】 Optimising MFCC parameters for the automatic detection of respiratory diseases
标题: 优化MFCC参数以自动检测呼吸道疾病
作者:Yuyang Yan,Sami O. Simons,Loes van Bemmel,Lauren Reinders,Frits M. E. Franssen,Visara Urovi
链接:点击下载PDF文件
摘要:源于呼吸道的语音信号被用作诊断和评估呼吸系统疾病的有价值的声学生物标志物。在所采用的声学特征中,Mel频率倒谱系数(MFCC)被广泛用于自动分析,MFCC提取通常依赖于默认参数。然而,没有全面的研究系统地研究了MFCC提取参数对呼吸系统疾病诊断的影响。在本研究中,我们通过检查关键参数(即系数数量、帧长度和帧间跳长)对呼吸状况检查的影响来解决这一差距。我们的调查使用了四个数据集:Cambridge COVID-19 Sound数据库、Coswara数据集、Saarbrucken Voice Disorders(SVD)数据库和TACTICAS数据集。支持向量机(SVM)被用作分类器,因为它被广泛采用和有效。我们的研究结果表明,MFCC的准确性随着跳长的增加而降低,并且观察到的最佳系数数量约为30。MFCC的性能随数据集的帧长度而变化:对于COVID-19数据集(Cambridge COVID-19 Sound数据库和Coswara数据集),性能随着帧长度的增加而下降,而对于SVD数据集,性能随着帧长度的增加而提高(从50 ms到500 ms)。此外,我们研究了这些参数的优化组合,并观察到大幅提高的准确性。与最差组合相比,SVM模型实现了81.1%,80.6%和71.7%的准确率,对于Cambridge COVID-19 Sound数据库,Coswara数据集和SVD数据集分别提高了19.6%,16.10%和14.90%。摘要:Voice signals originating from the respiratory tract are utilized as valuable acoustic biomarkers for the diagnosis and assessment of respiratory diseases. Among the employed acoustic features, Mel Frequency Cepstral Coefficients (MFCC) is widely used for automatic analysis, with MFCC extraction commonly relying on default parameters. However, no comprehensive study has systematically investigated the impact of MFCC extraction parameters on respiratory disease diagnosis. In this study, we address this gap by examining the effects of key parameters, namely the number of coefficients, frame length, and hop length between frames, on respiratory condition examination. Our investigation uses four datasets: the Cambridge COVID-19 Sound database, the Coswara dataset, the Saarbrucken Voice Disorders (SVD) database, and a TACTICAS dataset. The Support Vector Machine (SVM) is employed as the classifier, given its widespread adoption and efficacy. Our findings indicate that the accuracy of MFCC decreases as hop length increases, and the optimal number of coefficients is observed to be approximately 30. The performance of MFCC varies with frame length across the datasets: for the COVID-19 datasets (Cambridge COVID-19 Sound database and Coswara dataset), performance declines with longer frame lengths, while for the SVD dataset, performance improves with increasing frame length (from 50 ms to 500 ms). Furthermore, we investigate the optimized combination of these parameters and observe substantial enhancements in accuracy. Compared to the worst combination, the SVM model achieves an accuracy of 81.1%, 80.6%, and 71.7%, with improvements of 19.6%, 16.10%, and 14.90% for the Cambridge COVID-19 Sound database, the Coswara dataset, and the SVD dataset respectively.

【3】 DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement
标题: DPSNN:用于低延迟流语音增强的尖峰神经网络
作者:Tao Sun,Sander Bohté
链接:点击下载PDF文件
摘要:语音增强(SE)改善了嘈杂环境中的通信,影响了自动语音识别,助听器和电信等领域。由于这些域通常是功率受限和基于事件的,同时需要低延迟,因此尖峰神经网络(SNN)形式的神经形态算法具有很大的潜力。然而,当前有效的SNN解决方案需要上下文采样窗口,该上下文采样窗口施加了相当大的延迟,通常在32 ms左右,对于许多应用来说太长。受经典神经网络中双路径脉冲神经网络(DPSNN)的启发,提出了一种两阶段时域流SNN框架--双路径脉冲神经网络(DPSNN)。在DPSNN中,第一阶段使用尖峰卷积神经网络(SCNN)来捕获全局上下文信息,而第二阶段使用尖峰循环神经网络(SRNN)来关注频率相关特征。此外,正则化抑制激活,以进一步提高我们的DPSNN的能量效率。通过对VCTK和英特尔DNS数据集的评估,我们证明了我们的方法实现了助听器等应用所需的极低延迟(约5 ms),同时表现出出色的信噪比(SNR)、感知质量和能效。摘要:Speech enhancement (SE) improves communication in noisy environments, affecting areas such as automatic speech recognition, hearing aids, and telecommunications. With these domains typically being power-constrained and event-based while requiring low latency, neuromorphic algorithms in the form of spiking neural networks (SNNs) have great potential. Yet, current effective SNN solutions require a contextual sampling window imposing substantial latency, typically around 32ms, too long for many applications. Inspired by Dual-Path Spiking Neural Networks (DPSNNs) in classical neural networks, we develop a two-phase time-domain streaming SNN framework -- the Dual-Path Spiking Neural Network (DPSNN). In the DPSNN, the first phase uses Spiking Convolutional Neural Networks (SCNNs) to capture global contextual information, while the second phase uses Spiking Recurrent Neural Networks (SRNNs) to focus on frequency-related features. In addition, the regularizer suppresses activation to further enhance energy efficiency of our DPSNNs. Evaluating on the VCTK and Intel DNS Datasets, we demonstrate that our approach achieves the very low latency (approximately 5ms) required for applications like hearing aids, while demonstrating excellent signal-to-noise ratio (SNR), perceptual quality, and energy efficiency.

【4】 Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
标题: 语音与文字记录:语音总结中的人类注释者重要吗?
作者:Roshan Sharma,Suwon Shon,Mark Lindsey,Hira Dhamyal,Rita Singh,Bhiksha Raj
备注:Accepted to ACL 2024 Main Conference
链接:点击下载PDF文件
摘要:用于抽象语音摘要的参考摘要需要人工注释,这可以通过收听音频记录或通过阅读记录的文本转录来执行。在本文中,我们研究的基础上,注释者听录音的摘要是否不同于注释者阅读成绩单。使用现有的基于人工评估的内在评估,自动度量,基于LLM的评估和基于检索的无参考方法。我们发现,摘要确实是不同的源模态的基础上,和基于语音的摘要是更符合事实的一致性和信息选择性比基于成绩单的摘要。同时,基于成绩单的摘要受到源中识别错误的影响,专家撰写的摘要信息量更大,更可靠。我们将所有收集的数据和分析代码公开(https: github.com cmu-mlsp interview_humanssum),以方便复制我们的工作并推进该领域的研究。摘要:Reference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording. In this paper, we examine whether summaries based on annotators listening to the recordings differ from those based on annotators reading transcripts. Using existing intrinsic evaluation based on human evaluation, automatic metrics, LLM-based evaluation, and a retrieval-based reference-free method. We find that summaries are indeed different based on the source modality, and that speech-based summaries are more factually consistent and information-selective than transcript-based summaries. Meanwhile, transcript-based summaries are impacted by recognition errors in the source, and expert-written summaries are more informative and reliable. We make all the collected data and analysis code public(https: github.com cmu-mlsp interview_humanssum) to facilitate the reproduction of our work and advance research in this area.

【5】 Play Me Something Icy: Practical Challenges, Explainability and the Semantic Gap in Generative AI Music
标题: 为我播放一些冰冷的东西:生成性人工智能音乐中的实践挑战、解释性和语义差距
作者:Jesse Allison,Drew Farrar,Treya Nash,Carlos Román,Morgan Weeks,Fiona Xue Ju
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:这张图片旨在批判性地考虑可解释AI背景下文本到音频和文本到音乐生成工具的性质。作为一群实验音乐家和研究人员,我们对这些工具的创造潜力充满热情,并试图从即时创作,控制,可用性,可理解性,人工智能过程的可解释性以及结果的整体美学效果的角度来理解和评估它们。我们已经确定,这些工具没有明确解决的挑战之一是使用基于文本的工具来描述像音乐这样抽象的东西时固有的语义差距。其他差距包括可解释性与可用性,以及用户控制和输入与人类创造过程。这张图片的目的是提出一些问题供讨论,并就我们希望在生成式人工智能音乐工具中看到的改进提出一些一般性建议。摘要:This pictorial aims to critically consider the nature of text-to-audio and text-to-music generative tools in the context of explainable AI. As a group of experimental musicians and researchers, we are enthusiastic about the creative potential of these tools and have sought to understand and evaluate them from perspectives of prompt creation, control, usability, understandability, explainability of the AI process, and overall aesthetic effectiveness of the results. One of the challenges we have identified that is not explicitly addressed by these tools is the inherent semantic gap in using text-based tools to describe something as abstract as music. Other gaps include explainability vs. useability, and user control and input vs. the human creative process. The aim of this pictorial is to raise questions for discussion and make a few general suggestions on the kinds of improvements we would like to see in generative AI music tools.

【6】 A New Dataset, Notation Software, and Representation for Computational Schenkerian Analysis
标题: 计算申...的新数据集、标记软件和表示
作者:Stephen Ni-Hahn,Weihan Xu,Jerry Yin,Rico Zhu,Simon Mak,Yue Jiang,Cynthia Rudin
链接:点击下载PDF文件
摘要:申...法(Schenkerian Analysis,SchA)是一种独特的音乐分析方法,结合旋律、和声、对位和曲式等元素来描述支持音乐作品的层次结构。然而,尽管它强大的分析工具和潜力,以提高音乐的理解和生成,SchA很少被计算机音乐社区使用。这在很大程度上是由于缺乏计算机可读格式的高质量数据。有了更大的申克数据语料库,就有可能为机器学习模型注入对音乐结构更深入的理解,从而产生更“人性化”的结果。为了鼓励进一步研究申...及其对音乐信息学和生成的潜在好处,本文提出了三个主要贡献:1)一个新的和不断增长的SchAs数据集,是迄今为止人类和计算机可读格式中最大的数据集( 140个摘录),2)用于可视化和收集SchA数据的新颖软件,和3)新颖的,灵活地将SchA表示为异质边图数据结构。摘要:Schenkerian Analysis (SchA) is a uniquely expressive method of music analysis, combining elements of melody, harmony, counterpoint, and form to describe the hierarchical structure supporting a work of music. However, despite its powerful analytical utility and potential to improve music understanding and generation, SchA has rarely been utilized by the computer music community. This is in large part due to the paucity of available high-quality data in a computer-readable format. With a larger corpus of Schenkerian data, it may be possible to infuse machine learning models with a deeper understanding of musical structure, thus leading to more "human" results. To encourage further research in Schenkerian analysis and its potential benefits for music informatics and generation, this paper presents three main contributions: 1) a new and growing dataset of SchAs, the largest in human- and computer-readable formats to date ( 140 excerpts), 2) a novel software for visualization and collection of SchA data, and 3) a novel, flexible representation of SchA as a heterogeneous-edge graph data structure.

【7】 A Theory-Based Explainable Deep Learning Architecture for Music Emotion
标题: 基于理论的可解释的音乐情感深度学习架构
作者:Hortense Fong,Vineet Kumar,K. Sudhir
链接:点击下载PDF文件
摘要:本文开发了一种基于理论的、可解释的深度学习卷积神经网络(CNN)分类器,用于预测对音乐的时变情绪反应。我们设计了新颖的CNN滤波器,该滤波器利用声学物理学中已知的影响音乐特征感知的频率谐波结构。我们基于理论的模型更加简约,但提供了与非理论深度学习模型相当的预测性能,同时比使用手工制作的功能的模型表现更好。我们的模型可以补充手工制作的功能,但性能的改善是微不足道的。重要的是,放置在CNN滤波器上的基于谐波的结构为模型如何预测情绪反应(效价和唤醒)提供了更好的解释性,因为情绪与和谐密切相关-由谐波对齐定义的感知特征。最后,我们说明了我们的模型的实用程序涉及数字广告。受YouTube中置广告的启发,我们进行了一项实验室实验,在视频中的不同时间外源插入广告。我们发现,放置在情感相似的环境中的广告可以提高广告参与度(较低的跳过率,较高的品牌回忆率)。基于我们的理论模型预测的情感相似性指标的广告插入,相对于非理论模型产生可比或更好的参与。摘要:This paper paper develops a theory-based, explainable deep learning convolutional neural network (CNN) classifier to predict the time-varying emotional response to music. We design novel CNN filters that leverage the frequency harmonics structure from acoustic physics known to impact the perception of musical features. Our theory-based model is more parsimonious, but provides comparable predictive performance to atheoretical deep learning models, while performing better than models using handcrafted features. Our model can be complemented with handcrafted features, but the performance improvement is marginal. Importantly, the harmonics-based structure placed on the CNN filters provides better explainability for how the model predicts emotional response (valence and arousal), because emotion is closely related to consonance--a perceptual feature defined by the alignment of harmonics. Finally, we illustrate the utility of our model with an application involving digital advertising. Motivated by YouTube mid-roll ads, we conduct a lab experiment in which we exogenously insert ads at different times within videos. We find that ads placed in emotionally similar contexts increase ad engagement (lower skip rates, higher brand recall rates). Ad insertion based on emotional similarity metrics predicted by our theory-based, explainable model produces comparable or better engagement relative to atheoretical models.

【8】 Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models
标题: 无监督盲关节去回响和使用扩散模型的房间声学估计
作者:Jean-Marie Lemercier,Eloi Moliner,Simon Welker,Vesa Välimäki,Timo Gerkmann
备注:Submitted to IEEEACM Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:本文提出了一种无监督的单通道盲去混响和室内冲激响应(RIR)估计方法BUDDy。该算法是植根于贝叶斯后验抽样:它结合了一个似然模型,强制保真度的混响测量,和一个无回声的语音先验实现的无条件扩散模型。我们设计了一个参数滤波器代表RIR,与指数衰减的每个频率子带。房间声学估计和语音去混响联合进行,滤波器参数迭代估计和语音发音沿逆扩散轨迹细化。在房间脉冲响应未知的盲场景中,BUDDy成功地在各种声学场景中执行语音去混响,显著优于其他盲无监督基线。与通常难以推广的监督方法不同,BUDDy可以无缝地适应不同的声学条件。本文扩展了我们以前的工作,提供新的实验结果和见解的算法的性能和多功能性。我们首先调查的鲁棒性通知去混响方法RIR估计误差,激励联合声学估计和去混响范例。然后,我们证明了我们的方法的适应性,高分辨率歌声去混响,研究其在RIR估计的性能,并进行主观评价实验,以验证结果的感知质量,以及其他贡献。音频样本和代码可以在网上找到。摘要:This paper presents an unsupervised method for single-channel blind dereverberation and room impulse response (RIR) estimation, called BUDDy. The algorithm is rooted in Bayesian posterior sampling: it combines a likelihood model enforcing fidelity to the reverberant measurement, and an anechoic speech prior implemented by an unconditional diffusion model. We design a parametric filter representing the RIR, with exponential decay for each frequency subband. Room acoustics estimation and speech dereverberation are jointly carried out, as the filter parameters are iteratively estimated and the speech utterance refined along the reverse diffusion trajectory. In a blind scenario where the room impulse response is unknown, BUDDy successfully performs speech dereverberation in various acoustic scenarios, significantly outperforming other blind unsupervised baselines. Unlike supervised methods, which often struggle to generalize, BUDDy seamlessly adapts to different acoustic conditions. This paper extends our previous work by offering new experimental results and insights into the algorithm's performance and versatility. We first investigate the robustness of informed dereverberation methods to RIR estimation errors, to motivate the joint acoustic estimation and dereverberation paradigm. Then, we demonstrate the adaptability of our method to high-resolution singing voice dereverberation, study its performance in RIR estimation, and conduct subjective evaluation experiments to validate the perceptual quality of the results, among other contributions. Audio samples and code can be found online.


eess.AS音频处理
【1】 Unsupervised Blind Joint Dereverberation and Room Acoustics Estimation with Diffusion Models
标题: 无监督盲关节去回响和使用扩散模型的房间声学估计
作者:Jean-Marie Lemercier,Eloi Moliner,Simon Welker,Vesa Välimäki,Timo Gerkmann
备注:Submitted to IEEEACM Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:本文提出了一种无监督的单通道盲去混响和室内冲激响应(RIR)估计方法BUDDy。该算法是植根于贝叶斯后验抽样:它结合了一个似然模型,强制保真度的混响测量,和一个无回声的语音先验实现的无条件扩散模型。我们设计了一个参数滤波器代表RIR,与指数衰减的每个频率子带。房间声学估计和语音去混响联合进行,滤波器参数迭代估计和语音发音沿逆扩散轨迹细化。在房间脉冲响应未知的盲场景中,BUDDy成功地在各种声学场景中执行语音去混响,显著优于其他盲无监督基线。与通常难以推广的监督方法不同,BUDDy可以无缝地适应不同的声学条件。本文扩展了我们以前的工作,提供新的实验结果和见解的算法的性能和多功能性。我们首先调查的鲁棒性通知去混响方法RIR估计误差,激励联合声学估计和去混响范例。然后,我们证明了我们的方法的适应性,高分辨率歌声去混响,研究其在RIR估计的性能,并进行主观评价实验,以验证结果的感知质量,以及其他贡献。音频样本和代码可以在网上找到。摘要:This paper presents an unsupervised method for single-channel blind dereverberation and room impulse response (RIR) estimation, called BUDDy. The algorithm is rooted in Bayesian posterior sampling: it combines a likelihood model enforcing fidelity to the reverberant measurement, and an anechoic speech prior implemented by an unconditional diffusion model. We design a parametric filter representing the RIR, with exponential decay for each frequency subband. Room acoustics estimation and speech dereverberation are jointly carried out, as the filter parameters are iteratively estimated and the speech utterance refined along the reverse diffusion trajectory. In a blind scenario where the room impulse response is unknown, BUDDy successfully performs speech dereverberation in various acoustic scenarios, significantly outperforming other blind unsupervised baselines. Unlike supervised methods, which often struggle to generalize, BUDDy seamlessly adapts to different acoustic conditions. This paper extends our previous work by offering new experimental results and insights into the algorithm's performance and versatility. We first investigate the robustness of informed dereverberation methods to RIR estimation errors, to motivate the joint acoustic estimation and dereverberation paradigm. Then, we demonstrate the adaptability of our method to high-resolution singing voice dereverberation, study its performance in RIR estimation, and conduct subjective evaluation experiments to validate the perceptual quality of the results, among other contributions. Audio samples and code can be found online.

【2】 WavLM model ensemble for audio deepfake detection
标题: 用于音频深度伪造检测的WavLM模型集成
作者:David Combei,Adriana Stan,Dan Oneata,Horia Cucu
备注:Accepted at ASVspoof Workshop 2024
链接:点击下载PDF文件
摘要:在过去的几年里,音频深度伪造检测已经成为一项关键任务,因为许多最近的语音合成和语音克隆系统生成高度逼真的语音样本,从而使其能够用于恶意活动。在本文中,我们解决了ASVspoof5挑战中设置的音频深度伪造检测问题。首先,我们对10种预训练表示进行了基准测试,并表明来自wav2vec2和wavLM家族的自监督表示表现最好。在这两者中,wavLM在将预训练数据限制为LibriSpeech时更好,正如挑战规则所要求的那样。为了进一步提高性能,我们对deepfake检测任务的wavLM模型进行了微调。我们使用来自其他deepfake检测数据集的样本扩展ASVspoof5数据集,并应用数据增强。我们的最终挑战提交包括四个模型的后期融合组合,并在两个评估集上实现了6.56%和17.08%的相等错误率。摘要:Audio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples, thus enabling their use in malicious activities. In this paper we address the issue of audio deepfake detection as it was set in the ASVspoof5 challenge. First, we benchmark ten types of pretrained representations and show that the self-supervised representations stemming from the wav2vec2 and wavLM families perform best. Of the two, wavLM is better when restricting the pretraining data to LibriSpeech, as required by the challenge rules. To further improve performance, we finetune the wavLM model for the deepfake detection task. We extend the ASVspoof5 dataset with samples from other deepfake detection datasets and apply data augmentation. Our final challenge submission consists of a late fusion combination of four models and achieves an equal error rate of 6.56% and 17.08% on the two evaluation sets.

【3】 MorphFader: Enabling Fine-grained Controllable Morphing with Text-to-Audio Models
标题: MorphFader:通过文本到音频模型启用细粒度可控变形
作者:Purnima Kamath,Chitralekha Gupta,Suranga Nanayakkara
备注:Under Review
链接:点击下载PDF文件
摘要:声音变形是一种逐渐平滑地将一种声音转换为另一种声音的过程,以产生同时类似于两者的新颖且感知混合的声音。最近,基于扩散的文本到音频模型已经使用文本提示产生了高质量的声音。然而,精细地控制声音的语义,这是变形所必需的,使用文本可能具有挑战性。在本文中,我们提出了 textit {MorphFader},一种可控的方法,用于使用文本到音频模型对不同提示生成的声音进行变形。通过在扩散过程中截取和插值交叉注意层的分量,我们可以在不同文本提示生成的声音之间创建平滑的变形。使用客观指标和感知听力测试,我们证明了我们的方法的能力,粒度控制的声音中的语义和生成平滑的变形。摘要:Sound morphing is the process of gradually and smoothly transforming one sound into another to generate novel and perceptually hybrid sounds that simultaneously resemble both. Recently, diffusion-based text-to-audio models have produced high-quality sounds using text prompts. However, granularly controlling the semantics of the sound, which is necessary for morphing, can be challenging using text. In this paper, we propose textit{MorphFader}, a controllable method for morphing sounds generated by disparate prompts using text-to-audio models. By intercepting and interpolating the components of the cross-attention layers within the diffusion process, we can create smooth morphs between sounds generated by different text prompts. Using both objective metrics and perceptual listening tests, we demonstrate the ability of our method to granularly control the semantics in the sound and generate smooth morphs.

【4】 Direction of Arrival Correction through Speech Quality Feedback
标题: 通过语音质量反馈进行到达方向纠正
作者:Caleb Rascon
备注:Submitted to Digital Signal Processing
链接:点击下载PDF文件
摘要:实时语音增强的性能已经开始上升,Demucs降噪模型最近在多语音源场景中表现出了很强的性能,同时伴有基于位置的语音目标选择策略。然而,它已被证明是敏感的到达方向(DOA)估计中的错误。在这项工作中,提出了一种DOA校正方案,使用其增强输出的实时估计的语音质量作为观测变量,在基于亚当的优化反馈回路,以找到正确的DOA。尽管语音质量估计的高度可变性,所提出的系统能够仅使用语音质量作为其指导来实时校正高达150的误差。为所提出的系统的未来版本提供了几种见解,以加速收敛并进一步降低语音质量估计的可变性。摘要:Real-time speech enhancement has began to rise in performance, and the Demucs Denoiser model has recently demonstrated strong performance in multiple-speech-source scenarios when accompanied by a location-based speech target selection strategy. However, it has shown to be sensitive to errors in the direction-of-arrival (DOA) estimation. In this work, a DOA correction scheme is proposed that uses the real-time estimated speech quality of its enhanced output as the observed variable in an Adam-based optimization feedback loop to find the correct DOA. In spite of the high variability of the speech quality estimation, the proposed system is able to correct in real-time an error of up to 15$^o$ using only the speech quality as its guide. Several insights are provided for future versions of the proposed system to speed up convergence and further reduce the speech quality estimation variability.

【5】 Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
标题: 口语刻板印象:评估言语大语言模型中对说话者的社会偏见
作者:Yi-Cheng Lin,Wei-Chih Chen,Hung-yi Lee
链接:点击下载PDF文件
摘要:警告:本文可能包含令人不快的内容。 大型语言模型(LLM)在各种任务中取得了卓越的性能,包括涉及语音等多模态数据的任务。然而,这些模型由于其训练数据的性质而经常表现出偏差。最近,出现了更多的语音大语言模型(SLLM),强调了解决这些偏见的迫切需要。本研究介绍了口语Stereoset,一个专门设计用于评估SLLM中社会偏见的数据集。通过研究不同的模型如何对来自不同人口群体的言论做出反应,我们的目标是识别这些偏见。我们的实验揭示了他们的性能和偏见水平的显着见解。研究结果表明,虽然大多数模型显示最小的偏见,一些仍然表现出轻微的刻板印象或反刻板印象的倾向。摘要:Warning: This paper may contain texts with uncomfortable content. Large Language Models (LLMs) have achieved remarkable performance in various tasks, including those involving multimodal data like speech. However, these models often exhibit biases due to the nature of their training data. Recently, more Speech Large Language Models (SLLMs) have emerged, underscoring the urgent need to address these biases. This study introduces Spoken Stereoset, a dataset specifically designed to evaluate social biases in SLLMs. By examining how different models respond to speech from diverse demographic groups, we aim to identify these biases. Our experiments reveal significant insights into their performance and bias levels. The findings indicate that while most models show minimal bias, some still exhibit slightly stereotypical or anti-stereotypical tendencies.

【6】 Transformers and Large Language Models for Efficient Intrusion Detection Systems: A Comprehensive Survey
标题: 高效入侵检测系统的Transformer和大型语言模型:全面调查
作者:Hamza Kheddar
备注:arXiv admin note: text overlap with arXiv:2405.04760 by other authors
链接:点击下载PDF文件
摘要:随着Transformers LLM的显著进步,NLP由于其在文本生成和用户交互方面的增强功能,已将其范围扩展到许多研究领域。从这些进步中受益匪浅的一个领域是网络安全。在网络安全中,许多需要保护的参数和在接收者和接收者之间交换的参数都是以文本和表格数据的形式存在的,这使得NLP成为增强通信协议安全措施的重要工具。本调查论文对网络威胁检测系统中Transformers和LLM的利用进行了全面分析。论文的选择和文献计量分析的方法进行了概述,建立一个严格的框架,评估现有的研究。讨论了Transformers的基本原理,包括该领域常用的各种网络攻击和数据集的背景信息。该调查探讨了Transformers在IDS中的应用,重点关注不同的架构,如基于注意力的模型,LLM(如BERT和GPT),CNN LSTM-Transformer混合,新兴方法(如ViTs)等。此外,它还探讨了已实施基于Transformers和LLMS的IDS的各种环境和应用,包括计算机网络、物联网设备、关键基础设施保护、云计算、SDN以及自动驾驶汽车。本文还讨论了这一领域的研究挑战和未来方向,确定了可解释性、可扩展性和对不断变化的威胁的适应性等关键问题。最后,结论总结了研究结果,并强调了Transformers和LLM在增强网络威胁检测能力方面的重要性,同时还概述了进一步研究和开发的潜在途径。摘要:With significant advancements in Transformers LLMs, NLP has extended its reach into many research fields due to its enhanced capabilities in text generation and user interaction. One field benefiting greatly from these advancements is cybersecurity. In cybersecurity, many parameters that need to be protected and exchanged between senders and receivers are in the form of text and tabular data, making NLP a valuable tool in enhancing the security measures of communication protocols. This survey paper provides a comprehensive analysis of the utilization of Transformers and LLMs in cyber-threat detection systems. The methodology of paper selection and bibliometric analysis is outlined to establish a rigorous framework for evaluating existing research. The fundamentals of Transformers are discussed, including background information on various cyber-attacks and datasets commonly used in this field. The survey explores the application of Transformers in IDSs, focusing on different architectures such as Attention-based models, LLMs like BERT and GPT, CNN LSTM-Transformer hybrids, emerging approaches like ViTs, among others. Furthermore, it explores the diverse environments and applications where Transformers and LLMs-based IDS have been implemented, including computer networks, IoT devices, critical infrastructure protection, cloud computing, SDN, as well as in autonomous vehicles. The paper also addresses research challenges and future directions in this area, identifying key issues such as interpretability, scalability, and adaptability to evolving threats, and more. Finally, the conclusion summarizes the findings and highlights the significance of Transformers and LLMs in enhancing cyber-threat detection capabilities, while also outlining potential avenues for further research and development.

【7】 PeriodWave: Multi-Period Flow Matching for High-Fidelity Waveform Generation
标题: PeriodWave:用于高保真波生成的多周期流匹配
作者:Sang-Hoon Lee,Ha-Yeong Choi,Seong-Whan Lee
备注:24 pages, 16 tables, 4 figures
链接:点击下载PDF文件
摘要:最近,通用波形生成任务已被调查条件下的各种分布情况。虽然基于GAN的方法在快速波形生成方面表现出了强大的实力,但它们容易受到训练推理不匹配的情况,例如两阶段文本到语音。与此同时,基于扩散的模型在其他领域显示出强大的生成性能;然而,由于波形生成任务的推理速度较慢,它们一直处于聚光灯下。最重要的是,没有任何发生器架构可以显式地解开高分辨率波形信号的自然周期特征。在本文中,我们提出了PeriodWave,一种新的通用波形生成模型。首先,我们介绍了一个周期感知的流量匹配估计器,可以捕捉波形信号的周期特征时,估计的向量场。此外,我们利用一个多周期估计,避免重叠,以捕捉不同的周期特征的波形信号。虽然增加周期数可以显著提高性能,但这需要更多的计算成本。为了减少这个问题,我们还提出了一个单一的周期条件的通用估计,可以前馈并行周期分批推理。此外,本文还利用离散小波变换对波形信号的频率信息进行离散化去纠缠以实现高频建模,并引入FreeU算法对波形生成过程中的高频噪声进行去噪。实验结果表明,该模型在Mel-声谱图重建和文本到语音转换任务中的性能优于以前的模型。所有源代码都可以在 url{https: github.com sh-lee-prml PeriodWave}上找到。摘要:Recently, universal waveform generation tasks have been investigated conditioned on various out-of-distribution scenarios. Although GAN-based methods have shown their strength in fast waveform generation, they are vulnerable to train-inference mismatch scenarios such as two-stage text-to-speech. Meanwhile, diffusion-based models have shown their powerful generative performance in other domains; however, they stay out of the limelight due to slow inference speed in waveform generation tasks. Above all, there is no generator architecture that can explicitly disentangle the natural periodic features of high-resolution waveform signals. In this paper, we propose PeriodWave, a novel universal waveform generation model. First, we introduce a period-aware flow matching estimator that can capture the periodic features of the waveform signal when estimating the vector fields. Additionally, we utilize a multi-period estimator that avoids overlaps to capture different periodic features of waveform signals. Although increasing the number of periods can improve the performance significantly, this requires more computational costs. To reduce this issue, we also propose a single period-conditional universal estimator that can feed-forward parallel by period-wise batch inference. Additionally, we utilize discrete wavelet transform to losslessly disentangle the frequency information of waveform signals for high-frequency modeling, and introduce FreeU to reduce the high-frequency noise for waveform generation. The experimental results demonstrated that our model outperforms the previous models both in Mel-spectrogram reconstruction and text-to-speech tasks. All source code will be available at url{https: github.com sh-lee-prml PeriodWave}.

【8】 Optimising MFCC parameters for the automatic detection of respiratory diseases
标题: 优化MFCC参数以自动检测呼吸道疾病
作者:Yuyang Yan,Sami O. Simons,Loes van Bemmel,Lauren Reinders,Frits M. E. Franssen,Visara Urovi
链接:点击下载PDF文件
摘要:源于呼吸道的语音信号被用作诊断和评估呼吸系统疾病的有价值的声学生物标志物。在所采用的声学特征中,Mel频率倒谱系数(MFCC)被广泛用于自动分析,MFCC提取通常依赖于默认参数。然而,没有全面的研究系统地研究了MFCC提取参数对呼吸系统疾病诊断的影响。在这项研究中,我们通过检查关键参数的影响,即系数的数量,帧长度和帧之间的跳长,呼吸状况检查解决这个差距。我们的调查使用了四个数据集:Cambridge COVID-19 Sound数据库、Coswara数据集、Saarbrucken Voice Disorders(SVD)数据库和TACTICAS数据集。支持向量机(SVM)被用作分类器,因为它被广泛采用和有效。我们的研究结果表明,MFCC的准确性随着跳长的增加而降低,并且观察到的最佳系数数量约为30。MFCC的性能随着数据集的帧长度而变化:对于COVID-19数据集(剑桥COVID-19 Sound数据库和Coswara数据集),性能随着帧长度的增加而下降,而对于SVD数据集,性能随着帧长度的增加而提高(从50 ms到500 ms)。此外,我们研究了这些参数的优化组合,并观察到大幅提高的准确性。与最差组合相比,SVM模型的准确率为81.1%、80.6%和71.7%,剑桥COVID-19 Sound数据库、Coswara数据集和SVD数据集的准确率分别提高了19.6%、16.10%和14.90%。摘要:Voice signals originating from the respiratory tract are utilized as valuable acoustic biomarkers for the diagnosis and assessment of respiratory diseases. Among the employed acoustic features, Mel Frequency Cepstral Coefficients (MFCC) is widely used for automatic analysis, with MFCC extraction commonly relying on default parameters. However, no comprehensive study has systematically investigated the impact of MFCC extraction parameters on respiratory disease diagnosis. In this study, we address this gap by examining the effects of key parameters, namely the number of coefficients, frame length, and hop length between frames, on respiratory condition examination. Our investigation uses four datasets: the Cambridge COVID-19 Sound database, the Coswara dataset, the Saarbrucken Voice Disorders (SVD) database, and a TACTICAS dataset. The Support Vector Machine (SVM) is employed as the classifier, given its widespread adoption and efficacy. Our findings indicate that the accuracy of MFCC decreases as hop length increases, and the optimal number of coefficients is observed to be approximately 30. The performance of MFCC varies with frame length across the datasets: for the COVID-19 datasets (Cambridge COVID-19 Sound database and Coswara dataset), performance declines with longer frame lengths, while for the SVD dataset, performance improves with increasing frame length (from 50 ms to 500 ms). Furthermore, we investigate the optimized combination of these parameters and observe substantial enhancements in accuracy. Compared to the worst combination, the SVM model achieves an accuracy of 81.1%, 80.6%, and 71.7%, with improvements of 19.6%, 16.10%, and 14.90% for the Cambridge COVID-19 Sound database, the Coswara dataset, and the SVD dataset respectively.

【9】 DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement
标题: DPSNN:用于低延迟流语音增强的尖峰神经网络
作者:Tao Sun,Sander Bohté
链接:点击下载PDF文件
摘要:语音增强(SE)改善了嘈杂环境中的通信,影响了自动语音识别,助听器和电信等领域。由于这些域通常是功率受限和基于事件的,同时需要低延迟,因此尖峰神经网络(SNN)形式的神经形态算法具有很大的潜力。然而,当前有效的SNN解决方案需要上下文采样窗口,该上下文采样窗口施加了相当大的延迟,通常在32 ms左右,对于许多应用来说太长。受经典神经网络中双路径脉冲神经网络(DPSNN)的启发,提出了一种两阶段时域流SNN框架--双路径脉冲神经网络(DPSNN)。在DPSNN中,第一阶段使用尖峰卷积神经网络(SCNN)来捕获全局上下文信息,而第二阶段使用尖峰循环神经网络(SRNN)来关注频率相关特征。此外,正则化抑制激活,以进一步提高我们的DPSNN的能量效率。通过对VCTK和英特尔DNS数据集的评估,我们证明了我们的方法实现了助听器等应用所需的极低延迟(约5 ms),同时表现出出色的信噪比(SNR)、感知质量和能效。摘要:Speech enhancement (SE) improves communication in noisy environments, affecting areas such as automatic speech recognition, hearing aids, and telecommunications. With these domains typically being power-constrained and event-based while requiring low latency, neuromorphic algorithms in the form of spiking neural networks (SNNs) have great potential. Yet, current effective SNN solutions require a contextual sampling window imposing substantial latency, typically around 32ms, too long for many applications. Inspired by Dual-Path Spiking Neural Networks (DPSNNs) in classical neural networks, we develop a two-phase time-domain streaming SNN framework -- the Dual-Path Spiking Neural Network (DPSNN). In the DPSNN, the first phase uses Spiking Convolutional Neural Networks (SCNNs) to capture global contextual information, while the second phase uses Spiking Recurrent Neural Networks (SRNNs) to focus on frequency-related features. In addition, the regularizer suppresses activation to further enhance energy efficiency of our DPSNNs. Evaluating on the VCTK and Intel DNS Datasets, we demonstrate that our approach achieves the very low latency (approximately 5ms) required for applications like hearing aids, while demonstrating excellent signal-to-noise ratio (SNR), perceptual quality, and energy efficiency.

【10】 Speech vs. Transcript: Does It Matter for Human Annotators in Speech Summarization?
标题: 语音与文字记录:语音总结中的人类注释者重要吗?
作者:Roshan Sharma,Suwon Shon,Mark Lindsey,Hira Dhamyal,Rita Singh,Bhiksha Raj
备注:Accepted to ACL 2024 Main Conference
链接:点击下载PDF文件
摘要:用于抽象语音摘要的参考摘要需要人工注释,这可以通过收听音频记录或通过阅读记录的文本转录来执行。在本文中,我们研究的基础上,注释者听录音的摘要是否不同于注释者阅读成绩单。使用现有的基于人工评估的内在评估,自动度量,基于LLM的评估和基于检索的无参考方法。我们发现,摘要确实是不同的源模态的基础上,和基于语音的摘要是更符合事实的一致性和信息选择性比基于成绩单的摘要。同时,基于成绩单的摘要受到源中识别错误的影响,专家撰写的摘要信息量更大,更可靠。我们将所有收集的数据和分析代码公开(https: github.com cmu-mlsp interview_humanssum),以方便复制我们的工作并推进该领域的研究。摘要:Reference summaries for abstractive speech summarization require human annotation, which can be performed by listening to an audio recording or by reading textual transcripts of the recording. In this paper, we examine whether summaries based on annotators listening to the recordings differ from those based on annotators reading transcripts. Using existing intrinsic evaluation based on human evaluation, automatic metrics, LLM-based evaluation, and a retrieval-based reference-free method. We find that summaries are indeed different based on the source modality, and that speech-based summaries are more factually consistent and information-selective than transcript-based summaries. Meanwhile, transcript-based summaries are impacted by recognition errors in the source, and expert-written summaries are more informative and reliable. We make all the collected data and analysis code public(https: github.com cmu-mlsp interview_humanssum) to facilitate the reproduction of our work and advance research in this area.

【11】 Play Me Something Icy: Practical Challenges, Explainability and the Semantic Gap in Generative AI Music
标题: 为我播放一些冰冷的东西:生成性人工智能音乐中的实践挑战、解释性和语义差距
作者:Jesse Allison,Drew Farrar,Treya Nash,Carlos Román,Morgan Weeks,Fiona Xue Ju
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:这张图片旨在批判性地考虑可解释AI背景下文本到音频和文本到音乐生成工具的性质。作为一群实验音乐家和研究人员,我们对这些工具的创造潜力充满热情,并试图从即时创作,控制,可用性,可理解性,人工智能过程的可解释性以及结果的整体美学效果的角度来理解和评估它们。我们已经确定,这些工具没有明确解决的挑战之一是使用基于文本的工具来描述像音乐这样抽象的东西时固有的语义差距。其他差距包括可解释性与可用性,以及用户控制和输入与人类创造过程。这张图片的目的是提出一些问题供讨论,并就我们希望在生成式人工智能音乐工具中看到的改进提出一些一般性建议。摘要:This pictorial aims to critically consider the nature of text-to-audio and text-to-music generative tools in the context of explainable AI. As a group of experimental musicians and researchers, we are enthusiastic about the creative potential of these tools and have sought to understand and evaluate them from perspectives of prompt creation, control, usability, understandability, explainability of the AI process, and overall aesthetic effectiveness of the results. One of the challenges we have identified that is not explicitly addressed by these tools is the inherent semantic gap in using text-based tools to describe something as abstract as music. Other gaps include explainability vs. useability, and user control and input vs. the human creative process. The aim of this pictorial is to raise questions for discussion and make a few general suggestions on the kinds of improvements we would like to see in generative AI music tools.

【12】 A Theory-Based Explainable Deep Learning Architecture for Music Emotion
标题: 基于理论的可解释的音乐情感深度学习架构
作者:Hortense Fong,Vineet Kumar,K. Sudhir
链接:点击下载PDF文件
摘要:本文开发了一种基于理论的、可解释的深度学习卷积神经网络(CNN)分类器,用于预测对音乐的时变情绪反应。我们设计了新颖的CNN滤波器,该滤波器利用声学物理学中已知的影响音乐特征感知的频率谐波结构。我们基于理论的模型更加简约,但提供了与非理论深度学习模型相当的预测性能,同时比使用手工制作的功能的模型表现更好。我们的模型可以补充手工制作的功能,但性能的改善是微不足道的。重要的是,放置在CNN滤波器上的基于谐波的结构为模型如何预测情绪反应(效价和唤醒)提供了更好的解释性,因为情绪与和谐密切相关-由谐波对齐定义的感知特征。最后,我们说明了我们的模型的实用程序涉及数字广告。受YouTube中置广告的启发,我们进行了一项实验室实验,在视频中的不同时间外源插入广告。我们发现,放置在情感相似的环境中的广告可以提高广告参与度(较低的跳过率,较高的品牌回忆率)。基于我们的理论模型预测的情感相似性指标的广告插入,相对于非理论模型产生可比或更好的参与。摘要:This paper paper develops a theory-based, explainable deep learning convolutional neural network (CNN) classifier to predict the time-varying emotional response to music. We design novel CNN filters that leverage the frequency harmonics structure from acoustic physics known to impact the perception of musical features. Our theory-based model is more parsimonious, but provides comparable predictive performance to atheoretical deep learning models, while performing better than models using handcrafted features. Our model can be complemented with handcrafted features, but the performance improvement is marginal. Importantly, the harmonics-based structure placed on the CNN filters provides better explainability for how the model predicts emotional response (valence and arousal), because emotion is closely related to consonance--a perceptual feature defined by the alignment of harmonics. Finally, we illustrate the utility of our model with an application involving digital advertising. Motivated by YouTube mid-roll ads, we conduct a lab experiment in which we exogenously insert ads at different times within videos. We find that ads placed in emotionally similar contexts increase ad engagement (lower skip rates, higher brand recall rates). Ad insertion based on emotional similarity metrics predicted by our theory-based, explainable model produces comparable or better engagement relative to atheoretical models.


机器翻译,仅供参考