今日论文合集:cs.SD语音13篇,eess.AS音频处理18篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
标题: PSM:多尺度Zero-Shot声景映射的学习概率嵌入
作者:Subash Khanal,Eric Xing,Srikumar Sastry,Aayush Dhakal,Zhexiao Xiong,Adeel Ahmad,Nathan Jacobs
备注:Accepted at ACM MM 2024
链接:点击下载PDF文件
摘要:声景由人在一个位置感知的声学环境定义。在这项工作中,我们提出了一个框架,在地球上映射音景。由于音景涉及跨越不同空间尺度的声音分布,我们用多尺度卫星图像表示位置,并学习图像,音频和文本之间的联合表示。为了捕捉一个位置的声景中固有的不确定性,我们将表示空间设计为概率性的。我们还融合无处不在的元数据(包括地理位置,时间和数据源),使学习的空间和时间动态表示的音景。我们展示了我们的框架的效用,通过创建大规模的音景地图集成音频和文本与时间控制。为了方便未来的研究这项任务,我们还介绍了一个大规模的数据集,GeoSound,包含超过30万$的地理标记的音频样本配对的低分辨率和高分辨率的卫星图像。我们证明了我们的方法在GeoSound和现有的SoundingEarth数据集上的性能优于现有的最先进的方法。我们的数据集和代码可以在https: github.com mvrl PSM上找到。摘要:A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial scales, we represent locations with multi-scale satellite imagery and learn a joint representation among this imagery, audio, and text. To capture the inherent uncertainty in the soundscape of a location, we design the representation space to be probabilistic. We also fuse ubiquitous metadata (including geolocation, time, and data source) to enable learning of spatially and temporally dynamic representations of soundscapes. We demonstrate the utility of our framework by creating large-scale soundscape maps integrating both audio and text with temporal control. To facilitate future research on this task, we also introduce a large-scale dataset, GeoSound, containing over $300k$ geotagged audio samples paired with both low- and high-resolution satellite imagery. We demonstrate that our method outperforms the existing state-of-the-art on both GeoSound and the existing SoundingEarth dataset. Our dataset and code is available at https: github.com mvrl PSM.

【2】 Source Separation of Multi-source Raw Music using a Residual Quantized Variational Autoencoder
标题: 使用残余量化变分自动编码器进行多源原始音乐的源分离
作者:Leonardo Berti
备注:9 pages
链接:点击下载PDF文件
摘要:我开发了一个基于残差量化变分自编码器架构的神经音频编解码器模型。我在Slakh2100数据集上训练模型,这是一个用于音乐源分离的标准数据集,由多轨音频组成。该模型可以分离音频源,以更少的计算能力实现几乎SoTA的结果。该代码可在github.com LeonardoBerti00 Source-Separation-of-Multi-source-Music-using-Residual-Quantizad-Variational-Autoencoder上公开获取摘要:I developed a neural audio codec model based on the residual quantized variational autoencoder architecture. I train the model on the Slakh2100 dataset, a standard dataset for musical source separation, composed of multi-track audio. The model can separate audio sources, achieving almost SoTA results with much less computing power. The code is publicly available at github.com LeonardoBerti00 Source-Separation-of-Multi-source-Music-using-Residual-Quantizad-Variational-Autoencoder

【3】 Content and Style Aware Audio-Driven Facial Animation
标题: 内容和风格感知音频驱动的面部动画
作者:Qingju Liu,Hyeongwoo Kim,Gaurav Bharaj
备注:BMVC2024
链接:点击下载PDF文件
摘要:音频驱动的3D面部动画有几个虚拟人应用程序用于内容创建和编辑。虽然几种现有的方法提供了语音驱动的动画的解决方案,但对最终表演的内容(什么)和风格(如何)的精确控制仍然具有挑战性。我们提出了一种新的方法,需要作为输入的音频和相应的文本提取时间对齐的内容和解开的风格表示,以提供控制3D面部动画。我们的方法分为两个阶段进行训练,从音频突出风格(听起来如何)演变为视觉突出风格(看起来如何)。我们在阶段I中利用高资源音频数据集来学习在自监督学习框架中控制语音生成的风格,然后在阶段II中使用低资源音频 3D网格对来微调该模型以控制3D顶点生成。我们采用了非自回归seq 2seq公式来模拟口腔水平的依赖关系,以及更好的口腔发音。我们的方法提供了灵活性,参考音频的风格和源音频的内容可以组合以实现音频风格传输。类似地,可以修改内容,例如静音或交换单词,这使得能够进行风格保留的内容编辑。摘要:Audio-driven 3D facial animation has several virtual humans applications for content creation and editing. While several existing methods provide solutions for speech-driven animation, precise control over content (what) and style (how) of the final performance is still challenging. We propose a novel approach that takes as input an audio, and the corresponding text to extract temporally-aligned content and disentangled style representations, in order to provide controls over 3D facial animation. Our method is trained in two stages, that evolves from audio prominent styles (how it sounds) to visual prominent styles (how it looks). We leverage a high-resource audio dataset in stage I to learn styles that control speech generation in a self-supervised learning framework, and then fine-tune this model with low-resource audio 3D mesh pairs in stage II to control 3D vertex generation. We employ a non-autoregressive seq2seq formulation to model sentence-level dependencies, and better mouth articulations. Our method provides flexibility that the style of a reference audio and the content of a source audio can be combined to enable audio style transfer. Similarly, the content can be modified, e.g. muting or swapping words, that enables style-preserving content editing.

【4】 Neural Speech and Audio Coding
标题: 神经语音和音频编码
作者:Minje Kim,Jan Skoglund
备注:Accepted for publication in IEEE Signal Processing Magazine
链接:点击下载PDF文件
摘要:本文探讨了神经语音和音频编码系统领域内基于模型和数据驱动方法的集成。它强调了语音和音频编解码器的主观评估过程所带来的挑战,并讨论了纯数据驱动方法的局限性,这些方法通常需要低效的大型架构来匹配基于模型的方法的性能。该研究提出了混合系统作为一个可行的解决方案,通过精心选择的设计增强传统编解码器的性能提供显着的改善。具体而言,它引入了一种基于神经网络的信号增强器,旨在对现有编解码器的输出进行后处理,以及基于自动编码器的端到端模型和LPCNet--将线性预测编码(LPC)与神经网络相结合的混合系统。此外,本文深入研究了在自定义特征空间(TF-编解码器)或预定义变换域(MDCTNet)内操作的预测模型,并研究了使用心理声学校准的损失函数来训练端到端神经音频编解码器。通过这些调查,本文展示了混合系统的潜力,以推进语音和音频编码领域的弥合传统的基于模型的方法和现代数据驱动技术之间的差距。摘要:This paper explores the integration of model-based and data-driven approaches within the realm of neural speech and audio coding systems. It highlights the challenges posed by the subjective evaluation processes of speech and audio codecs and discusses the limitations of purely data-driven approaches, which often require inefficiently large architectures to match the performance of model-based methods. The study presents hybrid systems as a viable solution, offering significant improvements to the performance of conventional codecs through meticulously chosen design enhancements. Specifically, it introduces a neural network-based signal enhancer designed to post-process existing codecs' output, along with the autoencoder-based end-to-end models and LPCNet--hybrid systems that combine linear predictive coding (LPC) with neural networks. Furthermore, the paper delves into predictive models operating within custom feature spaces (TF-Codec) or predefined transform domains (MDCTNet) and examines the use of psychoacoustically calibrated loss functions to train end-to-end neural audio codecs. Through these investigations, the paper demonstrates the potential of hybrid systems to advance the field of speech and audio coding by bridging the gap between traditional model-based approaches and modern data-driven techniques.

【5】 Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge
标题: 时间变异性和多视图自我监督表示来应对ASVspoof 5 Deepfake挑战
作者:Yuankun Xie,Xiaopeng Wang,Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Haonan Cheng,Long Ye
链接:点击下载PDF文件
摘要:ASVspoof 5是ASVspoof系列的第五版,是全球最大的音频安全挑战之一。它旨在促进发展的对策(CM),以区分真正的和欺骗性的言语话语。在本文中,我们专注于解决开放域音频deepfake检测的问题,该问题直接对应于ASVspoof5 Track1开放条件。首先,我们全面研究了ASVspoof5上的各种CM,包括数据扩展,数据增强和自监督学习(SSL)功能。由于ASVspoof5数据集的高频间隙特性,我们引入了频率屏蔽,这是一种屏蔽特定频带以提高CM鲁棒性的数据增强方法。结合各种尺度的时间信息与多个SSL特征,我们的实验在ASVspoof 5 Track 1评估进度集上实现了0.0158的minDCF和0.55%的EER。摘要:ASVspoof5, the fifth edition of the ASVspoof series, is one of the largest global audio security challenges. It aims to advance the development of countermeasure (CM) to discriminate bonafide and spoofed speech utterances. In this paper, we focus on addressing the problem of open-domain audio deepfake detection, which corresponds directly to the ASVspoof5 Track1 open condition. At first, we comprehensively investigate various CM on ASVspoof5, including data expansion, data augmentation, and self-supervised learning (SSL) features. Due to the high-frequency gaps characteristic of the ASVspoof5 dataset, we introduce Frequency Mask, a data augmentation method that masks specific frequency bands to improve CM robustness. Combining various scale of temporal information with multiple SSL features, our experiments achieved a minDCF of 0.0158 and an EER of 0.55% on the ASVspoof 5 Track 1 evaluation progress set.

【6】 Deep Learning for Speaker Identification: Architectural Insights from AB-1 Corpus Analysis and Performance Evaluation
标题: 用于说话者识别的深度学习:来自AB-1 Corpus分析和性能评估的架构见解
作者:Matthias Bartolo
备注:Resultant work from Assignment, Department of AI, University of Malta. Code available at: this https URL
链接:点击下载PDF文件
摘要:在安全系统、法医调查和个性化服务领域,语音作为基本人类输入的重要性超过了基于文本的交互。本研究深入研究说话人识别(SID)的复杂领域,研究其基本组成部分,并强调梅尔频谱图和梅尔频率倒谱系数(MFCC)的特征提取。此外,本研究使用广泛的分析来评估六个略有不同的模型架构,以评估其性能,并将超参数调整应用于性能最佳的模型。这项工作进行了语言分析,以验证口音和性别的准确性,除了在AB-1语料库数据集内的偏见评估。摘要:In the fields of security systems, forensic investigations, and personalized services, the importance of speech as a fundamental human input outweighs text-based interactions. This research delves deeply into the complex field of Speaker Identification (SID), examining its essential components and emphasising Mel Spectrogram and Mel Frequency Cepstral Coefficients (MFCC) for feature extraction. Moreover, this study evaluates six slightly distinct model architectures using extensive analysis to evaluate their performance, with hyperparameter tuning applied to the best-performing model. This work performs a linguistic analysis to verify accent and gender accuracy, in addition to bias evaluation within the AB-1 Corpus dataset.

【7】 Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
标题: 使用细粒度Inbox检测视听Deepfakes
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in BMVC 2024
链接:点击下载PDF文件
摘要:现有的音频-视频深度伪造检测方法主要集中在用于对音频和视频数据之间的不一致进行建模的高级特征上。因此,这些方法通常会忽略更精细的视听伪影,这是deepfake所固有的。在这里,我们提出了引入细粒度的机制,用于检测空间和时间域中的细微工件。首先,我们引入了一个本地的视听模型,能够捕捉小的空间区域,容易与音频不一致。为此,采用了基于空间局部距离与注意力模块相结合的细粒度机制。其次,我们引入了一个时间局部伪假增强,以包括在我们的训练集中包含微妙的时间不一致的样本。在DFDC和FakeAVCeleb数据集上的实验表明,与数据集内和跨数据集设置下的最新技术相比,所提出的方法在泛化方面具有优越性。摘要:Existing methods on audio-visual deepfake detection mainly focus on high-level features for modeling inconsistencies between audio and visual data. As a result, these approaches usually overlook finer audio-visual artifacts, which are inherent to deepfakes. Herein, we propose the introduction of fine-grained mechanisms for detecting subtle artifacts in both spatial and temporal domains. First, we introduce a local audio-visual model capable of capturing small spatial regions that are prone to inconsistencies with audio. For that purpose, a fine-grained mechanism based on a spatially-local distance coupled with an attention module is adopted. Second, we introduce a temporally-local pseudo-fake augmentation to include samples incorporating subtle temporal inconsistencies in our training set. Experiments on the DFDC and the FakeAVCeleb datasets demonstrate the superiority of the proposed method in terms of generalization as compared to the state-of-the-art under both in-dataset and cross-dataset settings.

【8】 Exploring the anatomy of articulation rate in spontaneous English speech: relationships between utterance length effects and social factors
标题: 探索自发英语言语清晰率的解剖:话语长度效应与社会因素之间的关系
作者:James Tanner,Morgan Sonderegger,Jane Stuart-Smith,Tyler Kendall,Jeff Mielke,Robin Dodsworth,Erik Thomas
备注:Proceedings of Interspeech 2024. 5 pages, 4 figures
链接:点击下载PDF文件
摘要:言语速率已被证明是不同的社会类别,如性别,年龄和方言,同时也受到言语规划的属性。话语长度的影响,其中语速更快,更长的话语,也被证明可以减少社会因素的作用,一旦它已经被占,留下不清楚的社会因素和语音生产之间的关系调节语速。通过对13个英语语音语料库的语速建模,发现话语长度对语速的影响最大,尽管这种影响本身在语料库和说话人之间变化不大。虽然年龄和性别也会调节语速,但它们的影响程度要小得多。这些研究结果表明,话语长度的影响可能是有条件的发音和知觉的限制,并应在更广泛的背景下解释的语速变化是如何构建的社会影响。摘要:Speech rate has been shown to vary across social categories such as gender, age, and dialect, while also being conditioned by properties of speech planning. The effect of utterance length, where speech rate is faster and less variable for longer utterances, has also been shown to reduce the role of social factors once it has been accounted for, leaving unclear the relationship between social factors and speech production in conditioning speech rate. Through modelling of speech rate across 13 English speech corpora, it is found that utterance length has the largest effect on speech rate, though this effect itself varies little across corpora and speakers. While age and gender also modulate speech rate, their effects are much smaller in magnitude. These findings suggest utterance length effects may be conditioned by articulatory and perceptual constraints, and that social influences on speech rate should be interpreted in the broader context of how speech rate variation is structured.

【9】 Music2Latent: Consistency Autoencoders for Latent Audio Compression
标题: Music 2潜伏:用于潜伏音频压缩的一致性自动编码器
作者:Marco Pasini,Stefan Lattner,George Fazekas
备注:Accepted to ISMIR 2024
链接:点击下载PDF文件
摘要:压缩连续潜在空间中的高效音频表示对于生成音频建模和音乐信息检索(MIR)任务至关重要。然而,一些现有的音频自动编码器具有局限性,诸如多级训练过程、缓慢的迭代采样或低重建质量。我们介绍Music2Latent,一个音频自动编码器,它通过利用一致性模型克服了这些限制。Music2Latent在单个端到端训练过程中将样本编码到压缩的连续潜在空间中,同时实现高保真单步重建。关键创新包括通过交叉连接在所有级别上对上采样编码器输出调节一致性模型,使用频率方面的自注意力来捕获长范围频率依赖性,以及采用频率方面的学习缩放来处理不同噪声水平下频率之间的变化值分布。我们证明了Music2Latent在音质和重建精度方面优于现有的连续音频自动编码器,同时使用其潜在表示在下游MIR任务上实现了有竞争力的性能。据我们所知,这是训练端到端一致性自动编码器模型的第一次成功尝试。摘要:Efficient audio representations in a compressed continuous latent space are critical for generative audio modeling and Music Information Retrieval (MIR) tasks. However, some existing audio autoencoders have limitations, such as multi-stage training procedures, slow iterative sampling, or low reconstruction quality. We introduce Music2Latent, an audio autoencoder that overcomes these limitations by leveraging consistency models. Music2Latent encodes samples into a compressed continuous latent space in a single end-to-end training process while enabling high-fidelity single-step reconstruction. Key innovations include conditioning the consistency model on upsampled encoder outputs at all levels through cross connections, using frequency-wise self-attention to capture long-range frequency dependencies, and employing frequency-wise learned scaling to handle varying value distributions across frequencies at different noise levels. We demonstrate that Music2Latent outperforms existing continuous audio autoencoders in sound quality and reconstruction accuracy while achieving competitive performance on downstream MIR tasks using its latent representations. To our knowledge, this represents the first successful attempt at training an end-to-end consistency autoencoder model.

【10】 TOGGL: Transcribing Overlapping Speech with Staggered Labeling
标题: TOGGL:用错开标签转录重叠的语音
作者:Chak-Fai Li,William Hartmann,Matthew Snover
备注:5 pages
链接:点击下载PDF文件
摘要:转录多个重叠说话者的语音通常需要将音频分离成多个流并独立地识别每个流。最近的工作联合分离和转录,但需要一个单独的解码组件为每个扬声器。我们提出了TOGGL模型,同时转录多个扬声器的语音。TOGGL模型使用特殊的输出标记来将语音归属于每个说话者,只有一个解码器。我们的方法可以推广到两个扬声器之外,即使只在两个扬声器数据上训练。我们表现出优越的性能相比,竞争的方法在会话语音数据集。我们的方法还提高了单扬声器音频的性能。摘要:Transcribing the speech of multiple overlapping speakers typically requires separating the audio into multiple streams and recognizing each one independently. More recent work jointly separates and transcribes, but requires a separate decoding component for each speaker. We propose the TOGGL model to simultaneously transcribe the speech of multiple speakers. The TOGGL model uses special output tokens to attribute the speech to each speaker with only a single decoder. Our approach generalizes beyond two speakers, even when trained only on two-speaker data. We demonstrate superior performance compared to competing approaches on a conversational speech dataset. Our approach also improves performance on single-speaker audio.

【11】 FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
标题: FoVNet:可配置的视野语音增强,具有低计算和低失真的智能眼镜
作者:Zhongweiyang Xu,Ali Aroudi,Ke Tan,Ashutosh Pandey,Jung-Suk Lee,Buye Xu,Francesco Nesta
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:本文提出了一种新的多通道语音增强方法,FoVNet,它可以在智能眼镜用户的可配置视野(FoV)内实现高效的语音增强,而不需要特定的目标说话者方向。它通过增强任何给定FoV内的所有扬声器来改进先前的工作,采用混合信号处理和深度学习方法,设计具有高计算效率。神经网络组件的设计具有超低的计算量(约50 MMACS)。多通道维纳滤波器和后处理模块进一步用于改善感知质量。我们评估我们的算法与智能眼镜上的麦克风阵列,提供了一个可配置的,高效的解决方案,增强听力的能源受限的设备。FoVNet在多个场景中的计算效率和语音质量都很出色,使其成为智能眼镜应用的一个有前途的解决方案。摘要:This paper presents a novel multi-channel speech enhancement approach, FoVNet, that enables highly efficient speech enhancement within a configurable field of view (FoV) of a smart-glasses user without needing specific target-talker(s) directions. It advances over prior works by enhancing all speakers within any given FoV, with a hybrid signal processing and deep learning approach designed with high computational efficiency. The neural network component is designed with ultra-low computation (about 50 MMACS). A multi-channel Wiener filter and a post-processing module are further used to improve perceptual quality. We evaluate our algorithm with a microphone array on smart glasses, providing a configurable, efficient solution for augmented hearing on energy-constrained devices. FoVNet excels in both computational efficiency and speech quality across multiple scenarios, making it a promising solution for smart glasses applications.

【12】 Dilated Convolution with Learnable Spacings
标题: 可学习间隔的扩张卷积
作者:Ismail Khalfaoui-Hassani
备注:PhD Thesis
链接:点击下载PDF文件
摘要:本文提出并评价了具有可学习间隔的扩张卷积(DCLS)方法。通过在计算机视觉、音频和语音处理领域的各种监督学习实验,DCLS方法被证明优于标准和高级卷积技术。该研究分为几个步骤,首先分析文献和现有的卷积技术之前的DCLS方法的发展。我们特别感兴趣的是与我们自己的方法密切相关的方法,这些方法对于捕捉我们方法的细微差别和独特性仍然至关重要。我们研究的基石是将DCLS方法引入卷积神经网络(CNN),以及依赖卷积和视觉注意力方法的混合架构。DCLS被证明是特别有效的任务,如分类,语义分割和对象检测。最初使用双线性插值,该研究还探索了其他插值方法,发现高斯插值略微提高了性能。DCLS方法进一步应用于尖峰神经网络(SNN),以实现神经网络内的突触延迟学习,最终可以转移到所谓的神经形态芯片。结果表明,DCLS方法脱颖而出,作为一个新的国家的最先进的技术在SNN音频分类在这个领域的某些基准任务。这些任务涉及具有高时间分量的数据集。此外,我们表明,DCLS可以显着提高人工神经网络的多标签音频分类任务的准确性。最后,我们讨论了所选择的实验装置,其局限性,我们的方法的局限性,我们的结果。摘要:This thesis presents and evaluates the Dilated Convolution with Learnable Spacings (DCLS) method. Through various supervised learning experiments in the fields of computer vision, audio, and speech processing, the DCLS method proves to outperform both standard and advanced convolution techniques. The research is organized into several steps, starting with an analysis of the literature and existing convolution techniques that preceded the development of the DCLS method. We were particularly interested in the methods that are closely related to our own and that remain essential to capture the nuances and uniqueness of our approach. The cornerstone of our study is the introduction and application of the DCLS method to convolutional neural networks (CNNs), as well as to hybrid architectures that rely on both convolutional and visual attention approaches. DCLS is shown to be particularly effective in tasks such as classification, semantic segmentation, and object detection. Initially using bilinear interpolation, the study also explores other interpolation methods, finding that Gaussian interpolation slightly improves performance. The DCLS method is further applied to spiking neural networks (SNNs) to enable synaptic delay learning within a neural network that could eventually be transferred to so-called neuromorphic chips. The results show that the DCLS method stands out as a new state-of-the-art technique in SNN audio classification for certain benchmark tasks in this field. These tasks involve datasets with a high temporal component. In addition, we show that DCLS can significantly improve the accuracy of artificial neural networks for the multi-label audio classification task. We conclude with a discussion of the chosen experimental setup, its limitations, the limitations of our method, and our results.

【13】 Lyrics Transcription for Humans: A Readability-Aware Benchmark
标题: 歌词人类转录:可读性基准
作者:Ondřej Cífka,Hendrik Schreiber,Luke Miner,Fabian-Robert Stöter
备注:ISMIR 2024 camera-ready. 6 pages + references + supplementary material. Website this https URL Data this https URL Code this https URL arXiv admin note: text overlap with arXiv:2311.13987
链接:点击下载PDF文件
摘要:为人类消费而写下歌词不仅涉及准确地捕捉单词序列,而且还包括标点符号和格式,以便清晰并传达上下文信息。这包括歌曲结构,情感重点,以及主唱和背景人声之间的对比。虽然自动歌词转录(ALT)系统已经超越了产生非结构化的单词串,并且能够利用更广泛的上下文,但ALT基准并没有跟上步伐,并继续专注于单词。为了解决这一差距,我们引入了Jam-ALT,一个全面的歌词转录基准。该基准测试对JamendoLyrics数据集进行了全面修订,符合歌词转录和格式的行业标准,以及旨在捕捉和评估歌词特定细微差别的评估指标,为提高歌词的可读性奠定了基础。我们应用基准最近的转录系统,并提出了额外的错误分析,以及与古典音乐数据集的实验比较。摘要:Writing down lyrics for human consumption involves not only accurately capturing word sequences, but also incorporating punctuation and formatting for clarity and to convey contextual information. This includes song structure, emotional emphasis, and contrast between lead and background vocals. While automatic lyrics transcription (ALT) systems have advanced beyond producing unstructured strings of words and are able to draw on wider context, ALT benchmarks have not kept pace and continue to focus exclusively on words. To address this gap, we introduce Jam-ALT, a comprehensive lyrics transcription benchmark. The benchmark features a complete revision of the JamendoLyrics dataset, in adherence to industry standards for lyrics transcription and formatting, along with evaluation metrics designed to capture and assess the lyric-specific nuances, laying the foundation for improving the readability of lyrics. We apply the benchmark to recent transcription systems and present additional error analysis, as well as an experimental comparison with a classical music dataset.


eess.AS音频处理
【1】 Heterogeneous Space Fusion and Dual-Dimension Attention: A New Paradigm for Speech Enhancement
标题: 异类空间融合和二维注意力:语音增强的新范式
作者:Tao Zheng,Liejun Wang,Yinfeng Yu
备注:Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2024
链接:点击下载PDF文件
摘要:自监督学习在语音任务中表现出令人印象深刻的性能,但在语音增强研究领域仍然有很大的发展机会。在处理语音任务时,仅将注意力机制限制在时间维度上,从而限制了有效地关注关键语音特征。考虑到上述问题,我们的研究介绍了一种新的语音增强框架,HFSDA,巧妙地整合异构空间特征,并结合了双维注意机制,显着提高语音清晰度和质量在嘈杂的环境。通过将自监督学习嵌入与短时傅里叶变换(STFT)频谱图特征相结合,我们的模型在捕获高级语义信息和详细频谱数据方面表现出色,从而能够对语音信号进行更全面的分析和细化。此外,我们在频谱图输入分支中采用了创新的全维动态卷积(ODConv)技术,从而增强了对多个维度关键信息的提取和整合。此外,我们改进的构象模型,提高其特征提取能力,不仅在时间维度,而且在整个频谱域。在VCTK-DEMAND数据集上的大量实验表明,HFSDA与现有的最先进的模型相当,证实了我们方法的有效性。摘要:Self-supervised learning has demonstrated impressive performance in speech tasks, yet there remains ample opportunity for advancement in the realm of speech enhancement research. In addressing speech tasks, confining the attention mechanism solely to the temporal dimension poses limitations in effectively focusing on critical speech features. Considering the aforementioned issues, our study introduces a novel speech enhancement framework, HFSDA, which skillfully integrates heterogeneous spatial features and incorporates a dual-dimension attention mechanism to significantly enhance speech clarity and quality in noisy environments. By leveraging self-supervised learning embeddings in tandem with Short-Time Fourier Transform (STFT) spectrogram features, our model excels at capturing both high-level semantic information and detailed spectral data, enabling a more thorough analysis and refinement of speech signals. Furthermore, we employ the innovative Omni-dimensional Dynamic Convolution (ODConv) technology within the spectrogram input branch, enabling enhanced extraction and integration of crucial information across multiple dimensions. Additionally, we refine the Conformer model by enhancing its feature extraction capabilities not only in the temporal dimension but also across the spectral domain. Extensive experiments on the VCTK-DEMAND dataset show that HFSDA is comparable to existing state-of-the-art models, confirming the validity of our approach.

【2】 VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
标题: VNet:用于语音合成声码器的基于GAN的多层鉴别器网络
作者:Yubing Cao,Yongming Li,Liejun Wang,Yinfeng Yu
备注:Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2024
链接:点击下载PDF文件
摘要:自生成对抗网络(GANs)引入语音合成领域以来,语音合成领域取得了显著的成绩。在对声码器的深入探索中,人们发现可以以超过实时的速度生成音频波形,同时保持高保真度,这是通过利用基于GAN的模型实现的。通常,声码器的输入包括带限频谱信息,这不可避免地牺牲了高频细节。为了解决这个问题,我们采用全带梅尔频谱图信息作为输入,旨在提供最全面的信息可能的声码器。然而,以前的研究表明,使用全频带频谱信息作为输入可能会导致过度平滑的问题,损害合成语音的自然度。为了应对这一挑战,我们提出了VNet,这是一种基于GAN的神经声码器网络,它结合了全频带频谱信息,并引入了一个包括多个子鉴别器的多层鉴别器(MTD)来生成高分辨率信号。此外,我们引入了一种渐进约束方法,该方法修改了生成器和训练器的对抗损失,增强了训练过程的稳定性。通过严格的实验,我们证明了VNet模型能够产生高保真语音,并显着提高声码器的性能。摘要:Since the introduction of Generative Adversarial Networks (GANs) in speech synthesis, remarkable achievements have been attained. In a thorough exploration of vocoders, it has been discovered that audio waveforms can be generated at speeds exceeding real-time while maintaining high fidelity, achieved through the utilization of GAN-based models. Typically, the inputs to the vocoder consist of band-limited spectral information, which inevitably sacrifices high-frequency details. To address this, we adopt the full-band Mel spectrogram information as input, aiming to provide the vocoder with the most comprehensive information possible. However, previous studies have revealed that the use of full-band spectral information as input can result in the issue of over-smoothing, compromising the naturalness of the synthesized speech. To tackle this challenge, we propose VNet, a GAN-based neural vocoder network that incorporates full-band spectral information and introduces a Multi-Tier Discriminator (MTD) comprising multiple sub-discriminators to generate high-resolution signals. Additionally, we introduce an asymptotically constrained method that modifies the adversarial loss of the generator and discriminator, enhancing the stability of the training process. Through rigorous experiments, we demonstrate that the VNet model is capable of generating high-fidelity speech and significantly improving the performance of the vocoder.

【3】 SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis
标题: SaSLaW:具有视听自我中心信息的对话语音库,实现环境适应性对话语音合成
作者:Osamu Take,Shinnosuke Takamichi,Kentaro Seki,Yoshiaki Bando,Hiroshi Saruwatari
备注:5 pages, accepted for INTERSPEECH 2024
链接:点击下载PDF文件
摘要:本文介绍了SaSLaW,一个自发的对话语音语料库,包含同步录音的发言者说,听,看。在面对面的语音交流中,人们会考虑各种环境因素,从而控制自己的话语特征。能够适应这些音频环境的口语对话系统能够实现自然和无缝的通信。SaSLaW的开发是为了通过自发对话中的第一人称视听感知来模拟音频环境中的人类语音调整。我们提出了SaSLaW的构建方法并展示了语料库的分析结果。我们还进行了一项实验,使用SaSLaW开发文本到语音模型,并评估其适应音频环境的性能。结果表明,模型将听觉音频数据输出更合理的语音定制不同的音频环境比香草文本到语音模型。摘要:This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environmental factors and then control the features of their utterances in face-to-face voice communications. Spoken dialogue systems capable of this adaptation to these audio environments enable natural and seamless communications. SaSLaW was developed to model human-speech adjustment for audio environments via first-person audio-visual perceptions in spontaneous dialogues. We propose the construction methodology of SaSLaW and display the analysis result of the corpus. We additionally conducted an experiment to develop text-to-speech models using SaSLaW and evaluate their performance of adaptations to audio environments. The results indicate that models incorporating hearing-audio data output more plausible speech tailored to diverse audio environments than the vanilla text-to-speech model.

【4】 BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
标题: BSS-CFFMA:基于自监督嵌入的跨域特征融合和多关注语音增强网络
作者:Alimjan Mattursun,Liejun Wang,Yinfeng Yu
备注:Accepted for publication by IEEE International Conference on Systems, Man, and Cybernetics 2024
链接:点击下载PDF文件
摘要:语音自监督学习(SSL)代表了在多个下游任务中达到最先进(SOTA)的性能。然而,它在语音增强(SE)任务中的应用仍然不成熟,提供了改进的机会。在这项研究中,我们介绍了一种新的跨域特征融合和多注意力语音增强网络,称为BSS-CFFMA,它利用自监督嵌入。BSS-CFFMA包括多尺度跨域特征融合(MSCFF)块和残差混合多注意(RHMA)块。MSCFF块有效地集成了跨域特征,便于提取丰富的声学信息。RHMA块,作为主要的增强模块,利用三个不同的注意力模块来捕获不同的注意力表示和估计高质量的语音信号。 我们通过对VoiceBank-DEMAND数据集的比较和消融研究来评估BSS-CFFMA模型的性能,从而获得SOTA结果。此外,我们从WHAMR!数据集,一个专门为语音增强任务设计的集合,以评估BSS-CFFMA在仅去噪、仅去混响以及同时去噪和去混响等任务中的能力。这项研究标志着首次尝试探索基于自监督嵌入的语音增强方法在包括去混响和同时去噪和去混响的复杂任务中的有效性。BSS-CFFMA的演示实现可在线获得 footnote[2]{https:github.com AlimMat BSS-CFFMA. label{s1}}。摘要:Speech self-supervised learning (SSL) represents has achieved state-of-the-art (SOTA) performance in multiple downstream tasks. However, its application in speech enhancement (SE) tasks remains immature, offering opportunities for improvement. In this study, we introduce a novel cross-domain feature fusion and multi-attention speech enhancement network, termed BSS-CFFMA, which leverages self-supervised embeddings. BSS-CFFMA comprises a multi-scale cross-domain feature fusion (MSCFF) block and a residual hybrid multi-attention (RHMA) block. The MSCFF block effectively integrates cross-domain features, facilitating the extraction of rich acoustic information. The RHMA block, serving as the primary enhancement module, utilizes three distinct attention modules to capture diverse attention representations and estimate high-quality speech signals. We evaluate the performance of the BSS-CFFMA model through comparative and ablation studies on the VoiceBank-DEMAND dataset, achieving SOTA results. Furthermore, we select three types of data from the WHAMR! dataset, a collection specifically designed for speech enhancement tasks, to assess the capabilities of BSS-CFFMA in tasks such as denoising only, dereverberation only, and simultaneous denoising and dereverberation. This study marks the first attempt to explore the effectiveness of self-supervised embedding-based speech enhancement methods in complex tasks encompassing dereverberation and simultaneous denoising and dereverberation. The demo implementation of BSS-CFFMA is available online footnote[2]{https: github.com AlimMat BSS-CFFMA. label{s1}}.

【5】 PRESENT: Zero-Shot Text-to-Prosody Control
标题: 存在:Zero-Shot文本到韵律控制
作者:Perry Lam,Huayun Zhang,Nancy F. Chen,Berrak Sisman,Dorien Herremans
链接:点击下载PDF文件
摘要:目前在语音合成中实现细粒度韵律控制的策略需要提取额外的风格嵌入或采用更复杂的架构。为了实现预训练的文本到语音(TTS)模型的zero-shot应用,我们提出了PRESENT(没有样式嵌入或新训练的PRosody编辑),它通过直接修改推理过程来利用基于FastSpeech2的模型中的显式韵律预测。我们将我们的文本到韵律的框架zero-shot语言迁移使用JETS模型专门训练英语LJSpeech数据。我们获得的字符错误率(CER)分别为12.8%,18.7%和5.9%的德语,匈牙利语和西班牙语,超过2倍以上的所有三种语言的先前国家的最先进的CER。此外,我们允许子音素级控制,这是该领域的第一个。为了评估其有效性,我们表明,PRESENT可以提高问题的韵律,并使用它来生成普通话,一个音调的语言,元音音高变化在子音素水平。用JETS模型计算汉字的CER为25.3%,汉字的CER为13.0%。我们所有的代码和音频示例都可以在线获得。摘要:Current strategies for achieving fine-grained prosody control in speech synthesis entail extracting additional style embeddings or adopting more complex architectures. To enable zero-shot application of pretrained text-to-speech (TTS) models, we present PRESENT (PRosody Editing without Style Embeddings or New Training), which exploits explicit prosody prediction in FastSpeech2-based models by modifying the inference process directly. We apply our text-to-prosody framework to zero-shot language transfer using a JETS model exclusively trained on English LJSpeech data. We obtain character error rates (CER) of 12.8%, 18.7% and 5.9% for German, Hungarian and Spanish respectively, beating the previous state-of-the-art CER by over 2x for all three languages. Furthermore, we allow subphoneme-level control, a first in this field. To evaluate its effectiveness, we show that PRESENT can improve the prosody of questions, and use it to generate Mandarin, a tonal language where vowel pitch varies at subphoneme level. We attain 25.3% hanzi CER and 13.0% pinyin CER with the JETS model. All our code and audio samples are available online.

【6】 Lyrics Transcription for Humans: A Readability-Aware Benchmark
标题: 歌词人类转录:可读性基准
作者:Ondřej Cífka,Hendrik Schreiber,Luke Miner,Fabian-Robert Stöter
备注:ISMIR 2024 camera-ready. 6 pages + references + supplementary material. Website this https URL Data this https URL Code this https URL arXiv admin note: text overlap with arXiv:2311.13987
链接:点击下载PDF文件
摘要:为人类消费而写下歌词不仅涉及准确地捕捉单词序列,而且还包括标点符号和格式,以便清晰并传达上下文信息。这包括歌曲结构,情感重点,以及主唱和背景人声之间的对比。虽然自动歌词转录(ALT)系统已经超越了产生非结构化的单词串,并且能够利用更广泛的上下文,但ALT基准并没有跟上步伐,并继续专注于单词。为了解决这一差距,我们引入了Jam-ALT,一个全面的歌词转录基准。该基准测试对JamendoLyrics数据集进行了全面修订,符合歌词转录和格式的行业标准,以及旨在捕捉和评估歌词特定细微差别的评估指标,为提高歌词的可读性奠定了基础。我们应用基准最近的转录系统,并提出了额外的错误分析,以及与古典音乐数据集的实验比较。摘要:Writing down lyrics for human consumption involves not only accurately capturing word sequences, but also incorporating punctuation and formatting for clarity and to convey contextual information. This includes song structure, emotional emphasis, and contrast between lead and background vocals. While automatic lyrics transcription (ALT) systems have advanced beyond producing unstructured strings of words and are able to draw on wider context, ALT benchmarks have not kept pace and continue to focus exclusively on words. To address this gap, we introduce Jam-ALT, a comprehensive lyrics transcription benchmark. The benchmark features a complete revision of the JamendoLyrics dataset, in adherence to industry standards for lyrics transcription and formatting, along with evaluation metrics designed to capture and assess the lyric-specific nuances, laying the foundation for improving the readability of lyrics. We apply the benchmark to recent transcription systems and present additional error analysis, as well as an experimental comparison with a classical music dataset.

【7】 PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
标题: PSM:多尺度Zero-Shot声景映射的学习概率嵌入
作者:Subash Khanal,Eric Xing,Srikumar Sastry,Aayush Dhakal,Zhexiao Xiong,Adeel Ahmad,Nathan Jacobs
备注:Accepted at ACM MM 2024
链接:点击下载PDF文件
摘要:声景由人在一个位置感知的声学环境定义。在这项工作中,我们提出了一个绘制地球音景的框架。由于音景涉及跨越不同空间尺度的声音分布,我们用多尺度卫星图像表示位置,并学习图像,音频和文本之间的联合表示。为了捕捉一个位置的声景中固有的不确定性,我们将表示空间设计为概率性的。我们还融合无处不在的元数据(包括地理位置,时间和数据源),使学习的空间和时间动态表示的音景。我们展示了我们的框架的效用,通过创建大规模的音景地图集成音频和文本与时间控制。为了方便未来的研究这项任务,我们还介绍了一个大规模的数据集,GeoSound,包含超过30万$的地理标记的音频样本配对的低分辨率和高分辨率的卫星图像。我们证明了我们的方法在GeoSound和现有的SoundingEarth数据集上的性能优于现有的最先进的方法。我们的数据集和代码可以在https: github.com mvrl PSM上找到。摘要:A soundscape is defined by the acoustic environment a person perceives at a location. In this work, we propose a framework for mapping soundscapes across the Earth. Since soundscapes involve sound distributions that span varying spatial scales, we represent locations with multi-scale satellite imagery and learn a joint representation among this imagery, audio, and text. To capture the inherent uncertainty in the soundscape of a location, we design the representation space to be probabilistic. We also fuse ubiquitous metadata (including geolocation, time, and data source) to enable learning of spatially and temporally dynamic representations of soundscapes. We demonstrate the utility of our framework by creating large-scale soundscape maps integrating both audio and text with temporal control. To facilitate future research on this task, we also introduce a large-scale dataset, GeoSound, containing over $300k$ geotagged audio samples paired with both low- and high-resolution satellite imagery. We demonstrate that our method outperforms the existing state-of-the-art on both GeoSound and the existing SoundingEarth dataset. Our dataset and code is available at https: github.com mvrl PSM.

【8】 Source Separation of Multi-source Raw Music using a Residual Quantized Variational Autoencoder
标题: 使用残余量化变分自动编码器进行多源原始音乐的源分离
作者:Leonardo Berti
备注:9 pages
链接:点击下载PDF文件
摘要:我开发了一个基于残差量化变分自编码器架构的神经音频编解码器模型。我在Slakh2100数据集上训练模型,这是一个用于音乐源分离的标准数据集,由多轨音频组成。该模型可以分离音频源,以更少的计算能力实现几乎SoTA的结果。该代码可在github.com LeonardoBerti00 Source-Separation-of-Multi-source-Music-using-Residual-Quantizad-Variational-Autoencoder上公开获取摘要:I developed a neural audio codec model based on the residual quantized variational autoencoder architecture. I train the model on the Slakh2100 dataset, a standard dataset for musical source separation, composed of multi-track audio. The model can separate audio sources, achieving almost SoTA results with much less computing power. The code is publicly available at github.com LeonardoBerti00 Source-Separation-of-Multi-source-Music-using-Residual-Quantizad-Variational-Autoencoder

【9】 Content and Style Aware Audio-Driven Facial Animation
标题: 内容和风格感知音频驱动的面部动画
作者:Qingju Liu,Hyeongwoo Kim,Gaurav Bharaj
备注:BMVC2024
链接:点击下载PDF文件
摘要:音频驱动的3D面部动画有几个虚拟人应用程序用于内容创建和编辑。虽然几种现有的方法提供了语音驱动的动画的解决方案,但对最终表演的内容(什么)和风格(如何)的精确控制仍然具有挑战性。我们提出了一种新的方法,需要作为输入的音频和相应的文本提取时间对齐的内容和解开的风格表示,以提供控制3D面部动画。我们的方法分为两个阶段进行训练,从音频突出风格(听起来如何)演变为视觉突出风格(看起来如何)。我们在阶段I中利用高资源音频数据集来学习在自监督学习框架中控制语音生成的风格,然后在阶段II中使用低资源音频 3D网格对来微调该模型以控制3D顶点生成。我们采用了非自回归seq 2seq公式来模拟口腔水平的依赖关系,以及更好的口腔发音。我们的方法提供了灵活性,参考音频的风格和源音频的内容可以组合以实现音频风格传输。类似地,可以修改内容,例如静音或交换单词,这使得能够进行风格保留的内容编辑。摘要:Audio-driven 3D facial animation has several virtual humans applications for content creation and editing. While several existing methods provide solutions for speech-driven animation, precise control over content (what) and style (how) of the final performance is still challenging. We propose a novel approach that takes as input an audio, and the corresponding text to extract temporally-aligned content and disentangled style representations, in order to provide controls over 3D facial animation. Our method is trained in two stages, that evolves from audio prominent styles (how it sounds) to visual prominent styles (how it looks). We leverage a high-resource audio dataset in stage I to learn styles that control speech generation in a self-supervised learning framework, and then fine-tune this model with low-resource audio 3D mesh pairs in stage II to control 3D vertex generation. We employ a non-autoregressive seq2seq formulation to model sentence-level dependencies, and better mouth articulations. Our method provides flexibility that the style of a reference audio and the content of a source audio can be combined to enable audio style transfer. Similarly, the content can be modified, e.g. muting or swapping words, that enables style-preserving content editing.

【10】 Neural Speech and Audio Coding
标题: 神经语音和音频编码
作者:Minje Kim,Jan Skoglund
备注:Accepted for publication in IEEE Signal Processing Magazine
链接:点击下载PDF文件
摘要:本文探讨了神经语音和音频编码系统领域内基于模型和数据驱动方法的集成。它强调了语音和音频编解码器的主观评估过程所带来的挑战,并讨论了纯数据驱动方法的局限性,这些方法通常需要低效的大型架构来匹配基于模型的方法的性能。该研究提出了混合系统作为一个可行的解决方案,通过精心选择的设计增强传统编解码器的性能提供显着的改善。具体而言,它引入了一种基于神经网络的信号增强器,旨在对现有编解码器的输出进行后处理,以及基于自动编码器的端到端模型和LPCNet--将线性预测编码(LPC)与神经网络相结合的混合系统。此外,本文深入研究了在自定义特征空间(TF-编解码器)或预定义变换域(MDCTNet)内操作的预测模型,并研究了使用心理声学校准的损失函数来训练端到端神经音频编解码器。通过这些调查,本文展示了混合系统的潜力,以推进语音和音频编码领域的弥合传统的基于模型的方法和现代数据驱动技术之间的差距。摘要:This paper explores the integration of model-based and data-driven approaches within the realm of neural speech and audio coding systems. It highlights the challenges posed by the subjective evaluation processes of speech and audio codecs and discusses the limitations of purely data-driven approaches, which often require inefficiently large architectures to match the performance of model-based methods. The study presents hybrid systems as a viable solution, offering significant improvements to the performance of conventional codecs through meticulously chosen design enhancements. Specifically, it introduces a neural network-based signal enhancer designed to post-process existing codecs' output, along with the autoencoder-based end-to-end models and LPCNet--hybrid systems that combine linear predictive coding (LPC) with neural networks. Furthermore, the paper delves into predictive models operating within custom feature spaces (TF-Codec) or predefined transform domains (MDCTNet) and examines the use of psychoacoustically calibrated loss functions to train end-to-end neural audio codecs. Through these investigations, the paper demonstrates the potential of hybrid systems to advance the field of speech and audio coding by bridging the gap between traditional model-based approaches and modern data-driven techniques.

【11】 Temporal Variability and Multi-Viewed Self-Supervised Representations to Tackle the ASVspoof5 Deepfake Challenge
标题: 时间变异性和多视图自我监督表示来应对ASVspoof 5 Deepfake挑战
作者:Yuankun Xie,Xiaopeng Wang,Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Haonan Cheng,Long Ye
链接:点击下载PDF文件
摘要:ASVspoof 5是ASVspoof系列的第五版,是全球最大的音频安全挑战之一。它旨在促进发展的对策(CM),以区分真正的和欺骗性的言语话语。在本文中,我们专注于解决开放域音频deepfake检测的问题,该问题直接对应于ASVspoof5 Track1开放条件。首先,我们全面研究了ASVspoof5上的各种CM,包括数据扩展,数据增强和自监督学习(SSL)功能。由于ASVspoof5数据集的高频间隙特性,我们引入了频率屏蔽,这是一种屏蔽特定频带以提高CM鲁棒性的数据增强方法。结合各种尺度的时间信息与多个SSL特征,我们的实验在ASVspoof 5 Track 1评估进度集上实现了0.0158的minDCF和0.55%的EER。摘要:ASVspoof5, the fifth edition of the ASVspoof series, is one of the largest global audio security challenges. It aims to advance the development of countermeasure (CM) to discriminate bonafide and spoofed speech utterances. In this paper, we focus on addressing the problem of open-domain audio deepfake detection, which corresponds directly to the ASVspoof5 Track1 open condition. At first, we comprehensively investigate various CM on ASVspoof5, including data expansion, data augmentation, and self-supervised learning (SSL) features. Due to the high-frequency gaps characteristic of the ASVspoof5 dataset, we introduce Frequency Mask, a data augmentation method that masks specific frequency bands to improve CM robustness. Combining various scale of temporal information with multiple SSL features, our experiments achieved a minDCF of 0.0158 and an EER of 0.55% on the ASVspoof 5 Track 1 evaluation progress set.

【12】 Deep Learning for Speaker Identification: Architectural Insights from AB-1 Corpus Analysis and Performance Evaluation
标题: 用于说话者识别的深度学习:来自AB-1 Corpus分析和性能评估的架构见解
作者:Matthias Bartolo
备注:Resultant work from Assignment, Department of AI, University of Malta. Code available at: this https URL
链接:点击下载PDF文件
摘要:在安全系统、法医调查和个性化服务领域,语音作为基本人类输入的重要性超过了基于文本的交互。本研究深入研究说话人识别(SID)的复杂领域,研究其基本组成部分,并强调梅尔频谱图和梅尔频率倒谱系数(MFCC)的特征提取。此外,本研究使用广泛的分析来评估六个略有不同的模型架构,以评估其性能,并将超参数调整应用于性能最佳的模型。这项工作进行了语言分析,以验证口音和性别的准确性,除了在AB-1语料库数据集内的偏见评估。摘要:In the fields of security systems, forensic investigations, and personalized services, the importance of speech as a fundamental human input outweighs text-based interactions. This research delves deeply into the complex field of Speaker Identification (SID), examining its essential components and emphasising Mel Spectrogram and Mel Frequency Cepstral Coefficients (MFCC) for feature extraction. Moreover, this study evaluates six slightly distinct model architectures using extensive analysis to evaluate their performance, with hyperparameter tuning applied to the best-performing model. This work performs a linguistic analysis to verify accent and gender accuracy, in addition to bias evaluation within the AB-1 Corpus dataset.

【13】 Detecting Audio-Visual Deepfakes with Fine-Grained Inconsistencies
标题: 使用细粒度Inbox检测视听Deepfakes
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in BMVC 2024
链接:点击下载PDF文件
摘要:现有的音频-视频深度伪造检测方法主要集中在用于对音频和视频数据之间的不一致进行建模的高级特征上。因此,这些方法通常会忽略更精细的视听伪影,这是deepfake所固有的。在这里,我们提出了引入细粒度的机制,用于检测空间和时间域中的细微工件。首先,我们引入了一个本地的视听模型,能够捕捉小的空间区域,容易与音频不一致。为此,采用了一种基于空间局部距离的细粒度机制与注意力模块相结合。其次,我们引入了一个时间局部伪假增强,以包括在我们的训练集中包含微妙的时间不一致的样本。在DFDC和FakeAVCeleb数据集上的实验表明,与数据集内和跨数据集设置下的最新技术相比,所提出的方法在泛化方面具有优越性。摘要:Existing methods on audio-visual deepfake detection mainly focus on high-level features for modeling inconsistencies between audio and visual data. As a result, these approaches usually overlook finer audio-visual artifacts, which are inherent to deepfakes. Herein, we propose the introduction of fine-grained mechanisms for detecting subtle artifacts in both spatial and temporal domains. First, we introduce a local audio-visual model capable of capturing small spatial regions that are prone to inconsistencies with audio. For that purpose, a fine-grained mechanism based on a spatially-local distance coupled with an attention module is adopted. Second, we introduce a temporally-local pseudo-fake augmentation to include samples incorporating subtle temporal inconsistencies in our training set. Experiments on the DFDC and the FakeAVCeleb datasets demonstrate the superiority of the proposed method in terms of generalization as compared to the state-of-the-art under both in-dataset and cross-dataset settings.

【14】 Exploring the anatomy of articulation rate in spontaneous English speech: relationships between utterance length effects and social factors
标题: 探索自发英语言语清晰率的解剖:话语长度效应与社会因素之间的关系
作者:James Tanner,Morgan Sonderegger,Jane Stuart-Smith,Tyler Kendall,Jeff Mielke,Robin Dodsworth,Erik Thomas
备注:Proceedings of Interspeech 2024. 5 pages, 4 figures
链接:点击下载PDF文件
摘要:言语速率已被证明是不同的社会类别,如性别,年龄和方言,同时也受到言语规划的属性。话语长度的影响,其中语速更快,更长的话语,也被证明可以减少社会因素的作用,一旦它已经被占,留下不清楚的社会因素和语音生产之间的关系调节语速。通过对13个英语语音语料库的语速建模,发现话语长度对语速的影响最大,尽管这种影响本身在语料库和说话人之间变化不大。虽然年龄和性别也会调节语速,但它们的影响程度要小得多。这些研究结果表明,话语长度的影响可能是有条件的发音和知觉的限制,并应在更广泛的背景下解释的语速变化是如何构建的社会影响。摘要:Speech rate has been shown to vary across social categories such as gender, age, and dialect, while also being conditioned by properties of speech planning. The effect of utterance length, where speech rate is faster and less variable for longer utterances, has also been shown to reduce the role of social factors once it has been accounted for, leaving unclear the relationship between social factors and speech production in conditioning speech rate. Through modelling of speech rate across 13 English speech corpora, it is found that utterance length has the largest effect on speech rate, though this effect itself varies little across corpora and speakers. While age and gender also modulate speech rate, their effects are much smaller in magnitude. These findings suggest utterance length effects may be conditioned by articulatory and perceptual constraints, and that social influences on speech rate should be interpreted in the broader context of how speech rate variation is structured.

【15】 Music2Latent: Consistency Autoencoders for Latent Audio Compression
标题: Music 2潜伏:用于潜伏音频压缩的一致性自动编码器
作者:Marco Pasini,Stefan Lattner,George Fazekas
备注:Accepted to ISMIR 2024
链接:点击下载PDF文件
摘要:压缩连续潜在空间中的高效音频表示对于生成音频建模和音乐信息检索(MIR)任务至关重要。然而,一些现有的音频自动编码器具有局限性,诸如多级训练过程、缓慢的迭代采样或低重建质量。我们介绍Music2Latent,一个音频自动编码器,它通过利用一致性模型克服了这些限制。Music2Latent在单个端到端训练过程中将样本编码到压缩的连续潜在空间中,同时实现高保真单步重建。关键创新包括通过交叉连接在所有级别上对上采样编码器输出调节一致性模型,使用频率方面的自注意力来捕获长范围频率依赖性,以及采用频率方面的学习缩放来处理不同噪声水平下频率之间的变化值分布。我们证明了Music2Latent在音质和重建精度方面优于现有的连续音频自动编码器,同时使用其潜在表示在下游MIR任务上实现了有竞争力的性能。据我们所知,这是训练端到端一致性自动编码器模型的第一次成功尝试。摘要:Efficient audio representations in a compressed continuous latent space are critical for generative audio modeling and Music Information Retrieval (MIR) tasks. However, some existing audio autoencoders have limitations, such as multi-stage training procedures, slow iterative sampling, or low reconstruction quality. We introduce Music2Latent, an audio autoencoder that overcomes these limitations by leveraging consistency models. Music2Latent encodes samples into a compressed continuous latent space in a single end-to-end training process while enabling high-fidelity single-step reconstruction. Key innovations include conditioning the consistency model on upsampled encoder outputs at all levels through cross connections, using frequency-wise self-attention to capture long-range frequency dependencies, and employing frequency-wise learned scaling to handle varying value distributions across frequencies at different noise levels. We demonstrate that Music2Latent outperforms existing continuous audio autoencoders in sound quality and reconstruction accuracy while achieving competitive performance on downstream MIR tasks using its latent representations. To our knowledge, this represents the first successful attempt at training an end-to-end consistency autoencoder model.

【16】 TOGGL: Transcribing Overlapping Speech with Staggered Labeling
标题: TOGGL:用错开标签转录重叠的语音
作者:Chak-Fai Li,William Hartmann,Matthew Snover
备注:5 pages
链接:点击下载PDF文件
摘要:转录多个重叠说话者的语音通常需要将音频分离成多个流并独立地识别每个流。最近的工作联合分离和转录,但需要一个单独的解码组件为每个扬声器。我们提出了TOGGL模型,同时转录多个扬声器的语音。TOGGL模型使用特殊的输出标记来将语音归属于每个说话者,只有一个解码器。我们的方法可以推广到两个扬声器之外,即使只在两个扬声器数据上训练。我们表现出优越的性能相比,竞争的方法在会话语音数据集。我们的方法还提高了单扬声器音频的性能。摘要:Transcribing the speech of multiple overlapping speakers typically requires separating the audio into multiple streams and recognizing each one independently. More recent work jointly separates and transcribes, but requires a separate decoding component for each speaker. We propose the TOGGL model to simultaneously transcribe the speech of multiple speakers. The TOGGL model uses special output tokens to attribute the speech to each speaker with only a single decoder. Our approach generalizes beyond two speakers, even when trained only on two-speaker data. We demonstrate superior performance compared to competing approaches on a conversational speech dataset. Our approach also improves performance on single-speaker audio.

【17】 FoVNet: Configurable Field-of-View Speech Enhancement with Low Computation and Distortion for Smart Glasses
标题: FoVNet:可配置的视野语音增强,具有低计算和低失真的智能眼镜
作者:Zhongweiyang Xu,Ali Aroudi,Ke Tan,Ashutosh Pandey,Jung-Suk Lee,Buye Xu,Francesco Nesta
备注:Accepted by INTERSPEECH2024
链接:点击下载PDF文件
摘要:本文提出了一种新的多通道语音增强方法,FoVNet,它可以在智能眼镜用户的可配置视野(FoV)内实现高效的语音增强,而不需要特定的目标说话者方向。它通过增强任何给定FoV内的所有扬声器来改进先前的工作,采用混合信号处理和深度学习方法,设计具有高计算效率。神经网络组件的设计具有超低的计算量(约50 MMACS)。多通道维纳滤波器和后处理模块进一步用于改善感知质量。我们评估我们的算法与智能眼镜上的麦克风阵列,提供了一个可配置的,高效的解决方案,增强听力的能源受限的设备。FoVNet在多个场景中的计算效率和语音质量都很出色,使其成为智能眼镜应用的一个有前途的解决方案。摘要:This paper presents a novel multi-channel speech enhancement approach, FoVNet, that enables highly efficient speech enhancement within a configurable field of view (FoV) of a smart-glasses user without needing specific target-talker(s) directions. It advances over prior works by enhancing all speakers within any given FoV, with a hybrid signal processing and deep learning approach designed with high computational efficiency. The neural network component is designed with ultra-low computation (about 50 MMACS). A multi-channel Wiener filter and a post-processing module are further used to improve perceptual quality. We evaluate our algorithm with a microphone array on smart glasses, providing a configurable, efficient solution for augmented hearing on energy-constrained devices. FoVNet excels in both computational efficiency and speech quality across multiple scenarios, making it a promising solution for smart glasses applications.

【18】 Dilated Convolution with Learnable Spacings
标题: 可学习间隔的扩张卷积
作者:Ismail Khalfaoui-Hassani
备注:PhD Thesis
链接:点击下载PDF文件
摘要:本文提出并评价了具有可学习间隔的扩张卷积(DCLS)方法。通过在计算机视觉、音频和语音处理领域的各种监督学习实验,DCLS方法被证明优于标准和高级卷积技术。该研究分为几个步骤,首先分析文献和现有的卷积技术之前的DCLS方法的发展。我们特别感兴趣的是与我们自己的方法密切相关的方法,这些方法对于捕捉我们方法的细微差别和独特性仍然至关重要。我们研究的基石是将DCLS方法引入卷积神经网络(CNN),以及依赖卷积和视觉注意力方法的混合架构。DCLS被证明是特别有效的任务,如分类,语义分割和对象检测。最初使用双线性插值,该研究还探索了其他插值方法,发现高斯插值略微提高了性能。DCLS方法进一步应用于尖峰神经网络(SNN),以实现神经网络内的突触延迟学习,最终可以转移到所谓的神经形态芯片。结果表明,DCLS方法脱颖而出,作为一个新的国家的最先进的技术在SNN音频分类在这个领域的某些基准任务。这些任务涉及具有高时间分量的数据集。此外,我们表明,DCLS可以显着提高人工神经网络的多标签音频分类任务的准确性。最后,我们讨论了所选择的实验装置,其局限性,我们的方法的局限性,我们的结果。摘要:This thesis presents and evaluates the Dilated Convolution with Learnable Spacings (DCLS) method. Through various supervised learning experiments in the fields of computer vision, audio, and speech processing, the DCLS method proves to outperform both standard and advanced convolution techniques. The research is organized into several steps, starting with an analysis of the literature and existing convolution techniques that preceded the development of the DCLS method. We were particularly interested in the methods that are closely related to our own and that remain essential to capture the nuances and uniqueness of our approach. The cornerstone of our study is the introduction and application of the DCLS method to convolutional neural networks (CNNs), as well as to hybrid architectures that rely on both convolutional and visual attention approaches. DCLS is shown to be particularly effective in tasks such as classification, semantic segmentation, and object detection. Initially using bilinear interpolation, the study also explores other interpolation methods, finding that Gaussian interpolation slightly improves performance. The DCLS method is further applied to spiking neural networks (SNNs) to enable synaptic delay learning within a neural network that could eventually be transferred to so-called neuromorphic chips. The results show that the DCLS method stands out as a new state-of-the-art technique in SNN audio classification for certain benchmark tasks in this field. These tasks involve datasets with a high temporal component. In addition, we show that DCLS can significantly improve the accuracy of artificial neural networks for the multi-label audio classification task. We conclude with a discussion of the chosen experimental setup, its limitations, the limitations of our method, and our results.


机器翻译,仅供参考