今日论文合集:cs.SD语音13篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】 Improving Multimodal Learning with Multi-Loss Gradient Modulation

标题: 利用多损失梯度调制改善多模式学习

链接:https://arxiv.org/abs/2405.07930

作者:Konstantinos Kontras,Christos Chatzichristos,Matthew Blaschko,Maarten De Vos
摘要:从音频和视频等多种模态中学习,为利用互补信息、增强鲁棒性以及改善上下文理解和性能提供了机会。然而,组合这些模式带来了挑战,特别是当模式在数据结构,预测贡献和学习过程的复杂性方面不同时。已经观察到,一种模态可以潜在地主导学习过程,阻碍来自其他模态的信息的有效利用,并导致次优的模型性能。为了解决这个问题,绝大多数以前的工作建议,以评估单峰的贡献,并动态调整训练,以均衡它们。我们通过引入多损失目标并进一步完善平衡过程来改进以前的工作,使其能够在两个方向上动态调整每个模态的学习速度,加速和减速,并能够在收敛时逐步消除平衡效应。我们在三个音视频数据集上实现了卓越的结果:在CREMA-D上,使用ResNet骨干编码器的模型比之前的最佳模型提高了1.9%到12.4%,而Conformer骨干模型在不同的融合方法中提供了2.8%到14.1%的改进。在AVE上,改进范围从2.7%到7.7%,而在UCF 101上,增益高达6.1%。
摘要:Learning from multiple modalities, such as audio and video, offers opportunities for leveraging complementary information, enhancing robustness, and improving contextual understanding and performance. However, combining such modalities presents challenges, especially when modalities differ in data structure, predictive contribution, and the complexity of their learning processes. It has been observed that one modality can potentially dominate the learning process, hindering the effective utilization of information from other modalities and leading to sub-optimal model performance. To address this issue the vast majority of previous works suggest to assess the unimodal contributions and dynamically adjust the training to equalize them. We improve upon previous work by introducing a multi-loss objective and further refining the balancing process, allowing it to dynamically adjust the learning pace of each modality in both directions, acceleration and deceleration, with the ability to phase out balancing effects upon convergence. We achieve superior results across three audio-video datasets: on CREMA-D, models with ResNet backbone encoders surpass the previous best by 1.9% to 12.4%, and Conformer backbone models deliver improvements ranging from 2.8% to 14.1% across different fusion methods. On AVE, improvements range from 2.7% to 7.7%, while on UCF101, gains reach up to 6.1%.


【2】 Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech

标题: 儿童引导言语的对象相关分析和随机生成

链接:https://arxiv.org/abs/2405.07700

作者:Okko Räsänen,Daniil Kocharov
备注:Accepted for publication in Proc. 45th Annual Meeting of the Cognitive Science Society (CogSci-2024)
摘要:儿童导向言语(英语:Child-directed speech,简称CDS)是成年人在对幼儿说话时使用的一种特殊类型的言语。它的属性也会随着语言外因素的变化而变化,例如被称呼儿童的年龄。访问大量的代表性和不同的CDS将是有益的儿童语言研究,因为这将使控制计算建模实验的婴儿语言习得与现实的输入质量和数量。在这项研究中,我们描述了一种方法来模拟年龄相关的语言特性的CDS使用的语言模型(LM)训练的CDS成绩单和年龄的收件人儿童,从北美英语语料库的CHILDES数据库。然后,创建的LM可以用于以适合年龄的方式随机生成合成CDS转录物,从而在大小上扩展到原始数据集之外。我们比较了所生成的CDS对不同年龄儿童的真实语音的特性,表明LM能够捕获CDS中随年龄变化的变化,除了有效词汇量略有差异。作为一个副产品,我们还提供了一个系统的表征的年龄相关的语言特性的CDS在儿童,说明如何所有测量方面的CDS随儿童的年龄变化。
摘要:Child-directed speech (CDS) is a particular type of speech that adults use when addressing young children. Its properties also change as a function of extralinguistic factors, such as age of the child being addressed. Access to large amounts of representative and varied CDS would be useful for child language research, as this would enable controlled computational modeling experiments of infant language acquisition with realistic input in terms of quality and quantity. In this study, we describe an approach to model age-dependent linguistic properties of CDS using a language model (LM) trained on CDS transcripts and ages of the recipient children, as obtained from North American English corpora of the CHILDES database. The created LM can then be used to stochastically generate synthetic CDS transcripts in an age-appropriate manner, thereby scaling beyond the original datasets in size. We compare characteristics of the generated CDS against the real speech addressed at children of different ages, showing that the LM manages to capture age-dependent changes in CDS, except for a slight difference in the effective vocabulary size. As a side product, we also provide a systematic characterization of age-dependent linguistic properties of CDS in CHILDES, illustrating how all measured aspects of the CDS change with children's age.


【3】 FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

标题: FastSAG:迈向快速非自回归歌唱伴奏一代

链接:https://arxiv.org/abs/2405.07682

作者:Jianyi Chen,Wei Xue,Xu Tan,Zhen Ye,Qifeng Liu,Yike Guo
备注:IJCAI 2024
摘要:歌唱伴奏生成(SAG),即生成器乐来伴随输入的人声,对于开发人类-AI共生艺术创作系统至关重要。最先进的方法,SingSong,利用多阶段自回归(AR)模型的SAG,然而,这种方法是非常缓慢的,因为它产生的语义和声学令牌递归,这使得它不可能的实时应用。在本文中,我们的目标是开发一种快速SAG方法,可以创建高质量和连贯的图形。提出了一种基于非AR扩散的语音识别框架,通过精心设计从人声信号中推断出的条件,直接生成目标伴奏的Mel谱图。该方法通过扩散和Mel谱图建模,大大简化了基于AR标记的SingSong框架,并大大加快了生成速度。我们还设计了语义投影、先验投影块以及一组损失函数,以确保生成的伴奏与人声信号具有语义和节奏的一致性。通过大量的实验研究,我们证明了所提出的方法可以生成比SingSong更好的样本,并且将生成速度提高了至少30倍。音频样本和代码可在https://fastsag.github.io/上获得。
摘要:Singing Accompaniment Generation (SAG), which generates instrumental music to accompany input vocals, is crucial to developing human-AI symbiotic art creation systems. The state-of-the-art method, SingSong, utilizes a multi-stage autoregressive (AR) model for SAG, however, this method is extremely slow as it generates semantic and acoustic tokens recursively, and this makes it impossible for real-time applications. In this paper, we aim to develop a Fast SAG method that can create high-quality and coherent accompaniments. A non-AR diffusion-based framework is developed, which by carefully designing the conditions inferred from the vocal signals, generates the Mel spectrogram of the target accompaniment directly. With diffusion and Mel spectrogram modeling, the proposed method significantly simplifies the AR token-based SingSong framework, and largely accelerates the generation. We also design semantic projection, prior projection blocks as well as a set of loss functions, to ensure the generated accompaniment has semantic and rhythm coherence with the vocal signal. By intensive experimental studies, we demonstrate that the proposed method can generate better samples than SingSong, and accelerate the generation by at least 30 times. Audio samples and code are available at https://fastsag.github.io/.


【4】 Rene: A Pre-trained Multi-modal Architecture for Auscultation of Respiratory Diseases

标题: Rene:用于呼吸道疾病听诊的预训练多模式架构

链接:https://arxiv.org/abs/2405.07442

作者:Pengfei Zhang,Zhihang Zheng,Shichen Zhang,Minghao Yang,Shaojun Tang
摘要:This study presents a novel methodology utilizing a pre-trained speech recognition model for processing respiratory sound data. By incorporating medical record information, we introduce an innovative multi-modal deep-learning architecture, named Rene, which addresses the challenges of poor interpretability and underperformance in real-time clinical diagnostic response observed in previous respiratory disease-focused models. The proposed Rene architecture demonstrated significant improvements of 10.24%, 16.15%, 15.29%, and 18.90% respectively, compared to the baseline across four tasks related to respiratory event detection and audio record classification on the SPRSound database. In patient disease prediction tests on the ICBHI database, the architecture exhibited improvements of 23% in the mean of average score and harmonic score compared to the baseline. Furthermore, we developed a real-time respiratory sound discrimination system based on the Rene architecture, featuring a dual-thread design and compressed model parameters for simultaneous microphone recording and real-time dynamic decoding. Employing state-of-the-art Edge AI technology, this system enables rapid and accurate responses for respiratory sound auscultation, facilitating deployment on wearable clinical detection devices to capture incremental data, which can be synergistically evolved with large-scale models deployed on cloud servers for downstream tasks.


【5】 SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset

标题: SoccerNet-Echoes:足球比赛音频评论数据集

链接:https://arxiv.org/abs/2405.07354

作者:Sushant Gautam,Mehdi Houshmand Sarkhoosh,Jan Held,Cise Midoglu,Anthony Cioppa,Silvio Giancola,Vajira Thambawita,Michael A. Riegler,Pål Halvorsen,Mubarak Shah
摘要:自动语音识别(ASR)技术在足球中的应用为体育分析提供了许多机会。具体来说,使用ASR提取音频评论提供了对比赛事件的有价值的见解,并为自动高光生成等几个下游应用打开了大门。本文介绍了SoccerNet—Echoes,这是SoccerNet数据集的一个增强,它具有自动生成的足球比赛广播音频评论的传输,使用ASR从比赛音频中获得的丰富的文本信息层增强视频内容。这些使用Whisper模型生成并使用Google翻译翻译的文本评论扩展了SoccerNet数据集在各种应用程序中的有用性,例如增强的动作定位,自动字幕生成和游戏摘要。通过将文本数据与视觉和听觉内容相结合,SoccerNet—Echoes旨在成为专门用于捕获足球比赛动态的算法开发的综合资源。我们详细介绍了该数据集的管理和ASR集成所涉及的方法。我们还强调了多模态方法在体育分析中的意义,以及丰富的数据集如何支持不同的应用程序,从而扩大了体育分析领域的研究和开发范围。
摘要:The application of Automatic Speech Recognition (ASR) technology in soccer offers numerous opportunities for sports analytics. Specifically, extracting audio commentaries with ASR provides valuable insights into the events of the game, and opens the door to several downstream applications such as automatic highlight generation. This paper presents SoccerNet-Echoes, an augmentation of the SoccerNet dataset with automatically generated transcriptions of audio commentaries from soccer game broadcasts, enhancing video content with rich layers of textual information derived from the game audio using ASR. These textual commentaries, generated using the Whisper model and translated with Google Translate, extend the usefulness of the SoccerNet dataset in diverse applications such as enhanced action spotting, automatic caption generation, and game summarization. By incorporating textual data alongside visual and auditory content, SoccerNet-Echoes aims to serve as a comprehensive resource for the development of algorithms specialized in capturing the dynamics of soccer games. We detail the methods involved in the curation of this dataset and the integration of ASR. We also highlight the implications of a multimodal approach in sports analytics, and how the enriched dataset can support diverse applications, thus broadening the scope of research and development in the field of sports analytics.


【6】 Unified Video-Language Pre-training with Synchronized Audio

标题: 同步音频的统一视频语言预训练

链接:https://arxiv.org/abs/2405.07202

作者:Shentong Mo,Haofan Wang,Huaxia Li,Xu Tang
摘要:视频语言预训练是一个典型且具有挑战性的问题,旨在以自我监督的方式从大规模数据中学习视觉和文本表示。现有的预训练方法要么捕获图像—文本对的对应关系,要么利用帧的时间排序。然而,他们没有明确地探索音频和其他两种模态之间的自然同步。在这项工作中,我们提出了一个增强的框架,视频语言预训练与同步音频,称为VLSA,可以学习三模态表示在一个统一的自我监督的Transformer。具体来说,我们的VLSA联合聚合了视频、文本和音频的本地补丁和全局令牌的嵌入。此外,我们利用局部补丁掩蔽建模来学习模态感知特征,并利用全局音频匹配来捕获视频和文本的音频引导特征。我们进行了广泛的实验检索文本,视频和音频。我们的简单模型仅在0.9M数据上进行了预训练,实现了与最先进的基线相比的改进结果。此外,定性可视化生动地展示了我们的VLSA在学习有区别的视觉—文本表示方面的优势。
摘要:Video-language pre-training is a typical and challenging problem that aims at learning visual and textual representations from large-scale data in a self-supervised way. Existing pre-training approaches either captured the correspondence of image-text pairs or utilized temporal ordering of frames. However, they do not explicitly explore the natural synchronization between audio and the other two modalities. In this work, we propose an enhanced framework for Video-Language pre-training with Synchronized Audio, termed as VLSA, that can learn tri-modal representations in a unified self-supervised transformer. Specifically, our VLSA jointly aggregates embeddings of local patches and global tokens for video, text, and audio. Furthermore, we utilize local-patch masked modeling to learn modality-aware features, and leverage global audio matching to capture audio-guided features for video and text. We conduct extensive experiments on retrieval across text, video, and audio. Our simple model pre-trained on only 0.9M data achieves improving results against state-of-the-art baselines. In addition, qualitative visualizations vividly showcase the superiority of our VLSA in learning discriminative visual-textual representations.


【7】 Towards an Accessible and Rapidly Trainable Rhythm Sequencer Using a Generative Stacked Autoencoder

标题: 使用生成式堆叠自动编码器打造易于访问且可快速训练的节奏排序器

链接:https://arxiv.org/abs/2405.07034

作者:Alex Wastnidge
备注:7 pages, 7 figures
摘要:Neural networks and deep learning are often deployed for the sake of the most comprehensive music generation with as little involvement as possible from the human musician. Implementations in aid of, or being a tool for, music practitioners are sparse. This paper proposes the integration of generative stacked autoencoder structures for rhythm generation, within a conventional melodic step-sequencer. It further aims to work towards its implementation being accessible to the average electronic music practitioner. Several model architectures have been trained and tested for their creative potential. While the currently implementations do display limitations, they do represent viable creative solutions for music practitioners.


【8】 A framework of text-dependent speaker verification for chinese numerical string corpus

标题: 中文数字串库的文本相关说话人验证框架

链接:https://arxiv.org/abs/2405.07029

作者:Litong Zheng,Feng Hong,Weijie Xu,Wan Zheng
备注:arXiv admin note: text overlap with arXiv:2312.01645
摘要:The Chinese numerical string corpus, serves as a valuable resource for speaker verification, particularly in financial transactions. Researches indicate that in short speech scenarios, text-dependent speaker verification (TD-SV) consistently outperforms text-independent speaker verification (TI-SV). However, TD-SV potentially includes the validation of text information, that can be negatively impacted by reading rhythms and pauses. To address this problem, we propose an end-to-end speaker verification system that enhances TD-SV by decoupling speaker and text information. Our system consists of a text embedding extractor, a speaker embedding extractor and a fusion module. In the text embedding extractor, we employ an enhanced Transformer and introduce a triple loss including text classification loss, connectionist temporal classification (CTC) loss and decoder loss; while in the speaker embedding extractor, we create a multi-scale pooling method by combining sliding window attentive statistics pooling (SWASP) with attentive statistics pooling (ASP). To mitigate the scarcity of data, we have recorded a publicly available Chinese numerical corpus named SHALCAS22A (hereinafter called SHAL), which can be accessed on Open-SLR. Moreover, we employ data augmentation techniques using Tacotron2 and HiFi-GAN. Our method achieves an equal error rate (EER) performance improvement of 49.2% on Hi-Mia and 75.0% on SHAL, respectively.


【9】 Benchmarking Cross-Domain Audio-Visual Deception Detection

标题: 跨域视听欺骗检测基准

链接:https://arxiv.org/abs/2405.06995

作者:Xiaobao Guo,Zitong Yu,Nithish Muthuchamy Selvaraj,Bingquan Shen,Adams Wai-Kin Kong,Alex C. Kot
备注:10 pages
摘要:自动欺骗检测对于帮助人类准确评估真实性和识别欺骗行为至关重要。传统的基于接触的技术,如测谎仪,依赖于生理信号来确定个人陈述的真实性。然而,最近的发展,在自动欺骗检测已经证明,多模态的功能来自音频和视频模态可能会优于人类观察员公开可用的数据集。尽管有这些积极的发现,现有的视听欺骗检测方法在不同情况下的普遍性仍然在很大程度上未被探索。为了缩小这一差距,我们提出了第一个跨域视听欺骗检测基准,使我们能够评估这些方法在现实世界中的推广使用情况。我们使用广泛采用的音频和视觉特征以及不同的架构进行基准测试,比较单对单和多对单域泛化性能。为了进一步利用来自多个源域的数据进行训练的影响,我们研究了三种类型的域采样策略,包括域同步,域交替和逐域进行多到单域泛化评估。此外,我们提出了Attention—Mixer融合方法,以提高性能,我们相信,这种新的跨域基准将有助于未来的研究在视听欺骗检测。协议和源代码可在\href {https://github.com/Redaimao/cross_domain_DD}{https://github.com/Redaimao/cross\_domain\_DD}获得。
摘要:Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine the authenticity of an individual's statements. Nevertheless, recent developments in automated deception detection have demonstrated that multimodal features derived from both audio and video modalities may outperform human observers on publicly available datasets. Despite these positive findings, the generalizability of existing audio-visual deception detection approaches across different scenarios remains largely unexplored. To close this gap, we present the first cross-domain audio-visual deception detection benchmark, that enables us to assess how well these methods generalize for use in real-world scenarios. We used widely adopted audio and visual features and different architectures for benchmarking, comparing single-to-single and multi-to-single domain generalization performance. To further exploit the impacts using data from multiple source domains for training, we investigate three types of domain sampling strategies, including domain-simultaneous, domain-alternating, and domain-by-domain for multi-to-single domain generalization evaluation. Furthermore, we proposed the Attention-Mixer fusion method to improve performance, and we believe that this new cross-domain benchmark will facilitate future research in audio-visual deception detection. Protocols and source code are available at \href{https://github.com/Redaimao/cross_domain_DD}{https://github.com/Redaimao/cross\_domain\_DD}.


【10】 Time-of-arrival Estimation and Phase Unwrapping of Head-related Transfer Functions With Integer Linear Programming

标题: 基于ARCH线性规划的头部相关传递函数到达时间估计和阶段展开

链接:https://arxiv.org/abs/2405.06804

作者:Chin-Yun Yu,Johan Pauwels,György Fazekas
备注:Accepted to be presented at Audio Engineering Society 156th Convention, 2024 June, Madrid, Spain
摘要:In binaural audio synthesis, aligning head-related impulse responses (HRIRs) in time has been an important pre-processing step, enabling accurate spatial interpolation and efficient data compression. The maximum correlation time delay between spatially nearby HRIRs has previously been used to get accurate and smooth alignment by solving a matrix equation in which the solution has the minimum Euclidean distance to the time delay. However, the Euclidean criterion could lead to an over-smoothing solution in practice. In this paper, we solve the smoothing issue by formulating the task as solving an integer linear programming problem equivalent to minimising an $L^1$-norm. Moreover, we incorporate 1) the cross-correlation of inter-aural HRIRs, and 2) HRIRs with their minimum-phase responses to have more reference measurements for optimisation. We show the proposed method can get more accurate alignments than the Euclidean-based method by comparing the spectral reconstruction loss of time-aligned HRIRs using spherical harmonics representation on seven HRIRs consisting of human and dummy heads. The extra correlation features and the $L^1$-norm are also beneficial in extremely noisy conditions. In addition, this method can be applied to phase unwrapping of head-related transfer functions, where the unwrapped phase could be a compact feature for downstream tasks.


【11】 Music Emotion Prediction Using Recurrent Neural Networks

标题: 使用回归神经网络的音乐情感预测

链接:https://arxiv.org/abs/2405.06747

作者:Xinyu Chang,Xiangyu Zhang,Haoruo Zhang,Yulu Ran
备注:15 pages, 13 figures
摘要:本研究探讨了递归神经网络在识别音乐中传达的情感方面的应用,旨在通过定制音乐以适应听众的情绪状态来增强音乐推荐系统并支持治疗干预。我们利用罗素的情感象限将音乐分为四个不同的情感区域,并开发能够准确预测这些类别的模型。我们的方法涉及使用Librosa提取一组全面的音频特征,并应用各种递归神经网络架构,包括标准RNN,双向RNN和长短期记忆(LSTM)网络。使用900个音频片段的数据集进行初始实验,根据情感象限进行标记。我们将我们的神经网络模型的性能与一组基线分类器进行比较,并分析它们在捕捉音乐表达中固有的时间动态方面的有效性。结果表明,更简单的RNN架构可以执行更复杂的模型,特别是在较小的数据集上。我们还在更大的数据集上应用了以下实验:一个是基于我们的原始数据集进行增强的,另一个是来自其他来源的。这项研究不仅增强了我们对音乐情感影响的理解,还展示了神经网络在创建更具个性化和情感共鸣的音乐推荐和治疗系统方面的潜力。
摘要:This study explores the application of recurrent neural networks to recognize emotions conveyed in music, aiming to enhance music recommendation systems and support therapeutic interventions by tailoring music to fit listeners' emotional states. We utilize Russell's Emotion Quadrant to categorize music into four distinct emotional regions and develop models capable of accurately predicting these categories. Our approach involves extracting a comprehensive set of audio features using Librosa and applying various recurrent neural network architectures, including standard RNNs, Bidirectional RNNs, and Long Short-Term Memory (LSTM) networks. Initial experiments are conducted using a dataset of 900 audio clips, labeled according to the emotional quadrants. We compare the performance of our neural network models against a set of baseline classifiers and analyze their effectiveness in capturing the temporal dynamics inherent in musical expression. The results indicate that simpler RNN architectures may perform comparably or even superiorly to more complex models, particularly in smaller datasets. We've also applied the following experiments on larger datasets: one is augmented based on our original dataset, and the other is from other sources. This research not only enhances our understanding of the emotional impact of music but also demonstrates the potential of neural networks in creating more personalized and emotionally resonant music recommendation and therapy systems.


【12】 Evaluating Speech Enhancement Systems Through Listening Effort

标题: 通过倾听努力评估语音增强系统

链接:https://arxiv.org/abs/2405.07641

作者:Femke B. Gelderblom,Tron V. Tronstad,Iván López-Espejo
摘要:Understanding degraded speech is demanding, requiring increased listening effort (LE). Evaluating processed and unprocessed speech with respect to LE can objectively indicate if speech enhancement systems benefit listeners. However, existing methods for measuring LE are complex and not widely applicable. In this study, we propose a simple method to evaluate speech intelligibility and LE simultaneously without additional strain on subjects or operators. We assess this method using results from two independent studies in Norway and Denmark, testing 76 (50+26) subjects across 9 (6+3) processing conditions. Despite differences in evaluation setups, subject recruitment, and processing systems, trends are strikingly similar, demonstrating the proposed method's robustness and ease of implementation into existing practices.


【13】 IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization

标题: IPDnet:一种通用直接路径IPD估计网络,用于声音源定位

链接:https://arxiv.org/abs/2405.07021

作者:Yabo Wang,Bing Yang,Xiaofei Li
摘要:Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound sources from microphone array signals. The estimated DP-IPD can be easily translated to source location based on the known microphone array geometry. First, a full-band and narrow-band fusion network is proposed for DP-IPD estimation, in which alternating narrow-band and full-band layers are responsible for estimating the rough DP-IPD information in one frequency band and capturing the frequency correlations of DP-IPD, respectively. Second, a new multi-track DP-IPD learning target is proposed for the localization of flexible number of sound sources. Third, the IPDnet is extend to handling variable microphone arrays, once trained which is able to process arbitrary microphone arrays with different number of channels and array topology. Experiments of multiple-moving-speaker localization are conducted on both simulated and real-world data, which show that the proposed full-band and narrow-band fusion network and the proposed multi-track DP-IPD learning target together achieves excellent sound source localization performance. Moreover, the proposed variable-array model generalizes well to unseen microphone arrays.


eess.AS音频处理

【1】 Evaluating Speech Enhancement Systems Through Listening Effort

标题: 通过倾听努力评估语音增强系统

链接:https://arxiv.org/abs/2405.07641

作者:Femke B. Gelderblom,Tron V. Tronstad,Iván López-Espejo
摘要:理解退化的语音要求很高,需要增加听力努力(LE)。相对于LE评估处理的和未处理的语音可以客观地指示语音增强系统是否有益于收听者。然而,用于测量LE的现有方法是复杂的并且不广泛适用。在这项研究中,我们提出了一个简单的方法来评估语音清晰度和LE同时没有额外的压力的主题或操作员。我们使用挪威和丹麦的两项独立研究的结果评估了这种方法,在9(6+3)种处理条件下测试了76(50+26)名受试者。尽管在评估设置,受试者招募和处理系统的差异,趋势是惊人的相似,证明了所提出的方法的鲁棒性和易于实施到现有的做法。
摘要:Understanding degraded speech is demanding, requiring increased listening effort (LE). Evaluating processed and unprocessed speech with respect to LE can objectively indicate if speech enhancement systems benefit listeners. However, existing methods for measuring LE are complex and not widely applicable. In this study, we propose a simple method to evaluate speech intelligibility and LE simultaneously without additional strain on subjects or operators. We assess this method using results from two independent studies in Norway and Denmark, testing 76 (50+26) subjects across 9 (6+3) processing conditions. Despite differences in evaluation setups, subject recruitment, and processing systems, trends are strikingly similar, demonstrating the proposed method's robustness and ease of implementation into existing practices.


【2】 IPDnet: A Universal Direct-Path IPD Estimation Network for Sound Source Localization

标题: IPDnet:一种通用直接路径IPD估计网络,用于声音源定位

链接:https://arxiv.org/abs/2405.07021

作者:Yabo Wang,Bing Yang,Xiaofei Li
摘要:直接路径空间特征的提取是恶劣声环境下声源定位的关键。本文提出了IPDnet,一个神经网络,估计直接路径通道间相位差(DP-IPD)的声源从麦克风阵列信号。基于已知的麦克风阵列几何形状,可以容易地将估计的DP-IPD转换为源位置。首先,提出了一种用于DP-IPD估计的全带和窄带融合网络,其中交替的窄带和全带层分别负责估计一个频带中的粗略DP-IPD信息和捕获DP-IPD的频率相关性。其次,提出了一种新的多声道DP-IPD学习目标,用于灵活数量声源的定位。第三,IPDnet扩展到处理可变麦克风阵列,一旦训练,能够处理具有不同通道数和阵列拓扑的任意麦克风阵列。在仿真数据和真实数据上进行了多移动说话人定位实验,实验结果表明,所提出的全频带和窄带融合网络以及多声道DP-IPD学习目标共同实现了良好的声源定位性能。此外,所提出的可变阵列模型很好地推广到看不见的麦克风阵列。
摘要:Extracting direct-path spatial feature is crucial for sound source localization in adverse acoustic environments. This paper proposes the IPDnet, a neural network that estimates direct-path inter-channel phase difference (DP-IPD) of sound sources from microphone array signals. The estimated DP-IPD can be easily translated to source location based on the known microphone array geometry. First, a full-band and narrow-band fusion network is proposed for DP-IPD estimation, in which alternating narrow-band and full-band layers are responsible for estimating the rough DP-IPD information in one frequency band and capturing the frequency correlations of DP-IPD, respectively. Second, a new multi-track DP-IPD learning target is proposed for the localization of flexible number of sound sources. Third, the IPDnet is extend to handling variable microphone arrays, once trained which is able to process arbitrary microphone arrays with different number of channels and array topology. Experiments of multiple-moving-speaker localization are conducted on both simulated and real-world data, which show that the proposed full-band and narrow-band fusion network and the proposed multi-track DP-IPD learning target together achieves excellent sound source localization performance. Moreover, the proposed variable-array model generalizes well to unseen microphone arrays.


【3】 Improving Multimodal Learning with Multi-Loss Gradient Modulation

标题: 利用多损失梯度调制改善多模式学习

链接:https://arxiv.org/abs/2405.07930

作者:Konstantinos Kontras,Christos Chatzichristos,Matthew Blaschko,Maarten De Vos
摘要:从音频和视频等多种模态中学习,为利用互补信息、增强鲁棒性以及改善上下文理解和性能提供了机会。然而,组合这些模式带来了挑战,特别是当模式在数据结构,预测贡献和学习过程的复杂性方面不同时。已经观察到,一种模态可以潜在地主导学习过程,阻碍来自其他模态的信息的有效利用,并导致次优的模型性能。为了解决这个问题,绝大多数以前的工作建议,以评估单峰的贡献,并动态调整训练,以均衡它们。我们通过引入多损失目标并进一步完善平衡过程来改进以前的工作,使其能够在两个方向上动态调整每个模态的学习速度,加速和减速,并能够在收敛时逐步消除平衡效应。我们在三个音视频数据集上实现了卓越的结果:在CREMA-D上,使用ResNet骨干编码器的模型比之前的最佳模型提高了1.9%到12.4%,而Conformer骨干模型在不同的融合方法中提供了2.8%到14.1%的改进。在AVE上,改进范围从2.7%到7.7%,而在UCF 101上,增益高达6.1%。
摘要:Learning from multiple modalities, such as audio and video, offers opportunities for leveraging complementary information, enhancing robustness, and improving contextual understanding and performance. However, combining such modalities presents challenges, especially when modalities differ in data structure, predictive contribution, and the complexity of their learning processes. It has been observed that one modality can potentially dominate the learning process, hindering the effective utilization of information from other modalities and leading to sub-optimal model performance. To address this issue the vast majority of previous works suggest to assess the unimodal contributions and dynamically adjust the training to equalize them. We improve upon previous work by introducing a multi-loss objective and further refining the balancing process, allowing it to dynamically adjust the learning pace of each modality in both directions, acceleration and deceleration, with the ability to phase out balancing effects upon convergence. We achieve superior results across three audio-video datasets: on CREMA-D, models with ResNet backbone encoders surpass the previous best by 1.9% to 12.4%, and Conformer backbone models deliver improvements ranging from 2.8% to 14.1% across different fusion methods. On AVE, improvements range from 2.7% to 7.7%, while on UCF101, gains reach up to 6.1%.


【4】 Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech

标题: 儿童引导言语的对象相关分析和随机生成

链接:https://arxiv.org/abs/2405.07700

作者:Okko Räsänen,Daniil Kocharov
备注:Accepted for publication in Proc. 45th Annual Meeting of the Cognitive Science Society (CogSci-2024)
摘要:Child-directed speech (CDS) is a particular type of speech that adults use when addressing young children. Its properties also change as a function of extralinguistic factors, such as age of the child being addressed. Access to large amounts of representative and varied CDS would be useful for child language research, as this would enable controlled computational modeling experiments of infant language acquisition with realistic input in terms of quality and quantity. In this study, we describe an approach to model age-dependent linguistic properties of CDS using a language model (LM) trained on CDS transcripts and ages of the recipient children, as obtained from North American English corpora of the CHILDES database. The created LM can then be used to stochastically generate synthetic CDS transcripts in an age-appropriate manner, thereby scaling beyond the original datasets in size. We compare characteristics of the generated CDS against the real speech addressed at children of different ages, showing that the LM manages to capture age-dependent changes in CDS, except for a slight difference in the effective vocabulary size. As a side product, we also provide a systematic characterization of age-dependent linguistic properties of CDS in CHILDES, illustrating how all measured aspects of the CDS change with children's age.


【5】 FastSAG: Towards Fast Non-Autoregressive Singing Accompaniment Generation

标题: FastSAG:迈向快速非自回归歌唱伴奏一代

链接:https://arxiv.org/abs/2405.07682

作者:Jianyi Chen,Wei Xue,Xu Tan,Zhen Ye,Qifeng Liu,Yike Guo
备注:IJCAI 2024
摘要:歌唱伴奏生成(SAG),即生成器乐来伴随输入的人声,对于开发人类—AI共生艺术创作系统至关重要。最先进的方法,SingSong,利用多阶段自回归(AR)模型的SAG,然而,这种方法是非常缓慢的,因为它产生的语义和声学令牌递归,这使得它不可能的实时应用。在本文中,我们的目标是开发一种快速SAG方法,可以创建高质量和连贯的图形。提出了一种基于非AR扩散的语音识别框架,通过精心设计从人声信号中推断出的条件,直接生成目标伴奏的Mel谱图。该方法通过扩散和Mel谱图建模,大大简化了基于AR标记的SingSong框架,并大大加快了生成速度。我们还设计了语义投影、先验投影块以及一组损失函数,以确保生成的伴奏与人声信号具有语义和节奏的一致性。通过大量的实验研究,我们证明了所提出的方法可以生成比SingSong更好的样本,并且将生成速度提高了至少30倍。音频样本和代码可在www.example.com上获得。
摘要:Singing Accompaniment Generation (SAG), which generates instrumental music to accompany input vocals, is crucial to developing human-AI symbiotic art creation systems. The state-of-the-art method, SingSong, utilizes a multi-stage autoregressive (AR) model for SAG, however, this method is extremely slow as it generates semantic and acoustic tokens recursively, and this makes it impossible for real-time applications. In this paper, we aim to develop a Fast SAG method that can create high-quality and coherent accompaniments. A non-AR diffusion-based framework is developed, which by carefully designing the conditions inferred from the vocal signals, generates the Mel spectrogram of the target accompaniment directly. With diffusion and Mel spectrogram modeling, the proposed method significantly simplifies the AR token-based SingSong framework, and largely accelerates the generation. We also design semantic projection, prior projection blocks as well as a set of loss functions, to ensure the generated accompaniment has semantic and rhythm coherence with the vocal signal. By intensive experimental studies, we demonstrate that the proposed method can generate better samples than SingSong, and accelerate the generation by at least 30 times. Audio samples and code are available at https://fastsag.github.io/.


【6】 Rene: A Pre-trained Multi-modal Architecture for Auscultation of Respiratory Diseases

标题: Rene:用于呼吸道疾病听诊的预训练多模式架构

链接:https://arxiv.org/abs/2405.07442

作者:Pengfei Zhang,Zhihang Zheng,Shichen Zhang,Minghao Yang,Shaojun Tang
摘要:这项研究提出了一种新的方法,利用预先训练的语音识别模型处理呼吸声数据。通过整合医疗记录信息,我们引入了一种创新的多模式深度学习架构Rene,它解决了在以前的呼吸系统疾病模型中观察到的实时临床诊断响应的可解释性差和性能不佳的挑战。与SPRSound数据库上与呼吸事件检测和音频记录分类相关的四个任务的基线相比,所提出的Rene架构分别表现出10.24%、16.15%、15.29%和18.90%的显著改进。在ICBHI数据库的患者疾病预测测试中,与基线相比,该架构的平均评分和谐波评分平均值提高了23%。此外,我们开发了一个实时的呼吸声识别系统的基础上Rene架构,具有双线程设计和压缩模型参数,同时麦克风记录和实时动态解码。该系统采用最先进的Edge AI技术,能够快速准确地响应呼吸音听诊,便于部署在可穿戴临床检测设备上以捕获增量数据,这些数据可以与部署在云服务器上的大规模模型协同演进,用于下游任务。
摘要:This study presents a novel methodology utilizing a pre-trained speech recognition model for processing respiratory sound data. By incorporating medical record information, we introduce an innovative multi-modal deep-learning architecture, named Rene, which addresses the challenges of poor interpretability and underperformance in real-time clinical diagnostic response observed in previous respiratory disease-focused models. The proposed Rene architecture demonstrated significant improvements of 10.24%, 16.15%, 15.29%, and 18.90% respectively, compared to the baseline across four tasks related to respiratory event detection and audio record classification on the SPRSound database. In patient disease prediction tests on the ICBHI database, the architecture exhibited improvements of 23% in the mean of average score and harmonic score compared to the baseline. Furthermore, we developed a real-time respiratory sound discrimination system based on the Rene architecture, featuring a dual-thread design and compressed model parameters for simultaneous microphone recording and real-time dynamic decoding. Employing state-of-the-art Edge AI technology, this system enables rapid and accurate responses for respiratory sound auscultation, facilitating deployment on wearable clinical detection devices to capture incremental data, which can be synergistically evolved with large-scale models deployed on cloud servers for downstream tasks.


【7】 SoccerNet-Echoes: A Soccer Game Audio Commentary Dataset

标题: SoccerNet-Echoes:足球比赛音频评论数据集

链接:https://arxiv.org/abs/2405.07354

作者:Sushant Gautam,Mehdi Houshmand Sarkhoosh,Jan Held,Cise Midoglu,Anthony Cioppa,Silvio Giancola,Vajira Thambawita,Michael A. Riegler,Pål Halvorsen,Mubarak Shah
摘要:The application of Automatic Speech Recognition (ASR) technology in soccer offers numerous opportunities for sports analytics. Specifically, extracting audio commentaries with ASR provides valuable insights into the events of the game, and opens the door to several downstream applications such as automatic highlight generation. This paper presents SoccerNet-Echoes, an augmentation of the SoccerNet dataset with automatically generated transcriptions of audio commentaries from soccer game broadcasts, enhancing video content with rich layers of textual information derived from the game audio using ASR. These textual commentaries, generated using the Whisper model and translated with Google Translate, extend the usefulness of the SoccerNet dataset in diverse applications such as enhanced action spotting, automatic caption generation, and game summarization. By incorporating textual data alongside visual and auditory content, SoccerNet-Echoes aims to serve as a comprehensive resource for the development of algorithms specialized in capturing the dynamics of soccer games. We detail the methods involved in the curation of this dataset and the integration of ASR. We also highlight the implications of a multimodal approach in sports analytics, and how the enriched dataset can support diverse applications, thus broadening the scope of research and development in the field of sports analytics.


【8】 Unified Video-Language Pre-training with Synchronized Audio

标题: 同步音频的统一视频语言预训练

链接:https://arxiv.org/abs/2405.07202

作者:Shentong Mo,Haofan Wang,Huaxia Li,Xu Tang
摘要:视频语言预训练是一个典型且具有挑战性的问题,旨在以自我监督的方式从大规模数据中学习视觉和文本表示。现有的预训练方法要么捕获图像—文本对的对应关系,要么利用帧的时间排序。然而,他们没有明确地探索音频和其他两种模态之间的自然同步。在这项工作中,我们提出了一个增强的框架,视频语言预训练与同步音频,称为VLSA,可以学习三模态表示在一个统一的自我监督的Transformer。具体来说,我们的VLSA联合聚合了视频、文本和音频的本地补丁和全局令牌的嵌入。此外,我们利用局部补丁掩蔽建模来学习模态感知特征,并利用全局音频匹配来捕获视频和文本的音频引导特征。我们进行了广泛的实验检索文本,视频和音频。我们的简单模型仅在0.9M数据上进行了预训练,实现了与最先进的基线相比的改进结果。此外,定性可视化生动地展示了我们的VLSA在学习有区别的视觉—文本表示方面的优势。
摘要:Video-language pre-training is a typical and challenging problem that aims at learning visual and textual representations from large-scale data in a self-supervised way. Existing pre-training approaches either captured the correspondence of image-text pairs or utilized temporal ordering of frames. However, they do not explicitly explore the natural synchronization between audio and the other two modalities. In this work, we propose an enhanced framework for Video-Language pre-training with Synchronized Audio, termed as VLSA, that can learn tri-modal representations in a unified self-supervised transformer. Specifically, our VLSA jointly aggregates embeddings of local patches and global tokens for video, text, and audio. Furthermore, we utilize local-patch masked modeling to learn modality-aware features, and leverage global audio matching to capture audio-guided features for video and text. We conduct extensive experiments on retrieval across text, video, and audio. Our simple model pre-trained on only 0.9M data achieves improving results against state-of-the-art baselines. In addition, qualitative visualizations vividly showcase the superiority of our VLSA in learning discriminative visual-textual representations.


【9】 Towards an Accessible and Rapidly Trainable Rhythm Sequencer Using a Generative Stacked Autoencoder

标题: 使用生成式堆叠自动编码器打造易于访问且可快速训练的节奏排序器

链接:https://arxiv.org/abs/2405.07034

作者:Alex Wastnidge
备注:7 pages, 7 figures
摘要:神经网络和深度学习通常是为了最全面的音乐生成而部署的,尽可能少地涉及人类音乐家。帮助音乐从业者或作为音乐从业者工具的实现很少。本文提出了一个传统的旋律步进音序器的节奏生成生成堆叠自动编码器结构的集成。它还旨在努力使其实现可供普通电子音乐从业者使用。已经训练和测试了几个模型架构的创造潜力。虽然目前的实现确实显示了限制,但它们确实代表了音乐从业者的可行的创造性解决方案。
摘要:Neural networks and deep learning are often deployed for the sake of the most comprehensive music generation with as little involvement as possible from the human musician. Implementations in aid of, or being a tool for, music practitioners are sparse. This paper proposes the integration of generative stacked autoencoder structures for rhythm generation, within a conventional melodic step-sequencer. It further aims to work towards its implementation being accessible to the average electronic music practitioner. Several model architectures have been trained and tested for their creative potential. While the currently implementations do display limitations, they do represent viable creative solutions for music practitioners.


【10】 A framework of text-dependent speaker verification for chinese numerical string corpus

标题: 中文数字串库的文本相关说话人验证框架

链接:https://arxiv.org/abs/2405.07029

作者:Litong Zheng,Feng Hong,Weijie Xu,Wan Zheng
备注:arXiv admin note: text overlap with arXiv:2312.01645
摘要:None
摘要:The Chinese numerical string corpus, serves as a valuable resource for speaker verification, particularly in financial transactions. Researches indicate that in short speech scenarios, text-dependent speaker verification (TD-SV) consistently outperforms text-independent speaker verification (TI-SV). However, TD-SV potentially includes the validation of text information, that can be negatively impacted by reading rhythms and pauses. To address this problem, we propose an end-to-end speaker verification system that enhances TD-SV by decoupling speaker and text information. Our system consists of a text embedding extractor, a speaker embedding extractor and a fusion module. In the text embedding extractor, we employ an enhanced Transformer and introduce a triple loss including text classification loss, connectionist temporal classification (CTC) loss and decoder loss; while in the speaker embedding extractor, we create a multi-scale pooling method by combining sliding window attentive statistics pooling (SWASP) with attentive statistics pooling (ASP). To mitigate the scarcity of data, we have recorded a publicly available Chinese numerical corpus named SHALCAS22A (hereinafter called SHAL), which can be accessed on Open-SLR. Moreover, we employ data augmentation techniques using Tacotron2 and HiFi-GAN. Our method achieves an equal error rate (EER) performance improvement of 49.2% on Hi-Mia and 75.0% on SHAL, respectively.


【11】 Benchmarking Cross-Domain Audio-Visual Deception Detection

标题: 跨域视听欺骗检测基准

链接:https://arxiv.org/abs/2405.06995

作者:Xiaobao Guo,Zitong Yu,Nithish Muthuchamy Selvaraj,Bingquan Shen,Adams Wai-Kin Kong,Alex C. Kot
备注:10 pages
摘要:自动欺骗检测对于帮助人类准确评估真实性和识别欺骗行为至关重要。传统的基于接触的技术,如测谎仪,依赖于生理信号来确定个人陈述的真实性。然而,最近的发展,在自动欺骗检测已经证明,多模态的功能来自音频和视频模态可能会优于人类观察员公开可用的数据集。尽管有这些积极的发现,现有的视听欺骗检测方法在不同情况下的普遍性仍然在很大程度上未被探索。为了缩小这一差距,我们提出了第一个跨域视听欺骗检测基准,使我们能够评估这些方法在现实世界中的推广使用情况。我们使用广泛采用的音频和视觉特征以及不同的架构进行基准测试,比较单对单和多对单域泛化性能。为了进一步利用来自多个源域的数据进行训练的影响,我们研究了三种类型的域采样策略,包括域同步,域交替和逐域进行多到单域泛化评估。此外,我们提出了Attention-Mixer融合方法,以提高性能,我们相信,这种新的跨域基准将有助于未来的研究在视听欺骗检测。协议和源代码可在\href{https://github.com/Redaimao/cross_domain_DD}{https://github.com/Redaimao/cross\_domain\_DD}获得。
摘要:Automated deception detection is crucial for assisting humans in accurately assessing truthfulness and identifying deceptive behavior. Conventional contact-based techniques, like polygraph devices, rely on physiological signals to determine the authenticity of an individual's statements. Nevertheless, recent developments in automated deception detection have demonstrated that multimodal features derived from both audio and video modalities may outperform human observers on publicly available datasets. Despite these positive findings, the generalizability of existing audio-visual deception detection approaches across different scenarios remains largely unexplored. To close this gap, we present the first cross-domain audio-visual deception detection benchmark, that enables us to assess how well these methods generalize for use in real-world scenarios. We used widely adopted audio and visual features and different architectures for benchmarking, comparing single-to-single and multi-to-single domain generalization performance. To further exploit the impacts using data from multiple source domains for training, we investigate three types of domain sampling strategies, including domain-simultaneous, domain-alternating, and domain-by-domain for multi-to-single domain generalization evaluation. Furthermore, we proposed the Attention-Mixer fusion method to improve performance, and we believe that this new cross-domain benchmark will facilitate future research in audio-visual deception detection. Protocols and source code are available at \href{https://github.com/Redaimao/cross_domain_DD}{https://github.com/Redaimao/cross\_domain\_DD}.


【12】 Time-of-arrival Estimation and Phase Unwrapping of Head-related Transfer Functions With Integer Linear Programming

标题: 基于ARCH线性规划的头部相关传递函数到达时间估计和阶段展开

链接:https://arxiv.org/abs/2405.06804

作者:Chin-Yun Yu,Johan Pauwels,György Fazekas
备注:Accepted to be presented at Audio Engineering Society 156th Convention, 2024 June, Madrid, Spain
摘要:在双耳音频合成中,在时间上对准头部相关脉冲响应(HRIR)已经是重要的预处理步骤,从而实现精确的空间插值和高效的数据压缩。空间上邻近的HRIR之间的最大相关时间延迟先前已被用于通过求解矩阵方程来获得准确且平滑的对准,在该矩阵方程中,解具有到时间延迟的最小欧几里得距离。然而,欧几里德准则在实践中可能导致过度平滑的解决方案。在本文中,我们解决的平滑问题,制定的任务,解决一个整数线性规划问题,相当于最小化的$L^1$-范数。此外,我们结合了1)耳间HRIR的互相关,以及2)HRIR与它们的最小相位响应,以具有更多的参考测量用于优化。我们表明,该方法可以得到更准确的路线比基于欧几里得的方法,通过比较光谱重建损失的时间对准HRIR使用球面谐波表示7 HRIR组成的人和假人头部。额外的相关性特征和$L^1$-范数在极端噪声条件下也是有益的。此外,该方法可以应用于头部相关传递函数的相位展开,其中展开的相位可以是下游任务的紧凑特征。
摘要:In binaural audio synthesis, aligning head-related impulse responses (HRIRs) in time has been an important pre-processing step, enabling accurate spatial interpolation and efficient data compression. The maximum correlation time delay between spatially nearby HRIRs has previously been used to get accurate and smooth alignment by solving a matrix equation in which the solution has the minimum Euclidean distance to the time delay. However, the Euclidean criterion could lead to an over-smoothing solution in practice. In this paper, we solve the smoothing issue by formulating the task as solving an integer linear programming problem equivalent to minimising an $L^1$-norm. Moreover, we incorporate 1) the cross-correlation of inter-aural HRIRs, and 2) HRIRs with their minimum-phase responses to have more reference measurements for optimisation. We show the proposed method can get more accurate alignments than the Euclidean-based method by comparing the spectral reconstruction loss of time-aligned HRIRs using spherical harmonics representation on seven HRIRs consisting of human and dummy heads. The extra correlation features and the $L^1$-norm are also beneficial in extremely noisy conditions. In addition, this method can be applied to phase unwrapping of head-related transfer functions, where the unwrapped phase could be a compact feature for downstream tasks.


【13】 Music Emotion Prediction Using Recurrent Neural Networks

标题: 使用回归神经网络的音乐情感预测

链接:https://arxiv.org/abs/2405.06747

作者:Xinyu Chang,Xiangyu Zhang,Haoruo Zhang,Yulu Ran
备注:15 pages, 13 figures
摘要:本研究探讨了递归神经网络在识别音乐中传达的情感方面的应用,旨在通过定制音乐以适应听众的情绪状态来增强音乐推荐系统并支持治疗干预。我们利用罗素的情感象限将音乐分为四个不同的情感区域,并开发能够准确预测这些类别的模型。我们的方法涉及使用Librosa提取一组全面的音频特征,并应用各种递归神经网络架构,包括标准RNN,双向RNN和长短期记忆(LSTM)网络。使用900个音频片段的数据集进行初始实验,根据情感象限进行标记。我们将我们的神经网络模型的性能与一组基线分类器进行比较,并分析它们在捕捉音乐表达中固有的时间动态方面的有效性。结果表明,更简单的RNN架构可以执行更复杂的模型,特别是在较小的数据集上。我们还在更大的数据集上应用了以下实验:一个是基于我们的原始数据集进行增强的,另一个是来自其他来源的。这项研究不仅增强了我们对音乐情感影响的理解,还展示了神经网络在创建更具个性化和情感共鸣的音乐推荐和治疗系统方面的潜力。
摘要:This study explores the application of recurrent neural networks to recognize emotions conveyed in music, aiming to enhance music recommendation systems and support therapeutic interventions by tailoring music to fit listeners' emotional states. We utilize Russell's Emotion Quadrant to categorize music into four distinct emotional regions and develop models capable of accurately predicting these categories. Our approach involves extracting a comprehensive set of audio features using Librosa and applying various recurrent neural network architectures, including standard RNNs, Bidirectional RNNs, and Long Short-Term Memory (LSTM) networks. Initial experiments are conducted using a dataset of 900 audio clips, labeled according to the emotional quadrants. We compare the performance of our neural network models against a set of baseline classifiers and analyze their effectiveness in capturing the temporal dynamics inherent in musical expression. The results indicate that simpler RNN architectures may perform comparably or even superiorly to more complex models, particularly in smaller datasets. We've also applied the following experiments on larger datasets: one is augmented based on our original dataset, and the other is from other sources. This research not only enhances our understanding of the emotional impact of music but also demonstrates the potential of neural networks in creating more personalized and emotionally resonant music recommendation and therapy systems.


机器翻译由腾讯交互翻译提供,仅供参考