【1】Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset标题:音频能揭示音乐表演的困难吗?来自钢琴教学大纲数据集的见解链接:https://arxiv.org/abs/2403.03947作者:Pedro Ramoneda,Minhee Lee,Dasaem Jeong,J. J. Valero-Mas,Xavier Serra摘要:自动估计音乐作品的演奏难度是音乐教育中根据学生的个性化需求创建定制课程的关键过程。鉴于其相关性,音乐信息检索(MIR)领域描述了一些解决这一任务的概念验证工作,主要集中在高级音乐抽象,如机器可读的乐谱或乐谱图像。在这方面,直接分析录音的潜力通常被忽视,这阻止了学生探索可能没有正式符号级转录的各种音乐作品。这项工作在自动估计音频记录上的音乐作品的演奏难度方面具有两个精确的贡献:(i)第一个基于音频的难度估计数据集-即钢琴教学大纲(PSyllabus)数据集-包含来自1,233位作曲家的11个难度级别的7,901首钢琴作品;以及(ii)能够管理直接从音频导出的不同输入表示(单模态和多模态方式)以执行难度估计任务的识别框架。综合实验,包括不同的预训练方案,输入方式,和多任务的情况下证明了该建议的有效性,并建立PSyllabus作为参考数据集的音频为基础的难度估计在MIR领域。数据集以及开发的代码和训练的模型都是公开共享的,以促进该领域的进一步研究。摘要:Automatically estimating the performance difficulty of a music piece represents a key process in music education to create tailored curricula according to the individual needs of the students. Given its relevance, the Music Information Retrieval (MIR) field depicts some proof-of-concept works addressing this task that mainly focuses on high-level music abstractions such as machine-readable scores or music sheet images. In this regard, the potential of directly analyzing audio recordings has been generally neglected, which prevents students from exploring diverse music pieces that may not have a formal symbolic-level transcription. This work pioneers in the automatic estimation of performance difficulty of music pieces on audio recordings with two precise contributions: (i) the first audio-based difficulty estimation dataset -- namely, Piano Syllabus (PSyllabus) dataset -- featuring 7,901 piano pieces across 11 difficulty levels from 1,233 composers; and (ii) a recognition framework capable of managing different input representations -- both unimodal and multimodal manners -- directly derived from audio to perform the difficulty estimation task. The comprehensive experimentation comprising different pre-training schemes, input modalities, and multi-task scenarios prove the validity of the proposal and establishes PSyllabus as a reference dataset for audio-based difficulty estimation in the MIR field. The dataset as well as the developed code and trained models are publicly shared to promote further research in the field. 【2】 RADIA -- Radio Advertisement Detection with Intelligent Analytics标题:RADIA --无线电广告检测与智能分析链接:https://arxiv.org/abs/2403.03538作者:Jorge Álvarez,Juan Carlos Armenteros,Camilo Torrón,Miguel Ortega-Martín,Alfonso Ardoiz,Óscar García,Ignacio Arranz,Íñigo Galdeano,Ignacio Garrido,Adrián Alonso,Fernando Bayón,Oleg Vorontsov摘要:无线电广告仍然是现代营销战略的一个组成部分,其吸引力和潜在的有针对性的影响力是非常有效的。然而,无线电广播时间的动态性质和多个无线电点的上升趋势需要用于监测广告广播的有效系统。本研究探讨一种新的自动化的无线电广告检测技术,结合先进的语音识别和文本分类算法。RadIA的方法超越了传统方法,因为它不需要事先了解广播内容。该贡献允许检测即兴和新引入的广告,为无线电广播中的广告检测提供全面的解决方案。实验结果表明,在仔细分割和标记的文本数据上训练得到的模型,F1-macro得分为87.76,而理论最大值为89.33。本文提供了对超参数的选择及其对模型性能的影响的见解。这项研究表明,它的潜力,以确保遵守广告广播合同,并提供竞争监督。这项开创性的研究可能会从根本上改变广播广告的监控方式,并为营销优化打开新的大门。摘要:Radio advertising remains an integral part of modern marketing strategies, with its appeal and potential for targeted reach undeniably effective. However, the dynamic nature of radio airtime and the rising trend of multiple radio spots necessitates an efficient system for monitoring advertisement broadcasts. This study investigates a novel automated radio advertisement detection technique incorporating advanced speech recognition and text classification algorithms. RadIA's approach surpasses traditional methods by eliminating the need for prior knowledge of the broadcast content. This contribution allows for detecting impromptu and newly introduced advertisements, providing a comprehensive solution for advertisement detection in radio broadcasting. Experimental results show that the resulting model, trained on carefully segmented and tagged text data, achieves an F1-macro score of 87.76 against a theoretical maximum of 89.33. This paper provides insights into the choice of hyperparameters and their impact on the model's performance. This study demonstrates its potential to ensure compliance with advertising broadcast contracts and offer competitive surveillance. This groundbreaking research could fundamentally change how radio advertising is monitored and open new doors for marketing optimization. 【3】 Non-verbal information in spontaneous speech - towards a new framework of analysis标题:自发言语中的非语言信息--走向一个新的分析框架链接:https://arxiv.org/abs/2403.03522作者:Tirza Biron,Moshe Barboy,Eran Ben-Artzy,Alona Golubchik,Yanir Marmor,Smadar Szekely,Yaron Winter,David Harel摘要:语音中的非语言信号由韵律编码,并携带从会话动作到态度和情感的信息。尽管它的重要性,韵律结构的原则还没有得到充分的理解。本文为韵律信号的分类及其与意义的关联提供了一个分析图式和技术概念证明。该模式解释了多层次韵律事件的表面表征。作为实现的第一步,我们提出了一个分类过程,解开韵律现象的三个订单。它依赖于微调预训练的语音识别模型,从而实现同时的多类别/多标签检测。它概括了各种各样的自发数据,表现与人类注释相当或优于人类注释。除了韵律的标准化形式化,韵律模式的分离可以指导交际和言语组织理论。一个受欢迎的副产品是韵律的解释,这将增强语音和语言相关的技术。摘要:Non-verbal signals in speech are encoded by prosody and carry information that ranges from conversation action to attitude and emotion. Despite its importance, the principles that govern prosodic structure are not yet adequately understood. This paper offers an analytical schema and a technological proof-of-concept for the categorization of prosodic signals and their association with meaning. The schema interprets surface-representations of multi-layered prosodic events. As a first step towards implementation, we present a classification process that disentangles prosodic phenomena of three orders. It relies on fine-tuning a pre-trained speech recognition model, enabling the simultaneous multi-class/multi-label detection. It generalizes over a large variety of spontaneous data, performing on a par with, or superior to, human annotation. In addition to a standardized formalization of prosody, disentangling prosodic patterns can direct a theory of communication and speech organization. A welcome by-product is an interpretation of prosody that will enhance speech- and language-related technologies. 【4】 METAMAT 01: A semi-analytic Solution for Benchmarking Wave Propagation Simulations of homogeneous Absorbers in 1D/3D and 2D标题:METAMAT 01:一维/三维和二维均匀吸波体波传播数值模拟的半解析解链接:https://arxiv.org/abs/2403.03510作者:Stefan Schoder,Paul Maurerlehner备注:4摘要:在时域描述中声学仿真工作流程的发展对于预测航空声学或其他瞬态声学效应的声音是必不可少的。减少噪音的一种常见做法是使用吸收器。这些吸声器的建模通常在频域中提供。建立了几种方法来弥补这一差距,研究在时域中对吸收体进行建模的方法。因此,这篇短文描述了时域中的解析解,用于对具有无限1D、2D和3D域的吸收体模拟进行基准测试。与解析解相连接,提供了Matlab脚本以轻松获得参考解。在EAA TCCA基准测试数据库中,参考代码作为METAMAT 01的基准解决方案提供。摘要:The development of acoustic simulation workflows in the time-domain description is essential for predicting the sound of aeroacoustic or other transient acoustic effects. A common practice for noise mitigation is using absorbers. The modeling of these acoustic absorbers is typically provided in the frequency domain. Several, methods established bridging this gap, investigating methods to model absorber in the time domain. Therefore, this short article, describes the analytic solution in time-domain for benchmarking absorber simulations with infinite 1D, 2D, and 3D domains. Connected to the analytic solution, a Matlab script is provided to easily obtain the reference solution. The reference codes are provided as benchmark solution in the EAA TCCA Benchmarking database as METAMAT 01. 【5】 CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation标题:CrossNet:利用全局、跨频带、窄带和位置编码实现单声道和多声道说话人分离链接:https://arxiv.org/abs/2403.03411作者:Vahid Ahmadi Kalkhorani,DeLiang Wang备注:9 pages摘要:我们介绍CrossNet,一个复杂的频谱映射方法,在混响和噪声条件下的扬声器分离和增强。该架构包括编码层、全局多头自注意模块、跨频带模块、窄带模块和输出层。CrossNet捕获时频域中的全局、跨频带和窄带相关性。为了解决长话语中的性能下降,我们引入了随机块位置编码。在多个数据集上的实验结果证明了CrossNet的有效性和鲁棒性,在混响和噪声混响扬声器分离等任务中实现了最先进的性能。此外,与最近的基线相比,CrossNet的训练速度更快,更稳定。此外,CrossNet的高性能扩展到多麦克风条件,证明了其在各种声学场景中的多功能性。摘要:We introduce CrossNet, a complex spectral mapping approach to speaker separation and enhancement in reverberant and noisy conditions. The proposed architecture comprises an encoder layer, a global multi-head self-attention module, a cross-band module, a narrow-band module, and an output layer. CrossNet captures global, cross-band, and narrow-band correlations in the time-frequency domain. To address performance degradation in long utterances, we introduce a random chunk positional encoding. Experimental results on multiple datasets demonstrate the effectiveness and robustness of CrossNet, achieving state-of-the-art performance in tasks including reverberant and noisy-reverberant speaker separation. Furthermore, CrossNet exhibits faster and more stable training in comparison to recent baselines. Additionally, CrossNet's high performance extends to multi-microphone conditions, demonstrating its versatility in various acoustic scenarios. 【6】 Interactive Melody Generation System for Enhancing the Creativity of Musicians标题:互动式旋律产生系统以提升音乐家的创作力链接:https://arxiv.org/abs/2403.03395作者:So Hirawata,Noriko Otani摘要:本研究提出了一个系统,旨在列举人类之间的协作组成的过程中,使用自动音乐合成技术。通过集成多个递归神经网络(RNN)模型,该系统提供了类似于与多个作曲家合作的体验,从而培养了多样化的创造力。通过动态适应用户的创作意图,基于反馈,该系统增强了其生成符合用户偏好和创作需求的旋律的能力。该系统的有效性进行了评估,通过不同背景的作曲家的实验,揭示其潜力,以促进音乐的创造力,并提出进一步完善的途径。该研究强调了作曲家与人工智能之间互动的重要性,旨在使音乐创作更容易获得和个性化。该系统代表了将人工智能集成到创作过程中的一步,为作曲支持和协作艺术探索提供了新的工具。摘要:This study proposes a system designed to enumerate the process of collaborative composition among humans, using automatic music composition technology. By integrating multiple Recurrent Neural Network (RNN) models, the system provides an experience akin to collaborating with several composers, thereby fostering diverse creativity. Through dynamic adaptation to the user's creative intentions, based on feedback, the system enhances its capability to generate melodies that align with user preferences and creative needs. The system's effectiveness was evaluated through experiments with composers of varying backgrounds, revealing its potential to facilitate musical creativity and suggesting avenues for further refinement. The study underscores the importance of interaction between the composer and AI, aiming to make music composition more accessible and personalized. This system represents a step towards integrating AI into the creative process, offering a new tool for composition support and collaborative artistic exploration. 【7】 Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task标题:语谱图和标度图作为声学识别任务输入的性能比较链接:https://arxiv.org/abs/2403.03611作者:Dang Thoai Phan,Andre Jakob,Marcus Purat摘要:声学识别是近年来深度学习研究中的一项常见任务,采用短时傅立叶变换和小波变换等频谱特征提取。然而,并没有太多的研究发现,讨论的优点和缺点,以及性能比较之间的光谱特征提取。在这种情况下,本文的目的是比较这两种转换类型,称为频谱图和尺度图的属性。实现了一种用于声学故障识别的卷积神经网络,并记录了这两种谱提取器的性能以供比较。最新的研究相同的音频数据库被认为是基准测试,看看有多好的设计的声谱图和尺度图。分析了它们的优点和局限性。通过这样做,本文的结果提供了频谱图和尺度图的应用场景的指示,以及潜在的进一步研究方向在声学识别。摘要:Acoustic recognition is a common task for deep learning in recent researches, with the employment of spectral feature extraction such as Short-time Fourier transform and Wavelet transform. However, not many researches have found that discuss the advantages and drawbacks, as well as performance comparison amongst spectral feature extractors. In this consideration, this paper aims to comparing the attributes of these two transform types, called spectrogram and scalogram. A Convolutional Neural Networks for acoustic faults recognition is implemented, then the performance of these two types of spectral extractor is recorded for comparison. A latest research on the same audio database is considered for benchmarking to see how good the designed spectrogram and scalogram is. The advantages and limitations of them are also analyzed. By doing so, the results of this paper provide indications for application scenarios of spectrogram and scalogram, as well as potential further research directions in acoustic recognition.
【8】 Reinforcement Learning Jazz Improvisation: When Music Meets Game Theory标题:强化学习爵士即兴表演:当音乐与游戏理论相遇链接:https://arxiv.org/abs/2403.03224作者:Vedant Tapiavala,Joshua Piesner,Sourjyamoy Barman,Feng Fu备注:16 pages, 4 figures摘要:现场音乐表演总是很有魅力,由于音乐家之间的动态和与观众的互动,即兴创作具有不可预测性。爵士乐即兴创作是一个特别值得注意的例子,从理论的角度进一步研究。在这里,我们介绍了一种新的数学游戏理论模型的爵士乐即兴,提供了一个框架,研究音乐理论和即兴方法。我们使用计算建模,主要是强化学习,探索不同的随机即兴策略和他们的即兴表演配对。我们发现,最有效的策略对是一种策略,该策略对最近的收益(逐步变化)做出反应,强化学习策略仅限于给定和弦中的音符(和弦跟随强化学习)。相反,对合作伙伴的最后一个音符做出反应并试图与之协调的策略(和声预测)策略对产生最低的非控制回报和最高的标准差,这表明基于对合作伙伴的即时反应选择音符可能会产生不一致的结果。平均而言,和弦跟随强化学习策略的平均收益最高,而和声预测的平均收益最低。我们的工作为爵士乐之外的有前途的应用奠定了基础:包括使用人工智能(AI)模型从音频片段中提取数据以改进音乐奖励系统,以及在现有爵士乐独奏上训练机器学习(ML)模型以进一步改进游戏中的策略。摘要:Live performances of music are always charming, with the unpredictability of improvisation due to the dynamic between musicians and interactions with the audience. Jazz improvisation is a particularly noteworthy example for further investigation from a theoretical perspective. Here, we introduce a novel mathematical game theory model for jazz improvisation, providing a framework for studying music theory and improvisational methodologies. We use computational modeling, mainly reinforcement learning, to explore diverse stochastic improvisational strategies and their paired performance on improvisation. We find that the most effective strategy pair is a strategy that reacts to the most recent payoff (Stepwise Changes) with a reinforcement learning strategy limited to notes in the given chord (Chord-Following Reinforcement Learning). Conversely, a strategy that reacts to the partner's last note and attempts to harmonize with it (Harmony Prediction) strategy pair yields the lowest non-control payoff and highest standard deviation, indicating that picking notes based on immediate reactions to the partner player can yield inconsistent outcomes. On average, the Chord-Following Reinforcement Learning strategy demonstrates the highest mean payoff, while Harmony Prediction exhibits the lowest. Our work lays the foundation for promising applications beyond jazz: including the use of artificial intelligence (AI) models to extract data from audio clips to refine musical reward systems, and training machine learning (ML) models on existing jazz solos to further refine strategies within the game.
eess.AS音频处理【1】 Room Impulse Response Estimation using Optimal Transport: Simulation-Informed Inference标题:基于最优传输的房间脉冲响应估计:模拟推理链接:https://arxiv.org/abs/2403.03762作者:David Sundström,Anton Björkman,Andreas Jakobsson,Filip Elvander摘要:准确估计房间脉冲响应(RIR)的能力对于空间音频处理的许多应用是不可或缺的。遗憾的是,使用诸如语音或音乐的环境信号来估计RIR仍然是一个具有挑战性的问题,这是由于例如,低信噪比、有限的样本长度和差的光谱激发。通常,为了改善估计问题的条件,先验被放置在RIR的幅度上。虽然用作正则化器,但当仅可获得延迟结构的近似知识时,这种类型的先验通常是无用的,例如,当先验是来自房间几何形状的近似的模拟RIR时就是这种情况。在这项工作中,我们针对延迟结构本身,构建一个先验的基础上的概念,最佳运输。如使用模拟和测量数据所示,所得到的方法能够有利地结合甚至来自简单模拟模型的信息,对假设的房间尺寸及其温度的扰动显示出相当大的鲁棒性。摘要:The ability to accurately estimate room impulse responses (RIRs) is integral to many applications of spatial audio processing. Regrettably, estimating the RIR using ambient signals, such as speech or music, remains a challenging problem due to, e.g., low signal-to-noise ratios, finite sample lengths, and poor spectral excitation. Commonly, in order to improve the conditioning of the estimation problem, priors are placed on the amplitudes of the RIR. Although serving as a regularizer, this type of prior is generally not useful when only approximate knowledge of the delay structure is available, which, for example, is the case when the prior is a simulated RIR from an approximation of the room geometry. In this work, we target the delay structure itself, constructing a prior based on the concept of optimal transport. As illustrated using both simulated and measured data, the resulting method is able to beneficially incorporate information even from simple simulation models, displaying considerable robustness to perturbations in the assumed room dimensions and its temperature.
【2】 Comparison Performance of Spectrogram and Scalogram as Input of Acoustic Recognition Task标题:语谱图和标度图作为声学识别任务输入的性能比较链接:https://arxiv.org/abs/2403.03611作者:Dang Thoai Phan,Andre Jakob,Marcus Purat摘要:声学识别是近年来深度学习研究中的一项常见任务,采用短时傅立叶变换和小波变换等频谱特征提取。然而,并没有太多的研究发现,讨论的优点和缺点,以及性能比较之间的光谱特征提取。在这种情况下,本文的目的是比较这两种转换类型,称为频谱图和尺度图的属性。实现了一种用于声学故障识别的卷积神经网络,并记录了这两种谱提取器的性能以供比较。最新的研究相同的音频数据库被认为是基准测试,看看有多好的设计的声谱图和尺度图。分析了它们的优点和局限性。通过这样做,本文的结果提供了频谱图和尺度图的应用场景的指示,以及潜在的进一步研究方向在声学识别。摘要:Acoustic recognition is a common task for deep learning in recent researches, with the employment of spectral feature extraction such as Short-time Fourier transform and Wavelet transform. However, not many researches have found that discuss the advantages and drawbacks, as well as performance comparison amongst spectral feature extractors. In this consideration, this paper aims to comparing the attributes of these two transform types, called spectrogram and scalogram. A Convolutional Neural Networks for acoustic faults recognition is implemented, then the performance of these two types of spectral extractor is recorded for comparison. A latest research on the same audio database is considered for benchmarking to see how good the designed spectrogram and scalogram is. The advantages and limitations of them are also analyzed. By doing so, the results of this paper provide indications for application scenarios of spectrogram and scalogram, as well as potential further research directions in acoustic recognition. 【3】 Can Audio Reveal Music Performance Difficulty? Insights from the Piano Syllabus Dataset标题:音频能揭示音乐表现的困难吗?钢琴教学大纲数据集的见解链接:https://arxiv.org/abs/2403.03947作者:Pedro Ramoneda,Minhee Lee,Dasaem Jeong,J. J. Valero-Mas,Xavier Serra摘要:自动估计音乐作品的演奏难度是音乐教育中根据学生的个性化需求创建定制课程的关键过程。鉴于其相关性,音乐信息检索(MIR)领域描述了一些解决这一任务的概念验证工作,主要集中在高级音乐抽象,如机器可读的乐谱或乐谱图像。在这方面,直接分析录音的潜力通常被忽视,这阻止了学生探索可能没有正式符号级转录的各种音乐作品。这项工作在自动估计音频记录上的音乐作品的演奏难度方面具有两个精确的贡献:(i)第一个基于音频的难度估计数据集-即钢琴教学大纲(PSyllabus)数据集-包含来自1,233位作曲家的11个难度级别的7,901首钢琴作品;以及(ii)能够管理直接从音频导出的不同输入表示(单模态和多模态方式)以执行难度估计任务的识别框架。综合实验,包括不同的预训练方案,输入方式,和多任务的情况下证明了该建议的有效性,并建立PSyllabus作为参考数据集的音频为基础的难度估计在MIR领域。数据集以及开发的代码和训练的模型都是公开共享的,以促进该领域的进一步研究。摘要:Automatically estimating the performance difficulty of a music piece represents a key process in music education to create tailored curricula according to the individual needs of the students. Given its relevance, the Music Information Retrieval (MIR) field depicts some proof-of-concept works addressing this task that mainly focuses on high-level music abstractions such as machine-readable scores or music sheet images. In this regard, the potential of directly analyzing audio recordings has been generally neglected, which prevents students from exploring diverse music pieces that may not have a formal symbolic-level transcription. This work pioneers in the automatic estimation of performance difficulty of music pieces on audio recordings with two precise contributions: (i) the first audio-based difficulty estimation dataset -- namely, Piano Syllabus (PSyllabus) dataset -- featuring 7,901 piano pieces across 11 difficulty levels from 1,233 composers; and (ii) a recognition framework capable of managing different input representations -- both unimodal and multimodal manners -- directly derived from audio to perform the difficulty estimation task. The comprehensive experimentation comprising different pre-training schemes, input modalities, and multi-task scenarios prove the validity of the proposal and establishes PSyllabus as a reference dataset for audio-based difficulty estimation in the MIR field. The dataset as well as the developed code and trained models are publicly shared to promote further research in the field. 【4】 RADIA -- Radio Advertisement Detection with Intelligent Analytics标题:Radia--基于智能分析的广播广告检测链接:https://arxiv.org/abs/2403.03538作者:Jorge Álvarez,Juan Carlos Armenteros,Camilo Torrón,Miguel Ortega-Martín,Alfonso Ardoiz,Óscar García,Ignacio Arranz,Íñigo Galdeano,Ignacio Garrido,Adrián Alonso,Fernando Bayón,Oleg Vorontsov摘要:无线电广告仍然是现代营销战略的一个组成部分,其吸引力和潜在的有针对性的影响力是非常有效的。然而,无线电广播时间的动态性质和多个无线电点的上升趋势需要用于监测广告广播的有效系统。本研究探讨一种新的自动化的无线电广告检测技术,结合先进的语音识别和文本分类算法。RadIA的方法超越了传统方法,因为它不需要事先了解广播内容。该贡献允许检测即兴和新引入的广告,为无线电广播中的广告检测提供全面的解决方案。实验结果表明,在仔细分割和标记的文本数据上训练得到的模型,F1-macro得分为87.76,而理论最大值为89.33。本文提供了对超参数的选择及其对模型性能的影响的见解。这项研究表明,它的潜力,以确保遵守广告广播合同,并提供竞争监督。这项开创性的研究可能会从根本上改变广播广告的监控方式,并为营销优化打开新的大门。摘要:Radio advertising remains an integral part of modern marketing strategies, with its appeal and potential for targeted reach undeniably effective. However, the dynamic nature of radio airtime and the rising trend of multiple radio spots necessitates an efficient system for monitoring advertisement broadcasts. This study investigates a novel automated radio advertisement detection technique incorporating advanced speech recognition and text classification algorithms. RadIA's approach surpasses traditional methods by eliminating the need for prior knowledge of the broadcast content. This contribution allows for detecting impromptu and newly introduced advertisements, providing a comprehensive solution for advertisement detection in radio broadcasting. Experimental results show that the resulting model, trained on carefully segmented and tagged text data, achieves an F1-macro score of 87.76 against a theoretical maximum of 89.33. This paper provides insights into the choice of hyperparameters and their impact on the model's performance. This study demonstrates its potential to ensure compliance with advertising broadcast contracts and offer competitive surveillance. This groundbreaking research could fundamentally change how radio advertising is monitored and open new doors for marketing optimization. 【5】 Non-verbal information in spontaneous speech - towards a new framework of analysis标题:自发言语中的非语言信息--走向一个新的分析框架链接:https://arxiv.org/abs/2403.03522作者:Tirza Biron,Moshe Barboy,Eran Ben-Artzy,Alona Golubchik,Yanir Marmor,Smadar Szekely,Yaron Winter,David Harel摘要:语音中的非语言信号由韵律编码,并携带从会话动作到态度和情感的信息。尽管它的重要性,韵律结构的原则还没有得到充分的理解。本文为韵律信号的分类及其与意义的关联提供了一个分析图式和技术概念证明。该模式解释了多层次韵律事件的表面表征。作为实现的第一步,我们提出了一个分类过程,解开韵律现象的三个订单。它依赖于微调预训练的语音识别模型,从而实现同时的多类别/多标签检测。它概括了各种各样的自发数据,表现与人类注释相当或优于人类注释。除了韵律的标准化形式化,韵律模式的分离可以指导交际和言语组织理论。一个受欢迎的副产品是韵律的解释,这将增强语音和语言相关的技术。摘要:Non-verbal signals in speech are encoded by prosody and carry information that ranges from conversation action to attitude and emotion. Despite its importance, the principles that govern prosodic structure are not yet adequately understood. This paper offers an analytical schema and a technological proof-of-concept for the categorization of prosodic signals and their association with meaning. The schema interprets surface-representations of multi-layered prosodic events. As a first step towards implementation, we present a classification process that disentangles prosodic phenomena of three orders. It relies on fine-tuning a pre-trained speech recognition model, enabling the simultaneous multi-class/multi-label detection. It generalizes over a large variety of spontaneous data, performing on a par with, or superior to, human annotation. In addition to a standardized formalization of prosody, disentangling prosodic patterns can direct a theory of communication and speech organization. A welcome by-product is an interpretation of prosody that will enhance speech- and language-related technologies.
【6】 METAMAT 01: A semi-analytic Solution for Benchmarking Wave Propagation Simulations of homogeneous Absorbers in 1D/3D and 2D标题:METAMAT 01:一种半解析的解决方案,用于对均匀吸收体的波传播进行1D/3D和2D基准模拟链接:https://arxiv.org/abs/2403.03510作者:Stefan Schoder,Paul Maurerlehner备注:4摘要:在时域描述中声学仿真工作流程的发展对于预测航空声学或其他瞬态声学效应的声音是必不可少的。减少噪音的一种常见做法是使用吸收器。这些吸声器的建模通常在频域中提供。建立了几种方法来弥补这一差距,研究在时域中对吸收体进行建模的方法。因此,这篇短文描述了时域中的解析解,用于对具有无限1D、2D和3D域的吸收体模拟进行基准测试。与解析解相连接,提供了Matlab脚本以轻松获得参考解。在EAA TCCA基准测试数据库中,参考代码作为METAMAT 01的基准解决方案提供。摘要:The development of acoustic simulation workflows in the time-domain description is essential for predicting the sound of aeroacoustic or other transient acoustic effects. A common practice for noise mitigation is using absorbers. The modeling of these acoustic absorbers is typically provided in the frequency domain. Several, methods established bridging this gap, investigating methods to model absorber in the time domain. Therefore, this short article, describes the analytic solution in time-domain for benchmarking absorber simulations with infinite 1D, 2D, and 3D domains. Connected to the analytic solution, a Matlab script is provided to easily obtain the reference solution. The reference codes are provided as benchmark solution in the EAA TCCA Benchmarking database as METAMAT 01.
【7】 CrossNet: Leveraging Global, Cross-Band, Narrow-Band, and Positional Encoding for Single- and Multi-Channel Speaker Separation标题:CrossNet:利用全局、跨频带、窄带和位置编码实现单声道和多声道扬声器分离链接:https://arxiv.org/abs/2403.03411作者:Vahid Ahmadi Kalkhorani,DeLiang Wang备注:9 pages摘要:我们介绍CrossNet,一个复杂的频谱映射方法,在混响和噪声条件下的扬声器分离和增强。该架构包括编码层、全局多头自注意模块、跨频带模块、窄带模块和输出层。CrossNet捕获时频域中的全局、跨频带和窄带相关性。为了解决长话语中的性能下降,我们引入了随机块位置编码。在多个数据集上的实验结果证明了CrossNet的有效性和鲁棒性,在混响和噪声混响扬声器分离等任务中实现了最先进的性能。此外,与最近的基线相比,CrossNet的训练速度更快,更稳定。此外,CrossNet的高性能扩展到多麦克风条件,证明了其在各种声学场景中的多功能性。摘要:We introduce CrossNet, a complex spectral mapping approach to speaker separation and enhancement in reverberant and noisy conditions. The proposed architecture comprises an encoder layer, a global multi-head self-attention module, a cross-band module, a narrow-band module, and an output layer. CrossNet captures global, cross-band, and narrow-band correlations in the time-frequency domain. To address performance degradation in long utterances, we introduce a random chunk positional encoding. Experimental results on multiple datasets demonstrate the effectiveness and robustness of CrossNet, achieving state-of-the-art performance in tasks including reverberant and noisy-reverberant speaker separation. Furthermore, CrossNet exhibits faster and more stable training in comparison to recent baselines. Additionally, CrossNet's high performance extends to multi-microphone conditions, demonstrating its versatility in various acoustic scenarios. 【8】 Interactive Melody Generation System for Enhancing the Creativity of Musicians标题:用于增强音乐家创造力的交互式旋律生成系统链接:https://arxiv.org/abs/2403.03395作者:So Hirawata,Noriko Otani摘要:本研究提出了一个系统,旨在列举人类之间的协作组成的过程中,使用自动音乐合成技术。通过集成多个递归神经网络(RNN)模型,该系统提供了类似于与多个作曲家合作的体验,从而培养了多样化的创造力。通过动态适应用户的创作意图,基于反馈,该系统增强了其生成符合用户偏好和创作需求的旋律的能力。该系统的有效性进行了评估,通过不同背景的作曲家的实验,揭示其潜力,以促进音乐的创造力,并提出进一步完善的途径。该研究强调了作曲家与人工智能之间互动的重要性,旨在使音乐创作更容易获得和个性化。该系统代表了将人工智能集成到创作过程中的一步,为作曲支持和协作艺术探索提供了新的工具。摘要:This study proposes a system designed to enumerate the process of collaborative composition among humans, using automatic music composition technology. By integrating multiple Recurrent Neural Network (RNN) models, the system provides an experience akin to collaborating with several composers, thereby fostering diverse creativity. Through dynamic adaptation to the user's creative intentions, based on feedback, the system enhances its capability to generate melodies that align with user preferences and creative needs. The system's effectiveness was evaluated through experiments with composers of varying backgrounds, revealing its potential to facilitate musical creativity and suggesting avenues for further refinement. The study underscores the importance of interaction between the composer and AI, aiming to make music composition more accessible and personalized. This system represents a step towards integrating AI into the creative process, offering a new tool for composition support and collaborative artistic exploration.
【9】 Reinforcement Learning Jazz Improvisation: When Music Meets Game Theory标题:强化学习爵士即兴表演:当音乐与游戏理论相遇链接:https://arxiv.org/abs/2403.03224作者:Vedant Tapiavala,Joshua Piesner,Sourjyamoy Barman,Feng Fu备注:16 pages, 4 figures摘要:现场音乐表演总是很有魅力,由于音乐家之间的动态和与观众的互动,即兴创作具有不可预测性。爵士乐即兴创作是一个特别值得注意的例子,从理论的角度进一步研究。在这里,我们介绍了一种新的数学游戏理论模型的爵士乐即兴,提供了一个框架,研究音乐理论和即兴方法。我们使用计算建模,主要是强化学习,探索不同的随机即兴策略和他们的即兴表演配对。我们发现,最有效的策略对是一种策略,该策略对最近的收益(逐步变化)做出反应,强化学习策略仅限于给定和弦中的音符(和弦跟随强化学习)。相反,对合作伙伴的最后一个音符做出反应并试图与之协调的策略(和声预测)策略对产生最低的非控制回报和最高的标准差,这表明基于对合作伙伴的即时反应选择音符可能会产生不一致的结果。平均而言,和弦跟随强化学习策略的平均收益最高,而和声预测的平均收益最低。我们的工作为爵士乐之外的有前途的应用奠定了基础:包括使用人工智能(AI)模型从音频片段中提取数据以改进音乐奖励系统,以及在现有爵士乐独奏上训练机器学习(ML)模型以进一步改进游戏中的策略。摘要:Live performances of music are always charming, with the unpredictability of improvisation due to the dynamic between musicians and interactions with the audience. Jazz improvisation is a particularly noteworthy example for further investigation from a theoretical perspective. Here, we introduce a novel mathematical game theory model for jazz improvisation, providing a framework for studying music theory and improvisational methodologies. We use computational modeling, mainly reinforcement learning, to explore diverse stochastic improvisational strategies and their paired performance on improvisation. We find that the most effective strategy pair is a strategy that reacts to the most recent payoff (Stepwise Changes) with a reinforcement learning strategy limited to notes in the given chord (Chord-Following Reinforcement Learning). Conversely, a strategy that reacts to the partner's last note and attempts to harmonize with it (Harmony Prediction) strategy pair yields the lowest non-control payoff and highest standard deviation, indicating that picking notes based on immediate reactions to the partner player can yield inconsistent outcomes. On average, the Chord-Following Reinforcement Learning strategy demonstrates the highest mean payoff, while Harmony Prediction exhibits the lowest. Our work lays the foundation for promising applications beyond jazz: including the use of artificial intelligence (AI) models to extract data from audio clips to refine musical reward systems, and training machine learning (ML) models on existing jazz solos to further refine strategies within the game. 机器翻译由腾讯交互翻译提供,仅供参考