【1】MuPT: A Generative Symbolic Music Pretrained Transformer标题:MuPT:一个生成符号音乐预训练Transformer链接:https://arxiv.org/abs/2404.06393作者:Xingwei Qu,Yuelin Bai,Yinghao Ma,Ziya Zhou,Ka Man Lo,Jiaheng Liu,Ruibin Yuan,Lejun Min,Xueling Liu,Tianyu Zhang,Xinrun Du,Shuyue Guo,Yiming Liang,Yizhi Li,Shangda Wu,Junting Zhou,Tianyu Zheng,Ziyang Ma,Fengze Han,Wei Xue,Gus Xia,Emmanouil Benetos,Xiang Yue,Chenghua Lin,Xu Tan,Stephen W. Huang,Wenhu Chen,Jie Fu,Ge Zhang摘要:在本文中,我们探讨了大语言模型(LLM)的音乐预训练的应用。虽然LLM在音乐建模中的普遍使用是公认的,但我们的研究结果表明,LLM本质上与ABC符号更兼容,这与它们的设计和优势更紧密地结合在一起,从而提高了模型在音乐创作中的表现。为了解决在生成过程中与来自不同轨道的未对齐措施相关的挑战,我们建议开发一种\underline{S}简化\underline{M}ulti-\underline{T}rack ABC Notation(\textbf{SMT-ABC Notation}),旨在保持多个音乐轨道的一致性。我们的贡献包括一系列能够处理多达8192个令牌的模型,覆盖了我们训练集中90%的符号音乐数据。此外,我们探讨了\underline{S}符号\underline{M}usic \underline{S}标度律(\textbf{SMS Law})对模型性能的影响。这些结果为未来的音乐生成研究指明了一个有希望的方向,通过我们的开源贡献为社区主导的研究提供了广泛的资源。摘要:In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition. To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a \underline{S}ynchronized \underline{M}ulti-\underline{T}rack ABC Notation (\textbf{SMT-ABC Notation}), which aims to preserve coherence across multiple musical tracks. Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the \underline{S}ymbolic \underline{M}usic \underline{S}caling Law (\textbf{SMS Law}) on model performance. The results indicate a promising direction for future research in music generation, offering extensive resources for community-led research through our open-source contributions. 【2】 nEMO: Dataset of Emotional Speech in Polish标题:nEMO:波兰语情感言语数据集链接:https://arxiv.org/abs/2404.06292作者:Iwona Christop备注:Accepted for LREC-Coling 2024摘要:近年来,语音情感识别由于其在医疗保健、客户服务和个性化对话系统中的潜在应用而变得越来越重要。然而,这一领域的一个主要问题是缺乏能够充分代表不同语系基本情绪状态的数据集。由于涵盖斯拉夫语言的数据集非常罕见,因此需要解决这一研究空白。本文介绍了波兰语情感语音语料库nEMO的开发。该数据集包括超过3个小时的样本,由9名演员参与记录,描绘了6种情绪状态:愤怒,恐惧,快乐,悲伤,惊讶和中性状态。所用的文字材料经过精心挑选,充分体现了波兰语的语音。该语料库在知识共享许可证(CC BY-NC-SA 4.0)的条款下免费提供。摘要:Speech emotion recognition has become increasingly important in recent years due to its potential applications in healthcare, customer service, and personalization of dialogue systems. However, a major issue in this field is the lack of datasets that adequately represent basic emotional states across various language families. As datasets covering Slavic languages are rare, there is a need to address this research gap. This paper presents the development of nEMO, a novel corpus of emotional speech in Polish. The dataset comprises over 3 hours of samples recorded with the participation of nine actors portraying six emotional states: anger, fear, happiness, sadness, surprise, and a neutral state. The text material used was carefully selected to represent the phonetics of the Polish language adequately. The corpus is freely available under the terms of a Creative Commons license (CC BY-NC-SA 4.0). 【3】 Exploring Diverse Sounds: Identifying Outliers in a Music Corpus标题:探索不同的声音:识别音乐语料库中的异常值链接:https://arxiv.org/abs/2404.06103作者:Le Cai,Sam Ferguson,Gengfa Fang,Hani Alshamrani备注:None摘要:现有的音乐推荐系统的研究主要集中在推荐相似的音乐,从而往往忽略了多样性和独特的音乐记录。由于音乐本身固有的多样性,音乐离群值可以提供有价值的见解。在本文中,我们探讨音乐离群值,调查他们的音乐发现和推荐系统的潜在用途。我们认为,不是所有的离群值都应该被视为噪音,因为它们可以提供有趣的视角,并有助于更丰富地理解艺术家的作品。我们介绍了“真正的”音乐离群值的概念,并为他们提供了一个定义。这些真正的离群值可以揭示艺术家曲目的独特方面,并通过让听众体验新颖多样的音乐体验来增强音乐发现的潜力。摘要:Existing research on music recommendation systems primarily focuses on recommending similar music, thereby often neglecting diverse and distinctive musical recordings. Musical outliers can provide valuable insights due to the inherent diversity of music itself. In this paper, we explore music outliers, investigating their potential usefulness for music discovery and recommendation systems. We argue that not all outliers should be treated as noise, as they can offer interesting perspectives and contribute to a richer understanding of an artist's work. We introduce the concept of 'Genuine' music outliers and provide a definition for them. These genuine outliers can reveal unique aspects of an artist's repertoire and hold the potential to enhance music discovery by exposing listeners to novel and diverse musical experiences.
【4】 A Novel Bi-LSTM And Transformer Architecture For Generating Tabla Music标题:一种新的Bi—LSTM和Transformer结构的Tabla音乐生成链接:https://arxiv.org/abs/2404.05765作者:Roopa Mayya,Vivekanand Venkataraman,Anwesh P R,Narayana Darapaneni摘要:简介:音乐生成是一项复杂的任务,近年来受到了极大的关注,深度学习技术在这一领域取得了可喜的成果。目的:虽然在生成钢琴和其他西方音乐方面已经进行了大量的工作,但由于机器编码格式的印度音乐的稀缺性,对生成古典印度音乐的研究有限。在这篇技术论文中,提出了产生古典印度音乐,特别是塔布拉音乐的方法。首先,本文探讨了使用深度学习架构生成钢琴音乐。然后将基本原理扩展到生成塔布拉音乐。方法:使用Python中的librosa库对波形(.wav)文件中的Tabla音乐进行预处理。基于提取的特征和标签训练了一种新的具有注意力方法和Transformer模型的Bi-LSTM。结果:然后使用模型来预测下一个序列的塔布拉音乐。Bi-LSTM模型的损失为4.042,MAE为1.0814。在Transformer模型下,产生Tabla音乐的损耗为55.9278,MAE为3.5173。结论:由此产生的音乐体现了新颖性和熟悉性的和谐融合,将音乐创作的极限推向了新的视野。摘要:Introduction: Music generation is a complex task that has received significant attention in recent years, and deep learning techniques have shown promising results in this field. Objectives: While extensive work has been carried out on generating Piano and other Western music, there is limited research on generating classical Indian music due to the scarcity of Indian music in machine-encoded formats. In this technical paper, methods for generating classical Indian music, specifically tabla music, is proposed. Initially, this paper explores piano music generation using deep learning architectures. Then the fundamentals are extended to generating tabla music. Methods: Tabla music in waveform (.wav) files are pre-processed using the librosa library in Python. A novel Bi-LSTM with an Attention approach and a transformer model are trained on the extracted features and labels. Results: The models are then used to predict the next sequences of tabla music. A loss of 4.042 and MAE of 1.0814 are achieved with the Bi-LSTM model. With the transformer model, a loss of 55.9278 and MAE of 3.5173 are obtained for tabla music generation. Conclusion: The resulting music embodies a harmonious fusion of novelty and familiarity, pushing the limits of music composition to new horizons. 【5】 Masked Modeling Duo: Towards a Universal Audio Pre-training Framework标题:Masked Modeling Duo:迈向通用音频预训练框架链接:https://arxiv.org/abs/2404.06095作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Kunio Kashino备注:15 pages, 6 figures, 15 tables. Accepted by TASLP摘要:使用掩蔽预测的自监督学习(SSL)在通用音频表示方面取得了很大进展。这项研究提出了Masked Modeling Duo(M2 D),这是一种改进的掩码预测SSL,它通过预测作为训练信号的掩码输入信号的表示来学习。与传统方法不同,M2 D通过仅对掩码部分进行编码来获得训练信号,从而鼓励M2 D中的两个网络对输入进行建模。虽然M2 D改进了通用音频表示,但对于诸如工业和医疗领域等现实世界的应用,专业表示是必不可少的。这些领域中通常是机密和专有的数据通常大小有限,并且具有与预训练数据集不同的分布。因此,我们提出了M2 D for X(M2 D-X),它扩展了M2 D,以实现对应用程序X的专门表示的预训练。M2 D-X从M2 D和附加任务中学习,并输入背景噪声。我们使额外的任务可配置为服务于不同的应用程序,而背景噪声有助于在小数据上学习,并形成一个去噪任务,使表示鲁棒。有了这些设计选择,M2 D-X应该学习一种专门用于满足各种应用需求的表示。我们的实验证实,通用音频的表示,专门用于竞争激烈的AudioSet和语音领域,以及小数据医疗任务,实现了顶级性能,证明了使用我们的模型作为通用音频预训练框架的潜力。我们的代码可以在https://github.com/nttcslab/m2d上在线获得,以供将来研究。摘要:Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals. Unlike conventional methods, M2D obtains a training signal by encoding only the masked part, encouraging the two networks in M2D to model the input. While M2D improves general-purpose audio representations, a specialized representation is essential for real-world applications, such as in industrial and medical domains. The often confidential and proprietary data in such domains is typically limited in size and has a different distribution from that in pre-training datasets. Therefore, we propose M2D for X (M2D-X), which extends M2D to enable the pre-training of specialized representations for an application X. M2D-X learns from M2D and an additional task and inputs background noise. We make the additional task configurable to serve diverse applications, while the background noise helps learn on small data and forms a denoising task that makes representation robust. With these design choices, M2D-X should learn a representation specialized to serve various application needs. Our experiments confirmed that the representations for general-purpose audio, specialized for the highly competitive AudioSet and speech domain, and a small-data medical task achieve top-level performance, demonstrating the potential of using our models as a universal audio pre-training framework. Our code is available online for future studies at https://github.com/nttcslab/m2d eess.AS音频处理【1】 Masked Modeling Duo: Towards a Universal Audio Pre-training Framework标题:Masked Modeling Duo:迈向通用音频预训练框架链接:https://arxiv.org/abs/2404.06095作者:Daisuke Niizumi,Daiki Takeuchi,Yasunori Ohishi,Noboru Harada,Kunio Kashino备注:15 pages, 6 figures, 15 tables. Accepted by TASLP摘要:使用掩蔽预测的自监督学习(SSL)在通用音频表示方面取得了很大进展。这项研究提出了Masked Modeling Duo(M2 D),这是一种改进的掩码预测SSL,它通过预测作为训练信号的掩码输入信号的表示来学习。与传统方法不同,M2 D通过仅对掩码部分进行编码来获得训练信号,从而鼓励M2 D中的两个网络对输入进行建模。虽然M2 D改进了通用音频表示,但对于诸如工业和医疗领域等现实世界的应用,专业表示是必不可少的。这些领域中通常是机密和专有的数据通常大小有限,并且具有与预训练数据集不同的分布。因此,我们提出了M2 D for X(M2 D-X),它扩展了M2 D,以实现对应用程序X的专门表示的预训练。M2 D-X从M2 D和附加任务中学习,并输入背景噪声。我们使额外的任务可配置为服务于不同的应用程序,而背景噪声有助于在小数据上学习,并形成一个去噪任务,使表示鲁棒。有了这些设计选择,M2 D-X应该学习一种专门用于满足各种应用需求的表示。我们的实验证实,通用音频的表示,专门用于竞争激烈的AudioSet和语音领域,以及小数据医疗任务,实现了顶级性能,证明了使用我们的模型作为通用音频预训练框架的潜力。我们的代码可以在https://github.com/nttcslab/m2d上在线获得,以供将来研究。摘要:Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals. Unlike conventional methods, M2D obtains a training signal by encoding only the masked part, encouraging the two networks in M2D to model the input. While M2D improves general-purpose audio representations, a specialized representation is essential for real-world applications, such as in industrial and medical domains. The often confidential and proprietary data in such domains is typically limited in size and has a different distribution from that in pre-training datasets. Therefore, we propose M2D for X (M2D-X), which extends M2D to enable the pre-training of specialized representations for an application X. M2D-X learns from M2D and an additional task and inputs background noise. We make the additional task configurable to serve diverse applications, while the background noise helps learn on small data and forms a denoising task that makes representation robust. With these design choices, M2D-X should learn a representation specialized to serve various application needs. Our experiments confirmed that the representations for general-purpose audio, specialized for the highly competitive AudioSet and speech domain, and a small-data medical task achieve top-level performance, demonstrating the potential of using our models as a universal audio pre-training framework. Our code is available online for future studies at https://github.com/nttcslab/m2d 【2】 The X-LANCE Technical Report for Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge标题:Interspeech 2024年使用离散语音单元的语音处理技术报告链接:https://arxiv.org/abs/2404.06079作者:Yiwei Guo,Chenrun Wang,Yifan Yang,Hankun Wang,Ziyang Ma,Chenpeng Du,Shuai Wang,Hanzheng Li,Shuai Fan,Hui Zhang,Xie Chen,Kai Yu备注:5 pages, 3 figures. Report of a challenge摘要:离散语音标记在自动语音识别、文语转换和歌声合成等语音处理领域中得到了越来越广泛的应用。在本文中,我们描述了由上海交通大学X-LANCE小组开发的TTS(声学+声码器),SVS和ASR轨道在Interspeech 2024使用离散语音单元的语音处理挑战系统。值得注意的是,我们在TTS曲目的排行榜上获得了第一名,包括整个训练集和只有1h的训练数据,以及所有提交中最低的比特率。摘要:Discrete speech tokens have been more and more popular in multiple speech processing fields, including automatic speech recognition (ASR), text-to-speech (TTS) and singing voice synthesis (SVS). In this paper, we describe the systems developed by the SJTU X-LANCE group for the TTS (acoustic + vocoder), SVS, and ASR tracks in the Interspeech 2024 Speech Processing Using Discrete Speech Unit Challenge. Notably, we achieved 1st rank on the leaderboard in the TTS track both with the whole training set and only 1h training data, along with the lowest bitrate among all submissions. 【3】 MuPT: A Generative Symbolic Music Pretrained Transformer标题:MuPT:一个生成符号音乐预训练Transformer链接:https://arxiv.org/abs/2404.06393作者:Xingwei Qu,Yuelin Bai,Yinghao Ma,Ziya Zhou,Ka Man Lo,Jiaheng Liu,Ruibin Yuan,Lejun Min,Xueling Liu,Tianyu Zhang,Xinrun Du,Shuyue Guo,Yiming Liang,Yizhi Li,Shangda Wu,Junting Zhou,Tianyu Zheng,Ziyang Ma,Fengze Han,Wei Xue,Gus Xia,Emmanouil Benetos,Xiang Yue,Chenghua Lin,Xu Tan,Stephen W. Huang,Wenhu Chen,Jie Fu,Ge Zhang摘要:在本文中,我们探讨了大语言模型(LLM)的音乐预训练的应用。虽然LLM在音乐建模中的普遍使用是公认的,但我们的研究结果表明,LLM本质上与ABC符号更兼容,这与它们的设计和优势更紧密地结合在一起,从而提高了模型在音乐创作中的表现。为了解决在生成过程中与来自不同轨道的未对齐措施相关的挑战,我们建议开发一种\underline{S}简化\underline{M}ulti-\underline{T}rack ABC Notation(\textbf{SMT-ABC Notation}),旨在保持多个音乐轨道的一致性。我们的贡献包括一系列能够处理多达8192个令牌的模型,覆盖了我们训练集中90%的符号音乐数据。此外,我们探讨了\underline{S}符号\underline{M}usic \underline{S}标度律(\textbf{SMS Law})对模型性能的影响。这些结果为未来的音乐生成研究指明了一个有希望的方向,通过我们的开源贡献为社区主导的研究提供了广泛的资源。摘要:In this paper, we explore the application of Large Language Models (LLMs) to the pre-training of music. While the prevalent use of MIDI in music modeling is well-established, our findings suggest that LLMs are inherently more compatible with ABC Notation, which aligns more closely with their design and strengths, thereby enhancing the model's performance in musical composition. To address the challenges associated with misaligned measures from different tracks during generation, we propose the development of a \underline{S}ynchronized \underline{M}ulti-\underline{T}rack ABC Notation (\textbf{SMT-ABC Notation}), which aims to preserve coherence across multiple musical tracks. Our contributions include a series of models capable of handling up to 8192 tokens, covering 90\% of the symbolic music data in our training set. Furthermore, we explore the implications of the \underline{S}ymbolic \underline{M}usic \underline{S}caling Law (\textbf{SMS Law}) on model performance. The results indicate a promising direction for future research in music generation, offering extensive resources for community-led research through our open-source contributions. 【4】 nEMO: Dataset of Emotional Speech in Polish标题:nEMO:波兰语情感言语数据集链接:https://arxiv.org/abs/2404.06292作者:Iwona Christop备注:Accepted for LREC-Coling 2024摘要:近年来,语音情感识别由于其在医疗保健、客户服务和个性化对话系统中的潜在应用而变得越来越重要。然而,这一领域的一个主要问题是缺乏能够充分代表不同语系基本情绪状态的数据集。由于涵盖斯拉夫语言的数据集非常罕见,因此需要解决这一研究空白。本文介绍了波兰语情感语音语料库nEMO的开发。该数据集包括超过3个小时的样本,由9名演员参与记录,描绘了6种情绪状态:愤怒,恐惧,快乐,悲伤,惊讶和中性状态。所用的文字材料经过精心挑选,充分体现了波兰语的语音。该语料库在知识共享许可证(CC BY-NC-SA 4.0)的条款下免费提供。摘要:Speech emotion recognition has become increasingly important in recent years due to its potential applications in healthcare, customer service, and personalization of dialogue systems. However, a major issue in this field is the lack of datasets that adequately represent basic emotional states across various language families. As datasets covering Slavic languages are rare, there is a need to address this research gap. This paper presents the development of nEMO, a novel corpus of emotional speech in Polish. The dataset comprises over 3 hours of samples recorded with the participation of nine actors portraying six emotional states: anger, fear, happiness, sadness, surprise, and a neutral state. The text material used was carefully selected to represent the phonetics of the Polish language adequately. The corpus is freely available under the terms of a Creative Commons license (CC BY-NC-SA 4.0). 【5】 Exploring Diverse Sounds: Identifying Outliers in a Music Corpus标题:探索不同的声音:识别音乐语料库中的异常值链接:https://arxiv.org/abs/2404.06103作者:Le Cai,Sam Ferguson,Gengfa Fang,Hani Alshamrani备注:None摘要:现有的音乐推荐系统的研究主要集中在推荐相似的音乐,从而往往忽略了多样性和独特的音乐记录。由于音乐本身固有的多样性,音乐离群值可以提供有价值的见解。在本文中,我们探讨音乐离群值,调查他们的音乐发现和推荐系统的潜在用途。我们认为,不是所有的离群值都应该被视为噪音,因为它们可以提供有趣的视角,并有助于更丰富地理解艺术家的作品。我们介绍了“真正的”音乐离群值的概念,并为他们提供了一个定义。这些真正的离群值可以揭示艺术家曲目的独特方面,并通过让听众体验新颖多样的音乐体验来增强音乐发现的潜力。摘要:Existing research on music recommendation systems primarily focuses on recommending similar music, thereby often neglecting diverse and distinctive musical recordings. Musical outliers can provide valuable insights due to the inherent diversity of music itself. In this paper, we explore music outliers, investigating their potential usefulness for music discovery and recommendation systems. We argue that not all outliers should be treated as noise, as they can offer interesting perspectives and contribute to a richer understanding of an artist's work. We introduce the concept of 'Genuine' music outliers and provide a definition for them. These genuine outliers can reveal unique aspects of an artist's repertoire and hold the potential to enhance music discovery by exposing listeners to novel and diverse musical experiences. 【6】 A Novel Bi-LSTM And Transformer Architecture For Generating Tabla Music标题:一种新的Bi—LSTM和Transformer结构的Tabla音乐生成链接:https://arxiv.org/abs/2404.05765作者:Roopa Mayya,Vivekanand Venkataraman,Anwesh P R,Narayana Darapaneni摘要:简介:音乐生成是一项复杂的任务,近年来受到了极大的关注,深度学习技术在这一领域取得了可喜的成果。目的:虽然在生成钢琴和其他西方音乐方面已经进行了大量的工作,但由于机器编码格式的印度音乐的稀缺性,对生成古典印度音乐的研究有限。在这篇技术论文中,提出了产生古典印度音乐,特别是塔布拉音乐的方法。首先,本文探讨了使用深度学习架构生成钢琴音乐。然后将基本原理扩展到生成塔布拉音乐。方法:使用Python中的librosa库对波形(.wav)文件中的Tabla音乐进行预处理。基于提取的特征和标签训练了一种新的具有注意力方法和Transformer模型的Bi-LSTM。结果:然后使用模型来预测下一个序列的塔布拉音乐。Bi-LSTM模型的损失为4.042,MAE为1.0814。在Transformer模型下,产生Tabla音乐的损耗为55.9278,MAE为3.5173。结论:由此产生的音乐体现了新颖性和熟悉性的和谐融合,将音乐创作的极限推向了新的视野。摘要:Introduction: Music generation is a complex task that has received significant attention in recent years, and deep learning techniques have shown promising results in this field. Objectives: While extensive work has been carried out on generating Piano and other Western music, there is limited research on generating classical Indian music due to the scarcity of Indian music in machine-encoded formats. In this technical paper, methods for generating classical Indian music, specifically tabla music, is proposed. Initially, this paper explores piano music generation using deep learning architectures. Then the fundamentals are extended to generating tabla music. Methods: Tabla music in waveform (.wav) files are pre-processed using the librosa library in Python. A novel Bi-LSTM with an Attention approach and a transformer model are trained on the extracted features and labels. Results: The models are then used to predict the next sequences of tabla music. A loss of 4.042 and MAE of 1.0814 are achieved with the Bi-LSTM model. With the transformer model, a loss of 55.9278 and MAE of 3.5173 are obtained for tabla music generation. Conclusion: The resulting music embodies a harmonious fusion of novelty and familiarity, pushing the limits of music composition to new horizons.