本文经arXiv每日学术速递授权转载
【1】 GraphMuse: A Library for Symbolic Music Graph Processing
标题: GraphMuse:符号音乐图形处理库
作者:Emmanouil Karystinaios,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【2】 Audio Conditioning for Music Generation via Discrete Bottleneck Features
标题: 通过离散瓶颈功能实现音乐生成的音频调节
作者:Simon Rouard,Yossi Adi,Jade Copet,Axel Roebel,Alexandre Défossez
备注:6 pages, 2 figures, accepted at ISMIR 2024
链接:点击下载PDF文件
【3】 A Language Modeling Approach to Diacritic-Free Hebrew TTS
标题: 无变音希伯来语TTC的语言建模方法
作者:Amit Roth,Arnon Turetzky,Yossi Adi
备注:Accepted at Interspeech24
链接:点击下载PDF文件
【4】 The Kolmogorov Complexity of Irish traditional dance music
标题: 爱尔兰传统舞曲的科尔莫戈洛夫复杂性
作者:Michael McGettrick,Paul McGettrick
备注:6 pages
链接:点击下载PDF文件
【5】 Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
标题: 重新审视花朵:花朵的初步复制等1997
作者:Kajetan Enge,Liam Fabry,Robert Höldrich
备注:4 pages, 3 figures, to be presented as an extended abstract at the 29th International Conference on Auditory Display (2024) in Troy, New York, USA
链接:点击下载PDF文件
【6】 TTSDS -- Text-to-Speech Distribution Score
标题: TTSDs --文本到语音分布分数
作者:Christoph Minixhofer,Ondřej Klejch,Peter Bell
备注:Under review for SLT 2024
链接:点击下载PDF文件
【7】 PCQ: Emotion Recognition in Speech via Progressive Channel Querying
标题: PCQ:通过渐进式通道查询的语音情感识别
作者:Xincheng Wang,Liejun Wang,Yinfeng Yu,Xinxin Jiao
备注:Accepted for publication by International Conference On Intelligent Computing 2024. For data and code, see this https URL
链接:点击下载PDF文件
标题: TalTech-IRIT-LIS 2024年DISPLACE演讲者和语言扩展系统
作者:Joonas Kalda,Tanel Alumäe,Martin Lebourdais,Hervé Bredin,Séverin Baroudi,Ricard Marxer
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
【2】 TTSDS -- Text-to-Speech Distribution Score
标题: TTSDs --文本到语音分布分数
作者:Christoph Minixhofer,Ondřej Klejch,Peter Bell
备注:Under review for SLT 2024
链接:点击下载PDF文件
【3】 BSC-UPC at EmoSPeech-IberLEF2024: Attention Pooling for Emotion Recognition
标题: BSC-UPC出席CLARSPeech-IberLEF 2024:情感识别的注意力集中
作者:Marc Casals-Salvador,Federico Costa,Miquel India,Javier Hernando
链接:点击下载PDF文件
【4】 PCQ: Emotion Recognition in Speech via Progressive Channel Querying
标题: PCQ:通过渐进式通道查询的语音情感识别
作者:Xincheng Wang,Liejun Wang,Yinfeng Yu,Xinxin Jiao
备注:Accepted for publication by International Conference On Intelligent Computing 2024. For data and code, see this https URL
链接:点击下载PDF文件
【5】 Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
标题: 先笑后哭:控制基于流匹配的Zero-Shot文本到语音的时变情感状态
作者:Haibin Wu,Xiaofei Wang,Sefik Emre Eskimez,Manthan Thakker,Daniel Tompkins,Chung-Hsien Tsai,Canrun Li,Zhen Xiao,Sheng Zhao,Jinyu Li,Naoyuki Kanda
备注:See this https URL for demo samples
链接:点击下载PDF文件
【6】 Semantic Communication for the Internet of Sounds: Architecture, Design Principles, and Challenges
标题: 声音互联网的语义沟通:架构、设计原则和挑战
作者:Chengsi Liang,Yao Sun,Christo Kurisummoottil Thomas,Lina Mohjazi,Walid Saad
链接:点击下载PDF文件
【7】 ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024
标题: ICAGC 2024:鼓舞人心且令人信服的音频生成挑战2024
作者:Ruibo Fu,Rui Liu,Chunyu Qiang,Yingming Gao,Yi Lu,Tao Wang,Ya Li,Zhengqi Wen,Chen Zhang,Hui Bu,Yukun Liu,Shuchen Shi,Xin Qi,Guanjun Li
备注:ISCSLP 2024 Challenge description
链接:点击下载PDF文件
【8】 GraphMuse: A Library for Symbolic Music Graph Processing
标题: GraphMuse:符号音乐图形处理库
作者:Emmanouil Karystinaios,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【9】 Audio Conditioning for Music Generation via Discrete Bottleneck Features
标题: 通过离散瓶颈功能实现音乐生成的音频调节
作者:Simon Rouard,Yossi Adi,Jade Copet,Axel Roebel,Alexandre Défossez
备注:6 pages, 2 figures, accepted at ISMIR 2024
链接:点击下载PDF文件
【10】 A Language Modeling Approach to Diacritic-Free Hebrew TTS
标题: 无变音希伯来语TTC的语言建模方法
作者:Amit Roth,Arnon Turetzky,Yossi Adi
备注:Accepted at Interspeech24
链接:点击下载PDF文件
【11】 The Kolmogorov Complexity of Irish traditional dance music
标题: 爱尔兰传统舞曲的科尔莫戈洛夫复杂性
作者:Michael McGettrick,Paul McGettrick
备注:6 pages
链接:点击下载PDF文件
【12】 Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
标题: 重新审视花朵:花朵的初步复制等1997
作者:Kajetan Enge,Liam Fabry,Robert Höldrich
备注:4 pages, 3 figures, to be presented as an extended abstract at the 29th International Conference on Auditory Display (2024) in Troy, New York, USA
链接:点击下载PDF文件
标题: GraphMuse:符号音乐图形处理库
作者:Emmanouil Karystinaios,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:图神经网络(GNN)最近在符号音乐任务中获得了关注,但缺乏统一的框架阻碍了进展。为了解决这一差距,我们提出了GraphMuse,这是一个图形处理框架和库,可以促进高效的音乐图形处理和符号音乐任务的GNN训练。我们的贡献是一个新的邻居采样技术,专门针对有意义的行为,在乐谱。此外,GraphMuse集成了分层建模元素,增强了音乐任务的图形网络的表现力和能力。两个特定的音乐预测任务的实验-音高拼写和节奏检测-表现出显着的性能改善,比以前的方法。我们希望GraphMuse能够促进基于图形表示的符号音乐处理的标准化。该库可在https: github.com manoskary graphmuse上获得摘要:Graph Neural Networks (GNNs) have recently gained traction in symbolic music tasks, yet a lack of a unified framework impedes progress. Addressing this gap, we present GraphMuse, a graph processing framework and library that facilitates efficient music graph processing and GNN training for symbolic music tasks. Central to our contribution is a new neighbor sampling technique specifically targeted toward meaningful behavior in musical scores. Additionally, GraphMuse integrates hierarchical modeling elements that augment the expressivity and capabilities of graph networks for musical tasks. Experiments with two specific musical prediction tasks -- pitch spelling and cadence detection -- demonstrate significant performance improvement over previous methods. Our hope is that GraphMuse will lead to a boost in, and standardization of, symbolic music processing based on graph representations. The library is available at https: github.com manoskary graphmuse
【2】 Audio Conditioning for Music Generation via Discrete Bottleneck Features
标题: 通过离散瓶颈功能实现音乐生成的音频调节
作者:Simon Rouard,Yossi Adi,Jade Copet,Axel Roebel,Alexandre Défossez
备注:6 pages, 2 figures, accepted at ISMIR 2024
链接:点击下载PDF文件
摘要:虽然大多数音乐生成模型使用文本或参数条件(例如,节奏,和声,音乐流派),我们建议条件的语言模型为基础的音乐生成系统与音频输入。我们的探索涉及两种不同的策略。第一种策略,称为文本反转,利用预先训练的文本到音乐模型将音频输入映射到文本嵌入空间中相应的“伪词”。对于第二个模型,我们从零开始训练一个音乐语言模型与文本调节器和量化音频特征提取器。在推理时,我们可以混合文本和音频条件反射,并通过一种新的双分类器自由指导方法来平衡它们。我们进行自动和人类研究,以验证我们的方法。我们将发布代码,并在https: musicgenstyle.github.io上提供音乐样本,以显示我们模型的质量。摘要:While most music generation models use textual or parametric conditioning (e.g. tempo, harmony, musical genre), we propose to condition a language model based music generation system with audio input. Our exploration involves two distinct strategies. The first strategy, termed textual inversion, leverages a pre-trained text-to-music model to map audio input to corresponding "pseudowords" in the textual embedding space. For the second model we train a music language model from scratch jointly with a text conditioner and a quantized audio feature extractor. At inference time, we can mix textual and audio conditioning and balance them thanks to a novel double classifier free guidance method. We conduct automatic and human studies that validates our approach. We will release the code and we provide music samples on https: musicgenstyle.github.io in order to show the quality of our model.
【3】 A Language Modeling Approach to Diacritic-Free Hebrew TTS
标题: 无变音希伯来语TTC的语言建模方法
作者:Amit Roth,Arnon Turetzky,Yossi Adi
备注:Accepted at Interspeech24
链接:点击下载PDF文件
摘要:我们解决希伯来语的文本到语音(TTS)任务。传统希伯来语包含变音符号,这些变音符号规定了个人对给定单词的发音方式,然而,现代希伯来语很少使用它们。现代希伯来语缺乏变音符号,因此读者需要根据上下文得出正确的发音并理解要使用哪些音素。这对TTS系统在文本到语音之间准确映射提出了根本性挑战。在这项工作中,我们建议采用一种语言建模无变音符号的方法,希伯来文TTS的任务。该模型的离散语音表示和条件上的词片标记。我们优化了所提出的方法,使用在野生弱监督数据,并将其与几个基于变音的TTS系统。结果表明,该方法是优于评估基线同时考虑内容的保护和自然的生成的语音。示例可以在以下链接中找到:pages.cs.huji.ac.il adiyoss-lab HebTTS 摘要:We tackle the task of text-to-speech (TTS) in Hebrew. Traditional Hebrew contains Diacritics, which dictate the way individuals should pronounce given words, however, modern Hebrew rarely uses them. The lack of diacritics in modern Hebrew results in readers expected to conclude the correct pronunciation and understand which phonemes to use based on the context. This imposes a fundamental challenge on TTS systems to accurately map between text-to-speech. In this work, we propose to adopt a language modeling Diacritics-Free approach, for the task of Hebrew TTS. The model operates on discrete speech representations and is conditioned on a word-piece tokenizer. We optimize the proposed method using in-the-wild weakly supervised data and compare it to several diacritic-based TTS systems. Results suggest the proposed method is superior to the evaluated baselines considering both content preservation and naturalness of the generated speech. Samples can be found under the following link: pages.cs.huji.ac.il adiyoss-lab HebTTS
【4】 The Kolmogorov Complexity of Irish traditional dance music
标题: 爱尔兰传统舞曲的科尔莫戈洛夫复杂性
作者:Michael McGettrick,Paul McGettrick
备注:6 pages
链接:点击下载PDF文件
摘要:我们估计柯尔莫哥洛夫复杂性的旋律在爱尔兰传统的舞蹈音乐使用Lempel-Ziv压缩。音乐的“曲调”以所谓的“ABC记谱法”表示,仅仅是字母表中的字母序列:我们没有节奏变化,所有音符都是等长的。我们的算法复杂性的估计,可以用来区分“简单”或“容易”的曲调(更多的重复),从“困难”的(较少的重复),这应该证明是有用的学生学习曲调。我们进一步提出了两个调类别(卷轴和夹具)在其复杂性方面的比较。摘要:We estimate the Kolmogorov complexity of melodies in Irish traditional dance music using Lempel-Ziv compression. The "tunes" of the music are presented in so-called "ABC notation" as simply a sequence of letters from an alphabet: We have no rhythmic variation, with all notes being of equal length. Our estimation of algorithmic complexity can be used to distinguish "simple" or "easy" tunes (with more repetition) from "difficult" ones (with less repetition) which should prove useful for students learning tunes. We further present a comparison of two tune categories (reels and jigs) in terms of their complexity.
【5】 Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
标题: 重新审视花朵:花朵的初步复制等1997
作者:Kajetan Enge,Liam Fabry,Robert Höldrich
备注:4 pages, 3 figures, to be presented as an extended abstract at the 29th International Conference on Auditory Display (2024) in Troy, New York, USA
链接:点击下载PDF文件
摘要:1997年,Flowers,Buhman和Turnage发表了一篇题为“探索双变量数据样本的视觉和听觉散点图的跨模态等效性”的论文。"这篇论文研究了我们评估两个数据变量之间关系的能力,当通过视觉或听觉散点图呈现时。27年后,我们复制了这项有影响力的研究的第一部分,并提出了我们复制的初步结果,最初涉及21名参与者。除了纯粹的听觉和视觉散点图之外,我们还引入了视听散点图作为我们实验中的第三个条件。我们的初步研究结果反映了Flowers等人的研究结果。的原创研究。通过这个扩展的摘要,我们还旨在引发关于复制研究对我们的研究社区的意义的讨论。摘要:In 1997, Flowers, Buhman, and Turnage published a paper titled Cross-Modal Equivalence of Visual and Auditory Scatterplots for Exploring Bivariate Data Samples.'' This paper examined our capacity to assess the relationship between two data variables when presented through visual or auditory scatterplots. Twenty-seven years later, we have replicated the first part of this influential study and present the preliminary findings of our replication, initially involving 21 participants. In addition to purely auditory and visual scatterplots, we introduced audiovisual scatterplots as a third condition in our experiment. Our initial findings mirror those of Flowers et al.'s original research. With this extended abstract, we also aim to spark a discussion about the significance of replication studies for our research community in general.
【6】 TTSDS -- Text-to-Speech Distribution Score
标题: TTSDs --文本到语音分布分数
作者:Christoph Minixhofer,Ondřej Klejch,Peter Bell
备注:Under review for SLT 2024
链接:点击下载PDF文件
摘要:许多最近发布的文本到语音(TTS)系统产生接近真实语音的音频。然而,TTS评估需要重新审视,以使新的架构,方法和数据集所获得的结果的意义。我们建议评估合成语音的质量,如韵律,扬声器的身份和可懂度的多个因素的组合。我们的方法评估如何以及合成语音反映真实的语音,通过获得每个因素的相关性,并测量它们的距离从真实的语音数据集和噪声数据集。我们对2008年至2024年间开发的35个TTS系统进行了基准测试,结果表明,我们的得分计算为因素的未加权平均值,与每个时间段的人类评价密切相关。摘要:Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose evaluating the quality of synthetic speech as a combination of multiple factors such as prosody, speaker identity, and intelligibility. Our approach assesses how well synthetic speech mirrors real speech by obtaining correlates of each factor and measuring their distance from both real speech datasets and noise datasets. We benchmark 35 TTS systems developed between 2008 and 2024 and show that our score computed as an unweighted average of factors strongly correlates with the human evaluations from each time period.
【7】 PCQ: Emotion Recognition in Speech via Progressive Channel Querying
标题: PCQ:通过渐进式通道查询的语音情感识别
作者:Xincheng Wang,Liejun Wang,Yinfeng Yu,Xinxin Jiao
备注:Accepted for publication by International Conference On Intelligent Computing 2024. For data and code, see this https URL
链接:点击下载PDF文件
摘要:在人机交互(HCI)中,语音情感识别(SER)是理解人类意图和情感的关键技术。传统的SER方法难以有效地捕捉复杂情感表达中的长期时间相关性和动态变化。为了克服这些限制,我们引入了PCQ方法,这是一种通过 textbf{P}渐进 textbf{C}通道 textbf{Q}进行SER的开创性方法。该方法通过渠道查询技术,在渠道维度上逐层向下钻取,实现情感长期情境信息的动态建模。这种多层次的分析使PCQ方法在捕捉人类情感的细微差别方面具有优势。实验结果表明,在IEMOCAP和EMODB情感识别数据集上,该模型的加权平均(WA)准确率分别提高了3.98%和3.45%,未加权平均(UA)准确率分别提高了5.67%和5.83%,显著高于基线水平.摘要:In human-computer interaction (HCI), Speech Emotion Recognition (SER) is a key technology for understanding human intentions and emotions. Traditional SER methods struggle to effectively capture the long-term temporal correla-tions and dynamic variations in complex emotional expressions. To overcome these limitations, we introduce the PCQ method, a pioneering approach for SER via textbf{P}rogressive textbf{C}hannel textbf{Q}uerying. This method can drill down layer by layer in the channel dimension through the channel query technique to achieve dynamic modeling of long-term contextual information of emotions. This mul-ti-level analysis gives the PCQ method an edge in capturing the nuances of hu-man emotions. Experimental results show that our model improves the weighted average (WA) accuracy by 3.98 % and 3.45 % and the unweighted av-erage (UA) accuracy by 5.67 % and 5.83 % on the IEMOCAP and EMODB emotion recognition datasets, respectively, significantly exceeding the baseline levels.
eess.AS音频处理
【1】 TalTech-IRIT-LIS Speaker and Language Diarization Systems for DISPLACE 2024标题: TalTech-IRIT-LIS 2024年DISPLACE演讲者和语言扩展系统
作者:Joonas Kalda,Tanel Alumäe,Martin Lebourdais,Hervé Bredin,Séverin Baroudi,Ricard Marxer
备注:accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了TalTech-IRIT-LIS团队提交的DISPLACE 2024挑战。我们的团队参加了演讲者日记和语言日记的挑战。在扬声器日志化轨道中,我们最好的提交是基于pyannote.audio扬声器日志化管道的系统集成,该管道利用powerset训练和我们最近提出的PixIT方法执行联合日志化和语音分离。我们通过使用分离输出进行说话人嵌入提取来改进PixIT。我们的集成在评估数据集上实现了27.1%的日志化错误率。在语言日志化跟踪中,我们对域内数据进行了预训练的Wav 2 Vec 2-BERT语言嵌入模型进行了微调,并基于LDA PLDA的相似性得分使用AHC和VBx对短片段进行了聚类。这导致评估数据的语言日记错误率为27.6%。这两项成绩都在各自的挑战赛中排名第一。摘要:This paper describes the submissions of team TalTech-IRIT-LIS to the DISPLACE 2024 challenge. Our team participated in the speaker diarization and language diarization tracks of the challenge. In the speaker diarization track, our best submission was an ensemble of systems based on the pyannote.audio speaker diarization pipeline utilizing powerset training and our recently proposed PixIT method that performs joint diarization and speech separation. We improve upon PixIT by using the separation outputs for speaker embedding extraction. Our ensemble achieved a diarization error rate of 27.1% on the evaluation dataset. In the language diarization track, we fine-tuned a pre-trained Wav2Vec2-BERT language embedding model on in-domain data, and clustered short segments using AHC and VBx, based on similarity scores from LDA PLDA. This led to a language diarization error rate of 27.6% on the evaluation data. Both results were ranked first in their respective challenge tracks.
【2】 TTSDS -- Text-to-Speech Distribution Score
标题: TTSDs --文本到语音分布分数
作者:Christoph Minixhofer,Ondřej Klejch,Peter Bell
备注:Under review for SLT 2024
链接:点击下载PDF文件
摘要:许多最近发布的文本到语音(TTS)系统产生接近真实语音的音频。然而,TTS评估需要重新审视,以使新的架构,方法和数据集所获得的结果的意义。我们建议评估合成语音的质量,如韵律,扬声器的身份和可懂度的多个因素的组合。我们的方法评估如何以及合成语音反映真实的语音,通过获得每个因素的相关性,并测量它们的距离从真实的语音数据集和噪声数据集。我们对2008年至2024年间开发的35个TTS系统进行了基准测试,结果表明,我们的分数计算为因素的未加权平均值,与每个时间段的人类评估密切相关。摘要:Many recently published Text-to-Speech (TTS) systems produce audio close to real speech. However, TTS evaluation needs to be revisited to make sense of the results obtained with the new architectures, approaches and datasets. We propose evaluating the quality of synthetic speech as a combination of multiple factors such as prosody, speaker identity, and intelligibility. Our approach assesses how well synthetic speech mirrors real speech by obtaining correlates of each factor and measuring their distance from both real speech datasets and noise datasets. We benchmark 35 TTS systems developed between 2008 and 2024 and show that our score computed as an unweighted average of factors strongly correlates with the human evaluations from each time period.
【3】 BSC-UPC at EmoSPeech-IberLEF2024: Attention Pooling for Emotion Recognition
标题: BSC-UPC出席CLARSPeech-IberLEF 2024:情感识别的注意力集中
作者:Marc Casals-Salvador,Federico Costa,Miquel India,Javier Hernando
链接:点击下载PDF文件
摘要:None摘要:The domain of speech emotion recognition (SER) has persistently been a frontier within the landscape of machine learning. It is an active field that has been revolutionized in the last few decades and whose implementations are remarkable in multiple applications that could affect daily life. Consequently, the Iberian Languages Evaluation Forum (IberLEF) of 2024 held a competitive challenge to leverage the SER results with a Spanish corpus. This paper presents the approach followed with the goal of participating in this competition. The main architecture consists of different pre-trained speech and text models to extract features from both modalities, utilizing an attention pooling mechanism. The proposed system has achieved the first position in the challenge with an 86.69% in Macro F1-Score.
【4】 PCQ: Emotion Recognition in Speech via Progressive Channel Querying
标题: PCQ:通过渐进式通道查询的语音情感识别
作者:Xincheng Wang,Liejun Wang,Yinfeng Yu,Xinxin Jiao
备注:Accepted for publication by International Conference On Intelligent Computing 2024. For data and code, see this https URL
链接:点击下载PDF文件
摘要:在人机交互(HCI)中,语音情感识别(SER)是理解人类意图和情感的关键技术。传统的SER方法难以有效地捕捉复杂情感表达中的长期时间相关性和动态变化。为了克服这些限制,我们引入了PCQ方法,这是一种通过 textbf{P}渐进 textbf{C}通道 textbf{Q}进行SER的开创性方法。该方法通过渠道查询技术,在渠道维度上逐层向下钻取,实现情感长期情境信息的动态建模。这种多层次的分析使PCQ方法在捕捉人类情感的细微差别方面具有优势。实验结果表明,在IEMOCAP和EMODB情感识别数据集上,该模型的加权平均(WA)准确率分别提高了3.98%和3.45%,未加权平均(UA)准确率分别提高了5.67%和5.83%,显著高于基线水平.摘要:In human-computer interaction (HCI), Speech Emotion Recognition (SER) is a key technology for understanding human intentions and emotions. Traditional SER methods struggle to effectively capture the long-term temporal correla-tions and dynamic variations in complex emotional expressions. To overcome these limitations, we introduce the PCQ method, a pioneering approach for SER via textbf{P}rogressive textbf{C}hannel textbf{Q}uerying. This method can drill down layer by layer in the channel dimension through the channel query technique to achieve dynamic modeling of long-term contextual information of emotions. This mul-ti-level analysis gives the PCQ method an edge in capturing the nuances of hu-man emotions. Experimental results show that our model improves the weighted average (WA) accuracy by 3.98 % and 3.45 % and the unweighted av-erage (UA) accuracy by 5.67 % and 5.83 % on the IEMOCAP and EMODB emotion recognition datasets, respectively, significantly exceeding the baseline levels.
【5】 Laugh Now Cry Later: Controlling Time-Varying Emotional States of Flow-Matching-Based Zero-Shot Text-to-Speech
标题: 先笑后哭:控制基于流匹配的Zero-Shot文本到语音的时变情感状态
作者:Haibin Wu,Xiaofei Wang,Sefik Emre Eskimez,Manthan Thakker,Daniel Tompkins,Chung-Hsien Tsai,Canrun Li,Zhen Xiao,Sheng Zhao,Jinyu Li,Naoyuki Kanda
备注:See this https URL for demo samples
链接:点击下载PDF文件
摘要:人们改变他们的语调,通常伴随着非语言发声(NV),如笑声和哭声,以传达丰富的情感。然而,大多数文本到语音(TTS)系统缺乏生成具有丰富情感(包括NV)的语音的能力。本文介绍了一种情感可控的zero-shot TTS,它可以为任何说话者生成具有NV的高度情感的语音。TTS利用唤醒和效价值以及笑声嵌入来调节基于流匹配的zero-shot TTS。为了实现高质量的情感语音生成,TSCtrl-TTS使用基于伪标签策划的超过27,000小时的表达数据进行训练。综合评估表明,在语音到语音翻译场景中,语音控制-TTS在模仿音频提示的情感方面表现出色。我们还表明,该TTS可以捕捉情感变化,表达强烈的情感,并产生各种NV在zero-shot TTS。有关演示示例,请访问https: aka.ms emoctrl-tts。摘要:People change their tones of voice, often accompanied by nonverbal vocalizations (NVs) such as laughter and cries, to convey rich emotions. However, most text-to-speech (TTS) systems lack the capability to generate speech with rich emotions, including NVs. This paper introduces EmoCtrl-TTS, an emotion-controllable zero-shot TTS that can generate highly emotional speech with NVs for any speaker. EmoCtrl-TTS leverages arousal and valence values, as well as laughter embeddings, to condition the flow-matching-based zero-shot TTS. To achieve high-quality emotional speech generation, EmoCtrl-TTS is trained using more than 27,000 hours of expressive data curated based on pseudo-labeling. Comprehensive evaluations demonstrate that EmoCtrl-TTS excels in mimicking the emotions of audio prompts in speech-to-speech translation scenarios. We also show that EmoCtrl-TTS can capture emotion changes, express strong emotions, and generate various NVs in zero-shot TTS. See https: aka.ms emoctrl-tts for demo samples.
【6】 Semantic Communication for the Internet of Sounds: Architecture, Design Principles, and Challenges
标题: 声音互联网的语义沟通:架构、设计原则和挑战
作者:Chengsi Liang,Yao Sun,Christo Kurisummoottil Thomas,Lina Mohjazi,Walid Saad
链接:点击下载PDF文件
摘要:声音互联网(IoS)结合了声音感测、处理和传输技术,使各种声音设备之间能够进行协作。为了在IoS中实现声音同步的感知质量,有必要精确地同步三个关键因素:声音质量、定时和行为控制。然而,聚焦于比特再现的传统面向比特的通信可能无法在动态信道条件下满足这些同步要求。解决IoS同步挑战的一种有前途的方法是通过使用语义通信(SC),可以捕获和利用其源数据中的逻辑关系。因此,在本文中,我们提出了一个以物联网为中心的SC框架与收发器的设计。所设计的编码器从不同的源中提取语义信息,并将其传输到IoS侦听器。它还可以提取重要的语义信息,以减少传输延迟的定时同步。在接收端,解码器采用基于上下文和知识的推理技术来重建和整合声音,从而实现不同通信环境下的音质同步。此外,通过周期性地共享知识,IoS设备的SC模型可以被更新以优化它们的同步行为。最后,我们探讨了数学模型,资源分配和跨层协议的几个开放的问题。摘要:The Internet of Sounds (IoS) combines sound sensing, processing, and transmission techniques, enabling collaboration among diverse sound devices. To achieve perceptual quality of sound synchronization in the IoS, it is necessary to precisely synchronize three critical factors: sound quality, timing, and behavior control. However, conventional bit-oriented communication, which focuses on bit reproduction, may not be able to fulfill these synchronization requirements under dynamic channel conditions. One promising approach to address the synchronization challenges of the IoS is through the use of semantic communication (SC) that can capture and leverage the logical relationships in its source data. Consequently, in this paper, we propose an IoS-centric SC framework with a transceiver design. The designed encoder extracts semantic information from diverse sources and transmits it to IoS listeners. It can also distill important semantic information to reduce transmission latency for timing synchronization. At the receiver's end, the decoder employs context- and knowledge-based reasoning techniques to reconstruct and integrate sounds, which achieves sound quality synchronization across diverse communication environments. Moreover, by periodically sharing knowledge, SC models of IoS devices can be updated to optimize their synchronization behavior. Finally, we explore several open issues on mathematical models, resource allocation, and cross-layer protocols.
【7】 ICAGC 2024: Inspirational and Convincing Audio Generation Challenge 2024
标题: ICAGC 2024:鼓舞人心且令人信服的音频生成挑战2024
作者:Ruibo Fu,Rui Liu,Chunyu Qiang,Yingming Gao,Yi Lu,Tao Wang,Ya Li,Zhengqi Wen,Chen Zhang,Hui Bu,Yukun Liu,Shuchen Shi,Xin Qi,Guanjun Li
备注:ISCSLP 2024 Challenge description
链接:点击下载PDF文件
摘要:2024年鼓舞人心和令人信服的音频生成挑战赛(ICAGC 2024)是ISCSLP 2024竞赛和挑战赛的一部分。虽然目前的文本到语音(TTS)技术可以生成高质量的音频,但其传达复杂情感和受控细节内容的能力仍然有限。这种约束导致在实际应用中生成的音频与人类主观感知之间存在差异,例如儿童伴侣机器人和营销机器人。核心问题在于高质量音频生成与最终人类主观体验之间的不一致。因此,这项挑战旨在增强合成音频的说服力和可接受性,重点是人类对齐令人信服和鼓舞人心的音频生成。摘要:The Inspirational and Convincing Audio Generation Challenge 2024 (ICAGC 2024) is part of the ISCSLP 2024 Competitions and Challenges track. While current text-to-speech (TTS) technology can generate high-quality audio, its ability to convey complex emotions and controlled detail content remains limited. This constraint leads to a discrepancy between the generated audio and human subjective perception in practical applications like companion robots for children and marketing bots. The core issue lies in the inconsistency between high-quality audio generation and the ultimate human subjective experience. Therefore, this challenge aims to enhance the persuasiveness and acceptability of synthesized audio, focusing on human alignment convincing and inspirational audio generation.
【8】 GraphMuse: A Library for Symbolic Music Graph Processing
标题: GraphMuse:符号音乐图形处理库
作者:Emmanouil Karystinaios,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:图神经网络(GNN)最近在符号音乐任务中获得了关注,但缺乏统一的框架阻碍了进展。为了解决这一差距,我们提出了GraphMuse,这是一个图形处理框架和库,可以促进高效的音乐图形处理和符号音乐任务的GNN训练。我们的贡献是一个新的邻居采样技术,专门针对有意义的行为,在乐谱。此外,GraphMuse集成了分层建模元素,增强了音乐任务的图形网络的表现力和能力。两个特定的音乐预测任务的实验-音高拼写和节奏检测-表现出显着的性能改善,比以前的方法。我们希望GraphMuse能够促进基于图形表示的符号音乐处理的标准化。该库可在www.example.com上获得摘要:Graph Neural Networks (GNNs) have recently gained traction in symbolic music tasks, yet a lack of a unified framework impedes progress. Addressing this gap, we present GraphMuse, a graph processing framework and library that facilitates efficient music graph processing and GNN training for symbolic music tasks. Central to our contribution is a new neighbor sampling technique specifically targeted toward meaningful behavior in musical scores. Additionally, GraphMuse integrates hierarchical modeling elements that augment the expressivity and capabilities of graph networks for musical tasks. Experiments with two specific musical prediction tasks -- pitch spelling and cadence detection -- demonstrate significant performance improvement over previous methods. Our hope is that GraphMuse will lead to a boost in, and standardization of, symbolic music processing based on graph representations. The library is available at https: github.com manoskary graphmuse
【9】 Audio Conditioning for Music Generation via Discrete Bottleneck Features
标题: 通过离散瓶颈功能实现音乐生成的音频调节
作者:Simon Rouard,Yossi Adi,Jade Copet,Axel Roebel,Alexandre Défossez
备注:6 pages, 2 figures, accepted at ISMIR 2024
链接:点击下载PDF文件
摘要:虽然大多数音乐生成模型使用文本或参数条件(例如,节奏,和声,音乐流派),我们建议条件的语言模型为基础的音乐生成系统与音频输入。我们的探索涉及两种不同的策略。第一种策略,称为文本反转,利用预先训练的文本到音乐模型将音频输入映射到文本嵌入空间中相应的“伪词”。对于第二个模型,我们从零开始训练一个音乐语言模型与文本调节器和量化音频特征提取器。在推理时,我们可以混合文本和音频条件反射,并通过一种新的双分类器自由指导方法来平衡它们。我们进行自动和人类研究,以验证我们的方法。我们将发布代码,并在https: musicgenstyle.github.io上提供音乐样本,以显示我们模型的质量。摘要:While most music generation models use textual or parametric conditioning (e.g. tempo, harmony, musical genre), we propose to condition a language model based music generation system with audio input. Our exploration involves two distinct strategies. The first strategy, termed textual inversion, leverages a pre-trained text-to-music model to map audio input to corresponding "pseudowords" in the textual embedding space. For the second model we train a music language model from scratch jointly with a text conditioner and a quantized audio feature extractor. At inference time, we can mix textual and audio conditioning and balance them thanks to a novel double classifier free guidance method. We conduct automatic and human studies that validates our approach. We will release the code and we provide music samples on https: musicgenstyle.github.io in order to show the quality of our model.
【10】 A Language Modeling Approach to Diacritic-Free Hebrew TTS
标题: 无变音希伯来语TTC的语言建模方法
作者:Amit Roth,Arnon Turetzky,Yossi Adi
备注:Accepted at Interspeech24
链接:点击下载PDF文件
摘要:我们处理希伯来语的文本到语音(TTS)的任务。传统希伯来语包含变音符号,它规定了个人应该如何发音,然而,现代希伯来语很少使用它们。现代希伯来语缺乏变音符号,导致读者期望根据上下文得出正确的发音并理解使用哪些音素。这对TTS系统提出了一个基本的挑战,即在文本到语音之间准确地映射。在这项工作中,我们建议采用一种语言建模无变音符号的方法,希伯来文TTS的任务。该模型的离散语音表示和条件上的词片标记。我们优化了所提出的方法,使用在野生弱监督数据,并将其与几个基于变音的TTS系统。结果表明,该方法是优于评估基线同时考虑内容的保护和自然的生成的语音。示例可以在以下链接中找到:pages.cs.huji.ac.il adiyoss-lab HebTTS 摘要:We tackle the task of text-to-speech (TTS) in Hebrew. Traditional Hebrew contains Diacritics, which dictate the way individuals should pronounce given words, however, modern Hebrew rarely uses them. The lack of diacritics in modern Hebrew results in readers expected to conclude the correct pronunciation and understand which phonemes to use based on the context. This imposes a fundamental challenge on TTS systems to accurately map between text-to-speech. In this work, we propose to adopt a language modeling Diacritics-Free approach, for the task of Hebrew TTS. The model operates on discrete speech representations and is conditioned on a word-piece tokenizer. We optimize the proposed method using in-the-wild weakly supervised data and compare it to several diacritic-based TTS systems. Results suggest the proposed method is superior to the evaluated baselines considering both content preservation and naturalness of the generated speech. Samples can be found under the following link: pages.cs.huji.ac.il adiyoss-lab HebTTS
【11】 The Kolmogorov Complexity of Irish traditional dance music
标题: 爱尔兰传统舞曲的科尔莫戈洛夫复杂性
作者:Michael McGettrick,Paul McGettrick
备注:6 pages
链接:点击下载PDF文件
摘要:我们估计柯尔莫哥洛夫复杂性的旋律在爱尔兰传统的舞蹈音乐使用Lempel-Ziv压缩。音乐的“曲调”以所谓的“ABC记谱法”表示,仅仅是字母表中的字母序列:我们没有节奏变化,所有音符都是等长的。我们的算法复杂性的估计,可以用来区分“简单”或“容易”的曲调(更多的重复),从“困难”的(较少的重复),这应该证明是有用的学生学习曲调。我们进一步提出了两个调类别(卷轴和夹具)在其复杂性方面的比较。摘要:We estimate the Kolmogorov complexity of melodies in Irish traditional dance music using Lempel-Ziv compression. The "tunes" of the music are presented in so-called "ABC notation" as simply a sequence of letters from an alphabet: We have no rhythmic variation, with all notes being of equal length. Our estimation of algorithmic complexity can be used to distinguish "simple" or "easy" tunes (with more repetition) from "difficult" ones (with less repetition) which should prove useful for students learning tunes. We further present a comparison of two tune categories (reels and jigs) in terms of their complexity.
【12】 Flowers Revisited: A Preliminary Replication of Flowers et al. 1997
标题: 重新审视花朵:花朵的初步复制等1997
作者:Kajetan Enge,Liam Fabry,Robert Höldrich
备注:4 pages, 3 figures, to be presented as an extended abstract at the 29th International Conference on Auditory Display (2024) in Troy, New York, USA
链接:点击下载PDF文件
摘要:1997年,Flowers,Buhman和Turnage发表了一篇题为“探索双变量数据样本的视觉和听觉散点图的跨模态等效性”的论文。"这篇论文研究了我们评估两个数据变量之间关系的能力,当通过视觉或听觉散点图呈现时。27年后,我们复制了这项有影响力的研究的第一部分,并提出了我们复制的初步结果,最初涉及21名参与者。除了纯粹的听觉和视觉散点图之外,我们还引入了视听散点图作为我们实验中的第三个条件。我们的初步研究结果反映了Flowers等人的研究结果。的原创研究。通过这个扩展的摘要,我们还旨在引发关于复制研究对我们的研究社区的意义的讨论。摘要:In 1997, Flowers, Buhman, and Turnage published a paper titled Cross-Modal Equivalence of Visual and Auditory Scatterplots for Exploring Bivariate Data Samples.'' This paper examined our capacity to assess the relationship between two data variables when presented through visual or auditory scatterplots. Twenty-seven years later, we have replicated the first part of this influential study and present the preliminary findings of our replication, initially involving 21 participants. In addition to purely auditory and visual scatterplots, we introduced audiovisual scatterplots as a third condition in our experiment. Our initial findings mirror those of Flowers et al.'s original research. With this extended abstract, we also aim to spark a discussion about the significance of replication studies for our research community in general.
机器翻译,仅供参考
