本文经arXiv每日学术速递授权转载
【1】 AudioBERT: Audio Knowledge Augmented Language Model
标题: AudioBERT:音频知识增强语言模型
作者:Hyunjong Ok,Suho Yoo,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
【2】 The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
标题: Faetar基准:资源严重不足的语言中的语音识别
作者:Michael Ong,Sean Robertson,Leo Peckham,Alba Jorquera Jimenez de Aberasturi,Paula Arkhangorodsky,Robin Huo,Aman Sakhardande,Mark Hallap,Naomi Nagy,Ewan Dunbar
链接:点击下载PDF文件
【3】 Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
标题: Zero-Shot歌唱语音转换:基于集群的音素表示构建
作者:Wangjin Zhou,Fengrun Zhang,Yiming Liu,Wenhao Guan,Yi Zhao,He Qu
链接:点击下载PDF文件
【4】 Tidal MerzA: Combining affective modelling and autonomous code generation through Reinforcement Learning
标题: Tidal MerzA:通过强化学习结合情感建模和自主代码生成
作者:Elizabeth Wilson,György Fazekas,Geraint Wiggins
链接:点击下载PDF文件
【5】 A corpus-based investigation of pitch contours of monosyllabic words in conversational Taiwan Mandarin
标题: 台湾普通话对话中单音节词音调轮廓的研究
作者:Xiaoyun Jin,Mirjam Ernestus,R. Harald Baayen
链接:点击下载PDF文件
【6】 TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
标题: TSELM:使用离散标记和语言模型的目标说话人提取
作者:Beilong Tang,Bang Zeng,Ming Li
链接:点击下载PDF文件
【7】 Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
标题: 绘画与音乐的桥梁--通过绘画探索基于情感的音乐生成
作者:Tanisha Hisariya,Huan Zhang,Jinhua Liang
链接:点击下载PDF文件
【8】 FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
标题: FlowSep:采用整流流匹配的阈值查询声音分离
作者:Yi Yuan,Xubo Liu,Haohe Liu,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
【9】 Flexible Control in Symbolic Music Generation via Musical Metadata
标题: 通过音乐元数据灵活控制象征性音乐生成
作者:Sangjun Han,Jiwon Ham,Chaeeun Lee,Heejin Kim,Soojong Do,Sihyuk Yi,Jun Seo,Seoyoon Kim,Yountae Jung,Woohyung Lim
链接:点击下载PDF文件
【10】 Efficient Sparse Coding with the Adaptive Locally Competitive Algorithm for Speech Classification
标题: 采用自适应局部竞争算法的高效稀疏编码语音分类
作者:Soufiyan Bahadi,Eric Plourde,Jean Rouat
链接:点击下载PDF文件
【11】 Faster Speech-LLaMA Inference with Multi-token Prediction
标题: 更快的语音-具有多令牌预测的LLaMA推理
作者:Desh Raj,Gil Keren,Junteng Jia,Jay Mahadeokar,Ozlem Kalinli
备注:Submitted to IEEE ICASSP 2025
链接:点击下载PDF文件
【12】 Audio Decoding by Inverse Problem Solving
标题: 反问题求解的音频解码
作者:Pedro J. Villasana T.,Lars Villemoes,Janusz Klejsa,Per Hedelin
备注:5 pages, 4 figures, audio demo available at this https URL, pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
【13】 Music auto-tagging in the long tail: A few-shot approach
标题: 长尾音乐自动标记:几次拍摄的方法
作者:T. Aleksandra Ma,Alexander Lerch
备注:Published in Audio Engineering Society NY Show 2024 as a Peer Reviewed (Category 1) paper
链接:点击下载PDF文件
【14】 SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
标题: SSR-Speech:迈向稳定、安全和稳健的Zero-Shot基于文本的语音编辑和合成
作者:Helin Wang,Meng Yu,Jiarui Hai,Chen Chen,Yuchen Hu,Rilin Chen,Najim Dehak,Dong Yu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
标题: 采用自适应局部竞争算法的高效稀疏编码语音分类
作者:Soufiyan Bahadi,Eric Plourde,Jean Rouat
链接:点击下载PDF文件
【2】 Hierarchical Symbolic Pop Music Generation with Graph Neural Networks
标题: 利用图神经网络实现分层符号流行音乐生成
作者:Wen Qing Lim,Jinhua Liang,Huan Zhang
链接:点击下载PDF文件
【3】 Dark Experience for Incremental Keyword Spotting
标题: 增量关键词发现的黑暗体验
作者:Tianyi Peng,Yang Xiao
备注:submitted ICASSP 2025
链接:点击下载PDF文件
【4】 Faster Speech-LLaMA Inference with Multi-token Prediction
标题: 更快的语音-具有多令牌预测的LLaMA推理
作者:Desh Raj,Gil Keren,Junteng Jia,Jay Mahadeokar,Ozlem Kalinli
备注:Submitted to IEEE ICASSP 2025
链接:点击下载PDF文件
【5】 Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
标题: 自动地标:用于地标提取的声学地标数据集和开源工具包
作者:Xiangyu Zhang,Daijiao Liu,Tianyi Xiao,Cihan Xiao,Tuende Szalay,Mostafa Shahin,Beena Ahmed,Julien Epps
链接:点击下载PDF文件
【6】 Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion Models
标题: 利用扩散模型检测和防御自动语音识别中的对抗攻击
作者:Nikolai L. Kühne,Astrid H. F. Kitchen,Marie S. Jensen,Mikkel S. L. Brøndt,Martin Gonzalez,Christophe Biscio,Zheng-Hua Tan
备注:Under review at ICASSP 2025
链接:点击下载PDF文件
【7】 Audio Decoding by Inverse Problem Solving
标题: 反问题求解的音频解码
作者:Pedro J. Villasana T.,Lars Villemoes,Janusz Klejsa,Per Hedelin
备注:5 pages, 4 figures, audio demo available at this https URL, pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 Universal Pooling Method of Multi-layer Features from Pretrained Models for Speaker Verification
标题: 用于说话人验证的预训练模型多层特征的通用池化方法
作者:Jin Sob Kim,Hyun Joon Park,Wooseok Shin,Sung Won Han
备注:Preprint
链接:点击下载PDF文件
【9】 Music auto-tagging in the long tail: A few-shot approach
标题: 长尾音乐自动标记:几次拍摄的方法
作者:T. Aleksandra Ma,Alexander Lerch
备注:Published in Audio Engineering Society NY Show 2024 as a Peer Reviewed (Category 1) paper
链接:点击下载PDF文件
【10】 Super Monotonic Alignment Search
标题: 超级单调对齐搜索
作者:Junhyeok Lee,Hyeongju Kim
备注:Technical Report
链接:点击下载PDF文件
【11】 SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
标题: SSR-Speech:迈向稳定、安全和稳健的Zero-Shot基于文本的语音编辑和合成
作者:Helin Wang,Meng Yu,Jiarui Hai,Chen Chen,Yuchen Hu,Rilin Chen,Najim Dehak,Dong Yu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【12】 AudioBERT: Audio Knowledge Augmented Language Model
标题: AudioBERT:音频知识增强语言模型
作者:Hyunjong Ok,Suho Yoo,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
【13】 The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
标题: Faetar基准:资源严重不足的语言中的语音识别
作者:Michael Ong,Sean Robertson,Leo Peckham,Alba Jorquera Jimenez de Aberasturi,Paula Arkhangorodsky,Robin Huo,Aman Sakhardande,Mark Hallap,Naomi Nagy,Ewan Dunbar
链接:点击下载PDF文件
【14】 Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
标题: Zero-Shot歌唱语音转换:基于集群的音素表示构建
作者:Wangjin Zhou,Fengrun Zhang,Yiming Liu,Wenhao Guan,Yi Zhao,He Qu
链接:点击下载PDF文件
【15】 Tidal MerzA: Combining affective modelling and autonomous code generation through Reinforcement Learning
标题: Tidal MerzA:通过强化学习结合情感建模和自主代码生成
作者:Elizabeth Wilson,György Fazekas,Geraint Wiggins
链接:点击下载PDF文件
【16】 A corpus-based investigation of pitch contours of monosyllabic words in conversational Taiwan Mandarin
标题: 台湾普通话对话中单音节词音调轮廓的研究
作者:Xiaoyun Jin,Mirjam Ernestus,R. Harald Baayen
链接:点击下载PDF文件
【17】 Graph Neural Networks for Parkinsons Disease Detection
标题: 用于帕金森病检测的图神经网络
作者:Shakeel A. Sheikh,Yacouba Kaloga,Ina Kodrasi
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【18】 TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
标题: TSELM:使用离散标记和语言模型的目标说话人提取
作者:Beilong Tang,Bang Zeng,Ming Li
链接:点击下载PDF文件
【19】 Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
标题: 绘画与音乐的桥梁--通过绘画探索基于情感的音乐生成
作者:Tanisha Hisariya,Huan Zhang,Jinhua Liang
链接:点击下载PDF文件
【20】 Full-text Error Correction for Chinese Speech Recognition with Large Language Model
标题: 大语言模型中文语音识别的全文错误纠正
作者:Zhiyuan Tang,Dong Wang,Shen Huang,Shidong Shang
链接:点击下载PDF文件
【21】 FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
标题: FlowSep:采用整流流匹配的阈值查询声音分离
作者:Yi Yuan,Xubo Liu,Haohe Liu,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
【22】 Flexible Control in Symbolic Music Generation via Musical Metadata
标题: 通过音乐元数据灵活控制象征性音乐生成
作者:Sangjun Han,Jiwon Ham,Chaeeun Lee,Heejin Kim,Soojong Do,Sihyuk Yi,Jun Seo,Seoyoon Kim,Yountae Jung,Woohyung Lim
链接:点击下载PDF文件
【23】 Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
标题: 在真正的低资源环境中改进视觉化关键字本地化
作者:Leanne Nortje,Dan Oneata,Herman Kamper
链接:点击下载PDF文件
标题: AudioBERT:音频知识增强语言模型
作者:Hyunjong Ok,Suho Yoo,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
摘要:最近的研究发现,在纯文本数据集上预训练的语言模型通常缺乏基本的视觉知识, textit{e.g.}日常物品的颜色。受此观察的启发,我们问是否存在类似的缺点在 textit{听觉}知识。为了回答这个问题,我们构建了一个名为AuditoryBench的新数据集,它包括两个用于评估听觉知识的新任务。基于我们使用基准的分析,我们发现语言模型也严重缺乏听觉知识。为了解决这个问题,我们提出了AudioBERT,一种新的方法来增加听觉知识的BERT通过检索为基础的方法。首先,我们检测听觉知识跨度提示查询我们的检索模型有效。然后,我们将音频知识注入到BERT中,并在需要音频知识时打开低秩自适应以进行有效的自适应。我们的实验表明,AudioBERT是非常有效的,在AuditoryBench上实现了卓越的性能。数据集和代码可以在 bulurl{https: github.com HJ-Ok AudioBERT}上找到。摘要:Recent studies have identified that language models, pretrained on text-only datasets, often lack elementary visual knowledge, textit{e.g.,} colors of everyday objects. Motivated by this observation, we ask whether a similar shortcoming exists in terms of the textit{auditory} knowledge. To answer this question, we construct a new dataset called AuditoryBench, which consists of two novel tasks for evaluating auditory knowledge. Based on our analysis using the benchmark, we find that language models also suffer from a severe lack of auditory knowledge. To address this limitation, we propose AudioBERT, a novel method to augment the auditory knowledge of BERT through a retrieval-based approach. First, we detect auditory knowledge spans in prompts to query our retrieval model efficiently. Then, we inject audio knowledge into BERT and switch on low-rank adaptation for effective adaptation when audio knowledge is required. Our experiments demonstrate that AudioBERT is quite effective, achieving superior performance on the AuditoryBench. The dataset and code are available at bulurl{https: github.com HJ-Ok AudioBERT}.
【2】 The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
标题: Faetar基准:资源严重不足的语言中的语音识别
作者:Michael Ong,Sean Robertson,Leo Peckham,Alba Jorquera Jimenez de Aberasturi,Paula Arkhangorodsky,Robin Huo,Aman Sakhardande,Mark Hallap,Naomi Nagy,Ewan Dunbar
链接:点击下载PDF文件
摘要:我们介绍Faetar自动语音识别基准,一个基准语料库,旨在推动目前的低资源语音识别方法的限制。Faetar是一种主要在意大利使用的法语变体,没有标准的正字法,除了基准中包含的内容之外,几乎没有现有的文本或语音资源,并且与其他形式的法语变体有很大的不同。语料库来自现场录音,其中大部分是嘈杂的,其中只有5个小时具有匹配的音标,并且强制对准具有可变的质量。语料库包含额外的20小时的未标记的语音。我们报告了最先进的多语言语音基础模型的基线结果,最佳电话错误率为30.4%,使用管道继续使用未标记集对基础模型进行预训练。摘要:We introduce the Faetar Automatic Speech Recognition Benchmark, a benchmark corpus designed to push the limits of current approaches to low-resource speech recognition. Faetar, a Franco-Proven c{c}al variety spoken primarily in Italy, has no standard orthography, has virtually no existing textual or speech resources other than what is included in the benchmark, and is quite different from other forms of Franco-Proven c{c}al. The corpus comes from field recordings, most of which are noisy, for which only 5 hrs have matching transcriptions, and for which forced alignment is of variable quality. The corpus contains an additional 20 hrs of unlabelled speech. We report baseline results from state-of-the-art multilingual speech foundation models with a best phone error rate of 30.4%, using a pipeline that continues pre-training on the foundation model using the unlabelled set.
【3】 Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
标题: Zero-Shot歌唱语音转换:基于集群的音素表示构建
作者:Wangjin Zhou,Fengrun Zhang,Yiming Liu,Wenhao Guan,Yi Zhao,He Qu
链接:点击下载PDF文件
摘要:本研究提出了一种创新的Zero-Shot任意到任意歌唱语音转换(SVC)方法,利用一种新的基于聚类的音素表示有效地分离内容,音色和演唱风格。这种方法能够实现精确的语音特性操纵。我们发现,每个艺术家的录音较少的数据集更容易受到音色泄漏的影响。对超过10,000小时的演唱和用户反馈进行的广泛测试表明,我们的模型显着提高了音质和音色准确性,符合我们的目标,并推进了语音转换技术。此外,本研究提出了zero-shot SVC,并为未来的离散语音表示工作奠定了基础,强调了押韵的保留。摘要:This study presents an innovative Zero-Shot any-to-any Singing Voice Conversion (SVC) method, leveraging a novel clustering-based phoneme representation to effectively separate content, timbre, and singing style. This approach enables precise voice characteristic manipulation. We discovered that datasets with fewer recordings per artist are more susceptible to timbre leakage. Extensive testing on over 10,000 hours of singing and user feedback revealed our model significantly improves sound quality and timbre accuracy, aligning with our objectives and advancing voice conversion technology. Furthermore, this research advances zero-shot SVC and sets the stage for future work on discrete speech representation, emphasizing the preservation of rhyme.
【4】 Tidal MerzA: Combining affective modelling and autonomous code generation through Reinforcement Learning
标题: Tidal MerzA:通过强化学习结合情感建模和自主代码生成
作者:Elizabeth Wilson,György Fazekas,Geraint Wiggins
链接:点击下载PDF文件
摘要:本文介绍了Tidal-MerzA,这是一种新的系统,旨在在现场编码的背景下,人类和机器代理之间的协作表演,特别关注音乐模式的生成。Tidal-MerzA融合了两个基本模型:ALCAA(情感实时编码自治代理)和Tidal Fuzz,一个计算框架。通过将情感建模与计算生成相结合,该系统利用强化学习技术在TidalCycles框架内动态调整音乐作曲参数,确保模式的情感质量和语法正确性。Tidal-MerzA的开发引入了两个不同的代理:一个专注于生成用于音乐表达的迷你符号串,另一个专注于通过强化学习将音乐与目标情感状态对齐。这种方法增强了实时编码实践的适应性和创造性潜力,并允许探索人机创造性交互。Tidal-MerzA推进了计算音乐生成领域,提出了一种将人工智能融入艺术实践的新方法。摘要:This paper presents Tidal-MerzA, a novel system designed for collaborative performances between humans and a machine agent in the context of live coding, specifically focusing on the generation of musical patterns. Tidal-MerzA fuses two foundational models: ALCAA (Affective Live Coding Autonomous Agent) and Tidal Fuzz, a computational framework. By integrating affective modelling with computational generation, this system leverages reinforcement learning techniques to dynamically adapt music composition parameters within the TidalCycles framework, ensuring both affective qualities to the patterns and syntactical correctness. The development of Tidal-MerzA introduces two distinct agents: one focusing on the generation of mini-notation strings for musical expression, and another on the alignment of music with targeted affective states through reinforcement learning. This approach enhances the adaptability and creative potential of live coding practices and allows exploration of human-machine creative interactions. Tidal-MerzA advances the field of computational music generation, presenting a novel methodology for incorporating artificial intelligence into artistic practices.
【5】 A corpus-based investigation of pitch contours of monosyllabic words in conversational Taiwan Mandarin
标题: 台湾普通话对话中单音节词音调轮廓的研究
作者:Xiaoyun Jin,Mirjam Ernestus,R. Harald Baayen
链接:点击下载PDF文件
摘要:汉语单音节词的声调轮廓由四个声调组成:上声(T1)、升调(T2)、降调(T3)和降调(T4)。然而,在自发语音中,单音节词的实际音调实现可能会由于与相邻音调的音节内协同发音和音节间协同发音而显著偏离这些规范音调。此外,Chuang et al.(2024)最近报道,具有T2-T4声调模式的汉语双音节词的声调轮廓由它们的意义共同决定。在此基础上,本文利用语料库研究了汉语自然会话中单音节词的音高轮廓是如何实现的,一方面着重于语境预测因素的影响,另一方面着重于词义共同决定音高轮廓的方式。我们分析了自发的台湾普通话语料库中的3824个标记的63个不同的词类型的F0轮廓,使用广义加法(混合)模型分解成一组组件的音高轮廓观察到的音高轮廓。我们表明,音调的背景下,大大修改一个词的规范语气。一旦控制了音调环境的影响,T2和T3就以低平声出现,与T1作为高音形成对比,而T4作为高中降调。在标准描述中,基于前一个音调实现的轻声(T0)本身作为低音出现,以与标准音调T1、T2、T3和T4相同的方式被其他预测因子修改。我们还表明,词,甚至更重要的是,词的意义,共同决定词的F0轮廓。使用随机森林变量的重要性分析进一步支持音调上下文和词义的效果的实质性影响。摘要:In Mandarin, the tonal contours of monosyllabic words produced in isolation or in careful speech are characterized by four lexical tones: a high-level tone (T1), a rising tone (T2), a dipping tone (T3) and a falling tone (T4). However, in spontaneous speech, the actual tonal realization of monosyllabic words can deviate significantly from these canonical tones due to intra-syllabic co-articulation and inter-syllabic co-articulation with adjacent tones. In addition, Chuang et al. (2024) recently reported that the tonal contours of disyllabic Mandarin words with T2-T4 tone pattern are co-determined by their meanings. Following up on their research, we present a corpus-based investigation of how the pitch contours of monosyllabic words are realized in spontaneous conversational Mandarin, focusing on the effects of contextual predictors on the one hand, and the way in words' meanings co-determine pitch contours on the other hand. We analyze the F0 contours of 3824 tokens of 63 different word types in a spontaneous Taiwan Mandarin corpus, using the generalized additive (mixed) model to decompose a given observed pitch contour into a set of component pitch contours. We show that the tonal context substantially modify a word's canonical tone. Once the effect of tonal context is controlled for, T2 and T3 emerge as low flat tones, contrasting with T1 as a high tone, and with T4 as a high-to-mid falling tone. The neutral tone (T0), which in standard descriptions, is realized based on the preceding tone, emerges as a low tone in its own right, modified by the other predictors in the same way as the standard tones T1, T2, T3, and T4. We also show that word, and even more so, word sense, co-determine words' F0 contours. Analyses of variable importance using random forests further supported the substantial effect of tonal context and an effect of word sense.
【6】 TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
标题: TSELM:使用离散标记和语言模型的目标说话人提取
作者:Beilong Tang,Bang Zeng,Ming Li
链接:点击下载PDF文件
摘要:我们提出了TSELM,一种新的目标说话人提取网络,利用离散的令牌和语言模型。TSELM利用WavLM的多个离散层作为输入标记,并结合交叉注意机制来整合目标说话人信息。语言模型用于捕获序列依赖关系,而可扩展的HiFi-GAN用于从令牌重建音频。通过应用交叉熵损失,TSELM对输出令牌的概率分布进行建模,从而将音频生成的复杂回归问题转换为分类任务。实验结果表明,TSELM在语音质量和可懂度方面取得了很好的效果。摘要:We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate target speaker information. Language models are employed to capture the sequence dependencies, while a scalable HiFi-GAN is used to reconstruct the audio from the tokens. By applying a cross-entropy loss, TSELM models the probability distribution of output tokens, thus converting the complex regression problem of audio generation into a classification task. Experimental results show that TSELM achieves excellent results in speech quality and comparable results in speech intelligibility.
【7】 Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
标题: 绘画与音乐的桥梁--通过绘画探索基于情感的音乐生成
作者:Tanisha Hisariya,Huan Zhang,Jinhua Liang
链接:点击下载PDF文件
摘要:人工智能的快速发展大大增强了涉及音乐和图像的生成任务,采用了单峰和多峰方法。本研究开发了一种能够生成与视觉艺术中所描绘的情感产生共鸣的音乐的模型,将情感标签,图像字幕和语言模型集成在一起,将视觉输入转换为音乐作品。为了解决艺术和音乐数据的稀缺性,我们策划了情感绘画音乐数据集,将绘画与相应的音乐配对,以进行有效的训练和评估。我们的双阶段框架将图像转换为情感内容的文本描述,然后将这些描述转换为音乐,以最少的数据促进高效学习。性能使用诸如Fr 'echet音频距离(FAD)、总谐波失真(THD)、初始分数(IS)和KL发散度等指标进行评估,其中音频-情感文本相似性由预训练的CLAP模型确认,以证明生成的音乐和文本之间的高度对齐。这一综合工具将视觉艺术和音乐联系起来,通过提供丰富的多感官体验,增强了视障者的无障碍环境,并开辟了教育和治疗应用的途径。摘要:Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that resonates with the emotions depicted in visual arts, integrating emotion labeling, image captioning, and language models to transform visual inputs into musical compositions. Addressing the scarcity of aligned art and music data, we curated the Emotion Painting Music Dataset, pairing paintings with corresponding music for effective training and evaluation. Our dual-stage framework converts images to text descriptions of emotional content and then transforms these descriptions into music, facilitating efficient learning with minimal data. Performance is evaluated using metrics such as Fr 'echet Audio Distance (FAD), Total Harmonic Distortion (THD), Inception Score (IS), and KL divergence, with audio-emotion text similarity confirmed by the pre-trained CLAP model to demonstrate high alignment between generated music and text. This synthesis tool bridges visual art and music, enhancing accessibility for the visually impaired and opening avenues in educational and therapeutic applications by providing enriched multi-sensory experiences.
【8】 FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
标题: FlowSep:采用整流流匹配的阈值查询声音分离
作者:Yi Yuan,Xubo Liu,Haohe Liu,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
摘要:语音查询音频源分离(LASS)的重点是分离声音使用所需的源的文本描述。目前的方法主要使用判别方法,如时频掩蔽,以分离目标声音,并最大限度地减少来自其他来源的干扰。然而,这些模型在分离重叠的音轨时面临挑战,这可能导致诸如频谱孔或不完全分离的伪像。整流匹配(RFM),一种生成模型,建立数据和噪声的分布之间的线性关系,提供了优越的理论性能和简单性,但尚未在声音分离探索。在这项工作中,我们介绍FlowSep,一个新的生成模型RFM的基础上LASS任务。FlowSep在变分自动编码器(VAE)潜在空间内从噪声到目标源特征学习线性流轨迹。在推理过程中,RFM生成的潜在特征通过预先训练的VAE解码器重建成梅尔频谱图,然后通过预先训练的声码器合成波形。经过1,680小时的音频数据训练,FlowSep在多个基准测试中的表现优于最先进的模型,并通过主观和客观指标进行评估。此外,我们的研究结果表明,FlowSep在分离质量和推理效率方面都优于基于扩散的LASS模型,突出了其在音频源分离任务中的强大潜力。代码、预训练模型和演示可以在https: audio-agi.github.io FlowSep_demo 上找到。摘要:Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these models face challenges when separating overlapping soundtracks, which may lead to artifacts such as spectral holes or incomplete separation. Rectified flow matching (RFM), a generative model that establishes linear relations between the distribution of data and noise, offers superior theoretical properties and simplicity, but has not yet been explored in sound separation. In this work, we introduce FlowSep, a new generative model based on RFM for LASS tasks. FlowSep learns linear flow trajectories from noise to target source features within the variational autoencoder (VAE) latent space. During inference, the RFM-generated latent features are reconstructed into a mel-spectrogram via the pre-trained VAE decoder, followed by a pre-trained vocoder to synthesize the waveform. Trained on 1,680 hours of audio data, FlowSep outperforms the state-of-the-art models across multiple benchmarks, as evaluated with subjective and objective metrics. Additionally, our results show that FlowSep surpasses a diffusion-based LASS model in both separation quality and inference efficiency, highlighting its strong potential for audio source separation tasks. Code, pre-trained models and demos can be found at: https: audio-agi.github.io FlowSep_demo .
【9】 Flexible Control in Symbolic Music Generation via Musical Metadata
标题: 通过音乐元数据灵活控制象征性音乐生成
作者:Sangjun Han,Jiwon Ham,Chaeeun Lee,Heejin Kim,Soojong Do,Sihyuk Yi,Jun Seo,Seoyoon Kim,Yountae Jung,Woohyung Lim
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了象征性的音乐生成的示范,重点是提供短的音乐主题,作为叙事的中心主题。对于生成,我们采用了自回归模型,以音乐元数据作为输入,并生成4个酒吧的多轨音乐序列。在训练过程中,我们从音乐元数据中随机丢弃令牌,以确保灵活的控制。它为用户提供了选择输入类型的自由,同时保持生成性能,使音乐创作具有更大的灵活性。我们通过实验验证了该策略的有效性,在模型容量,音乐保真度,多样性和可控性。此外,我们放大模型,并通过主观测试与其他音乐生成模型进行比较。我们的研究结果表明,它的优势,在控制和音乐质量。我们提供一个URL链接https: www.youtube.com watch? v=-0drPrFJdMQ到我们的演示视频。摘要:In this work, we introduce the demonstration of symbolic music generation, focusing on providing short musical motifs that serve as the central theme of the narrative. For the generation, we adopt an autoregressive model which takes musical metadata as inputs and generates 4 bars of multitrack MIDI sequences. During training, we randomly drop tokens from the musical metadata to guarantee flexible control. It provides users with the freedom to select input types while maintaining generative performance, enabling greater flexibility in music composition. We validate the effectiveness of the strategy through experiments in terms of model capacity, musical fidelity, diversity, and controllability. Additionally, we scale up the model and compare it with other music generation model through a subjective test. Our results indicate its superiority in both control and music quality. We provide a URL link https: www.youtube.com watch?v=-0drPrFJdMQ to our demonstration video.
【10】 Efficient Sparse Coding with the Adaptive Locally Competitive Algorithm for Speech Classification
标题: 采用自适应局部竞争算法的高效稀疏编码语音分类
作者:Soufiyan Bahadi,Eric Plourde,Jean Rouat
链接:点击下载PDF文件
摘要:研究人员正在探索新的计算范式,如稀疏编码和神经形态计算,以弥合人脑和传统计算机在复杂任务中的效率差距。重点关注的一个关键领域是神经形态音频处理。虽然局部竞争算法已经成为稀疏编码的一个很有前途的解决方案,提供了潜在的实时和低功耗处理的神经形态硬件,其在神经形态语音分类的应用还没有得到彻底的研究。自适应局部竞争算法建立在局部竞争算法的基础上,通过动态调整滤波器组的调制参数来微调滤波器的灵敏度。这种适应性增强了侧抑制,提高了重建质量、稀疏性和收敛时间,这对于实时应用至关重要。本文证明了潜在的本地竞争算法及其自适应变种作为强大的特征提取神经形态语音分类。结果表明,局部竞争算法实现了更好的语音分类精度在较高的功耗为代价相比,用于基准测试的LAUSCHER耳蜗模型。另一方面,自适应局部竞争算法在不影响精度的情况下缓解了这种功耗问题。在神经形态硬件上,动态功耗降低到0.004到13毫瓦,比使用图形处理单元的设置低三个数量级。这些发现将自适应局部竞争算法定位为高效语音分类系统的引人注目的解决方案,有望在平衡语音分类准确性和功率效率方面取得实质性进展。摘要:Researchers are exploring novel computational paradigms such as sparse coding and neuromorphic computing to bridge the efficiency gap between the human brain and conventional computers in complex tasks. A key area of focus is neuromorphic audio processing. While the Locally Competitive Algorithm has emerged as a promising solution for sparse coding, offering potential for real-time and low-power processing on neuromorphic hardware, its applications in neuromorphic speech classification have not been thoroughly studied. The Adaptive Locally Competitive Algorithm builds upon the Locally Competitive Algorithm by dynamically adjusting the modulation parameters of the filter bank to fine-tune the filters' sensitivity. This adaptability enhances lateral inhibition, improving reconstruction quality, sparsity, and convergence time, which is crucial for real-time applications. This paper demonstrates the potential of the Locally Competitive Algorithm and its adaptive variant as robust feature extractors for neuromorphic speech classification. Results show that the Locally Competitive Algorithm achieves better speech classification accuracy at the expense of higher power consumption compared to the LAUSCHER cochlea model used for benchmarking. On the other hand, the Adaptive Locally Competitive Algorithm mitigates this power consumption issue without compromising the accuracy. The dynamic power consumption is reduced to a range of 0.004 to 13 milliwatts on neuromorphic hardware, three orders of magnitude less than setups using Graphics Processing Units. These findings position the Adaptive Locally Competitive Algorithm as a compelling solution for efficient speech classification systems, promising substantial advancements in balancing speech classification accuracy and power efficiency.
【11】 Faster Speech-LLaMA Inference with Multi-token Prediction
标题: 更快的语音-具有多令牌预测的LLaMA推理
作者:Desh Raj,Gil Keren,Junteng Jia,Jay Mahadeokar,Ozlem Kalinli
备注:Submitted to IEEE ICASSP 2025
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经能够熟练地解决各种各样的任务,包括涉及多模态输入的任务。特别是,用语音编码器实例化LLM(例如LLaMA)并在配对数据上训练它,将语音识别(ASR)能力赋予仅解码器模型,因此称为Speech-LLaMA。然而,由于自回归推理的顺序性质和相对较大的解码器,Speech-LLaMA模型需要相对较高的推理时间。在这项工作中,我们建议通过在同一解码步骤中预测多个令牌来加速Speech-LLaMA推断。我们探讨了几种模型架构,使这一点,并研究其性能使用基于阈值和基于验证的推理策略。我们还提出了一种基于前缀的波束搜索解码方法,该方法允许对此类模型进行有效的最小字错误率(MWER)训练。我们在各种公共基准测试中评估了我们的模型,在这些测试中,它们将解码器调用的数量减少了约3.2倍,同时保持或提高了WER性能。摘要:Large language models (LLMs) have become proficient at solving a wide variety of tasks, including those involving multi-modal inputs. In particular, instantiating an LLM (such as LLaMA) with a speech encoder and training it on paired data imparts speech recognition (ASR) abilities to the decoder-only model, hence called Speech-LLaMA. Nevertheless, due to the sequential nature of auto-regressive inference and the relatively large decoder, Speech-LLaMA models require relatively high inference time. In this work, we propose to speed up Speech-LLaMA inference by predicting multiple tokens in the same decoding step. We explore several model architectures that enable this, and investigate their performance using threshold-based and verification-based inference strategies. We also propose a prefix-based beam search decoding method that allows efficient minimum word error rate (MWER) training for such models. We evaluate our models on a variety of public benchmarks, where they reduce the number of decoder calls by ~3.2x while maintaining or improving WER performance.
【12】 Audio Decoding by Inverse Problem Solving
标题: 反问题求解的音频解码
作者:Pedro J. Villasana T.,Lars Villemoes,Janusz Klejsa,Per Hedelin
备注:5 pages, 4 figures, audio demo available at this https URL, pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:我们将音频解码视为一个逆问题,并通过扩散后验采样来解决它。显式调节函数被开发用于由变换域感知音频编解码器的示例提供的输入信号测量。通过评估一组比特率和任务不可知的先验模型的任意配对来证明可行性。例如,当语音模型被语音和钢琴训练的联合模型取代时,我们观察到钢琴的显着改善,同时保持语音性能。利用更通用的音乐模型,对于广泛的内容类型和比特率,获得了与传统方法相比改进的解码。的噪声平均值模型,基础上提出的推导条件,使扩散后采样的梯度评估显着减少,相比,基于Tweedie的平均值的方法。将Tweedie的平均值与我们的条件函数相结合,提高了客观性能。音频演示可在https: dpscodec-demo.github.io 上获得。摘要:We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio codec. Viability is demonstrated by evaluating arbitrary pairings of a set of bitrates and task-agnostic prior models. For instance, we observe significant improvements on piano while maintaining speech performance when a speech model is replaced by a joint model trained on both speech and piano. With a more general music model, improved decoding compared to legacy methods is obtained for a broad range of content types and bitrates. The noisy mean model, underlying the proposed derivation of conditioning, enables a significant reduction of gradient evaluations for diffusion posterior sampling, compared to methods based on Tweedie's mean. Combining Tweedie's mean with our conditioning functions improves the objective performance. An audio demo is available at https: dpscodec-demo.github.io .
【13】 Music auto-tagging in the long tail: A few-shot approach
标题: 长尾音乐自动标记:几次拍摄的方法
作者:T. Aleksandra Ma,Alexander Lerch
备注:Published in Audio Engineering Society NY Show 2024 as a Peer Reviewed (Category 1) paper
链接:点击下载PDF文件
摘要:在数字音乐领域,使用标签从大量数据库中有效地组织和检索音乐对音乐目录所有者至关重要。由专家进行的人工标记是劳动密集型的,但大多数是准确的,而通过监督学习进行的自动标记已经接近令人满意的准确性,但仅限于预定义的训练标签集。Few-Shot学习提供了一种可行的解决方案,通过使模型能够从仅几个人类提供的示例中学习来理解标签含义并随后自主应用这些标签,从而扩展到这一小组预定义标签之外。我们建议将Few-Shot学习方法集成到多标签音乐自动标记中,方法是使用来自预训练模型的特征作为轻量级线性分类器(也称为线性探针)的输入。我们研究了不同的流行的预训练特征,以及不同的Few-Shot参数化,其中每个类的类别和样本数量不同。我们的实验表明,具有预训练特征的简单模型可以实现接近最先进模型的性能,同时使用更少的训练数据,例如每个标签20个样本。此外,当在整个训练数据集上进行训练时,我们的线性探针与领先的模型具有竞争力。实验结果表明,这种基于迁移学习的Few-Shot方法可以有效解决在有限的标注数据下自动分配长尾标签的问题。摘要:In the realm of digital music, using tags to efficiently organize and retrieve music from extensive databases is crucial for music catalog owners. Human tagging by experts is labor-intensive but mostly accurate, whereas automatic tagging through supervised learning has approached satisfying accuracy but is restricted to a predefined set of training tags. Few-shot learning offers a viable solution to expand beyond this small set of predefined tags by enabling models to learn from only a few human-provided examples to understand tag meanings and subsequently apply these tags autonomously. We propose to integrate few-shot learning methodology into multi-label music auto-tagging by using features from pre-trained models as inputs to a lightweight linear classifier, also known as a linear probe. We investigate different popular pre-trained features, as well as different few-shot parametrizations with varying numbers of classes and samples per class. Our experiments demonstrate that a simple model with pre-trained features can achieve performance close to state-of-the-art models while using significantly less training data, such as 20 samples per tag. Additionally, our linear probe performs competitively with leading models when trained on the entire training dataset. The results show that this transfer learning-based few-shot approach could effectively address the issue of automatically assigning long-tail tags with only limited labeled data.
【14】 SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
标题: SSR-Speech:迈向稳定、安全和稳健的Zero-Shot基于文本的语音编辑和合成
作者:Helin Wang,Meng Yu,Jiarui Hai,Chen Chen,Yuchen Hu,Rilin Chen,Najim Dehak,Dong Yu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在本文中,我们介绍了SSR-Speech,一个神经编解码器自回归模型,设计用于稳定,安全和鲁棒的基于zero-shot文本的语音编辑和文本到语音合成。SSR-Speech是建立在一个Transformer解码器上的,并结合了无分类器的指导,以提高生成过程的稳定性。提出了一种水印编解码器,将帧级水印嵌入到语音的编辑区域中,以便检测哪些部分被编辑。此外,波形重建利用了原始的未编辑语音段,与Encodec模型相比,提供了更好的恢复。我们的方法在RealEdit语音编辑任务和LibriTTS文本到语音任务中实现了最先进的性能,超过了以前的方法。此外,SSR-Speech在多跨度语音编辑方面表现出色,并且对背景声音表现出显着的鲁棒性。源代码和演示已发布。摘要:In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves the state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. Source code and demos are released.
eess.AS音频处理
【1】 Efficient Sparse Coding with the Adaptive Locally Competitive Algorithm for Speech Classification标题: 采用自适应局部竞争算法的高效稀疏编码语音分类
作者:Soufiyan Bahadi,Eric Plourde,Jean Rouat
链接:点击下载PDF文件
摘要:研究人员正在探索新的计算范式,如稀疏编码和神经形态计算,以弥合人脑和传统计算机在复杂任务中的效率差距。重点关注的一个关键领域是神经形态音频处理。虽然局部竞争算法已经成为稀疏编码的一个很有前途的解决方案,提供了潜在的实时和低功耗处理的神经形态硬件,其在神经形态语音分类的应用还没有得到彻底的研究。自适应局部竞争算法建立在局部竞争算法的基础上,通过动态调整滤波器组的调制参数来微调滤波器的灵敏度。这种适应性增强了侧抑制,改善了重建质量,稀疏性和收敛时间,这对于实时应用至关重要。本文证明了潜在的本地竞争算法及其自适应变种作为强大的特征提取神经形态语音分类。结果表明,局部竞争算法实现了更好的语音分类精度在较高的功耗为代价相比,用于基准测试的LAUSCHER耳蜗模型。另一方面,自适应局部竞争算法在不影响精度的情况下缓解了这种功耗问题。在神经形态硬件上,动态功耗降低到0.004到13毫瓦,比使用图形处理单元的设置低三个数量级。这些发现将自适应局部竞争算法定位为高效语音分类系统的引人注目的解决方案,有望在平衡语音分类准确性和功率效率方面取得实质性进展。摘要:Researchers are exploring novel computational paradigms such as sparse coding and neuromorphic computing to bridge the efficiency gap between the human brain and conventional computers in complex tasks. A key area of focus is neuromorphic audio processing. While the Locally Competitive Algorithm has emerged as a promising solution for sparse coding, offering potential for real-time and low-power processing on neuromorphic hardware, its applications in neuromorphic speech classification have not been thoroughly studied. The Adaptive Locally Competitive Algorithm builds upon the Locally Competitive Algorithm by dynamically adjusting the modulation parameters of the filter bank to fine-tune the filters' sensitivity. This adaptability enhances lateral inhibition, improving reconstruction quality, sparsity, and convergence time, which is crucial for real-time applications. This paper demonstrates the potential of the Locally Competitive Algorithm and its adaptive variant as robust feature extractors for neuromorphic speech classification. Results show that the Locally Competitive Algorithm achieves better speech classification accuracy at the expense of higher power consumption compared to the LAUSCHER cochlea model used for benchmarking. On the other hand, the Adaptive Locally Competitive Algorithm mitigates this power consumption issue without compromising the accuracy. The dynamic power consumption is reduced to a range of 0.004 to 13 milliwatts on neuromorphic hardware, three orders of magnitude less than setups using Graphics Processing Units. These findings position the Adaptive Locally Competitive Algorithm as a compelling solution for efficient speech classification systems, promising substantial advancements in balancing speech classification accuracy and power efficiency.
【2】 Hierarchical Symbolic Pop Music Generation with Graph Neural Networks
标题: 利用图神经网络实现分层符号流行音乐生成
作者:Wen Qing Lim,Jinhua Liang,Huan Zhang
链接:点击下载PDF文件
摘要:音乐本质上是由复杂的结构组成的,将它们表示为图形有助于捕捉多层次的关系。虽然已经使用各种深度生成技术探索了音乐生成,但对与图形相关的音乐生成的研究很少。早期的基于图形的音乐生成仅用于生成旋律,而最近的作品生成复调音乐并不考虑长期结构。在本文中,我们探讨了一个多图的方法来表示的节奏模式和短语结构的中国流行音乐。因此,我们提出了一个两步的方法,旨在生成具有连贯节奏和长期结构的复调音乐。我们训练两个变分自动编码器网络-一个在一个数据集上生成4小节短语,另一个在歌曲结构标签上生成完整的歌曲结构。我们的工作表明,这些模型能够学习训练数据集中的大部分结构细微差别,包括和弦和音高频率分布以及短语属性。摘要:Music is inherently made up of complex structures, and representing them as graphs helps to capture multiple levels of relationships. While music generation has been explored using various deep generation techniques, research on graph-related music generation is sparse. Earlier graph-based music generation worked only on generating melodies, and recent works to generate polyphonic music do not account for longer-term structure. In this paper, we explore a multi-graph approach to represent both the rhythmic patterns and phrase structure of Chinese pop music. Consequently, we propose a two-step approach that aims to generate polyphonic music with coherent rhythm and long-term structure. We train two Variational Auto-Encoder networks - one on a MIDI dataset to generate 4-bar phrases, and another on song structure labels to generate full song structure. Our work shows that the models are able to learn most of the structural nuances in the training dataset, including chord and pitch frequency distributions, and phrase attributes.
【3】 Dark Experience for Incremental Keyword Spotting
标题: 增量关键词发现的黑暗体验
作者:Tianyi Peng,Yang Xiao
备注:submitted ICASSP 2025
链接:点击下载PDF文件
摘要:口语关键词识别(KWS)对于识别音频输入中的关键词至关重要,并广泛用于Apple Siri和Google Home等应用程序,特别是在边缘设备上。当前基于深度学习的KWS系统通常在有限的关键字集上进行训练,当遇到新领域时,可能会出现性能下降,这一挑战通常通过Few-Shot微调来解决。然而,这种适应经常导致灾难性的遗忘,其中模型对原始数据的性能恶化。已经提出了渐进式持续学习(CL)策略来克服这一点,但它们面临着诸如需要任务ID信息和增加存储等限制,使得它们对于轻量级设备不太实用。为了应对这些挑战,我们引入了关键词发现的黑暗经验(DE-KWS),这是一种新颖的CL方法,它利用暗知识在整个训练过程中提取过去的经验。DE-KWS结合了排练和蒸馏,使用存储在内存缓冲区中的地面真值标签和logits来维护跨任务的模型性能。对Google Speech Command数据集的评估表明,DE-KWS在平均准确度方面优于现有的CL基线,而不会增加模型大小,为资源受限的边缘设备提供了有效的解决方案。这些脚本可以在GitHub上找到,以供将来研究。摘要:Spoken keyword spotting (KWS) is crucial for identifying keywords within audio inputs and is widely used in applications like Apple Siri and Google Home, particularly on edge devices. Current deep learning-based KWS systems, which are typically trained on a limited set of keywords, can suffer from performance degradation when encountering new domains, a challenge often addressed through few-shot fine-tuning. However, this adaptation frequently leads to catastrophic forgetting, where the model's performance on original data deteriorates. Progressive continual learning (CL) strategies have been proposed to overcome this, but they face limitations such as the need for task-ID information and increased storage, making them less practical for lightweight devices. To address these challenges, we introduce Dark Experience for Keyword Spotting (DE-KWS), a novel CL approach that leverages dark knowledge to distill past experiences throughout the training process. DE-KWS combines rehearsal and distillation, using both ground truth labels and logits stored in a memory buffer to maintain model performance across tasks. Evaluations on the Google Speech Command dataset show that DE-KWS outperforms existing CL baselines in average accuracy without increasing model size, offering an effective solution for resource-constrained edge devices. The scripts are available on GitHub for the future research.
【4】 Faster Speech-LLaMA Inference with Multi-token Prediction
标题: 更快的语音-具有多令牌预测的LLaMA推理
作者:Desh Raj,Gil Keren,Junteng Jia,Jay Mahadeokar,Ozlem Kalinli
备注:Submitted to IEEE ICASSP 2025
链接:点击下载PDF文件
摘要:大型语言模型(LLM)已经能够熟练地解决各种各样的任务,包括涉及多模态输入的任务。特别是,用语音编码器实例化LLM(例如LLaMA)并在配对数据上训练它,将语音识别(ASR)能力赋予仅解码器模型,因此称为Speech-LLaMA。然而,由于自回归推理的顺序性质和相对较大的解码器,Speech-LLaMA模型需要相对较高的推理时间。在这项工作中,我们建议通过在同一解码步骤中预测多个令牌来加速Speech-LLaMA推断。我们探讨了几种模型架构,使这一点,并研究其性能使用基于阈值和基于验证的推理策略。我们还提出了一种基于前缀的波束搜索解码方法,该方法允许对此类模型进行有效的最小字错误率(MWER)训练。我们在各种公共基准测试中评估了我们的模型,在这些测试中,它们将解码器调用的数量减少了约3.2倍,同时保持或提高了WER性能。摘要:Large language models (LLMs) have become proficient at solving a wide variety of tasks, including those involving multi-modal inputs. In particular, instantiating an LLM (such as LLaMA) with a speech encoder and training it on paired data imparts speech recognition (ASR) abilities to the decoder-only model, hence called Speech-LLaMA. Nevertheless, due to the sequential nature of auto-regressive inference and the relatively large decoder, Speech-LLaMA models require relatively high inference time. In this work, we propose to speed up Speech-LLaMA inference by predicting multiple tokens in the same decoding step. We explore several model architectures that enable this, and investigate their performance using threshold-based and verification-based inference strategies. We also propose a prefix-based beam search decoding method that allows efficient minimum word error rate (MWER) training for such models. We evaluate our models on a variety of public benchmarks, where they reduce the number of decoder calls by ~3.2x while maintaining or improving WER performance.
【5】 Auto-Landmark: Acoustic Landmark Dataset and Open-Source Toolkit for Landmark Extraction
标题: 自动地标:用于地标提取的声学地标数据集和开源工具包
作者:Xiangyu Zhang,Daijiao Liu,Tianyi Xiao,Cihan Xiao,Tuende Szalay,Mostafa Shahin,Beena Ahmed,Julien Epps
链接:点击下载PDF文件
摘要:在语音信号中,声学地标识别语言动机的区别特征的声学表现最突出的时间。声学标志在语音识别、语音抑制检测、语音异常临床分析、语音障碍检测等领域有着广泛的应用。然而,目前还没有数据集可以为地标提供精确的定时信息,这对于涉及地标的下游应用来说是至关重要的。在本文中,我们选择了最有用的声学地标的基础上,以前的研究和注释TIMIT数据集,基于音素边界信息和人工检查相结合。此外,之前的地标提取工具不是开源的,也没有经过基准测试,因此为了解决这个问题,我们开发了一个基于Python的开源地标提取工具,并建立了一系列地标检测基线。这是第一个具有里程碑精确定时信息的数据集,里程碑提取工具和基线旨在支持各种未来研究。摘要:In the speech signal, acoustic landmarks identify times when the acoustic manifestations of the linguistically motivated distinctive features are most salient. Acoustic landmarks have been widely applied in various domains, including speech recognition, speech depression detection, clinical analysis of speech abnormalities, and the detection of disordered speech. However, there is currently no dataset available that provides precise timing information for landmarks, which has been proven to be crucial for downstream applications involving landmarks. In this paper, we selected the most useful acoustic landmarks based on previous research and annotated the TIMIT dataset with them, based on a combination of phoneme boundary information and manual inspection. Moreover, previous landmark extraction tools were not open source or benchmarked, so to address this, we developed an open source Python-based landmark extraction tool and established a series of landmark detection baselines. The first of their kinds, the dataset with landmark precise timing information, landmark extraction tool and baselines are designed to support a wide variety of future research.
【6】 Detecting and Defending Against Adversarial Attacks on Automatic Speech Recognition via Diffusion Models
标题: 利用扩散模型检测和防御自动语音识别中的对抗攻击
作者:Nikolai L. Kühne,Astrid H. F. Kitchen,Marie S. Jensen,Mikkel S. L. Brøndt,Martin Gonzalez,Christophe Biscio,Zheng-Hua Tan
备注:Under review at ICASSP 2025
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统容易受到对抗性攻击。本文讨论了检测和防御有针对性的白盒攻击语音信号的ASR系统。虽然现有的工作已经利用扩散模型(DM)来净化对抗性示例,在关键字识别任务中实现了最先进的结果,但它们对于更复杂的任务(如会话级ASR)的有效性仍未得到探索。此外,前向扩散步骤的数量对性能的影响还没有得到很好的理解。在本文中,我们系统地研究了使用DM防御对抗性攻击的句子,并检查不同的前向扩散步骤的效果。通过在Mozilla Common Voice数据集上的综合实验,我们证明了两个前向扩散步骤可以完全抵御对句子的对抗性攻击。此外,我们引入了一种新的,无需训练的方法,通过利用预先训练的DM来检测对抗性攻击。实验结果表明,该方法能够以较高的准确率检测出恶意攻击。摘要:Automatic speech recognition (ASR) systems are known to be vulnerable to adversarial attacks. This paper addresses detection and defence against targeted white-box attacks on speech signals for ASR systems. While existing work has utilised diffusion models (DMs) to purify adversarial examples, achieving state-of-the-art results in keyword spotting tasks, their effectiveness for more complex tasks such as sentence-level ASR remains unexplored. Additionally, the impact of the number of forward diffusion steps on performance is not well understood. In this paper, we systematically investigate the use of DMs for defending against adversarial attacks on sentences and examine the effect of varying forward diffusion steps. Through comprehensive experiments on the Mozilla Common Voice dataset, we demonstrate that two forward diffusion steps can completely defend against adversarial attacks on sentences. Moreover, we introduce a novel, training-free approach for detecting adversarial attacks by leveraging a pre-trained DM. Our experimental results show that this method can detect adversarial attacks with high accuracy.
【7】 Audio Decoding by Inverse Problem Solving
标题: 反问题求解的音频解码
作者:Pedro J. Villasana T.,Lars Villemoes,Janusz Klejsa,Per Hedelin
备注:5 pages, 4 figures, audio demo available at this https URL, pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:我们将音频解码视为一个逆问题,并通过扩散后验采样来解决它。显式调节函数被开发用于由变换域感知音频编解码器的示例提供的输入信号测量。通过评估一组比特率和任务不可知的先验模型的任意配对来证明可行性。例如,当语音模型被语音和钢琴训练的联合模型取代时,我们观察到钢琴的显着改善,同时保持语音性能。利用更通用的音乐模型,对于广泛的内容类型和比特率,获得了与传统方法相比改进的解码。的噪声平均值模型,基础上提出的推导条件,使扩散后采样的梯度评估显着减少,相比,基于Tweedie的平均值的方法。将Tweedie的平均值与我们的条件函数相结合,提高了客观性能。音频演示可在https: dpscodec-demo.github.io 上获得。摘要:We consider audio decoding as an inverse problem and solve it through diffusion posterior sampling. Explicit conditioning functions are developed for input signal measurements provided by an example of a transform domain perceptual audio codec. Viability is demonstrated by evaluating arbitrary pairings of a set of bitrates and task-agnostic prior models. For instance, we observe significant improvements on piano while maintaining speech performance when a speech model is replaced by a joint model trained on both speech and piano. With a more general music model, improved decoding compared to legacy methods is obtained for a broad range of content types and bitrates. The noisy mean model, underlying the proposed derivation of conditioning, enables a significant reduction of gradient evaluations for diffusion posterior sampling, compared to methods based on Tweedie's mean. Combining Tweedie's mean with our conditioning functions improves the objective performance. An audio demo is available at https: dpscodec-demo.github.io .
【8】 Universal Pooling Method of Multi-layer Features from Pretrained Models for Speaker Verification
标题: 用于说话人验证的预训练模型多层特征的通用池化方法
作者:Jin Sob Kim,Hyun Joon Park,Wooseok Shin,Sung Won Han
备注:Preprint
链接:点击下载PDF文件
摘要:自动说话人确认(ASV)研究的最新进展是通过利用大规模预训练网络实现的。在这项研究中,我们分析了实现这种范式的途径,并强调了层间信息加工的意义。因此,我们提出了一种新的方法来利用ASV预训练模型的多层性质,该方法包括一个层 帧级网络和每个层和帧轴的两个池化架构步骤。具体来说,我们让卷积架构直接处理一堆层输出。然后,我们提出了一种基于信道注意力的方法来衡量层的重要性,并用最具代表性的值挤压层级别。最后,在帧级表示上的注意统计产生单个矢量说话人嵌入。比较实验设计使用通用的数据环境和不同的预训练模型,以验证所提出的方法。实验结果证明了该方法在利用预训练架构时使用多层输出的稳定性。然后,我们验证了所提出的ASV后端结构,其中涉及逐层操作的优越性,在性能改善以及成本效益方面相比,传统的方法。消融研究显示了所提出的层间处理如何帮助最大限度地利用预训练模型的优势。摘要:Recent advancements in automatic speaker verification (ASV) studies have been achieved by leveraging large-scale pretrained networks. In this study, we analyze the approaches toward such a paradigm and underline the significance of interlayer information processing as a result. Accordingly, we present a novel approach for exploiting the multilayered nature of pretrained models for ASV, which comprises a layer frame-level network and two steps of pooling architectures for each layer and frame axis. Specifically, we let convolutional architecture directly processes a stack of layer outputs.Then, we present a channel attention-based scheme of gauging layer significance and squeeze the layer level with the most representative value. Finally, attentive statistics over frame-level representations yield a single vector speaker embedding. Comparative experiments are designed using versatile data environments and diverse pretraining models to validate the proposed approach. The experimental results demonstrate the stability of the approach using multi-layer outputs in leveraging pretrained architectures. Then, we verify the superiority of the proposed ASV backend structure, which involves layer-wise operations, in terms of performance improvement along with cost efficiency compared to the conventional method. The ablation study shows how the proposed interlayer processing aids in maximizing the advantage of utilizing pretrained models.
【9】 Music auto-tagging in the long tail: A few-shot approach
标题: 长尾音乐自动标记:几次拍摄的方法
作者:T. Aleksandra Ma,Alexander Lerch
备注:Published in Audio Engineering Society NY Show 2024 as a Peer Reviewed (Category 1) paper
链接:点击下载PDF文件
摘要:在数字音乐领域,使用标签从大量数据库中有效地组织和检索音乐对音乐目录所有者至关重要。由专家进行的人工标记是劳动密集型的,但大多数是准确的,而通过监督学习进行的自动标记已经接近令人满意的准确性,但仅限于预定义的训练标签集。Few-Shot学习提供了一种可行的解决方案,通过使模型能够从仅几个人类提供的示例中学习来理解标签含义并随后自主应用这些标签,从而扩展到这一小组预定义标签之外。我们建议将Few-Shot学习方法集成到多标签音乐自动标记中,方法是使用来自预训练模型的特征作为轻量级线性分类器(也称为线性探针)的输入。我们研究了不同的流行的预训练特征,以及不同的Few-Shot参数化,其中每个类的类别和样本数量不同。我们的实验表明,具有预训练特征的简单模型可以实现接近最先进模型的性能,同时使用更少的训练数据,例如每个标签20个样本。此外,当在整个训练数据集上进行训练时,我们的线性探针与领先的模型具有竞争力。实验结果表明,这种基于迁移学习的Few-Shot方法可以有效解决在有限的标注数据下自动分配长尾标签的问题。摘要:In the realm of digital music, using tags to efficiently organize and retrieve music from extensive databases is crucial for music catalog owners. Human tagging by experts is labor-intensive but mostly accurate, whereas automatic tagging through supervised learning has approached satisfying accuracy but is restricted to a predefined set of training tags. Few-shot learning offers a viable solution to expand beyond this small set of predefined tags by enabling models to learn from only a few human-provided examples to understand tag meanings and subsequently apply these tags autonomously. We propose to integrate few-shot learning methodology into multi-label music auto-tagging by using features from pre-trained models as inputs to a lightweight linear classifier, also known as a linear probe. We investigate different popular pre-trained features, as well as different few-shot parametrizations with varying numbers of classes and samples per class. Our experiments demonstrate that a simple model with pre-trained features can achieve performance close to state-of-the-art models while using significantly less training data, such as 20 samples per tag. Additionally, our linear probe performs competitively with leading models when trained on the entire training dataset. The results show that this transfer learning-based few-shot approach could effectively address the issue of automatically assigning long-tail tags with only limited labeled data.
【10】 Super Monotonic Alignment Search
标题: 超级单调对齐搜索
作者:Junhyeok Lee,Hyeongju Kim
备注:Technical Report
链接:点击下载PDF文件
摘要:Glow-TTS提出的单调对齐搜索(MAS)是TTS中最流行的一种算法,用于估计文本和语音之间的未知对齐。由于该算法需要通过缓存所有路径来用动态规划搜索最可能的对齐,因此该算法的时间复杂度为O(T times S)$。Glow-TTS的作者在CPU上运行该算法,虽然他们提到它很难并行化,但我们发现MAS可以在文本长度维度上并行化,并且CPU执行消耗大量的时间用于设备间复制。因此,我们实现了Triton内核和PyTorch JIT脚本来加速GPU上的MAS,而无需设备间复制。因此,Super-MAS Triton内核在极端长度情况下的速度快了72倍。代码可以在 url{https: github.com supertone-inc super-monotonic-align}上找到。摘要:Monotonic alignment search (MAS), introduced by Glow-TTS, is one of the most popular algorithm in TTS to estimate unknown alignments between text and speech. Since this algorithm needs to search for the most probable alignment with dynamic programming by caching all paths, the time complexity of the algorithm is $O(T times S)$. The authors of Glow-TTS run this algorithm on CPU, and while they mentioned it is difficult to parallelize, we found that MAS can be parallelized in text-length dimension and CPU execution consumes an inordinate amount of time for inter-device copy. Therefore, we implemented a Triton kernel and PyTorch JIT script to accelerate MAS on GPU without inter-device copy. As a result, Super-MAS Triton kernel is up to 72 times faster in the extreme-length case. The code is available at url{https: github.com supertone-inc super-monotonic-align}.
【11】 SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
标题: SSR-Speech:迈向稳定、安全和稳健的Zero-Shot基于文本的语音编辑和合成
作者:Helin Wang,Meng Yu,Jiarui Hai,Chen Chen,Yuchen Hu,Rilin Chen,Najim Dehak,Dong Yu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在本文中,我们介绍了SSR-Speech,一个神经编解码器自回归模型,设计用于稳定,安全和鲁棒的基于zero-shot文本的语音编辑和文本到语音合成。SSR-Speech是建立在一个Transformer解码器上的,并结合了无分类器的指导,以提高生成过程的稳定性。提出了一种水印编解码器,将帧级水印嵌入到语音的编辑区域中,以便检测哪些部分被编辑。此外,波形重建利用了原始的未编辑语音段,与Encodec模型相比,提供了更好的恢复。我们的方法在RealEdit语音编辑任务和LibriTTS文本到语音任务中实现了最先进的性能,超过了以前的方法。此外,SSR-Speech在多跨度语音编辑方面表现出色,并且对背景声音表现出显着的鲁棒性。源代码和演示已发布。摘要:In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves the state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. Source code and demos are released.
【12】 AudioBERT: Audio Knowledge Augmented Language Model
标题: AudioBERT:音频知识增强语言模型
作者:Hyunjong Ok,Suho Yoo,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
摘要:最近的研究发现,在纯文本数据集上预训练的语言模型通常缺乏基本的视觉知识, textit{e.g.}日常物品的颜色。受此观察的启发,我们问是否存在类似的缺点在 textit{听觉}知识。为了回答这个问题,我们构建了一个名为AuditoryBench的新数据集,它包括两个用于评估听觉知识的新任务。基于我们使用基准的分析,我们发现语言模型也严重缺乏听觉知识。为了解决这个问题,我们提出了AudioBERT,一种新的方法来增加听觉知识的BERT通过检索为基础的方法。首先,我们检测听觉知识跨度提示查询我们的检索模型有效。然后,我们将音频知识注入到BERT中,并在需要音频知识时打开低秩自适应以进行有效的自适应。我们的实验表明,AudioBERT是非常有效的,在AuditoryBench上实现了卓越的性能。数据集和代码可以在 bulurl{https: github.com HJ-Ok AudioBERT}上找到。摘要:Recent studies have identified that language models, pretrained on text-only datasets, often lack elementary visual knowledge, textit{e.g.,} colors of everyday objects. Motivated by this observation, we ask whether a similar shortcoming exists in terms of the textit{auditory} knowledge. To answer this question, we construct a new dataset called AuditoryBench, which consists of two novel tasks for evaluating auditory knowledge. Based on our analysis using the benchmark, we find that language models also suffer from a severe lack of auditory knowledge. To address this limitation, we propose AudioBERT, a novel method to augment the auditory knowledge of BERT through a retrieval-based approach. First, we detect auditory knowledge spans in prompts to query our retrieval model efficiently. Then, we inject audio knowledge into BERT and switch on low-rank adaptation for effective adaptation when audio knowledge is required. Our experiments demonstrate that AudioBERT is quite effective, achieving superior performance on the AuditoryBench. The dataset and code are available at bulurl{https: github.com HJ-Ok AudioBERT}.
【13】 The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
标题: Faetar基准:资源严重不足的语言中的语音识别
作者:Michael Ong,Sean Robertson,Leo Peckham,Alba Jorquera Jimenez de Aberasturi,Paula Arkhangorodsky,Robin Huo,Aman Sakhardande,Mark Hallap,Naomi Nagy,Ewan Dunbar
链接:点击下载PDF文件
摘要:我们介绍Faetar自动语音识别基准,一个基准语料库,旨在推动目前的低资源语音识别方法的限制。Faetar是一种主要在意大利使用的法语变体,没有标准的正字法,除了基准中包含的内容之外,几乎没有现有的文本或语音资源,并且与其他形式的法语变体有很大的不同。语料库来自现场录音,其中大部分是嘈杂的,其中只有5个小时具有匹配的音标,并且强制对准具有可变的质量。语料库包含额外的20小时的未标记的语音。我们报告了最先进的多语言语音基础模型的基线结果,最佳电话错误率为30.4%,使用管道继续使用未标记集对基础模型进行预训练。摘要:We introduce the Faetar Automatic Speech Recognition Benchmark, a benchmark corpus designed to push the limits of current approaches to low-resource speech recognition. Faetar, a Franco-Proven c{c}al variety spoken primarily in Italy, has no standard orthography, has virtually no existing textual or speech resources other than what is included in the benchmark, and is quite different from other forms of Franco-Proven c{c}al. The corpus comes from field recordings, most of which are noisy, for which only 5 hrs have matching transcriptions, and for which forced alignment is of variable quality. The corpus contains an additional 20 hrs of unlabelled speech. We report baseline results from state-of-the-art multilingual speech foundation models with a best phone error rate of 30.4%, using a pipeline that continues pre-training on the foundation model using the unlabelled set.
【14】 Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
标题: Zero-Shot歌唱语音转换:基于集群的音素表示构建
作者:Wangjin Zhou,Fengrun Zhang,Yiming Liu,Wenhao Guan,Yi Zhao,He Qu
链接:点击下载PDF文件
摘要:本研究提出了一种创新的Zero-Shot任意到任意歌唱语音转换(SVC)方法,利用一种新的基于聚类的音素表示有效地分离内容,音色和演唱风格。这种方法能够实现精确的语音特性操纵。我们发现,每个艺术家的录音较少的数据集更容易受到音色泄漏的影响。对超过10,000小时的演唱和用户反馈进行的广泛测试表明,我们的模型显着提高了音质和音色准确性,符合我们的目标,并推进了语音转换技术。此外,本研究提出了zero-shot SVC,并为未来的离散语音表示工作奠定了基础,强调了押韵的保留。摘要:This study presents an innovative Zero-Shot any-to-any Singing Voice Conversion (SVC) method, leveraging a novel clustering-based phoneme representation to effectively separate content, timbre, and singing style. This approach enables precise voice characteristic manipulation. We discovered that datasets with fewer recordings per artist are more susceptible to timbre leakage. Extensive testing on over 10,000 hours of singing and user feedback revealed our model significantly improves sound quality and timbre accuracy, aligning with our objectives and advancing voice conversion technology. Furthermore, this research advances zero-shot SVC and sets the stage for future work on discrete speech representation, emphasizing the preservation of rhyme.
【15】 Tidal MerzA: Combining affective modelling and autonomous code generation through Reinforcement Learning
标题: Tidal MerzA:通过强化学习结合情感建模和自主代码生成
作者:Elizabeth Wilson,György Fazekas,Geraint Wiggins
链接:点击下载PDF文件
摘要:本文介绍了Tidal-MerzA,这是一种新的系统,旨在在现场编码的背景下,人类和机器代理之间的协作表演,特别关注音乐模式的生成。Tidal-MerzA融合了两个基本模型:ALCAA(情感实时编码自治代理)和Tidal Fuzz,一个计算框架。通过将情感建模与计算生成相结合,该系统利用强化学习技术在TidalCycles框架内动态调整音乐作曲参数,确保模式的情感质量和语法正确性。Tidal-MerzA的开发引入了两个不同的代理:一个专注于生成用于音乐表达的迷你符号串,另一个专注于通过强化学习将音乐与目标情感状态对齐。这种方法增强了实时编码实践的适应性和创造性潜力,并允许探索人机创造性交互。Tidal-MerzA推进了计算音乐生成领域,提出了一种将人工智能融入艺术实践的新方法。摘要:This paper presents Tidal-MerzA, a novel system designed for collaborative performances between humans and a machine agent in the context of live coding, specifically focusing on the generation of musical patterns. Tidal-MerzA fuses two foundational models: ALCAA (Affective Live Coding Autonomous Agent) and Tidal Fuzz, a computational framework. By integrating affective modelling with computational generation, this system leverages reinforcement learning techniques to dynamically adapt music composition parameters within the TidalCycles framework, ensuring both affective qualities to the patterns and syntactical correctness. The development of Tidal-MerzA introduces two distinct agents: one focusing on the generation of mini-notation strings for musical expression, and another on the alignment of music with targeted affective states through reinforcement learning. This approach enhances the adaptability and creative potential of live coding practices and allows exploration of human-machine creative interactions. Tidal-MerzA advances the field of computational music generation, presenting a novel methodology for incorporating artificial intelligence into artistic practices.
【16】 A corpus-based investigation of pitch contours of monosyllabic words in conversational Taiwan Mandarin
标题: 台湾普通话对话中单音节词音调轮廓的研究
作者:Xiaoyun Jin,Mirjam Ernestus,R. Harald Baayen
链接:点击下载PDF文件
摘要:汉语单音节词的声调轮廓由四个声调组成:上声(T1)、升调(T2)、降调(T3)和降调(T4)。然而,在自发语音中,单音节词的实际音调实现可能会由于与相邻音调的音节内协同发音和音节间协同发音而显著偏离这些规范音调。此外,Chuang et al.(2024)最近报道,具有T2-T4声调模式的汉语双音节词的声调轮廓由它们的意义共同决定。在此基础上,本文利用语料库研究了汉语自然会话中单音节词的音高轮廓是如何实现的,一方面着重于语境预测因素的影响,另一方面着重于词义共同决定音高轮廓的方式。我们分析了自发的台湾普通话语料库中的3824个标记的63个不同的词类型的F0轮廓,使用广义加法(混合)模型分解成一组组件的音高轮廓观察到的音高轮廓。我们表明,音调的背景下,大大修改一个词的规范语气。一旦控制了音调环境的影响,T2和T3就以低平声出现,与T1作为高音形成对比,而T4作为高中降调。在标准描述中,基于前一个音调实现的轻声(T0)本身作为低音出现,以与标准音调T1、T2、T3和T4相同的方式被其他预测因子修改。我们还表明,词,甚至更重要的是,词的意义,共同决定词的F0轮廓。使用随机森林变量的重要性分析进一步支持音调上下文和词义的效果的实质性影响。摘要:In Mandarin, the tonal contours of monosyllabic words produced in isolation or in careful speech are characterized by four lexical tones: a high-level tone (T1), a rising tone (T2), a dipping tone (T3) and a falling tone (T4). However, in spontaneous speech, the actual tonal realization of monosyllabic words can deviate significantly from these canonical tones due to intra-syllabic co-articulation and inter-syllabic co-articulation with adjacent tones. In addition, Chuang et al. (2024) recently reported that the tonal contours of disyllabic Mandarin words with T2-T4 tone pattern are co-determined by their meanings. Following up on their research, we present a corpus-based investigation of how the pitch contours of monosyllabic words are realized in spontaneous conversational Mandarin, focusing on the effects of contextual predictors on the one hand, and the way in words' meanings co-determine pitch contours on the other hand. We analyze the F0 contours of 3824 tokens of 63 different word types in a spontaneous Taiwan Mandarin corpus, using the generalized additive (mixed) model to decompose a given observed pitch contour into a set of component pitch contours. We show that the tonal context substantially modify a word's canonical tone. Once the effect of tonal context is controlled for, T2 and T3 emerge as low flat tones, contrasting with T1 as a high tone, and with T4 as a high-to-mid falling tone. The neutral tone (T0), which in standard descriptions, is realized based on the preceding tone, emerges as a low tone in its own right, modified by the other predictors in the same way as the standard tones T1, T2, T3, and T4. We also show that word, and even more so, word sense, co-determine words' F0 contours. Analyses of variable importance using random forests further supported the substantial effect of tonal context and an effect of word sense.
【17】 Graph Neural Networks for Parkinsons Disease Detection
标题: 用于帕金森病检测的图神经网络
作者:Shakeel A. Sheikh,Yacouba Kaloga,Ina Kodrasi
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:尽管用于帕金森病(PD)检测的现有技术方法的性能有希望,但是这些方法通常孤立地分析单个语音段,这可能导致次优结果。表征来自PD患者的言语障碍的构音障碍线索预计与来自不同说话者的跨段相关。孤立的细分市场分析无法利用这些细分市场之间的关系。此外,并非PD患者的所有语音片段都表现出明显的构音障碍症状,从而引入标签噪音,这可能会对当前方法的性能和可推广性产生负面影响。为了解决这些挑战,我们提出了一种新的PD检测框架,利用图卷积网络(GCN)。通过将语音片段表示为节点并通过边缘捕获片段之间的相似性,我们的GCN模型有助于在整个图中聚合构音障碍线索,有效地利用片段关系并减轻标签噪声的影响。实验结果证明了所提出的GCN模型在局部放电检测中的优势,并提供了其潜在机制的见解摘要:Despite the promising performance of state of the art approaches for Parkinsons Disease (PD) detection, these approaches often analyze individual speech segments in isolation, which can lead to suboptimal results. Dysarthric cues that characterize speech impairments from PD patients are expected to be related across segments from different speakers. Isolated segment analysis fails to exploit these inter segment relationships. Additionally, not all speech segments from PD patients exhibit clear dysarthric symptoms, introducing label noise that can negatively affect the performance and generalizability of current approaches. To address these challenges, we propose a novel PD detection framework utilizing Graph Convolutional Networks (GCNs). By representing speech segments as nodes and capturing the similarity between segments through edges, our GCN model facilitates the aggregation of dysarthric cues across the graph, effectively exploiting segment relationships and mitigating the impact of label noise. Experimental results demonstrate theadvantages of the proposed GCN model for PD detection and provide insights into its underlying mechanisms
【18】 TSELM: Target Speaker Extraction using Discrete Tokens and Language Models
标题: TSELM:使用离散标记和语言模型的目标说话人提取
作者:Beilong Tang,Bang Zeng,Ming Li
链接:点击下载PDF文件
摘要:我们提出了TSELM,一种新的目标说话人提取网络,利用离散的令牌和语言模型。TSELM利用WavLM的多个离散层作为输入标记,并结合交叉注意机制来整合目标说话人信息。语言模型用于捕获序列依赖关系,而可扩展的HiFi-GAN用于从令牌重建音频。通过应用交叉熵损失,TSELM对输出令牌的概率分布进行建模,从而将音频生成的复杂回归问题转换为分类任务。实验结果表明,TSELM在语音质量方面取得了优异的效果,在语音清晰度方面也取得了相当的效果。摘要:We propose TSELM, a novel target speaker extraction network that leverages discrete tokens and language models. TSELM utilizes multiple discretized layers from WavLM as input tokens and incorporates cross-attention mechanisms to integrate target speaker information. Language models are employed to capture the sequence dependencies, while a scalable HiFi-GAN is used to reconstruct the audio from the tokens. By applying a cross-entropy loss, TSELM models the probability distribution of output tokens, thus converting the complex regression problem of audio generation into a classification task. Experimental results show that TSELM achieves excellent results in speech quality and comparable results in speech intelligibility.
【19】 Bridging Paintings and Music -- Exploring Emotion based Music Generation through Paintings
标题: 绘画与音乐的桥梁--通过绘画探索基于情感的音乐生成
作者:Tanisha Hisariya,Huan Zhang,Jinhua Liang
链接:点击下载PDF文件
摘要:人工智能的快速发展大大增强了涉及音乐和图像的生成任务,采用了单峰和多峰方法。本研究开发了一种能够生成与视觉艺术中所描绘的情感产生共鸣的音乐的模型,将情感标签,图像字幕和语言模型集成在一起,将视觉输入转换为音乐作品。为了解决艺术和音乐数据的稀缺性,我们策划了情感绘画音乐数据集,将绘画与相应的音乐配对,以进行有效的训练和评估。我们的双阶段框架将图像转换为情感内容的文本描述,然后将这些描述转换为音乐,以最少的数据促进高效学习。性能使用诸如Fr 'echet音频距离(FAD)、总谐波失真(THD)、初始分数(IS)和KL发散度等指标进行评估,其中音频-情感文本相似性由预训练的CLAP模型确认,以证明生成的音乐和文本之间的高度对齐。这一综合工具将视觉艺术和音乐联系起来,通过提供丰富的多感官体验,增强了视障者的无障碍环境,并开辟了教育和治疗应用的途径。摘要:Rapid advancements in artificial intelligence have significantly enhanced generative tasks involving music and images, employing both unimodal and multimodal approaches. This research develops a model capable of generating music that resonates with the emotions depicted in visual arts, integrating emotion labeling, image captioning, and language models to transform visual inputs into musical compositions. Addressing the scarcity of aligned art and music data, we curated the Emotion Painting Music Dataset, pairing paintings with corresponding music for effective training and evaluation. Our dual-stage framework converts images to text descriptions of emotional content and then transforms these descriptions into music, facilitating efficient learning with minimal data. Performance is evaluated using metrics such as Fr 'echet Audio Distance (FAD), Total Harmonic Distortion (THD), Inception Score (IS), and KL divergence, with audio-emotion text similarity confirmed by the pre-trained CLAP model to demonstrate high alignment between generated music and text. This synthesis tool bridges visual art and music, enhancing accessibility for the visually impaired and opening avenues in educational and therapeutic applications by providing enriched multi-sensory experiences.
【20】 Full-text Error Correction for Chinese Speech Recognition with Large Language Model
标题: 大语言模型中文语音识别的全文错误纠正
作者:Zhiyuan Tang,Dong Wang,Shen Huang,Shidong Shang
链接:点击下载PDF文件
摘要:大型语言模型(LLM)在自动语音识别(ASR)中具有很大的纠错潜力。然而,大多数研究都集中在短时间语音记录中的话语,这是监督ASR训练的主要语音数据形式。本文研究了LLM在ASR系统从较长的语音记录(如播客,新闻广播和会议的成绩单)生成的全文中纠错的有效性。首先,我们开发了一个用于全文纠错的中文数据集,名为ChFT,利用一个包含文本到语音合成,ASR和纠错对提取器的管道。该数据集使我们能够跨上下文纠正错误,包括全文和片段,并解决更广泛的错误类型,如标点符号恢复和反向文本规范化,从而使纠正过程全面。其次,我们使用一组不同的提示和目标格式在构建的数据集上微调预训练的LLM,并评估其在全文纠错方面的性能。具体来说,我们设计了基于全文和片段的提示,考虑了各种输出格式,例如直接校正的文本和基于JSON的错误校正对。通过各种测试设置,包括同质的,最新的,和硬测试集,我们发现,微调LLM表现良好,在全文设置不同的提示,每个呈现自己的优势和劣势。这为进一步的研究建立了一个有希望的基线。该数据集可在网站上查阅。摘要:Large Language Models (LLMs) have demonstrated substantial potential for error correction in Automatic Speech Recognition (ASR). However, most research focuses on utterances from short-duration speech recordings, which are the predominant form of speech data for supervised ASR training. This paper investigates the effectiveness of LLMs for error correction in full-text generated by ASR systems from longer speech recordings, such as transcripts from podcasts, news broadcasts, and meetings. First, we develop a Chinese dataset for full-text error correction, named ChFT, utilizing a pipeline that involves text-to-speech synthesis, ASR, and error-correction pair extractor. This dataset enables us to correct errors across contexts, including both full-text and segment, and to address a broader range of error types, such as punctuation restoration and inverse text normalization, thus making the correction process comprehensive. Second, we fine-tune a pre-trained LLM on the constructed dataset using a diverse set of prompts and target formats, and evaluate its performance on full-text error correction. Specifically, we design prompts based on full-text and segment, considering various output formats, such as directly corrected text and JSON-based error-correction pairs. Through various test settings, including homogeneous, up-to-date, and hard test sets, we find that the fine-tuned LLMs perform well in the full-text setting with different prompts, each presenting its own strengths and weaknesses. This establishes a promising baseline for further research. The dataset is available on the website.
【21】 FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
标题: FlowSep:采用整流流匹配的阈值查询声音分离
作者:Yi Yuan,Xubo Liu,Haohe Liu,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
摘要:语音查询音频源分离(LASS)的重点是分离声音使用所需的源的文本描述。目前的方法主要使用判别方法,如时频掩蔽,以分离目标声音,并最大限度地减少来自其他来源的干扰。然而,这些模型在分离重叠的音轨时面临挑战,这可能导致诸如频谱孔或不完全分离的伪像。整流匹配(RFM),一种生成模型,建立数据和噪声的分布之间的线性关系,提供了优越的理论性能和简单性,但尚未在声音分离探索。在这项工作中,我们介绍FlowSep,一个新的生成模型RFM的基础上LASS任务。FlowSep在变分自动编码器(VAE)潜在空间内从噪声到目标源特征学习线性流轨迹。在推理过程中,RFM生成的潜在特征通过预先训练的VAE解码器重建成梅尔频谱图,然后通过预先训练的声码器合成波形。经过1,680小时的音频数据训练,FlowSep在多个基准测试中的表现优于最先进的模型,并通过主观和客观指标进行评估。此外,我们的研究结果表明,FlowSep在分离质量和推理效率方面都优于基于扩散的LASS模型,突出了其在音频源分离任务中的强大潜力。代码、预训练模型和演示可以在https: audio-agi.github.io FlowSep_demo 上找到。摘要:Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these models face challenges when separating overlapping soundtracks, which may lead to artifacts such as spectral holes or incomplete separation. Rectified flow matching (RFM), a generative model that establishes linear relations between the distribution of data and noise, offers superior theoretical properties and simplicity, but has not yet been explored in sound separation. In this work, we introduce FlowSep, a new generative model based on RFM for LASS tasks. FlowSep learns linear flow trajectories from noise to target source features within the variational autoencoder (VAE) latent space. During inference, the RFM-generated latent features are reconstructed into a mel-spectrogram via the pre-trained VAE decoder, followed by a pre-trained vocoder to synthesize the waveform. Trained on 1,680 hours of audio data, FlowSep outperforms the state-of-the-art models across multiple benchmarks, as evaluated with subjective and objective metrics. Additionally, our results show that FlowSep surpasses a diffusion-based LASS model in both separation quality and inference efficiency, highlighting its strong potential for audio source separation tasks. Code, pre-trained models and demos can be found at: https: audio-agi.github.io FlowSep_demo .
【22】 Flexible Control in Symbolic Music Generation via Musical Metadata
标题: 通过音乐元数据灵活控制象征性音乐生成
作者:Sangjun Han,Jiwon Ham,Chaeeun Lee,Heejin Kim,Soojong Do,Sihyuk Yi,Jun Seo,Seoyoon Kim,Yountae Jung,Woohyung Lim
链接:点击下载PDF文件
摘要:在这项工作中,我们介绍了象征性的音乐生成的示范,重点是提供短的音乐主题,作为叙事的中心主题。对于生成,我们采用了自回归模型,以音乐元数据作为输入,并生成4个酒吧的多轨音乐序列。在训练过程中,我们从音乐元数据中随机丢弃令牌,以确保灵活的控制。它为用户提供了选择输入类型的自由,同时保持生成性能,使音乐创作具有更大的灵活性。我们通过实验验证了该策略的有效性,在模型容量,音乐保真度,多样性和可控性。此外,我们放大模型,并通过主观测试与其他音乐生成模型进行比较。我们的研究结果表明,它的优势,在控制和音乐质量。我们提供一个URL链接https: www.youtube.com watch? v=-0drPrFJdMQ到我们的演示视频。摘要:In this work, we introduce the demonstration of symbolic music generation, focusing on providing short musical motifs that serve as the central theme of the narrative. For the generation, we adopt an autoregressive model which takes musical metadata as inputs and generates 4 bars of multitrack MIDI sequences. During training, we randomly drop tokens from the musical metadata to guarantee flexible control. It provides users with the freedom to select input types while maintaining generative performance, enabling greater flexibility in music composition. We validate the effectiveness of the strategy through experiments in terms of model capacity, musical fidelity, diversity, and controllability. Additionally, we scale up the model and compare it with other music generation model through a subjective test. Our results indicate its superiority in both control and music quality. We provide a URL link https: www.youtube.com watch?v=-0drPrFJdMQ to our demonstration video.
【23】 Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
标题: 在真正的低资源环境中改进视觉化关键字本地化
作者:Leanne Nortje,Dan Oneata,Herman Kamper
链接:点击下载PDF文件
摘要:给定图像查询,视觉提示关键字本地化(VPKL)的目的是在语音集合中找到所描述的单词的出现。当transmittance不适用于低资源语言时(例如,如果它是非书面的),这可能很有用。以前的工作表明,VPKL可以使用在成对图像和未标记语音上训练的视觉接地语音模型来执行。但所有的实验都是在英语上进行的。此外,使用transmittance来获得对比损失的正对和负对。本文介绍了一种无需传输的Few-Shot学习算法。在英语方面,这只会导致成绩小幅下降。我们还首次考虑在真正的低资源语言约鲁巴语上使用VPKL。虽然分数是合理的,但与使用地面真值对相比,我们在这里看到了更大的性能下降,因为挖掘在约鲁巴语中不太准确。摘要:Given an image query, visually prompted keyword localisation (VPKL) aims to find occurrences of the depicted word in a speech collection. This can be useful when transcriptions are not available for a low-resource language (e.g. if it is unwritten). Previous work showed that VPKL can be performed with a visually grounded speech model trained on paired images and unlabelled speech. But all experiments were done on English. Moreover, transcriptions were used to get positive and negative pairs for the contrastive loss. This paper introduces a few-shot learning scheme to mine pairs automatically without transcriptions. On English, this results in only a small drop in performance. We also - for the first time - consider VPKL on a real low-resource language, Yoruba. While scores are reasonable, here we see a bigger drop in performance compared to using ground truth pairs because the mining is less accurate in Yoruba.
机器翻译,仅供参考
