本文经arXiv每日学术速递授权转载
【1】 Beat this! Accurate beat tracking without DBN postprocessing
标题: 打败这个!无需DBN后处理即可准确跟踪节拍
作者:Francesco Foscarin,Jan Schlüter,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
【2】 Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
标题: 通过LLM Agent实现端到端同步语音翻译的人类对等
作者:Shanbo Cheng,Zhichao Huang,Tom Ko,Hang Li,Ningxin Peng,Lu Xu,Qini Zhang
备注:Authors are listed in alphabetical order by last name. Demonstrations and human-annotated test sets are available at this https URL
链接:点击下载PDF文件
【3】 Between the AI and Me: Analysing Listeners' Perspectives on AI- and Human-Composed Progressive Metal Music
标题: 在人工智能和我之间:分析听众对人工智能和人类创作的进步金属音乐的看法
作者:Pedro Sarmento,Jackson Loth,Mathieu Barthet
备注:Reviewed pre-print accepted for publication at ISMIR 2024
链接:点击下载PDF文件
【4】 Enhancing Partially Spoofed Audio Localization with Boundary-aware Attention Mechanism
标题: 利用边界感知注意机制增强部分欺骗音频定位
作者:Jiafeng Zhong,Bin Li,Jiangyan Yi
链接:点击下载PDF文件
【5】 Robust Lossy Audio Compression Identification
标题: 稳健的有损音频压缩识别
作者:Hendrik Vincent Koops,Gianluca Micchi,Elio Quinton
备注:Accepted to be published in the Proceedings of the 25th International Society for Music Information Retrieval Conference 2024
链接:点击下载PDF文件
【6】 Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
作者:Ziya Zhou,Yuhang Wu,Zhiyue Wu,Xinyue Zhang,Ruibin Yuan,Yinghao Ma,Lu Wang,Emmanouil Benetos,Wei Xue,Yike Guo
备注:Accepted by ISMIR2024
链接:点击下载PDF文件
【7】 Generative Expressive Conversational Speech Synthesis
标题: 生成式表达式对话语音合成
作者:Rui Liu,Yifan Hu,Ren Yi,Yin Xiang,Haizhou Li
备注:14 pages, 6 figures, 8 tables. Accepted by ACM MM 2024
链接:点击下载PDF文件
【8】 On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
标题: 自动语音识别中合成数据生成的文本到语音模型选择问题
作者:Nick Rossenbach,Ralf Schlüter,Sakriani Sakti
备注:Accepted at the SynData4GenAI 2024 workshop
链接:点击下载PDF文件
【9】 TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
标题: TinyChirp:在低功耗无线声学传感器上使用TinyML模型进行鸟鸣识别
作者:Zhaolan Huang,Adrien Tousnakhoff,Polina Kozyr,Roman Rehausen,Felix Bießmann,Robert Lachlan,Cedric Adjih,Emmanuel Baccelli
链接:点击下载PDF文件
【10】 Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
标题: 基于多模式融合和深度学习的笑声识别系统的设计与开发
作者:Fuzheng Zhao,Yu Bai
备注:7 pages,2 figures
链接:点击下载PDF文件
【11】 Computational music analysis from first principles
标题: 从首要原则出发的计算音乐分析
作者:Dmitri Tymoczko,Mark Newman
链接:点击下载PDF文件
【12】 ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
标题: ELP适配器:针对各种语音处理任务的参数高效适配器调整
作者:Nakamasa Inoue,Shinta Otake,Takumi Hirose,Masanari Ohi,Rei Kawakami
链接:点击下载PDF文件
【13】 Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
标题: 使用CycleGAN和域间损失改善端到端ASB中低资源语言的嘈杂学生训练
作者:Chia-Yu Li,Ngoc Thang Vu
备注:10 pages (2 for references), 4 figures, published in SIGUL2024@LREC-COLING 2024
链接:点击下载PDF文件
【14】 Sentiment Reasoning for Healthcare
标题: 医疗保健的情感推理
作者:Khai Le-Duc,Khai-Nguyen Nguyen,Bach Phan Tat,Duy Le,Jerry Ngo,Long Vo-Dang,Anh Totti Nguyen,Truong-Son Hy
备注:Preprint, 18 pages
链接:点击下载PDF文件
【15】 Self-Supervised Models in Automatic Whispered Speech Recognition
标题: 自动耳语语音识别中的自我监督模型
作者:Aref Farhadipour,Homa Asadi,Volker Dellwo
备注:6 pages, 2 figures. Submitted to a conference
链接:点击下载PDF文件
标题: 使用置信度测量和提示将大型语言模型与ASB系统对接
作者:Maryam Naderi,Enno Hermann,Alexandre Nanchen,Sevada Hovsepyan,Mathew Magimai. -Doss
备注:5 pages, 3 figures, 5 tables. Accepted to Interspeech 2024
链接:点击下载PDF文件
【2】 Towards EMG-to-Speech with a Necklace Form Factor
标题: 通过项链外形实现EMG到言语
作者:Peter Wu,Ryan Kaveh,Raghav Nautiyal,Christine Zhang,Albert Guo,Anvitha Kachinthaya,Tavish Mishra,Bohan Yu,Alan W Black,Rikky Muller,Gopala Krishna Anumanchipalli
链接:点击下载PDF文件
【3】 Self-Supervised Models in Automatic Whispered Speech Recognition
标题: 自动耳语语音识别中的自我监督模型
作者:Aref Farhadipour,Homa Asadi,Volker Dellwo
备注:6 pages, 2 figures. Submitted to a conference
链接:点击下载PDF文件
【4】 Cluster and Separate: a GNN Approach to Voice and Staff Prediction for Score Engraving
标题: 集群和分离:分数雕刻的声音和人员预测的GNN方法
作者:Francesco Foscarin,Emmanouil Karystinaios,Eita Nakamura,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval (ISMIR) 2024
链接:点击下载PDF文件
【5】 Beat this! Accurate beat tracking without DBN postprocessing
标题: 打败这个!无需DBN后处理即可准确跟踪节拍
作者:Francesco Foscarin,Jan Schlüter,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
【6】 Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
标题: 通过LLM Agent实现端到端同步语音翻译的人类对等
作者:Shanbo Cheng,Zhichao Huang,Tom Ko,Hang Li,Ningxin Peng,Lu Xu,Qini Zhang
备注:Authors are listed in alphabetical order by last name. Demonstrations and human-annotated test sets are available at this https URL
链接:点击下载PDF文件
【7】 Between the AI and Me: Analysing Listeners' Perspectives on AI- and Human-Composed Progressive Metal Music
标题: 在人工智能和我之间:分析听众对人工智能和人类创作的进步金属音乐的看法
作者:Pedro Sarmento,Jackson Loth,Mathieu Barthet
备注:Reviewed pre-print accepted for publication at ISMIR 2024
链接:点击下载PDF文件
【8】 Enhancing Partially Spoofed Audio Localization with Boundary-aware Attention Mechanism
标题: 利用边界感知注意机制增强部分欺骗音频定位
作者:Jiafeng Zhong,Bin Li,Jiangyan Yi
链接:点击下载PDF文件
【9】 Robust Lossy Audio Compression Identification
标题: 稳健的有损音频压缩识别
作者:Hendrik Vincent Koops,Gianluca Micchi,Elio Quinton
备注:Accepted to be published in the Proceedings of the 25th International Society for Music Information Retrieval Conference 2024
链接:点击下载PDF文件
【10】 Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
作者:Ziya Zhou,Yuhang Wu,Zhiyue Wu,Xinyue Zhang,Ruibin Yuan,Yinghao Ma,Lu Wang,Emmanouil Benetos,Wei Xue,Yike Guo
备注:Accepted by ISMIR2024
链接:点击下载PDF文件
【11】 Generative Expressive Conversational Speech Synthesis
标题: 生成式表达式对话语音合成
作者:Rui Liu,Yifan Hu,Ren Yi,Yin Xiang,Haizhou Li
备注:14 pages, 6 figures, 8 tables. Accepted by ACM MM 2024
链接:点击下载PDF文件
【12】 On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
标题: 自动语音识别中合成数据生成的文本到语音模型选择问题
作者:Nick Rossenbach,Ralf Schlüter,Sakriani Sakti
备注:Accepted at the SynData4GenAI 2024 workshop
链接:点击下载PDF文件
【13】 TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
标题: TinyChirp:在低功耗无线声学传感器上使用TinyML模型进行鸟鸣识别
作者:Zhaolan Huang,Adrien Tousnakhoff,Polina Kozyr,Roman Rehausen,Felix Bießmann,Robert Lachlan,Cedric Adjih,Emmanuel Baccelli
链接:点击下载PDF文件
【14】 Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
标题: 基于多模式融合和深度学习的笑声识别系统的设计与开发
作者:Fuzheng Zhao,Yu Bai
备注:7 pages,2 figures
链接:点击下载PDF文件
【15】 AI Safety in Practice: Enhancing Adversarial Robustness in Multimodal Image Captioning
标题: 实践中的人工智能安全:增强多模式图像字幕中的对抗鲁棒性
作者:Maisha Binte Rashid,Pablo Rivas
备注:Accepted into KDD 2024 workshop on Ethical AI
链接:点击下载PDF文件
【16】 Computational music analysis from first principles
标题: 从首要原则出发的计算音乐分析
作者:Dmitri Tymoczko,Mark Newman
链接:点击下载PDF文件
【17】 ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
标题: ELP适配器:针对各种语音处理任务的参数高效适配器调整
作者:Nakamasa Inoue,Shinta Otake,Takumi Hirose,Masanari Ohi,Rei Kawakami
链接:点击下载PDF文件
【18】 Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
标题: 使用CycleGAN和域间损失改善端到端ASB中低资源语言的嘈杂学生训练
作者:Chia-Yu Li,Ngoc Thang Vu
备注:10 pages (2 for references), 4 figures, published in SIGUL2024@LREC-COLING 2024
链接:点击下载PDF文件
【19】 Sentiment Reasoning for Healthcare
标题: 医疗保健的情感推理
作者:Khai Le-Duc,Khai-Nguyen Nguyen,Bach Phan Tat,Duy Le,Jerry Ngo,Long Vo-Dang,Anh Totti Nguyen,Truong-Son Hy
备注:Preprint, 18 pages
链接:点击下载PDF文件
标题: 打败这个!无需DBN后处理即可准确跟踪节拍
作者:Francesco Foscarin,Jan Schlüter,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
摘要:我们提出了一个系统,用于跟踪节拍和重拍有两个目标:在不同的音乐范围的一般性,和高精度。我们通过在多个数据集上进行训练来实现通用性-包括独奏乐器录音,具有时间签名变化的作品以及具有高节奏变化的古典音乐-并通过删除常用的动态贝叶斯网络(DBN)后处理来实现通用性,该后处理引入了对节拍和节奏的限制。为了实现高准确性,除了其他改进之外,我们开发了一个耐受注释小时移的损失函数,以及一个在频率或时间上与Transformers交替卷积的架构。尽管没有使用DBN,但我们的系统在F1得分方面超越了当前的最先进水平。然而,它仍然可能失败,特别是对于困难和代表性不足的类型,并且在连续性指标上表现得更差,因此我们发布了我们的模型,代码和预处理数据集,并邀请其他人来击败它。摘要:We propose a system for tracking beats and downbeats with two objectives: generality across a diverse music range, and high accuracy. We achieve generality by training on multiple datasets -- including solo instrument recordings, pieces with time signature changes, and classical music with high tempo variations -- and by removing the commonly used Dynamic Bayesian Network (DBN) postprocessing, which introduces constraints on the meter and tempo. For high accuracy, among other improvements, we develop a loss function tolerant to small time shifts of annotations, and an architecture alternating convolutions with transformers either over frequency or time. Our system surpasses the current state of the art in F1 score despite using no DBN. However, it can still fail, especially for difficult and underrepresented genres, and performs worse on continuity metrics, so we publish our model, code, and preprocessed datasets, and invite others to beat this.
【2】 Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
标题: 通过LLM Agent实现端到端同步语音翻译的人类对等
作者:Shanbo Cheng,Zhichao Huang,Tom Ko,Hang Li,Ningxin Peng,Lu Xu,Qini Zhang
备注:Authors are listed in alphabetical order by last name. Demonstrations and human-annotated test sets are available at this https URL
链接:点击下载PDF文件
摘要:本文介绍了一个高质量的、类人的语音同声翻译系统--跨语言智能体--同声传译(CLASI)。受专业口译员的启发,我们采用了一种新颖的数据驱动读写策略来平衡翻译质量和延迟。为了解决翻译领域内术语的挑战,CLASI采用了多模态检索模块来获取相关信息,以增强翻译。在LLM的支持下,我们的方法可以通过考虑输入音频,历史上下文和检索信息来生成容错翻译。实验结果表明,我们的系统优于其他系统的显着利润率。与专业的人类口译员保持一致,我们使用更好的人类评估指标有效信息比例(VIP)来评估CLASI,该指标衡量可以成功传达给听众的信息量。在现实世界的场景中,演讲通常是不流利的,非正式的,不清楚的,CLASI实现了VIP的81.3%和78.0%的中文到英文和英文到中文的翻译方向,分别。相比之下,最先进的商业或开源系统仅达到35.4%和41.6%。在极硬的数据集上,其他系统的VIP低于13%,CLASI仍然可以达到70%的VIP。摘要:In this paper, we present Cross Language Agent -- Simultaneous Interpretation, CLASI, a high-quality and human-like Simultaneous Speech Translation (SiST) System. Inspired by professional human interpreters, we utilize a novel data-driven read-write strategy to balance the translation quality and latency. To address the challenge of translating in-domain terminologies, CLASI employs a multi-modal retrieving module to obtain relevant information to augment the translation. Supported by LLMs, our approach can generate error-tolerated translation by considering the input audio, historical context, and retrieved information. Experimental results show that our system outperforms other systems by significant margins. Aligned with professional human interpreters, we evaluate CLASI with a better human evaluation metric, valid information proportion (VIP), which measures the amount of information that can be successfully conveyed to the listeners. In the real-world scenarios, where the speeches are often disfluent, informal, and unclear, CLASI achieves VIP of 81.3% and 78.0% for Chinese-to-English and English-to-Chinese translation directions, respectively. In contrast, state-of-the-art commercial or open-source systems only achieve 35.4% and 41.6%. On the extremely hard dataset, where other systems achieve under 13% VIP, CLASI can still achieve 70% VIP.
【3】 Between the AI and Me: Analysing Listeners' Perspectives on AI- and Human-Composed Progressive Metal Music
标题: 在人工智能和我之间:分析听众对人工智能和人类创作的进步金属音乐的看法
作者:Pedro Sarmento,Jackson Loth,Mathieu Barthet
备注:Reviewed pre-print accepted for publication at ISMIR 2024
链接:点击下载PDF文件
摘要:生成式人工智能模型最近蓬勃发展,对艺术和音乐传统产生了重大影响。因此,研究人类如何与这些模型互动并认为这些模型至关重要。通过倾听和反思研究,我们探索了参与者对人工智能与人类生成的进步金属的看法,以象征性的形式,使用摇滚乐作为对照组。AI生成的示例由基于Transformer的模型ProgGP生成。我们提出了一种混合方法来评估生成类型(人类与人工智能),流派(进步金属与摇滚)和策展过程(随机与樱桃挑选)的影响。这结合了对体裁一致性、偏好、创造力、一致性、可玩性、人性化和可重复性的定量反馈,以及对听众体验的定性反馈。共有32名渐进金属风扇完成了这项研究。我们的研究结果验证了微调的使用,以实现AI音乐生成的特定类型的专业化,因为听众可以区分AI生成的摇滚和进步金属。尽管一些人工智能生成的摘录获得了与人类音乐相似的评级,但听众表现出对人类作品的偏好。主题分析确定了类型和AI与人类区别的关键特征。最后,我们认为,我们的工作在促进音乐数据的多样性MIR研究的伦理影响,专注于一个未开发的流派。摘要:Generative AI models have recently blossomed, significantly impacting artistic and musical traditions. Research investigating how humans interact with and deem these models is therefore crucial. Through a listening and reflection study, we explore participants' perspectives on AI- vs human-generated progressive metal, in symbolic format, using rock music as a control group. AI-generated examples were produced by ProgGP, a Transformer-based model. We propose a mixed methods approach to assess the effects of generation type (human vs. AI), genre (progressive metal vs. rock), and curation process (random vs. cherry-picked). This combines quantitative feedback on genre congruence, preference, creativity, consistency, playability, humanness, and repeatability, and qualitative feedback to provide insights into listeners' experiences. A total of 32 progressive metal fans completed the study. Our findings validate the use of fine-tuning to achieve genre-specific specialization in AI music generation, as listeners could distinguish between AI-generated rock and progressive metal. Despite some AI-generated excerpts receiving similar ratings to human music, listeners exhibited a preference for human compositions. Thematic analysis identified key features for genre and AI vs. human distinctions. Finally, we consider the ethical implications of our work in promoting musical data diversity within MIR research by focusing on an under-explored genre.
【4】 Enhancing Partially Spoofed Audio Localization with Boundary-aware Attention Mechanism
标题: 利用边界感知注意机制增强部分欺骗音频定位
作者:Jiafeng Zhong,Bin Li,Jiangyan Yi
链接:点击下载PDF文件
摘要:部分欺骗音频定位的任务旨在准确地确定帧级别的音频真实性。虽然一些工作已经取得了令人鼓舞的结果,利用边界信息在一个单一的模型仍然是一个未探索的研究课题。在这项工作中,我们提出了一种新的方法称为边界感知注意机制(BAM)。具体来说,它包括两个核心模块:边界增强和边界框架注意力。前者利用帧内和帧间信息提取具有鉴别力的边界特征,用于边界位置检测和真伪判定;后者利用边界预测结果显式控制帧间特征交互,实现真假帧的有效鉴别。在PartialSpoof数据库上的实验结果表明,该方法具有最佳的性能.该代码可在https: github.com media-sec-lab BAM上获得。摘要:The task of partially spoofed audio localization aims to accurately determine audio authenticity at a frame level. Although some works have achieved encouraging results, utilizing boundary information within a single model remains an unexplored research topic. In this work, we propose a novel method called Boundary-aware Attention Mechanism (BAM). Specifically, it consists of two core modules: Boundary Enhancement and Boundary Frame-wise Attention. The former assembles the intra-frame and inter-frame information to extract discriminative boundary features that are subsequently used for boundary position detection and authenticity decision, while the latter leverages boundary prediction results to explicitly control the feature interaction between frames, which achieves effective discrimination between real and fake frames. Experimental results on PartialSpoof database demonstrate our proposed method achieves the best performance. The code is available at https: github.com media-sec-lab BAM.
【5】 Robust Lossy Audio Compression Identification
标题: 稳健的有损音频压缩识别
作者:Hendrik Vincent Koops,Gianluca Micchi,Elio Quinton
备注:Accepted to be published in the Proceedings of the 25th International Society for Music Information Retrieval Conference 2024
链接:点击下载PDF文件
摘要:以前的研究贡献盲有损压缩识别报告接近完美的性能指标,在他们的测试集上,在各种编解码器和比特率。然而,我们表明,这样的结果可能是欺骗性的,可能无法准确地代表系统处理手头任务的真实能力。在这篇文章中,我们提出了一个有损音频识别模型的鲁棒性和泛化能力的调查。我们的贡献如下:(1)我们表现出缺乏鲁棒性的编解码器参数变化的模型相当于现有技术。特别是,当天真地训练一个有损压缩检测模型的音乐录音处理的数据集与一系列的编解码器及其无损同行,我们获得近乎完美的性能指标上举行的测试集,但严重退化的性能与编解码器参数没有看到在训练中产生的有损轨道。(2)我们提出并显示了改进的训练策略的有效性,以显着提高模型的鲁棒性和泛化能力,超越训练过程中看到的编解码器配置。也就是说,我们对输入频谱图应用随机掩码,以鼓励模型不要仅仅依赖于训练集的编解码器截止频率。摘要:Previous research contributions on blind lossy compression identification report near perfect performance metrics on their test set, across a variety of codecs and bit rates. However, we show that such results can be deceptive and may not accurately represent true ability of the system to tackle the task at hand. In this article, we present an investigation into the robustness and generalisation capability of a lossy audio identification model. Our contributions are as follows. (1) We show the lack of robustness to codec parameter variations of a model equivalent to prior art. In particular, when naively training a lossy compression detection model on a dataset of music recordings processed with a range of codecs and their lossless counterparts, we obtain near perfect performance metrics on the held-out test set, but severely degraded performance on lossy tracks produced with codec parameters not seen in training. (2) We propose and show the effectiveness of an improved training strategy to significantly increase the robustness and generalisation capability of the model beyond codec configurations seen during training. Namely we apply a random mask to the input spectrogram to encourage the model not to rely solely on the training set's codec cutoff frequency.
【6】 Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
作者:Ziya Zhou,Yuhang Wu,Zhiyue Wu,Xinyue Zhang,Ruibin Yuan,Yinghao Ma,Lu Wang,Emmanouil Benetos,Wei Xue,Yike Guo
备注:Accepted by ISMIR2024
链接:点击下载PDF文件
摘要:象征性的音乐,类似于语言,可以被编码成离散的符号。最近的研究已经将诸如GPT-4和Llama 2等大型语言模型(LLM)的应用扩展到符号音乐领域,包括理解和生成。然而,很少有研究探讨这些LLM如何在高级音乐理解和条件生成方面表现的细节,特别是从多步推理的角度来看,这是条件,可编辑和交互式人机共同创作过程中的一个关键方面。本研究旨在探讨学习记忆者在符号音乐处理上的能力与局限性。我们发现,目前的LLM在歌曲级多步音乐推理方面表现不佳,并且在处理复杂的音乐任务时通常无法利用所学的音乐知识。我们的研究结果表明,实现先进的音乐能力是不是固有的LLM获得的,未来的研究应该更多地集中在弥合音乐知识和推理之间的差距,以改善音乐家的共同创作体验。摘要:Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the symbolic music domain including understanding and generation. Yet scant research explores the details of how these LLMs perform on advanced music understanding and conditioned generation, especially from the multi-step reasoning perspective, which is a critical aspect in the conditioned, editable, and interactive human-computer co-creation process. This study conducts a thorough investigation of LLMs' capability and limitations in symbolic music processing. We identify that current LLMs exhibit poor performance in song-level multi-step music reasoning, and typically fail to leverage learned music knowledge when addressing complex musical tasks. An analysis of LLMs' responses highlights distinctly their pros and cons. Our findings suggest achieving advanced musical capability is not intrinsically obtained by LLMs, and future research should focus more on bridging the gap between music knowledge and reasoning, to improve the co-creation experience for musicians.
【7】 Generative Expressive Conversational Speech Synthesis
标题: 生成式表达式对话语音合成
作者:Rui Liu,Yifan Hu,Ren Yi,Yin Xiang,Haizhou Li
备注:14 pages, 6 figures, 8 tables. Accepted by ACM MM 2024
链接:点击下载PDF文件
摘要:会话语音合成(CSS)的目标是在用户-代理会话环境中以适当的说话风格表达目标话语。现有的CSS方法采用有效的多模态上下文建模技术来实现移情理解和表达。然而,他们通常需要设计复杂的网络架构,并精心优化其中的模块。此外,由于包含脚本记录风格的小规模数据集的限制,它们通常无法模拟真实的自然会话风格。为了解决上述问题,我们提出了一个新的生成式表达CSS系统,称为GPT-Talker,我们将多轮对话历史的多模态信息转换为离散的令牌序列,并将它们无缝集成,形成一个全面的用户代理对话上下文。利用GPT的力量,我们预测的令牌序列,其中包括语义和风格的知识,为代理的响应。在此基础上,我们提出了一个大规模的自然CSS数据集NCSSD,它既包含了自然录制的即兴风格的会话语音,也包含了从电视节目中提取的对话。我们对NCSSD的可靠性和GPT-Talker的有效性进行了全面的实验。主观和客观评估都表明,我们的模型在自然度和表现力方面显着优于其他最先进的CSS系统。代码、数据集和预训练模型可在https: github.com AI-S2-Lab GPT-Talker上获得。摘要:Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy understanding and expression. However, they often need to design complex network architectures and meticulously optimize the modules within them. In addition, due to the limitations of small-scale datasets containing scripted recording styles, they often fail to simulate real natural conversational styles. To address the above issues, we propose a novel generative expressive CSS system, termed GPT-Talker.We transform the multimodal information of the multi-turn dialogue history into discrete token sequences and seamlessly integrate them to form a comprehensive user-agent dialogue context. Leveraging the power of GPT, we predict the token sequence, that includes both semantic and style knowledge, of response for the agent. After that, the expressive conversational speech is synthesized by the conversation-enriched VITS to deliver feedback to the user.Furthermore, we propose a large-scale Natural CSS Dataset called NCSSD, that includes both naturally recorded conversational speech in improvised styles and dialogues extracted from TV shows. It encompasses both Chinese and English languages, with a total duration of 236 hours.We conducted comprehensive experiments on the reliability of the NCSSD and the effectiveness of our GPT-Talker. Both subjective and objective evaluations demonstrate that our model outperforms other state-of-the-art CSS systems significantly in terms of naturalness and expressiveness. The Code, Dataset, and Pre-trained Model are available at: https: github.com AI-S2-Lab GPT-Talker.
【8】 On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
标题: 自动语音识别中合成数据生成的文本到语音模型选择问题
作者:Nick Rossenbach,Ralf Schlüter,Sakriani Sakti
备注:Accepted at the SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:神经文本到语音(TTS)系统的快速发展使其能够用于自然语言处理的其他领域,如自动语音识别(ASR)或口语翻译(ESTA)。由于大量不同的TTS体系结构及其扩展,选择哪种TTS系统用于合成数据创建并不是一件容易的事情。我们使用五种不同的TTS解码器架构的合成数据生成的范围内的比较,以显示基于CTC的语音识别训练的影响。我们将识别结果与NISQA MOS和可懂度等可计算指标进行比较,发现与ASR性能没有明确的关系。我们还观察到,对于数据生成自回归解码比非自回归解码表现更好,并提出了一种方法来量化TTS泛化能力。摘要:The rapid development of neural text-to-speech (TTS) systems enabled its usage in other areas of natural language processing such as automatic speech recognition (ASR) or spoken language translation (SLT). Due to the large number of different TTS architectures and their extensions, selecting which TTS systems to use for synthetic data creation is not an easy task. We use the comparison of five different TTS decoder architectures in the scope of synthetic data generation to show the impact on CTC-based speech recognition training. We compare the recognition results to computable metrics like NISQA MOS and intelligibility, finding that there are no clear relations to the ASR performance. We also observe that for data generation auto-regressive decoding performs better than non-autoregressive decoding, and propose an approach to quantify TTS generalization capabilities.
【9】 TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
标题: TinyChirp:在低功耗无线声学传感器上使用TinyML模型进行鸟鸣识别
作者:Zhaolan Huang,Adrien Tousnakhoff,Polina Kozyr,Roman Rehausen,Felix Bießmann,Robert Lachlan,Cedric Adjih,Emmanuel Baccelli
链接:点击下载PDF文件
摘要:大规模监测生物多样性具有挑战性。在细粒度分类中检测和识别物种需要高度准确的机器学习(ML)方法。训练这样的模型需要大量高质量的数据集。将这些模型部署到低功耗设备需要新颖的压缩技术和模型架构。虽然物种分类方法受益于新的数据集和ML方法的进步,特别是神经网络,但将这些最先进的模型部署到低功耗设备仍然很困难。在这里,我们对用于物种分类的各种tinyML神经网络架构和压缩技术进行了全面的实证比较。我们专注于鸟鸣检测的例子,更具体地说,一个数据集策划研究玉米bunting鸟类。数据集与本研究的所有代码和实验一起发布。在我们的实验中,我们比较预测性能,内存和时间复杂度的经典的基于频谱图的方法和最近的方法操作的原始音频信号。我们的研究结果表明,个别鸟类物种可以稳健地检测到相对简单的架构,可以很容易地部署到低功耗设备。摘要:Monitoring biodiversity at scale is challenging. Detecting and identifying species in fine grained taxonomies requires highly accurate machine learning (ML) methods. Training such models requires large high quality data sets. And deploying these models to low power devices requires novel compression techniques and model architectures. While species classification methods have profited from novel data sets and advances in ML methods, in particular neural networks, deploying these state of the art models to low power devices remains difficult. Here we present a comprehensive empirical comparison of various tinyML neural network architectures and compression techniques for species classification. We focus on the example of bird song detection, more concretely a data set curated for studying the corn bunting bird species. The data set is released along with all code and experiments of this study. In our experiments we compare predictive performance, memory and time complexity of classical spectrogram based methods and recent approaches operating on raw audio signal. Our results indicate that individual bird species can be robustly detected with relatively simple architectures that can be readily deployed to low power devices.
【10】 Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
标题: 基于多模式融合和深度学习的笑声识别系统的设计与开发
作者:Fuzheng Zhao,Yu Bai
备注:7 pages,2 figures
链接:点击下载PDF文件
摘要:本研究旨在设计并实现一个基于多模态融合和深度学习的笑声识别系统,利用图像和音频处理技术实现准确的笑声识别和情感分析。首先,系统加载视频文件并使用OpenCV库提取面部信息,同时使用Librosa库处理音频功能,如MFCC。然后,使用多模态融合技术来整合图像和音频特征,然后使用深度学习模型进行训练和预测。评估结果表明,该模型在测试数据集上实现了80%的准确度、精确度和召回率,F1得分为80%,表现出强大的性能和处理真实世界数据变化的能力。该研究不仅验证了多模态融合方法在笑声识别中的有效性,而且突出了其在情感计算和人机交互中的潜在应用。未来的工作将集中在进一步优化特征提取和模型架构,以提高识别准确率和扩展应用场景,促进笑声识别技术在心理健康监测和教育活动评估等领域的发展摘要:This study aims to design and implement a laughter recognition system based on multimodal fusion and deep learning, leveraging image and audio processing technologies to achieve accurate laughter recognition and emotion analysis. First, the system loads video files and uses the OpenCV library to extract facial information while employing the Librosa library to process audio features such as MFCC. Then, multimodal fusion techniques are used to integrate image and audio features, followed by training and prediction using deep learning models. Evaluation results indicate that the model achieved 80% accuracy, precision, and recall on the test dataset, with an F1 score of 80%, demonstrating robust performance and the ability to handle real-world data variability. This study not only verifies the effectiveness of multimodal fusion methods in laughter recognition but also highlights their potential applications in affective computing and human-computer interaction. Future work will focus on further optimizing feature extraction and model architecture to improve recognition accuracy and expand application scenarios, promoting the development of laughter recognition technology in fields such as mental health monitoring and educational activity evaluation
【11】 Computational music analysis from first principles
标题: 从首要原则出发的计算音乐分析
作者:Dmitri Tymoczko,Mark Newman
链接:点击下载PDF文件
摘要:我们使用耦合隐马尔可夫模型自动注释的371巴赫合唱团在Riemenschneider版,语料库包含约100,000个音符和20,000和弦。我们给出了三个独立的分析,实现逐步更大的准确性,在对音乐句法越来越强的假设的成本。虽然我们的方法几乎不使用人工输入,但与专家人工分析相比,我们能够以85%或更高的准确度识别和弦和键,从而使注释足够准确,可用于一系列音乐理论目的,同时也不受主观人类判断的影响。我们的工作承担长期的争论的客观现实的结构所假设的标准西方和声理论,以及具体问题的性质西方和声句法。摘要:We use coupled hidden Markov models to automatically annotate the 371 Bach chorales in the Riemenschneider edition, a corpus containing approximately 100,000 notes and 20,000 chords. We give three separate analyses that achieve progressively greater accuracy at the cost of making increasingly strong assumptions about musical syntax. Although our method makes almost no use of human input, we are able to identify both chords and keys with an accuracy of 85% or greater when compared to an expert human analysis, resulting in annotations accurate enough to be used for a range of music-theoretical purposes, while also being free of subjective human judgments. Our work bears on longstanding debates about the objective reality of the structures postulated by standard Western harmonic theory, as well as on specific questions about the nature of Western harmonic syntax.
【12】 ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
标题: ELP适配器:针对各种语音处理任务的参数高效适配器调整
作者:Nakamasa Inoue,Shinta Otake,Takumi Hirose,Masanari Ohi,Rei Kawakami
链接:点击下载PDF文件
摘要:自监督学习已经成为从语音数据中学习通用表示的关键方法。尽管在语音识别、说话人验证和情感识别等下游任务中取得了令人鼓舞的结果,但需要大量参数,这使得对每个任务的微调效率低下。为了解决这一限制,我们引入ELP适配器调整,一种新的方法,使用三种类型的适配器,即编码器适配器(E-适配器),层适配器(L-适配器),和提示适配器(P-适配器)的参数有效的微调。E-adapter集成到基于transformer的编码器层中,有助于学习对语音识别有效的细粒度语音表示。L适配器创建从每个编码器层到下游头部的路径,并帮助从较低的编码器层提取非语言特征,这些特征对说话人验证和情感识别有效。P-adapter将伪特征附加到CNN特征,以进一步提高有效性和效率。有了这些适配器,模型可以快速适应各种语音处理任务。我们的评估在四个下游任务使用五个骨干模型证明了所提出的方法的有效性。使用WavLM主干,其性能与所有任务的完全微调相当或更好,同时需要减少90%的可学习参数。摘要:Self-supervised learning has emerged as a key approach for learning generic representations from speech data. Despite promising results in downstream tasks such as speech recognition, speaker verification, and emotion recognition, a significant number of parameters is required, which makes fine-tuning for each task memory-inefficient. To address this limitation, we introduce ELP-adapter tuning, a novel method for parameter-efficient fine-tuning using three types of adapter, namely encoder adapters (E-adapters), layer adapters (L-adapters), and a prompt adapter (P-adapter). The E-adapters are integrated into transformer-based encoder layers and help to learn fine-grained speech representations that are effective for speech recognition. The L-adapters create paths from each encoder layer to the downstream head and help to extract non-linguistic features from lower encoder layers that are effective for speaker verification and emotion recognition. The P-adapter appends pseudo features to CNN features to further improve effectiveness and efficiency. With these adapters, models can be quickly adapted to various speech processing tasks. Our evaluation across four downstream tasks using five backbone models demonstrated the effectiveness of the proposed method. With the WavLM backbone, its performance was comparable to or better than that of full fine-tuning on all tasks while requiring 90% fewer learnable parameters.
【13】 Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
标题: 使用CycleGAN和域间损失改善端到端ASB中低资源语言的嘈杂学生训练
作者:Chia-Yu Li,Ngoc Thang Vu
备注:10 pages (2 for references), 4 figures, published in SIGUL2024@LREC-COLING 2024
链接:点击下载PDF文件
摘要:使用噪声学生训练来训练半监督端到端语音识别系统显著提高了性能。然而,这种方法需要大量成对的语音文本和未标记的语音,这对于低资源语言来说是昂贵的。因此,本文考虑了半监督端到端自动语音识别的一个更极端的情况,其中有有限的配对语音文本,未标记的语音(少于五个小时),和丰富的外部文本。首先,我们通过使用我们以前在半监督学习“CycleGAN和域间损失”方面的工作单独使用外部文本来训练模型,从而观察到性能的提高。其次,我们通过结合自动超参数调整来增强“CycleGAN和域间损失”,称之为“增强的CycleGAN域间损失”。“第三,我们将其整合到低资源场景的嘈杂学生培训方法管道中。我们在Voxforge和Common Voice的六种非英语语言上进行的实验结果显示,与基线教师模型相比,单词错误率降低了20%,与基线最佳学生模型相比,单词错误率降低了10%,突出了通过我们提出的方法取得的显着改进。摘要:Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabeled speech, which is costly for low-resource languages. Therefore, this paper considers a more extreme case of semi-supervised end-to-end automatic speech recognition where there are limited paired speech-text, unlabeled speech (less than five hours), and abundant external text. Firstly, we observe improved performance by training the model using our previous work on semi-supervised learning "CycleGAN and inter-domain losses" solely with external text. Secondly, we enhance "CycleGAN and inter-domain losses" by incorporating automatic hyperparameter tuning, calling it "enhanced CycleGAN inter-domain losses." Thirdly, we integrate it into the noisy student training approach pipeline for low-resource scenarios. Our experimental results, conducted on six non-English languages from Voxforge and Common Voice, show a 20% word error rate reduction compared to the baseline teacher model and a 10% word error rate reduction compared to the baseline best student model, highlighting the significant improvements achieved through our proposed method.
【14】 Sentiment Reasoning for Healthcare
标题: 医疗保健的情感推理
作者:Khai Le-Duc,Khai-Nguyen Nguyen,Bach Phan Tat,Duy Le,Jerry Ngo,Long Vo-Dang,Anh Totti Nguyen,Truong-Son Hy
备注:Preprint, 18 pages
链接:点击下载PDF文件
摘要:由于错误的严重后果,人工智能决策的透明度在医疗保健中至关重要,这对于在情感分析任务中建立人工智能和用户之间的信任非常重要。增强推理能力有助于大型语言模型(LLM)在更广泛的背景下理解人类情感,处理细微差别和模糊的语言,并推断可能没有明确说明的潜在情感。在这项工作中,我们引入了一个新的任务-情感推理-语音和文本模态,以及我们提出的多模态多任务框架和数据集。我们的研究表明,理性增强训练增强了模型在人类转录本和ASR设置中的情感分类性能。此外,我们发现,生成的理由通常表现出不同的词汇相比,人类生成的理由,但保持类似的语义。所有代码、数据(英语翻译和越南语)和模型均在线发布:https: github.com leduckhai MultiMed摘要:Transparency in AI decision-making is crucial in healthcare due to the severe consequences of errors, and this is important for building trust among AI and users in sentiment analysis task. Incorporating reasoning capabilities helps Large Language Models (LLMs) understand human emotions within broader contexts, handle nuanced and ambiguous language, and infer underlying sentiments that may not be explicitly stated. In this work, we introduce a new task - Sentiment Reasoning - for both speech and text modalities, along with our proposed multimodal multitask framework and dataset. Our study showed that rationale-augmented training enhances model performance in sentiment classification across both human transcript and ASR settings. Also, we found that the generated rationales typically exhibit different vocabularies compared to human-generated rationales, but maintain similar semantics. All code, data (English-translated and Vietnamese) and models are published online: https: github.com leduckhai MultiMed
【15】 Self-Supervised Models in Automatic Whispered Speech Recognition
标题: 自动耳语语音识别中的自我监督模型
作者:Aref Farhadipour,Homa Asadi,Volker Dellwo
备注:6 pages, 2 figures. Submitted to a conference
链接:点击下载PDF文件
摘要:在自动语音识别中,任何改变语音声学特性的因素都会对系统的性能提出挑战。本文提出了一种新的方法自动耳语音识别爱尔兰方言使用自监督WavLM模型。传统的自动语音识别系统往往无法准确地识别耳语,由于其独特的声学特性和相关的训练数据的稀缺性。为了应对这一挑战,我们使用了一个预先训练好的WavLM模型,该模型结合了来自wTIMIT和CHAINS数据集的耳语和正常语音数据进行了微调,这些数据集分别包括新加坡和爱尔兰方言中的英语。我们使用OpenAI Whisper模型进行的基线评估突出了其局限性,在低声讲话中实现了18.8%的单词错误率(WER)。相比之下,所提出的基于WavLM的系统显着提高了性能,实现了9.22%的WER。这些结果表明,我们的方法在识别耳语语音的有效性,并强调了强大的自动语音识别系统定制的声学建模的重要性。这项研究为开发有效的自动语音识别解决方案提供了有价值的见解,以应对受耳语和方言影响的挑战性语音。本文的源代码是免费提供的。摘要:In automatic speech recognition, any factor that alters the acoustic properties of speech can pose a challenge to the system's performance. This paper presents a novel approach for automatic whispered speech recognition in the Irish dialect using the self-supervised WavLM model. Conventional automatic speech recognition systems often fail to accurately recognise whispered speech due to its distinct acoustic properties and the scarcity of relevant training data. To address this challenge, we utilized a pre-trained WavLM model, fine-tuned with a combination of whispered and normal speech data from the wTIMIT and CHAINS datasets, which include the English language in Singaporean and Irish dialects, respectively. Our baseline evaluation with the OpenAI Whisper model highlighted its limitations, achieving a Word Error Rate (WER) of 18.8% on whispered speech. In contrast, the proposed WavLM-based system significantly improved performance, achieving a WER of 9.22%. These results demonstrate the efficacy of our approach in recognising whispered speech and underscore the importance of tailored acoustic modeling for robust automatic speech recognition systems. This study provides valuable insights into developing effective automatic speech recognition solutions for challenging speech affected by whisper and dialect. The source codes for this paper are freely available.
eess.AS音频处理
【1】 Towards interfacing large language models with ASR systems using confidence measures and prompting标题: 使用置信度测量和提示将大型语言模型与ASB系统对接
作者:Maryam Naderi,Enno Hermann,Alexandre Nanchen,Sevada Hovsepyan,Mathew Magimai. -Doss
备注:5 pages, 3 figures, 5 tables. Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:随着大型语言模型(LLM)在参数大小和功能(例如通过提示进行交互)方面的增长,它们开辟了与自动语音识别(ASR)系统接口的新方法,而不仅仅是重新评分n-best列表。这项工作调查的ASR成绩单与LLM事后校正。为了避免引入错误到可能准确的成绩单,我们提出了一系列基于信心的过滤方法。我们的研究结果表明,这可以提高竞争力较低的ASR系统的性能。摘要:As large language models (LLMs) grow in parameter size and capabilities, such as interaction through prompting, they open up new ways of interfacing with automatic speech recognition (ASR) systems beyond rescoring n-best lists. This work investigates post-hoc correction of ASR transcripts with LLMs. To avoid introducing errors into likely accurate transcripts, we propose a range of confidence-based filtering methods. Our results indicate that this can improve the performance of less competitive ASR systems.
【2】 Towards EMG-to-Speech with a Necklace Form Factor
标题: 通过项链外形实现EMG到言语
作者:Peter Wu,Ryan Kaveh,Raghav Nautiyal,Christine Zhang,Albert Guo,Anvitha Kachinthaya,Tavish Mishra,Bohan Yu,Alan W Black,Rikky Muller,Gopala Krishna Anumanchipalli
链接:点击下载PDF文件
摘要:用于从肌电图(EMG)解码语音的电极通常放置在脸上,需要粘合剂,如果经常使用,粘合剂是不方便的并且刺激皮肤。我们探索了一种不同的设备形状因子,其中干电极被放置在颈部周围。在使用该设备记录的数据上训练的11个字,多说话者语音EMG分类器实现了92.7%的准确率。消融研究揭示了在颈部有两个以上电极的重要性,语音分析揭示了颈部和颈部和面部形状因素之间类似的分类混淆。最后,语音-EMG相关实验证明了许多EMG频谱图频率段和自监督语音表征维度之间的线性关系。摘要:Electrodes for decoding speech from electromyography (EMG) are typically placed on the face, requiring adhesives that are inconvenient and skin-irritating if used regularly. We explore a different device form factor, where dry electrodes are placed around the neck instead. 11-word, multi-speaker voiced EMG classifiers trained on data recorded with this device achieve 92.7% accuracy. Ablation studies reveal the importance of having more than two electrodes on the neck, and phonological analyses reveal similar classification confusions between neck-only and neck-and-face form factors. Finally, speech-EMG correlation experiments demonstrate a linear relationship between many EMG spectrogram frequency bins and self-supervised speech representation dimensions.
【3】 Self-Supervised Models in Automatic Whispered Speech Recognition
标题: 自动耳语语音识别中的自我监督模型
作者:Aref Farhadipour,Homa Asadi,Volker Dellwo
备注:6 pages, 2 figures. Submitted to a conference
链接:点击下载PDF文件
摘要:在自动语音识别中,任何改变语音声学特性的因素都会对系统的性能提出挑战。本文提出了一种新的方法自动耳语音识别爱尔兰方言使用自监督WavLM模型。传统的自动语音识别系统往往无法准确地识别耳语,由于其独特的声学特性和相关的训练数据的稀缺性。为了应对这一挑战,我们使用了一个预先训练好的WavLM模型,该模型结合了来自wTIMIT和CHAINS数据集的耳语和正常语音数据进行了微调,这些数据集分别包括新加坡和爱尔兰方言中的英语。我们使用OpenAI Whisper模型进行的基线评估突出了其局限性,在低声讲话中实现了18.8%的单词错误率(WER)。相比之下,所提出的基于WavLM的系统显着提高了性能,实现了9.22%的WER。这些结果表明,我们的方法在识别耳语语音的有效性,并强调了强大的自动语音识别系统定制的声学建模的重要性。这项研究为开发有效的自动语音识别解决方案提供了有价值的见解,以应对受耳语和方言影响的挑战性语音。本文的源代码是免费提供的。摘要:In automatic speech recognition, any factor that alters the acoustic properties of speech can pose a challenge to the system's performance. This paper presents a novel approach for automatic whispered speech recognition in the Irish dialect using the self-supervised WavLM model. Conventional automatic speech recognition systems often fail to accurately recognise whispered speech due to its distinct acoustic properties and the scarcity of relevant training data. To address this challenge, we utilized a pre-trained WavLM model, fine-tuned with a combination of whispered and normal speech data from the wTIMIT and CHAINS datasets, which include the English language in Singaporean and Irish dialects, respectively. Our baseline evaluation with the OpenAI Whisper model highlighted its limitations, achieving a Word Error Rate (WER) of 18.8% on whispered speech. In contrast, the proposed WavLM-based system significantly improved performance, achieving a WER of 9.22%. These results demonstrate the efficacy of our approach in recognising whispered speech and underscore the importance of tailored acoustic modeling for robust automatic speech recognition systems. This study provides valuable insights into developing effective automatic speech recognition solutions for challenging speech affected by whisper and dialect. The source codes for this paper are freely available.
【4】 Cluster and Separate: a GNN Approach to Voice and Staff Prediction for Score Engraving
标题: 集群和分离:分数雕刻的声音和人员预测的GNN方法
作者:Francesco Foscarin,Emmanouil Karystinaios,Eita Nakamura,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval (ISMIR) 2024
链接:点击下载PDF文件
摘要:本文探讨了从量化的符号音乐作品中分离音符的问题(例如,一个小文件)成多个声音和五线谱。这是乐谱雕刻(或乐谱排版)这一更大任务的基本部分,其目的是为人类表演者制作可读的乐谱。我们专注于钢琴音乐和支持谐音的声音,即,可以包含和弦的声音,以及跨五线谱的声音,这些都是在以前的研究中经常被忽视的非常困难的任务。我们提出了一个基于图神经网络的端到端系统,该系统将属于同一和弦的音符聚类,并将它们与语音的一部分连接起来。我们的研究结果表明,在两个不同风格的数据集上,与以前的方法相比,有了明显和一致的改进。为了帮助对结果进行定性分析,我们支持符号音乐格式的导出,并提供乐谱上输出图形的直接可视化。所有代码和预训练模型均可在https: github.com CPJKU piano_svsep上获得摘要:This paper approaches the problem of separating the notes from a quantized symbolic music piece (e.g., a MIDI file) into multiple voices and staves. This is a fundamental part of the larger task of music score engraving (or score typesetting), which aims to produce readable musical scores for human performers. We focus on piano music and support homophonic voices, i.e., voices that can contain chords, and cross-staff voices, which are notably difficult tasks that have often been overlooked in previous research. We propose an end-to-end system based on graph neural networks that clusters notes that belong to the same chord and connects them with edges if they are part of a voice. Our results show clear and consistent improvements over a previous approach on two datasets of different styles. To aid the qualitative analysis of our results, we support the export in symbolic music formats and provide a direct visualization of our outputs graph over the musical score. All code and pre-trained models are available at https: github.com CPJKU piano_svsep
【5】 Beat this! Accurate beat tracking without DBN postprocessing
标题: 打败这个!无需DBN后处理即可准确跟踪节拍
作者:Francesco Foscarin,Jan Schlüter,Gerhard Widmer
备注:Accepted at the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024
链接:点击下载PDF文件
摘要:我们提出了一个系统,用于跟踪节拍和重拍有两个目标:在不同的音乐范围的一般性,和高精度。我们通过在多个数据集上进行训练来实现通用性-包括独奏乐器录音,具有时间签名变化的作品以及具有高节奏变化的古典音乐-并通过删除常用的动态贝叶斯网络(DBN)后处理来实现通用性,该后处理引入了对节拍和节奏的限制。为了实现高准确性,除了其他改进之外,我们开发了一个耐受注释小时移的损失函数,以及一个在频率或时间上与Transformers交替卷积的架构。尽管没有使用DBN,但我们的系统在F1得分方面超越了当前的最先进水平。然而,它仍然可能失败,特别是对于困难和代表性不足的类型,并且在连续性指标上表现得更差,因此我们发布了我们的模型,代码和预处理数据集,并邀请其他人来击败它。摘要:We propose a system for tracking beats and downbeats with two objectives: generality across a diverse music range, and high accuracy. We achieve generality by training on multiple datasets -- including solo instrument recordings, pieces with time signature changes, and classical music with high tempo variations -- and by removing the commonly used Dynamic Bayesian Network (DBN) postprocessing, which introduces constraints on the meter and tempo. For high accuracy, among other improvements, we develop a loss function tolerant to small time shifts of annotations, and an architecture alternating convolutions with transformers either over frequency or time. Our system surpasses the current state of the art in F1 score despite using no DBN. However, it can still fail, especially for difficult and underrepresented genres, and performs worse on continuity metrics, so we publish our model, code, and preprocessed datasets, and invite others to beat this.
【6】 Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
标题: 通过LLM Agent实现端到端同步语音翻译的人类对等
作者:Shanbo Cheng,Zhichao Huang,Tom Ko,Hang Li,Ningxin Peng,Lu Xu,Qini Zhang
备注:Authors are listed in alphabetical order by last name. Demonstrations and human-annotated test sets are available at this https URL
链接:点击下载PDF文件
摘要:本文介绍了一个高质量的、类人的语音同声翻译系统--跨语言智能体--同声传译(CLASI)。受专业口译员的启发,我们采用了一种新颖的数据驱动读写策略来平衡翻译质量和延迟。为了解决翻译领域内术语的挑战,CLASI采用了多模态检索模块来获取相关信息,以增强翻译。在LLM的支持下,我们的方法可以通过考虑输入音频,历史上下文和检索信息来生成容错翻译。实验结果表明,我们的系统优于其他系统的显着利润率。与专业的人类口译员保持一致,我们使用更好的人类评估指标有效信息比例(VIP)来评估CLASI,该指标衡量可以成功传达给听众的信息量。在现实世界的场景中,演讲通常是不流利的,非正式的,不清楚的,CLASI实现了VIP的81.3%和78.0%的中文到英文和英文到中文的翻译方向,分别。相比之下,最先进的商业或开源系统仅达到35.4%和41.6%。在极硬的数据集上,其他系统的VIP低于13%,CLASI仍然可以达到70%的VIP。摘要:In this paper, we present Cross Language Agent -- Simultaneous Interpretation, CLASI, a high-quality and human-like Simultaneous Speech Translation (SiST) System. Inspired by professional human interpreters, we utilize a novel data-driven read-write strategy to balance the translation quality and latency. To address the challenge of translating in-domain terminologies, CLASI employs a multi-modal retrieving module to obtain relevant information to augment the translation. Supported by LLMs, our approach can generate error-tolerated translation by considering the input audio, historical context, and retrieved information. Experimental results show that our system outperforms other systems by significant margins. Aligned with professional human interpreters, we evaluate CLASI with a better human evaluation metric, valid information proportion (VIP), which measures the amount of information that can be successfully conveyed to the listeners. In the real-world scenarios, where the speeches are often disfluent, informal, and unclear, CLASI achieves VIP of 81.3% and 78.0% for Chinese-to-English and English-to-Chinese translation directions, respectively. In contrast, state-of-the-art commercial or open-source systems only achieve 35.4% and 41.6%. On the extremely hard dataset, where other systems achieve under 13% VIP, CLASI can still achieve 70% VIP.
【7】 Between the AI and Me: Analysing Listeners' Perspectives on AI- and Human-Composed Progressive Metal Music
标题: 在人工智能和我之间:分析听众对人工智能和人类创作的进步金属音乐的看法
作者:Pedro Sarmento,Jackson Loth,Mathieu Barthet
备注:Reviewed pre-print accepted for publication at ISMIR 2024
链接:点击下载PDF文件
摘要:生成式人工智能模型最近蓬勃发展,对艺术和音乐传统产生了重大影响。因此,研究人类如何与这些模型互动并认为这些模型至关重要。通过倾听和反思研究,我们探索了参与者对人工智能与人类生成的进步金属的看法,以象征性的形式,使用摇滚乐作为对照组。AI生成的示例由基于Transformer的模型ProgGP生成。我们提出了一种混合方法来评估生成类型(人类与人工智能),流派(进步金属与摇滚)和策展过程(随机与樱桃挑选)的影响。这结合了对体裁一致性、偏好、创造力、一致性、可玩性、人性化和可重复性的定量反馈,以及对听众体验的定性反馈。共有32名渐进金属风扇完成了这项研究。我们的研究结果验证了微调的使用,以实现AI音乐生成的特定类型的专业化,因为听众可以区分AI生成的摇滚和进步金属。尽管一些人工智能生成的摘录获得了与人类音乐相似的评级,但听众表现出对人类作品的偏好。主题分析确定了类型和AI与人类区别的关键特征。最后,我们认为,我们的工作在促进音乐数据的多样性MIR研究的伦理影响,专注于一个未开发的流派。摘要:Generative AI models have recently blossomed, significantly impacting artistic and musical traditions. Research investigating how humans interact with and deem these models is therefore crucial. Through a listening and reflection study, we explore participants' perspectives on AI- vs human-generated progressive metal, in symbolic format, using rock music as a control group. AI-generated examples were produced by ProgGP, a Transformer-based model. We propose a mixed methods approach to assess the effects of generation type (human vs. AI), genre (progressive metal vs. rock), and curation process (random vs. cherry-picked). This combines quantitative feedback on genre congruence, preference, creativity, consistency, playability, humanness, and repeatability, and qualitative feedback to provide insights into listeners' experiences. A total of 32 progressive metal fans completed the study. Our findings validate the use of fine-tuning to achieve genre-specific specialization in AI music generation, as listeners could distinguish between AI-generated rock and progressive metal. Despite some AI-generated excerpts receiving similar ratings to human music, listeners exhibited a preference for human compositions. Thematic analysis identified key features for genre and AI vs. human distinctions. Finally, we consider the ethical implications of our work in promoting musical data diversity within MIR research by focusing on an under-explored genre.
【8】 Enhancing Partially Spoofed Audio Localization with Boundary-aware Attention Mechanism
标题: 利用边界感知注意机制增强部分欺骗音频定位
作者:Jiafeng Zhong,Bin Li,Jiangyan Yi
链接:点击下载PDF文件
摘要:部分欺骗音频定位的任务旨在准确地确定帧级别的音频真实性。虽然一些工作已经取得了令人鼓舞的结果,利用边界信息在一个单一的模型仍然是一个未探索的研究课题。在这项工作中,我们提出了一种新的方法称为边界感知注意机制(BAM)。具体来说,它包括两个核心模块:边界增强和边界框架注意力。前者利用帧内和帧间信息提取具有鉴别力的边界特征,用于边界位置检测和真伪判定;后者利用边界预测结果显式控制帧间特征交互,实现真假帧的有效鉴别。在PartialSpoof数据库上的实验结果表明,该方法具有最佳的性能.该代码可在https: github.com media-sec-lab BAM上获得。摘要:The task of partially spoofed audio localization aims to accurately determine audio authenticity at a frame level. Although some works have achieved encouraging results, utilizing boundary information within a single model remains an unexplored research topic. In this work, we propose a novel method called Boundary-aware Attention Mechanism (BAM). Specifically, it consists of two core modules: Boundary Enhancement and Boundary Frame-wise Attention. The former assembles the intra-frame and inter-frame information to extract discriminative boundary features that are subsequently used for boundary position detection and authenticity decision, while the latter leverages boundary prediction results to explicitly control the feature interaction between frames, which achieves effective discrimination between real and fake frames. Experimental results on PartialSpoof database demonstrate our proposed method achieves the best performance. The code is available at https: github.com media-sec-lab BAM.
【9】 Robust Lossy Audio Compression Identification
标题: 稳健的有损音频压缩识别
作者:Hendrik Vincent Koops,Gianluca Micchi,Elio Quinton
备注:Accepted to be published in the Proceedings of the 25th International Society for Music Information Retrieval Conference 2024
链接:点击下载PDF文件
摘要:以前的研究贡献盲有损压缩识别报告接近完美的性能指标,在他们的测试集上,在各种编解码器和比特率。然而,我们表明,这样的结果可能是欺骗性的,可能无法准确地代表系统处理手头任务的真实能力。在这篇文章中,我们提出了一个有损音频识别模型的鲁棒性和泛化能力的调查。我们的贡献如下:(1)我们表现出缺乏鲁棒性的编解码器参数变化的模型相当于现有技术。特别是,当天真地训练一个有损压缩检测模型的音乐录音处理的数据集与一系列的编解码器及其无损同行,我们获得近乎完美的性能指标上举行的测试集,但严重退化的性能与编解码器参数没有看到在训练中产生的有损轨道。(2)我们提出并显示了改进的训练策略的有效性,以显着提高模型的鲁棒性和泛化能力,超越训练过程中看到的编解码器配置。也就是说,我们对输入频谱图应用随机掩码,以鼓励模型不要仅仅依赖于训练集的编解码器截止频率。摘要:Previous research contributions on blind lossy compression identification report near perfect performance metrics on their test set, across a variety of codecs and bit rates. However, we show that such results can be deceptive and may not accurately represent true ability of the system to tackle the task at hand. In this article, we present an investigation into the robustness and generalisation capability of a lossy audio identification model. Our contributions are as follows. (1) We show the lack of robustness to codec parameter variations of a model equivalent to prior art. In particular, when naively training a lossy compression detection model on a dataset of music recordings processed with a range of codecs and their lossless counterparts, we obtain near perfect performance metrics on the held-out test set, but severely degraded performance on lossy tracks produced with codec parameters not seen in training. (2) We propose and show the effectiveness of an improved training strategy to significantly increase the robustness and generalisation capability of the model beyond codec configurations seen during training. Namely we apply a random mask to the input spectrogram to encourage the model not to rely solely on the training set's codec cutoff frequency.
【10】 Can LLMs "Reason" in Music? An Evaluation of LLMs' Capability of Music Understanding and Generation
作者:Ziya Zhou,Yuhang Wu,Zhiyue Wu,Xinyue Zhang,Ruibin Yuan,Yinghao Ma,Lu Wang,Emmanouil Benetos,Wei Xue,Yike Guo
备注:Accepted by ISMIR2024
链接:点击下载PDF文件
摘要:象征性的音乐,类似于语言,可以被编码成离散的符号。最近的研究已经将诸如GPT-4和Llama 2等大型语言模型(LLM)的应用扩展到符号音乐领域,包括理解和生成。然而,很少有研究探讨这些LLM如何在高级音乐理解和条件生成方面表现的细节,特别是从多步推理的角度来看,这是条件,可编辑和交互式人机共同创作过程中的一个关键方面。本研究对LLM在符号音乐加工中的能力和局限性进行了深入的调查。我们发现,目前的LLM在歌曲级多步音乐推理方面表现不佳,并且在处理复杂的音乐任务时通常无法利用所学的音乐知识。我们的研究结果表明,实现先进的音乐能力是不是固有的LLM获得的,未来的研究应该更多地集中在弥合音乐知识和推理之间的差距,以改善音乐家的共同创作体验。摘要:Symbolic Music, akin to language, can be encoded in discrete symbols. Recent research has extended the application of large language models (LLMs) such as GPT-4 and Llama2 to the symbolic music domain including understanding and generation. Yet scant research explores the details of how these LLMs perform on advanced music understanding and conditioned generation, especially from the multi-step reasoning perspective, which is a critical aspect in the conditioned, editable, and interactive human-computer co-creation process. This study conducts a thorough investigation of LLMs' capability and limitations in symbolic music processing. We identify that current LLMs exhibit poor performance in song-level multi-step music reasoning, and typically fail to leverage learned music knowledge when addressing complex musical tasks. An analysis of LLMs' responses highlights distinctly their pros and cons. Our findings suggest achieving advanced musical capability is not intrinsically obtained by LLMs, and future research should focus more on bridging the gap between music knowledge and reasoning, to improve the co-creation experience for musicians.
【11】 Generative Expressive Conversational Speech Synthesis
标题: 生成式表达式对话语音合成
作者:Rui Liu,Yifan Hu,Ren Yi,Yin Xiang,Haizhou Li
备注:14 pages, 6 figures, 8 tables. Accepted by ACM MM 2024
链接:点击下载PDF文件
摘要:会话语音合成(CSS)的目标是在用户-代理会话环境中以适当的说话风格表达目标话语。现有的CSS方法采用有效的多模态上下文建模技术来实现移情理解和表达。然而,他们通常需要设计复杂的网络架构,并精心优化其中的模块。此外,由于包含脚本记录风格的小规模数据集的限制,它们通常无法模拟真实的自然会话风格。为了解决上述问题,我们提出了一个新的生成式表达CSS系统,称为GPT-Talker,我们将多轮对话历史的多模态信息转换为离散的令牌序列,并将它们无缝集成,形成一个全面的用户代理对话上下文。利用GPT的力量,我们预测的令牌序列,其中包括语义和风格的知识,为代理的响应。在此基础上,我们提出了一个大规模的自然CSS数据集NCSSD,它既包含了自然录制的即兴风格的会话语音,也包含了从电视节目中提取的对话。我们对NCSSD的可靠性和GPT-Talker的有效性进行了全面的实验。主观和客观的评价表明,我们的模型优于其他国家的最先进的CSS系统显着的自然性和表现力。代码、数据集和预训练模型可在https: github.com AI-S2-Lab GPT-Talker上获得。摘要:Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy understanding and expression. However, they often need to design complex network architectures and meticulously optimize the modules within them. In addition, due to the limitations of small-scale datasets containing scripted recording styles, they often fail to simulate real natural conversational styles. To address the above issues, we propose a novel generative expressive CSS system, termed GPT-Talker.We transform the multimodal information of the multi-turn dialogue history into discrete token sequences and seamlessly integrate them to form a comprehensive user-agent dialogue context. Leveraging the power of GPT, we predict the token sequence, that includes both semantic and style knowledge, of response for the agent. After that, the expressive conversational speech is synthesized by the conversation-enriched VITS to deliver feedback to the user.Furthermore, we propose a large-scale Natural CSS Dataset called NCSSD, that includes both naturally recorded conversational speech in improvised styles and dialogues extracted from TV shows. It encompasses both Chinese and English languages, with a total duration of 236 hours.We conducted comprehensive experiments on the reliability of the NCSSD and the effectiveness of our GPT-Talker. Both subjective and objective evaluations demonstrate that our model outperforms other state-of-the-art CSS systems significantly in terms of naturalness and expressiveness. The Code, Dataset, and Pre-trained Model are available at: https: github.com AI-S2-Lab GPT-Talker.
【12】 On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
标题: 自动语音识别中合成数据生成的文本到语音模型选择问题
作者:Nick Rossenbach,Ralf Schlüter,Sakriani Sakti
备注:Accepted at the SynData4GenAI 2024 workshop
链接:点击下载PDF文件
摘要:神经文本到语音(TTS)系统的快速发展使其能够用于自然语言处理的其他领域,如自动语音识别(ASR)或口语翻译(ESTA)。由于大量不同的TTS体系结构及其扩展,选择哪种TTS系统用于合成数据创建并不是一件容易的事情。我们使用五种不同的TTS解码器架构的合成数据生成的范围内的比较,以显示基于CTC的语音识别训练的影响。我们将识别结果与NISQA MOS和可懂度等可计算指标进行比较,发现与ASR性能没有明确的关系。我们还观察到,对于数据生成自回归解码比非自回归解码表现更好,并提出了一种方法来量化TTS泛化能力。摘要:The rapid development of neural text-to-speech (TTS) systems enabled its usage in other areas of natural language processing such as automatic speech recognition (ASR) or spoken language translation (SLT). Due to the large number of different TTS architectures and their extensions, selecting which TTS systems to use for synthetic data creation is not an easy task. We use the comparison of five different TTS decoder architectures in the scope of synthetic data generation to show the impact on CTC-based speech recognition training. We compare the recognition results to computable metrics like NISQA MOS and intelligibility, finding that there are no clear relations to the ASR performance. We also observe that for data generation auto-regressive decoding performs better than non-autoregressive decoding, and propose an approach to quantify TTS generalization capabilities.
【13】 TinyChirp: Bird Song Recognition Using TinyML Models on Low-power Wireless Acoustic Sensors
标题: TinyChirp:在低功耗无线声学传感器上使用TinyML模型进行鸟鸣识别
作者:Zhaolan Huang,Adrien Tousnakhoff,Polina Kozyr,Roman Rehausen,Felix Bießmann,Robert Lachlan,Cedric Adjih,Emmanuel Baccelli
链接:点击下载PDF文件
摘要:大规模监测生物多样性具有挑战性。在细粒度分类中检测和识别物种需要高度准确的机器学习(ML)方法。训练这样的模型需要大量高质量的数据集。将这些模型部署到低功耗设备需要新的压缩技术和模型架构。虽然物种分类方法受益于新的数据集和ML方法的进步,特别是神经网络,但将这些最先进的模型部署到低功耗设备仍然很困难。在这里,我们对用于物种分类的各种tinyML神经网络架构和压缩技术进行了全面的实证比较。我们专注于鸟鸣检测的例子,更具体地说,一个数据集策划研究玉米bunting鸟类。数据集与本研究的所有代码和实验一起发布。在我们的实验中,我们比较预测性能,内存和时间复杂度的经典的基于频谱图的方法和最近的方法操作的原始音频信号。我们的研究结果表明,个别鸟类物种可以稳健地检测到相对简单的架构,可以很容易地部署到低功耗设备。摘要:Monitoring biodiversity at scale is challenging. Detecting and identifying species in fine grained taxonomies requires highly accurate machine learning (ML) methods. Training such models requires large high quality data sets. And deploying these models to low power devices requires novel compression techniques and model architectures. While species classification methods have profited from novel data sets and advances in ML methods, in particular neural networks, deploying these state of the art models to low power devices remains difficult. Here we present a comprehensive empirical comparison of various tinyML neural network architectures and compression techniques for species classification. We focus on the example of bird song detection, more concretely a data set curated for studying the corn bunting bird species. The data set is released along with all code and experiments of this study. In our experiments we compare predictive performance, memory and time complexity of classical spectrogram based methods and recent approaches operating on raw audio signal. Our results indicate that individual bird species can be robustly detected with relatively simple architectures that can be readily deployed to low power devices.
【14】 Design and Development of Laughter Recognition System Based on Multimodal Fusion and Deep Learning
标题: 基于多模式融合和深度学习的笑声识别系统的设计与开发
作者:Fuzheng Zhao,Yu Bai
备注:7 pages,2 figures
链接:点击下载PDF文件
摘要:本研究旨在设计并实现一个基于多模态融合和深度学习的笑声识别系统,利用图像和音频处理技术实现准确的笑声识别和情感分析。首先,系统加载视频文件并使用OpenCV库提取面部信息,同时使用Librosa库处理音频功能,如MFCC。然后,使用多模态融合技术来整合图像和音频特征,然后使用深度学习模型进行训练和预测。评估结果表明,该模型在测试数据集上实现了80%的准确度、精确度和召回率,F1得分为80%,表现出强大的性能和处理真实世界数据变化的能力。该研究不仅验证了多模态融合方法在笑声识别中的有效性,而且突出了其在情感计算和人机交互中的潜在应用。未来的工作将集中在进一步优化特征提取和模型架构,以提高识别准确率和扩展应用场景,促进笑声识别技术在心理健康监测和教育活动评估等领域的发展摘要:This study aims to design and implement a laughter recognition system based on multimodal fusion and deep learning, leveraging image and audio processing technologies to achieve accurate laughter recognition and emotion analysis. First, the system loads video files and uses the OpenCV library to extract facial information while employing the Librosa library to process audio features such as MFCC. Then, multimodal fusion techniques are used to integrate image and audio features, followed by training and prediction using deep learning models. Evaluation results indicate that the model achieved 80% accuracy, precision, and recall on the test dataset, with an F1 score of 80%, demonstrating robust performance and the ability to handle real-world data variability. This study not only verifies the effectiveness of multimodal fusion methods in laughter recognition but also highlights their potential applications in affective computing and human-computer interaction. Future work will focus on further optimizing feature extraction and model architecture to improve recognition accuracy and expand application scenarios, promoting the development of laughter recognition technology in fields such as mental health monitoring and educational activity evaluation
【15】 AI Safety in Practice: Enhancing Adversarial Robustness in Multimodal Image Captioning
标题: 实践中的人工智能安全:增强多模式图像字幕中的对抗鲁棒性
作者:Maisha Binte Rashid,Pablo Rivas
备注:Accepted into KDD 2024 workshop on Ethical AI
链接:点击下载PDF文件
摘要:结合视觉和文本数据的多模态机器学习模型越来越多地被部署在关键应用程序中,由于它们容易受到对抗性攻击,因此引起了重大的安全和安全问题。本文提出了一种有效的策略,以提高对这种攻击的多模态图像字幕模型的鲁棒性。通过利用快速梯度符号方法(FGSM)生成对抗性示例并结合对抗性训练技术,我们在两个基准数据集上展示了改进的模型鲁棒性:Flickr8k和COCO。我们的研究结果表明,选择性地只训练多模态架构的文本解码器显示出与完全对抗训练相当的性能,同时提供更高的计算效率。这种有针对性的方法建议在鲁棒性和培训成本之间取得平衡,促进多模式人工智能系统在各个领域的道德部署。摘要:Multimodal machine learning models that combine visual and textual data are increasingly being deployed in critical applications, raising significant safety and security concerns due to their vulnerability to adversarial attacks. This paper presents an effective strategy to enhance the robustness of multimodal image captioning models against such attacks. By leveraging the Fast Gradient Sign Method (FGSM) to generate adversarial examples and incorporating adversarial training techniques, we demonstrate improved model robustness on two benchmark datasets: Flickr8k and COCO. Our findings indicate that selectively training only the text decoder of the multimodal architecture shows performance comparable to full adversarial training while offering increased computational efficiency. This targeted approach suggests a balance between robustness and training costs, facilitating the ethical deployment of multimodal AI systems across various domains.
【16】 Computational music analysis from first principles
标题: 从首要原则出发的计算音乐分析
作者:Dmitri Tymoczko,Mark Newman
链接:点击下载PDF文件
摘要:我们使用耦合隐马尔可夫模型自动注释的371巴赫合唱团在Riemenschneider版,语料库包含约100,000个音符和20,000和弦。我们给出了三个独立的分析,实现逐步更大的准确性,在对音乐句法越来越强的假设的成本。虽然我们的方法几乎不使用人工输入,但与专家人工分析相比,我们能够以85%或更高的准确度识别和弦和键,从而使注释足够准确,可用于一系列音乐理论目的,同时也不受主观人类判断的影响。我们的工作承担长期的争论的客观现实的结构所假设的标准西方和声理论,以及具体问题的性质西方和声句法。摘要:We use coupled hidden Markov models to automatically annotate the 371 Bach chorales in the Riemenschneider edition, a corpus containing approximately 100,000 notes and 20,000 chords. We give three separate analyses that achieve progressively greater accuracy at the cost of making increasingly strong assumptions about musical syntax. Although our method makes almost no use of human input, we are able to identify both chords and keys with an accuracy of 85% or greater when compared to an expert human analysis, resulting in annotations accurate enough to be used for a range of music-theoretical purposes, while also being free of subjective human judgments. Our work bears on longstanding debates about the objective reality of the structures postulated by standard Western harmonic theory, as well as on specific questions about the nature of Western harmonic syntax.
【17】 ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
标题: ELP适配器:针对各种语音处理任务的参数高效适配器调整
作者:Nakamasa Inoue,Shinta Otake,Takumi Hirose,Masanari Ohi,Rei Kawakami
链接:点击下载PDF文件
摘要:自监督学习已经成为从语音数据中学习通用表示的关键方法。尽管在语音识别、说话人验证和情感识别等下游任务中取得了令人鼓舞的结果,但需要大量的参数,这使得对每个任务的微调记忆效率低下。为了解决这一限制,我们引入ELP适配器调整,一种新的方法,使用三种类型的适配器,即编码器适配器(E-适配器),层适配器(L-适配器),和提示适配器(P-适配器)的参数有效的微调。E-adapter集成到基于transformer的编码器层中,有助于学习对语音识别有效的细粒度语音表示。L适配器创建从每个编码器层到下游头部的路径,并帮助从较低的编码器层提取非语言特征,这些特征对说话人验证和情感识别有效。P-adapter将伪特征附加到CNN特征,以进一步提高有效性和效率。有了这些适配器,模型可以快速适应各种语音处理任务。我们的评估在四个下游任务使用五个骨干模型证明了所提出的方法的有效性。使用WavLM主干,其性能与所有任务的完全微调相当或更好,同时需要减少90%的可学习参数。摘要:Self-supervised learning has emerged as a key approach for learning generic representations from speech data. Despite promising results in downstream tasks such as speech recognition, speaker verification, and emotion recognition, a significant number of parameters is required, which makes fine-tuning for each task memory-inefficient. To address this limitation, we introduce ELP-adapter tuning, a novel method for parameter-efficient fine-tuning using three types of adapter, namely encoder adapters (E-adapters), layer adapters (L-adapters), and a prompt adapter (P-adapter). The E-adapters are integrated into transformer-based encoder layers and help to learn fine-grained speech representations that are effective for speech recognition. The L-adapters create paths from each encoder layer to the downstream head and help to extract non-linguistic features from lower encoder layers that are effective for speaker verification and emotion recognition. The P-adapter appends pseudo features to CNN features to further improve effectiveness and efficiency. With these adapters, models can be quickly adapted to various speech processing tasks. Our evaluation across four downstream tasks using five backbone models demonstrated the effectiveness of the proposed method. With the WavLM backbone, its performance was comparable to or better than that of full fine-tuning on all tasks while requiring 90% fewer learnable parameters.
【18】 Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
标题: 使用CycleGAN和域间损失改善端到端ASB中低资源语言的嘈杂学生训练
作者:Chia-Yu Li,Ngoc Thang Vu
备注:10 pages (2 for references), 4 figures, published in SIGUL2024@LREC-COLING 2024
链接:点击下载PDF文件
摘要:使用噪声学生训练来训练半监督端到端语音识别系统显著提高了性能。然而,这种方法需要大量成对的语音文本和未标记的语音,这对于低资源语言来说是昂贵的。因此,本文考虑了半监督端到端自动语音识别的一个更极端的情况,其中有有限的配对语音文本,未标记的语音(少于五个小时),和丰富的外部文本。首先,我们通过使用我们以前在半监督学习“CycleGAN和域间损失”方面的工作单独使用外部文本来训练模型,从而观察到性能的提高。其次,我们通过结合自动超参数调整来增强“CycleGAN和域间损失”,称之为“增强的CycleGAN域间损失”。“第三,我们将其整合到低资源场景的嘈杂学生培训方法管道中。我们在Voxforge和Common Voice的六种非英语语言上进行的实验结果显示,与基线教师模型相比,单词错误率降低了20%,与基线最佳学生模型相比,单词错误率降低了10%,突出了通过我们提出的方法取得的显着改进。摘要:Training a semi-supervised end-to-end speech recognition system using noisy student training has significantly improved performance. However, this approach requires a substantial amount of paired speech-text and unlabeled speech, which is costly for low-resource languages. Therefore, this paper considers a more extreme case of semi-supervised end-to-end automatic speech recognition where there are limited paired speech-text, unlabeled speech (less than five hours), and abundant external text. Firstly, we observe improved performance by training the model using our previous work on semi-supervised learning "CycleGAN and inter-domain losses" solely with external text. Secondly, we enhance "CycleGAN and inter-domain losses" by incorporating automatic hyperparameter tuning, calling it "enhanced CycleGAN inter-domain losses." Thirdly, we integrate it into the noisy student training approach pipeline for low-resource scenarios. Our experimental results, conducted on six non-English languages from Voxforge and Common Voice, show a 20% word error rate reduction compared to the baseline teacher model and a 10% word error rate reduction compared to the baseline best student model, highlighting the significant improvements achieved through our proposed method.
【19】 Sentiment Reasoning for Healthcare
标题: 医疗保健的情感推理
作者:Khai Le-Duc,Khai-Nguyen Nguyen,Bach Phan Tat,Duy Le,Jerry Ngo,Long Vo-Dang,Anh Totti Nguyen,Truong-Son Hy
备注:Preprint, 18 pages
链接:点击下载PDF文件
摘要:由于错误的严重后果,人工智能决策的透明度在医疗保健中至关重要,这对于在情感分析任务中建立人工智能和用户之间的信任非常重要。增强推理能力有助于大型语言模型(LLM)在更广泛的背景下理解人类情感,处理细微差别和模糊的语言,并推断可能没有明确说明的潜在情感。在这项工作中,我们引入了一个新的任务-情感推理-语音和文本模态,以及我们提出的多模态多任务框架和数据集。我们的研究表明,理性增强训练增强了模型在人类转录本和ASR设置中的情感分类性能。此外,我们发现,生成的理由通常表现出不同的词汇相比,人类生成的理由,但保持类似的语义。所有代码、数据(英语翻译和越南语)和模型均在线发布:https: github.com leduckhai MultiMed摘要:Transparency in AI decision-making is crucial in healthcare due to the severe consequences of errors, and this is important for building trust among AI and users in sentiment analysis task. Incorporating reasoning capabilities helps Large Language Models (LLMs) understand human emotions within broader contexts, handle nuanced and ambiguous language, and infer underlying sentiments that may not be explicitly stated. In this work, we introduce a new task - Sentiment Reasoning - for both speech and text modalities, along with our proposed multimodal multitask framework and dataset. Our study showed that rationale-augmented training enhances model performance in sentiment classification across both human transcript and ASR settings. Also, we found that the generated rationales typically exhibit different vocabularies compared to human-generated rationales, but maintain similar semantics. All code, data (English-translated and Vietnamese) and models are published online: https: github.com leduckhai MultiMed
机器翻译,仅供参考
