今日论文合集:cs.SD语音8篇,eess.AS音频处理10篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 UniGlyph: A Seven-Segment Script for Universal Language Representation
标题: UniGlassis:通用语言表示的七段脚本
作者: G. V. Bency Sherin, A. Abijesh Euphrine, A. Lenora Moreen, L. Arun Jose
备注:This submission includes 23 pages and tables. No external funding has been received for this research. Acknowledgments to Jeseentha V. for contributions to the phonetic study
链接:点击下载PDF文件
摘要:UniGlands是一种构造语言(conlang),旨在使用从七段字符派生的脚本创建通用音译系统。UniGlancial的目标是通过提供一个灵活和一致的脚本来促进跨语言交流,该脚本可以表示各种语音。本文探讨了UniGlitter的设计,详细介绍了它的脚本结构,语音映射和音译规则。该系统解决了国际音标(IPA)和传统字符集的缺陷,提供了一个紧凑的,通用的方法来表示语音的多样性,跨语言。通过音高和长度标记,UniGlnut可确保准确的语音表示,同时保持较小的字符集。UniGlance的应用包括人工智能集成,如自然语言处理和多语言语音识别,增强不同语言之间的沟通。未来的扩展进行了讨论,包括添加动物语音的声音,其中独特的脚本被分配给不同的物种,扩大了UniGlands的范围超越人类的沟通。这项研究提出了开发这种通用脚本的挑战和解决方案,展示了UniGlanca在跨语言交流,教育语音学和AI驱动的应用程序中弥合语言差距的潜力。摘要:UniGlyph is a constructed language (conlang) designed to create a universal transliteration system using a script derived from seven-segment characters. The goal of UniGlyph is to facilitate cross-language communication by offering a flexible and consistent script that can represent a wide range of phonetic sounds. This paper explores the design of UniGlyph, detailing its script structure, phonetic mapping, and transliteration rules. The system addresses imperfections in the International Phonetic Alphabet (IPA) and traditional character sets by providing a compact, versatile method to represent phonetic diversity across languages. With pitch and length markers, UniGlyph ensures accurate phonetic representation while maintaining a small character set. Applications of UniGlyph include artificial intelligence integrations, such as natural language processing and multilingual speech recognition, enhancing communication across different languages. Future expansions are discussed, including the addition of animal phonetic sounds, where unique scripts are assigned to different species, broadening the scope of UniGlyph beyond human communication. This study presents the challenges and solutions in developing such a universal script, demonstrating the potential of UniGlyph to bridge linguistic gaps in cross-language communication, educational phonetics, and AI-driven applications.

【2】 Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
标题: 增强印度尼西亚自动语音识别:评估具有多样化语音变异性的多语言模型
作者: Aulia Adila, Dessi Lestari, Ayu Purwarianti, Dipta Tanaya, Kurniawati Azizah, Sakriani Sakti
链接:点击下载PDF文件
摘要:理想的语音识别模型具有在语音信号的各种特征下准确地转录语音的能力,所述语音信号的各种特征诸如说话风格(朗读和自发)、语音上下文(正式和非正式)以及背景噪声条件(干净和适度)。构建这样的模型需要大量具有不同语音特征的训练数据。目前,印尼语数据主要是阅读,正式和干净的语音,导致印尼语数据与其他语音变量的稀缺。为了开发印度尼西亚语自动语音识别(ASR),我们提出了我们的研究国家的最先进的语音识别模型,即大规模多语言语音(MMS)和耳语,以及编制一个数据集,包括印尼语语音的变异,以方便我们的研究。我们进一步研究了模型的预测能力,转录印尼语语音数据在不同的变异组。Whisper微调模型在具有各种特征的数据集上取得了最佳结果,如单词错误率(WER)和字符错误率(CER)的下降所示。此外,我们发现,说话风格的变异性影响模型的性能。摘要:An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise conditions (clean and moderate). Building such a model requires a significant amount of training data with diverse speech characteristics. Currently, Indonesian data is dominated by read, formal, and clean speech, leading to a scarcity of Indonesian data with other speech variabilities. To develop Indonesian automatic speech recognition (ASR), we present our research on state-of-the-art speech recognition models, namely Massively Multilingual Speech (MMS) and Whisper, as well as compiling a dataset comprising Indonesian speech with variabilities to facilitate our study. We further investigate the models' predictive ability to transcribe Indonesian speech data across different variability groups. The best results were achieved by the Whisper fine-tuned model across datasets with various characteristics, as indicated by the decrease in word error rate (WER) and character error rate (CER). Moreover, we found that speaking style variability affected model performance the most.

【3】 Small Tunes Transformer: Exploring Macro & Micro-Level Hierarchies for Skeleton-Conditioned Melody Generation
标题: 小型收件箱Transformer:探索普林斯顿条件旋律生成的宏观和微观层次结构
作者: Yishan Lv, Jing Luo, Boyuan Ju, Xinyu Yang
链接:点击下载PDF文件
摘要:近年来,符号音乐生成已经成为众多深度学习研究的焦点。结构作为音乐的重要组成部分,有助于提高音乐的质量,越来越多的作品开始研究层次结构。本研究从宏观层次和微观层次两个方面对音乐的多层次结构进行了探讨。在宏观层次上,通过乐句切分算法探索乐句对音乐整体发展的影响;在微观层次上,通过骨架音符提取策略探索乐句中的骨架音符对旋律生成的指导作用。此外,我们提出了一种新的短语级交叉注意机制,以捕捉宏观层次和微观层次之间的内在关系。此外,针对目前缺乏对中国风格音乐的研究,我们构建了我们的小样本数据集:一个由10088个小样本文件组成的大样本数据集,这是一个传统的中国民歌类别。该数据集是我们研究的重点。我们利用提取的骨架音符作为条件来生成小旋律歌曲,实验结果表明,我们提出的模型,小旋律Transformer,优于其他国家的最先进的模型。此外,我们设计了三个新的客观评价指标,从节奏和旋律两个维度来评价音乐。摘要:Recently, symbolic music generation has become a focus of numerous deep learning research. Structure as an important part of music, contributes to improving the quality of music, and an increasing number of works start to study the hierarchical structure. In this study, we delve into the multi-level structures within music from macro-level and micro-level hierarchies. At the macro-level hierarchy, we conduct phrase segmentation algorithm to explore how phrases influence the overall development of music, and at the micro-level hierarchy, we design skeleton notes extraction strategy to explore how skeleton notes within each phrase guide the melody generation. Furthermore, we propose a novel Phrase-level Cross-Attention mechanism to capture the intrinsic relationship between macro-level hierarchy and micro-level hierarchy. Moreover, in response to the current lack of research on Chinese-style music, we construct our Small Tunes Dataset: a substantial collection of MIDI files comprising 10088 Small Tunes, a category of traditional Chinese Folk Songs. This dataset serves as the focus of our study. We generate Small Tunes songs utilizing the extracted skeleton notes as conditions, and experiment results indicate that our proposed model, Small Tunes Transformer, outperforms other state-of-the-art models. Besides, we design three novel objective evaluation metrics to evaluate music from both rhythm and melody dimensions.

【4】 Symbolic Music Generation with Fine-grained Interactive Textural Guidance
标题: 具有细粒度交互文本指导的象征性音乐生成
作者: Tingyu Zhu, Haoyu Liu, Zhimin Jiang, Zeyu Zheng
链接:点击下载PDF文件
摘要:由于有限的数据可用性和对音符音高的高精度要求的结合,符号音乐生成的问题提出了独特的挑战。为了克服这些困难,我们在扩散模型中引入细粒度纹理指导(FTG)来纠正学习分布中的错误。通过结合FTG,扩散模型提高了音乐生成的准确性,使其非常适合渐进式音乐生成、即兴创作和交互式音乐创作等高级任务。我们得到的理论表征的挑战,在象征性的音乐生成和FTG方法的效果。我们提供了数值实验和一个演示页面,用于交互式音乐生成与用户输入,以展示我们的方法的有效性。摘要:The problem of symbolic music generation presents unique challenges due to the combination of limited data availability and the need for high precision in note pitch. To overcome these difficulties, we introduce Fine-grained Textural Guidance (FTG) within diffusion models to correct errors in the learned distributions. By incorporating FTG, the diffusion models improve the accuracy of music generation, which makes them well-suited for advanced tasks such as progressive music generation, improvisation and interactive music creation. We derive theoretical characterizations for both the challenges in symbolic music generation and the effect of the FTG approach. We provide numerical experiments and a demo page for interactive music generation with user input to showcase the effectiveness of our approach.

【5】 The language of sound search: Examining User Queries in Audio Search Engines
标题: 声音搜索语言:检查音频搜索引擎中的用户收件箱
作者: Benno Weck, Frederic Font
备注:Accepted at DCASE 2024. Supplementary materials at this https URL
链接:点击下载PDF文件
摘要:本研究探讨文字,用户编写的搜索查询的背景下,声音搜索引擎,包括各种应用程序,如福利,音效,和一般的音频检索。当前的研究在设计基于文本的音频检索系统时没有充分解决现实世界的用户需求和行为。为了弥补这一差距,我们分析了两个来源的搜索查询:自定义调查和Freesound网站查询日志。该调查旨在收集一个不受限制的,假设的声音搜索引擎的查询,从而产生一个数据集,捕捉用户的意图,而不受现有系统的限制。该数据集也可与研究界共享。相比之下,Freesound查询日志包含大约900万个搜索请求,提供了真实世界使用模式的全面视图。我们的研究结果表明,调查查询一般比Freesound查询,这表明用户更喜欢详细的查询时,不受系统限制。这两个数据集主要以基于关键词的查询为特征,很少有调查参与者使用完整的句子。影响调查查询的关键因素包括主要声源、预期用途、感知位置和声源数量。这些见解对于开发以用户为中心,有效的基于文本的音频检索系统至关重要,增强了我们对用户在声音搜索环境中行为的理解。摘要:This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequately addresses real-world user needs and behaviours in designing text-based audio retrieval systems. To bridge this gap, we analysed search queries from two sources: a custom survey and Freesound website query logs. The survey was designed to collect queries for an unrestricted, hypothetical sound search engine, resulting in a dataset that captures user intentions without the constraints of existing systems. This dataset is also made available for sharing with the research community. In contrast, the Freesound query logs encompass approximately 9 million search requests, providing a comprehensive view of real-world usage patterns. Our findings indicate that survey queries are generally longer than Freesound queries, suggesting users prefer detailed queries when not limited by system constraints. Both datasets predominantly feature keyword-based queries, with few survey participants using full sentences. Key factors influencing survey queries include the primary sound source, intended usage, perceived location, and the number of sound sources. These insights are crucial for developing user-centred, effective text-based audio retrieval systems, enhancing our understanding of user behaviour in sound search contexts.

【6】 Music Genre Classification using Large Language Models
标题: 使用大型语言模型的音乐流派分类
作者: Mohamed El Amine Meguenani, Alceu de Souza Britto Jr., Alessandro Lameiras Koerich
备注:7 pages
链接:点击下载PDF文件
摘要:本文利用预训练的大型语言模型(LLM)的zero-shot能力进行音乐流派分类。所提出的方法将音频信号分成20 ms的块,并通过卷积特征编码器、Transformer编码器和用于编码音频单元和生成特征向量的附加层来处理它们。提取的特征向量用于训练分类头。在推断期间,对各个块的预测被聚合以用于最终的流派分类。我们对LLM(包括WavLM、HuBERT和wav2vec 2.0)与传统的深度学习架构(如1D和2D卷积神经网络(CNN)和音频频谱图Transformer(AST))进行了全面的比较。我们的研究结果表明,AST模型的性能优越,总体准确率达到85.5%,超过了所有其他评估的模型。这些结果突出了LLM和基于变压器的架构在推进音乐信息检索任务方面的潜力,即使在zero-shot场景中也是如此。摘要:This paper exploits the zero-shot capabilities of pre-trained large language models (LLMs) for music genre classification. The proposed approach splits audio signals into 20 ms chunks and processes them through convolutional feature encoders, a transformer encoder, and additional layers for coding audio units and generating feature vectors. The extracted feature vectors are used to train a classification head. During inference, predictions on individual chunks are aggregated for a final genre classification. We conducted a comprehensive comparison of LLMs, including WavLM, HuBERT, and wav2vec 2.0, with traditional deep learning architectures like 1D and 2D convolutional neural networks (CNNs) and the audio spectrogram transformer (AST). Our findings demonstrate the superior performance of the AST model, achieving an overall accuracy of 85.5%, surpassing all other models evaluated. These results highlight the potential of LLMs and transformer-based architectures for advancing music information retrieval tasks, even in zero-shot scenarios.

【7】 A Recurrent Neural Network Approach to the Answering Machine Detection Problem
标题: 解决志愿服务机检测问题的回归神经网络方法
作者: Kemal Altwlkany, Sead Delalic, Elmedin Selmanovic, Adis Alihodzic, Ivica Lovric
备注:6 pages, 2 figures, 2024 47th MIPRO ICT and Electronics Convention (MIPRO)
链接:点击下载PDF文件
摘要:在电信和云通信领域,准确且实时地检测是人还是应答机应答了呼出呼叫是至关重要的。这个问题在竞选期间特别重要,因为它通过精确的呼叫者识别来提高服务质量、效率和降低成本。尽管该领域的重要性,它仍然没有充分探讨现有的文献。本文提出了一种创新的应答机检测方法,该方法通过YAMNet模型利用迁移学习进行特征提取。YAMNet架构有助于训练基于递归的分类器,从而实现音频流的实时处理,而不是固定长度的记录。结果表明,在测试集上的准确率超过96%。此外,我们进行了深入的分析误分类的样本,并揭示了超过98%的准确度可以实现与集成的沉默检测算法,如FFmpeg提供的一个。摘要:In the field of telecommunications and cloud communications, accurately and in real-time detecting whether a human or an answering machine has answered an outbound call is of paramount importance. This problem is of particular significance during campaigns as it enhances service quality, efficiency and cost reduction through precise caller identification. Despite the significance of the field, it remains inadequately explored in the existing literature. This paper presents an innovative approach to answering machine detection that leverages transfer learning through the YAMNet model for feature extraction. The YAMNet architecture facilitates the training of a recurrent-based classifier, enabling real-time processing of audio streams, as opposed to fixed-length recordings. The results demonstrate an accuracy of over 96% on the test set. Furthermore, we conduct an in-depth analysis of misclassified samples and reveal that an accuracy exceeding 98% can be achieved with the integration of a silence detection algorithm, such as the one provided by FFmpeg.

【8】 Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
标题: 探索基于SVR的Wave 2Vec 2用于自动言语障碍评估:见解和分析
作者: Tuan Nguyen, Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard
备注:Accepted at the Spoken Language Technology (SLT) Conference 2024
链接:点击下载PDF文件
摘要:随着SSL和ASR技术的兴起,基于Wav2Vec2 ASR的模型已经针对自动语音障碍质量评估任务进行了微调,产生了令人印象深刻的结果,并为头颈癌语音环境设定了新的基线。这表明来自Wav2Vec2的ASR维度与评估维度紧密一致。尽管其有效性,该系统仍然是一个黑匣子,没有明确的解释模型ASR维度和临床评估之间的联系。本文提出了第一次分析的语音质量评估的基线模型,专注于可懂度和严重性的任务。我们进行了逐层分析,以识别关键层,并根据预先训练的数据比较不同的SSL和ASR Wav2Vec2模型。此外,事后XAI方法,包括典型相关分析(CCA)和可视化技术,用于跟踪模型演变和可视化嵌入,以增强可解释性。摘要:With the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pre-trained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability.


eess.AS音频处理
【1】 Low-complexity Attention-based Unsupervised Anomalous Sound Detection exploiting Separable Convolutions and Angular Loss
标题: 利用可分离卷积和角度损失的低复杂度基于注意力的无监督异常声音检测
作者: Michael Neri, Marco Carli
备注:Accepted for publication in IEEE Sensors Letters. 4 pages, 4 figures
链接:点击下载PDF文件
摘要:在这项工作中,提出了一种新型的深度神经网络,旨在提高无监督声音异常检测的效率和有效性。该模型利用注意力模块和可分离卷积来识别音频数据中的显著时频模式,以降低计算复杂度来区分正常和异常声音。该方法通过使用DCASE 2020挑战的任务2数据集进行的大量实验进行了验证。结果表明,在异常检测精度方面具有优异的性能,同时具有比最先进的方法更少的参数。实现细节、代码和预训练模型可在https: github.com michaelneri unsupervised-audio-anomaly-detection中获得。摘要:In this work, a novel deep neural network, designed to enhance the efficiency and effectiveness of unsupervised sound anomaly detection, is presented. The proposed model exploits an attention module and separable convolutions to identify salient time-frequency patterns in audio data to discriminate between normal and anomalous sounds with reduced computational complexity. The approach is validated through extensive experiments using the Task 2 dataset of the DCASE 2020 challenge. Results demonstrate superior performance in terms of anomaly detection accuracy while having fewer parameters than state-of-the-art methods. Implementation details, code, and pre-trained models are available in https: github.com michaelneri unsupervised-audio-anomaly-detection.

【2】 Low Bitrate High-Quality RVQGAN-based Discrete Speech Tokenizer
标题: 低比特率高质量基于RVQGAN的离散语音令牌器
作者: Slava Shechtman, Avihu Dekel
备注:You can download the model from https:huggingface.coibmDAC.speech.v1.0
Journal-ref:Proc. Interspeech 2024, 4174-4178
链接:点击下载PDF文件
摘要:离散音频编解码器(或音频标记器)最近重新受到关注,因为大型语言模型(LLM)能够学习其压缩的声学表示。最近,各种公开可用的可训练离散标记器在音频标记化方面表现出了令人印象深刻的结果,但它们大多需要高标记速率来获得高质量的重建。在这项研究中,我们使用不同的开源语音数据,考虑到各种录音条件和质量水平,微调了开源通用音频RVQGAN模型。由此产生的宽带(24 kHz)语音模型实现语音重建,这几乎是无法区分的PCM(脉冲编码调制)与150-300令牌每秒(1500-3000 bps)的速率。评估使用了全面的英语语音数据,包括不同的录音条件,包括录音室设置。语音样本可在http: ibm.biz IS24SpeechRVQ上公开获取。该模型在https: huggingface.co ibm DAC.speech.v1.0上正式发布摘要:Discrete Audio codecs (or audio tokenizers) have recently regained interest due to the ability of Large Language Models (LLMs) to learn their compressed acoustic representations. Various publicly available trainable discrete tokenizers recently demonstrated impressive results for audio tokenization, yet they mostly require high token rates to gain high-quality reconstruction. In this study, we fine-tuned an open-source general audio RVQGAN model using diverse open-source speech data, considering various recording conditions and quality levels. The resulting wideband (24kHz) speech-only model achieves speech reconstruction, which is nearly indistinguishable from PCM (pulse-code modulation) with a rate of 150-300 tokens per second (1500-3000 bps). The evaluation used comprehensive English speech data encompassing different recording conditions, including studio settings. Speech samples are made publicly available in http: ibm.biz IS24SpeechRVQ . The model is officially released in https: huggingface.co ibm DAC.speech.v1.0

【3】 Exploring ASR-Based Wav2Vec2 for Automated Speech Disorder Assessment: Insights and Analysis
标题: 探索基于SVR的Wave 2Vec 2用于自动言语障碍评估:见解和分析
作者: Tuan Nguyen, Corinne Fredouille, Alain Ghio, Mathieu Balaguer, Virginie Woisard
备注:Accepted at the Spoken Language Technology (SLT) Conference 2024
链接:点击下载PDF文件
摘要:随着SSL和ASR技术的兴起,基于Wav2 Vec 2 ASR的模型已针对自动语音障碍质量评估任务进行了微调,产生了令人印象深刻的结果,并为头颈癌语音背景设定了新的基线。这表明来自Wav2Vec2的ASR维度与评估维度紧密一致。尽管该系统很有效,但它仍然是一个黑匣子,无法明确解释模型ASR维度与临床评估之间的联系。本文提出了第一次分析的语音质量评估的基线模型,专注于可懂度和严重性的任务。我们进行了逐层分析,以识别关键层,并根据预先训练的数据比较不同的SSL和ASR Wav2Vec2模型。此外,事后XAI方法,包括典型相关分析(CCA)和可视化技术,用于跟踪模型演变和可视化嵌入,以增强可解释性。摘要:With the rise of SSL and ASR technologies, the Wav2Vec2 ASR-based model has been fine-tuned for automated speech disorder quality assessment tasks, yielding impressive results and setting a new baseline for Head and Neck Cancer speech contexts. This demonstrates that the ASR dimension from Wav2Vec2 closely aligns with assessment dimensions. Despite its effectiveness, this system remains a black box with no clear interpretation of the connection between the model ASR dimension and clinical assessments. This paper presents the first analysis of this baseline model for speech quality assessment, focusing on intelligibility and severity tasks. We conduct a layer-wise analysis to identify key layers and compare different SSL and ASR Wav2Vec2 models based on pre-trained data. Additionally, post-hoc XAI methods, including Canonical Correlation Analysis (CCA) and visualization techniques, are used to track model evolution and visualize embeddings for enhanced interpretability.

【4】 UniGlyph: A Seven-Segment Script for Universal Language Representation
标题: UniGlassis:通用语言表示的七段脚本
作者: G. V. Bency Sherin, A. Abijesh Euphrine, A. Lenora Moreen, L. Arun Jose
备注:This submission includes 23 pages and tables. No external funding has been received for this research. Acknowledgments to Jeseentha V. for contributions to the phonetic study
链接:点击下载PDF文件
摘要:UniGlands是一种构造语言(conlang),旨在使用从七段字符派生的脚本创建通用音译系统。UniGlancial的目标是通过提供一个灵活和一致的脚本来促进跨语言交流,该脚本可以表示各种语音。本文探讨了UniGlitter的设计,详细介绍了它的脚本结构,语音映射和音译规则。该系统解决了国际音标(IPA)和传统字符集的缺陷,提供了一个紧凑的,通用的方法来表示语音的多样性,跨语言。通过音高和长度标记,UniGlnut可确保准确的语音表示,同时保持较小的字符集。UniGlance的应用包括人工智能集成,如自然语言处理和多语言语音识别,增强不同语言之间的沟通。未来的扩展进行了讨论,包括添加动物语音的声音,其中独特的脚本被分配给不同的物种,扩大了UniGlands的范围超越人类的沟通。这项研究提出了开发这种通用脚本的挑战和解决方案,展示了UniGlanca在跨语言交流,教育语音学和AI驱动的应用程序中弥合语言差距的潜力。摘要:UniGlyph is a constructed language (conlang) designed to create a universal transliteration system using a script derived from seven-segment characters. The goal of UniGlyph is to facilitate cross-language communication by offering a flexible and consistent script that can represent a wide range of phonetic sounds. This paper explores the design of UniGlyph, detailing its script structure, phonetic mapping, and transliteration rules. The system addresses imperfections in the International Phonetic Alphabet (IPA) and traditional character sets by providing a compact, versatile method to represent phonetic diversity across languages. With pitch and length markers, UniGlyph ensures accurate phonetic representation while maintaining a small character set. Applications of UniGlyph include artificial intelligence integrations, such as natural language processing and multilingual speech recognition, enhancing communication across different languages. Future expansions are discussed, including the addition of animal phonetic sounds, where unique scripts are assigned to different species, broadening the scope of UniGlyph beyond human communication. This study presents the challenges and solutions in developing such a universal script, demonstrating the potential of UniGlyph to bridge linguistic gaps in cross-language communication, educational phonetics, and AI-driven applications.

【5】 Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
标题: 增强印度尼西亚自动语音识别:评估具有多样化语音变异性的多语言模型
作者: Aulia Adila, Dessi Lestari, Ayu Purwarianti, Dipta Tanaya, Kurniawati Azizah, Sakriani Sakti
链接:点击下载PDF文件
摘要:理想的语音识别模型具有在语音信号的各种特征下准确地转录语音的能力,所述语音信号的各种特征诸如说话风格(朗读和自发)、语音上下文(正式和非正式)以及背景噪声条件(干净和适度)。构建这样的模型需要大量具有不同语音特征的训练数据。目前,印尼语数据主要是阅读,正式和干净的语音,导致印尼语数据与其他语音变量的稀缺。为了开发印度尼西亚语自动语音识别(ASR),我们提出了我们的研究国家的最先进的语音识别模型,即大规模多语言语音(MMS)和耳语,以及编制一个数据集,包括印尼语语音的变异,以方便我们的研究。我们进一步研究了模型的预测能力,转录印尼语语音数据在不同的变异组。Whisper微调模型在具有各种特征的数据集上取得了最佳结果,如单词错误率(WER)和字符错误率(CER)的下降所示。此外,我们发现,说话风格的变异性影响模型的性能。摘要:An ideal speech recognition model has the capability to transcribe speech accurately under various characteristics of speech signals, such as speaking style (read and spontaneous), speech context (formal and informal), and background noise conditions (clean and moderate). Building such a model requires a significant amount of training data with diverse speech characteristics. Currently, Indonesian data is dominated by read, formal, and clean speech, leading to a scarcity of Indonesian data with other speech variabilities. To develop Indonesian automatic speech recognition (ASR), we present our research on state-of-the-art speech recognition models, namely Massively Multilingual Speech (MMS) and Whisper, as well as compiling a dataset comprising Indonesian speech with variabilities to facilitate our study. We further investigate the models' predictive ability to transcribe Indonesian speech data across different variability groups. The best results were achieved by the Whisper fine-tuned model across datasets with various characteristics, as indicated by the decrease in word error rate (WER) and character error rate (CER). Moreover, we found that speaking style variability affected model performance the most.

【6】 Small Tunes Transformer: Exploring Macro & Micro-Level Hierarchies for Skeleton-Conditioned Melody Generation
标题: 小型收件箱Transformer:探索普林斯顿条件旋律生成的宏观和微观层次结构
作者: Yishan Lv, Jing Luo, Boyuan Ju, Xinyu Yang
链接:点击下载PDF文件
摘要:近年来,符号音乐生成已经成为众多深度学习研究的焦点。结构作为音乐的重要组成部分,有助于提高音乐的质量,越来越多的作品开始研究层次结构。本研究从宏观层次和微观层次两个方面对音乐的多层次结构进行了探讨。在宏观层次上,通过乐句切分算法探索乐句对音乐整体发展的影响;在微观层次上,通过骨架音符提取策略探索乐句中的骨架音符对旋律生成的指导作用。此外,我们提出了一种新的短语级交叉注意机制,以捕捉宏观层次和微观层次之间的内在关系。此外,针对目前缺乏对中国风格音乐的研究,我们构建了我们的小样本数据集:一个由10088个小样本文件组成的大样本数据集,这是一个传统的中国民歌类别。该数据集是我们研究的重点。我们利用提取的骨架音符作为条件来生成小旋律歌曲,实验结果表明,我们提出的模型,小旋律Transformer,优于其他国家的最先进的模型。此外,我们设计了三个新的客观评价指标,从节奏和旋律两个维度来评价音乐。摘要:Recently, symbolic music generation has become a focus of numerous deep learning research. Structure as an important part of music, contributes to improving the quality of music, and an increasing number of works start to study the hierarchical structure. In this study, we delve into the multi-level structures within music from macro-level and micro-level hierarchies. At the macro-level hierarchy, we conduct phrase segmentation algorithm to explore how phrases influence the overall development of music, and at the micro-level hierarchy, we design skeleton notes extraction strategy to explore how skeleton notes within each phrase guide the melody generation. Furthermore, we propose a novel Phrase-level Cross-Attention mechanism to capture the intrinsic relationship between macro-level hierarchy and micro-level hierarchy. Moreover, in response to the current lack of research on Chinese-style music, we construct our Small Tunes Dataset: a substantial collection of MIDI files comprising 10088 Small Tunes, a category of traditional Chinese Folk Songs. This dataset serves as the focus of our study. We generate Small Tunes songs utilizing the extracted skeleton notes as conditions, and experiment results indicate that our proposed model, Small Tunes Transformer, outperforms other state-of-the-art models. Besides, we design three novel objective evaluation metrics to evaluate music from both rhythm and melody dimensions.

【7】 Symbolic Music Generation with Fine-grained Interactive Textural Guidance
标题: 具有细粒度交互文本指导的象征性音乐生成
作者: Tingyu Zhu, Haoyu Liu, Zhimin Jiang, Zeyu Zheng
链接:点击下载PDF文件
摘要:由于有限的数据可用性和对音符音高的高精度要求的结合,符号音乐生成的问题提出了独特的挑战。为了克服这些困难,我们在扩散模型中引入细粒度纹理指导(FTG)来纠正学习分布中的错误。通过结合FTG,扩散模型提高了音乐生成的准确性,使其非常适合渐进式音乐生成、即兴创作和交互式音乐创作等高级任务。我们得到的理论表征的挑战,在象征性的音乐生成和FTG方法的效果。我们提供了数值实验和一个演示页面,用于交互式音乐生成与用户输入,以展示我们的方法的有效性。摘要:The problem of symbolic music generation presents unique challenges due to the combination of limited data availability and the need for high precision in note pitch. To overcome these difficulties, we introduce Fine-grained Textural Guidance (FTG) within diffusion models to correct errors in the learned distributions. By incorporating FTG, the diffusion models improve the accuracy of music generation, which makes them well-suited for advanced tasks such as progressive music generation, improvisation and interactive music creation. We derive theoretical characterizations for both the challenges in symbolic music generation and the effect of the FTG approach. We provide numerical experiments and a demo page for interactive music generation with user input to showcase the effectiveness of our approach.

【8】 The language of sound search: Examining User Queries in Audio Search Engines
标题: 声音搜索语言:检查音频搜索引擎中的用户收件箱
作者: Benno Weck, Frederic Font
备注:Accepted at DCASE 2024. Supplementary materials at this https URL
链接:点击下载PDF文件
摘要:本研究探讨文字,用户编写的搜索查询的背景下,声音搜索引擎,包括各种应用程序,如福利,音效,和一般的音频检索。当前的研究在设计基于文本的音频检索系统时没有充分解决现实世界的用户需求和行为。为了弥补这一差距,我们分析了两个来源的搜索查询:自定义调查和Freesound网站查询日志。该调查旨在收集一个不受限制的,假设的声音搜索引擎的查询,从而产生一个数据集,捕捉用户的意图,而不受现有系统的限制。该数据集也可与研究界共享。相比之下,Freesound查询日志包含大约900万个搜索请求,提供了真实世界使用模式的全面视图。我们的研究结果表明,调查查询一般比Freesound查询,这表明用户更喜欢详细的查询时,不受系统限制。这两个数据集主要以基于关键词的查询为特征,很少有调查参与者使用完整的句子。影响调查查询的关键因素包括主要声源、预期用途、感知位置和声源数量。这些见解对于开发以用户为中心,有效的基于文本的音频检索系统至关重要,增强了我们对用户在声音搜索环境中行为的理解。摘要:This study examines textual, user-written search queries within the context of sound search engines, encompassing various applications such as foley, sound effects, and general audio retrieval. Current research inadequately addresses real-world user needs and behaviours in designing text-based audio retrieval systems. To bridge this gap, we analysed search queries from two sources: a custom survey and Freesound website query logs. The survey was designed to collect queries for an unrestricted, hypothetical sound search engine, resulting in a dataset that captures user intentions without the constraints of existing systems. This dataset is also made available for sharing with the research community. In contrast, the Freesound query logs encompass approximately 9 million search requests, providing a comprehensive view of real-world usage patterns. Our findings indicate that survey queries are generally longer than Freesound queries, suggesting users prefer detailed queries when not limited by system constraints. Both datasets predominantly feature keyword-based queries, with few survey participants using full sentences. Key factors influencing survey queries include the primary sound source, intended usage, perceived location, and the number of sound sources. These insights are crucial for developing user-centred, effective text-based audio retrieval systems, enhancing our understanding of user behaviour in sound search contexts.

【9】 Music Genre Classification using Large Language Models
标题: 使用大型语言模型的音乐流派分类
作者: Mohamed El Amine Meguenani, Alceu de Souza Britto Jr., Alessandro Lameiras Koerich
备注:7 pages
链接:点击下载PDF文件
摘要:本文利用预训练的大型语言模型(LLM)的zero-shot能力进行音乐流派分类。所提出的方法将音频信号分成20 ms的块,并通过卷积特征编码器、Transformer编码器和用于编码音频单元和生成特征向量的附加层来处理它们。提取的特征向量用于训练分类头。在推断期间,对各个块的预测被聚合以用于最终的流派分类。我们对LLM(包括WavLM、HuBERT和wav2vec 2.0)与传统的深度学习架构(如1D和2D卷积神经网络(CNN)和音频频谱图Transformer(AST))进行了全面的比较。我们的研究结果表明,AST模型的性能优越,总体准确率达到85.5%,超过了所有其他评估的模型。这些结果突出了LLM和基于变压器的架构在推进音乐信息检索任务方面的潜力,即使在zero-shot场景中也是如此。摘要:This paper exploits the zero-shot capabilities of pre-trained large language models (LLMs) for music genre classification. The proposed approach splits audio signals into 20 ms chunks and processes them through convolutional feature encoders, a transformer encoder, and additional layers for coding audio units and generating feature vectors. The extracted feature vectors are used to train a classification head. During inference, predictions on individual chunks are aggregated for a final genre classification. We conducted a comprehensive comparison of LLMs, including WavLM, HuBERT, and wav2vec 2.0, with traditional deep learning architectures like 1D and 2D convolutional neural networks (CNNs) and the audio spectrogram transformer (AST). Our findings demonstrate the superior performance of the AST model, achieving an overall accuracy of 85.5%, surpassing all other models evaluated. These results highlight the potential of LLMs and transformer-based architectures for advancing music information retrieval tasks, even in zero-shot scenarios.

【10】 A Recurrent Neural Network Approach to the Answering Machine Detection Problem
标题: 解决志愿服务机检测问题的回归神经网络方法
作者: Kemal Altwlkany, Sead Delalic, Elmedin Selmanovic, Adis Alihodzic, Ivica Lovric
备注:6 pages, 2 figures, 2024 47th MIPRO ICT and Electronics Convention (MIPRO)
链接:点击下载PDF文件
摘要:在电信和云通信领域,准确且实时地检测是人还是应答机应答了呼出呼叫是至关重要的。这个问题在竞选期间特别重要,因为它通过精确的呼叫者识别来提高服务质量、效率和降低成本。尽管该领域的重要性,它仍然没有充分探讨现有的文献。本文提出了一种创新的应答机检测方法,该方法通过YAMNet模型利用迁移学习进行特征提取。YAMNet架构有助于训练基于递归的分类器,从而实现音频流的实时处理,而不是固定长度的记录。结果表明,在测试集上的准确率超过96%。此外,我们进行了深入的分析误分类的样本,并揭示了超过98%的准确度可以实现与集成的沉默检测算法,如FFmpeg提供的一个。摘要:In the field of telecommunications and cloud communications, accurately and in real-time detecting whether a human or an answering machine has answered an outbound call is of paramount importance. This problem is of particular significance during campaigns as it enhances service quality, efficiency and cost reduction through precise caller identification. Despite the significance of the field, it remains inadequately explored in the existing literature. This paper presents an innovative approach to answering machine detection that leverages transfer learning through the YAMNet model for feature extraction. The YAMNet architecture facilitates the training of a recurrent-based classifier, enabling real-time processing of audio streams, as opposed to fixed-length recordings. The results demonstrate an accuracy of over 96% on the test set. Furthermore, we conduct an in-depth analysis of misclassified samples and reveal that an accuracy exceeding 98% can be achieved with the integration of a silence detection algorithm, such as the one provided by FFmpeg.


机器翻译,仅供参考