今日论文合集:cs.SD语音17篇,eess.AS音频处理18篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 dMel: Speech Tokenization made Simple
标题: dMel:语音代币化变得简单
作者:He Bai,Tatiana Likhomanenko,Ruixiang Zhang,Zijin Gu,Zakaria Aldeneh,Navdeep Jaitly
备注:under review
链接:点击下载PDF文件
摘要:大型语言模型通过利用大量文本数据的自我监督预训练,彻底改变了自然语言处理。受这一成功的启发,研究人员研究了复杂的语音标记化方法来离散连续语音信号,以便语言建模技术可以应用于语音数据。然而,现有的方法要么模型语义令牌,潜在地丢失声学信息,或者模型声学令牌,有丢失语义信息的风险。拥有多个令牌类型也会使架构复杂化,并需要额外的预训练。在这里,我们表明,离散梅尔滤波器组通道到离散强度箱产生一个简单的表示(dMel),比其他现有的语音标记化方法的性能更好。使用一个Transformer解码器的语音-文本模型,我们全面评估不同的语音标记化方法的语音识别(ASR),语音合成(TTS)。我们的研究结果表明,dMel在一个统一的框架内实现两个任务的高性能的有效性,为高效和有效的语音和文本的联合建模铺平了道路。摘要:Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have investigated complicated speech tokenization methods to discretize continuous speech signals so that language modeling techniques can be applied to speech data. However, existing approaches either model semantic tokens, potentially losing acoustic information, or model acoustic tokens, risking the loss of semantic information. Having multiple token types also complicates the architecture and requires additional pretraining. Here we show that discretizing mel-filterbank channels into discrete intensity bins produces a simple representation (dMel), that performs better than other existing speech tokenization methods. Using a transformer decoder-only architecture for speech-text modeling, we comprehensively evaluate different speech tokenization methods on speech recognition (ASR), speech synthesis (TTS). Our results demonstrate the effectiveness of dMel in achieving high performance on both tasks within a unified framework, paving the way for efficient and effective joint modeling of speech and text.

【2】 J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
标题: J-CHAT:用于口语对话语言建模的日本大型口语对话库
作者:Wataru Nakata,Kentaro Seki,Hitomi Yanaka,Yuki Saito,Shinnosuke Takamichi,Hiroshi Saruwatari
备注:8 pages, 6 figures
链接:点击下载PDF文件
摘要:口语对话在人类与人工智能的交互中起着至关重要的作用,因此需要面向对话的口语模型(SLM)。为了开发多功能的SLM,大规模和多样化的语音数据集是必不可少的。此外,为了确保高质量的语音生成,数据必须像野生数据一样是自发的,并且必须在去除噪声的情况下声学上干净。尽管迫切需要,但没有满足所有这些标准的开源语料库。本研究通过构建和发布一个大规模的口语对话语料库来解决这一差距,该语料库名为Japanese Corpus for Human-AI Talks(J-CHAT),可公开访问。此外,本文提出了一种独立于语言的语料库建设方法,并描述了使用J-CHAT训练的SLM对话生成实验。实验结果表明,该方法收集的多个领域的数据提高了对话生成的自然性和意义。摘要:Spoken dialogue plays a crucial role in human-AI interactions, necessitating dialogue-oriented spoken language models (SLMs). To develop versatile SLMs, large-scale and diverse speech datasets are essential. Additionally, to ensure hiqh-quality speech generation, the data must be spontaneous like in-wild data and must be acoustically clean with noise removed. Despite the critical need, no open-source corpus meeting all these criteria has been available. This study addresses this gap by constructing and releasing a large-scale spoken dialogue corpus, named Japanese Corpus for Human-AI Talks (J-CHAT), which is publicly accessible. Furthermore, this paper presents a language-independent method for corpus construction and describes experiments on dialogue generation using SLMs trained on J-CHAT. Experimental results indicate that the collected data from multiple domains by our method improve the naturalness and meaningfulness of dialogue generation.

【3】 Computer Audition: From Task-Specific Machine Learning to Foundation Models
标题: 计算机试镜:从特定任务的机器学习到基础模型
作者:Andreas Triantafyllopoulos,Iosif Tsangko,Alexander Gebhard,Annamaria Mesaros,Tuomas Virtanen,Björn Schuller
链接:点击下载PDF文件
摘要:基础模型(FM)越来越多地引领着计算机听觉范围内的各种任务的最新进展-使用机器来理解声音。与传统管道相比,它们具有几个优势:其中包括将多个任务整合到单个模型中的能力,利用其他模式的知识的选择,以及与人类用户的现成交互。自然,这些承诺在音频社区中引起了极大的兴奋,并导致了一波早期尝试构建新的通用音频基础模型的浪潮。在目前的贡献中,我们给出了一个概述的计算音频分析,因为它从传统的管道过渡到听觉基础模型。我们的工作突出了支撑这些模型的关键操作原则,并展示了它们如何适应音频社区以前单独处理的多个任务。摘要:Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand sounds. They feature several advantages over traditional pipelines: among others, the ability to consolidate multiple tasks in a single model, the option to leverage knowledge from other modalities, and the readily-available interaction with human users. Naturally, these promises have created substantial excitement in the audio community, and have led to a wave of early attempts to build new, general-purpose foundation models for audio. In the present contribution, we give an overview of computational audio analysis as it transitions from traditional pipelines towards auditory foundation models. Our work highlights the key operating principles that underpin those models, and showcases how they can accommodate multiple tasks that the audio community previously tackled separately.

【4】 Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing
标题: Annealed Multiple Choice Learning:通过Annealed克服赢家通吃的局限性
作者:David Perera,Victor Letzelter,Théo Mariotte,Adrien Cortés,Mickael Chen,Slim Essid,Gaël Richard
链接:点击下载PDF文件
摘要:我们介绍了退火多选择学习(aMCL),它结合了模拟退火和MCL。MCL是一个学习框架,通过预测一小部分合理的假设来处理模糊的任务。这些假设使用Winner-takes-all(WTA)方案进行训练,该方案促进了预测的多样性。然而,由于WTA的贪婪性质,该方案可能会收敛到任意次优的局部最小值。我们使用退火克服了这一限制,这增强了训练过程中对假设空间的探索。我们利用统计物理学和信息论的见解来提供模型训练轨迹的详细描述。此外,我们验证了我们的算法通过广泛的实验合成数据集,标准的UCI基准,语音分离。摘要:We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversity of the predictions. However, this scheme may converge toward an arbitrarily suboptimal local minimum, due to the greedy nature of WTA. We overcome this limitation using annealing, which enhances the exploration of the hypothesis space during training. We leverage insights from statistical physics and information theory to provide a detailed description of the model training trajectory. Additionally, we validate our algorithm by extensive experiments on synthetic datasets, on the standard UCI benchmark, and on speech separation.

【5】 SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
标题: SELM:增强域外场景的语音情感识别
作者:Hazim Bukhari,Soham Deshmukh,Hira Dhamyal,Bhiksha Raj,Rita Singh
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音情感识别(SER)传统上被制定为一个分类任务。然而,情绪通常是一个频谱,其分布因情况而异,导致域外(OOD)性能差。我们从自动语音识别(ASR)的统计公式中获得灵感,并将SER任务制定为生成最可能的文本标记序列来推断情感。该公式将SER分解为预测由语言模型预测加权的声学模型特征。作为这种方法的一个实例,我们提出了SELM,一个音频条件的语言模型SER,预测不同的情感观点。我们训练SELM策展语音情感语料库和测试三个OOD数据集(RAVDESS,CREMAD,IEMOCAP)没有在训练中使用。SELM在最先进的基线上实现了显着的改进,RAVDESS和CREMA-D的相对精度分别提高了17%和7%。此外,SELM可以通过使用一些注释示例的Few-Shot学习来进一步提高其性能。结果突出了我们的SER配方的有效性,特别是在OOD场景中提高性能。摘要:Speech Emotion Recognition (SER) has been traditionally formulated as a classification task. However, emotions are generally a spectrum whose distribution varies from situation to situation leading to poor Out-of-Domain (OOD) performance. We take inspiration from statistical formulation of Automatic Speech Recognition (ASR) and formulate the SER task as generating the most likely sequence of text tokens to infer emotion. The formulation breaks SER into predicting acoustic model features weighted by language model prediction. As an instance of this approach, we present SELM, an audio-conditioned language model for SER that predicts different emotion views. We train SELM on curated speech emotion corpus and test it on three OOD datasets (RAVDESS, CREMAD, IEMOCAP) not used in training. SELM achieves significant improvements over the state-of-the-art baselines, with 17% and 7% relative accuracy gains for RAVDESS and CREMA-D, respectively. Moreover, SELM can further boost its performance by Few-Shot Learning using a few annotated examples. The results highlight the effectiveness of our SER formulation, especially to improve performance in OOD scenarios.

【6】 Explainability Paths for Sustained Artistic Practice with AI
标题: 人工智能持续艺术实践的解释途径
作者:Austin Tecks,Thomas Peschlow,Gabriel Vigliensoni
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:人工智能驱动的生成音频的发展反映了更广泛的人工智能趋势,通常以牺牲可解释性为代价优先考虑即时可访问性。因此,将这些工具纳入持续的艺术实践仍然是一个重大挑战。在本文中,我们探索了几条提高可解释性的途径,主要是从我们在训练和实现生成音频模型方面的研究-创作实践中得出的。作为提高可解释性的实际规定,我们强调了培训材料的人类代理,小规模数据集的可行性,迭代创造过程的便利化,以及交互式机器学习作为映射工具的集成。重要的是,这些步骤旨在增强人类在生成AI系统上的代理能力,不仅在模型推理过程中,而且在管理和预处理训练数据以及模型训练阶段。摘要:The development of AI-driven generative audio mirrors broader AI trends, often prioritizing immediate accessibility at the expense of explainability. Consequently, integrating such tools into sustained artistic practice remains a significant challenge. In this paper, we explore several paths to improve explainability, drawing primarily from our research-creation practice in training and implementing generative audio models. As practical provisions for improved explainability, we highlight human agency over training materials, the viability of small-scale datasets, the facilitation of the iterative creative process, and the integration of interactive machine learning as a mapping tool. Importantly, these steps aim to enhance human agency over generative AI systems not only during model inference, but also when curating and preprocessing training data as well as during the training phase of models.

【7】 MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation
标题: MusiConGen:基于转换器的文本到音乐生成的节奏和和弦控制
作者:Yun-Han Lan,Wen-Yi Hsiao,Hao-Chung Cheng,Yi-Hsuan Yang
备注:Accepted by the 25th International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:现有的文本到音乐模型可以产生具有极大多样性的高质量音频。然而,单独的文本提示不能精确地控制所生成的音乐的时间音乐特征,诸如和弦和节奏。为了应对这一挑战,我们引入了MusiConGen,这是一种基于时间条件的基于transformer的文本到音乐模型,它建立在预训练的MusicGen框架之上。我们的创新在于为消费级GPU量身定制的高效微调机制,该机制集成了自动提取的节奏和和弦作为条件信号。在推断期间,条件可以是从参考音频信号提取的音乐特征,或者是用户定义的符号和弦序列、BPM和文本提示。我们对两个数据集(一个来自提取的特征,另一个来自用户创建的输入)的性能评估表明,MusiConGen可以生成与指定条件一致的真实背景音乐。我们开源代码和模型检查点,并在线提供音频示例,https: musicongen.github.io musicongen_demo 。摘要:Existing text-to-music models can produce high-quality audio with great diversity. However, textual prompts alone cannot precisely control temporal musical features such as chords and rhythm of the generated music. To address this challenge, we introduce MusiConGen, a temporally-conditioned Transformer-based text-to-music model that builds upon the pretrained MusicGen framework. Our innovation lies in an efficient finetuning mechanism, tailored for consumer-grade GPUs, that integrates automatically-extracted rhythm and chords as the condition signal. During inference, the condition can either be musical features extracted from a reference audio signal, or be user-defined symbolic chord sequence, BPM, and textual prompts. Our performance evaluation on two datasets -- one derived from extracted features and the other from user-created inputs -- demonstrates that MusiConGen can generate realistic backing track music that aligns well with the specified conditions. We open-source the code and model checkpoints, and provide audio examples online, https: musicongen.github.io musicongen_demo .

【8】 Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
标题: 对话鲁BERT用于检测SVR转录对话中的竞争中断
作者:Dmitrii Galimzianov,Viacheslav Vyshegorodtsev
Journal-ref:Computer Science & Information Technology (CS & IT), ISSN : 2231 - 5403, Volume 14, Number 13, June 2024
链接:点击下载PDF文件
摘要:对话中的中断发生在听者在当前说话者结束讲话之前开始讲话时。打断可以大致分为两类:合作型(当听者想要支持说话者时)和竞争型(当听者试图违背说话者的意愿控制谈话时)。自动分类中断的系统可以用于呼叫中心,特别是在客户满意度监控和代理监控的任务中。在这项研究中,我们开发了一个基于文本的中断分类模型,准备一个内部数据集组成的ASR转录的客户支持电话对话在俄罗斯。我们在我们的数据集上微调了对话式RuBERT,并优化了超参数,模型表现良好。通过进一步的改进,该模型可以应用于自动监控系统。摘要:Interruption in a dialogue occurs when the listener begins their speech before the current speaker finishes speaking. Interruptions can be broadly divided into two groups: cooperative (when the listener wants to support the speaker), and competitive (when the listener tries to take control of the conversation against the speaker's will). A system that automatically classifies interruptions can be used in call centers, specifically in the tasks of customer satisfaction monitoring and agent monitoring. In this study, we developed a text-based interruption classification model by preparing an in-house dataset consisting of ASR-transcribed customer support telephone dialogues in Russian. We fine-tuned Conversational RuBERT on our dataset and optimized hyperparameters, and the model performed well. With further improvements, the proposed model can be applied to automatic monitoring systems.

【9】 Composer's Assistant 2: Interactive Multi-Track MIDI Infilling with Fine-Grained User Control
标题: 作曲家助理2:交互式多轨收件箱填充细粒度用户控制
作者:Martin E. Malandro
备注:8 pages, 6 figures, 2 tables. To be published in ISMIR 2024
链接:点击下载PDF文件
摘要:我们介绍作曲家的助手2,在收割机数字音频工作站的交互式人机合成系统。我们的工作升级了作曲家的助手系统(它执行多轨道填充的符号音乐在轨道措施的水平)与广泛的新控件,让用户细粒度控制系统的输出。在这项工作中引入的控制包括两种类型的节奏调节控制,水平和垂直注意开始密度控制,几种类型的音高控制,和节奏感兴趣的控制。我们训练一个类似T5的Transformer模型来实现这些控制,并作为我们系统的骨干。有了这些控制,我们实现了显着的改善,在原来的系统的客观指标。我们还研究了我们的模型对控件含义的理解程度,并且我们进行了一项听力研究,没有发现真实音乐和与我们的系统以共同创作的方式创作的音乐之间存在显着差异。我们发布了完整的系统,包括源代码,预训练模型和REAPER脚本。摘要:We introduce Composer's Assistant 2, a system for interactive human-computer composition in the REAPER digital audio workstation. Our work upgrades the Composer's Assistant system (which performs multi-track infilling of symbolic music at the track-measure level) with a wide range of new controls to give users fine-grained control over the system's outputs. Controls introduced in this work include two types of rhythmic conditioning controls, horizontal and vertical note onset density controls, several types of pitch controls, and a rhythmic interest control. We train a T5-like transformer model to implement these controls and to serve as the backbone of our system. With these controls, we achieve a dramatic improvement in objective metrics over the original system. We also study how well our model understands the meaning of our controls, and we conduct a listening study that does not find a significant difference between real music and music composed in a co-creative fashion with our system. We release our complete system, consisting of source code, pretrained models, and REAPER scripts.

【10】 Morse Code-Enabled Speech Recognition for Individuals with Visual and Hearing Impairments
标题: 针对视觉和听力障碍的个人的莫尔斯码语音识别
作者:Ritabrata Roy Choudhury
备注:10 pages, 11 figures, 4 tables
链接:点击下载PDF文件
摘要:该模型旨在为听力、语言或认知残疾人开发语音识别技术。语音识别领域的所有可用技术都没有为有听力、语言或认知障碍的人提供交流界面。所提出的模型提出了从用户的语音,被发送到语音识别层,在那里它被转换成文本,然后该文本被发送到莫尔斯码转换层,在那里相应的语音的莫尔斯码作为输出。该模型的准确性完全依赖于语音识别,因为莫尔斯码转换是一个过程。该模型进行了测试,录制的音频文件具有不同的参数。所提出的模型的WER和准确度都被确定为10.18%和89.82%,分别。摘要:The proposed model aims to develop a speech recognition technology for hearing, speech, or cognitively disabled people. All the available technology in the field of speech recognition doesn't come with an interface for communication for people with hearing, speech, or cognitive disabilities. The proposed model proposes the speech from the user, is transmitted to the speech recognition layer where it is converted into text and then that text is then transmitted to the morse code conversion layer where the morse code of the corresponding speech is given as the output. The accuracy of the model is completely dependent on speech recognition, as the morse code conversion is a process. The model is tested with recorded audio files with different parameters. The proposed model's WER and accuracy are both determined to be 10.18% and 89.82%, respectively.

【11】 Generating Sample-Based Musical Instruments Using Neural Audio Codec Language Models
标题: 使用神经音频编解码器语言模型生成基于样本的乐器
作者:Shahan Nercessian,Johannes Imort,Ninon Devis,Frederik Blang
备注:8 pages, 2 figures. Accepted to the 25th Conference of the International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:在本文中,我们提出并研究了使用神经音频编解码器语言模型的自动生成基于文本或参考音频提示的样本为基础的乐器。我们的方法扩展了一个生成音频框架,以88键频谱,速度和组合的文本 音频嵌入的音高为条件。我们确定保持音色的一致性所产生的文书作为一个主要的挑战。为了解决这个问题,我们引入三种不同的条件反射方案。我们通过客观的指标和人类听力测试来分析我们的方法,证明我们的方法可以制作出令人信服的乐器。具体来说,我们引入了一个新的客观指标来评估生成的乐器的音色一致性,并调整文本到乐器的情况下的平均对比音频预训练(CLAP)分数,注意到它的天真的应用程序是不适合评估这项任务。我们的研究结果揭示了音色一致性,生成的样本的质量,以及它们与输入提示的对应关系之间复杂的相互作用。摘要:In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio framework to condition on pitch across an 88-key spectrum, velocity, and a combined text audio embedding. We identify maintaining timbral consistency within the generated instruments as a major challenge. To tackle this issue, we introduce three distinct conditioning schemes. We analyze our methods through objective metrics and human listening tests, demonstrating that our approach can produce compelling musical instruments. Specifically, we introduce a new objective metric to evaluate the timbral consistency of the generated instruments and adapt the average Contrastive Language-Audio Pretraining (CLAP) score for the text-to-instrument case, noting that its naive application is unsuitable for assessing this task. Our findings reveal a complex interplay between timbral consistency, the quality of generated samples, and their correspondence to the input prompt.

【12】 DSP-informed bandwidth extension using locally-conditioned excitation and linear time-varying filter subnetworks
标题: 使用局部条件激励和线性时变过滤器子网络进行DSP通知的带宽扩展
作者:Shahan Nercessian,Alexey Lukin,Johannes Imort
备注:5 pages, 3 figures. Accepted to the 18th International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个双级架构的带宽扩展(BWE)增加有效的语音信号的采样率从8 kHz到48 kHz。与现有的端到端深度学习模型不同,我们提出的方法使用激励和线性时变(LTV)滤波器阶段显式地对BWE进行建模。激励阶段加宽输入的频谱,而滤波阶段基于来自声学特征预测器的输出对其进行适当整形。为此,声学特征损失项可以隐含地促进激励子网络在要合成的上频带中产生白谱。实验结果表明,通过我们的方法提供的附加电感偏置可以使用来自SEANet或HiFi-GAN的发生器作为激励器来改善BWE结果,并且我们使用声学特征预测来适应处理的方法比HiFi-GAN-2中使用的方法更有效。次要贡献包括扩展的SEANet模型,以适应当地的条件信息,以及应用HiFi-GAN-2斗轮挖掘机问题。摘要:In this paper, we propose a dual-stage architecture for bandwidth extension (BWE) increasing the effective sampling rate of speech signals from 8 kHz to 48 kHz. Unlike existing end-to-end deep learning models, our proposed method explicitly models BWE using excitation and linear time-varying (LTV) filter stages. The excitation stage broadens the spectrum of the input, while the filtering stage properly shapes it based on outputs from an acoustic feature predictor. To this end, an acoustic feature loss term can implicitly promote the excitation subnetwork to produce white spectra in the upper frequency band to be synthesized. Experimental results demonstrate that the added inductive bias provided by our approach can improve upon BWE results using the generators from both SEANet or HiFi-GAN as exciters, and that our means of adapting processing with acoustic feature predictions is more effective than that used in HiFi-GAN-2. Secondary contributions include extensions of the SEANet model to accommodate local conditioning information, as well as the application of HiFi-GAN-2 for the BWE problem.

【13】 EMO-Codec: A Depth Look at Emotion Preservation Capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
标题: EMO-Codec:通过主观和客观评估深入研究传统和神经Codec模型的情绪保存能力
作者:Wenze Ren,Yi-Cheng Lin,Huang-Cheng Chou,Haibin Wu,Yi-Chiao Wu,Chi-Chun Lee,Hung-yi Lee,Yu Tsao
链接:点击下载PDF文件
摘要:神经编解码器模型减少了语音数据传输延迟,并作为语音语言模型(语音LM)的基础标记。在编解码器中保留情感信息对于有效沟通和上下文理解至关重要。然而,在现有的编解码器中缺乏对情感损失的研究。本文使用主观和客观方法对IEMOCAP等情感数据集进行评估。我们的研究确定了哪些编解码器在各种比特率情况下最好地保留了情感信息。我们发现,用英文和中文数据训练编解码器模型在保留中文情感信息方面取得的成功有限。此外,通过这些编解码器重新合成语音会降低语音情感识别(SER)的性能,特别是对于悲伤,抑郁,恐惧和厌恶等情绪。人类听力测试证实了这些发现。这项工作指导未来的语音技术发展,以确保新的编解码器保持语音中情感信息的完整性。摘要:The neural codec model reduces speech data transmission delay and serves as the foundational tokenizer for speech language models (speech LMs). Preserving emotional information in codecs is crucial for effective communication and context understanding. However, there is a lack of studies on emotion loss in existing codecs. This paper evaluates neural and legacy codecs using subjective and objective methods on emotion datasets like IEMOCAP. Our study identifies which codecs best preserve emotional information under various bitrate scenarios. We found that training codec models with both English and Chinese data had limited success in retaining emotional information in Chinese. Additionally, resynthesizing speech through these codecs degrades the performance of speech emotion recognition (SER), particularly for emotions like sadness, depression, fear, and disgust. Human listening tests confirmed these findings. This work guides future speech technology developments to ensure new codecs maintain the integrity of emotional information in speech.

【14】 Integrating IP Broadcasting with Audio Tags- Workflow and Challenges
标题: 将IP广播与音频标签集成-工作流程和挑战
作者:Rhys Burchett-Vass,Arshdeep Singh,Gabriel Bibbó,Mark D. Plumbley
备注:Submitted to DCASE 2024 Workshop
链接:点击下载PDF文件
摘要:广播业越来越多地采用IP技术,从新闻采集到现场音乐活动,彻底改变了现场和预先录制的内容制作。IP广播允许以易于配置的方式传输音频和视频信号,与现代网络技术保持一致。这种向IP工作流的转变不仅在路由信号方面,而且在使用标准Web开发技术的工具集成方面,都具有更大的灵活性。一个可能的工具可以包括使用现场音频标记,这在内容制作中有许多用途。这些包括从自动隐藏字幕到识别场景中不需要的声音事件。在本文中,我们描述了将音频标记模型容器化到微服务中的过程,微服务是一个可以集成到多种不同网络设置中的小型隔离代码模块。我们的目标是开发一个模块化的,可访问的,灵活的工具,能够无缝部署到各种规模的广播工作流程,从小制作到大公司。所选择的音频标记模型的延迟及其对最终产品的有用性的影响周围的挑战进行了讨论。摘要:The broadcasting industry is increasingly adopting IP techniques, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allows for the transport of audio and video signals in an easily configurable way, aligning with modern networking techniques. This shift towards an IP workflow allows for much greater flexibility, not only in routing signals but with the integration of tools using standard web development techniques. One possible tool could include the use of live audio tagging, which has a number of uses in the production of content. These include from automated closed captioning to identifying unwanted sound events within a scene. In this paper, we describe the process of containerising an audio tagging model into a microservice, a small segregated code module that can be integrated into a multitude of different network setups. The goal is to develop a modular, accessible, and flexible tool capable of seamless deployment into broadcasting workflows of all sizes, from small productions to large corporations. Challenges surrounding latency of the selected audio tagging model and its effect on the usefulness of the end product are discussed.

【15】 Can all variations within the unified mask-based beamformer framework achieve identical peak extraction performance?
标题: 统一的基于屏蔽的射束形成器框架内的所有变体能否实现相同的峰值提取性能?
作者:Atsuo Hiroe,Katsutoshi Itoyama,Kazuhiro Nakadai
备注:Submitted to EURASIP journal on Audio, Speech, and Music Processing
链接:点击下载PDF文件
摘要:本文研究了基于掩模的波束形成器(BF),它使用时频掩模估计目标声提取(TSE)的滤波器。虽然已经提出了多个基于掩模的BF,但对于目标提取性能的最佳BF尚未达成共识。以前,我们发现,最大信噪比和最小均方误差(MSE)BF可以实现相同的提取性能的理论上限的性能,每个BF包含一个不同的最佳掩模。然而,这些显着的研究结果留下了两个问题没有解决:只有两个BF被覆盖,不包括最小方差无失真响应BF;和理想缩放(IS)被用来理想地调整输出规模,这是不适用于现实场景。为了解决这些覆盖和缩放问题,本研究提出了一个统一的框架,基于掩码的BF包括两个过程:滤波器估计,可以覆盖所有BF和缩放适用于现实的情况下,通过采用掩码生成一个缩放参考。我们还提出了一种方法来枚举所有可能的BF,并得出12个变化。通过最小化目标和BF输出之间的MSE来获得这两个过程的最佳掩模。使用CHiME-4数据集的实验结果表明:1)所有12种变化都可以达到理论上的上限性能,2)基于掩模的缩放可以表现为IS。这些结果可以通过考虑掩模的实际参数计数来解释。这些发现有助于1)设计TSE系统,2)估计BF的提取性能,以及3)结合基于掩模的缩放来提高缩放精度。这些贡献也适用于基于独立分量分析的TSE方法,因为统一框架也涵盖了它们。摘要:This study investigates mask-based beamformers (BFs), which estimate filters for target sound extraction (TSE) using time-frequency masks. Although multiple mask-based BFs have been proposed, no consensus has been established on the best one for target-extracting performance. Previously, we found that maximum signal-to-noise ratio and minimum mean square error (MSE) BFs can achieve the same extraction performance as the theoretical upper-bound performance, with each BF containing a different optimal mask. However, these remarkable findings left two issues unsolved: only two BFs were covered, excluding the minimum variance distortionless response BF; and ideal scaling (IS) was employed to ideally adjust the output scale, which is not applicable to realistic scenarios. To address these coverage and scaling issues, this study proposes a unified framework for mask-based BFs comprising two processes: filter estimation that can cover all BFs and scaling applicable to realistic scenarios by employing a mask to generate a scaling reference. We also propose a methodology to enumerate all possible BFs and derive 12 variations. Optimal masks for both processes are obtained by minimizing the MSE between the target and BF output. The experimental results using the CHiME-4 dataset suggested that 1) all 12 variations can achieve the theoretical upper-bound performance, and 2) mask-based scaling can behave as IS. These results can be explained by considering the practical parameter count of the masks. These findings contribute to 1) designing a TSE system, 2) estimating the extraction performance of a BF, and 3) improving scaling accuracy combined with mask-based scaling. The contributions also apply to TSE methods based on independent component analysis, as the unified framework covers them too.

【16】 Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
标题: 说话人建模及其应用概述:从深度说话人表示学习的角度
作者:Shuai Wang,Zhengyang Chen,Kong Aik Lee,Yanmin Qian,Haizhou Li
链接:点击下载PDF文件
摘要:说话人个性信息是语音信号中最重要的信息之一。通过对这些信息进行全面而准确的建模,它可以用于各种智能语音应用,例如说话人识别,说话人日记,语音合成和目标说话人提取。在这篇文章中,我们的目标是,从一个独特的角度来看,发展历史,范式转变,以及在深度表示学习框架的背景下,扬声器建模技术的应用领域。这篇综述旨在为说话人建模领域的研究人员以及那些希望将说话人建模技术应用于特定下游任务的人提供明确的参考。摘要:Speaker individuality information is among the most critical elements within speech signals. By thoroughly and accurately modeling this information, it can be utilized in various intelligent speech applications, such as speaker recognition, speaker diarization, speech synthesis, and target speaker extraction. In this article, we aim to present, from a unique perspective, the developmental history, paradigm shifts, and application domains of speaker modeling technologies within the context of deep representation learning framework. This review is designed to provide a clear reference for researchers in the speaker modeling field, as well as for those who wish to apply speaker modeling techniques to specific downstream tasks.

【17】 Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
标题: 使用可控情绪强度实现现实情绪语音转换
作者:Tianhua Qi,Shiyan Wang,Cheng Lu,Yan Zhao,Yuan Zong,Wenming Zheng
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
摘要:真实情感语音转换(EVC)旨在增强转换音频的情感多样性,使合成的语音更加真实自然。为此,我们提出了情绪强度感知网络(EINet),动态调整语调和节奏,通过纳入可控的情绪强度。为了更好地捕捉情感强度的细微差别,我们超越了声学特征之间的距离测量。相反,情感评估器是用来精确量化扬声器的情绪状态。通过使用强度映射器,获得强度伪标签,以弥合情感语音强度建模和运行时转换之间的差距。为了确保高的语音质量,同时保持可控性,情感渲染器用于在帧级平滑地将语言特征与操纵的情感特征相结合。此外,我们采用了持续时间预测,以促进自适应预测的节奏变化条件下指定的强度值。实验结果表明,与最先进的EVC方法相比,EINet在情感表达的自然性和多样性方面具有优越的性能。摘要:Realistic emotional voice conversion (EVC) aims to enhance emotional diversity of converted audios, making the synthesized voices more authentic and natural. To this end, we propose Emotional Intensity-aware Network (EINet), dynamically adjusting intonation and rhythm by incorporating controllable emotional intensity. To better capture nuances in emotional intensity, we go beyond mere distance measurements among acoustic features. Instead, an emotion evaluator is utilized to precisely quantify speaker's emotional state. By employing an intensity mapper, intensity pseudo-labels are obtained to bridge the gap between emotional speech intensity modeling and run-time conversion. To ensure high speech quality while retaining controllability, an emotion renderer is used for combining linguistic features smoothly with manipulated emotional features at frame level. Furthermore, we employ a duration predictor to facilitate adaptive prediction of rhythm changes condition on specifying intensity value. Experimental results show EINet's superior performance in naturalness and diversity of emotional expression compared to state-of-the-art EVC methods.


eess.AS音频处理
【1】 Robustness of Speech Separation Models for Similar-pitch Speakers
标题: 相似音调说话者语音分离模型的鲁棒性
作者:Bunlong Lay,Sebastian Zaczek,Kristina Tesch,Timo Gerkmann
链接:点击下载PDF文件
摘要:单通道语音分离是提高多说话人环境下语音识别系统性能的一个关键问题。本文研究了最先进的神经网络模型在说话人之间的音高差异最小的情况下的鲁棒性。基于Ditter和Gerkmann的早期发现,在类似的音高条件下,2018年Chimera++的性能显著下降,我们的研究将分析扩展到更近期和更复杂的神经网络模型。我们的实验表明,现代模型大大减少了匹配的训练和测试条件下的性能差距。然而,在不匹配的条件下,一个实质性的性能差距仍然存在,模型在大的音高差异下表现良好,但如果扬声器的音高相似,则表现较差。这些发现激发了进一步研究语音分离模型对相似音高扬声器和看不见的数据的可推广性。摘要:Single-channel speech separation is a crucial task for enhancing speech recognition systems in multi-speaker environments. This paper investigates the robustness of state-of-the-art Neural Network models in scenarios where the pitch differences between speakers are minimal. Building on earlier findings by Ditter and Gerkmann, which identified a significant performance drop for the 2018 Chimera++ under similar-pitch conditions, our study extends the analysis to more recent and sophisticated Neural Network models. Our experiments reveal that modern models have substantially reduced the performance gap for matched training and testing conditions. However, a substantial performance gap persists under mismatched conditions, with models performing well for large pitch differences but showing worse performance if the speakers' pitches are similar. These findings motivate further research into the generalizability of speech separation models to similar-pitch speakers and unseen data.

【2】 Generating Sample-Based Musical Instruments Using Neural Audio Codec Language Models
标题: 使用神经音频编解码器语言模型生成基于样本的乐器
作者:Shahan Nercessian,Johannes Imort,Ninon Devis,Frederik Blang
备注:8 pages, 2 figures. Accepted to the 25th Conference of the International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:在本文中,我们提出并研究了使用神经音频编解码器语言模型的自动生成基于文本或参考音频提示的样本为基础的乐器。我们的方法扩展了一个生成音频框架,以88键频谱,速度和组合的文本 音频嵌入的音高为条件。我们确定保持音色的一致性所产生的文书作为一个主要的挑战。为了解决这个问题,我们引入三种不同的条件反射方案。我们通过客观的指标和人类听力测试来分析我们的方法,证明我们的方法可以制作出令人信服的乐器。具体来说,我们引入了一个新的客观指标来评估生成的乐器的音色一致性,并调整文本到乐器的情况下的平均对比音频预训练(CLAP)分数,注意到它的天真的应用程序是不适合评估这项任务。我们的研究结果揭示了音色一致性,生成的样本的质量,以及它们与输入提示的对应关系之间复杂的相互作用。摘要:In this paper, we propose and investigate the use of neural audio codec language models for the automatic generation of sample-based musical instruments based on text or reference audio prompts. Our approach extends a generative audio framework to condition on pitch across an 88-key spectrum, velocity, and a combined text audio embedding. We identify maintaining timbral consistency within the generated instruments as a major challenge. To tackle this issue, we introduce three distinct conditioning schemes. We analyze our methods through objective metrics and human listening tests, demonstrating that our approach can produce compelling musical instruments. Specifically, we introduce a new objective metric to evaluate the timbral consistency of the generated instruments and adapt the average Contrastive Language-Audio Pretraining (CLAP) score for the text-to-instrument case, noting that its naive application is unsuitable for assessing this task. Our findings reveal a complex interplay between timbral consistency, the quality of generated samples, and their correspondence to the input prompt.

【3】 DSP-informed bandwidth extension using locally-conditioned excitation and linear time-varying filter subnetworks
标题: 使用局部条件激励和线性时变过滤器子网络进行DSP通知的带宽扩展
作者:Shahan Nercessian,Alexey Lukin,Johannes Imort
备注:5 pages, 3 figures. Accepted to the 18th International Workshop on Acoustic Signal Enhancement (IWAENC 2024)
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个双级架构的带宽扩展(BWE)增加有效的语音信号的采样率从8 kHz到48 kHz。与现有的端到端深度学习模型不同,我们提出的方法使用激励和线性时变(LTV)滤波器阶段显式地对BWE进行建模。激励阶段加宽输入的频谱,而滤波阶段基于来自声学特征预测器的输出对其进行适当整形。为此,声学特征损失项可以隐含地促进激励子网络在要合成的上频带中产生白谱。实验结果表明,通过我们的方法提供的附加电感偏置可以使用来自SEANet或HiFi-GAN的发生器作为激励器来改善BWE结果,并且我们使用声学特征预测来适应处理的方法比HiFi-GAN-2中使用的方法更有效。次要贡献包括扩展的SEANet模型,以适应当地的条件信息,以及应用HiFi-GAN-2斗轮挖掘机问题。摘要:In this paper, we propose a dual-stage architecture for bandwidth extension (BWE) increasing the effective sampling rate of speech signals from 8 kHz to 48 kHz. Unlike existing end-to-end deep learning models, our proposed method explicitly models BWE using excitation and linear time-varying (LTV) filter stages. The excitation stage broadens the spectrum of the input, while the filtering stage properly shapes it based on outputs from an acoustic feature predictor. To this end, an acoustic feature loss term can implicitly promote the excitation subnetwork to produce white spectra in the upper frequency band to be synthesized. Experimental results demonstrate that the added inductive bias provided by our approach can improve upon BWE results using the generators from both SEANet or HiFi-GAN as exciters, and that our means of adapting processing with acoustic feature predictions is more effective than that used in HiFi-GAN-2. Secondary contributions include extensions of the SEANet model to accommodate local conditioning information, as well as the application of HiFi-GAN-2 for the BWE problem.

【4】 EMO-Codec: A Depth Look at Emotion Preservation Capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
标题: EMO-Codec:通过主观和客观评估深入研究传统和神经Codec模型的情绪保存能力
作者:Wenze Ren,Yi-Cheng Lin,Huang-Cheng Chou,Haibin Wu,Yi-Chiao Wu,Chi-Chun Lee,Hung-yi Lee,Yu Tsao
链接:点击下载PDF文件
摘要:神经编解码器模型减少了语音数据传输延迟,并作为语音语言模型(语音LM)的基础标记。在编解码器中保留情感信息对于有效沟通和上下文理解至关重要。然而,在现有的编解码器中缺乏对情感损失的研究。本文使用主观和客观方法对IEMOCAP等情感数据集进行评估。我们的研究确定了哪些编解码器在各种比特率情况下最好地保留了情感信息。我们发现,用英文和中文数据训练编解码器模型在保留中文情感信息方面取得的成功有限。此外,通过这些编解码器重新合成语音会降低语音情感识别(SER)的性能,特别是对于悲伤,抑郁,恐惧和厌恶等情绪。人类听力测试证实了这些发现。这项工作指导未来的语音技术发展,以确保新的编解码器保持语音中情感信息的完整性。摘要:The neural codec model reduces speech data transmission delay and serves as the foundational tokenizer for speech language models (speech LMs). Preserving emotional information in codecs is crucial for effective communication and context understanding. However, there is a lack of studies on emotion loss in existing codecs. This paper evaluates neural and legacy codecs using subjective and objective methods on emotion datasets like IEMOCAP. Our study identifies which codecs best preserve emotional information under various bitrate scenarios. We found that training codec models with both English and Chinese data had limited success in retaining emotional information in Chinese. Additionally, resynthesizing speech through these codecs degrades the performance of speech emotion recognition (SER), particularly for emotions like sadness, depression, fear, and disgust. Human listening tests confirmed these findings. This work guides future speech technology developments to ensure new codecs maintain the integrity of emotional information in speech.

【5】 Integrating IP Broadcasting with Audio Tags- Workflow and Challenges
标题: 将IP广播与音频标签集成-工作流程和挑战
作者:Rhys Burchett-Vass,Arshdeep Singh,Gabriel Bibbó,Mark D. Plumbley
备注:Submitted to DCASE 2024 Workshop
链接:点击下载PDF文件
摘要:广播业越来越多地采用IP技术,从新闻采集到现场音乐活动,彻底改变了现场和预先录制的内容制作。IP广播允许以易于配置的方式传输音频和视频信号,与现代网络技术保持一致。这种向IP工作流的转变不仅在路由信号方面,而且在使用标准Web开发技术的工具集成方面,都具有更大的灵活性。一个可能的工具可以包括使用现场音频标记,这在内容制作中有许多用途。这些包括从自动隐藏字幕到识别场景中不需要的声音事件。在本文中,我们描述了将音频标记模型容器化到微服务中的过程,微服务是一个可以集成到多种不同网络设置中的小型隔离代码模块。我们的目标是开发一个模块化的,可访问的,灵活的工具,能够无缝部署到各种规模的广播工作流程,从小型制作到大型公司。所选择的音频标记模型的延迟及其对最终产品的有用性的影响周围的挑战进行了讨论。摘要:The broadcasting industry is increasingly adopting IP techniques, revolutionising both live and pre-recorded content production, from news gathering to live music events. IP broadcasting allows for the transport of audio and video signals in an easily configurable way, aligning with modern networking techniques. This shift towards an IP workflow allows for much greater flexibility, not only in routing signals but with the integration of tools using standard web development techniques. One possible tool could include the use of live audio tagging, which has a number of uses in the production of content. These include from automated closed captioning to identifying unwanted sound events within a scene. In this paper, we describe the process of containerising an audio tagging model into a microservice, a small segregated code module that can be integrated into a multitude of different network setups. The goal is to develop a modular, accessible, and flexible tool capable of seamless deployment into broadcasting workflows of all sizes, from small productions to large corporations. Challenges surrounding latency of the selected audio tagging model and its effect on the usefulness of the end product are discussed.

【6】 Can all variations within the unified mask-based beamformer framework achieve identical peak extraction performance?
标题: 统一的基于屏蔽的射束形成器框架内的所有变体能否实现相同的峰值提取性能?
作者:Atsuo Hiroe,Katsutoshi Itoyama,Kazuhiro Nakadai
备注:Submitted to EURASIP journal on Audio, Speech, and Music Processing
链接:点击下载PDF文件
摘要:本文研究了基于掩模的波束形成器(BF),它使用时频掩模估计目标声提取(TSE)的滤波器。虽然已经提出了多个基于掩模的BF,但对于目标提取性能的最佳BF尚未达成共识。以前,我们发现,最大信噪比和最小均方误差(MSE)BF可以实现相同的提取性能的理论上限的性能,每个BF包含一个不同的最佳掩模。然而,这些显着的研究结果留下了两个问题没有解决:只有两个BF被覆盖,不包括最小方差无失真响应BF;和理想缩放(IS)被用来理想地调整输出规模,这是不适用于现实场景。为了解决这些覆盖和缩放问题,本研究提出了一个统一的框架,基于掩码的BF包括两个过程:滤波器估计,可以覆盖所有BF和缩放适用于现实的情况下,通过采用掩码生成一个缩放参考。我们还提出了一种方法来枚举所有可能的BF,并得出12个变化。通过最小化目标和BF输出之间的MSE来获得这两个过程的最佳掩模。使用CHiME-4数据集的实验结果表明:1)所有12种变化都可以达到理论上的上限性能,2)基于掩模的缩放可以表现为IS。这些结果可以通过考虑掩模的实际参数计数来解释。这些发现有助于1)设计TSE系统,2)估计BF的提取性能,以及3)结合基于掩模的缩放来提高缩放精度。这些贡献也适用于基于独立分量分析的TSE方法,因为统一框架也涵盖了它们。摘要:This study investigates mask-based beamformers (BFs), which estimate filters for target sound extraction (TSE) using time-frequency masks. Although multiple mask-based BFs have been proposed, no consensus has been established on the best one for target-extracting performance. Previously, we found that maximum signal-to-noise ratio and minimum mean square error (MSE) BFs can achieve the same extraction performance as the theoretical upper-bound performance, with each BF containing a different optimal mask. However, these remarkable findings left two issues unsolved: only two BFs were covered, excluding the minimum variance distortionless response BF; and ideal scaling (IS) was employed to ideally adjust the output scale, which is not applicable to realistic scenarios. To address these coverage and scaling issues, this study proposes a unified framework for mask-based BFs comprising two processes: filter estimation that can cover all BFs and scaling applicable to realistic scenarios by employing a mask to generate a scaling reference. We also propose a methodology to enumerate all possible BFs and derive 12 variations. Optimal masks for both processes are obtained by minimizing the MSE between the target and BF output. The experimental results using the CHiME-4 dataset suggested that 1) all 12 variations can achieve the theoretical upper-bound performance, and 2) mask-based scaling can behave as IS. These results can be explained by considering the practical parameter count of the masks. These findings contribute to 1) designing a TSE system, 2) estimating the extraction performance of a BF, and 3) improving scaling accuracy combined with mask-based scaling. The contributions also apply to TSE methods based on independent component analysis, as the unified framework covers them too.

【7】 Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation Learning
标题: 说话人建模及其应用概述:从深度说话人表示学习的角度
作者:Shuai Wang,Zhengyang Chen,Kong Aik Lee,Yanmin Qian,Haizhou Li
链接:点击下载PDF文件
摘要:说话人个性信息是语音信号中最重要的信息之一。通过对这些信息进行全面而准确的建模,它可以用于各种智能语音应用,例如说话人识别,说话人日记,语音合成和目标说话人提取。在这篇文章中,我们的目标是,从一个独特的角度来看,发展历史,范式转变,以及在深度表示学习框架的背景下,扬声器建模技术的应用领域。这篇综述旨在为说话人建模领域的研究人员以及那些希望将说话人建模技术应用于特定下游任务的人提供明确的参考。摘要:Speaker individuality information is among the most critical elements within speech signals. By thoroughly and accurately modeling this information, it can be utilized in various intelligent speech applications, such as speaker recognition, speaker diarization, speech synthesis, and target speaker extraction. In this article, we aim to present, from a unique perspective, the developmental history, paradigm shifts, and application domains of speaker modeling technologies within the context of deep representation learning framework. This review is designed to provide a clear reference for researchers in the speaker modeling field, as well as for those who wish to apply speaker modeling techniques to specific downstream tasks.

【8】 Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
标题: 使用可控情绪强度实现现实情绪语音转换
作者:Tianhua Qi,Shiyan Wang,Cheng Lu,Yan Zhao,Yuan Zong,Wenming Zheng
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
摘要:真实情感语音转换(EVC)旨在增强转换音频的情感多样性,使合成的语音更加真实自然。为此,我们提出了情绪强度感知网络(EINet),动态调整语调和节奏,通过纳入可控的情绪强度。为了更好地捕捉情感强度的细微差别,我们超越了声学特征之间的距离测量。相反,情感评估器被用来精确地量化扬声器的情绪状态。通过使用强度映射器,获得强度伪标签,以弥合情感语音强度建模和运行时转换之间的差距。为了确保高的语音质量,同时保持可控性,情感渲染器用于在帧级平滑地将语言特征与操纵的情感特征相结合。此外,我们采用了持续时间预测,以促进自适应预测的节奏变化条件下指定的强度值。实验结果表明,与最先进的EVC方法相比,EINet在情感表达的自然性和多样性方面具有优越的性能。摘要:Realistic emotional voice conversion (EVC) aims to enhance emotional diversity of converted audios, making the synthesized voices more authentic and natural. To this end, we propose Emotional Intensity-aware Network (EINet), dynamically adjusting intonation and rhythm by incorporating controllable emotional intensity. To better capture nuances in emotional intensity, we go beyond mere distance measurements among acoustic features. Instead, an emotion evaluator is utilized to precisely quantify speaker's emotional state. By employing an intensity mapper, intensity pseudo-labels are obtained to bridge the gap between emotional speech intensity modeling and run-time conversion. To ensure high speech quality while retaining controllability, an emotion renderer is used for combining linguistic features smoothly with manipulated emotional features at frame level. Furthermore, we employ a duration predictor to facilitate adaptive prediction of rhythm changes condition on specifying intensity value. Experimental results show EINet's superior performance in naturalness and diversity of emotional expression compared to state-of-the-art EVC methods.

【9】 Multi-label audio classification with a noisy zero-shot teacher
标题: 具有噪音Zero-Shot教师的多标签音频分类
作者:Sebastian Braun,Hannes Gamper
备注:accepted at IEEE IWAENC 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的训练方案,使用自标签校正和数据增强方法,旨在处理嘈杂的标签,提高现实世界的准确性的复调音频内容检测任务。增强方法通过混合多个音频片段并加入它们的标签来减少标签噪声,同时与多个活动标签兼容。我们还表明,使用相同的预训练模型,通过自标记校正方法可以提高性能。最后,我们表明使用CLAP等强零次(zero-shot)模型为未标记数据生成标签并使用所提出的训练和标签增强方法改进结果是可行的。所得到的模型执行类似于CLAP,同时是一个有效的移动终端友好的架构,可以快速适应未标记的声音类。摘要:We propose a novel training scheme using self-label correction and data augmentation methods designed to deal with noisy labels and improve real-world accuracy on a polyphonic audio content detection task. The augmentation method reduces label noise by mixing multiple audio clips and joining their labels, while being compatible with multiple active labels. We additionally show that performance can be improved by a self-label correction method using the same pretrained model. Finally, we show that it is feasible to use a strong zero-shot model such as CLAP to generate labels for unlabeled data and improve the results using the proposed training and label enhancement methods. The resulting model performs similar to CLAP while being an efficient mobile device friendly architecture and can be quickly adapted to unlabeled sound classes.

【10】 dMel: Speech Tokenization made Simple
标题: dMel:语音代币化变得简单
作者:He Bai,Tatiana Likhomanenko,Ruixiang Zhang,Zijin Gu,Zakaria Aldeneh,Navdeep Jaitly
备注:under review
链接:点击下载PDF文件
摘要:大型语言模型通过利用大量文本数据的自我监督预训练,彻底改变了自然语言处理。受这一成功的启发,研究人员研究了复杂的语音标记化方法来离散连续语音信号,以便语言建模技术可以应用于语音数据。然而,现有的方法要么模型语义令牌,潜在地丢失声学信息,或者模型声学令牌,有丢失语义信息的风险。拥有多个令牌类型也会使架构复杂化,并需要额外的预训练。在这里,我们表明,离散梅尔滤波器组通道到离散强度箱产生一个简单的表示(dMel),比其他现有的语音标记化方法的性能更好。使用一个Transformer解码器的语音-文本模型,我们全面评估不同的语音标记化方法的语音识别(ASR),语音合成(TTS)。我们的研究结果表明,dMel在一个统一的框架内实现两个任务的高性能的有效性,为高效和有效的语音和文本的联合建模铺平了道路。摘要:Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have investigated complicated speech tokenization methods to discretize continuous speech signals so that language modeling techniques can be applied to speech data. However, existing approaches either model semantic tokens, potentially losing acoustic information, or model acoustic tokens, risking the loss of semantic information. Having multiple token types also complicates the architecture and requires additional pretraining. Here we show that discretizing mel-filterbank channels into discrete intensity bins produces a simple representation (dMel), that performs better than other existing speech tokenization methods. Using a transformer decoder-only architecture for speech-text modeling, we comprehensively evaluate different speech tokenization methods on speech recognition (ASR), speech synthesis (TTS). Our results demonstrate the effectiveness of dMel in achieving high performance on both tasks within a unified framework, paving the way for efficient and effective joint modeling of speech and text.

【11】 J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
标题: J-CHAT:用于口语对话语言建模的日本大型口语对话库
作者:Wataru Nakata,Kentaro Seki,Hitomi Yanaka,Yuki Saito,Shinnosuke Takamichi,Hiroshi Saruwatari
备注:8 pages, 6 figures
链接:点击下载PDF文件
摘要:口语对话在人类与人工智能的交互中起着至关重要的作用,因此需要面向对话的口语模型(SLM)。为了开发多功能的SLM,大规模和多样化的语音数据集是必不可少的。此外,为了确保高质量的语音生成,数据必须像野生数据一样是自发的,并且必须在去除噪声的情况下声学上干净。尽管迫切需要,但没有满足所有这些标准的开源语料库。本研究通过构建和发布一个大规模的口语对话语料库来解决这一差距,该语料库名为Japanese Corpus for Human-AI Talks(J-CHAT),可公开访问。此外,本文提出了一种独立于语言的语料库建设方法,并描述了使用J-CHAT训练的SLM对话生成实验。实验结果表明,该方法收集的多个领域的数据提高了对话生成的自然性和意义。摘要:Spoken dialogue plays a crucial role in human-AI interactions, necessitating dialogue-oriented spoken language models (SLMs). To develop versatile SLMs, large-scale and diverse speech datasets are essential. Additionally, to ensure hiqh-quality speech generation, the data must be spontaneous like in-wild data and must be acoustically clean with noise removed. Despite the critical need, no open-source corpus meeting all these criteria has been available. This study addresses this gap by constructing and releasing a large-scale spoken dialogue corpus, named Japanese Corpus for Human-AI Talks (J-CHAT), which is publicly accessible. Furthermore, this paper presents a language-independent method for corpus construction and describes experiments on dialogue generation using SLMs trained on J-CHAT. Experimental results indicate that the collected data from multiple domains by our method improve the naturalness and meaningfulness of dialogue generation.

【12】 Computer Audition: From Task-Specific Machine Learning to Foundation Models
标题: 计算机试镜:从特定任务的机器学习到基础模型
作者:Andreas Triantafyllopoulos,Iosif Tsangko,Alexander Gebhard,Annamaria Mesaros,Tuomas Virtanen,Björn Schuller
链接:点击下载PDF文件
摘要:基础模型(FM)越来越多地引领着计算机听觉范围内的各种任务的最新进展-使用机器来理解声音。与传统管道相比,它们具有几个优势:其中包括将多个任务整合到单个模型中的能力,利用其他模式的知识的选择,以及与人类用户的现成交互。自然,这些承诺在音频社区中引起了极大的兴奋,并导致了一波早期尝试构建新的通用音频基础模型的浪潮。在目前的贡献中,我们给出了一个概述的计算音频分析,因为它从传统的管道过渡到听觉基础模型。我们的工作突出了支撑这些模型的关键操作原则,并展示了它们如何适应音频社区以前单独处理的多个任务。摘要:Foundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition -- the use of machines to understand sounds. They feature several advantages over traditional pipelines: among others, the ability to consolidate multiple tasks in a single model, the option to leverage knowledge from other modalities, and the readily-available interaction with human users. Naturally, these promises have created substantial excitement in the audio community, and have led to a wave of early attempts to build new, general-purpose foundation models for audio. In the present contribution, we give an overview of computational audio analysis as it transitions from traditional pipelines towards auditory foundation models. Our work highlights the key operating principles that underpin those models, and showcases how they can accommodate multiple tasks that the audio community previously tackled separately.

【13】 Annealed Multiple Choice Learning: Overcoming limitations of Winner-takes-all with annealing
标题: Annealed Multiple Choice Learning:通过Annealed克服赢家通吃的局限性
作者:David Perera,Victor Letzelter,Théo Mariotte,Adrien Cortés,Mickael Chen,Slim Essid,Gaël Richard
链接:点击下载PDF文件
摘要:我们介绍了退火多选择学习(aMCL),它结合了模拟退火和MCL。MCL是一个学习框架,通过预测一小部分合理的假设来处理模糊的任务。这些假设使用Winner-takes-all(WTA)方案进行训练,该方案促进了预测的多样性。然而,由于WTA的贪婪性质,该方案可能会收敛到任意次优的局部最小值。我们使用退火克服了这一限制,这增强了训练过程中对假设空间的探索。我们利用统计物理学和信息论的见解来提供模型训练轨迹的详细描述。此外,我们验证了我们的算法通过广泛的实验合成数据集,标准的UCI基准,语音分离。摘要:We introduce Annealed Multiple Choice Learning (aMCL) which combines simulated annealing with MCL. MCL is a learning framework handling ambiguous tasks by predicting a small set of plausible hypotheses. These hypotheses are trained using the Winner-takes-all (WTA) scheme, which promotes the diversity of the predictions. However, this scheme may converge toward an arbitrarily suboptimal local minimum, due to the greedy nature of WTA. We overcome this limitation using annealing, which enhances the exploration of the hypothesis space during training. We leverage insights from statistical physics and information theory to provide a detailed description of the model training trajectory. Additionally, we validate our algorithm by extensive experiments on synthetic datasets, on the standard UCI benchmark, and on speech separation.

【14】 SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
标题: SELM:增强域外场景的语音情感识别
作者:Hazim Bukhari,Soham Deshmukh,Hira Dhamyal,Bhiksha Raj,Rita Singh
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:语音情感识别(SER)传统上被制定为一个分类任务。然而,情绪通常是一个频谱,其分布因情况而异,导致域外(OOD)性能差。我们从自动语音识别(ASR)的统计公式中获得灵感,并将SER任务制定为生成最可能的文本标记序列来推断情感。该公式将SER分解为预测由语言模型预测加权的声学模型特征。作为这种方法的一个实例,我们提出了SELM,一个音频条件的语言模型SER,预测不同的情感观点。我们训练SELM策展语音情感语料库和测试三个OOD数据集(RAVDESS,CREMAD,IEMOCAP)没有在训练中使用。SELM在最先进的基线上实现了显着的改进,RAVDESS和CREMA-D的相对精度分别提高了17%和7%。此外,SELM可以通过使用一些注释示例的Few-Shot学习来进一步提高其性能。结果突出了我们的SER配方的有效性,特别是在OOD场景中提高性能。摘要:Speech Emotion Recognition (SER) has been traditionally formulated as a classification task. However, emotions are generally a spectrum whose distribution varies from situation to situation leading to poor Out-of-Domain (OOD) performance. We take inspiration from statistical formulation of Automatic Speech Recognition (ASR) and formulate the SER task as generating the most likely sequence of text tokens to infer emotion. The formulation breaks SER into predicting acoustic model features weighted by language model prediction. As an instance of this approach, we present SELM, an audio-conditioned language model for SER that predicts different emotion views. We train SELM on curated speech emotion corpus and test it on three OOD datasets (RAVDESS, CREMAD, IEMOCAP) not used in training. SELM achieves significant improvements over the state-of-the-art baselines, with 17% and 7% relative accuracy gains for RAVDESS and CREMA-D, respectively. Moreover, SELM can further boost its performance by Few-Shot Learning using a few annotated examples. The results highlight the effectiveness of our SER formulation, especially to improve performance in OOD scenarios.

【15】 Explainability Paths for Sustained Artistic Practice with AI
标题: 人工智能持续艺术实践的解释途径
作者:Austin Tecks,Thomas Peschlow,Gabriel Vigliensoni
备注:In Proceedings of Explainable AI for the Arts Workshop 2024 (XAIxArts 2024) arXiv:2406.14485
链接:点击下载PDF文件
摘要:人工智能驱动的生成音频的开发反映了更广泛的人工智能趋势,通常优先考虑即时可访问性,而牺牲了可解释性。因此,将这些工具纳入持续的艺术实践仍然是一个重大挑战。在本文中,我们探索了几条提高可解释性的途径,主要是从我们在训练和实现生成音频模型方面的研究-创作实践中得出的。作为提高可解释性的实际规定,我们强调了培训材料的人类代理,小规模数据集的可行性,迭代创造过程的便利化,以及交互式机器学习作为映射工具的集成。重要的是,这些步骤旨在增强人类在生成AI系统上的代理能力,不仅在模型推理过程中,而且在管理和预处理训练数据以及模型训练阶段。摘要:The development of AI-driven generative audio mirrors broader AI trends, often prioritizing immediate accessibility at the expense of explainability. Consequently, integrating such tools into sustained artistic practice remains a significant challenge. In this paper, we explore several paths to improve explainability, drawing primarily from our research-creation practice in training and implementing generative audio models. As practical provisions for improved explainability, we highlight human agency over training materials, the viability of small-scale datasets, the facilitation of the iterative creative process, and the integration of interactive machine learning as a mapping tool. Importantly, these steps aim to enhance human agency over generative AI systems not only during model inference, but also when curating and preprocessing training data as well as during the training phase of models.

【16】 MusiConGen: Rhythm and Chord Control for Transformer-Based Text-to-Music Generation
标题: MusiConGen:基于转换器的文本到音乐生成的节奏和和弦控制
作者:Yun-Han Lan,Wen-Yi Hsiao,Hao-Chung Cheng,Yi-Hsuan Yang
备注:Accepted by the 25th International Society for Music Information Retrieval (ISMIR)
链接:点击下载PDF文件
摘要:现有的文本到音乐模型可以产生具有极大多样性的高质量音频。然而,单独的文本提示不能精确地控制所生成的音乐的时间音乐特征,诸如和弦和节奏。为了应对这一挑战,我们引入了MusiConGen,这是一种基于时间条件的基于transformer的文本到音乐模型,它建立在预训练的MusicGen框架之上。我们的创新在于为消费级GPU量身定制的高效微调机制,该机制集成了自动提取的节奏和和弦作为条件信号。在推断期间,条件可以是从参考音频信号提取的音乐特征,或者是用户定义的符号和弦序列、BPM和文本提示。我们对两个数据集(一个来自提取的特征,另一个来自用户创建的输入)的性能评估表明,MusiConGen可以生成与指定条件一致的真实背景音乐。我们开源代码和模型检查点,并在线提供音频示例,www.example.com。摘要:Existing text-to-music models can produce high-quality audio with great diversity. However, textual prompts alone cannot precisely control temporal musical features such as chords and rhythm of the generated music. To address this challenge, we introduce MusiConGen, a temporally-conditioned Transformer-based text-to-music model that builds upon the pretrained MusicGen framework. Our innovation lies in an efficient finetuning mechanism, tailored for consumer-grade GPUs, that integrates automatically-extracted rhythm and chords as the condition signal. During inference, the condition can either be musical features extracted from a reference audio signal, or be user-defined symbolic chord sequence, BPM, and textual prompts. Our performance evaluation on two datasets -- one derived from extracted features and the other from user-created inputs -- demonstrates that MusiConGen can generate realistic backing track music that aligns well with the specified conditions. We open-source the code and model checkpoints, and provide audio examples online, https: musicongen.github.io musicongen_demo .

【17】 Conversational Rubert for Detecting Competitive Interruptions in ASR-Transcribed Dialogues
标题: 对话鲁BERT用于检测SVR转录对话中的竞争中断
作者:Dmitrii Galimzianov,Viacheslav Vyshegorodtsev
Journal-ref:Computer Science & Information Technology (CS & IT), ISSN : 2231 - 5403, Volume 14, Number 13, June 2024
链接:点击下载PDF文件
摘要:对话中的中断发生在听者在当前说话者结束讲话之前开始讲话时。打断可以大致分为两类:合作型(当听者想要支持说话者时)和竞争型(当听者试图违背说话者的意愿控制谈话时)。自动分类中断的系统可以用于呼叫中心,特别是在客户满意度监控和代理监控的任务中。在这项研究中,我们开发了一个基于文本的中断分类模型,准备一个内部数据集组成的ASR转录的客户支持电话对话在俄罗斯。我们在我们的数据集上微调了对话式RuBERT,并优化了超参数,模型表现良好。通过进一步的改进,该模型可以应用于自动监控系统。摘要:Interruption in a dialogue occurs when the listener begins their speech before the current speaker finishes speaking. Interruptions can be broadly divided into two groups: cooperative (when the listener wants to support the speaker), and competitive (when the listener tries to take control of the conversation against the speaker's will). A system that automatically classifies interruptions can be used in call centers, specifically in the tasks of customer satisfaction monitoring and agent monitoring. In this study, we developed a text-based interruption classification model by preparing an in-house dataset consisting of ASR-transcribed customer support telephone dialogues in Russian. We fine-tuned Conversational RuBERT on our dataset and optimized hyperparameters, and the model performed well. With further improvements, the proposed model can be applied to automatic monitoring systems.

【18】 Composer's Assistant 2: Interactive Multi-Track MIDI Infilling with Fine-Grained User Control
标题: 作曲家助理2:交互式多轨收件箱填充细粒度用户控制
作者:Martin E. Malandro
备注:8 pages, 6 figures, 2 tables. To be published in ISMIR 2024
链接:点击下载PDF文件
摘要:我们介绍作曲家的助手2,在收割机数字音频工作站的交互式人机合成系统。我们的工作升级了Composer's Assistant系统(在音轨级别执行符号音乐的多音轨填充),并提供了广泛的新控件,让用户可以对系统的输出进行细粒度控制。在这项工作中引入的控制包括两种类型的节奏调节控制,水平和垂直注意开始密度控制,几种类型的音高控制,和节奏感兴趣的控制。我们训练一个类似T5的Transformer模型来实现这些控制,并作为我们系统的骨干。有了这些控制,我们实现了显着的改善,在原来的系统的客观指标。我们还研究了我们的模型对控件含义的理解程度,并且我们进行了一项听力研究,没有发现真实音乐和与我们的系统以共同创作的方式创作的音乐之间存在显着差异。我们发布了完整的系统,包括源代码,预训练模型和REAPER脚本。摘要:We introduce Composer's Assistant 2, a system for interactive human-computer composition in the REAPER digital audio workstation. Our work upgrades the Composer's Assistant system (which performs multi-track infilling of symbolic music at the track-measure level) with a wide range of new controls to give users fine-grained control over the system's outputs. Controls introduced in this work include two types of rhythmic conditioning controls, horizontal and vertical note onset density controls, several types of pitch controls, and a rhythmic interest control. We train a T5-like transformer model to implement these controls and to serve as the backbone of our system. With these controls, we achieve a dramatic improvement in objective metrics over the original system. We also study how well our model understands the meaning of our controls, and we conduct a listening study that does not find a significant difference between real music and music composed in a co-creative fashion with our system. We release our complete system, consisting of source code, pretrained models, and REAPER scripts.


机器翻译,仅供参考