本文经arXiv每日学术速递授权转载
链接:https://arxiv.org/abs/2411.18497
摘要:在有监督的环境中训练语音分离模型提出了一个置换问题:找到模型预测和地面真实分离信号之间的最佳分配。这个固有的模糊任务通常使用Permutation Invariant Training(PIT)来解决。在本文中,我们考虑使用多项选择学习(MCL)框架,该框架最初是为了解决模糊任务而引入的。我们在流行的WSJ 0-mix和LibriMix基准测试上通过实验证明,MCL与PIT的性能相匹配,同时在计算上具有优势。这为一个有前途的研究方向打开了大门,因为MCL可以自然地扩展到处理可变数量的扬声器,或者在无监督环境中处理语音分离。
摘要:Training speech separation models in the supervised setting raises apermutation problem: finding the best assignation between the model predictionsand the ground truth separated signals. This inherently ambiguous task iscustomarily solved using Permutation Invariant Training (PIT). In this article,we instead consider using the Multiple Choice Learning (MCL) framework, whichwas originally introduced to tackle ambiguous tasks. We demonstrateexperimentally on the popular WSJ0-mix and LibriMix benchmarks that MCL matchesthe performances of PIT, while being computationally advantageous. This opensthe door to a promising research direction, as MCL can be naturally extended tohandle a variable number of speakers, or to tackle speech separation in theunsupervised setting.
标题:带噪音增强的连续自回归模型避免误差累积
链接:https://arxiv.org/abs/2411.18447
备注:Accepted to NeurIPS 2024 - Audio Imagination Workshop
摘要:自回归模型通常应用于离散令牌序列,但最近的研究表明,以自回归方式生成连续嵌入序列也是可行的。然而,这样的连续自回归模型(CAM)可能会遭受由于推断期间的错误累积而导致的扩展序列的生成质量下降。我们引入了一种新的方法来解决这个问题,在训练过程中将随机噪声注入输入嵌入。这个过程使模型在推理时对不同的误差水平具有鲁棒性。我们通过引入低水平噪声的推理过程进一步减少了误差积累。音乐音频生成的实验表明,CAM大大优于现有的自回归和非自回归方法,同时保持音频质量超过扩展序列。这项工作为在纯自回归环境中生成连续嵌入铺平了道路,为实时和交互式生成应用程序开辟了新的可能性。
摘要:Autoregressive models are typically applied to sequences of discrete tokens,but recent research indicates that generating sequences of continuousembeddings in an autoregressive manner is also feasible. However, suchContinuous Autoregressive Models (CAMs) can suffer from a decline in generationquality over extended sequences due to error accumulation during inference. Weintroduce a novel method to address this issue by injecting random noise intothe input embeddings during training. This procedure makes the model robustagainst varying error levels at inference. We further reduce error accumulationthrough an inference procedure that introduces low-level noise. Experiments onmusical audio generation show that CAM substantially outperforms existingautoregressive and non-autoregressive approaches while preserving audio qualityover extended sequences. This work paves the way for generating continuousembeddings in a purely autoregressive setting, opening new possibilities forreal-time and interactive generative applications.
标题:如何学习一门新语言?自我监督学习模型的有效解决方案低资源场景下的隐形语言适应
链接:https://arxiv.org/abs/2411.18217
摘要:语音自监督学习(SSL)模型的使用在自动语音识别(ASR)上取得了令人印象深刻的性能。然而,在低资源语言ASR中,他们遇到了预训练语言和低资源语言之间的域不匹配问题。典型的解决方案,如微调SSL模型,遭受高计算成本,而使用冻结SSL模型作为特征提取器,性能较差。为了处理这些问题,我们扩展了传统的高效的微调方案的基础上的适配器。我们添加了一个额外的中间适配来预热适配器和下游模型初始化。值得注意的是,我们只更新总模型参数的1-5%来实现自适应。ML-SUPERB数据集上的实验结果表明,我们的解决方案优于传统的高效微调。在适应看不见的语言时,它在字符/音素错误率方面实现了高达28%的相对改善。
摘要:The utilization of speech Self-Supervised Learning (SSL) models achievesimpressive performance on Automatic Speech Recognition (ASR). However, inlow-resource language ASR, they encounter the domain mismatch problem betweenpre-trained and low-resource languages. Typical solutions like fine-tuning theSSL model suffer from high computation costs while using frozen SSL models asfeature extractors comes with poor performance. To handle these issues, weextend a conventional efficient fine-tuning scheme based on the adapter. We addan extra intermediate adaptation to warm up the adapter and downstream modelinitialization. Remarkably, we update only 1-5% of the total model parametersto achieve the adaptation. Experimental results on the ML-SUPERB dataset showthat our solution outperforms conventional efficient fine-tuning. It achievesup to a 28% relative improvement in the Character/Phoneme error rate whenadapting to unseen languages.
标题:MSA-ASB:使用冻结ASB模型的高效多语言说话者归因
链接:https://arxiv.org/abs/2411.18152
摘要:说话人属性自动语音识别(SA-ASR)的目标是在转录语音的同时准确地将转录本分配给相应的说话人。现有方法通常依赖于复杂的模块化系统或需要对关节模块进行大量微调,从而限制了其适应性和总体效率。本文介绍了一种新的方法,利用冻结的多语言ASR模型,将说话人属性到transmittance,只使用标准的单语ASR数据集。我们的方法涉及训练扬声器模块,以基于弱标签预测扬声器嵌入,而不需要额外的ASR模型修改。尽管只使用非重叠的单语数据进行训练,但我们的方法可以有效地在不同的多语言数据集中提取说话人属性,包括那些具有重叠语音的数据集。实验结果表明,与强基线相比,具有竞争力的性能,突出了模型的鲁棒性和实际应用的潜力。
摘要:Speaker-attributed automatic speech recognition (SA-ASR) aims to transcribespeech while assigning transcripts to the corresponding speakers accurately.Existing methods often rely on complex modular systems or require extensivefine-tuning of joint modules, limiting their adaptability and generalefficiency. This paper introduces a novel approach, leveraging a frozenmultilingual ASR model to incorporate speaker attribution into thetranscriptions, using only standard monolingual ASR datasets. Our methodinvolves training a speaker module to predict speaker embeddings based on weaklabels without requiring additional ASR model modifications. Despite beingtrained exclusively with non-overlapping monolingual data, our approacheffectively extracts speaker attributes across diverse multilingual datasets,including those with overlapping speech. Experimental results demonstratecompetitive performance compared to strong baselines, highlighting the model'srobustness and potential for practical applications.
标题:多语言自动语音识别的离散表示和自增强表示的融合
链接:https://arxiv.org/abs/2411.18107
备注:SLT 2024
摘要:自监督学习(SSL)模型在各种语音处理任务中表现出卓越的能力。连续SSL表示是有效的,但遭受高计算和存储需求。另一方面,离散的SSL表示,虽然性能下降,降低了传输和存储成本,并通过重复数据删除和子字建模提高输入序列的效率。为了提高ASR的离散表示的性能,我们引入了一种新的融合机制,集成了两个离散表示。融合机制保留了离散表示的所有优点,同时通过集成互补信息来增强模型的性能。此外,我们还探索了“自增强”离散表示,它将变换应用于单个连续SSL表示,消除了融合机制对多个SSL模型的依赖,并进一步降低了其推理成本。在LibriSpeech和ML-SUPERB等基准测试上的实验结果表明,与非融合基线相比,相对字符错误率分别提高了19%和24%,验证了本文方法的有效性。
摘要:Self-supervised learning (SSL) models have shown exceptional capabilitiesacross various speech-processing tasks. Continuous SSL representations areeffective but suffer from high computational and storage demands. On the otherhand, discrete SSL representations, although with degraded performance, reducetransmission and storage costs, and improve input sequence efficiency throughde-duplication and subword-modeling. To boost the performance of discreterepresentations for ASR, we introduce a novel fusion mechanism that integratestwo discrete representations. The fusion mechanism preserves all the benefitsof discrete representation while enhancing the model's performance byintegrating complementary information. Additionally, we explore"self-augmented'' discrete representations, which apply transformations to asingle continuous SSL representation, eliminating the fusion mechanism'sdependency on multiple SSL models and further decreasing its inference costs.Experimental results on benchmarks, including LibriSpeech and ML-SUPERB,indicate up to 19% and 24% relative character error rate improvement comparedwith the non-fusion baseline, validating the effectiveness of our proposedmethods.
标题:Music 2 Fail:将音乐转移到失败的录音机风格
链接:https://arxiv.org/abs/2411.18075
备注:Accepted by APSIPA 2024
摘要:音乐风格转换的目的是将一种乐器演奏的音乐转换为另一种乐器演奏的音乐,同时保持音乐内容不变。在本文中,我们研究了另一种风格转移的情况下,所谓的“失败的音乐风格转移”。与通常的音乐风格转移不同,在通常的音乐风格转移中,内容保持不变,只有乐器的特征被改变,这种情况试图将音乐从源乐器转移到目标乐器,这是故意偏离音高执行的。我们的工作试图将正常播放的音乐转换为非音高录音机音乐,我们称之为“失败式录音机”,并研究转换的结果。为了完成这项工作,我们还提出了一个失败式记录器的数据集,称为“FR109数据集”。这样的实验在一个更有表现力的环境中探索音乐风格转移任务,因为生成的音频听起来应该像一个“偏离音高的录音机”,同时保持一定程度的自然性。
摘要:The goal of music style transfer is to convert a music performance by oneinstrument into another while keeping the musical contents unchanged. In thispaper, we investigate another style transfer scenario called ``failed-musicstyle transfer''. Unlike the usual music style transfer where the contentremains the same and only the instrumental characteristics are changed, thisscenario seeks to transfer the music from the source instrument to the targetinstrument which is deliberately performed off-pitch. Our work attempts totransfer normally played music into off-pitch recorder music, which we call``failed-style recorder'', and study the results of the conversion. To carryout this work, we have also proposed a dataset of failed-style recorders forthis task, called ``FR109 Dataset''. Such an experiment explores the musicstyle transfer task in a more expressive setting, as the generated audio shouldsound like an ``off-pitch recorder'' while maintaining a certain degree ofnaturalness.
标题:可穿戴智能喉可使患有构音障碍的中风患者能够自然说话
链接:https://arxiv.org/abs/2411.18266
备注:5 figures, 45 references
摘要:可穿戴无声语音系统具有恢复言语障碍患者沟通的巨大潜力。然而,无缝、连贯的语音仍然难以实现,临床疗效也尚未得到证实。在这里,我们提出了一个人工智能驱动的智能喉咙(IT)系统,该系统将喉咙肌肉振动和颈动脉脉搏信号传感器与大型语言模型(LLM)处理集成在一起,以实现流畅,情感表达的沟通。该系统利用超灵敏的纺织应变传感器来捕获来自颈部区域的高质量信号,并支持令牌级处理,以进行实时、连续的语音解码,从而实现无缝、无延迟的通信。在对五名患有构音障碍的中风患者进行的测试中,IT的LLM代理智能地纠正了标记错误,并丰富了情绪和逻辑连贯性,实现了低错误率(4.2%的单词错误率,2.9%的句子错误率)和55%的用户满意度。这项工作为构音障碍患者建立了一个便携式、直观的沟通平台,有可能广泛应用于不同的神经系统疾病和多语言支持系统。
摘要:Wearable silent speech systems hold significant potential for restoringcommunication in patients with speech impairments. However, seamless, coherentspeech remains elusive, and clinical efficacy is still unproven. Here, wepresent an AI-driven intelligent throat (IT) system that integrates throatmuscle vibrations and carotid pulse signal sensors with large language model(LLM) processing to enable fluent, emotionally expressive communication. Thesystem utilizes ultrasensitive textile strain sensors to capture high-qualitysignals from the neck area and supports token-level processing for real-time,continuous speech decoding, enabling seamless, delay-free communication. Intests with five stroke patients with dysarthria, IT's LLM agents intelligentlycorrected token errors and enriched sentence-level emotional and logicalcoherence, achieving low error rates (4.2% word error rate, 2.9% sentence errorrate) and a 55% increase in user satisfaction. This work establishes aportable, intuitive communication platform for patients with dysarthria withthe potential to be applied broadly across different neurological conditionsand in multi-language support systems.
标题:迈向改进的客观感知音频质量评估--第1部分:一种新型的数据驱动认知模型
链接:https://arxiv.org/abs/2411.18222
备注:Accepter for publication in in IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:高效的音频质量评估对于简化音频编解码器开发至关重要。随着时间的推移,客观评估工具已经开发出来,可以通过算法从主观评估中预测质量评级,这是质量判断的黄金标准。这些工具中的许多使用感知听觉模型来提取音频特征,这些音频特征被映射到使用机器学习算法和主观分数作为训练数据的基本音频质量分数预测。然而,现有的工具在质量预测中难以泛化,特别是当面对未知的信号和失真类型时。这在使用非波形保持参数技术编码的信号的存在下特别明显。为了应对这些挑战,这两部分工作提出了对音频质量感知评估(PEAQ-ITU-R BS. 1387 -1)建议的扩展。第1部分着重于提高泛化能力,而第2部分则针对音频编码中的精确空间音频质量测量。 为了提高预测的泛化能力,本文(第1部分)介绍了一种新的机器学习方法,该方法使用主观数据对音频质量感知的认知方面进行建模。该方法通过自适应加权不同的失真度量来模拟听觉失真的感知严重程度。权重使用捕获失真显著性和认知效应之间的关系的交互成本函数来确定。与其他机器学习方法和已建立的工具相比,所提出的架构在以前看不见的主观质量分数的大型数据库上实现了更高的预测精度。感知激励模型为通用机器学习算法提供了更易于管理的替代方案,允许对多维质量测量进行潜在的扩展和改进,而无需完全重新训练。
摘要:Efficient audio quality assessment is vital for streamlining audio codecdevelopment. Objective assessment tools have been developed over time toalgorithmically predict quality ratings from subjective assessments, the goldstandard for quality judgment. Many of these tools use perceptual auditorymodels to extract audio features that are mapped to a basic audio quality scoreprediction using machine learning algorithms and subjective scores as trainingdata. However, existing tools struggle with generalization in qualityprediction, especially when faced with unknown signal and distortion types.This is particularly evident in the presence of signals coded usingnon-waveform-preserving parametric techniques. Addressing these challenges,this two-part work proposes extensions to the Perceptual Evaluation of AudioQuality (PEAQ - ITU-R BS.1387-1) recommendation. Part 1 focuses on increasinggeneralization, while Part 2 targets accurate spatial audio quality measurementin audio coding. To enhance prediction generalization, this paper (Part 1) introduces a novelmachine learning approach that uses subjective data to model cognitive aspectsof audio quality perception. The proposed method models the perceived severityof audible distortions by adaptively weighting different distortion metrics.The weights are determined using an interaction cost function that capturesrelationships between distortion salience and cognitive effects. Compared toother machine learning methods and established tools, the proposed architectureachieves higher prediction accuracy on large databases of previously unseensubjective quality scores. The perceptually-motivated model offers a moremanageable alternative to general-purpose machine learning algorithms, allowingpotential extensions and improvements to multi-dimensional quality measurementwithout complete retraining.
标题:SALMON-omni:一种用于双环语音理解和生成的无编解码器LLM
链接:https://arxiv.org/abs/2411.18138
备注:Technical report
摘要:全双工多模态大型语言模型(LLM)为解决不同的语音理解和生成任务提供了统一的框架,从而实现了更自然、更无缝的人机对话。与传统的模块化会话AI系统不同,它将语音识别,理解和文本到语音生成分离为不同的组件,多模式LLM作为单一的端到端模型运行。这种简化的设计消除了跨组件的错误传播,并充分利用了嵌入在输入语音信号中的丰富的非语言信息。我们介绍SALMONN-omni,一种无编解码器的全双工语音理解和生成模型,能够在说话时同时收听自己生成的语音和背景声音。为了支持这种能力,我们提出了一种新的双工口语对话框架,结合了一个“思考”的机制,有利于异步文本和语音生成依赖于嵌入,而不是编解码器(量化的语音和音频令牌)。实验结果表明,SALMONN-omni的通用性在广泛的流语音任务,包括语音识别,语音增强,和口语问答。此外,SALMONN-omni在管理轮对、打断和回声消除场景方面表现出色,从而确立了其作为全双工对话式AI系统的强大原型的潜力。据我们所知,SALMONN-omni是同类产品中第一个无编解码器的型号。一份完整的技术报告以及模型检查点将很快发布。
摘要:Full-duplex multimodal large language models (LLMs) provide a unifiedframework for addressing diverse speech understanding and generation tasks,enabling more natural and seamless human-machine conversations. Unliketraditional modularised conversational AI systems, which separate speechrecognition, understanding, and text-to-speech generation into distinctcomponents, multimodal LLMs operate as single end-to-end models. Thisstreamlined design eliminates error propagation across components and fullyleverages the rich non-verbal information embedded in input speech signals. Weintroduce SALMONN-omni, a codec-free, full-duplex speech understanding andgeneration model capable of simultaneously listening to its own generatedspeech and background sounds while speaking. To support this capability, wepropose a novel duplex spoken dialogue framework incorporating a ``thinking''mechanism that facilitates asynchronous text and speech generation relying onembeddings instead of codecs (quantized speech and audio tokens). Experimentalresults demonstrate SALMONN-omni's versatility across a broad range ofstreaming speech tasks, including speech recognition, speech enhancement, andspoken question answering. Additionally, SALMONN-omni excels at managingturn-taking, barge-in, and echo cancellation scenarios, establishing itspotential as a robust prototype for full-duplex conversational AI systems. Tothe best of our knowledge, SALMONN-omni is the first codec-free model of itskind. A full technical report along with model checkpoints will be releasedsoon.
标题:JPPO:加速大型语言模型服务的联合动力和即时优化
链接:https://arxiv.org/abs/2411.18010
摘要:大型语言模型(LLM)在各种任务中表现出卓越的能力,导致它们在无线网络中越来越多地部署,以提供各种各样的用户服务。然而,越来越长的提示设置突出了计算资源需求和巨大的通信负载的关键问题。为了应对这一挑战,我们提出了联合功率和即时优化(JPPO),这是一个将基于小语言模型(SLM)的即时压缩与无线功率分配优化相结合的框架。通过在用户设备上部署SLM以进行快速压缩,并采用深度强化学习来联合优化压缩比和传输功率,JPPO有效地平衡了服务质量与资源效率。实验结果表明,我们的框架实现了高服务保真度和低误码率,同时优化无线LLM服务的功率使用。该系统将响应时间缩短了约17%,改进程度取决于原始提示的长度。
摘要:Large Language Models (LLMs) have demonstrated remarkable capabilities invarious tasks, leading to their increasing deployment in wireless networks fora wide variety of user services. However, the growing longer prompt settinghighlights the crucial issue of computational resource demands and hugecommunication load. To address this challenge, we propose Joint Power andPrompt Optimization (JPPO), a framework that combines Small Language Model(SLM)-based prompt compression with wireless power allocation optimization. Bydeploying SLM at user devices for prompt compression and employing DeepReinforcement Learning for joint optimization of compression ratio andtransmission power, JPPO effectively balances service quality with resourceefficiency. Experimental results demonstrate that our framework achieves highservice fidelity and low bit error rates while optimizing power usage inwireless LLM services. The system reduces response time by about 17%, with theimprovement varying based on the length of the original prompt.
标题:开放集Raga分类的新型类发现
链接:https://arxiv.org/abs/2411.18611
备注:Under Review at ICASSP-25
摘要:印度艺术音乐(IAM)中的Raga分类任务受到标记数据集有限可用性的限制,导致许多Raga在机器学习模型的训练过程中无法表示。传统的Raga分类方法依赖于监督学习,并假设要通过Raga分类模型进行分类的测试音频必须在训练数据中表示,这限制了它们在真实世界场景中的有效性,其中可能会出现新的,看不见的Raga。为了解决这个问题,我们提出了一种基于新类发现(NCD)的方法来检测和分类以前看不见的Ragas。我们的方法利用以监督方式训练的特征提取器来生成嵌入,然后在对比学习框架内进行自监督训练,从而能够识别以前看不见的Raga类。结果表明,所提出的方法可以准确地检测与这些新型Ragas对应的音频样本,为利用在线可用的大量未标记音乐数据提供了一个鲁棒的解决方案。这种方法减少了手动标记的需要,同时扩展了已识别的Ragas和音乐信息检索(MIR)中的其他音乐数据的曲目。
摘要:The task of Raga classification in Indian Art Music (IAM) is constrained bythe limited availability of labeled datasets, resulting in many Ragas beingunrepresented during the training of machine learning models. Traditional Ragaclassification methods rely on supervised learning, and assume that for a testaudio to be classified by a Raga classification model, it must have beenrepresented in the training data, which limits their effectiveness inreal-world scenarios where novel, unseen Ragas may appear. To address thislimitation, we propose a method based on Novel Class Discovery (NCD) to detectand classify previously unseen Ragas. Our approach utilizes a feature extractortrained in a supervised manner to generate embeddings, which are then employedwithin a contrastive learning framework for self-supervised training, enablingthe identification of previously unseen Raga classes. The results demonstratethat the proposed method can accurately detect audio samples corresponding tothese novel Ragas, offering a robust solution for utilizing the vast amount ofunlabeled music data available online. This approach reduces the need formanual labeling while expanding the repertoire of recognized Ragas, and othermusic data in Music Information Retrieval (MIR).
标题:可穿戴智能喉可使患有构音障碍的中风患者能够自然说话
链接:https://arxiv.org/abs/2411.18266
备注:5 figures, 45 references
摘要:可穿戴无声语音系统具有恢复言语障碍患者沟通的巨大潜力。然而,无缝、连贯的语音仍然难以实现,临床疗效也尚未得到证实。在这里,我们提出了一个人工智能驱动的智能喉咙(IT)系统,该系统将喉咙肌肉振动和颈动脉脉搏信号传感器与大型语言模型(LLM)处理集成在一起,以实现流畅,情感表达的沟通。该系统利用超灵敏的纺织应变传感器来捕获来自颈部区域的高质量信号,并支持令牌级处理,以进行实时、连续的语音解码,从而实现无缝、无延迟的通信。在对五名患有构音障碍的中风患者进行的测试中,IT的LLM代理智能地纠正了标记错误,并丰富了情绪和逻辑连贯性,实现了低错误率(4.2%的单词错误率,2.9%的句子错误率)和55%的用户满意度。这项工作为构音障碍患者建立了一个便携式,直观的沟通平台,有可能广泛应用于不同的神经系统疾病和多语言支持系统。
摘要:Wearable silent speech systems hold significant potential for restoringcommunication in patients with speech impairments. However, seamless, coherentspeech remains elusive, and clinical efficacy is still unproven. Here, wepresent an AI-driven intelligent throat (IT) system that integrates throatmuscle vibrations and carotid pulse signal sensors with large language model(LLM) processing to enable fluent, emotionally expressive communication. Thesystem utilizes ultrasensitive textile strain sensors to capture high-qualitysignals from the neck area and supports token-level processing for real-time,continuous speech decoding, enabling seamless, delay-free communication. Intests with five stroke patients with dysarthria, IT's LLM agents intelligentlycorrected token errors and enriched sentence-level emotional and logicalcoherence, achieving low error rates (4.2% word error rate, 2.9% sentence errorrate) and a 55% increase in user satisfaction. This work establishes aportable, intuitive communication platform for patients with dysarthria withthe potential to be applied broadly across different neurological conditionsand in multi-language support systems.
标题:迈向改进的客观感知音频质量评估--第1部分:一种新型的数据驱动认知模型
链接:https://arxiv.org/abs/2411.18222
备注:Accepter for publication in in IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:高效的音频质量评估对于简化音频编解码器开发至关重要。随着时间的推移,客观评估工具已经开发出来,可以通过算法从主观评估中预测质量评级,这是质量判断的黄金标准。这些工具中的许多使用感知听觉模型来提取音频特征,这些音频特征被映射到使用机器学习算法和主观分数作为训练数据的基本音频质量分数预测。然而,现有的工具在质量预测中难以泛化,特别是当面对未知的信号和失真类型时。这在使用非波形保持参数技术编码的信号的存在下特别明显。为了应对这些挑战,这两部分工作提出了对音频质量感知评估(PEAQ-ITU-R BS. 1387 -1)建议的扩展。第1部分着重于提高泛化能力,而第2部分则针对音频编码中的精确空间音频质量测量。 为了提高预测的泛化能力,本文(第1部分)介绍了一种新的机器学习方法,该方法使用主观数据对音频质量感知的认知方面进行建模。该方法通过自适应加权不同的失真度量来模拟听觉失真的感知严重程度。权重使用捕获失真显著性和认知效应之间的关系的交互成本函数来确定。与其他机器学习方法和已建立的工具相比,所提出的架构在以前看不见的主观质量分数的大型数据库上实现了更高的预测精度。感知激励模型为通用机器学习算法提供了更易于管理的替代方案,允许对多维质量测量进行潜在的扩展和改进,而无需完全重新训练。
摘要:Efficient audio quality assessment is vital for streamlining audio codecdevelopment. Objective assessment tools have been developed over time toalgorithmically predict quality ratings from subjective assessments, the goldstandard for quality judgment. Many of these tools use perceptual auditorymodels to extract audio features that are mapped to a basic audio quality scoreprediction using machine learning algorithms and subjective scores as trainingdata. However, existing tools struggle with generalization in qualityprediction, especially when faced with unknown signal and distortion types.This is particularly evident in the presence of signals coded usingnon-waveform-preserving parametric techniques. Addressing these challenges,this two-part work proposes extensions to the Perceptual Evaluation of AudioQuality (PEAQ - ITU-R BS.1387-1) recommendation. Part 1 focuses on increasinggeneralization, while Part 2 targets accurate spatial audio quality measurementin audio coding. To enhance prediction generalization, this paper (Part 1) introduces a novelmachine learning approach that uses subjective data to model cognitive aspectsof audio quality perception. The proposed method models the perceived severityof audible distortions by adaptively weighting different distortion metrics.The weights are determined using an interaction cost function that capturesrelationships between distortion salience and cognitive effects. Compared toother machine learning methods and established tools, the proposed architectureachieves higher prediction accuracy on large databases of previously unseensubjective quality scores. The perceptually-motivated model offers a moremanageable alternative to general-purpose machine learning algorithms, allowingpotential extensions and improvements to multi-dimensional quality measurementwithout complete retraining.
标题:SALMON-omni:一种用于双环语音理解和生成的无编解码器LLM
链接:https://arxiv.org/abs/2411.18138
备注:Technical report
摘要:全双工多模态大型语言模型(LLM)为解决不同的语音理解和生成任务提供了统一的框架,从而实现了更自然、更无缝的人机对话。与传统的模块化会话AI系统不同,它将语音识别,理解和文本到语音生成分离为不同的组件,多模式LLM作为单一的端到端模型运行。这种简化的设计消除了跨组件的错误传播,并充分利用了嵌入在输入语音信号中的丰富的非语言信息。我们介绍SALMONN-omni,一种无编解码器的全双工语音理解和生成模型,能够在说话时同时收听自己生成的语音和背景声音。为了支持这种能力,我们提出了一种新的双工口语对话框架,结合了一个“思考”的机制,有利于异步文本和语音生成依赖于嵌入,而不是编解码器(量化的语音和音频令牌)。实验结果表明,SALMONN-omni的通用性在广泛的流语音任务,包括语音识别,语音增强,和口语问答。此外,SALMONN-omni在管理轮对、打断和回声消除场景方面表现出色,从而确立了其作为全双工对话式AI系统的强大原型的潜力。据我们所知,SALMONN-omni是同类产品中第一个无编解码器的型号。一份完整的技术报告以及模型检查点将很快发布。
摘要:Full-duplex multimodal large language models (LLMs) provide a unifiedframework for addressing diverse speech understanding and generation tasks,enabling more natural and seamless human-machine conversations. Unliketraditional modularised conversational AI systems, which separate speechrecognition, understanding, and text-to-speech generation into distinctcomponents, multimodal LLMs operate as single end-to-end models. Thisstreamlined design eliminates error propagation across components and fullyleverages the rich non-verbal information embedded in input speech signals. Weintroduce SALMONN-omni, a codec-free, full-duplex speech understanding andgeneration model capable of simultaneously listening to its own generatedspeech and background sounds while speaking. To support this capability, wepropose a novel duplex spoken dialogue framework incorporating a ``thinking''mechanism that facilitates asynchronous text and speech generation relying onembeddings instead of codecs (quantized speech and audio tokens). Experimentalresults demonstrate SALMONN-omni's versatility across a broad range ofstreaming speech tasks, including speech recognition, speech enhancement, andspoken question answering. Additionally, SALMONN-omni excels at managingturn-taking, barge-in, and echo cancellation scenarios, establishing itspotential as a robust prototype for full-duplex conversational AI systems. Tothe best of our knowledge, SALMONN-omni is the first codec-free model of itskind. A full technical report along with model checkpoints will be releasedsoon.
标题:JPPO:加速大型语言模型服务的联合动力和即时优化
链接:https://arxiv.org/abs/2411.18010
摘要:大型语言模型(LLM)在各种任务中表现出卓越的能力,导致它们在无线网络中越来越多地部署,以提供各种各样的用户服务。然而,越来越长的提示设置突出了计算资源需求和巨大的通信负载的关键问题。为了应对这一挑战,我们提出了联合功率和即时优化(JPPO),这是一个将基于小语言模型(SLM)的即时压缩与无线功率分配优化相结合的框架。通过在用户设备上部署SLM进行快速压缩,并采用深度强化学习对压缩比和传输功率进行联合优化,JPPO有效地平衡了服务质量和资源效率。实验结果表明,我们的框架实现了高服务保真度和低误码率,同时优化无线LLM服务的功率使用。该系统将响应时间缩短了约17%,改进程度取决于原始提示的长度。
摘要:Large Language Models (LLMs) have demonstrated remarkable capabilities invarious tasks, leading to their increasing deployment in wireless networks fora wide variety of user services. However, the growing longer prompt settinghighlights the crucial issue of computational resource demands and hugecommunication load. To address this challenge, we propose Joint Power andPrompt Optimization (JPPO), a framework that combines Small Language Model(SLM)-based prompt compression with wireless power allocation optimization. Bydeploying SLM at user devices for prompt compression and employing DeepReinforcement Learning for joint optimization of compression ratio andtransmission power, JPPO effectively balances service quality with resourceefficiency. Experimental results demonstrate that our framework achieves highservice fidelity and low bit error rates while optimizing power usage inwireless LLM services. The system reduces response time by about 17%, with theimprovement varying based on the length of the original prompt.
标题:使用具有嵌入损失的神经音频编解码器的语音分离
链接:https://arxiv.org/abs/2411.17998
备注:Accepted by APSIPA ASC 2024
摘要:神经音频编解码器通过使语音任务能够在高度压缩的表示上执行而彻底改变了音频处理。最近的工作表明,语音分离可以在这些压缩域中实现,从而提供更快的训练和更低的推理成本。然而,目前的方法仍然依赖于基于波形的损失函数,在训练期间需要不必要的解码步骤。我们提出了一种新的基于神经音频编解码器的语音分离的嵌入损失,它直接对压缩的音频表示进行操作,无需在训练过程中进行解码。为了验证我们的方法,我们使用客观指标和感知评估技术(包括侵入性和非侵入性方法)进行全面评估。我们的研究结果表明,嵌入损失可用于训练基于编解码器的语音分离模型,训练速度和计算成本提高了2倍,同时在WSJ 0 - 2 mix数据集上跨3种不同的预训练编解码器实现了更好的DNSMOS和STOI性能。
摘要:Neural audio codecs have revolutionized audio processing by enabling speechtasks to be performed on highly compressed representations. Recent work hasshown that speech separation can be achieved within these compressed domains,offering faster training and reduced inference costs. However, currentapproaches still rely on waveform-based loss functions, necessitatingunnecessary decoding steps during training. We propose a novel embedding lossfor neural audio codec-based speech separation that operates directly oncompressed audio representations, eliminating the need for decoding duringtraining. To validate our approach, we conduct comprehensive evaluations usingboth objective metrics and perceptual assessment techniques, includingintrusive and non-intrusive methods. Our results demonstrate that embeddingloss can be used to train codec-based speech separation models with a 2ximprovement in training speed and computational cost while achieving betterDNSMOS and STOI performance on the WSJ0-2mix dataset across 3 differentpre-trained codecs.
标题:Disentangled-Transformer:一种具有语音内容-上下文分离的可解释的端到端自动语音识别模型
链接:https://arxiv.org/abs/2411.17846
备注:Accepted by the 6th IEEE International Conference on Image Processing Applications and Systems
摘要:端到端基于transformer的自动语音识别(ASR)系统通常在其学习的表示中捕获多个语音特征,这些特征高度纠缠,导致缺乏可解释性。在本研究中,我们提出了可解释的Disentangled-Transformer,它根据不同的时间分辨率将内部表示分解为具有明确内容和说话者特征的子嵌入。实验结果表明,所提出的分离变压器产生一个明确的说话人身份,从语音内容分离,说话人日记,同时提高ASR性能。
摘要:End-to-end transformer-based automatic speech recognition (ASR) systems oftencapture multiple speech traits in their learned representations that are highlyentangled, leading to a lack of interpretability. In this study, we propose theexplainable Disentangled-Transformer, which disentangles the internalrepresentations into sub-embeddings with explicit content and speaker traitsbased on varying temporal resolutions. Experimental results show that theproposed Disentangled-Transformer produces a clear speaker identity, separatedfrom the speech content, for speaker diarization while improving ASRperformance.
标题:嵌入式声发射分析的微型机器学习技术比较
链接:https://arxiv.org/abs/2411.17733
备注:Conference Presentations (Accepted) at IEEE 10th World Forum on Internet of Things. "this https URL"
摘要:本文比较了机器学习方法与不同的输入数据格式的声发射(AE)信号的分类。声发射信号是一种很有前途的监测技术,在许多结构健康监测应用。机器学习已被证明是一种有效的数据分析方法,可以根据不同的AE信号所代表的损伤机制对其进行分类。这些分类可以基于整个AE波形或从中提取的特定特征来执行。然而,目前还不知道这些方法中的哪一种是优选的。以资源受限的嵌入式物联网(IoT)系统的模型部署为目标,这项工作评估和比较两种方法的分类精度,内存需求,处理时间和能耗。为了实现这一目标,提取并仔细选择特征,为每个输入数据场景设计和优化神经网络模型,并将这些模型部署在低功耗物联网节点上。对比分析表明,所有模型都能达到99%以上的高分类精度,但嵌入式特征提取计算量大。因此,利用原始AE信号作为输入的模型具有最快的处理速度,从而具有最低的能耗,这是以较大的存储器需求为代价的。
摘要:This paper compares machine learning approaches with different input dataformats for the classification of acoustic emission (AE) signals. AE signalsare a promising monitoring technique in many structural health monitoringapplications. Machine learning has been demonstrated as an effective dataanalysis method, classifying different AE signals according to the damagemechanism they represent. These classifications can be performed based on theentire AE waveform or specific features that have been extracted from it.However, it is currently unknown which of these approaches is preferred. Withthe goal of model deployment on resource-constrained embedded Internet ofThings (IoT) systems, this work evaluates and compares both approaches in termsof classification accuracy, memory requirement, processing time, and energyconsumption. To accomplish this, features are extracted and carefully selected,neural network models are designed and optimized for each input data scenario,and the models are deployed on a low-power IoT node. The comparative analysisreveals that all models can achieve high classification accuracies of over99\%, but that embedded feature extraction is computationally expensive.Consequently, models utilizing the raw AE signal as input have the fastestprocessing speed and thus the lowest energy consumption, which comes at thecost of a larger memory requirement.
标题:多项选择学习,实现多个扬声器的高效语音分离
链接:https://arxiv.org/abs/2411.18497
摘要:在有监督的环境中训练语音分离模型提出了一个置换问题:找到模型预测和地面真实分离信号之间的最佳分配。这个固有的模糊任务通常使用Permutation Invariant Training(PIT)来解决。在本文中,我们考虑使用多项选择学习(MCL)框架,该框架最初是为了解决模糊任务而引入的。我们在流行的WSJ 0-mix和LibriMix基准测试上通过实验证明,MCL与PIT的性能相匹配,同时在计算上具有优势。这为一个有前途的研究方向打开了大门,因为MCL可以自然地扩展到处理可变数量的扬声器,或者在无监督环境中处理语音分离。
摘要:Training speech separation models in the supervised setting raises apermutation problem: finding the best assignation between the model predictionsand the ground truth separated signals. This inherently ambiguous task iscustomarily solved using Permutation Invariant Training (PIT). In this article,we instead consider using the Multiple Choice Learning (MCL) framework, whichwas originally introduced to tackle ambiguous tasks. We demonstrateexperimentally on the popular WSJ0-mix and LibriMix benchmarks that MCL matchesthe performances of PIT, while being computationally advantageous. This opensthe door to a promising research direction, as MCL can be naturally extended tohandle a variable number of speakers, or to tackle speech separation in theunsupervised setting.
标题:带噪音增强的连续自回归模型避免误差累积
链接:https://arxiv.org/abs/2411.18447
备注:Accepted to NeurIPS 2024 - Audio Imagination Workshop
摘要:自回归模型通常应用于离散令牌序列,但最近的研究表明,以自回归方式生成连续嵌入序列也是可行的。然而,这样的连续自回归模型(CAM)可能会遭受由于推断期间的错误累积而导致的扩展序列的生成质量下降。我们引入了一种新的方法来解决这个问题,在训练过程中将随机噪声注入输入嵌入。这个过程使模型在推理时对不同的误差水平具有鲁棒性。我们通过引入低水平噪声的推理过程进一步减少了误差积累。音乐音频生成的实验表明,CAM大大优于现有的自回归和非自回归方法,同时保持音频质量超过扩展序列。这项工作为在纯自回归设置中生成连续嵌入铺平了道路,为实时和交互式生成应用开辟了新的可能性。
摘要:Autoregressive models are typically applied to sequences of discrete tokens,but recent research indicates that generating sequences of continuousembeddings in an autoregressive manner is also feasible. However, suchContinuous Autoregressive Models (CAMs) can suffer from a decline in generationquality over extended sequences due to error accumulation during inference. Weintroduce a novel method to address this issue by injecting random noise intothe input embeddings during training. This procedure makes the model robustagainst varying error levels at inference. We further reduce error accumulationthrough an inference procedure that introduces low-level noise. Experiments onmusical audio generation show that CAM substantially outperforms existingautoregressive and non-autoregressive approaches while preserving audio qualityover extended sequences. This work paves the way for generating continuousembeddings in a purely autoregressive setting, opening new possibilities forreal-time and interactive generative applications.
标题:AMPS:具有多模式解释监督的ASB
链接:https://arxiv.org/abs/2411.18368
摘要:自发或会话式多语言语音对最先进的自动语音识别(ASR)系统提出了许多挑战。在这项工作中,我们提出了一种新的技术AMPS,增强了多语言的多模式ASR系统与基于释义的监督,以改善会话ASR在多种语言,包括印地语,马拉地语,马拉雅拉姆语,卡纳达语和Nyanja。我们在训练多模态ASR模型时使用参考转录的释义作为额外的监督,并选择性地为ASR性能较差的话语调用此释义目标。使用AMPS和最先进的多模式模型SeamlessM4 T,我们可以将字错误率(WER)显着相对降低高达5%。我们提出了详细的分析,我们的系统使用客观和人为的评价指标。
摘要:Spontaneous or conversational multilingual speech presents many challengesfor state-of-the-art automatic speech recognition (ASR) systems. In this work,we present a new technique AMPS that augments a multilingual multimodal ASRsystem with paraphrase-based supervision for improved conversational ASR inmultiple languages, including Hindi, Marathi, Malayalam, Kannada, and Nyanja.We use paraphrases of the reference transcriptions as additional supervisionwhile training the multimodal ASR model and selectively invoke this paraphraseobjective for utterances with poor ASR performance. Using AMPS with astate-of-the-art multimodal model SeamlessM4T, we obtain significant relativereductions in word error rates (WERs) of up to 5%. We present detailed analysesof our system using both objective and human evaluation metrics.
标题:利用梯度情景记忆进行机器语音链中的连续学习
链接:https://arxiv.org/abs/2411.18320
备注:Published as a conference paper at O-COCOSDA 2024. 6 pages; 2 figures
摘要:自动语音识别(ASR)系统的持续学习提出了一个挑战,特别是需要避免灾难性遗忘,同时保持先前学习任务的性能。本文介绍了一种新的方法,利用机器语音链框架,使持续学习的ASR使用梯度情景记忆(GEM)。通过将文本到语音(TTS)的机器语音链中的组件,我们支持GEM必不可少的重放机制,允许ASR模型顺序学习新的任务,而不会显着降低早期任务的性能。我们在LJ Speech数据集上进行的实验表明,我们的方法优于传统的微调和多任务学习方法,实现了大幅降低错误率,同时在不同的噪声条件下保持高性能。我们展示了半监督机器语音链方法在语音识别中有效且高效地持续学习的潜力。
摘要:Continual learning for automatic speech recognition (ASR) systems poses achallenge, especially with the need to avoid catastrophic forgetting whilemaintaining performance on previously learned tasks. This paper introduces anovel approach leveraging the machine speech chain framework to enablecontinual learning in ASR using gradient episodic memory (GEM). Byincorporating a text-to-speech (TTS) component within the machine speech chain,we support the replay mechanism essential for GEM, allowing the ASR model tolearn new tasks sequentially without significant performance degradation onearlier tasks. Our experiments, conducted on the LJ Speech dataset, demonstratethat our method outperforms traditional fine-tuning and multitask learningapproaches, achieving a substantial error rate reduction while maintaining highperformance across varying noise conditions. We showed the potential of oursemi-supervised machine speech chain approach for effective and efficientcontinual learning in speech recognition.
标题:如何学习一门新语言?自我监督学习模型的有效解决方案低资源场景下的隐形语言适应
链接:https://arxiv.org/abs/2411.18217
摘要:语音自监督学习(SSL)模型的使用在自动语音识别(ASR)上取得了令人印象深刻的性能。然而,在低资源语言ASR中,他们遇到了预训练语言和低资源语言之间的域不匹配问题。典型的解决方案,如微调SSL模型,遭受高计算成本,而使用冻结SSL模型作为特征提取器,性能较差。为了处理这些问题,我们扩展了传统的高效的微调方案的基础上的适配器。我们添加了一个额外的中间适配器来预热适配器和下游模型初始化。值得注意的是,我们只更新总模型参数的1-5%来实现自适应。ML-SUPERB数据集上的实验结果表明,我们的解决方案优于传统的高效微调。在适应看不见的语言时,它在字符/音素错误率方面实现了高达28%的相对改善。
摘要:The utilization of speech Self-Supervised Learning (SSL) models achievesimpressive performance on Automatic Speech Recognition (ASR). However, inlow-resource language ASR, they encounter the domain mismatch problem betweenpre-trained and low-resource languages. Typical solutions like fine-tuning theSSL model suffer from high computation costs while using frozen SSL models asfeature extractors comes with poor performance. To handle these issues, weextend a conventional efficient fine-tuning scheme based on the adapter. We addan extra intermediate adaptation to warm up the adapter and downstream modelinitialization. Remarkably, we update only 1-5% of the total model parametersto achieve the adaptation. Experimental results on the ML-SUPERB dataset showthat our solution outperforms conventional efficient fine-tuning. It achievesup to a 28% relative improvement in the Character/Phoneme error rate whenadapting to unseen languages.
标题:MSA-ASB:使用冻结ASB模型的高效多语言说话者归因
链接:https://arxiv.org/abs/2411.18152
摘要:说话人属性自动语音识别(SA-ASR)的目标是在转录语音的同时准确地将转录本分配给相应的说话人。现有方法通常依赖于复杂的模块化系统或需要对关节模块进行大量微调,从而限制了它们的适应性和总体效率。本文介绍了一种新的方法,利用冻结的多语言ASR模型,将说话人属性到transmittance,只使用标准的单语ASR数据集。我们的方法涉及训练扬声器模块,以基于弱标签预测扬声器嵌入,而不需要额外的ASR模型修改。尽管只使用非重叠的单语数据进行训练,但我们的方法可以有效地在不同的多语言数据集中提取说话人属性,包括那些具有重叠语音的数据集。实验结果表明,与强基线相比,具有竞争力的性能,突出了模型的鲁棒性和实际应用的潜力。
摘要:Speaker-attributed automatic speech recognition (SA-ASR) aims to transcribespeech while assigning transcripts to the corresponding speakers accurately.Existing methods often rely on complex modular systems or require extensivefine-tuning of joint modules, limiting their adaptability and generalefficiency. This paper introduces a novel approach, leveraging a frozenmultilingual ASR model to incorporate speaker attribution into thetranscriptions, using only standard monolingual ASR datasets. Our methodinvolves training a speaker module to predict speaker embeddings based on weaklabels without requiring additional ASR model modifications. Despite beingtrained exclusively with non-overlapping monolingual data, our approacheffectively extracts speaker attributes across diverse multilingual datasets,including those with overlapping speech. Experimental results demonstratecompetitive performance compared to strong baselines, highlighting the model'srobustness and potential for practical applications.
标题:多语言自动语音识别的离散表示和自增强表示的融合
链接:https://arxiv.org/abs/2411.18107
备注:SLT 2024
摘要:自监督学习(SSL)模型在各种语音处理任务中表现出卓越的能力。连续SSL表示是有效的,但遭受高计算和存储需求。另一方面,离散的SSL表示,虽然性能下降,降低了传输和存储成本,并通过重复数据删除和子字建模提高输入序列的效率。为了提高ASR的离散表示的性能,我们引入了一种新的融合机制,集成了两个离散表示。融合机制保留了离散表示的所有优点,同时通过集成互补信息来增强模型的性能。此外,我们还探索了“自增强”离散表示,它将变换应用于单个连续SSL表示,消除了融合机制对多个SSL模型的依赖,并进一步降低了其推理成本。在LibriSpeech和ML-SUPERB等基准测试上的实验结果表明,与非融合基线相比,相对字符错误率分别提高了19%和24%,验证了本文方法的有效性。
摘要:Self-supervised learning (SSL) models have shown exceptional capabilitiesacross various speech-processing tasks. Continuous SSL representations areeffective but suffer from high computational and storage demands. On the otherhand, discrete SSL representations, although with degraded performance, reducetransmission and storage costs, and improve input sequence efficiency throughde-duplication and subword-modeling. To boost the performance of discreterepresentations for ASR, we introduce a novel fusion mechanism that integratestwo discrete representations. The fusion mechanism preserves all the benefitsof discrete representation while enhancing the model's performance byintegrating complementary information. Additionally, we explore"self-augmented'' discrete representations, which apply transformations to asingle continuous SSL representation, eliminating the fusion mechanism'sdependency on multiple SSL models and further decreasing its inference costs.Experimental results on benchmarks, including LibriSpeech and ML-SUPERB,indicate up to 19% and 24% relative character error rate improvement comparedwith the non-fusion baseline, validating the effectiveness of our proposedmethods.
标题:Music 2 Fail:将音乐转移到失败的录音机风格
链接:https://arxiv.org/abs/2411.18075
备注:Accepted by APSIPA 2024
摘要:音乐风格转换的目的是将一种乐器演奏的音乐转换为另一种乐器演奏的音乐,同时保持音乐内容不变。在本文中,我们研究了另一种风格转移的情况下,所谓的“失败的音乐风格转移”。与通常的音乐风格转移不同,在通常的音乐风格转移中,内容保持不变,只有乐器的特征被改变,这种情况试图将音乐从源乐器转移到目标乐器,这是故意偏离音高执行的。我们的工作试图将正常播放的音乐转换为非音高录音机音乐,我们称之为“失败式录音机”,并研究转换的结果。为了完成这项工作,我们还提出了一个失败式记录器的数据集,称为“FR109数据集”。这样的实验在一个更有表现力的环境中探索音乐风格转移任务,因为生成的音频听起来应该像一个“偏离音高的录音机”,同时保持一定程度的自然性。
摘要:The goal of music style transfer is to convert a music performance by oneinstrument into another while keeping the musical contents unchanged. In thispaper, we investigate another style transfer scenario called ``failed-musicstyle transfer''. Unlike the usual music style transfer where the contentremains the same and only the instrumental characteristics are changed, thisscenario seeks to transfer the music from the source instrument to the targetinstrument which is deliberately performed off-pitch. Our work attempts totransfer normally played music into off-pitch recorder music, which we call``failed-style recorder'', and study the results of the conversion. To carryout this work, we have also proposed a dataset of failed-style recorders forthis task, called ``FR109 Dataset''. Such an experiment explores the musicstyle transfer task in a more expressive setting, as the generated audio shouldsound like an ``off-pitch recorder'' while maintaining a certain degree ofnaturalness.
