今日论文合集:cs.SD语音7篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
标题: BLSP-KD:通过知识提炼的Bootstrapping语音预训练
作者:Chen Wang,Minpeng Liao,Zhongqiang Huang,Jiajun Zhang
链接:点击下载PDF文件
摘要:最近的端到端方法在将大型语言模型(LLM)扩展到语音输入方面表现出了希望,但在直接评估和优化对齐质量方面面临限制,并且由于语音-文本长度不匹配而无法实现细粒度对齐。我们介绍了BLSP-KD,一种新的方法,通过知识蒸馏的Bootstrapping语音预训练,它通过两个关键技术来解决这些限制。首先,它通过使用知识蒸馏最小化LLM的语音和文本输入的下一个令牌预测分布之间的差异来优化语音-文本对齐。其次,它采用了一个continuous-integrate-andfire策略,将语音分割成与文本标记一一对应的标记,从而实现细粒度对齐。我们还介绍了部分LoRA(PLoRA),一种新的适应方法,支持LLM微调语音输入下的知识蒸馏。定量评估表明,BLSP-KD优于以前的端到端基线和级联系统,具有可比的参数规模,有利于语音输入的LLM的一般跟踪能力。这种方法为将LLM扩展到口语交互提供了新的可能性。摘要:Recent end-to-end approaches have shown promise in extending large language models (LLMs) to speech inputs, but face limitations in directly assessing and optimizing alignment quality and fail to achieve fine-grained alignment due to speech-text length mismatch. We introduce BLSP-KD, a novel approach for Bootstrapping Language-Speech Pretraining via Knowledge Distillation, which addresses these limitations through two key techniques. First, it optimizes speech-text alignment by minimizing the divergence between the LLM's next-token prediction distributions for speech and text inputs using knowledge distillation. Second, it employs a continuous-integrate-andfire strategy to segment speech into tokens that correspond one-to-one with text tokens, enabling fine-grained alignment. We also introduce Partial LoRA (PLoRA), a new adaptation method supporting LLM finetuning for speech inputs under knowledge distillation. Quantitative evaluation shows that BLSP-KD outperforms previous end-to-end baselines and cascaded systems with comparable scale of parameters, facilitating general instruction-following capabilities for LLMs with speech inputs. This approach provides new possibilities for extending LLMs to spoken language interactions.

【2】 Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
标题: 逆转听觉处理途径:从fMRI进行粗到细的音频重建
作者:Che Liu,Changde Du,Xiaoyu Chen,Huiguang He
链接:点击下载PDF文件
摘要:从人类听觉系统的分层处理,将声音从低级别的声学特征转换到高级别的语义理解的灵感,我们介绍了一种新的粗到细的音频重建方法。利用非侵入性功能磁共振成像(fMRI)数据,我们的方法模仿听觉处理的逆通路。首先,我们利用CLAP解码fMRI数据粗到一个低维的语义空间,然后由一个细粒度的解码到高维AudioMAE潜在的空间的语义特征的指导下。这些细粒度的神经特征作为通过潜在扩散模型(LDM)进行音频重建的条件。三个公共fMRI测试集Brain 2Sound,Brain 2 Music和Brain 2Speech的验证强调了我们的粗到细解码方法优于独立的细粒度方法,展示了FD,FAD和KL等指标的最新性能。此外,通过在解码过程中采用语义提示,我们提高了重建音频的质量时,语义特征是次优的。我们的模型在不同刺激下的多功能性突出了其作为通用脑-音频框架的潜力。这项研究有助于理解人类听觉系统,推动神经解码和音频重建方法的界限。摘要:Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel coarse-to-fine audio reconstruction method. Leveraging non-invasive functional Magnetic Resonance Imaging (fMRI) data, our approach mimics the inverse pathway of auditory processing. Initially, we utilize CLAP to decode fMRI data coarsely into a low-dimensional semantic space, followed by a fine-grained decoding into the high-dimensional AudioMAE latent space guided by semantic features. These fine-grained neural features serve as conditions for audio reconstruction through a Latent Diffusion Model (LDM). Validation on three public fMRI datasets-Brain2Sound, Brain2Music, and Brain2Speech-underscores the superiority of our coarse-to-fine decoding method over stand-alone fine-grained approaches, showcasing state-of-the-art performance in metrics like FD, FAD, and KL. Moreover, by employing semantic prompts during decoding, we enhance the quality of reconstructed audio when semantic features are suboptimal. The demonstrated versatility of our model across diverse stimuli highlights its potential as a universal brain-to-audio framework. This research contributes to the comprehension of the human auditory system, pushing boundaries in neural decoding and audio reconstruction methodologies.

【3】 SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
标题: SoundChem:将基于分数的模型和一致性模型结合起来用于文本到声音的生成
作者:Koichi Saito,Dongjun Kim,Takashi Shibuya,Chieh-Hsin Lai,Zhi Zhong,Yuhta Takida,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:声音内容是视频游戏、音乐和电影等多媒体作品不可或缺的元素。最近的高质量的基于扩散的声音生成模型可以作为创作者的宝贵工具。然而,尽管产生高质量的声音,这些模型往往受到缓慢的推理速度。这一缺点给创作者带来了负担,他们通常通过试验和错误来完善他们的声音,使其与他们的艺术意图保持一致。为了解决这个问题,我们引入了声音一致性轨迹模型(SoundCTM)。我们的模型可以在高质量的一步声音生成和通过多步生成的卓越音质之间灵活过渡。这允许创作者在通过多步生成进行优化之前,首先使用一步采样控制声音。虽然CTM从根本上实现了灵活的一步和多步生成,但其令人印象深刻的性能在很大程度上取决于额外的预训练特征提取器和对抗性损失,这对于训练来说是昂贵的,并且在其他领域并不总是可用。因此,我们重新构建CTM的训练框架,并通过利用教师网络的蒸馏损失引入一个新的特征距离。此外,在提取无分类器引导轨迹的同时,我们同时训练有条件和无条件学生模型,并在推理过程中在这些模型之间进行插值。我们还提出了SoundCTM的免训练可控框架,利用其灵活的采样能力。SoundCTM实现了有前途的一步和多步实时声音生成,而无需使用任何额外的现成网络。此外,我们证明了SoundCTM的能力,可控的声音生成在一个培训的方式。摘要:Sound content is an indispensable element for multimedia works such as video games, music, and films. Recent high-quality diffusion-based sound generation models can serve as valuable tools for the creators. However, despite producing high-quality sounds, these models often suffer from slow inference speeds. This drawback burdens creators, who typically refine their sounds through trial and error to align them with their artistic intentions. To address this issue, we introduce Sound Consistency Trajectory Models (SoundCTM). Our model enables flexible transitioning between high-quality 1-step sound generation and superior sound quality through multi-step generation. This allows creators to initially control sounds with 1-step samples before refining them through multi-step generation. While CTM fundamentally achieves flexible 1-step and multi-step generation, its impressive performance heavily depends on an additional pretrained feature extractor and an adversarial loss, which are expensive to train and not always available in other domains. Thus, we reframe CTM's training framework and introduce a novel feature distance by utilizing the teacher's network for a distillation loss. Additionally, while distilling classifier-free guided trajectories, we train conditional and unconditional student models simultaneously and interpolate between these models during inference. We also propose training-free controllable frameworks for SoundCTM, leveraging its flexible sampling capability. SoundCTM achieves both promising 1-step and multi-step real-time sound generation without using any extra off-the-shelf networks. Furthermore, we demonstrate SoundCTM's capability of controllable sound generation in a training-free manner.

【4】 Improving Speech Decoding from ECoG with Self-Supervised Pretraining
标题: 通过自我监督预训练改进ECoG的语音解码
作者:Brian A. Yuan,Joseph G. Makin
链接:点击下载PDF文件
摘要:最近关于颅内脑机接口的研究表明,可以高精度地解码语音,基本上是通过将问题视为监督学习的实例,并训练深度神经网络将神经活动映射到文本。然而,这样的网络用大量的标记数据来支付它们的表现力,这对于从人类患者获得的侵入性神经记录来说是特别繁重的要求。另一方面,这些患者通常在用于训练解码器的实验块之外产生语音。利用这些数据和其他患者的数据来改善解码将减轻数据收集的负担-特别是对疾病和关节炎患者来说是繁重的。在这里,我们证明了这是可能的,通过重新设计wav 2 vec-一个简单的,自我监督的,完全卷积的模型,使用噪声对比损失来学习音频的潜在表示-用于皮层电图(ECoG)数据。我们在未标记的ECoG记录上训练这个模型,然后使用它将ECoG从标记的语音会话转换到wav 2 vec的表示空间,最后训练一个有监督的编码器-解码器将这些表示映射到文本。我们用不同数量的标记块进行了实验;对于几乎所有的选择,新的表示都比原始ECoG数据产生更好的解码性能,并且在任何情况下都不会产生更差的解码性能。在某些情况下,还可以通过在另一个患者的数据上预训练wav 2 vec来提高性能。在最好的情况下,wav 2 vec的表示将原始数据的单词错误率降低了50%以上。摘要:Recent work on intracranial brain-machine interfaces has demonstrated that spoken speech can be decoded with high accuracy, essentially by treating the problem as an instance of supervised learning and training deep neural networks to map from neural activity to text. However, such networks pay for their expressiveness with very large numbers of labeled data, a requirement that is particularly burdensome for invasive neural recordings acquired from human patients. On the other hand, these patients typically produce speech outside of the experimental blocks used for training decoders. Making use of such data, and data from other patients, to improve decoding would ease the burden of data collection -- especially onerous for dys- and anarthric patients. Here we demonstrate that this is possible, by reengineering wav2vec -- a simple, self-supervised, fully convolutional model that learns latent representations of audio using a noise-contrastive loss -- for electrocorticographic (ECoG) data. We train this model on unlabelled ECoG recordings, and subsequently use it to transform ECoG from labeled speech sessions into wav2vec's representation space, before finally training a supervised encoder-decoder to map these representations to text. We experiment with various numbers of labeled blocks; for almost all choices, the new representations yield superior decoding performance to the original ECoG data, and in no cases do they yield worse. Performance can also be improved in some cases by pretraining wav2vec on another patient's data. In the best cases, wav2vec's representations decrease word error rates over the original data by upwards of 50%.


eess.AS音频处理
【1】 BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
标题: BLSP-KD:通过知识提炼的Bootstrapping语音预训练
作者:Chen Wang,Minpeng Liao,Zhongqiang Huang,Jiajun Zhang
链接:点击下载PDF文件
摘要:最近的端到端方法在将大型语言模型(LLM)扩展到语音输入方面表现出了希望,但在直接评估和优化对齐质量方面面临限制,并且由于语音-文本长度不匹配而无法实现细粒度对齐。我们介绍了BLSP-KD,一种新的方法,通过知识蒸馏的Bootstrapping语音预训练,它通过两个关键技术来解决这些限制。首先,它通过使用知识蒸馏最小化LLM的语音和文本输入的下一个令牌预测分布之间的差异来优化语音-文本对齐。其次,它采用了一个continuous-integrate-andfire策略,将语音分割成与文本标记一一对应的标记,从而实现细粒度对齐。我们还介绍了部分LoRA(PLoRA),一种新的适应方法,支持LLM微调语音输入下的知识蒸馏。定量评估表明,BLSP-KD优于以前的端到端基线和级联系统,具有可比的参数规模,有利于语音输入的LLM的一般跟踪能力。这种方法为将LLM扩展到口语交互提供了新的可能性。摘要:Recent end-to-end approaches have shown promise in extending large language models (LLMs) to speech inputs, but face limitations in directly assessing and optimizing alignment quality and fail to achieve fine-grained alignment due to speech-text length mismatch. We introduce BLSP-KD, a novel approach for Bootstrapping Language-Speech Pretraining via Knowledge Distillation, which addresses these limitations through two key techniques. First, it optimizes speech-text alignment by minimizing the divergence between the LLM's next-token prediction distributions for speech and text inputs using knowledge distillation. Second, it employs a continuous-integrate-andfire strategy to segment speech into tokens that correspond one-to-one with text tokens, enabling fine-grained alignment. We also introduce Partial LoRA (PLoRA), a new adaptation method supporting LLM finetuning for speech inputs under knowledge distillation. Quantitative evaluation shows that BLSP-KD outperforms previous end-to-end baselines and cascaded systems with comparable scale of parameters, facilitating general instruction-following capabilities for LLMs with speech inputs. This approach provides new possibilities for extending LLMs to spoken language interactions.

【2】 Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
标题: 逆转听觉处理途径:从fMRI进行粗到细的音频重建
作者:Che Liu,Changde Du,Xiaoyu Chen,Huiguang He
链接:点击下载PDF文件
摘要:从人类听觉系统的分层处理,将声音从低级别的声学特征转换到高级别的语义理解的灵感,我们介绍了一种新的粗到细的音频重建方法。利用非侵入性功能磁共振成像(fMRI)数据,我们的方法模仿听觉处理的逆通路。首先,我们利用CLAP解码fMRI数据粗到一个低维的语义空间,然后由一个细粒度的解码到高维AudioMAE潜在的空间的语义特征的指导下。这些细粒度的神经特征作为通过潜在扩散模型(LDM)进行音频重建的条件。三个公共fMRI测试集Brain 2Sound,Brain 2 Music和Brain 2Speech的验证强调了我们的粗到细解码方法优于独立的细粒度方法,展示了FD,FAD和KL等指标的最新性能。此外,通过在解码过程中采用语义提示,我们提高了重建音频的质量时,语义特征是次优的。我们的模型在不同刺激下的多功能性突出了其作为通用脑-音频框架的潜力。这项研究有助于理解人类听觉系统,推动神经解码和音频重建方法的界限。摘要:Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we introduce a novel coarse-to-fine audio reconstruction method. Leveraging non-invasive functional Magnetic Resonance Imaging (fMRI) data, our approach mimics the inverse pathway of auditory processing. Initially, we utilize CLAP to decode fMRI data coarsely into a low-dimensional semantic space, followed by a fine-grained decoding into the high-dimensional AudioMAE latent space guided by semantic features. These fine-grained neural features serve as conditions for audio reconstruction through a Latent Diffusion Model (LDM). Validation on three public fMRI datasets-Brain2Sound, Brain2Music, and Brain2Speech-underscores the superiority of our coarse-to-fine decoding method over stand-alone fine-grained approaches, showcasing state-of-the-art performance in metrics like FD, FAD, and KL. Moreover, by employing semantic prompts during decoding, we enhance the quality of reconstructed audio when semantic features are suboptimal. The demonstrated versatility of our model across diverse stimuli highlights its potential as a universal brain-to-audio framework. This research contributes to the comprehension of the human auditory system, pushing boundaries in neural decoding and audio reconstruction methodologies.

【3】 Zipper: A Multi-Tower Decoder Architecture for Fusing Modalities
标题: 拉链:用于融合模式的多塔解码器架构
作者:Vicky Zayats,Peter Chen,Melissa Merrari,Dirk Padfield
备注:Under review at NeurIPS
链接:点击下载PDF文件
摘要:将多个生成基础模型,特别是那些在不同模式下训练的模型,整合成比其部分总和更大的东西,这是一个巨大的挑战。两个关键障碍是对齐数据(包含相似含义但在不同模态中表达不同的概念)的可用性,以及在跨域生成任务中有效利用单峰表示,而不影响其原始单峰能力。 我们提出了Zipper,一个多塔解码器架构,通过使用交叉注意力来灵活地从独立的预训练的单峰解码器组成多模态生成模型,从而解决了这些问题。在我们的实验融合语音和文本模态,我们表明,所提出的架构执行非常有竞争力的方案,有限的对齐文本语音数据。我们还展示了我们的模型的灵活性,以选择性地保持单峰(例如,文本到文本生成)生成性能。在跨模态的任务,如自动语音识别(ASR)的输出模态是文本,我们表明,冻结的文本骨干可以忽略不计的性能下降。在跨模态的任务,如文本到语音生成(TTS)的输出模态是语音,我们表明,使用预先训练的语音骨干的结果在性能优于基线。摘要:Integrating multiple generative foundation models, especially those trained on different modalities, into something greater than the sum of its parts poses significant challenges. Two key hurdles are the availability of aligned data (concepts that contain similar meaning but is expressed differently in different modalities), and effectively leveraging unimodal representations in cross-domain generative tasks, without compromising their original unimodal capabilities. We propose Zipper, a multi-tower decoder architecture that addresses these concerns by using cross-attention to flexibly compose multimodal generative models from independently pre-trained unimodal decoders. In our experiments fusing speech and text modalities, we show the proposed architecture performs very competitively in scenarios with limited aligned text-speech data. We also showcase the flexibility of our model to selectively maintain unimodal (e.g., text-to-text generation) generation performance by freezing the corresponding modal tower (e.g. text). In cross-modal tasks such as automatic speech recognition (ASR) where the output modality is text, we show that freezing the text backbone results in negligible performance degradation. In cross-modal tasks such as text-to-speech generation (TTS) where the output modality is speech, we show that using a pre-trained speech backbone results in superior performance to the baseline.

【4】 Improving Speech Decoding from ECoG with Self-Supervised Pretraining
标题: 通过自我监督预训练改进ECoG的语音解码
作者:Brian A. Yuan,Joseph G. Makin
链接:点击下载PDF文件
摘要:最近关于颅内脑机接口的研究表明,可以高精度地解码语音,基本上是通过将问题视为监督学习的实例,并训练深度神经网络将神经活动映射到文本。然而,这样的网络用大量的标记数据来支付它们的表现力,这对于从人类患者获得的侵入性神经记录来说是特别繁重的要求。另一方面,这些患者通常在用于训练解码器的实验块之外产生语音。利用这些数据和其他患者的数据来改善解码将减轻数据收集的负担-特别是对疾病和关节炎患者来说是繁重的。在这里,我们证明了这是可能的,通过重新设计wav 2 vec-一个简单的,自我监督的,完全卷积的模型,使用噪声对比损失来学习音频的潜在表示-用于皮层电图(ECoG)数据。我们在未标记的ECoG记录上训练这个模型,然后使用它将ECoG从标记的语音会话转换到wav 2 vec的表示空间,最后训练一个有监督的编码器-解码器将这些表示映射到文本。我们用不同数量的标记块进行了实验;对于几乎所有的选择,新的表示都比原始ECoG数据产生更好的解码性能,并且在任何情况下都不会产生更差的解码性能。在某些情况下,还可以通过在另一个患者的数据上预训练wav 2 vec来提高性能。在最好的情况下,wav 2 vec的表示将原始数据的单词错误率降低了50%以上。摘要:Recent work on intracranial brain-machine interfaces has demonstrated that spoken speech can be decoded with high accuracy, essentially by treating the problem as an instance of supervised learning and training deep neural networks to map from neural activity to text. However, such networks pay for their expressiveness with very large numbers of labeled data, a requirement that is particularly burdensome for invasive neural recordings acquired from human patients. On the other hand, these patients typically produce speech outside of the experimental blocks used for training decoders. Making use of such data, and data from other patients, to improve decoding would ease the burden of data collection -- especially onerous for dys- and anarthric patients. Here we demonstrate that this is possible, by reengineering wav2vec -- a simple, self-supervised, fully convolutional model that learns latent representations of audio using a noise-contrastive loss -- for electrocorticographic (ECoG) data. We train this model on unlabelled ECoG recordings, and subsequently use it to transform ECoG from labeled speech sessions into wav2vec's representation space, before finally training a supervised encoder-decoder to map these representations to text. We experiment with various numbers of labeled blocks; for almost all choices, the new representations yield superior decoding performance to the original ECoG data, and in no cases do they yield worse. Performance can also be improved in some cases by pretraining wav2vec on another patient's data. In the best cases, wav2vec's representations decrease word error rates over the original data by upwards of 50%.

【5】 SoundCTM: Uniting Score-based and Consistency Models for Text-to-Sound Generation
标题: SoundChem:将基于分数的模型和一致性模型结合起来用于文本到声音的生成
作者:Koichi Saito,Dongjun Kim,Takashi Shibuya,Chieh-Hsin Lai,Zhi Zhong,Yuhta Takida,Yuki Mitsufuji
链接:点击下载PDF文件
摘要:声音内容是视频游戏、音乐和电影等多媒体作品不可或缺的元素。最近的高质量的基于扩散的声音生成模型可以作为创作者的宝贵工具。然而,尽管产生高质量的声音,这些模型往往受到缓慢的推理速度。这一缺点给创作者带来了负担,他们通常通过试验和错误来完善他们的声音,使其与他们的艺术意图保持一致。为了解决这个问题,我们引入了声音一致性轨迹模型(SoundCTM)。我们的模型可以在高质量的一步声音生成和通过多步生成的卓越音质之间灵活过渡。这允许创作者在通过多步生成进行优化之前,首先使用一步采样控制声音。虽然CTM从根本上实现了灵活的一步和多步生成,但其令人印象深刻的性能在很大程度上取决于额外的预训练特征提取器和对抗性损失,这对于训练来说是昂贵的,并且在其他领域并不总是可用。因此,我们重新构建CTM的训练框架,并通过利用教师网络的蒸馏损失引入一个新的特征距离。此外,在提取无分类器引导轨迹的同时,我们同时训练有条件和无条件学生模型,并在推理过程中在这些模型之间进行插值。我们还提出了SoundCTM的免训练可控框架,利用其灵活的采样能力。SoundCTM实现了有前途的一步和多步实时声音生成,而无需使用任何额外的现成网络。此外,我们证明了SoundCTM的能力,可控的声音生成在一个培训的方式。摘要:Sound content is an indispensable element for multimedia works such as video games, music, and films. Recent high-quality diffusion-based sound generation models can serve as valuable tools for the creators. However, despite producing high-quality sounds, these models often suffer from slow inference speeds. This drawback burdens creators, who typically refine their sounds through trial and error to align them with their artistic intentions. To address this issue, we introduce Sound Consistency Trajectory Models (SoundCTM). Our model enables flexible transitioning between high-quality 1-step sound generation and superior sound quality through multi-step generation. This allows creators to initially control sounds with 1-step samples before refining them through multi-step generation. While CTM fundamentally achieves flexible 1-step and multi-step generation, its impressive performance heavily depends on an additional pretrained feature extractor and an adversarial loss, which are expensive to train and not always available in other domains. Thus, we reframe CTM's training framework and introduce a novel feature distance by utilizing the teacher's network for a distillation loss. Additionally, while distilling classifier-free guided trajectories, we train conditional and unconditional student models simultaneously and interpolate between these models during inference. We also propose training-free controllable frameworks for SoundCTM, leveraging its flexible sampling capability. SoundCTM achieves both promising 1-step and multi-step real-time sound generation without using any extra off-the-shelf networks. Furthermore, we demonstrate SoundCTM's capability of controllable sound generation in a training-free manner.


机器翻译,仅供参考