今日论文合集:cs.SD语音13篇,eess.AS音频处理14篇。本文经arXiv每日学术速递授权转载
【1】Not All Weights Are Created Equal: Enhancing Energy Efficiency in On-Device Streaming Speech Recognition
标题:并非所有权重都是相等的:提高设备流传输语音识别的能效
链接:https://arxiv.org/abs/2402.13076
作者:Yang Li,Yuan Shangguan,Yuhao Wang,Liangzhen Lai,Ernie Chang,Changsheng Zhao,Yangyang Shi,Vikas Chandra
摘要:功耗在设备上流式语音识别中起着重要作用,因为它直接影响用户体验。本研究深入探讨语音识别模型中的权重参数如何影响这些模型的整体功耗。我们发现,权重参数对功耗的影响各不相同,受调用频率及其在内存中的位置等因素的影响。基于这一见解,我们开发了旨在优化设备上语音识别模型的设计指南。这些准则侧重于在不影响准确性的情况下最大限度地减少功耗。我们的方法,它采用有针对性的压缩的基础上的权重参数的变化的敏感性,表现出优越的性能相比,国家的最先进的压缩方法。它实现了高达47%的能源使用减少,同时保持类似的模型精度和提高实时系数。
摘要:Power consumption plays an important role in on-device streaming speech recognition, as it has a direct impact on the user experience. This study delves into how weight parameters in speech recognition models influence the overall power consumption of these models. We discovered that the impact of weight parameters on power consumption varies, influenced by factors including how often they are invoked and their placement in memory. Armed with this insight, we developed design guidelines aimed at optimizing on-device speech recognition models. These guidelines focus on minimizing power use without substantially affecting accuracy. Our method, which employs targeted compression based on the varying sensitivities of weight parameters, demonstrates superior performance compared to state-of-the-art compression methods. It achieves a reduction in energy usage of up to 47% while maintaining similar model accuracy and improving the real-time factor.
【2】SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion标题:SingVisio:歌声转换扩散模型的可视化分析作者:Liumeng Xue,Chaoren Wang,Mingxuan Wang,Xueyao Zhang,Jun Han,Zhizheng Wu摘要:在这项研究中,我们提出了Singlephon,一个交互式的视觉分析系统,旨在解释的扩散模型中使用的歌声转换。Singapolis提供了扩散模型中生成过程的可视化显示,展示了噪声频谱的逐步去噪及其转换为捕获所需歌手音色的干净频谱。该系统还有助于对不同条件(如源内容、旋律和目标音色)进行并排比较,突出显示这些条件对扩散生成过程和转换的影响。通过全面评估,新加坡证明了其在系统设计、功能、可解释性和用户友好性方面的有效性。它为不同背景的用户提供了宝贵的学习经验和对歌唱声音转换扩散模型的见解。摘要:In this study, we present SingVisio, an interactive visual analysis system that aims to explain the diffusion model used in singing voice conversion. SingVisio provides a visual display of the generation process in diffusion models, showcasing the step-by-step denoising of the noisy spectrum and its transformation into a clean spectrum that captures the desired singer's timbre. The system also facilitates side-by-side comparisons of different conditions, such as source content, melody, and target timbre, highlighting the impact of these conditions on the diffusion generation process and resulting conversions. Through comprehensive evaluations, SingVisio demonstrates its effectiveness in terms of system design, functionality, explainability, and user-friendliness. It offers users of various backgrounds valuable learning experiences and insights into the diffusion model for singing voice conversion.【3】Guiding the underwater acoustic target recognition with interpretable contrastive learning作者:Yuan Xie,Jiawei Ren,Ji Xu摘要:由于复杂的海洋环境和多变的水下信道,从声信号中识别水下目标是一项具有挑战性的任务。虽然基于深度学习的系统已成为水声目标识别的主流方法,但它们在实际应用中因缺乏可解释性和泛化性能弱而受到批评。在这项工作中,我们应用类激活映射(CAM)生成基于频谱的识别系统的预测的视觉解释。CAM可以通过突出显示对预测贡献最大的输入特征区域来帮助理解识别模型的行为。我们的探索表明,识别模型往往侧重于低频线谱和高频周期调制信息的水下信号。基于观察,我们提出了一个可解释的对比学习(ICL)策略,采用两个编码器学习不同的重点(线频谱和调制信息)的声学特征。通过在编码器之间施加约束,所提出的策略可以提高识别系统的泛化性能。我们的实验表明,所提出的对比学习方法可以提高识别精度,并在各种水下数据库带来显着的改善。摘要:Recognizing underwater targets from acoustic signals is a challenging task owing to the intricate ocean environments and variable underwater channels. While deep learning-based systems have become the mainstream approach for underwater acoustic target recognition, they have faced criticism for their lack of interpretability and weak generalization performance in practical applications. In this work, we apply the class activation mapping (CAM) to generate visual explanations for the predictions of a spectrogram-based recognition system. CAM can help to understand the behavior of recognition models by highlighting the regions of the input features that contribute the most to the prediction. Our explorations reveal that recognition models tend to focus on the low-frequency line spectrum and high-frequency periodic modulation information of underwater signals. Based on the observation, we propose an interpretable contrastive learning (ICL) strategy that employs two encoders to learn from acoustic features with different emphases (line spectrum and modulation information). By imposing constraints between encoders, the proposed strategy can enhance the generalization performance of the recognition system. Our experiments demonstrate that the proposed contrastive learning approach can improve the recognition accuracy and bring significant improvements across various underwater databases.
【4】OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification标题:OWSM-CTC:一种用于语音识别、翻译和语言识别的开放编码器语音基础模型作者:Yifan Peng,Yui Sudo,Muhammad Shakeel,Shinji Watanabe摘要:人们对能够在单个模型中执行多个语音处理任务的大型语音模型越来越感兴趣。这些模型通常采用编解码器或仅解码器架构,因为它们在许多领域中的流行性和良好性能。然而,与非自回归模型相比,自回归模型在推理过程中可能会更慢,并且也有潜在的幻觉风险。虽然先前的研究观察到非自回归模型在小规模的某些任务中有很好的结果,但目前还不清楚它们是否可以扩展到不同语言和任务的语音到文本生成。受开放式耳语风格语音模型(OWSM)项目的启发,我们提出了OWSM-CTC,一种新的基于连接主义时间分类(CTC)的仅编码器语音基础模型。它在18万小时的公共音频数据上进行训练,用于多语言自动语音识别(ASR),语音翻译(ST)和语言识别(LID)。与编码器-解码器OWSM相比,我们的OWSM-CTC在ASR上取得了有竞争力的结果,在ST上的相对改进高达25%,同时它更健壮,推理速度快3到4倍。OWSM-CTC还以20倍的速度改进了长格式ASR结果。我们将公开发布我们的代码库、预训练模型和训练日志,以促进语音基础模型的开放科学。摘要:There has been an increasing interest in large speech models that can perform multiple speech processing tasks in a single model. Such models usually adopt the encoder-decoder or decoder-only architecture due to their popularity and good performance in many domains. However, autoregressive models can be slower during inference compared to non-autoregressive models and also have potential risks of hallucination. Though prior studies observed promising results of non-autoregressive models for certain tasks at small scales, it remains unclear if they can be scaled to speech-to-text generation in diverse languages and tasks. Inspired by the Open Whisper-style Speech Model (OWSM) project, we propose OWSM-CTC, a novel encoder-only speech foundation model based on Connectionist Temporal Classification (CTC). It is trained on 180k hours of public audio data for multilingual automatic speech recognition (ASR), speech translation (ST), and language identification (LID). Compared to encoder-decoder OWSM, our OWSM-CTC achieves competitive results on ASR and up to 25% relative improvement on ST, while it is more robust and 3 to 4 times faster for inference. OWSM-CTC also improves the long-form ASR result with 20x speed-up. We will publicly release our codebase, pre-trained model, and training logs to promote open science in speech foundation models.【5】SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech标题:SECP:一种基于语音增强的可扩展清洁语音获取流水线作者:Adam Sabra,Cyprian Wronka,Michelle Mao,Samer Hijazi备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024摘要:随着越来越多的语音技术依赖于有监督的深度学习方法,将干净的语音作为基础事实,需要一种方法来大规模地加载所述语音。然而,这种方法需要最大限度地减少对人类听力和注释的依赖,仅在需要时需要人工参与。在本文中,我们解决这个问题,概述了语音增强为基础的策展管道(SECP),作为一个框架,板载干净的语音。然后,这个干净的语音可以训练语音增强模型,该模型可以进一步细化原始数据集,从而关闭迭代循环。通过运行两轮迭代,我们观察到,根据本文中使用的指标$\Delta_{PESQ}$,用作地面实况的增强输出不会降低模型性能。我们还表明,通过比较平均意见得分(CMOS)的主观测试,最高和最低限度的细化数据是感知优于原始数据。摘要:As more speech technologies rely on a supervised deep learning approach with clean speech as the ground truth, a methodology to onboard said speech at scale is needed. However, this approach needs to minimize the dependency on human listening and annotation, only requiring a human-in-the-loop when needed. In this paper, we address this issue by outlining Speech Enhancement-based Curation Pipeline (SECP) which serves as a framework to onboard clean speech. This clean speech can then train a speech enhancement model, which can further refine the original dataset and thus close the iterative loop. By running two iterative rounds, we observe that enhanced output used as ground truth does not degrade model performance according to $\Delta_{PESQ}$, a metric used in this paper. We also show through comparative mean opinion score (CMOS) based subjective tests that the highest and lowest bound of refined data is perceptually better than the original data.【6】On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models作者:Miri Varshavsky Hassid,Roy Hirsch,Regev Cohen,Tomer Golany,Daniel Freedman,Ehud Rivlin摘要:将去噪扩散模型(DDMs)应用于文语转换(TTS)领域的研究正在兴起,为合成高质量语音提供了重要的价值。虽然它们表现出令人印象深刻的音频质量,但它们的语义能力的程度是未知的,并且控制它们的合成语音的声音特性仍然是一个挑战。受图像合成最新进展的启发,我们探索了冻结TTS模型的潜在空间,该空间由DDM的去噪器的潜在瓶颈激活组成。我们确定,这个空间包含丰富的语义信息,并概述了几种新的方法来寻找语义方向,监督和无监督。然后,我们将展示这些如何实现现成的音频编辑,而无需任何进一步的培训,架构更改或数据要求。我们提供了编辑音频的语义和声学质量的证据,并提供了补充样本:https://latent-analysis-grad-tts.github.io/speech-samples/。摘要:The incorporation of Denoising Diffusion Models (DDMs) in the Text-to-Speech (TTS) domain is rising, providing great value in synthesizing high quality speech. Although they exhibit impressive audio quality, the extent of their semantic capabilities is unknown, and controlling their synthesized speech's vocal properties remains a challenge. Inspired by recent advances in image synthesis, we explore the latent space of frozen TTS models, which is composed of the latent bottleneck activations of the DDM's denoiser. We identify that this space contains rich semantic information, and outline several novel methods for finding semantic directions within it, both supervised and unsupervised. We then demonstrate how these enable off-the-shelf audio editing, without any further training, architectural changes or data requirements. We present evidence of the semantic and acoustic qualities of the edited audio, and provide supplemental samples: https://latent-analysis-grad-tts.github.io/speech-samples/.【7】Towards audio language modeling - an overview作者:Haibin Wu,Xuanjun Chen,Yi-Cheng Lin,Kai-wei Chang,Ho-Lam Chung,Alexander H. Liu,Hung-yi Lee摘要:神经音频编解码器最初被引入以将音频数据压缩成紧凑代码以减少传输延迟。研究人员最近发现编解码器作为合适的标记器的潜力,用于将连续音频转换为离散代码,可用于开发音频语言模型(LM)。已经开发了许多高性能神经音频编解码器和基于编解码器的LM。本文旨在对神经音频编解码器模型和基于编解码器的LM进行全面而系统的概述。摘要:Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as suitable tokenizers for converting continuous audio into discrete codes, which can be employed to develop audio language models (LMs). Numerous high-performance neural audio codecs and codec-based LMs have been developed. The paper aims to provide a thorough and systematic overview of the neural audio codec models and codec-based LMs.
【8】Probing Self-supervised Learning Models with Target Speech Extraction作者:Junyi Peng,Marc Delcroix,Tsubasa Ochiai,Oldrich Plchot,Takanori Ashihara,Shoko Araki,Jan Cernocky备注:Accepted to ICASSP 2024, Self-supervision in Audio, Speech, and Beyond (SASB) workshop摘要:大规模预训练的自监督学习(SSL)模型在语音相关任务中表现出显着的进步。然而,这些模型在复杂的多说话者场景中的利用,例如在混合物中提取目标说话者,尚未得到充分评估。在本文中,我们引入目标语音提取(TSE)作为一种新的下游任务,以评估预训练SSL模型的特征提取能力。TSE独特地需要说话人识别和语音分离,将其与语音处理通用性能基准(SUPERB)评估中的其他任务区分开来。具体来说,我们提出了一个TSE下游模型组成的两个轻量级的面向任务的模块基于相同的冻结SSL模型。一个模块作为一个扬声器编码器的功能,以获得目标扬声器的信息,从注册语音,而另一个估计目标扬声器的掩模,以提取其语音的混合。Libri2mix数据集上的实验结果揭示了TSE下游任务与探测SSL模型的相关性,因为它的性能不能简单地从其他相关任务(如说话人验证和分离)中推断出来。摘要:Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-talker scenarios, such as extracting a target speaker in a mixture, is yet to be fully evaluated. In this paper, we introduce target speech extraction (TSE) as a novel downstream task to evaluate the feature extraction capabilities of pre-trained SSL models. TSE uniquely requires both speaker identification and speech separation, distinguishing it from other tasks in the Speech processing Universal PERformance Benchmark (SUPERB) evaluation. Specifically, we propose a TSE downstream model composed of two lightweight task-oriented modules based on the same frozen SSL model. One module functions as a speaker encoder to obtain target speaker information from an enrollment speech, while the other estimates the target speaker's mask to extract its speech from the mixture. Experimental results on the Libri2mix datasets reveal the relevance of the TSE downstream task to probe SSL models, as its performance cannot be simply deduced from other related tasks such as speaker verification and separation.【9】Target Speech Extraction with Pre-trained Self-supervised Learning Models作者:Junyi Peng,Marc Delcroix,Tsubasa Ochiai,Oldrich Plchot,Shoko Araki,Jan Cernocky备注:Accepted to ICASSP 2024摘要:预训练的自监督学习(SSL)模型在各种语音任务中取得了显着的成功。然而,它们在目标语音提取(TSE)中的潜力尚未得到充分利用。TSE的目标是在注册话语的引导下从混合语音中提取目标说话人的语音。我们在TSE框架内出于两个目的利用预训练的SSL模型,即,以处理输入混合并从登记中导出说话者嵌入。在本文中,我们专注于如何有效地使用SSL模型的TSE。我们首先介绍了一种新的TSE下游任务的SUPERB原则。这个简单的实验显示了TSE的SSL模型的潜力,但提取性能仍然远远落后于最先进的。然后,我们扩展了一个强大的TSE架构,结合两个基于SSL的模块:自适应输入增强器(AIE)和扬声器编码器。具体而言,所提出的AIE通过渐进式上采样调整CNN编码器和Transformer块的时间分辨率来利用来自CNN编码器的中间表示,从而捕获细粒度和分层特征。我们的方法优于当前TSE系统,在LibriMix上实现了14.0 dB的SI-SDR改进。此外,我们还可以通过微调包括SSL模型参数在内的整个模型,将性能进一步提高0.7 dB。摘要:Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exploit pre-trained SSL models for two purposes within a TSE framework, i.e., to process the input mixture and to derive speaker embeddings from the enrollment. In this paper, we focus on how to effectively use SSL models for TSE. We first introduce a novel TSE downstream task following the SUPERB principles. This simple experiment shows the potential of SSL models for TSE, but extraction performance remains far behind the state-of-the-art. We then extend a powerful TSE architecture by incorporating two SSL-based modules: an Adaptive Input Enhancer (AIE) and a speaker encoder. Specifically, the proposed AIE utilizes intermediate representations from the CNN encoder by adjusting the time resolution of CNN encoder and transformer blocks through progressive upsampling, capturing both fine-grained and hierarchical features. Our method outperforms current TSE systems achieving a SI-SDR improvement of 14.0 dB on LibriMix. Moreover, we can further improve performance by 0.7 dB by fine-tuning the whole model including the SSL model parameters.【10】HiRIS: an Airborne Sonar Sensor with a 1024 Channel Microphone Array for In-Air Acoustic Imaging标题:HIRIS:一种用于空中声成像的1024通道麦克风阵列机载声纳传感器作者:Dennis Laurijssen,Walter Daems,Jan Steckel摘要:使用超声的机载3D成像是用于恶劣环境中的机器人应用的有前途的感测模态。在过去的十年中,在文献中已经提出了几个高性能系统。这些传感器中的大多数使用减小孔径的麦克风阵列,导致在所得到的声学图像中的伪影。本文提出了一种新型的空气中的超声波传感器,它采用1024麦克风,在一个32 × 32均匀的矩形阵列,结合分布式嵌入式硬件设计来执行数据采集。该传感器采用宽带最小方差无失真响应(MVDR)波束形成器和前后向空间平滑(FB-SS)技术,能够以高达70 dB的主瓣与旁瓣比创建具有高角度精度的全额叶2D和3D超声图像。本文描述了获得这种高度详细的声学图像所需的硬件基础设施,以及将原始声学数据转换为所述图像所需的信号处理链。利用这种新型的高分辨率超声成像传感器,我们希望通过利用这种几乎无伪影的成像方式来研究被动和主动机载超声传感的局限性。摘要:Airborne 3D imaging using ultrasound is a promising sensing modality for robotic applications in harsh environments. Over the last decade, several high-performance systems have been proposed in the literature. Most of these sensors use a reduced aperture microphone array, leading to artifacts in the resulting acoustic images. This paper presents a novel in-air ultrasound sensor that incorporates 1024 microphones, in a 32-by- 32 uniform rectangular array, in combination with a distributed embedded hardware design to perform the data acquisition. Using a broadband Minimum Variance Distortionless Response (MVDR) beamformer with Forward-Backward Spatial Smoothing (FB-SS), the sensor is able to create both 2D and 3D ultrasound images of the full-frontal hemisphere with high angular accuracy with up to 70dB main lobe to side lobe ratio. This paper describes both the hardware infrastructure needed to obtain such highly detailed acoustical images, as well as the signal processing chain needed to convert the raw acoustic data into said images. Utilizing this novel high-resolution ultrasound imaging sensor, we wish to investigate the limits of both passive and active airborne ultrasound sensing by utilizing this virtually artifact-free imaging modality.【11】Codec-SUPERB: An In-Depth Analysis of Sound Codec Models标题:Codec-Superb:声音编解码器模型的深入分析作者:Haibin Wu,Ho-Lam Chung,Yi-Cheng Lin,Yuan-Kuei Wu,Xuanjun Chen,Yu-Chi Pai,Hsiu-Hsuan Wang,Kai-Wei Chang,Alexander H. Liu,Hung-yi Lee备注:Github: this https URL摘要:声音编解码器在最小化数据传输延迟和充当标记器方面的双重角色强调了其至关重要性。近年来,编解码器模型有了重大的发展。理想的声音编解码器应该保留内容,语言,扬声器和音频信息。然而,哪种编解码器实现最佳声音信息保存的问题仍然没有答案,因为在不同的论文中,模型在其选择的实验设置上进行评估。本研究介绍了编解码器SUPERB,一个缩写的编解码器声音处理通用性能基准。Codec-SUPERB是一个生态系统,旨在评估代表性声音应用中的编解码器模型和植根于声音领域知识的信号级指标。Codec-SUPERB通过在线排行榜简化了结果共享,促进了社区驱动的基准数据库内的协作,从而刺激了编解码器的新开发周期。此外,我们进行了深入的分析,从应用和信号的角度提供对编解码器模型的见解,与以前的编解码器论文主要集中在信号电平比较不同。最后,我们将发布代码、排行榜和数据,以加速社区内的进展。摘要:The sound codec's dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance. Recent years have witnessed significant developments in codec models. The ideal sound codec should preserve content, paralinguistics, speakers, and audio information. However, the question of which codec achieves optimal sound information preservation remains unanswered, as in different papers, models are evaluated on their selected experimental settings. This study introduces Codec-SUPERB, an acronym for Codec sound processing Universal PERformance Benchmark. It is an ecosystem designed to assess codec models across representative sound applications and signal-level metrics rooted in sound domain knowledge.Codec-SUPERB simplifies result sharing through an online leaderboard, promoting collaboration within a community-driven benchmark database, thereby stimulating new development cycles for codecs. Furthermore, we undertake an in-depth analysis to offer insights into codec models from both application and signal perspectives, diverging from previous codec papers mainly concentrating on signal-level comparisons. Finally, we will release codes, the leaderboard, and data to accelerate progress within the community.【12】EMO-SUPERB: An In-depth Look at Speech Emotion Recognition作者:Haibin Wu,Huang-Cheng Chou,Kai-Wei Chang,Lucas Goncalves,Jiawei Du,Jyh-Shing Roger Jang,Chi-Chun Lee,Hung-Yi Lee备注:webpage: this https URL摘要:语音情感识别是人机交互系统的关键技术。然而,80.77%的SER论文产生的结果无法复制。我们开发了EMO-SUPERB,EMOtion Speech Universal Perception Benchmark的缩写,旨在增强SER的开源计划。EMO-SUPERB包括一个用户友好的代码库,可以利用15个最先进的语音自监督学习模型(SSLM)对六个开源SER数据集进行详尽的评估。EMO-SUPERB通过在线排行榜简化了结果共享,促进了社区驱动基准内的协作,从而增强了SER的开发。平均而言,2.58%的注释使用自然语言进行注释。SER依赖于分类模型,无法处理自然语言,导致丢弃这些有价值的注释。我们提示ChatGPT模仿注释器,理解自然语言注释,然后重新标记数据。通过利用ChatGPT生成的标签,我们在所有设置中始终实现了3.08%的平均相对增益。摘要:Speech emotion recognition (SER) is a pivotal technology for human-computer interaction systems. However, 80.77% of SER papers yield results that cannot be reproduced. We develop EMO-SUPERB, short for EMOtion Speech Universal PERformance Benchmark, which aims to enhance open-source initiatives for SER. EMO-SUPERB includes a user-friendly codebase to leverage 15 state-of-the-art speech self-supervised learning models (SSLMs) for exhaustive evaluation across six open-source SER datasets. EMO-SUPERB streamlines result sharing via an online leaderboard, fostering collaboration within a community-driven benchmark and thereby enhancing the development of SER. On average, 2.58% of annotations are annotated using natural language. SER relies on classification models and is unable to process natural languages, leading to the discarding of these valuable annotations. We prompt ChatGPT to mimic annotators, comprehend natural language annotations, and subsequently re-label the data. By utilizing labels generated by ChatGPT, we consistently achieve an average relative gain of 3.08% across all settings.
【13】Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network标题:插件语音增强:一种受动态神经网络启发的通用语音增强框架作者:Yanan Chen,Zihao Cui,Yingying Gao,Junlan Feng,Chao Deng,Shilei Zhang摘要:部署通用神经网络进行语音增强的期望,其目的是提高不同语音处理任务的噪声鲁棒性,由于静态语音增强框架内对下游模块中的预期语音缺乏认识,因此面临挑战。这些限制阻碍了静态语音增强方法在实现一系列语音处理任务的最佳性能方面的有效性,从而挑战了普遍适用性的概念。实现通用语音增强的根本问题在于有效地将下游模块的特征通知语音增强模块。在这项研究中,我们提出了一种新的加权预测方法,该方法从下游训练信息中显式学习任务关系,以解决通用语音增强的核心挑战。我们发现了决定是否采用数据增强技术作为关键下游训练信息的作用。该决定显著影响预期语音和语音增强模块的性能。此外,我们介绍了一种新的语音增强网络,插件语音增强(插件SE)。Plugin-SE是一个动态神经网络,包括语音增强模块、门模块和权重预测模块。实验结果表明,所提出的插件SE方法是竞争力或优于其他联合训练方法在各种下游任务。摘要:The expectation to deploy a universal neural network for speech enhancement, with the aim of improving noise robustness across diverse speech processing tasks, faces challenges due to the existing lack of awareness within static speech enhancement frameworks regarding the expected speech in downstream modules. These limitations impede the effectiveness of static speech enhancement approaches in achieving optimal performance for a range of speech processing tasks, thereby challenging the notion of universal applicability. The fundamental issue in achieving universal speech enhancement lies in effectively informing the speech enhancement module about the features of downstream modules. In this study, we present a novel weighting prediction approach, which explicitly learns the task relationships from downstream training information to address the core challenge of universal speech enhancement. We found the role of deciding whether to employ data augmentation techniques as crucial downstream training information. This decision significantly impacts the expected speech and the performance of the speech enhancement module. Moreover, we introduce a novel speech enhancement network, the Plugin Speech Enhancement (Plugin-SE). The Plugin-SE is a dynamic neural network that includes the speech enhancement module, gate module, and weight prediction module. Experimental results demonstrate that the proposed Plugin-SE approach is competitive or superior to other joint training methods across various downstream tasks.
【1】Towards audio language modeling - an overview作者:Haibin Wu,Xuanjun Chen,Yi-Cheng Lin,Kai-wei Chang,Ho-Lam Chung,Alexander H. Liu,Hung-yi Lee摘要:神经音频编解码器最初被引入以将音频数据压缩成紧凑代码以减少传输延迟。研究人员最近发现编解码器作为合适的标记器的潜力,用于将连续音频转换为离散代码,可用于开发音频语言模型(LM)。已经开发了许多高性能神经音频编解码器和基于编解码器的LM。本文旨在对神经音频编解码器模型和基于编解码器的LM进行全面而系统的概述。摘要:Neural audio codecs are initially introduced to compress audio data into compact codes to reduce transmission latency. Researchers recently discovered the potential of codecs as suitable tokenizers for converting continuous audio into discrete codes, which can be employed to develop audio language models (LMs). Numerous high-performance neural audio codecs and codec-based LMs have been developed. The paper aims to provide a thorough and systematic overview of the neural audio codec models and codec-based LMs.
【2】Probing Self-supervised Learning Models with Target Speech Extraction作者:Junyi Peng,Marc Delcroix,Tsubasa Ochiai,Oldrich Plchot,Takanori Ashihara,Shoko Araki,Jan Cernocky备注:Accepted to ICASSP 2024, Self-supervision in Audio, Speech, and Beyond (SASB) workshop摘要:大规模预训练的自监督学习(SSL)模型在语音相关任务中表现出显着的进步。然而,这些模型在复杂的多说话者场景中的利用,例如在混合物中提取目标说话者,尚未得到充分评估。在本文中,我们引入目标语音提取(TSE)作为一种新的下游任务,以评估预训练SSL模型的特征提取能力。TSE独特地需要说话人识别和语音分离,将其与语音处理通用性能基准(SUPERB)评估中的其他任务区分开来。具体来说,我们提出了一个TSE下游模型组成的两个轻量级的面向任务的模块基于相同的冻结SSL模型。一个模块作为一个扬声器编码器的功能,以获得目标扬声器的信息,从注册语音,而另一个估计目标扬声器的掩模,以提取其语音的混合。Libri2mix数据集上的实验结果揭示了TSE下游任务与探测SSL模型的相关性,因为它的性能不能简单地从其他相关任务(如说话人验证和分离)中推断出来。摘要:Large-scale pre-trained self-supervised learning (SSL) models have shown remarkable advancements in speech-related tasks. However, the utilization of these models in complex multi-talker scenarios, such as extracting a target speaker in a mixture, is yet to be fully evaluated. In this paper, we introduce target speech extraction (TSE) as a novel downstream task to evaluate the feature extraction capabilities of pre-trained SSL models. TSE uniquely requires both speaker identification and speech separation, distinguishing it from other tasks in the Speech processing Universal PERformance Benchmark (SUPERB) evaluation. Specifically, we propose a TSE downstream model composed of two lightweight task-oriented modules based on the same frozen SSL model. One module functions as a speaker encoder to obtain target speaker information from an enrollment speech, while the other estimates the target speaker's mask to extract its speech from the mixture. Experimental results on the Libri2mix datasets reveal the relevance of the TSE downstream task to probe SSL models, as its performance cannot be simply deduced from other related tasks such as speaker verification and separation.【3】Target Speech Extraction with Pre-trained Self-supervised Learning Models作者:Junyi Peng,Marc Delcroix,Tsubasa Ochiai,Oldrich Plchot,Shoko Araki,Jan Cernocky备注:Accepted to ICASSP 2024摘要:预训练的自监督学习(SSL)模型在各种语音任务中取得了显着的成功。然而,它们在目标语音提取(TSE)中的潜力尚未得到充分利用。TSE的目标是在注册话语的引导下从混合语音中提取目标说话人的语音。我们在TSE框架内出于两个目的利用预训练的SSL模型,即,以处理输入混合并从登记中导出说话者嵌入。在本文中,我们专注于如何有效地使用SSL模型的TSE。我们首先介绍了一种新的TSE下游任务的SUPERB原则。这个简单的实验显示了TSE的SSL模型的潜力,但提取性能仍然远远落后于最先进的。然后,我们扩展了一个强大的TSE架构,结合两个基于SSL的模块:自适应输入增强器(AIE)和扬声器编码器。具体而言,所提出的AIE通过渐进式上采样调整CNN编码器和Transformer块的时间分辨率来利用来自CNN编码器的中间表示,从而捕获细粒度和分层特征。我们的方法优于当前TSE系统,在LibriMix上实现了14.0 dB的SI-SDR改进。此外,我们还可以通过微调包括SSL模型参数在内的整个模型,将性能进一步提高0.7 dB。摘要:Pre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exploit pre-trained SSL models for two purposes within a TSE framework, i.e., to process the input mixture and to derive speaker embeddings from the enrollment. In this paper, we focus on how to effectively use SSL models for TSE. We first introduce a novel TSE downstream task following the SUPERB principles. This simple experiment shows the potential of SSL models for TSE, but extraction performance remains far behind the state-of-the-art. We then extend a powerful TSE architecture by incorporating two SSL-based modules: an Adaptive Input Enhancer (AIE) and a speaker encoder. Specifically, the proposed AIE utilizes intermediate representations from the CNN encoder by adjusting the time resolution of CNN encoder and transformer blocks through progressive upsampling, capturing both fine-grained and hierarchical features. Our method outperforms current TSE systems achieving a SI-SDR improvement of 14.0 dB on LibriMix. Moreover, we can further improve performance by 0.7 dB by fine-tuning the whole model including the SSL model parameters.【4】HiRIS: an Airborne Sonar Sensor with a 1024 Channel Microphone Array for In-Air Acoustic Imaging标题:HIRIS:一种用于空中声成像的1024通道麦克风阵列机载声纳传感器作者:Dennis Laurijssen,Walter Daems,Jan Steckel摘要:使用超声的机载3D成像是用于恶劣环境中的机器人应用的有前途的感测模态。在过去的十年中,在文献中已经提出了几个高性能系统。这些传感器中的大多数使用减小孔径的麦克风阵列,导致在所得到的声学图像中的伪影。本文提出了一种新型的空气中的超声波传感器,它采用1024麦克风,在一个32 × 32均匀的矩形阵列,结合分布式嵌入式硬件设计来执行数据采集。该传感器采用宽带最小方差无失真响应(MVDR)波束形成器和前后向空间平滑(FB-SS)技术,能够以高达70 dB的主瓣与旁瓣比创建具有高角度精度的全额叶2D和3D超声图像。本文描述了获得这种高度详细的声学图像所需的硬件基础设施,以及将原始声学数据转换为所述图像所需的信号处理链。利用这种新型的高分辨率超声成像传感器,我们希望通过利用这种几乎无伪影的成像方式来研究被动和主动机载超声传感的局限性。摘要:Airborne 3D imaging using ultrasound is a promising sensing modality for robotic applications in harsh environments. Over the last decade, several high-performance systems have been proposed in the literature. Most of these sensors use a reduced aperture microphone array, leading to artifacts in the resulting acoustic images. This paper presents a novel in-air ultrasound sensor that incorporates 1024 microphones, in a 32-by- 32 uniform rectangular array, in combination with a distributed embedded hardware design to perform the data acquisition. Using a broadband Minimum Variance Distortionless Response (MVDR) beamformer with Forward-Backward Spatial Smoothing (FB-SS), the sensor is able to create both 2D and 3D ultrasound images of the full-frontal hemisphere with high angular accuracy with up to 70dB main lobe to side lobe ratio. This paper describes both the hardware infrastructure needed to obtain such highly detailed acoustical images, as well as the signal processing chain needed to convert the raw acoustic data into said images. Utilizing this novel high-resolution ultrasound imaging sensor, we wish to investigate the limits of both passive and active airborne ultrasound sensing by utilizing this virtually artifact-free imaging modality.
【5】Codec-SUPERB: An In-Depth Analysis of Sound Codec Models标题:Codec-Superb:声音编解码器模型的深入分析作者:Haibin Wu,Ho-Lam Chung,Yi-Cheng Lin,Yuan-Kuei Wu,Xuanjun Chen,Yu-Chi Pai,Hsiu-Hsuan Wang,Kai-Wei Chang,Alexander H. Liu,Hung-yi Lee备注:Github: this https URL摘要:声音编解码器在最小化数据传输延迟和充当标记器方面的双重角色强调了其至关重要性。近年来,编解码器模型有了重大的发展。理想的声音编解码器应该保留内容,语言,扬声器和音频信息。然而,哪种编解码器实现最佳声音信息保存的问题仍然没有答案,因为在不同的论文中,模型在其选择的实验设置上进行评估。本研究介绍了编解码器SUPERB,一个缩写的编解码器声音处理通用性能基准。Codec-SUPERB是一个生态系统,旨在评估代表性声音应用中的编解码器模型和植根于声音领域知识的信号级指标。Codec-SUPERB通过在线排行榜简化了结果共享,促进了社区驱动的基准数据库内的协作,从而刺激了编解码器的新开发周期。此外,我们进行了深入的分析,从应用和信号的角度提供对编解码器模型的见解,与以前的编解码器论文主要集中在信号电平比较不同。最后,我们将发布代码、排行榜和数据,以加速社区内的进展。摘要:The sound codec's dual roles in minimizing data transmission latency and serving as tokenizers underscore its critical importance. Recent years have witnessed significant developments in codec models. The ideal sound codec should preserve content, paralinguistics, speakers, and audio information. However, the question of which codec achieves optimal sound information preservation remains unanswered, as in different papers, models are evaluated on their selected experimental settings. This study introduces Codec-SUPERB, an acronym for Codec sound processing Universal PERformance Benchmark. It is an ecosystem designed to assess codec models across representative sound applications and signal-level metrics rooted in sound domain knowledge.Codec-SUPERB simplifies result sharing through an online leaderboard, promoting collaboration within a community-driven benchmark database, thereby stimulating new development cycles for codecs. Furthermore, we undertake an in-depth analysis to offer insights into codec models from both application and signal perspectives, diverging from previous codec papers mainly concentrating on signal-level comparisons. Finally, we will release codes, the leaderboard, and data to accelerate progress within the community.【6】EMO-SUPERB: An In-depth Look at Speech Emotion Recognition作者:Haibin Wu,Huang-Cheng Chou,Kai-Wei Chang,Lucas Goncalves,Jiawei Du,Jyh-Shing Roger Jang,Chi-Chun Lee,Hung-Yi Lee备注:webpage: this https URL摘要:语音情感识别是人机交互系统的关键技术。然而,80.77%的SER论文产生的结果无法复制。我们开发了EMO-SUPERB,EMOtion Speech Universal Perception Benchmark的缩写,旨在增强SER的开源计划。EMO-SUPERB包括一个用户友好的代码库,可以利用15个最先进的语音自监督学习模型(SSLM)对六个开源SER数据集进行详尽的评估。EMO-SUPERB通过在线排行榜简化了结果共享,促进了社区驱动基准内的协作,从而增强了SER的开发。平均而言,2.58%的注释使用自然语言进行注释。SER依赖于分类模型,无法处理自然语言,导致丢弃这些有价值的注释。我们提示ChatGPT模仿注释器,理解自然语言注释,然后重新标记数据。通过利用ChatGPT生成的标签,我们在所有设置中始终实现了3.08%的平均相对增益。摘要:Speech emotion recognition (SER) is a pivotal technology for human-computer interaction systems. However, 80.77% of SER papers yield results that cannot be reproduced. We develop EMO-SUPERB, short for EMOtion Speech Universal PERformance Benchmark, which aims to enhance open-source initiatives for SER. EMO-SUPERB includes a user-friendly codebase to leverage 15 state-of-the-art speech self-supervised learning models (SSLMs) for exhaustive evaluation across six open-source SER datasets. EMO-SUPERB streamlines result sharing via an online leaderboard, fostering collaboration within a community-driven benchmark and thereby enhancing the development of SER. On average, 2.58% of annotations are annotated using natural language. SER relies on classification models and is unable to process natural languages, leading to the discarding of these valuable annotations. We prompt ChatGPT to mimic annotators, comprehend natural language annotations, and subsequently re-label the data. By utilizing labels generated by ChatGPT, we consistently achieve an average relative gain of 3.08% across all settings.【7】Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network标题:插件语音增强:一种受动态神经网络启发的通用语音增强框架作者:Yanan Chen,Zihao Cui,Yingying Gao,Junlan Feng,Chao Deng,Shilei Zhang摘要:部署通用神经网络进行语音增强的期望,其目的是提高不同语音处理任务的噪声鲁棒性,由于静态语音增强框架内对下游模块中的预期语音缺乏认识,因此面临挑战。这些限制阻碍了静态语音增强方法在实现一系列语音处理任务的最佳性能方面的有效性,从而挑战了普遍适用性的概念。实现通用语音增强的根本问题在于有效地将下游模块的特征通知语音增强模块。在这项研究中,我们提出了一种新的加权预测方法,该方法从下游训练信息中显式学习任务关系,以解决通用语音增强的核心挑战。我们发现了决定是否采用数据增强技术作为关键下游训练信息的作用。该决定显著影响预期语音和语音增强模块的性能。此外,我们介绍了一种新的语音增强网络,插件语音增强(插件SE)。Plugin-SE是一个动态神经网络,包括语音增强模块、门模块和权重预测模块。实验结果表明,所提出的插件SE方法是竞争力或优于其他联合训练方法在各种下游任务。摘要:The expectation to deploy a universal neural network for speech enhancement, with the aim of improving noise robustness across diverse speech processing tasks, faces challenges due to the existing lack of awareness within static speech enhancement frameworks regarding the expected speech in downstream modules. These limitations impede the effectiveness of static speech enhancement approaches in achieving optimal performance for a range of speech processing tasks, thereby challenging the notion of universal applicability. The fundamental issue in achieving universal speech enhancement lies in effectively informing the speech enhancement module about the features of downstream modules. In this study, we present a novel weighting prediction approach, which explicitly learns the task relationships from downstream training information to address the core challenge of universal speech enhancement. We found the role of deciding whether to employ data augmentation techniques as crucial downstream training information. This decision significantly impacts the expected speech and the performance of the speech enhancement module. Moreover, we introduce a novel speech enhancement network, the Plugin Speech Enhancement (Plugin-SE). The Plugin-SE is a dynamic neural network that includes the speech enhancement module, gate module, and weight prediction module. Experimental results demonstrate that the proposed Plugin-SE approach is competitive or superior to other joint training methods across various downstream tasks.【8】Not All Weights Are Created Equal: Enhancing Energy Efficiency in On-Device Streaming Speech Recognition标题:并非所有权重都是相等的:提高设备流传输语音识别的能效作者:Yang Li,Yuan Shangguan,Yuhao Wang,Liangzhen Lai,Ernie Chang,Changsheng Zhao,Yangyang Shi,Vikas Chandra摘要:功耗在设备上流式语音识别中起着重要作用,因为它直接影响用户体验。本研究深入探讨语音识别模型中的权重参数如何影响这些模型的整体功耗。我们发现,权重参数对功耗的影响各不相同,受调用频率及其在内存中的位置等因素的影响。基于这一见解,我们开发了旨在优化设备上语音识别模型的设计指南。这些准则侧重于在不影响准确性的情况下最大限度地减少功耗。我们的方法,它采用有针对性的压缩的基础上的权重参数的变化的敏感性,表现出优越的性能相比,国家的最先进的压缩方法。它实现了高达47%的能源使用减少,同时保持类似的模型精度和提高实时系数。摘要:Power consumption plays an important role in on-device streaming speech recognition, as it has a direct impact on the user experience. This study delves into how weight parameters in speech recognition models influence the overall power consumption of these models. We discovered that the impact of weight parameters on power consumption varies, influenced by factors including how often they are invoked and their placement in memory. Armed with this insight, we developed design guidelines aimed at optimizing on-device speech recognition models. These guidelines focus on minimizing power use without substantially affecting accuracy. Our method, which employs targeted compression based on the varying sensitivities of weight parameters, demonstrates superior performance compared to state-of-the-art compression methods. It achieves a reduction in energy usage of up to 47% while maintaining similar model accuracy and improving the real-time factor.【9】Advancing Large Language Models to Capture Varied Speaking Styles and Respond Properly in Spoken Conversations标题:提出大型语言模型,以捕捉不同的说话风格,并在口语对话中做出正确反应作者:Guan-Ting Lin,Cheng-Han Chiang,Hung-yi Lee摘要:在口语对话中,即使两个当前的话轮是同一个句子,当他们以不同的风格说话时,他们的反应仍然可能不同。口语语体是语篇和言语模态之间最显著的区别,它包含了语言学和韵律学的信息。当使用纯文本LLM对口语对话进行建模时,纯文本LLM不能基于当前回合的说话风格给出不同的响应。在本文中,我们专注于使LLM听说话的风格,并作出适当的反应。我们的目标是教法学硕士,“即使句子是相同的,如果他们说在不同的风格,他们相应的反应可能是不同的”。由于没有合适的数据集来实现这一目标,我们收集了一个语音到语音的数据集StyleTalk,具有以下所需的特征:当两个当前的语音具有相同的内容,但以不同的风格说话时,它们的反应将是不同的。为了教LLM理解并正确应对演讲风格,我们提出了口语LLM框架,可以模拟语言内容和演讲风格。我们使用StyleTalk数据集训练Spoken-LLM,并设计了一个两阶段的训练管道,以帮助Spoken-LLM更好地学习说话风格。基于大量的实验,我们表明,口语LLM优于纯文本基线和先前的语音LLM方法。摘要:In spoken dialogue, even if two current turns are the same sentence, their responses might still differ when they are spoken in different styles. The spoken styles, containing paralinguistic and prosodic information, mark the most significant difference between text and speech modality. When using text-only LLMs to model spoken dialogue, text-only LLMs cannot give different responses based on the speaking style of the current turn. In this paper, we focus on enabling LLMs to listen to the speaking styles and respond properly. Our goal is to teach the LLM that "even if the sentences are identical if they are spoken in different styles, their corresponding responses might be different". Since there is no suitable dataset for achieving this goal, we collect a speech-to-speech dataset, StyleTalk, with the following desired characteristics: when two current speeches have the same content but are spoken in different styles, their responses will be different. To teach LLMs to understand and respond properly to the speaking styles, we propose the Spoken-LLM framework that can model the linguistic content and the speaking styles. We train Spoken-LLM using the StyleTalk dataset and devise a two-stage training pipeline to help the Spoken-LLM better learn the speaking styles. Based on extensive experiments, we show that Spoken-LLM outperforms text-only baselines and prior speech LLMs methods.
【10】SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion标题:SingVisio:歌声转换扩散模型的可视化分析作者:Liumeng Xue,Chaoren Wang,Mingxuan Wang,Xueyao Zhang,Jun Han,Zhizheng Wu摘要:在这项研究中,我们提出了Singlephon,一个交互式的视觉分析系统,旨在解释的扩散模型中使用的歌声转换。Singapolis提供了扩散模型中生成过程的可视化显示,展示了噪声频谱的逐步去噪及其转换为捕获所需歌手音色的干净频谱。该系统还有助于对不同条件(如源内容、旋律和目标音色)进行并排比较,突出显示这些条件对扩散生成过程和转换的影响。通过全面评估,新加坡证明了其在系统设计、功能、可解释性和用户友好性方面的有效性。它为不同背景的用户提供了宝贵的学习经验和对歌唱声音转换扩散模型的见解。摘要:In this study, we present SingVisio, an interactive visual analysis system that aims to explain the diffusion model used in singing voice conversion. SingVisio provides a visual display of the generation process in diffusion models, showcasing the step-by-step denoising of the noisy spectrum and its transformation into a clean spectrum that captures the desired singer's timbre. The system also facilitates side-by-side comparisons of different conditions, such as source content, melody, and target timbre, highlighting the impact of these conditions on the diffusion generation process and resulting conversions. Through comprehensive evaluations, SingVisio demonstrates its effectiveness in terms of system design, functionality, explainability, and user-friendliness. It offers users of various backgrounds valuable learning experiences and insights into the diffusion model for singing voice conversion.
【11】Guiding the underwater acoustic target recognition with interpretable contrastive learning作者:Yuan Xie,Jiawei Ren,Ji Xu摘要:由于复杂的海洋环境和多变的水下信道,从声信号中识别水下目标是一项具有挑战性的任务。虽然基于深度学习的系统已成为水声目标识别的主流方法,但它们在实际应用中因缺乏可解释性和泛化性能弱而受到批评。在这项工作中,我们应用类激活映射(CAM)生成基于频谱的识别系统的预测的视觉解释。CAM可以通过突出显示对预测贡献最大的输入特征区域来帮助理解识别模型的行为。我们的探索表明,识别模型往往侧重于低频线谱和高频周期调制信息的水下信号。基于观察,我们提出了一个可解释的对比学习(ICL)策略,采用两个编码器学习不同的重点(线频谱和调制信息)的声学特征。通过在编码器之间施加约束,所提出的策略可以提高识别系统的泛化性能。我们的实验表明,所提出的对比学习方法可以提高识别精度,并在各种水下数据库带来显着的改善。摘要:Recognizing underwater targets from acoustic signals is a challenging task owing to the intricate ocean environments and variable underwater channels. While deep learning-based systems have become the mainstream approach for underwater acoustic target recognition, they have faced criticism for their lack of interpretability and weak generalization performance in practical applications. In this work, we apply the class activation mapping (CAM) to generate visual explanations for the predictions of a spectrogram-based recognition system. CAM can help to understand the behavior of recognition models by highlighting the regions of the input features that contribute the most to the prediction. Our explorations reveal that recognition models tend to focus on the low-frequency line spectrum and high-frequency periodic modulation information of underwater signals. Based on the observation, we propose an interpretable contrastive learning (ICL) strategy that employs two encoders to learn from acoustic features with different emphases (line spectrum and modulation information). By imposing constraints between encoders, the proposed strategy can enhance the generalization performance of the recognition system. Our experiments demonstrate that the proposed contrastive learning approach can improve the recognition accuracy and bring significant improvements across various underwater databases.
【12】OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification标题:OWSM-CTC:一种用于语音识别、翻译和语言识别的开放编码器语音基础模型作者:Yifan Peng,Yui Sudo,Muhammad Shakeel,Shinji Watanabe摘要:人们对能够在单个模型中执行多个语音处理任务的大型语音模型越来越感兴趣。这些模型通常采用编解码器或仅解码器架构,因为它们在许多领域中的流行性和良好性能。然而,与非自回归模型相比,自回归模型在推理过程中可能会更慢,并且也有潜在的幻觉风险。虽然先前的研究观察到非自回归模型在小规模的某些任务中有很好的结果,但目前还不清楚它们是否可以扩展到不同语言和任务的语音到文本生成。受开放式耳语风格语音模型(OWSM)项目的启发,我们提出了OWSM-CTC,一种新的基于连接主义时间分类(CTC)的仅编码器语音基础模型。它在18万小时的公共音频数据上进行训练,用于多语言自动语音识别(ASR),语音翻译(ST)和语言识别(LID)。与编码器-解码器OWSM相比,我们的OWSM-CTC在ASR上取得了有竞争力的结果,在ST上的相对改进高达25%,同时它更健壮,推理速度快3到4倍。OWSM-CTC还以20倍的速度改进了长格式ASR结果。我们将公开发布我们的代码库、预训练模型和训练日志,以促进语音基础模型的开放科学。摘要:There has been an increasing interest in large speech models that can perform multiple speech processing tasks in a single model. Such models usually adopt the encoder-decoder or decoder-only architecture due to their popularity and good performance in many domains. However, autoregressive models can be slower during inference compared to non-autoregressive models and also have potential risks of hallucination. Though prior studies observed promising results of non-autoregressive models for certain tasks at small scales, it remains unclear if they can be scaled to speech-to-text generation in diverse languages and tasks. Inspired by the Open Whisper-style Speech Model (OWSM) project, we propose OWSM-CTC, a novel encoder-only speech foundation model based on Connectionist Temporal Classification (CTC). It is trained on 180k hours of public audio data for multilingual automatic speech recognition (ASR), speech translation (ST), and language identification (LID). Compared to encoder-decoder OWSM, our OWSM-CTC achieves competitive results on ASR and up to 25% relative improvement on ST, while it is more robust and 3 to 4 times faster for inference. OWSM-CTC also improves the long-form ASR result with 20x speed-up. We will publicly release our codebase, pre-trained model, and training logs to promote open science in speech foundation models.【13】SECP: A Speech Enhancement-Based Curation Pipeline For Scalable Acquisition Of Clean Speech标题:SECP:一种基于语音增强的可扩展清洁语音获取流水线作者:Adam Sabra,Cyprian Wronka,Michelle Mao,Samer Hijazi备注:Accepted to the International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2024摘要:随着越来越多的语音技术依赖于有监督的深度学习方法,将干净的语音作为基础事实,需要一种方法来大规模地加载所述语音。然而,这种方法需要最大限度地减少对人类听力和注释的依赖,仅在需要时需要人工参与。在本文中,我们解决这个问题,概述了语音增强为基础的策展管道(SECP),作为一个框架,板载干净的语音。然后,这个干净的语音可以训练语音增强模型,该模型可以进一步细化原始数据集,从而关闭迭代循环。通过运行两轮迭代,我们观察到,根据本文中使用的指标$\Delta_{PESQ}$,用作地面实况的增强输出不会降低模型性能。我们还表明,通过比较平均意见得分(CMOS)的主观测试,最高和最低限度的细化数据是感知优于原始数据。摘要:As more speech technologies rely on a supervised deep learning approach with clean speech as the ground truth, a methodology to onboard said speech at scale is needed. However, this approach needs to minimize the dependency on human listening and annotation, only requiring a human-in-the-loop when needed. In this paper, we address this issue by outlining Speech Enhancement-based Curation Pipeline (SECP) which serves as a framework to onboard clean speech. This clean speech can then train a speech enhancement model, which can further refine the original dataset and thus close the iterative loop. By running two iterative rounds, we observe that enhanced output used as ground truth does not degrade model performance according to $\Delta_{PESQ}$, a metric used in this paper. We also show through comparative mean opinion score (CMOS) based subjective tests that the highest and lowest bound of refined data is perceptually better than the original data.【14】On the Semantic Latent Space of Diffusion-Based Text-to-Speech Models作者:Miri Varshavsky Hassid,Roy Hirsch,Regev Cohen,Tomer Golany,Daniel Freedman,Ehud Rivlin摘要:将去噪扩散模型(DDMs)应用于文语转换(TTS)领域的研究正在兴起,为合成高质量语音提供了重要的价值。虽然它们表现出令人印象深刻的音频质量,但它们的语义能力的程度是未知的,并且控制它们的合成语音的声音特性仍然是一个挑战。受图像合成最新进展的启发,我们探索了冻结TTS模型的潜在空间,该空间由DDM的去噪器的潜在瓶颈激活组成。我们确定,这个空间包含丰富的语义信息,并概述了几种新的方法来寻找语义方向,监督和无监督。然后,我们将展示这些如何实现现成的音频编辑,而无需任何进一步的培训,架构更改或数据要求。我们提供了编辑音频的语义和声学质量的证据,并提供了补充样本:https://latent-analysis-grad-tts.github.io/speech-samples/。摘要:The incorporation of Denoising Diffusion Models (DDMs) in the Text-to-Speech (TTS) domain is rising, providing great value in synthesizing high quality speech. Although they exhibit impressive audio quality, the extent of their semantic capabilities is unknown, and controlling their synthesized speech's vocal properties remains a challenge. Inspired by recent advances in image synthesis, we explore the latent space of frozen TTS models, which is composed of the latent bottleneck activations of the DDM's denoiser. We identify that this space contains rich semantic information, and outline several novel methods for finding semantic directions within it, both supervised and unsupervised. We then demonstrate how these enable off-the-shelf audio editing, without any further training, architectural changes or data requirements. We present evidence of the semantic and acoustic qualities of the edited audio, and provide supplemental samples: https://latent-analysis-grad-tts.github.io/speech-samples/.