今日论文合集:cs.SD语音15篇,eess.AS音频处理22篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】The Solution for Temporal Sound Localisation Task of ICCV 1st Perception  Test Challenge 2023

标题:ICCV 2023年第一届感知测试挑战赛时间声音定位任务的解决方案

链接:https://arxiv.org/abs/2407.02318

作者:Yurui Huang,Yang Yang,Shou Chen,Xiangyu Wu,Qingguo Chen,Jianfeng Lu

摘要:在本文中,我们提出了一种解决方案,以提高质量的时间声音定位。我们采用多模态融合的方法来结合视觉和音频功能。使用最先进的自监督预训练网络提取高质量的视觉特征,从而实现高效的视频特征表示。同时,音频特征作为补充信息,帮助模型更好地定位声音的开始和结束。融合后的特征在多尺度Transformer中进行训练。在最终的测试数据集中,我们实现了0.33的平均精度(mAP),在该赛道中获得了第二好的性能。

摘要:In this paper, we propose a solution for improving the quality of temporal sound localization. We employ a multimodal fusion approach to combine visual and audio features. High-quality visual features are extracted using a state-of-the-art self-supervised pre-training network, resulting in efficient video feature representations. At the same time, audio features serve as complementary information to help the model better localize the start and end of sounds. The fused features are trained in a multi-scale Transformer for training. In the final test dataset, we achieved a mean average precision (mAP) of 0.33, obtaining the second-best performance in this track.


【2】 MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music  Processing
标题:MelodyT5:用于符号音乐处理的统一乐谱Transformer
链接:https://arxiv.org/abs/2407.02277
作者:Shangda Wu,Yashan Wang,Xiaobing Li,Feng Yu,Maosong Sun
备注:9 pages, 2 figures, 3 tables, accepted by ISMIR 2024
摘要:在符号音乐研究领域,由于缺乏可用的训练数据和对针对特定任务定制的模型的需求,开发可扩展系统的进展受到了显着阻碍。为了解决这些问题,我们提出了MelodyT 5,一个新的统一框架,利用编码器-解码器架构,专为符号音乐处理在ABC符号。这个框架挑战了传统的任务特定的方法,考虑各种符号音乐任务的分数到分数的转换。因此,它集成了七个以旋律为中心的任务,从生成到协调和分割,在一个单一的模型。MelodyHub是一个新策划的集合,包含超过261 K以ABC符号编码的独特旋律,并包含超过一百万个任务实例,MelodyT 5通过多任务迁移学习在符号音乐处理方面表现出卓越的性能。我们的研究结果强调了多任务迁移学习在符号音乐处理中的有效性,特别是对于数据稀缺的任务,挑战了流行的特定任务范式,并为该领域的未来探索提供了全面的数据集和框架。
摘要:In the domain of symbolic music research, the progress of developing scalable systems has been notably hindered by the scarcity of available training data and the demand for models tailored to specific tasks. To address these issues, we propose MelodyT5, a novel unified framework that leverages an encoder-decoder architecture tailored for symbolic music processing in ABC notation. This framework challenges the conventional task-specific approach, considering various symbolic music tasks as score-to-score transformations. Consequently, it integrates seven melody-centric tasks, from generation to harmonization and segmentation, within a single model. Pre-trained on MelodyHub, a newly curated collection featuring over 261K unique melodies encoded in ABC notation and encompassing more than one million task instances, MelodyT5 demonstrates superior performance in symbolic music processing via multi-task transfer learning. Our findings highlight the efficacy of multi-task transfer learning in symbolic music processing, particularly for data-scarce tasks, challenging the prevailing task-specific paradigms and offering a comprehensive dataset and framework for future explorations in this domain.

【3】 SOAF: Scene Occlusion-aware Neural Acoustic Field
标题:SOAF:场景遮挡感知神经声学场
链接:https://arxiv.org/abs/2407.02264
作者:Huiyu Gao,Jiahao Ma,David Ahmedt-Aristizabal,Chuong Nguyen,Miaomiao Liu
摘要:None
摘要:This paper tackles the problem of novel view audio-visual synthesis along an arbitrary trajectory in an indoor scene, given the audio-video recordings from other known trajectories of the scene. Existing methods often overlook the effect of room geometry, particularly wall occlusion to sound propagation, making them less accurate in multi-room environments. In this work, we propose a new approach called Scene Occlusion-aware Acoustic Field (SOAF) for accurate sound generation. Our approach derives a prior for sound energy field using distance-aware parametric sound-propagation modelling and then transforms it based on scene transmittance learned from the input video. We extract features from the local acoustic field centred around the receiver using a Fibonacci Sphere to generate binaural audio for novel views with a direction-aware attention mechanism. Extensive experiments on the real dataset~\emph{RWAVS} and the synthetic dataset~\emph{SoundSpaces} demonstrate that our method outperforms previous state-of-the-art techniques in audio generation. Project page: https://github.com/huiyu-gao/SOAF/.

【4】 Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference  Optimization
标题:具有反向推理优化的鲁棒Zero-Shot文本到语音合成
链接:https://arxiv.org/abs/2407.02243
作者:Yuchen Hu,Chen Chen,Siyin Wang,Eng Siong Chng,Chao Zhang备注:12 pages, Work in progress
摘要:在本文中,我们提出了反向推理优化(RIO),一个简单而有效的方法,旨在提高鲁棒性的自回归模型为基础的zero-shot文本到语音(TTS)系统,采用强化学习从人类反馈(RLHF)。为了评估语音质量的TTS系统所产生的没有人的注释,RIO介绍了一种新的概念,称为反向推理的基础上贝叶斯原理,这表明,高质量的生成的语音应该能够被用来作为一个提示,为后续的一代使用相同的TTS模型。通过利用反向推理作为从TTS系统自身生成的语音样本中选择RLHF中使用的样本的标准,RIO将随后的优化转向增强TTS鲁棒性的方向。RIO框架包括采样、自动标注和学习,避免了对奖励模型或成对偏好数据的需要,并通过减少训练和推理条件之间的差异来显著提高zero-shot TTS性能的稳定性。我们的实验结果验证了RIO可以有效地提高主观和客观指标,包括平均意见分数,单词错误率和说话人相似度。值得注意的是,RIO还可以将不良输出的发生率降低到几乎为零,与使用地面实况语音作为提示时的鲁棒性相媲美。摘要:In this paper, we propose reverse inference optimization (RIO), a simple and effective method designed to enhance the robustness of autoregressive-model-based zero-shot text-to-speech (TTS) systems using reinforcement learning from human feedback (RLHF). To assess the quality of speech produced by the TTS system without human annotations, RIO introduces a novel concept termed as reverse inference based on the Bayesian principle, which suggests that a high-quality generated speech should be able to be used as a prompt for subsequent generation using the same TTS model. By leveraging reverse inference as the standard to select exemplars used in RLHF from the speech samples generated by the TTS system itself, RIO steers the subsequent optimization towards a direction of enhancing the TTS robustness. The RIO framework, comprising sampling, automatic annotating, and learning, obviates the need for a reward model or pairwise preference data, and significantly improves the stability of zero-shot TTS performance by reducing the discrepancies between training and inference conditions. Our experimental results verify that RIO can effectively improve both subjective and objective metrics, including mean opinion scores, word error rates, and speaker similarity. Remarkably, RIO can also diminish the incidence of bad outputs to nearly zero percent, rivalling the robustness when using ground-truth speech as the prompt.

【5】 GMM-ResNet2: Ensemble of Group ResNet Networks for Synthetic Speech  Detection
标题:GMM-ResNet 2:用于合成语音检测的群组ResNet网络的扩展
链接:https://arxiv.org/abs/2407.02170
作者:Zhenchun Lei,Hui Yan,Changhong Liu,Yong Zhou,Minglei Ma
摘要:深度学习模型广泛用于说话人识别和欺骗语音检测。我们提出了GMM-ResNet 2的合成语音检测。GMM-ResNet 2与之前的GMM-ResNet模型相比,有四个方面的改进。首先,不同阶数的高斯分布具有不同的平滑逼近能力,多阶高斯分布用于提取多尺度对数高斯概率特征。其次,分组技术被用来提高分类精度,暴露组基数,同时减少参数的数量和训练时间。使用平均方法通过所有组分类器输出的集成获得最终分数。第三,通过包括一个激活函数和一个批归一化层来改进残差块。最后,提出了一个集成感知的损失函数,以整合所有集成成员的独立损失函数。在ASVspoof 2019 LA任务中,GMM-ResNet 2实现了0.0227的最小t-DCF和0.79\%的EER。在ASVspoof 2021 LA任务中,GMM-ResNet 2实现了0.2362的最小t-DCF和2.19\%的EER,与LFCC-LCNN基线相比,相对减少了31.4\%和76.3\%。
摘要:Deep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly, the different order GMMs have different capabilities to form smooth approximations to the feature distribution, and multiple GMMs are used to extract multi-scale Log Gaussian Probability features. Secondly, the grouping technique is used to improve the classification accuracy by exposing the group cardinality while reducing both the number of parameters and the training time. The final score is obtained by ensemble of all group classifier outputs using the averaging method. Thirdly, the residual block is improved by including one activation function and one batch normalization layer. Finally, an ensemble-aware loss function is proposed to integrate the independent loss functions of all ensemble members. On the ASVspoof 2019 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.0227 and an EER of 0.79\%. On the ASVspoof 2021 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.2362 and an EER of 2.19\%, and represents a relative reductions of 31.4\% and 76.3\% compared with the LFCC-LCNN baseline.

【6】 Towards Training Music Taggers on Synthetic Data
标题:利用合成数据训练音乐标记者
链接:https://arxiv.org/abs/2407.02156
作者:Nadine Kroher,Steven Manangu,Aggelos Pikrakis
备注:6 pages, 3 figures, accepted to 21st International Conference on Content-based Multimedia Indexing (CBMI) 2024, code available this https URL
摘要:大多数当代音乐标记系统依赖于大量的注释数据。作为一种替代方案,我们调查在何种程度上合成生成的音乐摘录可以提高标签系统时,只有小的注释集合。为此,我们发布了GTZAN-synth,这是一个合成数据集,它遵循了著名的GTZAN数据集的分类法,同时数据量是它的十倍。我们首先观察到,简单地将这个合成数据集添加到GTZAN的训练分割中并不会导致性能提高。然后,我们继续调查领域适应,迁移学习和微调策略,为手头的任务,并得出结论,最后两个选项产生的准确性增加。总体而言,所提出的方法可以被认为是在一个有前途的领域,为未来的研究的第一个指南。
摘要:Most contemporary music tagging systems rely on large volumes of annotated data. As an alternative, we investigate the extent to which synthetically generated music excerpts can improve tagging systems when only small annotated collections are available. To this end, we release GTZAN-synth, a synthetic dataset that follows the taxonomy of the well-known GTZAN dataset while being ten times larger in data volume. We first observe that simply adding this synthetic dataset to the training split of GTZAN does not result into performance improvements. We then proceed to investigating domain adaptation, transfer learning and fine-tuning strategies for the task at hand and draw the conclusion that the last two options yield an increase in accuracy. Overall, the proposed approach can be considered as a first guide in a promising field for future research.

【7】 An End-to-End Speech Summarization Using Large Language Model
标题:使用大型语言模型的端到端语音摘要
链接:https://arxiv.org/abs/2407.02005
作者:Hengchao Shang,Zongyao Li,Jiaxin Guo,Shaojun Li,Zhiqiang Rao,Yuanchang Luo,Daimeng Wei,Hao Yang
备注:InterSpeech 2024
摘要:摘要语音摘要(SSum)旨在从语音内容中生成类似于人类的文本摘要。它在处理长语音输入和捕获长语音输入与短文本摘要之间复杂的跨模态映射方面遇到了困难。大型语言模型(LLM)和多模态信息融合的研究为解决这些挑战提供了新的见解。在本文中,我们提出了一个端到端的SSUM模型,利用Q-Former作为音频文本模态的连接器,并采用LLM直接从语音特征生成文本摘要。我们采用了一种多阶段的训练方法,包括基于LLM的ASR和文本摘要(TSum)任务作为辅助任务。ASR任务用于对齐特征空间并增强LLM处理较长语音的能力。然后,我们利用课程学习策略,以促进模型的过渡,从TSum到SSum。最后,我们的模型在How-2数据集上实现了具有竞争力的性能。
摘要:Abstractive Speech Summarization (SSum) aims to generate human-like text summaries from spoken content. It encounters difficulties in handling long speech input and capturing the intricate cross-modal mapping between long speech inputs and short text summaries. Research on large language models (LLMs) and multimodal information fusion has provided new insights for addressing these challenges. In this paper, we propose an end-to-end SSum model that utilizes Q-Former as a connector for the audio-text modality and employs LLMs to generate text summaries directly from speech features. We adopt a multi-stage training approach that includes LLM based ASR and Text Summarization (TSum) tasks as auxiliary tasks. ASR tasks are used to align feature spaces and enhance the LLM's ability to handle longer speech. Then, we utilize a curriculum learning strategy to facilitate the model's transition from TSum to SSum. Finally, our model achieves competitive performance on the How-2 dataset.

【8】 SAVE: Segment Audio-Visual Easy way using Segment Anything Mode
l标题:保存:使用Segment Anything模型轻松地分段视听
链接:https://arxiv.org/abs/2407.02004
作者:Khanh-Binh Nguyen,Chae Jung Park
摘要:视听分割(AVS)的主要目的是通过在像素级准确预测分割掩模来精确识别和定位视觉场景中的听觉元素。实现这一目标需要全面考虑数据和模型方面,以有效地解决这一任务。这项研究提出了一种轻量级的方法,SAVE,它有效地适应预先训练的段任何模型(SAM)的AVS任务。通过将图像编码器适配器结合到Transformer块中以更好地捕获不同的数据集信息,并提出残余音频编码器适配器以将音频特征编码为稀疏提示,我们提出的模型在编码阶段实现了有效的视听融合和交互。我们提出的方法通过将输入分辨率从1024像素降低到256像素来加快训练和推理速度,同时与以前的SOTA相比实现了更高的性能。大量的实验验证了我们的方法,表明我们提出的模型优于其他SOTA方法显着。此外,利用合成数据的预训练模型增强了真实AVSBench数据的性能,在S4(V1S)子集上实现了84.59 mIoU,在MS3(V1M)集上实现了70.28 mIoU,输入图像只有256个像素。这在S4(V1S)上增加到86.16 mIoU,在MS3(V1M)上增加到70.83 mIoU,输入为1024像素。
摘要:The primary aim of Audio-Visual Segmentation (AVS) is to precisely identify and locate auditory elements within visual scenes by accurately predicting segmentation masks at the pixel level. Achieving this involves comprehensively considering data and model aspects to address this task effectively. This study presents a lightweight approach, SAVE, which efficiently adapts the pre-trained segment anything model (SAM) to the AVS task. By incorporating an image encoder adapter into the transformer blocks to better capture the distinct dataset information and proposing a residual audio encoder adapter to encode the audio features as a sparse prompt, our proposed model achieves effective audio-visual fusion and interaction during the encoding stage. Our proposed method accelerates the training and inference speed by reducing the input resolution from 1024 to 256 pixels while achieving higher performance compared with the previous SOTA. Extensive experimentation validates our approach, demonstrating that our proposed model outperforms other SOTA methods significantly. Moreover, leveraging the pre-trained model on synthetic data enhances performance on real AVSBench data, achieving 84.59 mIoU on the S4 (V1S) subset and 70.28 mIoU on the MS3 (V1M) set with only 256 pixels for input images. This increases up to 86.16 mIoU on the S4 (V1S) and 70.83 mIoU on the MS3 (V1M) with inputs of 1024 pixels.


【9】 Investigating the Effects of Large-Scale Pseudo-Stereo Data and  Different Speech Foundation Model on Dialogue Generative Spoken Language  Model
标题:研究大规模伪立体声数据和不同语音基础模型对对话生成口语模型的影响
链接:https://arxiv.org/abs/2407.01911
作者:Yu-Kuan Fu,Cheng-Kuang Lee,Hsiu-Hsuan Wang,Hung-yi Lee
备注:submitted to interspeech 2024
摘要:最近的努力,在口语对话建模的目的是合成口语对话,而不需要直接转录,从而保留了丰富的非文本信息固有的讲话。然而,当说话者同时说话时,这种方法面临挑战,需要在单独的通道上记录说话者的立体声对话数据,这是一种非常稀缺的资源。为了解决这个问题,我们开发了一种创新的管道,能够将单声道对话数据转换为伪立体声数据。这将我们的训练数据集从仅仅2,000小时扩展到令人印象深刻的17,600小时,大大丰富了可用训练示例的多样性和质量。包含该伪立体声数据已被证明在提高口语对话语言模型的性能方面是有效的。此外,我们探讨了使用不同的语音基础模型的离散单元的口语对话生成。
摘要:Recent efforts in Spoken Dialogue Modeling aim to synthesize spoken dialogue without the need for direct transcription, thereby preserving the wealth of non-textual information inherent in speech. However, this approach faces a challenge when speakers talk simultaneously, requiring stereo dialogue data with speakers recorded on separate channels, a notably scarce resource. To address this, we have developed an innovative pipeline capable of transforming single-channel dialogue data into pseudo-stereo data. This expanded our training dataset from a mere 2,000 to an impressive 17,600 hours, significantly enriching the diversity and quality of the training examples available. The inclusion of this pseudo-stereo data has proven to be effective in improving the performance of spoken dialogue language models. Additionally, we explored the use of discrete units of different speech foundation models for spoken dialogue generation.


【10】 Pinyin Regularization in Error Correction for Chinese Speech Recognition  with Large Language Models
标题:大语言模型中文语音识别错误纠正中的拼音正规化
链接:https://arxiv.org/abs/2407.01909
作者:Zhiyuan Tang,Dong Wang,Shen Huang,Shidong Shang
备注:Interspeech 2024
摘要:最近的研究已经证明了大语言模型(LLM)在自动语音识别(ASR)纠错中的有效性。然而,大部分研究都集中在英语语言上。本文将注意力转向中文。首先,我们构建了一个专门的基准数据集,旨在为724 K假设-转录对的中文ASR纠错,命名为中国假设天堂数据集(ChineseHP),它包含了广泛的场景,并提出了重大的挑战。随后,我们使用数据集对直接提示和微调预训练的LLM进行了初步评估。此外,我们提出了一个简单的方法,拼音正则化的提示,这涉及到拼音的转录直接从文本的假设。实验结果表明,与未正则化的LLM相比,拼音正则化一致地增强了LLM的纠错能力。该数据集可在网站上查阅。
摘要:Recent studies have demonstrated the efficacy of large language models (LLMs) in error correction for automatic speech recognition (ASR). However, much of the research focuses on the English language. This paper redirects the attention to Chinese. Firstly, we construct a specialized benchmark dataset aimed at error correction for Chinese ASR with 724K hypotheses-transcription pairs, named the Chinese Hypotheses Paradise dataset (ChineseHP), which contains a wide range of scenarios and presents significant challenges. Subsequently, we conduct a preliminary evaluation using the dataset for both direct-prompting and fine-tuning pre-trained LLMs. Furthermore, we propose a straightforward method of Pinyin regularization for prompts, which involves the transcription of Pinyin directly from text hypotheses. The experimental results reveal that Pinyin regularization consistently enhances the error-correcting ability of LLMs when compared with those without regularization. The dataset is available on the website.

【11】 Constant Directivity Loudspeaker Beamforming
标题:恒定方向性扬声器射束成形
链接:https://arxiv.org/abs/2407.01860
作者:Yuancheng Luo
备注:Accepted at EUSIPCO 2024
摘要:扬声器阵列波束成形是用于声学指向性控制和鲁棒音频再现的常见信号处理技术。与它们的麦克风对应物不同,扬声器约束通常是异质的,这是由于阵列换能器在频率、声电灵敏度、效率和方向性方面具有不同的操作范围。这项工作提出了一种频率正则化方法的广义瑞利商方向性规范和两个新的波束形成器的设计,优化最大效率恒定方向性(MECD)和最大灵敏度恒定方向性(MSCD)。我们得到快速收敛和解析解,从他们的二次等式约束的二次规划公式。实验优化广义方向性指数约束的波束形成器设计的全频带异构阵列。
摘要:Loudspeaker array beamforming is a common signal processing technique for acoustic directivity control and robust audio reproduction. Unlike their microphone counterpart, loudspeaker constraints are often heterogeneous due to arrayed transducers with varying operating ranges in frequency, acoustic-electrical sensitivity, efficiency, and directivity. This work proposes a frequency-regularization method for generalized Rayleigh quotient directivity specifications and two novel beamformer designs that optimize for maximum efficiency constant directivity (MECD) and maximum sensitivity constant directivity (MSCD). We derive fast converging and analytic solutions from their quadratic equality constrained quadratic program formulations. Experiments optimize generalized directivity index constrained beamformer designs for a full-band heterogeneous array.

【12】 Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of  Deep Learning Models
标题:使用基于谱图的特征和深度学习模型集成的Deepfake音频检测
链接:https://arxiv.org/abs/2407.01777
作者:Lam Pham,Phat Lam,Truong Nguyen,Huyen Nguyen,Alexander Schindler
摘要:在本文中,我们提出了一个基于深度学习的系统,用于deepfake音频检测任务。具体地,首先使用短时傅立叶变换(STFT)、恒定Q变换(CQT)、小波变换(WT)这三种变换方法与Mel、Gammatone、线性滤波器(LF)和离散余弦变换(DCT)的不同的基于分类的滤波器组合将绘制输入音频变换为各种频谱图。鉴于频谱图,我们基于三种深度学习方法评估了各种分类模型。第一种方法是使用我们提出的基于CNN的模型(CNN基线),基于RNN的模型(RNN基线),C-RNN模型(C-RNN基线)的基线模型直接训练谱图。同时,第二种方法是从ResNet-18,MobileNet-V3,EfficientNet-B 0,DenseNet-121,SuffleNet-V2,Swint,Convnext-Tiny,GoogLeNet,MNASsnet,RegNet等计算机视觉模型中进行迁移学习。在第三种方法中,我们利用最先进的音频预训练模型Whisper,Seamless,Speechbrain和Pyannote从输入频谱图中提取音频嵌入。然后利用多层感知器(MLP)模型对音频嵌入进行探索,以检测出虚假或真实的音频样本。最后,将这些方法的高性能深度学习模型融合在一起,以实现最佳性能。我们在ASVspoof 2019基准数据集上评估了我们提出的模型。我们最好的集成模型实现了0.03的等错误率(EER),这与ASVspoofing 2019挑战中的顶级系统具有很强的竞争力。实验结果还强调了选择性频谱图和深度学习方法在增强音频deepfake检测任务方面的潜力。
摘要:In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier Transform (STFT), Constant-Q Transform (CQT), Wavelet Transform (WT) combined with different auditory-based filters of Mel, Gammatone, linear filters (LF), and discrete cosine transform (DCT). Given the spectrograms, we evaluate a wide range of classification models based on three deep learning approaches. The first approach is to train directly the spectrograms using our proposed baseline models of CNN-based model (CNN-baseline), RNN-based model (RNN-baseline), C-RNN model (C-RNN baseline). Meanwhile, the second approach is transfer learning from computer vision models such as ResNet-18, MobileNet-V3, EfficientNet-B0, DenseNet-121, SuffleNet-V2, Swint, Convnext-Tiny, GoogLeNet, MNASsnet, RegNet. In the third approach, we leverage the state-of-the-art audio pre-trained models of Whisper, Seamless, Speechbrain, and Pyannote to extract audio embeddings from the input spectrograms. Then, the audio embeddings are explored by a Multilayer perceptron (MLP) model to detect the fake or real audio samples. Finally, high-performance deep learning models from these approaches are fused to achieve the best performance. We evaluated our proposed models on ASVspoof 2019 benchmark dataset. Our best ensemble model achieved an Equal Error Rate (EER) of 0.03, which is highly competitive to top-performing systems in the ASVspoofing 2019 challenge. Experimental results also highlight the potential of selective spectrograms and deep learning approaches to enhance the task of audio deepfake detection.


【13】 The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
标题:ICMC-ASB挑战赛的USC-NERCSLIP系统
链接:https://arxiv.org/abs/2407.02052
作者:Minghui Wu,Luzhen Xu,Jie Zhang,Haitao Tang,Yanyan Yue,Ruizhi Liao,Jintao Zhao,Zhengzhe Zhang,Yichi Wang,Haoyin Yan,Hongliang Yu,Tongle Ma,Jiachen Liu,Chongliang Wu,Yongchao Li,Yanyong Zhang,Xin Fang,Yue Zhang
备注:Accepted at ICASSP 2024
摘要:本报告描述了提交的系统,以车内多通道自动语音识别(ICMC-ASR)的挑战,其中考虑了ASR任务与多扬声器重叠和普通话口音动态的ICMC的情况下。我们分别使用基于自监督学习表示的多说话人嵌入和使用说话人位置的波束形成来实现前端说话人日志化。对于ASR,我们采用一种基于融合模型的迭代伪标签生成方法来获得无监督数据的文本标签。为了减轻口音的影响,Accent-ASR框架提出,它捕获发音相关的口音特征在细粒度的水平和语言信息在粗粒度的水平。在ICMC-ASR评估集上,所提出的系统在轨道1上实现了13.16%的CER,在轨道2上实现了21.48%的cpCER,显著优于官方基线系统,并在两个轨道上均获得第一名。
摘要:This report describes the submitted system to the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) challenge, which considers the ASR task with multi-speaker overlapping and Mandarin accent dynamics in the ICMC case. We implement the front-end speaker diarization using the self-supervised learning representation based multi-speaker embedding and beamforming using the speaker position, respectively. For ASR, we employ an iterative pseudo-label generation method based on fusion model to obtain text labels of unsupervised data. To mitigate the impact of accent, an Accent-ASR framework is proposed, which captures pronunciation-related accent features at a fine-grained level and linguistic information at a coarse-grained level. On the ICMC-ASR eval set, the proposed system achieves a CER of 13.16% on track 1 and a cpCER of 21.48% on track 2, which significantly outperforms the official baseline system and obtains the first rank on both tracks.

【14】 Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
标题:具有完全文本控制旋律的伴奏歌唱声音合成
链接:https://arxiv.org/abs/2407.02049
作者:Ruiqi Li,Zhiqing Hong,Yongqi Wang,Lichao Zhang,Rongjie Huang,Siqi Zheng,Zhou Zhao备注:Working in progress
摘要:文本到歌曲(TTSong)是一个音乐生成任务,合成伴随的歌声。目前的TTSong方法继承自歌唱声音合成(SVS),需要旋律相关的信息,有时可能是不切实际的,如乐谱或录音序列。我们提出了MelodyLM,第一个TTSong模型,它可以生成具有完全文本控制旋律的高质量歌曲片段,实现最低的用户要求和最大的控制灵活性。MelodyLM显式地将音频建模为中间旋律相关特征,并以语言模型的方式顺序地生成声乐曲目,以文本和声乐提示为条件。伴奏音乐随后合成一个潜在的扩散模型与混合条件的时间对齐。用户只需输入歌词和参考语音就可以合成一首歌曲样本。要实现完全控制,只需输入文本提示,甚至直接输入“”即可。实验结果表明,MelodyLM在客观和主观指标方面都取得了优异的性能。音频样本可在https://melodylm666.github.io上获得。
摘要:Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be impractical, such as music scores or MIDI sequences. We present MelodyLM, the first TTSong model that generates high-quality song pieces with fully text-controlled melodies, achieving minimal user requirements and maximum control flexibility. MelodyLM explicitly models MIDI as the intermediate melody-related feature and sequentially generates vocal tracks in a language model manner, conditioned on textual and vocal prompts. The accompaniment music is subsequently synthesized by a latent diffusion model with hybrid conditioning for temporal alignment. With minimal requirements, users only need to input lyrics and a reference voice to synthesize a song sample. For full control, just input textual prompts or even directly input MIDI. Experimental results indicate that MelodyLM achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://melodylm666.github.io.

【15】 SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight  Conv-TasNet and State Space Modeling
标题:SpeakerBeam-SS:使用轻量级Conv-TasNet和状态空间建模的实时目标说话人提取
链接:https://arxiv.org/abs/2407.01857
作者:Hiroshi Sato,Takafumi Moriya,Masato Mimura,Shota Horiguchi,Tsubasa Ochiai,Takanori Ashihara,Atsushi Ando,Kentaro Shinayama,Marc Delcroix
备注:Accepted to Interspeech 2024
摘要:实时目标说话人提取(TSE)旨在以流式方式从观察到的多个说话人的混合中提取期望说话人的语音。实现实时TSE是具有挑战性的,因为必须降低计算复杂度以提供实时操作。这项工作介绍了基于Conv-TasNet的TSE的状态空间建模(SSM)的基础上,已被证明可以有效地模拟长期依赖的新架构。由于SSM,需要更少的膨胀卷积层来捕获Conv-TasNet中的时间依赖性,从而降低了模型复杂度。我们还扩大了卷积(TasNet)前端编码器的窗口长度和移位,以进一步降低计算成本;通过前端编码器的过度参数化来补偿性能下降。所提出的方法从传统的因果Conv-TasNet的TSE减少了78%的实时因素,同时匹配其性能。
摘要:Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real-time TSE is challenging as the computational complexity must be reduced to provide real-time operation. This work introduces to Conv-TasNet-based TSE a new architecture based on state space modeling (SSM) that has been shown to model long-term dependency effectively. Owing to SSM, fewer dilated convolutional layers are required to capture temporal dependency in Conv-TasNet, resulting in the reduction of model complexity. We also enlarge the window length and shift of the convolutional (TasNet) frontend encoder to reduce the computational cost further; the performance decline is compensated by over-parameterization of the frontend encoder. The proposed method reduces the real-time factor by 78% from the conventional causal Conv-TasNet-based TSE while matching its performance.

eess.AS音频处理
【1】 The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
标题:ICMC-ASB挑战赛的USC-NERCSLIP系统
链接:https://arxiv.org/abs/2407.02052
作者:Minghui Wu,Luzhen Xu,Jie Zhang,Haitao Tang,Yanyan Yue,Ruizhi Liao,Jintao Zhao,Zhengzhe Zhang,Yichi Wang,Haoyin Yan,Hongliang Yu,Tongle Ma,Jiachen Liu,Chongliang Wu,Yongchao Li,Yanyong Zhang,Xin Fang,Yue Zhang
备注:Accepted at ICASSP 2024
摘要:本报告描述了提交的系统,以车内多通道自动语音识别(ICMC-ASR)的挑战,其中考虑了ASR任务与多扬声器重叠和普通话口音动态的ICMC的情况下。我们分别使用基于自监督学习表示的多说话人嵌入和使用说话人位置的波束形成来实现前端说话人日志化。对于ASR,我们采用一种基于融合模型的迭代伪标签生成方法来获得无监督数据的文本标签。为了减轻口音的影响,Accent-ASR框架提出,它捕获发音相关的口音特征在细粒度的水平和语言信息在粗粒度的水平。在ICMC-ASR评估集上,所提出的系统在轨道1上实现了13.16%的CER,在轨道2上实现了21.48%的cpCER,显著优于官方基线系统,并在两个轨道上均获得第一名。
摘要:This report describes the submitted system to the In-Car Multi-Channel Automatic Speech Recognition (ICMC-ASR) challenge, which considers the ASR task with multi-speaker overlapping and Mandarin accent dynamics in the ICMC case. We implement the front-end speaker diarization using the self-supervised learning representation based multi-speaker embedding and beamforming using the speaker position, respectively. For ASR, we employ an iterative pseudo-label generation method based on fusion model to obtain text labels of unsupervised data. To mitigate the impact of accent, an Accent-ASR framework is proposed, which captures pronunciation-related accent features at a fine-grained level and linguistic information at a coarse-grained level. On the ICMC-ASR eval set, the proposed system achieves a CER of 13.16% on track 1 and a cpCER of 21.48% on track 2, which significantly outperforms the official baseline system and obtains the first rank on both tracks.

【2】 Accompanied Singing Voice Synthesis with Fully Text-controlled Melody
标题:具有完全文本控制旋律的伴奏歌唱声音合成
链接:https://arxiv.org/abs/2407.02049
作者:Ruiqi Li,Zhiqing Hong,Yongqi Wang,Lichao Zhang,Rongjie Huang,Siqi Zheng,Zhou Zhao
备注:Working in progress
摘要:文本到歌曲(TTSong)是一个音乐生成任务,合成伴随的歌声。目前的TTSong方法继承自歌唱声音合成(SVS),需要旋律相关的信息,有时可能是不切实际的,如乐谱或录音序列。我们提出了MelodyLM,第一个TTSong模型,它可以生成具有完全文本控制旋律的高质量歌曲片段,实现最低的用户要求和最大的控制灵活性。MelodyLM显式地将音频建模为中间旋律相关特征,并以语言模型的方式顺序地生成声乐曲目,以文本和声乐提示为条件。伴奏音乐随后合成一个潜在的扩散模型与混合条件的时间对齐。用户只需输入歌词和参考语音就可以合成一首歌曲样本。要实现完全控制,只需输入文本提示,甚至直接输入“”即可。实验结果表明,MelodyLM在客观和主观指标方面都取得了优异的性能。音频样本可在https://melodylm666.github.io上获得。
摘要:Text-to-song (TTSong) is a music generation task that synthesizes accompanied singing voices. Current TTSong methods, inherited from singing voice synthesis (SVS), require melody-related information that can sometimes be impractical, such as music scores or MIDI sequences. We present MelodyLM, the first TTSong model that generates high-quality song pieces with fully text-controlled melodies, achieving minimal user requirements and maximum control flexibility. MelodyLM explicitly models MIDI as the intermediate melody-related feature and sequentially generates vocal tracks in a language model manner, conditioned on textual and vocal prompts. The accompaniment music is subsequently synthesized by a latent diffusion model with hybrid conditioning for temporal alignment. With minimal requirements, users only need to input lyrics and a reference voice to synthesize a song sample. For full control, just input textual prompts or even directly input MIDI. Experimental results indicate that MelodyLM achieves superior performance in terms of both objective and subjective metrics. Audio samples are available at https://melodylm666.github.io.

【3】 SOT Triggered Neural Clustering for Speaker Attributed ASR
标题:SOT触发的说话者归因ASB的神经集群
链接:https://arxiv.org/abs/2407.02007
作者:Xianrui Zheng,Guangzhi Sun,Chao Zhang,Philip C. Woodland
备注:To appear in Interspeech 2024
摘要:本文介绍了一种新的方法,说话人属性的ASR转录使用神经聚类方法。通过并行处理机制,可以同时应用diarisation和ASR,有助于防止级联系统中从一个子系统到下一个子系统的错误积累。这是通过使用ASR来实现的,ASR使用序列化输出训练方法进行训练,以及分段级判别神经聚类(SDNC)来分配扬声器标签。使用SDNC,我们的系统不需要额外的非神经聚类方法来分配说话者标签,从而允许整个系统基于神经网络。AMI会议数据集上的实验结果表明,SDNC优于谱聚类(SC)的19%相对diarisation错误率(DER)减少AMI Eval集。当与SC的级联系统相比,SDNC的并联系统给出了一个7%/4%的相对改善,在cpWER的Dev/Eval集。
摘要:This paper introduces a novel approach to speaker-attributed ASR transcription using a neural clustering method. With a parallel processing mechanism, diarisation and ASR can be applied simultaneously, helping to prevent the accumulation of errors from one sub-system to the next in a cascaded system. This is achieved by the use of ASR, trained using a serialised output training method, together with segment-level discriminative neural clustering (SDNC) to assign speaker labels. With SDNC, our system does not require an extra non-neural clustering method to assign speaker labels, thus allowing the entire system to be based on neural networks. Experimental results on the AMI meeting dataset demonstrate that SDNC outperforms spectral clustering (SC) by a 19% relative diarisation error rate (DER) reduction on the AMI Eval set. When compared with the cascaded system with SC, the parallel system with SDNC gives a 7%/4% relative improvement in cpWER on the Dev/Eval set.


【4】 Towards Unsupervised Speaker Diarization System for Multilingual  Telephone Calls Using Pre-trained Whisper Model and Mixture of Sparse  Autoencoders
标题:使用预训练的Whisper模型和稀疏自动编码器混合的多语言电话呼叫的无监督说话者拨号系统
链接:https://arxiv.org/abs/2407.01963
作者:Phat Lam,Lam Pham,Tin Nguyen,Thinh Pham,Loi Khanh Nguyen,Alexander Schindler
备注:8 pages, 7 figures
摘要:现有的说话人日志系统严重依赖于大量的人工标注数据,这是劳动密集型的,具有挑战性的,以收集在现实世界中的场景。此外,特定语言的限制,在发言人日记系统显着阻碍其适用性和可扩展性,在多语言环境。因此,在本文中,我们提出了一个基于聚类的扬声器日记系统的多语言电话呼叫应用程序。所提出的系统支持多种语言,并且不需要大规模的注释数据用于训练过程,因为利用多语言Whisper模型来提取说话人嵌入,并提出了一种用于无监督说话人聚类的新型混合稀疏自动编码器(Mix-SAE)网络架构。在CALLHOME和CALLFRIEND电话语音语料库的两说话人子集上的实验结果表明,与其他基于自动编码器的聚类方法相比,所提出的Mix-SAE网络具有更高的效率。我们所提出的系统的整体性能也表明了我们的方法在有限的注释数据的背景下开发无监督的多语言扬声器日记应用程序,并提高集成能力到全面的多任务语音分析系统(即语音到文本,语言检测,扬声器日记集成在一个低复杂度的系统的多个任务)的潜力。
摘要:Existing speaker diarization systems heavily rely on large amounts of manually annotated data, which is labor-intensive and challenging to collect in real-world scenarios. Additionally, the language-specific constraint in speaker diarization systems significantly hinders their applicability and scalability in multilingual settings. In this paper, we therefore propose a cluster-based speaker diarization system for multilingual telephone call applications. The proposed system supports multiple languages and does not require large-scale annotated data for the training process as leveraging the multilingual Whisper model to extract speaker embeddings and proposing a novel Mixture of Sparse Autoencoders (Mix-SAE) network architecture for unsupervised speaker clustering. Experimental results on the evaluating dataset derived from two-speaker subsets of CALLHOME and CALLFRIEND telephonic speech corpora demonstrate superior efficiency of the proposed Mix-SAE network to other autoencoder-based clustering methods. The overall performance of our proposed system also indicates the promising potential of our approach in developing unsupervised multilingual speaker diarization applications within the context of limited annotated data and enhancing the integration ability into comprehensive multi-task speech analysis systems (i.e. multiple tasks of speech-to-text, language detection, speaker diarization integrated in a low-complexity system).

【5】 Unsupervised Face-Mask Speech Enhancement Using Generative Adversarial  Networks with Human-in-the-Loop Assessment Metrics
标题:使用具有人在环评估子索的生成对抗网络的无监督面罩语音增强
链接:https://arxiv.org/abs/2407.01939
作者:Syu-Siang Wang,Jia-Yang Chen,Bo-Ren Bai,Shih-Hau Fang,Yu Tsao
备注:None
摘要:使用口罩是一项重要的医疗保健措施,特别是在流行病期间,但它可能会对我们日常生活中的沟通带来挑战。为了解决这个问题,我们提出了一种新的方法,称为人在环StarGAN(HL-StarGAN)的人脸掩蔽语音增强方法。HL-StarGAN包括搜索引擎、分类器、度量评估预测器和利用注意力机制的生成器。指标评估预测器(称为MaskQSS)在其开发过程中纳入了人类参与者,并在HL-StarGAN的学习过程中充当“人在回路”模块。整个HL-StarGAN模型使用无监督学习策略进行训练,该策略同时关注原始干净语音的重建和人类感知的优化。为了实现HL-StarGAN,我们策划了一个名为“FMVD”的蒙面语音数据库,其中包括来自34个扬声器的录音,这些录音来自三种不同的蒙面场景和一个干净的条件。我们使用该数据库对拟议的HL-StarGAN进行了主观和客观测试。测试结果如下:(1)MaskQSS成功地预测了面具语音的质量分数,优于现有的几种语音评估方法。(2)MaskQSS预测器的集成增强了HL-StarGAN将面罩语音转换为高质量语音的能力;这种增强在客观和主观测试中都很明显,优于传统的StarGAN和基于CycleGAN的系统。
摘要:The utilization of face masks is an essential healthcare measure, particularly during times of pandemics, yet it can present challenges in communication in our daily lives. To address this problem, we propose a novel approach known as the human-in-the-loop StarGAN (HL-StarGAN) face-masked speech enhancement method. HL-StarGAN comprises discriminator, classifier, metric assessment predictor, and generator that leverages an attention mechanism. The metric assessment predictor, referred to as MaskQSS, incorporates human participants in its development and serves as a "human-in-the-loop" module during the learning process of HL-StarGAN. The overall HL-StarGAN model was trained using an unsupervised learning strategy that simultaneously focuses on the reconstruction of the original clean speech and the optimization of human perception. To implement HL-StarGAN, we curated a face-masked speech database named "FMVD," which comprises recordings from 34 speakers in three distinct face-masked scenarios and a clean condition. We conducted subjective and objective tests on the proposed HL-StarGAN using this database. The outcomes of the test results are as follows: (1) MaskQSS successfully predicted the quality scores of face mask voices, outperforming several existing speech assessment methods. (2) The integration of the MaskQSS predictor enhanced the ability of HL-StarGAN to transform face mask voices into high-quality speech; this enhancement is evident in both objective and subjective tests, outperforming conventional StarGAN and CycleGAN-based systems.

【6】 TTSlow: Slow Down Text-to-Speech with Efficiency Robustness Evaluations
标题:TTSlow:通过效率稳健性评估减缓文本转语音
链接:https://arxiv.org/abs/2407.01927
作者:Xiaoxue Gao,Yiming Chen,Xianghu Yue,Yu Tsao,Nancy F. Chen
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
摘要:文语转换(TTS)技术在各种实时应用中发挥着重要作用,已被广泛研究用于生成具有文本输入的高质量语音。对于现实世界的部署,确保稳定和及时生成TTS模型对微小的输入扰动是至关重要的。因此,评估TTS模型对这种扰动(通常称为对抗性攻击)的鲁棒性是非常必要的。在本文中,我们提出了TTSlow,这是一种专门用于减缓TTS系统中语音生成过程的新型对抗方法。为了诱导长的TTS等待时间,我们设计了新的效率为导向的对抗性损失,以鼓励无休止的生成过程。TTSlow包含针对文本输入和说话人嵌入的两种攻击策略。具体来说,我们提出了TTSlow-text,它利用了基于homoglyphs和基于swap的扰动的组合,以及TTSlow-spk,它采用了梯度优化攻击方法进行说话人嵌入。TTSlow是针对各种TTS模型的第一种攻击方法,包括自回归和非自回归TTS模型,从而推进了音频安全的探索。进行了大量的实验,以评估TTS模型的推理效率,并使用Gemini生成的语音可懂度进行深入分析。结果表明,TTSlow可以有效地减慢三个公开数据集上的两个TTS模型。我们承诺在接受后发布源代码,以促进该领域的进一步研究和基准测试。
摘要:Text-to-speech (TTS) has been extensively studied for generating high-quality speech with textual inputs, playing a crucial role in various real-time applications. For real-world deployment, ensuring stable and timely generation in TTS models against minor input perturbations is of paramount importance. Therefore, evaluating the robustness of TTS models against such perturbations, commonly known as adversarial attacks, is highly desirable. In this paper, we propose TTSlow, a novel adversarial approach specifically tailored to slow down the speech generation process in TTS systems. To induce long TTS waiting time, we design novel efficiency-oriented adversarial loss to encourage endless generation process. TTSlow encompasses two attack strategies targeting both text inputs and speaker embedding. Specifically, we propose TTSlow-text, which utilizes a combination of homoglyphs-based and swap-based perturbations, along with TTSlow-spk, which employs a gradient optimization attack approach for speaker embedding. TTSlow serves as the first attack approach targeting a wide range of TTS models, including autoregressive and non-autoregressive TTS ones, thereby advancing exploration in audio security. Extensive experiments are conducted to evaluate the inference efficiency of TTS models, and in-depth analysis of generated speech intelligibility is performed using Gemini. The results demonstrate that TTSlow can effectively slow down two TTS models across three publicly available datasets. We are committed to releasing the source code upon acceptance, facilitating further research and benchmarking in this domain.

【7】 SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight  Conv-TasNet and State Space Modeling
标题:SpeakerBeam-SS:使用轻量级Conv-TasNet和状态空间建模的实时目标说话人提取
链接:https://arxiv.org/abs/2407.01857
作者:Hiroshi Sato,Takafumi Moriya,Masato Mimura,Shota Horiguchi,Tsubasa Ochiai,Takanori Ashihara,Atsushi Ando,Kentaro Shinayama,Marc Delcroix
备注:Accepted to Interspeech 2024
摘要:实时目标说话人提取(TSE)旨在以流式方式从观察到的多个说话人的混合中提取期望说话人的语音。实现实时TSE是具有挑战性的,因为必须降低计算复杂度以提供实时操作。这项工作介绍了基于Conv-TasNet的TSE的状态空间建模(SSM)的基础上,已被证明可以有效地模拟长期依赖的新架构。由于SSM,需要更少的膨胀卷积层来捕获Conv-TasNet中的时间依赖性,从而降低了模型复杂度。我们还扩大了卷积(TasNet)前端编码器的窗口长度和移位,以进一步降低计算成本;通过前端编码器的过度参数化来补偿性能下降。所提出的方法从传统的因果Conv-TasNet的TSE减少了78%的实时因素,同时匹配其性能。
摘要:Real-time target speaker extraction (TSE) is intended to extract the desired speaker's voice from the observed mixture of multiple speakers in a streaming manner. Implementing real-time TSE is challenging as the computational complexity must be reduced to provide real-time operation. This work introduces to Conv-TasNet-based TSE a new architecture based on state space modeling (SSM) that has been shown to model long-term dependency effectively. Owing to SSM, fewer dilated convolutional layers are required to capture temporal dependency in Conv-TasNet, resulting in the reduction of model complexity. We also enlarge the window length and shift of the convolutional (TasNet) frontend encoder to reduce the computational cost further; the performance decline is compensated by over-parameterization of the frontend encoder. The proposed method reduces the real-time factor by 78% from the conventional causal Conv-TasNet-based TSE while matching its performance.

【8】 peerRTF: Robust MVDR Beamforming Using Graph Convolutional Network
标题:peerCTF:使用图卷积网络的鲁棒MVDR束形成
链接:https://arxiv.org/abs/2407.01779
作者:Amit Sofer,Daniel Levi,Sharon Gannot
摘要:准确可靠地识别麦克风之间相对于所需源的RTF是麦克风阵列波束形成器设计中的重要组成部分,特别是MVDR标准。由于在嘈杂和混响环境中准确估计RTF是一项繁琐的任务,我们的目标是利用声学外壳的先验知识,通过学习RTF流形来鲁棒RTF估计。在本文中,我们提出了一种新的鲁棒RTF识别方法,测试和训练的真实记录,这依赖于学习RTF流形使用GCN推断一个强大的RTF表示在一个有限的区域,从而提高波束形成器的性能。
摘要:Accurate and reliable identification of the RTF between microphones with respect to a desired source is an essential component in the design of microphone array beamformers, specifically the MVDR criterion. Since an accurate estimation of the RTF in a noisy and reverberant environment is a cumbersome task, we aim at leveraging prior knowledge of the acoustic enclosure to robustify the RTF estimation by learning the RTF manifold. In this paper, we present a novel robust RTF identification method, tested and trained with real recordings, which relies on learning the RTF manifold using a GCN to infer a robust representation of the RTF in a confined area, and consequently enhance the beamformer's performance.

【9】 Audio-Visual Approach For Multimodal Concurrent Speaker Detection
标题:多模式并发说话人检测的视听方法
链接:https://arxiv.org/abs/2407.01774
作者:Amit Eliav,Sharon Gannot
摘要:并发说话人检测(CSD)是识别音频信号中活跃说话人的存在和重叠的任务,对于诸如会议转录、说话人日记和语音分离等许多音频任务至关重要。这项研究介绍了一种利用音频和视觉信息的多模态深度学习方法。该模型采用早期融合策略,通过跨模态注意机制结合音频和视觉特征,并使用可学习的[CLS]标记捕获相关的视听关系。  该模型在两个真实世界的数据集上进行了广泛的评估,AMI和最近推出的EasyCom数据集。实验验证了多模态融合策略的有效性。消融研究进一步支持模型的设计选择和训练过程。由于这是第一次在具有挑战性的EasyCom数据集上报告CSD结果,因此研究结果证明了所提出的CSD多模态方法在现实世界场景中的潜力。
摘要:Concurrent Speaker Detection (CSD), the task of identifying the presence and overlap of active speakers in an audio signal, is crucial for many audio tasks such as meeting transcription, speaker diarization, and speech separation. This study introduces a multimodal deep learning approach that leverages both audio and visual information. The proposed model employs an early fusion strategy combining audio and visual features through cross-modal attention mechanisms, with a learnable [CLS] token capturing the relevant audio-visual relationships.  The model is extensively evaluated on two real-world datasets, AMI and the recently introduced EasyCom dataset. Experiments validate the effectiveness of the multimodal fusion strategy. Ablation studies further support the design choices and the training procedure of the model. As this is the first work reporting CSD results on the challenging EasyCom dataset, the findings demonstrate the potential of the proposed multimodal approach for CSD in real-world scenarios.


【10】 The Solution for Temporal Sound Localisation Task of ICCV 1st Perception  Test Challenge 2023
标题:ICCV 2023年第一届感知测试挑战赛时间声音定位任务的解决方案
链接:https://arxiv.org/abs/2407.02318
作者:Yurui Huang,Yang Yang,Shou Chen,Xiangyu Wu,Qingguo Chen,Jianfeng Lu
摘要:在本文中,我们提出了一种解决方案,以提高质量的时间声音定位。我们采用多模态融合的方法来结合视觉和音频功能。使用最先进的自监督预训练网络提取高质量的视觉特征,从而实现高效的视频特征表示。同时,音频特征作为补充信息,帮助模型更好地定位声音的开始和结束。融合后的特征在多尺度Transformer中进行训练。在最终的测试数据集中,我们实现了0.33的平均精度(mAP),在该赛道中获得了第二好的性能。
摘要:In this paper, we propose a solution for improving the quality of temporal sound localization. We employ a multimodal fusion approach to combine visual and audio features. High-quality visual features are extracted using a state-of-the-art self-supervised pre-training network, resulting in efficient video feature representations. At the same time, audio features serve as complementary information to help the model better localize the start and end of sounds. The fused features are trained in a multi-scale Transformer for training. In the final test dataset, we achieved a mean average precision (mAP) of 0.33, obtaining the second-best performance in this track.


【11】 MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music  Processing
标题:MelodyT5:用于符号音乐处理的统一乐谱Transformer
链接:https://arxiv.org/abs/2407.02277
作者:Shangda Wu,Yashan Wang,Xiaobing Li,Feng Yu,Maosong Sun
备注:9 pages, 2 figures, 3 tables, accepted by ISMIR 2024
摘要:在符号音乐研究领域,由于缺乏可用的训练数据和对针对特定任务定制的模型的需求,开发可扩展系统的进展受到了显着阻碍。为了解决这些问题,我们提出了MelodyT 5,一个新的统一框架,利用编码器-解码器架构,专为符号音乐处理在ABC符号。这个框架挑战了传统的任务特定的方法,考虑各种符号音乐任务的分数到分数的转换。因此,它集成了七个以旋律为中心的任务,从生成到协调和分割,在一个单一的模型。MelodyHub是一个新策划的集合,包含超过261 K以ABC符号编码的独特旋律,并包含超过一百万个任务实例,MelodyT 5通过多任务迁移学习在符号音乐处理方面表现出卓越的性能。我们的研究结果强调了多任务迁移学习在符号音乐处理中的有效性,特别是对于数据稀缺的任务,挑战了流行的特定任务范式,并为该领域的未来探索提供了全面的数据集和框架。
摘要:In the domain of symbolic music research, the progress of developing scalable systems has been notably hindered by the scarcity of available training data and the demand for models tailored to specific tasks. To address these issues, we propose MelodyT5, a novel unified framework that leverages an encoder-decoder architecture tailored for symbolic music processing in ABC notation. This framework challenges the conventional task-specific approach, considering various symbolic music tasks as score-to-score transformations. Consequently, it integrates seven melody-centric tasks, from generation to harmonization and segmentation, within a single model. Pre-trained on MelodyHub, a newly curated collection featuring over 261K unique melodies encoded in ABC notation and encompassing more than one million task instances, MelodyT5 demonstrates superior performance in symbolic music processing via multi-task transfer learning. Our findings highlight the efficacy of multi-task transfer learning in symbolic music processing, particularly for data-scarce tasks, challenging the prevailing task-specific paradigms and offering a comprehensive dataset and framework for future explorations in this domain.

【12】 SOAF: Scene Occlusion-aware Neural Acoustic Field
标题:SOAF:场景遮挡感知神经声学场
链接:https://arxiv.org/abs/2407.02264
作者:Huiyu Gao,Jiahao Ma,David Ahmedt-Aristizabal,Chuong Nguyen,Miaomiao Liu
摘要:本文解决了在室内场景中,给定来自场景的其他已知轨迹的音视频记录,沿着任意轨迹的新视图音视频合成的问题。现有的方法往往忽略了房间几何形状的影响,特别是墙壁遮挡声音传播,使他们在多房间环境中不太准确。在这项工作中,我们提出了一种新的方法,称为场景遮挡感知声场(SOAF)准确的声音生成。我们的方法推导出一个先验的声能场使用距离感知参数声音传播建模,然后将其转换的基础上从输入视频的场景透射率。我们使用斐波那契球从以接收器为中心的局部声场中提取特征,以产生具有方向感知注意机制的新颖视图的双耳音频。在真实数据集RWAVS和合成数据集SoundSpaces上的大量实验表明,我们的方法在音频生成方面优于以前的最先进技术。项目页面:https://github.com/huiyu-gao/SOAF/。
摘要:This paper tackles the problem of novel view audio-visual synthesis along an arbitrary trajectory in an indoor scene, given the audio-video recordings from other known trajectories of the scene. Existing methods often overlook the effect of room geometry, particularly wall occlusion to sound propagation, making them less accurate in multi-room environments. In this work, we propose a new approach called Scene Occlusion-aware Acoustic Field (SOAF) for accurate sound generation. Our approach derives a prior for sound energy field using distance-aware parametric sound-propagation modelling and then transforms it based on scene transmittance learned from the input video. We extract features from the local acoustic field centred around the receiver using a Fibonacci Sphere to generate binaural audio for novel views with a direction-aware attention mechanism. Extensive experiments on the real dataset~\emph{RWAVS} and the synthetic dataset~\emph{SoundSpaces} demonstrate that our method outperforms previous state-of-the-art techniques in audio generation. Project page: https://github.com/huiyu-gao/SOAF/.


【13】 Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference  Optimization
标题:具有反向推理优化的鲁棒Zero-Shot文本到语音合成
链接:https://arxiv.org/abs/2407.02243
作者:Yuchen Hu,Chen Chen,Siyin Wang,Eng Siong Chng,Chao Zhang
备注:12 pages, Work in progress
摘要:在本文中,我们提出了反向推理优化(RIO),一个简单而有效的方法,旨在提高鲁棒性的自回归模型为基础的zero-shot文本到语音(TTS)系统,采用强化学习从人类反馈(RLHF)。为了评估语音质量的TTS系统所产生的没有人的注释,RIO介绍了一种新的概念,称为反向推理的基础上贝叶斯原理,这表明,高质量的生成的语音应该能够被用来作为一个提示,为后续的一代使用相同的TTS模型。通过利用反向推理作为从TTS系统自身生成的语音样本中选择RLHF中使用的样本的标准,RIO将随后的优化转向增强TTS鲁棒性的方向。RIO框架包括采样、自动标注和学习,避免了对奖励模型或成对偏好数据的需要,并通过减少训练和推理条件之间的差异来显著提高zero-shot TTS性能的稳定性。我们的实验结果验证了RIO可以有效地提高主观和客观指标,包括平均意见分数,单词错误率和说话人相似度。值得注意的是,RIO还可以将不良输出的发生率降低到几乎为零,与使用地面实况语音作为提示时的鲁棒性相媲美。
摘要:In this paper, we propose reverse inference optimization (RIO), a simple and effective method designed to enhance the robustness of autoregressive-model-based zero-shot text-to-speech (TTS) systems using reinforcement learning from human feedback (RLHF). To assess the quality of speech produced by the TTS system without human annotations, RIO introduces a novel concept termed as reverse inference based on the Bayesian principle, which suggests that a high-quality generated speech should be able to be used as a prompt for subsequent generation using the same TTS model. By leveraging reverse inference as the standard to select exemplars used in RLHF from the speech samples generated by the TTS system itself, RIO steers the subsequent optimization towards a direction of enhancing the TTS robustness. The RIO framework, comprising sampling, automatic annotating, and learning, obviates the need for a reward model or pairwise preference data, and significantly improves the stability of zero-shot TTS performance by reducing the discrepancies between training and inference conditions. Our experimental results verify that RIO can effectively improve both subjective and objective metrics, including mean opinion scores, word error rates, and speaker similarity. Remarkably, RIO can also diminish the incidence of bad outputs to nearly zero percent, rivalling the robustness when using ground-truth speech as the prompt.


【14】 GMM-ResNet2: Ensemble of Group ResNet Networks for Synthetic Speech  Detection
标题:GMM-ResNet 2:用于合成语音检测的群组ResNet网络的扩展
链接:https://arxiv.org/abs/2407.02170
作者:Zhenchun Lei,Hui Yan,Changhong Liu,Yong Zhou,Minglei Ma
摘要:深度学习模型广泛用于说话人识别和欺骗语音检测。我们提出了GMM-ResNet 2的合成语音检测。GMM-ResNet 2与之前的GMM-ResNet模型相比,有四个方面的改进。首先,不同阶数的高斯分布具有不同的平滑逼近能力,多阶高斯分布用于提取多尺度对数高斯概率特征。其次,分组技术被用来提高分类精度,暴露组基数,同时减少参数的数量和训练时间。使用平均方法通过所有组分类器输出的集成获得最终分数。第三,通过包括一个激活函数和一个批归一化层来改进残差块。最后,提出了一个集成感知的损失函数,以整合所有集成成员的独立损失函数。在ASVspoof 2019 LA任务中,GMM-ResNet 2实现了0.0227的最小t-DCF和0.79\%的EER。在ASVspoof 2021 LA任务中,GMM-ResNet 2实现了0.2362的最小t-DCF和2.19\%的EER,与LFCC-LCNN基线相比,相对减少了31.4\%和76.3\%。
摘要:Deep learning models are widely used for speaker recognition and spoofing speech detection. We propose the GMM-ResNet2 for synthesis speech detection. Compared with the previous GMM-ResNet model, GMM-ResNet2 has four improvements. Firstly, the different order GMMs have different capabilities to form smooth approximations to the feature distribution, and multiple GMMs are used to extract multi-scale Log Gaussian Probability features. Secondly, the grouping technique is used to improve the classification accuracy by exposing the group cardinality while reducing both the number of parameters and the training time. The final score is obtained by ensemble of all group classifier outputs using the averaging method. Thirdly, the residual block is improved by including one activation function and one batch normalization layer. Finally, an ensemble-aware loss function is proposed to integrate the independent loss functions of all ensemble members. On the ASVspoof 2019 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.0227 and an EER of 0.79\%. On the ASVspoof 2021 LA task, the GMM-ResNet2 achieves a minimum t-DCF of 0.2362 and an EER of 2.19\%, and represents a relative reductions of 31.4\% and 76.3\% compared with the LFCC-LCNN baseline.


【15】 Towards Training Music Taggers on Synthetic Data
标题:利用合成数据训练音乐标记者
链接:https://arxiv.org/abs/2407.02156
作者:Nadine Kroher,Steven Manangu,Aggelos Pikrakis
备注:6 pages, 3 figures, accepted to 21st International Conference on Content-based Multimedia Indexing (CBMI) 2024, code available this https URL
摘要:大多数当代音乐标记系统依赖于大量的注释数据。作为一种替代方案,我们调查在何种程度上合成生成的音乐摘录可以提高标签系统时,只有小的注释集合。为此,我们发布了GTZAN-synth,这是一个合成数据集,它遵循了著名的GTZAN数据集的分类法,同时数据量是它的十倍。我们首先观察到,简单地将这个合成数据集添加到GTZAN的训练分割中并不会导致性能提高。然后,我们继续调查领域适应,迁移学习和微调策略,为手头的任务,并得出结论,最后两个选项产生的准确性增加。总体而言,所提出的方法可以被认为是在一个有前途的领域,为未来的研究的第一个指南。
摘要:Most contemporary music tagging systems rely on large volumes of annotated data. As an alternative, we investigate the extent to which synthetically generated music excerpts can improve tagging systems when only small annotated collections are available. To this end, we release GTZAN-synth, a synthetic dataset that follows the taxonomy of the well-known GTZAN dataset while being ten times larger in data volume. We first observe that simply adding this synthetic dataset to the training split of GTZAN does not result into performance improvements. We then proceed to investigating domain adaptation, transfer learning and fine-tuning strategies for the task at hand and draw the conclusion that the last two options yield an increase in accuracy. Overall, the proposed approach can be considered as a first guide in a promising field for future research.

【16】 An End-to-End Speech Summarization Using Large Language Model
标题:使用大型语言模型的端到端语音摘要
链接:https://arxiv.org/abs/2407.02005
作者:Hengchao Shang,Zongyao Li,Jiaxin Guo,Shaojun Li,Zhiqiang Rao,Yuanchang Luo,Daimeng Wei,Hao Yang
备注:InterSpeech 2024
摘要:摘要语音摘要(SSum)旨在从语音内容中生成类似于人类的文本摘要。它在处理长语音输入和捕获长语音输入与短文本摘要之间复杂的跨模态映射方面遇到了困难。大型语言模型(LLM)和多模态信息融合的研究为解决这些挑战提供了新的见解。在本文中,我们提出了一个端到端的SSUM模型,利用Q-Former作为音频文本模态的连接器,并采用LLM直接从语音特征生成文本摘要。我们采用了一种多阶段的训练方法,包括基于LLM的ASR和文本摘要(TSum)任务作为辅助任务。ASR任务用于对齐特征空间并增强LLM处理较长语音的能力。然后,我们利用课程学习策略,以促进模型的过渡,从TSum到SSum。最后,我们的模型在How-2数据集上实现了具有竞争力的性能。
摘要:Abstractive Speech Summarization (SSum) aims to generate human-like text summaries from spoken content. It encounters difficulties in handling long speech input and capturing the intricate cross-modal mapping between long speech inputs and short text summaries. Research on large language models (LLMs) and multimodal information fusion has provided new insights for addressing these challenges. In this paper, we propose an end-to-end SSum model that utilizes Q-Former as a connector for the audio-text modality and employs LLMs to generate text summaries directly from speech features. We adopt a multi-stage training approach that includes LLM based ASR and Text Summarization (TSum) tasks as auxiliary tasks. ASR tasks are used to align feature spaces and enhance the LLM's ability to handle longer speech. Then, we utilize a curriculum learning strategy to facilitate the model's transition from TSum to SSum. Finally, our model achieves competitive performance on the How-2 dataset.


【17】 SAVE: Segment Audio-Visual Easy way using Segment Anything Model
标题:保存:使用Segment Anything模型轻松地分段视听
链接:https://arxiv.org/abs/2407.02004
作者:Khanh-Binh Nguyen,Chae Jung Park
摘要:视听分割(AVS)的主要目的是通过在像素级准确预测分割掩模来精确识别和定位视觉场景中的听觉元素。实现这一目标需要全面考虑数据和模型方面,以有效地解决这一任务。这项研究提出了一种轻量级的方法,SAVE,它有效地适应预先训练的段任何模型(SAM)的AVS任务。通过将图像编码器适配器结合到Transformer块中以更好地捕获不同的数据集信息,并提出残余音频编码器适配器以将音频特征编码为稀疏提示,我们提出的模型在编码阶段实现了有效的视听融合和交互。我们提出的方法通过将输入分辨率从1024像素降低到256像素来加快训练和推理速度,同时与以前的SOTA相比实现了更高的性能。大量的实验验证了我们的方法,表明我们提出的模型优于其他SOTA方法显着。此外,利用合成数据的预训练模型增强了真实AVSBench数据的性能,在S4(V1S)子集上实现了84.59 mIoU,在MS3(V1M)集上实现了70.28 mIoU,输入图像只有256个像素。这在S4(V1S)上增加到86.16 mIoU,在MS3(V1M)上增加到70.83 mIoU,输入为1024像素。
摘要:The primary aim of Audio-Visual Segmentation (AVS) is to precisely identify and locate auditory elements within visual scenes by accurately predicting segmentation masks at the pixel level. Achieving this involves comprehensively considering data and model aspects to address this task effectively. This study presents a lightweight approach, SAVE, which efficiently adapts the pre-trained segment anything model (SAM) to the AVS task. By incorporating an image encoder adapter into the transformer blocks to better capture the distinct dataset information and proposing a residual audio encoder adapter to encode the audio features as a sparse prompt, our proposed model achieves effective audio-visual fusion and interaction during the encoding stage. Our proposed method accelerates the training and inference speed by reducing the input resolution from 1024 to 256 pixels while achieving higher performance compared with the previous SOTA. Extensive experimentation validates our approach, demonstrating that our proposed model outperforms other SOTA methods significantly. Moreover, leveraging the pre-trained model on synthetic data enhances performance on real AVSBench data, achieving 84.59 mIoU on the S4 (V1S) subset and 70.28 mIoU on the MS3 (V1M) set with only 256 pixels for input images. This increases up to 86.16 mIoU on the S4 (V1S) and 70.83 mIoU on the MS3 (V1M) with inputs of 1024 pixels.


【18】 Investigating the Effects of Large-Scale Pseudo-Stereo Data and  Different Speech Foundation Model on Dialogue Generative Spoken Language  Model
标题:研究大规模伪立体声数据和不同语音基础模型对对话生成口语模型的影响
链接:https://arxiv.org/abs/2407.01911
作者:Yu-Kuan Fu,Cheng-Kuang Lee,Hsiu-Hsuan Wang,Hung-yi Lee
备注:submitted to interspeech 2024
摘要:最近的努力,在口语对话建模的目的是合成口语对话,而不需要直接转录,从而保留了丰富的非文本信息固有的讲话。然而,当说话者同时说话时,这种方法面临挑战,需要在单独的通道上记录说话者的立体声对话数据,这是一种非常稀缺的资源。为了解决这个问题,我们开发了一种创新的管道,能够将单声道对话数据转换为伪立体声数据。这将我们的训练数据集从仅仅2,000小时扩展到令人印象深刻的17,600小时,大大丰富了可用训练示例的多样性和质量。包含该伪立体声数据已被证明在提高口语对话语言模型的性能方面是有效的。此外,我们探讨了使用不同的语音基础模型的离散单元的口语对话生成。
摘要:Recent efforts in Spoken Dialogue Modeling aim to synthesize spoken dialogue without the need for direct transcription, thereby preserving the wealth of non-textual information inherent in speech. However, this approach faces a challenge when speakers talk simultaneously, requiring stereo dialogue data with speakers recorded on separate channels, a notably scarce resource. To address this, we have developed an innovative pipeline capable of transforming single-channel dialogue data into pseudo-stereo data. This expanded our training dataset from a mere 2,000 to an impressive 17,600 hours, significantly enriching the diversity and quality of the training examples available. The inclusion of this pseudo-stereo data has proven to be effective in improving the performance of spoken dialogue language models. Additionally, we explored the use of discrete units of different speech foundation models for spoken dialogue generation.

【19】 Pinyin Regularization in Error Correction for Chinese Speech Recognition  with Large Language Models
标题:大语言模型中文语音识别错误纠正中的拼音正规化
链接:https://arxiv.org/abs/2407.01909
作者:Zhiyuan Tang,Dong Wang,Shen Huang,Shidong Shang
备注:Interspeech 2024
摘要:最近的研究已经证明了大语言模型(LLM)在自动语音识别(ASR)纠错中的有效性。然而,大部分研究都集中在英语语言上。本文将注意力转向中文。首先,我们构建了一个专门的基准数据集,旨在为724 K假设-转录对的中文ASR纠错,命名为中国假设天堂数据集(ChineseHP),它包含了广泛的场景,并提出了重大的挑战。随后,我们使用数据集对直接提示和微调预训练的LLM进行了初步评估。此外,我们提出了一个简单的方法,拼音正则化的提示,这涉及到拼音的转录直接从文本的假设。实验结果表明,与未正则化的LLM相比,拼音正则化一致地增强了LLM的纠错能力。该数据集可在网站上查阅。
摘要:Recent studies have demonstrated the efficacy of large language models (LLMs) in error correction for automatic speech recognition (ASR). However, much of the research focuses on the English language. This paper redirects the attention to Chinese. Firstly, we construct a specialized benchmark dataset aimed at error correction for Chinese ASR with 724K hypotheses-transcription pairs, named the Chinese Hypotheses Paradise dataset (ChineseHP), which contains a wide range of scenarios and presents significant challenges. Subsequently, we conduct a preliminary evaluation using the dataset for both direct-prompting and fine-tuning pre-trained LLMs. Furthermore, we propose a straightforward method of Pinyin regularization for prompts, which involves the transcription of Pinyin directly from text hypotheses. The experimental results reveal that Pinyin regularization consistently enhances the error-correcting ability of LLMs when compared with those without regularization. The dataset is available on the website.

【20】 Constant Directivity Loudspeaker Beamforming
标题:恒定方向性扬声器射束成形
链接:https://arxiv.org/abs/2407.01860
作者:Yuancheng Luo
备注:Accepted at EUSIPCO 2024
摘要:扬声器阵列波束成形是用于声学指向性控制和鲁棒音频再现的常见信号处理技术。与它们的麦克风对应物不同,扬声器约束通常是异质的,这是由于阵列换能器在频率、声电灵敏度、效率和方向性方面具有不同的操作范围。这项工作提出了一种频率正则化方法的广义瑞利商方向性规范和两个新的波束形成器的设计,优化最大效率恒定方向性(MECD)和最大灵敏度恒定方向性(MSCD)。我们得到快速收敛和解析解,从他们的二次等式约束的二次规划公式。实验优化广义方向性指数约束的波束形成器设计的全频带异构阵列。
摘要:Loudspeaker array beamforming is a common signal processing technique for acoustic directivity control and robust audio reproduction. Unlike their microphone counterpart, loudspeaker constraints are often heterogeneous due to arrayed transducers with varying operating ranges in frequency, acoustic-electrical sensitivity, efficiency, and directivity. This work proposes a frequency-regularization method for generalized Rayleigh quotient directivity specifications and two novel beamformer designs that optimize for maximum efficiency constant directivity (MECD) and maximum sensitivity constant directivity (MSCD). We derive fast converging and analytic solutions from their quadratic equality constrained quadratic program formulations. Experiments optimize generalized directivity index constrained beamformer designs for a full-band heterogeneous array.

【21】 Meerkat: Audio-Visual Large Language Model for Grounding in Space and  Time
标题:Meerkat:基于时空的视听大型语言模型
链接:https://arxiv.org/abs/2407.01851
作者:Sanjoy Chowdhury,Sayan Nag,Subhrajyoti Dasgupta,Jun Chen,Mohamed Elhoseiny,Ruohan Gao,Dinesh Manocha
备注:Accepted at ECCV 2024
摘要:利用大型语言模型在基于文本的任务中的出色能力,最近关于多模态LLM(MLLM)的工作将其扩展到其他模态,如视觉和音频。然而,在这些方向上的进展主要集中在只需要对视听语义进行粗粒度理解的任务上。我们提出了Meerkat,一个视听LLM配备了图像和音频在空间和时间上的细粒度的理解。凭借基于最佳传输的新模态对齐模块和加强视听一致性的交叉注意模块,Meerkat可以解决具有挑战性的任务,例如音频参考图像接地,图像引导音频时间定位和视听事实检查。此外,我们精心策划了一个大型数据集AVFIT,其中包括从开源数据集收集的3M指令调优样本,并引入了统一了五个具有挑战性的视听任务的MeerkatBench。我们在所有这些下游任务上都实现了最先进的性能,相对改善高达37.12%。
摘要:Leveraging Large Language Models' remarkable proficiency in text-based tasks, recent works on Multi-modal LLMs (MLLMs) extend them to other modalities like vision and audio. However, the progress in these directions has been mostly focused on tasks that only require a coarse-grained understanding of the audio-visual semantics. We present Meerkat, an audio-visual LLM equipped with a fine-grained understanding of image and audio both spatially and temporally. With a new modality alignment module based on optimal transport and a cross-attention module that enforces audio-visual consistency, Meerkat can tackle challenging tasks such as audio referred image grounding, image guided audio temporal localization, and audio-visual fact-checking. Moreover, we carefully curate a large dataset AVFIT that comprises 3M instruction tuning samples collected from open-source datasets, and introduce MeerkatBench that unifies five challenging audio-visual tasks. We achieve state-of-the-art performance on all these downstream tasks with a relative improvement of up to 37.12%.

【22】 Deepfake Audio Detection Using Spectrogram-based Feature and Ensemble of  Deep Learning Models
标题:使用基于谱图的特征和深度学习模型集成的Deepfake音频检测
链接:https://arxiv.org/abs/2407.01777
作者:Lam Pham,Phat Lam,Truong Nguyen,Huyen Nguyen,Alexander Schindler
摘要:None
摘要:In this paper, we propose a deep learning based system for the task of deepfake audio detection. In particular, the draw input audio is first transformed into various spectrograms using three transformation methods of Short-time Fourier Transform (STFT), Constant-Q Transform (CQT), Wavelet Transform (WT) combined with different auditory-based filters of Mel, Gammatone, linear filters (LF), and discrete cosine transform (DCT). Given the spectrograms, we evaluate a wide range of classification models based on three deep learning approaches. The first approach is to train directly the spectrograms using our proposed baseline models of CNN-based model (CNN-baseline), RNN-based model (RNN-baseline), C-RNN model (C-RNN baseline). Meanwhile, the second approach is transfer learning from computer vision models such as ResNet-18, MobileNet-V3, EfficientNet-B0, DenseNet-121, SuffleNet-V2, Swint, Convnext-Tiny, GoogLeNet, MNASsnet, RegNet. In the third approach, we leverage the state-of-the-art audio pre-trained models of Whisper, Seamless, Speechbrain, and Pyannote to extract audio embeddings from the input spectrograms. Then, the audio embeddings are explored by a Multilayer perceptron (MLP) model to detect the fake or real audio samples. Finally, high-performance deep learning models from these approaches are fused to achieve the best performance. We evaluated our proposed models on ASVspoof 2019 benchmark dataset. Our best ensemble model achieved an Equal Error Rate (EER) of 0.03, which is highly competitive to top-performing systems in the ASVspoofing 2019 challenge. Experimental results also highlight the potential of selective spectrograms and deep learning approaches to enhance the task of audio deepfake detection.

机器翻译由腾讯交互翻译提供,仅供参考