今年入选 ICASSP 2023 的论文中,说话人识别(声纹识别)方向约有64篇,初步划分为Speaker Verification(31篇)、Speaker Recognition(9篇)、Speaker Diarization(17篇)、Anti-Spoofing(4篇)、others(3篇)五种类型。

本文是 ICASSP 2023说话人识别方向论文合集系列第二期,整理了Speaker Verification后16篇和Speaker Diarization部分的17篇。
Speaker Verification
16.Margin-Mixup: A Method For Robust Speaker Verification In Multi-Speaker Audio
标题:边缘混合:一种多说话人音频中的鲁棒性说话人验证方法
作者:Jenthe Thienpondt;Nilesh Madhu,;Kris Demuynck
单位:IDLab, Department of Electronics and Information Systems,Ghent University - imec, Belgium
链接:https://ieeexplore.ieee.org/document/10095305
摘要:本文涉及对多个重叠说话人的音频进行说话人验证的任务。大多数说话人验证系统的设计都假设给定音频段中存在单个说话人。然而,在现实世界中,这个假设并不总是成立。在本文中,我们证明当前的说话人验证系统对于具有明显说话人重叠的音频并不稳健。为了缓解这个问题,我们提出了 margin-mixup,这是一种简单的训练策略,可以轻松地被现有的说话人验证管道采用,以使生成的说话人嵌入对多说话人音频具有鲁棒性。与其他方法相比,边缘混合不需要改变常规说话人验证架构,同时获得更好的结果。在我们基于 VoxCeleb1 的多说话人测试集上,相对于我们最先进的说话人验证基线系统,所提出的边际混合策略将 EER 平均提高了 44.4%。
This paper is concerned with the task of speaker verification on audio with multiple overlapping speakers. Most speaker verification systems are designed with the assumption of a single speaker being present in a given audio segment. However, in a real-world setting this assumption does not always hold. In this paper, we demonstrate that current speaker verification systems are not robust against audio with noticeable speaker overlap. To alleviate this issue, we propose marginmixup, a simple training strategy that can easily be adopted by existing speaker verification pipelines to make the resulting speaker embeddings robust against multi-speaker audio. In contrast to other methods, margin-mixup requires no alterations to regular speaker verification architectures, while attaining better results. On our multi-speaker test set based on VoxCeleb1, the proposed margin-mixup strategy improves the EER on average with 44.4% relative to our state-of-the-art speaker verification baseline systems.
17.Noise-Disentanglement Metric Learning for Robust Speaker Verification
标题:用于鲁棒说话者验证的噪声解缠度量学习
作者:Yao Sun 1 , Hanyi Zhang 1 , Longbiao Wang 1,∗ , Kong Aik Lee 2,∗ , Meng Liu 1 , Jianwu Dang 1
单位:1Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University, Tianjin, China
2 Institute for Infocomm Research, A ⋆STAR, Singapore
链接:https://ieeexplore.ieee.org/document/10096848
摘要:自动说话人验证 (ASV) 在嘈杂的环境中会出现性能下降的问题。为了解决这个问题,我们提出了噪声解缠度量学习来减少与说话人无关的噪声分量并构建噪声不变的嵌入空间。具体来说,解耦模块包括说话人编码器和重构模块,专门用于解耦语音信号。说话人编码器用于解开与说话人相关的组件,重构模块通过重构信号来增加模型约束噪声信息的能力。此外,还引入了分布优化来监督噪声环境下说话人嵌入的空间结构。Vox-Celeb1 上的实验表明,所提出的方法提高了说话人验证系统在干净和噪声条件下的性能。
Automatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglement module, including the speaker encoder and reconstruction module, is dedicated to decoupling speech signals. The speaker encoder is used to disentangle speaker related components, and the reconstruction module increases the model’s ability to constrain the noise information by reconstructing the signal. In addition, distribution optimization is introduced to supervise the spatial structure of speaker embeddings under noisy environments. Experiments on VoxCeleb1 indicate that the proposed method improves the performance of the speaker verification system in both clean and noisy conditions.
18.Optimal Transport With A Diversified Memory Bank For Cross-Domain Speaker Verification
标题:跨域多样化内存库优化传输说话人验证
作者:Ruiteng Zhang 1 , Jianguo Wei 1,2 , Xugang Lu 3 , Wenhuan Lu 1 , Di Jin 1 , Lin Zhang 4 , Junhai Xu 1
单位:1College of Intelligence and Computing, Tianjin University, Tianjin, China
2Computer College, Qinghai Nationalities University, Xining, China
3National Institute of Information and Communications Technology, Kyoto, Japan
4National Institute of Informatics, Tokyo, Japan
链接:https://ieeexplore.ieee.org/document/10095876
摘要:通过将说话人的概率分布从源域转换到目标域,最优传输(OT)可以应用于说话人验证(SV)中的跨域适应。然而,在涉及过大类别(说话者)或难以区分的样本的场景中,OT 通常难以计算有效的传输。为了应对这一挑战,我们提出了一种基于 OT 的无监督域适应(UDA)框架,用于 SV、OT 并具有多样化的内存库,称为 DMB-OT,它通过两种策略确保传输的准确性:(1)规范化解决方案 OT 空间,尝试以高置信度规划来自同一说话人的音频样本之间的转换;(2)集成了动态课程学习算法,避免了OT在UDA前期基于硬判别样本计算传输耦合。不同目标领域下的实验表明,我们的无监督 DMB-OT 可以显着提高基于 OT 的 UDA 的性能,甚至可以与基于 PLDA 的监督自适应的性能相匹配。
Optimal transport (OT) can be applied to cross-domain adaptation in speaker verification (SV) by converting speakers’ probability distributions from source to target domains. However, in scenarios involving over-massive categories (speakers) or difficult samples in discrimination, OT often has difficulty computing effective transports. To address this challenge, we propose an OT-based unsupervised domain adaptation (UDA) framework for SV, OT with a diversified memory bank, called DMB-OT, which ensures the accuracy of transfers by two strategies: (1) It regularizes the solution space of OT, which attempts to plan transformations between audio samples from the same speaker with high confidence; (2) it integrates a dynamic curriculum learning algorithm, preventing OT from calculating transport couplings based on hard-discriminative samples in the early stage of UDA. Experiments under different target domains showed that our unsupervised DMB-OT could significantly improve the performance of OT-based UDA and could even match the performance of the supervised PLDA-based adaptation.
19.Parameter-Efficient Transfer Learning Of Pre-Trained Transformer Models For Speaker Verification Using Adapters
标题:使用适配器的用于的说话人验证的预训练Transformer模型的参数高效迁移学习
作者:Junyi Peng 1 , Themos Stafylakis 2 , Rongzhi Gu 3 , Oldrˇich Plchot 1 , Ladislav Mosˇner 1
Luka´sˇ Burget 1 , Jan Cˇernocky´ 1
单位:1Brno University of Technology, Faculty of Information Technology, Speech@FIT, Czechia
2Omilia - Conversational Intelligence, Athens, Greece
3Tencent AI Lab, Shenzhen, China
链接:https://ieeexplore.ieee.org/document/10094795
摘要:最近,预训练的 Transformer 模型由于在各种下游任务中取得了巨大成功,在语音处理领域受到了越来越多的关注。然而,大多数微调方法都会更新预训练模型的所有参数,随着模型大小的增长,这会变得令人望而却步,有时会导致小数据集上的过度拟合。在本文中,我们对应用参数高效迁移学习(PETL)方法来减少适应说话人验证任务所需的可学习参数进行了全面分析。具体来说,在微调过程中,预训练的模型被冻结,只有插入每个 Transformer 块中的轻量级模块才可训练(这种方法称为适配器)。此外,为了提高跨语言低资源场景下的性能,Transformer 模型在大型中间数据集上进一步调优,然后直接在小型数据集上进行微调。通过更新不到 4% 的参数,(我们提出的)基于 PETL 的方法实现了与完全微调方法相当的性能(Vox1-O:0.55%,Vox1-E:0.82%,Vox1-H:1.73%)。
Recently, the pre-trained Transformer models have received a rising interest in the field of speech processing thanks to their great success in various downstream tasks. However, most fine-tuning approaches update all the parameters of the pre-trained model, which becomes prohibitive as the model size grows and sometimes results in overfitting on small datasets. In this paper, we conduct a comprehensive analysis of applying parameter-efficient transfer learning (PETL) methods to reduce the required learnable parameters for adapting to speaker verification tasks. Specifically, during the fine-tuning process, the pre-trained models are frozen, and only lightweight modules inserted in each Transformer block are trainable (a method known as adapters). Moreover, to boost the performance in a cross language low-resource scenario, the Transformer model is further tuned on a large intermediate dataset before directly fine-tuning it on a small dataset. With updating fewer than 4% of parameters, (our proposed) PETL-based methods achieve comparable performances with full fine-tuning methods (Vox1-O: 0.55%, Vox1-E: 0.82%, Vox1-H:1.73%).
20.Pcf: Ecapa-Tdnn With Progressive Channel Fusion For Speaker Verification
标题:PCF:具有渐进式通道融合、用于说话人验证的ECAPA-TDNN
作者:Zhenduo Zhao 1,2 , Zhuo Li 1,2 , Wenchao Wang 1 , Pengyuan Zhang 1,2
单位:1Key Laboratory of Speech Acoustics and Content Understanding, Institute of Acoustics, Chinese Academy of Sciences, Beijing, China
2University of Chinese Academy of Sciences, Beijing, China
链接:https://ieeexplore.ieee.org/document/10095051
摘要:ECAPA-TDNN是目前最流行的用于说话人验证的TDNN系列模型,它刷新了TDNN模型的最先进(SOTA)性能。然而,一维卷积在特征通道上具有全局感受野。它破坏了频谱图的时频相关性。此外,由于 ECAPA-TDNN 只有五层,与 ResNet 相比,其更浅的结构限制了生成深度表示的能力。为了进一步改进 ECAPA-TDNN,我们提出了一种渐进通道融合策略,该策略将频谱图跨特征通道分割,并通过网络逐渐扩大感受野。其次,我们通过扩展深度和添加分支来扩大模型。我们提出的模型在vox1o上实现了0.718的EER和0.0858的minDCF(0.01),与ECAPA-TDNN(C=1024)相比相对提高了16.1%和19.5%。
ECAPA-TDNN is currently the most popular TDNN-series model for speaker verification, which refreshed the state-of-the-art (SOTA) performance of TDNN models. However, one-dimensional convolution has a global receptive field over the feature channel. It destroys the time-frequency relevance of the spectrogram. Besides, as ECAPA-TDNN only has five layers, a much shallower structure compared to ResNet restricts the capability to generate deep representations. To further improve ECAPA-TDNN, we propose a progressive channel fusion strategy that splits the spectrogram across the feature channel and gradually expands the receptive field through the network. Secondly, we enlarge the model by extending the depth and adding branches. Our proposed model achieves EER with 0.718 and minDCF(0.01) with 0.0858 on vox1o, relatively improved 16.1% and 19.5% compared with ECAPA-TDNN(C=1024).
21.Pretraining Conformer With Asr For Speaker Verification
标题:使用ASR预训练Conformer以进行说话人验证
作者:Danwei Cai 1, Weiqing Wang1, Ming Li 1,2, Rui Xia3, Chuanzeng Huang3
单位:1Department of Electrical and Computer Engineering, Duke University, Durham, USA
2Data Science Research Center, Duke Kunshan University, Kunshan, China
3Speech, Audio, and Music (SAMI) Group, Bytedance, China
链接:https://ieeexplore.ieee.org/document/10096659
摘要:本文提出使用自动语音识别(ASR)任务来预训练 Conformer以进行说话人验证。Conformer结合了卷积神经网络 (CNN) 和Transformer模型,分别用于建模局部和全局特征。最近,多尺度特征聚合Conformer(MFA-Conformer)被提出用于自动说话人验证。MFA-Conformer连接所有Conformer块的帧级输出以进行进一步池化。然而,我们的实验表明,Conformer很容易因有限的说话人识别训练数据而过度拟合。为了避免过度拟合,我们建议将从 ASR 中学到的知识转移到说话人验证中。具体来说,使用 ASR 预训练的Conformer来初始化MFA-Conformer、的训练,以进行说话人验证。我们的实验表明,使用ASR预训练 Conformer可以在不同模型大小上带来显着的性能提升。最佳模型在Voxceleb1-O、Voxceleb1-E和Voxceleb1-H上分别实现了0.48%、0.71% 和1.54% EER。
This paper proposes to pretrain Conformer with automatic speech recognition (ASR) task for speaker verification. Conformer combines convolution neural network (CNN) and Transformer model for modeling local and global features, respectively. Recently, multi-scale feature aggregation Conformer (MFA-Conformer) has been proposed for automatic speaker verification. MFA-Conformer concatenates framelevel outputs from all Conformer blocks for further pooling. However, our experiments show that Conformer can be easily overfitted with limited speaker recognition training data. To avoid overfitting, we propose to transfer the knowledge learned from ASR to speaker verification. Specifically, an ASR pretrained Conformer is used to initialize the training of MFA-Conformer for speaker verification. Our experiments show that pretraining Conformer with ASR leads to significant performance gains across model sizes. The best model achieves 0.48%, 0.71% and 1.54% EER on Voxceleb1-O, Voxceleb1-E, and Voxceleb1-H, respectively.
22.Probabilistic Back-Ends For Online Speaker Recognition And Clustering
标题:用于在线说话者识别和聚类的概率后端
作者:Alexey Sholokhov 1∗ , Nikita Kuzmin 2,3∗ , Kong Aik Lee 3 , Eng Siong Chng 2
单位:1Federal Research Center “Computer Science and Control” of the Russian Academy of Sciences, Moscow, Russia
2Nanyang Technological University, Singapore
3 Institute for Infocomm Research, A ⋆STAR, Singapore
链接:https://ieeexplore.ieee.org/document/10097032
摘要:本文重点关注在线说话人聚类任务中自然发生的多注册说话人识别,并研究该场景中不同评分后端的特性。首先,我们表明,流行的余弦评分因注册话语数量不同而导致分数校准不佳。其次,我们提出了基于概率线性判别分析(PLDA)的极其受限版本的余弦评分的简单替代。所提出的模型改进了多注册识别的余弦评分,同时在一对一比较的情况下保持相同的性能。最后,我们考虑一个在线说话人聚类任务,其中每个步骤自然涉及多注册识别。我们提出了一种在线聚类算法,使我们能够从 PLDA 模型中受益,例如处理不确定性和更好的分数校准的能力。我们的实验证明了所提出算法的有效性。
This paper focuses on multi-enrollment speaker recognition which naturally occurs in the task of online speaker clustering, and studies the properties of different scoring back-ends in this scenario. First, we show that popular cosine scoring suffers from poor score calibration with a varying number of enrollment utterances. Second, we propose a simple replacement for cosine scoring based on an extremely constrained version of probabilistic linear discriminant analysis (PLDA). The proposed model improves over the cosine scoring for multi-enrollment recognition while keeping the same performance in the case of one-to-one comparisons. Finally, we consider an online speaker clustering task where each step naturally involves multi enrollment recognition. We propose an online clustering algorithm allowing us to take benefits from the PLDA model such as the ability to handle uncertainty and better score calibration. Our experiments demonstrate the effectiveness of the proposed algorithm.
23.Pushing The Limits Of Self-Supervised Speaker Verification Using Regularized Distillation Framework
标题:使用正则化蒸馏框架突破自监督扬声器验证的极限
作者:Yafeng Chen;Siqi Zheng;Hui Wang;Luyao Cheng;Qian Chen
单位:Speech Lab of DAMO Academy, Alibaba Group
链接:https://ieeexplore.ieee.org/document/10096915
摘要:在没有说话人标签的情况下训练强大的说话人验证系统一直是一项具有挑战性的任务。先前的研究观察到自监督方法和完全监督方法之间存在巨大的性能差距。在本文中,我们应用了一种称为“带 NO 标签的蒸馏”(DINO) 的非对比自监督学习框架,并提出了两个应用于 DINO 中嵌入的正则化项。一个正则化项保证了嵌入的多样性,而另一个正则化项则使每个嵌入的变量去相关。在时域和频域上探索了各种数据增强技术的有效性。在 VoxCeleb 数据集上进行的一系列实验证明了正则化 DINO 框架在说话人验证方面的优越性。我们的方法在 VoxCeleb 上的单级自监督设置下实现了最先进的说话人验证性能。
Training robust speaker verification systems without speaker labels has long been a challenging task. Previous studies observed a large performance gap between self-supervised and fully supervised methods. In this paper, we apply a noncontrastive self-supervised learning framework called DIstillation with NO labels (DINO) and propose two regularization terms applied to embeddings in DINO. One regularization term guarantees the diversity of the embeddings, while the other regularization term decorrelates the variables of each embedding. The effectiveness of various data augmentation techniques are explored, on both time and frequency domain. A range of experiments conducted on the VoxCeleb datasets demonstrate the superiority of the regularized DINO framework in speaker verification.Our method achieves the stateof-the-art speaker verification performance under a singlestage self-supervised setting on VoxCeleb.
24.Self-Supervised Audio-Visual Speaker Representation With Co-Meta Learning
标题:使用 CO-META 学习进行自监督视听演讲者表示
作者:Hui Chen 1,2 , Hanyi Zhang 2 , Longbiao Wang 2,∗ , Kong Aik Lee 3,∗ , Meng Liu 2 , Jianwu Dang 2
单位:1Tianjin International Engineering Institute, Tianjin University, Tianjin, China
2Tianjin Key Laboratory of Cognitive Computing and Application, College of Intelligence and Computing, Tianjin University, Tianjin, China
3 Institute for Infocomm Research, A ⋆STAR, Singapore
链接:https://ieeexplore.ieee.org/document/10096925
摘要:在自监督说话人验证中,伪标签的质量决定了其性能的上限,并且最终出现大量不可靠的伪标签的情况并不罕见。我们观察到,不同方式的互补信息确保了音频和视觉表示学习的强大监督信号。这促使我们提出一种名为Co-Meta Learning的视听自监督学习框架。受协同教学+的启发,我们设计了一种策略,允许通过分歧更新来协调两种模式的信息。此外,我们使用与模型无关的元学习(MAML)的思想来更新网络参数,这使得两种模态的硬样本可以通过梯度正则化被另一种模态更好地解决。与基线相比,我们提出的方法在Voxceleb1评估数据集的 Vox-O、Vox-E 和 Vox-H 试验上分别实现了29.8%、11.7%和12.9%的相对改进。
In self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory signal for audio and visual representation learning. This motivates us to propose an audio-visual self-supervised learning framework named Co-Meta Learning. Inspired by the Coteaching+, we design a strategy that allows the information of two modalities to be coordinated through the Update by Disagreement. Moreover, we use the idea of model-agnostic meta learning (MAML) to update the network parameters, which makes the hard samples of two modalities to be better resolved by the other modality through gradient regularization. Compared to the baseline, our proposed method achieves a 29.8%, 11.7% and 12.9% relative improvement on Vox-O, Vox-E and Vox-H trials of Voxceleb1 evaluation dataset respectively.
25.Short-Segment Speaker Verification Using Ecapa-Tdnn With Multi-Resolution Encoder
标题:使用带有多分辨率编码器的 ECAPA-TDNN 进行短段说话人验证
作者:Sangwook Han;Youngdo Ahn;Kyeongmuk Kang;Jong Won Shin
单位:School of Electrical Engineering and Computer Science, Gwangju Institute of Science and Technology, Gwangju, Korea
链接:https://ieeexplore.ieee.org/document/10096839
摘要:时域方法已显示出提高说话人验证性能的潜力,但主要方法仍然利用手工制作的特征,例如梅尔滤波器组能量。尽管这些功能基于语音感知模型并表现出令人印象深刻的性能,但固定的帧大小无法同时实现良好的时间和频谱分辨率,并且在获取幅度谱和频率重新缩放期间会存在信息丢失。在本文中,我们建议将多分辨率时域信息合并到 ECAPA-TDNN 说话人验证系统中。我们构建了一个多分辨率编码器来提取不同时间分辨率的多个特征,并让提取的特征驱动适配器模块。实验结果表明,当VoxCeleb数据集的输入长度为 2 秒或更短时,所提出的方法优于其他最近提出的方法。所提出的方法还在Google Speech Commands数据集v2上显示出卓越的性能。
Time-domain approaches have shown the potential to improve the performance of speaker verification, but still predominant approaches utilize hand-crafted features such as the mel filterbank energies. Although these features are based on speech perception models and exhibited impressive performances, the fixed frame size does not allow good temporal and spectral resolutions at the same time and there is information loss when taking the magnitude spectrum and during frequency rescaling. In this paper, we propose to incorporate multi resolution time-domain information into the ECAPA-TDNN speaker verification system. We construct a multi-resolution encoder to extract multiple features in different temporal resolutions, and let the extracted features drive the adapter modules. Experimental results showed that the proposed method outperformed other recently proposed approaches when the input length was 2 seconds or shorter for the VoxCeleb dataset. The proposed approach also showed superior performance on the Google Speech Commands dataset v2.
26.Stargan-Vc Based Cross-Domain Data Augmentation For Speaker Verification
标题:基于 STARGAN-VC 的跨域数据增强,用于说话人验证
作者:Hang-Rui Hu 1,2 , Yan Song 1 , Jian-Tao Zhang 1 , Li-Rong Dai 1 , Ian McLoughlin 1
Zhu Zhuo 2 , Yu Zhou 2 , Yu-Hong Li 2 , Hui Xue 2
单位:1National Engineering Research Center for Speech and Language Information Processing, University of Science and Technology of China, Hefei, China.
2Alibaba Group, China.
链接:https://ieeexplore.ieee.org/document/10094698
摘要:自动说话人验证(ASV)在实际应用中面临着由于录音设备和说话风格等内在因素和外在因素不匹配而导致的域转移,从而导致严重的性能下降。由于单说话人多条件(SSMC)数据在实践中很难收集,现有的域自适应方法很难保证同一类不同域的特征一致性。为此,我们提出了一种跨域数据生成方法来获得域不变的 ASV 系统。受语音转换(VC)任务的启发,基于 StarGAN 的生成模型首先从 SSMC 数据中学习跨域映射,然后为所有说话人生成缺失的域数据,从而增加训练集的类内多样性。考虑到ASV和VC任务的差异,我们更新了相应的训练目标和网络结构,使适应任务具有针对性。评估结果显示,在 minDCF 和 EER 方面,相对基准性能提高了约 5-8%,优于同等规模的 CNSRC 获胜者系统。
Automatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors, such as recording device and speaking style, in real-world applications, which leads to severe performance degradation. Since single-speaker multi-condition (SSMC) data is difficult to collect in practice, existing domain adaptation methods are hard to ensure the feature consistency of the same class but different domains. To this end, we propose a cross-domain data generation method to obtain a domain-invariant ASV system. Inspired by voice conversion (VC) task, a StarGAN based generative model first learns cross-domain mappings from SSMC data, and then generates missing domain data for all speakers, thus increasing the intra-class diversity of the training set. Considering the difference between ASV and VC task, we renovate the corresponding training objectives and network structure to make the adaptation task-specific. Evaluations on achieve a relative performance improvement of about 5-8% over the baseline in terms of minDCF and EER, outperforming the CNSRC winner’s system of the equivalent scale.
27.Step Restriction For Improving Adversarial Attacks
标题:改善对抗性攻击的步骤限制
作者:Keita Goto;Shinta Otake;Rei Kawakami;Nakamasa Inoue
单位:Tokyo Institute of Technology, Japan
链接:https://ieeexplore.ieee.org/document/10094644
摘要:我们提出了一种算法,可以自动限制迭代优化过程中的步长,并应用于对说话人验证模型的对抗性攻击。该算法动态地确定一个限制半径为r的子空间,在每次迭代时对其应用泰勒近似,然后使用投影梯度法求解该子空间内的线性问题。在实验中,我们演示了对三种说话人验证模型的对抗性攻击:i-vectors、SE-ResNet-34 和 ECAPATDNN。我们表明,所提出的算法产生的对抗扰动程度小于传统攻击方法产生的对抗扰动程度。
We propose an algorithm to automatically restrict the step size in the iterative optimization process with an application to adversarial attacks on speaker verification models. The proposed algorithm dynamically determines a subspace with a restriction radius r to which the Taylor approximation is applied at each iteration and then solves a linear problem within the subspace by using the projected gradient method. In experiments, we demonstrate adversarial attacks on three speaker verification models: i-vectors, SE-ResNet-34, and ECAPATDNN. We show that the degree of adversarial perturbations generated by the proposed algorithm is smaller than that generated by the conventional attack method.
28.Study On The Fairness Of Speaker Verification Systems Across Accent And Gender Groups
标题:跨口音和性别群体的说话人验证系统的公平性研究
作者:Mariel Estevez;Luciana Ferrer
单位:Instituto de Investigación en Ciencias de la Computación, UBA-CONICET, Argentina
链接:https://ieeexplore.ieee.org/document/10095150
摘要:说话人验证 (SV) 系统目前用于执行诸如授予银行帐户访问权限或做出取证决策等后续任务。确保这些系统公平且不偏袒任何特定群体至关重要。在这项工作中,我们分析了两个基于 X 向量的 SV 系统在不同群体中的性能,这些群体是由性别和说话者口音定义的。为此,我们通过从不同国家口音的说话者中选择样本,基于 VoxCeleb 语料库创建了一个新的数据集。我们使用该数据集来评估使用 VoxCeleb 数据训练的 SV 系统的系统性能。我们发现,在训练中代表性不足的群体(女性和非母语英语口音的说话者)中,使用校准敏感指标测量的表现明显下降。最后,我们证明了一种简单的数据平衡方法可以减轻少数群体的这种不良偏见,而不会降低多数群体的表现。
Speaker verification (SV) systems are currently used for consequential tasks like giving access to bank accounts or making forensic decisions. Ensuring that these systems are fair and do not disfavor any particular group is crucial. In this work, we analyze the performance of two X-vector-based SV systems across groups defined by gender and accent of the speakers when speaking English. To this end, we created a new dataset based on the VoxCeleb corpus by selecting samples from speakers with accents from different countries. We used this dataset to evaluate system performance of SV systems trained with VoxCeleb data. We show that performance, measured with a calibration-sensitive metric, is markedly degraded on groups that are underrepresented in training: females and speakers with nonnative accents in English. Finally, we show that a simple data balancing approach mitigates this undesirable bias on the minority groups without degrading performance on the majority groups.
29.Towards A Unified Conformer Structure: From Asr To Asv Task
标题:迈向统一的整合者结构:从 ASR 到 ASV 任务
作者:Dexin Liao 1 , Tao Jiang †3 , Feng Wang 1 , Lin Li 2 , Qingyang Hong ∗1
单位:1School of Informatics, Xiamen University, China
2School of Electronic Science and Engineering, Xiamen University, China
3Xiamen Talentedsoft Co., Ltd., China
链接:https://ieeexplore.ieee.org/document/10095433
摘要:Transformer 凭借其强大的自注意力机制,在自然语言处理和计算机视觉任务中取得了非凡的表现,其变体 Conformer 已成为自动语音识别(ASR)领域最先进的架构。然而,自动说话人验证(ASV)的主流架构是卷积神经网络,基于 Conformer 的 ASV 仍有很大的研究空间。在本文中,我们首先将 Conformer 架构从 ASR 修改为 ASV,并进行了非常小的更改。采用长度尺度注意力(LSA)方法和锐度感知最小化(SAM)来提高模型泛化能力。在 VoxCeleb 和 CN-Celeb 上进行的实验表明,与流行的 ECAPA-TDNN 相比,我们基于 Conformer 的 ASV 实现了具有竞争力的性能。其次,受迁移学习策略的启发,ASV Conformer 很自然地从预训练的 ASR 模型中初始化。通过参数传递,自注意力机制可以更好地关注序列特征之间的关系,在 VoxCeleb 和 CN-Celeb 测试集上的 EER 相对提高了 11%,这揭示了 Conformer 统一 ASV 和 ASR 任务的潜力。最后,我们在 ASV-Subtools 中提供了一个运行时来评估其在生产场景中的推理速度。我们的代码发布于https://github.com/Snowdar/asv- subtools/tree/master/doc/papers/conformer.md。
Transformer has achieved extraordinary performance in Natural Language Processing and Computer Vision tasks thanks to its powerful self-attention mechanism, and its variant Conformer has become a state-of-the-art architecture in the field of Automatic Speech Recognition (ASR). However, the main-stream architecture for Automatic Speaker Verification (ASV) is convolutional Neural Networks, and there is still much room for research on the Conformer based ASV. In this paper, firstly, we modify the Conformer architecture from ASR to ASV with very minor changes. Length-Scaled Attention (LSA) method and Sharpness-Aware Minimization (SAM) are adopted to improve model generalization. Experiments conducted on VoxCeleb and CN-Celeb show that our Conformer based ASV achieves competitive performance com-pared with the popular ECAPA-TDNN. Secondly, inspired by the transfer learning strategy, ASV Conformer is natural to be initialized from the pretrained ASR model. Via parameter transferring, self-attention mechanism could better focus on the relationship between sequence features, brings about 11% relative improvement in EER on test set of VoxCeleb and CN-Celeb, which reveals the potential of Conformer to unify ASV and ASR task. Finally, we provide a runtime in ASV-Subtools to evaluate its inference speed in production scenario. Our code is released at :
https://github.com/Snowdar/asvsubtools/tree/master/doc/papers/conformer.md.
30.Unsupervised Speaker Verification Using Pre-Trained Model And Label Correction
标题:使用预训练模型和标签校正进行无监督说话者验证
作者:Zhicong Chen †1 , Jie Wang †1 , Wenxuan Hu 2 , Lin Li ∗1 , Qingyang Hong 2
单位:1School of Electronic Science and Engineering, Xiamen University, China
2School of Informatics, Xiamen University, China
链接:https://ieeexplore.ieee.org/document/10094610
摘要:最近,微调预训练模型框架已成为语音处理任务的一个有前途的范例。在本研究中,我们提出了一种使用预训练模型子结构 (Sub-PTM) 进行无监督说话人验证的新策略,该策略由基于 CNN 的特征提取器和多个 Transformer 块组成。为了获得初始伪标签,我们利用 Infomap 对从 Sub-PTM 中提取的表示进行聚类。然后,利用生成的伪标签来训练包含 Sub-PTM 和下游网络的说话人验证模型。我们还提出了一种在线和离线标签校正(OAO-LC)方法来减轻不正确的伪标签的影响。通过结合这些技术,我们的系统与监督基线相比取得了有竞争力的结果。
Recently, the fine-tuning pre-trained model framework has emerged as a promising paradigm for speech-processing tasks. In this study, we present a novel strategy for unsupervised speaker verification using the Sub-structure of Pre-Trained Model (Sub-PTM), which consists of a CNN-based feature extractor and several Transformer blocks. To obtain the initial pseudo labels, we utilize Infomap to perform clustering on the representations extracted from the Sub-PTM. The generated pseudo labels are then leveraged to train a speaker verification model containing a Sub-PTM and a downstream network. We also propose an Online and Offline Label Correction (OAO-LC) method to alleviate the effects of incorrect pseudo labels. By incorporating these techniques, our system achieves competitive results compared to the supervised baseline.
31.Wespeaker: A Research And Production Oriented Speaker Embedding Learning Toolkit
标题:WESPAKER:一款面向研发和生产的说话人嵌入学习工具包
作者:Hongji Wang 1,3,8 , Chengdong Liang 3,4,7 , Shuai Wang 1,2,8,∗ , Zhengyang Chen 1 , Binbin Zhang 5,8 , Xu Xiang 6 , Yanlei Deng 7 , Yanmin Qian 1,∗
单位:1MoE Key Lab of AI, X-LANCE Lab, CSE Dept, Shanghai Jiao Tong University, Shanghai, China
2Shenzhen Research Institute of Big Data, The Chinese University of Hong Kong, Shenzhen
3Tencent Ethereal Audio Lab, Tencent Corporation, Shenzhen, China
4School of Marine Science and Technology, Northwestern Polytechnical University, Xi’an, China
5Horizon Robotics, Beijing, China, 6AISpeech Ltd, Suzhou, China, 7NVIDIA, Santa Clara, USA
8WeNet Open Source Community
链接:
https://ieeexplore.ieee.org/document/10096626
摘要:说话人建模对于许多相关任务至关重要,例如说话人识别和说话人二值化。主要的建模方法是固定维向量表示,即说话人嵌入。本文介绍了一种面向研究和生产的扬声器嵌入学习工具包Wespeaker。Wespeaker 包含可扩展数据管理、最先进的说话人嵌入模型、损失函数和评分后端的实施,通过在多个说话人验证挑战赛的获胜系统中采用的结构化配方取得了极具竞争力的结果。相关配方中还展示了对其他下游任务(例如说话人分类)的应用。此外,集成了CPU和GPU兼容的部署代码,以进行面向生产的开发。该工具包可在 https://github.com/wenet-e2e/wespeaker 上公开获取。
Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research and production oriented speaker embedding learning toolkit, Wespeaker. Wespeaker contains the implementation of scalable data management, state-of-the-art speaker embedding models, loss functions, and scoring back-ends, with highly competitive results achieved by structured recipes which were adopted in the winning systems in several speaker verification challenges. The application to other downstream tasks such as speaker diarization is also exhibited in the related recipe. Moreover, CPU- and GPU-compatible deployment codes are integrated for production-oriented development. The toolkit is publicly available at https://github.com/wenet-e2e/wespeaker.
Speaker Diarization
1.Advancing the Dimensionality Deduction Of Speaker Embeddings For Speaker Diarisation: Disentangling Noise And Informing Speech Activity
标题:推进说话人嵌入的降维以实现说话人分类:解开噪声并告知语音活动
作者:You Jin Kim1;Hee-Soo Heo1;Jee-Weon Jung1;Youngki Kwon1;Bong-Jin Lee1;Joon Son Chung2
单位:1 Naver Cloud Corporation, South Korea;
2 Korea Advanced Institute of Science and Technology, South Korea
链接:https://ieeexplore.ieee.org/document/10095530
摘要: 这项工作的目的是训练适合说话人二值化的抗噪声说话人嵌入。说话人嵌入在分类系统的性能中起着至关重要的作用,但它们经常捕获噪声等虚假信息,从而对性能产生不利影响。我们之前的工作提出了一种基于自动编码器的降维模块来帮助去除冗余信息。然而,它们没有明确分离这些信息,并且还被发现对超参数值敏感。为此,我们提出了两个贡献来克服这些问题:(i)一种新颖的降维框架,可以从说话人嵌入中分离出虚假信息;(ii) 使用语音活动向量来防止说话者代码代表背景噪声。通过对四个数据集进行的一系列实验,我们的方法一致地证明了无需系统融合的模型中最先进的性能。
The objective of this work is to train noise-robust speaker embeddings adapted for speaker diarisation. Speaker embeddings play a crucial role in the performance of diarisation systems, but they often capture spurious information such as noise, adversely affecting performance. Our previous work has proposed an auto-encoder-based dimensionality reduction module to help remove the redundant information. However, they do not explicitly separate such information and have also been found to be sensitive to hyper-parameter values. To this end, we propose two contributions to overcome these issues: (i) a novel dimensionality reduction framework that can disentangle spurious information from the speaker embeddings; (ii) the use of speech activity vector to prevent the speaker code from representing the background noise. Through a range of experiments conducted on four datasets, our approach consistently demonstrates the state-of-the-art performance among models without system fusion.
2.Absolute Decision Corrupts Absolutely: Conservative Online Speaker Diarisation
标题:绝对的决定绝对会腐败:保守的在线演讲者分类
作者:Youngki Kwon;Hee-Soo Heo;Bong-Jin Lee;You Jin Kim;Jee-Weon Jung
单位:NAVER Cloud Corporation, South Korea
链接:https://ieeexplore.ieee.org/document/10096276
摘要: 我们的重点在于开发一个在线说话人分类框架,该框架在不同领域表现出强大的性能。在在线说话人分类中,实时生成的输出是不可逆的,输入会话早期阶段的一些误判可能会导致灾难性的结果。我们假设,在许多其他因素中,谨慎增加估计发言者的数量是最重要的。因此,我们提出的框架包括当系统判断过去的增加是错误的时,将发言者的数量减少1。我们还采用双缓冲区、检查点和质心,其中检查点与轮廓系数相结合来估计说话者的数量,质心代表说话者。同样,我们相信一个说话者可以生成多个质心。因此,我们设计了一种基于聚类的标签匹配技术来实时分配标签。由此产生的系统虽然重量轻,但效果却出奇地好。该系统在 DIHARD II 和 III 数据集上展示了最先进的性能,在 AMI 和 VoxConverse 测试集上也具有竞争力。
Our focus lies in developing an online speaker diarisation framework which demonstrates robust performance across diverse domains. In online speaker diarisation, outputs generated in real-time are irreversible, and a few misjudgements in the early phase of an input session can lead to catastrophic results. We hypothesise that cautiously increasing the number of estimated speakers is of paramount importance among many other factors. Thus, our proposed framework includes decreasing the number of speakers by one when the system judges that an increase in the past was faulty. We also adopt dual buffers, checkpoints and centroids, where checkpoints are combined with silhouette coefficients to estimate the number of speakers and centroids represent speakers. Again, we believe that more than one centroid can be generated from one speaker. Thus we design a clustering-based label matching technique to assign labels in realtime. The resulting system is lightweight yet surprisingly effective. The system demonstrates state-of-the-art performance on DIHARD II and III datasets, where it is also competitive in AMI and VoxConverse test sets.
3.Audio-Visual Speaker Diarization In The Framework Of Multi-User Human-Robot Interaction
标题:多用户人机交互框架下的视听说话人分类
作者:Timothée Dhaussy1;Bassam Jabaian1;Fabrice Lefèvre1;Radu Horaud2
单位:1 LIA-CERI, Avignon University;
2 Inria at Université Grenoble Alpes
链接:https://ieeexplore.ieee.org/document/10096295
摘要: 说话者分类任务回答“在给定时间谁在说话?”的问题。它代表了机器人等领域场景分析的宝贵信息。在本文中,我们介绍了一种用于多用户说话者二值化的时间视听融合模型,该模型具有计算要求低、鲁棒性良好且无需训练阶段的特点。所提出的方法通过测量声音位置和视觉存在之间的空间重合来识别主要说话者并随着时间的推移跟踪他们。该模型是生成式的,参数是在线估计的,不需要训练。其有效性是使用两个数据集进行评估的,一个是公共数据集,另一个是由 Pepper 人形机器人内部收集的数据集。
The speaker diarization task answers the question "who is speaking at a given time?". It represents valuable information for scene analysis in a domain such as robotics. In this paper, we introduce a temporal audio-visual fusion model for multiusers speaker diarization, with low computing requirement, a good robustness and an absence of training phase. The proposed method identifies the dominant speakers and tracks them over time by measuring the spatial coincidence between sound locations and visual presence. The model is generative, parameters are estimated online, and does not require training. Its effectiveness was assessed using two datasets, a public one and one collected in-house with the Pepper humanoid robot.
4.Community Detection Graph Convolutional Network For Overlap-Aware Speaker Diarization
标题:用于重叠感知说话者分类的社区检测图卷积网络
作者:Jie Wang2;Zhicong Chen2;Haodong Zhou2;Lin Li2;Qingyang Hong1
单位:1 School of Informatics, Xiamen University, China;
2 School of Electronic Science and Engineering, Xiamen University, China
链接:https://ieeexplore.ieee.org/document/10095143
摘要: 聚类算法在说话人分类系统中起着至关重要的作用。然而,传统的聚类算法受到说话人嵌入分布复杂以及缺乏挖掘会话中说话人之间潜在关系的困扰。我们提出了一种新颖的基于图的聚类方法,称为社区检测图卷积网络(CDGCN),以提高说话人二值化系统的性能。基于CDGCN的聚类方法由图生成、子图检测和基于图的重叠语音检测(Graph-OSD)组成。首先,图生成细化了语音片段之间的局部联系。其次,子图检测找到说话者图的最佳全局划分。最后,我们将重叠感知说话人二值化的说话人聚类视为重叠社区检测任务,并设计一个 Graph-OSD 组件来输出重叠感知标签。通过捕获本地和全局信息,具有 CDGCN 聚类的说话人二值化系统在 DIHARD III 语料库上优于传统的基于聚类的说话人二值化 (CSD) 系统。
The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationships between speakers in a session. We propose a novel graph-based clustering approach called Community Detection Graph Convolutional Network (CDGCN) to improve the performance of the speaker diarization system. The CDGCN-based clustering method consists of graph generation, sub-graph detection, and Graph-based Overlapped Speech Detection (Graph-OSD). Firstly, the graph generation refines the local linkages among speech segments. Secondly the sub-graph detection finds the optimal global partition of the speaker graph. Finally, we view speaker clustering for overlap-aware speaker diarization as an overlapped community detection task and design a Graph-OSD component to output overlap-aware labels. By capturing local and global information, the speaker diarization system with CDGCN clustering outperforms the traditional Clustering-based Speaker Diarization (CSD) systems on the DIHARD III corpus.
5.Frame-Wise And Overlap-Robust Speaker Embeddings For Meeting Diarization
标题:用于会议分类的逐帧和重叠鲁棒说话人嵌入
作者:Tobias Cord-Landwehr1;Christoph Boeddeker1;Cătălin Zorilă2;Rama Doddipatla2;Reinhold Haeb-Umbach1
单位:1 Department of Communications Engineering, Paderborn University, Paderborn, Germany;
2 Toshiba Cambridge Research Laboratories, Cambridge, United Kingdom
链接:https://ieeexplore.ieee.org/document/10095370
摘要: 使用师生训练方法,我们开发了一种说话人嵌入提取系统,该系统以帧速率输出嵌入。考虑到这种高时间分辨率以及学生即使对于具有语音重叠的片段也能生成合理的说话人嵌入这一事实,逐帧嵌入可以作为输入语音信号的适当表示,用于端到端神经会议二值化 (EEND) 系统。我们在实验中表明,这种表示有助于缓解 EEND 系统的一个众所周知的问题:当增加说话人数量时,二值化性能下降会显着减少。我们还引入了分块处理,以便能够对任意长的会议进行日志记录。
Using a Teacher-Student training approach we developed a speaker embedding extraction system that outputs embeddings at frame rate. Given this high temporal resolution and the fact that the student produces sensible speaker embeddings even for segments with speech overlap, the frame-wise embeddings serve as an appropriate representation of the input speech signal for an end-to-end neural meeting diarization (EEND) system. We show in experiments that this representation helps mitigate a well-known problem of EEND systems: when increasing the number of speakers the diarization performance drop is significantly reduced. We also introduce block-wise processing to be able to diarize arbitrarily long meetings.
6.High-Resolution Embedding Extractor For Speaker Diarisation
标题:用于说话人分类的高分辨率嵌入提取器
作者:Hee-Soo Heo;Youngki Kwon;Bong-Jin Lee;You Jin Kim;Jee-Weon Jung
单位:NAVER Cloud Corporation, South Korea
链接:https://ieeexplore.ieee.org/document/10097190
摘要: 说话人嵌入提取器显着影响基于聚类的说话人二值化系统的性能。传统上,仅从每个语音片段中提取一个嵌入。然而,由于滑动窗口方法,由于说话人变化点,一个片段很容易包括两个或更多说话人。本研究提出了一种新颖的嵌入提取器架构,称为高分辨率嵌入提取器(HEE),它从每个语音片段中提取多个高分辨率嵌入。Hee由特征图提取器和增强器组成,其中具有自注意力机制的增强器是成功的关键。HEE的增强器取代了聚合过程;增强器不是全局池化层,而是通过利用全局上下文的注意力将相关信息组合到每个帧。提取的密集帧级嵌入每个都可以代表一个说话者。因此,多个说话者可以由每个片段中的不同帧级特征来表示。我们还提出了一个人工生成的混合数据训练框架来训练所提出的 HEE。通过对五个评估集(包括四个公共数据集)的实验,所提出的 HEE 在每个评估集上都显示出至少 10% 的改进,但一个数据集除外,我们分析该数据集不存在快速的说话人变化。
Speaker embedding extractors significantly influence the performance of clustering-based speaker diarisation systems. Conventionally, only one embedding is extracted from each speech segment. However, because of the sliding window approach, a segment easily includes two or more speakers owing to speaker change points. This study proposes a novel embedding extractor architecture, referred to as a high-resolution embedding extractor (HEE), which extracts multiple high-resolution embeddings from each speech segment. Hee consists of a feature-map extractor and an enhancer, where the enhancer with the self-attention mechanism is the key to success. The enhancer of HEE replaces the aggregation process; instead of a global pooling layer, the enhancer combines relative information to each frame via attention leveraging the global context. Extracted dense frame-level embeddings can each represent a speaker. Thus, multiple speakers can be represented by different frame-level features in each segment. We also propose an artificially generating mixture data training framework to train the proposed HEE. Through experiments on five evaluation sets, including four public datasets, the proposed HEE demonstrates at least 10% improvement on each evaluation set, except for one dataset, which we analyse that rapid speaker changes less exist.
7.Improving Transformer-Based End-To-End Speaker Diarization By Assigning Auxiliary Losses To Attention Heads
标题:通过将辅助损耗分配给注意头来改进基于Transformer的端到端说话人二值化
作者:Ye-Rin Jeoung;Joon-Young Yang;Jeong-Hwan Choi;Joon-Hyuk Chang
单位:Department of Electronic Engineering, Hanyang University, Seoul, Republic of Korea
链接:https://ieeexplore.ieee.org/document/10095589
摘要: 基于 Transformer 的端到端神经说话人二值化 (EEND) 模型利用多头自注意力 (SA) 机制来实现重叠语音区域中准确的说话人标签预测。在本研究中,为了提高 SA-EEND 模型的训练效果,我们建议对Transformer层的 SA 头使用辅助损耗。具体来说,我们假设 SA 层的注意力权重矩阵是冗余的,如果它们的模式与单位矩阵的模式相似。然后,我们明确约束此类矩阵以展示与语音活动检测或重叠语音检测任务相关的特定说话者活动模式。因此,我们期望所提出的辅助损失能够引导Transformer层在注意力权重中表现出更多样化的模式,从而减少 SA 头中假设的冗余。使用模拟数据集和 CALLHOME 数据集进行双说话人二值化任务证明了该方法的有效性,将传统 SA-EEND 模型的二值化错误率分别降低了 32.58% 和 17.11%。
Transformer-based end-to-end neural speaker diarization (EEND) models utilize the multi-head self-attention (SA) mechanism to enable accurate speaker label prediction in overlapped speech regions. In this study, to enhance the training effectiveness of SA-EEND models, we propose the use of auxiliary losses for the SA heads of the transformer layers. Specifically, we assume that the attention weight matrices of an SA layer are redundant if their patterns are similar to those of the identity matrix. We then explicitly constrain such matrices to exhibit specific speaker activity patterns relevant to voice activity detection or overlapped speech detection tasks. Consequently, we expect the proposed auxiliary losses to guide the transformer layers to exhibit more diverse patterns in the attention weights, thereby reducing the assumed redundancies in the SA heads. The effectiveness of the proposed method is demonstrated using the simulated and CALLHOME datasets for two-speaker diarization tasks, reducing the diarization error rate of the conventional SA-EEND model by 32.58% and 17.11%, respectively.
8.In Search Of Strong Embedding Extractors For Speaker Diarisation
标题:寻找用于说话人分类的强大嵌入提取器
作者:Jee-Weon Jung2;Hee-Soo Heo2;Bong-Jin Lee2;Jaesung Huh4;Andrew Brown4;Youngki Kwon2;Shinji Watanabe3;Joon Son Chung1
单位:1 Korea Advanced Institute of Science and Technology, South Korea;
2 NAVER Corporation, South Korea;
3 Carnegie Mellon University, Pittsburgh, PA, USA;
4 Department of Engineering Science, Visual Geometry Group, University of Oxford, UK
链接:https://ieeexplore.ieee.org/document/10096449
摘要: 说话人嵌入提取器(EE)将输入音频映射到说话人判别潜在空间,在说话人二值化中至关重要。然而,采用 EE 进行二值化时存在一些挑战,我们从中解决了两个关键问题。首先,评估并不简单,因为说话者验证和分类之间所需的功能不同。我们表明,广泛采用的说话人验证评估协议的更好性能并不会带来更好的二值化性能。其次,嵌入提取器还没有发现存在多个说话者的话语。由于语音重叠和说话人变化,这些输入不可避免地出现在说话人分类中;它们会降低性能。为了缓解第一个问题,我们生成了更好地模拟二值化场景的说话人验证评估协议。我们提出了两种数据增强技术来缓解第二个问题,使嵌入提取器意识到重叠的语音或说话者变化的输入。一种技术生成重叠的语音片段,另一种技术生成两个说话者顺序说话的片段。使用三个最先进的说话人嵌入提取器进行的大量实验结果表明,所提出的两种方法都是有效的。
Speaker embedding extractors (EEs), which map input audio to a speaker discriminant latent space, are of paramount importance in speaker diarisation. However, there are several challenges when adopting EEs for diarisation, from which we tackle two key problems. First, the evaluation is not straightforward because the required features differ between speaker verification and diarisation. We show that better performance on widely adopted speaker verification evaluation protocols does not lead to better diarisation performance. Second, embedding extractors have not seen utterances in which multiple speakers exist. These inputs are inevitably present in speaker diarisation because of overlapped speech and speaker changes; they degrade the performance. To mitigate the first problem, we generate speaker verification evaluation protocols that better mimic the diarisation scenario. We propose two data augmentation techniques to alleviate the second problem, making embedding extractors aware of overlapped speech or speaker change input. One technique generates overlapped speech segments, and the other generates segments where two speakers utter sequentially. Extensive experimental results using three state-of-the-art speaker embedding extractors demonstrate that both proposed approaches are effective.
9.Multi-Speaker And Wide-Band Simulated Conversations As Training Data For End-To-End Neural Diarization
标题:多说话人和宽带模拟对话作为端到端神经二值化的训练数据
作者:ederico Landini1;Mireia Diez1;Alicia Lozano-Diez2;Lukáš Burget1
单位:1 Faculty of Information Technology, Speech@FIT, Brno University of Technology, Czechia;
2 AUDIAS (Audio, Data Intelligence and Speech), Universidad Autónoma de Madrid, Spain
链接:https://ieeexplore.ieee.org/document/10097049
摘要: 端到端二值化为标准级联二值化系统提供了一种有吸引力的替代方案,因为单个系统可以同时处理任务的所有方面。人们已经提出了许多类型的端到端模型,但所有这些模型都需要(到目前为止还不存在)大量带注释的数据进行训练。折衷的解决方案在于生成合成数据,最近提出的模拟对话(SC)比原始模拟混合物(SM)有了显着的改进。在这项工作中,我们创建了每次对话有多个说话人的 SC,并表明它们比 SM 具有更好的性能,同时也减少了对微调阶段的依赖。我们还使用宽带公共音频源创建 SC,并对几个评估集进行分析。与本出版物一起,我们发布了生成此类数据的方法和在公共集上训练的模型,以及有效处理每个对话的多个发言者和辅助语音活动检测损失的实现。
End-to-end diarization presents an attractive alternative to standard cascaded diarization systems because a single system can handle all aspects of the task at once. Many flavors of end-to-end models have been proposed but all of them require (so far non-existing) large amounts of annotated data for training. The compromise solution consists in generating synthetic data and the recently proposed simulated conversations (SC) have shown remarkable improvements over the original simulated mixtures (SM). In this work, we create SC with multiple speakers per conversation and show that they allow for substantially better performance than SM, also reducing the dependence on a fine-tuning stage. We also create SC with wide-band public audio sources and present an analysis on several evaluation sets. Together with this publication, we release the recipes for generating such data and models trained on public sets as well as the implementation to efficiently handle multiple speakers per conversation and an auxiliary voice activity detection loss.
10.Multi-Speaker End-To-End Multi-Modal Speaker Diarization System For The Misp 2022 Challenge
标题:应对 MISP 2022 挑战的多说话人端到端多模态说话人分类系统
作者:Tao Liu;Zhengyang Chen;Yanmin Qian;Kai Yu
单位:MoE Key Lab of Artificial Intelligence, AI Institute, X-LANCE Lab, Shanghai Jiao Tong University
链接:https://ieeexplore.ieee.org/document/10096327
摘要: 本文介绍了我们的系统在基于多模态信息的语音处理 (MISP) 2022 挑战赛第 1 场的设计和实现。我们设计了一个基于端到端Transformer的多说话者系统。Transformer 主干非常适合捕获长期特征,这对于缺少时间模态的情况下的多模态说话人二值化至关重要。此外,我们采用了多种损失函数和图像数据增强技术来防止训练过程中的过度拟合。此外,为了进一步提高系统性能,我们结合通道间相位差(IPD)来对位置特征进行建模,并预训练基于 ECAPA-TDNN 的模型来提取说话人嵌入特征。我们的系统在评估集上实现了 10.82% 的二值化错误率 (DER),这使我们在 MISP 2022 挑战赛的视听说话人二值化任务中获得第二名。
This paper presents the design and implementation of our system for Track 1 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We design an end-to-end transformer-based multi-talker system. The transformer backbone is well-suited to capture long-term features, which is crucial for multi-modal speaker diarization in cases where temporal modalities are missing. Besides, we employ several loss functions and image data augmentation techniques to prevent over-fitting during training. Moreover, to further improve the system’s performance, we incorporate Interchannel Phase Difference (IPD) to model the location features and pre-train an ECAPA-TDNN-based model to extract speaker embedding features. Our system achieved a diarization error rate (DER) of 10.82% on the evaluation set, which earned us second place in the audio-visual speaker diarization task of the MISP 2022 challenge.
11.Privacy-Preserving Automatic Speaker Diarization
标题:保护隐私的自动说话人分类
作者:Francisco Teixeira1;Alberto Abad1;Bhiksha Raj23;Isabel Trancoso 1
单位:1 INESC-ID/IST, University of Lisbon, Portugal;
2 LTI, Carnegie Mellon University, USA;
3 Mohammed bin Zayed University of AI, UAE
链接:https://ieeexplore.ieee.org/document/10096113
摘要: 自动说话人分类 (ASD) 是一种具有多种应用的支持技术,它处理多个说话人的录音,引起了隐私方面的特别关注。事实上,在远程设置中,录音与服务器共享,客户不仅放弃了对话的隐私,还放弃了可以从他们的声音推断出的所有信息。然而,据我们所知,隐私保护 ASD 系统的开发迄今为止一直被忽视。在这项工作中,我们结合使用安全多方计算 (SMC) 和安全模块化哈希这两种加密技术来解决这个问题,并将它们应用于级联 ASD 系统的两个主要步骤:说话人嵌入提取和凝聚层次聚类。我们的系统能够在性能和效率之间实现合理的权衡,对于两种不同的 SMC 安全设置,实时因子分别为 1.1 和 1.6。
Automatic Speaker Diarization (ASD) is an enabling technology with numerous applications, which deals with recordings of multiple speakers, raising special concerns in terms of privacy. In fact, in remote settings, where recordings are shared with a server, clients relinquish not only the privacy of their conversation, but also of all the information that can be inferred from their voices. However, to the best of our knowledge, the development of privacy-preserving ASD systems has been overlooked thus far. In this work, we tackle this problem using a combination of two cryptographic techniques, Secure Multiparty Computation (SMC) and Secure Modular Hashing, and apply them to the two main steps of a cascaded ASD system: speaker embedding extraction and agglomerative hierarchical clustering. Our system is able to achieve a reasonable trade-off between performance and efficiency, presenting real-time factors of 1.1 and 1.6, for two different SMC security settings.
12.Spectral Clustering-Aware Learning Of Embeddings For Speaker Diarisation
标题:用于说话人分类的嵌入的频谱聚类感知学习
作者:Evonne P.C. Lee;Guangzhi Sun;Chao Zhang;Philip C. Woodland
单位:Engineering Dept., Cambridge University, Cambridge, U.K.
链接:https://ieeexplore.ieee.org/document/10096993
摘要:在说话人二值化中,说话人嵌入提取模型经常会遇到训练损失函数与说话人聚类方法不匹配的问题。在本文中,我们提出了频谱聚类感知嵌入学习(SCALE)的方法来解决不匹配问题。具体来说,除了角度原型(AP)损失之外,SCALE还使用了一种新颖的亲和力矩阵损失,它直接最小化了从说话者嵌入估计的亲和力矩阵与参考之间的误差。SCALE 还包括 p 百分位数阈值和高斯模糊,作为训练中谱聚类的两个重要超参数。AMI 数据集上的实验表明,与使用基于 AP 损失的强基线相比,通过 SCALE 获得的说话人嵌入使用 Oracle 分割实现了超过 50% 的相对说话人错误率降低,使用自动分割实现了超过 30% 的相对二值化错误率降低 说话人嵌入。
In speaker diarisation, speaker embedding extraction models often suffer from the mismatch between their training loss functions and the speaker clustering method. In this paper, we propose the method of spectral clustering-aware learning of embeddings (SCALE) to address the mismatch. Specifically, besides an angular prototypical (AP) loss, SCALE uses a novel affinity matrix loss which directly minimises the error between the affinity matrix estimated from speaker embeddings and the reference. SCALE also includes p-percentile thresholding and Gaussian blur as two important hyper-parameters for spectral clustering in training. Experiments on the AMI dataset showed that speaker embeddings obtained with SCALE achieved over 50% relative speaker error rate reductions using oracle segmentation, and over 30% relative diarisation error rate reductions using automatic segmentation when compared to a strong baseline with the AP-loss-based speaker embeddings.
13.Supervised Hierarchical Clustering Using Graph Neural Networks For Speaker Diarization
标题:使用图神经网络进行说话人二值化的有监督层次聚类
作者:Prachi Singh;Amrit Kaul;Sriram Ganapathy
单位:LEAP Lab, Electrical Engineering, Indian Institute of Science, Bangalore
链接:https://ieeexplore.ieee.org/document/10095372
摘要:说话人二值化的传统方法涉及将音频文件窗口化为短片段以提取说话人嵌入,然后对嵌入进行无监督聚类。这种多步骤方法为每个片段生成发言人分配。在本文中,我们提出了一种新颖的用于说话人二值化的监督分层图聚类算法(SHARC),其中我们引入了使用图神经网络(GNN)来执行监督聚类的分层结构。监督允许模型更新表示并直接提高聚类性能,从而实现单步二值化方法。在所提出的工作中,输入段嵌入被视为图的节点,其边权重对应于节点之间的相似性得分。我们还提出了一种联合更新嵌入提取器和 GNN 模型以执行端到端说话人二值化(E2E-SHARC)的方法。在推理过程中,使用节点密度和边缘存在概率来执行层次聚类,以合并片段直至收敛。在二值化实验中,我们表明所提出的 E2E-SHARC 方法在 AMI 和 Voxconverse 等基准数据集上比基准系统分别实现了 53% 和 44% 的相对改进。
Conventional methods for speaker diarization involve windowing an audio file into short segments to extract speaker embeddings, followed by an unsupervised clustering of the embeddings. This multistep approach generates speaker assignments for each segment. In this paper, we propose a novel Supervised HierArchical gRaph Clustering algorithm (SHARC) for speaker diarization where we introduce a hierarchical structure using Graph Neural Network (GNN) to perform supervised clustering. The supervision allows the model to update the representations and directly improve the clustering performance, thus enabling a single-step approach for diarization. In the proposed work, the input segment embeddings are treated as nodes of a graph with the edge weights corresponding to the similarity scores between the nodes. We also propose an approach to jointly update the embedding extractor and the GNN model to perform end-to-end speaker diarization (E2E-SHARC). During inference, the hierarchical clustering is performed using node densities and edge existence probabilities to merge the segments until convergence. In the diarization experiments, we illustrate that the proposed E2E-SHARC approach achieves 53% and 44% relative improvements over the baseline systems on benchmark datasets like AMI and Voxconverse, respectively.
14.Target Speaker Voice Activity Detection with Transformers and Its Integration with End-To-End Neural Diarization
标题:通过序列到序列预测进行目标说话人语音活动检测
作者:Dongmei Wang;Xiong Xiao;Naoyuki Kanda;Takuya Yoshioka;Jian Wu
单位:Microsoft, One Microsoft Way, Redmond, WA, USA
链接:https://ieeexplore.ieee.org/document/10095185
摘要:目标说话人语音活动检测目前是复杂声学环境中说话人二值化的一种很有前途的方法。本文提出了一种新颖的序列到序列目标说话人语音活动检测(Seq2Seq-TSVAD)方法,该方法可以有效解决大规模说话人的联合建模并预测高分辨率语音活动。实验结果表明,更大的说话人容量和更高的输出分辨率可以显着降低二值化错误率(DER),在 VoxConverse 测试集上实现了 4.55% 的新的 state-of-the-art 性能,在 Track 1 上实现了 10.77% 的新的 state-of-the-art 性能。广泛使用的评估指标下的DIHARD-III评估集。
Target-speaker voice activity detection is currently a promising approach for speaker diarization in complex acoustic environments. This paper presents a novel Sequence-to-Sequence Target-Speaker Voice Activity Detection (Seq2Seq-TSVAD) method that can efficiently address the joint modeling of large-scale speakers and predict high-resolution voice activities. Experimental results show that larger speaker capacity and higher output resolution can significantly reduce the diarization error rate (DER), which achieves the new state-of-the-art performance of 4.55% on the VoxConverse test set and 10.77% on Track 1 of the DIHARD-III evaluation set under the widely-used evaluation metrics.
15.The Multimodal Information Based Speech Processing (Misp) 2022 Challenge: Audio-Visual Diarization And Recognition
标题:基于多模态信息的语音处理 (Misp) 2022 挑战:视听二值化和识别
作者:Zhe Wang1;Shilong Wu1;Hang Chen1;Mao-Kui He1;Jun Du1;Chin-Hui Lee2;Jingdong Chen7;Shinji Watanabe2;Sabato Siniscalchi34;Odette Scharenborg5;Diyuan Liu6;Baocai Yin6;Jia Pan6;Jianqing Gao6;Cong Liu6
单位:1 University of Science and Technology of China, China;
2 Carnegie Mellon University, USA;
3 Georgia Institute of Technology, USA;
4 Kore University of Enna, Italy;
5.Delft University of Technology, The Netherlands;
6.iFlytek, China;
7.Northwestern Polytechnical University, China
链接:https://ieeexplore.ieee.org/document/10094836
摘要:基于多模态信息的语音处理(MISP)挑战赛旨在通过推动唤醒词、说话人二值化、语音识别等技术的研究,扩展信号处理技术在特定场景中的应用。MISP2022 挑战赛有两个赛道:1)视听说话人分类(AVSD),旨在使用音频和视频数据解决“谁在何时说话”;2)一种新颖的视听二值化和识别(AVDR)任务,重点是通过视听说话者二值化结果解决“谁在何时说了什么”。这两首曲目都以中文为重点,并在真实的家庭电视场景中使用远场音频和视频:2-6 人在背景中的电视噪音中相互交流。本文介绍了 MISP2022 挑战赛的数据集、赛道设置和基线。我们对实验和示例的分析表明 AVDR 基线系统具有良好的性能,以及由于远场视频质量、背景中电视噪声的存在以及无法区分的说话人等原因而导致的挑战中的潜在困难。
The Multi-modal Information based Speech Processing (MISP) challenge aims to extend the application of signal processing technology in specific scenarios by promoting the research into wake-up words, speaker diarization, speech recognition, and other technologies. The MISP2022 challenge has two tracks: 1) audio-visual speaker diarization (AVSD), aiming to solve "who spoken when" using both audio and visual data; 2) a novel audio-visual diarization and recognition (AVDR) task that focuses on addressing "who spoken what when" with audio-visual speaker diarization results. Both tracks focus on the Chinese language, and use far-field audio and video in real home-tv scenarios: 2-6 people communicating each other with TV noise in the background. This paper introduces the dataset, track settings, and baselines of the MISP2022 challenge. Our analyses of experiments and examples indicate the good performance of AVDR baseline system, and the potential difficulties in this challenge due to, e.g., the far-field video quality, the presence of TV noise in the background, and the indistinguishable speakers.
16.The WHU-Alibaba Audio-Visual Speaker Diarization System For The MISP 2022 Challenge
标题:武汉大学-阿里巴巴针对MISP 2022挑战赛的视听说话人分类系统
作者:Ming Cheng12;Haoxu Wang12;Ziteng Wang3;Qiang Fu3;Ming Li12
单位:1 School of Computer Science, Wuhan University, Wuhan, China;
2 Data Science Research Center, Duke Kunshan University, Kunshan, China;
3 Alibaba Group, China
链接:https://ieeexplore.ieee.org/document/10095802
摘要:本文介绍了武汉大学-阿里巴巴团队为基于多模态信息的语音处理 (MISP) 2022 挑战赛开发的系统。我们扩展了序列到序列目标说话者语音活动检测框架,以同时从视听信号中检测多个说话者的语音活动。最终系统在竞赛数据库评估集上实现了8.82%的二值化错误率(DER),在MISP 2022 ICASSP信号处理大赛的说话人二值化赛道中排名第一。
This paper describes the system developed by the WHU-Alibaba team for the Multimodal Information Based Speech Processing (MISP) 2022 Challenge. We extend the Sequence-to-Sequence Target-Speaker Voice Activity Detection framework to simultaneously detect multiple speakers’ voice activities from audio-visual signals. The final system achieves a diarization error rate (DER) of 8.82% on the evaluation set of the competition database, which ranks 1st in the speaker diarization track of the MISP 2022, ICASSP Signal Processing Grand Challenge.
17.TOLD: A Novel Two-Stage Overlap-Aware Framework For Speaker Diarization
标题:TOLD:一种新颖的两阶段重叠感知框架,用于说话人二值化
作者:Jiaming Wang;Zhihao Du;Shiliang Zhang
单位:Alibaba Group, Speech Lab of DAMO Academy, China
链接:https://ieeexplore.ieee.org/document/10096436
摘要:最近,端到端神经二值化(EEND)被引入,并在说话者重叠的场景中取得了可喜的结果。在 EEND 中,说话人二值化被表述为多标签预测问题,其中说话人活动是独立估计的,并且没有很好地考虑它们的依赖性。为了克服这些缺点,我们采用幂集编码将说话人二值化重新表述为单标签分类问题,并提出重叠感知 EEND (EEND-OLA) 模型,其中可以明确地对说话人重叠和依赖关系进行建模。受两阶段混合系统成功的启发,我们进一步提出了一种新颖的两阶段 OverLap 感知二值化框架(TOLD),通过涉及说话者重叠感知后处理(SOAP)模型来迭代地细化 EEND 的二值化结果 奥拉。实验结果表明,与原始EEND相比,所提出的EEND-OLA在分类错误率(DER)方面实现了14.39%的相对改进,并且利用SOAP又提供了19.33%的相对改进。结果,我们的方法 TOLD 在 CALLHOME 数据集上实现了 10.14% 的 DER,据我们所知,这是该基准测试的最新结果。
Recently, end-to-end neural diarization (EEND) is introduced and achieves promising results in speaker-overlapped scenarios. In EEND, speaker diarization is formulated as a multi-label prediction problem, where speaker activities are estimated independently and their dependency are not well considered. To overcome these disadvantages, we employ the power set encoding to reformulate speaker diarization as a single-label classification problem and propose the overlap-aware EEND (EEND-OLA) model, in which speaker overlaps and dependency can be modeled explicitly. Inspired by the success of two-stage hybrid systems, we further propose a novel Two-stage OverLap-aware Diarization framework (TOLD) by involving a speaker overlap-aware post-processing (SOAP) model to iteratively refine the diarization results of EEND-OLA. Experimental results show that, compared with the original EEND, the proposed EEND-OLA achieves a 14.39% relative improvement in terms of diarization error rates (DER), and utilizing SOAP provides another 19.33% relative improvement. As a result, our method TOLD achieves a DER of 10.14% on the CALLHOME dataset, which is a new state-of-the-art result on this benchmark to the best of our knowledge.
