本文经arXiv每日学术速递授权转载
【1】 PDAF: A Phonetic Debiasing Attention Framework For Speaker Verification
标题: PDAF:用于说话者验证的语音去偏注意框架
作者:Massa Baali,Abdulhamid Aldoobi,Hira Dhamyal,Rita Singh,Bhiksha Raj
备注:Accepted to SLT
链接:点击下载PDF文件
【2】 Vector Quantized Diffusion Model Based Speech Bandwidth Extension
标题: 基于量化扩散模型的语音带宽扩展
作者:Yuan Fang,Jiajie Wang,Xueliang Zhang
备注:4pages
链接:点击下载PDF文件
【3】 Evaluation of real-time transcriptions using end-to-end ASR models
标题: 使用端到端ASB模型评估实时传输
作者:Carlos Arriaga,Alejandro Pozo,Javier Conde,Alvaro Alonso
备注:15 pages, 4 figures
链接:点击下载PDF文件
【4】 Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
标题: 视听演讲者日记化:当前的数据库、方法和挑战
作者:Victoria Mingote,Alfonso Ortega,Antonio Miguel,Eduardo Lleida
链接:点击下载PDF文件
【5】 Harmonic Reasoning in Large Language Models
标题: 大型语言模型中的和谐推理
作者:Anna Kruspe
链接:点击下载PDF文件
【6】 IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
标题: IndicVoices-R:解锁大规模多语言多说话者语音库,用于扩展印度TTC
作者:Ashwin Sankar,Srija Anand,Praveen Srinivasa Varadhan,Sherry Thomas,Mehak Singal,Shridhar Kumar,Deovrat Mehendale,Aditi Krishana,Giri Raju,Mitesh Khapra
链接:点击下载PDF文件
【7】 Machine Anomalous Sound Detection Using Spectral-temporal Modulation Representations Derived from Machine-specific Filterbanks
标题: 使用从机器特定过滤器组获得的谱-时间调制表示进行机器异常声音检测
作者:Kai Li,Khalid Zaman,Xingfeng Li,Masato Akagi,Masashi Unoki
链接:点击下载PDF文件
【8】 Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
标题: 更好的野外西班牙情感识别:关注深频谱语音分析
作者:Elena Ortega-Beltrán,Josep Cabacas-Maso,Ismael Benito-Altamirano,Carles Ventura
链接:点击下载PDF文件
【9】 The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
标题: 第一个Cadenza挑战:利用机器学习竞赛为听力丧失的听众改善音乐
作者:Gerardo Roa Dabike,Michael A. Akeroyd,Scott Bannister,Jon P. Barker,Trevor J. Cox,Bruno Fazenda,Jennifer Firth,Simone Graetzer,Alinka Greasley,Rebecca R. Vos,William M. Whitmer
链接:点击下载PDF文件
【10】 From Computation to Consumption: Exploring the Compute-Energy Link for Training and Testing Neural Networks for SED Systems
标题: 从计算到消费:探索计算-能量链接以训练和测试MED系统的神经网络
作者:Constance Douwes,Romain Serizel
链接:点击下载PDF文件
【11】 Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
标题: 用于域广义异常声音检测的深度通用表示
作者:Phurich Saengthong,Takahiro Shinozaki
链接:点击下载PDF文件
【12】 Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
标题: 利用声学适应和视觉对齐来改善多模式情绪识别
作者:Zhixian zhao,Haifeng Chen,Xi Li,Dongmei Jiang,Lei Xie
链接:点击下载PDF文件
【13】 Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
标题: 多模式情绪分析的音频引导融合技术
作者:Pujin Shi,Fei Gao
链接:点击下载PDF文件
【14】 Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
标题: 用预训练模型解开韵律和语义信息,用于基于上下文学习的Zero-Shot语音转换
作者:Zhengyang Chen,Shuai Wang,Mingyang Zhang,Xuechen Liu,Junichi Yamagishi,Yanmin Qian
链接:点击下载PDF文件
【15】 Evaluating Neural Networks Architectures for Spring Reverb Modelling
标题: 评估春季回响建模的神经网络架构
作者:Francesco Papaleo,Xavier Lizarraga-Seijas,Frederic Font
备注:8 pages, 7 figures, 2 tables
链接:点击下载PDF文件
【16】 Attention-Based Efficient Breath Sound Removal in Studio Audio Recordings
标题: 录音室录音中基于注意力的高效呼吸声去除
作者:Nidula Elgiriyewithana,N. D. Kodikara
Journal-ref:CS & IT Conference Proceedings, vol. 14, no. 6, 2024
链接:点击下载PDF文件
【17】 Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
标题: Flow-TSVAD:通过潜在流匹配进行目标说话者语音活动检测
作者:Zhengyang Chen,Bing Han,Shuai Wang,Yidi Jiang,Yanmin Qian
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【18】 PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
标题: PB-LRDDWWS系统,用于SYS 2024低资源发音障碍唤醒单词发现挑战赛
作者:Shiyao Wang,Jiaming Zhou,Shiwan Zhao,Yong Qin
备注:accept by SLT 2024
链接:点击下载PDF文件
【19】 Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
标题: Mel-Roformer用于人声分离和人声旋律转录
作者:Ju-Chiang Wang,Wei-Tsung Lu,Jitong Chen
备注:Accepted to appear in ISMIR 2024
链接:点击下载PDF文件
【20】 Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples
标题: 利用对比学习和自我训练以有限的标记样本进行多模式情感识别
作者:Qi Fan,Yutong Li,Yi Xin,Xinyu Cheng,Guanglai Gao,Miao Ma
备注:Accepted by ACM MM Workshop 2024
链接:点击下载PDF文件
【21】 Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
标题: 2024年普通话口吃事件检测和自动语音识别挑战赛的结果
作者:Hongfei Xue,Rong Gong,Mingchen Shao,Xin Xu,Lezhi Wang,Lei Xie,Hui Bu,Jiaming Zhou,Yong Qin,Jun Du,Ming Li,Binbin Zhang,Bin Jia
备注:8 pages, 2 figures, accepted by SLT 2024
链接:点击下载PDF文件
【22】 BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
标题: BigCodec:突破低比特率神经语音编解码器的极限
作者:Detai Xin,Xu Tan,Shinnosuke Takamichi,Hiroshi Saruwatari
备注:4 pages, 1 figure. Audio samples available at: this https URL
链接:点击下载PDF文件
【23】 SS-BRPE: Self-Supervised Blind Room Parameter Estimation Using Attention Mechanisms
标题: SS-BRPE:使用注意力机制的自我监督盲点参数估计
作者:Chunxi Wang,Maoshen Jia,Meiran Li,Changchun Bao,Wenyu Jin
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【24】 Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
标题: 使用薛定格桥和对称噪音表的基于扩散的语音增强
作者:Siyi Wang,Siyi Liu,Andrew Harper,Paul Kendrick,Mathieu Salzmann,Milos Cernak
链接:点击下载PDF文件
【25】 TF-Mamba: A Time-Frequency Network for Sound Source Localization
标题: TF-Mamba:一种用于光源定位的时频网络
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【26】 Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
标题: 探索WavLM后台进行语音欺骗和Deepfake检测
作者:Theophile Stourbe,Victor Miara,Theo Lepage,Reda Dehak
链接:点击下载PDF文件
【27】 Leveraging Moving Sound Source Trajectories for Universal Sound Separation
标题: 利用移动光源轨迹实现通用声音分离
作者:Donghang Wu,Xihong Wu,Tianshu Qu
备注:9 pages,7 figures,submitted to IEEEACM Transactions on Audio, Speech and Language Processing(TASLP)
链接:点击下载PDF文件
【28】 Cross-attention Inspired Selective State Space Models for Target Sound Extraction
标题: 交叉注意启发的选择性状态空间模型用于目标声音提取
作者:Donghang Wu,Yiwen Wang,Xihong Wu,Tianshu Qu
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
标题: 联合说话人拨号和识别工具包及其应用于说话人归因的ASB
作者:Giovanni Morrone,Enrico Zovato,Fabio Brugnara,Enrico Sartori,Leonardo Badino
Journal-ref:Proceedings of Interspeech 2024, pp. 3652--3653
链接:点击下载PDF文件
【2】 AS-Speech: Adaptive Style For Speech Synthesis
标题: AS-Speech:语音合成的自适应风格
作者:Zhipeng Li,Xiaofen Xing,Jun Wang,Shuaiqi Chen,Guoqiao Yu,Guanglu Wan,Xiangmin Xu
备注:Accepted by SLT 2024
链接:点击下载PDF文件
【3】 Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
标题: 越长(不一定)越强:用于增强语音识别和翻译的间断长序列训练
作者:Nithin Rao Koluguri,Travis Bartley,Hainan Xu,Oleksii Hrinchuk,Jagadeesh Balam,Boris Ginsburg,Georg Kucsko
备注:Accepted at SLT 2024
链接:点击下载PDF文件
【4】 An investigation of modularity for noise robustness in conformer-based ASR
标题: 基于一致性的ASB中噪音鲁棒性的模块化研究
作者:Louise Coppieters de Gibson,Philip N. Garner,Pierre-Edouard Honnet
备注:5 pages, 3 figures
链接:点击下载PDF文件
【5】 Leveraging Content and Acoustic Representations for Efficient Speech Emotion Recognition
标题: 利用内容和声学表示实现高效的语音情感识别
作者:Soumya Dutta,Sriram Ganapathy
备注:10 pages, 4 figures, 7 tables
链接:点击下载PDF文件
【6】 NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
标题: 用于CHiME-8挑战赛DSVR任务的NTT多扬声器ASB系统
作者:Naoyuki Kamo,Naohiro Tawara,Atsushi Ando,Takatomo Kano,Hiroshi Sato,Rintaro Ikeshita,Takafumi Moriya,Shota Horiguchi,Kohei Matsuura,Atsunori Ogawa,Alexis Plaquet,Takanori Ashihara,Tsubasa Ochiai,Masato Mimura,Marc Delcroix,Tomohiro Nakatani,Taichi Asami,Shoko Araki
备注:5 pages, 4 figures, CHiME8 challenge
链接:点击下载PDF文件
【7】 Transferable Selective Virtual Sensing Active Noise Control Technique Based on Metric Learning
标题: 基于度量学习的可转移选择性虚拟感知主动噪音控制技术
作者:Boxiang Wang,Dongyuan Shi,Zhengding Luo,Xiaoyi Shen,Junwei Ji,Woon-Seng Gan
链接:点击下载PDF文件
【8】 Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
标题: 2024年普通话口吃事件检测和自动语音识别挑战赛的结果
作者:Hongfei Xue,Rong Gong,Mingchen Shao,Xin Xu,Lezhi Wang,Lei Xie,Hui Bu,Jiaming Zhou,Yong Qin,Jun Du,Ming Li,Binbin Zhang,Bin Jia
备注:8 pages, 2 figures, accepted by SLT 2024
链接:点击下载PDF文件
【9】 BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
标题: BigCodec:突破低比特率神经语音编解码器的极限
作者:Detai Xin,Xu Tan,Shinnosuke Takamichi,Hiroshi Saruwatari
备注:4 pages, 1 figure. Audio samples available at: this https URL
链接:点击下载PDF文件
【10】 SS-BRPE: Self-Supervised Blind Room Parameter Estimation Using Attention Mechanisms
标题: SS-BRPE:使用注意力机制的自我监督盲点参数估计
作者:Chunxi Wang,Maoshen Jia,Meiran Li,Changchun Bao,Wenyu Jin
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【11】 Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
标题: 使用薛定格桥和对称噪音表的基于扩散的语音增强
作者:Siyi Wang,Siyi Liu,Andrew Harper,Paul Kendrick,Mathieu Salzmann,Milos Cernak
链接:点击下载PDF文件
【12】 TF-Mamba: A Time-Frequency Network for Sound Source Localization
标题: TF-Mamba:一种用于光源定位的时频网络
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【13】 Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
标题: 探索WavLM后台进行语音欺骗和Deepfake检测
作者:Theophile Stourbe,Victor Miara,Theo Lepage,Reda Dehak
链接:点击下载PDF文件
【14】 Leveraging Moving Sound Source Trajectories for Universal Sound Separation
标题: 利用移动光源轨迹实现通用声音分离
作者:Donghang Wu,Xihong Wu,Tianshu Qu
备注:9 pages,7 figures,submitted to IEEEACM Transactions on Audio, Speech and Language Processing(TASLP)
链接:点击下载PDF文件
【15】 Cross-attention Inspired Selective State Space Models for Target Sound Extraction
标题: 交叉注意启发的选择性状态空间模型用于目标声音提取
作者:Donghang Wu,Yiwen Wang,Xihong Wu,Tianshu Qu
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【16】 Vector Quantized Diffusion Model Based Speech Bandwidth Extension
标题: 基于量化扩散模型的语音带宽扩展
作者:Yuan Fang,Jiajie Wang,Xueliang Zhang
备注:4pages
链接:点击下载PDF文件
【17】 Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
标题: 视听演讲者日记化:当前的数据库、方法和挑战
作者:Victoria Mingote,Alfonso Ortega,Antonio Miguel,Eduardo Lleida
链接:点击下载PDF文件
【18】 Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
标题: 更好的野外西班牙情感识别:关注深频谱语音分析
作者:Elena Ortega-Beltrán,Josep Cabacas-Maso,Ismael Benito-Altamirano,Carles Ventura
链接:点击下载PDF文件
【19】 The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
标题: 第一个Cadenza挑战:利用机器学习竞赛为听力丧失的听众改善音乐
作者:Gerardo Roa Dabike,Michael A. Akeroyd,Scott Bannister,Jon P. Barker,Trevor J. Cox,Bruno Fazenda,Jennifer Firth,Simone Graetzer,Alinka Greasley,Rebecca R. Vos,William M. Whitmer
链接:点击下载PDF文件
【20】 Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
标题: 用于域广义异常声音检测的深度通用表示
作者:Phurich Saengthong,Takahiro Shinozaki
链接:点击下载PDF文件
【21】 Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
标题: 利用声学适应和视觉对齐来改善多模式情绪识别
作者:Zhixian zhao,Haifeng Chen,Xi Li,Dongmei Jiang,Lei Xie
链接:点击下载PDF文件
【22】 Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
标题: 多模式情绪分析的音频引导融合技术
作者:Pujin Shi,Fei Gao
链接:点击下载PDF文件
【23】 Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
标题: 用预训练模型解开韵律和语义信息,用于基于上下文学习的Zero-Shot语音转换
作者:Zhengyang Chen,Shuai Wang,Mingyang Zhang,Xuechen Liu,Junichi Yamagishi,Yanmin Qian
链接:点击下载PDF文件
【24】 Attention-Based Efficient Breath Sound Removal in Studio Audio Recordings
标题: 录音室录音中基于注意力的高效呼吸声去除
作者:Nidula Elgiriyewithana,N. D. Kodikara
Journal-ref:CS & IT Conference Proceedings, vol. 14, no. 6, 2024
链接:点击下载PDF文件
【25】 Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
标题: 只是ASB + LLM吗?言语大语言模型识别和理解口语对话中说话人的能力研究
作者:Junkai Wu,Xulin Fan,Bo-Ru Lu,Xilin Jiang,Nima Mesgarani,Mark Hasegawa-Johnson,Mari Ostendorf
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【26】 Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
标题: Flow-TSVAD:通过潜在流匹配进行目标说话者语音活动检测
作者:Zhengyang Chen,Bing Han,Shuai Wang,Yidi Jiang,Yanmin Qian
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【27】 PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
标题: PB-LRDDWWS系统,用于SYS 2024低资源发音障碍唤醒单词发现挑战赛
作者:Shiyao Wang,Jiaming Zhou,Shiwan Zhao,Yong Qin
备注:accept by SLT 2024
链接:点击下载PDF文件
【28】 Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
标题: Mel-Roformer用于人声分离和人声旋律转录
作者:Ju-Chiang Wang,Wei-Tsung Lu,Jitong Chen
备注:Accepted to appear in ISMIR 2024
链接:点击下载PDF文件
【29】 Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples
标题: 利用对比学习和自我训练以有限的标记样本进行多模式情感识别
作者:Qi Fan,Yutong Li,Yi Xin,Xinyu Cheng,Guanglai Gao,Miao Ma
备注:Accepted by ACM MM Workshop 2024
链接:点击下载PDF文件
标题: PDAF:用于说话者验证的语音去偏注意框架
作者:Massa Baali,Abdulhamid Aldoobi,Hira Dhamyal,Rita Singh,Bhiksha Raj
备注:Accepted to SLT
链接:点击下载PDF文件
摘要:说话人验证系统对于通过语音进行身份验证至关重要。传统上,这些系统专注于比较特征向量,忽略了语音的内容。然而,本文提出了挑战,强调语音优势的重要性,衡量音素的频率或持续时间,作为说话人验证的一个关键线索。提出了一种新的音素去偏注意框架(PDAF),该框架与现有的注意框架相结合,以消除语音优势引起的偏误。PDAF调整每个音素的权重并影响特征提取,从而允许对语音进行更细致的分析。这种方法为通过语音进行更准确和可靠的身份认证铺平了道路。此外,通过采用不同的加权策略,我们评估的语音特征的说话人确认系统的功效的影响。摘要:Speaker verification systems are crucial for authenticating identity through voice. Traditionally, these systems focus on comparing feature vectors, overlooking the speech's content. However, this paper challenges this by highlighting the importance of phonetic dominance, a measure of the frequency or duration of phonemes, as a crucial cue in speaker verification. A novel Phoneme Debiasing Attention Framework (PDAF) is introduced, integrating with existing attention frameworks to mitigate biases caused by phonetic dominance. PDAF adjusts the weighting for each phoneme and influences feature extraction, allowing for a more nuanced analysis of speech. This approach paves the way for more accurate and reliable identity authentication through voice. Furthermore, by employing various weighting strategies, we evaluate the influence of phonetic features on the efficacy of the speaker verification system.
【2】 Vector Quantized Diffusion Model Based Speech Bandwidth Extension
标题: 基于量化扩散模型的语音带宽扩展
作者:Yuan Fang,Jiajie Wang,Xueliang Zhang
备注:4pages
链接:点击下载PDF文件
摘要:神经音频编解码器(NAC)的最新进展开启了音频信号处理的新潜力。越来越多的研究探索利用NAC的潜在功能进行各种语音信号处理任务。本文介绍了第一种方法,语音带宽扩展(BWE),利用从NAC获得的离散功能。通过在高度压缩的离散令牌中恢复高频细节,这种方法增强了语音的可懂度和自然度。基于矢量量化扩散,该框架结合了先进的NAC,扩散模型和Mamba-2的优势,重建高频语音成分。大量的实验表明,该方法在对数谱距离和ViSQOL方面表现出优异的性能,显着提高了语音质量。摘要:Recent advancements in neural audio codec (NAC) unlock new potential in audio signal processing. Studies have increasingly explored leveraging the latent features of NAC for various speech signal processing tasks. This paper introduces the first approach to speech bandwidth extension (BWE) that utilizes the discrete features obtained from NAC. By restoring high-frequency details within highly compressed discrete tokens, this approach enhances speech intelligibility and naturalness. Based on Vector Quantized Diffusion, the proposed framework combines the strengths of advanced NAC, diffusion models, and Mamba-2 to reconstruct high-frequency speech components. Extensive experiments demonstrate that this method exhibits superior performance across both log-spectral distance and ViSQOL, significantly improving speech quality.
【3】 Evaluation of real-time transcriptions using end-to-end ASR models
标题: 使用端到端ASB模型评估实时传输
作者:Carlos Arriaga,Alejandro Pozo,Javier Conde,Alvaro Alonso
备注:15 pages, 4 figures
链接:点击下载PDF文件
摘要:自动语音识别(ASR)或语音到文本(STT)在过去几年中有了很大的发展。基于管道的传统架构已被联合端到端(E2 E)架构所取代,该架构简化了模型训练过程。此外,新的人工智能训练方法,如弱监督学习,减少了对用于模型训练的高质量音频数据集的需求。然而,尽管有这些进步,很少或没有研究已经做了实时转录。在实时场景中,音频不是预先录制的,并且输入音频必须被分段以由ASR系统处理。为了实现实时要求,这些片段必须尽可能短以减少延迟。然而,音频不能在任何点被分割,因为将话语分割成两个单独的片段将生成不正确的转录。此外,较短的片段为ASR模型提供较少的上下文。出于这个原因,有必要设计和测试不同的分割算法,以优化最终转录的质量和延迟。在本文中,三个音频分割算法进行了评估与不同的ASR模型,以确定其对转录的质量和端到端的延迟的影响。这些算法是固定间隔分段、语音活动检测(VAD)和反馈分段。将结果与没有音频碎片的相同模型的性能进行比较,以确定这种划分的效果。结果表明,VAD分段提供了最好的质量和最高的延迟,而以固定的时间间隔分段提供了最低的质量和最低的延迟。新提出的反馈算法交换了2-4%的WER增加减少1.5-2s的延迟,分别到VAD分裂。摘要:Automatic Speech Recognition (ASR) or Speech-to-text (STT) has greatly evolved in the last few years. Traditional architectures based on pipelines have been replaced by joint end-to-end (E2E) architectures that simplify and streamline the model training process. In addition, new AI training methods, such as weak-supervised learning have reduced the need for high-quality audio datasets for model training. However, despite all these advancements, little to no research has been done on real-time transcription. In real-time scenarios, the audio is not pre-recorded, and the input audio must be fragmented to be processed by the ASR systems. To achieve real-time requirements, these fragments must be as short as possible to reduce latency. However, audio cannot be split at any point as dividing an utterance into two separate fragments will generate an incorrect transcription. Also, shorter fragments provide less context for the ASR model. For this reason, it is necessary to design and test different splitting algorithms to optimize the quality and delay of the resulting transcription. In this paper, three audio splitting algorithms are evaluated with different ASR models to determine their impact on both the quality of the transcription and the end-to-end delay. The algorithms are fragmentation at fixed intervals, voice activity detection (VAD), and fragmentation with feedback. The results are compared to the performance of the same model, without audio fragmentation, to determine the effects of this division. The results show that VAD fragmentation provides the best quality with the highest delay, whereas fragmentation at fixed intervals provides the lowest quality and the lowest delay. The newly proposed feedback algorithm exchanges a 2-4% increase in WER for a reduction of 1.5-2s delay, respectively, to the VAD splitting.
【4】 Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
标题: 视听演讲者日记化:当前的数据库、方法和挑战
作者:Victoria Mingote,Alfonso Ortega,Antonio Miguel,Eduardo Lleida
链接:点击下载PDF文件
摘要:如今,大量的视听内容已经促进了开发新的鲁棒的自动说话人日志系统来分析和验证它的需求。这种系统有助于降低手动执行此过程的成本,并且允许将说话人信息用于不同的应用,因为存在大量的信息,例如,面部图像或音频记录。因此,本文旨在解决说话人日志化系统领域的一个关键领域,即不同领域视听内容的整合。本文旨在通过开发一个强大的视听演讲者日记框架来超越当前最先进的实践,该框架适用于各种数据域,包括电视场景,会议和日常活动。与大多数现有的视听扬声器日记系统不同,该框架还将包括一种方法的建议,以引导在名人出现的电视场景中精确分配特定身份。此外,在这项工作中,我们已经进行了广泛的汇编,目前的国家的最先进的方法和现有的数据库开发视听扬声器日记。摘要:Nowadays, the large amount of audio-visual content available has fostered the need to develop new robust automatic speaker diarization systems to analyse and characterise it. This kind of system helps to reduce the cost of doing this process manually and allows the use of the speaker information for different applications, as a huge quantity of information is present, for example, images of faces, or audio recordings. Therefore, this paper aims to address a critical area in the field of speaker diarization systems, the integration of audio-visual content of different domains. This paper seeks to push beyond current state-of-the-art practices by developing a robust audio-visual speaker diarization framework adaptable to various data domains, including TV scenarios, meetings, and daily activities. Unlike most of the existing audio-visual speaker diarization systems, this framework will also include the proposal of an approach to lead the precise assignment of specific identities in TV scenarios where celebrities appear. In addition, in this work, we have conducted an extensive compilation of the current state-of-the-art approaches and the existing databases for developing audio-visual speaker diarization.
【5】 Harmonic Reasoning in Large Language Models
标题: 大型语言模型中的和谐推理
作者:Anna Kruspe
链接:点击下载PDF文件
摘要:大型语言模型(LLM)正变得非常流行,并用于许多不同的目的,包括艺术中的创造性任务。然而,这些模型有时会在特定的推理任务中遇到麻烦,特别是那些涉及逻辑思维和计数的任务。本文着眼于如何以及法学硕士理解和处理音乐任务时,如找出音符从间隔和识别和弦和音阶的原因。我们测试了GPT-3.5和GPT-4 o,看看它们如何处理这些任务。我们的研究结果表明,虽然LLM在音符间隔方面做得很好,但它们在识别和弦和音阶等更复杂的任务中表现不佳。这指出了当前LLM能力的明确限制,并显示了我们需要使其更好的地方,这可能有助于改善他们在艺术和其他复杂领域的思维和工作方式。我们还为所描述的任务提供了自动生成的基准数据集。摘要:Large Language Models (LLMs) are becoming very popular and are used for many different purposes, including creative tasks in the arts. However, these models sometimes have trouble with specific reasoning tasks, especially those that involve logical thinking and counting. This paper looks at how well LLMs understand and reason when dealing with musical tasks like figuring out notes from intervals and identifying chords and scales. We tested GPT-3.5 and GPT-4o to see how they handle these tasks. Our results show that while LLMs do well with note intervals, they struggle with more complicated tasks like recognizing chords and scales. This points out clear limits in current LLM abilities and shows where we need to make them better, which could help improve how they think and work in both artistic and other complex areas. We also provide an automatically generated benchmark data set for the described tasks.
【6】 IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
标题: IndicVoices-R:解锁大规模多语言多说话者语音库,用于扩展印度TTC
作者:Ashwin Sankar,Srija Anand,Praveen Srinivasa Varadhan,Sherry Thomas,Mehak Singal,Shridhar Kumar,Deovrat Mehendale,Aditi Krishana,Giri Raju,Mitesh Khapra
链接:点击下载PDF文件
摘要:文本到语音(TTS)合成的最新进展表明,使用大量Web数据训练的大规模模型可以产生高度自然的输出。然而,由于LibriVox或YouTube等平台上缺乏高质量的手动字幕数据,印度语言的此类数据很少。为了解决这一差距,我们增强了现有的大规模ASR数据集,其中包含在低质量环境中收集的自然对话,以生成高质量的TTS训练数据。我们的管道利用了在英语上训练并应用于印度语言的去噪和语音增强模型的跨语言泛化。这导致了IndicVoices-R(IV-R),这是来自ASR数据集的最大的多语言印度TTS数据集,拥有来自22种印度语言的10,496名发言者的1,704小时的高质量语音。IV-R与LJSpeech、LibriTTS和IndicTTS等黄金标准TTS数据集的质量相匹配。我们还介绍了IV-R基准,第一个评估印度语音TTS模型的zero-shot,Few-Shot和man-shot扬声器泛化能力,确保年龄,性别和风格的多样性。我们证明,与单独对IndicTTS数据集进行微调相比,在高质量IndicTTS和IV-R数据集的组合数据集上对英语预训练模型进行微调,可以获得更好的zero-shot扬声器泛化。此外,我们的评估显示,在先前数据集上训练的TTS模型中,印度语音的zero-shot泛化有限,我们通过对包含不同语系扬声器的数据进行微调来改进模型。我们开源了所有数据和代码,为所有22种印度官方语言发布了第一个TTS模型。摘要:Recent advancements in text-to-speech (TTS) synthesis show that large-scale models trained with extensive web data produce highly natural-sounding output. However, such data is scarce for Indian languages due to the lack of high-quality, manually subtitled data on platforms like LibriVox or YouTube. To address this gap, we enhance existing large-scale ASR datasets containing natural conversations collected in low-quality environments to generate high-quality TTS training data. Our pipeline leverages the cross-lingual generalization of denoising and speech enhancement models trained on English and applied to Indian languages. This results in IndicVoices-R (IV-R), the largest multilingual Indian TTS dataset derived from an ASR dataset, with 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. IV-R matches the quality of gold-standard TTS datasets like LJSpeech, LibriTTS, and IndicTTS. We also introduce the IV-R Benchmark, the first to assess zero-shot, few-shot, and many-shot speaker generalization capabilities of TTS models on Indian voices, ensuring diversity in age, gender, and style. We demonstrate that fine-tuning an English pre-trained model on a combined dataset of high-quality IndicTTS and our IV-R dataset results in better zero-shot speaker generalization compared to fine-tuning on the IndicTTS dataset alone. Further, our evaluation reveals limited zero-shot generalization for Indian voices in TTS models trained on prior datasets, which we improve by fine-tuning the model on our data containing diverse set of speakers across language families. We open-source all data and code, releasing the first TTS model for all 22 official Indian languages.
【7】 Machine Anomalous Sound Detection Using Spectral-temporal Modulation Representations Derived from Machine-specific Filterbanks
标题: 使用从机器特定过滤器组获得的谱-时间调制表示进行机器异常声音检测
作者:Kai Li,Khalid Zaman,Xingfeng Li,Masato Akagi,Masashi Unoki
链接:点击下载PDF文件
摘要:工厂机械故障的早期检测在工业应用中至关重要。在机器异常声音检测(ASD)中,不同的机器根据其物理特性表现出独特的振动频率范围。同时,人类听觉系统擅长跟踪机器声音的时间和频谱动态。因此,将人类听觉系统的计算听觉模型与机器特定的属性相结合可以是机器ASD的有效方法。我们首先量化的频率重要性的四种类型的机器使用Fisher比率(F比)。量化的频率重要性,然后被用来设计特定于机器的非均匀滤波器组(NUFB),提取对数非均匀谱(LNS)的功能。所设计的NUFB在相对高F比的频率区域具有较窄的带宽和较高的滤波器分布密度。最后,提出了基于LNS特征的频谱和时间调制表示方法。这些提出的LNS特征和调制表示被输入到ASD的基于自动编码器神经网络的检测器中。来自具有6 dB信噪比(SNR)的故障工业机器调查和检测数据集的训练集的量化结果表明,不同机器的正常和异常声音之间的区分信息在频域中被非均匀地编码。通过使用NUFB突出显示这些重要的频率区域,LNS功能可以在各种SNR条件下使用AUC(接收器工作特征曲线下的面积)的度量来显著增强性能。此外,调制表示可以进一步提高性能。具体而言,时间调制对风扇、泵和滑块有效,而光谱调制对阀门特别有效。摘要:Early detection of factory machinery malfunctions is crucial in industrial applications. In machine anomalous sound detection (ASD), different machines exhibit unique vibration-frequency ranges based on their physical properties. Meanwhile, the human auditory system is adept at tracking both temporal and spectral dynamics of machine sounds. Consequently, integrating the computational auditory models of the human auditory system with machine-specific properties can be an effective approach to machine ASD. We first quantified the frequency importances of four types of machines using the Fisher ratio (F-ratio). The quantified frequency importances were then used to design machine-specific non-uniform filterbanks (NUFBs), which extract the log non-uniform spectrum (LNS) feature. The designed NUFBs have a narrower bandwidth and higher filter distribution density in frequency regions with relatively high F-ratios. Finally, spectral and temporal modulation representations derived from the LNS feature were proposed. These proposed LNS feature and modulation representations are input into an autoencoder neural-network-based detector for ASD. The quantification results from the training set of the Malfunctioning Industrial Machine Investigation and Inspection dataset with a signal-to-noise (SNR) of 6 dB reveal that the distinguishing information between normal and anomalous sounds of different machines is encoded non-uniformly in the frequency domain. By highlighting these important frequency regions using NUFBs, the LNS feature can significantly enhance performance using the metric of AUC (area under the receiver operating characteristic curve) under various SNR conditions. Furthermore, modulation representations can further improve performance. Specifically, temporal modulation is effective for fans, pumps, and sliders, while spectral modulation is particularly effective for valves.
【8】 Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
标题: 更好的野外西班牙情感识别:关注深频谱语音分析
作者:Elena Ortega-Beltrán,Josep Cabacas-Maso,Ismael Benito-Altamirano,Carles Ventura
链接:点击下载PDF文件
摘要:在创造新的社会辅助机器人的背景下,情感识别已经成为一个关键的发展因素,因为它允许机器人适应用户在野外的情绪状态。在这项工作中,我们重点分析了两个语音记录西班牙语数据集:ELRA-S 0329和ESPRITHMatchSpanishDB。具体地说,我们的工作集中在语言,e。G.伴随信息并阐明含义的声音特征。我们提出了使用DeepSpectrum方法,该方法包括提取音轨的视觉表示并将其馈送到预训练的CNN模型。对于分类任务,DeepSpectrum通常与支持向量分类器(DS-SVC)或全连接深度学习分类器(DS-FC)配对。我们将DS-SVC和DS-FC架构的结果与ELRA-S 0329和TMS 320 MatchSpanishDB的最新技术(SOTA)进行了比较。此外,我们提出了我们自己的分类器的基础上的注意机制,即DS-AM。我们针对这两个数据集训练了所有模型,我们发现我们的DS-AM模型在数据集和SOTA DeepSpectrum架构上优于SOTA模型。最后,我们在一个数据集中训练了我们的DS-AM模型,并在另一个数据集中对其进行了测试,以模拟真实世界条件下模型对数据集的偏差。摘要:Within the context of creating new Socially Assistive Robots, emotion recognition has become a key development factor, as it allows the robot to adapt to the user's emotional state in the wild. In this work, we focused on the analysis of two voice recording Spanish datasets: ELRA-S0329 and EmoMatchSpanishDB. Specifically, we centered our work in the paralanguage, e.~g. the vocal characteristics that go along with the message and clarifies the meaning. We proposed the use of the DeepSpectrum method, which consists of extracting a visual representation of the audio tracks and feeding them to a pretrained CNN model. For the classification task, DeepSpectrum is often paired with a Support Vector Classifier --DS-SVC--, or a Fully-Connected deep-learning classifier --DS-FC--. We compared the results of the DS-SVC and DS-FC architectures with the state-of-the-art (SOTA) for ELRA-S0329 and EmoMatchSpanishDB. Moreover, we proposed our own classifier based upon Attention Mechanisms, namely DS-AM. We trained all models against both datasets, and we found that our DS-AM model outperforms the SOTA models for the datasets and the SOTA DeepSpectrum architectures. Finally, we trained our DS-AM model in one dataset and tested it in the other, to simulate real-world conditions on how biased is the model to the dataset.
【9】 The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
标题: 第一个Cadenza挑战:利用机器学习竞赛为听力丧失的听众改善音乐
作者:Gerardo Roa Dabike,Michael A. Akeroyd,Scott Bannister,Jon P. Barker,Trevor J. Cox,Bruno Fazenda,Jennifer Firth,Simone Graetzer,Alinka Greasley,Rebecca R. Vos,William M. Whitmer
链接:点击下载PDF文件
摘要:众所周知,听音乐是听力损失患者的一个问题,助听器并不是一个通用的解决方案。如何使用机器学习来解决这个问题?本文详细介绍了开放挑战方法的首次应用,即使用机器学习来改善听力损失患者的音乐音频质量。第一个挑战是一个独立的竞争(CAD 1),有9个参赛者。第二个是2024年ICASSP大挑战赛(ICASSP 24),吸引了17名参赛者。挑战任务涉及分离和混音流行 摇滚音乐,以允许在混音中个性化地重新平衡乐器,以及放大以校正提高的听力阈值。为参赛者提供的软件基线使用了两种最先进的demix算法:混合Demucs和开放Unmix。系统评估使用客观指标HAAQI(助听器音频质量指数)进行。没有参赛者在CAD 1的最佳基线上有所改善,因为没有足够的改进空间。因此,对于ICASSP 24,通过使用扬声器再现和在重新混合之前应用的指定增益,使场景变得更加困难。这也使得该场景对于通过助听器收听更有用。9名参赛者得分高于ICASSP 24最佳基线。大多数参赛者使用了改进版的混合Demucs和NAL-R放大。最高评分系统结合了几个解混算法的输出,在一个集成的方法。这些挑战现在是未来研究的开放基准,软件和数据可以免费获得。摘要:It is well established that listening to music is an issue for those with hearing loss, and hearing aids are not a universal solution. How can machine learning be used to address this? This paper details the first application of the open challenge methodology to use machine learning to improve audio quality of music for those with hearing loss. The first challenge was a stand-alone competition (CAD1) and had 9 entrants. The second was an 2024 ICASSP grand challenge (ICASSP24) and attracted 17 entrants. The challenge tasks concerned demixing and remixing pop rock music to allow a personalised rebalancing of the instruments in the mix, along with amplification to correct for raised hearing thresholds. The software baselines provided for entrants to build upon used two state-of-the-art demix algorithms: Hybrid Demucs and Open-Unmix. Evaluation of systems was done using the objective metric HAAQI, the Hearing-Aid Audio Quality Index. No entrants improved on the best baseline in CAD1 because there was insufficient room for improvement. Consequently, for ICASSP24 the scenario was made more difficult by using loudspeaker reproduction and specified gains to be applied before remixing. This also made the scenario more useful for listening through hearing aids. 9 entrants scored better than the the best ICASSP24 baseline. Most entrants used a refined version of Hybrid Demucs and NAL-R amplification. The highest scoring system combined the outputs of several demixing algorithms in an ensemble approach. These challenges are now open benchmarks for future research with the software and data being freely available.
【10】 From Computation to Consumption: Exploring the Compute-Energy Link for Training and Testing Neural Networks for SED Systems
标题: 从计算到消费:探索计算-能量链接以训练和测试MED系统的神经网络
作者:Constance Douwes,Romain Serizel
链接:点击下载PDF文件
摘要:机器学习模型的大量使用,特别是神经网络,引起了人们对其环境影响的严重担忧。事实上,在过去的几年里,我们已经看到了与培训和部署这些系统相关的计算成本的爆炸式增长。因此,至关重要的是要了解他们的能源需求,以便更好地将其纳入模型的评估,迄今为止主要侧重于性能。在本文中,我们研究了几个神经网络架构的声音事件检测系统的关键组成部分,使用音频标记任务作为一个例子。我们测量了训练和测试小型到大型架构的能耗,并在能耗、浮点运算次数、参数数量和GPU 内存利用率之间建立了复杂的关系。摘要:The massive use of machine learning models, particularly neural networks, has raised serious concerns about their environmental impact. Indeed, over the last few years we have seen an explosion in the computing costs associated with training and deploying these systems. It is, therefore, crucial to understand their energy requirements in order to better integrate them into the evaluation of models, which has so far focused mainly on performance. In this paper, we study several neural network architectures that are key components of sound event detection systems, using an audio tagging task as an example. We measure the energy consumption for training and testing small to large architectures and establish complex relationships between the energy consumption, the number of floating-point operations, the number of parameters, and the GPU memory utilization.
【11】 Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
标题: 用于域广义异常声音检测的深度通用表示
作者:Phurich Saengthong,Takahiro Shinozaki
链接:点击下载PDF文件
摘要:开发一个可靠的异常声音检测(ASD)系统需要对噪声的鲁棒性,对域偏移的适应性,以及在有限的训练数据下的有效性能。当前的主要方法依赖于每个目标机器类型的大量标记数据来使用离群暴露(OE)技术来训练特征提取器,然而它们在目标域上的性能仍然是次优的。在本文中,我们提出了 textit{GenRep},它利用了一个强大的,大规模的预训练特征提取器与kNN相结合的通用特征表示,用于域广义ASD,而不需要微调。 textit{GenRep}集成了MemMixup,这是一种使用最近的源样本增强目标内存库的简单方法,并结合域归一化技术来解决源域和目标域之间的不平衡。 textit{GenRep}优于基于OE的最佳方法,无需标记数据,在DCASE 2023 T2 Eval集上的官方评分为73.79 %,并在有限的数据场景下证明了稳健性。该代码是开放源代码。摘要:Developing a reliable anomalous sound detection (ASD) system requires robustness to noise, adaptation to domain shifts, and effective performance with limited training data. Current leading methods rely on extensive labeled data for each target machine type to train feature extractors using Outlier-Exposure (OE) techniques, yet their performance on the target domain remains sub-optimal. In this paper, we present textit{GenRep}, which utilizes generic feature representations from a robust, large-scale pre-trained feature extractor combined with kNN for domain-generalized ASD, without the need for fine-tuning. textit{GenRep} incorporates MemMixup, a simple approach for augmenting the target memory bank using nearest source samples, paired with a domain normalization technique to address the imbalance between source and target domains. textit{GenRep} outperforms the best OE-based approach without a need for labeled data with an Official Score of 73.79 % on the DCASE2023T2 Eval set and demonstrates robustness under limited data scenarios. The code is available open-source.
【12】 Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
标题: 利用声学适应和视觉对齐来改善多模式情绪识别
作者:Zhixian zhao,Haifeng Chen,Xi Li,Dongmei Jiang,Lei Xie
链接:点击下载PDF文件
摘要:多模态情感识别(MER)旨在通过整合来自各种模态的信息来自动识别和理解人类的情感状态。然而,带注释的多模态数据的稀缺性极大地阻碍了这一研究领域的发展。本文介绍了我们针对MER 2024的MER-SEMI子挑战的解决方案。首先,为了更好地适应MER任务的声学模态特征,我们通过实验评估了预训练语音模型HuBERT的不同层在情感识别中的贡献。基于这些观察结果,我们对被确定为对情感识别任务最有效的层进行参数高效微调(PEFT),从而以最少数量的可学习参数实现情感识别的最佳适应。其次,利用声学模态的优势,我们提出了一种特征对齐预训练方法。这种方法使用大规模的未标记的数据来训练视觉编码器,从而促进声学特征空间内的视觉特征的语义对齐。最后,使用适应的声学特征,对齐的视觉特征,和词汇特征,我们采用了一个注意力机制的特征融合。在MER 2024-SEMI测试集上,该方法的加权F1得分为88.90%,在所有参与团队中排名第四,验证了我们方法的有效性。摘要:Multimodal Emotion Recognition (MER) aims to automatically identify and understand human emotional states by integrating information from various modalities. However, the scarcity of annotated multimodal data significantly hinders the advancement of this research field. This paper presents our solution for the MER-SEMI sub-challenge of MER 2024. First, to better adapt acoustic modality features for the MER task, we experimentally evaluate the contributions of different layers of the pre-trained speech model HuBERT in emotion recognition. Based on these observations, we perform Parameter-Efficient Fine-Tuning (PEFT) on the layers identified as most effective for emotion recognition tasks, thereby achieving optimal adaptation for emotion recognition with a minimal number of learnable parameters. Second, leveraging the strengths of the acoustic modality, we propose a feature alignment pre-training method. This approach uses large-scale unlabeled data to train a visual encoder, thereby promoting the semantic alignment of visual features within the acoustic feature space. Finally, using the adapted acoustic features, aligned visual features, and lexical features, we employ an attention mechanism for feature fusion. On the MER2024-SEMI test set, the proposed method achieves a weighted F1 score of 88.90%, ranking fourth among all participating teams, validating the effectiveness of our approach.
【13】 Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
标题: 多模式情绪分析的音频引导融合技术
作者:Pujin Shi,Fei Gao
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个解决方案,半监督学习跟踪(MER-SEMI)在MER 2024。首先,为了增强特征提取器在情感分类任务上的性能,我们使用标记数据微调了视频和文本特征提取器,特别是CLIP-vit-large和Baichuan-13 B。这种方法有效地保留了视频中传达的原始情感信息。其次,我们提出了一个音频引导的Transformer(AGT)融合机制,它利用了Hubert-large的鲁棒性,在融合通道间和通道内信息方面表现出优越的有效性。第三,为了提高模型的准确性,我们通过使用高置信度的未标记数据作为伪标签来迭代地应用自监督学习。最后,通过黑盒探测,我们发现训练集和测试集之间的数据分布不平衡。因此,我们采用了基于先验知识的投票机制。结果证明了我们战略的有效性,最终为我们赢得了MER-SEMI赛道的第三名。摘要:In this paper, we propose a solution for the semi-supervised learning track (MER-SEMI) in MER2024. First, in order to enhance the performance of the feature extractor on sentiment classification tasks,we fine-tuned video and text feature extractors, specifically CLIP-vit-large and Baichuan-13B, using labeled data. This approach effectively preserves the original emotional information conveyed in the videos. Second, we propose an Audio-Guided Transformer (AGT) fusion mechanism, which leverages the robustness of Hubert-large, showing superior effectiveness in fusing both inter-channel and intra-channel information. Third, To enhance the accuracy of the model, we iteratively apply self-supervised learning by using high-confidence unlabeled data as pseudo-labels. Finally, through black-box probing, we discovered an imbalanced data distribution between the training and test sets. Therefore, We adopt a prior-knowledge-based voting mechanism. The results demonstrate the effectiveness of our strategy, ultimately earning us third place in the MER-SEMI track.
【14】 Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
标题: 用预训练模型解开韵律和语义信息,用于基于上下文学习的Zero-Shot语音转换
作者:Zhengyang Chen,Shuai Wang,Mingyang Zhang,Xuechen Liu,Junichi Yamagishi,Yanmin Qian
链接:点击下载PDF文件
摘要:语音转换(VC)的目的是修改扬声器的音色,同时保留语音内容。先前的方法已经将自监督的输出标记为语义标记,从而促进语音内容信息的解开。最近,上下文学习(ICL)出现在文本到语音(TTS)系统中,用于通过上下文条件反射有效地对音色等特定特征进行建模。本文提出了一种ICL能力增强的VC系统(ICL-VC),采用基于流匹配生成模型的掩码和重建训练策略。我们在LibriTTS数据集上的实验表明,ICL-VC增加了语义标记,提高了说话人相似性。此外,我们发现k-means是一种通用的标记化方法,适用于各种预训练模型。然而,ICL-VC系统在保持源语音的韵律方面面临挑战。为了缓解这个问题,我们建议将从预训练的情感识别模型中提取的韵律嵌入到我们的系统中。韵律嵌入的集成显着提高了系统的能力,以保持源语音韵律,情感语音数据库上验证。摘要:Voice conversion (VC) aims to modify the speaker's timbre while retaining speech content. Previous approaches have tokenized the outputs from self-supervised into semantic tokens, facilitating disentanglement of speech content information. Recently, in-context learning (ICL) has emerged in text-to-speech (TTS) systems for effectively modeling specific characteristics such as timbre through context conditioning. This paper proposes an ICL capability enhanced VC system (ICL-VC) employing a mask and reconstruction training strategy based on flow-matching generative models. Augmented with semantic tokens, our experiments on the LibriTTS dataset demonstrate that ICL-VC improves speaker similarity. Additionally, we find that k-means is a versatile tokenization method applicable to various pre-trained models. However, the ICL-VC system faces challenges in preserving the prosody of the source speech. To mitigate this issue, we propose incorporating prosody embeddings extracted from a pre-trained emotion recognition model into our system. Integration of prosody embeddings notably enhances the system's capability to preserve source speech prosody, as validated on the Emotional Speech Database.
【15】 Evaluating Neural Networks Architectures for Spring Reverb Modelling
标题: 评估春季回响建模的神经网络架构
作者:Francesco Papaleo,Xavier Lizarraga-Seijas,Frederic Font
备注:8 pages, 7 figures, 2 tables
链接:点击下载PDF文件
摘要:混响是空间音频感知中的一个关键要素,历史上使用模拟设备(如板和弹簧混响)实现,在过去几十年中使用数字信号处理技术实现,这些技术允许采用不同的方法进行虚拟混响建模(VAM)。弹簧混响的机电功能使其成为一个非线性系统,难以用白盒建模技术在数字域中完全仿真。在这项研究中,我们比较了五种不同的神经网络架构,包括卷积和递归模型,以评估它们在复制这种音频效果特征方面的有效性。在16 kHz和48 kHz的采样率下对两个数据集进行评估。本文特别关注提供参数控制的神经音频架构,旨在推进当前黑箱建模技术在春季混响领域的边界。摘要:Reverberation is a key element in spatial audio perception, historically achieved with the use of analogue devices, such as plate and spring reverb, and in the last decades with digital signal processing techniques that have allowed different approaches for Virtual Analogue Modelling (VAM). The electromechanical functioning of the spring reverb makes it a nonlinear system that is difficult to fully emulate in the digital domain with white-box modelling techniques. In this study, we compare five different neural network architectures, including convolutional and recurrent models, to assess their effectiveness in replicating the characteristics of this audio effect. The evaluation is conducted on two datasets at sampling rates of 16 kHz and 48 kHz. This paper specifically focuses on neural audio architectures that offer parametric control, aiming to advance the boundaries of current black-box modelling techniques in the domain of spring reverberation.
【16】 Attention-Based Efficient Breath Sound Removal in Studio Audio Recordings
标题: 录音室录音中基于注意力的高效呼吸声去除
作者:Nidula Elgiriyewithana,N. D. Kodikara
Journal-ref:CS & IT Conference Proceedings, vol. 14, no. 6, 2024
链接:点击下载PDF文件
摘要:在这项研究中,我们提出了一个创新的,参数高效的模型,利用注意力U-Net架构的自动检测和消除非语音声乐的声音,特别是呼吸声,在声乐录音。这项任务在声音工程领域是至关重要的,尽管相对来说还没有得到充分的探索。检测和消除这些声音的传统手动过程需要大量的专业知识,并且非常耗时。现有的自动检测和移除方法在效率和精度方面往往不足。我们提出的模型通过应用先进的深度学习技术,提供简化的流程和卓越的准确性,解决了这些限制。一个独特的数据集,来自设备和生产语音(DAPS),用于此目的。模型的训练阶段强调对数谱图,并集成了早期停止机制以防止过拟合。我们的模式不仅为音响工程师节省了宝贵的时间,还提高了音频制作的质量和一致性。这是一个重大突破,其相对效率证明了这一点,仅需要190万个参数和3.2小时的训练时间-明显低于该领域的顶级模型。该模型能够生成与先前模型相同的输出,并大幅提高精度,使其成为最佳选择。摘要:In this research, we present an innovative, parameter-efficient model that utilizes the attention U-Net architecture for the automatic detection and eradication of non-speech vocal sounds, specifically breath sounds, in vocal recordings. This task is of paramount importance in the field of sound engineering, despite being relatively under-explored. The conventional manual process for detecting and eliminating these sounds requires significant expertise and is extremely time-intensive. Existing automated detection and removal methods often fall short in terms of efficiency and precision. Our proposed model addresses these limitations by offering a streamlined process and superior accuracy, achieved through the application of advanced deep learning techniques. A unique dataset, derived from Device and Produced Speech (DAPS), was employed for this purpose. The training phase of the model emphasizes a log spectrogram and integrates an early stopping mechanism to prevent overfitting. Our model not only conserves precious time for sound engineers but also enhances the quality and consistency of audio production. This constitutes a significant breakthrough, as evidenced by its comparative efficiency, necessitating only 1.9M parameters and a training duration of 3.2 hours - markedly less than the top-performing models in this domain. The model is capable of generating identical outputs as previous models with drastically improved precision, making it an optimal choice.
【17】 Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
标题: Flow-TSVAD:通过潜在流匹配进行目标说话者语音活动检测
作者:Zhengyang Chen,Bing Han,Shuai Wang,Yidi Jiang,Yanmin Qian
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:说话者日记通常被认为是一项区分任务,使用区分方法来产生固定的日记结果。在本文中,我们首次探索使用基于神经网络的生成方法进行说话人日记化。我们实现了一个基于流匹配(FM)的生成算法内的序列到序列的目标说话人语音活动检测(Seq 2Seq-TSVAD)日记系统。我们的实验表明,直接将生成方法应用于TS-VAD输出的原始二进制标签序列空间是无效的。为了解决这个问题,我们建议在应用生成算法之前将二进制标签序列映射到密集的潜在空间中,并且我们提出的Flow-TSVAD方法优于Seq 2Seq-TSVAD系统。此外,我们观察到,FM算法收敛迅速,在推理阶段,只需要两个推理步骤,以实现有希望的结果。作为一个生成模型,Flow-TSVAD允许通过多次运行模型来采样不同的日志结果。此外,来自各种采样实例的集成结果进一步增强了日志化性能。摘要:Speaker diarization is typically considered a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore the use of neural network-based generative methods for speaker diarization for the first time. We implement a Flow-Matching (FM) based generative algorithm within the sequence-to-sequence target speaker voice activity detection (Seq2Seq-TSVAD) diarization system. Our experiments reveal that applying the generative method directly to the original binary label sequence space of the TS-VAD output is ineffective. To address this issue, we propose mapping the binary label sequence into a dense latent space before applying the generative algorithm and our proposed Flow-TSVAD method outperforms the Seq2Seq-TSVAD system. Additionally, we observe that the FM algorithm converges rapidly during the inference stage, requiring only two inference steps to achieve promising results. As a generative model, Flow-TSVAD allows for sampling different diarization results by running the model multiple times. Moreover, ensembling results from various sampling instances further enhances diarization performance.
【18】 PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
标题: PB-LRDDWWS系统,用于SYS 2024低资源发音障碍唤醒单词发现挑战赛
作者:Shiyao Wang,Jiaming Zhou,Shiwan Zhao,Yong Qin
备注:accept by SLT 2024
链接:点击下载PDF文件
摘要:对于2024年的低资源构音障碍唤醒词发现(LRDWWS)挑战赛,我们介绍了PB-LRDWWS系统。该系统将用于原型构建的构音障碍语音内容特征提取器与基于原型的分类方法相结合。特征提取器是一个微调的HuBERT模型,通过使用交叉熵损失的三阶段微调过程获得。这个经过微调的HuBERT从目标构音障碍演讲者的注册演讲中提取特征来构建原型。通过计算构音障碍说话人评价语音的HuBERT特征与原型之间的余弦相似度来实现分类。尽管它的简单性,我们的方法通过实验结果证明了有效性。我们的系统在LRDWWS挑战赛的最终测试B中获得第二名。摘要:For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extractor for prototype construction with a prototype-based classification method. The feature extractor is a fine-tuned HuBERT model obtained through a three-stage fine-tuning process using cross-entropy loss. This fine-tuned HuBERT extracts features from the target dysarthric speaker's enrollment speech to build prototypes. Classification is achieved by calculating the cosine similarity between the HuBERT features of the target dysarthric speaker's evaluation speech and prototypes. Despite its simplicity, our method demonstrates effectiveness through experimental results. Our system achieves second place in the final Test-B of the LRDWWS Challenge.
【19】 Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
标题: Mel-Roformer用于人声分离和人声旋律转录
作者:Ju-Chiang Wang,Wei-Tsung Lu,Jitong Chen
备注:Accepted to appear in ISMIR 2024
链接:点击下载PDF文件
摘要:开发一个多功能的深度神经网络来对音乐音频进行建模在MIR中至关重要。由于音乐信号中固有的复杂频谱变化,这项任务是具有挑战性的,这些音乐信号传达了不同乐器的旋律,和声和音色。在本文中,我们介绍了梅尔RoFormer,一个基于频谱图的模型具有两个关键的设计:一个新的梅尔波段投影模块在前端,以提高模型的能力,捕捉信息功能在多个频带,和交错的RoPE Transformers明确建模的频率和时间维度作为两个单独的序列。我们应用Mel-RoFormer来解决两个基本的MIR任务:声乐分离和声乐旋律转录,旨在将歌声从音频混合中分离出来,并分别转录其主旋律。尽管它们都关注于信号,但这些任务具有不同的优化目标。我们没有训练一个统一的模型,而是采用了两步走的方法。首先,我们训练一个声乐分离模型,随后作为声乐旋律转录微调的基础模型。通过在基准数据集上进行的大量实验,我们展示了我们的模型在声乐分离和旋律转录任务中实现了最先进的性能,强调了Mel-RoFormer在模拟复杂音乐音频信号方面的有效性和多功能性。摘要:Developing a versatile deep neural network to model music audio is crucial in MIR. This task is challenging due to the intricate spectral variations inherent in music signals, which convey melody, harmonics, and timbres of diverse instruments. In this paper, we introduce Mel-RoFormer, a spectrogram-based model featuring two key designs: a novel Mel-band Projection module at the front-end to enhance the model's capability to capture informative features across multiple frequency bands, and interleaved RoPE Transformers to explicitly model the frequency and time dimensions as two separate sequences. We apply Mel-RoFormer to tackle two essential MIR tasks: vocal separation and vocal melody transcription, aimed at isolating singing voices from audio mixtures and transcribing their lead melodies, respectively. Despite their shared focus on singing signals, these tasks possess distinct optimization objectives. Instead of training a unified model, we adopt a two-step approach. Initially, we train a vocal separation model, which subsequently serves as a foundation model for fine-tuning for vocal melody transcription. Through extensive experiments conducted on benchmark datasets, we showcase that our models achieve state-of-the-art performance in both vocal separation and melody transcription tasks, underscoring the efficacy and versatility of Mel-RoFormer in modeling complex music audio signals.
【20】 Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples
标题: 利用对比学习和自我训练以有限的标记样本进行多模式情感识别
作者:Qi Fan,Yutong Li,Yi Xin,Xinyu Cheng,Guanglai Gao,Miao Ma
备注:Accepted by ACM MM Workshop 2024
链接:点击下载PDF文件
摘要:多模态情感识别挑战MER 2024专注于使用音频,语言和视觉信号识别情感。在本文中,我们提出了半监督学习子挑战赛(MER 2024-SEMI)的提交解决方案,该挑战赛解决了情感识别中有限注释数据的问题。首先,为了解决类的不平衡,我们采用了过采样策略。其次,我们提出了一个模态表示组合对比学习(MR-CCL)框架上的三模态输入数据建立鲁棒的初始模型。第三,我们探索了一种自我训练的方法来扩展训练集。最后,我们通过多分类器加权软投票策略增强预测鲁棒性。我们提出的方法在MER 2024-SEMI挑战赛上被验证是有效的,实现了88.25%的加权平均F分数,在排行榜上排名第6。我们的项目可以在https: github.com WooyoohL MER2024-SEMI上找到。摘要:The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions for the Semi-Supervised Learning Sub-Challenge (MER2024-SEMI), which tackles the issue of limited annotated data in emotion recognition. Firstly, to address the class imbalance, we adopt an oversampling strategy. Secondly, we propose a modality representation combinatorial contrastive learning (MR-CCL) framework on the trimodal input data to establish robust initial models. Thirdly, we explore a self-training approach to expand the training set. Finally, we enhance prediction robustness through a multi-classifier weighted soft voting strategy. Our proposed method is validated to be effective on the MER2024-SEMI Challenge, achieving a weighted average F-score of 88.25% and ranking 6th on the leaderboard. Our project is available at https: github.com WooyoohL MER2024-SEMI.
【21】 Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
标题: 2024年普通话口吃事件检测和自动语音识别挑战赛的结果
作者:Hongfei Xue,Rong Gong,Mingchen Shao,Xin Xu,Lezhi Wang,Lei Xie,Hui Bu,Jiaming Zhou,Yong Qin,Jun Du,Ming Li,Binbin Zhang,Bin Jia
备注:8 pages, 2 figures, accepted by SLT 2024
链接:点击下载PDF文件
摘要:StutteringSpeech Challenge专注于为口吃者提供先进的语音技术,特别针对普通话中的口吃事件检测(SED)和自动语音识别(ASR)。该挑战包括三个轨道:(1)SED,旨在开发用于检测口吃事件的系统;(2)ASR,专注于创建用于识别口吃语音的强大系统;以及(3)利用所提供的数据集的创新方法的研究轨道。我们使用了一个开源的汉语口吃数据集AS-70,该数据集已经被分成了新的训练集和测试集。本文介绍了数据集,详细介绍了挑战跟踪,并分析了顶级系统的性能,突出了检测准确性的提高和识别错误率的降低。我们的研究结果强调了专门的模型和增强策略在开发口吃语音技术方面的潜力。摘要:The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recognition (ASR) in Mandarin. The challenge comprises three tracks: (1) SED, which aims to develop systems for detection of stuttering events; (2) ASR, which focuses on creating robust systems for recognizing stuttered speech; and (3) Research track for innovative approaches utilizing the provided dataset. We utilizes an open-source Mandarin stuttering dataset AS-70, which has been split into new training and test sets for the challenge. This paper presents the dataset, details the challenge tracks, and analyzes the performance of the top systems, highlighting improvements in detection accuracy and reductions in recognition error rates. Our findings underscore the potential of specialized models and augmentation strategies in developing stuttered speech technologies.
【22】 BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
标题: BigCodec:突破低比特率神经语音编解码器的极限
作者:Detai Xin,Xu Tan,Shinnosuke Takamichi,Hiroshi Saruwatari
备注:4 pages, 1 figure. Audio samples available at: this https URL
链接:点击下载PDF文件
摘要:我们提出了BigCodec,一个低比特率的神经语音编解码器。虽然最近的神经语音编解码器已经取得了令人印象深刻的进展,但它们的性能在低比特率(约1 kbps)下会显着恶化。虽然低比特率本质上限制了性能,但其他因素(如模型容量)也阻碍了进一步的改进。为了解决这个问题,我们将模型大小扩展到159 M参数,这是大约10 M参数的流行编解码器的10倍以上。此外,我们将序列模型集成到传统的卷积架构中,以更好地捕获时间依赖性,并采用低维矢量量化来确保高代码利用率。综合的客观和主观评估表明,BigCodec,1.04 kbps的比特率,显着优于现有的几个低比特率编解码器。此外,BigCodec实现了与以高出4-6倍的比特率运行的流行编解码器相当的客观性能,甚至提供了比地面真实更好的主观感知质量。摘要:We present BigCodec, a low-bitrate neural speech codec. While recent neural speech codecs have shown impressive progress, their performance significantly deteriorates at low bitrates (around 1 kbps). Although a low bitrate inherently restricts performance, other factors, such as model capacity, also hinder further improvements. To address this problem, we scale up the model size to 159M parameters that is more than 10 times larger than popular codecs with about 10M parameters. Besides, we integrate sequential models into traditional convolutional architectures to better capture temporal dependency and adopt low-dimensional vector quantization to ensure a high code utilization. Comprehensive objective and subjective evaluations show that BigCodec, with a bitrate of 1.04 kbps, significantly outperforms several existing low-bitrate codecs. Furthermore, BigCodec achieves objective performance comparable to popular codecs operating at 4-6 times higher bitrates, and even delivers better subjective perceptual quality than the ground truth.
【23】 SS-BRPE: Self-Supervised Blind Room Parameter Estimation Using Attention Mechanisms
标题: SS-BRPE:使用注意力机制的自我监督盲点参数估计
作者:Chunxi Wang,Maoshen Jia,Meiran Li,Changchun Bao,Wenyu Jin
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:近年来,声学环境的动态参数化在音频处理中引起了关注。这个焦点包括房间体积和混响时间(RT 60),它们定义了独立于声源和接收器方向的局部声学。以往的研究表明,纯粹的注意力为基础的模型可以实现先进的结果,在房间参数估计。然而,他们的成功依赖于监督预训练,这需要大量的房间参数和复杂训练管道的标记真值。鉴于此,我们提出了一种新的自监督盲房间参数估计(SS-BRPE)系统。该系统将纯粹基于注意力的模型与自监督学习相结合,从单通道有噪语音信号中估计房间声学参数。通过利用未标记的音频数据进行预训练,所提出的系统显着降低了对昂贵的标记数据集的依赖性。我们的模型还在微调过程中加入了动态特征增强,以增强适应性和泛化能力。实验结果表明,SS-BRPE系统不仅实现了更优越的性能,在估计房间参数比国家的最先进的(SOTA)的方法,但也有效地保持高精度的条件下,有限的标记数据。代码可在https: github.com bjut-chunxiwang SS-BRPE上获得。摘要:In recent years, dynamic parameterization of acoustic environments has garnered attention in audio processing. This focus includes room volume and reverberation time (RT60), which define local acoustics independent of sound source and receiver orientation. Previous studies show that purely attention-based models can achieve advanced results in room parameter estimation. However, their success relies on supervised pretrainings that require a large amount of labeled true values for room parameters and complex training pipelines. In light of this, we propose a novel Self-Supervised Blind Room Parameter Estimation (SS-BRPE) system. This system combines a purely attention-based model with self-supervised learning to estimate room acoustic parameters, from single-channel noisy speech signals. By utilizing unlabeled audio data for pretraining, the proposed system significantly reduces dependencies on costly labeled datasets. Our model also incorporates dynamic feature augmentation during fine-tuning to enhance adaptability and generalizability. Experimental results demonstrate that the SS-BRPE system not only achieves more superior performance in estimating room parameters than state-of-the-art (SOTA) methods but also effectively maintains high accuracy under conditions with limited labeled data. Code available at https: github.com bjut-chunxiwang SS-BRPE.
【24】 Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
标题: 使用薛定格桥和对称噪音表的基于扩散的语音增强
作者:Siyi Wang,Siyi Liu,Andrew Harper,Paul Kendrick,Mathieu Salzmann,Milos Cernak
链接:点击下载PDF文件
摘要:最近,基于扩散的生成模型在语音增强任务中表现出显着的性能。然而,这些方法仍然面临挑战,包括缺乏结构信息和低信噪比(SNR)情况下的性能差。为了克服这些挑战,我们提出了基于Schr“oodinger桥的语音增强(SBSE)方法,该方法直接学习噪声输入和干净分布之间的扩散过程,而不像传统的基于扩散的语音增强系统那样学习数据的高斯分布。为了提高性能,在非常嘈杂的条件下,我们引入了一个两阶段的系统,将比率掩模信息到基于扩散的生成模型。我们的实验结果表明,我们提出的SBSE方法优于所有的基线模型,并达到最先进的性能,特别是在低信噪比条件下。重要的是,只需要几个推理步骤就可以获得最佳结果。摘要:Recently, diffusion-based generative models have demonstrated remarkable performance in speech enhancement tasks. However, these methods still encounter challenges, including the lack of structural information and poor performance in low Signal-to-Noise Ratio (SNR) scenarios. To overcome these challenges, we propose the Schr "oodinger Bridge-based Speech Enhancement (SBSE) method, which learns the diffusion processes directly between the noisy input and the clean distribution, unlike conventional diffusion-based speech enhancement systems that learn data to Gaussian distributions. To enhance performance in extremely noisy conditions, we introduce a two-stage system incorporating ratio mask information into the diffusion-based generative model. Our experimental results show that our proposed SBSE method outperforms all the baseline models and achieves state-of-the-art performance, especially in low SNR conditions. Importantly, only a few inference steps are required to achieve the best result.
【25】 TF-Mamba: A Time-Frequency Network for Sound Source Localization
标题: TF-Mamba:一种用于光源定位的时频网络
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:声源定位(SSL)使用多声道音频数据确定声源的位置。它通常用于改善语音增强和分离。提取空间特征对于SSL至关重要,特别是在具有挑战性的声学环境中。以前的研究基于长短期记忆模型进行得很好。最近,一种称为Mamba的新型可扩展SSM在各种基于序列的模式(包括音频和语音)中表现出显着的性能。本研究介绍了用于SSL任务的Mamba。我们考虑基于Mamba的模型,通过融合时间和频率特征来分析语音信号的空间特征,并开发了一个SSL系统TF-Mamba。该系统集成了时间和频率融合,Bidirectional Mamba管理时间和频率处理。我们在模拟数据集和LOCATA数据集上进行了实验。实验表明,TF-Mamba在模拟和真实数据上的性能明显优于其他先进方法。摘要:Sound source localization (SSL) determines the position of sound sources using multi-channel audio data. It is commonly used to improve speech enhancement and separation. Extracting spatial features is crucial for SSL, especially in challenging acoustic environments. Previous studies performed well based on long short-term memory models. Recently, a novel scalable SSM referred to as Mamba demonstrated notable performance across various sequence-based modalities, including audio and speech. This study introduces the Mamba for SSL tasks. We consider the Mamba-based model to analyze spatial features from speech signals by fusing both time and frequency features, and we develop an SSL system called TF-Mamba. This system integrates time and frequency fusion, with Bidirectional Mamba managing both time-wise and frequency-wise processing. We conduct the experiments on the simulated dataset and the LOCATA dataset. Experiments show that TF-Mamba significantly outperforms other advanced methods on simulated and real-world data.
【26】 Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
标题: 探索WavLM后台进行语音欺骗和Deepfake检测
作者:Theophile Stourbe,Victor Miara,Theo Lepage,Reda Dehak
链接:点击下载PDF文件
摘要:本文描述了我们提交给ASVspoof 5挑战赛的系统:Speech Deepfake Detection - Open Condition,其中包括一个独立的语音deepfake(bonafide vs spoof)检测任务。最近,大规模自监督模型成为自动语音识别(ASR)和其他语音处理任务的标准。因此,我们利用预先训练的WavLM作为前端模型,并将其表示与不同的后端技术相结合。完整的框架仅使用挑战的训练数据集进行微调,类似于关闭条件。此外,我们采用数据增强,通过添加噪声和混响使用MUSAN噪声和RIR数据集。我们还尝试了编解码器增强,以提高我们的方法的性能。最终,我们使用Bosaris工具包进行分数校准和系统融合,以获得更好的Cllr分数。我们的融合系统达到0.0937 minDCF,3.42%EER,0.1927 Cllr,和0.1375 actDCF。摘要:This paper describes our submitted systems to the ASVspoof 5 Challenge Track 1: Speech Deepfake Detection - Open Condition, which consists of a stand-alone speech deepfake (bonafide vs spoof) detection task. Recently, large-scale self-supervised models become a standard in Automatic Speech Recognition (ASR) and other speech processing tasks. Thus, we leverage a pre-trained WavLM as a front-end model and pool its representations with different back-end techniques. The complete framework is fine-tuned using only the trained dataset of the challenge, similar to the close condition. Besides, we adopt data-augmentation by adding noise and reverberation using MUSAN noise and RIR datasets. We also experiment with codec augmentations to increase the performance of our method. Ultimately, we use the Bosaris toolkit for score calibration and system fusion to get better Cllr scores. Our fused system achieves 0.0937 minDCF, 3.42% EER, 0.1927 Cllr, and 0.1375 actDCF.
【27】 Leveraging Moving Sound Source Trajectories for Universal Sound Separation
标题: 利用移动光源轨迹实现通用声音分离
作者:Donghang Wu,Xihong Wu,Tianshu Qu
备注:9 pages,7 figures,submitted to IEEEACM Transactions on Audio, Speech and Language Processing(TASLP)
链接:点击下载PDF文件
摘要:利用空间信息进行声源分离的现有方法需要源的到达方向(DOA)的先验知识,或者利用估计但不精确的定位结果,这损害了分离性能,特别是当声源移动时。事实上,声源定位和分离是相互关联的问题,即声源定位有助于声音分离,而声音分离有助于更精确的声源定位。本文提出了一种利用声源定位和运动声源分离之间相互促进机制的方法。最初,使用粗略的初步声源跟踪结果进行声音分离。然后对分离的信号执行声源跟踪,从而跟踪结果可以变得更精确。精确的弹道可以进一步提高分离性能。这种相互促进的过程可以在几次迭代中执行。混响条件下和运动声源下的仿真实验表明,该方法可以实现更准确的分离的基础上更精确的跟踪结果。摘要:Existing methods utilizing spatial information for sound source separation require prior knowledge of the direction of arrival (DOA) of the source or utilize estimated but imprecise localization results, which impairs the separation performance, especially when the sound sources are moving. In fact, sound source localization and separation are interconnected problems, that is, sound source localization facilitates sound separation while sound separation contributes to more precise source localization. This paper proposes a method utilizing the mutual facilitation mechanism between sound source localization and separation for moving sources. Initially, sound separation is conducted using rough preliminary sound source tracking results. Sound source tracking is then performed on the separated signals thus the tracking results can become more precise. The precise trajectory can further enhances the separation performance. This mutual facilitation process can be performed over several iterations. Simulation experiments conducted under reverberation conditions and with moving sound sources demonstrate that the proposed method can achieve more accurate separation based on more precise tracking results.
【28】 Cross-attention Inspired Selective State Space Models for Target Sound Extraction
标题: 交叉注意启发的选择性状态空间模型用于目标声音提取
作者:Donghang Wu,Yiwen Wang,Xihong Wu,Tianshu Qu
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:Transformer模型,特别是其交叉注意模块,被广泛用于目标声音提取中的特征融合,其基于给定的线索提取感兴趣的信号。尽管它的有效性,这种方法遭受低的计算效率。状态空间模型的最新进展,特别是最新的工作Mamba,已经显示出与基于Transformer的方法相当的性能,同时显着降低了各种任务的计算复杂性。然而,Mamba在目标声音提取中的适用性是有限的,因为它不能像交叉注意那样捕获不同序列之间的依赖关系。在本文中,我们提出了CrossMamba的目标声音提取,它利用隐藏的注意力机制的Mamba来计算给定的线索和音频混合之间的依赖关系。Mamba的计算可以分为查询、键和值。我们利用线索来生成查询和音频混合来获得键和值,遵循Transformers中的交叉注意机制的原理。两种典型目标声提取方法的实验结果验证了CrossMamba算法的有效性。摘要:The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba's applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.
eess.AS音频处理
【1】 A Toolkit for Joint Speaker Diarization and Identification with Application to Speaker-Attributed ASR标题: 联合说话人拨号和识别工具包及其应用于说话人归因的ASB
作者:Giovanni Morrone,Enrico Zovato,Fabio Brugnara,Enrico Sartori,Leonardo Badino
Journal-ref:Proceedings of Interspeech 2024, pp. 3652--3653
链接:点击下载PDF文件
摘要:我们提出了一个模块化的工具包来执行联合说话人日记和说话人识别。该工具包可以利用在配置文件中定义的多个模型和算法。这种灵活性允许我们的系统在各种条件下正常工作(例如,多个注册的说话者的集合、声学条件和语言)和跨应用领域(例如,媒体监控、机构、语音分析)。在这个演示中,我们展示了一个实际的用例,其中与说话者相关的信息与自动语音识别引擎联合使用,以生成说话者属性transmittance。为了实现这一点,我们采用了一个用户友好的基于Web的界面来处理音频和视频输入与所选的配置。摘要:We present a modular toolkit to perform joint speaker diarization and speaker identification. The toolkit can leverage on multiple models and algorithms which are defined in a configuration file. Such flexibility allows our system to work properly in various conditions (e.g., multiple registered speakers' sets, acoustic conditions and languages) and across application domains (e.g. media monitoring, institutional, speech analytics). In this demonstration we show a practical use-case in which speaker-related information is used jointly with automatic speech recognition engines to generate speaker-attributed transcriptions. To achieve that, we employ a user-friendly web-based interface to process audio and video inputs with the chosen configuration.
【2】 AS-Speech: Adaptive Style For Speech Synthesis
标题: AS-Speech:语音合成的自适应风格
作者:Zhipeng Li,Xiaofen Xing,Jun Wang,Shuaiqi Chen,Guoqiao Yu,Guanglu Wan,Xiangmin Xu
备注:Accepted by SLT 2024
链接:点击下载PDF文件
摘要:近年来,文本到语音(TTS)合成技术取得了重大进展,使得能够在常见场景中高质量地合成语音。在不可见的情况下,自适应TTS需要对说话人风格特征有很强的泛化能力。然而,现有的自适应方法只能分别提取和整合粗粒度的音色或混合节奏属性。在本文中,我们提出了AS-Speech,一种自适应风格的方法,它集成了扬声器的音色特征和节奏属性到一个统一的框架,用于文本到语音合成。具体而言,AS-Speech通过细粒度的基于文本的音色特征和全局节奏信息,准确模拟风格特征,并通过扩散模型实现高保真语音合成。实验表明,该模型产生的语音具有更高的自然度和相似性的音色和节奏相比,一系列的自适应TTS模型。摘要:In recent years, there has been significant progress in Text-to-Speech (TTS) synthesis technology, enabling the high-quality synthesis of voices in common scenarios. In unseen situations, adaptive TTS requires a strong generalization capability to speaker style characteristics. However, the existing adaptive methods can only extract and integrate coarse-grained timbre or mixed rhythm attributes separately. In this paper, we propose AS-Speech, an adaptive style methodology that integrates the speaker timbre characteristics and rhythmic attributes into a unified framework for text-to-speech synthesis. Specifically, AS-Speech can accurately simulate style characteristics through fine-grained text-based timbre features and global rhythm information, and achieve high-fidelity speech synthesis through the diffusion model. Experiments show that the proposed model produces voices with higher naturalness and similarity in terms of timbre and rhythm compared to a series of adaptive TTS models.
【3】 Longer is (Not Necessarily) Stronger: Punctuated Long-Sequence Training for Enhanced Speech Recognition and Translation
标题: 越长(不一定)越强:用于增强语音识别和翻译的间断长序列训练
作者:Nithin Rao Koluguri,Travis Bartley,Hainan Xu,Oleksii Hrinchuk,Jagadeesh Balam,Boris Ginsburg,Georg Kucsko
备注:Accepted at SLT 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的方法来训练序列到序列模型的语音识别和翻译任务。与传统的仅包含标点符号或部分标点符号和大写(PnC)句子的短片段训练模型的方法不同,我们建议对包含完整句子的较长话语进行训练,这些句子具有适当的标点符号和大写。我们通过使用FastConformer架构来实现这一点,该架构允许训练10亿个参数模型,其序列长达60秒,并具有完全的注意力。然而,虽然使用PnC进行训练可以提高整体性能,但我们观察到,在各种评估设置中,当训练超过40秒的序列时,准确率会达到平台。我们提出的方法显着提高标点符号和大写的准确性,显示了25%的相对字错误率(WER)的收益-21和收益-22基准改善。此外,在较长的音频段上进行训练可以提高语音识别和翻译基准的整体模型准确性。模型权重和训练代码通过NVIDIA NeMo开源。摘要:This paper presents a new method for training sequence-to-sequence models for speech recognition and translation tasks. Instead of the traditional approach of training models on short segments containing only lowercase or partial punctuation and capitalization (PnC) sentences, we propose training on longer utterances that include complete sentences with proper punctuation and capitalization. We achieve this by using the FastConformer architecture which allows training 1 Billion parameter models with sequences up to 60 seconds long with full attention. However, while training with PnC enhances the overall performance, we observed that accuracy plateaus when training on sequences longer than 40 seconds across various evaluation settings. Our proposed method significantly improves punctuation and capitalization accuracy, showing a 25% relative word error rate (WER) improvement on the Earnings-21 and Earnings-22 benchmarks. Additionally, training on longer audio segments increases the overall model accuracy across speech recognition and translation benchmarks. The model weights and training code are open-sourced though NVIDIA NeMo.
【4】 An investigation of modularity for noise robustness in conformer-based ASR
标题: 基于一致性的ASB中噪音鲁棒性的模块化研究
作者:Louise Coppieters de Gibson,Philip N. Garner,Pierre-Edouard Honnet
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:虽然最先进的自动语音识别(ASR)可以表现良好,但当暴露于与训练模型时使用的声学环境不同的声学环境时,它仍然会降级。对于一个给定的模型来说,不熟悉的环境可能是先验的,但产生的适应数据相对较少。在这项实验研究中,我们调查在何种程度上最近的形式化的模块化可以帮助适应新的声学环境的ASR。使用基于一致性的模型和固定路由,我们确认环境感知确实可以提高已知环境中的性能。然而,至少在研究中的(CHIME)数据集上,分类器模块很难区分不同的嘈杂环境,嘈杂和干净语音之间的简单区分是最佳配置。这些结果对于在特定环境中部署大型模型具有明确的意义,无论是否具有环境噪声的先验知识。摘要:Whilst state of the art automatic speech recognition (ASR) can perform well, it still degrades when exposed to acoustic environments that differ from those used when training the model. Unfamiliar environments for a given model may well be known a-priori, but yield comparatively small amounts of adaptation data. In this experimental study, we investigate to what extent recent formalisations of modularity can aid adaptation of ASR to new acoustic environments. Using a conformer based model and fixed routing, we confirm that environment awareness can indeed lead to improved performance in known environments. However, at least on the (CHIME) datasets in the study, it is difficult for a classifier module to distinguish different noisy environments, a simpler distinction between noisy and clean speech being the optimal configuration. The results have clear implications for deploying large models in particular environments with or without a-priori knowledge of the environmental noise.
【5】 Leveraging Content and Acoustic Representations for Efficient Speech Emotion Recognition
标题: 利用内容和声学表示实现高效的语音情感识别
作者:Soumya Dutta,Sriram Ganapathy
备注:10 pages, 4 figures, 7 tables
链接:点击下载PDF文件
摘要:语音情感识别(SER)是从语音内容中识别情感表达的任务,由于难以从语音中提取捕获情感属性的表示而具有挑战性。大型标记数据集的稀缺性使大型模型容易过度拟合的挑战进一步复杂化。在本文中,我们提出了照顾(内容和声学表示的情绪),我们设计了一个双重编码方案,强调语音的语义和声学因素。语义编码器通过提取话语级文本表示模型进行训练,而声学编码器则通过训练预测语音信号的低层帧特征。提出的双重编码方案是一个基本大小的模型,只在无监督的原始语音训练。通过在下游任务上训练的简单轻量级分类模型,我们证明了CARE嵌入在各种任务上提供了有效的情感识别。我们将该建议与其他几种自监督模型以及最近的基于大语言模型的方法进行了比较。在这些评估中,基于8个不同数据集的平均性能,所提出的CARE模型被证明是性能最好的模型。我们还进行了几项消融研究,以分析各种设计选择的重要性。摘要:Speech emotion recognition (SER), the task of identifying the expression of emotion from spoken content, is challenging due to the difficulty in extracting representations that capture emotional attributes from speech. The scarcity of large labeled datasets further complicates the challenge where large models are prone to over-fitting. In this paper, we propose CARE (Content and Acoustic Representations of Emotions), where we design a dual encoding scheme which emphasizes semantic and acoustic factors of speech. While the semantic encoder is trained with the distillation of utterance-level text representation model, the acoustic encoder is trained to predict low-level frame-wise features of the speech signal. The proposed dual encoding scheme is a base-sized model trained only on unsupervised raw speech. With a simple light-weight classification model trained on the downstream task, we show that the CARE embeddings provide effective emotion recognition on a variety of tasks. We compare the proposal with several other self-supervised models as well as recent large-language model based approaches. In these evaluations, the proposed CARE model is shown to be the best performing model based on average performance across 8 diverse datasets. We also conduct several ablation studies to analyze the importance of various design choices.
【6】 NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
标题: 用于CHiME-8挑战赛DSVR任务的NTT多扬声器ASB系统
作者:Naoyuki Kamo,Naohiro Tawara,Atsushi Ando,Takatomo Kano,Hiroshi Sato,Rintaro Ikeshita,Takafumi Moriya,Shota Horiguchi,Kohei Matsuura,Atsunori Ogawa,Alexis Plaquet,Takanori Ashihara,Tsubasa Ochiai,Masato Mimura,Marc Delcroix,Tomohiro Nakatani,Taichi Asami,Shoko Araki
备注:5 pages, 4 figures, CHiME8 challenge
链接:点击下载PDF文件
摘要:我们提出了一个远程自动语音识别(DASR)系统开发的CHiME-8 DASR轨道。它包括一个日志化的第一管道。对于日志化,我们使用端到端日志化与向量聚类(EEND-VC),然后目标说话人语音活动检测(TS-VAD)细化。为了处理不同数量的说话人,我们开发了一种新的多通道说话人计数方法。然后,我们应用引导源分离(GSS),并对基线系统进行了几项改进。最后,我们使用由强大的预训练模型构建的系统组合来执行ASR。我们提出的系统在开发集上实现了21.3%的宏tcpWER,这比基线相对提高了57%。摘要:We present a distant automatic speech recognition (DASR) system developed for the CHiME-8 DASR track. It consists of a diarization first pipeline. For diarization, we use end-to-end diarization with vector clustering (EEND-VC) followed by target speaker voice activity detection (TS-VAD) refinement. To deal with various numbers of speakers, we developed a new multi-channel speaker counting approach. We then apply guided source separation (GSS) with several improvements to the baseline system. Finally, we perform ASR using a combination of systems built from strong pre-trained models. Our proposed system achieves a macro tcpWER of 21.3 % on the dev set, which is a 57 % relative improvement over the baseline.
【7】 Transferable Selective Virtual Sensing Active Noise Control Technique Based on Metric Learning
标题: 基于度量学习的可转移选择性虚拟感知主动噪音控制技术
作者:Boxiang Wang,Dongyuan Shi,Zhengding Luo,Xiaoyi Shen,Junwei Ji,Woon-Seng Gan
链接:点击下载PDF文件
摘要:虚拟感测(VS)技术使有源噪声控制(ANC)系统能够衰减远离物理误差麦克风的虚拟位置处的噪声。适当的辅助滤波器(AF)可以显着提高VS方法的有效性。可以使用卷积神经网络(CNN)自动实现针对各种类型的噪声选择适当的AF。然而,训练用于不同ANC系统的CNN模型通常是劳动密集型和耗时的。为了解决这个问题,我们提出了一种新的方法,可转移的选择性VS,通过将度量学习技术集成到基于CNN的VS方法。可转移选择性VS方法允许将预先训练的CNN直接应用于新的ANC系统,而无需重新训练,并且可以处理不可见的噪声类型。数值仿真结果表明,该方法在抑制突变宽带噪声和真实噪声方面是有效的。摘要:Virtual sensing (VS) technology enables active noise control (ANC) systems to attenuate noise at virtual locations distant from the physical error microphones. Appropriate auxiliary filters (AF) can significantly enhance the effectiveness of VS approaches. The selection of appropriate AF for various types of noise can be automatically achieved using convolutional neural networks (CNNs). However, training the CNN model for different ANC systems is often labour-intensive and time-consuming. To tackle this problem, we propose a novel method, Transferable Selective VS, by integrating metric-learning technology into CNN-based VS approaches. The Transferable Selective VS method allows a pre-trained CNN to be applied directly to new ANC systems without requiring retraining, and it can handle unseen noise types. Numerical simulations demonstrate the effectiveness of the proposed method in attenuating sudden-varying broadband noises and real-world noises.
【8】 Findings of the 2024 Mandarin Stuttering Event Detection and Automatic Speech Recognition Challenge
标题: 2024年普通话口吃事件检测和自动语音识别挑战赛的结果
作者:Hongfei Xue,Rong Gong,Mingchen Shao,Xin Xu,Lezhi Wang,Lei Xie,Hui Bu,Jiaming Zhou,Yong Qin,Jun Du,Ming Li,Binbin Zhang,Bin Jia
备注:8 pages, 2 figures, accepted by SLT 2024
链接:点击下载PDF文件
摘要:StutteringSpeech Challenge专注于为口吃者提供先进的语音技术,特别针对普通话中的口吃事件检测(SED)和自动语音识别(ASR)。该挑战包括三个轨道:(1)SED,旨在开发用于检测口吃事件的系统;(2)ASR,专注于创建用于识别口吃语音的强大系统;以及(3)利用所提供的数据集的创新方法的研究轨道。我们使用了一个开源的汉语口吃数据集AS-70,该数据集已经被分成了新的训练集和测试集。本文介绍了数据集,详细介绍了挑战跟踪,并分析了顶级系统的性能,突出了检测准确性的提高和识别错误率的降低。我们的研究结果强调了专门的模型和增强策略在开发口吃语音技术方面的潜力。摘要:The StutteringSpeech Challenge focuses on advancing speech technologies for people who stutter, specifically targeting Stuttering Event Detection (SED) and Automatic Speech Recognition (ASR) in Mandarin. The challenge comprises three tracks: (1) SED, which aims to develop systems for detection of stuttering events; (2) ASR, which focuses on creating robust systems for recognizing stuttered speech; and (3) Research track for innovative approaches utilizing the provided dataset. We utilizes an open-source Mandarin stuttering dataset AS-70, which has been split into new training and test sets for the challenge. This paper presents the dataset, details the challenge tracks, and analyzes the performance of the top systems, highlighting improvements in detection accuracy and reductions in recognition error rates. Our findings underscore the potential of specialized models and augmentation strategies in developing stuttered speech technologies.
【9】 BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
标题: BigCodec:突破低比特率神经语音编解码器的极限
作者:Detai Xin,Xu Tan,Shinnosuke Takamichi,Hiroshi Saruwatari
备注:4 pages, 1 figure. Audio samples available at: this https URL
链接:点击下载PDF文件
摘要:我们提出了BigCodec,一个低比特率的神经语音编解码器。虽然最近的神经语音编解码器已经取得了令人印象深刻的进展,但它们的性能在低比特率(约1 kbps)下会显着恶化。虽然低比特率本质上限制了性能,但其他因素(如模型容量)也阻碍了进一步的改进。为了解决这个问题,我们将模型大小扩展到159 M参数,这是大约10 M参数的流行编解码器的10倍以上。此外,我们将序列模型集成到传统的卷积架构中,以更好地捕获时间依赖性,并采用低维矢量量化来确保高代码利用率。综合的客观和主观评估表明,BigCodec,1.04 kbps的比特率,显着优于现有的几个低比特率编解码器。此外,BigCodec实现了与以高出4-6倍的比特率运行的流行编解码器相当的客观性能,甚至提供了比地面真实更好的主观感知质量。摘要:We present BigCodec, a low-bitrate neural speech codec. While recent neural speech codecs have shown impressive progress, their performance significantly deteriorates at low bitrates (around 1 kbps). Although a low bitrate inherently restricts performance, other factors, such as model capacity, also hinder further improvements. To address this problem, we scale up the model size to 159M parameters that is more than 10 times larger than popular codecs with about 10M parameters. Besides, we integrate sequential models into traditional convolutional architectures to better capture temporal dependency and adopt low-dimensional vector quantization to ensure a high code utilization. Comprehensive objective and subjective evaluations show that BigCodec, with a bitrate of 1.04 kbps, significantly outperforms several existing low-bitrate codecs. Furthermore, BigCodec achieves objective performance comparable to popular codecs operating at 4-6 times higher bitrates, and even delivers better subjective perceptual quality than the ground truth.
【10】 SS-BRPE: Self-Supervised Blind Room Parameter Estimation Using Attention Mechanisms
标题: SS-BRPE:使用注意力机制的自我监督盲点参数估计
作者:Chunxi Wang,Maoshen Jia,Meiran Li,Changchun Bao,Wenyu Jin
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:近年来,声学环境的动态参数化在音频处理中引起了关注。这个焦点包括房间体积和混响时间(RT 60),它们定义了独立于声源和接收器方向的局部声学。以往的研究表明,纯粹的注意力为基础的模型可以实现先进的结果,在房间参数估计。然而,他们的成功依赖于监督预训练,这需要大量的房间参数和复杂训练管道的标记真值。鉴于此,我们提出了一种新的自监督盲房间参数估计(SS-BRPE)系统。该系统结合了一个纯粹的基于注意力的模型与自监督学习,估计房间的声学参数,从单通道嘈杂的语音信号。通过利用未标记的音频数据进行预训练,所提出的系统显着降低了对昂贵的标记数据集的依赖性。我们的模型还在微调过程中加入了动态特征增强,以增强适应性和泛化能力。实验结果表明,SS-BRPE系统不仅实现了更优越的性能,在估计房间参数比国家的最先进的(SOTA)的方法,但也有效地保持高精度的条件下,有限的标记数据。代码可在https: github.com bjut-chunxiwang SS-BRPE上获得。摘要:In recent years, dynamic parameterization of acoustic environments has garnered attention in audio processing. This focus includes room volume and reverberation time (RT60), which define local acoustics independent of sound source and receiver orientation. Previous studies show that purely attention-based models can achieve advanced results in room parameter estimation. However, their success relies on supervised pretrainings that require a large amount of labeled true values for room parameters and complex training pipelines. In light of this, we propose a novel Self-Supervised Blind Room Parameter Estimation (SS-BRPE) system. This system combines a purely attention-based model with self-supervised learning to estimate room acoustic parameters, from single-channel noisy speech signals. By utilizing unlabeled audio data for pretraining, the proposed system significantly reduces dependencies on costly labeled datasets. Our model also incorporates dynamic feature augmentation during fine-tuning to enhance adaptability and generalizability. Experimental results demonstrate that the SS-BRPE system not only achieves more superior performance in estimating room parameters than state-of-the-art (SOTA) methods but also effectively maintains high accuracy under conditions with limited labeled data. Code available at https: github.com bjut-chunxiwang SS-BRPE.
【11】 Diffusion-based Speech Enhancement with Schrödinger Bridge and Symmetric Noise Schedule
标题: 使用薛定格桥和对称噪音表的基于扩散的语音增强
作者:Siyi Wang,Siyi Liu,Andrew Harper,Paul Kendrick,Mathieu Salzmann,Milos Cernak
链接:点击下载PDF文件
摘要:最近,基于扩散的生成模型在语音增强任务中表现出显着的性能。然而,这些方法仍然面临挑战,包括缺乏结构信息和低信噪比(SNR)情况下的性能差。为了克服这些挑战,我们提出了基于Schr“oodinger桥的语音增强(SBSE)方法,该方法直接学习噪声输入和干净分布之间的扩散过程,而不像传统的基于扩散的语音增强系统那样学习数据的高斯分布。为了提高性能,在非常嘈杂的条件下,我们引入了一个两阶段的系统,将比率掩模信息到基于扩散的生成模型。我们的实验结果表明,我们提出的SBSE方法优于所有的基线模型,并达到最先进的性能,特别是在低信噪比条件下。重要的是,只需要几个推理步骤就可以获得最佳结果。摘要:Recently, diffusion-based generative models have demonstrated remarkable performance in speech enhancement tasks. However, these methods still encounter challenges, including the lack of structural information and poor performance in low Signal-to-Noise Ratio (SNR) scenarios. To overcome these challenges, we propose the Schr "oodinger Bridge-based Speech Enhancement (SBSE) method, which learns the diffusion processes directly between the noisy input and the clean distribution, unlike conventional diffusion-based speech enhancement systems that learn data to Gaussian distributions. To enhance performance in extremely noisy conditions, we introduce a two-stage system incorporating ratio mask information into the diffusion-based generative model. Our experimental results show that our proposed SBSE method outperforms all the baseline models and achieves state-of-the-art performance, especially in low SNR conditions. Importantly, only a few inference steps are required to achieve the best result.
【12】 TF-Mamba: A Time-Frequency Network for Sound Source Localization
标题: TF-Mamba:一种用于光源定位的时频网络
作者:Yang Xiao,Rohan Kumar Das
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:声源定位(SSL)使用多声道音频数据确定声源的位置。它通常用于改善语音增强和分离。提取空间特征对于SSL至关重要,特别是在具有挑战性的声学环境中。以前的研究基于长短期记忆模型进行得很好。最近,一种称为Mamba的新型可扩展SSM在各种基于序列的模式(包括音频和语音)中表现出显着的性能。本研究介绍了用于SSL任务的Mamba。我们考虑基于Mamba的模型,通过融合时间和频率特征来分析语音信号的空间特征,并开发了一个SSL系统TF-Mamba。该系统集成了时间和频率融合,Bidirectional Mamba管理时间和频率处理。我们在模拟数据集和LOCATA数据集上进行了实验。实验表明,TF-Mamba在模拟和真实数据上的性能明显优于其他先进方法。摘要:Sound source localization (SSL) determines the position of sound sources using multi-channel audio data. It is commonly used to improve speech enhancement and separation. Extracting spatial features is crucial for SSL, especially in challenging acoustic environments. Previous studies performed well based on long short-term memory models. Recently, a novel scalable SSM referred to as Mamba demonstrated notable performance across various sequence-based modalities, including audio and speech. This study introduces the Mamba for SSL tasks. We consider the Mamba-based model to analyze spatial features from speech signals by fusing both time and frequency features, and we develop an SSL system called TF-Mamba. This system integrates time and frequency fusion, with Bidirectional Mamba managing both time-wise and frequency-wise processing. We conduct the experiments on the simulated dataset and the LOCATA dataset. Experiments show that TF-Mamba significantly outperforms other advanced methods on simulated and real-world data.
【13】 Exploring WavLM Back-ends for Speech Spoofing and Deepfake Detection
标题: 探索WavLM后台进行语音欺骗和Deepfake检测
作者:Theophile Stourbe,Victor Miara,Theo Lepage,Reda Dehak
链接:点击下载PDF文件
摘要:本文描述了我们提交给ASVspoof 5挑战赛的系统:Speech Deepfake Detection - Open Condition,其中包括一个独立的语音deepfake(bonafide vs spoof)检测任务。近年来,大规模自监督模型成为自动语音识别(ASR)和其他语音处理任务的标准。因此,我们利用预先训练的WavLM作为前端模型,并将其表示与不同的后端技术相结合。完整的框架仅使用挑战的训练数据集进行微调,类似于关闭条件。此外,我们采用数据增强,通过添加噪声和混响使用MUSAN噪声和RIR数据集。我们还尝试了编解码器增强,以提高我们的方法的性能。最终,我们使用Bosaris工具包进行分数校准和系统融合,以获得更好的Cllr分数。我们的融合系统达到0.0937 minDCF,3.42%EER,0.1927 Cllr,和0.1375 actDCF。摘要:This paper describes our submitted systems to the ASVspoof 5 Challenge Track 1: Speech Deepfake Detection - Open Condition, which consists of a stand-alone speech deepfake (bonafide vs spoof) detection task. Recently, large-scale self-supervised models become a standard in Automatic Speech Recognition (ASR) and other speech processing tasks. Thus, we leverage a pre-trained WavLM as a front-end model and pool its representations with different back-end techniques. The complete framework is fine-tuned using only the trained dataset of the challenge, similar to the close condition. Besides, we adopt data-augmentation by adding noise and reverberation using MUSAN noise and RIR datasets. We also experiment with codec augmentations to increase the performance of our method. Ultimately, we use the Bosaris toolkit for score calibration and system fusion to get better Cllr scores. Our fused system achieves 0.0937 minDCF, 3.42% EER, 0.1927 Cllr, and 0.1375 actDCF.
【14】 Leveraging Moving Sound Source Trajectories for Universal Sound Separation
标题: 利用移动光源轨迹实现通用声音分离
作者:Donghang Wu,Xihong Wu,Tianshu Qu
备注:9 pages,7 figures,submitted to IEEEACM Transactions on Audio, Speech and Language Processing(TASLP)
链接:点击下载PDF文件
摘要:利用空间信息进行声源分离的现有方法需要源的到达方向(DOA)的先验知识,或者利用估计但不精确的定位结果,这损害了分离性能,特别是当声源移动时。事实上,声源定位和分离是相互关联的问题,即声源定位有助于声音分离,而声音分离有助于更精确的声源定位。本文提出了一种利用声源定位和运动声源分离之间相互促进机制的方法。最初,使用粗略的初步声源跟踪结果进行声音分离。然后对分离的信号执行声源跟踪,从而跟踪结果可以变得更精确。精确的轨迹可以进一步提高分离性能。这种相互促进的过程可以在几个迭代中执行。混响条件下和运动声源下的仿真实验表明,该方法可以实现更准确的分离的基础上更精确的跟踪结果。摘要:Existing methods utilizing spatial information for sound source separation require prior knowledge of the direction of arrival (DOA) of the source or utilize estimated but imprecise localization results, which impairs the separation performance, especially when the sound sources are moving. In fact, sound source localization and separation are interconnected problems, that is, sound source localization facilitates sound separation while sound separation contributes to more precise source localization. This paper proposes a method utilizing the mutual facilitation mechanism between sound source localization and separation for moving sources. Initially, sound separation is conducted using rough preliminary sound source tracking results. Sound source tracking is then performed on the separated signals thus the tracking results can become more precise. The precise trajectory can further enhances the separation performance. This mutual facilitation process can be performed over several iterations. Simulation experiments conducted under reverberation conditions and with moving sound sources demonstrate that the proposed method can achieve more accurate separation based on more precise tracking results.
【15】 Cross-attention Inspired Selective State Space Models for Target Sound Extraction
标题: 交叉注意启发的选择性状态空间模型用于目标声音提取
作者:Donghang Wu,Yiwen Wang,Xihong Wu,Tianshu Qu
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:Transformer模型,特别是其交叉注意模块,被广泛用于目标声音提取中的特征融合,其基于给定的线索提取感兴趣的信号。尽管它的有效性,这种方法遭受低的计算效率。状态空间模型的最新进展,特别是最新的工作Mamba,已经显示出与基于Transformer的方法相当的性能,同时显着降低了各种任务的计算复杂性。然而,Mamba在目标声音提取中的适用性是有限的,因为它不能像交叉注意那样捕获不同序列之间的依赖关系。在本文中,我们提出了CrossMamba的目标声音提取,它利用隐藏的注意力机制的Mamba来计算给定的线索和音频混合之间的依赖关系。Mamba的计算可以分为查询、键和值。我们利用线索来生成查询和音频混合来获得键和值,遵循Transformers中的交叉注意机制的原理。两种典型目标声提取方法的实验结果验证了CrossMamba算法的有效性。摘要:The Transformer model, particularly its cross-attention module, is widely used for feature fusion in target sound extraction which extracts the signal of interest based on given clues. Despite its effectiveness, this approach suffers from low computational efficiency. Recent advancements in state space models, notably the latest work Mamba, have shown comparable performance to Transformer-based methods while significantly reducing computational complexity in various tasks. However, Mamba's applicability in target sound extraction is limited due to its inability to capture dependencies between different sequences as the cross-attention does. In this paper, we propose CrossMamba for target sound extraction, which leverages the hidden attention mechanism of Mamba to compute dependencies between the given clues and the audio mixture. The calculation of Mamba can be divided to the query, key and value. We utilize the clue to generate the query and the audio mixture to derive the key and value, adhering to the principle of the cross-attention mechanism in Transformers. Experimental results from two representative target sound extraction methods validate the efficacy of the proposed CrossMamba.
【16】 Vector Quantized Diffusion Model Based Speech Bandwidth Extension
标题: 基于量化扩散模型的语音带宽扩展
作者:Yuan Fang,Jiajie Wang,Xueliang Zhang
备注:4pages
链接:点击下载PDF文件
摘要:神经音频编解码器(NAC)的最新进展开启了音频信号处理的新潜力。越来越多的研究探索利用NAC的潜在功能进行各种语音信号处理任务。本文介绍了第一种方法,语音带宽扩展(BWE),利用从NAC获得的离散功能。通过在高度压缩的离散令牌中恢复高频细节,这种方法增强了语音的可懂度和自然度。该框架基于矢量量化扩散,结合了先进的NAC、扩散模型和Mamba-2的优点,以重建高频语音分量。大量的实验表明,该方法在对数谱距离和ViSQOL方面表现出优异的性能,显着提高了语音质量。摘要:Recent advancements in neural audio codec (NAC) unlock new potential in audio signal processing. Studies have increasingly explored leveraging the latent features of NAC for various speech signal processing tasks. This paper introduces the first approach to speech bandwidth extension (BWE) that utilizes the discrete features obtained from NAC. By restoring high-frequency details within highly compressed discrete tokens, this approach enhances speech intelligibility and naturalness. Based on Vector Quantized Diffusion, the proposed framework combines the strengths of advanced NAC, diffusion models, and Mamba-2 to reconstruct high-frequency speech components. Extensive experiments demonstrate that this method exhibits superior performance across both log-spectral distance and ViSQOL, significantly improving speech quality.
【17】 Audio-Visual Speaker Diarization: Current Databases, Approaches and Challenges
标题: 视听演讲者日记化:当前的数据库、方法和挑战
作者:Victoria Mingote,Alfonso Ortega,Antonio Miguel,Eduardo Lleida
链接:点击下载PDF文件
摘要:如今,大量的视听内容已经促进了开发新的鲁棒的自动说话人日志系统来分析和验证它的需求。这种系统有助于降低手动执行此过程的成本,并且允许将说话人信息用于不同的应用,因为存在大量的信息,例如,面部图像或音频记录。因此,本文旨在解决说话人日志化系统领域的一个关键领域,即不同领域视听内容的整合。本文旨在通过开发一个强大的视听演讲者日记框架来超越当前最先进的实践,该框架适用于各种数据域,包括电视场景,会议和日常活动。与大多数现有的视听扬声器日记系统不同,该框架还将包括一种方法的建议,以引导在名人出现的电视场景中精确分配特定身份。此外,在这项工作中,我们已经进行了广泛的汇编,目前的国家的最先进的方法和现有的数据库开发视听扬声器日记。摘要:Nowadays, the large amount of audio-visual content available has fostered the need to develop new robust automatic speaker diarization systems to analyse and characterise it. This kind of system helps to reduce the cost of doing this process manually and allows the use of the speaker information for different applications, as a huge quantity of information is present, for example, images of faces, or audio recordings. Therefore, this paper aims to address a critical area in the field of speaker diarization systems, the integration of audio-visual content of different domains. This paper seeks to push beyond current state-of-the-art practices by developing a robust audio-visual speaker diarization framework adaptable to various data domains, including TV scenarios, meetings, and daily activities. Unlike most of the existing audio-visual speaker diarization systems, this framework will also include the proposal of an approach to lead the precise assignment of specific identities in TV scenarios where celebrities appear. In addition, in this work, we have conducted an extensive compilation of the current state-of-the-art approaches and the existing databases for developing audio-visual speaker diarization.
【18】 Better Spanish Emotion Recognition In-the-wild: Bringing Attention to Deep Spectrum Voice Analysis
标题: 更好的野外西班牙情感识别:关注深频谱语音分析
作者:Elena Ortega-Beltrán,Josep Cabacas-Maso,Ismael Benito-Altamirano,Carles Ventura
链接:点击下载PDF文件
摘要:在创造新的社会辅助机器人的背景下,情感识别已经成为一个关键的发展因素,因为它允许机器人适应用户在野外的情绪状态。在这项工作中,我们重点分析了两个语音记录西班牙语数据集:ELRA-S 0329和ESPRITHMatchSpanishDB。具体地说,我们的工作集中在语言,e。G.伴随信息并阐明含义的声音特征。我们提出了使用DeepSpectrum方法,该方法包括提取音轨的视觉表示并将其馈送到预训练的CNN模型。对于分类任务,DeepSpectrum通常与支持向量分类器(DS-SVC)或全连接深度学习分类器(DS-FC)配对。我们将DS-SVC和DS-FC架构的结果与ELRA-S 0329和TMS 320 MatchSpanishDB的最新技术(SOTA)进行了比较。此外,我们提出了我们自己的分类器的基础上的注意机制,即DS-AM。我们针对这两个数据集训练了所有模型,我们发现我们的DS-AM模型在数据集和SOTA DeepSpectrum架构上优于SOTA模型。最后,我们在一个数据集中训练了我们的DS-AM模型,并在另一个数据集中对其进行了测试,以模拟真实世界条件下模型对数据集的偏差。摘要:Within the context of creating new Socially Assistive Robots, emotion recognition has become a key development factor, as it allows the robot to adapt to the user's emotional state in the wild. In this work, we focused on the analysis of two voice recording Spanish datasets: ELRA-S0329 and EmoMatchSpanishDB. Specifically, we centered our work in the paralanguage, e.~g. the vocal characteristics that go along with the message and clarifies the meaning. We proposed the use of the DeepSpectrum method, which consists of extracting a visual representation of the audio tracks and feeding them to a pretrained CNN model. For the classification task, DeepSpectrum is often paired with a Support Vector Classifier --DS-SVC--, or a Fully-Connected deep-learning classifier --DS-FC--. We compared the results of the DS-SVC and DS-FC architectures with the state-of-the-art (SOTA) for ELRA-S0329 and EmoMatchSpanishDB. Moreover, we proposed our own classifier based upon Attention Mechanisms, namely DS-AM. We trained all models against both datasets, and we found that our DS-AM model outperforms the SOTA models for the datasets and the SOTA DeepSpectrum architectures. Finally, we trained our DS-AM model in one dataset and tested it in the other, to simulate real-world conditions on how biased is the model to the dataset.
【19】 The first Cadenza challenges: using machine learning competitions to improve music for listeners with a hearing loss
标题: 第一个Cadenza挑战:利用机器学习竞赛为听力丧失的听众改善音乐
作者:Gerardo Roa Dabike,Michael A. Akeroyd,Scott Bannister,Jon P. Barker,Trevor J. Cox,Bruno Fazenda,Jennifer Firth,Simone Graetzer,Alinka Greasley,Rebecca R. Vos,William M. Whitmer
链接:点击下载PDF文件
摘要:众所周知,听音乐是听力损失患者的一个问题,助听器并不是一个通用的解决方案。如何使用机器学习来解决这个问题?本文详细介绍了开放挑战方法的首次应用,即使用机器学习来改善听力损失患者的音乐音频质量。第一个挑战是一场独立比赛(CAD 1),有9名参赛者。第二个是2024年ICASSP大挑战赛(ICASSP 24),吸引了17名参赛者。挑战任务涉及分离和混音流行 摇滚音乐,以允许在混音中个性化地重新平衡乐器,以及放大以校正提高的听力阈值。为参赛者提供的软件基线使用了两种最先进的demix算法:混合Demucs和开放Unmix。使用客观指标HAAQI(助听器音频质量指数)对系统进行评估。没有参赛者在CAD 1的最佳基线上有所改善,因为没有足够的改进空间。因此,对于ICASSP 24,通过使用扬声器再现和在重新混合之前应用的指定增益,使场景变得更加困难。这也使得该场景对于通过助听器收听更有用。9名参赛者得分高于ICASSP 24最佳基线。大多数参赛者使用了改进版的混合Demucs和NAL-R放大。最高评分系统结合了几个解混算法的输出,在一个集成的方法。这些挑战现在是未来研究的开放基准,软件和数据可以免费获得。摘要:It is well established that listening to music is an issue for those with hearing loss, and hearing aids are not a universal solution. How can machine learning be used to address this? This paper details the first application of the open challenge methodology to use machine learning to improve audio quality of music for those with hearing loss. The first challenge was a stand-alone competition (CAD1) and had 9 entrants. The second was an 2024 ICASSP grand challenge (ICASSP24) and attracted 17 entrants. The challenge tasks concerned demixing and remixing pop rock music to allow a personalised rebalancing of the instruments in the mix, along with amplification to correct for raised hearing thresholds. The software baselines provided for entrants to build upon used two state-of-the-art demix algorithms: Hybrid Demucs and Open-Unmix. Evaluation of systems was done using the objective metric HAAQI, the Hearing-Aid Audio Quality Index. No entrants improved on the best baseline in CAD1 because there was insufficient room for improvement. Consequently, for ICASSP24 the scenario was made more difficult by using loudspeaker reproduction and specified gains to be applied before remixing. This also made the scenario more useful for listening through hearing aids. 9 entrants scored better than the the best ICASSP24 baseline. Most entrants used a refined version of Hybrid Demucs and NAL-R amplification. The highest scoring system combined the outputs of several demixing algorithms in an ensemble approach. These challenges are now open benchmarks for future research with the software and data being freely available.
【20】 Deep Generic Representations for Domain-Generalized Anomalous Sound Detection
标题: 用于域广义异常声音检测的深度通用表示
作者:Phurich Saengthong,Takahiro Shinozaki
链接:点击下载PDF文件
摘要:开发可靠的异常声音检测(ASD)系统需要对噪声的鲁棒性、对域偏移的适应性以及在有限训练数据的情况下的有效性能。当前的主要方法依赖于每个目标机器类型的大量标记数据来使用离群暴露(OE)技术来训练特征提取器,然而它们在目标域上的性能仍然是次优的。在本文中,我们提出了 textit{GenRep},它利用了一个强大的,大规模的预训练特征提取器与kNN相结合的通用特征表示,用于域广义ASD,而不需要微调。 textit{GenRep}集成了MemMixup,这是一种使用最近的源样本增强目标内存库的简单方法,并结合域归一化技术来解决源域和目标域之间的不平衡。 textit{GenRep}优于基于OE的最佳方法,无需标记数据,在DCASE 2023 T2 Eval集上的官方评分为73.79 %,并在有限的数据场景下证明了稳健性。该代码是开放源代码。摘要:Developing a reliable anomalous sound detection (ASD) system requires robustness to noise, adaptation to domain shifts, and effective performance with limited training data. Current leading methods rely on extensive labeled data for each target machine type to train feature extractors using Outlier-Exposure (OE) techniques, yet their performance on the target domain remains sub-optimal. In this paper, we present textit{GenRep}, which utilizes generic feature representations from a robust, large-scale pre-trained feature extractor combined with kNN for domain-generalized ASD, without the need for fine-tuning. textit{GenRep} incorporates MemMixup, a simple approach for augmenting the target memory bank using nearest source samples, paired with a domain normalization technique to address the imbalance between source and target domains. textit{GenRep} outperforms the best OE-based approach without a need for labeled data with an Official Score of 73.79 % on the DCASE2023T2 Eval set and demonstrates robustness under limited data scenarios. The code is available open-source.
【21】 Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
标题: 利用声学适应和视觉对齐来改善多模式情绪识别
作者:Zhixian zhao,Haifeng Chen,Xi Li,Dongmei Jiang,Lei Xie
链接:点击下载PDF文件
摘要:多模态情感识别(MER)旨在通过整合来自各种模态的信息来自动识别和理解人类的情感状态。然而,带注释的多模态数据的稀缺性极大地阻碍了这一研究领域的发展。本文介绍了我们针对MER 2024的MER-SEMI子挑战的解决方案。首先,为了更好地适应MER任务的声学模态特征,我们通过实验评估了预训练语音模型HuBERT的不同层在情感识别中的贡献。基于这些观察结果,我们对被确定为对情感识别任务最有效的层进行参数高效微调(PEFT),从而以最少数量的可学习参数实现情感识别的最佳适应。其次,利用声学模态的优势,我们提出了一种特征对齐预训练方法。这种方法使用大规模的未标记的数据来训练视觉编码器,从而促进声学特征空间内的视觉特征的语义对齐。最后,使用适应的声学特征,对齐的视觉特征,和词汇特征,我们采用了一个注意力机制的特征融合。在MER 2024-SEMI测试集上,该方法的加权F1得分为88.90%,在所有参与团队中排名第四,验证了我们方法的有效性。摘要:Multimodal Emotion Recognition (MER) aims to automatically identify and understand human emotional states by integrating information from various modalities. However, the scarcity of annotated multimodal data significantly hinders the advancement of this research field. This paper presents our solution for the MER-SEMI sub-challenge of MER 2024. First, to better adapt acoustic modality features for the MER task, we experimentally evaluate the contributions of different layers of the pre-trained speech model HuBERT in emotion recognition. Based on these observations, we perform Parameter-Efficient Fine-Tuning (PEFT) on the layers identified as most effective for emotion recognition tasks, thereby achieving optimal adaptation for emotion recognition with a minimal number of learnable parameters. Second, leveraging the strengths of the acoustic modality, we propose a feature alignment pre-training method. This approach uses large-scale unlabeled data to train a visual encoder, thereby promoting the semantic alignment of visual features within the acoustic feature space. Finally, using the adapted acoustic features, aligned visual features, and lexical features, we employ an attention mechanism for feature fusion. On the MER2024-SEMI test set, the proposed method achieves a weighted F1 score of 88.90%, ranking fourth among all participating teams, validating the effectiveness of our approach.
【22】 Audio-Guided Fusion Techniques for Multimodal Emotion Analysis
标题: 多模式情绪分析的音频引导融合技术
作者:Pujin Shi,Fei Gao
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个解决方案,半监督学习跟踪(MER-SEMI)在MER 2024。首先,为了增强特征提取器在情感分类任务上的性能,我们使用标记数据微调了视频和文本特征提取器,特别是CLIP-vit-large和Baichuan-13 B。这种方法有效地保留了视频中传达的原始情感信息。其次,我们提出了一个音频引导的Transformer(AGT)融合机制,它利用了Hubert-large的鲁棒性,在融合通道间和通道内信息方面表现出优越的有效性。第三,为了提高模型的准确性,我们通过使用高置信度的未标记数据作为伪标签来迭代地应用自监督学习。最后,通过黑盒探测,我们发现训练集和测试集之间的数据分布不平衡。因此,我们采用了基于先验知识的投票机制。结果证明了我们战略的有效性,最终为我们赢得了MER-SEMI赛道的第三名。摘要:In this paper, we propose a solution for the semi-supervised learning track (MER-SEMI) in MER2024. First, in order to enhance the performance of the feature extractor on sentiment classification tasks,we fine-tuned video and text feature extractors, specifically CLIP-vit-large and Baichuan-13B, using labeled data. This approach effectively preserves the original emotional information conveyed in the videos. Second, we propose an Audio-Guided Transformer (AGT) fusion mechanism, which leverages the robustness of Hubert-large, showing superior effectiveness in fusing both inter-channel and intra-channel information. Third, To enhance the accuracy of the model, we iteratively apply self-supervised learning by using high-confidence unlabeled data as pseudo-labels. Finally, through black-box probing, we discovered an imbalanced data distribution between the training and test sets. Therefore, We adopt a prior-knowledge-based voting mechanism. The results demonstrate the effectiveness of our strategy, ultimately earning us third place in the MER-SEMI track.
【23】 Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
标题: 用预训练模型解开韵律和语义信息,用于基于上下文学习的Zero-Shot语音转换
作者:Zhengyang Chen,Shuai Wang,Mingyang Zhang,Xuechen Liu,Junichi Yamagishi,Yanmin Qian
链接:点击下载PDF文件
摘要:语音转换(VC)的目的是修改扬声器的音色,同时保留语音内容。先前的方法已经将自监督的输出标记为语义标记,从而促进语音内容信息的解开。最近,在上下文学习(ICL)已经出现在文本到语音(TTS)系统,有效地建模特定的特征,如通过上下文条件反射的音色。本文提出了一种ICL能力增强的VC系统(ICL-VC),采用基于流匹配生成模型的掩码和重建训练策略。我们在LibriTTS数据集上的实验表明,ICL-VC增加了语义标记,提高了说话人相似性。此外,我们发现k-means是一种通用的标记化方法,适用于各种预训练模型。然而,ICL-VC系统在保持源语音的韵律方面面临挑战。为了缓解这个问题,我们建议将从预训练的情感识别模型中提取的韵律嵌入到我们的系统中。韵律嵌入的集成显着提高了系统的能力,以保持源语音韵律,情感语音数据库上验证。摘要:Voice conversion (VC) aims to modify the speaker's timbre while retaining speech content. Previous approaches have tokenized the outputs from self-supervised into semantic tokens, facilitating disentanglement of speech content information. Recently, in-context learning (ICL) has emerged in text-to-speech (TTS) systems for effectively modeling specific characteristics such as timbre through context conditioning. This paper proposes an ICL capability enhanced VC system (ICL-VC) employing a mask and reconstruction training strategy based on flow-matching generative models. Augmented with semantic tokens, our experiments on the LibriTTS dataset demonstrate that ICL-VC improves speaker similarity. Additionally, we find that k-means is a versatile tokenization method applicable to various pre-trained models. However, the ICL-VC system faces challenges in preserving the prosody of the source speech. To mitigate this issue, we propose incorporating prosody embeddings extracted from a pre-trained emotion recognition model into our system. Integration of prosody embeddings notably enhances the system's capability to preserve source speech prosody, as validated on the Emotional Speech Database.
【24】 Attention-Based Efficient Breath Sound Removal in Studio Audio Recordings
标题: 录音室录音中基于注意力的高效呼吸声去除
作者:Nidula Elgiriyewithana,N. D. Kodikara
Journal-ref:CS & IT Conference Proceedings, vol. 14, no. 6, 2024
链接:点击下载PDF文件
摘要:在这项研究中,我们提出了一个创新的,参数高效的模型,利用注意力U-Net架构的自动检测和消除非语音声乐的声音,特别是呼吸声,在声乐录音。这项任务在声音工程领域是至关重要的,尽管相对来说还没有得到充分的探索。检测和消除这些声音的传统手动过程需要大量的专业知识,并且非常耗时。现有的自动检测和移除方法在效率和精度方面往往不足。我们提出的模型通过应用先进的深度学习技术,提供简化的流程和卓越的准确性,解决了这些限制。一个独特的数据集,来自设备和生产语音(DAPS),用于此目的。模型的训练阶段强调对数谱图,并集成了早期停止机制以防止过拟合。我们的模式不仅为音响工程师节省了宝贵的时间,而且还提高了音频制作的质量和一致性。这是一个重大突破,其相对效率证明了这一点,仅需要190万个参数和3.2小时的训练时间-明显低于该领域的顶级模型。该模型能够生成与先前模型相同的输出,并大幅提高精度,使其成为最佳选择。摘要:In this research, we present an innovative, parameter-efficient model that utilizes the attention U-Net architecture for the automatic detection and eradication of non-speech vocal sounds, specifically breath sounds, in vocal recordings. This task is of paramount importance in the field of sound engineering, despite being relatively under-explored. The conventional manual process for detecting and eliminating these sounds requires significant expertise and is extremely time-intensive. Existing automated detection and removal methods often fall short in terms of efficiency and precision. Our proposed model addresses these limitations by offering a streamlined process and superior accuracy, achieved through the application of advanced deep learning techniques. A unique dataset, derived from Device and Produced Speech (DAPS), was employed for this purpose. The training phase of the model emphasizes a log spectrogram and integrates an early stopping mechanism to prevent overfitting. Our model not only conserves precious time for sound engineers but also enhances the quality and consistency of audio production. This constitutes a significant breakthrough, as evidenced by its comparative efficiency, necessitating only 1.9M parameters and a training duration of 3.2 hours - markedly less than the top-performing models in this domain. The model is capable of generating identical outputs as previous models with drastically improved precision, making it an optimal choice.
【25】 Just ASR + LLM? A Study on Speech Large Language Models' Ability to Identify and Understand Speaker in Spoken Dialogue
标题: 只是ASB + LLM吗?言语大语言模型识别和理解口语对话中说话人的能力研究
作者:Junkai Wu,Xulin Fan,Bo-Ru Lu,Xilin Jiang,Nima Mesgarani,Mark Hasegawa-Johnson,Mari Ostendorf
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:近年来,我们观察到语音语言模型(SpeechLLM)的快速发展,赶上了人类的听力和推理能力。值得注意的是,SpeechLLM在中国高考英语听力测试中表现出令人印象深刻的口语对话问答(SQA)表现,这似乎需要理解对话中说话者的口语内容和语音特征。然而,在仔细研究高考试题后,我们发现许多问题的正确答案可以仅从对话上下文推断出来,而无需识别问题中的说话者。我们在高考中对最先进的Qwen-Audio和WavLLM模型的评估以及我们提出的“你喜欢什么?“数据集显示,这些基于上下文的问题的准确率明显高于身份关键问题,后者只能通过正确的说话人识别来正确回答。我们的研究结果和分析表明,在解决SQA时,当前的SpeechLLM从音频中表现出有限的扬声器意识,并且表现出与没有声音的对话转录的LLM推理相似的行为。我们建议,我们的定义和基于上下文和身份关键问题的自动分类可以提供一个更准确的评估框架的SpeechLLM在SQA任务。摘要:In recent years, we have observed a rapid advancement in speech language models (SpeechLLMs), catching up with humans' listening and reasoning abilities. Remarkably, SpeechLLMs have demonstrated impressive spoken dialogue question-answering (SQA) performance in benchmarks like Gaokao, the English listening test of the college entrance exam in China, which seemingly requires understanding both the spoken content and voice characteristics of speakers in a conversation. However, after carefully examining Gaokao's questions, we find the correct answers to many questions can be inferred from the conversation context alone without identifying the speaker asked in the question. Our evaluation of state-of-the-art models Qwen-Audio and WavLLM in both Gaokao and our proposed "What Do You Like?" dataset shows a significantly higher accuracy in these context-based questions than in identity-critical questions, which can only be answered correctly with correct speaker identification. Our results and analysis suggest that when solving SQA, the current SpeechLLMs exhibit limited speaker awareness from the audio and behave similarly to an LLM reasoning from the conversation transcription without sound. We propose that our definitions and automated classification of context-based and identity-critical questions could offer a more accurate evaluation framework of SpeechLLMs in SQA tasks.
【26】 Flow-TSVAD: Target-Speaker Voice Activity Detection via Latent Flow Matching
标题: Flow-TSVAD:通过潜在流匹配进行目标说话者语音活动检测
作者:Zhengyang Chen,Bing Han,Shuai Wang,Yidi Jiang,Yanmin Qian
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:说话人日记化通常被认为是一个判别任务,使用判别方法来产生固定的日记化结果。在本文中,我们首次探索使用基于神经网络的生成方法进行说话人日记化。我们实现了一个基于流匹配(FM)的生成算法内的序列到序列的目标说话人语音活动检测(Seq 2Seq-TSVAD)日记系统。我们的实验表明,直接将生成方法应用于TS-VAD输出的原始二进制标签序列空间是无效的。为了解决这个问题,我们建议在应用生成算法之前将二进制标签序列映射到密集的潜在空间中,并且我们提出的Flow-TSVAD方法优于Seq 2Seq-TSVAD系统。此外,我们观察到,FM算法收敛迅速,在推理阶段,只需要两个推理步骤,以实现有希望的结果。作为一个生成模型,Flow-TSVAD允许通过多次运行模型来采样不同的日志结果。此外,来自各种采样实例的集成结果进一步增强了日志化性能。摘要:Speaker diarization is typically considered a discriminative task, using discriminative approaches to produce fixed diarization results. In this paper, we explore the use of neural network-based generative methods for speaker diarization for the first time. We implement a Flow-Matching (FM) based generative algorithm within the sequence-to-sequence target speaker voice activity detection (Seq2Seq-TSVAD) diarization system. Our experiments reveal that applying the generative method directly to the original binary label sequence space of the TS-VAD output is ineffective. To address this issue, we propose mapping the binary label sequence into a dense latent space before applying the generative algorithm and our proposed Flow-TSVAD method outperforms the Seq2Seq-TSVAD system. Additionally, we observe that the FM algorithm converges rapidly during the inference stage, requiring only two inference steps to achieve promising results. As a generative model, Flow-TSVAD allows for sampling different diarization results by running the model multiple times. Moreover, ensembling results from various sampling instances further enhances diarization performance.
【27】 PB-LRDWWS System for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge
标题: PB-LRDDWWS系统,用于SYS 2024低资源发音障碍唤醒单词发现挑战赛
作者:Shiyao Wang,Jiaming Zhou,Shiwan Zhao,Yong Qin
备注:accept by SLT 2024
链接:点击下载PDF文件
摘要:对于2024年的低资源构音障碍唤醒词发现(LRDWWS)挑战赛,我们介绍了PB-LRDWWS系统。该系统结合了构音障碍的语音内容特征提取器的原型建设与基于原型的分类方法。特征提取器是一个微调的HuBERT模型,通过使用交叉熵损失的三阶段微调过程获得。这个经过微调的HuBERT从目标构音障碍演讲者的注册演讲中提取特征来构建原型。通过计算构音障碍说话人评价语音的HuBERT特征与原型之间的余弦相似度来实现分类。尽管它的简单性,我们的方法通过实验结果证明了有效性。我们的系统在LRDWWS挑战赛的最终测试B中获得第二名。摘要:For the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting (LRDWWS) Challenge, we introduce the PB-LRDWWS system. This system combines a dysarthric speech content feature extractor for prototype construction with a prototype-based classification method. The feature extractor is a fine-tuned HuBERT model obtained through a three-stage fine-tuning process using cross-entropy loss. This fine-tuned HuBERT extracts features from the target dysarthric speaker's enrollment speech to build prototypes. Classification is achieved by calculating the cosine similarity between the HuBERT features of the target dysarthric speaker's evaluation speech and prototypes. Despite its simplicity, our method demonstrates effectiveness through experimental results. Our system achieves second place in the final Test-B of the LRDWWS Challenge.
【28】 Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
标题: Mel-Roformer用于人声分离和人声旋律转录
作者:Ju-Chiang Wang,Wei-Tsung Lu,Jitong Chen
备注:Accepted to appear in ISMIR 2024
链接:点击下载PDF文件
摘要:开发一个多功能的深度神经网络来对音乐音频进行建模在MIR中至关重要。由于音乐信号中固有的复杂频谱变化,这项任务是具有挑战性的,这些音乐信号传达了不同乐器的旋律,和声和音色。在本文中,我们介绍了梅尔RoFormer,一个基于频谱图的模型具有两个关键的设计:一个新的梅尔波段投影模块在前端,以提高模型的能力,捕捉信息功能在多个频带,和交错的RoPE Transformers明确建模的频率和时间维度作为两个单独的序列。我们应用Mel-RoFormer来解决两个基本的MIR任务:声乐分离和声乐旋律转录,旨在将歌声从音频混合中分离出来,并分别转录其主旋律。尽管它们都关注于信号,但这些任务具有不同的优化目标。我们没有训练一个统一的模型,而是采用了两步走的方法。首先,我们训练一个声乐分离模型,随后作为声乐旋律转录微调的基础模型。通过在基准数据集上进行的大量实验,我们展示了我们的模型在声乐分离和旋律转录任务中实现了最先进的性能,强调了Mel-RoFormer在模拟复杂音乐音频信号方面的有效性和多功能性。摘要:Developing a versatile deep neural network to model music audio is crucial in MIR. This task is challenging due to the intricate spectral variations inherent in music signals, which convey melody, harmonics, and timbres of diverse instruments. In this paper, we introduce Mel-RoFormer, a spectrogram-based model featuring two key designs: a novel Mel-band Projection module at the front-end to enhance the model's capability to capture informative features across multiple frequency bands, and interleaved RoPE Transformers to explicitly model the frequency and time dimensions as two separate sequences. We apply Mel-RoFormer to tackle two essential MIR tasks: vocal separation and vocal melody transcription, aimed at isolating singing voices from audio mixtures and transcribing their lead melodies, respectively. Despite their shared focus on singing signals, these tasks possess distinct optimization objectives. Instead of training a unified model, we adopt a two-step approach. Initially, we train a vocal separation model, which subsequently serves as a foundation model for fine-tuning for vocal melody transcription. Through extensive experiments conducted on benchmark datasets, we showcase that our models achieve state-of-the-art performance in both vocal separation and melody transcription tasks, underscoring the efficacy and versatility of Mel-RoFormer in modeling complex music audio signals.
【29】 Leveraging Contrastive Learning and Self-Training for Multimodal Emotion Recognition with Limited Labeled Samples
标题: 利用对比学习和自我训练以有限的标记样本进行多模式情感识别
作者:Qi Fan,Yutong Li,Yi Xin,Xinyu Cheng,Guanglai Gao,Miao Ma
备注:Accepted by ACM MM Workshop 2024
链接:点击下载PDF文件
摘要:多模态情感识别挑战MER 2024专注于使用音频,语言和视觉信号识别情感。在本文中,我们提出了半监督学习子挑战赛(MER 2024-SEMI)的提交解决方案,该挑战赛解决了情感识别中有限注释数据的问题。首先,为了解决类的不平衡,我们采用了过采样策略。其次,我们提出了一个模态表示组合对比学习(MR-CCL)框架上的三模态输入数据建立鲁棒的初始模型。第三,我们探索了一种自我训练的方法来扩展训练集。最后,我们通过多分类器加权软投票策略增强预测鲁棒性。我们提出的方法在MER 2024-SEMI挑战赛上被验证是有效的,实现了88.25%的加权平均F分数,在排行榜上排名第6。我们的项目可在https: github.com WooyoohL MER2024-SEMI上获得。摘要:The Multimodal Emotion Recognition challenge MER2024 focuses on recognizing emotions using audio, language, and visual signals. In this paper, we present our submission solutions for the Semi-Supervised Learning Sub-Challenge (MER2024-SEMI), which tackles the issue of limited annotated data in emotion recognition. Firstly, to address the class imbalance, we adopt an oversampling strategy. Secondly, we propose a modality representation combinatorial contrastive learning (MR-CCL) framework on the trimodal input data to establish robust initial models. Thirdly, we explore a self-training approach to expand the training set. Finally, we enhance prediction robustness through a multi-classifier weighted soft voting strategy. Our proposed method is validated to be effective on the MER2024-SEMI Challenge, achieving a weighted average F-score of 88.25% and ranking 6th on the leaderboard. Our project is available at https: github.com WooyoohL MER2024-SEMI.
机器翻译,仅供参考
