本文经arXiv每日学术速递授权转载
【1】 A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
标题: Deepfake语音检测的全面调查和批判性分析
作者:Lam Pham,Phat Lam,Tin Nguyen,Hieu Tang,Huyen Nguyen,Alexander Schindler,Canh Vu
备注:Journal preprint
链接:点击下载PDF文件
【2】 Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
标题: 通过负选择策略的自适应学习用于Few-Shot生物声学事件检测
作者:Yaxiong Chen,Xueping Zhang,Yunfei Zi,Shengwu Xiong
链接:点击下载PDF文件
【3】 LoVA: Long-form Video-to-Audio Generation
标题: LoVA:长格式视频到音频生成
作者:Xin Cheng,Xihua Wang,Yihan Wu,Yuyue Wang,Ruihua Song
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【4】 GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
标题: GALD-SE:引导各向异性轻量级扩散,实现高效语音增强
作者:Chengzhong Wang,Jianjun Gu,Dingding Yao,Zelin Qiu,Jiale Zhao,Junfeng Li
备注:5 pages, 2 figures
链接:点击下载PDF文件
【5】 Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
标题: 根据单独的房间和场景特定信息生成盲空间脉冲响应
作者:Francesc Lluís,Nils Meyer-Kahlen
链接:点击下载PDF文件
【6】 Voice Conversion-based Privacy through Adversarial Information Hiding
标题: 通过对抗性信息隐藏实现基于语音转换的隐私
作者:Jacob J Webber,Oliver Watts,Gustav Eje Henter,Jennifer Williams,Simon King
备注:Accepted for publication in proceedings of 4th symposium on security and privacy in speech communication
链接:点击下载PDF文件
【7】 HiFi-Glot: Neural Formant Synthesis with Differentiable Resonant Filters
标题: HiFi-Glot:使用可区分共振过滤器的神经Forces合成
作者:Lauri Juvela,Pablo Pérez Zarazaga,Gustav Eje Henter,Zofia Malisz
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 SongTrans: An unified song transcription and alignment method for lyrics and notes
标题: SongTrans:歌词和音符的统一歌曲转录和对齐方法
作者:Siwei Wu,Jinzheng He,Ruibin Yuan,Haojie Wei,Xipin Wei,Chenghua Lin,Jin Xu,Junyang Lin
链接:点击下载PDF文件
【9】 What Are They Doing? Joint Audio-Speech Co-Reasoning
标题: 他们在做什么?音频-语音联合推理
作者:Yingzhi Wang,Pooneh Mousavi,Artem Ploujnikov,Mirco Ravanelli
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【10】 CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
标题: CPD增强Wav2vec2.0:面向课堂环境的噪音稳健语音识别
作者:Ahmed Adel Attia,Dorottya Demszky,Tolulope Ogunremi,Jing Liu,Carol Espy-Wilson
备注:arXiv admin note: substantial text overlap with arXiv:2405.13018
链接:点击下载PDF文件
【11】 Self-Supervised Audio-Visual Soundscape Stylization
标题: 自我监督的视听声景风格化
作者:Tingle Li,Renhao Wang,Po-Yao Huang,Andrew Owens,Gopala Anumanchipalli
备注:ECCV 2024
链接:点击下载PDF文件
【12】 AMT-APC: Automatic Piano Cover by Fine-Tuning an Automatic Music Transcription Model
标题: AMT-IPC:通过微调自动音乐转录模型实现自动钢琴翻唱
作者:Kazuma Komiya,Yoshihisa Fukuhara
链接:点击下载PDF文件
【13】 MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
标题: MultiMed:通过注意力编码器解码器的多语言医学语音识别
作者:Khai Le-Duc,Phuc Phan,Tan-Hanh Pham,Bach Phan Tat,Minh-Huong Ngo,Truong-Son Hy
备注:Preprint
链接:点击下载PDF文件
【14】 ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
标题: ECHO:采用分层实体指导的半监督学习的环境声音分类
作者:Pranav Gupta,Raunak Sharma,Rashmi Kumari,Sri Krishna Aditya,Shwetank Choudhary,Sumit Kumar,Kanchana M,Thilagavathy R
备注:IEEE CONECCT 2024, Signal Processing and Pattern Recognition, Environmental Sound Classification, ESC
链接:点击下载PDF文件
【15】 Training Large ASR Encoders with Differential Privacy
标题: 训练具有差异隐私的大型ASB编码器
作者:Geeticka Chauhan,Steve Chien,Om Thakkar,Abhradeep Thakurta,Arun Narayanan
备注:In proceedings of the IEEE Spoken Language Technologies Workshop, 2024
链接:点击下载PDF文件
【16】 Target word activity detector: An approach to obtain ASR word boundaries without lexicon
标题: 目标词活动检测器:一种无需词典即可获得ASB词边界的方法
作者:Sunit Sivasankaran,Eric Sun,Jinyu Li,Yan Huang,Jing Pan
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【17】 PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
标题: PTQ4ADM:高效文本条件音频扩散模型的训练后量化
作者:Jayneel Vora,Aditya Krishnan,Nader Bouacida,Prabhu RV Shankar,Prasant Mohapatra
链接:点击下载PDF文件
【18】 Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
标题: 基于柱状图层延时神经网络的时频特征组合研究
作者:Amirmohammad Mohammadi,Iren'e Masabarakiza,Ethan Barnes,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 14 figures. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【19】 Transfer Learning for Passive Sonar Classification using Pre-trained Audio and ImageNet Models
标题: 使用预训练的音频和ImageNet模型进行被动声纳分类的迁移学习
作者:Amirmohammad Mohammadi,Tejashri Kelhe,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 6 figures, This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【20】 A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
标题: 随机信封波动对噪音音素感知影响的微观研究
作者:Alejandro Osses,Léo Varnet
Journal-ref:Journal of the Acoustical Society of America, 2024, 155 (2), pp.1469-1485
链接:点击下载PDF文件
【21】 Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
标题: 优化歌曲创作过程:使用深度学习模型基于流派的歌词生成
作者:Tracy Cai,Wilson Liang,Donte Townes
链接:点击下载PDF文件
【22】 Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
标题: 通过原生数据库训练增强库尔德语文本到语音:高质量WaveGlow声码器方法
作者:Abdulhady Abas Abdullah,Sabat Salih Muhamad,Hadi Veisi
链接:点击下载PDF文件
【23】 Lightweight Transducer Based on Frame-Level Criterion
标题: 基于框架级标准的轻型传感器
作者:Genshun Wan,Mengzhi Wang,Tingzhi Mao,Hang Chen,Zhongfu Ye
备注:Accepted by Interspeech 2024, code repository: this https URL
链接:点击下载PDF文件
【24】 CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
标题: CA-MHFA:用于基于SSL的说话人验证的上下文感知多头因子化注意力池
作者:Junyi Peng,Ladislav Mošner,Lin Zhang,Oldřich Plchot,Themos Stafylakis,Lukáš Burget,Jan Černocký
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【25】 LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
标题: LlamaPartialSpoof:一个LLM驱动的模拟虚假信息生成的假语音数据集
作者:Hieu-Thi Luong,Haoyang Li,Lin Zhang,Kong Aik Lee,Eng Siong Chng
备注:5 pages, submitted to ICASSP 2025
链接:点击下载PDF文件
【26】 Room Impulse Responses help attackers to evade Deep Fake Detection
标题: 房间冲动响应帮助攻击者逃避深度伪造检测
作者:Hieu-Thi Luong,Duc-Tuan Truong,Kong Aik Lee,Eng Siong Chng
备注:7 pages, to be presented at SLT 2024
链接:点击下载PDF文件
【27】 Video-to-Audio Generation with Fine-grained Temporal Semantics
标题: 具有细粒度时间语义的视频到音频生成
作者:Yuchen Hu,Yu Gu,Chenxing Li,Rilin Chen,Dong Yu
链接:点击下载PDF文件
【28】 Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
标题: 稳健的视听语音增强:利用高级后处理纠正复杂环境中的误分配
作者:Wenze Ren,Kuo-Hsuan Hung,Chao Rong,YouJin Li,Hsin-Min Wang,Tsao Yu
链接:点击下载PDF文件
【29】 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
标题: 无监督单词发现:使用集群进行边界检测与动态编程
作者:Simon Malan,Benjamin van Niekerk,Herman Kamper
备注:3 figures, 3 tables
链接:点击下载PDF文件
【30】 A Feature Engineering Approach for Literary and Colloquial Tamil Speech Classification using 1D-CNN
标题: 使用1D-CNN进行文学和口语泰米尔语语音分类的特征工程方法
作者:M. Nanmalar,S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
链接:点击下载PDF文件
【31】 Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
标题: 通过可靠性加权,使用可穿戴麦克风阵列改善动态环境的到达方向估计
作者:Daniel A. Mitchell,Boaz Rafaely,Anurag Kumar,Vladimir Tourbabin
链接:点击下载PDF文件
【32】 Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
标题: 复仇者联盟集结:抑郁症检测的非语义特征融合
作者:Orchid Chetia Phukan,Swarup Ranjan Behera,Shubham Singh,Muskaan Singh,Vandana Rajan,Arun Balaji Buduru,Rajesh Sharma,S. R. Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【33】 Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
标题: 单独坚强,共同坚强:协同情态约束基础模型与非言语情感识别的最佳传输
作者:Orchid Chetia Phukan,Mohd Mujtaba Akhtar,Girish,Swarup Ranjan Behera,Sishir Kalita,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【34】 Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
标题: 音乐基金会模型在演唱声音Deepfake检测方面更好吗? 更好地使用Speech Foundation模型来简化它们
作者:Orchid Chetia Phukan,Sarthak Jain,Swarup Ranjan Behera,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【35】 Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
标题: Codec-SURB @ SYS 2024:神经音频编解码器模型的轻量级基准
作者:Haibin Wu,Xuanjun Chen,Yi-Cheng Lin,Kaiwei Chang,Jiawei Du,Ke-Han Lu,Alexander H. Liu,Ho-Lam Chung,Yuan-Kuei Wu,Dongchao Yang,Songxiang Liu,Yi-Chiao Wu,Xu Tan,James Glass,Shinji Watanabe,Hung-yi Lee
链接:点击下载PDF文件
【36】 Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
标题: 半侵入式音频评估:将非侵入式评估视为多模式文本预测任务
作者:Jozef Coldenhoff,Milos Cernak
链接:点击下载PDF文件
【37】 Zero-shot Cross-lingual Voice Transfer for TTS
标题: 针对TTC的零攻击跨语言语音传输
作者:Fadi Biadsy,Youzheng Chen,Isaac Elias,Kyle Kastner,Gary Wang,Andrew Rosenberg,Bhuvana Ramabhadran
备注:Submitted to ICASSP
链接:点击下载PDF文件
【38】 GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
标题: GT Singer:全球多技术歌唱数据库,为所有歌唱任务提供现实音乐分数
作者:Yu Zhang,Changhao Pan,Wenxiang Guo,Ruiqi Li,Zhiyuan Zhu,Jialei Wang,Wenhao Xu,Jingyu Lu,Zhiqing Hong,Chuxin Wang,LiChao Zhang,Jinzheng He,Ziyue Jiang,Yuxin Chen,Chen Yang,Jiecheng Zhou,Xinyu Cheng,Zhou Zhao
备注:under processing
链接:点击下载PDF文件
标题: CA-MHFA:用于基于SSL的说话人验证的上下文感知多头因子化注意力池
作者:Junyi Peng,Ladislav Mošner,Lin Zhang,Oldřich Plchot,Themos Stafylakis,Lukáš Burget,Jan Černocký
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
标题: LlamaPartialSpoof:一个LLM驱动的模拟虚假信息生成的假语音数据集
作者:Hieu-Thi Luong,Haoyang Li,Lin Zhang,Kong Aik Lee,Eng Siong Chng
备注:5 pages, submitted to ICASSP 2025
链接:点击下载PDF文件
【3】 Room Impulse Responses help attackers to evade Deep Fake Detection
标题: 房间冲动响应帮助攻击者逃避深度伪造检测
作者:Hieu-Thi Luong,Duc-Tuan Truong,Kong Aik Lee,Eng Siong Chng
备注:7 pages, to be presented at SLT 2024
链接:点击下载PDF文件
【4】 Video-to-Audio Generation with Fine-grained Temporal Semantics
标题: 具有细粒度时间语义的视频到音频生成
作者:Yuchen Hu,Yu Gu,Chenxing Li,Rilin Chen,Dong Yu
链接:点击下载PDF文件
【5】 Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
标题: 稳健的视听语音增强:利用高级后处理纠正复杂环境中的误分配
作者:Wenze Ren,Kuo-Hsuan Hung,Chao Rong,YouJin Li,Hsin-Min Wang,Tsao Yu
链接:点击下载PDF文件
【6】 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
标题: 无监督单词发现:使用集群进行边界检测与动态编程
作者:Simon Malan,Benjamin van Niekerk,Herman Kamper
备注:3 figures, 3 tables
链接:点击下载PDF文件
【7】 A Feature Engineering Approach for Literary and Colloquial Tamil Speech Classification using 1D-CNN
标题: 使用1D-CNN进行文学和口语泰米尔语语音分类的特征工程方法
作者:M. Nanmalar,S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
链接:点击下载PDF文件
【8】 Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
标题: 通过可靠性加权,使用可穿戴麦克风阵列改善动态环境的到达方向估计
作者:Daniel A. Mitchell,Boaz Rafaely,Anurag Kumar,Vladimir Tourbabin
链接:点击下载PDF文件
【9】 Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
标题: 复仇者联盟集结:抑郁症检测的非语义特征融合
作者:Orchid Chetia Phukan,Swarup Ranjan Behera,Shubham Singh,Muskaan Singh,Vandana Rajan,Arun Balaji Buduru,Rajesh Sharma,S. R. Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【10】 Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
标题: 单独坚强,共同坚强:协同情态约束基础模型与非言语情感识别的最佳传输
作者:Orchid Chetia Phukan,Mohd Mujtaba Akhtar,Girish,Swarup Ranjan Behera,Sishir Kalita,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【11】 Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
标题: 音乐基金会模型在演唱声音Deepfake检测方面更好吗? 更好地使用Speech Foundation模型来简化它们
作者:Orchid Chetia Phukan,Sarthak Jain,Swarup Ranjan Behera,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【12】 Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
标题: Codec-SURB @ SYS 2024:神经音频编解码器模型的轻量级基准
作者:Haibin Wu,Xuanjun Chen,Yi-Cheng Lin,Kaiwei Chang,Jiawei Du,Ke-Han Lu,Alexander H. Liu,Ho-Lam Chung,Yuan-Kuei Wu,Dongchao Yang,Songxiang Liu,Yi-Chiao Wu,Xu Tan,James Glass,Shinji Watanabe,Hung-yi Lee
链接:点击下载PDF文件
【13】 Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
标题: 半侵入式音频评估:将非侵入式评估视为多模式文本预测任务
作者:Jozef Coldenhoff,Milos Cernak
链接:点击下载PDF文件
【14】 Zero-shot Cross-lingual Voice Transfer for TTS
标题: 针对TTC的零攻击跨语言语音传输
作者:Fadi Biadsy,Youzheng Chen,Isaac Elias,Kyle Kastner,Gary Wang,Andrew Rosenberg,Bhuvana Ramabhadran
备注:Submitted to ICASSP
链接:点击下载PDF文件
【15】 GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
标题: GT Singer:全球多技术歌唱数据库,为所有歌唱任务提供现实音乐分数
作者:Yu Zhang,Changhao Pan,Wenxiang Guo,Ruiqi Li,Zhiyuan Zhu,Jialei Wang,Wenhao Xu,Jingyu Lu,Zhiqing Hong,Chuxin Wang,LiChao Zhang,Jinzheng He,Ziyue Jiang,Yuxin Chen,Chen Yang,Jiecheng Zhou,Xinyu Cheng,Zhou Zhao
备注:under processing
链接:点击下载PDF文件
【16】 A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
标题: Deepfake语音检测的全面调查和批判性分析
作者:Lam Pham,Phat Lam,Tin Nguyen,Hieu Tang,Huyen Nguyen,Alexander Schindler,Canh Vu
备注:Journal preprint
链接:点击下载PDF文件
【17】 Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
标题: 通过负选择策略的自适应学习用于Few-Shot生物声学事件检测
作者:Yaxiong Chen,Xueping Zhang,Yunfei Zi,Shengwu Xiong
链接:点击下载PDF文件
【18】 LoVA: Long-form Video-to-Audio Generation
标题: LoVA:长格式视频到音频生成
作者:Xin Cheng,Xihua Wang,Yihan Wu,Yuyue Wang,Ruihua Song
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【19】 GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
标题: GALD-SE:引导各向异性轻量级扩散,实现高效语音增强
作者:Chengzhong Wang,Jianjun Gu,Dingding Yao,Zelin Qiu,Jiale Zhao,Junfeng Li
备注:5 pages, 2 figures
链接:点击下载PDF文件
【20】 Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
标题: 根据单独的房间和场景特定信息生成盲空间脉冲响应
作者:Francesc Lluís,Nils Meyer-Kahlen
链接:点击下载PDF文件
【21】 Voice Conversion-based Privacy through Adversarial Information Hiding
标题: 通过对抗性信息隐藏实现基于语音转换的隐私
作者:Jacob J Webber,Oliver Watts,Gustav Eje Henter,Jennifer Williams,Simon King
备注:Accepted for publication in proceedings of 4th symposium on security and privacy in speech communication
链接:点击下载PDF文件
【22】 HiFi-Glot: Neural Formant Synthesis with Differentiable Resonant Filters
标题: HiFi-Glot:使用可区分共振过滤器的神经Forces合成
作者:Lauri Juvela,Pablo Pérez Zarazaga,Gustav Eje Henter,Zofia Malisz
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【23】 SongTrans: An unified song transcription and alignment method for lyrics and notes
标题: SongTrans:歌词和音符的统一歌曲转录和对齐方法
作者:Siwei Wu,Jinzheng He,Ruibin Yuan,Haojie Wei,Xipin Wei,Chenghua Lin,Jin Xu,Junyang Lin
链接:点击下载PDF文件
【24】 What Are They Doing? Joint Audio-Speech Co-Reasoning
标题: 他们在做什么?音频-语音联合推理
作者:Yingzhi Wang,Pooneh Mousavi,Artem Ploujnikov,Mirco Ravanelli
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【25】 CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
标题: CPD增强Wav2vec2.0:面向课堂环境的噪音稳健语音识别
作者:Ahmed Adel Attia,Dorottya Demszky,Tolulope Ogunremi,Jing Liu,Carol Espy-Wilson
备注:arXiv admin note: substantial text overlap with arXiv:2405.13018
链接:点击下载PDF文件
【26】 Self-Supervised Audio-Visual Soundscape Stylization
标题: 自我监督的视听声景风格化
作者:Tingle Li,Renhao Wang,Po-Yao Huang,Andrew Owens,Gopala Anumanchipalli
备注:ECCV 2024
链接:点击下载PDF文件
【27】 AMT-APC: Automatic Piano Cover by Fine-Tuning an Automatic Music Transcription Model
标题: AMT-IPC:通过微调自动音乐转录模型实现自动钢琴翻唱
作者:Kazuma Komiya,Yoshihisa Fukuhara
链接:点击下载PDF文件
【28】 MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
标题: MultiMed:通过注意力编码器解码器的多语言医学语音识别
作者:Khai Le-Duc,Phuc Phan,Tan-Hanh Pham,Bach Phan Tat,Minh-Huong Ngo,Truong-Son Hy
备注:Preprint
链接:点击下载PDF文件
【29】 ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
标题: ECHO:采用分层实体指导的半监督学习的环境声音分类
作者:Pranav Gupta,Raunak Sharma,Rashmi Kumari,Sri Krishna Aditya,Shwetank Choudhary,Sumit Kumar,Kanchana M,Thilagavathy R
备注:IEEE CONECCT 2024, Signal Processing and Pattern Recognition, Environmental Sound Classification, ESC
链接:点击下载PDF文件
【30】 Training Large ASR Encoders with Differential Privacy
标题: 训练具有差异隐私的大型ASB编码器
作者:Geeticka Chauhan,Steve Chien,Om Thakkar,Abhradeep Thakurta,Arun Narayanan
备注:In proceedings of the IEEE Spoken Language Technologies Workshop, 2024
链接:点击下载PDF文件
【31】 Target word activity detector: An approach to obtain ASR word boundaries without lexicon
标题: 目标词活动检测器:一种无需词典即可获得ASB词边界的方法
作者:Sunit Sivasankaran,Eric Sun,Jinyu Li,Yan Huang,Jing Pan
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【32】 PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
标题: PTQ4ADM:高效文本条件音频扩散模型的训练后量化
作者:Jayneel Vora,Aditya Krishnan,Nader Bouacida,Prabhu RV Shankar,Prasant Mohapatra
链接:点击下载PDF文件
【33】 Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
标题: 基于柱状图层延时神经网络的时频特征组合研究
作者:Amirmohammad Mohammadi,Iren'e Masabarakiza,Ethan Barnes,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 14 figures. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【34】 Transfer Learning for Passive Sonar Classification using Pre-trained Audio and ImageNet Models
标题: 使用预训练的音频和ImageNet模型进行被动声纳分类的迁移学习
作者:Amirmohammad Mohammadi,Tejashri Kelhe,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 6 figures, This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【35】 On the Feasibility of Fully AI-automated Vishing Attacks
标题: 全人工智能自动Vising攻击的可行性
作者:João Figueiredo,Afonso Carvalho,Daniel Castro,Daniel Gonçalves,Nuno Santos
链接:点击下载PDF文件
【36】 A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
标题: 随机信封波动对噪音音素感知影响的微观研究
作者:Alejandro Osses,Léo Varnet
Journal-ref:Journal of the Acoustical Society of America, 2024, 155 (2), pp.1469-1485
链接:点击下载PDF文件
【37】 Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
标题: 优化歌曲创作过程:使用深度学习模型基于流派的歌词生成
作者:Tracy Cai,Wilson Liang,Donte Townes
链接:点击下载PDF文件
【38】 Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
标题: 通过原生数据库训练增强库尔德语文本到语音:高质量WaveGlow声码器方法
作者:Abdulhady Abas Abdullah,Sabat Salih Muhamad,Hadi Veisi
链接:点击下载PDF文件
【39】 Lightweight Transducer Based on Frame-Level Criterion
标题: 基于框架级标准的轻型传感器
作者:Genshun Wan,Mengzhi Wang,Tingzhi Mao,Hang Chen,Zhongfu Ye
备注:Accepted by Interspeech 2024, code repository: this https URL
链接:点击下载PDF文件
标题: Deepfake语音检测的全面调查和批判性分析
作者:Lam Pham,Phat Lam,Tin Nguyen,Hieu Tang,Huyen Nguyen,Alexander Schindler,Canh Vu
备注:Journal preprint
链接:点击下载PDF文件
摘要:由于深度学习的进步,语音生成系统现在为各种现实世界的应用提供了动力,例如语音障碍患者的文本到语音,呼叫中心的语音聊天机器人,跨语言语音翻译等。这促使研究团体开发用于检测合成语音的模型(例如,由基于深度学习的模型生成的假语音(Deepfake Speech Detection)任务。由于Deepfake语音检测任务是近年来出现的,因此针对该任务提出的调查论文并不多。此外,针对Deepfake语音检测任务的现有调查倾向于总结用于构建Deepfake语音检测系统的技术,而不是提供全面的分析。这一差距促使我们进行了一次全面的调查,对Deepfake语音检测的挑战和发展进行了批判性分析。我们的调查是创新性的结构,提供了对当前挑战比赛,公共数据集和深度学习技术的深入分析,这些技术提供了增强的解决方案,以解决该领域的现有挑战。根据我们的分析,我们提出了利用和结合特定深度学习技术来提高Deepfake语音检测系统有效性的假设。除了进行调查外,我们还进行了大量的实验来验证这些假设,并为Deepfake语音检测任务提出了一个极具竞争力的模型。鉴于分析和实验结果,我们最终指出了Deepfake语音检测任务潜在且有前途的研究方向。摘要:Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech translation, etc. While these systems can autonomously generate human-like speech and replicate specific voices, they also pose risks when misused for malicious purposes. This motivates the research community to develop models for detecting synthesized speech (e.g., fake speech) generated by deep-learning-based models, referred to as the Deepfake Speech Detection task. As the Deepfake Speech Detection task has emerged in recent years, there are not many survey papers proposed for this task. Additionally, existing surveys for the Deepfake Speech Detection task tend to summarize techniques used to construct a Deepfake Speech Detection system rather than providing a thorough analysis. This gap motivated us to conduct a comprehensive survey, providing a critical analysis of the challenges and developments in Deepfake Speech Detection. Our survey is innovatively structured, offering an in-depth analysis of current challenge competitions, public datasets, and the deep-learning techniques that provide enhanced solutions to address existing challenges in the field. From our analysis, we propose hypotheses on leveraging and combining specific deep learning techniques to improve the effectiveness of Deepfake Speech Detection systems. Beyond conducting a survey, we perform extensive experiments to validate these hypotheses and propose a highly competitive model for the task of Deepfake Speech Detection. Given the analysis and the experimental results, we finally indicate potential and promising research directions for the Deepfake Speech Detection task.
【2】 Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
标题: 通过负选择策略的自适应学习用于Few-Shot生物声学事件检测
作者:Yaxiong Chen,Xueping Zhang,Yunfei Zi,Shengwu Xiong
链接:点击下载PDF文件
摘要:虽然原型网络(ProtoNet)已经证明了在Few-Shot生物事件检测中的有效性,但仍然存在两个持续的问题。首先,由于缺乏明确注释的阴性样本,很难构建具有代表性的阴性原型。其次,目标生物发声的持续时间在不同的任务中各不相同,这使得模型在所有任务中始终产生最佳结果具有挑战性。为了解决这些问题,我们提出了一种新的自适应学习框架,具有自适应学习损失来指导分类器更新。此外,我们提出了一个否定选择策略,以构建一个更有代表性的否定原型的ProtoNet。所有实验均在DCASE 2023 TASK 5 Few-Shot生物声学事件检测数据集上进行。结果表明,我们提出的方法实现了0.703的F-措施,提高了12.84%。摘要:Although the Prototypical Network (ProtoNet) has demonstrated effectiveness in few-shot biological event detection, two persistent issues remain. Firstly, there is difficulty in constructing a representative negative prototype due to the absence of explicitly annotated negative samples. Secondly, the durations of the target biological vocalisations vary across tasks, making it challenging for the model to consistently yield optimal results across all tasks. To address these issues, we propose a novel adaptive learning framework with an adaptive learning loss to guide classifier updates. Additionally, we propose a negative selection strategy to construct a more representative negative prototype for ProtoNet. All experiments ware performed on the DCASE 2023 TASK5 few-shot bioacoustic event detection dataset. The results show that our proposed method achieves an F-measure of 0.703, an improvement of 12.84%.
【3】 LoVA: Long-form Video-to-Audio Generation
标题: LoVA:长格式视频到音频生成
作者:Xin Cheng,Xihua Wang,Yihan Wu,Yuyue Wang,Ruihua Song
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:视频到音频(V2A)生成对于视频编辑和后处理非常重要,可以为无声视频创建语义对齐的音频。然而,大多数现有的方法集中于为短视频片段(小于10秒)生成短格式音频,而很少关注长格式视频输入的场景。对于当前基于UNet的扩散V2A模型,在处理长格式音频生成时不可避免的问题是最终级联音频内的不一致性。在本文中,我们首先强调了长形式V2A问题的重要性。此外,我们提出了LoVA,一个新的模型,长格式视频到音频生成。事实证明,与现有的自回归模型和基于UNet的扩散模型相比,LoVA基于扩散Transformer(DiT)架构,在生成长格式音频方面更有效。大量的客观和主观实验表明,LoVA在10秒V2A基准测试中实现了相当的性能,并且在长格式视频输入的基准测试中优于所有其他基准。摘要:Video-to-audio (V2A) generation is important for video editing and post-processing, enabling the creation of semantics-aligned audio for silent video. However, most existing methods focus on generating short-form audio for short video segment (less than 10 seconds), while giving little attention to the scenario of long-form video inputs. For current UNet-based diffusion V2A models, an inevitable problem when handling long-form audio generation is the inconsistencies within the final concatenated audio. In this paper, we first highlight the importance of long-form V2A problem. Besides, we propose LoVA, a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models. Extensive objective and subjective experiments demonstrate that LoVA achieves comparable performance on 10-second V2A benchmark and outperforms all other baselines on a benchmark with long-form video input.
【4】 GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
标题: GALD-SE:引导各向异性轻量级扩散,实现高效语音增强
作者:Chengzhong Wang,Jianjun Gu,Dingding Yao,Zelin Qiu,Jiale Zhao,Junfeng Li
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:语音增强的目的是在各种噪声条件下提高语音的清晰度和质量。近年来,扩散模型在语音增强领域受到了广泛的关注,取得了令人瞩目的成果。当前基于扩散的方法用各向同性高斯噪声模糊初始信号,并从先前恢复干净的语音。然而,这些方法通常遭受大量的计算负担。我们认为,低效率源于忽视语音增强不是纯粹的生成任务;它主要涉及降噪和缺失信息的完成,而原始混合物中的干净线索不需要重新生成。在本文中,我们提出了一种方法,在扩散过程中引入具有各向异性指导的噪声,使神经网络能够专注于噪声记录中的干净线索。该方法对各种类型的噪声干扰和语音失真具有鲁棒性,并显著降低了计算量。实验表明,该方法实现了国家的最先进的结果,只有大约450万个参数,显着少于其他扩散方法所需的参数。这有效地缩小了基于扩散和预测语音增强方法之间的模型大小的差距。此外,所提出的方法在非常嘈杂的情况下表现良好,证明了其在极具挑战性的环境中的应用潜力。摘要:Speech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion model has gained lots of attention in speech enhancement area, achieving competitive results. Current diffusion-based methods blur the initial signal with isotropic Gaussian noise and recover clean speech from the prior. However, these methods often suffer from a substantial computational burden. We argue that the inefficiency stems from the oversight that speech enhancement is not purely a generative task; it primarily involves noise reduction and completion of missing information, while the clean clues in the original mixture do not need to be regenerated. In this paper, we propose a method that introduces noise with anisotropic guidance during the diffusion process, allowing the neural network to focus on clean clues within noisy recordings. This approach is robust against various types of noise interference and speech distortion, and significantly reduces the computational load. Experiments demonstrate that the proposed method achieves state-of-the-art results with only approximately 4.5 million parameters, significantly fewer than those required by other diffusion methods. This effectively narrows the gap in model size between diffusion-based and predictive speech enhancement approaches. Additionally, the proposed method performs well in very noisy scenarios, demonstrating its potential for applications in highly challenging environments.
【5】 Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
标题: 根据单独的房间和场景特定信息生成盲空间脉冲响应
作者:Francesc Lluís,Nils Meyer-Kahlen
链接:点击下载PDF文件
摘要:对于增强现实(AR)中的音频,用户真实声学环境的知识对于渲染无缝融入环境的虚拟声音至关重要。由于声学测量在实际的AR应用中通常是不可行的,因此需要从可用的声源推断出关于房间的信息。然后,可以用相同的房间声学质量来呈现附加的声源。至关重要的是,这些被放置在与可用于估计的源不同的位置。在这里,我们建议使用一个编码器网络,该网络使用对比度损失进行训练,将输入声音映射到仅表示房间特定信息的低维特征空间。然后,基于扩散的空间房间脉冲响应发生器被训练以获取潜在空间并生成新的响应,给定新的源-接收器位置。我们展示了如何在最终输出中考虑房间和位置特定的参数。摘要:For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR applications, information about the room needs to be inferred from available sound sources. Then, additional sound sources can be rendered with the same room acoustic qualities. Crucially, these are placed at different positions than the sources available for estimation. Here, we propose to use an encoder network trained using a contrastive loss that maps input sounds to a low-dimensional feature space representing only room-specific information. Then, a diffusion-based spatial room impulse response generator is trained to take the latent space and generate a new response, given a new source-receiver position. We show how both room- and position-specific parameters are considered in the final output.
【6】 Voice Conversion-based Privacy through Adversarial Information Hiding
标题: 通过对抗性信息隐藏实现基于语音转换的隐私
作者:Jacob J Webber,Oliver Watts,Gustav Eje Henter,Jennifer Williams,Simon King
备注:Accepted for publication in proceedings of 4th symposium on security and privacy in speech communication
链接:点击下载PDF文件
摘要:隐私保护语音转换的目标是只删除语音音频中传达身份信息的属性,而保持其他语音特征不变。本文提出了一种保护隐私的语音转换机制,允许使用对抗性信息隐藏来控制身份承载信息的泄漏。这使得能够在保持源语音特征和修改说话者身份之间进行慎重的权衡。因此,该方法改进了CycleGAN和StarGAN等语音转换技术,这些技术不是为隐私而设计的,这意味着转换后的语音可能会以不可预测的方式泄露个人信息。我们的方法也比ASR-TTS语音转换管道更灵活,ASR-TTS语音转换管道通过设计丢弃与文本内容相关的所有韵律信息。评估结果表明,该系统成功地修改感知扬声器的身份,同时很好地保持源词汇内容。摘要:Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion that allows controlling the leakage of identity-bearing information using adversarial information hiding. This enables a deliberate trade-off between maintaining source-speech characteristics and modification of speaker identity. As such, the approach improves on voice-conversion techniques like CycleGAN and StarGAN, which were not designed for privacy, meaning that converted speech may leak personal information in unpredictable ways. Our approach is also more flexible than ASR-TTS voice conversion pipelines, which by design discard all prosodic information linked to textual content. Evaluations show that the proposed system successfully modifies perceived speaker identity whilst well maintaining source lexical content.
【7】 HiFi-Glot: Neural Formant Synthesis with Differentiable Resonant Filters
标题: HiFi-Glot:使用可区分共振过滤器的神经Forces合成
作者:Lauri Juvela,Pablo Pérez Zarazaga,Gustav Eje Henter,Zofia Malisz
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:我们介绍了一个使用语音产生的源过滤器模型的端到端神经语音合成系统。具体来说,我们应用微分谐振滤波器的声门波形产生的神经声码器。其目的是获得一个可控的合成器,类似于经典的共振峰合成,但具有更高的感知质量-填补了当前神经波形发生器的研究空白,并响应语音科学迄今未满足的需求。我们的设置从语音上有意义的语音参数的核心集生成音频,滤波器在合成中提供对共振峰频率共振的直接控制。直接合成控制是重要语音科学实验中可靠刺激产生的关键特征。我们表明,所提出的源滤波器方法给出了比共振峰操纵的行业标准更好的感知质量(即,Praat),同时在共振峰频率控制精度方面具有竞争力。摘要:We introduce an end-to-end neural speech synthesis system that uses the source-filter model of speech production. Specifically, we apply differentiable resonant filters to a glottal waveform generated by a neural vocoder. The aim is to obtain a controllable synthesiser, similar to classic formant synthesis, but with much higher perceptual quality - filling a research gap in current neural waveform generators and responding to hitherto unmet needs in the speech sciences. Our setup generates audio from a core set of phonetically meaningful speech parameters, with the filters providing direct control over formant frequency resonances in synthesis. Direct synthesis control is a key feature for reliable stimulus creation in important speech science experiments. We show that the proposed source-filter method gives better perceptual quality than the industry standard for formant manipulation (i.e., Praat), whilst being competitive in terms of formant frequency control accuracy.
【8】 SongTrans: An unified song transcription and alignment method for lyrics and notes
标题: SongTrans:歌词和音符的统一歌曲转录和对齐方法
作者:Siwei Wu,Jinzheng He,Ruibin Yuan,Haojie Wei,Xipin Wei,Chenghua Lin,Jin Xu,Junyang Lin
链接:点击下载PDF文件
摘要:处理后的数据量是提高歌唱声音合成领域的关键。虽然存在可用于歌词或音符转录任务的工具,但它们都需要相对耗时的预处理数据(例如,声乐和伴奏分离)。此外,这些工具中的大多数都是为了解决一个单一的任务,并努力对齐歌词和音符(即,识别歌词中每个单词的相应音符)。为了应对这些挑战,我们首先通过优化现有工具并注释大量歌曲的歌词-音符对来设计管道。然后,基于标注的数据,我们训练一个统一的SongTrans模型,该模型可以直接转录歌词和音符,同时对齐它们,而不需要对歌曲进行预处理。我们的SongTrans模型由两个模块组成:(1) textbf{自回归模块}预测歌词,以及歌词中每个单词对应的持续时间和音符数量。(2) textbf{非自回归模块}预测音符的音高和持续时间。我们的实验表明,SongTrans实现了最先进的(SOTA)结果在歌词和笔记转录任务。此外,它是第一个能够将歌词与音符对齐的模型。实验结果表明,SongTrans模型可以有效地适应不同类型的歌曲(例如,歌曲与伴奏),展示了其多功能性的现实世界的应用。摘要:The quantity of processed data is crucial for advancing the field of singing voice synthesis. While there are tools available for lyric or note transcription tasks, they all need pre-processed data which is relatively time-consuming (e.g., vocal and accompaniment separation). Besides, most of these tools are designed to address a single task and struggle with aligning lyrics and notes (i.e., identifying the corresponding notes of each word in lyrics). To address those challenges, we first design a pipeline by optimizing existing tools and annotating numerous lyric-note pairs of songs. Then, based on the annotated data, we train a unified SongTrans model that can directly transcribe lyrics and notes while aligning them simultaneously, without requiring pre-processing songs. Our SongTrans model consists of two modules: (1) the textbf{Autoregressive module} predicts the lyrics, along with the duration and note number corresponding to each word in a lyric. (2) the textbf{Non-autoregressive module} predicts the pitch and duration of the notes. Our experiments demonstrate that SongTrans achieves state-of-the-art (SOTA) results in both lyric and note transcription tasks. Furthermore, it is the first model capable of aligning lyrics with notes. Experimental results demonstrate that the SongTrans model can effectively adapt to different types of songs (e.g., songs with accompaniment), showcasing its versatility for real-world applications.
【9】 What Are They Doing? Joint Audio-Speech Co-Reasoning
标题: 他们在做什么?音频-语音联合推理
作者:Yingzhi Wang,Pooneh Mousavi,Artem Ploujnikov,Mirco Ravanelli
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在音频和语音处理中,任务通常集中在音频或语音模态上,即使声音和人类语音都存在于同一音频片段中。最近的听觉大语言模型(ALLM)使得在单个模型中同时处理音频和语音成为可能,从而进一步考虑联合音频语音任务。 在本文中,我们调查如何以及ALLM可以执行联合音频语音处理。具体来说,我们介绍了联合音频语音协同推理(JASCO),一种新的任务,统一的音频和语音处理,严格要求跨两种模式的协同推理。我们发布了一个名为“他们在做什么”的场景推理数据集,并建立了一个联合音频语音基准来评估流行的ALLM的联合推理能力。此外,我们通过分析模型对每种模态的依赖性,对模型的行为提供了更深入的了解。摘要:In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have made it possible to process audio and speech simultaneously within a single model, leading to further considerations of joint audio-speech tasks. In this paper, we investigate how well ALLMs can perform joint audio-speech processing. Specifically, we introduce Joint Audio-Speech Co-Reasoning (JASCO), a novel task that unifies audio and speech processing, strictly requiring co-reasoning across both modalities. We release a scene-reasoning dataset called "What Are They Doing" and establish a joint audio-speech benchmark to evaluate the joint reasoning capability of popular ALLMs. Additionally, we provide deeper insights into the models' behaviors by analyzing their dependence on each modality.
【10】 CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
标题: CPD增强Wav2vec2.0:面向课堂环境的噪音稳健语音识别
作者:Ahmed Adel Attia,Dorottya Demszky,Tolulope Ogunremi,Jing Liu,Carol Espy-Wilson
备注:arXiv admin note: substantial text overlap with arXiv:2405.13018
链接:点击下载PDF文件
摘要:创建强大且适应课堂条件的自动语音识别(ASR)系统对于开发帮助教师和学生的人工智能工具至关重要。在这项工作中,我们研究了持续预训练(CPT)在适应Wav2vec2.0课堂领域的功效。我们表明,CPT在这方面是一个强大的工具,并将基于Wav2vec2.0的模型的字错误率(WER)降低了10%以上。更具体地说,CPT提高了模型对不同噪声、麦克风和教室条件的鲁棒性。摘要:Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model's robustness to different noises, microphones and classroom conditions.
【11】 Self-Supervised Audio-Visual Soundscape Stylization
标题: 自我监督的视听声景风格化
作者:Tingle Li,Renhao Wang,Po-Yao Huang,Andrew Owens,Gopala Anumanchipalli
备注:ECCV 2024
链接:点击下载PDF文件
摘要:语音声音传达了大量关于场景的信息,从而产生从混响到附加环境声音的各种效果。在本文中,我们操纵输入语音的声音,好像它是记录在一个不同的场景,给定的视听条件的例子记录从该场景。我们的模型通过自我监督来学习,利用自然视频包含重复出现的声音事件和纹理的事实。我们从视频中提取音频片段并应用语音增强。然后,我们训练一个潜在的扩散模型来恢复原始语音,使用从视频中其他地方拍摄的另一个视听剪辑作为条件提示。通过这个过程,模型学习将条件示例的声音属性转移到输入语音。我们证明了我们的模型可以使用未标记的野外视频成功训练,并且额外的视觉信号可以提高其声音预测能力。请参阅我们的项目网页视频结果:https: tinglok.netlify.app files avsoundscape 摘要:Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input speech to sound as though it was recorded within a different scene, given an audio-visual conditional example recorded from that scene. Our model learns through self-supervision, taking advantage of the fact that natural video contains recurring sound events and textures. We extract an audio clip from a video and apply speech enhancement. We then train a latent diffusion model to recover the original speech, using another audio-visual clip taken from elsewhere in the video as a conditional hint. Through this process, the model learns to transfer the conditional example's sound properties to the input speech. We show that our model can be successfully trained using unlabeled, in-the-wild videos, and that an additional visual signal can improve its sound prediction abilities. Please see our project webpage for video results: https: tinglok.netlify.app files avsoundscape
【12】 AMT-APC: Automatic Piano Cover by Fine-Tuning an Automatic Music Transcription Model
标题: AMT-IPC:通过微调自动音乐转录模型实现自动钢琴翻唱
作者:Kazuma Komiya,Yoshihisa Fukuhara
链接:点击下载PDF文件
摘要:已经有几项关于自动生成钢琴封面的研究,最近深度学习的进步使得能够创建更复杂的封面。然而,现有的自动钢琴盖模型在表现力和对原作的保真度方面仍有改进的空间。为了解决这些问题,我们提出了一种称为AMT-APC的学习算法,该算法利用了自动音乐转录模型的功能。通过利用完善的自动音乐转录模型的优势,我们的目标是提高钢琴覆盖生成的准确性。我们的实验表明,AMT-APC模型比任何现有模型更准确地再现原始轨迹。摘要:There have been several studies on automatically generating piano covers, and recent advancements in deep learning have enabled the creation of more sophisticated covers. However, existing automatic piano cover models still have room for improvement in terms of expressiveness and fidelity to the original. To address these issues, we propose a learning algorithm called AMT-APC, which leverages the capabilities of automatic music transcription models. By utilizing the strengths of well-established automatic music transcription models, we aim to improve the accuracy of piano cover generation. Our experiments demonstrate that the AMT-APC model reproduces original tracks more accurately than any existing models.
【13】 MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
标题: MultiMed:通过注意力编码器解码器的多语言医学语音识别
作者:Khai Le-Duc,Phuc Phan,Tan-Hanh Pham,Bach Phan Tat,Minh-Huong Ngo,Truong-Son Hy
备注:Preprint
链接:点击下载PDF文件
摘要:医学领域的多语言自动语音识别(ASR)是语音翻译、口语理解和语音激活助理等各种下游应用的基础任务。该技术通过跨越语言障碍实现高效沟通、缓解专业劳动力短缺以及促进改善诊断和治疗(特别是在大流行期间)来增强患者护理。在这项工作中,我们介绍了MultiMed,这是一个针对医疗领域的小型到大型端到端ASR模型集合,涵盖五种语言:越南语,英语,德语,法语和汉语普通话,以及相应的真实世界ASR数据集。据我们所知,MultiMed是最大和第一个多语言医学ASR数据集,包括总持续时间,说话者数量,疾病多样性,记录条件,说话者角色,独特的医学术语,口音和ICD-10代码。其次,我们建立了经验基线,提出了第一个可重复的医学ASR多语言研究,进行了端到端ASR培训的逐层消融研究,并为多语言医学ASR提供了第一个语言分析。所有代码、数据和模型均可在https: github.com leduckhai MultiMed tree master MultiMed上获得摘要:Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants. This technology enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we introduce MultiMed, a collection of small-to-large end-to-end ASR models for the medical domain, spanning five languages: Vietnamese, English, German, French, and Mandarin Chinese, together with the corresponding real-world ASR dataset. To our best knowledge, MultiMed stands as the largest and the first multilingual medical ASR dataset, in terms of total duration, number of speakers, diversity of diseases, recording conditions, speaker roles, unique medical terms, accents, and ICD-10 codes. Secondly, we establish the empirical baselines, present the first reproducible study of multilinguality in medical ASR, conduct a layer-wise ablation study for end-to-end ASR training, and provide the first linguistic analysis for multilingual medical ASR. All code, data, and models are available online https: github.com leduckhai MultiMed tree master MultiMed
【14】 ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
标题: ECHO:采用分层实体指导的半监督学习的环境声音分类
作者:Pranav Gupta,Raunak Sharma,Rashmi Kumari,Sri Krishna Aditya,Shwetank Choudhary,Sumit Kumar,Kanchana M,Thilagavathy R
备注:IEEE CONECCT 2024, Signal Processing and Pattern Recognition, Environmental Sound Classification, ESC
链接:点击下载PDF文件
摘要:环境声分类是信号处理领域的一个研究热点,目前研究的重点是全监督分类。在过去的几年里,焦点已经转向半监督方法,专注于利用未标记的数据,和自我监督的方法,通过借口任务或对比学习学习的中间表示。然而,这两种方法都需要大量未标记的数据来提高性能。在这项工作中,我们提出了一个新的框架,称为环境声音分类与层次本体指导的半监督学习(ECHO),利用标签本体为基础的层次结构,通过定义一个新的借口任务来学习语义表示。在prefect任务中,该模型尝试基于地面真值标签本体来预测由大型语言模型(LLM)定义的粗略标签。经过训练的模型以监督的方式进一步微调,以预测实际任务。我们提出的新的半监督框架在三个数据集(即UrbanSound 8 K,ESC-10和ESC-50)上实现了比基线系统1%至8%的精度提高。摘要:Environment Sound Classification has been a well-studied research problem in the field of signal processing and up till now more focus has been laid on fully supervised approaches. Over the last few years, focus has moved towards semi-supervised methods which concentrate on the utilization of unlabeled data, and self-supervised methods which learn the intermediate representation through pretext task or contrastive learning. However, both approaches require a vast amount of unlabelled data to improve performance. In this work, we propose a novel framework called Environmental Sound Classification with Hierarchical Ontology-guided semi-supervised Learning (ECHO) that utilizes label ontology-based hierarchy to learn semantic representation by defining a novel pretext task. In the pretext task, the model tries to predict coarse labels defined by the Large Language Model (LLM) based on ground truth label ontology. The trained model is further fine-tuned in a supervised way to predict the actual task. Our proposed novel semi-supervised framework achieves an accuracy improvement in the range of 1 % to 8 % over baseline systems across three datasets namely UrbanSound8K, ESC-10, and ESC-50.
【15】 Training Large ASR Encoders with Differential Privacy
标题: 训练具有差异隐私的大型ASB编码器
作者:Geeticka Chauhan,Steve Chien,Om Thakkar,Abhradeep Thakurta,Arun Narayanan
备注:In proceedings of the IEEE Spoken Language Technologies Workshop, 2024
链接:点击下载PDF文件
摘要:大型语音模型的自监督学习(SSL)方法已被证明在ASR中非常有效。随着对公共部署大型预训练模型的兴趣,人们越来越担心训练数据中的敏感数据点的意外记忆和泄漏。在本文中,我们将差分私有(DP)预训练应用于基于SOTA Conformer的编码器,并研究其在下游ASR任务上的性能,假设微调数据是公开的。本文首次将DP应用于SSL以实现ASR,研究了BEST-RQ预训练方法的DP噪声容忍度。值得注意的是,我们引入了一种新的模型修剪变体,称为基于梯度的层冻结,它在隐私-效用-计算权衡方面提供了很大的改进。我们的方法得出的LibriSpeech测试-干净 其他WER(%)为3.78 8.41,外推到低数据集尺度时为($10$,1 e ^-9)-DP,外推到高尺度时为2.81 5.89,外推到(10,7.9e^-11)-DP。摘要:Self-supervised learning (SSL) methods for large speech models have proven to be highly effective at ASR. With the interest in public deployment of large pre-trained models, there is a rising concern for unintended memorization and leakage of sensitive data points from the training data. In this paper, we apply differentially private (DP) pre-training to a SOTA Conformer-based encoder, and study its performance on a downstream ASR task assuming the fine-tuning data is public. This paper is the first to apply DP to SSL for ASR, investigating the DP noise tolerance of the BEST-RQ pre-training method. Notably, we introduce a novel variant of model pruning called gradient-based layer freezing that provides strong improvements in privacy-utility-compute trade-offs. Our approach yields a LibriSpeech test-clean other WER (%) of 3.78 8.41 with ($10$, 1e^-9)-DP for extrapolation towards low dataset scales, and 2.81 5.89 with (10, 7.9e^-11)-DP for extrapolation towards high scales.
【16】 Target word activity detector: An approach to obtain ASR word boundaries without lexicon
标题: 目标词活动检测器:一种无需词典即可获得ASB词边界的方法
作者:Sunit Sivasankaran,Eric Sun,Jinyu Li,Yan Huang,Jing Pan
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:由于在训练期间缺乏明确的时间对齐,从端到端(E2E)ASR模型中获取单词时间戳信息仍然具有挑战性。这个问题在多语言模型中更加复杂。现有的方法要么依赖于词典,要么引入额外的令牌,导致可扩展性问题和增加的计算成本。在这项工作中,我们提出了一种新的方法来估计词的边界,而不依赖于词典。我们的方法利用了来自子单词标记单元的单词嵌入和预训练的ASR模型,在训练过程中只需要单词对齐信息。我们提出的方法可以扩展到任何数量的语言,而不会产生任何额外的成本。我们使用在五种语言上训练的多语言ASR模型来验证我们的方法,并证明了其对强大基线的有效性。摘要:Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complicated in multilingual models. Existing methods, either rely on lexicons or introduce additional tokens, leading to scalability issues and increased computational costs. In this work, we propose a new approach to estimate word boundaries without relying on lexicons. Our method leverages word embeddings from sub-word token units and a pretrained ASR model, requiring only word alignment information during training. Our proposed method can scale-up to any number of languages without incurring any additional cost. We validate our approach using a multilingual ASR model trained on five languages and demonstrate its effectiveness against a strong baseline.
【17】 PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
标题: PTQ4ADM:高效文本条件音频扩散模型的训练后量化
作者:Jayneel Vora,Aditya Krishnan,Nader Bouacida,Prabhu RV Shankar,Prasant Mohapatra
链接:点击下载PDF文件
摘要:去噪扩散模型已成为图像、音频和视频领域生成任务的最新技术,可生成高质量、多样化和上下文相关的数据。然而,它们的广泛采用受到高计算成本和大内存占用的限制。后训练量化(PTQ)通过低带宽参数降低模型复杂度,提供了一种有前途的方法来缓解这些挑战。然而,直接将PTQ应用于扩散模型可能会降低合成质量,这是由于多个去噪步骤中累积的量化噪声,特别是在文本到音频合成等条件任务中。本文介绍了一种新的音频扩散模型量化框架PTQ 4ADM。我们的主要贡献包括(1)覆盖率驱动的提示增强方法和(2)激活感知的文本条件ADM的校准集生成算法。这些技术确保全面覆盖音频方面和模态,同时保持合成保真度。我们验证了我们的方法TANGO,Make-An-Audio和AudioLDM模型的文本条件音频生成。大量的实验证明PTQ 4ADM的能力,以减少模型的大小高达70%,同时实现合成质量指标与全精度模型(FD分数增加$<5%)。我们表明,骨干网络中的特定层可以量化为4位权重和8位激活,而不会有显着的质量损失。这项工作为在资源受限的环境中更有效地部署ADM铺平了道路。摘要:Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is limited by high computational costs and large memory footprints. Post-training quantization (PTQ) offers a promising approach to mitigate these challenges by reducing model complexity through low-bandwidth parameters. Yet, direct application of PTQ to diffusion models can degrade synthesis quality due to accumulated quantization noise across multiple denoising steps, particularly in conditional tasks like text-to-audio synthesis. This work introduces PTQ4ADM, a novel framework for quantizing audio diffusion models(ADMs). Our key contributions include (1) a coverage-driven prompt augmentation method and (2) an activation-aware calibration set generation algorithm for text-conditional ADMs. These techniques ensure comprehensive coverage of audio aspects and modalities while preserving synthesis fidelity. We validate our approach on TANGO, Make-An-Audio, and AudioLDM models for text-conditional audio generation. Extensive experiments demonstrate PTQ4ADM's capability to reduce the model size by up to 70 % while achieving synthesis quality metrics comparable to full-precision models($<$5 % increase in FD scores). We show that specific layers in the backbone network can be quantized to 4-bit weights and 8-bit activations without significant quality loss. This work paves the way for more efficient deployment of ADMs in resource-constrained environments.
【18】 Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
标题: 基于柱状图层延时神经网络的时频特征组合研究
作者:Amirmohammad Mohammadi,Iren'e Masabarakiza,Ethan Barnes,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 14 figures. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:虽然深度学习减少了手动特征提取的流行,但通过特征工程进行数据转换对于提高模型性能仍然至关重要,特别是对于水声信号。将音频信号转换为时频表示的方法以及随后对这些频谱图的处理可以显著影响性能。这项工作演示了在直方图层时间延迟神经网络中使用不同的时频特征组合对性能的影响。一组最佳的功能识别的结果表明,特定的功能组合优于单一的数据功能。摘要:While deep learning has reduced the prevalence of manual feature extraction, transformation of data via feature engineering remains essential for improving model performance, particularly for underwater acoustic signals. The methods by which audio signals are converted into time-frequency representations and the subsequent handling of these spectrograms can significantly impact performance. This work demonstrates the performance impact of using different combinations of time-frequency features in a histogram layer time delay neural network. An optimal set of features is identified with results indicating that specific feature combinations outperform single data features.
【19】 Transfer Learning for Passive Sonar Classification using Pre-trained Audio and ImageNet Models
标题: 使用预训练的音频和ImageNet模型进行被动声纳分类的迁移学习
作者:Amirmohammad Mohammadi,Tejashri Kelhe,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 6 figures, This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:迁移学习通常用于利用大型预训练模型并对下游任务进行微调。最流行的预训练模型最初使用ImageNet进行训练。然而,它们的泛化能力在不同的数据模式中可能会有所不同。本研究在水下声学目标识别(UATR)的背景下比较了预训练的音频神经网络(PANN)和ImageNet预训练模型。据观察,ImageNet预训练模型在被动声纳分类中的表现略优于预训练音频模型。我们还分析了音频采样率对模型预训练和微调的影响。这项研究有助于UATR的迁移学习应用,说明了预训练模型在解决UATR领域稀缺的标记数据所造成的局限性方面的潜力。摘要:Transfer learning is commonly employed to leverage large, pre-trained models and perform fine-tuning for downstream tasks. The most prevalent pre-trained models are initially trained using ImageNet. However, their ability to generalize can vary across different data modalities. This study compares pre-trained Audio Neural Networks (PANNs) and ImageNet pre-trained models within the context of underwater acoustic target recognition (UATR). It was observed that the ImageNet pre-trained models slightly out-perform pre-trained audio models in passive sonar classification. We also analyzed the impact of audio sampling rates for model pre-training and fine-tuning. This study contributes to transfer learning applications of UATR, illustrating the potential of pre-trained models to address limitations caused by scarce, labeled data in the UATR domain.
【20】 A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
标题: 随机信封波动对噪音音素感知影响的微观研究
作者:Alejandro Osses,Léo Varnet
Journal-ref:Journal of the Acoustical Society of America, 2024, 155 (2), pp.1469-1485
链接:点击下载PDF文件
摘要:在这项研究中,我们研究了特定的噪音实现对两个辅音, b 和 d 的歧视的影响。为此,我们收集了12名参与者的数据,他们听了嵌入在三种背景噪音中的 aba 或 ada 。所有的噪声都有相同的长期频谱,但随机包络波动的数量不同。使用反向相关法在逐个试验的基础上对数据进行分析。结果表明,它是有可能预测的分类响应具有比机会更好的准确性纯粹基于相应的噪声的随机包络波动的频谱-时间分布,而不考虑实际的目标或在试验中使用的信噪比。噪声波动的影响平均解释了参与者在白噪声中8.1%的反应,对于波动量较大的噪声,这一比例增加到13.3%。估计的时频权重表明,测量的效果源于噪声波动和目标词的相关声学线索之间的混淆。从使用人工...的模拟中获得了基本相似的结论。我们认为,这种标记特定的噪声效应是一种形式的信息掩蔽。摘要:In this study, we investigated the effect of specific noise realizations on the discrimination of two consonants, b and d . For this purpose, we collected data from twelve participants, who listened to the words aba or ada embedded in one of three background noises. All noises had the same long-term spectrum but differed in the amount of random envelope fluctuations. The data were analyzed on a trial-by-trial basis using the reverse-correlation method. The results revealed that it is possible to predict the categorical responses with better-than-chance accuracy purely based on the spectro-temporal distribution of the random envelope fluctuations of the corresponding noises, without taking into account the actual targets or the signal-to-noise ratios used in the trials. The effect of the noise fluctuations explained on average 8.1% of the participants' responses in white noise, a proportion that increased up to 13.3% for noises with a larger amount of fluctuations. The estimated time-frequency weights revealed that the measured effect originated from confusions between noise fluctuations and relevant acoustic cues from the target words. Substantially similar conclusions were obtained from simulations using an artificial listener. We argue that this token-specific effect of noise is a form of informational masking.
【21】 Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
标题: 优化歌曲创作过程:使用深度学习模型基于流派的歌词生成
作者:Tracy Cai,Wilson Liang,Donte Townes
链接:点击下载PDF文件
摘要:传统的歌曲创作过程是相当复杂的,这是显而易见的时间,它需要产生的歌词,适合体裁和形式全面的诗句。我们的项目旨在通过深度学习技术简化这一过程,从而优化歌曲创作过程,并使艺术家能够通过保持流派来达到目标受众。使用Spotify上的18,000首歌曲的数据集,我们开发了一种独特的预处理格式,使用令牌将歌词解析为单独的诗句。这些结果被用来训练基线预训练seq2seq模型,和基于LSTM的神经网络模型,根据歌曲流派。我们发现,在基线模型中,生成产生了更高的召回率(ROUGE),但两个模型的精度(BLEU)相似。从质量上看,我们发现原始模型生成的许多抒情短语仍然可以理解和辨别,尽管它们不一定与真正的歌词完全相同。总的来说,我们的研究结果表明,歌词生成可以合理地加快产生基于体裁的歌词,并有助于加快歌曲创作过程。摘要:The traditional songwriting process is rather complex and this is evident in the time it takes to produce lyrics that fit the genre and form comprehensive verses. Our project aims to simplify this process with deep learning techniques, thus optimizing the songwriting process and enabling an artist to hit their target audience by staying in genre. Using a dataset of 18,000 songs off Spotify, we developed a unique preprocessing format using tokens to parse lyrics into individual verses. These results were used to train a baseline pretrained seq2seq model, and a LSTM-based neural network models according to song genres. We found that generation yielded higher recall (ROUGE) in the baseline model, but similar precision (BLEU) for both models. Qualitatively, we found that many of the lyrical phrases generated by the original model were still comprehensible and discernible between which genres they fit into, despite not necessarily being the exact the same as the true lyrics. Overall, our results yielded that lyric generation can reasonably be sped up to produce genre-based lyrics and aid in hastening the songwriting process.
【22】 Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
标题: 通过原生数据库训练增强库尔德语文本到语音:高质量WaveGlow声码器方法
作者:Abdulhady Abas Abdullah,Sabat Salih Muhamad,Hadi Veisi
链接:点击下载PDF文件
摘要:随着文本转语音技术的进步,从文本合成口语的能力极大地便利了数字内容的获取。然而,有效的TTS开发低资源的语言,如中央库尔德语(CKB),仍然面临着许多挑战,主要是由于缺乏语言信息和专用资源。在本文中,我们改进了基于Tacotron的库尔德文语转换系统,通过在21小时的中央库尔德语语音语料库上训练库尔德语WaveGlow声码器,而不是使用预先训练的英语声码器WaveGlow。为了准确、流畅地适应库尔德语语音和韵律的变化,需要在目标语语料库上进行声码器训练。这些增强的有效性在于,我们的模型明显优于使用英语预训练模型的基线系统。特别是,我们的自适应WaveGlow模型达到了令人印象深刻的4.91 MOS,为库尔德语语音合成设定了新的基准。一方面,这项研究赋予中央库尔德语TTS系统的先进功能,另一方面,它为库尔德语和其他相关语言的其他方言的进一步发展打开了大门。摘要:The ability to synthesize spoken language from text has greatly facilitated access to digital content with the advances in text-to-speech technology. However, effective TTS development for low-resource languages, such as Central Kurdish (CKB), still faces many challenges due mainly to the lack of linguistic information and dedicated resources. In this paper, we improve the Kurdish TTS system based on Tacotron by training the Kurdish WaveGlow vocoder on a 21-hour central Kurdish speech corpus instead of using a pre-trained English vocoder WaveGlow. Vocoder training on the target language corpus is required to accurately and fluently adapt phonetic and prosodic changes in Kurdish language. The effectiveness of these enhancements is that our model is significantly better than the baseline system with English pretrained models. In particular, our adaptive WaveGlow model achieves an impressive MOS of 4.91, which sets a new benchmark for Kurdish speech synthesis. On one hand, this study empowers the advanced features of the TTS system for Central Kurdish, and on the other hand, it opens the doors for other dialects in Kurdish and other related languages to further develop.
【23】 Lightweight Transducer Based on Frame-Level Criterion
标题: 基于框架级标准的轻型传感器
作者:Genshun Wan,Mengzhi Wang,Tingzhi Mao,Hang Chen,Zhongfu Ye
备注:Accepted by Interspeech 2024, code repository: this https URL
链接:点击下载PDF文件
摘要:基于序列级准则训练的换能器模型由于产生大概率矩阵而需要大量的内存。提出了一种基于帧级准则的轻量级传感器模型,该模型利用CTC强制对齐算法的结果来确定每帧的标签。然后,编码器输出可以在相应的时间与解码器输出组合,而不是像在换能器中那样将编码器输出的每个元素添加到解码器输出的每个元素。这大大降低了内存和计算需求。为了解决标签中过多空白导致的分类不平衡问题,我们将空白和非空白概率解耦,并将空白分类器的梯度截断到主网络。这使得轻质换能器能够实现与换能器类似的结果。此外,我们使用更丰富的信息来预测空白的概率,取得了优于传感器的结果。摘要:The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer model based on frame-level criterion, which uses the results of the CTC forced alignment algorithm to determine the label for each frame. Then the encoder output can be combined with the decoder output at the corresponding time, rather than adding each element output by the encoder to each element output by the decoder as in the transducer. This significantly reduces memory and computation requirements. To address the problem of imbalanced classification caused by excessive blanks in the label, we decouple the blank and non-blank probabilities and truncate the gradient of the blank classifier to the main network. This enables the lightweight transducer achieving similar results to transducer. Additionally, we use richer information to predict the probability of blank, achieving superior results to transducer.
【24】 CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification
标题: CA-MHFA:用于基于SSL的说话人验证的上下文感知多头因子化注意力池
作者:Junyi Peng,Ladislav Mošner,Lin Zhang,Oldřich Plchot,Themos Stafylakis,Lukáš Burget,Jan Černocký
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:近年来,用于说话人确认(SV)的自监督学习(SSL)模型受到了极大的关注。然而,现有的基于SSL的SV系统往往很难捕捉本地时间依赖性和概括不同的任务。在本文中,我们提出了上下文感知的多头因式分解注意池(CA-MHFA),一个轻量级的框架,结合周围的帧的上下文信息。CA-MHFA利用分组的、可学习的查询来有效地建模上下文依赖关系,同时通过在组之间共享键和值来保持效率。在VoxCeleb数据集上的实验结果表明,CA-MHFA在Vox 1-O、Vox 1-E和Vox 1-H上的EER分别为0.42 %、0.48 %和0.96 %,优于WavLM-TDNN等复杂模型,参数较少,收敛速度更快。此外,CA-MHFA在多个SSL模型和任务中表现出强大的泛化能力,包括情感识别和反欺骗,突出了其鲁棒性和多功能性。摘要:Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42 %, 0.48 %, and 0.96 % on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.
【25】 LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
标题: LlamaPartialSpoof:一个LLM驱动的模拟虚假信息生成的假语音数据集
作者:Hieu-Thi Luong,Haoyang Li,Lin Zhang,Kong Aik Lee,Eng Siong Chng
备注:5 pages, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:以前的假语音数据集是从防御者的角度构建的,以开发对策(CM)系统,而不考虑攻击者的不同动机。为了更好地与现实生活中的场景保持一致,我们创建了LlamaPartialSpoof,这是一个130小时的数据集,包含完全和部分虚假的语音,使用大型语言模型(LLM)和语音克隆技术来评估CM的鲁棒性。通过检查对攻击者和防御者都有价值的信息,我们确定了当前CM系统中的几个关键漏洞,这些漏洞可以被利用来提高攻击成功率,包括对某些文本到语音模型或拼接方法的偏见。我们的实验结果表明,目前的虚假语音检测系统的斗争,以推广到看不见的情况下,实现了24.44%的等错误率的最佳性能。摘要:Previous fake speech datasets were constructed from a defender's perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining information valuable to both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.44% equal error rate.
【26】 Room Impulse Responses help attackers to evade Deep Fake Detection
标题: 房间冲动响应帮助攻击者逃避深度伪造检测
作者:Hieu-Thi Luong,Duc-Tuan Truong,Kong Aik Lee,Eng Siong Chng
备注:7 pages, to be presented at SLT 2024
链接:点击下载PDF文件
摘要:ASVspoof 2021基准测试是一个广泛使用的反欺骗评估框架,由两个子集组成:逻辑访问(LA)和Deepfake(DF),具有不同编码特征和压缩伪影的样本。值得注意的是,当前最先进的(SOTA)系统拥有令人印象深刻的性能,在LA子集上实现了0.87%的等错误率(EER),在DF上实现了2.58%的等错误率(EER)。然而,基准测试的准确性并不能保证真实场景中的鲁棒性。本文研究了利用房间脉冲响应(RIR)来增强假语音和增加其逃避假语音检测系统的可能性的有效性。我们的研究结果表明,这种简单的方法显着提高了逃避率,使SOTA系统的EER翻了一番。为了应对这种类型的攻击,我们使用大规模的合成 模拟RIR数据集来增强训练数据。结果表明,混响假语音和原始样本的显着改善,降低DF任务EER为2.13%。摘要:The ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the current state-of-the-art (SOTA) system boasts impressive performance, achieving an Equal Error Rate (EER) of 0.87% on the LA subset and 2.58% on the DF. However, benchmark accuracy is no guarantee of robustness in real-world scenarios. This paper investigates the effectiveness of utilizing room impulse responses (RIRs) to enhance fake speech and increase their likelihood of evading fake speech detection systems. Our findings reveal that this simple approach significantly improves the evasion rate, doubling the SOTA system's EER. To counter this type of attack, We augmented training data with a large-scale synthetic simulated RIR dataset. The results demonstrate significant improvement on both reverberated fake speech and original samples, reducing DF task EER to 2.13%.
【27】 Video-to-Audio Generation with Fine-grained Temporal Semantics
标题: 具有细粒度时间语义的视频到音频生成
作者:Yuchen Hu,Yu Gu,Chenxing Li,Rilin Chen,Dong Yu
链接:点击下载PDF文件
摘要:随着AIGC的最新进展,视频生成在学术界和工业界都获得了激增的研究兴趣(例如,Sora)。然而,它仍然是一个挑战,以产生时间上对齐的音频同步所生成的视频,考虑到复杂的语义信息包括在后者。在这项工作中,受到最近文本到音频(TTA)生成成功的启发,我们首先研究了基于潜在扩散模型(LDM)的视频到音频(VTA)生成框架。与VTA的最新探索类似,我们的初步结果也显示了LDM在VTA任务中的巨大潜力,但它仍然存在次优的时间对齐。为此,我们建议增强与帧级语义信息的VTA的时间对齐。利用最近流行的接地段任何模型(接地SAM),我们可以提取视频帧中的细粒度语义,使VTA产生更好的对齐音频信号。大量的实验表明,我们的系统的有效性的客观和主观的评价指标,这表明更好的音频质量和细粒度的时间对齐。摘要:With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video, considering the complicated semantic information included in the latter. In this work, inspired by the recent success of text-to-audio (TTA) generation, we first investigate the video-to-audio (VTA) generation framework based on latent diffusion model (LDM). Similar to latest pioneering exploration in VTA, our preliminary results also show great potentials of LDM in VTA task, but it still suffers from sub-optimal temporal alignment. To this end, we propose to enhance the temporal alignment of VTA with frame-level semantic information. With the recently popular grounding segment anything model (Grounding SAM), we can extract the fine-grained semantics in video frames to enable VTA to produce better-aligned audio signal. Extensive experiments demonstrate the effectiveness of our system on both objective and subjective evaluation metrics, which shows both better audio quality and fine-grained temporal alignment.
【28】 Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
标题: 稳健的视听语音增强:利用高级后处理纠正复杂环境中的误分配
作者:Wenze Ren,Kuo-Hsuan Hung,Chao Rong,YouJin Li,Hsin-Min Wang,Tsao Yu
链接:点击下载PDF文件
摘要:针对音视频语音增强(AVSE)系统中普遍存在的视频质量差、训练和测试数据不匹配等问题,提出了一种新的语音增强方法。我们引入了一个后处理分类器(PPC)来纠正这些错误的输出,确保增强的语音准确地对应于预期的扬声器。在PPC训练中,我们还采用了混合策略来提高其鲁棒性。在AVSE-challenge数据集上的实验结果表明,将PPC集成到AVSE模型中可以显着提高AVSE性能,并且将PPC与使用置换不变训练(PIT)训练的AVSE模型相结合可以产生最佳性能。所提出的方法大大优于基线模型的大幅度。这项工作突出了在各种模式和架构中更广泛应用的潜力,为该领域的未来研究提供了一个有希望的方向。摘要:This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched training and test data. We introduce a post-processing classifier (PPC) to rectify these erroneous outputs, ensuring that the enhanced speech corresponds accurately to the intended speaker. We also adopt a mixup strategy in PPC training to improve its robustness. Experimental results on the AVSE-challenge dataset show that integrating PPC into the AVSE model can significantly improve AVSE performance, and combining PPC with the AVSE model trained with permutation invariant training (PIT) yields the best performance. The proposed method substantially outperforms the baseline model by a large margin. This work highlights the potential for broader applications across various modalities and architectures, providing a promising direction for future research in this field.
【29】 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
标题: 无监督单词发现:使用集群进行边界检测与动态编程
作者:Simon Malan,Benjamin van Niekerk,Herman Kamper
备注:3 figures, 3 tables
链接:点击下载PDF文件
摘要:我们看看长期存在的问题,分割成词样的部分和聚类这些词的语音。之前的几种方法使用评分模型与动态规划相结合来找到最佳分割。在这里,我们提出了一个更简单的策略:我们预测词边界使用相邻的自监督功能之间的相异性,然后我们聚类预测段构建一个词典。为了进行公平的比较,我们更新了旧的ES-KMeans动态规划方法,具有更好的功能和边界约束。在五种语言的ZeroSpeech基准测试中,与新的ES-KMeans+方法相比,我们的简单方法给出了类似的最先进的结果,同时速度快了近五倍。摘要:We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. Here we propose a much simpler strategy: we predict word boundaries using the dissimilarity between adjacent self-supervised features, then we cluster the predicted segments to construct a lexicon. For a fair comparison, we update the older ES-KMeans dynamic programming method with better features and boundary constraints. On the five-language ZeroSpeech benchmarks, our simple approach gives similar state-of-the-art results compared to the new ES-KMeans+ method, while being almost five times faster.
【30】 A Feature Engineering Approach for Literary and Colloquial Tamil Speech Classification using 1D-CNN
标题: 使用1D-CNN进行文学和口语泰米尔语语音分类的特征工程方法
作者:M. Nanmalar,S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
链接:点击下载PDF文件
摘要:在理想的人机交互(HCI)中,大多数用户更喜欢口语形式的语言,因为这是他们日常对话中使用的形式。然而,也有一个不可否认的必要性,以保持正式的文学形式。通过接受新的和保留旧的,既可以为普通人服务(实用性),也可以为语言本身服务(保护)。因此,理想的情况是计算机能够根据需要接受、处理和交谈两种形式的语言。为了解决这个问题,首先需要识别输入语音的形式,在目前的工作中,输入语音的形式介于泰米尔语的书面和口头之间。这样的前端系统必须包括一个简单、有效和轻量级的分类器,该分类器在一些有效的特征上进行训练,这些特征能够捕获语音信号的潜在模式。为了实现这一点,提出了一种一维卷积神经网络(1D-CNN),它可以学习随时间变化的特征包络。该网络最初在选定数量的手工特征上进行训练,然后在Mel频率倒谱系数(MFCC)上进行比较。选择手工制作的特征来解决语音的各个方面,例如频谱和时间特征、韵律和语音质量。通过考虑十个平行的话语和观察每个特征相对于时间的趋势,初步分析的功能。使用手工特征训练的1D-CNN的F1得分为0.9803,而在MFCC上训练的1D-CNN的F1得分为0.9895。在此基础上,对特征切除和特征组合进行了探索。当将特征消融研究中排名最高的手工特征与MFCC相结合时,它们提供了最佳结果,F1得分为0.9946。摘要:In ideal human computer interaction (HCI), the colloquial form of a language would be preferred by most users, since it is the form used in their day-to-day conversations. However, there is also an undeniable necessity to preserve the formal literary form. By embracing the new and preserving the old, both service to the common man (practicality) and service to the language itself (conservation) can be rendered. Hence, it is ideal for computers to have the ability to accept, process, and converse in both forms of the language, as required. To address this, it is first necessary to identify the form of the input speech, which in the current work is between literary and colloquial Tamil speech. Such a front-end system must consist of a simple, effective, and lightweight classifier that is trained on a few effective features that are capable of capturing the underlying patterns of the speech signal. To accomplish this, a one-dimensional convolutional neural network (1D-CNN) that learns the envelope of features across time, is proposed. The network is trained on a select number of handcrafted features initially, and then on Mel frequency cepstral coefficients (MFCC) for comparison. The handcrafted features were selected to address various aspects of speech such as the spectral and temporal characteristics, prosody, and voice quality. The features are initially analyzed by considering ten parallel utterances and observing the trend of each feature with respect to time. The proposed 1D-CNN, trained using the handcrafted features, offers an F1 score of 0.9803, while that trained on the MFCC offers an F1 score of 0.9895. In light of this, feature ablation and feature combination are explored. When the best ranked handcrafted features, from the feature ablation study, are combined with the MFCC, they offer the best results with an F1 score of 0.9946.
【31】 Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
标题: 通过可靠性加权,使用可穿戴麦克风阵列改善动态环境的到达方向估计
作者:Daniel A. Mitchell,Boaz Rafaely,Anurag Kumar,Vladimir Tourbabin
链接:点击下载PDF文件
摘要:室内多个说话人的波达方向估计是一项重要的研究课题,具有广泛的应用前景。特别是,具有移动扬声器、混响和噪声的挑战性环境导致当前方法的显著性能下降。为了更好地理解影响性能的因素并改进现有方法,本文研究了在噪声、动态和混响环境下,采用可穿戴麦克风阵列的局部空间域距离(LSDD)算法的改进版本对多说话人波达方向(DOA)估计的影响。这项研究利用了最近发表的EasyCom语音数据集,使用安装在眼镜上的可穿戴麦克风阵列记录。虽然原始的LSDD算法在静态环境中表现出很强的性能,但在EasyCom数据集的动态设置中,其有效性显着降低。几个增强的LSDD算法的开发后,全面的性能和系统分析,使这些具有挑战性的条件下,改进的DOA估计。这些改进包括将加权可靠性方法和引入一个新的质量措施,可靠地识别更准确的DOA估计,从而提高了算法在具有挑战性的环境中的鲁棒性和准确性。摘要:Direction-of-arrival estimation of multiple speakers in a room is an important task for a wide range of applications. In particular, challenging environments with moving speakers, reverberation and noise, lead to significant performance degradation for current methods. With the aim of better understanding factors affecting performance and improving current methods, in this paper multi-speaker direction-of-arrival (DOA) estimation is investigated using a modified version of the local space domain distance (LSDD) algorithm in a noisy, dynamic and reverberant environment employing a wearable microphone array. This study utilizes the recently published EasyCom speech dataset, recorded using a wearable microphone array mounted on eyeglasses. While the original LSDD algorithm demonstrates strong performance in static environments, its efficacy significantly diminishes in the dynamic settings of the EasyCom dataset. Several enhancements to the LSDD algorithm are developed following a comprehensive performance and system analysis, which enable improved DOA estimation under these challenging conditions. These improvements include incorporating a weighted reliability approach and introducing a new quality measure that reliably identifies the more accurate DOA estimates, thereby enhancing both the robustness and accuracy of the algorithm in challenging environments.
【32】 Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
标题: 复仇者联盟集结:抑郁症检测的非语义特征融合
作者:Orchid Chetia Phukan,Swarup Ranjan Behera,Shubham Singh,Muskaan Singh,Vandana Rajan,Arun Balaji Buduru,Rajesh Sharma,S. R. Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项研究中,我们解决了从语音中检测抑郁症的挑战,重点是非语义特征(NSFs)捕捉抑郁症的微妙标记的潜力。虽然先前的研究已经利用了这项任务的各种功能,但从为非语义任务设计的预训练模型(PTM)中提取的NSF,如非语言语音处理(TRILLsson),说话人识别(x-vector)和情感识别(emoHuBERT),已经显示出显着的前景。然而,将这些不同的功能结合起来的潜力尚未得到充分探索。在这项工作中,我们证明了NSFs的合并导致互补行为,从而增强了抑郁症检测性能。此外,为了我们的目的,我们引入了一个简单的新框架,FuSeR,旨在有效地结合这些功能。我们的研究结果表明,FuSeR优于利用单个NSF以及基线融合技术的模型,并在E-DAIC基准测试中获得了最先进的(SOTA)性能,RMSE为5.51,MAE为4.48,将其确立为抑郁症检测的鲁棒方法。摘要:In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. While prior research has leveraged various features for this task, NSFs-extracted from pre-trained models (PTMs) designed for non-semantic tasks such as paralinguistic speech processing (TRILLsson), speaker recognition (x-vector), and emotion recognition (emoHuBERT)-have shown significant promise. However, the potential of combining these diverse features has not been fully explored. In this work, we demonstrate that the amalgamation of NSFs results in complementary behavior, leading to enhanced depression detection performance. Furthermore, to our end, we introduce a simple novel framework, FuSeR, designed to effectively combine these features. Our results show that FuSeR outperforms models utilizing individual NSFs as well as baseline fusion techniques and obtains state-of-the-art (SOTA) performance in E-DAIC benchmark with RMSE of 5.51 and MAE of 4.48, establishing it as a robust approach for depression detection.
【33】 Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
标题: 单独坚强,共同坚强:协同情态约束基础模型与非言语情感识别的最佳传输
作者:Orchid Chetia Phukan,Mohd Mujtaba Akhtar,Girish,Swarup Ranjan Behera,Sishir Kalita,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项研究中,我们调查多模态基础模型(MFM)的情感识别从非语言的声音。我们假设,MFM,与他们的联合预训练跨多种形式,将更有效地在非语言的声音情感识别(NVER),更好地解释和区分微妙的情感线索,可能是模糊的音频基础模型(AFMs)。为了验证我们的假设,我们从最先进的(SOTA)MFMs和AFMs中提取表示,并在基准NVER数据集上对其进行评估。我们还研究了结合选定的基础模型表示来增强NVER的潜力,进一步受到语音识别和音频deepfake检测研究的启发。为了实现这一点,我们提出了一个框架,称为MATA(通过运输注意力的内部通道对齐)。通过MATA加上FM的组合:ImageBind和ImageBind,我们报告了ASVP-ESD,JNV和VIVAE数据集相对于单个FM和基线融合技术的最高性能,准确率为76.47%,77.40%,75.12%,F1分数为70.35%,76.19%,74.63%,并报告了基准数据集的SOTA。摘要:In this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal sounds emotion recognition (NVER) by better interpreting and differentiating subtle emotional cues that may be ambiguous in audio-only foundation models (AFMs). To validate our hypothesis, we extract representations from state-of-the-art (SOTA) MFMs and AFMs and evaluated them on benchmark NVER datasets. We also investigate the potential of combining selected foundation model representations to enhance NVER further inspired by research in speech recognition and audio deepfake detection. To achieve this, we propose a framework called MATA (Intra-Modality Alignment through Transport Attention). Through MATA coupled with the combination of MFMs: LanguageBind and ImageBind, we report the topmost performance with accuracies of 76.47%, 77.40%, 75.12% and F1-scores of 70.35%, 76.19%, 74.63% for ASVP-ESD, JNV, and VIVAE datasets against individual FMs and baseline fusion techniques and report SOTA on the benchmark datasets.
【34】 Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
标题: 音乐基金会模型在演唱声音Deepfake检测方面更好吗? 更好地使用Speech Foundation模型来简化它们
作者:Orchid Chetia Phukan,Sarthak Jain,Swarup Ranjan Behera,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项研究中,我们第一次广泛调查了音乐基础模型(MFM)或语音基础模型(SFM)是否更好地用于歌声深度假检测(SVDD),这最近引起了研究界的关注。为此,我们对最先进的(SOTA)MFM(MERT变体和music 2 vec)和SFM(针对一般语音表示学习和说话人识别进行预训练)进行了全面的比较研究。我们表明,说话人识别SFM表示在所有基础模型(FM)中表现最好,这种表现可以归因于其在捕获音高,音调,强度等方面的更高功效,存在于歌声中的特征。为了我们的目的,我们还探讨了融合功能模块,利用其互补的行为,以改善SVDD,我们提出了一个新的框架,FIONA相同。使用FIONA,通过x矢量(说话人识别SFM)和MERT-v1- 330 M(MFM)的同步,我们报告了最佳性能,最低等错误率(EER)为13.74%,击败了所有单独的FM以及基线FM融合,并实现了SOTA结果。摘要:In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detection (SVDD), which has recently attracted attention in the research community. For this, we perform a comprehensive comparative study of state-of-the-art (SOTA) MFMs (MERT variants and music2vec) and SFMs (pre-trained for general speech representation learning as well as speaker recognition). We show that speaker recognition SFM representations perform the best amongst all the foundation models (FMs), and this performance can be attributed to its higher efficacy in capturing the pitch, tone, intensity, etc, characteristics present in singing voices. To our end, we also explore the fusion of FMs for exploiting their complementary behavior for improved SVDD, and we propose a novel framework, FIONA for the same. With FIONA, through the synchronization of x-vector (speaker recognition SFM) and MERT-v1-330M (MFM), we report the best performance with the lowest Equal Error Rate (EER) of 13.74 %, beating all the individual FMs as well as baseline FM fusions and achieving SOTA results.
【35】 Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
标题: Codec-SURB @ SYS 2024:神经音频编解码器模型的轻量级基准
作者:Haibin Wu,Xuanjun Chen,Yi-Cheng Lin,Kaiwei Chang,Jiawei Du,Ke-Han Lu,Alexander H. Liu,Ho-Lam Chung,Yuan-Kuei Wu,Dongchao Yang,Songxiang Liu,Yi-Chiao Wu,Xu Tan,James Glass,Shinji Watanabe,Hung-yi Lee
链接:点击下载PDF文件
摘要:神经音频编解码器模型变得越来越重要,因为它们作为音频的标记器,能够实现有效的传输或促进语音语言建模。理想的神经音频编解码器即使在低比特率下也应该保持内容、语音、扬声器特征和音频信息。最近,已经提出了许多先进的神经编解码器模型。然而,编解码器模型通常在不同的实验条件下进行测试。因此,我们在2024年国际编码标准大会上推出了Codec-SUPERB挑战赛,旨在促进现有编解码器模型之间的公平和轻量级比较,并激励该领域的进步。这一挑战将代表性的语音应用和客观指标结合在一起,并仔细选择免许可证的数据集,将其抽样为小集合,以降低评估计算成本。本文介绍了挑战的规则,数据集,五个参与者系统,结果和发现。摘要:Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 2024, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge's rules, datasets, five participant systems, results, and findings.
【36】 Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
标题: 半侵入式音频评估:将非侵入式评估视为多模式文本预测任务
作者:Jozef Coldenhoff,Milos Cernak
链接:点击下载PDF文件
摘要:人类对音频的评估具有独特的能力,可以在混合信号中处理特定的源。模仿这种人类的能力,我们提出了一个半侵入式的评估,我们框架的音频评估任务作为一个文本预测任务与音频文本输入。为此,我们利用多模态PENGI模型的指令微调。我们使用真实和模拟数据对语音和音乐进行MOS预测的实验表明,平均而言,所提出的方法优于在单个任务上操作的基线。为了证明模型的可生成性,我们提出了一种新的半侵入式SNR估计器,能够估计任意信号类的SNR在不同类别的信号的混合物。摘要:Assessment of audio by humans possesses the unique ability to attend to specific sources in a mixture of signals. Mimicking this human ability, we propose a semi-intrusive assessment where we frame the audio assessment task as a text prediction task with audio-text input. To this end we leverage instruction fine-tuning of the multi-modal PENGI model. Our experiments on MOS prediction for speech and music using both real and simulated data show that the proposed method, on average, outperforms baselines that operate on a single task. To justify the model generability, we propose a new semi-intrusive SNR estimator that is able to estimate the SNR of arbitrary signal classes in a mixture of signals with different classes.
【37】 Zero-shot Cross-lingual Voice Transfer for TTS
标题: 针对TTC的零攻击跨语言语音传输
作者:Fadi Biadsy,Youzheng Chen,Isaac Elias,Kyle Kastner,Gary Wang,Andrew Rosenberg,Bhuvana Ramabhadran
备注:Submitted to ICASSP
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个zero-shot语音传输(VT)模块,可以无缝地集成到一个多语言的文本到语音(TTS)系统传输个人的语音跨语言。我们建议的VT模块包括一个扬声器编码器,处理参考语音,瓶颈层,和残余适配器,连接到预先存在的TTS层。我们比较了这些组件的各种配置的性能,并报告了不同语言的平均意见得分(MOS)和说话人相似性。使用一个单一的英语参考语音每个扬声器,我们实现了在九个目标语言的平均语音传输相似性得分为73%。声音特征对个体身份的建构和感知有着重要的作用。由于身体或神经疾病而失去声音,可能会导致深刻的失落感,影响一个人的核心身份。作为一个案例研究,我们证明,我们的方法不仅可以传输典型的语音,但也恢复与构音障碍的个人的声音,即使只有非典型的语音样本-对于那些谁从来没有典型的语音或银行他们的声音有价值的实用程序。跨语言的典型音频样本,以及演示构音障碍说话者语音恢复的视频可以在这里找到(google.github.io tacotron publications zero_shot_voice_transfer)。摘要:In this paper, we introduce a zero-shot Voice Transfer (VT) module that can be seamlessly integrated into a multi-lingual Text-to-speech (TTS) system to transfer an individual's voice across languages. Our proposed VT module comprises a speaker-encoder that processes reference speech, a bottleneck layer, and residual adapters, connected to preexisting TTS layers. We compare the performance of various configurations of these components and report Mean Opinion Score (MOS) and Speaker Similarity across languages. Using a single English reference speech per speaker, we achieve an average voice transfer similarity score of 73% across nine target languages. Vocal characteristics contribute significantly to the construction and perception of individual identity. The loss of one's voice, due to physical or neurological conditions, can lead to a profound sense of loss, impacting one's core identity. As a case study, we demonstrate that our approach can not only transfer typical speech but also restore the voices of individuals with dysarthria, even when only atypical speech samples are available - a valuable utility for those who have never had typical speech or banked their voice. Cross-lingual typical audio samples, plus videos demonstrating voice restoration for dysarthric speakers are available here (google.github.io tacotron publications zero_shot_voice_transfer).
【38】 GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
标题: GT Singer:全球多技术歌唱数据库,为所有歌唱任务提供现实音乐分数
作者:Yu Zhang,Changhao Pan,Wenxiang Guo,Ruiqi Li,Zhiyuan Zhu,Jialei Wang,Wenhao Xu,Jingyu Lu,Zhiqing Hong,Chuxin Wang,LiChao Zhang,Jinzheng He,Ziyue Jiang,Yuxin Chen,Chen Yang,Jiecheng Zhou,Xinyu Cheng,Zhou Zhao
备注:under processing
链接:点击下载PDF文件
摘要:高质量和多任务的歌唱数据集的稀缺严重阻碍了多样化可控和个性化歌唱任务的发展,因为现有的歌唱数据集质量低,语言和歌手的多样性有限,缺乏多技术信息和真实的乐谱,以及任务适用性差。为了解决这些问题,我们提出了 textbf{GTSinger},一个大型的 textbf{G}库,多 textbf{T}技术,免费使用,高质量的歌唱语料库与现实的乐谱,设计用于所有的歌唱任务,以及它的基准。特别是,(1)我们收集了80.59小时的高质量歌唱声音,形成了最大的录音歌唱数据集;(2)20名专业歌手,横跨9种广泛使用的语言,提供了不同的音色和风格;(3)我们提供了六种常用歌唱技术的受控比较和音素级注释,帮助技术建模和控制;(4)GTSinger提供逼真的乐谱,辅助真实世界的音乐创作;(5)歌声伴随着手动音素到音频对齐,全球风格标签,以及16.16小时的配对语音,用于各种歌唱任务。此外,为了方便使用GTSinger,我们进行了四个基准实验:技术可控的歌唱声音合成,技术识别,风格转移,语音到歌唱转换。语料库和演示可以在http: gtsinger.github.io上找到。我们在https: huggingface.co datasets GTSinger GTSinger和https: github.com GTSinger GTSinger上提供数据集和处理数据和进行基准测试的代码。摘要:The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and realistic music scores, and poor task suitability. To tackle these problems, we present textbf{GTSinger}, a large textbf{G}lobal, multi- textbf{T}echnique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks. Particularly, (1) we collect 80.59 hours of high-quality singing voices, forming the largest recorded singing dataset; (2) 20 professional singers across nine widely spoken languages offer diverse timbres and styles; (3) we provide controlled comparison and phoneme-level annotations of six commonly used singing techniques, helping technique modeling and control; (4) GTSinger offers realistic music scores, assisting real-world musical composition; (5) singing voices are accompanied by manual phoneme-to-audio alignments, global style labels, and 16.16 hours of paired speech for various singing tasks. Moreover, to facilitate the use of GTSinger, we conduct four benchmark experiments: technique-controllable singing voice synthesis, technique recognition, style transfer, and speech-to-singing conversion. The corpus and demos can be found at http: gtsinger.github.io. We provide the dataset and the code for processing data and conducting benchmarks at https: huggingface.co datasets GTSinger GTSinger and https: github.com GTSinger GTSinger.
eess.AS音频处理
【1】 CA-MHFA: A Context-Aware Multi-Head Factorized Attentive Pooling for SSL-Based Speaker Verification标题: CA-MHFA:用于基于SSL的说话人验证的上下文感知多头因子化注意力池
作者:Junyi Peng,Ladislav Mošner,Lin Zhang,Oldřich Plchot,Themos Stafylakis,Lukáš Burget,Jan Černocký
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:近年来,用于说话人确认(SV)的自监督学习(SSL)模型受到了极大的关注。然而,现有的基于SSL的SV系统往往很难捕捉本地时间依赖性和概括不同的任务。在本文中,我们提出了上下文感知的多头因式分解注意池(CA-MHFA),一个轻量级的框架,结合周围的帧的上下文信息。CA-MHFA利用分组的、可学习的查询来有效地建模上下文依赖关系,同时通过在组之间共享键和值来保持效率。在VoxCeleb数据集上的实验结果表明,CA-MHFA在Vox 1-O、Vox 1-E和Vox 1-H上的EER分别为0.42 %、0.48 %和0.96 %,优于WavLM-TDNN等复杂模型,参数较少,收敛速度更快。此外,CA-MHFA在多个SSL模型和任务中表现出强大的泛化能力,包括情感识别和反欺骗,突出了其鲁棒性和多功能性。摘要:Self-supervised learning (SSL) models for speaker verification (SV) have gained significant attention in recent years. However, existing SSL-based SV systems often struggle to capture local temporal dependencies and generalize across different tasks. In this paper, we propose context-aware multi-head factorized attentive pooling (CA-MHFA), a lightweight framework that incorporates contextual information from surrounding frames. CA-MHFA leverages grouped, learnable queries to effectively model contextual dependencies while maintaining efficiency by sharing keys and values across groups. Experimental results on the VoxCeleb dataset show that CA-MHFA achieves EERs of 0.42 %, 0.48 %, and 0.96 % on Vox1-O, Vox1-E, and Vox1-H, respectively, outperforming complex models like WavLM-TDNN with fewer parameters and faster convergence. Additionally, CA-MHFA demonstrates strong generalization across multiple SSL models and tasks, including emotion recognition and anti-spoofing, highlighting its robustness and versatility.
【2】 LlamaPartialSpoof: An LLM-Driven Fake Speech Dataset Simulating Disinformation Generation
标题: LlamaPartialSpoof:一个LLM驱动的模拟虚假信息生成的假语音数据集
作者:Hieu-Thi Luong,Haoyang Li,Lin Zhang,Kong Aik Lee,Eng Siong Chng
备注:5 pages, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:以前的假语音数据集是从防御者的角度构建的,以开发对策(CM)系统,而不考虑攻击者的不同动机。为了更好地与现实生活中的场景保持一致,我们创建了LlamaPartialSpoof,这是一个130小时的数据集,包含完全和部分虚假的语音,使用大型语言模型(LLM)和语音克隆技术来评估CM的鲁棒性。通过检查对攻击者和防御者都有价值的信息,我们发现了当前CM系统中的几个关键漏洞,可以利用这些漏洞来提高攻击成功率,包括对某些文本到语音模型或拼接方法的偏见。我们的实验结果表明,目前的虚假语音检测系统的斗争,以推广到看不见的情况下,实现了24.44%的等错误率的最佳性能。摘要:Previous fake speech datasets were constructed from a defender's perspective to develop countermeasure (CM) systems without considering diverse motivations of attackers. To better align with real-life scenarios, we created LlamaPartialSpoof, a 130-hour dataset contains both fully and partially fake speech, using a large language model (LLM) and voice cloning technologies to evaluate the robustness of CMs. By examining information valuable to both attackers and defenders, we identify several key vulnerabilities in current CM systems, which can be exploited to enhance attack success rates, including biases toward certain text-to-speech models or concatenation methods. Our experimental results indicate that current fake speech detection system struggle to generalize to unseen scenarios, achieving a best performance of 24.44% equal error rate.
【3】 Room Impulse Responses help attackers to evade Deep Fake Detection
标题: 房间冲动响应帮助攻击者逃避深度伪造检测
作者:Hieu-Thi Luong,Duc-Tuan Truong,Kong Aik Lee,Eng Siong Chng
备注:7 pages, to be presented at SLT 2024
链接:点击下载PDF文件
摘要:ASVspoof 2021基准测试是一个广泛使用的反欺骗评估框架,由两个子集组成:逻辑访问(LA)和Deepfake(DF),具有不同编码特征和压缩伪影的样本。值得注意的是,当前最先进的(SOTA)系统拥有令人印象深刻的性能,在LA子集上实现了0.87%的等错误率(EER),在DF上实现了2.58%的等错误率(EER)。然而,基准测试的准确性并不能保证真实场景中的鲁棒性。本文研究了利用房间脉冲响应(RIR)来增强假语音和增加其逃避假语音检测系统的可能性的有效性。我们的研究结果表明,这种简单的方法显着提高了逃避率,使SOTA系统的EER翻了一番。为了应对这种类型的攻击,我们使用大规模的合成 模拟RIR数据集来增强训练数据。结果表明,混响假语音和原始样本的显着改善,降低DF任务EER为2.13%。摘要:The ASVspoof 2021 benchmark, a widely-used evaluation framework for anti-spoofing, consists of two subsets: Logical Access (LA) and Deepfake (DF), featuring samples with varied coding characteristics and compression artifacts. Notably, the current state-of-the-art (SOTA) system boasts impressive performance, achieving an Equal Error Rate (EER) of 0.87% on the LA subset and 2.58% on the DF. However, benchmark accuracy is no guarantee of robustness in real-world scenarios. This paper investigates the effectiveness of utilizing room impulse responses (RIRs) to enhance fake speech and increase their likelihood of evading fake speech detection systems. Our findings reveal that this simple approach significantly improves the evasion rate, doubling the SOTA system's EER. To counter this type of attack, We augmented training data with a large-scale synthetic simulated RIR dataset. The results demonstrate significant improvement on both reverberated fake speech and original samples, reducing DF task EER to 2.13%.
【4】 Video-to-Audio Generation with Fine-grained Temporal Semantics
标题: 具有细粒度时间语义的视频到音频生成
作者:Yuchen Hu,Yu Gu,Chenxing Li,Rilin Chen,Dong Yu
链接:点击下载PDF文件
摘要:随着AIGC的最新进展,视频生成在学术界和工业界都获得了激增的研究兴趣(例如,Sora)。然而,它仍然是一个挑战,以产生时间上对齐的音频同步所生成的视频,考虑到复杂的语义信息包括在后者。在这项工作中,受到最近成功的文本到音频(TTA)生成的启发,我们首先研究了基于潜在扩散模型(LDM)的视频到音频(VTA)生成框架。与VTA的最新探索类似,我们的初步结果也显示了LDM在VTA任务中的巨大潜力,但它仍然存在次优的时间对齐。为此,我们建议增强与帧级语义信息的VTA的时间对齐。利用最近流行的接地段任何模型(接地SAM),我们可以提取视频帧中的细粒度语义,使VTA产生更好的对齐音频信号。大量的实验表明,我们的系统的有效性的客观和主观的评价指标,这表明更好的音频质量和细粒度的时间对齐。摘要:With recent advances of AIGC, video generation have gained a surge of research interest in both academia and industry (e.g., Sora). However, it remains a challenge to produce temporally aligned audio to synchronize the generated video, considering the complicated semantic information included in the latter. In this work, inspired by the recent success of text-to-audio (TTA) generation, we first investigate the video-to-audio (VTA) generation framework based on latent diffusion model (LDM). Similar to latest pioneering exploration in VTA, our preliminary results also show great potentials of LDM in VTA task, but it still suffers from sub-optimal temporal alignment. To this end, we propose to enhance the temporal alignment of VTA with frame-level semantic information. With the recently popular grounding segment anything model (Grounding SAM), we can extract the fine-grained semantics in video frames to enable VTA to produce better-aligned audio signal. Extensive experiments demonstrate the effectiveness of our system on both objective and subjective evaluation metrics, which shows both better audio quality and fine-grained temporal alignment.
【5】 Robust Audio-Visual Speech Enhancement: Correcting Misassignments in Complex Environments with Advanced Post-Processing
标题: 稳健的视听语音增强:利用高级后处理纠正复杂环境中的误分配
作者:Wenze Ren,Kuo-Hsuan Hung,Chao Rong,YouJin Li,Hsin-Min Wang,Tsao Yu
链接:点击下载PDF文件
摘要:针对音视频语音增强(AVSE)系统中普遍存在的视频质量差、训练和测试数据不匹配等问题,提出了一种新的语音增强方法。我们引入了一个后处理分类器(PPC)来纠正这些错误的输出,确保增强的语音准确地对应于预期的扬声器。在PPC训练中,我们还采用了混合策略来提高其鲁棒性。在AVSE-challenge数据集上的实验结果表明,将PPC集成到AVSE模型中可以显着提高AVSE性能,并且将PPC与使用置换不变训练(PIT)训练的AVSE模型相结合可以产生最佳性能。所提出的方法大大优于基线模型的大幅度。这项工作突出了在各种模式和架构中更广泛应用的潜力,为该领域的未来研究提供了一个有希望的方向。摘要:This paper addresses the prevalent issue of incorrect speech output in audio-visual speech enhancement (AVSE) systems, which is often caused by poor video quality and mismatched training and test data. We introduce a post-processing classifier (PPC) to rectify these erroneous outputs, ensuring that the enhanced speech corresponds accurately to the intended speaker. We also adopt a mixup strategy in PPC training to improve its robustness. Experimental results on the AVSE-challenge dataset show that integrating PPC into the AVSE model can significantly improve AVSE performance, and combining PPC with the AVSE model trained with permutation invariant training (PIT) yields the best performance. The proposed method substantially outperforms the baseline model by a large margin. This work highlights the potential for broader applications across various modalities and architectures, providing a promising direction for future research in this field.
【6】 Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
标题: 无监督单词发现:使用集群进行边界检测与动态编程
作者:Simon Malan,Benjamin van Niekerk,Herman Kamper
备注:3 figures, 3 tables
链接:点击下载PDF文件
摘要:我们看看长期存在的问题,分割成词样的部分和聚类这些词的语音。之前的几种方法使用评分模型与动态规划相结合来找到最佳分割。在这里,我们提出了一个更简单的策略:我们预测词边界使用相邻的自监督功能之间的相异性,然后我们聚类预测段构建一个词典。为了进行公平的比较,我们更新了旧的ES-KMeans动态规划方法,具有更好的功能和边界约束。在五种语言的ZeroSpeech基准测试中,与新的ES-KMeans+方法相比,我们的简单方法给出了类似的最先进的结果,同时速度快了近五倍。摘要:We look at the long-standing problem of segmenting unlabeled speech into word-like segments and clustering these into a lexicon. Several previous methods use a scoring model coupled with dynamic programming to find an optimal segmentation. Here we propose a much simpler strategy: we predict word boundaries using the dissimilarity between adjacent self-supervised features, then we cluster the predicted segments to construct a lexicon. For a fair comparison, we update the older ES-KMeans dynamic programming method with better features and boundary constraints. On the five-language ZeroSpeech benchmarks, our simple approach gives similar state-of-the-art results compared to the new ES-KMeans+ method, while being almost five times faster.
【7】 A Feature Engineering Approach for Literary and Colloquial Tamil Speech Classification using 1D-CNN
标题: 使用1D-CNN进行文学和口语泰米尔语语音分类的特征工程方法
作者:M. Nanmalar,S. Johanan Joysingh,P. Vijayalakshmi,T. Nagarajan
链接:点击下载PDF文件
摘要:在理想的人机交互(HCI)中,大多数用户更喜欢口语形式的语言,因为这是他们日常对话中使用的形式。然而,也有一个不可否认的必要性,以保持正式的文学形式。通过接受新的和保留旧的,既可以为普通人服务(实用性),也可以为语言本身服务(保护)。因此,理想的情况是计算机能够根据需要接受、处理和交谈两种形式的语言。为了解决这个问题,首先需要识别输入语音的形式,在目前的工作中,输入语音的形式介于泰米尔语的书面和口头之间。这样的前端系统必须包括一个简单、有效和轻量级的分类器,该分类器在一些有效的特征上进行训练,这些特征能够捕获语音信号的潜在模式。为了实现这一点,提出了一种一维卷积神经网络(1D-CNN),它可以学习随时间变化的特征包络。该网络最初在选定数量的手工特征上进行训练,然后在Mel频率倒谱系数(MFCC)上进行比较。选择手工制作的特征来解决语音的各个方面,例如频谱和时间特征、韵律和语音质量。通过考虑十个平行的话语和观察每个特征相对于时间的趋势,初步分析的功能。使用手工特征训练的1D-CNN的F1得分为0.9803,而在MFCC上训练的1D-CNN的F1得分为0.9895。在此基础上,对特征切除和特征组合进行了探索。当将特征消融研究中排名最高的手工特征与MFCC相结合时,它们提供了最佳结果,F1得分为0.9946。摘要:In ideal human computer interaction (HCI), the colloquial form of a language would be preferred by most users, since it is the form used in their day-to-day conversations. However, there is also an undeniable necessity to preserve the formal literary form. By embracing the new and preserving the old, both service to the common man (practicality) and service to the language itself (conservation) can be rendered. Hence, it is ideal for computers to have the ability to accept, process, and converse in both forms of the language, as required. To address this, it is first necessary to identify the form of the input speech, which in the current work is between literary and colloquial Tamil speech. Such a front-end system must consist of a simple, effective, and lightweight classifier that is trained on a few effective features that are capable of capturing the underlying patterns of the speech signal. To accomplish this, a one-dimensional convolutional neural network (1D-CNN) that learns the envelope of features across time, is proposed. The network is trained on a select number of handcrafted features initially, and then on Mel frequency cepstral coefficients (MFCC) for comparison. The handcrafted features were selected to address various aspects of speech such as the spectral and temporal characteristics, prosody, and voice quality. The features are initially analyzed by considering ten parallel utterances and observing the trend of each feature with respect to time. The proposed 1D-CNN, trained using the handcrafted features, offers an F1 score of 0.9803, while that trained on the MFCC offers an F1 score of 0.9895. In light of this, feature ablation and feature combination are explored. When the best ranked handcrafted features, from the feature ablation study, are combined with the MFCC, they offer the best results with an F1 score of 0.9946.
【8】 Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
标题: 通过可靠性加权,使用可穿戴麦克风阵列改善动态环境的到达方向估计
作者:Daniel A. Mitchell,Boaz Rafaely,Anurag Kumar,Vladimir Tourbabin
链接:点击下载PDF文件
摘要:房间中多个扬声器的到达方向估计对于广泛的应用来说是一项重要任务。特别是,具有移动扬声器、混响和噪声的挑战性环境导致当前方法的显著性能下降。为了更好地理解影响性能的因素并改进现有方法,本文研究了在噪声、动态和混响环境下,采用可穿戴麦克风阵列的局部空间域距离(LSDD)算法的改进版本对多说话人波达方向(DOA)估计的影响。这项研究利用了最近发表的EasyCom语音数据集,使用安装在眼镜上的可穿戴麦克风阵列记录。虽然原始的LSDD算法在静态环境中表现出很强的性能,但在EasyCom数据集的动态设置中,其有效性显着降低。几个增强的LSDD算法的开发后,全面的性能和系统分析,使这些具有挑战性的条件下,改进的DOA估计。这些改进包括将加权可靠性方法和引入一个新的质量措施,可靠地识别更准确的DOA估计,从而提高了算法在具有挑战性的环境中的鲁棒性和准确性。摘要:Direction-of-arrival estimation of multiple speakers in a room is an important task for a wide range of applications. In particular, challenging environments with moving speakers, reverberation and noise, lead to significant performance degradation for current methods. With the aim of better understanding factors affecting performance and improving current methods, in this paper multi-speaker direction-of-arrival (DOA) estimation is investigated using a modified version of the local space domain distance (LSDD) algorithm in a noisy, dynamic and reverberant environment employing a wearable microphone array. This study utilizes the recently published EasyCom speech dataset, recorded using a wearable microphone array mounted on eyeglasses. While the original LSDD algorithm demonstrates strong performance in static environments, its efficacy significantly diminishes in the dynamic settings of the EasyCom dataset. Several enhancements to the LSDD algorithm are developed following a comprehensive performance and system analysis, which enable improved DOA estimation under these challenging conditions. These improvements include incorporating a weighted reliability approach and introducing a new quality measure that reliably identifies the more accurate DOA estimates, thereby enhancing both the robustness and accuracy of the algorithm in challenging environments.
【9】 Avengers Assemble: Amalgamation of Non-Semantic Features for Depression Detection
标题: 复仇者联盟集结:抑郁症检测的非语义特征融合
作者:Orchid Chetia Phukan,Swarup Ranjan Behera,Shubham Singh,Muskaan Singh,Vandana Rajan,Arun Balaji Buduru,Rajesh Sharma,S. R. Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项研究中,我们解决了从语音中检测抑郁症的挑战,重点是非语义特征(NSFs)捕捉抑郁症的微妙标记的潜力。虽然先前的研究已经利用了这项任务的各种功能,但从为非语义任务设计的预训练模型(PTM)中提取的NSF,如非语言语音处理(TRILLsson),说话人识别(x-vector)和情感识别(emoHuBERT),已经显示出显着的前景。然而,将这些不同的功能结合起来的潜力尚未得到充分探索。在这项工作中,我们证明了NSFs的合并导致互补行为,从而增强了抑郁症检测性能。此外,为了我们的目的,我们引入了一个简单的新框架,FuSeR,旨在有效地结合这些功能。我们的研究结果表明,FuSeR优于利用单个NSF以及基线融合技术的模型,并在E-DAIC基准测试中获得了最先进的(SOTA)性能,RMSE为5.51,MAE为4.48,将其确立为抑郁症检测的鲁棒方法。摘要:In this study, we address the challenge of depression detection from speech, focusing on the potential of non-semantic features (NSFs) to capture subtle markers of depression. While prior research has leveraged various features for this task, NSFs-extracted from pre-trained models (PTMs) designed for non-semantic tasks such as paralinguistic speech processing (TRILLsson), speaker recognition (x-vector), and emotion recognition (emoHuBERT)-have shown significant promise. However, the potential of combining these diverse features has not been fully explored. In this work, we demonstrate that the amalgamation of NSFs results in complementary behavior, leading to enhanced depression detection performance. Furthermore, to our end, we introduce a simple novel framework, FuSeR, designed to effectively combine these features. Our results show that FuSeR outperforms models utilizing individual NSFs as well as baseline fusion techniques and obtains state-of-the-art (SOTA) performance in E-DAIC benchmark with RMSE of 5.51 and MAE of 4.48, establishing it as a robust approach for depression detection.
【10】 Strong Alone, Stronger Together: Synergizing Modality-Binding Foundation Models with Optimal Transport for Non-Verbal Emotion Recognition
标题: 单独坚强,共同坚强:协同情态约束基础模型与非言语情感识别的最佳传输
作者:Orchid Chetia Phukan,Mohd Mujtaba Akhtar,Girish,Swarup Ranjan Behera,Sishir Kalita,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项研究中,我们调查多模态基础模型(MFM)的情感识别从非语言的声音。我们假设,MFM,与他们的联合预训练跨多种形式,将更有效地在非语言的声音情感识别(NVER),更好地解释和区分微妙的情感线索,可能是模糊的音频基础模型(AFMs)。为了验证我们的假设,我们从最先进的(SOTA)MFMs和AFMs中提取表示,并在基准NVER数据集上对其进行评估。我们还研究了结合选定的基础模型表示来增强NVER的潜力,进一步受到语音识别和音频deepfake检测研究的启发。为了实现这一点,我们提出了一个框架,称为MATA(通过运输注意力的内部通道对齐)。通过MATA加上FM的组合:ImageBind和ImageBind,我们报告了ASVP-ESD,JNV和VIVAE数据集相对于单个FM和基线融合技术的最高性能,准确率为76.47%,77.40%,75.12%,F1分数为70.35%,76.19%,74.63%,并报告了基准数据集的SOTA。摘要:In this study, we investigate multimodal foundation models (MFMs) for emotion recognition from non-verbal sounds. We hypothesize that MFMs, with their joint pre-training across multiple modalities, will be more effective in non-verbal sounds emotion recognition (NVER) by better interpreting and differentiating subtle emotional cues that may be ambiguous in audio-only foundation models (AFMs). To validate our hypothesis, we extract representations from state-of-the-art (SOTA) MFMs and AFMs and evaluated them on benchmark NVER datasets. We also investigate the potential of combining selected foundation model representations to enhance NVER further inspired by research in speech recognition and audio deepfake detection. To achieve this, we propose a framework called MATA (Intra-Modality Alignment through Transport Attention). Through MATA coupled with the combination of MFMs: LanguageBind and ImageBind, we report the topmost performance with accuracies of 76.47%, 77.40%, 75.12% and F1-scores of 70.35%, 76.19%, 74.63% for ASVP-ESD, JNV, and VIVAE datasets against individual FMs and baseline fusion techniques and report SOTA on the benchmark datasets.
【11】 Are Music Foundation Models Better at Singing Voice Deepfake Detection? Far-Better Fuse them with Speech Foundation Models
标题: 音乐基金会模型在演唱声音Deepfake检测方面更好吗? 更好地使用Speech Foundation模型来简化它们
作者:Orchid Chetia Phukan,Sarthak Jain,Swarup Ranjan Behera,Arun Balaji Buduru,Rajesh Sharma,S. R Mahadeva Prasanna
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在这项研究中,我们第一次广泛调查了音乐基础模型(MFM)或语音基础模型(SFM)是否更好地用于歌声深度假检测(SVDD),这最近引起了研究界的关注。为此,我们对最先进的(SOTA)MFM(MERT变体和music 2 vec)和SFM(针对一般语音表示学习和说话人识别进行预训练)进行了全面的比较研究。我们表明,说话人识别SFM表示在所有基础模型(FM)中表现最好,这种表现可以归因于其在捕获音高,音调,强度等方面的更高功效,存在于歌声中的特征。为了我们的目的,我们还探讨了融合功能模块,利用其互补的行为,以改善SVDD,我们提出了一个新的框架,FIONA相同。使用FIONA,通过x矢量(说话人识别SFM)和MERT-v1- 330 M(MFM)的同步,我们报告了最佳性能,最低等错误率(EER)为13.74%,击败了所有单独的FM以及基线FM融合,并实现了SOTA结果。摘要:In this study, for the first time, we extensively investigate whether music foundation models (MFMs) or speech foundation models (SFMs) work better for singing voice deepfake detection (SVDD), which has recently attracted attention in the research community. For this, we perform a comprehensive comparative study of state-of-the-art (SOTA) MFMs (MERT variants and music2vec) and SFMs (pre-trained for general speech representation learning as well as speaker recognition). We show that speaker recognition SFM representations perform the best amongst all the foundation models (FMs), and this performance can be attributed to its higher efficacy in capturing the pitch, tone, intensity, etc, characteristics present in singing voices. To our end, we also explore the fusion of FMs for exploiting their complementary behavior for improved SVDD, and we propose a novel framework, FIONA for the same. With FIONA, through the synchronization of x-vector (speaker recognition SFM) and MERT-v1-330M (MFM), we report the best performance with the lowest Equal Error Rate (EER) of 13.74 %, beating all the individual FMs as well as baseline FM fusions and achieving SOTA results.
【12】 Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models
标题: Codec-SURB @ SYS 2024:神经音频编解码器模型的轻量级基准
作者:Haibin Wu,Xuanjun Chen,Yi-Cheng Lin,Kaiwei Chang,Jiawei Du,Ke-Han Lu,Alexander H. Liu,Ho-Lam Chung,Yuan-Kuei Wu,Dongchao Yang,Songxiang Liu,Yi-Chiao Wu,Xu Tan,James Glass,Shinji Watanabe,Hung-yi Lee
链接:点击下载PDF文件
摘要:神经音频编解码器模型变得越来越重要,因为它们作为音频的标记器,能够实现有效的传输或促进语音语言建模。理想的神经音频编解码器即使在低比特率下也应该保持内容、语音、扬声器特征和音频信息。最近,已经提出了许多先进的神经编解码器模型。然而,编解码器模型通常在不同的实验条件下进行测试。因此,我们在2024年国际编码标准大会上推出了Codec-SUPERB挑战赛,旨在促进现有编解码器模型之间的公平和轻量级比较,并激励该领域的进步。这一挑战将代表性的语音应用和客观指标结合在一起,并仔细选择免许可证的数据集,将其抽样为小集合,以降低评估计算成本。本文介绍了挑战的规则,数据集,五个参与者系统,结果和发现。摘要:Neural audio codec models are becoming increasingly important as they serve as tokenizers for audio, enabling efficient transmission or facilitating speech language modeling. The ideal neural audio codec should maintain content, paralinguistics, speaker characteristics, and audio information even at low bitrates. Recently, numerous advanced neural codec models have been proposed. However, codec models are often tested under varying experimental conditions. As a result, we introduce the Codec-SUPERB challenge at SLT 2024, designed to facilitate fair and lightweight comparisons among existing codec models and inspire advancements in the field. This challenge brings together representative speech applications and objective metrics, and carefully selects license-free datasets, sampling them into small sets to reduce evaluation computation costs. This paper presents the challenge's rules, datasets, five participant systems, results, and findings.
【13】 Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction task
标题: 半侵入式音频评估:将非侵入式评估视为多模式文本预测任务
作者:Jozef Coldenhoff,Milos Cernak
链接:点击下载PDF文件
摘要:人类对音频的评估具有独特的能力,可以在混合信号中处理特定的源。模仿这种人类的能力,我们提出了一个半侵入式的评估,我们框架的音频评估任务作为一个文本预测任务与音频文本输入。为此,我们利用多模态PENGI模型的指令微调。我们使用真实和模拟数据对语音和音乐进行MOS预测的实验表明,平均而言,所提出的方法优于在单个任务上操作的基线。为了证明模型的可生成性,我们提出了一种新的半侵入式SNR估计器,能够估计任意信号类的SNR在不同类别的信号的混合物。摘要:Assessment of audio by humans possesses the unique ability to attend to specific sources in a mixture of signals. Mimicking this human ability, we propose a semi-intrusive assessment where we frame the audio assessment task as a text prediction task with audio-text input. To this end we leverage instruction fine-tuning of the multi-modal PENGI model. Our experiments on MOS prediction for speech and music using both real and simulated data show that the proposed method, on average, outperforms baselines that operate on a single task. To justify the model generability, we propose a new semi-intrusive SNR estimator that is able to estimate the SNR of arbitrary signal classes in a mixture of signals with different classes.
【14】 Zero-shot Cross-lingual Voice Transfer for TTS
标题: 针对TTC的零攻击跨语言语音传输
作者:Fadi Biadsy,Youzheng Chen,Isaac Elias,Kyle Kastner,Gary Wang,Andrew Rosenberg,Bhuvana Ramabhadran
备注:Submitted to ICASSP
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一个zero-shot语音传输(VT)模块,可以无缝地集成到一个多语言的文本到语音(TTS)系统传输个人的语音跨语言。我们建议的VT模块包括一个扬声器编码器,处理参考语音,瓶颈层,和残余适配器,连接到预先存在的TTS层。我们比较了这些组件的各种配置的性能,并报告了不同语言的平均意见得分(MOS)和说话人相似性。使用一个单一的英语参考语音每个扬声器,我们实现了在九个目标语言的平均语音传输相似性得分为73%。声音特征对个体身份的建构和感知有着重要的作用。由于身体或神经疾病而失去声音,可能会导致深刻的失落感,影响一个人的核心身份。作为一个案例研究,我们证明,我们的方法不仅可以传输典型的语音,但也恢复与构音障碍的个人的声音,即使只有非典型的语音样本-对于那些谁从来没有典型的语音或银行他们的声音有价值的实用程序。跨语言的典型音频样本,以及演示构音障碍说话者语音恢复的视频可以在这里找到(google.github.io tacotron publications zero_shot_voice_transfer)。摘要:In this paper, we introduce a zero-shot Voice Transfer (VT) module that can be seamlessly integrated into a multi-lingual Text-to-speech (TTS) system to transfer an individual's voice across languages. Our proposed VT module comprises a speaker-encoder that processes reference speech, a bottleneck layer, and residual adapters, connected to preexisting TTS layers. We compare the performance of various configurations of these components and report Mean Opinion Score (MOS) and Speaker Similarity across languages. Using a single English reference speech per speaker, we achieve an average voice transfer similarity score of 73% across nine target languages. Vocal characteristics contribute significantly to the construction and perception of individual identity. The loss of one's voice, due to physical or neurological conditions, can lead to a profound sense of loss, impacting one's core identity. As a case study, we demonstrate that our approach can not only transfer typical speech but also restore the voices of individuals with dysarthria, even when only atypical speech samples are available - a valuable utility for those who have never had typical speech or banked their voice. Cross-lingual typical audio samples, plus videos demonstrating voice restoration for dysarthric speakers are available here (google.github.io tacotron publications zero_shot_voice_transfer).
【15】 GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks
标题: GT Singer:全球多技术歌唱数据库,为所有歌唱任务提供现实音乐分数
作者:Yu Zhang,Changhao Pan,Wenxiang Guo,Ruiqi Li,Zhiyuan Zhu,Jialei Wang,Wenhao Xu,Jingyu Lu,Zhiqing Hong,Chuxin Wang,LiChao Zhang,Jinzheng He,Ziyue Jiang,Yuxin Chen,Chen Yang,Jiecheng Zhou,Xinyu Cheng,Zhou Zhao
备注:under processing
链接:点击下载PDF文件
摘要:高质量和多任务的歌唱数据集的稀缺严重阻碍了多样化可控和个性化歌唱任务的发展,因为现有的歌唱数据集质量低,语言和歌手的多样性有限,缺乏多技术信息和真实的乐谱,以及任务适用性差。为了解决这些问题,我们提出了 textbf{GTSinger},一个大型的 textbf{G}库,多 textbf{T}技术,免费使用,高质量的歌唱语料库与现实的乐谱,设计用于所有的歌唱任务,以及它的基准。特别是,(1)我们收集了80.59小时的高质量歌唱声音,形成了最大的录音歌唱数据集;(2)20名专业歌手,横跨9种广泛使用的语言,提供了不同的音色和风格;(3)我们提供了六种常用歌唱技术的受控比较和音素级注释,帮助技术建模和控制;(4)GTSinger提供逼真的乐谱,辅助真实世界的音乐创作;(5)歌声伴随着手动音素到音频对齐,全球风格标签,以及16.16小时的配对语音,用于各种歌唱任务。此外,为了方便使用GTSinger,我们进行了四个基准实验:技术可控的歌唱声音合成,技术识别,风格转移,语音到歌唱转换。语料库和演示可以在http: gtsinger.github.io上找到。我们在https: huggingface.co datasets GTSinger GTSinger和https: github.com GTSinger GTSinger上提供数据集和处理数据和进行基准测试的代码。摘要:The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages and singers, absence of multi-technique information and realistic music scores, and poor task suitability. To tackle these problems, we present textbf{GTSinger}, a large textbf{G}lobal, multi- textbf{T}echnique, free-to-use, high-quality singing corpus with realistic music scores, designed for all singing tasks, along with its benchmarks. Particularly, (1) we collect 80.59 hours of high-quality singing voices, forming the largest recorded singing dataset; (2) 20 professional singers across nine widely spoken languages offer diverse timbres and styles; (3) we provide controlled comparison and phoneme-level annotations of six commonly used singing techniques, helping technique modeling and control; (4) GTSinger offers realistic music scores, assisting real-world musical composition; (5) singing voices are accompanied by manual phoneme-to-audio alignments, global style labels, and 16.16 hours of paired speech for various singing tasks. Moreover, to facilitate the use of GTSinger, we conduct four benchmark experiments: technique-controllable singing voice synthesis, technique recognition, style transfer, and speech-to-singing conversion. The corpus and demos can be found at http: gtsinger.github.io. We provide the dataset and the code for processing data and conducting benchmarks at https: huggingface.co datasets GTSinger GTSinger and https: github.com GTSinger GTSinger.
【16】 A Comprehensive Survey with Critical Analysis for Deepfake Speech Detection
标题: Deepfake语音检测的全面调查和批判性分析
作者:Lam Pham,Phat Lam,Tin Nguyen,Hieu Tang,Huyen Nguyen,Alexander Schindler,Canh Vu
备注:Journal preprint
链接:点击下载PDF文件
摘要:由于深度学习的进步,语音生成系统现在为各种现实世界的应用提供动力,例如语音障碍患者的文本到语音,呼叫中心的语音聊天机器人,跨语言语音翻译等。这促使研究团体开发用于检测合成语音的模型(例如,由基于深度学习的模型生成的假语音(Deepfake Speech Detection)任务。由于Deepfake语音检测任务是近年来出现的,因此针对该任务提出的调查论文并不多。此外,针对Deepfake语音检测任务的现有调查倾向于总结用于构建Deepfake语音检测系统的技术,而不是提供全面的分析。这一差距促使我们进行了一次全面的调查,对Deepfake语音检测的挑战和发展进行了批判性分析。我们的调查是创新性的结构,提供了对当前挑战比赛,公共数据集和深度学习技术的深入分析,这些技术提供了增强的解决方案,以解决该领域的现有挑战。根据我们的分析,我们提出了利用和结合特定深度学习技术来提高Deepfake语音检测系统有效性的假设。除了进行调查外,我们还进行了大量的实验来验证这些假设,并为Deepfake语音检测任务提出了一个极具竞争力的模型。鉴于分析和实验结果,我们最终指出了Deepfake语音检测任务潜在且有前途的研究方向。摘要:Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech translation, etc. While these systems can autonomously generate human-like speech and replicate specific voices, they also pose risks when misused for malicious purposes. This motivates the research community to develop models for detecting synthesized speech (e.g., fake speech) generated by deep-learning-based models, referred to as the Deepfake Speech Detection task. As the Deepfake Speech Detection task has emerged in recent years, there are not many survey papers proposed for this task. Additionally, existing surveys for the Deepfake Speech Detection task tend to summarize techniques used to construct a Deepfake Speech Detection system rather than providing a thorough analysis. This gap motivated us to conduct a comprehensive survey, providing a critical analysis of the challenges and developments in Deepfake Speech Detection. Our survey is innovatively structured, offering an in-depth analysis of current challenge competitions, public datasets, and the deep-learning techniques that provide enhanced solutions to address existing challenges in the field. From our analysis, we propose hypotheses on leveraging and combining specific deep learning techniques to improve the effectiveness of Deepfake Speech Detection systems. Beyond conducting a survey, we perform extensive experiments to validate these hypotheses and propose a highly competitive model for the task of Deepfake Speech Detection. Given the analysis and the experimental results, we finally indicate potential and promising research directions for the Deepfake Speech Detection task.
【17】 Adaptive Learning via a Negative Selection Strategy for Few-Shot Bioacoustic Event Detection
标题: 通过负选择策略的自适应学习用于Few-Shot生物声学事件检测
作者:Yaxiong Chen,Xueping Zhang,Yunfei Zi,Shengwu Xiong
链接:点击下载PDF文件
摘要:虽然原型网络(ProtoNet)已经证明了在Few-Shot生物事件检测中的有效性,但仍然存在两个持续的问题。首先,由于缺乏明确注释的阴性样本,很难构建具有代表性的阴性原型。其次,目标生物发声的持续时间在不同的任务中各不相同,这使得模型在所有任务中始终产生最佳结果具有挑战性。为了解决这些问题,我们提出了一种新的自适应学习框架,具有自适应学习损失来指导分类器更新。此外,我们提出了一个否定选择策略,以构建一个更有代表性的否定原型的ProtoNet。所有实验均在DCASE 2023 TASK 5 Few-Shot生物声学事件检测数据集上进行。结果表明,我们提出的方法实现了0.703的F-措施,提高了12.84%。摘要:Although the Prototypical Network (ProtoNet) has demonstrated effectiveness in few-shot biological event detection, two persistent issues remain. Firstly, there is difficulty in constructing a representative negative prototype due to the absence of explicitly annotated negative samples. Secondly, the durations of the target biological vocalisations vary across tasks, making it challenging for the model to consistently yield optimal results across all tasks. To address these issues, we propose a novel adaptive learning framework with an adaptive learning loss to guide classifier updates. Additionally, we propose a negative selection strategy to construct a more representative negative prototype for ProtoNet. All experiments ware performed on the DCASE 2023 TASK5 few-shot bioacoustic event detection dataset. The results show that our proposed method achieves an F-measure of 0.703, an improvement of 12.84%.
【18】 LoVA: Long-form Video-to-Audio Generation
标题: LoVA:长格式视频到音频生成
作者:Xin Cheng,Xihua Wang,Yihan Wu,Yuyue Wang,Ruihua Song
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:视频到音频(V2A)生成对于视频编辑和后处理非常重要,可以为无声视频创建语义对齐的音频。然而,大多数现有的方法集中于为短视频片段(小于10秒)生成短格式音频,而很少关注长格式视频输入的场景。对于当前基于UNet的扩散V2A模型,在处理长格式音频生成时不可避免的问题是最终级联音频内的不一致性。在本文中,我们首先强调了长形式V2A问题的重要性。此外,我们提出了LoVA,一个新的模型,长格式视频到音频生成。事实证明,与现有的自回归模型和基于UNet的扩散模型相比,LoVA基于扩散Transformer(DiT)架构,在生成长格式音频方面更有效。大量的客观和主观实验表明,LoVA在10秒V2A基准测试中实现了相当的性能,并且在长格式视频输入的基准测试中优于所有其他基准。摘要:Video-to-audio (V2A) generation is important for video editing and post-processing, enabling the creation of semantics-aligned audio for silent video. However, most existing methods focus on generating short-form audio for short video segment (less than 10 seconds), while giving little attention to the scenario of long-form video inputs. For current UNet-based diffusion V2A models, an inevitable problem when handling long-form audio generation is the inconsistencies within the final concatenated audio. In this paper, we first highlight the importance of long-form V2A problem. Besides, we propose LoVA, a novel model for Long-form Video-to-Audio generation. Based on the Diffusion Transformer (DiT) architecture, LoVA proves to be more effective at generating long-form audio compared to existing autoregressive models and UNet-based diffusion models. Extensive objective and subjective experiments demonstrate that LoVA achieves comparable performance on 10-second V2A benchmark and outperforms all other baselines on a benchmark with long-form video input.
【19】 GALD-SE: Guided Anisotropic Lightweight Diffusion for Efficient Speech Enhancement
标题: GALD-SE:引导各向异性轻量级扩散,实现高效语音增强
作者:Chengzhong Wang,Jianjun Gu,Dingding Yao,Zelin Qiu,Jiale Zhao,Junfeng Li
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:语音增强的目的是在各种噪声条件下提高语音的清晰度和质量。近年来,扩散模型在语音增强领域受到了广泛的关注,取得了令人瞩目的成果。当前基于扩散的方法用各向同性高斯噪声模糊初始信号,并从先前恢复干净的语音。然而,这些方法通常遭受大量的计算负担。我们认为,低效率源于忽视语音增强不是纯粹的生成任务;它主要涉及降噪和缺失信息的完成,而原始混合物中的干净线索不需要重新生成。在本文中,我们提出了一种方法,在扩散过程中引入具有各向异性指导的噪声,使神经网络能够专注于噪声记录中的干净线索。该方法对各种类型的噪声干扰和语音失真具有鲁棒性,并显著降低了计算量。实验表明,该方法实现了国家的最先进的结果,只有大约450万个参数,显着少于其他扩散方法所需的参数。这有效地缩小了基于扩散和预测语音增强方法之间的模型大小的差距。此外,所提出的方法在非常嘈杂的情况下表现良好,证明了其在极具挑战性的环境中的应用潜力。摘要:Speech enhancement is designed to enhance the intelligibility and quality of speech across diverse noise conditions. Recently, diffusion model has gained lots of attention in speech enhancement area, achieving competitive results. Current diffusion-based methods blur the initial signal with isotropic Gaussian noise and recover clean speech from the prior. However, these methods often suffer from a substantial computational burden. We argue that the inefficiency stems from the oversight that speech enhancement is not purely a generative task; it primarily involves noise reduction and completion of missing information, while the clean clues in the original mixture do not need to be regenerated. In this paper, we propose a method that introduces noise with anisotropic guidance during the diffusion process, allowing the neural network to focus on clean clues within noisy recordings. This approach is robust against various types of noise interference and speech distortion, and significantly reduces the computational load. Experiments demonstrate that the proposed method achieves state-of-the-art results with only approximately 4.5 million parameters, significantly fewer than those required by other diffusion methods. This effectively narrows the gap in model size between diffusion-based and predictive speech enhancement approaches. Additionally, the proposed method performs well in very noisy scenarios, demonstrating its potential for applications in highly challenging environments.
【20】 Blind Spatial Impulse Response Generation from Separate Room- and Scene-Specific Information
标题: 根据单独的房间和场景特定信息生成盲空间脉冲响应
作者:Francesc Lluís,Nils Meyer-Kahlen
链接:点击下载PDF文件
摘要:对于增强现实(AR)中的音频,用户真实声学环境的知识对于渲染无缝融入环境的虚拟声音至关重要。由于声学测量在实际的AR应用中通常是不可行的,因此需要从可用的声源推断出关于房间的信息。然后,可以用相同的房间声学质量来呈现附加的声源。至关重要的是,这些被放置在与可用于估计的源不同的位置。在这里,我们建议使用一个编码器网络,该网络使用对比度损失进行训练,将输入声音映射到仅表示房间特定信息的低维特征空间。然后,基于扩散的空间房间脉冲响应发生器被训练以获取潜在空间并生成新的响应,给定新的源-接收器位置。我们展示了如何在最终输出中考虑房间和位置特定的参数。摘要:For audio in augmented reality (AR), knowledge of the users' real acoustic environment is crucial for rendering virtual sounds that seamlessly blend into the environment. As acoustic measurements are usually not feasible in practical AR applications, information about the room needs to be inferred from available sound sources. Then, additional sound sources can be rendered with the same room acoustic qualities. Crucially, these are placed at different positions than the sources available for estimation. Here, we propose to use an encoder network trained using a contrastive loss that maps input sounds to a low-dimensional feature space representing only room-specific information. Then, a diffusion-based spatial room impulse response generator is trained to take the latent space and generate a new response, given a new source-receiver position. We show how both room- and position-specific parameters are considered in the final output.
【21】 Voice Conversion-based Privacy through Adversarial Information Hiding
标题: 通过对抗性信息隐藏实现基于语音转换的隐私
作者:Jacob J Webber,Oliver Watts,Gustav Eje Henter,Jennifer Williams,Simon King
备注:Accepted for publication in proceedings of 4th symposium on security and privacy in speech communication
链接:点击下载PDF文件
摘要:隐私保护语音转换的目标是只删除语音音频中传达身份信息的属性,而保持其他语音特征不变。本文提出了一种保护隐私的语音转换机制,允许使用对抗性信息隐藏来控制身份承载信息的泄漏。这使得能够在保持源语音特征和修改说话者身份之间进行慎重的权衡。因此,该方法改进了CycleGAN和StarGAN等语音转换技术,这些技术不是为隐私而设计的,这意味着转换后的语音可能会以不可预测的方式泄露个人信息。我们的方法也比ASR-TTS语音转换管道更灵活,ASR-TTS语音转换管道通过设计丢弃与文本内容相关的所有韵律信息。评估结果表明,该系统成功地修改感知扬声器的身份,同时很好地保持源词汇内容。摘要:Privacy-preserving voice conversion aims to remove only the attributes of speech audio that convey identity information, keeping other speech characteristics intact. This paper presents a mechanism for privacy-preserving voice conversion that allows controlling the leakage of identity-bearing information using adversarial information hiding. This enables a deliberate trade-off between maintaining source-speech characteristics and modification of speaker identity. As such, the approach improves on voice-conversion techniques like CycleGAN and StarGAN, which were not designed for privacy, meaning that converted speech may leak personal information in unpredictable ways. Our approach is also more flexible than ASR-TTS voice conversion pipelines, which by design discard all prosodic information linked to textual content. Evaluations show that the proposed system successfully modifies perceived speaker identity whilst well maintaining source lexical content.
【22】 HiFi-Glot: Neural Formant Synthesis with Differentiable Resonant Filters
标题: HiFi-Glot:使用可区分共振过滤器的神经Forces合成
作者:Lauri Juvela,Pablo Pérez Zarazaga,Gustav Eje Henter,Zofia Malisz
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:我们介绍了一个使用语音产生的源过滤器模型的端到端神经语音合成系统。具体来说,我们应用微分谐振滤波器的声门波形产生的神经声码器。其目的是获得一个可控的合成器,类似于经典的共振峰合成,但具有更高的感知质量-填补了当前神经波形发生器的研究空白,并响应语音科学迄今未满足的需求。我们的设置从语音上有意义的语音参数的核心集生成音频,滤波器在合成中提供对共振峰频率共振的直接控制。直接合成控制是重要语音科学实验中可靠刺激产生的关键特征。我们表明,所提出的源滤波器方法给出了比共振峰操纵的行业标准更好的感知质量(即,Praat),同时在共振峰频率控制精度方面具有竞争力。摘要:We introduce an end-to-end neural speech synthesis system that uses the source-filter model of speech production. Specifically, we apply differentiable resonant filters to a glottal waveform generated by a neural vocoder. The aim is to obtain a controllable synthesiser, similar to classic formant synthesis, but with much higher perceptual quality - filling a research gap in current neural waveform generators and responding to hitherto unmet needs in the speech sciences. Our setup generates audio from a core set of phonetically meaningful speech parameters, with the filters providing direct control over formant frequency resonances in synthesis. Direct synthesis control is a key feature for reliable stimulus creation in important speech science experiments. We show that the proposed source-filter method gives better perceptual quality than the industry standard for formant manipulation (i.e., Praat), whilst being competitive in terms of formant frequency control accuracy.
【23】 SongTrans: An unified song transcription and alignment method for lyrics and notes
标题: SongTrans:歌词和音符的统一歌曲转录和对齐方法
作者:Siwei Wu,Jinzheng He,Ruibin Yuan,Haojie Wei,Xipin Wei,Chenghua Lin,Jin Xu,Junyang Lin
链接:点击下载PDF文件
摘要:处理后的数据量是提高歌唱声音合成领域的关键。虽然存在可用于歌词或音符转录任务的工具,但它们都需要相对耗时的预处理数据(例如,声乐和伴奏分离)。此外,这些工具中的大多数都是为了解决一个单一的任务,并努力对齐歌词和音符(即,识别歌词中每个单词的相应音符)。为了应对这些挑战,我们首先通过优化现有工具并注释大量歌曲的歌词-音符对来设计管道。然后,基于标注的数据,我们训练一个统一的SongTrans模型,该模型可以直接转录歌词和音符,同时对齐它们,而不需要对歌曲进行预处理。我们的SongTrans模型由两个模块组成:(1) textbf{自回归模块}预测歌词,以及歌词中每个单词对应的持续时间和音符数量。(2) textbf{非自回归模块}预测音符的音高和持续时间。我们的实验表明,SongTrans实现了最先进的(SOTA)结果在歌词和笔记转录任务。此外,它是第一个能够将歌词与音符对齐的模型。实验结果表明,SongTrans模型可以有效地适应不同类型的歌曲(例如,歌曲与伴奏),展示了其多功能性的现实世界的应用。摘要:The quantity of processed data is crucial for advancing the field of singing voice synthesis. While there are tools available for lyric or note transcription tasks, they all need pre-processed data which is relatively time-consuming (e.g., vocal and accompaniment separation). Besides, most of these tools are designed to address a single task and struggle with aligning lyrics and notes (i.e., identifying the corresponding notes of each word in lyrics). To address those challenges, we first design a pipeline by optimizing existing tools and annotating numerous lyric-note pairs of songs. Then, based on the annotated data, we train a unified SongTrans model that can directly transcribe lyrics and notes while aligning them simultaneously, without requiring pre-processing songs. Our SongTrans model consists of two modules: (1) the textbf{Autoregressive module} predicts the lyrics, along with the duration and note number corresponding to each word in a lyric. (2) the textbf{Non-autoregressive module} predicts the pitch and duration of the notes. Our experiments demonstrate that SongTrans achieves state-of-the-art (SOTA) results in both lyric and note transcription tasks. Furthermore, it is the first model capable of aligning lyrics with notes. Experimental results demonstrate that the SongTrans model can effectively adapt to different types of songs (e.g., songs with accompaniment), showcasing its versatility for real-world applications.
【24】 What Are They Doing? Joint Audio-Speech Co-Reasoning
标题: 他们在做什么?音频-语音联合推理
作者:Yingzhi Wang,Pooneh Mousavi,Artem Ploujnikov,Mirco Ravanelli
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在音频和语音处理中,任务通常集中在音频或语音模态上,即使声音和人类语音都存在于同一音频片段中。最近的听觉大语言模型(ALLM)使得在单个模型中同时处理音频和语音成为可能,从而进一步考虑联合音频语音任务。 在本文中,我们调查如何以及ALLM可以执行联合音频语音处理。具体来说,我们介绍了联合音频语音协同推理(JASCO),一种新的任务,统一的音频和语音处理,严格要求跨两种模式的协同推理。我们发布了一个名为“他们在做什么”的场景推理数据集,并建立了一个联合音频语音基准来评估流行的ALLM的联合推理能力。此外,我们通过分析模型对每种模态的依赖性,对模型的行为提供了更深入的了解。摘要:In audio and speech processing, tasks usually focus on either the audio or speech modality, even when both sounds and human speech are present in the same audio clip. Recent Auditory Large Language Models (ALLMs) have made it possible to process audio and speech simultaneously within a single model, leading to further considerations of joint audio-speech tasks. In this paper, we investigate how well ALLMs can perform joint audio-speech processing. Specifically, we introduce Joint Audio-Speech Co-Reasoning (JASCO), a novel task that unifies audio and speech processing, strictly requiring co-reasoning across both modalities. We release a scene-reasoning dataset called "What Are They Doing" and establish a joint audio-speech benchmark to evaluate the joint reasoning capability of popular ALLMs. Additionally, we provide deeper insights into the models' behaviors by analyzing their dependence on each modality.
【25】 CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
标题: CPD增强Wav2vec2.0:面向课堂环境的噪音稳健语音识别
作者:Ahmed Adel Attia,Dorottya Demszky,Tolulope Ogunremi,Jing Liu,Carol Espy-Wilson
备注:arXiv admin note: substantial text overlap with arXiv:2405.13018
链接:点击下载PDF文件
摘要:创建对课堂条件具有鲁棒性和弹性的自动语音识别(ASR)系统对于开发AI工具以帮助教师和学生至关重要。在这项工作中,我们研究了持续预训练(CPT)在适应Wav2vec2.0课堂领域的功效。我们表明,CPT在这方面是一个强大的工具,并将基于Wav2vec2.0的模型的字错误率(WER)降低了10%以上。更具体地说,CPT提高了模型对不同噪声、麦克风和教室条件的鲁棒性。摘要:Creating Automatic Speech Recognition (ASR) systems that are robust and resilient to classroom conditions is paramount to the development of AI tools to aid teachers and students. In this work, we study the efficacy of continued pretraining (CPT) in adapting Wav2vec2.0 to the classroom domain. We show that CPT is a powerful tool in that regard and reduces the Word Error Rate (WER) of Wav2vec2.0-based models by upwards of 10%. More specifically, CPT improves the model's robustness to different noises, microphones and classroom conditions.
【26】 Self-Supervised Audio-Visual Soundscape Stylization
标题: 自我监督的视听声景风格化
作者:Tingle Li,Renhao Wang,Po-Yao Huang,Andrew Owens,Gopala Anumanchipalli
备注:ECCV 2024
链接:点击下载PDF文件
摘要:语音声音传达了大量关于场景的信息,从而产生从混响到附加环境声音的各种效果。在本文中,我们操纵输入语音的声音,好像它是记录在一个不同的场景,给定的视听条件的例子记录从该场景。我们的模型通过自我监督来学习,利用自然视频包含重复出现的声音事件和纹理的事实。我们从视频中提取音频片段并应用语音增强。然后,我们训练一个潜在的扩散模型来恢复原始语音,使用从视频中其他地方拍摄的另一个视听剪辑作为条件提示。通过这个过程,模型学习将条件示例的声音属性转移到输入语音。我们证明了我们的模型可以使用未标记的野外视频成功训练,并且额外的视觉信号可以提高其声音预测能力。请参阅我们的项目网页视频结果:https: tinglok.netlify.app files avsoundscape 摘要:Speech sounds convey a great deal of information about the scenes, resulting in a variety of effects ranging from reverberation to additional ambient sounds. In this paper, we manipulate input speech to sound as though it was recorded within a different scene, given an audio-visual conditional example recorded from that scene. Our model learns through self-supervision, taking advantage of the fact that natural video contains recurring sound events and textures. We extract an audio clip from a video and apply speech enhancement. We then train a latent diffusion model to recover the original speech, using another audio-visual clip taken from elsewhere in the video as a conditional hint. Through this process, the model learns to transfer the conditional example's sound properties to the input speech. We show that our model can be successfully trained using unlabeled, in-the-wild videos, and that an additional visual signal can improve its sound prediction abilities. Please see our project webpage for video results: https: tinglok.netlify.app files avsoundscape
【27】 AMT-APC: Automatic Piano Cover by Fine-Tuning an Automatic Music Transcription Model
标题: AMT-IPC:通过微调自动音乐转录模型实现自动钢琴翻唱
作者:Kazuma Komiya,Yoshihisa Fukuhara
链接:点击下载PDF文件
摘要:已经有几项关于自动生成钢琴封面的研究,最近深度学习的进步使得能够创建更复杂的封面。然而,现有的自动钢琴盖模型在表现力和对原件的保真度方面仍有改进的空间。为了解决这些问题,我们提出了一种称为AMT-APC的学习算法,该算法利用了自动音乐转录模型的功能。通过利用完善的自动音乐转录模型的优势,我们的目标是提高钢琴覆盖生成的准确性。我们的实验表明,AMT-APC模型比任何现有模型更准确地再现原始轨迹。摘要:There have been several studies on automatically generating piano covers, and recent advancements in deep learning have enabled the creation of more sophisticated covers. However, existing automatic piano cover models still have room for improvement in terms of expressiveness and fidelity to the original. To address these issues, we propose a learning algorithm called AMT-APC, which leverages the capabilities of automatic music transcription models. By utilizing the strengths of well-established automatic music transcription models, we aim to improve the accuracy of piano cover generation. Our experiments demonstrate that the AMT-APC model reproduces original tracks more accurately than any existing models.
【28】 MultiMed: Multilingual Medical Speech Recognition via Attention Encoder Decoder
标题: MultiMed:通过注意力编码器解码器的多语言医学语音识别
作者:Khai Le-Duc,Phuc Phan,Tan-Hanh Pham,Bach Phan Tat,Minh-Huong Ngo,Truong-Son Hy
备注:Preprint
链接:点击下载PDF文件
摘要:医学领域的多语言自动语音识别(ASR)是语音翻译、口语理解和语音激活助理等各种下游应用的基础任务。这项技术通过跨越语言障碍实现有效沟通,缓解专业劳动力短缺,并促进改善诊断和治疗,特别是在大流行期间,从而增强了患者护理。在这项工作中,我们介绍了MultiMed,这是一个针对医疗领域的小型到大型端到端ASR模型集合,涵盖五种语言:越南语,英语,德语,法语和汉语普通话,以及相应的真实世界ASR数据集。据我们所知,MultiMed是最大和第一个多语言医学ASR数据集,包括总持续时间,说话者数量,疾病多样性,记录条件,说话者角色,独特的医学术语,口音和ICD-10代码。其次,我们建立了经验基线,提出了第一个可重复的医学ASR多语言研究,进行了端到端ASR培训的逐层消融研究,并为多语言医学ASR提供了第一个语言分析。所有代码、数据和模型均可在https: github.com leduckhai MultiMed tree master MultiMed上获得摘要:Multilingual automatic speech recognition (ASR) in the medical domain serves as a foundational task for various downstream applications such as speech translation, spoken language understanding, and voice-activated assistants. This technology enhances patient care by enabling efficient communication across language barriers, alleviating specialized workforce shortages, and facilitating improved diagnosis and treatment, particularly during pandemics. In this work, we introduce MultiMed, a collection of small-to-large end-to-end ASR models for the medical domain, spanning five languages: Vietnamese, English, German, French, and Mandarin Chinese, together with the corresponding real-world ASR dataset. To our best knowledge, MultiMed stands as the largest and the first multilingual medical ASR dataset, in terms of total duration, number of speakers, diversity of diseases, recording conditions, speaker roles, unique medical terms, accents, and ICD-10 codes. Secondly, we establish the empirical baselines, present the first reproducible study of multilinguality in medical ASR, conduct a layer-wise ablation study for end-to-end ASR training, and provide the first linguistic analysis for multilingual medical ASR. All code, data, and models are available online https: github.com leduckhai MultiMed tree master MultiMed
【29】 ECHO: Environmental Sound Classification with Hierarchical Ontology-guided Semi-Supervised Learning
标题: ECHO:采用分层实体指导的半监督学习的环境声音分类
作者:Pranav Gupta,Raunak Sharma,Rashmi Kumari,Sri Krishna Aditya,Shwetank Choudhary,Sumit Kumar,Kanchana M,Thilagavathy R
备注:IEEE CONECCT 2024, Signal Processing and Pattern Recognition, Environmental Sound Classification, ESC
链接:点击下载PDF文件
摘要:环境声音分类一直是信号处理领域一个备受关注的研究问题,迄今为止,人们更多地关注全监督方法。在过去的几年里,焦点已经转向半监督方法,专注于利用未标记的数据,和自我监督的方法,通过借口任务或对比学习学习的中间表示。然而,这两种方法都需要大量未标记的数据来提高性能。在这项工作中,我们提出了一个新的框架,称为环境声音分类与层次本体指导的半监督学习(ECHO),利用标签本体为基础的层次结构,通过定义一个新的借口任务来学习语义表示。在prefect任务中,该模型尝试基于地面真值标签本体来预测由大型语言模型(LLM)定义的粗略标签。经过训练的模型以监督的方式进一步微调,以预测实际任务。我们提出的新的半监督框架在三个数据集(即UrbanSound 8 K,ESC-10和ESC-50)上实现了比基线系统1%至8%的精度提高。摘要:Environment Sound Classification has been a well-studied research problem in the field of signal processing and up till now more focus has been laid on fully supervised approaches. Over the last few years, focus has moved towards semi-supervised methods which concentrate on the utilization of unlabeled data, and self-supervised methods which learn the intermediate representation through pretext task or contrastive learning. However, both approaches require a vast amount of unlabelled data to improve performance. In this work, we propose a novel framework called Environmental Sound Classification with Hierarchical Ontology-guided semi-supervised Learning (ECHO) that utilizes label ontology-based hierarchy to learn semantic representation by defining a novel pretext task. In the pretext task, the model tries to predict coarse labels defined by the Large Language Model (LLM) based on ground truth label ontology. The trained model is further fine-tuned in a supervised way to predict the actual task. Our proposed novel semi-supervised framework achieves an accuracy improvement in the range of 1 % to 8 % over baseline systems across three datasets namely UrbanSound8K, ESC-10, and ESC-50.
【30】 Training Large ASR Encoders with Differential Privacy
标题: 训练具有差异隐私的大型ASB编码器
作者:Geeticka Chauhan,Steve Chien,Om Thakkar,Abhradeep Thakurta,Arun Narayanan
备注:In proceedings of the IEEE Spoken Language Technologies Workshop, 2024
链接:点击下载PDF文件
摘要:大型语音模型的自监督学习(SSL)方法已被证明在ASR中非常有效。随着对公共部署大型预训练模型的兴趣,人们越来越担心训练数据中的敏感数据点的意外记忆和泄漏。在本文中,我们将差分私有(DP)预训练应用于基于SOTA Conformer的编码器,并研究其在下游ASR任务上的性能,假设微调数据是公开的。本文首次将DP应用于SSL以实现ASR,研究了BEST-RQ预训练方法的DP噪声容忍度。值得注意的是,我们引入了一种新的模型修剪变体,称为基于梯度的层冻结,它在隐私-效用-计算权衡方面提供了很大的改进。我们的方法得出的LibriSpeech测试-干净 其他WER(%)为3.78 8.41,外推到低数据集尺度时为($10$,1 e ^-9)-DP,外推到高尺度时为2.81 5.89,外推到(10,7.9e^-11)-DP。摘要:Self-supervised learning (SSL) methods for large speech models have proven to be highly effective at ASR. With the interest in public deployment of large pre-trained models, there is a rising concern for unintended memorization and leakage of sensitive data points from the training data. In this paper, we apply differentially private (DP) pre-training to a SOTA Conformer-based encoder, and study its performance on a downstream ASR task assuming the fine-tuning data is public. This paper is the first to apply DP to SSL for ASR, investigating the DP noise tolerance of the BEST-RQ pre-training method. Notably, we introduce a novel variant of model pruning called gradient-based layer freezing that provides strong improvements in privacy-utility-compute trade-offs. Our approach yields a LibriSpeech test-clean other WER (%) of 3.78 8.41 with ($10$, 1e^-9)-DP for extrapolation towards low dataset scales, and 2.81 5.89 with (10, 7.9e^-11)-DP for extrapolation towards high scales.
【31】 Target word activity detector: An approach to obtain ASR word boundaries without lexicon
标题: 目标词活动检测器:一种无需词典即可获得ASB词边界的方法
作者:Sunit Sivasankaran,Eric Sun,Jinyu Li,Yan Huang,Jing Pan
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:由于在训练期间缺乏明确的时间对齐,从端到端(E2E)ASR模型中获取单词时间戳信息仍然具有挑战性。这个问题在多语言模型中更加复杂。现有的方法要么依赖于词典,要么引入额外的令牌,导致可扩展性问题和增加的计算成本。在这项工作中,我们提出了一种新的方法来估计词的边界,而不依赖于词典。我们的方法利用了来自子单词标记单元的单词嵌入和预训练的ASR模型,在训练过程中只需要单词对齐信息。我们提出的方法可以扩展到任何数量的语言,而不会产生任何额外的成本。我们使用经过五种语言训练的多语言ASR模型来验证我们的方法,并根据强基线证明其有效性。摘要:Obtaining word timestamp information from end-to-end (E2E) ASR models remains challenging due to the lack of explicit time alignment during training. This issue is further complicated in multilingual models. Existing methods, either rely on lexicons or introduce additional tokens, leading to scalability issues and increased computational costs. In this work, we propose a new approach to estimate word boundaries without relying on lexicons. Our method leverages word embeddings from sub-word token units and a pretrained ASR model, requiring only word alignment information during training. Our proposed method can scale-up to any number of languages without incurring any additional cost. We validate our approach using a multilingual ASR model trained on five languages and demonstrate its effectiveness against a strong baseline.
【32】 PTQ4ADM: Post-Training Quantization for Efficient Text Conditional Audio Diffusion Models
标题: PTQ4ADM:高效文本条件音频扩散模型的训练后量化
作者:Jayneel Vora,Aditya Krishnan,Nader Bouacida,Prabhu RV Shankar,Prasant Mohapatra
链接:点击下载PDF文件
摘要:去噪扩散模型已成为图像、音频和视频领域生成任务的最新技术,可生成高质量、多样化和上下文相关的数据。然而,它们的广泛采用受到高计算成本和大内存占用的限制。后训练量化(PTQ)通过低带宽参数降低模型复杂度,提供了一种有前途的方法来缓解这些挑战。然而,直接将PTQ应用于扩散模型可能会降低合成质量,这是由于多个去噪步骤中累积的量化噪声,特别是在文本到音频合成等条件任务中。本文介绍了一种新的音频扩散模型量化框架PTQ 4ADM。我们的主要贡献包括(1)覆盖率驱动的提示增强方法和(2)激活感知的文本条件ADM的校准集生成算法。这些技术确保全面覆盖音频方面和模态,同时保持合成保真度。我们验证了我们的方法TANGO,Make-An-Audio和AudioLDM模型的文本条件音频生成。大量的实验证明PTQ 4ADM的能力,以减少模型的大小高达70%,同时实现合成质量指标与全精度模型(FD分数增加$<5%)。我们表明,骨干网络中的特定层可以量化为4位权重和8位激活,而不会有显着的质量损失。这项工作为在资源受限的环境中更有效地部署ADM铺平了道路。摘要:Denoising diffusion models have emerged as state-of-the-art in generative tasks across image, audio, and video domains, producing high-quality, diverse, and contextually relevant data. However, their broader adoption is limited by high computational costs and large memory footprints. Post-training quantization (PTQ) offers a promising approach to mitigate these challenges by reducing model complexity through low-bandwidth parameters. Yet, direct application of PTQ to diffusion models can degrade synthesis quality due to accumulated quantization noise across multiple denoising steps, particularly in conditional tasks like text-to-audio synthesis. This work introduces PTQ4ADM, a novel framework for quantizing audio diffusion models(ADMs). Our key contributions include (1) a coverage-driven prompt augmentation method and (2) an activation-aware calibration set generation algorithm for text-conditional ADMs. These techniques ensure comprehensive coverage of audio aspects and modalities while preserving synthesis fidelity. We validate our approach on TANGO, Make-An-Audio, and AudioLDM models for text-conditional audio generation. Extensive experiments demonstrate PTQ4ADM's capability to reduce the model size by up to 70 % while achieving synthesis quality metrics comparable to full-precision models($<$5 % increase in FD scores). We show that specific layers in the backbone network can be quantized to 4-bit weights and 8-bit activations without significant quality loss. This work paves the way for more efficient deployment of ADMs in resource-constrained environments.
【33】 Investigation of Time-Frequency Feature Combinations with Histogram Layer Time Delay Neural Networks
标题: 基于柱状图层延时神经网络的时频特征组合研究
作者:Amirmohammad Mohammadi,Iren'e Masabarakiza,Ethan Barnes,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 14 figures. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:虽然深度学习减少了手动特征提取的流行,但通过特征工程进行数据转换对于提高模型性能仍然至关重要,特别是对于水声信号。将音频信号转换为时频表示的方法以及随后对这些频谱图的处理可以显著影响性能。这项工作演示了在直方图层时间延迟神经网络中使用不同的时频特征组合对性能的影响。一组最佳的功能识别的结果表明,特定的功能组合优于单一的数据功能。摘要:While deep learning has reduced the prevalence of manual feature extraction, transformation of data via feature engineering remains essential for improving model performance, particularly for underwater acoustic signals. The methods by which audio signals are converted into time-frequency representations and the subsequent handling of these spectrograms can significantly impact performance. This work demonstrates the performance impact of using different combinations of time-frequency features in a histogram layer time delay neural network. An optimal set of features is identified with results indicating that specific feature combinations outperform single data features.
【34】 Transfer Learning for Passive Sonar Classification using Pre-trained Audio and ImageNet Models
标题: 使用预训练的音频和ImageNet模型进行被动声纳分类的迁移学习
作者:Amirmohammad Mohammadi,Tejashri Kelhe,Davelle Carreiro,Alexandra Van Dine,Joshua Peeples
备注:5 pages, 6 figures, This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:迁移学习通常用于利用大型的预训练模型,并对下游任务进行微调。最流行的预训练模型最初使用ImageNet进行训练。然而,它们的泛化能力在不同的数据模式中可能会有所不同。本研究在水下声学目标识别(UATR)的背景下比较了预训练的音频神经网络(PANN)和ImageNet预训练模型。据观察,ImageNet预训练模型在被动声纳分类中的表现略优于预训练音频模型。我们还分析了音频采样率对模型预训练和微调的影响。这项研究有助于UATR的迁移学习应用,说明了预训练模型在解决UATR领域稀缺的标记数据所造成的局限性方面的潜力。摘要:Transfer learning is commonly employed to leverage large, pre-trained models and perform fine-tuning for downstream tasks. The most prevalent pre-trained models are initially trained using ImageNet. However, their ability to generalize can vary across different data modalities. This study compares pre-trained Audio Neural Networks (PANNs) and ImageNet pre-trained models within the context of underwater acoustic target recognition (UATR). It was observed that the ImageNet pre-trained models slightly out-perform pre-trained audio models in passive sonar classification. We also analyzed the impact of audio sampling rates for model pre-training and fine-tuning. This study contributes to transfer learning applications of UATR, illustrating the potential of pre-trained models to address limitations caused by scarce, labeled data in the UATR domain.
【35】 On the Feasibility of Fully AI-automated Vishing Attacks
标题: 全人工智能自动Vising攻击的可行性
作者:João Figueiredo,Afonso Carvalho,Daniel Castro,Daniel Gonçalves,Nuno Santos
链接:点击下载PDF文件
摘要:钓鱼攻击是一种社会工程,攻击者使用电话欺骗个人泄露敏感信息,如个人数据、财务信息或安全凭证。攻击者利用语音通信的紧迫性和真实性来操纵受害者,通常冒充银行或技术支持等合法实体。网络钓鱼是一种特别严重的威胁,因为它绕过了旨在保护信息的安全控制。在这项工作中,我们研究了随着人工智能的出现,网络钓鱼攻击升级的可能性。从理论上讲,人工智能驱动的软件机器人可能有能力通过电话与潜在受害者进行对话并欺骗他们披露敏感信息来自动化这些攻击。为了验证这一理论,我们介绍了Viking,这是一个使用公开可用的AI技术开发的AI驱动的vishing系统。它依赖于大型语言模型(LLM)作为其核心认知处理器来引导与受害者的对话,并辅之以语音到文本和文本到语音模块的管道,以促进电话中的音频-文本转换。通过一项涉及240名参与者的受控社会实验,我们发现Viking成功地说服了许多参与者透露敏感信息,即使是那些被明确警告过钓鱼活动风险的人。与维京机器人的互动通常被认为是现实的。从这些发现中,我们得出结论,像Viking这样的工具可能已经可以被潜在的恶意行为者使用,同时也是网络意识计划的宝贵资源。摘要:A vishing attack is a form of social engineering where attackers use phone calls to deceive individuals into disclosing sensitive information, such as personal data, financial information, or security credentials. Attackers exploit the perceived urgency and authenticity of voice communication to manipulate victims, often posing as legitimate entities like banks or tech support. Vishing is a particularly serious threat as it bypasses security controls designed to protect information. In this work, we study the potential for vishing attacks to escalate with the advent of AI. In theory, AI-powered software bots may have the ability to automate these attacks by initiating conversations with potential victims via phone calls and deceiving them into disclosing sensitive information. To validate this thesis, we introduce ViKing, an AI-powered vishing system developed using publicly available AI technology. It relies on a Large Language Model (LLM) as its core cognitive processor to steer conversations with victims, complemented by a pipeline of speech-to-text and text-to-speech modules that facilitate audio-text conversion in phone calls. Through a controlled social experiment involving 240 participants, we discovered that ViKing has successfully persuaded many participants to reveal sensitive information, even those who had been explicitly warned about the risk of vishing campaigns. Interactions with ViKing's bots were generally considered realistic. From these findings, we conclude that tools like ViKing may already be accessible to potential malicious actors, while also serving as an invaluable resource for cyber awareness programs.
【36】 A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
标题: 随机信封波动对噪音音素感知影响的微观研究
作者:Alejandro Osses,Léo Varnet
Journal-ref:Journal of the Acoustical Society of America, 2024, 155 (2), pp.1469-1485
链接:点击下载PDF文件
摘要:在这项研究中,我们研究了特定的噪音实现对两个辅音, b 和 d 的歧视的影响。为此,我们收集了12名参与者的数据,他们听了嵌入在三种背景噪音中的 aba 或 ada 。所有的噪声都有相同的长期频谱,但随机包络波动的数量不同。使用反向相关法在逐个试验的基础上对数据进行分析。结果表明,它是有可能预测的分类响应具有比机会更好的准确性纯粹基于相应的噪声的随机包络波动的频谱-时间分布,而不考虑实际的目标或在试验中使用的信噪比。噪声波动的影响平均解释了参与者在白噪声中8.1%的反应,对于波动量较大的噪声,这一比例增加到13.3%。估计的时频权重显示,测量的效果源自噪声波动和目标词的相关声学线索之间的混淆。从使用人工...的模拟中获得了基本相似的结论。我们认为,这种标记特定的噪声效应是一种形式的信息掩蔽。摘要:In this study, we investigated the effect of specific noise realizations on the discrimination of two consonants, b and d . For this purpose, we collected data from twelve participants, who listened to the words aba or ada embedded in one of three background noises. All noises had the same long-term spectrum but differed in the amount of random envelope fluctuations. The data were analyzed on a trial-by-trial basis using the reverse-correlation method. The results revealed that it is possible to predict the categorical responses with better-than-chance accuracy purely based on the spectro-temporal distribution of the random envelope fluctuations of the corresponding noises, without taking into account the actual targets or the signal-to-noise ratios used in the trials. The effect of the noise fluctuations explained on average 8.1% of the participants' responses in white noise, a proportion that increased up to 13.3% for noises with a larger amount of fluctuations. The estimated time-frequency weights revealed that the measured effect originated from confusions between noise fluctuations and relevant acoustic cues from the target words. Substantially similar conclusions were obtained from simulations using an artificial listener. We argue that this token-specific effect of noise is a form of informational masking.
【37】 Optimizing the Songwriting Process: Genre-Based Lyric Generation Using Deep Learning Models
标题: 优化歌曲创作过程:使用深度学习模型基于流派的歌词生成
作者:Tracy Cai,Wilson Liang,Donte Townes
链接:点击下载PDF文件
摘要:传统的歌曲创作过程是相当复杂的,这是显而易见的时间,它需要产生的歌词,适合体裁和形式全面的诗句。我们的项目旨在通过深度学习技术简化这一过程,从而优化歌曲创作过程,并使艺术家能够通过保持流派来达到目标受众。使用Spotify上的18,000首歌曲的数据集,我们开发了一种独特的预处理格式,使用令牌将歌词解析为单独的诗句。这些结果被用来训练基线预训练seq2seq模型,和基于LSTM的神经网络模型根据歌曲流派。我们发现,在基线模型中,生成产生了更高的召回率(ROUGE),但两个模型的精度(BLEU)相似。从质量上看,我们发现原始模型生成的许多抒情短语仍然可以理解和辨别,尽管它们不一定与真正的歌词完全相同。总的来说,我们的研究结果表明,歌词生成可以合理地加快产生基于体裁的歌词,并有助于加快歌曲创作过程。摘要:The traditional songwriting process is rather complex and this is evident in the time it takes to produce lyrics that fit the genre and form comprehensive verses. Our project aims to simplify this process with deep learning techniques, thus optimizing the songwriting process and enabling an artist to hit their target audience by staying in genre. Using a dataset of 18,000 songs off Spotify, we developed a unique preprocessing format using tokens to parse lyrics into individual verses. These results were used to train a baseline pretrained seq2seq model, and a LSTM-based neural network models according to song genres. We found that generation yielded higher recall (ROUGE) in the baseline model, but similar precision (BLEU) for both models. Qualitatively, we found that many of the lyrical phrases generated by the original model were still comprehensible and discernible between which genres they fit into, despite not necessarily being the exact the same as the true lyrics. Overall, our results yielded that lyric generation can reasonably be sped up to produce genre-based lyrics and aid in hastening the songwriting process.
【38】 Enhancing Kurdish Text-to-Speech with Native Corpus Training: A High-Quality WaveGlow Vocoder Approach
标题: 通过原生数据库训练增强库尔德语文本到语音:高质量WaveGlow声码器方法
作者:Abdulhady Abas Abdullah,Sabat Salih Muhamad,Hadi Veisi
链接:点击下载PDF文件
摘要:随着文本转语音技术的进步,从文本合成口语的能力极大地便利了数字内容的获取。然而,有效的TTS开发低资源的语言,如中央库尔德语(CKB),仍然面临着许多挑战,主要是由于缺乏语言信息和专用资源。在本文中,我们改进了基于Tacotron的库尔德文语转换系统,通过在21小时的中央库尔德语语音语料库上训练库尔德语WaveGlow声码器,而不是使用预先训练的英语声码器WaveGlow。为了准确、流畅地适应库尔德语语音和韵律的变化,需要在目标语语料库上进行声码器训练。这些增强的有效性在于,我们的模型明显优于使用英语预训练模型的基线系统。特别是,我们的自适应WaveGlow模型达到了令人印象深刻的4.91 MOS,为库尔德语语音合成设定了新的基准。一方面,这项研究赋予中央库尔德语TTS系统的先进功能,另一方面,它为库尔德语和其他相关语言的其他方言的进一步发展打开了大门。摘要:The ability to synthesize spoken language from text has greatly facilitated access to digital content with the advances in text-to-speech technology. However, effective TTS development for low-resource languages, such as Central Kurdish (CKB), still faces many challenges due mainly to the lack of linguistic information and dedicated resources. In this paper, we improve the Kurdish TTS system based on Tacotron by training the Kurdish WaveGlow vocoder on a 21-hour central Kurdish speech corpus instead of using a pre-trained English vocoder WaveGlow. Vocoder training on the target language corpus is required to accurately and fluently adapt phonetic and prosodic changes in Kurdish language. The effectiveness of these enhancements is that our model is significantly better than the baseline system with English pretrained models. In particular, our adaptive WaveGlow model achieves an impressive MOS of 4.91, which sets a new benchmark for Kurdish speech synthesis. On one hand, this study empowers the advanced features of the TTS system for Central Kurdish, and on the other hand, it opens the doors for other dialects in Kurdish and other related languages to further develop.
【39】 Lightweight Transducer Based on Frame-Level Criterion
标题: 基于框架级标准的轻型传感器
作者:Genshun Wan,Mengzhi Wang,Tingzhi Mao,Hang Chen,Zhongfu Ye
备注:Accepted by Interspeech 2024, code repository: this https URL
链接:点击下载PDF文件
摘要:基于序列级准则训练的换能器模型由于产生大概率矩阵而需要大量的内存。提出了一种基于帧级准则的轻量级传感器模型,该模型利用CTC强制对齐算法的结果来确定每帧的标签。然后,编码器输出可以在相应的时间与解码器输出组合,而不是像在换能器中那样将编码器输出的每个元素添加到解码器输出的每个元素。这大大降低了内存和计算需求。为了解决标签中过多空白导致的分类不平衡问题,我们将空白和非空白概率解耦,并将空白分类器的梯度截断到主网络。这使得轻质换能器能够实现与换能器类似的结果。此外,我们使用更丰富的信息来预测空白的概率,取得了优于传感器的结果。摘要:The transducer model trained based on sequence-level criterion requires a lot of memory due to the generation of the large probability matrix. We proposed a lightweight transducer model based on frame-level criterion, which uses the results of the CTC forced alignment algorithm to determine the label for each frame. Then the encoder output can be combined with the decoder output at the corresponding time, rather than adding each element output by the encoder to each element output by the decoder as in the transducer. This significantly reduces memory and computation requirements. To address the problem of imbalanced classification caused by excessive blanks in the label, we decouple the blank and non-blank probabilities and truncate the gradient of the blank classifier to the main network. This enables the lightweight transducer achieving similar results to transducer. Additionally, we use richer information to predict the probability of blank, achieving superior results to transducer.
机器翻译,仅供参考
