本文经arXiv每日学术速递授权转载
【1】 MusicLIME: Explainable Multimodal Music Understanding
标题: MusicLIME:可解释的多模式音乐理解
作者:Theodoros Sotirou,Vassilis Lyberatos,Orfeas Menis Mastromichalakis,Giorgos Stamou
备注:GitHub repository: this https URL
链接:点击下载PDF文件
【2】 2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
标题: 2D或不是2D:手势表示的抽象性如何影响3D同声手势生成?
作者:Téo Guichoux,Laure Soulier,Nicolas Obin,Catherine Pelachaud
链接:点击下载PDF文件
【3】 DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
标题: DreamHead:通过分层扩散学习时空对应性,以实现音频驱动的会说话的头部合成
作者:Fa-Ting Hong,Yunfei Liu,Yu Li,Changyin Zhou,Fei Yu,Dan Xu
链接:点击下载PDF文件
【4】 Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
标题: 基于扬声器分离HuBERT的自监督音节发现
作者:Ryota Komatsu,Takahiro Shinozaki
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【5】 Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
标题: 优化构音障碍唤醒词定位:SEARCH 2024 LRDDWWS挑战赛的端到端方法
作者:Shuiyun Liu,Yuxiang Kong,Pengcheng Guo,Weiji Zhuang,Peng Gao,Yujun Wang,Lei Xie
备注:8 pages, Accepted to SLT 2024
链接:点击下载PDF文件
【6】 Speaker Contrastive Learning for Source Speaker Tracing
标题: 用于源说话人追踪的说话人对比学习
作者:Qing Wang,Hongmei Guo,Jian Kang,Mengjie Du,Jie Li,Xiao-Lei Zhang,Lei Xie
备注:7 pages, 2 figures, accepted by SLT
链接:点击下载PDF文件
【7】 Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
标题: 自然环境中用于头部定向的音频驱动强化学习
作者:Wessel Ledder,Yuzhen Qin,Kiki van der Heijden
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
标题: DiffTAR:音频文本检索的基于扩散的生成建模
作者:Yifei Xin,Xuxin Cheng,Zhihong Zhu,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
【9】 Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
标题: 通过多任务学习从转录的语音音频中获取发音知识
作者:Siqi Sun,Korin Richmond
备注:5 pages
链接:点击下载PDF文件
【10】 Constructing a Singing Style Caption Dataset
标题: 构建歌唱风格字幕数据集
作者:Hyunjong Ok,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
【11】 Efficient Video to Audio Mapper with Visual Scene Detection
标题: 具有视觉场景检测的高效视频到音频映射器
作者:Mingjing Yi,Ming Li
链接:点击下载PDF文件
【12】 Large Language Model Based Generative Error Correction: A Challenge and Baselines forSpeech Recognition, Speaker Tagging, and Emotion Recognition
标题: 基于大语言模型的生成式错误纠正:语音识别、说话人标记和情感识别的挑战和基线
作者:Chao-Han Huck Yang,Taejin Park,Yuan Gong,Yuanchao Li,Zhehuai Chen,Yen-Ting Lin,Chen Chen,Yuchen Hu,Kunal Dhawan,Piotr Żelasko,Chao Zhang,Yun-Nung Chen,Yu Tsao,Jagadeesh Balam,Boris Ginsburg,Sabato Marco Siniscalchi,Eng Siong Chng,Peter Bell,Catherine Lai,Shinji Watanabe,Andreas Stolcke
备注:IEEE SLT 2024. The initial draft version has been done in December 2023. Post-ASR Text Processing and Understanding Community: this https URL
链接:点击下载PDF文件
【13】 Self-supervised Learning for Acoustic Few-Shot Classification
标题: 声学Few-Shot分类的自我监督学习
作者:Jingyong Liang,Bernd Meyer,Issac Ning Lee,Thanh-Toan Do
链接:点击下载PDF文件
【14】 Compositional Audio Representation Learning
标题: 合成音频表示学习
作者:Sripathi Sridhar,Mark Cartwright
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【15】 Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
标题: 整合音频叙述加强多模式第一人称动作识别中的领域概括
作者:Cagri Gungor,Adriana Kovashka
链接:点击下载PDF文件
【16】 A Survey of Foundation Models for Music Understanding
标题: 音乐理解的基础模型综述
作者:Wenjun Li,Ying Cai,Ziyang Wu,Wenyi Zhang,Yifan Chen,Rundong Qi,Mengqi Dong,Peigen Chen,Xiao Dong,Fenghao Shi,Lei Guo,Junwei Han,Bao Ge,Tianming Liu,Lin Gan,Tuo Zhang
备注:20 pages, 2 figures
链接:点击下载PDF文件
【17】 On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
标题: 关于目标说话人提取的注册语音增强的有效性
作者:Junjie Li,Ke Zhang,Shuai Wang,Haizhou Li,Man-Wai Mak,Kong Aik Lee
备注:Accepted by SLT2024
链接:点击下载PDF文件
【18】 ASR Error Correction using Large Language Models
标题: 使用大型语言模型的ASB错误纠正
作者:Rao Ma,Mengjie Qian,Mark Gales,Kate Knill
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
【19】 Multi-Microphone and Multi-Modal Emotion Recognition in Reverbrant Enviroment
标题: 可逆环境中的多麦克风和多模式情感识别
作者:Ohad Cohen,Gershon Hazan,Sharon Gannot
链接:点击下载PDF文件
【20】 Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
标题: 通过预测可解释的声学特征来解释语音情感识别的深度学习嵌入
作者:Satvik Dixit,Daniel M. Low,Gasser Elbanna,Fabio Catania,Satrajit S. Ghosh
链接:点击下载PDF文件
【21】 ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
标题: ESPnet-ZZ:仅使用Python的ESPnet,易于微调和集成
作者:Masao Someki,Kwanghee Choi,Siddhant Arora,William Chen,Samuele Cornell,Jionghao Han,Yifan Peng,Jiatong Shi,Vaibhav Srivastav,Shinji Watanabe
备注:Accepted to SLT 2024
链接:点击下载PDF文件
【22】 Prevailing Research Areas for Music AI in the Era of Foundation Models
标题: 基础模型时代音乐人工智能的主流研究领域
作者:Megan Wei,Mateusz Modrzejewski,Aswin Sivaraman,Dorien Herremans
链接:点击下载PDF文件
【23】 Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
标题: 联合语义知识提取和掩蔽声学建模用于提高可理解度的全频段语音恢复
作者:Xiaoyu Liu,Xu Li,Joan Serrà,Santiago Pascual
备注:Demo link this https URL
链接:点击下载PDF文件
【24】 MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
标题: MacST:通过文本音译进行多口音语音合成以实现口音转换
作者:Sho Inoue,Shuai Wang,Wanxing Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
备注:Project page with Speech Demo: this https URL
链接:点击下载PDF文件
【25】 Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
标题: 儿童与成人二元互动中的自我中心说话者分类:从感知到计算建模
作者:Tiantian Feng,Anfeng Xu,Xuan Shi,Somer Bishop,Shrikanth Narayanan
备注:pre-print under review
链接:点击下载PDF文件
【26】 The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
标题: 2024年VoiceMOS挑战赛的T05系统:从深度图像分类器转移学习到高质量合成语音的Naturalness MOS预测
作者:Kaito Baba,Wataru Nakata,Yuki Saito,Hiroshi Saruwatari
备注:Accepted by IEEE SLT 2024. Our MOS prediction system (UTMOSv2) is available in this https URL
链接:点击下载PDF文件
【27】 Subband Splitting: Simple, Efficient and Effective Technique for Solving Block Permutation Problem in Determined Blind Source Separation
标题: 子带分裂:解决确定盲源分离中块排列问题的简单、高效且有效的技术
作者:Kazuki Matsumoto,Kohei Yatabe
链接:点击下载PDF文件
【28】 DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
标题: DSCSYS:特定领域对比音频预训练
作者:Shengqiang Liu,Da Liu,Anna Wang,Zhiyu Zhang,Jie Gao,Yali Li
链接:点击下载PDF文件
【29】 M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
标题: M$^{3}$V:用于设备引导语音检测的多模式多视图方法
作者:Anna Wang,Da Liu,Zhiyu Zhang,Shengqiang Liu,Jie Gao,Yali Li
链接:点击下载PDF文件
【30】 SafeEar: Content Privacy-Preserving Audio Deepfake Detection
标题: SafeEar:内容隐私保护音频Deepfake检测
作者:Xinfeng Li,Kai Li,Yifan Zheng,Chen Yan,Xiaoyu Ji,Wenyuan Xu
备注:Accepted by ACM CCS 2024. Please cite this paper as "Xinfeng Li, Kai Li, Yifan Zheng, Chen Yan, Xiaoyu Ji, Wenyuan Xu. SafeEar: Content Privacy-Preserving Audio Deepfake Detection. In Proceedings of ACM Conference on Computer and Communications Security (CCS), 2024."
链接:点击下载PDF文件
【31】 Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
标题: 基于转换器的分层对齐和解纠缠跨模式表示的音频文本检索
作者:Yifei Xin,Zhihong Zhu,Xuxin Cheng,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
【32】 Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
标题: 多模式语音Transformer解码器:多模式何时可以提高准确性?
作者:Yiwen Guan,Viet Anh Trinh,Vivek Voleti,Jacob Whitehill
链接:点击下载PDF文件
【33】 Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
标题: Seed-Music:高质量和受控音乐生成的统一框架
作者:Ye Bai,Haonan Chen,Jitong Chen,Zhuo Chen,Yi Deng,Xiaohong Dong,Lamtharn Hantrakul,Weituo Hao,Qingqing Huang,Zhongyi Huang,Dongya Jia,Feihu La,Duc Le,Bochen Li,Chumin Li,Hui Li,Xingxing Li,Shouda Liu,Wei-Tsung Lu,Yiqing Lu,Andrew Shaw,Janne Spijkervet,Yakun Sun,Bo Wang,Ju-Chiang Wang,Yuping Wang,Yuxuan Wang,Ling Xu,Yifeng Yang,Chao Yao,Shuo Zhang,Yang Zhang,Yilin Zhang,Hang Zhao,Ziyi Zhao,Dejian Zhong,Shicen Zhou,Pei Zou
备注:Seed-Music technical report, 20 pages, 5 figures
链接:点击下载PDF文件
【34】 AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
标题: AccentBox:迈向高保真Zero-Shot口音一代
作者:Jinzuomu Zhong,Korin Richmond,Zhiba Su,Siqi Sun
链接:点击下载PDF文件
【35】 An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems
标题: 交互式口语对话系统的高效自学习框架
作者:Hitesh Tulsiani,David M. Chan,Shalini Ghosh,Garima Lalwani,Prabhat Pandey,Ankish Bansal,Sri Garimella,Ariya Rastrow,Björn Hoffmeister
备注:Presented at ICML 2024
链接:点击下载PDF文件
【36】 Meta-Whisper: Speech-Based Meta-ICL for ASR on Low-Resource Languages
标题: Meta-Whisper:基于语音的Meta-ICL,用于低资源语言上的ASB
作者:Ming-Hao Hsu,Kuan Po Huang,Hung-yi Lee
链接:点击下载PDF文件
【37】 Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
标题: 利用MAMBA联合频谱和空间学习进行多通道语音增强
作者:Wenze Ren,Haibin Wu,Yi-Cheng Lin,Xuanjun Chen,Rong Chao,Kuo-Hsuan Hung,You-Jin Li,Wen-Yuan Ting,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
【38】 Ultra-Low Latency Speech Enhancement - A Comprehensive Study
标题: 超低延迟语音增强-综合研究
作者:Haibin Wu,Sebastian Braun
链接:点击下载PDF文件
【39】 oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
标题: oboVox远场说话人识别:一种采用预训练模型的新型数据增强方法
作者:Muhammad Sudipto Siam Dip,Md Anik Hasan,Sapnil Sarker Bipro,Md Abdur Raiyan,Mohammod Abdul Motin
备注:5 pages, 2 figures
链接:点击下载PDF文件
【40】 Speech as a Biomarker for Disease Detection
标题: 言语作为疾病检测的生物标志物
作者:Catarina Botelho,Alberto Abad,Tanja Schultz,Isabel Trancoso
链接:点击下载PDF文件
【41】 RF-GML: Reference-Free Generative Machine Listener
标题: RF-GML:无参考生成机器收件箱
作者:Arijit Biswas,Guanxin Jiang
备注:Pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
【42】 Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
标题: CLAR-DPO:通过直接偏好优化的可控情感语音合成
作者:Xiaoxue Gao,Chen Zhang,Yiming Chen,Huayun Zhang,Nancy F. Chen
备注:5 pages
链接:点击下载PDF文件
【43】 Room impulse response prototyping using receiver distance estimations for high quality room equalisation algorithms
标题: 使用接收器距离估计的房间脉冲响应原型用于高质量房间均衡算法
作者:James Brooks-Park,Martin Bo Møller,Jan Østergaard,Søren Bech,Steven van de Par
链接:点击下载PDF文件
【44】 StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
标题: StyleTTS-ZZ:具有蒸馏时变风格扩散的高效高质量Zero-Shot文本到语音合成
作者:Yinghao Aaron Li,Xilin Jiang,Cong Han,Nima Mesgarani
链接:点击下载PDF文件
【45】 TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
标题: TBDM-Net:用于语音情感识别的具有性别信息的双向密集网络
作者:Vlad Striletchi,Cosmin Striletchi,Adriana Stan
备注:In Proceedings of 2024 IEEE International Workshop on Machine Learning for Signal Processing, London, UK
链接:点击下载PDF文件
【46】 DNN-based ensemble singing voice synthesis with interactions between singers
标题: 基于DNN的合奏歌唱声音合成,具有歌手之间的互动
作者:Hiroaki Hyodo,Shinnosuke Takamichi,Tomohiko Nakamura,Junya Koguchi,Hiroshi Saruwatari
链接:点击下载PDF文件
【47】 A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
标题: 使用大型语言模型的Zero-Shot非侵入性语音评估研究
作者:Ryandhimas E. Zezario,Sabato M. Siniscalchi,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
【48】 Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
标题: 自我监督的多模式语音表示用于评估精神分裂症症状
作者:Gowtham Premananth,Carol Espy-Wilson
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【49】 Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
标题: 提取和扩散:用于改进基于扩散的语音和人声增强的潜在集成
作者:Yudong Yang,Zhan Liu,Wenyi Yu,Guangzhi Sun,Qiuqiang Kong,Chao Zhang
链接:点击下载PDF文件
【50】 Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
标题: 口吃解决器:端到端多语言流利检测
作者:Xuanru Zhou,Cheol Jun Cho,Ayati Sharma,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Boon Lead Tee,Maria Luisa Gorno Tempini,Jiachen Lian,Gopala Anumanchipalli
备注:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【51】 Effective Pre-Training of Audio Transformers for Sound Event Detection
标题: 音频Transformer的有效预训练以进行声音事件检测
作者:Florian Schmid,Tobias Morocutti,Francesco Foscarin,Jan Schlüter,Paul Primus,Gerhard Widmer
备注:Submitted to ICASSP'25. Source code available: this https URL
链接:点击下载PDF文件
【52】 Target Speaker ASR with Whisper
标题: 目标说话者ASB与Whisper
作者:Alexander Polok,Dominik Klement,Matthew Wiesner,Sanjeev Khudanpur,Jan Černocký,Lukáš Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【53】 Leveraging Self-Supervised Learning for Speaker Diarization
标题: 利用自我监督学习进行发言者日记化
作者:Jiangyu Han,Federico Landini,Johan Rohdin,Anna Silnova,Mireia Diez,Lukas Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【54】 Language-Queried Target Sound Extraction Without Parallel Training Data
标题: 无需并行训练数据的数据查询目标声音提取
作者:Hao Ma,Zhiyuan Peng,Xu Li,Yukai Li,Mingjie Shao,Qiuqiang Kong,Ju Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【55】 Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
标题: 使用带伪标签的最佳传输进行说话人验证的通道自适应
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Lei Li,Xugang Lu
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【56】 Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
标题: 用于增强说话人验证的集成多层知识提炼
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Xugang Lu,Lei Li
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【57】 Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
标题: 文本提示还不够:用于目标风格音频生成的声音事件增强提示适配器
作者:Chenxu Xiong,Ruibo Fu,Shuchen Shi,Zhengqi Wen,Jianhua Tao,Tao Wang,Chenxing Li,Chunyu Qiang,Yuankun Xie,Xin Qi,Guanjun Li,Zizheng Yang
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【58】 E1 TTS: Simple and Fast Non-Autoregressive TTS
标题: E1 TTC:简单快速的非自回归TTC
作者:Zhijun Liu,Shuai Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
链接:点击下载PDF文件
【59】 Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
标题: Wave-U-Mamba:一个端到端框架,实现高质量和高效的语音超分辨率
作者:Yongjoon Lee,Chanwoo Kim
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【60】 Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
标题: 未标记条件下异常声音检测的鉴别特征空间训练的改进
作者:Takuya Fujimura,Ibuki Kuroyanagi,Tomoki Toda
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
【61】 Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
标题: 通过稳定的Forces生成提高基于扩散的零激发语音合成的鲁棒性
作者:Changjin Han,Seokgi Lee,Gyuhyeon Nam,Gyeongsu Chae
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【62】 ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
标题: ReCLAP:通过描述声音改进Zero-Shot音频分类
作者:Sreyan Ghosh,Sonal Kumar,Chandra Kiran Reddy Evuru,Oriol Nieto,Ramani Duraiswami,Dinesh Manocha
备注:Code and Checkpoints: this https URL
链接:点击下载PDF文件
【63】 Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
标题: 从策划值得信赖、注释良好且有用的无序英语言语数据集中吸取的教训
作者:Pan-Pan Jiang,Jimmy Tobin,Katrin Tomanek,Robert L. MacDonald,Katie Seaver,Richard Cave,Marilyn Ladewig,Rus Heywood,Jordan R. Green
备注:Interspeech 2024
链接:点击下载PDF文件
【64】 MambaFoley: Foley Sound Generation using Selective State-Space Models
标题: MambaFoley:使用选择性状态空间模型生成Foley声音
作者:Marco Furio Colombo,Francesca Ronchini,Luca Comanducci,Fabio Antonacci
链接:点击下载PDF文件
【65】 SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
标题: SLiCK:利用子序列进行长度限制的关键字发现
作者:Kumari Nishu,Minsik Cho,Devang Naik
链接:点击下载PDF文件
标题: 交互式口语对话系统的高效自学习框架
作者:Hitesh Tulsiani,David M. Chan,Shalini Ghosh,Garima Lalwani,Prabhat Pandey,Ankish Bansal,Sri Garimella,Ariya Rastrow,Björn Hoffmeister
备注:Presented at ICML 2024
链接:点击下载PDF文件
【2】 Meta-Whisper: Speech-Based Meta-ICL for ASR on Low-Resource Languages
标题: Meta-Whisper:基于语音的Meta-ICL,用于低资源语言上的ASB
作者:Ming-Hao Hsu,Kuan Po Huang,Hung-yi Lee
链接:点击下载PDF文件
【3】 Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
标题: 利用MAMBA联合频谱和空间学习进行多通道语音增强
作者:Wenze Ren,Haibin Wu,Yi-Cheng Lin,Xuanjun Chen,Rong Chao,Kuo-Hsuan Hung,You-Jin Li,Wen-Yuan Ting,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
【4】 Ultra-Low Latency Speech Enhancement - A Comprehensive Study
标题: 超低延迟语音增强-综合研究
作者:Haibin Wu,Sebastian Braun
链接:点击下载PDF文件
【5】 oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
标题: oboVox远场说话人识别:一种采用预训练模型的新型数据增强方法
作者:Muhammad Sudipto Siam Dip,Md Anik Hasan,Sapnil Sarker Bipro,Md Abdur Raiyan,Mohammod Abdul Motin
备注:5 pages, 2 figures
链接:点击下载PDF文件
【6】 Speech as a Biomarker for Disease Detection
标题: 言语作为疾病检测的生物标志物
作者:Catarina Botelho,Alberto Abad,Tanja Schultz,Isabel Trancoso
链接:点击下载PDF文件
【7】 RF-GML: Reference-Free Generative Machine Listener
标题: RF-GML:无参考生成机器收件箱
作者:Arijit Biswas,Guanxin Jiang
备注:Pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
【8】 Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
标题: CLAR-DPO:通过直接偏好优化的可控情感语音合成
作者:Xiaoxue Gao,Chen Zhang,Yiming Chen,Huayun Zhang,Nancy F. Chen
备注:5 pages
链接:点击下载PDF文件
【9】 Room impulse response prototyping using receiver distance estimations for high quality room equalisation algorithms
标题: 使用接收器距离估计的房间脉冲响应原型用于高质量房间均衡算法
作者:James Brooks-Park,Martin Bo Møller,Jan Østergaard,Søren Bech,Steven van de Par
链接:点击下载PDF文件
【10】 StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
标题: StyleTTS-ZZ:具有蒸馏时变风格扩散的高效高质量Zero-Shot文本到语音合成
作者:Yinghao Aaron Li,Xilin Jiang,Cong Han,Nima Mesgarani
链接:点击下载PDF文件
【11】 TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
标题: TBDM-Net:用于语音情感识别的具有性别信息的双向密集网络
作者:Vlad Striletchi,Cosmin Striletchi,Adriana Stan
备注:In Proceedings of 2024 IEEE International Workshop on Machine Learning for Signal Processing, London, UK
链接:点击下载PDF文件
【12】 DNN-based ensemble singing voice synthesis with interactions between singers
标题: 基于DNN的合奏歌唱声音合成,具有歌手之间的互动
作者:Hiroaki Hyodo,Shinnosuke Takamichi,Tomohiko Nakamura,Junya Koguchi,Hiroshi Saruwatari
链接:点击下载PDF文件
【13】 A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
标题: 使用大型语言模型的Zero-Shot非侵入性语音评估研究
作者:Ryandhimas E. Zezario,Sabato M. Siniscalchi,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
【14】 Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
标题: 自我监督的多模式语音表示用于评估精神分裂症症状
作者:Gowtham Premananth,Carol Espy-Wilson
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【15】 Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
标题: 提取和扩散:用于改进基于扩散的语音和人声增强的潜在集成
作者:Yudong Yang,Zhan Liu,Wenyi Yu,Guangzhi Sun,Qiuqiang Kong,Chao Zhang
链接:点击下载PDF文件
【16】 Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
标题: 口吃解决器:端到端多语言流利检测
作者:Xuanru Zhou,Cheol Jun Cho,Ayati Sharma,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Boon Lead Tee,Maria Luisa Gorno Tempini,Jiachen Lian,Gopala Anumanchipalli
备注:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
【17】 Effective Pre-Training of Audio Transformers for Sound Event Detection
标题: 音频Transformer的有效预训练以进行声音事件检测
作者:Florian Schmid,Tobias Morocutti,Francesco Foscarin,Jan Schlüter,Paul Primus,Gerhard Widmer
备注:Submitted to ICASSP'25. Source code available: this https URL
链接:点击下载PDF文件
【18】 Target Speaker ASR with Whisper
标题: 目标说话者ASB与Whisper
作者:Alexander Polok,Dominik Klement,Matthew Wiesner,Sanjeev Khudanpur,Jan Černocký,Lukáš Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【19】 Leveraging Self-Supervised Learning for Speaker Diarization
标题: 利用自我监督学习进行发言者日记化
作者:Jiangyu Han,Federico Landini,Johan Rohdin,Anna Silnova,Mireia Diez,Lukas Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【20】 Language-Queried Target Sound Extraction Without Parallel Training Data
标题: 无需并行训练数据的数据查询目标声音提取
作者:Hao Ma,Zhiyuan Peng,Xu Li,Yukai Li,Mingjie Shao,Qiuqiang Kong,Ju Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【21】 Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
标题: 使用带伪标签的最佳传输进行说话人验证的通道自适应
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Lei Li,Xugang Lu
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【22】 Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
标题: 用于增强说话人验证的集成多层知识提炼
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Xugang Lu,Lei Li
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【23】 Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
标题: 文本提示还不够:用于目标风格音频生成的声音事件增强提示适配器
作者:Chenxu Xiong,Ruibo Fu,Shuchen Shi,Zhengqi Wen,Jianhua Tao,Tao Wang,Chenxing Li,Chunyu Qiang,Yuankun Xie,Xin Qi,Guanjun Li,Zizheng Yang
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
【24】 E1 TTS: Simple and Fast Non-Autoregressive TTS
标题: E1 TTC:简单快速的非自回归TTC
作者:Zhijun Liu,Shuai Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
链接:点击下载PDF文件
【25】 Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
标题: Wave-U-Mamba:一个端到端框架,实现高质量和高效的语音超分辨率
作者:Yongjoon Lee,Chanwoo Kim
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
【26】 Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
标题: 未标记条件下异常声音检测的鉴别特征空间训练的改进
作者:Takuya Fujimura,Ibuki Kuroyanagi,Tomoki Toda
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
【27】 Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
标题: 通过稳定的Forces生成提高基于扩散的零激发语音合成的鲁棒性
作者:Changjin Han,Seokgi Lee,Gyuhyeon Nam,Gyeongsu Chae
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【28】 ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
标题: ReCLAP:通过描述声音改进Zero-Shot音频分类
作者:Sreyan Ghosh,Sonal Kumar,Chandra Kiran Reddy Evuru,Oriol Nieto,Ramani Duraiswami,Dinesh Manocha
备注:Code and Checkpoints: this https URL
链接:点击下载PDF文件
【29】 Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
标题: 从策划值得信赖、注释良好且有用的无序英语言语数据集中吸取的教训
作者:Pan-Pan Jiang,Jimmy Tobin,Katrin Tomanek,Robert L. MacDonald,Katie Seaver,Richard Cave,Marilyn Ladewig,Rus Heywood,Jordan R. Green
备注:Interspeech 2024
链接:点击下载PDF文件
【30】 MambaFoley: Foley Sound Generation using Selective State-Space Models
标题: MambaFoley:使用选择性状态空间模型生成Foley声音
作者:Marco Furio Colombo,Francesca Ronchini,Luca Comanducci,Fabio Antonacci
链接:点击下载PDF文件
【31】 SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
标题: SLiCK:利用子序列进行长度限制的关键字发现
作者:Kumari Nishu,Minsik Cho,Devang Naik
链接:点击下载PDF文件
【32】 MusicLIME: Explainable Multimodal Music Understanding
标题: MusicLIME:可解释的多模式音乐理解
作者:Theodoros Sotirou,Vassilis Lyberatos,Orfeas Menis Mastromichalakis,Giorgos Stamou
备注:GitHub repository: this https URL
链接:点击下载PDF文件
【33】 2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
标题: 2D或不是2D:手势表示的抽象性如何影响3D同声手势生成?
作者:Téo Guichoux,Laure Soulier,Nicolas Obin,Catherine Pelachaud
链接:点击下载PDF文件
【34】 DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
标题: DreamHead:通过分层扩散学习时空对应性,以实现音频驱动的会说话的头部合成
作者:Fa-Ting Hong,Yunfei Liu,Yu Li,Changyin Zhou,Fei Yu,Dan Xu
链接:点击下载PDF文件
【35】 Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
标题: 基于扬声器分离HuBERT的自监督音节发现
作者:Ryota Komatsu,Takahiro Shinozaki
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
【36】 Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
标题: 优化构音障碍唤醒词定位:SEARCH 2024 LRDDWWS挑战赛的端到端方法
作者:Shuiyun Liu,Yuxiang Kong,Pengcheng Guo,Weiji Zhuang,Peng Gao,Yujun Wang,Lei Xie
备注:8 pages, Accepted to SLT 2024
链接:点击下载PDF文件
【37】 Speaker Contrastive Learning for Source Speaker Tracing
标题: 用于源说话人追踪的说话人对比学习
作者:Qing Wang,Hongmei Guo,Jian Kang,Mengjie Du,Jie Li,Xiao-Lei Zhang,Lei Xie
备注:7 pages, 2 figures, accepted by SLT
链接:点击下载PDF文件
【38】 Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
标题: 自然环境中用于头部定向的音频驱动强化学习
作者:Wessel Ledder,Yuzhen Qin,Kiki van der Heijden
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
【39】 DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
标题: DiffTAR:音频文本检索的基于扩散的生成建模
作者:Yifei Xin,Xuxin Cheng,Zhihong Zhu,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
【40】 Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
标题: 通过多任务学习从转录的语音音频中获取发音知识
作者:Siqi Sun,Korin Richmond
备注:5 pages
链接:点击下载PDF文件
【41】 Constructing a Singing Style Caption Dataset
标题: 构建歌唱风格字幕数据集
作者:Hyunjong Ok,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
【42】 Efficient Video to Audio Mapper with Visual Scene Detection
标题: 具有视觉场景检测的高效视频到音频映射器
作者:Mingjing Yi,Ming Li
链接:点击下载PDF文件
【43】 Large Language Model Based Generative Error Correction: A Challenge and Baselines forSpeech Recognition, Speaker Tagging, and Emotion Recognition
标题: 基于大语言模型的生成式错误纠正:语音识别、说话人标记和情感识别的挑战和基线
作者:Chao-Han Huck Yang,Taejin Park,Yuan Gong,Yuanchao Li,Zhehuai Chen,Yen-Ting Lin,Chen Chen,Yuchen Hu,Kunal Dhawan,Piotr Żelasko,Chao Zhang,Yun-Nung Chen,Yu Tsao,Jagadeesh Balam,Boris Ginsburg,Sabato Marco Siniscalchi,Eng Siong Chng,Peter Bell,Catherine Lai,Shinji Watanabe,Andreas Stolcke
备注:IEEE SLT 2024. The initial draft version has been done in December 2023. Post-ASR Text Processing and Understanding Community: this https URL
链接:点击下载PDF文件
【44】 Self-supervised Learning for Acoustic Few-Shot Classification
标题: 声学Few-Shot分类的自我监督学习
作者:Jingyong Liang,Bernd Meyer,Issac Ning Lee,Thanh-Toan Do
链接:点击下载PDF文件
【45】 Compositional Audio Representation Learning
标题: 合成音频表示学习
作者:Sripathi Sridhar,Mark Cartwright
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【46】 Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
标题: 整合音频叙述加强多模式第一人称动作识别中的领域概括
作者:Cagri Gungor,Adriana Kovashka
链接:点击下载PDF文件
【47】 A Survey of Foundation Models for Music Understanding
标题: 音乐理解的基础模型综述
作者:Wenjun Li,Ying Cai,Ziyang Wu,Wenyi Zhang,Yifan Chen,Rundong Qi,Mengqi Dong,Peigen Chen,Xiao Dong,Fenghao Shi,Lei Guo,Junwei Han,Bao Ge,Tianming Liu,Lin Gan,Tuo Zhang
备注:20 pages, 2 figures
链接:点击下载PDF文件
【48】 On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
标题: 关于目标说话人提取的注册语音增强的有效性
作者:Junjie Li,Ke Zhang,Shuai Wang,Haizhou Li,Man-Wai Mak,Kong Aik Lee
备注:Accepted by SLT2024
链接:点击下载PDF文件
【49】 ASR Error Correction using Large Language Models
标题: 使用大型语言模型的ASB错误纠正
作者:Rao Ma,Mengjie Qian,Mark Gales,Kate Knill
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
【50】 Multi-Microphone and Multi-Modal Emotion Recognition in Reverbrant Enviroment
标题: 可逆环境中的多麦克风和多模式情感识别
作者:Ohad Cohen,Gershon Hazan,Sharon Gannot
链接:点击下载PDF文件
【51】 Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
标题: 通过预测可解释的声学特征来解释语音情感识别的深度学习嵌入
作者:Satvik Dixit,Daniel M. Low,Gasser Elbanna,Fabio Catania,Satrajit S. Ghosh
链接:点击下载PDF文件
【52】 ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
标题: ESPnet-ZZ:仅使用Python的ESPnet,易于微调和集成
作者:Masao Someki,Kwanghee Choi,Siddhant Arora,William Chen,Samuele Cornell,Jionghao Han,Yifan Peng,Jiatong Shi,Vaibhav Srivastav,Shinji Watanabe
备注:Accepted to SLT 2024
链接:点击下载PDF文件
【53】 Prevailing Research Areas for Music AI in the Era of Foundation Models
标题: 基础模型时代音乐人工智能的主流研究领域
作者:Megan Wei,Mateusz Modrzejewski,Aswin Sivaraman,Dorien Herremans
链接:点击下载PDF文件
【54】 Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
标题: 联合语义知识提取和掩蔽声学建模用于提高可理解度的全频段语音恢复
作者:Xiaoyu Liu,Xu Li,Joan Serrà,Santiago Pascual
备注:Demo link this https URL
链接:点击下载PDF文件
【55】 MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
标题: MacST:通过文本音译进行多口音语音合成以实现口音转换
作者:Sho Inoue,Shuai Wang,Wanxing Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
备注:Project page with Speech Demo: this https URL
链接:点击下载PDF文件
【56】 Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
标题: 儿童与成人二元互动中的自我中心说话者分类:从感知到计算建模
作者:Tiantian Feng,Anfeng Xu,Xuan Shi,Somer Bishop,Shrikanth Narayanan
备注:pre-print under review
链接:点击下载PDF文件
【57】 The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
标题: 2024年VoiceMOS挑战赛的T05系统:从深度图像分类器转移学习到高质量合成语音的Naturalness MOS预测
作者:Kaito Baba,Wataru Nakata,Yuki Saito,Hiroshi Saruwatari
备注:Accepted by IEEE SLT 2024. Our MOS prediction system (UTMOSv2) is available in this https URL
链接:点击下载PDF文件
【58】 Subband Splitting: Simple, Efficient and Effective Technique for Solving Block Permutation Problem in Determined Blind Source Separation
标题: 子带分裂:解决确定盲源分离中块排列问题的简单、高效且有效的技术
作者:Kazuki Matsumoto,Kohei Yatabe
链接:点击下载PDF文件
【59】 DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
标题: DSCSYS:特定领域对比音频预训练
作者:Shengqiang Liu,Da Liu,Anna Wang,Zhiyu Zhang,Jie Gao,Yali Li
链接:点击下载PDF文件
【60】 M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
标题: M$^{3}$V:用于设备引导语音检测的多模式多视图方法
作者:Anna Wang,Da Liu,Zhiyu Zhang,Shengqiang Liu,Jie Gao,Yali Li
链接:点击下载PDF文件
【61】 SafeEar: Content Privacy-Preserving Audio Deepfake Detection
标题: SafeEar:内容隐私保护音频Deepfake检测
作者:Xinfeng Li,Kai Li,Yifan Zheng,Chen Yan,Xiaoyu Ji,Wenyuan Xu
备注:Accepted by ACM CCS 2024. Please cite this paper as "Xinfeng Li, Kai Li, Yifan Zheng, Chen Yan, Xiaoyu Ji, Wenyuan Xu. SafeEar: Content Privacy-Preserving Audio Deepfake Detection. In Proceedings of ACM Conference on Computer and Communications Security (CCS), 2024."
链接:点击下载PDF文件
【62】 Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
标题: 基于转换器的分层对齐和解纠缠跨模式表示的音频文本检索
作者:Yifei Xin,Zhihong Zhu,Xuxin Cheng,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
【63】 Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
标题: 多模式语音Transformer解码器:多模式何时可以提高准确性?
作者:Yiwen Guan,Viet Anh Trinh,Vivek Voleti,Jacob Whitehill
链接:点击下载PDF文件
【64】 Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
标题: Seed-Music:高质量和受控音乐生成的统一框架
作者:Ye Bai,Haonan Chen,Jitong Chen,Zhuo Chen,Yi Deng,Xiaohong Dong,Lamtharn Hantrakul,Weituo Hao,Qingqing Huang,Zhongyi Huang,Dongya Jia,Feihu La,Duc Le,Bochen Li,Chumin Li,Hui Li,Xingxing Li,Shouda Liu,Wei-Tsung Lu,Yiqing Lu,Andrew Shaw,Janne Spijkervet,Yakun Sun,Bo Wang,Ju-Chiang Wang,Yuping Wang,Yuxuan Wang,Ling Xu,Yifeng Yang,Chao Yao,Shuo Zhang,Yang Zhang,Yilin Zhang,Hang Zhao,Ziyi Zhao,Dejian Zhong,Shicen Zhou,Pei Zou
备注:Seed-Music technical report, 20 pages, 5 figures
链接:点击下载PDF文件
【65】 AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
标题: AccentBox:迈向高保真Zero-Shot口音一代
作者:Jinzuomu Zhong,Korin Richmond,Zhiba Su,Siqi Sun
链接:点击下载PDF文件
【66】 Estimating the Completeness of Discrete Speech Units
标题: 估计离散语音单元的完整性
作者:Sung-Lin Yeh,Hao Tang
备注:SLT2024
链接:点击下载PDF文件
标题: MusicLIME:可解释的多模式音乐理解
作者:Theodoros Sotirou,Vassilis Lyberatos,Orfeas Menis Mastromichalakis,Giorgos Stamou
备注:GitHub repository: this https URL
链接:点击下载PDF文件
摘要:多模态模型对于音乐理解任务至关重要,因为它们捕捉了音频和歌词之间复杂的相互作用。然而,随着这些模型变得越来越普遍,对可解释性的需求也在增长理解这些系统如何做出决策对于确保公平、减少偏见和培养信任至关重要。在本文中,我们介绍了MusicLIME,一个模型无关的特征重要性解释方法,专为多模态音乐模型。与传统的单峰方法不同,传统的单峰方法单独分析每种模态,而不考虑它们之间的相互作用,通常会导致不完整或误导性的解释,MusicLIME揭示了音频和抒情特征如何相互作用并有助于预测,提供了模型决策的整体视图。此外,我们通过将局部解释聚合为全局解释来增强局部解释,为用户提供更广泛的模型行为视角。通过这项工作,我们有助于提高多模态音乐模型的可解释性,使用户能够做出明智的选择,并促进更公平,公正和透明的音乐理解系统。摘要:Multimodal models are critical for music understanding tasks, as they capture the complex interplay between audio and lyrics. However, as these models become more prevalent, the need for explainability grows-understanding how these systems make decisions is vital for ensuring fairness, reducing bias, and fostering trust. In this paper, we introduce MusicLIME, a model-agnostic feature importance explanation method designed for multimodal music models. Unlike traditional unimodal methods, which analyze each modality separately without considering the interaction between them, often leading to incomplete or misleading explanations, MusicLIME reveals how audio and lyrical features interact and contribute to predictions, providing a holistic view of the model's decision-making. Additionally, we enhance local explanations by aggregating them into global explanations, giving users a broader perspective of model behavior. Through this work, we contribute to improving the interpretability of multimodal music models, empowering users to make informed choices, and fostering more equitable, fair, and transparent music understanding systems.
【2】 2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
标题: 2D或不是2D:手势表示的抽象性如何影响3D同声手势生成?
作者:Téo Guichoux,Laure Soulier,Nicolas Obin,Catherine Pelachaud
链接:点击下载PDF文件
摘要:共同语言手势是沟通的基础。最近深度学习技术的出现促进了为嵌入式会话代理创建逼真、同步的协同语音手势。“野外”数据集,通过人体姿势检测技术聚合来自YouTube等平台的视频内容,通过提供与语音对齐的2D骨架序列提供了一个可行的解决方案。提升模型的同时发展使得这些2D序列能够转换为3D手势数据库。然而,重要的是要注意,从2D提取的姿态估计的3D姿态本质上是地面实况的近似,其保持在2D域中。这种区别提出了关于手势表示维度对生成的运动质量的影响的问题-据我们所知,这个话题在很大程度上仍未被探索。我们的研究考察了使用2D或3D关节坐标作为训练数据对语音到手势深度生成模型性能的影响。我们采用提升模型将生成的2D姿势序列转换为3D,并评估直接在3D中创建的手势如何与最初在2D中生成的手势叠加,然后转换为3D。我们使用广泛使用的指标在手势生成领域以及用户研究进行客观的评价,定性评估不同的方法。摘要:Co-speech gestures are fundamental for communication. The advent of recent deep learning techniques has facilitated the creation of lifelike, synchronous co-speech gestures for Embodied Conversational Agents. "In-the-wild" datasets, aggregating video content from platforms like YouTube via human pose detection technologies, provide a feasible solution by offering 2D skeletal sequences aligned with speech. Concurrent developments in lifting models enable the conversion of these 2D sequences into 3D gesture databases. However, it is important to note that the 3D poses estimated from the 2D extracted poses are, in essence, approximations of the ground-truth, which remains in the 2D domain. This distinction raises questions about the impact of gesture representation dimensionality on the quality of generated motions - a topic that, to our knowledge, remains largely unexplored. Our study examines the effect of using either 2D or 3D joint coordinates as training data on the performance of speech-to-gesture deep generative models. We employ a lifting model for converting generated 2D pose sequences into 3D and assess how gestures created directly in 3D stack up against those initially generated in 2D and then converted to 3D. We perform an objective evaluation using widely used metrics in the gesture generation field as well as a user study to qualitatively evaluate the different approaches.
【3】 DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
标题: DreamHead:通过分层扩散学习时空对应性,以实现音频驱动的会说话的头部合成
作者:Fa-Ting Hong,Yunfei Liu,Yu Li,Changyin Zhou,Fei Yu,Dan Xu
链接:点击下载PDF文件
摘要:音频驱动的说话头合成努力从提供的音频生成逼真的视频肖像。扩散模型,公认其优越的质量和强大的泛化,已探讨了这项任务。然而,建立一个强大的时间音频线索和相应的空间面部表情与扩散模型之间的对应关系仍然是一个重大的挑战,在说话的头部生成。为了弥合这一差距,我们提出了DreamHead,这是一个分层扩散框架,它可以在不影响模型内在质量和适应性的情况下学习说话头部合成中的时空对应关系。DreamHead学习从音频中预测密集的面部标志作为中间信号,以模拟空间和时间的对应关系。具体地,第一层次的音频到界标扩散首先被设计为在给定音频序列信号的情况下预测时间上平滑且准确的界标序列。然后,第二层次的地标到图像的扩散,进一步提出了产生空间一致的人脸肖像视频,通过建模密集的面部地标和外观之间的空间对应关系。大量的实验表明,提出的DreamHead可以有效地学习时空一致性与设计的分层扩散,并产生高保真音频驱动的多个身份的说话头视频。摘要:Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. However, establishing a robust correspondence between temporal audio cues and corresponding spatial facial expressions with diffusion models remains a significant challenge in talking head generation. To bridge this gap, we present DreamHead, a hierarchical diffusion framework that learns spatial-temporal correspondences in talking head synthesis without compromising the model's intrinsic quality and adaptability.~DreamHead learns to predict dense facial landmarks from audios as intermediate signals to model the spatial and temporal correspondences.~Specifically, a first hierarchy of audio-to-landmark diffusion is first designed to predict temporally smooth and accurate landmark sequences given audio sequence signals. Then, a second hierarchy of landmark-to-image diffusion is further proposed to produce spatially consistent facial portrait videos, by modeling spatial correspondences between the dense facial landmark and appearance. Extensive experiments show that proposed DreamHead can effectively learn spatial-temporal consistency with the designed hierarchical diffusion and produce high-fidelity audio-driven talking head videos for multiple identities.
【4】 Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
标题: 基于扬声器分离HuBERT的自监督音节发现
作者:Ryota Komatsu,Takahiro Shinozaki
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:自监督语音表示学习对于从未转录音频中提取有意义的特征至关重要。最近的进展突出了从与语言单位相关的特征中导出离散符号的潜力,这使得在不同的任务中进行无文本训练成为可能。特别地,预训练的HuBERT(SD-HuBERT)的重复级自蒸馏在从中间Transformer层提取的潜在语音帧表示内诱导音节结构。在SD-HuBERT中,句子级表示是使用特殊的CLS令牌通过自注意层从语音帧特征中累积的。然而,我们观察到,在CLS令牌中聚集的信息与说话者身份的相关性比与语言内容的相关性更高。为了解决这个问题,我们提出了一个语音只自我监督微调方法,分离音节单位从扬声器信息。我们的方法引入说话人扰动作为数据增强,并采用帧级训练目标来防止CLS令牌聚集语言信息。实验结果表明,我们的方法超过了目前最先进的方法在大多数音节分割和音节单元质量指标Libripeech,强调其有效性,促进音节组织内的语音模型。摘要:Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features correlated with linguistic units, which enables text-less training across diverse tasks. In particular, sentence-level Self-Distillation of the pretrained HuBERT (SD-HuBERT) induces syllabic structures within latent speech frame representations extracted from an intermediate Transformer layer. In SD-HuBERT, sentence-level representation is accumulated from speech frame features through self-attention layers using a special CLS token. However, we observe that the information aggregated in the CLS token correlates more with speaker identity than with linguistic content. To address this, we propose a speech-only self-supervised fine-tuning approach that separates syllabic units from speaker information. Our method introduces speaker perturbation as data augmentation and adopts a frame-level training objective to prevent the CLS token from aggregating paralinguistic information. Experimental results show that our approach surpasses the current state-of-the-art method in most syllable segmentation and syllabic unit quality metrics on Librispeech, underscoring its effectiveness in promoting syllabic organization within speech-only models.
【5】 Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
标题: 优化构音障碍唤醒词定位:SEARCH 2024 LRDDWWS挑战赛的端到端方法
作者:Shuiyun Liu,Yuxiang Kong,Pengcheng Guo,Weiji Zhuang,Peng Gao,Yujun Wang,Lei Xie
备注:8 pages, Accepted to SLT 2024
链接:点击下载PDF文件
摘要:语音已经成为跨各种应用程序的广泛接受的用户界面。然而,对于患有构音障碍的个体,其言语的固有可变性构成了重大挑战。本文提出了一种端到端的基于预训练的双过滤器构音障碍唤醒词识别(PD-DWS)系统,用于2024年低资源构音障碍唤醒词识别挑战赛。具体来说,我们的系统从两个关键方面提高了性能:音频建模和双滤波器策略。对于音频建模,我们提出了一种基于预训练的data 2 vec 2(d2 v2)的创新2branch-d2 v2模型,该模型可以通过统一的多任务微调范式同时对自动语音识别(ASR)和唤醒单词定位(WWS)任务进行建模。此外,一个双过滤器的策略,以减少错误接受率(FAR),同时保持相同的错误拒绝率(FRR)。实验结果表明,我们的PD-DWS系统实现了0.00321的FAR和0.005的FRR,在测试B评估集上的总得分为0.00821,在挑战中获得第一名。摘要:Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses significant challenges. This paper presents an end-to-end Pretrain-based Dual-filter Dysarthria Wake-up word Spotting (PD-DWS) system for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge. Specifically, our system improves performance from two key perspectives: audio modeling and dual-filter strategy. For audio modeling, we propose an innovative 2branch-d2v2 model based on the pre-trained data2vec2 (d2v2), which can simultaneously model automatic speech recognition (ASR) and wake-up word spotting (WWS) tasks through a unified multi-task finetuning paradigm. Additionally, a dual-filter strategy is introduced to reduce the false accept rate (FAR) while maintaining the same false reject rate (FRR). Experimental results demonstrate that our PD-DWS system achieves an FAR of 0.00321 and an FRR of 0.005, with a total score of 0.00821 on the test-B eval set, securing first place in the challenge.
【6】 Speaker Contrastive Learning for Source Speaker Tracing
标题: 用于源说话人追踪的说话人对比学习
作者:Qing Wang,Hongmei Guo,Jian Kang,Mengjie Du,Jie Li,Xiao-Lei Zhang,Lei Xie
备注:7 pages, 2 figures, accepted by SLT
链接:点击下载PDF文件
摘要:作为生物识别技术的一种形式,说话人验证系统的安全性至关重要。然而,SV系统本质上容易受到各种类型的攻击,这些攻击可能会损害其准确性和可靠性。其中一种攻击是语音转换,它通过改变各种声音特征来修改一个人的语音,使其听起来像另一个人。这对SV系统构成了重大威胁。为了解决这个问题,IEEE SLT 2024中的源说话人跟踪挑战旨在识别被操纵的语音信号中的源说话人信息。具体来说,SSTC专注于针对语音转换的源说话人验证,以确定两个转换后的语音样本是否来自同一个源说话人。在这项研究中,我们提出了一种基于说话人对比学习的源说话人跟踪方法来学习转换语音中潜在的源说话人信息。为了学习更多的源说话人相关的表示,我们在嵌入提取器的训练过程中使用说话人对比度损失。这种说话人对比损失有助于在几个干扰说话人嵌入中识别真正的源说话人嵌入,使嵌入提取器能够学习转换后的语音中存在的潜在拥有源说话人信息。实验表明,我们提出的说话人对比学习系统在挑战测试集上达到了最低的EER 16.788%,在挑战中获得第一名。摘要:As a form of biometric authentication technology, the security of speaker verification systems is of utmost importance. However, SV systems are inherently vulnerable to various types of attacks that can compromise their accuracy and reliability. One such attack is voice conversion, which modifies a persons speech to sound like another person by altering various vocal characteristics. This poses a significant threat to SV systems. To address this challenge, the Source Speaker Tracing Challenge in IEEE SLT2024 aims to identify the source speaker information in manipulated speech signals. Specifically, SSTC focuses on source speaker verification against voice conversion to determine whether two converted speech samples originate from the same source speaker. In this study, we propose a speaker contrastive learning-based approach for source speaker tracing to learn the latent source speaker information in converted speech. To learn a more source-speaker-related representation, we employ speaker contrastive loss during the training of the embedding extractor. This speaker contrastive loss helps identify the true source speaker embedding among several distractor speaker embeddings, enabling the embedding extractor to learn the potentially possessing source speaker information present in the converted speech. Experiments demonstrate that our proposed speaker contrastive learning system achieves the lowest EER of 16.788% on the challenge test set, securing first place in the challenge.
【7】 Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
标题: 自然环境中用于头部定向的音频驱动强化学习
作者:Wessel Ledder,Yuzhen Qin,Kiki van der Heijden
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:尽管近年来音频信号处理中的深度强化学习(DRL)方法取得了实质性进展,但在人机交互的背景下,用于导航,凝视控制和头部方向控制等任务的音频驱动DRL很少受到关注。在这里,我们提出了一个音频驱动的DRL框架,在该框架中,我们利用深度Q学习来开发一个自主代理,该代理基于立体声语音记录在声学环境中面向说话者。我们的研究结果表明,当在无回声环境(即没有混响)中对语音段进行训练时,智能体学会了以近乎完美的水平执行任务。自然声学环境中混响的存在影响了代理的性能,尽管代理仍然大大优于基线随机代理。最后,我们量化了所提出的DRL方法在自然声学环境中的泛化程度。我们的实验表明,在中等或高混响环境中训练的代理学习的政策推广到低混响环境,但在消声或低混响环境中训练的代理学习的政策没有推广到中等或高混响环境。总而言之,这项研究证明了音频驱动的日间行车学习在头部方向控制等任务中的潜力,并强调了需要培训策略,以便能够在现实世界音频驱动的日间行车学习应用的环境中实现强大的泛化。摘要:Although deep reinforcement learning (DRL) approaches in audio signal processing have seen substantial progress in recent years, audio-driven DRL for tasks such as navigation, gaze control and head-orientation control in the context of human-robot interaction have received little attention. Here, we propose an audio-driven DRL framework in which we utilise deep Q-learning to develop an autonomous agent that orients towards a talker in the acoustic environment based on stereo speech recordings. Our results show that the agent learned to perform the task at a near perfect level when trained on speech segments in anechoic environments (that is, without reverberation). The presence of reverberation in naturalistic acoustic environments affected the agent's performance, although the agent still substantially outperformed a baseline, randomly acting agent. Finally, we quantified the degree of generalization of the proposed DRL approach across naturalistic acoustic environments. Our experiments revealed that policies learned by agents trained on medium or high reverb environments generalized to low reverb environments, but policies learned by agents trained on anechoic or low reverb environments did not generalize to medium or high reverb environments. Taken together, this study demonstrates the potential of audio-driven DRL for tasks such as head-orientation control and highlights the need for training strategies that enable robust generalization across environments for real-world audio-driven DRL applications.
【8】 DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
标题: DiffTAR:音频文本检索的基于扩散的生成建模
作者:Yifei Xin,Xuxin Cheng,Zhihong Zhu,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:现有的音频文本检索(ATR)方法本质上是一种判别模型,其目标是最大化条件似然,表示为p(候选|查询)。然而,这种方法没有考虑内在的数据分布p(查询),导致难以辨别的分布数据。在本研究中,我们尝试透过产生式的观点来解决这个限制,并将声音与文字之间的关系模型化为它们的联合概率p(候选者,查询)。为此,我们提出了一个基于扩散的ATR框架(DiffATR),该框架将ATR建模为一个迭代过程,逐步从噪声中生成联合分布。在整个训练阶段,DiffATR从生成和判别两个角度进行优化:生成器通过生成损失来细化,而特征提取器则从对比损失中受益,从而将两种方法的优点结合起来。在AudioCaps和Clotho数据集上的实验结果表明,该方法具有较好的性能.值得注意的是,在没有任何改变的情况下,我们的DiffATR在域外检索设置中始终表现出强大的性能。摘要:Existing audio-text retrieval (ATR) methods are essentially discriminative models that aim to maximize the conditional likelihood, represented as p(candidates|query). Nevertheless, this methodology fails to consider the intrinsic data distribution p(query), leading to difficulties in discerning out-of-distribution data. In this work, we attempt to tackle this constraint through a generative perspective and model the relationship between audio and text as their joint probability p(candidates,query). To this end, we present a diffusion-based ATR framework (DiffATR), which models ATR as an iterative procedure that progressively generates joint distribution from noise. Throughout its training phase, DiffATR is optimized from both generative and discriminative viewpoints: the generator is refined through a generation loss, while the feature extractor benefits from a contrastive loss, thus combining the merits of both methodologies. Experiments on the AudioCaps and Clotho datasets with superior performances, verify the effectiveness of our approach. Notably, without any alterations, our DiffATR consistently exhibits strong performance in out-of-domain retrieval settings.
【9】 Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
标题: 通过多任务学习从转录的语音音频中获取发音知识
作者:Siqi Sun,Korin Richmond
备注:5 pages
链接:点击下载PDF文件
摘要:最近的工作表明,从传统的基于管道的文本到语音(TTS)前端引导集成的序列到序列(Seq 2 Seq)语言前端的可行性和好处。为了克服自举训练数据的固定词汇覆盖,先前的工作已经提出利用容易访问的转录语音音频作为用于获取未覆盖单词的新发音知识的额外训练源,其依赖于辅助ASR模型作为繁琐的实现流程的一部分。在这项工作中,我们提出了一种替代方法,利用转录的语音音频作为额外的训练源,基于多任务学习(MTL)。实验表明,与基线Seq2Seq前端相比,所提出的基于MTL的方法对于转录语音音频中专门覆盖的单词类型将PER从2.5%降低到1.6%,实现了与先前方法相似的性能,但实现流程要简单得多。摘要:Recent work has shown the feasibility and benefit of bootstrapping an integrated sequence-to-sequence (Seq2Seq) linguistic frontend from a traditional pipeline-based frontend for text-to-speech (TTS). To overcome the fixed lexical coverage of bootstrapping training data, previous work has proposed to leverage easily accessible transcribed speech audio as an additional training source for acquiring novel pronunciation knowledge for uncovered words, which relies on an auxiliary ASR model as part of a cumbersome implementation flow. In this work, we propose an alternative method to leverage transcribed speech audio as an additional training source, based on multi-task learning (MTL). Experiments show that, compared to a baseline Seq2Seq frontend, the proposed MTL-based method reduces PER from 2.5% to 1.6% for those word types covered exclusively in transcribed speech audio, achieving a similar performance to the previous method but with a much simpler implementation flow.
【10】 Constructing a Singing Style Caption Dataset
标题: 构建歌唱风格字幕数据集
作者:Hyunjong Ok,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
摘要:歌唱嗓音的合成与转换已成为嗓音生成的重要子领域,对非条件生成提出了更高的要求。与普通的语音数据不同,生成歌声需要了解各种相关的声乐和音乐特征,例如歌手的音调或情感表达。然而,现有的用于语音生成的开源音频文本数据集往往只捕获非常有限的属性范围,通常缺少音频的音乐特征。为了填补这一空白,我们引入了S2 Cap,这是一个具有不同属性集的音频-文本对数据集。S2 Cap由成对的文本提示和音乐音频样本组成,具有广泛的声乐和音乐属性,包括音高,音量,节奏,情绪,歌手的性别和年龄,音乐类型和情感表达。利用S2 Cap,我们提出了一个有效的新的基线算法的演唱风格的字幕。演唱风格字幕是一个相对于语音生成的任务,生成文本描述的声音特征,这是我们首先提出的。首先,为了减轻音频编码器和文本解码器之间的不对齐,我们提出了一种名为CRESCENDO的新机制,该机制利用正对相似性学习来同步预训练音频编码器的嵌入空间,以获得与文本编码器相似的嵌入。我们还使用歌手的声音来监督模型,歌手的声音被伴奏分离。这种监督允许模型更准确地捕捉声音特征,从而改进演唱风格字幕,更好地反映歌手的风格。数据集和代码可以在 bulurl{https: github.com HJ-Ok S2cap}上找到。摘要:Singing voice synthesis and conversion have emerged as significant subdomains of voice generation, leading to much demands on prompt-conditioned generation. Unlike common voice data, generating a singing voice requires an understanding of various associated vocal and musical characteristics, such as the vocal tone of the singer or emotional expressions. However, existing open-source audio-text datasets for voice generation tend to capture only a very limited range of attributes, often missing musical characteristics of the audio. To fill this gap, we introduce S2Cap, an audio-text pair dataset with a diverse set of attributes. S2Cap consists of pairs of textual prompts and music audio samples with a wide range of vocal and musical attributes, including pitch, volume, tempo, mood, singer's gender and age, and musical genre and emotional expression. Utilizing S2Cap, we suggest an effective novel baseline algorithm for singing style captioning. Singing style captioning is a relative task to voice generation that generates text descriptions of vocal characteristics, which we first suggested. First, to mitigate the misalignment between the audio encoder and the text decoder, we present a novel mechanism called CRESCENDO, which utilizes positive-pair similarity learning to synchronize the embedding spaces of a pretrained audio encoder to get similar embeddings with a text encoder. We additionally supervise the model using the singer's voice, which is demixed by the accompaniment. This supervision allows the model to more accurately capture vocal characteristics, leading to improved singing style captions that better reflect the style of the singer. The dataset and the codes are available at bulurl{https: github.com HJ-Ok S2cap}.
【11】 Efficient Video to Audio Mapper with Visual Scene Detection
标题: 具有视觉场景检测的高效视频到音频映射器
作者:Mingjing Yi,Ming Li
链接:点击下载PDF文件
摘要:视频到音频(V2A)生成的目的是在给定无声视频输入的情况下产生对应的音频。这项任务是特别具有挑战性的,由于跨模态和连续性的视听功能所涉及的。最近的作品在弥合视频和音频之间的域差距方面取得了重大进展,生成与视频内容语义一致的音频。然而,这些方法的一个关键限制是它们不能有效地识别和处理视频中的多个场景,在这种情况下通常导致次优的音频生成。在本文中,我们首先重新实现了一个最先进的V2A模型,稍微修改了轻量级架构,实现了优于基线的结果。然后,我们提出了一个改进的V2A模型,它结合了场景检测器,以解决多个视觉场景之间切换的挑战。在VGGSound上的结果表明,我们的模型可以识别和处理视频中的多个场景,并在保真度和相关性方面取得了优于基线的性能。摘要:Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature of the audio-visual features involved. Recent works have made significant progress in bridging the domain gap between video and audio, generating audio that is semantically aligned with the video content. However, a critical limitation of these approaches is their inability to effectively recognize and handle multiple scenes within a video, often leading to suboptimal audio generation in such cases. In this paper, we first reimplement a state-of-the-art V2A model with a slightly modified light-weight architecture, achieving results that outperform the baseline. We then propose an improved V2A model that incorporates a scene detector to address the challenge of switching between multiple visual scenes. Results on VGGSound show that our model can recognize and handle multiple scenes within a video and achieve superior performance against the baseline for both fidelity and relevance.
【12】 Large Language Model Based Generative Error Correction: A Challenge and Baselines forSpeech Recognition, Speaker Tagging, and Emotion Recognition
标题: 基于大语言模型的生成式错误纠正:语音识别、说话人标记和情感识别的挑战和基线
作者:Chao-Han Huck Yang,Taejin Park,Yuan Gong,Yuanchao Li,Zhehuai Chen,Yen-Ting Lin,Chen Chen,Yuchen Hu,Kunal Dhawan,Piotr Żelasko,Chao Zhang,Yun-Nung Chen,Yu Tsao,Jagadeesh Balam,Boris Ginsburg,Sabato Marco Siniscalchi,Eng Siong Chng,Peter Bell,Catherine Lai,Shinji Watanabe,Andreas Stolcke
备注:IEEE SLT 2024. The initial draft version has been done in December 2023. Post-ASR Text Processing and Understanding Community: this https URL
链接:点击下载PDF文件
摘要:鉴于生成式人工智能技术的最新进展,一个关键问题是大型语言模型(LLM)如何使用来自冻结的预训练自动语音识别(ASR)模型的文本解码结果来增强声学建模任务。为了探索语音处理语言建模的新功能,我们引入了生成式语音转录错误纠正(GenSEC)挑战。这个挑战包括三个后ASR语言建模任务:(i)后ASR转录校正,(ii)说话人标记,以及(iii)情感识别。这些任务旨在模拟未来基于LLM的代理处理基于语音的界面,同时通过利用开放的预训练语言模型或基于代理的API保持对广泛受众的访问。我们还讨论了从基线评估的见解,以及设计未来的评估经验教训。摘要:Given recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations.
【13】 Self-supervised Learning for Acoustic Few-Shot Classification
标题: 声学Few-Shot分类的自我监督学习
作者:Jingyong Liang,Bernd Meyer,Issac Ning Lee,Thanh-Toan Do
链接:点击下载PDF文件
摘要:标记数据有限,自我监督学习是减少标记要求的最重要方法之一。虽然它已经在图像领域得到了广泛的探索,但到目前为止,它在声学领域还没有得到同样多的关注。然而,减少标签是许多声学应用的关键要求。特别是在生物声学中,很少有足够的标签用于完全监督学习。这导致了声学识别器的广泛使用,这些识别器已经在生物声学任务的不相关数据上进行了预训练。我们认为,在实际任务数据上进行训练,并将自我监督的预训练与Few-Shot分类相结合是一种优越的方法,即使只有少数标签可用,也能够提供高精度。为此,我们引入并评估了一种新的架构,该架构将基于CNN的预处理与基于状态空间模型(SSM)的特征提取相结合。这种组合的动机是,基于CNN的网络很难有效地捕获时间信息,这对于分类声学信号至关重要。另一方面,SSM,特别是S4和Mamba,已被证明具有捕获序列数据中的长程依赖性的出色能力。我们使用实际任务数据上的对比学习和随后的微调来预训练这个架构,这些数据是非常少量的标记数据。我们在标准基准测试以及真实数据上评估了该拟议架构的($n$-shot,$n$-class)分类性能。我们的评估表明,它优于国家的最先进的架构上的Few-Shot分类问题。摘要:Labelled data are limited and self-supervised learning is one of the most important approaches for reducing labelling requirements. While it has been extensively explored in the image domain, it has so far not received the same amount of attention in the acoustic domain. Yet, reducing labelling is a key requirement for many acoustic applications. Specifically in bioacoustic, there are rarely sufficient labels for fully supervised learning available. This has led to the widespread use of acoustic recognisers that have been pre-trained on unrelated data for bioacoustic tasks. We posit that training on the actual task data and combining self-supervised pre-training with few-shot classification is a superior approach that has the ability to deliver high accuracy even when only a few labels are available. To this end, we introduce and evaluate a new architecture that combines CNN-based preprocessing with feature extraction based on state space models (SSMs). This combination is motivated by the fact that CNN-based networks alone struggle to capture temporal information effectively, which is crucial for classifying acoustic signals. SSMs, specifically S4 and Mamba, on the other hand, have been shown to have an excellent ability to capture long-range dependencies in sequence data. We pre-train this architecture using contrastive learning on the actual task data and subsequent fine-tuning with an extremely small amount of labelled data. We evaluate the performance of this proposed architecture for ($n$-shot, $n$-class) classification on standard benchmarks as well as real-world data. Our evaluation shows that it outperforms state-of-the-art architectures on the few-shot classification problem.
【14】 Compositional Audio Representation Learning
标题: 合成音频表示学习
作者:Sripathi Sridhar,Mark Cartwright
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:人类的听觉感知本质上是合成的--我们从具有多个声音事件的听觉场景中识别听觉流。然而,这样的听觉场景通常使用不解开组成声源的剪辑级表示来表示。在这项工作中,我们学习了以源为中心的音频表示,其中每个声源都使用嵌入在音频表示中的不同的、分离的源来表示。我们提出了两种新的方法来学习以源为中心的音频表示:分类指导的监督模型和特征重建指导的无监督模型,这两种方法都优于基线。我们彻底评估这两种方法的设计选择使用音频分类任务。我们发现,监督有利于学习以源为中心的表示,并且重建音频特征比重建频谱图更有用,以学习无监督的以源为中心的表示。利用以源为中心的模型可以帮助释放机器听力中更大的可解释性和更灵活的解码潜力。摘要:Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not disentangle the constituent sound sources. In this work, we learn source-centric audio representations where each sound source is represented using a distinct, disentangled source embedding in the audio representation. We propose two novel approaches to learning source-centric audio representations: a supervised model guided by classification and an unsupervised model guided by feature reconstruction, both of which outperform the baselines. We thoroughly evaluate the design choices of both approaches using an audio classification task. We find that supervision is beneficial to learn source-centric representations, and that reconstructing audio features is more useful than reconstructing spectrograms to learn unsupervised source-centric representations. Leveraging source-centric models can help unlock the potential of greater interpretability and more flexible decoding in machine listening.
【15】 Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
标题: 整合音频叙述加强多模式第一人称动作识别中的领域概括
作者:Cagri Gungor,Adriana Kovashka
链接:点击下载PDF文件
摘要:由于可穿戴摄像头的广泛使用,第一人称活动识别正在迅速发展,但面临着不同环境中域转移的挑战,例如不同的对象或背景场景。我们提出了一个多模态框架,通过整合运动,音频和外观特征,提高域泛化。主要贡献包括分析音频和运动特征对域转移的弹性,使用音频叙述来增强音频-文本对齐,以及在音频和视觉叙述之间应用一致性评级来优化音频在训练过程中识别的影响。我们的方法在ARGO 1 M数据集上实现了最先进的性能,有效地概括了看不见的场景和位置。摘要:First-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multimodal framework that improves domain generalization by integrating motion, audio, and appearance features. Key contributions include analyzing the resilience of audio and motion features to domain shifts, using audio narrations for enhanced audio-text alignment, and applying consistency ratings between audio and visual narrations to optimize the impact of audio in recognition during training. Our approach achieves state-of-the-art performance on the ARGO1M dataset, effectively generalizing across unseen scenarios and locations.
【16】 A Survey of Foundation Models for Music Understanding
标题: 音乐理解的基础模型综述
作者:Wenjun Li,Ying Cai,Ziyang Wu,Wenyi Zhang,Yifan Chen,Rundong Qi,Mengqi Dong,Peigen Chen,Xiao Dong,Fenghao Shi,Lei Guo,Junwei Han,Bao Ge,Tianming Liu,Lin Gan,Tuo Zhang
备注:20 pages, 2 figures
链接:点击下载PDF文件
摘要:音乐在日常生活中至关重要,满足情感和娱乐需求,并将我们个人,社会和文化联系起来。更好地理解音乐可以增强我们的情感,认知技能和文化联系。人工智能(AI)的快速发展引入了分析音乐的新方法,旨在复制人类对音乐的理解并提供相关服务。传统的模型主要关注音频特征和简单的任务,而最近发展起来的大型语言模型(LLM)和基础模型(FM)通过整合语义信息和展示强大的推理能力,在各个领域表现出色,可以捕获复杂的音乐特征和模式,将音乐与语言结合起来,并包含丰富的音乐,情感和心理知识。因此,它们有潜力从语义的角度处理复杂的音乐理解任务,产生更接近人类感知的输出。据我们所知,这项工作是人工智能技术和音乐理解交叉的早期评论之一。我们调查,分析,并测试最近的大型音乐基础模型在他们的音乐理解能力。我们还讨论了它们的局限性,并提出了未来可能的发展方向,为该领域的研究人员提供了见解。摘要:Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance our emotions, cognitive skills, and cultural connections. The rapid advancement of artificial intelligence (AI) has introduced new ways to analyze music, aiming to replicate human understanding of music and provide related services. While the traditional models focused on audio features and simple tasks, the recent development of large language models (LLMs) and foundation models (FMs), which excel in various fields by integrating semantic information and demonstrating strong reasoning abilities, could capture complex musical features and patterns, integrate music with language and incorporate rich musical, emotional and psychological knowledge. Therefore, they have the potential in handling complex music understanding tasks from a semantic perspective, producing outputs closer to human perception. This work, to our best knowledge, is one of the early reviews of the intersection of AI techniques and music understanding. We investigated, analyzed, and tested recent large-scale music foundation models in respect of their music comprehension abilities. We also discussed their limitations and proposed possible future directions, offering insights for researchers in this field.
【17】 On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
标题: 关于目标说话人提取的注册语音增强的有效性
作者:Junjie Li,Ke Zhang,Shuai Wang,Haizhou Li,Man-Wai Mak,Kong Aik Lee
备注:Accepted by SLT2024
链接:点击下载PDF文件
摘要:深度学习技术显著提高了目标说话人提取(TSE)任务的性能。为了提高这些算法在训练数据不足时的泛化能力和鲁棒性,数据增强是一种常用的技术。不同于典型的数据增强应用于语音混合,这项工作彻底调查的有效性,增强注册语音空间。我们发现,对于预训练和联合优化的扬声器编码器,直接增强注册语音会导致一致的性能改善。除了传统的方法,如噪声和混响添加,我们提出了一种新的增强方法称为自估计语音增强(SSA)。Libri2Mix测试集上的实验结果表明,我们提出的方法可以实现高达2.5 dB的改善。摘要:Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a commonly adopted technique. Unlike typical data augmentation applied to speech mixtures, this work thoroughly investigates the effectiveness of augmenting the enrollment speech space. We found that for both pretrained and jointly optimized speaker encoders, directly augmenting the enrollment speech leads to consistent performance improvement. In addition to conventional methods such as noise and reverberation addition, we propose a novel augmentation method called self-estimated speech augmentation (SSA). Experimental results on the Libri2Mix test set show that our proposed method can achieve an improvement of up to 2.5 dB.
【18】 ASR Error Correction using Large Language Models
标题: 使用大型语言模型的ASB错误纠正
作者:Rao Ma,Mengjie Qian,Mark Gales,Kate Knill
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:纠错(EC)模型在改进自动语音识别(ASR)译文、提高译文的可读性和质量方面起着至关重要的作用。在不需要访问底层代码或模型权重的情况下,EC可以提高性能并为黑盒ASR系统提供域自适应。这项工作研究了使用大型语言模型(LLM)在不同的场景中进行纠错。1-最佳ASR假设通常用作EC模型的输入。我们建议使用ASR N-最佳列表来构建高性能的EC模型,该列表应该为校正过程提供更多的上下文信息。此外,标准EC模型的生成过程在可以生成任何输出序列的意义上是不受限制的。对于某些场景,例如看不见的域,这种灵活性可能会影响性能。为了解决这个问题,我们引入了一种基于N-最佳列表或ASR格的约束解码方法。最后,大多数EC模型都是针对特定的ASR系统进行训练的,每当底层ASR系统发生变化时,都需要重新训练。本文探讨了EC模型对不同ASR系统的输出进行操作的能力。该概念进一步扩展到使用LLM(诸如ChatGPT)的zero-shot纠错。在三个标准数据集上的实验证明了我们提出的方法对换能器和基于注意力的编码器-解码器ASR系统的有效性。此外,所提出的方法可以作为一种有效的方法,模型集成。摘要:Error correction (EC) models play a crucial role in refining Automatic Speech Recognition (ASR) transcriptions, enhancing the readability and quality of transcriptions. Without requiring access to the underlying code or model weights, EC can improve performance and provide domain adaptation for black-box ASR systems. This work investigates the use of large language models (LLMs) for error correction across diverse scenarios. 1-best ASR hypotheses are commonly used as the input to EC models. We propose building high-performance EC models using ASR N-best lists which should provide more contextual information for the correction process. Additionally, the generation process of a standard EC model is unrestricted in the sense that any output sequence can be generated. For some scenarios, such as unseen domains, this flexibility may impact performance. To address this, we introduce a constrained decoding approach based on the N-best list or an ASR lattice. Finally, most EC models are trained for a specific ASR system requiring retraining whenever the underlying ASR system is changed. This paper explores the ability of EC models to operate on the output of different ASR systems. This concept is further extended to zero-shot error correction using LLMs, such as ChatGPT. Experiments on three standard datasets demonstrate the efficacy of our proposed methods for both Transducer and attention-based encoder-decoder ASR systems. In addition, the proposed method can serve as an effective method for model ensembling.
【19】 Multi-Microphone and Multi-Modal Emotion Recognition in Reverbrant Enviroment
标题: 可逆环境中的多麦克风和多模式情感识别
作者:Ohad Cohen,Gershon Hazan,Sharon Gannot
链接:点击下载PDF文件
摘要:本文提出了一种多模态情感识别(MER)系统,旨在提高情感识别的准确性,在具有挑战性的声学条件。我们的方法结合了修改和扩展的层次令牌语义音频Transformer(HTS-AT)的多通道音频处理与R(2+1)D卷积神经网络(CNN)模型的视频分析。我们评估我们提出的方法上的混响版本的瑞尔森视听数据库的情感语音和歌曲(RAVDESS)数据集使用合成和真实世界的房间脉冲响应(RIR)。我们的研究结果表明,整合音频和视频模态产生优越的性能相比,单模态的方法,特别是在具有挑战性的声学条件。此外,我们表明,多模态(视听)的方法,利用多个麦克风优于其单麦克风对应。摘要:This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio Transformer (HTS-AT) for multi-channel audio processing with an R(2+1)D Convolutional Neural Networks (CNN) model for video analysis. We evaluate our proposed method on a reverberated version of the Ryerson audio-visual database of emotional speech and song (RAVDESS) dataset using synthetic and real-world Room Impulse Responsess (RIRs). Our results demonstrate that integrating audio and video modalities yields superior performance compared to uni-modal approaches, especially in challenging acoustic conditions. Moreover, we show that the multimodal (audiovisual) approach that utilizes multiple microphones outperforms its single-microphone counterpart.
【20】 Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
标题: 通过预测可解释的声学特征来解释语音情感识别的深度学习嵌入
作者:Satvik Dixit,Daniel M. Low,Gasser Elbanna,Fabio Catania,Satrajit S. Ghosh
链接:点击下载PDF文件
摘要:预训练的深度学习嵌入在语音情感识别(SER)中一直表现出优于手工制作的声学特征的性能。然而,与具有明确物理意义的声学特征不同,这些嵌入缺乏明确的可解释性。解释这些嵌入对于在医疗保健和安全应用中建立信任以及推进对其中编码的声学信息的科学理解至关重要。本文提出了一种改进的探测方法来解释SER空间中的深度学习嵌入。我们预测可解释的声学特征(例如,f0,响度)从(i)嵌入的完整集合和(ii)被识别为对于预测每个情感最重要的嵌入维度的子集。如果最重要维度的子集比所有维度更好地预测给定的情感,并且还更准确地预测特定的声学特征,则我们推断这些声学特征对于给定任务的嵌入模型很重要。我们使用WavLM嵌入和eGeMAPS声学特征作为音频表示进行了实验,将我们的方法应用于RAVDESS和SAVEE情感语音数据集。基于此评估,我们证明了能量,频率,频谱和时间类别的声学特征提供了减少的信息,以SER在该顺序,演示了实用程序的探测分类器方法相关的嵌入到可解释的声学特征。摘要:Pre-trained deep learning embeddings have consistently shown superior performance over handcrafted acoustic features in speech emotion recognition (SER). However, unlike acoustic features with clear physical meaning, these embeddings lack clear interpretability. Explaining these embeddings is crucial for building trust in healthcare and security applications and advancing the scientific understanding of the acoustic information that is encoded in them. This paper proposes a modified probing approach to explain deep learning embeddings in the SER space. We predict interpretable acoustic features (e.g., f0, loudness) from (i) the complete set of embeddings and (ii) a subset of the embedding dimensions identified as most important for predicting each emotion. If the subset of the most important dimensions better predicts a given emotion than all dimensions and also predicts specific acoustic features more accurately, we infer those acoustic features are important for the embedding model for the given task. We conducted experiments using the WavLM embeddings and eGeMAPS acoustic features as audio representations, applying our method to the RAVDESS and SAVEE emotional speech datasets. Based on this evaluation, we demonstrate that Energy, Frequency, Spectral, and Temporal categories of acoustic features provide diminishing information to SER in that order, demonstrating the utility of the probing classifier method to relate embeddings to interpretable acoustic features.
【21】 ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
标题: ESPnet-ZZ:仅使用Python的ESPnet,易于微调和集成
作者:Masao Someki,Kwanghee Choi,Siddhant Arora,William Chen,Samuele Cornell,Jionghao Han,Yifan Peng,Jiatong Shi,Vaibhav Srivastav,Shinji Watanabe
备注:Accepted to SLT 2024
链接:点击下载PDF文件
摘要:我们介绍ESPnet-EZ,一个开源的语音处理工具包ESPnet的扩展,旨在快速,方便地开发语音模型。ESPnet-EZ专注于两个主要方面:(i)在各种任务上对现有ESPnet模型进行简单的微调和推理,以及(ii)与流行的深度神经网络框架(如PyTorch-Lightning,Hugging Face Transformers和datasets以及Lhotse)轻松集成。通过将继承自Kaldi的ESPnet设计选择替换为仅使用Python的无Bash接口,我们大大减少了构建、调试和使用新模型所需的工作量。例如,为了微调语音基础模型,与ESPnet相比,ESPnet-EZ将新编写的代码数量减少了2.7倍,相关代码数量减少了6.7倍,同时大大减少了Bash脚本依赖性。ESPnet-EZ的代码库是公开的。摘要:We introduce ESPnet-EZ, an extension of the open-source speech processing toolkit ESPnet, aimed at quick and easy development of speech models. ESPnet-EZ focuses on two major aspects: (i) easy fine-tuning and inference of existing ESPnet models on various tasks and (ii) easy integration with popular deep neural network frameworks such as PyTorch-Lightning, Hugging Face transformers and datasets, and Lhotse. By replacing ESPnet design choices inherited from Kaldi with a Python-only, Bash-free interface, we dramatically reduce the effort required to build, debug, and use a new model. For example, to fine-tune a speech foundation model, ESPnet-EZ, compared to ESPnet, reduces the number of newly written code by 2.7x and the amount of dependent code by 6.7x while dramatically reducing the Bash script dependencies. The codebase of ESPnet-EZ is publicly available.
【22】 Prevailing Research Areas for Music AI in the Era of Foundation Models
标题: 基础模型时代音乐人工智能的主流研究领域
作者:Megan Wei,Mateusz Modrzejewski,Aswin Sivaraman,Dorien Herremans
链接:点击下载PDF文件
摘要:随着基础模型研究的最新进展,在过去几年中,生成音乐AI应用程序激增。随着人工智能生成或人工智能增强音乐的想法变得越来越主流,音乐人工智能社区的许多研究人员可能想知道还剩下什么研究途径。关于音乐生成模型,我们概述了目前的研究领域,有显着的探索空间。首先,我们提出的问题,这些生成模型的基本表示和调查的可解释性的方法。接下来,我们将讨论音乐数据集的现状及其局限性。然后,我们概述了不同的生成模型,评估这些模型的形式,以及它们的计算约束 限制。随后,我们强调了这些生成模型的应用程序扩展到多种形式和艺术家的工作流程以及音乐教育系统的集成。最后,我们调查了生成音乐的潜在版权影响,并讨论了保护音乐家权利的策略。虽然这并不意味着是详尽的,但我们的调查引起了人们对音乐基金会模型所支持的各种研究方向的关注。摘要:In tandem with the recent advancements in foundation model research, there has been a surge of generative music AI applications within the past few years. As the idea of AI-generated or AI-augmented music becomes more mainstream, many researchers in the music AI community may be wondering what avenues of research are left. With regards to music generative models, we outline the current areas of research with significant room for exploration. Firstly, we pose the question of foundational representation of these generative models and investigate approaches towards explainability. Next, we discuss the current state of music datasets and their limitations. We then overview different generative models, forms of evaluating these models, and their computational constraints limitations. Subsequently, we highlight applications of these generative models towards extensions to multiple modalities and integration with artists' workflow as well as music education systems. Finally, we survey the potential copyright implications of generative music and discuss strategies for protecting the rights of musicians. While it is not meant to be exhaustive, our survey calls to attention a variety of research directions enabled by music foundation models.
【23】 Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
标题: 联合语义知识提取和掩蔽声学建模用于提高可理解度的全频段语音恢复
作者:Xiaoyu Liu,Xu Li,Joan Serrà,Santiago Pascual
备注:Demo link this https URL
链接:点击下载PDF文件
摘要:语音恢复的目标是在考虑各种失真的情况下,恢复具有高质量和可懂度的全频带语音。MaskSR是最近提出的用于此任务的生成模型。与其他同类模型一样,MaskSR达到了高质量,但正如我们所展示的那样,可理解性可以大大提高。我们通过使用预先训练的自监督教师模型,通过预测目标语音的语义表示来提升MaskSR的语音编码器组件。然后,掩蔽的语言模型的条件下学习的语义特征,以预测声学令牌编码的目标语音的低级别的频谱细节。我们发现,在相同的MaskSR模型容量和推理时间下,所提出的模型MaskSR2显着降低了单词错误率,这是可理解性的典型指标。MaskSR2在提供卓越质量的同时,还实现了与其他型号相比具有竞争力的字错误率。消融研究显示了各种语义表征的有效性。摘要:Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substantially improved. We do so by boosting the speech encoder component of MaskSR with predictions of semantic representations of the target speech, using a pre-trained self-supervised teacher model. Then, a masked language model is conditioned on the learned semantic features to predict acoustic tokens that encode low level spectral details of the target speech. We show that, with the same MaskSR model capacity and inference time, the proposed model, MaskSR2, significantly reduces the word error rate, a typical metric for intelligibility. MaskSR2 also achieves competitive word error rate among other models, while providing superior quality. An ablation study shows the effectiveness of various semantic representations.
【24】 MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
标题: MacST:通过文本音译进行多口音语音合成以实现口音转换
作者:Sho Inoue,Shuai Wang,Wanxing Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
备注:Project page with Speech Demo: this https URL
链接:点击下载PDF文件
摘要:在口音语音转换或口音转换中,我们寻求在语音中将口音彼此转换,同时保留说话者身份和语义内容。在这项研究中,我们制定了一个新的方法来创建多口音的语音样本,从而对口音的语音样本由同一扬声器,通过文本音译训练口音转换系统。我们首先使用大型语言模型(LLM)生成音译文本,然后将其输入多语言TTS模型以合成带口音的英语语音。作为参考系统,我们在合成并行语料库上建立了一个序列到序列模型来进行口音转换。我们验证了所提出的方法为母语和非母语的英语。主观和客观的评价进一步验证了我们的数据集在口音转换研究中的有效性。摘要:In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies.
【25】 Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
标题: 儿童与成人二元互动中的自我中心说话者分类:从感知到计算建模
作者:Tiantian Feng,Anfeng Xu,Xuan Shi,Somer Bishop,Shrikanth Narayanan
备注:pre-print under review
链接:点击下载PDF文件
摘要:自闭症谱系障碍(ASD)是一种神经发育状况,其特征在于社交,重复行为和感觉处理方面的挑战。ASD的一个重要研究领域是评估儿童在治疗期间随时间的行为变化。具有此目标的标准协议是BOSCC,其涉及儿童和执行预定义的一组活动的临床医生之间的二元交互。理解儿童在这些互动中的行为的一个基本方面是自动语音理解,特别是识别谁在说话以及何时说话。在这方面的传统方法严重依赖于从旁观者的角度记录的语音样本,并且对以自我为中心的语音建模的研究有限。在这项研究中,我们设计了一个实验,从自我中心的角度使用可穿戴传感器在BOSCC采访中进行语音采样,并探索预训练Ego 4D语音样本,以提高儿童-成人说话者分类的二元互动。我们的研究结果突出了以自我为中心的语音收集和预训练,以提高说话人分类的准确性的潜力。摘要:Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral changes over time during treatment. The standard protocol with this objective is BOSCC, which involves dyadic interactions between a child and clinicians performing a pre-defined set of activities. A fundamental aspect of understanding children's behavior in these interactions is automatic speech understanding, particularly identifying who speaks and when. Conventional approaches in this area heavily rely on speech samples recorded from a spectator perspective, and there is limited research on egocentric speech modeling. In this study, we design an experiment to perform speech sampling in BOSCC interviews from an egocentric perspective using wearable sensors and explore pre-training Ego4D speech samples to enhance child-adult speaker classification in dyadic interactions. Our findings highlight the potential of egocentric speech collection and pre-training to improve speaker classification accuracy.
【26】 The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
标题: 2024年VoiceMOS挑战赛的T05系统:从深度图像分类器转移学习到高质量合成语音的Naturalness MOS预测
作者:Kaito Baba,Wataru Nakata,Yuki Saito,Hiroshi Saruwatari
备注:Accepted by IEEE SLT 2024. Our MOS prediction system (UTMOSv2) is available in this https URL
链接:点击下载PDF文件
摘要:我们为2024年的VoiceMOS挑战赛(VMC)展示了我们的系统(表示为T05)。我们的系统是为VMC 2024 Track 1设计的,该系统专注于准确预测高质量合成语音的自然度平均意见得分(MOS)。除了预训练的基于自监督学习(SSL)的语音特征提取器外,我们的系统还集成了预训练的图像特征提取器,以捕获语音频谱图中观察到的合成语音的差异。我们首先分别训练两个MOS预测器,使用基于SSL或基于频谱的功能。然后,我们使用两个提取的特征的融合来微调两个预测器以获得更好的MOS预测。在VMC 2024 Track 1中,我们的T05系统在16个评估指标中的7个指标中获得第一名,在其余9个指标中获得第二名,与排名第三及以下的系统相比有显着差异。我们还报告了我们的消融研究的结果,以调查我们的系统的基本因素。摘要:We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-quality synthetic speech. In addition to a pretrained self-supervised learning (SSL)-based speech feature extractor, our system incorporates a pretrained image feature extractor to capture the difference of synthetic speech observed in speech spectrograms. We first separately train two MOS predictors that use either of an SSL-based or spectrogram-based feature. Then, we fine-tune the two predictors for better MOS prediction using the fusion of two extracted features. In the VMC 2024 Track 1, our T05 system achieved first place in 7 out of 16 evaluation metrics and second place in the remaining 9 metrics, with a significant difference compared to those ranked third and below. We also report the results of our ablation study to investigate essential factors of our system.
【27】 Subband Splitting: Simple, Efficient and Effective Technique for Solving Block Permutation Problem in Determined Blind Source Separation
标题: 子带分裂:解决确定盲源分离中块排列问题的简单、高效且有效的技术
作者:Kazuki Matsumoto,Kohei Yatabe
链接:点击下载PDF文件
摘要:置换问题的求解是确定性盲源分离的关键。现有的方法,如独立向量分析(IVA)和独立低秩矩阵分析(ILRMA),解决置换问题的源信号的频率分量的同现建模。这些方法中的剩余挑战之一是块置换问题,这可能导致差的分离结果。在本文中,我们提出了一个简单而有效的技术解决块置换问题。所提出的技术将整个频率分成重叠的子带,并顺序应用BSS方法(例如,IVA、ILRMA或任何其它方法)到每个子带。由于问题的大小减少了分裂,BSS方法可以有效地工作在每个子带。然后,通过使用一个子带中的分离结果作为其他子带的初始值来对齐子带之间的排列。实验结果表明,该方法在不增加总计算量的情况下,显著提高了分离性能。摘要:Solving the permutation problem is essential for determined blind source separation (BSS). Existing methods, such as independent vector analysis (IVA) and independent low-rank matrix analysis (ILRMA), tackle the permutation problem by modeling the co-occurrence of the frequency components of source signals. One of the remaining challenges in these methods is the block permutation problem, which may lead to poor separation results. In this paper, we propose a simple and effective technique for solving the block permutation problem. The proposed technique splits the entire frequencies into overlapping subbands and sequentially applies a BSS method (e.g., IVA, ILRMA, or any other method) to each subband. Since the problem size is reduced by the splitting, the BSS method can effectively work in each subband. Then, the permutations between the subbands are aligned by using the separation result in one subband as the initial values for the other subbands. Experimental results showed that the proposed technique remarkably improved the separation performance without increasing the total computational cost.
【28】 DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
标题: DSCSYS:特定领域对比音频预训练
作者:Shengqiang Liu,Da Liu,Anna Wang,Zhiyu Zhang,Jie Gao,Yali Li
链接:点击下载PDF文件
摘要:分析真实世界的多模态信号是智能语音助理(IVA)的一项重要且具有挑战性的任务。主流方法在使用预训练的音频模型和文本模型的IVA的各种下游任务上取得了显着的性能。然而,这些模型是独立地预先训练的,并且通常针对与目标域不同的任务,从而导致下游任务的次优模态表示。此外,在许多领域,收集足够的语言-音频对是极其困难的,并且转录原始音频也需要很高的专业技能,使得联合预训练变得困难甚至不可行。为了解决这些痛点,我们提出了DSCEMP 3,这是一个简单有效的框架,可以只使用原始音频信号输入进行语言音频预训练。具体而言,DSCCast通过ASR系统将原始音频信号转换为文本,并结合对比学习目标和语言-音频匹配目标来对齐音频和ASR传输。我们在12,107小时的车载域音频上预训练DSCEMP 3。两个下游任务的实证结果表明,虽然概念上简单,DSCERAGE显着优于基线模型在所有指标,显示特定领域的IVA应用程序的巨大潜力。摘要:Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on various downstream tasks of IVAs with pre-trained audio models and text models. However, these models are pre-trained independently and usually on tasks different from target domains, resulting in sub-optimal modality representations for downstream tasks. Moreover, in many domains, collecting enough language-audio pairs is extremely hard, and transcribing raw audio also requires high professional skills, making it difficult or even infeasible to joint pre-training. To address these painpoints, we propose DSCLAP, a simple and effective framework that enables language-audio pre-training with only raw audio signal input. Specifically, DSCLAP converts raw audio signals into text via an ASR system and combines a contrastive learning objective and a language-audio matching objective to align the audio and ASR transcriptions. We pre-train DSCLAP on 12,107 hours of in-vehicle domain audio. Empirical results on two downstream tasks show that while conceptually simple, DSCLAP significantly outperforms the baseline models in all metrics, showing great promise for domain-specific IVAs applications.
【29】 M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
标题: M$^{3}$V:用于设备引导语音检测的多模式多视图方法
作者:Anna Wang,Da Liu,Zhiyu Zhang,Shengqiang Liu,Jie Gao,Yali Li
链接:点击下载PDF文件
摘要:为了与虚拟语音助手进行更自然、更人性化的交互,该领域最近的研究集中在全双工交互模式上,而不依赖于重复的唤醒词。这要求在具有复杂声源的场景中,语音助理必须将话语分类为面向设备或非面向设备。由文本和语音共同建模的双编码器结构已成为面向设备的语音检测的典范。然而,在实践中,由于自动语音识别(ASR)不可避免的错误,这些模型经常对未对齐的输入对产生不正确的预测。为了解决这一挑战,我们提出了M$^{3}$V,一种用于设备定向语音检测的多模态多视图方法,我们把这个问题定义为一个多视角的学习任务,它引入了单峰视角和文本,音频对齐视图在网络中除了多模态。实验结果表明,M$^{3}$V显著优于仅使用单一或多模态训练的模型,并首次超过人类对ASR错误数据的判断性能。摘要:With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with complex sound sources, the voice assistant must classify utterances as device-oriented or non-device-oriented. The dual-encoder structure, which is jointly modeled by text and speech, has become the paradigm of device-directed speech detection. However, in practice, these models often produce incorrect predictions for unaligned input pairs due to the unavoidable errors of automatic speech recognition (ASR).To address this challenge, we propose M$^{3}$V, a multi-modal multi-view approach for device-directed speech detection, which frames we frame the problem as a multi-view learning task that introduces unimodal views and a text-audio alignment view in the network besides the multi-modal. Experimental results show that M$^{3}$V significantly outperforms models trained using only single or multi-modality and surpasses human judgment performance on ASR error data for the first time.
【30】 SafeEar: Content Privacy-Preserving Audio Deepfake Detection
标题: SafeEar:内容隐私保护音频Deepfake检测
作者:Xinfeng Li,Kai Li,Yifan Zheng,Chen Yan,Xiaoyu Ji,Wenyuan Xu
备注:Accepted by ACM CCS 2024. Please cite this paper as "Xinfeng Li, Kai Li, Yifan Zheng, Chen Yan, Xiaoyu Ji, Wenyuan Xu. SafeEar: Content Privacy-Preserving Audio Deepfake Detection. In Proceedings of ACM Conference on Computer and Communications Security (CCS), 2024."
链接:点击下载PDF文件
摘要:文本到语音(TTS)和语音转换(VC)模型在生成逼真和自然的音频方面表现出了卓越的性能。然而,他们的黑暗面,音频deepfake对社会和个人都构成了重大威胁。现有的对策主要集中在基于完整的原始音频记录来确定语音的私密性,然而,原始音频记录通常包含隐私内容。这种疏忽可能会抑制许多应用程序的deepfake检测,特别是在涉及商业秘密等敏感信息的情况下。在本文中,我们提出了SafeEar,这是一个新的框架,旨在检测deepfake音频,而不依赖于访问其中的语音内容。我们的关键思想是将神经音频编解码器设计成一种新颖的解耦模型,该模型很好地将语义和声学信息从音频样本中分离出来,并且仅使用声学信息(例如,韵律和音色)用于深度伪造检测。通过这种方式,没有语义内容将暴露给检测器。为了克服在没有语义线索的情况下识别各种deepfake音频的挑战,我们用真实世界的编解码器增强来增强我们的deepfake检测器。在四个基准数据集上进行的广泛实验证明了SafeEar在检测各种深度伪造技术方面的有效性,其等错误率(EER)降至2.02%。同时,它屏蔽了五种语言的语音内容,使其不被机器和人类的听觉分析破译,这一点在我们的用户研究和单词错误率(WER)中都超过了93.93%。此外,我们为反deepfake和反内容恢复评估构建的基准有助于为音频隐私保护和deepfake检测领域的未来研究提供基础。摘要:Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world codec augmentation. Extensive experiments conducted on four benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.02%. Simultaneously, it shields five-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.93% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection.
【31】 Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
标题: 基于转换器的分层对齐和解纠缠跨模式表示的音频文本检索
作者:Yifei Xin,Zhihong Zhu,Xuxin Cheng,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:大多数现有的音频-文本检索(ATR)方法通常依赖于单级交互来关联音频和文本,限制了它们对齐不同模态的能力,并导致次优匹配。在这项工作中,我们提出了一种新的ATR框架,利用两个流的Transformers结合层次对齐(THA)模块,以确定音频和文本之间的不同Transformer块的多级对应关系。此外,目前的ATR方法主要集中在学习全局级表示,错过了复杂的细节,以捕捉对应于文本语义的音频出现。为了弥合这一差距,我们引入了一个解开跨模态表示(DCR)的方法,解开高维特征紧凑的潜在因素,把握细粒度的音频文本语义相关性。此外,我们开发了一个置信度感知(CA)模块来估计每个潜在因素对的置信度,并自适应地聚合跨模态潜在因素,以实现局部语义对齐。实验表明,我们的THA有效地提高ATR性能,与DCR方法进一步有助于一致的性能增益。摘要:Most existing audio-text retrieval (ATR) approaches typically rely on a single-level interaction to associate audio and text, limiting their ability to align different modalities and leading to suboptimal matches. In this work, we present a novel ATR framework that leverages two-stream Transformers in conjunction with a Hierarchical Alignment (THA) module to identify multi-level correspondences of different Transformer blocks between audio and text. Moreover, current ATR methods mainly focus on learning a global-level representation, missing out on intricate details to capture audio occurrences that correspond to textual semantics. To bridge this gap, we introduce a Disentangled Cross-modal Representation (DCR) approach that disentangles high-dimensional features into compact latent factors to grasp fine-grained audio-text semantic correlations. Additionally, we develop a confidence-aware (CA) module to estimate the confidence of each latent factor pair and adaptively aggregate cross-modal latent factors to achieve local semantic alignment. Experiments show that our THA effectively boosts ATR performance, with the DCR approach further contributing to consistent performance gains.
【32】 Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
标题: 多模式语音Transformer解码器:多模式何时可以提高准确性?
作者:Yiwen Guan,Viet Anh Trinh,Vivek Voleti,Jacob Whitehill
链接:点击下载PDF文件
摘要:仅解码器的离散令牌语言模型最近在自动语音识别中取得了显著的成功。然而,对不同模式如何影响特定情景下的绩效的系统分析仍然有限。在本文中,我们研究了多种模态对合成和真实世界数据集识别准确性的影响。我们的实验表明:(1)整合更多的模态可以提高准确性;特别是,据我们所知,我们的论文是第一个展示结合音频,图像上下文和嘴唇信息的好处的论文;(2)图像作为语音识别的补充模态在中等噪声水平下提供最大的好处,此外,与固有同步的模态(如嘴唇运动)相比,它们表现出不同的趋势;(3)当最相关的视觉信息被过滤作为预处理步骤时,性能在合成和真实世界数据集上都有所提高。摘要:Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In this paper, we investigate the effects of multiple modalities on recognition accuracy on both synthetic and real-world datasets. Our experiments suggest that: (1) Integrating more modalities can increase accuracy; in particular, our paper is, to our best knowledge, the first to show the benefit of combining audio, image context, and lip information; (2) Images as a supplementary modality for speech recognition provide the greatest benefit at moderate noise levels, moreover, they exhibit a different trend compared to inherently synchronized modalities like lip movements; (3) Performance improves on both synthetic and real-world datasets when the most relevant visual information is filtered as a preprocessing step.
【33】 Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
标题: Seed-Music:高质量和受控音乐生成的统一框架
作者:Ye Bai,Haonan Chen,Jitong Chen,Zhuo Chen,Yi Deng,Xiaohong Dong,Lamtharn Hantrakul,Weituo Hao,Qingqing Huang,Zhongyi Huang,Dongya Jia,Feihu La,Duc Le,Bochen Li,Chumin Li,Hui Li,Xingxing Li,Shouda Liu,Wei-Tsung Lu,Yiqing Lu,Andrew Shaw,Janne Spijkervet,Yakun Sun,Bo Wang,Ju-Chiang Wang,Yuping Wang,Yuxuan Wang,Ling Xu,Yifeng Yang,Chao Yao,Shuo Zhang,Yang Zhang,Yilin Zhang,Hang Zhao,Ziyi Zhao,Dejian Zhong,Shicen Zhou,Pei Zou
备注:Seed-Music technical report, 20 pages, 5 figures
链接:点击下载PDF文件
摘要:我们推出了Seed-Music,这是一套音乐生成系统,能够生成具有细粒度风格控制的高质量音乐。我们的统一框架利用自回归语言建模和扩散方法来支持两个关键的音乐创作工作流程: textit{受控音乐生成}和 textit{后期制作编辑}。对于受控的音乐生成,我们的系统使声乐生成与性能控制从多模态输入,包括风格描述,音频参考,乐谱,和语音提示。对于后期制作编辑,它提供了直接在生成的音频中编辑歌词和声乐旋律的交互式工具。 我们鼓励读者在https: team.doubao.com seed-music上收听演示音频示例。摘要:We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control. Our unified framework leverages both auto-regressive language modeling and diffusion approaches to support two key music creation workflows: textit{controlled music generation} and textit{post-production editing}. For controlled music generation, our system enables vocal music generation with performance controls from multi-modal inputs, including style descriptions, audio references, musical scores, and voice prompts. For post-production editing, it offers interactive tools for editing lyrics and vocal melodies directly in the generated audio. We encourage readers to listen to demo audio examples at https: team.doubao.com seed-music .
【34】 AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
标题: AccentBox:迈向高保真Zero-Shot口音一代
作者:Jinzuomu Zhong,Korin Richmond,Zhiba Su,Siqi Sun
链接:点击下载PDF文件
摘要:虽然最近的零拍文本到语音(Zero-Shot Text-to-Speech,TTS)模型已经实现了高自然度和说话人相似度,但它们在口音保真度和控制方面存在不足。为了解决这个问题,我们提出了zero-shot口音生成,统一外国口音转换(FAC),重音TTS,和语音TTS,一个新的两阶段的管道。在第一阶段,我们实现了最先进的(SOTA)口音识别(AID)与0.56 f1分数看不见的扬声器。在第二阶段,我们的条件下的预训练的说话人不可知的口音嵌入提取的AID模型的TTS系统。所提出的系统实现了更高的口音保真度的固有 交叉口音生成,并使看不见的口音生成。摘要:While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation that unifies Foreign Accent Conversion (FAC), accented TTS, and ZS-TTS, with a novel two-stage pipeline. In the first stage, we achieve state-of-the-art (SOTA) on Accent Identification (AID) with 0.56 f1 score on unseen speakers. In the second stage, we condition ZS-TTS system on the pretrained speaker-agnostic accent embeddings extracted by the AID model. The proposed system achieves higher accent fidelity on inherent cross accent generation, and enables unseen accent generation.
【35】 An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems
标题: 交互式口语对话系统的高效自学习框架
作者:Hitesh Tulsiani,David M. Chan,Shalini Ghosh,Garima Lalwani,Prabhat Pandey,Ankish Bansal,Sri Garimella,Ariya Rastrow,Björn Hoffmeister
备注:Presented at ICML 2024
链接:点击下载PDF文件
摘要:对话系统,如语音助理,预计将与用户进行复杂的,不断发展的对话。不幸的是,在这样的应用中部署的传统自动语音识别(ASR)系统通常被训练为独立地识别每个回合,并且缺乏适应会话上下文或结合用户反馈的能力。在这项工作中,我们介绍了一个通用的框架ASR对话系统,可以超越学习单轮话语,并随着时间的推移学习如何适应显式监督和隐式用户反馈中存在的多轮对话。我们通过利用学生-教师学习和上下文感知对话处理的进步,并设计对比自我监督方法欧姆,一种新的在线硬否定挖掘方法。我们表明,与传统训练相比,利用我们的新框架,在现实世界的对话系统中,相对WER减少了近10%,在公共合成数据中减少了高达26%。摘要:Dialog systems, such as voice assistants, are expected to engage with users in complex, evolving conversations. Unfortunately, traditional automatic speech recognition (ASR) systems deployed in such applications are usually trained to recognize each turn independently and lack the ability to adapt to the conversational context or incorporate user feedback. In this work, we introduce a general framework for ASR in dialog systems that can go beyond learning from single-turn utterances and learn over time how to adapt to both explicit supervision and implicit user feedback present in multi-turn conversations. We accomplish that by leveraging advances in student-teacher learning and context-aware dialog processing, and designing contrastive self-supervision approaches with Ohm, a new online hard-negative mining approach. We show that leveraging our new framework compared to traditional training leads to relative WER reductions of close to 10% in real-world dialog systems, and up to 26% on public synthetic data.
【36】 Meta-Whisper: Speech-Based Meta-ICL for ASR on Low-Resource Languages
标题: Meta-Whisper:基于语音的Meta-ICL,用于低资源语言上的ASB
作者:Ming-Hao Hsu,Kuan Po Huang,Hung-yi Lee
链接:点击下载PDF文件
摘要:本文提出了元耳语,一种新的方法来提高自动语音识别(ASR)的低资源语言使用耳语模型。通过利用Meta上下文学习(Meta-ICL)和k最近邻(KNN)算法进行样本选择,Meta-Whisper增强了Whisper在不熟悉的语言中识别语音的能力,而无需进行广泛的微调。在ML-SUPERB数据集上的实验表明,与原始Whisper模型相比,Meta-Whisper显著降低了低资源语言的字符错误率(CER)。这种方法为开发适应性更强的多语言ASR系统提供了一种很有前途的解决方案,特别是对于资源有限的语言。摘要:This paper presents Meta-Whisper, a novel approach to improve automatic speech recognition (ASR) for low-resource languages using the Whisper model. By leveraging Meta In-Context Learning (Meta-ICL) and a k-Nearest Neighbors (KNN) algorithm for sample selection, Meta-Whisper enhances Whisper's ability to recognize speech in unfamiliar languages without extensive fine-tuning. Experiments on the ML-SUPERB dataset show that Meta-Whisper significantly reduces the Character Error Rate (CER) for low-resource languages compared to the original Whisper model. This method offers a promising solution for developing more adaptable multilingual ASR systems, particularly for languages with limited resources.
【37】 Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
标题: 利用MAMBA联合频谱和空间学习进行多通道语音增强
作者:Wenze Ren,Haibin Wu,Yi-Cheng Lin,Xuanjun Chen,Rong Chao,Kuo-Hsuan Hung,You-Jin Li,Wen-Yuan Ting,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
摘要:在多通道语音增强中,有效地捕获不同麦克风之间的空间和频谱信息对于降噪至关重要。传统方法(例如CNN或LSTM)试图对全波段和子波段光谱和空间特征的时间动态进行建模。然而,这些方法在完全建模复杂的时间依赖性方面面临限制,特别是在动态声学环境中。为了克服这些挑战,我们修改了目前的先进模型McNet通过引入改进版本的Mamba,一个状态空间模型,并进一步提出MCMamba。MCMAamba已经完全重新设计,将全波段和窄带空间信息与子波段和全波段光谱特征集成在一起,为空间和光谱信息建模提供了更全面的方法。我们的实验结果表明,MCMamba显著提高了多通道语音增强中的空间和频谱特征建模,优于McNet,并在CHiME-3数据集上实现了最先进的性能。此外,我们发现,曼巴表现非常好的建模光谱信息。摘要:In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, these approaches face limitations in fully modeling complex temporal dependencies, especially in dynamic acoustic environments. To overcome these challenges, we modify the current advanced model McNet by introducing an improved version of Mamba, a state-space model, and further propose MCMamba. MCMamba has been completely reengineered to integrate full-band and narrow-band spatial information with sub-band and full-band spectral features, providing a more comprehensive approach to modeling spatial and spectral information. Our experimental results demonstrate that MCMamba significantly improves the modeling of spatial and spectral features in multichannel speech enhancement, outperforming McNet and achieving state-of-the-art performance on the CHiME-3 dataset. Additionally, we find that Mamba performs exceptionally well in modeling spectral information.
【38】 Ultra-Low Latency Speech Enhancement - A Comprehensive Study
标题: 超低延迟语音增强-综合研究
作者:Haibin Wu,Sebastian Braun
链接:点击下载PDF文件
摘要:语音增强模型应满足非常低的延迟要求,通常小于5毫秒的听力辅助设备。虽然已经提出了各种低延迟技术,但在使用DNN的受控设置中比较这些方法仍然是空白。以前的论文在任务、训练数据、脚本和评估设置方面存在差异,这使得公平的比较变得不可能。此外,所有方法都是在小型模拟数据集上进行测试的,因此很难公平地评估它们在真实世界条件下的性能,这可能会影响科学发现的可靠性。为了解决这些问题,我们使用大规模数据的一致训练来全面研究各种低延迟技术,并使用真实世界数据的更相关指标进行评估。具体来说,我们探讨了非对称窗口,可学习的窗口,自适应时域滤波器组和未来帧预测技术的有效性。此外,我们还研究了增加模型大小是否可以补偿减小的窗口大小,以及低延迟环境中的新型Mamba架构。摘要:Speech enhancement models should meet very low latency requirements typically smaller than 5 ms for hearing assistive devices. While various low-latency techniques have been proposed, comparing these methods in a controlled setup using DNNs remains blank. Previous papers have variations in task, training data, scripts, and evaluation settings, which make fair comparison impossible. Moreover, all methods are tested on small, simulated datasets, making it difficult to fairly assess their performance in real-world conditions, which could impact the reliability of scientific findings. To address these issues, we comprehensively investigate various low-latency techniques using consistent training on large-scale data and evaluate with more relevant metrics on real-world data. Specifically, we explore the effectiveness of asymmetric windows, learnable windows, adaptive time domain filterbanks, and the future-frame prediction technique. Additionally, we examine whether increasing the model size can compensate for the reduced window size, as well as the novel Mamba architecture in low-latency environments.
【39】 oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
标题: oboVox远场说话人识别:一种采用预训练模型的新型数据增强方法
作者:Muhammad Sudipto Siam Dip,Md Anik Hasan,Sapnil Sarker Bipro,Md Abdur Raiyan,Mohammod Abdul Motin
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:在这项研究中,我们解决了说话人识别的挑战,使用一种新的数据增强技术,增加噪音的注册文件。这种技术有效地对齐了测试和注册文件的来源,提高了可比性。使用了各种预训练模型,其中resnet模型实现了最高的DCF 0.84和EER 13.44。增强技术显着改善这些结果为0.75 DCF和12.79 EER的resnet模型。比较分析表明,resnet的优势,如ECPA,梅尔光谱图,Payonnet,和Titanet大模型。结果,以及不同的增强方案,有助于本文的RoboVox远场说话人识别的成功摘要:In this study, we address the challenge of speaker recognition using a novel data augmentation technique of adding noise to enrollment files. This technique efficiently aligns the sources of test and enrollment files, enhancing comparability. Various pre-trained models were employed, with the resnet model achieving the highest DCF of 0.84 and an EER of 13.44. The augmentation technique notably improved these results to 0.75 DCF and 12.79 EER for the resnet model. Comparative analysis revealed the superiority of resnet over models such as ECPA, Mel-spectrogram, Payonnet, and Titanet large. Results, along with different augmentation schemes, contribute to the success of RoboVox far-field speaker recognition in this paper
【40】 Speech as a Biomarker for Disease Detection
标题: 言语作为疾病检测的生物标志物
作者:Catarina Botelho,Alberto Abad,Tanja Schultz,Isabel Trancoso
链接:点击下载PDF文件
摘要:语音是一种丰富的生物标志物,它编码了有关说话者健康的大量信息,因此已被提出用于检测许多疾病,取得了可喜的成果。然而,关于为自动检测这些疾病而训练的模型实际上在学习什么以及它们预测的基础仍然存在问题,这可能会对患者的生活产生重大影响。这项工作倡导一种可解释的健康模型,适合检测几种疾病,其动机是观察到影响言语的疾病往往对言语信号产生重叠影响。提出了一个框架,首先定义“参考语音”,然后利用该定义进行疾病检测。参考语音通过参考间隔来表征,即,来自参考人群的具有临床意义的声学和语言特征的典型值。这种在语音领域作为生物标志物的新方法受到临床实验室科学中使用参考区间的启发。新的扬声器从这个参考模型的偏差进行量化,并作为输入检测阿尔茨海默氏症和帕金森氏症。探索的分类策略是基于神经加法模型,一种玻璃盒神经网络,它可以解释。建议的参考语音表征和疾病检测框架的目的是支持医学界提供临床上有意义的解释,可以作为一个有价值的第二意见。摘要:Speech is a rich biomarker that encodes substantial information about the health of a speaker, and thus it has been proposed for the detection of numerous diseases, achieving promising results. However, questions remain about what the models trained for the automatic detection of these diseases are actually learning and the basis for their predictions, which can significantly impact patients' lives. This work advocates for an interpretable health model, suitable for detecting several diseases, motivated by the observation that speech-affecting disorders often have overlapping effects on speech signals. A framework is presented that first defines "reference speech" and then leverages this definition for disease detection. Reference speech is characterized through reference intervals, i.e., the typical values of clinically meaningful acoustic and linguistic features derived from a reference population. This novel approach in the field of speech as a biomarker is inspired by the use of reference intervals in clinical laboratory science. Deviations of new speakers from this reference model are quantified and used as input to detect Alzheimer's and Parkinson's disease. The classification strategy explored is based on Neural Additive Models, a type of glass-box neural network, which enables interpretability. The proposed framework for reference speech characterization and disease detection is designed to support the medical community by providing clinically meaningful explanations that can serve as a valuable second opinion.
【41】 RF-GML: Reference-Free Generative Machine Listener
标题: RF-GML:无参考生成机器收件箱
作者:Arijit Biswas,Guanxin Jiang
备注:Pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本文介绍了一种新的无参考(RF)的音频质量度量称为RF生成机器的音频质量(RF GML),旨在评估编码的单声道,立体声和双耳音频在48 kHz的采样率。RF-GML利用了来自最先进的全参考(FR)生成机器学习(GML)的迁移学习,只需最小的架构修改。术语“生成”是指模型生成任意数量的模拟听力分数的能力。与现有的RF模型不同,RF-GML可以准确预测各种内容类型和编解码器的主观质量分数。广泛的评估表明,它的优势,在评级未编码的音频和区分不同层次的编码文物。RF-GML的性能和多功能性使其成为各种应用中编码音频质量评估和监控的宝贵工具,所有这些都不需要参考信号。摘要:This paper introduces a novel reference-free (RF) audio quality metric called the RF-Generative Machine Listener (RF-GML), designed to evaluate coded mono, stereo, and binaural audio at a 48 kHz sample rate. RF-GML leverages transfer learning from a state-of-the-art full-reference (FR) Generative Machine Listener (GML) with minimal architectural modifications. The term "generative" refers to the model's ability to generate an arbitrary number of simulated listening scores. Unlike existing RF models, RF-GML accurately predicts subjective quality scores across diverse content types and codecs. Extensive evaluations demonstrate its superiority in rating unencoded audio and distinguishing different levels of coding artifacts. RF-GML's performance and versatility make it a valuable tool for coded audio quality assessment and monitoring in various applications, all without the need for a reference signal.
【42】 Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
标题: CLAR-DPO:通过直接偏好优化的可控情感语音合成
作者:Xiaoxue Gao,Chen Zhang,Yiming Chen,Huayun Zhang,Nancy F. Chen
备注:5 pages
链接:点击下载PDF文件
摘要:当前的情感文本到语音(TTS)模型主要进行监督训练,以学习从文本和期望的情感到其情感语音的转换,专注于每个文本-语音对的单个情感。这些模型只能学习正确的情感输出,而不能完全理解其他情感特征,这限制了它们捕捉不同情感之间细微差别的能力。我们提出了一个可控的DPO方法,它采用直接偏好优化,以区分微妙的情感之间的细微差别,通过优化对首选的情感,而不是不太喜欢的情感。我们建议利用情感感知LLM-TTS神经架构来利用LLM的上下文学习和推理跟随能力,而不是依赖于现有情感TTS模型中使用的传统神经架构。综合实验证实,我们提出的方法优于现有的基线。摘要:Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs' in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines.
【43】 Room impulse response prototyping using receiver distance estimations for high quality room equalisation algorithms
标题: 使用接收器距离估计的房间脉冲响应原型用于高质量房间均衡算法
作者:James Brooks-Park,Martin Bo Møller,Jan Østergaard,Søren Bech,Steven van de Par
链接:点击下载PDF文件
摘要:房间均衡旨在提高混响环境中扬声器再现的质量,补偿由不完美的房间反射和频率相关扬声器方向性引起的着色。房间均衡领域中的一种常见技术是反转原型房间脉冲响应(RIR)。原型响应由分布在收听区域周围的几个响应组成,而不是在收听位置反转单个RIR。本文提出了一种脉冲响应原型的方法,使用估计的接收机位置,形成一个加权平均原型响应。描述了一种接收机距离估计的方法,支持原型RIR的实现。建议的原型制作方法相比,其他方法通过测量后均衡光谱偏差在几个位置在一个模拟的房间。摘要:Room equalisation aims to increase the quality of loudspeaker reproduction in reverberant environments, compensating for colouration caused by imperfect room reflections and frequency dependant loudspeaker directivity. A common technique in the field of room equalisation, is to invert a prototype Room Impulse Response (RIR). Rather than inverting a single RIR at the listening position, a prototype response is composed of several responses distributed around the listening area. This paper proposes a method of impulse response prototyping, using estimated receiver positions, to form a weighted average prototype response. A method of receiver distance estimation is described, supporting the implementation of the prototype RIR. The proposed prototyping method is compared to other methods by measuring their post equalisation spectral deviation at several positions in a simulated room.
【44】 StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
标题: StyleTTS-ZZ:具有蒸馏时变风格扩散的高效高质量Zero-Shot文本到语音合成
作者:Yinghao Aaron Li,Xilin Jiang,Cong Han,Nima Mesgarani
链接:点击下载PDF文件
摘要:大规模文本到语音(TTS)模型的快速发展,导致了显着的进步,建模不同的说话人韵律和声音。然而,这些模型通常面临推理速度慢、依赖复杂的预训练神经编解码器表示以及难以实现自然度和与参考说话人的高相似度等问题。为了解决这些挑战,这项工作介绍了StyleTTS-ZS,一个有效的zero-shot TTS模型,利用蒸馏时变风格扩散,以捕捉不同的扬声器身份和韵律。我们提出了一种新的方法,代表人类语音使用输入文本和固定长度的时变离散风格代码来捕捉不同的韵律变化,训练对抗多模态判别。然后建立一个扩散模型,对这种时变风格的代码进行采样,以实现有效的潜在扩散。在风格扩散过程中,StyleTTS-ZS使用无分类器的指导,实现了与参考说话人的高相似度。此外,为了加快采样速度,风格扩散模型仅使用10 k个样本进行感知损失提取,保持语音质量和相似性,同时将推理速度降低90%。该模型在自然度和相似度方面均优于现有的大规模zero-shot TTS模型,采样速度提高了10-20倍,是大规模zero-shot TTS系统的理想选择。音频演示,代码和模型可在https: styletts-zs.github.io 。摘要:The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex pre-trained neural codec representations, and difficulties in achieving naturalness and high similarity to reference speakers. To address these challenges, this work introduces StyleTTS-ZS, an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies. We propose a novel approach that represents human speech using input text and fixed-length time-varying discrete style codes to capture diverse prosodic variations, trained adversarially with multi-modal discriminators. A diffusion model is then built to sample this time-varying style code for efficient latent diffusion. Using classifier-free guidance, StyleTTS-ZS achieves high similarity to the reference speaker in the style diffusion process. Furthermore, to expedite sampling, the style diffusion model is distilled with perceptual loss using only 10k samples, maintaining speech quality and similarity while reducing inference speed by 90%. Our model surpasses previous state-of-the-art large-scale zero-shot TTS models in both naturalness and similarity, offering a 10-20 faster sampling speed, making it an attractive alternative for efficient large-scale zero-shot TTS systems. The audio demo, code and models are available at https: styletts-zs.github.io .
【45】 TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
标题: TBDM-Net:用于语音情感识别的具有性别信息的双向密集网络
作者:Vlad Striletchi,Cosmin Striletchi,Adriana Stan
备注:In Proceedings of 2024 IEEE International Workshop on Machine Learning for Signal Processing, London, UK
链接:点击下载PDF文件
摘要:本文提出了一种新的基于深度神经网络的语音情感识别(SER)架构。该架构利用了多层双向扩张卷积之间的密集互连。线性内核动态融合这些层的输出,以产生最终的情感类预测。这种创新的架构被表示为TBDM-Net:时间感知双向密集多尺度网络。我们对TBDM-Net进行了全面的性能评估,包括消融研究,在六个广泛认可的SER数据集上进行单峰语音情感识别。此外,我们探索的影响,性别知情的情绪预测附加黄金或预测的性别标签的架构的输入或预测。TBDM-Net的实施情况可在以下网址查阅:https: github.com adrianastan tbdm-net摘要:This paper presents a novel deep neural network-based architecture tailored for Speech Emotion Recognition (SER). The architecture capitalises on dense interconnections among multiple layers of bidirectional dilated convolutions. A linear kernel dynamically fuses the outputs of these layers to yield the final emotion class prediction. This innovative architecture is denoted as TBDM-Net: Temporally-Aware Bi-directional Dense Multi-Scale Network. We conduct a comprehensive performance evaluation of TBDM-Net, including an ablation study, across six widely-acknowledged SER datasets for unimodal speech emotion recognition. Additionally, we explore the influence of gender-informed emotion prediction by appending either golden or predicted gender labels to the architecture's inputs or predictions. The implementation of TBDM-Net is accessible at: https: github.com adrianastan tbdm-net
【46】 DNN-based ensemble singing voice synthesis with interactions between singers
标题: 基于DNN的合奏歌唱声音合成,具有歌手之间的互动
作者:Hiroaki Hyodo,Shinnosuke Takamichi,Tomohiko Nakamura,Junya Koguchi,Hiroshi Saruwatari
链接:点击下载PDF文件
摘要:我们提出了一个歌唱声音合成(SVS)的方法更统一的合奏歌唱声音建模歌手之间的相互作用。大多数现有的SVS方法旨在合成独唱声音,并且不考虑歌手之间的交互,即,调整自己的声音来适应别人的声音由于从独唱声音到合奏声音的产生忽略了相互作用,它会降低声乐合奏的统一性。因此,我们提出了一个SVS再现的相互作用。它是基于一个架构,使用多个语音部分的乐谱,和损失函数,模拟相互作用的影响,声学特征。实验结果表明,我们的方法提高了声乐合奏的统一性。摘要:We propose a singing voice synthesis (SVS) method for a more unified ensemble singing voice by modeling interactions between singers. Most existing SVS methods aim to synthesize a solo voice, and do not consider interactions between singers, i.e., adjusting one's own voice to the others' voices. Since the production of ensemble voices from solo singing voices ignores the interactions, it can degrade the unity of the vocal ensemble. Therefore, we propose a SVS that reproduces the interactions. It is based on an architecture that uses musical scores of multiple voice parts, and loss functions that simulate the interactions' effect to acoustic features. Experimental results show that our methods improve the unity of the vocal ensemble.
【47】 A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
标题: 使用大型语言模型的Zero-Shot非侵入性语音评估研究
作者:Ryandhimas E. Zezario,Sabato M. Siniscalchi,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
摘要:本文研究了两种利用大型语言模型的zero-shot非侵入式语音评估策略。首先,我们探索GPT-4 o的音频分析功能。其次,我们提出了GPT-Whisper,它使用Whisper作为音频到文本的模块,并通过有针对性的提示工程评估文本的自然度。我们评估评估指标预测的GPT-4 o和GPT-Whisper检查其与基于人类的质量和可懂度评估,以及自动语音识别的字符错误率(CER)的相关性。实验结果表明,单独使用GPT-4 o对音频分析无效;而GPT-Whisper具有更高的预测能力,与语音质量和可懂度具有中等相关性,与CER具有较高的相关性。与有监督的非侵入式神经语音评估模型(即MOS-SSL和MTI-Net)相比,GPT-Whisper与Whisper的CER产生了显着更高的Spearman秩相关性。这些发现验证了GPT-Whisper作为准确的zero-shot语音评估的可靠方法,而不需要额外的训练数据(语音数据和相应的评估分数)。摘要:This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the naturalness of text via targeted prompt engineering. We evaluate assessment metrics predicted by GPT-4o and GPT-Whisper examining their correlations with human-based quality and intelligibility assessments, and character error rate (CER) of automatic speech recognition. Experimental results show that GPT-4o alone is not effective for audio analysis; whereas, GPT-Whisper demonstrates higher prediction, showing moderate correlation with speech quality and intelligibility, and high correlation with CER. Compared to supervised non-intrusive neural speech assessment models, namely MOS-SSL and MTI-Net, GPT-Whisper yields a notably higher Spearman's rank correlation with the CER of Whisper. These findings validate GPT-Whisper as a reliable method for accurate zero-shot speech assessment without requiring additional training data (speech data and corresponding assessment scores).
【48】 Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
标题: 自我监督的多模式语音表示用于评估精神分裂症症状
作者:Gowtham Premananth,Carol Espy-Wilson
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:多模式精神分裂症评估系统在过去几年中获得了牵引力。这项工作介绍了精神分裂症评估系统,以区分精神分裂症的突出症状类别,并预测整体精神分裂症的严重程度评分。我们开发了一个基于矢量量化变分自动编码器(VQ-VAE)的多模态表示学习(MRL)模型,从声道变量(TV)和面部动作单元(FAU)生成任务无关的语音表示。然后,这些表示用于基于多任务学习(MTL)的下游预测模型,以获得类别标签和总体严重性评分。所提出的框架在所有评估指标(加权F1得分,AUC-ROC得分和加权准确度)上都优于以前的多类分类任务。此外,它还估计了精神分裂症的严重程度评分,这是一项早期方法没有解决的任务。摘要:Multimodal schizophrenia assessment systems have gained traction over the last few years. This work introduces a schizophrenia assessment system to discern between prominent symptom classes of schizophrenia and predict an overall schizophrenia severity score. We develop a Vector Quantized Variational Auto-Encoder (VQ-VAE) based Multimodal Representation Learning (MRL) model to produce task-agnostic speech representations from vocal Tract Variables (TVs) and Facial Action Units (FAUs). These representations are then used in a Multi-Task Learning (MTL) based downstream prediction model to obtain class labels and an overall severity score. The proposed framework outperforms the previous works on the multi-class classification task across all evaluation metrics (Weighted F1 score, AUC-ROC score, and Weighted Accuracy). Additionally, it estimates the schizophrenia severity score, a task not addressed by earlier approaches.
【49】 Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
标题: 提取和扩散:用于改进基于扩散的语音和人声增强的潜在集成
作者:Yudong Yang,Zhan Liu,Wenyi Yu,Guangzhi Sun,Qiuqiang Kong,Chao Zhang
链接:点击下载PDF文件
摘要:基于扩散的生成模型最近在语音和声音增强方面取得了显着的成果,因为它们能够模拟复杂的语音数据分布。虽然这些模型很好地推广到看不见的声学环境,但它们可能无法实现与专门训练以增强特定声学条件的区分模型相同的保真度水平。在本文中,我们提出了Ex-Diff,一种新的基于分数的扩散模型,它集成了由判别模型产生的潜在表示,以改善语音和声音增强,它结合了生成模型和判别模型的优势。在广泛使用的MUSDB数据集上的实验结果显示,与语音和声音增强任务的基线扩散模型相比,SI-SDR和SI-SIR的相对改善分别为3.7%和10.0%。此外,案例研究提供了进一步说明和分析在这种情况下生成和歧视性模型的互补性。摘要:Diffusion-based generative models have recently achieved remarkable results in speech and vocal enhancement due to their ability to model complex speech data distributions. While these models generalize well to unseen acoustic environments, they may not achieve the same level of fidelity as the discriminative models specifically trained to enhance particular acoustic conditions. In this paper, we propose Ex-Diff, a novel score-based diffusion model that integrates the latent representations produced by a discriminative model to improve speech and vocal enhancement, which combines the strengths of both generative and discriminative models. Experimental results on the widely used MUSDB dataset show relative improvements of 3.7% in SI-SDR and 10.0% in SI-SIR compared to the baseline diffusion model for speech and vocal enhancement tasks, respectively. Additionally, case studies are provided to further illustrate and analyze the complementary nature of generative and discriminative models in this context.
【50】 Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
标题: 口吃解决器:端到端多语言流利检测
作者:Xuanru Zhou,Cheol Jun Cho,Ayati Sharma,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Boon Lead Tee,Maria Luisa Gorno Tempini,Jiachen Lian,Gopala Anumanchipalli
备注:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:当前的事实上的不流利建模方法利用模板匹配算法,其不能推广到跨语言的域外真实世界的不流利,并且不能随着训练数据量的增加而扩展。为了解决这些问题,我们提出了Stutter-Solver:一个端到端的框架,它可以通过精确的类型和时间转录来检测不流畅,灵感来自YOLO对象检测算法。Stutter-Solver可以处理不流利的情况,是一个自然的多语言不流利检测器。为了利用可扩展性和提高性能,我们还引入了三个新的不流利语料库:VCTK-Pro,VCTK-Art和AISHELL 3-Pro,通过发音编码器和基于TTS的方法模拟自然的口语不流利,包括重复,阻塞,缺失,替换和延长。我们的方法在所有可用的不流利语料库上实现了最先进的性能。代码和数据集在https: github.com eureka235 Stutter-Solver上开源摘要:Current de-facto dysfluency modeling methods utilize template matching algorithms which are not generalizable to out-of-domain real-world dysfluencies across languages, and are not scalable with increasing amounts of training data. To handle these problems, we propose Stutter-Solver: an end-to-end framework that detects dysfluency with accurate type and time transcription, inspired by the YOLO object detection algorithm. Stutter-Solver can handle co-dysfluencies and is a natural multi-lingual dysfluency detector. To leverage scalability and boost performance, we also introduce three novel dysfluency corpora: VCTK-Pro, VCTK-Art, and AISHELL3-Pro, simulating natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation through articulatory-encodec and TTS-based methods. Our approach achieves state-of-the-art performance on all available dysfluency corpora. Code and datasets are open-sourced at https: github.com eureka235 Stutter-Solver
【51】 Effective Pre-Training of Audio Transformers for Sound Event Detection
标题: 音频Transformer的有效预训练以进行声音事件检测
作者:Florian Schmid,Tobias Morocutti,Francesco Foscarin,Jan Schlüter,Paul Primus,Gerhard Widmer
备注:Submitted to ICASSP'25. Source code available: this https URL
链接:点击下载PDF文件
摘要:我们提出了一种用于音频频谱图Transformers的预训练管道,用于帧级声音事件检测任务。在常见的预训练步骤之上,我们在AudioSet帧级注释上添加了精心设计的训练例程。这包括一个平衡的采样器,积极的数据增强,和合奏知识蒸馏。对于五个Transformers,我们在AudioSet帧级预测和帧级声音事件检测下游任务上获得了比以前可用的检查点更大的性能改进,证实了我们的管道的有效性。我们发布了由此产生的检查点,研究人员可以直接微调,以构建用于声音事件检测任务的高性能模型。摘要:We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-level annotations. This includes a balanced sampler, aggressive data augmentation, and ensemble knowledge distillation. For five transformers, we obtain a substantial performance improvement over previously available checkpoints both on AudioSet frame-level predictions and on frame-level sound event detection downstream tasks, confirming our pipeline's effectiveness. We publish the resulting checkpoints that researchers can directly fine-tune to build high-performance models for sound event detection tasks.
【52】 Target Speaker ASR with Whisper
标题: 目标说话者ASB与Whisper
作者:Alexander Polok,Dominik Klement,Matthew Wiesner,Sanjeev Khudanpur,Jan Černocký,Lukáš Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,使使用大型,单扬声器ASR模型,如耳语,目标扬声器ASR。该方法的关键见解是,通过学习以帧级日记化输出为条件来建模说话者之间的相对差异比学习所有说话者嵌入的空间要容易得多。我们发现,在第一个Transformer块之前添加每个日志输出类型的单个偏置项可以将单个扬声器ASR模型转换为目标扬声器ASR模型。我们的目标说话人ASR模型可以用于扬声器归因的ASR生产,在序列中,每个假设的扬声器在日记输出的成绩单。这个简化的模型,扬声器归因ASR使用一个麦克风,仅优于级联的语音分离和日记的11%的绝对ORC-WER的NOTSOFAR-1数据集。摘要:We propose a novel approach to enable the use of large, single speaker ASR models, such as Whisper, for target speaker ASR. The key insight of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs, than to learn the space of all speaker embeddings. We find that adding even a single bias term per diarization output type before the first transformer block can transform single speaker ASR models, into target speaker ASR models. Our target-speaker ASR model can be used for speaker attributed ASR by producing, in sequence, a transcript for each hypothesized speaker in a diarization output. This simplified model for speaker attributed ASR using only a single microphone outperforms cascades of speech separation and diarization by 11% absolute ORC-WER on the NOTSOFAR-1 dataset.
【53】 Leveraging Self-Supervised Learning for Speaker Diarization
标题: 利用自我监督学习进行发言者日记化
作者:Jiangyu Han,Federico Landini,Johan Rohdin,Anna Silnova,Mireia Diez,Lukas Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在过去的几年里,端到端神经日志化已经有了很大的发展,但数据稀缺仍然是进一步改进的主要障碍。WavLM等自监督学习方法在一些下游任务上表现出良好的性能,但它们在说话人日记化方面的应用受到一定限制。在这项工作中,我们探索使用WavLM来缓解神经日志化训练的数据稀缺问题。我们使用与Pyannote相同的流水线,并使用WavLM和Conformer改进了局部端到端神经日志化。在远场AMI,AISHELL-4和AliMeeting数据集上的实验表明,我们的方法大大优于Pyannote基线,并达到了与AMI和AISHELL-4上最先进的结果相当的性能。此外,通过分析不同数据量场景下的系统性能,我们发现WavLM表示比滤波器组特征更能抵抗数据稀缺,从而实现更少的数据饥饿训练策略。此外,我们发现,通常用于训练端到端日志模型的模拟数据在我们的实验中使用WavLM时没有帮助。此外,我们还在最近的CHiME 8 NOTSOFAR-1任务上评估了我们的模型,在该任务中,它实现了比Pyannote基线更好的性能。我们的源代码可在https: github.com BUTSpeechFIT DiariZen上公开获取。摘要:End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several downstream tasks, but their application on speaker diarization is somehow limited. In this work, we explore using WavLM to alleviate the problem of data scarcity for neural diarization training. We use the same pipeline as Pyannote and improve the local end-to-end neural diarization with WavLM and Conformer. Experiments on far-field AMI, AISHELL-4, and AliMeeting datasets show that our method substantially outperforms the Pyannote baseline and achieves performance comparable to the state-of-the-art results on AMI and AISHELL-4. In addition, by analyzing the system performance under different data quantity scenarios, we show that WavLM representations are much more robust against data scarcity than filterbank features, enabling less data hungry training strategies. Furthermore, we found that simulated data, usually used to train endto-end diarization models, does not help when using WavLM in our experiments. Additionally, we also evaluate our model on the recent CHiME8 NOTSOFAR-1 task where it achieves better performance than the Pyannote baseline. Our source code is publicly available at https: github.com BUTSpeechFIT DiariZen.
【54】 Language-Queried Target Sound Extraction Without Parallel Training Data
标题: 无需并行训练数据的数据查询目标声音提取
作者:Hao Ma,Zhiyuan Peng,Xu Li,Yukai Li,Mingjie Shao,Qiuqiang Kong,Ju Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:混合查询目标声音提取(TSE)的目的是从混合的语言查询的基础上提取特定的声音。传统的全监督训练方案需要大量注释的并行音频文本数据,这是劳动密集型的。我们引入了一种无语言训练方案,通过利用对比语言音频预训练模型(CLAP)的多模态表示对齐特性,仅需要未标记的音频片段进行TSE模型训练。在无语言训练阶段,使用预训练的CLAP音频编码器对目标音频进行编码,以形成TSE模型的条件嵌入,而在推理期间,用户语言查询由CLAP文本编码器进行编码。由于训练和推理查询之间的模态差距以及在训练期间直接暴露于目标音频的信息泄漏,这种直接的方法面临挑战。为了解决这个问题,我们提出了一个检索增强策略。具体来说,我们使用由大型语言模型(LLM)生成的音频字幕创建嵌入缓存。在训练期间,目标音频嵌入从该缓存中检索文本嵌入以用作条件嵌入,确保训练和推理之间的一致模态并消除信息泄漏。大量的实验结果表明,我们的检索增强的方法实现了一致的和显着的性能改善现有的国家的最先进的更好的泛化能力。摘要:Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are labor-intensive. We introduce a language-free training scheme, requiring only unlabelled audio clips for TSE model training by utilizing the multi-modal representation alignment nature of the contrastive language-audio pre-trained model (CLAP). In a vanilla language-free training stage, target audio is encoded using the pre-trained CLAP audio encoder to form a condition embedding for the TSE model, while during inference, user language queries are encoded by CLAP text encoder. This straightforward approach faces challenges due to the modality gap between training and inference queries and information leakage from direct exposure to target audio during training. To address this, we propose a retrieval-augmented strategy. Specifically, we create an embedding cache using audio captions generated by a large language model (LLM). During training, target audio embeddings retrieve text embeddings from this cache to use as condition embeddings, ensuring consistent modalities between training and inference and eliminating information leakage. Extensive experiment results show that our retrieval-augmented approach achieves consistent and notable performance improvements over existing state-of-the-art with better generalizability.
【55】 Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
标题: 使用带伪标签的最佳传输进行说话人验证的通道自适应
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Lei Li,Xugang Lu
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在说话人确认系统中,当训练数据和真实测试语音的统计分布不匹配时,域间隙会降低系统的性能。信道变化是造成这种差距的主要因素,它比其他问题(例如,噪声)。虽然各种领域自适应算法可以用来处理这种领域间隙问题,但大多数算法都不能采用判别式学习进行领域对齐时复杂的分布结构。在本文中,我们提出了一种新的无监督域自适应方法,即,联合部分最优传输与伪标签(JPOT-PL),以缓解信道失配问题。利用分布对齐中最优传输的几何感知距离度量,我们进一步设计了一种基于伪标签的判别学习,其中伪标签可以被视为来自最优耦合的新型软扬声器标签。利用JPOT-PL,以VoxCeleb为基础语料,对SV通道自适应任务进行了实验。实验表明,我们的方法减少EER超过10%,与几个国家的最先进的信道自适应算法相比。摘要:Domain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation, a primary factor causing this gap, is less addressed than other issues (e.g., noise). Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the channel mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of soft speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel adaptation task with VoxCeleb as the basis corpus. Experiments show our method reduces EER by over 10% compared with several state-of-the-art channel adaptation algorithms.
【56】 Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
标题: 用于增强说话人验证的集成多层知识提炼
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Xugang Lu,Lei Li
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:知识蒸馏(KD)广泛用于音频任务,如说话人验证(SV),通过将知识从训练有素的大型模型(教师)转移到更小,更紧凑的模型(学生),以提高效率和可移植性。SV的现有KD方法通常反映了图像处理中使用的方法,专注于近似预测概率和隐藏表示。然而,这些方法未能考虑到语音音频的多级时间特性。在本文中,我们提出了一种新的KD方法,集成的多级知识蒸馏(IML-KD),将语音的各种时间尺度特征的知识从教师模型转移到学生模型。在IML-KD中,来自教师模型的时间上下文信息被集成到来自具有各种持续时间的语音片段的新的集成的基于一致性的输入敏感表示中,并且学生模型被训练以推断这些表示,并对输出进行多级对齐。我们在VoxCeleb 1数据集上进行SV实验来评估所提出的方法。实验结果表明,IML-KD显著提高了KD性能,将等错误率(EER)降低了5%。摘要:Knowledge distillation (KD) is widely used in audio tasks, such as speaker verification (SV), by transferring knowledge from a well-trained large model (the teacher) to a smaller, more compact model (the student) for efficiency and portability. Existing KD methods for SV often mirror those used in image processing, focusing on approximating predicted probabilities and hidden representations. However, these methods fail to account for the multi-level temporal properties of speech audio. In this paper, we propose a novel KD method, i.e., Integrated Multi-level Knowledge Distillation (IML-KD), to transfer knowledge of various temporal-scale features of speech from a teacher model to a student model. In the IML-KD, temporal context information from the teacher model is integrated into novel Integrated Gradient-based input-sensitive representations from speech segments with various durations, and the student model is trained to infer these representations with multi-level alignment for the output. We conduct SV experiments on the VoxCeleb1 dataset to evaluate the proposed method. Experimental results demonstrate that IML-KD significantly enhances KD performance, reducing the Equal Error Rate (EER) by 5%.
【57】 Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
标题: 文本提示还不够:用于目标风格音频生成的声音事件增强提示适配器
作者:Chenxu Xiong,Ruibo Fu,Shuchen Shi,Zhengqi Wen,Jianhua Tao,Tao Wang,Chenxing Li,Chunyu Qiang,Yuankun Xie,Xin Qi,Guanjun Li,Zizheng Yang
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:目前主流的音频生成方法主要依赖于简单的文本提示,通常无法捕捉多风格音频生成所需的细微差别。为了解决这一限制,提出了声音事件增强提示适配器。与传统的静态全局风格转移不同,该方法通过文本和参考音频之间的交叉注意提取风格嵌入,以实现自适应风格控制。然后利用自适应层规范化来增强模型表达多种风格的能力。此外,声音事件参考风格传输数据集(SERST)被引入用于所提出的目标风格音频生成任务,从而使用文本和音频参考来实现双提示音频生成。实验结果表明,该模型的鲁棒性,实现了最先进的Fr 'echet距离为26.94和KL发散度为1.82,超过了Tango,AudioLDM和AudioGen。此外,所生成的音频显示出与其对应的音频参考的高相似性。演示、代码和数据集都是公开的。摘要:Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To address this limitation, the Sound Event Enhanced Prompt Adapter is proposed. Unlike traditional static global style transfer, this method extracts style embedding through cross-attention between text and reference audio for adaptive style control. Adaptive layer normalization is then utilized to enhance the model's capacity to express multiple styles. Additionally, the Sound Event Reference Style Transfer Dataset (SERST) is introduced for the proposed target style audio generation task, enabling dual-prompt audio generation using both text and audio references. Experimental results demonstrate the robustness of the model, achieving state-of-the-art Fr 'echet Distance of 26.94 and KL Divergence of 1.82, surpassing Tango, AudioLDM, and AudioGen. Furthermore, the generated audio shows high similarity to its corresponding audio reference. The demo, code, and dataset are publicly available.
【58】 E1 TTS: Simple and Fast Non-Autoregressive TTS
标题: E1 TTC:简单快速的非自回归TTC
作者:Zhijun Liu,Shuai Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
链接:点击下载PDF文件
摘要:介绍了一种基于去噪扩散预训练和分布匹配蒸馏的高效非自回归zero-shot文语转换系统--Easy One-Step Text-to-Speech(E1 TTS)。E1 TTS的训练很简单;它不需要文本和音频对之间的显式单调对齐。E1 TTS的推理是有效的,只需要一个神经网络评估每个话语。尽管其采样效率,E1 TTS实现的自然度和说话人相似度可与各种强基线模型相媲美。音频样本可在http: e1tts.github.io 上获得。摘要:This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is straightforward; it does not require explicit monotonic alignment between the text and audio pairs. The inference of E1 TTS is efficient, requiring only one neural network evaluation for each utterance. Despite its sampling efficiency, E1 TTS achieves naturalness and speaker similarity comparable to various strong baseline models. Audio samples are available at http: e1tts.github.io .
【59】 Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
标题: Wave-U-Mamba:一个端到端框架,实现高质量和高效的语音超分辨率
作者:Yongjoon Lee,Chanwoo Kim
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:语音超分辨率(SSR)是通过恢复丢失的高频分量来增强低分辨率语音信号的任务。传统的方法通常重建对数梅尔特征,随后是在波形域中生成高分辨率语音的声码器。然而,由于对数梅尔特征缺乏相位信息,这可能导致重建阶段期间的性能下降。最近的进展与选择性状态空间模型(SSM)的动机,我们提出了一种方法,称为波U-曼巴,直接在时域进行SSR。在我们的比较研究中,包括WSRGlow,NU-Wave 2和AudioSR等模型,Wave-U-Mamba表现出卓越的性能,在各种低分辨率采样率(从8 kHz到24 kHz)中实现了最低的对数谱距离(LSD)。此外,使用平均意见得分(MOS)进行的主观人类评价表明,我们的方法产生了具有自然和人类品质的SSR。此外,Wave-U-Mamba实现了这些结果,同时在单个A100 GPU上生成高分辨率语音的速度比基线模型快九倍,参数大小不到基线模型的2%。摘要:Speech Super-Resolution (SSR) is a task of enhancing low-resolution speech signals by restoring missing high-frequency components. Conventional approaches typically reconstruct log-mel features, followed by a vocoder that generates high-resolution speech in the waveform domain. However, as log-mel features lack phase information, this can result in performance degradation during the reconstruction phase. Motivated by recent advances with Selective State Spaces Models (SSMs), we propose a method, referred to as Wave-U-Mamba that directly performs SSR in time domain. In our comparative study, including models such as WSRGlow, NU-Wave 2, and AudioSR, Wave-U-Mamba demonstrates superior performance, achieving the lowest Log-Spectral Distance (LSD) across various low-resolution sampling rates, ranging from 8 kHz to 24 kHz. Additionally, subjective human evaluations, scored using Mean Opinion Score (MOS) reveal that our method produces SSR with natural and human-like quality. Furthermore, Wave-U-Mamba achieves these results while generating high-resolution speech over nine times faster than baseline models on a single A100 GPU, with parameter sizes less than 2% of those in the baseline models.
【60】 Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
标题: 未标记条件下异常声音检测的鉴别特征空间训练的改进
作者:Takuya Fujimura,Ibuki Kuroyanagi,Tomoki Toda
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:在异常声音检测中,判别方法表现出优越的性能。该方法通过对正常声音元信息标签的分类,构造了一个有区别的特征空间。该特征空间反映了机器声音的差异,并有效地捕获异常声音。然而,当元信息标签丢失时,其性能显着降低。在本文中,我们通过两种方法来提高未标记条件下的判别方法的性能。首先,我们增强了特征提取器,使其在未标记条件下表现更好。我们的增强型特征提取器采用多分辨率频谱图和新的训练策略。其次,我们提出了各种伪标记的方法,有效地训练的特征提取器。实验结果表明,所提出的特征提取器和伪标记方法显着提高性能,在未标记的条件下。摘要:In anomalous sound detection, the discriminative method has demonstrated superior performance. This approach constructs a discriminative feature space through the classification of the meta-information labels for normal sounds. This feature space reflects the differences in machine sounds and effectively captures anomalous sounds. However, its performance significantly degrades when the meta-information labels are missing. In this paper, we improve the performance of a discriminative method under unlabeled conditions by two approaches. First, we enhance the feature extractor to perform better under unlabeled conditions. Our enhanced feature extractor utilizes multi-resolution spectrograms with a new training strategy. Second, we propose various pseudo-labeling methods to effectively train the feature extractor. The experimental evaluations show that the proposed feature extractor and pseudo-labeling methods significantly improve performance under unlabeled conditions.
【61】 Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
标题: 通过稳定的Forces生成提高基于扩散的零激发语音合成的鲁棒性
作者:Changjin Han,Seokgi Lee,Gyuhyeon Nam,Gyeongsu Chae
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:扩散模型在文本到语音(TTS),即使在zero-shot的情况下取得了显着的成功。最近的努力旨在解决推理速度和声音质量之间的权衡,通常被认为是扩散模型的主要缺点。然而,我们发现一个关键的发音错误问题被忽视了。我们的初步研究揭示了不稳定的发音扩散过程中产生的。基于这一观察,我们介绍了StableForm-TTS,一种新的zero-shot语音合成框架,旨在产生强大的发音,同时保持扩散建模的优势。通过开创性地采用源过滤器理论的扩散TTS,我们提出了一个精心设计的架构,稳定的共振峰生成。对未知说话人的实验结果表明,我们的模型在发音准确性和自然度方面优于最先进的方法,具有可比的说话人相似性。此外,随着数据和模型大小的增加,我们的模型表现出有效的可扩展性。摘要:Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of diffusion models. However, we find a critical mispronunciation issue is being overlooked. Our preliminary study reveals the unstable pronunciation resulting from the diffusion process. Based on this observation, we introduce StableForm-TTS, a novel zero-shot speech synthesis framework designed to produce robust pronunciation while maintaining the advantages of diffusion modeling. By pioneering the adoption of source-filter theory in diffusion TTS, we propose an elaborate architecture for stable formant generation. Experimental results on unseen speakers show that our model outperforms the state-of-the-art method in terms of pronunciation accuracy and naturalness, with comparable speaker similarity. Moreover, our model demonstrates effective scalability as both data and model sizes increase.
【62】 ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
标题: ReCLAP:通过描述声音改进Zero-Shot音频分类
作者:Sreyan Ghosh,Sonal Kumar,Chandra Kiran Reddy Evuru,Oriol Nieto,Ramani Duraiswami,Dinesh Manocha
备注:Code and Checkpoints: this https URL
链接:点击下载PDF文件
摘要:开放词汇音频语言模型,如CLAP,提供了一个很有前途的方法,zero-shot音频分类(ZSAC),使分类与自然语言提示指定的任何任意一组类别。在本文中,我们提出了一个简单而有效的方法来改善ZSAC与CLAP。具体来说,我们从使用带有抽象类别标签的提示的传统方法(例如,器官的声音)到在不同的上下文中使用声音的固有描述性特征来描述声音的提示(例如,风琴低沉而洪亮的音调充满了大教堂。为了实现这一目标,我们首先提出了ReCLAP,这是一个用重写的音频字幕训练的CLAP模型,用于提高对野外声音的理解。这些重写的字幕使用其独特的辨别特征来描述原始字幕中的每个声音事件。ReCLAP在多模态音频文本检索和ZSAC方面都优于所有基线。接下来,为了使用ReCLAP改进zero-shot音频分类,我们提出了提示增强。与采用手写模板提示的传统方法相比,我们为数据集中的每个唯一标签生成自定义提示。这些自定义提示首先在标签中描述声音事件,然后在不同的场景中使用它们。我们提出的方法将ReCLAP在ZSAC上的性能提高了1%-18%,并优于所有基线1% -55%。摘要:Open-vocabulary audio-language models, like CLAP, offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper, we propose a simple but effective method to improve ZSAC with CLAP. Specifically, we shift from the conventional method of using prompts with abstract category labels (e.g., Sound of an organ) to prompts that describe sounds using their inherent descriptive features in a diverse context (e.g.,The organ's deep and resonant tones filled the cathedral.). To achieve this, we first propose ReCLAP, a CLAP model trained with rewritten audio captions for improved understanding of sounds in the wild. These rewritten captions describe each sound event in the original caption using their unique discriminative characteristics. ReCLAP outperforms all baselines on both multi-modal audio-text retrieval and ZSAC. Next, to improve zero-shot audio classification with ReCLAP, we propose prompt augmentation. In contrast to the traditional method of employing hand-written template prompts, we generate custom prompts for each unique label in the dataset. These custom prompts first describe the sound event in the label and then employ them in diverse scenes. Our proposed method improves ReCLAP's performance on ZSAC by 1%-18% and outperforms all baselines by 1% - 55%.
【63】 Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
标题: 从策划值得信赖、注释良好且有用的无序英语言语数据集中吸取的教训
作者:Pan-Pan Jiang,Jimmy Tobin,Katrin Tomanek,Robert L. MacDonald,Katie Seaver,Richard Cave,Marilyn Ladewig,Rus Heywood,Jordan R. Green
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:Project Euphonia是Google的一项计划,致力于改善无序语音的自动语音识别(ASR)。该项目的一个中心目标是创建一个大型,高质量和多样化的语音语料库。该报告描述了该项目在数据收集和注释方法方面的最新进展,例如扩大数据库中的说话者多样性,为35万(总计120万)音频记录添加人工审核的转录校正和音频质量标签,并为数据库中超过75%的说话者积累了一套全面的元数据(包括40多个语音特征标签)。我们报告了转录更正对我们的机器学习(ML)研究的影响,评估者之间的变异性评估的混乱的语音模式,以及我们收集语音元数据的理由。我们还考虑了使用自动化现成的注释方法来评估无序语音的局限性。摘要:Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. This report describes the project's latest advancements in data collection and annotation methodologies, such as expanding speaker diversity in the database, adding human-reviewed transcript corrections and audio quality tags to 350K (of the 1.2M total) audio recordings, and amassing a comprehensive set of metadata (including more than 40 speech characteristic labels) for over 75 % of the speakers in the database. We report on the impact of transcript corrections on our machine-learning (ML) research, inter-rater variability of assessments of disordered speech patterns, and our rationale for gathering speech metadata. We also consider the limitations of using automated off-the-shelf annotation methods for assessing disordered speech.
【64】 MambaFoley: Foley Sound Generation using Selective State-Space Models
标题: MambaFoley:使用选择性状态空间模型生成Foley声音
作者:Marco Furio Colombo,Francesca Ronchini,Luca Comanducci,Fabio Antonacci
链接:点击下载PDF文件
摘要:深度学习的最新进展导致音频内容生成技术的广泛使用,特别是在各种任务中采用去噪扩散概率模型(DDPM)。其中,Foley声音合成因其在创建多媒体内容的应用中的作用而特别感兴趣。鉴于声音的时间依赖性,设计能够有效处理音频样本的顺序建模的生成模型至关重要。选择性状态空间模型(SSM)最近被提出作为一个有效的替代方案,以前提出的技术,表现出较低的计算复杂度的竞争力的性能。在本文中,我们介绍了MambaFoley,一个基于扩散的模型,据我们所知,是第一个利用最近提出的SSM称为曼巴的福利声音生成任务。为了评估所提出的方法的有效性,我们比较它与一个国家的最先进的福利声音生成模型,使用客观和主观的分析。摘要:Recent advancements in deep learning have led to widespread use of techniques for audio content generation, notably employing Denoising Diffusion Probabilistic Models (DDPM) across various tasks. Among these, Foley Sound Synthesis is of particular interest for its role in applications for the creation of multimedia content. Given the temporal-dependent nature of sound, it is crucial to design generative models that can effectively handle the sequential modeling of audio samples. Selective State Space Models (SSMs) have recently been proposed as a valid alternative to previously proposed techniques, demonstrating competitive performance with lower computational complexity. In this paper, we introduce MambaFoley, a diffusion-based model that, to the best of our knowledge, is the first to leverage the recently proposed SSM known as Mamba for the Foley sound generation task. To evaluate the effectiveness of the proposed method, we compare it with a state-of-the-art Foley sound generative model using both objective and subjective analyses.
【65】 SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
标题: SLiCK:利用子序列进行长度限制的关键字发现
作者:Kumari Nishu,Minsik Cho,Devang Naik
链接:点击下载PDF文件
摘要:在资源受限的边缘设备上进行用户定义的关键字识别是一项挑战。然而,关键字通常受到最大关键字长度的限制,这在以前的作品中基本上没有得到充分利用。我们对关键字长度分布的分析表明,用户定义的关键字定位可以被视为一个长度受限的问题,消除了对可变文本长度聚合的需要。这导致了我们提出的高效关键字定位方法SLiCK(利用子序列进行长度约束关键字定位)。我们还引入了一个语义级匹配方案,以更细的粒度学习音频-文本关系,从而通过增强上下文更有效地区分相似的关键字。在SLiCK中,使用两个模块的多任务学习方法来训练模型:Matcher(话语级匹配任务,新颖的语义级匹配任务)和Encoder(音素识别任务)。所提出的方法改进了Librphrase硬数据集上的基线结果,将AUC从88.52 $增加到94.9 $,并将EER从18.82 $降低到11.1 $。摘要:User-defined keyword spotting on a resource-constrained edge device is challenging. However, keywords are often bounded by a maximum keyword length, which has been largely under-leveraged in prior works. Our analysis of keyword-length distribution shows that user-defined keyword spotting can be treated as a length-constrained problem, eliminating the need for aggregation over variable text length. This leads to our proposed method for efficient keyword spotting, SLiCK (exploiting Subsequences for Length-Constrained Keyword spotting). We further introduce a subsequence-level matching scheme to learn audio-text relations at a finer granularity, thus distinguishing similar-sounding keywords more effectively through enhanced context. In SLiCK, the model is trained with a multi-task learning approach using two modules: Matcher (utterance-level matching task, novel subsequence-level matching task) and Encoder (phoneme recognition task). The proposed method improves the baseline results on Libriphrase hard dataset, increasing AUC from $88.52$ to $94.9$ and reducing EER from $18.82$ to $11.1$.
eess.AS音频处理
【1】 An Efficient Self-Learning Framework For Interactive Spoken Dialog Systems标题: 交互式口语对话系统的高效自学习框架
作者:Hitesh Tulsiani,David M. Chan,Shalini Ghosh,Garima Lalwani,Prabhat Pandey,Ankish Bansal,Sri Garimella,Ariya Rastrow,Björn Hoffmeister
备注:Presented at ICML 2024
链接:点击下载PDF文件
摘要:对话系统,如语音助理,预计将与用户进行复杂的,不断发展的对话。不幸的是,在这样的应用中部署的传统自动语音识别(ASR)系统通常被训练为独立地识别每个回合,并且缺乏适应会话上下文或结合用户反馈的能力。在这项工作中,我们介绍了一个通用的框架ASR对话系统,可以超越学习单轮话语,并随着时间的推移学习如何适应显式监督和隐式用户反馈中存在的多轮对话。我们通过利用学生-教师学习和上下文感知对话处理的进步,并设计对比自我监督方法欧姆,一种新的在线硬否定挖掘方法。我们表明,与传统训练相比,利用我们的新框架,在现实世界的对话系统中,相对WER减少了近10%,在公共合成数据中减少了高达26%。摘要:Dialog systems, such as voice assistants, are expected to engage with users in complex, evolving conversations. Unfortunately, traditional automatic speech recognition (ASR) systems deployed in such applications are usually trained to recognize each turn independently and lack the ability to adapt to the conversational context or incorporate user feedback. In this work, we introduce a general framework for ASR in dialog systems that can go beyond learning from single-turn utterances and learn over time how to adapt to both explicit supervision and implicit user feedback present in multi-turn conversations. We accomplish that by leveraging advances in student-teacher learning and context-aware dialog processing, and designing contrastive self-supervision approaches with Ohm, a new online hard-negative mining approach. We show that leveraging our new framework compared to traditional training leads to relative WER reductions of close to 10% in real-world dialog systems, and up to 26% on public synthetic data.
【2】 Meta-Whisper: Speech-Based Meta-ICL for ASR on Low-Resource Languages
标题: Meta-Whisper:基于语音的Meta-ICL,用于低资源语言上的ASB
作者:Ming-Hao Hsu,Kuan Po Huang,Hung-yi Lee
链接:点击下载PDF文件
摘要:本文提出了元耳语,一种新的方法来提高自动语音识别(ASR)的低资源语言使用耳语模型。通过利用Meta上下文学习(Meta-ICL)和k最近邻(KNN)算法进行样本选择,Meta-Whisper增强了Whisper在不熟悉的语言中识别语音的能力,而无需进行广泛的微调。在ML-SUPERB数据集上的实验表明,与原始Whisper模型相比,Meta-Whisper显著降低了低资源语言的字符错误率(CER)。这种方法为开发适应性更强的多语言ASR系统提供了一种很有前途的解决方案,特别是对于资源有限的语言。摘要:This paper presents Meta-Whisper, a novel approach to improve automatic speech recognition (ASR) for low-resource languages using the Whisper model. By leveraging Meta In-Context Learning (Meta-ICL) and a k-Nearest Neighbors (KNN) algorithm for sample selection, Meta-Whisper enhances Whisper's ability to recognize speech in unfamiliar languages without extensive fine-tuning. Experiments on the ML-SUPERB dataset show that Meta-Whisper significantly reduces the Character Error Rate (CER) for low-resource languages compared to the original Whisper model. This method offers a promising solution for developing more adaptable multilingual ASR systems, particularly for languages with limited resources.
【3】 Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement
标题: 利用MAMBA联合频谱和空间学习进行多通道语音增强
作者:Wenze Ren,Haibin Wu,Yi-Cheng Lin,Xuanjun Chen,Rong Chao,Kuo-Hsuan Hung,You-Jin Li,Wen-Yuan Ting,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
摘要:在多通道语音增强中,有效地捕获不同麦克风之间的空间和频谱信息对于降噪至关重要。传统方法,如CNN或LSTM,试图对全波段和子波段光谱和空间特征的时间动态进行建模。然而,这些方法在完全建模复杂的时间依赖性方面面临限制,特别是在动态声学环境中。为了克服这些挑战,我们修改了目前的先进模型McNet通过引入改进版本的Mamba,一个状态空间模型,并进一步提出MCMamba。MCMAamba已经完全重新设计,将全波段和窄带空间信息与子波段和全波段光谱特征集成在一起,为空间和光谱信息建模提供了更全面的方法。我们的实验结果表明,MCMamba显著提高了多通道语音增强中的空间和频谱特征建模,优于McNet,并在CHiME-3数据集上实现了最先进的性能。此外,我们发现,曼巴表现非常好的建模光谱信息。摘要:In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, these approaches face limitations in fully modeling complex temporal dependencies, especially in dynamic acoustic environments. To overcome these challenges, we modify the current advanced model McNet by introducing an improved version of Mamba, a state-space model, and further propose MCMamba. MCMamba has been completely reengineered to integrate full-band and narrow-band spatial information with sub-band and full-band spectral features, providing a more comprehensive approach to modeling spatial and spectral information. Our experimental results demonstrate that MCMamba significantly improves the modeling of spatial and spectral features in multichannel speech enhancement, outperforming McNet and achieving state-of-the-art performance on the CHiME-3 dataset. Additionally, we find that Mamba performs exceptionally well in modeling spectral information.
【4】 Ultra-Low Latency Speech Enhancement - A Comprehensive Study
标题: 超低延迟语音增强-综合研究
作者:Haibin Wu,Sebastian Braun
链接:点击下载PDF文件
摘要:语音增强模型应满足非常低的延迟要求,通常小于5毫秒的听力辅助设备。虽然已经提出了各种低延迟技术,但在使用DNN的受控设置中比较这些方法仍然是空白。以前的论文在任务、训练数据、脚本和评估设置方面存在差异,这使得公平的比较变得不可能。此外,所有方法都是在小型模拟数据集上进行测试的,因此很难公平地评估它们在真实世界条件下的性能,这可能会影响科学发现的可靠性。为了解决这些问题,我们使用大规模数据的一致训练来全面研究各种低延迟技术,并使用真实世界数据的更相关指标进行评估。具体来说,我们探讨了非对称窗口,可学习的窗口,自适应时域滤波器组和未来帧预测技术的有效性。此外,我们还研究了增加模型大小是否可以补偿减小的窗口大小,以及低延迟环境中的新型Mamba架构。摘要:Speech enhancement models should meet very low latency requirements typically smaller than 5 ms for hearing assistive devices. While various low-latency techniques have been proposed, comparing these methods in a controlled setup using DNNs remains blank. Previous papers have variations in task, training data, scripts, and evaluation settings, which make fair comparison impossible. Moreover, all methods are tested on small, simulated datasets, making it difficult to fairly assess their performance in real-world conditions, which could impact the reliability of scientific findings. To address these issues, we comprehensively investigate various low-latency techniques using consistent training on large-scale data and evaluate with more relevant metrics on real-world data. Specifically, we explore the effectiveness of asymmetric windows, learnable windows, adaptive time domain filterbanks, and the future-frame prediction technique. Additionally, we examine whether increasing the model size can compensate for the reduced window size, as well as the novel Mamba architecture in low-latency environments.
【5】 oboVox Far Field Speaker Recognition: A Novel Data Augmentation Approach with Pretrained Models
标题: oboVox远场说话人识别:一种采用预训练模型的新型数据增强方法
作者:Muhammad Sudipto Siam Dip,Md Anik Hasan,Sapnil Sarker Bipro,Md Abdur Raiyan,Mohammod Abdul Motin
备注:5 pages, 2 figures
链接:点击下载PDF文件
摘要:在这项研究中,我们解决了说话人识别的挑战,使用一种新的数据增强技术,增加噪音的注册文件。这种技术有效地对齐了测试和注册文件的来源,提高了可比性。使用了各种预训练模型,其中resnet模型实现了最高的DCF 0.84和EER 13.44。增强技术显着改善这些结果为0.75 DCF和12.79 EER的resnet模型。比较分析表明,resnet的优势,如ECPA,梅尔光谱图,Payonnet,和Titanet大模型。结果,以及不同的增强方案,有助于本文的RoboVox远场说话人识别的成功摘要:In this study, we address the challenge of speaker recognition using a novel data augmentation technique of adding noise to enrollment files. This technique efficiently aligns the sources of test and enrollment files, enhancing comparability. Various pre-trained models were employed, with the resnet model achieving the highest DCF of 0.84 and an EER of 13.44. The augmentation technique notably improved these results to 0.75 DCF and 12.79 EER for the resnet model. Comparative analysis revealed the superiority of resnet over models such as ECPA, Mel-spectrogram, Payonnet, and Titanet large. Results, along with different augmentation schemes, contribute to the success of RoboVox far-field speaker recognition in this paper
【6】 Speech as a Biomarker for Disease Detection
标题: 言语作为疾病检测的生物标志物
作者:Catarina Botelho,Alberto Abad,Tanja Schultz,Isabel Trancoso
链接:点击下载PDF文件
摘要:语音是一种丰富的生物标志物,它编码了有关说话者健康的大量信息,因此已被提出用于检测许多疾病,取得了可喜的成果。然而,关于为自动检测这些疾病而训练的模型实际上在学习什么以及它们预测的基础仍然存在问题,这可能会对患者的生活产生重大影响。这项工作倡导一种可解释的健康模型,适合检测几种疾病,其动机是观察到影响言语的疾病往往对言语信号产生重叠影响。提出了一个框架,首先定义“参考语音”,然后利用该定义进行疾病检测。参考语音通过参考间隔来表征,即,来自参考人群的具有临床意义的声学和语言特征的典型值。这种在语音领域作为生物标志物的新方法受到临床实验室科学中使用参考区间的启发。新的扬声器从这个参考模型的偏差进行量化,并作为输入检测阿尔茨海默氏症和帕金森氏症。探索的分类策略是基于神经加法模型,一种玻璃盒神经网络,它可以解释。建议的参考语音表征和疾病检测框架的目的是支持医学界提供临床上有意义的解释,可以作为一个有价值的第二意见。摘要:Speech is a rich biomarker that encodes substantial information about the health of a speaker, and thus it has been proposed for the detection of numerous diseases, achieving promising results. However, questions remain about what the models trained for the automatic detection of these diseases are actually learning and the basis for their predictions, which can significantly impact patients' lives. This work advocates for an interpretable health model, suitable for detecting several diseases, motivated by the observation that speech-affecting disorders often have overlapping effects on speech signals. A framework is presented that first defines "reference speech" and then leverages this definition for disease detection. Reference speech is characterized through reference intervals, i.e., the typical values of clinically meaningful acoustic and linguistic features derived from a reference population. This novel approach in the field of speech as a biomarker is inspired by the use of reference intervals in clinical laboratory science. Deviations of new speakers from this reference model are quantified and used as input to detect Alzheimer's and Parkinson's disease. The classification strategy explored is based on Neural Additive Models, a type of glass-box neural network, which enables interpretability. The proposed framework for reference speech characterization and disease detection is designed to support the medical community by providing clinically meaningful explanations that can serve as a valuable second opinion.
【7】 RF-GML: Reference-Free Generative Machine Listener
标题: RF-GML:无参考生成机器收件箱
作者:Arijit Biswas,Guanxin Jiang
备注:Pre-review version submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:本文介绍了一种新的无参考(RF)的音频质量度量称为RF生成机器的音频质量(RF GML),旨在评估编码的单声道,立体声和双耳音频在48 kHz的采样率。RF-GML利用了来自最先进的全参考(FR)生成机器学习(GML)的迁移学习,只需最小的架构修改。术语“生成”是指模型生成任意数量的模拟听力分数的能力。与现有的RF模型不同,RF-GML可以准确预测各种内容类型和编解码器的主观质量分数。广泛的评估表明,它的优势,在评级未编码的音频和区分不同层次的编码文物。RF-GML的性能和多功能性使其成为各种应用中编码音频质量评估和监控的宝贵工具,所有这些都不需要参考信号。摘要:This paper introduces a novel reference-free (RF) audio quality metric called the RF-Generative Machine Listener (RF-GML), designed to evaluate coded mono, stereo, and binaural audio at a 48 kHz sample rate. RF-GML leverages transfer learning from a state-of-the-art full-reference (FR) Generative Machine Listener (GML) with minimal architectural modifications. The term "generative" refers to the model's ability to generate an arbitrary number of simulated listening scores. Unlike existing RF models, RF-GML accurately predicts subjective quality scores across diverse content types and codecs. Extensive evaluations demonstrate its superiority in rating unencoded audio and distinguishing different levels of coding artifacts. RF-GML's performance and versatility make it a valuable tool for coded audio quality assessment and monitoring in various applications, all without the need for a reference signal.
【8】 Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
标题: CLAR-DPO:通过直接偏好优化的可控情感语音合成
作者:Xiaoxue Gao,Chen Zhang,Yiming Chen,Huayun Zhang,Nancy F. Chen
备注:5 pages
链接:点击下载PDF文件
摘要:当前的情感文本到语音(TTS)模型主要进行监督训练,以学习从文本和期望的情感到其情感语音的转换,专注于每个文本-语音对的单个情感。这些模型只能学习正确的情感输出,而不能完全理解其他情感特征,这限制了它们捕捉不同情感之间细微差别的能力。我们提出了一个可控的DPO方法,它采用直接偏好优化,以区分微妙的情感之间的细微差别,通过优化对首选的情感,而不是不太喜欢的情感。我们建议利用情感感知LLM-TTS神经架构来利用LLM的上下文学习和推理跟随能力,而不是依赖于现有情感TTS模型中使用的传统神经架构。综合实验证实,我们提出的方法优于现有的基线。摘要:Current emotional text-to-speech (TTS) models predominantly conduct supervised training to learn the conversion from text and desired emotion to its emotional speech, focusing on a single emotion per text-speech pair. These models only learn the correct emotional outputs without fully comprehending other emotion characteristics, which limits their capabilities of capturing the nuances between different emotions. We propose a controllable Emo-DPO approach, which employs direct preference optimization to differentiate subtle emotional nuances between emotions through optimizing towards preferred emotions over less preferred emotional ones. Instead of relying on traditional neural architectures used in existing emotional TTS models, we propose utilizing the emotion-aware LLM-TTS neural architecture to leverage LLMs' in-context learning and instruction-following capabilities. Comprehensive experiments confirm that our proposed method outperforms the existing baselines.
【9】 Room impulse response prototyping using receiver distance estimations for high quality room equalisation algorithms
标题: 使用接收器距离估计的房间脉冲响应原型用于高质量房间均衡算法
作者:James Brooks-Park,Martin Bo Møller,Jan Østergaard,Søren Bech,Steven van de Par
链接:点击下载PDF文件
摘要:房间均衡旨在提高混响环境中扬声器再现的质量,补偿由不完美的房间反射和频率相关扬声器方向性引起的着色。房间均衡领域中的一种常见技术是反转原型房间脉冲响应(RIR)。原型响应由分布在收听区域周围的几个响应组成,而不是在收听位置反转单个RIR。本文提出了一种脉冲响应原型的方法,使用估计的接收机位置,形成一个加权平均原型响应。描述了一种接收机距离估计的方法,支持原型RIR的实现。建议的原型制作方法相比,其他方法通过测量后均衡光谱偏差在几个位置在一个模拟的房间。摘要:Room equalisation aims to increase the quality of loudspeaker reproduction in reverberant environments, compensating for colouration caused by imperfect room reflections and frequency dependant loudspeaker directivity. A common technique in the field of room equalisation, is to invert a prototype Room Impulse Response (RIR). Rather than inverting a single RIR at the listening position, a prototype response is composed of several responses distributed around the listening area. This paper proposes a method of impulse response prototyping, using estimated receiver positions, to form a weighted average prototype response. A method of receiver distance estimation is described, supporting the implementation of the prototype RIR. The proposed prototyping method is compared to other methods by measuring their post equalisation spectral deviation at several positions in a simulated room.
【10】 StyleTTS-ZS: Efficient High-Quality Zero-Shot Text-to-Speech Synthesis with Distilled Time-Varying Style Diffusion
标题: StyleTTS-ZZ:具有蒸馏时变风格扩散的高效高质量Zero-Shot文本到语音合成
作者:Yinghao Aaron Li,Xilin Jiang,Cong Han,Nima Mesgarani
链接:点击下载PDF文件
摘要:大规模文本到语音(TTS)模型的快速发展,导致了显着的进步,建模不同的说话人韵律和声音。然而,这些模型通常面临推理速度慢、依赖复杂的预训练神经编解码器表示以及难以实现自然度和与参考说话人的高相似度等问题。为了解决这些挑战,这项工作介绍了StyleTTS-ZS,一个有效的zero-shot TTS模型,利用蒸馏时变风格扩散,以捕捉不同的扬声器身份和韵律。我们提出了一种新的方法,代表人类语音使用输入文本和固定长度的时变离散风格代码来捕捉不同的韵律变化,训练对抗多模态判别。然后建立一个扩散模型,对这种时变风格的代码进行采样,以实现有效的潜在扩散。在风格扩散过程中,StyleTTS-ZS使用无分类器的指导,实现了与参考说话人的高相似度。此外,为了加快采样速度,风格扩散模型仅使用10 k个样本进行感知损失提取,保持语音质量和相似性,同时将推理速度降低90%。该模型在自然度和相似度方面均优于现有的大规模zero-shot TTS模型,采样速度提高了10-20倍,是大规模zero-shot TTS系统的理想选择。音频演示,代码和模型可在https: styletts-zs.github.io 。摘要:The rapid development of large-scale text-to-speech (TTS) models has led to significant advancements in modeling diverse speaker prosody and voices. However, these models often face issues such as slow inference speeds, reliance on complex pre-trained neural codec representations, and difficulties in achieving naturalness and high similarity to reference speakers. To address these challenges, this work introduces StyleTTS-ZS, an efficient zero-shot TTS model that leverages distilled time-varying style diffusion to capture diverse speaker identities and prosodies. We propose a novel approach that represents human speech using input text and fixed-length time-varying discrete style codes to capture diverse prosodic variations, trained adversarially with multi-modal discriminators. A diffusion model is then built to sample this time-varying style code for efficient latent diffusion. Using classifier-free guidance, StyleTTS-ZS achieves high similarity to the reference speaker in the style diffusion process. Furthermore, to expedite sampling, the style diffusion model is distilled with perceptual loss using only 10k samples, maintaining speech quality and similarity while reducing inference speed by 90%. Our model surpasses previous state-of-the-art large-scale zero-shot TTS models in both naturalness and similarity, offering a 10-20 faster sampling speed, making it an attractive alternative for efficient large-scale zero-shot TTS systems. The audio demo, code and models are available at https: styletts-zs.github.io .
【11】 TBDM-Net: Bidirectional Dense Networks with Gender Information for Speech Emotion Recognition
标题: TBDM-Net:用于语音情感识别的具有性别信息的双向密集网络
作者:Vlad Striletchi,Cosmin Striletchi,Adriana Stan
备注:In Proceedings of 2024 IEEE International Workshop on Machine Learning for Signal Processing, London, UK
链接:点击下载PDF文件
摘要:本文提出了一种新的基于深度神经网络的语音情感识别(SER)架构。该架构利用了多层双向扩张卷积之间的密集互连。线性内核动态融合这些层的输出,以产生最终的情感类预测。这种创新的架构被表示为TBDM-Net:时间感知双向密集多尺度网络。我们对TBDM-Net进行了全面的性能评估,包括消融研究,在六个广泛认可的SER数据集上进行单峰语音情感识别。此外,我们探索的影响,性别知情的情绪预测附加黄金或预测的性别标签的架构的输入或预测。TBDM-Net的实施情况可在以下网址查阅:https: github.com adrianastan tbdm-net摘要:This paper presents a novel deep neural network-based architecture tailored for Speech Emotion Recognition (SER). The architecture capitalises on dense interconnections among multiple layers of bidirectional dilated convolutions. A linear kernel dynamically fuses the outputs of these layers to yield the final emotion class prediction. This innovative architecture is denoted as TBDM-Net: Temporally-Aware Bi-directional Dense Multi-Scale Network. We conduct a comprehensive performance evaluation of TBDM-Net, including an ablation study, across six widely-acknowledged SER datasets for unimodal speech emotion recognition. Additionally, we explore the influence of gender-informed emotion prediction by appending either golden or predicted gender labels to the architecture's inputs or predictions. The implementation of TBDM-Net is accessible at: https: github.com adrianastan tbdm-net
【12】 DNN-based ensemble singing voice synthesis with interactions between singers
标题: 基于DNN的合奏歌唱声音合成,具有歌手之间的互动
作者:Hiroaki Hyodo,Shinnosuke Takamichi,Tomohiko Nakamura,Junya Koguchi,Hiroshi Saruwatari
链接:点击下载PDF文件
摘要:我们提出了一个歌唱声音合成(SVS)的方法更统一的合奏歌唱声音建模歌手之间的相互作用。大多数现有的SVS方法旨在合成独唱声音,并且不考虑歌手之间的交互,即,调整自己的声音以适应其他人的声音。由于从独唱声音到合奏声音的产生忽略了相互作用,它会降低声乐合奏的统一性。因此,我们提出了一个SVS再现的相互作用。它是基于一个架构,使用多个语音部分的乐谱,和损失函数,模拟相互作用的影响,声学特征。实验结果表明,我们的方法提高了声乐合奏的统一性。摘要:We propose a singing voice synthesis (SVS) method for a more unified ensemble singing voice by modeling interactions between singers. Most existing SVS methods aim to synthesize a solo voice, and do not consider interactions between singers, i.e., adjusting one's own voice to the others' voices. Since the production of ensemble voices from solo singing voices ignores the interactions, it can degrade the unity of the vocal ensemble. Therefore, we propose a SVS that reproduces the interactions. It is based on an architecture that uses musical scores of multiple voice parts, and loss functions that simulate the interactions' effect to acoustic features. Experimental results show that our methods improve the unity of the vocal ensemble.
【13】 A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models
标题: 使用大型语言模型的Zero-Shot非侵入性语音评估研究
作者:Ryandhimas E. Zezario,Sabato M. Siniscalchi,Hsin-Min Wang,Yu Tsao
链接:点击下载PDF文件
摘要:本文研究了两种利用大型语言模型的zero-shot非侵入式语音评估策略。首先,我们探索GPT-4 o的音频分析功能。其次,我们提出了GPT-Whisper,它使用Whisper作为音频到文本的模块,并通过有针对性的提示工程评估文本的自然度。我们评估评估指标预测的GPT-4 o和GPT-Whisper检查其与基于人类的质量和可懂度评估,以及自动语音识别的字符错误率(CER)的相关性。实验结果表明,单独的GPT-4 o是不是有效的音频分析,而GPT-Whisper表现出更高的预测,表现出中等相关性的语音质量和可懂度,并与CER的高相关性。与有监督的非侵入式神经语音评估模型(即MOS-SSL和MTI-Net)相比,GPT-Whisper与Whisper的CER产生了显着更高的Spearman秩相关性。这些发现验证了GPT-Whisper作为准确的zero-shot语音评估的可靠方法,而不需要额外的训练数据(语音数据和相应的评估分数)。摘要:This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the naturalness of text via targeted prompt engineering. We evaluate assessment metrics predicted by GPT-4o and GPT-Whisper examining their correlations with human-based quality and intelligibility assessments, and character error rate (CER) of automatic speech recognition. Experimental results show that GPT-4o alone is not effective for audio analysis; whereas, GPT-Whisper demonstrates higher prediction, showing moderate correlation with speech quality and intelligibility, and high correlation with CER. Compared to supervised non-intrusive neural speech assessment models, namely MOS-SSL and MTI-Net, GPT-Whisper yields a notably higher Spearman's rank correlation with the CER of Whisper. These findings validate GPT-Whisper as a reliable method for accurate zero-shot speech assessment without requiring additional training data (speech data and corresponding assessment scores).
【14】 Self-supervised Multimodal Speech Representations for the Assessment of Schizophrenia Symptoms
标题: 自我监督的多模式语音表示用于评估精神分裂症症状
作者:Gowtham Premananth,Carol Espy-Wilson
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:多模式精神分裂症评估系统在过去几年中获得了牵引力。这项工作介绍了精神分裂症评估系统,以区分精神分裂症的突出症状类别,并预测整体精神分裂症的严重程度评分。我们开发了一个基于矢量量化变分自动编码器(VQ-VAE)的多模态表示学习(MRL)模型,从声道变量(TV)和面部动作单元(FAU)生成任务无关的语音表示。然后,这些表示用于基于多任务学习(MTL)的下游预测模型,以获得类别标签和总体严重性评分。所提出的框架在所有评估指标(加权F1得分,AUC-ROC得分和加权准确度)上都优于以前的多类分类任务。此外,它还估计了精神分裂症的严重程度评分,这是一项早期方法没有解决的任务。摘要:Multimodal schizophrenia assessment systems have gained traction over the last few years. This work introduces a schizophrenia assessment system to discern between prominent symptom classes of schizophrenia and predict an overall schizophrenia severity score. We develop a Vector Quantized Variational Auto-Encoder (VQ-VAE) based Multimodal Representation Learning (MRL) model to produce task-agnostic speech representations from vocal Tract Variables (TVs) and Facial Action Units (FAUs). These representations are then used in a Multi-Task Learning (MTL) based downstream prediction model to obtain class labels and an overall severity score. The proposed framework outperforms the previous works on the multi-class classification task across all evaluation metrics (Weighted F1 score, AUC-ROC score, and Weighted Accuracy). Additionally, it estimates the schizophrenia severity score, a task not addressed by earlier approaches.
【15】 Extract and Diffuse: Latent Integration for Improved Diffusion-based Speech and Vocal Enhancement
标题: 提取和扩散:用于改进基于扩散的语音和人声增强的潜在集成
作者:Yudong Yang,Zhan Liu,Wenyi Yu,Guangzhi Sun,Qiuqiang Kong,Chao Zhang
链接:点击下载PDF文件
摘要:基于扩散的生成模型最近在语音和声音增强方面取得了显着的成果,因为它们能够模拟复杂的语音数据分布。虽然这些模型很好地推广到看不见的声学环境,但它们可能无法实现与专门训练以增强特定声学条件的区分模型相同的保真度水平。在本文中,我们提出了Ex-Diff,一种新的基于分数的扩散模型,它集成了由判别模型产生的潜在表示,以改善语音和声音增强,它结合了生成模型和判别模型的优势。在广泛使用的MUSDB数据集上的实验结果显示,与语音和声音增强任务的基线扩散模型相比,SI-SDR和SI-SIR的相对改善分别为3.7%和10.0%。此外,案例研究提供了进一步说明和分析在这种情况下生成和歧视性模型的互补性。摘要:Diffusion-based generative models have recently achieved remarkable results in speech and vocal enhancement due to their ability to model complex speech data distributions. While these models generalize well to unseen acoustic environments, they may not achieve the same level of fidelity as the discriminative models specifically trained to enhance particular acoustic conditions. In this paper, we propose Ex-Diff, a novel score-based diffusion model that integrates the latent representations produced by a discriminative model to improve speech and vocal enhancement, which combines the strengths of both generative and discriminative models. Experimental results on the widely used MUSDB dataset show relative improvements of 3.7% in SI-SDR and 10.0% in SI-SIR compared to the baseline diffusion model for speech and vocal enhancement tasks, respectively. Additionally, case studies are provided to further illustrate and analyze the complementary nature of generative and discriminative models in this context.
【16】 Stutter-Solver: End-to-end Multi-lingual Dysfluency Detection
标题: 口吃解决器:端到端多语言流利检测
作者:Xuanru Zhou,Cheol Jun Cho,Ayati Sharma,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Boon Lead Tee,Maria Luisa Gorno Tempini,Jiachen Lian,Gopala Anumanchipalli
备注:IEEE Spoken Language Technology Workshop 2024
链接:点击下载PDF文件
摘要:当前的事实上的不流利建模方法利用模板匹配算法,其不能推广到跨语言的域外真实世界的不流利,并且不能随着训练数据量的增加而扩展。为了解决这些问题,我们提出了Stutter-Solver:一个端到端的框架,它可以通过精确的类型和时间转录来检测不流畅,灵感来自YOLO对象检测算法。Stutter-Solver可以处理不流利的情况,是一个自然的多语言不流利检测器。为了利用可扩展性和提高性能,我们还引入了三个新的不流利语料库:VCTK-Pro,VCTK-Art和AISHELL 3-Pro,通过发音编码器和基于TTS的方法模拟自然的口语不流利,包括重复,阻塞,缺失,替换和延长。我们的方法在所有可用的不流利语料库上实现了最先进的性能。代码和数据集在https: github.com eureka235 Stutter-Solver上开源摘要:Current de-facto dysfluency modeling methods utilize template matching algorithms which are not generalizable to out-of-domain real-world dysfluencies across languages, and are not scalable with increasing amounts of training data. To handle these problems, we propose Stutter-Solver: an end-to-end framework that detects dysfluency with accurate type and time transcription, inspired by the YOLO object detection algorithm. Stutter-Solver can handle co-dysfluencies and is a natural multi-lingual dysfluency detector. To leverage scalability and boost performance, we also introduce three novel dysfluency corpora: VCTK-Pro, VCTK-Art, and AISHELL3-Pro, simulating natural spoken dysfluencies including repetition, block, missing, replacement, and prolongation through articulatory-encodec and TTS-based methods. Our approach achieves state-of-the-art performance on all available dysfluency corpora. Code and datasets are open-sourced at https: github.com eureka235 Stutter-Solver
【17】 Effective Pre-Training of Audio Transformers for Sound Event Detection
标题: 音频Transformer的有效预训练以进行声音事件检测
作者:Florian Schmid,Tobias Morocutti,Francesco Foscarin,Jan Schlüter,Paul Primus,Gerhard Widmer
备注:Submitted to ICASSP'25. Source code available: this https URL
链接:点击下载PDF文件
摘要:我们提出了一种用于音频频谱图Transformers的预训练管道,用于帧级声音事件检测任务。在常见的预训练步骤之上,我们在AudioSet帧级注释上添加了精心设计的训练例程。这包括一个平衡的采样器,积极的数据增强,和合奏知识蒸馏。对于五个Transformers,我们在AudioSet帧级预测和帧级声音事件检测下游任务上获得了比以前可用的检查点更大的性能改进,证实了我们的管道的有效性。我们发布了由此产生的检查点,研究人员可以直接微调,以构建用于声音事件检测任务的高性能模型。摘要:We propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-level annotations. This includes a balanced sampler, aggressive data augmentation, and ensemble knowledge distillation. For five transformers, we obtain a substantial performance improvement over previously available checkpoints both on AudioSet frame-level predictions and on frame-level sound event detection downstream tasks, confirming our pipeline's effectiveness. We publish the resulting checkpoints that researchers can directly fine-tune to build high-performance models for sound event detection tasks.
【18】 Target Speaker ASR with Whisper
标题: 目标说话者ASB与Whisper
作者:Alexander Polok,Dominik Klement,Matthew Wiesner,Sanjeev Khudanpur,Jan Černocký,Lukáš Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,使使用大型,单扬声器ASR模型,如耳语,目标扬声器ASR。该方法的关键见解是,通过学习以帧级日记化输出为条件来建模说话者之间的相对差异比学习所有说话者嵌入的空间要容易得多。我们发现,在第一个Transformer块之前添加每个日志输出类型的单个偏置项可以将单个扬声器ASR模型转换为目标扬声器ASR模型。我们的目标说话人ASR模型可以用于扬声器归因的ASR生产,在序列中,每个假设的扬声器在日记输出的成绩单。这个简化的模型,扬声器归因ASR使用一个麦克风,仅优于级联的语音分离和日记的11%的绝对ORC-WER的NOTSOFAR-1数据集。摘要:We propose a novel approach to enable the use of large, single speaker ASR models, such as Whisper, for target speaker ASR. The key insight of this method is that it is much easier to model relative differences among speakers by learning to condition on frame-level diarization outputs, than to learn the space of all speaker embeddings. We find that adding even a single bias term per diarization output type before the first transformer block can transform single speaker ASR models, into target speaker ASR models. Our target-speaker ASR model can be used for speaker attributed ASR by producing, in sequence, a transcript for each hypothesized speaker in a diarization output. This simplified model for speaker attributed ASR using only a single microphone outperforms cascades of speech separation and diarization by 11% absolute ORC-WER on the NOTSOFAR-1 dataset.
【19】 Leveraging Self-Supervised Learning for Speaker Diarization
标题: 利用自我监督学习进行发言者日记化
作者:Jiangyu Han,Federico Landini,Johan Rohdin,Anna Silnova,Mireia Diez,Lukas Burget
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在过去的几年里,端到端神经日志化已经有了很大的发展,但数据稀缺仍然是进一步改进的主要障碍。WavLM等自监督学习方法在一些下游任务上表现出良好的性能,但它们在说话人日记化方面的应用受到一定限制。在这项工作中,我们探索使用WavLM来缓解神经日志化训练的数据稀缺问题。我们使用与Pyannote相同的流水线,并使用WavLM和Conformer改进了局部端到端神经日志化。在远场AMI,AISHELL-4和AliMeeting数据集上的实验表明,我们的方法大大优于Pyannote基线,并达到了与AMI和AISHELL-4上最先进的结果相当的性能。此外,通过分析不同数据量场景下的系统性能,我们发现WavLM表示比滤波器组特征更能抵抗数据稀缺,从而实现更少的数据饥饿训练策略。此外,我们发现,通常用于训练端到端日志模型的模拟数据在我们的实验中使用WavLM时没有帮助。此外,我们还在最近的CHiME 8 NOTSOFAR-1任务上评估了我们的模型,它比Pyannote基线实现了更好的性能。我们的源代码可在www.example.com上公开获得。摘要:End-to-end neural diarization has evolved considerably over the past few years, but data scarcity is still a major obstacle for further improvements. Self-supervised learning methods such as WavLM have shown promising performance on several downstream tasks, but their application on speaker diarization is somehow limited. In this work, we explore using WavLM to alleviate the problem of data scarcity for neural diarization training. We use the same pipeline as Pyannote and improve the local end-to-end neural diarization with WavLM and Conformer. Experiments on far-field AMI, AISHELL-4, and AliMeeting datasets show that our method substantially outperforms the Pyannote baseline and achieves performance comparable to the state-of-the-art results on AMI and AISHELL-4. In addition, by analyzing the system performance under different data quantity scenarios, we show that WavLM representations are much more robust against data scarcity than filterbank features, enabling less data hungry training strategies. Furthermore, we found that simulated data, usually used to train endto-end diarization models, does not help when using WavLM in our experiments. Additionally, we also evaluate our model on the recent CHiME8 NOTSOFAR-1 task where it achieves better performance than the Pyannote baseline. Our source code is publicly available at https: github.com BUTSpeechFIT DiariZen.
【20】 Language-Queried Target Sound Extraction Without Parallel Training Data
标题: 无需并行训练数据的数据查询目标声音提取
作者:Hao Ma,Zhiyuan Peng,Xu Li,Yukai Li,Mingjie Shao,Qiuqiang Kong,Ju Liu
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:混合查询目标声音提取(TSE)的目的是从混合的语言查询的基础上提取特定的声音。传统的全监督训练方案需要大量注释的并行音频文本数据,这是劳动密集型的。我们引入了一种无语言训练方案,通过利用对比语言音频预训练模型(CLAP)的多模态表示对齐特性,仅需要未标记的音频片段进行TSE模型训练。在无语言训练阶段,使用预训练的CLAP音频编码器对目标音频进行编码,以形成TSE模型的条件嵌入,而在推理期间,用户语言查询由CLAP文本编码器进行编码。由于训练和推理查询之间的模态差距以及在训练期间直接暴露于目标音频的信息泄漏,这种直接的方法面临挑战。为了解决这个问题,我们提出了一个检索增强策略。具体来说,我们使用由大型语言模型(LLM)生成的音频字幕创建嵌入缓存。在训练期间,目标音频嵌入从该缓存中检索文本嵌入以用作条件嵌入,确保训练和推理之间的一致模态并消除信息泄漏。大量的实验结果表明,我们的检索增强的方法实现了一致的和显着的性能改善现有的国家的最先进的更好的泛化能力。摘要:Language-queried target sound extraction (TSE) aims to extract specific sounds from mixtures based on language queries. Traditional fully-supervised training schemes require extensively annotated parallel audio-text data, which are labor-intensive. We introduce a language-free training scheme, requiring only unlabelled audio clips for TSE model training by utilizing the multi-modal representation alignment nature of the contrastive language-audio pre-trained model (CLAP). In a vanilla language-free training stage, target audio is encoded using the pre-trained CLAP audio encoder to form a condition embedding for the TSE model, while during inference, user language queries are encoded by CLAP text encoder. This straightforward approach faces challenges due to the modality gap between training and inference queries and information leakage from direct exposure to target audio during training. To address this, we propose a retrieval-augmented strategy. Specifically, we create an embedding cache using audio captions generated by a large language model (LLM). During training, target audio embeddings retrieve text embeddings from this cache to use as condition embeddings, ensuring consistent modalities between training and inference and eliminating information leakage. Extensive experiment results show that our retrieval-augmented approach achieves consistent and notable performance improvements over existing state-of-the-art with better generalizability.
【21】 Channel Adaptation for Speaker Verification Using Optimal Transport with Pseudo Label
标题: 使用带伪标签的最佳传输进行说话人验证的通道自适应
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Lei Li,Xugang Lu
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:在说话人确认系统中,当训练数据和真实测试语音的统计分布不匹配时,域间隙会降低系统的性能。信道变化是造成这种差距的主要因素,它比其他问题(例如,噪声)。虽然各种领域自适应算法可以用来处理这种领域间隙问题,但大多数算法都不能采用判别式学习进行领域对齐时复杂的分布结构。在本文中,我们提出了一种新的无监督域自适应方法,即,联合部分最优传输与伪标签(JPOT-PL),以缓解信道失配问题。利用分布对齐中最优传输的几何感知距离度量,我们进一步设计了一种基于伪标签的判别学习,其中伪标签可以被视为来自最优耦合的新型软扬声器标签。利用JPOT-PL,以VoxCeleb为基础语料,对SV通道自适应任务进行了实验。实验表明,我们的方法减少EER超过10%,与几个国家的最先进的信道自适应算法相比。摘要:Domain gap often degrades the performance of speaker verification (SV) systems when the statistical distributions of training data and real-world test speech are mismatched. Channel variation, a primary factor causing this gap, is less addressed than other issues (e.g., noise). Although various domain adaptation algorithms could be applied to handle this domain gap problem, most algorithms could not take the complex distribution structure in domain alignment with discriminative learning. In this paper, we propose a novel unsupervised domain adaptation method, i.e., Joint Partial Optimal Transport with Pseudo Label (JPOT-PL), to alleviate the channel mismatch problem. Leveraging the geometric-aware distance metric of optimal transport in distribution alignment, we further design a pseudo label-based discriminative learning where the pseudo label can be regarded as a new type of soft speaker label derived from the optimal coupling. With the JPOT-PL, we carry out experiments on the SV channel adaptation task with VoxCeleb as the basis corpus. Experiments show our method reduces EER by over 10% compared with several state-of-the-art channel adaptation algorithms.
【22】 Integrated Multi-Level Knowledge Distillation for Enhanced Speaker Verification
标题: 用于增强说话人验证的集成多层知识提炼
作者:Wenhao Yang,Jianguo Wei,Wenhuan Lu,Xugang Lu,Lei Li
备注:5 pages, 3 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:知识蒸馏(KD)广泛用于音频任务,如说话人验证(SV),通过将知识从训练有素的大型模型(教师)转移到更小,更紧凑的模型(学生),以提高效率和可移植性。SV的现有KD方法通常反映了图像处理中使用的方法,专注于近似预测概率和隐藏表示。然而,这些方法未能考虑到语音音频的多级时间特性。在本文中,我们提出了一种新的KD方法,集成的多级知识蒸馏(IML-KD),将语音的各种时间尺度特征的知识从教师模型转移到学生模型。在IML-KD中,来自教师模型的时间上下文信息被集成到来自具有各种持续时间的语音片段的新的集成的基于一致性的输入敏感表示中,并且学生模型被训练以推断这些表示,并对输出进行多级对齐。我们在VoxCeleb 1数据集上进行SV实验来评估所提出的方法。实验结果表明,IML-KD显著提高了KD性能,将等错误率(EER)降低了5%。摘要:Knowledge distillation (KD) is widely used in audio tasks, such as speaker verification (SV), by transferring knowledge from a well-trained large model (the teacher) to a smaller, more compact model (the student) for efficiency and portability. Existing KD methods for SV often mirror those used in image processing, focusing on approximating predicted probabilities and hidden representations. However, these methods fail to account for the multi-level temporal properties of speech audio. In this paper, we propose a novel KD method, i.e., Integrated Multi-level Knowledge Distillation (IML-KD), to transfer knowledge of various temporal-scale features of speech from a teacher model to a student model. In the IML-KD, temporal context information from the teacher model is integrated into novel Integrated Gradient-based input-sensitive representations from speech segments with various durations, and the student model is trained to infer these representations with multi-level alignment for the output. We conduct SV experiments on the VoxCeleb1 dataset to evaluate the proposed method. Experimental results demonstrate that IML-KD significantly enhances KD performance, reducing the Equal Error Rate (EER) by 5%.
【23】 Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
标题: 文本提示还不够:用于目标风格音频生成的声音事件增强提示适配器
作者:Chenxu Xiong,Ruibo Fu,Shuchen Shi,Zhengqi Wen,Jianhua Tao,Tao Wang,Chenxing Li,Chunyu Qiang,Yuankun Xie,Xin Qi,Guanjun Li,Zizheng Yang
备注:5 pages, 2 figures, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:目前主流的音频生成方法主要依赖于简单的文本提示,通常无法捕捉多风格音频生成所需的细微差别。为了解决这一限制,提出了声音事件增强提示适配器。与传统的静态全局风格转移不同,该方法通过文本和参考音频之间的交叉注意提取风格嵌入,以实现自适应风格控制。然后利用自适应层规范化来增强模型表达多种风格的能力。此外,声音事件参考风格传输数据集(SERST)被引入用于所提出的目标风格音频生成任务,从而使用文本和音频参考来实现双提示音频生成。实验结果表明,该模型的鲁棒性,实现了最先进的Fr 'echet距离为26.94和KL发散度为1.82,超过了Tango,AudioLDM和AudioGen。此外,所生成的音频显示出与其对应的音频参考的高相似性。演示、代码和数据集都是公开的。摘要:Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To address this limitation, the Sound Event Enhanced Prompt Adapter is proposed. Unlike traditional static global style transfer, this method extracts style embedding through cross-attention between text and reference audio for adaptive style control. Adaptive layer normalization is then utilized to enhance the model's capacity to express multiple styles. Additionally, the Sound Event Reference Style Transfer Dataset (SERST) is introduced for the proposed target style audio generation task, enabling dual-prompt audio generation using both text and audio references. Experimental results demonstrate the robustness of the model, achieving state-of-the-art Fr 'echet Distance of 26.94 and KL Divergence of 1.82, surpassing Tango, AudioLDM, and AudioGen. Furthermore, the generated audio shows high similarity to its corresponding audio reference. The demo, code, and dataset are publicly available.
【24】 E1 TTS: Simple and Fast Non-Autoregressive TTS
标题: E1 TTC:简单快速的非自回归TTC
作者:Zhijun Liu,Shuai Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
链接:点击下载PDF文件
摘要:介绍了一种基于去噪扩散预训练和分布匹配蒸馏的高效非自回归zero-shot文语转换系统--Easy One-Step Text-to-Speech(E1 TTS)。E1 TTS的训练很简单;它不需要文本和音频对之间的显式单调对齐。E1 TTS的推理是有效的,只需要一个神经网络评估每个话语。尽管其采样效率,E1 TTS实现的自然度和说话人相似度可与各种强基线模型相媲美。音频样本可在http: e1tts.github.io 上获得。摘要:This paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is straightforward; it does not require explicit monotonic alignment between the text and audio pairs. The inference of E1 TTS is efficient, requiring only one neural network evaluation for each utterance. Despite its sampling efficiency, E1 TTS achieves naturalness and speaker similarity comparable to various strong baseline models. Audio samples are available at http: e1tts.github.io .
【25】 Wave-U-Mamba: An End-To-End Framework For High-Quality And Efficient Speech Super Resolution
标题: Wave-U-Mamba:一个端到端框架,实现高质量和高效的语音超分辨率
作者:Yongjoon Lee,Chanwoo Kim
备注:This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible
链接:点击下载PDF文件
摘要:语音超分辨率(SSR)是通过恢复丢失的高频分量来增强低分辨率语音信号的任务。传统的方法通常重建对数梅尔特征,随后是在波形域中生成高分辨率语音的声码器。然而,由于对数梅尔特征缺乏相位信息,这可能导致重建阶段期间的性能下降。最近的进展与选择性状态空间模型(SSM)的动机,我们提出了一种方法,称为波U-曼巴,直接在时域进行SSR。在我们的比较研究中,包括WSRGlow,NU-Wave 2和AudioSR等模型,Wave-U-Mamba表现出卓越的性能,在各种低分辨率采样率(从8 kHz到24 kHz)中实现了最低的对数谱距离(LSD)。此外,使用平均意见得分(MOS)进行的主观人类评价表明,我们的方法产生了具有自然和人类品质的SSR。此外,Wave-U-Mamba实现了这些结果,同时在单个A100 GPU上生成高分辨率语音的速度比基线模型快9倍,参数大小不到基线模型的2%。摘要:Speech Super-Resolution (SSR) is a task of enhancing low-resolution speech signals by restoring missing high-frequency components. Conventional approaches typically reconstruct log-mel features, followed by a vocoder that generates high-resolution speech in the waveform domain. However, as log-mel features lack phase information, this can result in performance degradation during the reconstruction phase. Motivated by recent advances with Selective State Spaces Models (SSMs), we propose a method, referred to as Wave-U-Mamba that directly performs SSR in time domain. In our comparative study, including models such as WSRGlow, NU-Wave 2, and AudioSR, Wave-U-Mamba demonstrates superior performance, achieving the lowest Log-Spectral Distance (LSD) across various low-resolution sampling rates, ranging from 8 kHz to 24 kHz. Additionally, subjective human evaluations, scored using Mean Opinion Score (MOS) reveal that our method produces SSR with natural and human-like quality. Furthermore, Wave-U-Mamba achieves these results while generating high-resolution speech over nine times faster than baseline models on a single A100 GPU, with parameter sizes less than 2% of those in the baseline models.
【26】 Improvements of Discriminative Feature Space Training for Anomalous Sound Detection in Unlabeled Conditions
标题: 未标记条件下异常声音检测的鉴别特征空间训练的改进
作者:Takuya Fujimura,Ibuki Kuroyanagi,Tomoki Toda
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:在异常声音检测中,判别方法表现出优越的性能。该方法通过对正常声音元信息标签的分类,构造了一个有区别的特征空间。该特征空间反映了机器声音的差异,并有效地捕获异常声音。然而,当元信息标签丢失时,其性能显着降低。在本文中,我们通过两种方法来提高未标记条件下的判别方法的性能。首先,我们增强了特征提取器,使其在未标记条件下表现更好。我们的增强型特征提取器采用多分辨率频谱图和新的训练策略。其次,我们提出了各种伪标记的方法,有效地训练的特征提取器。实验结果表明,所提出的特征提取器和伪标记方法显着提高性能,在未标记的条件下。摘要:In anomalous sound detection, the discriminative method has demonstrated superior performance. This approach constructs a discriminative feature space through the classification of the meta-information labels for normal sounds. This feature space reflects the differences in machine sounds and effectively captures anomalous sounds. However, its performance significantly degrades when the meta-information labels are missing. In this paper, we improve the performance of a discriminative method under unlabeled conditions by two approaches. First, we enhance the feature extractor to perform better under unlabeled conditions. Our enhanced feature extractor utilizes multi-resolution spectrograms with a new training strategy. Second, we propose various pseudo-labeling methods to effectively train the feature extractor. The experimental evaluations show that the proposed feature extractor and pseudo-labeling methods significantly improve performance under unlabeled conditions.
【27】 Improving Robustness of Diffusion-Based Zero-Shot Speech Synthesis via Stable Formant Generation
标题: 通过稳定的Forces生成提高基于扩散的零激发语音合成的鲁棒性
作者:Changjin Han,Seokgi Lee,Gyuhyeon Nam,Gyeongsu Chae
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:扩散模型在文本到语音(TTS),即使在zero-shot的情况下取得了显着的成功。最近的努力旨在解决推理速度和声音质量之间的权衡,通常被认为是扩散模型的主要缺点。然而,我们发现一个关键的发音错误问题被忽视了。我们的初步研究揭示了不稳定的发音扩散过程中产生的。基于这一观察,我们介绍了StableForm-TTS,一种新的zero-shot语音合成框架,旨在产生强大的发音,同时保持扩散建模的优势。通过开创性地采用源过滤器理论的扩散TTS,我们提出了一个精心设计的架构,稳定的共振峰生成。对未知说话人的实验结果表明,我们的模型在发音准确性和自然度方面优于最先进的方法,具有可比的说话人相似性。此外,随着数据和模型大小的增加,我们的模型表现出有效的可扩展性。摘要:Diffusion models have achieved remarkable success in text-to-speech (TTS), even in zero-shot scenarios. Recent efforts aim to address the trade-off between inference speed and sound quality, often considered the primary drawback of diffusion models. However, we find a critical mispronunciation issue is being overlooked. Our preliminary study reveals the unstable pronunciation resulting from the diffusion process. Based on this observation, we introduce StableForm-TTS, a novel zero-shot speech synthesis framework designed to produce robust pronunciation while maintaining the advantages of diffusion modeling. By pioneering the adoption of source-filter theory in diffusion TTS, we propose an elaborate architecture for stable formant generation. Experimental results on unseen speakers show that our model outperforms the state-of-the-art method in terms of pronunciation accuracy and naturalness, with comparable speaker similarity. Moreover, our model demonstrates effective scalability as both data and model sizes increase.
【28】 ReCLAP: Improving Zero Shot Audio Classification by Describing Sounds
标题: ReCLAP:通过描述声音改进Zero-Shot音频分类
作者:Sreyan Ghosh,Sonal Kumar,Chandra Kiran Reddy Evuru,Oriol Nieto,Ramani Duraiswami,Dinesh Manocha
备注:Code and Checkpoints: this https URL
链接:点击下载PDF文件
摘要:开放词汇音频语言模型,如CLAP,提供了一个很有前途的方法,zero-shot音频分类(ZSAC),使分类与自然语言提示指定的任何任意一组类别。在本文中,我们提出了一个简单而有效的方法来改善ZSAC与CLAP。具体来说,我们从使用带有抽象类别标签的提示的传统方法(例如,器官的声音)到在不同的上下文中使用声音的固有描述性特征来描述声音的提示(例如,风琴低沉而洪亮的音调充满了大教堂。为了实现这一目标,我们首先提出了ReCLAP,这是一个用重写的音频字幕训练的CLAP模型,用于提高对野外声音的理解。这些重写的字幕使用其独特的辨别特征来描述原始字幕中的每个声音事件。ReCLAP在多模态音频文本检索和ZSAC方面都优于所有基线。接下来,为了使用ReCLAP改进zero-shot音频分类,我们提出了提示增强。与采用手写模板提示的传统方法相比,我们为数据集中的每个唯一标签生成自定义提示。这些自定义提示首先在标签中描述声音事件,然后在不同的场景中使用它们。我们提出的方法将ReCLAP在ZSAC上的性能提高了1%-18%,并优于所有基线1% -55%。摘要:Open-vocabulary audio-language models, like CLAP, offer a promising approach for zero-shot audio classification (ZSAC) by enabling classification with any arbitrary set of categories specified with natural language prompts. In this paper, we propose a simple but effective method to improve ZSAC with CLAP. Specifically, we shift from the conventional method of using prompts with abstract category labels (e.g., Sound of an organ) to prompts that describe sounds using their inherent descriptive features in a diverse context (e.g.,The organ's deep and resonant tones filled the cathedral.). To achieve this, we first propose ReCLAP, a CLAP model trained with rewritten audio captions for improved understanding of sounds in the wild. These rewritten captions describe each sound event in the original caption using their unique discriminative characteristics. ReCLAP outperforms all baselines on both multi-modal audio-text retrieval and ZSAC. Next, to improve zero-shot audio classification with ReCLAP, we propose prompt augmentation. In contrast to the traditional method of employing hand-written template prompts, we generate custom prompts for each unique label in the dataset. These custom prompts first describe the sound event in the label and then employ them in diverse scenes. Our proposed method improves ReCLAP's performance on ZSAC by 1%-18% and outperforms all baselines by 1% - 55%.
【29】 Learnings from curating a trustworthy, well-annotated, and useful dataset of disordered English speech
标题: 从策划值得信赖、注释良好且有用的无序英语言语数据集中吸取的教训
作者:Pan-Pan Jiang,Jimmy Tobin,Katrin Tomanek,Robert L. MacDonald,Katie Seaver,Richard Cave,Marilyn Ladewig,Rus Heywood,Jordan R. Green
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:Project Euphonia是Google的一项计划,致力于改善无序语音的自动语音识别(ASR)。该项目的一个中心目标是创建一个大型,高质量和多样化的语音语料库。该报告描述了该项目在数据收集和注释方法方面的最新进展,例如扩大数据库中的说话者多样性,为35万(总计120万)音频记录添加人工审核的转录校正和音频质量标签,并为数据库中超过75%的说话者积累了一套全面的元数据(包括40多个语音特征标签)。我们报告了转录更正对我们的机器学习(ML)研究的影响,评估者之间的变异性评估的混乱的语音模式,以及我们收集语音元数据的理由。我们还考虑了使用自动化现成的注释方法来评估无序语音的局限性。摘要:Project Euphonia, a Google initiative, is dedicated to improving automatic speech recognition (ASR) of disordered speech. A central objective of the project is to create a large, high-quality, and diverse speech corpus. This report describes the project's latest advancements in data collection and annotation methodologies, such as expanding speaker diversity in the database, adding human-reviewed transcript corrections and audio quality tags to 350K (of the 1.2M total) audio recordings, and amassing a comprehensive set of metadata (including more than 40 speech characteristic labels) for over 75 % of the speakers in the database. We report on the impact of transcript corrections on our machine-learning (ML) research, inter-rater variability of assessments of disordered speech patterns, and our rationale for gathering speech metadata. We also consider the limitations of using automated off-the-shelf annotation methods for assessing disordered speech.
【30】 MambaFoley: Foley Sound Generation using Selective State-Space Models
标题: MambaFoley:使用选择性状态空间模型生成Foley声音
作者:Marco Furio Colombo,Francesca Ronchini,Luca Comanducci,Fabio Antonacci
链接:点击下载PDF文件
摘要:深度学习的最新进展导致音频内容生成技术的广泛使用,特别是在各种任务中采用去噪扩散概率模型(DDPM)。其中,Foley声音合成因其在创建多媒体内容的应用中的作用而特别感兴趣。鉴于声音的时间依赖性,设计能够有效处理音频样本的顺序建模的生成模型至关重要。选择性状态空间模型(SSM)最近被提出作为一个有效的替代方案,以前提出的技术,表现出较低的计算复杂度的竞争力的性能。在本文中,我们介绍了MambaFoley,一个基于扩散的模型,据我们所知,是第一个利用最近提出的SSM称为曼巴的福利声音生成任务。为了评估所提出的方法的有效性,我们比较它与一个国家的最先进的福利声音生成模型,使用客观和主观的分析。摘要:Recent advancements in deep learning have led to widespread use of techniques for audio content generation, notably employing Denoising Diffusion Probabilistic Models (DDPM) across various tasks. Among these, Foley Sound Synthesis is of particular interest for its role in applications for the creation of multimedia content. Given the temporal-dependent nature of sound, it is crucial to design generative models that can effectively handle the sequential modeling of audio samples. Selective State Space Models (SSMs) have recently been proposed as a valid alternative to previously proposed techniques, demonstrating competitive performance with lower computational complexity. In this paper, we introduce MambaFoley, a diffusion-based model that, to the best of our knowledge, is the first to leverage the recently proposed SSM known as Mamba for the Foley sound generation task. To evaluate the effectiveness of the proposed method, we compare it with a state-of-the-art Foley sound generative model using both objective and subjective analyses.
【31】 SLiCK: Exploiting Subsequences for Length-Constrained Keyword Spotting
标题: SLiCK:利用子序列进行长度限制的关键字发现
作者:Kumari Nishu,Minsik Cho,Devang Naik
链接:点击下载PDF文件
摘要:在资源受限的边缘设备上进行用户定义的关键字识别是一项挑战。然而,关键字通常受到最大关键字长度的限制,这在以前的作品中基本上没有得到充分利用。我们对关键字长度分布的分析表明,用户定义的关键字定位可以被视为一个长度受限的问题,消除了对可变文本长度聚合的需要。这导致了我们提出的高效关键字识别方法SLiCK(利用子序列进行长度约束关键字识别)。我们还引入了一个语义级匹配方案,以更细的粒度学习音频-文本关系,从而通过增强上下文更有效地区分相似的关键字。在SLiCK中,使用两个模块的多任务学习方法来训练模型:Matcher(话语级匹配任务,新颖的语义级匹配任务)和Encoder(音素识别任务)。所提出的方法改进了Librphrase硬数据集上的基线结果,将AUC从88.52 $增加到94.9 $,并将EER从18.82 $降低到11.1 $。摘要:User-defined keyword spotting on a resource-constrained edge device is challenging. However, keywords are often bounded by a maximum keyword length, which has been largely under-leveraged in prior works. Our analysis of keyword-length distribution shows that user-defined keyword spotting can be treated as a length-constrained problem, eliminating the need for aggregation over variable text length. This leads to our proposed method for efficient keyword spotting, SLiCK (exploiting Subsequences for Length-Constrained Keyword spotting). We further introduce a subsequence-level matching scheme to learn audio-text relations at a finer granularity, thus distinguishing similar-sounding keywords more effectively through enhanced context. In SLiCK, the model is trained with a multi-task learning approach using two modules: Matcher (utterance-level matching task, novel subsequence-level matching task) and Encoder (phoneme recognition task). The proposed method improves the baseline results on Libriphrase hard dataset, increasing AUC from $88.52$ to $94.9$ and reducing EER from $18.82$ to $11.1$.
【32】 MusicLIME: Explainable Multimodal Music Understanding
标题: MusicLIME:可解释的多模式音乐理解
作者:Theodoros Sotirou,Vassilis Lyberatos,Orfeas Menis Mastromichalakis,Giorgos Stamou
备注:GitHub repository: this https URL
链接:点击下载PDF文件
摘要:多模态模型对于音乐理解任务至关重要,因为它们捕捉了音频和歌词之间复杂的相互作用。然而,随着这些模型变得越来越普遍,对可解释性的需求也在增长理解这些系统如何做出决策对于确保公平、减少偏见和培养信任至关重要。在本文中,我们介绍了MusicLIME,一个模型无关的特征重要性解释方法,专为多模态音乐模型。与传统的单峰方法不同,传统的单峰方法单独分析每种模态,而不考虑它们之间的相互作用,通常会导致不完整或误导性的解释,MusicLIME揭示了音频和抒情特征如何相互作用并有助于预测,提供了模型决策的整体视图。此外,我们通过将局部解释聚合为全局解释来增强局部解释,为用户提供更广泛的模型行为视角。通过这项工作,我们有助于提高多模态音乐模型的可解释性,使用户能够做出明智的选择,并促进更公平,公正和透明的音乐理解系统。摘要:Multimodal models are critical for music understanding tasks, as they capture the complex interplay between audio and lyrics. However, as these models become more prevalent, the need for explainability grows-understanding how these systems make decisions is vital for ensuring fairness, reducing bias, and fostering trust. In this paper, we introduce MusicLIME, a model-agnostic feature importance explanation method designed for multimodal music models. Unlike traditional unimodal methods, which analyze each modality separately without considering the interaction between them, often leading to incomplete or misleading explanations, MusicLIME reveals how audio and lyrical features interact and contribute to predictions, providing a holistic view of the model's decision-making. Additionally, we enhance local explanations by aggregating them into global explanations, giving users a broader perspective of model behavior. Through this work, we contribute to improving the interpretability of multimodal music models, empowering users to make informed choices, and fostering more equitable, fair, and transparent music understanding systems.
【33】 2D or not 2D: How Does the Dimensionality of Gesture Representation Affect 3D Co-Speech Gesture Generation?
标题: 2D或不是2D:手势表示的抽象性如何影响3D同声手势生成?
作者:Téo Guichoux,Laure Soulier,Nicolas Obin,Catherine Pelachaud
链接:点击下载PDF文件
摘要:共同语言手势是沟通的基础。最近深度学习技术的出现促进了为嵌入式会话代理创建逼真、同步的协同语音手势。“野外”数据集,通过人体姿势检测技术聚合来自YouTube等平台的视频内容,通过提供与语音对齐的2D骨架序列提供了一个可行的解决方案。提升模型的同时发展使得这些2D序列能够转换为3D手势数据库。然而,重要的是要注意,从2D提取的姿态估计的3D姿态本质上是地面实况的近似,其保持在2D域中。这种区别提出了关于手势表示维度对生成的运动质量的影响的问题-据我们所知,这个话题在很大程度上仍未被探索。我们的研究考察了使用2D或3D关节坐标作为训练数据对语音到手势深度生成模型性能的影响。我们采用提升模型将生成的2D姿势序列转换为3D,并评估直接在3D中创建的手势如何与最初在2D中生成的手势叠加,然后转换为3D。我们使用广泛使用的指标在手势生成领域以及用户研究进行客观的评价,定性评估不同的方法。摘要:Co-speech gestures are fundamental for communication. The advent of recent deep learning techniques has facilitated the creation of lifelike, synchronous co-speech gestures for Embodied Conversational Agents. "In-the-wild" datasets, aggregating video content from platforms like YouTube via human pose detection technologies, provide a feasible solution by offering 2D skeletal sequences aligned with speech. Concurrent developments in lifting models enable the conversion of these 2D sequences into 3D gesture databases. However, it is important to note that the 3D poses estimated from the 2D extracted poses are, in essence, approximations of the ground-truth, which remains in the 2D domain. This distinction raises questions about the impact of gesture representation dimensionality on the quality of generated motions - a topic that, to our knowledge, remains largely unexplored. Our study examines the effect of using either 2D or 3D joint coordinates as training data on the performance of speech-to-gesture deep generative models. We employ a lifting model for converting generated 2D pose sequences into 3D and assess how gestures created directly in 3D stack up against those initially generated in 2D and then converted to 3D. We perform an objective evaluation using widely used metrics in the gesture generation field as well as a user study to qualitatively evaluate the different approaches.
【34】 DreamHead: Learning Spatial-Temporal Correspondence via Hierarchical Diffusion for Audio-driven Talking Head Synthesis
标题: DreamHead:通过分层扩散学习时空对应性,以实现音频驱动的会说话的头部合成
作者:Fa-Ting Hong,Yunfei Liu,Yu Li,Changyin Zhou,Fei Yu,Dan Xu
链接:点击下载PDF文件
摘要:音频驱动的说话头合成努力从提供的音频生成逼真的视频肖像。扩散模型,公认其优越的质量和强大的泛化,已探讨了这项任务。然而,建立一个强大的时间音频线索和相应的空间面部表情与扩散模型之间的对应关系仍然是一个重大的挑战,在说话的头部生成。为了弥合这一差距,我们提出了DreamHead,这是一个分层扩散框架,它可以在不影响模型内在质量和适应性的情况下学习说话头部合成中的时空对应关系。DreamHead学习从音频中预测密集的面部标志作为中间信号,以模拟空间和时间的对应关系。具体地,第一层次的音频到界标扩散首先被设计为在给定音频序列信号的情况下预测时间上平滑且准确的界标序列。然后,第二层次的地标到图像的扩散,进一步提出了产生空间一致的人脸肖像视频,通过建模密集的面部地标和外观之间的空间对应关系。大量的实验表明,提出的DreamHead可以有效地学习时空一致性与设计的分层扩散,并产生高保真音频驱动的多个身份的说话头视频。摘要:Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. However, establishing a robust correspondence between temporal audio cues and corresponding spatial facial expressions with diffusion models remains a significant challenge in talking head generation. To bridge this gap, we present DreamHead, a hierarchical diffusion framework that learns spatial-temporal correspondences in talking head synthesis without compromising the model's intrinsic quality and adaptability.~DreamHead learns to predict dense facial landmarks from audios as intermediate signals to model the spatial and temporal correspondences.~Specifically, a first hierarchy of audio-to-landmark diffusion is first designed to predict temporally smooth and accurate landmark sequences given audio sequence signals. Then, a second hierarchy of landmark-to-image diffusion is further proposed to produce spatially consistent facial portrait videos, by modeling spatial correspondences between the dense facial landmark and appearance. Extensive experiments show that proposed DreamHead can effectively learn spatial-temporal consistency with the designed hierarchical diffusion and produce high-fidelity audio-driven talking head videos for multiple identities.
【35】 Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
标题: 基于扬声器分离HuBERT的自监督音节发现
作者:Ryota Komatsu,Takahiro Shinozaki
备注:Accepted by IEEE SLT 2024
链接:点击下载PDF文件
摘要:自监督语音表示学习对于从未转录音频中提取有意义的特征至关重要。最近的进展突出了从与语言单位相关的特征中导出离散符号的潜力,这使得在不同的任务中进行无文本训练成为可能。特别地,预训练的HuBERT(SD-HuBERT)的重复级自蒸馏在从中间Transformer层提取的潜在语音帧表示内诱导音节结构。在SD-HuBERT中,使用特殊的CLS令牌从语音帧特征通过自注意层积累语音级表示。然而,我们观察到,在CLS令牌中聚集的信息与说话者身份的相关性比与语言内容的相关性更高。为了解决这个问题,我们提出了一个语音只自我监督微调方法,分离音节单位从扬声器信息。我们的方法引入说话人扰动作为数据增强,并采用帧级训练目标来防止CLS令牌聚集语言信息。实验结果表明,我们的方法超过了目前最先进的方法在大多数音节分割和音节单元质量指标Libripeech,强调其有效性,促进音节组织内的语音模型。摘要:Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features correlated with linguistic units, which enables text-less training across diverse tasks. In particular, sentence-level Self-Distillation of the pretrained HuBERT (SD-HuBERT) induces syllabic structures within latent speech frame representations extracted from an intermediate Transformer layer. In SD-HuBERT, sentence-level representation is accumulated from speech frame features through self-attention layers using a special CLS token. However, we observe that the information aggregated in the CLS token correlates more with speaker identity than with linguistic content. To address this, we propose a speech-only self-supervised fine-tuning approach that separates syllabic units from speaker information. Our method introduces speaker perturbation as data augmentation and adopts a frame-level training objective to prevent the CLS token from aggregating paralinguistic information. Experimental results show that our approach surpasses the current state-of-the-art method in most syllable segmentation and syllabic unit quality metrics on Librispeech, underscoring its effectiveness in promoting syllabic organization within speech-only models.
【36】 Optimizing Dysarthria Wake-Up Word Spotting: An End-to-End Approach for SLT 2024 LRDWWS Challenge
标题: 优化构音障碍唤醒词定位:SEARCH 2024 LRDDWWS挑战赛的端到端方法
作者:Shuiyun Liu,Yuxiang Kong,Pengcheng Guo,Weiji Zhuang,Peng Gao,Yujun Wang,Lei Xie
备注:8 pages, Accepted to SLT 2024
链接:点击下载PDF文件
摘要:语音已经成为跨各种应用程序的广泛接受的用户界面。然而,对于患有构音障碍的个体,其言语的固有可变性构成了重大挑战。本文提出了一种端到端的基于预训练的双过滤器构音障碍唤醒词识别(PD-DWS)系统,用于2024年低资源构音障碍唤醒词识别挑战赛。具体来说,我们的系统从两个关键方面提高了性能:音频建模和双滤波器策略。对于音频建模,我们提出了一种基于预训练的data 2 vec 2(d2 v2)的创新2branch-d2 v2模型,该模型可以通过统一的多任务微调范式同时对自动语音识别(ASR)和唤醒单词定位(WWS)任务进行建模。此外,一个双过滤器的策略,以减少错误接受率(FAR),同时保持相同的错误拒绝率(FRR)。实验结果表明,我们的PD-DWS系统实现了0.00321的FAR和0.005的FRR,在测试B评估集上的总得分为0.00821,在挑战中获得第一名。摘要:Speech has emerged as a widely embraced user interface across diverse applications. However, for individuals with dysarthria, the inherent variability in their speech poses significant challenges. This paper presents an end-to-end Pretrain-based Dual-filter Dysarthria Wake-up word Spotting (PD-DWS) system for the SLT 2024 Low-Resource Dysarthria Wake-Up Word Spotting Challenge. Specifically, our system improves performance from two key perspectives: audio modeling and dual-filter strategy. For audio modeling, we propose an innovative 2branch-d2v2 model based on the pre-trained data2vec2 (d2v2), which can simultaneously model automatic speech recognition (ASR) and wake-up word spotting (WWS) tasks through a unified multi-task finetuning paradigm. Additionally, a dual-filter strategy is introduced to reduce the false accept rate (FAR) while maintaining the same false reject rate (FRR). Experimental results demonstrate that our PD-DWS system achieves an FAR of 0.00321 and an FRR of 0.005, with a total score of 0.00821 on the test-B eval set, securing first place in the challenge.
【37】 Speaker Contrastive Learning for Source Speaker Tracing
标题: 用于源说话人追踪的说话人对比学习
作者:Qing Wang,Hongmei Guo,Jian Kang,Mengjie Du,Jie Li,Xiao-Lei Zhang,Lei Xie
备注:7 pages, 2 figures, accepted by SLT
链接:点击下载PDF文件
摘要:作为生物识别技术的一种形式,说话人验证系统的安全性至关重要。然而,SV系统本质上容易受到各种类型的攻击,这些攻击可能会损害其准确性和可靠性。其中一种攻击是语音转换,它通过改变各种声音特征来修改一个人的语音,使其听起来像另一个人。这对SV系统构成了重大威胁。为了解决这个问题,IEEE SLT 2024中的源说话人跟踪挑战旨在识别被操纵的语音信号中的源说话人信息。具体来说,SSTC专注于针对语音转换的源说话人验证,以确定两个转换后的语音样本是否来自同一个源说话人。在这项研究中,我们提出了一种基于说话人对比学习的源说话人跟踪方法来学习转换语音中潜在的源说话人信息。为了学习更多的源说话人相关的表示,我们在嵌入提取器的训练过程中使用说话人对比度损失。这种说话人对比损失有助于在几个干扰说话人嵌入中识别真正的源说话人嵌入,使嵌入提取器能够学习转换后的语音中存在的潜在拥有源说话人信息。实验表明,我们提出的说话人对比学习系统在挑战测试集上达到了最低的EER 16.788%,在挑战中获得第一名。摘要:As a form of biometric authentication technology, the security of speaker verification systems is of utmost importance. However, SV systems are inherently vulnerable to various types of attacks that can compromise their accuracy and reliability. One such attack is voice conversion, which modifies a persons speech to sound like another person by altering various vocal characteristics. This poses a significant threat to SV systems. To address this challenge, the Source Speaker Tracing Challenge in IEEE SLT2024 aims to identify the source speaker information in manipulated speech signals. Specifically, SSTC focuses on source speaker verification against voice conversion to determine whether two converted speech samples originate from the same source speaker. In this study, we propose a speaker contrastive learning-based approach for source speaker tracing to learn the latent source speaker information in converted speech. To learn a more source-speaker-related representation, we employ speaker contrastive loss during the training of the embedding extractor. This speaker contrastive loss helps identify the true source speaker embedding among several distractor speaker embeddings, enabling the embedding extractor to learn the potentially possessing source speaker information present in the converted speech. Experiments demonstrate that our proposed speaker contrastive learning system achieves the lowest EER of 16.788% on the challenge test set, securing first place in the challenge.
【38】 Audio-Driven Reinforcement Learning for Head-Orientation in Naturalistic Environments
标题: 自然环境中用于头部定向的音频驱动强化学习
作者:Wessel Ledder,Yuzhen Qin,Kiki van der Heijden
备注:submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:尽管近年来音频信号处理中的深度强化学习(DRL)方法取得了实质性进展,但在人机交互的背景下,用于导航,凝视控制和头部方向控制等任务的音频驱动DRL很少受到关注。在这里,我们提出了一个音频驱动的DRL框架,在该框架中,我们利用深度Q学习来开发一个自主代理,该代理基于立体声语音记录在声学环境中面向说话者。我们的研究结果表明,当在无回声环境(即没有混响)中对语音段进行训练时,智能体学会了以近乎完美的水平执行任务。自然声学环境中混响的存在影响了代理的性能,尽管代理仍然大大优于基线随机代理。最后,我们量化了所提出的DRL方法在自然声学环境中的泛化程度。我们的实验表明,在中等或高混响环境中训练的代理学习的政策推广到低混响环境,但在消声或低混响环境中训练的代理学习的政策没有推广到中等或高混响环境。总之,这项研究表明了音频驱动的DRL的潜力,如头部方向控制的任务,并强调需要培训策略,使强大的泛化跨环境的真实世界的音频驱动的DRL应用。摘要:Although deep reinforcement learning (DRL) approaches in audio signal processing have seen substantial progress in recent years, audio-driven DRL for tasks such as navigation, gaze control and head-orientation control in the context of human-robot interaction have received little attention. Here, we propose an audio-driven DRL framework in which we utilise deep Q-learning to develop an autonomous agent that orients towards a talker in the acoustic environment based on stereo speech recordings. Our results show that the agent learned to perform the task at a near perfect level when trained on speech segments in anechoic environments (that is, without reverberation). The presence of reverberation in naturalistic acoustic environments affected the agent's performance, although the agent still substantially outperformed a baseline, randomly acting agent. Finally, we quantified the degree of generalization of the proposed DRL approach across naturalistic acoustic environments. Our experiments revealed that policies learned by agents trained on medium or high reverb environments generalized to low reverb environments, but policies learned by agents trained on anechoic or low reverb environments did not generalize to medium or high reverb environments. Taken together, this study demonstrates the potential of audio-driven DRL for tasks such as head-orientation control and highlights the need for training strategies that enable robust generalization across environments for real-world audio-driven DRL applications.
【39】 DiffATR: Diffusion-based Generative Modeling for Audio-Text Retrieval
标题: DiffTAR:音频文本检索的基于扩散的生成建模
作者:Yifei Xin,Xuxin Cheng,Zhihong Zhu,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:现有的音频文本检索(ATR)方法本质上是一种判别模型,其目标是最大化条件似然,表示为p(候选|查询)。然而,这种方法没有考虑内在的数据分布p(查询),导致难以辨别的分布数据。在这项工作中,我们试图通过生成的角度来解决这个约束,并将音频和文本之间的关系建模为它们的联合概率p(候选人,查询)。为此,我们提出了一个基于扩散的ATR框架(DiffATR),该框架将ATR建模为一个迭代过程,逐步从噪声中生成联合分布。在整个训练阶段,DiffATR从生成和判别两个角度进行优化:生成器通过生成损失来细化,而特征提取器则从对比损失中受益,从而将两种方法的优点结合起来。在AudioCaps和Clotho数据集上的实验结果表明,该方法具有较好的性能.值得注意的是,在没有任何改变的情况下,我们的DiffATR在域外检索设置中始终表现出强大的性能。摘要:Existing audio-text retrieval (ATR) methods are essentially discriminative models that aim to maximize the conditional likelihood, represented as p(candidates|query). Nevertheless, this methodology fails to consider the intrinsic data distribution p(query), leading to difficulties in discerning out-of-distribution data. In this work, we attempt to tackle this constraint through a generative perspective and model the relationship between audio and text as their joint probability p(candidates,query). To this end, we present a diffusion-based ATR framework (DiffATR), which models ATR as an iterative procedure that progressively generates joint distribution from noise. Throughout its training phase, DiffATR is optimized from both generative and discriminative viewpoints: the generator is refined through a generation loss, while the feature extractor benefits from a contrastive loss, thus combining the merits of both methodologies. Experiments on the AudioCaps and Clotho datasets with superior performances, verify the effectiveness of our approach. Notably, without any alterations, our DiffATR consistently exhibits strong performance in out-of-domain retrieval settings.
【40】 Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
标题: 通过多任务学习从转录的语音音频中获取发音知识
作者:Siqi Sun,Korin Richmond
备注:5 pages
链接:点击下载PDF文件
摘要:最近的工作已经表明了从传统的基于流水线的文本到语音(TTS)前端引导集成序列到序列(Seq2Seq)语言前端的可行性和好处。为了克服自举训练数据的固定词汇覆盖,先前的工作已经提出利用容易访问的转录语音音频作为用于获取未覆盖单词的新发音知识的额外训练源,其依赖于辅助ASR模型作为繁琐的实现流程的一部分。在这项工作中,我们提出了一种替代方法,利用转录的语音音频作为额外的训练源,基于多任务学习(MTL)。实验表明,与基线Seq2Seq前端相比,对于转录语音音频中专门覆盖的单词类型,提出的基于MTL的方法将PER从2.5%降低到1.6%,实现了与之前方法相似的性能,但实现流程要简单得多。摘要:Recent work has shown the feasibility and benefit of bootstrapping an integrated sequence-to-sequence (Seq2Seq) linguistic frontend from a traditional pipeline-based frontend for text-to-speech (TTS). To overcome the fixed lexical coverage of bootstrapping training data, previous work has proposed to leverage easily accessible transcribed speech audio as an additional training source for acquiring novel pronunciation knowledge for uncovered words, which relies on an auxiliary ASR model as part of a cumbersome implementation flow. In this work, we propose an alternative method to leverage transcribed speech audio as an additional training source, based on multi-task learning (MTL). Experiments show that, compared to a baseline Seq2Seq frontend, the proposed MTL-based method reduces PER from 2.5% to 1.6% for those word types covered exclusively in transcribed speech audio, achieving a similar performance to the previous method but with a much simpler implementation flow.
【41】 Constructing a Singing Style Caption Dataset
标题: 构建歌唱风格字幕数据集
作者:Hyunjong Ok,Jaeho Lee
备注:Preprint
链接:点击下载PDF文件
摘要:歌唱嗓音的合成与转换已成为嗓音生成的重要子领域,对非条件生成提出了更高的要求。与普通的语音数据不同,生成歌声需要了解各种相关的声乐和音乐特征,例如歌手的音调或情感表达。然而,现有的用于语音生成的开源音频文本数据集往往只捕获非常有限的属性范围,通常缺少音频的音乐特征。为了填补这一空白,我们引入了S2 Cap,这是一个具有不同属性集的音频-文本对数据集。S2 Cap由成对的文本提示和音乐音频样本组成,具有广泛的声乐和音乐属性,包括音高,音量,节奏,情绪,歌手的性别和年龄,音乐类型和情感表达。利用S2 Cap,我们提出了一个有效的新的基线算法的演唱风格的字幕。演唱风格字幕是一个相对于语音生成的任务,生成文本描述的声音特征,这是我们首先提出的。首先,为了减轻音频编码器和文本解码器之间的不对齐,我们提出了一种名为CRESCENDO的新机制,该机制利用正对相似性学习来同步预训练音频编码器的嵌入空间,以获得与文本编码器相似的嵌入。我们还使用歌手的声音来监督模型,歌手的声音被伴奏分离。这种监督允许模型更准确地捕捉声音特征,从而改进演唱风格字幕,更好地反映歌手的风格。数据集和代码可以在 bulurl{https: github.com HJ-Ok S2cap}上找到。摘要:Singing voice synthesis and conversion have emerged as significant subdomains of voice generation, leading to much demands on prompt-conditioned generation. Unlike common voice data, generating a singing voice requires an understanding of various associated vocal and musical characteristics, such as the vocal tone of the singer or emotional expressions. However, existing open-source audio-text datasets for voice generation tend to capture only a very limited range of attributes, often missing musical characteristics of the audio. To fill this gap, we introduce S2Cap, an audio-text pair dataset with a diverse set of attributes. S2Cap consists of pairs of textual prompts and music audio samples with a wide range of vocal and musical attributes, including pitch, volume, tempo, mood, singer's gender and age, and musical genre and emotional expression. Utilizing S2Cap, we suggest an effective novel baseline algorithm for singing style captioning. Singing style captioning is a relative task to voice generation that generates text descriptions of vocal characteristics, which we first suggested. First, to mitigate the misalignment between the audio encoder and the text decoder, we present a novel mechanism called CRESCENDO, which utilizes positive-pair similarity learning to synchronize the embedding spaces of a pretrained audio encoder to get similar embeddings with a text encoder. We additionally supervise the model using the singer's voice, which is demixed by the accompaniment. This supervision allows the model to more accurately capture vocal characteristics, leading to improved singing style captions that better reflect the style of the singer. The dataset and the codes are available at bulurl{https: github.com HJ-Ok S2cap}.
【42】 Efficient Video to Audio Mapper with Visual Scene Detection
标题: 具有视觉场景检测的高效视频到音频映射器
作者:Mingjing Yi,Ming Li
链接:点击下载PDF文件
摘要:视频到音频(V2A)生成的目的是在给定无声视频输入的情况下产生对应的音频。这项任务是特别具有挑战性的,由于跨模态和连续性的视听功能所涉及的。最近的作品在弥合视频和音频之间的域差距方面取得了重大进展,生成与视频内容语义一致的音频。然而,这些方法的一个关键限制是它们不能有效地识别和处理视频中的多个场景,在这种情况下通常导致次优的音频生成。在本文中,我们首先重新实现了一个最先进的V2A模型,稍微修改了轻量级架构,实现了优于基线的结果。然后,我们提出了一个改进的V2A模型,它结合了场景检测器,以解决多个视觉场景之间切换的挑战。在VGGSound上的结果表明,我们的模型可以识别和处理视频中的多个场景,并在保真度和相关性方面取得了优于基线的性能。摘要:Video-to-audio (V2A) generation aims to produce corresponding audio given silent video inputs. This task is particularly challenging due to the cross-modality and sequential nature of the audio-visual features involved. Recent works have made significant progress in bridging the domain gap between video and audio, generating audio that is semantically aligned with the video content. However, a critical limitation of these approaches is their inability to effectively recognize and handle multiple scenes within a video, often leading to suboptimal audio generation in such cases. In this paper, we first reimplement a state-of-the-art V2A model with a slightly modified light-weight architecture, achieving results that outperform the baseline. We then propose an improved V2A model that incorporates a scene detector to address the challenge of switching between multiple visual scenes. Results on VGGSound show that our model can recognize and handle multiple scenes within a video and achieve superior performance against the baseline for both fidelity and relevance.
【43】 Large Language Model Based Generative Error Correction: A Challenge and Baselines forSpeech Recognition, Speaker Tagging, and Emotion Recognition
标题: 基于大语言模型的生成式错误纠正:语音识别、说话人标记和情感识别的挑战和基线
作者:Chao-Han Huck Yang,Taejin Park,Yuan Gong,Yuanchao Li,Zhehuai Chen,Yen-Ting Lin,Chen Chen,Yuchen Hu,Kunal Dhawan,Piotr Żelasko,Chao Zhang,Yun-Nung Chen,Yu Tsao,Jagadeesh Balam,Boris Ginsburg,Sabato Marco Siniscalchi,Eng Siong Chng,Peter Bell,Catherine Lai,Shinji Watanabe,Andreas Stolcke
备注:IEEE SLT 2024. The initial draft version has been done in December 2023. Post-ASR Text Processing and Understanding Community: this https URL
链接:点击下载PDF文件
摘要:鉴于生成式人工智能技术的最新进展,一个关键问题是大型语言模型(LLM)如何使用来自冻结的预训练自动语音识别(ASR)模型的文本解码结果来增强声学建模任务。为了探索语音处理语言建模的新功能,我们引入了生成式语音转录纠错(GenSEC)挑战。这个挑战包括三个后ASR语言建模任务:(i)后ASR转录校正,(ii)说话人标记,以及(iii)情感识别。这些任务旨在模拟未来基于LLM的代理处理基于语音的界面,同时通过利用开放的预训练语言模型或基于代理的API保持对广泛受众的访问。我们还讨论了从基线评估的见解,以及设计未来的评估经验教训。摘要:Given recent advances in generative AI technology, a key question is how large language models (LLMs) can enhance acoustic modeling tasks using text decoding results from a frozen, pretrained automatic speech recognition (ASR) model. To explore new capabilities in language modeling for speech processing, we introduce the generative speech transcription error correction (GenSEC) challenge. This challenge comprises three post-ASR language modeling tasks: (i) post-ASR transcription correction, (ii) speaker tagging, and (iii) emotion recognition. These tasks aim to emulate future LLM-based agents handling voice-based interfaces while remaining accessible to a broad audience by utilizing open pretrained language models or agent-based APIs. We also discuss insights from baseline evaluations, as well as lessons learned for designing future evaluations.
【44】 Self-supervised Learning for Acoustic Few-Shot Classification
标题: 声学Few-Shot分类的自我监督学习
作者:Jingyong Liang,Bernd Meyer,Issac Ning Lee,Thanh-Toan Do
链接:点击下载PDF文件
摘要:标记数据有限,自我监督学习是减少标记要求的最重要方法之一。虽然它已经在图像领域得到了广泛的探索,但到目前为止,它在声学领域还没有得到同样多的关注。然而,减少标签是许多声学应用的关键要求。特别是在生物声学中,很少有足够的标签用于完全监督学习。这导致了声学识别器的广泛使用,这些识别器已经在生物声学任务的不相关数据上进行了预训练。我们认为,在实际任务数据上进行训练,并将自我监督的预训练与Few-Shot分类相结合是一种优越的方法,即使只有少数标签可用,也能够提供高精度。为此,我们引入并评估了一种新的架构,该架构将基于CNN的预处理与基于状态空间模型(SSM)的特征提取相结合。这种组合的动机是,基于CNN的网络很难有效地捕获时间信息,这对于分类声学信号至关重要。另一方面,SSM,特别是S4和Mamba,已被证明具有捕获序列数据中的长程依赖性的出色能力。我们使用实际任务数据上的对比学习和随后的微调来预训练这个架构,这些数据是非常少量的标记数据。我们评估的性能,这个建议的架构($n$杆,$n$类)分类标准的基准以及现实世界的数据。我们的评估表明,它优于国家的最先进的架构上的Few-Shot分类问题。摘要:Labelled data are limited and self-supervised learning is one of the most important approaches for reducing labelling requirements. While it has been extensively explored in the image domain, it has so far not received the same amount of attention in the acoustic domain. Yet, reducing labelling is a key requirement for many acoustic applications. Specifically in bioacoustic, there are rarely sufficient labels for fully supervised learning available. This has led to the widespread use of acoustic recognisers that have been pre-trained on unrelated data for bioacoustic tasks. We posit that training on the actual task data and combining self-supervised pre-training with few-shot classification is a superior approach that has the ability to deliver high accuracy even when only a few labels are available. To this end, we introduce and evaluate a new architecture that combines CNN-based preprocessing with feature extraction based on state space models (SSMs). This combination is motivated by the fact that CNN-based networks alone struggle to capture temporal information effectively, which is crucial for classifying acoustic signals. SSMs, specifically S4 and Mamba, on the other hand, have been shown to have an excellent ability to capture long-range dependencies in sequence data. We pre-train this architecture using contrastive learning on the actual task data and subsequent fine-tuning with an extremely small amount of labelled data. We evaluate the performance of this proposed architecture for ($n$-shot, $n$-class) classification on standard benchmarks as well as real-world data. Our evaluation shows that it outperforms state-of-the-art architectures on the few-shot classification problem.
【45】 Compositional Audio Representation Learning
标题: 合成音频表示学习
作者:Sripathi Sridhar,Mark Cartwright
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:人类的听觉感知本质上是合成的--我们从具有多个声音事件的听觉场景中识别听觉流。然而,这样的听觉场景通常使用不解开组成声源的剪辑级表示来表示。在这项工作中,我们学习了以源为中心的音频表示,其中每个声源都使用嵌入在音频表示中的不同的、分离的源来表示。我们提出了两种新的方法来学习以源为中心的音频表示:分类指导的监督模型和特征重建指导的无监督模型,这两种方法都优于基线。我们彻底评估这两种方法的设计选择使用音频分类任务。我们发现,监督有利于学习以源为中心的表示,并且重建音频特征比重建频谱图更有用,以学习无监督的以源为中心的表示。利用以源为中心的模型可以帮助释放机器听力中更大的可解释性和更灵活的解码潜力。摘要:Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not disentangle the constituent sound sources. In this work, we learn source-centric audio representations where each sound source is represented using a distinct, disentangled source embedding in the audio representation. We propose two novel approaches to learning source-centric audio representations: a supervised model guided by classification and an unsupervised model guided by feature reconstruction, both of which outperform the baselines. We thoroughly evaluate the design choices of both approaches using an audio classification task. We find that supervision is beneficial to learn source-centric representations, and that reconstructing audio features is more useful than reconstructing spectrograms to learn unsupervised source-centric representations. Leveraging source-centric models can help unlock the potential of greater interpretability and more flexible decoding in machine listening.
【46】 Integrating Audio Narrations to Strengthen Domain Generalization in Multimodal First-Person Action Recognition
标题: 整合音频叙述加强多模式第一人称动作识别中的领域概括
作者:Cagri Gungor,Adriana Kovashka
链接:点击下载PDF文件
摘要:由于可穿戴摄像头的广泛使用,第一人称活动识别正在迅速发展,但面临着不同环境中域转移的挑战,例如不同的对象或背景场景。我们提出了一个多模态框架,通过整合运动,音频和外观特征,提高域泛化。主要贡献包括分析音频和运动特征对域转移的弹性,使用音频叙述来增强音频-文本对齐,以及在音频和视觉叙述之间应用一致性评级来优化音频在训练过程中识别的影响。我们的方法在ARGO 1 M数据集上实现了最先进的性能,有效地概括了看不见的场景和位置。摘要:First-person activity recognition is rapidly growing due to the widespread use of wearable cameras but faces challenges from domain shifts across different environments, such as varying objects or background scenes. We propose a multimodal framework that improves domain generalization by integrating motion, audio, and appearance features. Key contributions include analyzing the resilience of audio and motion features to domain shifts, using audio narrations for enhanced audio-text alignment, and applying consistency ratings between audio and visual narrations to optimize the impact of audio in recognition during training. Our approach achieves state-of-the-art performance on the ARGO1M dataset, effectively generalizing across unseen scenarios and locations.
【47】 A Survey of Foundation Models for Music Understanding
标题: 音乐理解的基础模型综述
作者:Wenjun Li,Ying Cai,Ziyang Wu,Wenyi Zhang,Yifan Chen,Rundong Qi,Mengqi Dong,Peigen Chen,Xiao Dong,Fenghao Shi,Lei Guo,Junwei Han,Bao Ge,Tianming Liu,Lin Gan,Tuo Zhang
备注:20 pages, 2 figures
链接:点击下载PDF文件
摘要:音乐在日常生活中至关重要,满足情感和娱乐需求,并将我们个人,社会和文化联系起来。更好地理解音乐可以增强我们的情感,认知技能和文化联系。人工智能(AI)的快速发展引入了分析音乐的新方法,旨在复制人类对音乐的理解并提供相关服务。传统的模型主要关注音频特征和简单的任务,而最近发展起来的大型语言模型(LLM)和基础模型(FM)通过整合语义信息和展示强大的推理能力,在各个领域表现出色,可以捕获复杂的音乐特征和模式,将音乐与语言结合起来,并包含丰富的音乐,情感和心理知识。因此,它们有潜力从语义的角度处理复杂的音乐理解任务,产生更接近人类感知的输出。据我们所知,这项工作是人工智能技术和音乐理解交叉的早期评论之一。我们调查,分析,并测试最近的大型音乐基础模型在他们的音乐理解能力。我们还讨论了它们的局限性,并提出了未来可能的发展方向,为该领域的研究人员提供了见解。摘要:Music is essential in daily life, fulfilling emotional and entertainment needs, and connecting us personally, socially, and culturally. A better understanding of music can enhance our emotions, cognitive skills, and cultural connections. The rapid advancement of artificial intelligence (AI) has introduced new ways to analyze music, aiming to replicate human understanding of music and provide related services. While the traditional models focused on audio features and simple tasks, the recent development of large language models (LLMs) and foundation models (FMs), which excel in various fields by integrating semantic information and demonstrating strong reasoning abilities, could capture complex musical features and patterns, integrate music with language and incorporate rich musical, emotional and psychological knowledge. Therefore, they have the potential in handling complex music understanding tasks from a semantic perspective, producing outputs closer to human perception. This work, to our best knowledge, is one of the early reviews of the intersection of AI techniques and music understanding. We investigated, analyzed, and tested recent large-scale music foundation models in respect of their music comprehension abilities. We also discussed their limitations and proposed possible future directions, offering insights for researchers in this field.
【48】 On the effectiveness of enrollment speech augmentation for Target Speaker Extraction
标题: 关于目标说话人提取的注册语音增强的有效性
作者:Junjie Li,Ke Zhang,Shuai Wang,Haizhou Li,Man-Wai Mak,Kong Aik Lee
备注:Accepted by SLT2024
链接:点击下载PDF文件
摘要:深度学习技术显著提高了目标说话人提取(TSE)任务的性能。为了提高这些算法在训练数据不足时的泛化能力和鲁棒性,数据增强是一种常用的技术。不同于典型的数据增强应用于语音混合,这项工作彻底调查的有效性,增强注册语音空间。我们发现,对于预训练和联合优化的扬声器编码器,直接增强注册语音会导致一致的性能改善。除了传统的方法,如噪声和混响添加,我们提出了一种新的增强方法称为自估计语音增强(SSA)。Libri2Mix测试集上的实验结果表明,我们提出的方法可以实现高达2.5 dB的改善。摘要:Deep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a commonly adopted technique. Unlike typical data augmentation applied to speech mixtures, this work thoroughly investigates the effectiveness of augmenting the enrollment speech space. We found that for both pretrained and jointly optimized speaker encoders, directly augmenting the enrollment speech leads to consistent performance improvement. In addition to conventional methods such as noise and reverberation addition, we propose a novel augmentation method called self-estimated speech augmentation (SSA). Experimental results on the Libri2Mix test set show that our proposed method can achieve an improvement of up to 2.5 dB.
【49】 ASR Error Correction using Large Language Models
标题: 使用大型语言模型的ASB错误纠正
作者:Rao Ma,Mengjie Qian,Mark Gales,Kate Knill
备注:Submitted to IEEE Transactions on Audio, Speech and Language Processing
链接:点击下载PDF文件
摘要:纠错(EC)模型在改进自动语音识别(ASR)译文、提高译文的可读性和质量方面起着至关重要的作用。在不需要访问底层代码或模型权重的情况下,EC可以提高性能并为黑盒ASR系统提供域自适应。这项工作研究了使用大型语言模型(LLM)在不同的场景中进行纠错。1-最佳ASR假设通常用作EC模型的输入。我们建议使用ASR N-最佳列表来构建高性能的EC模型,该列表应该为校正过程提供更多的上下文信息。此外,标准EC模型的生成过程是不受限制的,因为可以生成任何输出序列。对于某些场景(例如看不见的域),这种灵活性可能会影响性能。为了解决这个问题,我们引入了一种基于N-最佳列表或ASR格的约束解码方法。最后,大多数EC模型都是针对特定的ASR系统进行训练的,每当底层ASR系统发生变化时,都需要重新训练。本文探讨了EC模型对不同ASR系统的输出进行操作的能力。该概念进一步扩展到使用LLM(诸如ChatGPT)的zero-shot纠错。在三个标准数据集上的实验证明了我们提出的方法对换能器和基于注意力的编码器-解码器ASR系统的有效性。此外,所提出的方法可以作为一种有效的方法,模型集成。摘要:Error correction (EC) models play a crucial role in refining Automatic Speech Recognition (ASR) transcriptions, enhancing the readability and quality of transcriptions. Without requiring access to the underlying code or model weights, EC can improve performance and provide domain adaptation for black-box ASR systems. This work investigates the use of large language models (LLMs) for error correction across diverse scenarios. 1-best ASR hypotheses are commonly used as the input to EC models. We propose building high-performance EC models using ASR N-best lists which should provide more contextual information for the correction process. Additionally, the generation process of a standard EC model is unrestricted in the sense that any output sequence can be generated. For some scenarios, such as unseen domains, this flexibility may impact performance. To address this, we introduce a constrained decoding approach based on the N-best list or an ASR lattice. Finally, most EC models are trained for a specific ASR system requiring retraining whenever the underlying ASR system is changed. This paper explores the ability of EC models to operate on the output of different ASR systems. This concept is further extended to zero-shot error correction using LLMs, such as ChatGPT. Experiments on three standard datasets demonstrate the efficacy of our proposed methods for both Transducer and attention-based encoder-decoder ASR systems. In addition, the proposed method can serve as an effective method for model ensembling.
【50】 Multi-Microphone and Multi-Modal Emotion Recognition in Reverbrant Enviroment
标题: 可逆环境中的多麦克风和多模式情感识别
作者:Ohad Cohen,Gershon Hazan,Sharon Gannot
链接:点击下载PDF文件
摘要:本文提出了一种多模态情感识别(MER)系统,旨在提高情感识别的准确性,在具有挑战性的声学条件。我们的方法结合了修改和扩展的层次令牌语义音频Transformer(HTS-AT)的多通道音频处理与R(2+1)D卷积神经网络(CNN)模型的视频分析。我们评估我们提出的方法上的混响版本的瑞尔森视听数据库的情感语音和歌曲(RAVDESS)数据集使用合成和真实世界的房间脉冲响应(RIR)。我们的研究结果表明,整合音频和视频模态产生优越的性能相比,单模态的方法,特别是在具有挑战性的声学条件。此外,我们表明,多模态(视听)的方法,利用多个麦克风优于其单麦克风对应。摘要:This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-semantic Audio Transformer (HTS-AT) for multi-channel audio processing with an R(2+1)D Convolutional Neural Networks (CNN) model for video analysis. We evaluate our proposed method on a reverberated version of the Ryerson audio-visual database of emotional speech and song (RAVDESS) dataset using synthetic and real-world Room Impulse Responsess (RIRs). Our results demonstrate that integrating audio and video modalities yields superior performance compared to uni-modal approaches, especially in challenging acoustic conditions. Moreover, we show that the multimodal (audiovisual) approach that utilizes multiple microphones outperforms its single-microphone counterpart.
【51】 Explaining Deep Learning Embeddings for Speech Emotion Recognition by Predicting Interpretable Acoustic Features
标题: 通过预测可解释的声学特征来解释语音情感识别的深度学习嵌入
作者:Satvik Dixit,Daniel M. Low,Gasser Elbanna,Fabio Catania,Satrajit S. Ghosh
链接:点击下载PDF文件
摘要:预训练的深度学习嵌入在语音情感识别(SER)中一直表现出优于手工制作的声学特征的性能。然而,与具有明确物理意义的声学特征不同,这些嵌入缺乏明确的可解释性。解释这些嵌入对于在医疗保健和安全应用中建立信任以及推进对其中编码的声学信息的科学理解至关重要。本文提出了一种改进的探测方法来解释SER空间中的深度学习嵌入。我们预测可解释的声学特征(例如,f0,响度)从(i)嵌入的完整集合和(ii)被识别为对于预测每个情感最重要的嵌入维度的子集。如果最重要维度的子集比所有维度更好地预测给定的情感,并且还更准确地预测特定的声学特征,则我们推断这些声学特征对于给定任务的嵌入模型是重要的。我们使用WavLM嵌入和eGeMAPS声学特征作为音频表示进行了实验,将我们的方法应用于RAVDESS和SAVEE情感语音数据集。基于此评估,我们证明了能量,频率,频谱和时间类别的声学特征提供了减少的信息,以SER在该顺序,演示了实用程序的探测分类器方法相关的嵌入到可解释的声学特征。摘要:Pre-trained deep learning embeddings have consistently shown superior performance over handcrafted acoustic features in speech emotion recognition (SER). However, unlike acoustic features with clear physical meaning, these embeddings lack clear interpretability. Explaining these embeddings is crucial for building trust in healthcare and security applications and advancing the scientific understanding of the acoustic information that is encoded in them. This paper proposes a modified probing approach to explain deep learning embeddings in the SER space. We predict interpretable acoustic features (e.g., f0, loudness) from (i) the complete set of embeddings and (ii) a subset of the embedding dimensions identified as most important for predicting each emotion. If the subset of the most important dimensions better predicts a given emotion than all dimensions and also predicts specific acoustic features more accurately, we infer those acoustic features are important for the embedding model for the given task. We conducted experiments using the WavLM embeddings and eGeMAPS acoustic features as audio representations, applying our method to the RAVDESS and SAVEE emotional speech datasets. Based on this evaluation, we demonstrate that Energy, Frequency, Spectral, and Temporal categories of acoustic features provide diminishing information to SER in that order, demonstrating the utility of the probing classifier method to relate embeddings to interpretable acoustic features.
【52】 ESPnet-EZ: Python-only ESPnet for Easy Fine-tuning and Integration
标题: ESPnet-ZZ:仅使用Python的ESPnet,易于微调和集成
作者:Masao Someki,Kwanghee Choi,Siddhant Arora,William Chen,Samuele Cornell,Jionghao Han,Yifan Peng,Jiatong Shi,Vaibhav Srivastav,Shinji Watanabe
备注:Accepted to SLT 2024
链接:点击下载PDF文件
摘要:我们介绍ESPnet-EZ,一个开源的语音处理工具包ESPnet的扩展,旨在快速,方便地开发语音模型。ESPnet-EZ专注于两个主要方面:(i)在各种任务上对现有ESPnet模型进行简单的微调和推理,以及(ii)与流行的深度神经网络框架(如PyTorch-Lightning,Hugging Face Transformers和datasets以及Lhotse)轻松集成。通过将继承自Kaldi的ESPnet设计选择替换为仅使用Python的无Bash接口,我们大大减少了构建、调试和使用新模型所需的工作量。例如,为了微调语音基础模型,与ESPnet相比,ESPnet-EZ将新编写的代码数量减少了2.7倍,相关代码数量减少了6.7倍,同时大大减少了Bash脚本依赖性。ESPnet-EZ的代码库是公开的。摘要:We introduce ESPnet-EZ, an extension of the open-source speech processing toolkit ESPnet, aimed at quick and easy development of speech models. ESPnet-EZ focuses on two major aspects: (i) easy fine-tuning and inference of existing ESPnet models on various tasks and (ii) easy integration with popular deep neural network frameworks such as PyTorch-Lightning, Hugging Face transformers and datasets, and Lhotse. By replacing ESPnet design choices inherited from Kaldi with a Python-only, Bash-free interface, we dramatically reduce the effort required to build, debug, and use a new model. For example, to fine-tune a speech foundation model, ESPnet-EZ, compared to ESPnet, reduces the number of newly written code by 2.7x and the amount of dependent code by 6.7x while dramatically reducing the Bash script dependencies. The codebase of ESPnet-EZ is publicly available.
【53】 Prevailing Research Areas for Music AI in the Era of Foundation Models
标题: 基础模型时代音乐人工智能的主流研究领域
作者:Megan Wei,Mateusz Modrzejewski,Aswin Sivaraman,Dorien Herremans
链接:点击下载PDF文件
摘要:随着基础模型研究的最新进展,在过去几年中,生成音乐AI应用程序激增。随着人工智能生成或人工智能增强音乐的想法变得越来越主流,音乐人工智能社区的许多研究人员可能想知道还剩下什么研究途径。关于音乐生成模型,我们概述了目前的研究领域,有显着的探索空间。首先,我们提出的问题,这些生成模型的基本表示和调查的可解释性的方法。接下来,我们将讨论音乐数据集的现状及其局限性。然后,我们概述了不同的生成模型,评估这些模型的形式,以及它们的计算约束 限制。随后,我们强调了这些生成模型的应用程序扩展到多种形式和艺术家的工作流程以及音乐教育系统的集成。最后,我们调查了生成音乐的潜在版权影响,并讨论了保护音乐家权利的策略。虽然这并不意味着是详尽的,但我们的调查引起了人们对音乐基金会模型所支持的各种研究方向的关注。摘要:In tandem with the recent advancements in foundation model research, there has been a surge of generative music AI applications within the past few years. As the idea of AI-generated or AI-augmented music becomes more mainstream, many researchers in the music AI community may be wondering what avenues of research are left. With regards to music generative models, we outline the current areas of research with significant room for exploration. Firstly, we pose the question of foundational representation of these generative models and investigate approaches towards explainability. Next, we discuss the current state of music datasets and their limitations. We then overview different generative models, forms of evaluating these models, and their computational constraints limitations. Subsequently, we highlight applications of these generative models towards extensions to multiple modalities and integration with artists' workflow as well as music education systems. Finally, we survey the potential copyright implications of generative music and discuss strategies for protecting the rights of musicians. While it is not meant to be exhaustive, our survey calls to attention a variety of research directions enabled by music foundation models.
【54】 Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
标题: 联合语义知识提取和掩蔽声学建模用于提高可理解度的全频段语音恢复
作者:Xiaoyu Liu,Xu Li,Joan Serrà,Santiago Pascual
备注:Demo link this https URL
链接:点击下载PDF文件
摘要:语音恢复的目标是在考虑各种失真的情况下,恢复具有高质量和可懂度的全频带语音。MaskSR是最近提出的用于此任务的生成模型。与其他同类模型一样,MaskSR达到了高质量,但正如我们所展示的那样,可理解性可以大大提高。我们通过使用预先训练的自监督教师模型,通过预测目标语音的语义表示来提升MaskSR的语音编码器组件。然后,掩蔽的语言模型的条件下学习的语义特征,以预测声学令牌编码的目标语音的低级别的频谱细节。我们发现,在相同的MaskSR模型容量和推理时间下,所提出的模型MaskSR2显着降低了单词错误率,这是可理解性的典型指标。MaskSR2在提供卓越质量的同时,还实现了与其他型号相比具有竞争力的字错误率。消融研究显示了各种语义表征的有效性。摘要:Speech restoration aims at restoring full-band speech with high quality and intelligibility, considering a diverse set of distortions. MaskSR is a recently proposed generative model for this task. As other models of its kind, MaskSR attains high quality but, as we show, intelligibility can be substantially improved. We do so by boosting the speech encoder component of MaskSR with predictions of semantic representations of the target speech, using a pre-trained self-supervised teacher model. Then, a masked language model is conditioned on the learned semantic features to predict acoustic tokens that encode low level spectral details of the target speech. We show that, with the same MaskSR model capacity and inference time, the proposed model, MaskSR2, significantly reduces the word error rate, a typical metric for intelligibility. MaskSR2 also achieves competitive word error rate among other models, while providing superior quality. An ablation study shows the effectiveness of various semantic representations.
【55】 MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
标题: MacST:通过文本音译进行多口音语音合成以实现口音转换
作者:Sho Inoue,Shuai Wang,Wanxing Wang,Pengcheng Zhu,Mengxiao Bi,Haizhou Li
备注:Project page with Speech Demo: this https URL
链接:点击下载PDF文件
摘要:在口音语音转换或口音转换中,我们寻求在语音中将口音彼此转换,同时保留说话者身份和语义内容。在这项研究中,我们制定了一个新的方法来创建多口音的语音样本,从而对口音的语音样本由同一扬声器,通过文本音译训练口音转换系统。我们首先使用大型语言模型(LLM)生成音译文本,然后将其输入多语言TTS模型以合成带口音的英语语音。作为参考系统,我们在合成并行语料库上建立了一个序列到序列模型来进行口音转换。我们验证了所提出的方法为母语和非母语的英语。主观和客观的评价进一步验证了我们的数据集在口音转换研究中的有效性。摘要:In accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset's effectiveness in accent conversion studies.
【56】 Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
标题: 儿童与成人二元互动中的自我中心说话者分类:从感知到计算建模
作者:Tiantian Feng,Anfeng Xu,Xuan Shi,Somer Bishop,Shrikanth Narayanan
备注:pre-print under review
链接:点击下载PDF文件
摘要:自闭症谱系障碍(ASD)是一种神经发育状况,其特征在于社交,重复行为和感觉处理方面的挑战。ASD的一个重要研究领域是评估儿童在治疗期间随时间的行为变化。具有此目标的标准协议是BOSCC,其涉及儿童和执行预定义的一组活动的临床医生之间的二元交互。理解儿童在这些互动中的行为的一个基本方面是自动语音理解,特别是识别谁在说话以及何时说话。在这方面的传统方法严重依赖于从旁观者的角度记录的语音样本,并且对以自我为中心的语音建模的研究有限。在这项研究中,我们设计了一个实验,从自我中心的角度使用可穿戴传感器在BOSCC采访中进行语音采样,并探索预训练Ego 4D语音样本,以提高儿童-成人说话者分类的二元互动。我们的研究结果突出了以自我为中心的语音收集和预训练,以提高说话人分类的准确性的潜力。摘要:Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral changes over time during treatment. The standard protocol with this objective is BOSCC, which involves dyadic interactions between a child and clinicians performing a pre-defined set of activities. A fundamental aspect of understanding children's behavior in these interactions is automatic speech understanding, particularly identifying who speaks and when. Conventional approaches in this area heavily rely on speech samples recorded from a spectator perspective, and there is limited research on egocentric speech modeling. In this study, we design an experiment to perform speech sampling in BOSCC interviews from an egocentric perspective using wearable sensors and explore pre-training Ego4D speech samples to enhance child-adult speaker classification in dyadic interactions. Our findings highlight the potential of egocentric speech collection and pre-training to improve speaker classification accuracy.
【57】 The T05 System for The VoiceMOS Challenge 2024: Transfer Learning from Deep Image Classifier to Naturalness MOS Prediction of High-Quality Synthetic Speech
标题: 2024年VoiceMOS挑战赛的T05系统:从深度图像分类器转移学习到高质量合成语音的Naturalness MOS预测
作者:Kaito Baba,Wataru Nakata,Yuki Saito,Hiroshi Saruwatari
备注:Accepted by IEEE SLT 2024. Our MOS prediction system (UTMOSv2) is available in this https URL
链接:点击下载PDF文件
摘要:我们为2024年的VoiceMOS挑战赛(VMC)展示了我们的系统(表示为T05)。我们的系统是为VMC 2024 Track 1设计的,该系统专注于准确预测高质量合成语音的自然度平均意见得分(MOS)。除了预训练的基于自监督学习(SSL)的语音特征提取器外,我们的系统还集成了预训练的图像特征提取器,以捕获语音频谱图中观察到的合成语音的差异。我们首先分别训练两个MOS预测器,使用基于SSL或基于频谱的功能。然后,我们使用两个提取的特征的融合来微调两个预测器以获得更好的MOS预测。在VMC 2024 Track 1中,我们的T05系统在16项评估指标中的7项中获得第一名,在其余9项指标中获得第二名,与排名第三及以下的指标相比有显著差异。我们还报告了我们的消融研究的结果,以调查我们的系统的基本因素。摘要:We present our system (denoted as T05) for the VoiceMOS Challenge (VMC) 2024. Our system was designed for the VMC 2024 Track 1, which focused on the accurate prediction of naturalness mean opinion score (MOS) for high-quality synthetic speech. In addition to a pretrained self-supervised learning (SSL)-based speech feature extractor, our system incorporates a pretrained image feature extractor to capture the difference of synthetic speech observed in speech spectrograms. We first separately train two MOS predictors that use either of an SSL-based or spectrogram-based feature. Then, we fine-tune the two predictors for better MOS prediction using the fusion of two extracted features. In the VMC 2024 Track 1, our T05 system achieved first place in 7 out of 16 evaluation metrics and second place in the remaining 9 metrics, with a significant difference compared to those ranked third and below. We also report the results of our ablation study to investigate essential factors of our system.
【58】 Subband Splitting: Simple, Efficient and Effective Technique for Solving Block Permutation Problem in Determined Blind Source Separation
标题: 子带分裂:解决确定盲源分离中块排列问题的简单、高效且有效的技术
作者:Kazuki Matsumoto,Kohei Yatabe
链接:点击下载PDF文件
摘要:置换问题的求解是确定性盲源分离的关键。现有的方法,如独立向量分析(IVA)和独立低秩矩阵分析(ILRMA),解决置换问题的源信号的频率分量的同现建模。这些方法中的剩余挑战之一是块置换问题,这可能导致差的分离结果。在本文中,我们提出了一个简单而有效的技术解决块置换问题。所提出的技术将整个频率分成重叠的子带,并顺序应用BSS方法(例如,IVA、ILRMA或任何其它方法)到每个子带。由于分裂减少了问题规模,因此BSS方法可以有效地在每个子带中工作。然后,通过使用一个子带中的分离结果作为其他子带的初始值来对齐子带之间的排列。实验结果表明,该方法在不增加总计算量的情况下,显著提高了分离性能。摘要:Solving the permutation problem is essential for determined blind source separation (BSS). Existing methods, such as independent vector analysis (IVA) and independent low-rank matrix analysis (ILRMA), tackle the permutation problem by modeling the co-occurrence of the frequency components of source signals. One of the remaining challenges in these methods is the block permutation problem, which may lead to poor separation results. In this paper, we propose a simple and effective technique for solving the block permutation problem. The proposed technique splits the entire frequencies into overlapping subbands and sequentially applies a BSS method (e.g., IVA, ILRMA, or any other method) to each subband. Since the problem size is reduced by the splitting, the BSS method can effectively work in each subband. Then, the permutations between the subbands are aligned by using the separation result in one subband as the initial values for the other subbands. Experimental results showed that the proposed technique remarkably improved the separation performance without increasing the total computational cost.
【59】 DSCLAP: Domain-Specific Contrastive Language-Audio Pre-Training
标题: DSCSYS:特定领域对比音频预训练
作者:Shengqiang Liu,Da Liu,Anna Wang,Zhiyu Zhang,Jie Gao,Yali Li
链接:点击下载PDF文件
摘要:分析真实世界的多模态信号是智能语音助理(IVA)的一项重要且具有挑战性的任务。主流方法在使用预训练的音频模型和文本模型的IVA的各种下游任务上取得了显着的性能。然而,这些模型是独立地预先训练的,并且通常针对与目标域不同的任务,从而导致下游任务的次优模态表示。此外,在许多领域,收集足够的语言-音频对是极其困难的,并且转录原始音频也需要很高的专业技能,使得联合预训练变得困难甚至不可行。为了解决这些痛点,我们提出了DSCEMP 3,这是一个简单有效的框架,可以只使用原始音频信号输入进行语言音频预训练。具体而言,DSCCast通过ASR系统将原始音频信号转换为文本,并结合对比学习目标和语言-音频匹配目标来对齐音频和ASR传输。我们在12,107小时的车载域音频上预训练DSCEMP 3。两个下游任务的实证结果表明,虽然概念上简单,DSCERAGE显着优于基线模型在所有指标,显示特定领域的IVA应用程序的巨大潜力。摘要:Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on various downstream tasks of IVAs with pre-trained audio models and text models. However, these models are pre-trained independently and usually on tasks different from target domains, resulting in sub-optimal modality representations for downstream tasks. Moreover, in many domains, collecting enough language-audio pairs is extremely hard, and transcribing raw audio also requires high professional skills, making it difficult or even infeasible to joint pre-training. To address these painpoints, we propose DSCLAP, a simple and effective framework that enables language-audio pre-training with only raw audio signal input. Specifically, DSCLAP converts raw audio signals into text via an ASR system and combines a contrastive learning objective and a language-audio matching objective to align the audio and ASR transcriptions. We pre-train DSCLAP on 12,107 hours of in-vehicle domain audio. Empirical results on two downstream tasks show that while conceptually simple, DSCLAP significantly outperforms the baseline models in all metrics, showing great promise for domain-specific IVAs applications.
【60】 M$^{3}$V: A multi-modal multi-view approach for Device-Directed Speech Detection
标题: M$^{3}$V:用于设备引导语音检测的多模式多视图方法
作者:Anna Wang,Da Liu,Zhiyu Zhang,Shengqiang Liu,Jie Gao,Yali Li
链接:点击下载PDF文件
摘要:为了与虚拟语音助手进行更自然、更人性化的交互,该领域最近的研究集中在全双工交互模式上,而不依赖于重复的唤醒词。这要求在具有复杂声源的场景中,语音助理必须将话语分类为面向设备或非面向设备。由文本和语音共同建模的双编码器结构已成为面向设备的语音检测的典范。然而,在实践中,由于自动语音识别(ASR)不可避免的错误,这些模型经常对未对齐的输入对产生不正确的预测。为了解决这一挑战,我们提出了M$^{3}$V,一种用于设备定向语音检测的多模态多视图方法,我们把这个问题定义为一个多视角的学习任务,它引入了单峰视角和文本,音频对齐视图在网络中除了多模态。实验结果表明,M$^{3}$V显著优于仅使用单一或多模态训练的模型,并首次超过人类对ASR错误数据的判断性能。摘要:With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with complex sound sources, the voice assistant must classify utterances as device-oriented or non-device-oriented. The dual-encoder structure, which is jointly modeled by text and speech, has become the paradigm of device-directed speech detection. However, in practice, these models often produce incorrect predictions for unaligned input pairs due to the unavoidable errors of automatic speech recognition (ASR).To address this challenge, we propose M$^{3}$V, a multi-modal multi-view approach for device-directed speech detection, which frames we frame the problem as a multi-view learning task that introduces unimodal views and a text-audio alignment view in the network besides the multi-modal. Experimental results show that M$^{3}$V significantly outperforms models trained using only single or multi-modality and surpasses human judgment performance on ASR error data for the first time.
【61】 SafeEar: Content Privacy-Preserving Audio Deepfake Detection
标题: SafeEar:内容隐私保护音频Deepfake检测
作者:Xinfeng Li,Kai Li,Yifan Zheng,Chen Yan,Xiaoyu Ji,Wenyuan Xu
备注:Accepted by ACM CCS 2024. Please cite this paper as "Xinfeng Li, Kai Li, Yifan Zheng, Chen Yan, Xiaoyu Ji, Wenyuan Xu. SafeEar: Content Privacy-Preserving Audio Deepfake Detection. In Proceedings of ACM Conference on Computer and Communications Security (CCS), 2024."
链接:点击下载PDF文件
摘要:文本到语音(TTS)和语音转换(VC)模型在生成逼真和自然的音频方面表现出了卓越的性能。然而,他们的黑暗面,音频deepfake对社会和个人都构成了重大威胁。现有的对策主要集中在基于完整的原始音频记录来确定语音的私密性,然而,原始音频记录通常包含隐私内容。这种疏忽可能会抑制许多应用程序的deepfake检测,特别是在涉及商业秘密等敏感信息的情况下。在本文中,我们提出了SafeEar,这是一个新的框架,旨在检测deepfake音频,而不依赖于访问其中的语音内容。我们的关键思想是将神经音频编解码器设计成一种新颖的解耦模型,该模型很好地将语义和声学信息从音频样本中分离出来,并且仅使用声学信息(例如,韵律和音色)用于深度伪造检测。通过这种方式,没有语义内容将暴露给检测器。为了克服在没有语义线索的情况下识别各种deepfake音频的挑战,我们用真实世界的编解码器增强来增强我们的deepfake检测器。在四个基准数据集上进行的广泛实验证明了SafeEar在检测各种深度伪造技术方面的有效性,其等错误率(EER)降至2.02%。同时,它屏蔽了五种语言的语音内容,使其不被机器和人类的听觉分析破译,这一点在我们的用户研究和单词错误率(WER)中都超过了93.93%。此外,我们为反deepfake和反内容恢复评估构建的基准有助于为音频隐私保护和deepfake检测领域的未来研究提供基础。摘要:Text-to-Speech (TTS) and Voice Conversion (VC) models have exhibited remarkable performance in generating realistic and natural audio. However, their dark side, audio deepfake poses a significant threat to both society and individuals. Existing countermeasures largely focus on determining the genuineness of speech based on complete original audio recordings, which however often contain private content. This oversight may refrain deepfake detection from many applications, particularly in scenarios involving sensitive information like business secrets. In this paper, we propose SafeEar, a novel framework that aims to detect deepfake audios without relying on accessing the speech content within. Our key idea is to devise a neural audio codec into a novel decoupling model that well separates the semantic and acoustic information from audio samples, and only use the acoustic information (e.g., prosody and timbre) for deepfake detection. In this way, no semantic content will be exposed to the detector. To overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world codec augmentation. Extensive experiments conducted on four benchmark datasets demonstrate SafeEar's effectiveness in detecting various deepfake techniques with an equal error rate (EER) down to 2.02%. Simultaneously, it shields five-language speech content from being deciphered by both machine and human auditory analysis, demonstrated by word error rates (WERs) all above 93.93% and our user study. Furthermore, our benchmark constructed for anti-deepfake and anti-content recovery evaluation helps provide a basis for future research in the realms of audio privacy preservation and deepfake detection.
【62】 Audio-text Retrieval with Transformer-based Hierarchical Alignment and Disentangled Cross-modal Representation
标题: 基于转换器的分层对齐和解纠缠跨模式表示的音频文本检索
作者:Yifei Xin,Zhihong Zhu,Xuxin Cheng,Xusheng Yang,Yuexian Zou
备注:Accepted by Interspeech2024
链接:点击下载PDF文件
摘要:大多数现有的音频-文本检索(ATR)方法通常依赖于单级交互来关联音频和文本,限制了它们对齐不同模态的能力,并导致次优匹配。在这项工作中,我们提出了一种新的ATR框架,利用两个流的Transformers结合层次对齐(THA)模块,以确定音频和文本之间的不同Transformer块的多级对应关系。此外,目前的ATR方法主要集中在学习全局级表示,错过了复杂的细节,以捕捉对应于文本语义的音频出现。为了弥合这一差距,我们引入了一个解开跨模态表示(DCR)的方法,解开高维特征紧凑的潜在因素,把握细粒度的音频文本语义相关性。此外,我们开发了一个置信度感知(CA)模块来估计每个潜在因素对的置信度,并自适应地聚合跨模态潜在因素,以实现局部语义对齐。实验表明,我们的THA有效地提高ATR性能,与DCR方法进一步有助于一致的性能增益。摘要:Most existing audio-text retrieval (ATR) approaches typically rely on a single-level interaction to associate audio and text, limiting their ability to align different modalities and leading to suboptimal matches. In this work, we present a novel ATR framework that leverages two-stream Transformers in conjunction with a Hierarchical Alignment (THA) module to identify multi-level correspondences of different Transformer blocks between audio and text. Moreover, current ATR methods mainly focus on learning a global-level representation, missing out on intricate details to capture audio occurrences that correspond to textual semantics. To bridge this gap, we introduce a Disentangled Cross-modal Representation (DCR) approach that disentangles high-dimensional features into compact latent factors to grasp fine-grained audio-text semantic correlations. Additionally, we develop a confidence-aware (CA) module to estimate the confidence of each latent factor pair and adaptively aggregate cross-modal latent factors to achieve local semantic alignment. Experiments show that our THA effectively boosts ATR performance, with the DCR approach further contributing to consistent performance gains.
【63】 Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?
标题: 多模式语音Transformer解码器:多模式何时可以提高准确性?
作者:Yiwen Guan,Viet Anh Trinh,Vivek Voleti,Jacob Whitehill
链接:点击下载PDF文件
摘要:仅解码器的离散令牌语言模型最近在自动语音识别中取得了显著的成功。然而,对不同模式如何影响特定情景下的绩效的系统分析仍然有限。在本文中,我们研究了多种模态对合成和真实世界数据集识别准确性的影响。我们的实验表明:(1)整合更多的模态可以提高准确性;特别是,据我们所知,我们的论文是第一个展示结合音频,图像上下文和嘴唇信息的好处的论文;(2)图像作为语音识别的补充模态在中等噪声水平下提供最大的好处,此外,与固有同步的模态(如嘴唇运动)相比,它们表现出不同的趋势;(3)当最相关的视觉信息被过滤作为预处理步骤时,性能在合成和真实世界数据集上都有所提高。摘要:Decoder-only discrete-token language models have recently achieved significant success in automatic speech recognition. However, systematic analyses of how different modalities impact performance in specific scenarios remain limited. In this paper, we investigate the effects of multiple modalities on recognition accuracy on both synthetic and real-world datasets. Our experiments suggest that: (1) Integrating more modalities can increase accuracy; in particular, our paper is, to our best knowledge, the first to show the benefit of combining audio, image context, and lip information; (2) Images as a supplementary modality for speech recognition provide the greatest benefit at moderate noise levels, moreover, they exhibit a different trend compared to inherently synchronized modalities like lip movements; (3) Performance improves on both synthetic and real-world datasets when the most relevant visual information is filtered as a preprocessing step.
【64】 Seed-Music: A Unified Framework for High Quality and Controlled Music Generation
标题: Seed-Music:高质量和受控音乐生成的统一框架
作者:Ye Bai,Haonan Chen,Jitong Chen,Zhuo Chen,Yi Deng,Xiaohong Dong,Lamtharn Hantrakul,Weituo Hao,Qingqing Huang,Zhongyi Huang,Dongya Jia,Feihu La,Duc Le,Bochen Li,Chumin Li,Hui Li,Xingxing Li,Shouda Liu,Wei-Tsung Lu,Yiqing Lu,Andrew Shaw,Janne Spijkervet,Yakun Sun,Bo Wang,Ju-Chiang Wang,Yuping Wang,Yuxuan Wang,Ling Xu,Yifeng Yang,Chao Yao,Shuo Zhang,Yang Zhang,Yilin Zhang,Hang Zhao,Ziyi Zhao,Dejian Zhong,Shicen Zhou,Pei Zou
备注:Seed-Music technical report, 20 pages, 5 figures
链接:点击下载PDF文件
摘要:我们推出了Seed-Music,这是一套音乐生成系统,能够生成具有细粒度风格控制的高质量音乐。我们的统一框架利用自回归语言建模和扩散方法来支持两个关键的音乐创作工作流程: textit{受控音乐生成}和 textit{后期制作编辑}。对于受控的音乐生成,我们的系统使声乐生成与性能控制从多模态输入,包括风格描述,音频参考,乐谱,和语音提示。对于后期制作编辑,它提供了直接在生成的音频中编辑歌词和声乐旋律的交互式工具。 我们鼓励读者在https: team.doubao.com seed-music上收听演示音频示例。摘要:We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control. Our unified framework leverages both auto-regressive language modeling and diffusion approaches to support two key music creation workflows: textit{controlled music generation} and textit{post-production editing}. For controlled music generation, our system enables vocal music generation with performance controls from multi-modal inputs, including style descriptions, audio references, musical scores, and voice prompts. For post-production editing, it offers interactive tools for editing lyrics and vocal melodies directly in the generated audio. We encourage readers to listen to demo audio examples at https: team.doubao.com seed-music .
【65】 AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
标题: AccentBox:迈向高保真Zero-Shot口音一代
作者:Jinzuomu Zhong,Korin Richmond,Zhiba Su,Siqi Sun
链接:点击下载PDF文件
摘要:虽然最近的零拍文本到语音(Zero-Shot Text-to-Speech,TTS)模型已经实现了高自然度和说话人相似度,但它们在口音保真度和控制方面存在不足。为了解决这个问题,我们提出了zero-shot口音生成,统一外国口音转换(FAC),重音TTS,和语音TTS,一个新的两阶段的管道。在第一阶段,我们实现了最先进的(SOTA)口音识别(AID)与0.56 f1分数看不见的扬声器。在第二阶段,我们的条件下的预训练的说话人不可知的口音嵌入提取的AID模型的TTS系统。所提出的系统实现了更高的口音保真度的固有 交叉口音生成,并使看不见的口音生成。摘要:While recent Zero-Shot Text-to-Speech (ZS-TTS) models have achieved high naturalness and speaker similarity, they fall short in accent fidelity and control. To address this issue, we propose zero-shot accent generation that unifies Foreign Accent Conversion (FAC), accented TTS, and ZS-TTS, with a novel two-stage pipeline. In the first stage, we achieve state-of-the-art (SOTA) on Accent Identification (AID) with 0.56 f1 score on unseen speakers. In the second stage, we condition ZS-TTS system on the pretrained speaker-agnostic accent embeddings extracted by the AID model. The proposed system achieves higher accent fidelity on inherent cross accent generation, and enables unseen accent generation.
【66】 Estimating the Completeness of Discrete Speech Units
标题: 估计离散语音单元的完整性
作者:Sung-Lin Yeh,Hao Tang
备注:SLT2024
链接:点击下载PDF文件
摘要:用离散单元表示语音在语音编解码和语音生成中有着广泛的应用。然而,有几个未经验证的索赔自我监督离散单元,如解开语音和说话人信息与k均值,或假设信息丢失后k均值。在这项工作中,我们采取信息理论的角度来回答有多少信息是存在的(信息完整性)和有多少信息是可访问的(信息可访问性),之前和之后的残差矢量量化。我们展示了一个下界的信息完整性和估计完整性离散化后的休伯特表示残差矢量量化。我们发现,扬声器信息是充分存在于休伯特离散单元,语音信息是充分存在于残差,表明矢量量化不实现解纠缠。我们的研究结果提供了一个全面的评估离散单元的选择,并建议更多的信息,在剩余应被挖掘,而不是丢弃。摘要:Representing speech with discrete units has been widely used in speech codec and speech generation. However, there are several unverified claims about self-supervised discrete units, such as disentangling phonetic and speaker information with k-means, or assuming information loss after k-means. In this work, we take an information-theoretic perspective to answer how much information is present (information completeness) and how much information is accessible (information accessibility), before and after residual vector quantization. We show a lower bound for information completeness and estimate completeness on discretized HuBERT representations after residual vector quantization. We find that speaker information is sufficiently present in HuBERT discrete units, and that phonetic information is sufficiently present in the residual, showing that vector quantization does not achieve disentanglement. Our results offer a comprehensive assessment on the choice of discrete units, and suggest that a lot more information in the residual should be mined rather than discarded.
机器翻译,仅供参考
