本文经arXiv每日学术速递授权转载
【1】 Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
标题: 用于受控语音生成的TTC生成语音中的韵律参数操纵
作者:Podakanti Satyajith Chary
备注:9 pages, 4 figures, International Summer School on NLP 2024 at IIIT Hyderabad
链接:点击下载PDF文件
【2】 The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
标题: 不同记录条件下阿尔茨海默氏症语音数据集中声学系统的不可靠性
作者:Lara Gauder,Pablo Riera,Andrea Slachevsky,Gonzalo Forno,Adolfo M. Garcia,Luciana Ferrer
备注:5 pages, 1 figure, 1 table
链接:点击下载PDF文件
【3】 Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
标题: Takin:一群高质量Zero-Shot语音生成模型
作者:EverestAI,:,Sijin Chen,Yuan Feng,Laipeng He,Tianwei He,Wendi He,Yanni Hu,Bin Lin,Yiting Lin,Pengfei Tan,Chengwei Tian,Chen Wang,Zhicheng Wang,Ruoye Xie,Jingjing Yin,Jianhao Ye,Jixun Yao,Quanlei Yan,Yuguang Yang
链接:点击下载PDF文件
【4】 WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
标题: WMCCodec:具有深度水印用于真实性验证的端到端神经语音编解码器
作者:Junzuo Zhou,Jiangyan Yi,Yong Ren,Jianhua Tao,Tao Wang,Chu Yuan Zhang
链接:点击下载PDF文件
【5】 Pareto Data Framework: Steps Towards Resource-Efficient Decision Making Using Minimum Viable Data (MVD)
标题: 帕累托数据框架:使用最少可行数据(MVD)实现资源高效决策的步骤
作者:Tashfain Ahmed,Josh Siegel
链接:点击下载PDF文件
【6】 ASR Benchmarking: Need for a More Representative Conversational Dataset
标题: ASB基准:需要更具代表性的对话数据集
作者:Gaurav Maheshwari,Dmitry Ivanov,Théo Johannet,Kevin El Haddad
链接:点击下载PDF文件
【7】 Data Efficient Acoustic Scene Classification using Teacher-Informed Confusing Class Instruction
标题: 使用教师知情的混淆课堂教学进行数据高效的声学场景分类
作者:Jin Jie Sean Yeo,Ee-Leng Tan,Jisheng Bai,Santi Peksi,Woon-Seng Gan
备注:5 pages, 3 figures
链接:点击下载PDF文件
【8】 Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
标题: 使用Frozen wav2vec 2.0进行虚假音频检测的混合专家融合
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Xiaopeng Wang,Yuankun Xie,Xin Qi,Shuchen Shi,Yi Lu,Yukun Liu,Chenxing Li,Xuefei Liu,Guanjun Li
备注:submitted to ICASSP2025
链接:点击下载PDF文件
【9】 M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
标题: M2 R-Whisper:用于增强Whisper的多阶段、多规模检索增强
作者:Jiaming Zhou,Shiwan Zhao,Jiabei He,Hui Wang,Wenjia Zeng,Yong Chen,Haoqin Sun,Aobo Kong,Yong Qin
链接:点击下载PDF文件
【10】 DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
标题: DPI-TTC:用于文本到语音中快速收敛和风格时态建模的定向补丁交互
作者:Xin Qi,Ruibo Fu,Zhengqi Wen,Tao Wang,Chunyu Qiang,Jianhua Tao,Chenxing Li,Yi Lu,Shuchen Shi,Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Yukun Liu,Xuefei Liu,Guanjun Li
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
【11】 Spin Detection Using Racket Bounce Sounds in Table Tennis
标题: 乒乓球运动中利用球拍弹跳声进行旋转检测
作者:Thomas Gossard,Julian Schmalzl,Andreas Ziegler,Andreas Zell
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【12】 METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation
标题: METEOR:旋律感知、文本可控的象征性中音音乐生成
作者:Dinh-Viet-Toan Le,Yi-Hsuan Yang
备注:this https URL
链接:点击下载PDF文件
【13】 SALT: Standardized Audio event Label Taxonomy
标题: SALT:标准化音频事件标签分类
作者:Paraskevas Stamatiadis,Michel Olvera,Slim Essid
Journal-ref:DCASE, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
【14】 Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
标题: 模拟母语说话者阴影以进行具有潜在语音表示的非母语语音评估
作者:Haopeng Geng,Daisuke Saito,Minematsu Nobuaki
备注:Submitted to ICASSP2025 Demo available: this https URL
链接:点击下载PDF文件
【15】 DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
标题: 检测:利用对象信息增强视听表示学习
作者:Shota Nakada,Taichi Nishimura,Hokuto Munakata,Masayoshi Kondo,Tatsuya Komatsu
备注:under review
链接:点击下载PDF文件
【16】 Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
标题: 从粗到细:通过多尺度语音编码和生成改进神经编解码器语言模型
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【17】 Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
标题: 增强、删除和交换:改善LLM字幕的多样性,以实现高效的音乐文本表示学习
作者:Ilaria Manco,Justin Salamon,Oriol Nieto
备注:To appear in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【18】 Evaluation of pretrained language models on music understanding
标题: 预训练语言模型对音乐理解的评估
作者:Yannis Vasilakis,Rachel Bittner,Johan Pauwels
链接:点击下载PDF文件
【19】 Machine listening in a neonatal intensive care unit
标题: 新生儿重症监护室中的机器监听
作者:Modan Tailleur,Vincent Lostanlen,Jean-Philippe Rivière,Pierre Aumond
Journal-ref:DCASE2024 Workshop, Nobutaka Ono; Noboru Harada; Yohei Kawaguchi, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
【20】 Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
标题: 低帧率语音编解码器:专为快速高质量语音LLM训练和推理而设计的编解码器
作者:Edresson Casanova,Ryan Langman,Paarth Neekhara,Shehzeen Hussain,Jason Li,Subhankar Ghosh,Ante Jukić,Sang-gil Lee
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【21】 Conformal Prediction for Manifold-based Source Localization with Gaussian Processes
标题: 基于高斯过程的基于Manifold的源定位保形预测
作者:Vadim Rozenfeld,Bracha Laufer Goldshtein
链接:点击下载PDF文件
【22】 Insights into the Incorporation of Signal Information in Binaural Signal Matching with Wearable Microphone Arrays
标题: 对使用可穿戴麦克风阵列进行双耳信号匹配时信号信息的融入的见解
作者:Ami Berger,Vladimir Tourbabin,Jacob Donley,Zamir Ben-Hur,Boaz Rafaely
链接:点击下载PDF文件
【23】 Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement
标题: Dense-TSNet:用于超轻量级语音增强的密集连接两级结构
作者:Zizhen Lin,Yuanle Li,Junyu Wang,Ruili Li
链接:点击下载PDF文件
【24】 Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
标题: 基于离散单元的掩蔽改善语音转换中的解纠缠
作者:Philip H. Lee,Ismail Rasim Ulgen,Berrak Sisman
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【25】 M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
标题: M-BEST-PQ:智能眼镜的多通道语音基础模型
作者:Yufeng Yang,Desh Raj,Ju Lin,Niko Moritz,Junteng Jia,Gil Keren,Egor Lakomkin,Yiteng Huang,Jacob Donley,Jay Mahadeokar,Ozlem Kalinli
备注:In submission to IEEE ICASSP 2025
链接:点击下载PDF文件
标题: 低帧率语音编解码器:专为快速高质量语音LLM训练和推理而设计的编解码器
作者:Edresson Casanova,Ryan Langman,Paarth Neekhara,Shehzeen Hussain,Jason Li,Subhankar Ghosh,Ante Jukić,Sang-gil Lee
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【2】 Exploring an Inter-Pausal Unit (IPU) based Approach for Indic End-to-End TTS Systems
标题: 探索基于间歇单元(IPU)的印度端到端TTC系统方法
作者:Anusha Prakash,Hema A Murthy
链接:点击下载PDF文件
【3】 Conformal Prediction for Manifold-based Source Localization with Gaussian Processes
标题: 基于高斯过程的基于Manifold的源定位保形预测
作者:Vadim Rozenfeld,Bracha Laufer Goldshtein
链接:点击下载PDF文件
【4】 Insights into the Incorporation of Signal Information in Binaural Signal Matching with Wearable Microphone Arrays
标题: 对使用可穿戴麦克风阵列进行双耳信号匹配时信号信息的融入的见解
作者:Ami Berger,Vladimir Tourbabin,Jacob Donley,Zamir Ben-Hur,Boaz Rafaely
链接:点击下载PDF文件
【5】 Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement
标题: Dense-TSNet:用于超轻量级语音增强的密集连接两级结构
作者:Zizhen Lin,Yuanle Li,Junyu Wang,Ruili Li
链接:点击下载PDF文件
【6】 Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
标题: 基于离散单元的掩蔽改善语音转换中的解纠缠
作者:Philip H. Lee,Ismail Rasim Ulgen,Berrak Sisman
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
【7】 M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
标题: M-BEST-PQ:智能眼镜的多通道语音基础模型
作者:Yufeng Yang,Desh Raj,Ju Lin,Niko Moritz,Junteng Jia,Gil Keren,Egor Lakomkin,Yiteng Huang,Jacob Donley,Jay Mahadeokar,Ozlem Kalinli
备注:In submission to IEEE ICASSP 2025
链接:点击下载PDF文件
【8】 Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
标题: 用于受控语音生成的TTC生成语音中的韵律参数操纵
作者:Podakanti Satyajith Chary
备注:9 pages, 4 figures, International Summer School on NLP 2024 at IIIT Hyderabad
链接:点击下载PDF文件
【9】 The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
标题: 不同记录条件下阿尔茨海默氏症语音数据集中声学系统的不可靠性
作者:Lara Gauder,Pablo Riera,Andrea Slachevsky,Gonzalo Forno,Adolfo M. Garcia,Luciana Ferrer
备注:5 pages, 1 figure, 1 table
链接:点击下载PDF文件
【10】 Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
标题: Takin:一群高质量Zero-Shot语音生成模型
作者:EverestAI,:,Sijin Chen,Yuan Feng,Laipeng He,Tianwei He,Wendi He,Yanni Hu,Bin Lin,Yiting Lin,Pengfei Tan,Chengwei Tian,Chen Wang,Zhicheng Wang,Ruoye Xie,Jingjing Yin,Jianhao Ye,Jixun Yao,Quanlei Yan,Yuguang Yang
链接:点击下载PDF文件
【11】 WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
标题: WMCCodec:具有深度水印用于真实性验证的端到端神经语音编解码器
作者:Junzuo Zhou,Jiangyan Yi,Yong Ren,Jianhua Tao,Tao Wang,Chu Yuan Zhang
链接:点击下载PDF文件
【12】 Pareto Data Framework: Steps Towards Resource-Efficient Decision Making Using Minimum Viable Data (MVD)
标题: 帕累托数据框架:使用最少可行数据(MVD)实现资源高效决策的步骤
作者:Tashfain Ahmed,Josh Siegel
链接:点击下载PDF文件
【13】 ASR Benchmarking: Need for a More Representative Conversational Dataset
标题: ASB基准:需要更具代表性的对话数据集
作者:Gaurav Maheshwari,Dmitry Ivanov,Théo Johannet,Kevin El Haddad
链接:点击下载PDF文件
【14】 Data Efficient Acoustic Scene Classification using Teacher-Informed Confusing Class Instruction
标题: 使用教师知情的混淆课堂教学进行数据高效的声学场景分类
作者:Jin Jie Sean Yeo,Ee-Leng Tan,Jisheng Bai,Santi Peksi,Woon-Seng Gan
备注:5 pages, 3 figures
链接:点击下载PDF文件
【15】 Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
标题: 使用Frozen wav2vec 2.0进行虚假音频检测的混合专家融合
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Xiaopeng Wang,Yuankun Xie,Xin Qi,Shuchen Shi,Yi Lu,Yukun Liu,Chenxing Li,Xuefei Liu,Guanjun Li
备注:submitted to ICASSP2025
链接:点击下载PDF文件
【16】 M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
标题: M2 R-Whisper:用于增强Whisper的多阶段、多规模检索增强
作者:Jiaming Zhou,Shiwan Zhao,Jiabei He,Hui Wang,Wenjia Zeng,Yong Chen,Haoqin Sun,Aobo Kong,Yong Qin
链接:点击下载PDF文件
【17】 DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
标题: DPI-TTC:用于文本到语音中快速收敛和风格时态建模的定向补丁交互
作者:Xin Qi,Ruibo Fu,Zhengqi Wen,Tao Wang,Chunyu Qiang,Jianhua Tao,Chenxing Li,Yi Lu,Shuchen Shi,Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Yukun Liu,Xuefei Liu,Guanjun Li
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
【18】 Spin Detection Using Racket Bounce Sounds in Table Tennis
标题: 乒乓球运动中利用球拍弹跳声进行旋转检测
作者:Thomas Gossard,Julian Schmalzl,Andreas Ziegler,Andreas Zell
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
【19】 METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation
标题: METEOR:旋律感知、文本可控的象征性中音音乐生成
作者:Dinh-Viet-Toan Le,Yi-Hsuan Yang
备注:this https URL
链接:点击下载PDF文件
【20】 SALT: Standardized Audio event Label Taxonomy
标题: SALT:标准化音频事件标签分类
作者:Paraskevas Stamatiadis,Michel Olvera,Slim Essid
Journal-ref:DCASE, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
【21】 Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
标题: 模拟母语说话者阴影以进行具有潜在语音表示的非母语语音评估
作者:Haopeng Geng,Daisuke Saito,Minematsu Nobuaki
备注:Submitted to ICASSP2025 Demo available: this https URL
链接:点击下载PDF文件
【22】 DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
标题: 检测:利用对象信息增强视听表示学习
作者:Shota Nakada,Taichi Nishimura,Hokuto Munakata,Masayoshi Kondo,Tatsuya Komatsu
备注:under review
链接:点击下载PDF文件
【23】 Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
标题: 从粗到细:通过多尺度语音编码和生成改进神经编解码器语言模型
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Xixin Wu,Helen Meng
链接:点击下载PDF文件
【24】 Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
标题: 根据人类对语言、言语和视觉任务的反馈调整偏好:一项调查
作者:Genta Indra Winata,Hanyang Zhao,Anirban Das,Wenpin Tang,David D. Yao,Shi-Xiong Zhang,Sambit Sahu
备注:Survey paper
链接:点击下载PDF文件
【25】 Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
标题: 增强、删除和交换:改善LLM字幕的多样性,以实现高效的音乐文本表示学习
作者:Ilaria Manco,Justin Salamon,Oriol Nieto
备注:To appear in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【26】 Evaluation of pretrained language models on music understanding
标题: 预训练语言模型对音乐理解的评估
作者:Yannis Vasilakis,Rachel Bittner,Johan Pauwels
链接:点击下载PDF文件
【27】 Machine listening in a neonatal intensive care unit
标题: 新生儿重症监护室中的机器监听
作者:Modan Tailleur,Vincent Lostanlen,Jean-Philippe Rivière,Pierre Aumond
Journal-ref:DCASE2024 Workshop, Nobutaka Ono; Noboru Harada; Yohei Kawaguchi, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
标题: 用于受控语音生成的TTC生成语音中的韵律参数操纵
作者:Podakanti Satyajith Chary
备注:9 pages, 4 figures, International Summer School on NLP 2024 at IIIT Hyderabad
链接:点击下载PDF文件
摘要:本文探讨了文本到语音(TTS)系统中的韵律参数的操作,以实现受控的语音生成。通过利用先进的语音处理技术,我们将TTS生成的音频与人类录制的语音进行比较,以分析音调,持续时间和能量的差异。使用PyWorld和Librosa等工具提取关键特征,然后调整这些特征以符合自然人类语音的韵律特征。修改后的功能进行合成,产生增强的TTS输出,更密切地反映人类语音的自然韵律。该方法旨在通过提供一个精确的韵律参数调整框架来增强TTS系统的自然度和表现力。我们的方法包括特征提取,韵律操作和合成,然后进行全面的评估,以确保与人类语音模式的一致性。研究结果表明,韵律参数操纵的可行性和有效性的控制语音生成,突出其潜力显着提高TTS的应用。摘要:This paper explores the manipulation of prosodic parameters in Text-to-Speech (TTS) systems to achieve controlled speech generation. By leveraging advanced speech processing techniques, we compare TTS-generated audio with human-recorded speech to analyze differences in pitch, duration, and energy. Key features are extracted using tools like PyWorld and Librosa, which are then adjusted to align with the prosodic characteristics of natural human speech. The modified features undergo synthesis, producing enhanced TTS outputs that more closely mirror the natural prosody of human speech. This approach aims to enhance the naturalness and expressiveness of TTS systems by providing a framework for precise prosodic parameter adjustments. Our methodology involves feature extraction, prosodic manipulation, and synthesis, followed by comprehensive evaluations to ensure consistency with human speech patterns. The findings demonstrate the feasibility and effectiveness of prosodic parameter manipulation for controlled speech generation, highlighting its potential to significantly improve TTS applications.
【2】 The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
标题: 不同记录条件下阿尔茨海默氏症语音数据集中声学系统的不可靠性
作者:Lara Gauder,Pablo Riera,Andrea Slachevsky,Gonzalo Forno,Adolfo M. Garcia,Luciana Ferrer
备注:5 pages, 1 figure, 1 table
链接:点击下载PDF文件
摘要:自动语音分析是检测阿尔茨海默病(AD)早期标志物的一种蓬勃发展的方法。然而,大多数AD数据集中的记录条件是异质的,患者和对照通常在不同的声学设置中进行评估。虽然这对于基于语音转录或从手动对齐获得的特征的分析来说不是问题,但它确实对声学特征的有效性产生了严重的怀疑,这些声学特征受到采集条件的强烈影响。我们在ADreSSo数据集中研究了这个问题,该数据集来自广泛使用的Pitt语料库。我们表明,基于两个声学特征,MFCC和Wav2vec 2.0嵌入的系统,可以区分AD患者与控制上述机会的性能时,只使用非语音部分的音频信号。我们在一个单独的西班牙语使用者数据集中复制了这一发现。因此,在这些数据集中,可以通过记录条件部分地预测类别。我们的研究结果是对使用声学系统识别基于非标准化记录的患者的警告。我们建议,声学异构数据集的痴呆症研究应(a)分析仅使用转录本或其他功能来自手动注释,或(b)取代收集的数据集,严格控制的声学条件。摘要:Automated speech analysis is a thriving approach to detect early markers of Alzheimer's disease (AD). Yet, recording conditions in most AD datasets are heterogeneous, with patients and controls often evaluated in different acoustic settings. While this is not a problem for analyses based on speech transcription or features obtained from manual alignment, it does cast serious doubts on the validity of acoustic features, which are strongly influenced by acquisition conditions. We examined this issue in the ADreSSo dataset, derived from the widely used Pitt corpus. We show that systems based on two acoustic features, MFCCs and Wav2vec 2.0 embeddings, can discriminate AD patients from controls with above-chance performance when using only the non-speech part of the audio signals. We replicated this finding in a separate dataset of Spanish speakers. Thus, in these datasets, the class can be partly predicted by recording conditions. Our results are a warning against the use of acoustic systems for identifying patients based on non-standardized recordings. We propose that acoustically heterogeneous datasets for dementia studies should be either (a) analyzed using only transcripts or other features derived from manual annotations, or (b) replaced by datasets collected with strictly controlled acoustic conditions.
【3】 Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
标题: Takin:一群高质量Zero-Shot语音生成模型
作者:EverestAI,:,Sijin Chen,Yuan Feng,Laipeng He,Tianwei He,Wendi He,Yanni Hu,Bin Lin,Yiting Lin,Pengfei Tan,Chengwei Tian,Chen Wang,Zhicheng Wang,Ruoye Xie,Jingjing Yin,Jianhao Ye,Jixun Yao,Quanlei Yan,Yuguang Yang
链接:点击下载PDF文件
摘要:随着大数据和大语言模型时代的到来,zero-shot个性化快速定制已经成为一个显著的趋势。在这份报告中,我们介绍了Takin AudioLLM,一系列技术和模型,主要包括Takin TTS,Takin VC和Takin Morphing,专门为有声读物制作而设计。这些模型能够进行zero-shot语音生成,生成与真实人类语音几乎无法区分的高质量语音,并便于个人根据自己的需求定制语音内容。具体来说,我们首先介绍Takin TTS,一种神经编解码器语言模型,它建立在增强的神经语音编解码器和多任务训练框架之上,能够以zero-shot方式生成高保真自然语音。对于Takin VC,我们提倡一种有效的内容和音色联合建模方法来提高说话人相似性,同时提倡一种基于条件流匹配的解码器来进一步增强其自然度和表现力。最后,我们提出了Takin Morphing系统,该系统具有高度解耦和先进的音色和韵律建模方法,使个人能够以精确和可控的方式自定义语音生产。大量的实验验证了我们的Takin AudioLLM系列模型的有效性和鲁棒性。有关详细演示,请参阅https: takinaudiollm.github.io。摘要:With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin AudioLLM, a series of techniques and models, mainly including Takin TTS, Takin VC, and Takin Morphing, specifically designed for audiobook production. These models are capable of zero-shot speech production, generating high-quality speech that is nearly indistinguishable from real human speech and facilitating individuals to customize the speech content according to their own needs. Specifically, we first introduce Takin TTS, a neural codec language model that builds upon an enhanced neural speech codec and a multi-task training framework, capable of generating high-fidelity natural speech in a zero-shot way. For Takin VC, we advocate an effective content and timbre joint modeling approach to improve the speaker similarity, while advocating for a conditional flow matching based decoder to further enhance its naturalness and expressiveness. Last, we propose the Takin Morphing system with highly decoupled and advanced timbre and prosody modeling approaches, which enables individuals to customize speech production with their preferred timbre and prosody in a precise and controllable manner. Extensive experiments validate the effectiveness and robustness of our Takin AudioLLM series models. For detailed demos, please refer to https: takinaudiollm.github.io.
【4】 WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
标题: WMCCodec:具有深度水印用于真实性验证的端到端神经语音编解码器
作者:Junzuo Zhou,Jiangyan Yi,Yong Ren,Jianhua Tao,Tao Wang,Chu Yuan Zhang
链接:点击下载PDF文件
摘要:语音欺骗的最新进展需要神经语音编解码器中更强的验证机制来确保真实性。目前的方法在压缩前嵌入数字水印,并从重建的语音中提取数字水印进行验证,但面临着水印和编解码器的单独训练过程,以及跨模态信息集成不足等限制,导致水印不可感知性,提取精度和容量降低。为了解决这些问题,我们提出了WMCodec,第一个神经语音编解码器联合训练压缩重建和水印嵌入提取在一个端到端的方式,优化不可感知性和可提取的水印。此外,我们还设计了一个迭代的注意力印记单元(AIU)来更深层次地融合水印和语音的特征,降低量化噪声对水印的影响。实验结果表明,WMCodec优于AudioSeal与Encodec在大多数质量指标的水印不可见性和一致超过AudioSeal与Encodec和增强TraceableSpeech的水印提取精度。WMCodec在6kbps带宽下,水印容量为16bps,在常见攻击下提取准确率保持在99%以上,具有较强的鲁棒性。摘要:Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training processes for the watermark and codec, and insufficient cross-modal information integration, leading to reduced watermark imperceptibility, extraction accuracy, and capacity. To address these issues, we propose WMCodec, the first neural speech codec to jointly train compression-reconstruction and watermark embedding-extraction in an end-to-end manner, optimizing both imperceptibility and extractability of the watermark. Furthermore, We design an iterative Attention Imprint Unit (AIU) for deeper feature integration of watermark and speech, reducing the impact of quantization noise on the watermark. Experimental results show WMCodec outperforms AudioSeal with Encodec in most quality metrics for watermark imperceptibility and consistently exceeds both AudioSeal with Encodec and reinforced TraceableSpeech in extraction accuracy of watermark. At bandwidth of 6 kbps with a watermark capacity of 16 bps, WMCodec maintains over 99% extraction accuracy under common attacks, demonstrating strong robustness.
【5】 Pareto Data Framework: Steps Towards Resource-Efficient Decision Making Using Minimum Viable Data (MVD)
标题: 帕累托数据框架:使用最少可行数据(MVD)实现资源高效决策的步骤
作者:Tashfain Ahmed,Josh Siegel
链接:点击下载PDF文件
摘要:本文介绍了Pareto数据框架,这是一种识别和选择在受限平台(如嵌入式系统,移动设备和物联网(IoT)设备)上启用机器学习应用程序所需的最小可行数据(MVD)的方法。我们证明,战略性的数据减少可以保持高性能,同时显着降低带宽,能源,计算和存储成本。该框架确定了最小可行数据(MVD),以在资源受限的环境中优化效率,而不会牺牲性能。它解决了物联网应用中常见的低效做法,例如传感器的过度配置和过精度,以及信号的过采样,为最佳传感器选择,信号提取和传输以及数据表示提出了可扩展的解决方案。实验方法证明了在下采样、量化和截断后有效的声学数据表征,以模拟保真度降低的传感器和网络和存储约束;结果表明,性能可以保持高达95%,采样率降低75%,位深度和剪辑长度减少50%,这意味着大量的成本和资源减少。这些发现对约束系统的设计和开发具有一定的影响。本文还讨论了该框架的更广泛影响,包括在物联网应用和农业、交通运输和制造业等领域实现先进人工智能技术民主化的潜力,以改善数据驱动的洞察力的可访问性并增加其效益。摘要:This paper introduces the Pareto Data Framework, an approach for identifying and selecting the Minimum Viable Data (MVD) required for enabling machine learning applications on constrained platforms such as embedded systems, mobile devices, and Internet of Things (IoT) devices. We demonstrate that strategic data reduction can maintain high performance while significantly reducing bandwidth, energy, computation, and storage costs. The framework identifies Minimum Viable Data (MVD) to optimize efficiency across resource-constrained environments without sacrificing performance. It addresses common inefficient practices in an IoT application such as overprovisioning of sensors and overprecision, and oversampling of signals, proposing scalable solutions for optimal sensor selection, signal extraction and transmission, and data representation. An experimental methodology demonstrates effective acoustic data characterization after downsampling, quantization, and truncation to simulate reduced-fidelity sensors and network and storage constraints; results shows that performance can be maintained up to 95 % with sample rates reduced by 75 % and bit depths and clip length reduced by 50 % which translates into substantial cost and resource reduction. These findings have implications on the design and development of constrained systems. The paper also discusses broader implications of the framework, including the potential to democratize advanced AI technologies across IoT applications and sectors such as agriculture, transportation, and manufacturing to improve access and multiply the benefits of data-driven insights.
【6】 ASR Benchmarking: Need for a More Representative Conversational Dataset
标题: ASB基准:需要更具代表性的对话数据集
作者:Gaurav Maheshwari,Dmitry Ivanov,Théo Johannet,Kevin El Haddad
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统在LibriSpeech和Fleurs等广泛使用的基准测试中取得了卓越的性能。然而,这些基准并不能充分反映现实世界对话环境的复杂性,其中语音通常是非结构化的,并且包含停顿、中断和不同口音等不流利现象。在这项研究中,我们介绍了一个多语言的会话数据集,来自TalkBank,由成人之间的非结构化电话交谈。我们的研究结果显示,在会话设置中进行测试时,各种最先进的ASR模型的性能显着下降。此外,我们观察到的单词错误率和语音不流利的存在之间的相关性,突出了更现实的,会话ASR基准的迫切需要。摘要:Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational environments, where speech is often unstructured and contains disfluencies such as pauses, interruptions, and diverse accents. In this study, we introduce a multilingual conversational dataset, derived from TalkBank, consisting of unstructured phone conversation between adults. Our results show a significant performance drop across various state-of-the-art ASR models when tested in conversational settings. Furthermore, we observe a correlation between Word Error Rate and the presence of speech disfluencies, highlighting the critical need for more realistic, conversational ASR benchmarks.
【7】 Data Efficient Acoustic Scene Classification using Teacher-Informed Confusing Class Instruction
标题: 使用教师知情的混淆课堂教学进行数据高效的声学场景分类
作者:Jin Jie Sean Yeo,Ee-Leng Tan,Jisheng Bai,Santi Peksi,Woon-Seng Gan
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:在本技术报告中,我们描述了SNTL-NTU团队提交的任务1数据高效低复杂度声学场景分类的声学场景和事件的检测和分类(DCASE)2024挑战。引入了三个系统来处理不同大小的训练分割。对于小的训练分割,我们探索了通过减少基本通道的数量来降低所提供的基线模型的复杂性。我们以混合的形式引入数据增强,以增加训练样本的多样性。对于较大的训练分割,我们使用FocusNet为多个Patchout faSt Spectrogram Transformer(PaSST)模型和在原始采样率44.1 kHz上训练的基线模型的集合提供混淆的类信息。我们使用知识蒸馏将集成模型提取为基线学生模型。在TAU Urban Acoustic Scene 2022 Mobile开发数据集上训练系统,三个系统的平均测试准确率分别为(62.21,59.82,56.81,53.03,47.97)%(100,50,25,10,5)%。摘要:In this technical report, we describe the SNTL-NTU team's submission for Task 1 Data-Efficient Low-Complexity Acoustic Scene Classification of the detection and classification of acoustic scenes and events (DCASE) 2024 challenge. Three systems are introduced to tackle training splits of different sizes. For small training splits, we explored reducing the complexity of the provided baseline model by reducing the number of base channels. We introduce data augmentation in the form of mixup to increase the diversity of training samples. For the larger training splits, we use FocusNet to provide confusing class information to an ensemble of multiple Patchout faSt Spectrogram Transformer (PaSST) models and baseline models trained on the original sampling rate of 44.1 kHz. We use Knowledge Distillation to distill the ensemble model to the baseline student model. Training the systems on the TAU Urban Acoustic Scene 2022 Mobile development dataset yielded the highest average testing accuracy of (62.21, 59.82, 56.81, 53.03, 47.97)% on split (100, 50, 25, 10, 5)% respectively over the three systems.
【8】 Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
标题: 使用Frozen wav2vec 2.0进行虚假音频检测的混合专家融合
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Xiaopeng Wang,Yuankun Xie,Xin Qi,Shuchen Shi,Yi Lu,Yukun Liu,Chenxing Li,Xuefei Liu,Guanjun Li
备注:submitted to ICASSP2025
链接:点击下载PDF文件
摘要:语音合成技术对说话人确认系统构成了严重威胁。 目前,最有效的虚假音频检测方法利用预训练模型,并从预训练模型的各个层中集成特征进一步提高检测性能。 然而,大多数先前提出的融合方法需要微调的预训练模型,导致训练时间过长,并阻碍模型迭代时,面对新的语音合成技术。 针对这一问题,提出了一种基于混合专家的特征融合方法,该方法在基于最后一层特征的门控网络的指导下,从层特征中提取并融合与虚假音频检测相关的特征,同时冻结预训练模型。 在ASVspoof2019和ASVspoof2021数据集上进行的实验表明,与需要微调的方法相比,所提出的方法具有竞争力的性能。摘要:Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning.
【9】 M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
标题: M2 R-Whisper:用于增强Whisper的多阶段、多规模检索增强
作者:Jiaming Zhou,Shiwan Zhao,Jiabei He,Hui Wang,Wenjia Zeng,Yong Chen,Haoqin Sun,Aobo Kong,Yong Qin
链接:点击下载PDF文件
摘要:像OpenAI的Whisper这样的最先进的模型在多语言自动语音识别(ASR)方面表现出了强大的性能,但它们在准确识别不同的次方言方面仍然面临挑战。在本文中,我们提出了M2 R-whisper,一种新的多阶段和多尺度检索增强方法,旨在提高低资源设置中的ASR性能。基于上下文学习(ICL)和检索增强技术的原理,我们的方法在预处理阶段采用标记级ICL来利用上下文信息,同时将标记级k最近邻(kNN)检索作为后处理步骤,以进一步细化最终的输出分布。通过协同结合文本级和标记级检索策略,M2 R-whisper有效地减少了各种类型的识别错误。在普通话和亚方言数据集(包括AISHELL-1和KeSpeech)上进行的实验表明,ASR准确性有了实质性的提高,所有这些都是在没有任何参数更新的情况下实现的。摘要:State-of-the-art models like OpenAI's Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-whisper, a novel multi-stage and multi-scale retrieval augmentation approach designed to enhance ASR performance in low-resource settings. Building on the principles of in-context learning (ICL) and retrieval-augmented techniques, our method employs sentence-level ICL in the pre-processing stage to harness contextual information, while integrating token-level k-Nearest Neighbors (kNN) retrieval as a post-processing step to further refine the final output distribution. By synergistically combining sentence-level and token-level retrieval strategies, M2R-whisper effectively mitigates various types of recognition errors. Experiments conducted on Mandarin and subdialect datasets, including AISHELL-1 and KeSpeech, demonstrate substantial improvements in ASR accuracy, all achieved without any parameter updates.
【10】 DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
标题: DPI-TTC:用于文本到语音中快速收敛和风格时态建模的定向补丁交互
作者:Xin Qi,Ruibo Fu,Zhengqi Wen,Tao Wang,Chunyu Qiang,Jianhua Tao,Chenxing Li,Yi Lu,Shuchen Shi,Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Yukun Liu,Xuefei Liu,Guanjun Li
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:近年来,语音扩散模型发展迅速。除了广泛使用的U-Net架构,基于变压器的模型,如扩散Transformer(DiT)也得到了关注。然而,目前的DiT语音模型将Mel频谱图视为一般图像,忽略了语音的特定声学特性。为了解决这些限制,我们提出了一种称为文本到语音定向补丁交互(DPI-TTS)的方法,该方法建立在DiT的基础上,可以在不影响准确性的情况下实现快速训练。值得注意的是,DPI-TTS采用了一种从低到高的频率,逐帧渐进式推理方法,该方法与声学特性更紧密地结合在一起,从而增强了所生成语音的自然度。此外,我们引入了一个细粒度的风格时间建模方法,进一步提高说话人风格的相似性。实验结果表明,我们的方法将训练速度提高了近2倍,并显着优于基线模型。摘要:In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
【11】 Spin Detection Using Racket Bounce Sounds in Table Tennis
标题: 乒乓球运动中利用球拍弹跳声进行旋转检测
作者:Thomas Gossard,Julian Schmalzl,Andreas Ziegler,Andreas Zell
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:虽然乒乓球运动员主要依靠视觉线索,但声音提供了有价值的信息。当球撞击球拍时产生的声音可以帮助预测球的轨迹,特别是确定旋转。虽然职业球员可以通过这些听觉线索来区分旋转,但未经训练的球员往往不会注意到。在本文中,我们证明了不同的球拍产生不同的声音,这可以用来识别球拍类型。此外,我们表明,由球拍产生的声音可以表明是否旋转被施加到球,或没有。为了实现这一目标,我们创建了一个综合数据集,其中包含来自10种球拍配置的反弹声音,每种配置都对球施加不同的旋转。为了实现毫秒级的时间精度,我们首先检测可能对应于乒乓球反弹的高频峰值。然后,我们使用基于CNN的分类器来改进这些结果,该分类器可以准确预测所使用的球拍类型以及是否应用了旋转。摘要:While table tennis players primarily rely on visual cues, sound provides valuable information. The sound generated when the ball strikes the racket can assist in predicting the ball's trajectory, especially in determining the spin. While professional players can distinguish spin through these auditory cues, they often go unnoticed by untrained players. In this paper, we demonstrate that different rackets produce distinct sounds, which can be used to identify the racket type. In addition, we show that the sound generated by the racket can indicate whether spin was applied to the ball, or not. To achieve this, we created a comprehensive dataset featuring bounce sounds from 10 racket configurations, each applying various spins to the ball. To achieve millisecond level temporal accuracy, we first detect high frequency peaks that may correspond to table tennis ball bounces. We then refine these results using a CNN based classifier that accurately predicts both the type of racket used and whether spin was applied.
【12】 METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation
标题: METEOR:旋律感知、文本可控的象征性中音音乐生成
作者:Dinh-Viet-Toan Le,Yi-Hsuan Yang
备注:this https URL
链接:点击下载PDF文件
摘要:西方音乐的特点通常是谐音织体,其中音乐内容可以组织成旋律和伴奏。在管弦乐中,特别是,作曲家可以选择特定的特征,为每个乐器的部分内的伴奏,同时也需要适应的旋律,以适应能力的仪器执行it.In这项工作中,我们提出了METEOR,一个模型的旋律感知纹理可控的管弦乐音乐生成。该模型执行符号化的多轨音乐风格转移,重点是旋律保真度。我们允许酒吧和轨道水平的控制性的伴奏与各种纹理属性,同时保持一个谐音纹理。我们表明,该模型可以实现类似于强基线的可控性性能,同时大大提高旋律保真度。摘要:Western music is often characterized by a homophonic texture, in which the musical content can be organized into a melody and an accompaniment. In orchestral music, in particular, the composer can select specific characteristics for each instrument's part within the accompaniment, while also needing to adapt the melody to suit the capabilities of the instruments performing it. In this work, we propose METEOR, a model for Melody-aware Texture-controllable Orchestral music generation. This model performs symbolic multi-track music style transfer with a focus on melodic fidelity. We allow bar- and track-level controllability of the accompaniment with various textural attributes while keeping a homophonic texture. We show that the model can achieve controllability performances similar to strong baselines while greatly improve melodic fidelity.
【13】 SALT: Standardized Audio event Label Taxonomy
标题: SALT:标准化音频事件标签分类
作者:Paraskevas Stamatiadis,Michel Olvera,Slim Essid
Journal-ref:DCASE, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
摘要:机器监听系统通常依赖于固定的分类来组织和标记音频数据,这是训练和评估深度神经网络(DNN)和其他监督算法的关键。然而,这样的分类法面临着重大的限制:它们由依赖于应用程序的预定义类别组成,这阻碍了新的或不同的声音的整合,并且由于不一致的标签标准而表现出有限的跨数据集兼容性。为了克服这些限制,我们引入了SALT:标准化音频事件标签分类。AudioSet的本体论的层次结构的基础上,我们的分类扩展,并在24个公开可用的环境声音数据集的标签,允许从不同的数据集的类标签映射到一个统一的系统。我们的提案附带了一个新的Python包,旨在导航和利用这种分类法,简化跨数据集标签搜索和分层探索。值得注意的是,我们的软件包允许轻松地从不同来源聚合数据,因此可以轻松地对组合数据集进行实验。摘要:Machine listening systems often rely on fixed taxonomies to organize and label audio data, key for training and evaluating deep neural networks (DNNs) and other supervised algorithms. However, such taxonomies face significant constraints: they are composed of application-dependent predefined categories, which hinders the integration of new or varied sounds, and exhibits limited cross-dataset compatibility due to inconsistent labeling standards. To overcome these limitations, we introduce SALT: Standardized Audio event Label Taxonomy. Building upon the hierarchical structure of AudioSet's ontology, our taxonomy extends and standardizes labels across 24 publicly available environmental sound datasets, allowing the mapping of class labels from diverse datasets to a unified system. Our proposal comes with a new Python package designed for navigating and utilizing this taxonomy, easing cross-dataset label searching and hierarchical exploration. Notably, our package allows effortless data aggregation from diverse sources, hence easy experimentation with combined datasets.
【14】 Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
标题: 模拟母语说话者阴影以进行具有潜在语音表示的非母语语音评估
作者:Haopeng Geng,Daisuke Saito,Minematsu Nobuaki
备注:Submitted to ICASSP2025 Demo available: this https URL
链接:点击下载PDF文件
摘要:语音清晰度评价是计算机辅助语言学习系统中的一项重要任务。传统方法通常依赖于自动语音识别(ASR)提供的单词错误率(WER)作为可懂度分数。然而,由于人类语音识别(HSR)和ASR之间的显著差异,这种方法具有显著的局限性。一个有希望的替代方法是让母语(L1)说话者参与非母语(L2)说话者所说的话。L1说话者的阴影话语中的故障或错误发音可以作为评估L2语音可懂度的指标。在这项研究中,我们提出了一个语音生成系统,该系统使用语音转换(VC)技术和潜在语音表示来模拟L1阴影过程。我们的实验结果表明,该方法有效地复制了L1的阴影过程,提供了一个创新的工具来评估L2语音可懂度。值得注意的是,利用自监督语音表示(S3R)的系统在语言准确性和自然性方面与真实的L1阴影话语具有更高的相似度。摘要:Evaluating speech intelligibility is a critical task in computer-aided language learning systems. Traditional methods often rely on word error rates (WER) provided by automatic speech recognition (ASR) as intelligibility scores. However, this approach has significant limitations due to notable differences between human speech recognition (HSR) and ASR. A promising alternative is to involve a native (L1) speaker in shadowing what nonnative (L2) speakers say. Breakdowns or mispronunciations in the L1 speaker's shadowing utterance can serve as indicators for assessing L2 speech intelligibility. In this study, we propose a speech generation system that simulates the L1 shadowing process using voice conversion (VC) techniques and latent speech representations. Our experimental results demonstrate that this method effectively replicates the L1 shadowing process, offering an innovative tool to evaluate L2 speech intelligibility. Notably, systems that utilize self-supervised speech representations (S3R) show a higher degree of similarity to real L1 shadowing utterances in both linguistic accuracy and naturalness.
【15】 DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
标题: 检测:利用对象信息增强视听表示学习
作者:Shota Nakada,Taichi Nishimura,Hokuto Munakata,Masayoshi Kondo,Tatsuya Komatsu
备注:under review
链接:点击下载PDF文件
摘要:当前的视听表示学习可以捕获粗略的对象类别(例如,“动物”和“乐器”),但它缺乏识别细粒度细节的能力,例如动物和乐器中的“狗”和“长笛”等特定类别。为了解决这个问题,我们引入了DETECTORM,这是一种利用对象信息增强视听表示学习的方法。我们的主要思想是在现有的对比视听掩蔽自动编码器中引入视听标签预测损失,以增强其对象感知能力。为了避免昂贵的手动注释,我们使用最先进的语言音频模型和对象检测器从音频和视觉输入中准备对象标签。我们使用VGGSound和AudioSet20K数据集评估视听检索和分类的方法。我们的方法在recall@10方面分别实现了音频到视频和视频到音频检索的+1.5%和+1.2%的改进,并且在视听分类方面提高了+0.6%的准确性。摘要:Current audio-visual representation learning can capture rough object categories (e.g., animals'' and instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like dogs'' and flutes'' within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification.
【16】 Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
标题: 从粗到细:通过多尺度语音编码和生成改进神经编解码器语言模型
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:神经编解码语言模型(CLM)在文语转换(TTS)合成中表现出了卓越的性能。然而,受“近因偏差”的困扰,CLM缺乏对更高时间尺度上的粗粒度信息的足够关注,经常产生不自然甚至无法理解的语音。这项工作提出了CoFi-Speech,一种从粗到细的CLM-TTS方法,采用多尺度语音编码和生成来解决这个问题。我们训练多尺度神经编解码器,CoFi-Codec,将语音编码成多尺度离散表示,包括具有不同时间分辨率的多个令牌序列。然后,我们提出了CoFi-LM,它可以在两种模式下生成这种表示:基于单LM的规模链生成和基于多LM的规模堆栈生成。在实验中,CoFi-Speech在zero-shot TTS中的自然度和说话者相似性方面显着优于单尺度基线系统。多尺度编码的分析证明了CoFi-Codec在学习多尺度离散语音表示的同时保持高质量语音重建的有效性。从粗到细的多尺度生成,特别是对于尺度堆栈方法,也被验证为追求高质量的TTS神经编解码器语言模型的关键方法。摘要:The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by recency bias", CLM lacks sufficient attention to coarse-grained information at a higher temporal scale, often producing unnatural or even unintelligible speech. This work proposes CoFi-Speech, a coarse-to-fine CLM-TTS approach, employing multi-scale speech coding and generation to address this issue. We train a multi-scale neural codec, CoFi-Codec, to encode speech into a multi-scale discrete representation, comprising multiple token sequences with different time resolutions. Then, we propose CoFi-LM that can generate this representation in two modes: the single-LM-based chain-of-scale generation and the multiple-LM-based stack-of-scale generation. In experiments, CoFi-Speech significantly outperforms single-scale baseline systems on naturalness and speaker similarity in zero-shot TTS. The analysis of multi-scale coding demonstrates the effectiveness of CoFi-Codec in learning multi-scale discrete speech representations while keeping high-quality speech reconstruction. The coarse-to-fine multi-scale generation, especially for the stack-of-scale approach, is also validated as a crucial approach in pursuing a high-quality neural codec language model for TTS.
【17】 Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
标题: 增强、删除和交换:改善LLM字幕的多样性,以实现高效的音乐文本表示学习
作者:Ilaria Manco,Justin Salamon,Oriol Nieto
备注:To appear in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:音频文本对比模型已经成为音乐表征学习的一种强有力的方法。尽管他们的经验的成功,但是,鲜为人知的是通过这个框架学习的音乐文本表示的质量的关键设计选择的影响。在这项工作中,我们在有限的数据和计算预算的约束下暴露了这些设计选择,并沿着三个轴建立了基于经验观察的对其影响的更坚实的理解:基本编码器的选择,训练数据的策展水平,以及文本增强的使用。我们发现,在资源受限的情况下,数据策展是音乐文本对比训练的最重要因素。基于这一认识,我们引入了两种新技术,增强视图丢弃和文本交换,它们增加了训练中文本输入的多样性和连续性。通过我们的实验,我们证明了这些方法可以有效地提高不同预训练机制、模型架构和下游数据分布的性能,而不会产生更高的计算成本或需要额外的训练数据。摘要:Audio-text contrastive models have become a powerful approach in music representation learning. Despite their empirical success, however, little is known about the influence of key design choices on the quality of music-text representations learnt through this framework. In this work, we expose these design choices within the constraints of limited data and computation budgets, and establish a more solid understanding of their impact grounded in empirical observations along three axes: the choice of base encoders, the level of curation in training data, and the use of text augmentation. We find that data curation is the single most important factor for music-text contrastive training in resource-constrained scenarios. Motivated by this insight, we introduce two novel techniques, Augmented View Dropout and TextSwap, which increase the diversity and descriptiveness of text inputs seen in training. Through our experiments we demonstrate that these are effective at boosting performance across different pre-training regimes, model architectures, and downstream data distributions, without incurring higher computational costs or requiring additional training data.
【18】 Evaluation of pretrained language models on music understanding
标题: 预训练语言模型对音乐理解的评估
作者:Yannis Vasilakis,Rachel Bittner,Johan Pauwels
链接:点击下载PDF文件
摘要:音乐文本多模态系统已经使新的方法,音乐信息研究(MIR)的应用,如音频到文本和文本到音频检索,基于文本的歌曲生成,和音乐字幕。尽管报告的成功,很少有人投入到评估大语言模型(LLM)的音乐知识。在本文中,我们证明了LLM遭受1)提示敏感性,2)无法模拟否定(例如“没有吉他的摇滚歌曲”),以及3)对特定单词的敏感性。我们将这些属性量化为基于三元组的准确性,评估在分层本体中对标签的相对相似性进行建模的能力。我们利用Audioset本体来生成由流派和乐器子树的锚、正(相关)标签和负(不太相关)标签组成的三元组。我们评估了基于三元组的音乐知识的六个通用的基于变压器的模型。通过这种方法获得的三联体需要过滤,因为有些三联体难以判断,因此对于评价目的而言相对缺乏信息。尽管报告的准确性相对较高,但所有六种模型都存在明显的不一致性,这表明现成的LLM在使用前需要适应音乐。摘要:Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the reported success, little effort has been put into evaluating the musical knowledge of Large Language Models (LLM). In this paper, we demonstrate that LLMs suffer from 1) prompt sensitivity, 2) inability to model negation (e.g. 'rock song without guitar'), and 3) sensitivity towards the presence of specific words. We quantified these properties as a triplet-based accuracy, evaluating the ability to model the relative similarity of labels in a hierarchical ontology. We leveraged the Audioset ontology to generate triplets consisting of an anchor, a positive (relevant) label, and a negative (less relevant) label for the genre and instruments sub-tree. We evaluated the triplet-based musical knowledge for six general-purpose Transformer-based models. The triplets obtained through this methodology required filtering, as some were difficult to judge and therefore relatively uninformative for evaluation purposes. Despite the relatively high accuracy reported, inconsistencies are evident in all six models, suggesting that off-the-shelf LLMs need adaptation to music before use.
【19】 Machine listening in a neonatal intensive care unit
标题: 新生儿重症监护室中的机器监听
作者:Modan Tailleur,Vincent Lostanlen,Jean-Philippe Rivière,Pierre Aumond
Journal-ref:DCASE2024 Workshop, Nobutaka Ono; Noboru Harada; Yohei Kawaguchi, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
摘要:氧合器、报警装置和脚步声是医院中最常见的声源。检测它们对环境心理学具有科学价值,但也带来了自身的挑战:即隐私保护和有限的标记数据。在本文中,我们通过边缘计算和云计算的结合来解决这两个挑战。为了保护隐私,我们设计了一种声学传感器,它可以实时计算第三倍频程频谱图,而不是记录音频波形。为了实现样本高效的机器学习,我们通过频谱转码和标签空间自适应来重新利用预训练的音频神经网络(PANN)。在神经病学重症监护室(NICU)中的小规模研究证实,检测到的事件的时间序列与另一种测量方式一致:即,为父母和医疗保健专业人员提供电子徽章。因此,本文论证了在医院病房中使用复音机收听的可行性,同时通过设计来保证隐私。摘要:Oxygenators, alarm devices, and footsteps are some of the most common sound sources in a hospital. Detecting them has scientific value for environmental psychology but comes with challenges of its own: namely, privacy preservation and limited labeled data. In this paper, we address these two challenges via a combination of edge computing and cloud computing. For privacy preservation, we have designed an acoustic sensor which computes third-octave spectrograms on the fly instead of recording audio waveforms. For sample-efficient machine learning, we have repurposed a pretrained audio neural network (PANN) via spectral transcoding and label space adaptation. A small-scale study in a neonatological intensive care unit (NICU) confirms that the time series of detected events align with another modality of measurement: i.e., electronic badges for parents and healthcare professionals. Hence, this paper demonstrates the feasibility of polyphonic machine listening in a hospital ward while guaranteeing privacy by design.
【20】 Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
标题: 低帧率语音编解码器:专为快速高质量语音LLM训练和推理而设计的编解码器
作者:Edresson Casanova,Ryan Langman,Paarth Neekhara,Shehzeen Hussain,Jason Li,Subhankar Ghosh,Ante Jukić,Sang-gil Lee
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:大型语言模型(LLM)通过将音频转换为离散令牌的音频编解码器显著地改进了音频处理,从而能够将语言建模技术应用于音频数据。然而,音频编解码器通常以高帧速率操作,导致缓慢的训练和推断,特别是对于自回归模型。为了应对这一挑战,我们提出了低帧率语音编解码器(LFSC):一种神经音频编解码器,利用有限标量量化和大型语音语言模型的对抗训练,以实现1.89 kbps比特率和每秒21.5帧的高质量音频压缩。我们证明,我们的新型编解码器可以使基于LLM的文本到语音模型的推理速度提高三倍左右,同时提高可懂度和生产质量,与以前的模型相比。摘要:Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models.
【21】 Conformal Prediction for Manifold-based Source Localization with Gaussian Processes
标题: 基于高斯过程的基于Manifold的源定位保形预测
作者:Vadim Rozenfeld,Bracha Laufer Goldshtein
链接:点击下载PDF文件
摘要:我们解决了在不利的声学环境中声源定位的不确定性量化的挑战。估计源的位置受到诸如噪声和混响的各种因素的影响,导致显著的不确定性。量化这种不确定性至关重要,特别是当定位结果影响关键决策过程时,例如在机器人试听中,位置估计的准确性直接影响后续行动。尽管如此,许多定位方法通常提供点估计而不量化估计不确定性。为了解决这个问题,我们采用了共形预测(CP)-一个框架,提供统计上有效的预测区间与有限样本的保证,独立于数据分布。然而,常用的感应CP(ICP)方法需要大量的标记数据,这在定位设置中可能难以获得。为了减轻这一限制,我们将一个基于流形的定位方法,使用高斯过程回归(GPR),一个高效的Transductive CP(TCP)技术专门为GPR设计。我们证明了我们的方法在不同的声学条件下产生统计上有效的不确定性区间。摘要:We tackle the challenge of uncertainty quantification in the localization of a sound source within adverse acoustic environments. Estimating the position of the source is influenced by various factors such as noise and reverberation, leading to significant uncertainty. Quantifying this uncertainty is essential, particularly when localization outcomes impact critical decision-making processes, such as in robot audition, where the accuracy of location estimates directly influences subsequent actions. Despite this, many localization methods typically offer point estimates without quantifying the estimation uncertainty. To address this, we employ conformal prediction (CP)-a framework that delivers statistically valid prediction intervals with finite-sample guarantees, independent of the data distribution. However, commonly used Inductive CP (ICP) methods require a substantial amount of labeled data, which can be difficult to obtain in the localization setting. To mitigate this limitation, we incorporate a manifold-based localization method using Gaussian process regression (GPR), with an efficient Transductive CP (TCP) technique specifically designed for GPR. We demonstrate that our method generates statistically valid uncertainty intervals across different acoustic conditions.
【22】 Insights into the Incorporation of Signal Information in Binaural Signal Matching with Wearable Microphone Arrays
标题: 对使用可穿戴麦克风阵列进行双耳信号匹配时信号信息的融入的见解
作者:Ami Berger,Vladimir Tourbabin,Jacob Donley,Zamir Ben-Hur,Boaz Rafaely
链接:点击下载PDF文件
摘要:空间音频在诸如电话会议、娱乐和虚拟现实等应用中的日益普及导致了双耳再现方法的最近发展。然而,这些方法中只有少数非常适合可穿戴和移动阵列,这些阵列通常由少量麦克风组成。一种这样的方法是双耳信号匹配(BSM),其已经被示出为可穿戴阵列产生高质量的双耳信号。然而,BSM可能是次优的情况下,高直达混响比(DRR),因为它是基于扩散声场假设。为了克服这一局限性,以前的研究采用了声场模型,而不是扩散。然而,这一方法没有得到全面研究。本文广泛研究了两种基于BSM的方法设计的高DRR的情况下。该方法结合了由直达分量和混响分量组成的声场模型,对该方法进行了数学和仿真研究,最后通过听音测试进行了验证。结果表明,所提出的方法可以显着提高BSM的性能,特别是在源的方向上,而在其他方向上只呈现出微不足道的退化。此外,当源方向估计是不准确的,这些方法的性能下降到等于BSM,呈现出所需的鲁棒性质量。摘要:The increasing popularity of spatial audio in applications such as teleconferencing, entertainment, and virtual reality has led to the recent developments of binaural reproduction methods. However, only a few of these methods are well-suited for wearable and mobile arrays, which typically consist of a small number of microphones. One such method is binaural signal matching (BSM), which has been shown to produce high-quality binaural signals for wearable arrays. However, BSM may be suboptimal in cases of high direct-to-reverberant ratio (DRR) as it is based on the diffuse sound field assumption. To overcome this limitation, previous studies incorporated sound-field models other than diffuse. However, this approach was not studied comprehensively. This paper extensively investigates two BSM-based methods designed for high DRR scenarios. The methods incorporate a sound field model composed of direct and reverberant components.The methods are investigated both mathematically and using simulations, finally validated by a listening test. The results show that the proposed methods can significantly improve the performance of BSM , in particular in the direction of the source, while presenting only a negligible degradation in other directions. Furthermore, when source direction estimation is inaccurate, performance of these methods degrade to equal that of the BSM, presenting a desired robustness quality.
【23】 Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement
标题: Dense-TSNet:用于超轻量级语音增强的密集连接两级结构
作者:Zizhen Lin,Yuanle Li,Junyu Wang,Ruili Li
链接:点击下载PDF文件
摘要:语音增强的目的是提高语音质量和清晰度在嘈杂的环境。最近的进展集中在深度神经网络上,特别是采用两阶段(TS)架构来增强特征提取。然而,这些模型的复杂性和规模仍然很大,这限制了它们在资源有限的情况下的适用性。设计适用于边缘设备的模型面临着一系列挑战。窄的轻量级模型经常遇到性能瓶颈,由于不均匀的损失景观。此外,诸如Transformers或Mamba等高级运营商可能缺乏卷积神经网络(CNN)在现实世界部署中提供的实际适应性和效率。为了应对这些挑战,我们提出了一种创新的超轻量级语音增强网络Dense-TSNet。我们的方法采用了一种新的密集两阶段(密集TS)架构,与经典的两阶段架构相比,它确保了在后期训练阶段对目标函数进行更强大的细化。这导致改进的最终性能,解决了基线模型的早期收敛限制。我们还介绍了多视图凝视块(MVGB),它通过卷积神经网络(CNN)整合全局,通道和局部视角来增强特征提取。此外,我们讨论了损失函数的选择如何影响感知质量。Dense-TSNet具有约14 K参数的紧凑模型大小,表现出良好的性能,使其特别适合在资源受限的环境中部署。摘要:Speech enhancement aims to improve speech quality and intelligibility in noisy environments. Recent advancements have concentrated on deep neural networks, particularly employing the Two-Stage (TS) architecture to enhance feature extraction. However, the complexity and size of these models remain significant, which limits their applicability in resource-constrained scenarios. Designing models suitable for edge devices presents its own set of challenges. Narrow lightweight models often encounter performance bottlenecks due to uneven loss landscapes. Additionally, advanced operators such as Transformers or Mamba may lack the practical adaptability and efficiency that convolutional neural networks (CNNs) offer in real-world deployments. To address these challenges, we propose Dense-TSNet, an innovative ultra-lightweight speech enhancement network. Our approach employs a novel Dense Two-Stage (Dense-TS) architecture, which, compared to the classic Two-Stage architecture, ensures more robust refinement of the objective function in the later training stages. This leads to improved final performance, addressing the early convergence limitations of the baseline model. We also introduce the Multi-View Gaze Block (MVGB), which enhances feature extraction by incorporating global, channel, and local perspectives through convolutional neural networks (CNNs). Furthermore, we discuss how the choice of loss function impacts perceptual quality. Dense-TSNet demonstrates promising performance with a compact model size of around 14K parameters, making it particularly well-suited for deployment in resource-constrained environments.
【24】 Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
标题: 基于离散单元的掩蔽改善语音转换中的解纠缠
作者:Philip H. Lee,Ismail Rasim Ulgen,Berrak Sisman
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:语音转换(VC)的目的是修改说话人的身份,同时保持语言内容。通常,VC方法使用编码器-解码器架构,其中将说话者的身份从语言信息中分离出来至关重要。然而,在这些方法中使用的解纠缠方法是有限的,因为说话人的功能取决于语音内容的话语,妥协解纠缠。这种依赖性被基于注意力的方法放大了。为了解决这个问题,我们在扬声器编码之前的输入中引入了一种新的掩蔽机制,掩蔽了与音素类高度对应的某些离散语音单元。我们的工作旨在通过限制对某些语音信息的访问来减少说话人特征的语音依赖。此外,由于我们的方法是在输入级,它适用于任何基于VC框架的编码器-解码器。我们的方法提高了多个VC方法的解纠缠和转换性能,显示出显着的有效性,特别是在基于注意力的方法,与44%的相对改善客观可懂度。摘要:Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic information is crucial. However, the disentanglement approaches used in these methods are limited as the speaker features depend on the phonetic content of the utterance, compromising disentanglement. This dependency is amplified with attention-based methods. To address this, we introduce a novel masking mechanism in the input before speaker encoding, masking certain discrete speech units that correspond highly with phoneme classes. Our work aims to reduce the phonetic dependency of speaker features by restricting access to some phonetic information. Furthermore, since our approach is at the input level, it is applicable to any encoder-decoder based VC framework. Our approach improves disentanglement and conversion performance across multiple VC methods, showing significant effectiveness, particularly in attention-based method, with 44% relative improvement in objective intelligibility.
【25】 M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
标题: M-BEST-PQ:智能眼镜的多通道语音基础模型
作者:Yufeng Yang,Desh Raj,Ju Lin,Niko Moritz,Junteng Jia,Gil Keren,Egor Lakomkin,Yiteng Huang,Jacob Donley,Jay Mahadeokar,Ozlem Kalinli
备注:In submission to IEEE ICASSP 2025
链接:点击下载PDF文件
摘要:多通道可穿戴设备(如智能眼镜)的日益普及,导致了定向语音识别和增强听力等应用的激增。然而,目前解决这些任务的方法使用独立训练的模型,这可能不会从大量未标记的数据中受益。在本文中,我们提出了M-BEST-RQ,这是智能眼镜的第一个多通道语音基础模型,旨在利用大规模自监督学习(SSL)的阵列几何不可知方法。虽然之前的多通道语音SSL工作仅在模拟设置上进行了评估,但我们策划了一套真实的下游任务来评估我们的模型,即(i)会话自动语音识别(ASR),(ii)球形主动源定位,以及(iii)眼镜佩戴者语音活动检测,这些任务来自MMCSG和EasyCom数据集。我们证明了通用的M-BEST-RQ编码器能够在所有任务中匹配或超越监督模型。特别是对于会话式ASR任务,仅使用8小时的标记语音,我们的模型优于在2000小时的标记数据上训练的监督ASR基线,这证明了我们方法的有效性。摘要:The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of unlabeled data. In this paper, we propose M-BEST-RQ, the first multi-channel speech foundation model for smart glasses, which is designed to leverage large-scale self-supervised learning (SSL) in an array-geometry agnostic approach. While prior work on multi-channel speech SSL only evaluated on simulated settings, we curate a suite of real downstream tasks to evaluate our model, namely (i) conversational automatic speech recognition (ASR), (ii) spherical active source localization, and (iii) glasses wearer voice activity detection, which are sourced from the MMCSG and EasyCom datasets. We show that a general-purpose M-BEST-RQ encoder is able to match or surpass supervised models across all tasks. For the conversational ASR task in particular, using only 8 hours of labeled speech, our model outperforms a supervised ASR baseline that is trained on 2000 hours of labeled data, which demonstrates the effectiveness of our approach.
eess.AS音频处理
【1】 Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference标题: 低帧率语音编解码器:专为快速高质量语音LLM训练和推理而设计的编解码器
作者:Edresson Casanova,Ryan Langman,Paarth Neekhara,Shehzeen Hussain,Jason Li,Subhankar Ghosh,Ante Jukić,Sang-gil Lee
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:大型语言模型(LLM)通过将音频转换为离散令牌的音频编解码器显著地改进了音频处理,从而能够将语言建模技术应用于音频数据。然而,音频编解码器通常以高帧速率操作,导致缓慢的训练和推断,特别是对于自回归模型。为了应对这一挑战,我们提出了低帧率语音编解码器(LFSC):一种神经音频编解码器,利用有限标量量化和大型语音语言模型的对抗训练,以实现1.89 kbps比特率和每秒21.5帧的高质量音频压缩。我们证明,我们的新型编解码器可以使基于LLM的文本到语音模型的推理速度提高三倍左右,同时提高可懂度和生产质量,与以前的模型相比。摘要:Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models.
【2】 Exploring an Inter-Pausal Unit (IPU) based Approach for Indic End-to-End TTS Systems
标题: 探索基于间歇单元(IPU)的印度端到端TTC系统方法
作者:Anusha Prakash,Hema A Murthy
链接:点击下载PDF文件
摘要:印度语言中的句子通常比英语中的句子长。印度语言也被认为是基于短语的,其中语义上完整的短语被连接起来组成句子。长的话语导致文本到语音模型的训练较差,并且在合成期间导致较差的韵律。在这项工作中,我们探索了端到端(E2 E)框架中基于停顿单元(IPU)的方法,重点关注合成对话风格的文本。在我们的研究中,我们考虑自回归Tacotron2和非自回归FastSpeech2架构,并使用三种印度语言进行实验,即印地语,泰米尔语和泰卢固语。使用基于IPU的Tacotron 2方法,我们看到合成音频中的插入和删除错误减少,在减少错误方面为FastSpeech(2)网络提供了另一种方法。基于IPU的方法需要更少的计算资源,并产生韵律更丰富的合成相比,传统的基于韵律的系统。摘要:Sentences in Indian languages are generally longer than those in English. Indian languages are also considered to be phrase-based, wherein semantically complete phrases are concatenated to make up sentences. Long utterances lead to poor training of text-to-speech models and result in poor prosody during synthesis. In this work, we explore an inter-pausal unit (IPU) based approach in the end-to-end (E2E) framework, focusing on synthesising conversational-style text. We consider both autoregressive Tacotron2 and non-autoregressive FastSpeech2 architectures in our study and perform experiments with three Indian languages, namely, Hindi, Tamil and Telugu. With the IPU-based Tacotron2 approach, we see a reduction in insertion and deletion errors in the synthesised audio, providing an alternative approach to the FastSpeech(2) network in terms of error reduction. The IPU-based approach requires less computational resources and produces prosodically richer synthesis compared to conventional sentence-based systems.
【3】 Conformal Prediction for Manifold-based Source Localization with Gaussian Processes
标题: 基于高斯过程的基于Manifold的源定位保形预测
作者:Vadim Rozenfeld,Bracha Laufer Goldshtein
链接:点击下载PDF文件
摘要:我们解决了在不利的声学环境中声源定位的不确定性量化的挑战。估计源的位置受到诸如噪声和混响的各种因素的影响,导致显著的不确定性。量化这种不确定性至关重要,特别是当定位结果影响关键决策过程时,例如在机器人试听中,位置估计的准确性直接影响后续行动。尽管如此,许多定位方法通常提供点估计而不量化估计不确定性。为了解决这个问题,我们采用了共形预测(CP)-一个框架,提供统计上有效的预测区间与有限样本的保证,独立于数据分布。然而,常用的感应CP(ICP)方法需要大量的标记数据,这在定位设置中可能难以获得。为了减轻这一限制,我们将一个基于流形的定位方法,使用高斯过程回归(GPR),一个高效的Transductive CP(TCP)技术专门为GPR设计。我们证明了我们的方法在不同的声学条件下产生统计上有效的不确定性区间。摘要:We tackle the challenge of uncertainty quantification in the localization of a sound source within adverse acoustic environments. Estimating the position of the source is influenced by various factors such as noise and reverberation, leading to significant uncertainty. Quantifying this uncertainty is essential, particularly when localization outcomes impact critical decision-making processes, such as in robot audition, where the accuracy of location estimates directly influences subsequent actions. Despite this, many localization methods typically offer point estimates without quantifying the estimation uncertainty. To address this, we employ conformal prediction (CP)-a framework that delivers statistically valid prediction intervals with finite-sample guarantees, independent of the data distribution. However, commonly used Inductive CP (ICP) methods require a substantial amount of labeled data, which can be difficult to obtain in the localization setting. To mitigate this limitation, we incorporate a manifold-based localization method using Gaussian process regression (GPR), with an efficient Transductive CP (TCP) technique specifically designed for GPR. We demonstrate that our method generates statistically valid uncertainty intervals across different acoustic conditions.
【4】 Insights into the Incorporation of Signal Information in Binaural Signal Matching with Wearable Microphone Arrays
标题: 对使用可穿戴麦克风阵列进行双耳信号匹配时信号信息的融入的见解
作者:Ami Berger,Vladimir Tourbabin,Jacob Donley,Zamir Ben-Hur,Boaz Rafaely
链接:点击下载PDF文件
摘要:空间音频在诸如电话会议、娱乐和虚拟现实等应用中的日益普及导致了双耳再现方法的最近发展。然而,这些方法中只有少数非常适合可穿戴和移动阵列,这些阵列通常由少量麦克风组成。一种这样的方法是双耳信号匹配(BSM),其已经被示出为可穿戴阵列产生高质量的双耳信号。然而,BSM可能是次优的情况下,高直达混响比(DRR),因为它是基于扩散声场假设。为了克服这一局限性,以前的研究采用了声场模型,而不是扩散。然而,这一方法没有得到全面研究。本文广泛研究了两种基于BSM的方法设计的高DRR的情况下。该方法结合了由直达分量和混响分量组成的声场模型,对该方法进行了数学和仿真研究,最后通过听音测试进行了验证。结果表明,所提出的方法可以显着提高BSM的性能,特别是在源的方向上,而在其他方向上只呈现出微不足道的退化。此外,当源方向估计是不准确的,这些方法的性能下降到等于BSM,呈现出所需的鲁棒性质量。摘要:The increasing popularity of spatial audio in applications such as teleconferencing, entertainment, and virtual reality has led to the recent developments of binaural reproduction methods. However, only a few of these methods are well-suited for wearable and mobile arrays, which typically consist of a small number of microphones. One such method is binaural signal matching (BSM), which has been shown to produce high-quality binaural signals for wearable arrays. However, BSM may be suboptimal in cases of high direct-to-reverberant ratio (DRR) as it is based on the diffuse sound field assumption. To overcome this limitation, previous studies incorporated sound-field models other than diffuse. However, this approach was not studied comprehensively. This paper extensively investigates two BSM-based methods designed for high DRR scenarios. The methods incorporate a sound field model composed of direct and reverberant components.The methods are investigated both mathematically and using simulations, finally validated by a listening test. The results show that the proposed methods can significantly improve the performance of BSM , in particular in the direction of the source, while presenting only a negligible degradation in other directions. Furthermore, when source direction estimation is inaccurate, performance of these methods degrade to equal that of the BSM, presenting a desired robustness quality.
【5】 Dense-TSNet: Dense Connected Two-Stage Structure for Ultra-Lightweight Speech Enhancement
标题: Dense-TSNet:用于超轻量级语音增强的密集连接两级结构
作者:Zizhen Lin,Yuanle Li,Junyu Wang,Ruili Li
链接:点击下载PDF文件
摘要:语音增强的目的是提高语音质量和清晰度在嘈杂的环境。最近的进展集中在深度神经网络上,特别是采用两阶段(TS)架构来增强特征提取。然而,这些模型的复杂性和规模仍然很大,这限制了它们在资源有限的情况下的适用性。设计适用于边缘设备的模型面临着一系列挑战。窄的轻量级模型经常遇到性能瓶颈,由于不均匀的损失景观。此外,诸如Transformers或Mamba等高级运营商可能缺乏卷积神经网络(CNN)在现实世界部署中提供的实际适应性和效率。为了应对这些挑战,我们提出了一种创新的超轻量级语音增强网络Dense-TSNet。我们的方法采用了一种新的密集两阶段(密集TS)架构,与经典的两阶段架构相比,它确保了在后期训练阶段对目标函数进行更强大的细化。这导致改进的最终性能,解决了基线模型的早期收敛限制。我们还介绍了多视图凝视块(MVGB),它通过卷积神经网络(CNN)整合全局,通道和局部视角来增强特征提取。此外,我们讨论了损失函数的选择如何影响感知质量。Dense-TSNet具有约14 K参数的紧凑模型大小,表现出良好的性能,使其特别适合在资源受限的环境中部署。摘要:Speech enhancement aims to improve speech quality and intelligibility in noisy environments. Recent advancements have concentrated on deep neural networks, particularly employing the Two-Stage (TS) architecture to enhance feature extraction. However, the complexity and size of these models remain significant, which limits their applicability in resource-constrained scenarios. Designing models suitable for edge devices presents its own set of challenges. Narrow lightweight models often encounter performance bottlenecks due to uneven loss landscapes. Additionally, advanced operators such as Transformers or Mamba may lack the practical adaptability and efficiency that convolutional neural networks (CNNs) offer in real-world deployments. To address these challenges, we propose Dense-TSNet, an innovative ultra-lightweight speech enhancement network. Our approach employs a novel Dense Two-Stage (Dense-TS) architecture, which, compared to the classic Two-Stage architecture, ensures more robust refinement of the objective function in the later training stages. This leads to improved final performance, addressing the early convergence limitations of the baseline model. We also introduce the Multi-View Gaze Block (MVGB), which enhances feature extraction by incorporating global, channel, and local perspectives through convolutional neural networks (CNNs). Furthermore, we discuss how the choice of loss function impacts perceptual quality. Dense-TSNet demonstrates promising performance with a compact model size of around 14K parameters, making it particularly well-suited for deployment in resource-constrained environments.
【6】 Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
标题: 基于离散单元的掩蔽改善语音转换中的解纠缠
作者:Philip H. Lee,Ismail Rasim Ulgen,Berrak Sisman
备注:Accepted to IEEE SLT 2024
链接:点击下载PDF文件
摘要:语音转换(VC)的目的是修改说话人的身份,同时保持语言内容。通常,VC方法使用编码器-解码器架构,其中将说话者的身份从语言信息中分离出来至关重要。然而,在这些方法中使用的解纠缠方法是有限的,因为说话人的功能取决于语音内容的话语,妥协解纠缠。这种依赖性被基于注意力的方法放大了。为了解决这个问题,我们在扬声器编码之前的输入中引入了一种新的掩蔽机制,掩蔽了与音素类高度对应的某些离散语音单元。我们的工作旨在通过限制对某些语音信息的访问来减少说话人特征的语音依赖。此外,由于我们的方法是在输入级,它适用于任何基于VC框架的编码器-解码器。我们的方法提高了多个VC方法的解纠缠和转换性能,显示出显着的有效性,特别是在基于注意力的方法,与44%的相对改善客观可懂度。摘要:Voice conversion (VC) aims to modify the speaker's identity while preserving the linguistic content. Commonly, VC methods use an encoder-decoder architecture, where disentangling the speaker's identity from linguistic information is crucial. However, the disentanglement approaches used in these methods are limited as the speaker features depend on the phonetic content of the utterance, compromising disentanglement. This dependency is amplified with attention-based methods. To address this, we introduce a novel masking mechanism in the input before speaker encoding, masking certain discrete speech units that correspond highly with phoneme classes. Our work aims to reduce the phonetic dependency of speaker features by restricting access to some phonetic information. Furthermore, since our approach is at the input level, it is applicable to any encoder-decoder based VC framework. Our approach improves disentanglement and conversion performance across multiple VC methods, showing significant effectiveness, particularly in attention-based method, with 44% relative improvement in objective intelligibility.
【7】 M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
标题: M-BEST-PQ:智能眼镜的多通道语音基础模型
作者:Yufeng Yang,Desh Raj,Ju Lin,Niko Moritz,Junteng Jia,Gil Keren,Egor Lakomkin,Yiteng Huang,Jacob Donley,Jay Mahadeokar,Ozlem Kalinli
备注:In submission to IEEE ICASSP 2025
链接:点击下载PDF文件
摘要:多通道可穿戴设备(如智能眼镜)的日益普及,导致了定向语音识别和增强听力等应用的激增。然而,目前解决这些任务的方法使用独立训练的模型,这可能不会从大量未标记的数据中受益。在本文中,我们提出了M-BEST-RQ,这是智能眼镜的第一个多通道语音基础模型,旨在利用大规模自监督学习(SSL)的阵列几何不可知方法。虽然之前的多通道语音SSL工作仅在模拟设置上进行了评估,但我们策划了一套真实的下游任务来评估我们的模型,即(i)会话自动语音识别(ASR),(ii)球形主动源定位,以及(iii)眼镜佩戴者语音活动检测,这些任务来自MMCSG和EasyCom数据集。我们证明了通用的M-BEST-RQ编码器能够在所有任务中匹配或超越监督模型。特别是对于会话式ASR任务,仅使用8小时的标记语音,我们的模型优于在2000小时的标记数据上训练的监督ASR基线,这证明了我们方法的有效性。摘要:The growing popularity of multi-channel wearable devices, such as smart glasses, has led to a surge of applications such as targeted speech recognition and enhanced hearing. However, current approaches to solve these tasks use independently trained models, which may not benefit from large amounts of unlabeled data. In this paper, we propose M-BEST-RQ, the first multi-channel speech foundation model for smart glasses, which is designed to leverage large-scale self-supervised learning (SSL) in an array-geometry agnostic approach. While prior work on multi-channel speech SSL only evaluated on simulated settings, we curate a suite of real downstream tasks to evaluate our model, namely (i) conversational automatic speech recognition (ASR), (ii) spherical active source localization, and (iii) glasses wearer voice activity detection, which are sourced from the MMCSG and EasyCom datasets. We show that a general-purpose M-BEST-RQ encoder is able to match or surpass supervised models across all tasks. For the conversational ASR task in particular, using only 8 hours of labeled speech, our model outperforms a supervised ASR baseline that is trained on 2000 hours of labeled data, which demonstrates the effectiveness of our approach.
【8】 Prosodic Parameter Manipulation in TTS generated speech for Controlled Speech Generation
标题: 用于受控语音生成的TTC生成语音中的韵律参数操纵
作者:Podakanti Satyajith Chary
备注:9 pages, 4 figures, International Summer School on NLP 2024 at IIIT Hyderabad
链接:点击下载PDF文件
摘要:本文探讨了文本到语音(TTS)系统中的韵律参数的操作,以实现受控的语音生成。通过利用先进的语音处理技术,我们将TTS生成的音频与人类录制的语音进行比较,以分析音调,持续时间和能量的差异。使用PyWorld和Librosa等工具提取关键特征,然后调整这些特征以符合自然人类语音的韵律特征。修改后的功能进行合成,产生增强的TTS输出,更密切地反映人类语音的自然韵律。该方法旨在通过提供一个精确的韵律参数调整框架来增强TTS系统的自然度和表现力。我们的方法包括特征提取,韵律操作和合成,然后进行全面的评估,以确保与人类语音模式的一致性。研究结果表明,韵律参数操纵的可行性和有效性的控制语音生成,突出其潜力显着提高TTS的应用。摘要:This paper explores the manipulation of prosodic parameters in Text-to-Speech (TTS) systems to achieve controlled speech generation. By leveraging advanced speech processing techniques, we compare TTS-generated audio with human-recorded speech to analyze differences in pitch, duration, and energy. Key features are extracted using tools like PyWorld and Librosa, which are then adjusted to align with the prosodic characteristics of natural human speech. The modified features undergo synthesis, producing enhanced TTS outputs that more closely mirror the natural prosody of human speech. This approach aims to enhance the naturalness and expressiveness of TTS systems by providing a framework for precise prosodic parameter adjustments. Our methodology involves feature extraction, prosodic manipulation, and synthesis, followed by comprehensive evaluations to ensure consistency with human speech patterns. The findings demonstrate the feasibility and effectiveness of prosodic parameter manipulation for controlled speech generation, highlighting its potential to significantly improve TTS applications.
【9】 The Unreliability of Acoustic Systems in Alzheimer's Speech Datasets with Heterogeneous Recording Conditions
标题: 不同记录条件下阿尔茨海默氏症语音数据集中声学系统的不可靠性
作者:Lara Gauder,Pablo Riera,Andrea Slachevsky,Gonzalo Forno,Adolfo M. Garcia,Luciana Ferrer
备注:5 pages, 1 figure, 1 table
链接:点击下载PDF文件
摘要:自动语音分析是检测阿尔茨海默病(AD)早期标志物的一种蓬勃发展的方法。然而,大多数AD数据集中的记录条件是异质的,患者和对照通常在不同的声学设置中进行评估。虽然这对于基于语音转录或从手动对齐获得的特征的分析来说不是问题,但它确实对声学特征的有效性产生了严重的怀疑,这些声学特征受到采集条件的强烈影响。我们在ADreSSo数据集中研究了这个问题,该数据集来自广泛使用的Pitt语料库。我们表明,基于两个声学特征,MFCC和Wav2vec 2.0嵌入的系统,可以区分AD患者与控制上述机会的性能时,只使用非语音部分的音频信号。我们在一个单独的西班牙语使用者数据集中复制了这一发现。因此,在这些数据集中,可以通过记录条件部分地预测类别。我们的研究结果是对使用声学系统识别基于非标准化记录的患者的警告。我们建议,声学异构数据集的痴呆症研究应(a)分析仅使用转录本或其他功能来自手动注释,或(b)取代收集的数据集,严格控制的声学条件。摘要:Automated speech analysis is a thriving approach to detect early markers of Alzheimer's disease (AD). Yet, recording conditions in most AD datasets are heterogeneous, with patients and controls often evaluated in different acoustic settings. While this is not a problem for analyses based on speech transcription or features obtained from manual alignment, it does cast serious doubts on the validity of acoustic features, which are strongly influenced by acquisition conditions. We examined this issue in the ADreSSo dataset, derived from the widely used Pitt corpus. We show that systems based on two acoustic features, MFCCs and Wav2vec 2.0 embeddings, can discriminate AD patients from controls with above-chance performance when using only the non-speech part of the audio signals. We replicated this finding in a separate dataset of Spanish speakers. Thus, in these datasets, the class can be partly predicted by recording conditions. Our results are a warning against the use of acoustic systems for identifying patients based on non-standardized recordings. We propose that acoustically heterogeneous datasets for dementia studies should be either (a) analyzed using only transcripts or other features derived from manual annotations, or (b) replaced by datasets collected with strictly controlled acoustic conditions.
【10】 Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models
标题: Takin:一群高质量Zero-Shot语音生成模型
作者:EverestAI,:,Sijin Chen,Yuan Feng,Laipeng He,Tianwei He,Wendi He,Yanni Hu,Bin Lin,Yiting Lin,Pengfei Tan,Chengwei Tian,Chen Wang,Zhicheng Wang,Ruoye Xie,Jingjing Yin,Jianhao Ye,Jixun Yao,Quanlei Yan,Yuguang Yang
链接:点击下载PDF文件
摘要:随着大数据和大语言模型时代的到来,zero-shot个性化快速定制已经成为一个显著的趋势。在这份报告中,我们介绍了Takin AudioLLM,一系列技术和模型,主要包括Takin TTS,Takin VC和Takin Morphing,专门为有声读物制作而设计。这些模型能够进行zero-shot语音生成,生成与真实人类语音几乎无法区分的高质量语音,并便于个人根据自己的需求定制语音内容。具体来说,我们首先介绍Takin TTS,一种神经编解码器语言模型,它建立在增强的神经语音编解码器和多任务训练框架之上,能够以zero-shot方式生成高保真自然语音。对于Takin VC,我们提倡一种有效的内容和音色联合建模方法来提高说话人相似性,同时提倡一种基于条件流匹配的解码器来进一步增强其自然度和表现力。最后,我们提出了Takin Morphing系统,该系统具有高度解耦和先进的音色和韵律建模方法,使个人能够以精确和可控的方式自定义语音生产。大量的实验验证了我们的Takin AudioLLM系列模型的有效性和鲁棒性。有关详细演示,请参阅https: takinaudiollm.github.io。摘要:With the advent of the big data and large language model era, zero-shot personalized rapid customization has emerged as a significant trend. In this report, we introduce Takin AudioLLM, a series of techniques and models, mainly including Takin TTS, Takin VC, and Takin Morphing, specifically designed for audiobook production. These models are capable of zero-shot speech production, generating high-quality speech that is nearly indistinguishable from real human speech and facilitating individuals to customize the speech content according to their own needs. Specifically, we first introduce Takin TTS, a neural codec language model that builds upon an enhanced neural speech codec and a multi-task training framework, capable of generating high-fidelity natural speech in a zero-shot way. For Takin VC, we advocate an effective content and timbre joint modeling approach to improve the speaker similarity, while advocating for a conditional flow matching based decoder to further enhance its naturalness and expressiveness. Last, we propose the Takin Morphing system with highly decoupled and advanced timbre and prosody modeling approaches, which enables individuals to customize speech production with their preferred timbre and prosody in a precise and controllable manner. Extensive experiments validate the effectiveness and robustness of our Takin AudioLLM series models. For detailed demos, please refer to https: takinaudiollm.github.io.
【11】 WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
标题: WMCCodec:具有深度水印用于真实性验证的端到端神经语音编解码器
作者:Junzuo Zhou,Jiangyan Yi,Yong Ren,Jianhua Tao,Tao Wang,Chu Yuan Zhang
链接:点击下载PDF文件
摘要:语音欺骗的最新进展需要神经语音编解码器中更强的验证机制来确保真实性。目前的方法在压缩前嵌入数字水印,并从重建的语音中提取数字水印进行验证,但面临着水印和编解码器的单独训练过程,以及跨模态信息集成不足等限制,导致水印不可感知性,提取精度和容量降低。为了解决这些问题,我们提出了WMCodec,第一个神经语音编解码器联合训练压缩重建和水印嵌入提取在一个端到端的方式,优化不可感知性和可提取的水印。此外,我们还设计了一个迭代的注意力印记单元(AIU)来更深层次地融合水印和语音的特征,降低量化噪声对水印的影响。实验结果表明,WMCodec优于AudioSeal与Encodec在大多数质量指标的水印不可见性和一致超过AudioSeal与Encodec和增强TraceableSpeech的水印提取精度。WMCodec在6kbps带宽下,水印容量为16bps,在常见攻击下提取准确率保持在99%以上,具有较强的鲁棒性。摘要:Recent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training processes for the watermark and codec, and insufficient cross-modal information integration, leading to reduced watermark imperceptibility, extraction accuracy, and capacity. To address these issues, we propose WMCodec, the first neural speech codec to jointly train compression-reconstruction and watermark embedding-extraction in an end-to-end manner, optimizing both imperceptibility and extractability of the watermark. Furthermore, We design an iterative Attention Imprint Unit (AIU) for deeper feature integration of watermark and speech, reducing the impact of quantization noise on the watermark. Experimental results show WMCodec outperforms AudioSeal with Encodec in most quality metrics for watermark imperceptibility and consistently exceeds both AudioSeal with Encodec and reinforced TraceableSpeech in extraction accuracy of watermark. At bandwidth of 6 kbps with a watermark capacity of 16 bps, WMCodec maintains over 99% extraction accuracy under common attacks, demonstrating strong robustness.
【12】 Pareto Data Framework: Steps Towards Resource-Efficient Decision Making Using Minimum Viable Data (MVD)
标题: 帕累托数据框架:使用最少可行数据(MVD)实现资源高效决策的步骤
作者:Tashfain Ahmed,Josh Siegel
链接:点击下载PDF文件
摘要:本文介绍了Pareto数据框架,这是一种识别和选择在受限平台(如嵌入式系统,移动设备和物联网(IoT)设备)上启用机器学习应用程序所需的最小可行数据(MVD)的方法。我们证明,战略性的数据减少可以保持高性能,同时显着降低带宽,能源,计算和存储成本。该框架确定了最小可行数据(MVD),以在资源受限的环境中优化效率,而不会牺牲性能。它解决了物联网应用中常见的低效做法,例如传感器的过度配置和过精度,以及信号的过采样,为最佳传感器选择,信号提取和传输以及数据表示提出了可扩展的解决方案。实验方法证明了在下采样、量化和截断后有效的声学数据表征,以模拟保真度降低的传感器和网络和存储约束;结果表明,性能可以保持高达95%,采样率降低75%,位深度和剪辑长度减少50%,这意味着大量的成本和资源减少。这些发现对约束系统的设计和开发具有影响。本文还讨论了该框架的更广泛影响,包括在物联网应用和农业、交通运输和制造业等领域实现先进人工智能技术民主化的潜力,以改善数据驱动的洞察力的可访问性并增加其效益。摘要:This paper introduces the Pareto Data Framework, an approach for identifying and selecting the Minimum Viable Data (MVD) required for enabling machine learning applications on constrained platforms such as embedded systems, mobile devices, and Internet of Things (IoT) devices. We demonstrate that strategic data reduction can maintain high performance while significantly reducing bandwidth, energy, computation, and storage costs. The framework identifies Minimum Viable Data (MVD) to optimize efficiency across resource-constrained environments without sacrificing performance. It addresses common inefficient practices in an IoT application such as overprovisioning of sensors and overprecision, and oversampling of signals, proposing scalable solutions for optimal sensor selection, signal extraction and transmission, and data representation. An experimental methodology demonstrates effective acoustic data characterization after downsampling, quantization, and truncation to simulate reduced-fidelity sensors and network and storage constraints; results shows that performance can be maintained up to 95 % with sample rates reduced by 75 % and bit depths and clip length reduced by 50 % which translates into substantial cost and resource reduction. These findings have implications on the design and development of constrained systems. The paper also discusses broader implications of the framework, including the potential to democratize advanced AI technologies across IoT applications and sectors such as agriculture, transportation, and manufacturing to improve access and multiply the benefits of data-driven insights.
【13】 ASR Benchmarking: Need for a More Representative Conversational Dataset
标题: ASB基准:需要更具代表性的对话数据集
作者:Gaurav Maheshwari,Dmitry Ivanov,Théo Johannet,Kevin El Haddad
链接:点击下载PDF文件
摘要:自动语音识别(ASR)系统在LibriSpeech和Fleurs等广泛使用的基准测试中取得了卓越的性能。然而,这些基准不能充分反映现实世界对话环境的复杂性,其中语音通常是非结构化的,并且包含诸如停顿、中断和不同口音的不流利。在这项研究中,我们介绍了一个多语言的会话数据集,来自TalkBank,由成人之间的非结构化电话交谈。我们的研究结果显示,在会话设置中进行测试时,各种最先进的ASR模型的性能显着下降。此外,我们观察到的单词错误率和语音不流利的存在之间的相关性,突出了更现实的,会话ASR基准的迫切需要。摘要:Automatic Speech Recognition (ASR) systems have achieved remarkable performance on widely used benchmarks such as LibriSpeech and Fleurs. However, these benchmarks do not adequately reflect the complexities of real-world conversational environments, where speech is often unstructured and contains disfluencies such as pauses, interruptions, and diverse accents. In this study, we introduce a multilingual conversational dataset, derived from TalkBank, consisting of unstructured phone conversation between adults. Our results show a significant performance drop across various state-of-the-art ASR models when tested in conversational settings. Furthermore, we observe a correlation between Word Error Rate and the presence of speech disfluencies, highlighting the critical need for more realistic, conversational ASR benchmarks.
【14】 Data Efficient Acoustic Scene Classification using Teacher-Informed Confusing Class Instruction
标题: 使用教师知情的混淆课堂教学进行数据高效的声学场景分类
作者:Jin Jie Sean Yeo,Ee-Leng Tan,Jisheng Bai,Santi Peksi,Woon-Seng Gan
备注:5 pages, 3 figures
链接:点击下载PDF文件
摘要:在本技术报告中,我们描述了SNTL-NTU团队提交的任务1数据高效低复杂度声学场景分类的声学场景和事件的检测和分类(DCASE)2024挑战。引入了三个系统来处理不同大小的训练分割。对于小的训练分割,我们探索了通过减少基本通道的数量来降低所提供的基线模型的复杂性。我们以混合的形式引入数据增强,以增加训练样本的多样性。对于较大的训练分割,我们使用FocusNet为多个Patchout faSt Spectrogram Transformer(PaSST)模型和在原始采样率44.1 kHz上训练的基线模型的集合提供混淆的类信息。我们使用知识蒸馏将集成模型提取为基线学生模型。在TAU Urban Acoustic Scene 2022 Mobile开发数据集上训练系统,三个系统的平均测试准确率分别为(62.21,59.82,56.81,53.03,47.97)%(100,50,25,10,5)%。摘要:In this technical report, we describe the SNTL-NTU team's submission for Task 1 Data-Efficient Low-Complexity Acoustic Scene Classification of the detection and classification of acoustic scenes and events (DCASE) 2024 challenge. Three systems are introduced to tackle training splits of different sizes. For small training splits, we explored reducing the complexity of the provided baseline model by reducing the number of base channels. We introduce data augmentation in the form of mixup to increase the diversity of training samples. For the larger training splits, we use FocusNet to provide confusing class information to an ensemble of multiple Patchout faSt Spectrogram Transformer (PaSST) models and baseline models trained on the original sampling rate of 44.1 kHz. We use Knowledge Distillation to distill the ensemble model to the baseline student model. Training the systems on the TAU Urban Acoustic Scene 2022 Mobile development dataset yielded the highest average testing accuracy of (62.21, 59.82, 56.81, 53.03, 47.97)% on split (100, 50, 25, 10, 5)% respectively over the three systems.
【15】 Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
标题: 使用Frozen wav2vec 2.0进行虚假音频检测的混合专家融合
作者:Zhiyong Wang,Ruibo Fu,Zhengqi Wen,Jianhua Tao,Xiaopeng Wang,Yuankun Xie,Xin Qi,Shuchen Shi,Yi Lu,Yukun Liu,Chenxing Li,Xuefei Liu,Guanjun Li
备注:submitted to ICASSP2025
链接:点击下载PDF文件
摘要:语音合成技术对说话人确认系统构成了严重威胁。 目前,最有效的虚假音频检测方法利用预训练模型,并从预训练模型的各个层中集成特征进一步提高检测性能。 然而,大多数先前提出的融合方法需要微调的预训练模型,导致训练时间过长,并阻碍模型迭代时,面对新的语音合成技术。 针对这一问题,提出了一种基于混合专家的特征融合方法,该方法在基于最后一层特征的门控网络的指导下,从层特征中提取并融合与虚假音频检测相关的特征,同时冻结预训练模型。 在ASVspoof2019和ASVspoof2021数据集上进行的实验表明,与需要微调的方法相比,所提出的方法具有竞争力的性能。摘要:Speech synthesis technology has posed a serious threat to speaker verification systems. Currently, the most effective fake audio detection methods utilize pretrained models, and integrating features from various layers of pretrained model further enhances detection performance. However, most of the previously proposed fusion methods require fine-tuning the pretrained models, resulting in excessively long training times and hindering model iteration when facing new speech synthesis technology. To address this issue, this paper proposes a feature fusion method based on the Mixture of Experts, which extracts and integrates features relevant to fake audio detection from layer features, guided by a gating network based on the last layer feature, while freezing the pretrained model. Experiments conducted on the ASVspoof2019 and ASVspoof2021 datasets demonstrate that the proposed method achieves competitive performance compared to those requiring fine-tuning.
【16】 M2R-Whisper: Multi-stage and Multi-scale Retrieval Augmentation for Enhancing Whisper
标题: M2 R-Whisper:用于增强Whisper的多阶段、多规模检索增强
作者:Jiaming Zhou,Shiwan Zhao,Jiabei He,Hui Wang,Wenjia Zeng,Yong Chen,Haoqin Sun,Aobo Kong,Yong Qin
链接:点击下载PDF文件
摘要:像OpenAI的Whisper这样的最先进的模型在多语言自动语音识别(ASR)方面表现出了强大的性能,但它们在准确识别不同的次方言方面仍然面临挑战。在本文中,我们提出了M2 R-whisper,一种新的多阶段和多尺度检索增强方法,旨在提高低资源设置中的ASR性能。基于上下文学习(ICL)和检索增强技术的原理,我们的方法在预处理阶段采用标记级ICL来利用上下文信息,同时将标记级k最近邻(kNN)检索作为后处理步骤,以进一步细化最终的输出分布。通过协同结合文本级和标记级检索策略,M2 R-whisper有效地减少了各种类型的识别错误。在普通话和亚方言数据集(包括AISHELL-1和KeSpeech)上进行的实验表明,ASR准确性有了实质性的提高,所有这些都是在没有任何参数更新的情况下实现的。摘要:State-of-the-art models like OpenAI's Whisper exhibit strong performance in multilingual automatic speech recognition (ASR), but they still face challenges in accurately recognizing diverse subdialects. In this paper, we propose M2R-whisper, a novel multi-stage and multi-scale retrieval augmentation approach designed to enhance ASR performance in low-resource settings. Building on the principles of in-context learning (ICL) and retrieval-augmented techniques, our method employs sentence-level ICL in the pre-processing stage to harness contextual information, while integrating token-level k-Nearest Neighbors (kNN) retrieval as a post-processing step to further refine the final output distribution. By synergistically combining sentence-level and token-level retrieval strategies, M2R-whisper effectively mitigates various types of recognition errors. Experiments conducted on Mandarin and subdialect datasets, including AISHELL-1 and KeSpeech, demonstrate substantial improvements in ASR accuracy, all achieved without any parameter updates.
【17】 DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
标题: DPI-TTC:用于文本到语音中快速收敛和风格时态建模的定向补丁交互
作者:Xin Qi,Ruibo Fu,Zhengqi Wen,Tao Wang,Chunyu Qiang,Jianhua Tao,Chenxing Li,Yi Lu,Shuchen Shi,Zhiyong Wang,Xiaopeng Wang,Yuankun Xie,Yukun Liu,Xuefei Liu,Guanjun Li
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:近年来,语音扩散模型发展迅速。除了广泛使用的U-Net架构,基于变压器的模型,如扩散Transformer(DiT)也得到了关注。然而,目前的DiT语音模型将Mel频谱图视为一般图像,忽略了语音的特定声学特性。为了解决这些限制,我们提出了一种称为文本到语音定向补丁交互(DPI-TTS)的方法,该方法建立在DiT的基础上,可以在不影响准确性的情况下实现快速训练。值得注意的是,DPI-TTS采用了一种从低到高的频率,逐帧渐进式推理方法,该方法与声学特性更紧密地结合在一起,从而增强了所生成语音的自然度。此外,我们引入了一个细粒度的风格时间建模方法,进一步提高说话人风格的相似性。实验结果表明,我们的方法将训练速度提高了近2倍,并显着优于基线模型。摘要:In recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models.
【18】 Spin Detection Using Racket Bounce Sounds in Table Tennis
标题: 乒乓球运动中利用球拍弹跳声进行旋转检测
作者:Thomas Gossard,Julian Schmalzl,Andreas Ziegler,Andreas Zell
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:虽然乒乓球运动员主要依靠视觉线索,但声音提供了有价值的信息。当球撞击球拍时产生的声音可以帮助预测球的轨迹,特别是确定旋转。虽然职业球员可以通过这些听觉线索来区分旋转,但未经训练的球员往往不会注意到。在本文中,我们证明了不同的球拍产生不同的声音,这可以用来识别球拍类型。此外,我们表明,由球拍产生的声音可以表明是否旋转被施加到球,或没有。为了实现这一目标,我们创建了一个综合数据集,其中包含来自10种球拍配置的反弹声音,每种配置都对球施加不同的旋转。为了实现毫秒级的时间精度,我们首先检测可能对应于乒乓球反弹的高频峰值。然后,我们使用基于CNN的分类器来细化这些结果,该分类器可以准确预测所使用的球拍类型和是否应用了旋转。摘要:While table tennis players primarily rely on visual cues, sound provides valuable information. The sound generated when the ball strikes the racket can assist in predicting the ball's trajectory, especially in determining the spin. While professional players can distinguish spin through these auditory cues, they often go unnoticed by untrained players. In this paper, we demonstrate that different rackets produce distinct sounds, which can be used to identify the racket type. In addition, we show that the sound generated by the racket can indicate whether spin was applied to the ball, or not. To achieve this, we created a comprehensive dataset featuring bounce sounds from 10 racket configurations, each applying various spins to the ball. To achieve millisecond level temporal accuracy, we first detect high frequency peaks that may correspond to table tennis ball bounces. We then refine these results using a CNN based classifier that accurately predicts both the type of racket used and whether spin was applied.
【19】 METEOR: Melody-aware Texture-controllable Symbolic Orchestral Music Generation
标题: METEOR:旋律感知、文本可控的象征性中音音乐生成
作者:Dinh-Viet-Toan Le,Yi-Hsuan Yang
备注:this https URL
链接:点击下载PDF文件
摘要:西方音乐的特点通常是谐音织体,其中音乐内容可以组织成旋律和伴奏。在管弦乐中,特别是,作曲家可以选择特定的特征,为每个乐器的部分内的伴奏,同时也需要适应的旋律,以适应能力的仪器执行it.In这项工作中,我们提出了METEOR,一个模型的旋律感知纹理可控的管弦乐音乐生成。该模型执行符号化的多轨音乐风格转移,重点是旋律保真度。我们允许酒吧和轨道水平的控制性的伴奏与各种纹理属性,同时保持一个谐音纹理。我们表明,该模型可以实现类似于强基线的可控性性能,同时大大提高旋律保真度。摘要:Western music is often characterized by a homophonic texture, in which the musical content can be organized into a melody and an accompaniment. In orchestral music, in particular, the composer can select specific characteristics for each instrument's part within the accompaniment, while also needing to adapt the melody to suit the capabilities of the instruments performing it. In this work, we propose METEOR, a model for Melody-aware Texture-controllable Orchestral music generation. This model performs symbolic multi-track music style transfer with a focus on melodic fidelity. We allow bar- and track-level controllability of the accompaniment with various textural attributes while keeping a homophonic texture. We show that the model can achieve controllability performances similar to strong baselines while greatly improve melodic fidelity.
【20】 SALT: Standardized Audio event Label Taxonomy
标题: SALT:标准化音频事件标签分类
作者:Paraskevas Stamatiadis,Michel Olvera,Slim Essid
Journal-ref:DCASE, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
摘要:机器监听系统通常依赖于固定的分类来组织和标记音频数据,这是训练和评估深度神经网络(DNN)和其他监督算法的关键。然而,这样的分类法面临着重大的限制:它们由依赖于应用程序的预定义类别组成,这阻碍了新的或不同的声音的整合,并且由于不一致的标签标准而表现出有限的跨数据集兼容性。为了克服这些限制,我们引入了SALT:标准化音频事件标签分类。AudioSet的本体论的层次结构的基础上,我们的分类扩展,并在24个公开可用的环境声音数据集的标签,允许从不同的数据集的类标签映射到一个统一的系统。我们的提案附带了一个新的Python包,旨在导航和利用这种分类法,简化跨数据集标签搜索和分层探索。值得注意的是,我们的软件包允许轻松地从不同来源聚合数据,因此可以轻松地对组合数据集进行实验。摘要:Machine listening systems often rely on fixed taxonomies to organize and label audio data, key for training and evaluating deep neural networks (DNNs) and other supervised algorithms. However, such taxonomies face significant constraints: they are composed of application-dependent predefined categories, which hinders the integration of new or varied sounds, and exhibits limited cross-dataset compatibility due to inconsistent labeling standards. To overcome these limitations, we introduce SALT: Standardized Audio event Label Taxonomy. Building upon the hierarchical structure of AudioSet's ontology, our taxonomy extends and standardizes labels across 24 publicly available environmental sound datasets, allowing the mapping of class labels from diverse datasets to a unified system. Our proposal comes with a new Python package designed for navigating and utilizing this taxonomy, easing cross-dataset label searching and hierarchical exploration. Notably, our package allows effortless data aggregation from diverse sources, hence easy experimentation with combined datasets.
【21】 Simulating Native Speaker Shadowing for Nonnative Speech Assessment with Latent Speech Representations
标题: 模拟母语说话者阴影以进行具有潜在语音表示的非母语语音评估
作者:Haopeng Geng,Daisuke Saito,Minematsu Nobuaki
备注:Submitted to ICASSP2025 Demo available: this https URL
链接:点击下载PDF文件
摘要:语音清晰度评价是计算机辅助语言学习系统中的一项重要任务。传统方法通常依赖于由自动语音识别(ASR)提供的单词错误率(WER)作为可懂度分数。然而,由于人类语音识别(HSR)和ASR之间的显著差异,这种方法具有显著的局限性。一个有希望的替代方法是让母语(L1)说话者参与非母语(L2)说话者所说的话。L1说话者的阴影话语中的故障或错误发音可以作为评估L2语音可懂度的指标。在这项研究中,我们提出了一个语音生成系统,该系统使用语音转换(VC)技术和潜在语音表示来模拟L1阴影过程。我们的实验结果表明,该方法有效地复制了L1的阴影过程,提供了一个创新的工具来评估L2语音可懂度。值得注意的是,利用自监督语音表示(S3R)的系统在语言准确性和自然性方面与真实的L1阴影话语具有更高的相似度。摘要:Evaluating speech intelligibility is a critical task in computer-aided language learning systems. Traditional methods often rely on word error rates (WER) provided by automatic speech recognition (ASR) as intelligibility scores. However, this approach has significant limitations due to notable differences between human speech recognition (HSR) and ASR. A promising alternative is to involve a native (L1) speaker in shadowing what nonnative (L2) speakers say. Breakdowns or mispronunciations in the L1 speaker's shadowing utterance can serve as indicators for assessing L2 speech intelligibility. In this study, we propose a speech generation system that simulates the L1 shadowing process using voice conversion (VC) techniques and latent speech representations. Our experimental results demonstrate that this method effectively replicates the L1 shadowing process, offering an innovative tool to evaluate L2 speech intelligibility. Notably, systems that utilize self-supervised speech representations (S3R) show a higher degree of similarity to real L1 shadowing utterances in both linguistic accuracy and naturalness.
【22】 DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
标题: 检测:利用对象信息增强视听表示学习
作者:Shota Nakada,Taichi Nishimura,Hokuto Munakata,Masayoshi Kondo,Tatsuya Komatsu
备注:under review
链接:点击下载PDF文件
摘要:当前的视听表示学习可以捕获粗略的对象类别(例如,“动物”和“乐器”),但它缺乏识别细粒度细节的能力,例如动物和乐器中的“狗”和“长笛”等特定类别。为了解决这个问题,我们引入了DETECTORM,这是一种利用对象信息增强视听表示学习的方法。我们的主要想法是向现有的对比视听掩蔽自动编码器引入视听标签预测损失,以增强其对象感知能力。为了避免昂贵的手动注释,我们使用最先进的语言音频模型和对象检测器从音频和视觉输入中准备对象标签。我们使用VGGSound和AudioSet20K数据集评估视听检索和分类的方法。我们的方法在recall@10方面分别实现了音频到视频和视频到音频检索的+1.5%和+1.2%的改进,并且在视听分类方面提高了+0.6%的准确性。摘要:Current audio-visual representation learning can capture rough object categories (e.g., animals'' and instruments''), but it lacks the ability to recognize fine-grained details, such as specific categories like dogs'' and flutes'' within animals and instruments. To address this issue, we introduce DETECLAP, a method to enhance audio-visual representation learning with object information. Our key idea is to introduce an audio-visual label prediction loss to the existing Contrastive Audio-Visual Masked AutoEncoder to enhance its object awareness. To avoid costly manual annotations, we prepare object labels from both audio and visual inputs using state-of-the-art language-audio models and object detectors. We evaluate the method of audio-visual retrieval and classification using the VGGSound and AudioSet20K datasets. Our method achieves improvements in recall@10 of +1.5% and +1.2% for audio-to-visual and visual-to-audio retrieval, respectively, and an improvement in accuracy of +0.6% for audio-visual classification.
【23】 Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation
标题: 从粗到细:通过多尺度语音编码和生成改进神经编解码器语言模型
作者:Haohan Guo,Fenglong Xie,Dongchao Yang,Xixin Wu,Helen Meng
链接:点击下载PDF文件
摘要:神经编解码语言模型(CLM)在文语转换(TTS)合成中表现出了卓越的性能。然而,受“近因偏差”的困扰,CLM缺乏对更高时间尺度上的粗粒度信息的足够关注,经常产生不自然甚至无法理解的语音。这项工作提出了CoFi-Speech,一种从粗到细的CLM-TTS方法,采用多尺度语音编码和生成来解决这个问题。我们训练多尺度神经编解码器,CoFi-Codec,将语音编码成多尺度离散表示,包括具有不同时间分辨率的多个令牌序列。然后,我们提出了CoFi-LM,它可以在两种模式下生成这种表示:基于单LM的规模链生成和基于多LM的规模堆栈生成。实验结果表明,在zero-shot TTS中,CoFi-Speech在自然度和说话人相似度方面明显优于单尺度基线系统。多尺度编码的分析证明了CoFi-Codec在学习多尺度离散语音表示的同时保持高质量语音重建的有效性。从粗到细的多尺度生成,特别是对于尺度堆栈方法,也被验证为追求高质量的TTS神经编解码器语言模型的关键方法。摘要:The neural codec language model (CLM) has demonstrated remarkable performance in text-to-speech (TTS) synthesis. However, troubled by recency bias", CLM lacks sufficient attention to coarse-grained information at a higher temporal scale, often producing unnatural or even unintelligible speech. This work proposes CoFi-Speech, a coarse-to-fine CLM-TTS approach, employing multi-scale speech coding and generation to address this issue. We train a multi-scale neural codec, CoFi-Codec, to encode speech into a multi-scale discrete representation, comprising multiple token sequences with different time resolutions. Then, we propose CoFi-LM that can generate this representation in two modes: the single-LM-based chain-of-scale generation and the multiple-LM-based stack-of-scale generation. In experiments, CoFi-Speech significantly outperforms single-scale baseline systems on naturalness and speaker similarity in zero-shot TTS. The analysis of multi-scale coding demonstrates the effectiveness of CoFi-Codec in learning multi-scale discrete speech representations while keeping high-quality speech reconstruction. The coarse-to-fine multi-scale generation, especially for the stack-of-scale approach, is also validated as a crucial approach in pursuing a high-quality neural codec language model for TTS.
【24】 Preference Tuning with Human Feedback on Language, Speech, and Vision Tasks: A Survey
标题: 根据人类对语言、言语和视觉任务的反馈调整偏好:一项调查
作者:Genta Indra Winata,Hanyang Zhao,Anirban Das,Wenpin Tang,David D. Yao,Shi-Xiong Zhang,Sambit Sahu
备注:Survey paper
链接:点击下载PDF文件
摘要:偏好调整是将深度生成模型与人类偏好相匹配的关键过程。这项调查提供了一个全面的概述偏好调整和人类反馈的整合的最新进展。本文分为三个主要部分:1)引言和预备:介绍强化学习框架、偏好调整任务、模型和各种模式的数据集:语言、言语和视觉,以及不同的政策方法,2)深入研究每种偏好调整方法:对偏好调整中使用的方法的详细分析,以及3)应用、讨论和未来方向:探索偏好调整在下游任务中的应用,包括不同模式的评估方法,并展望了未来的研究方向。我们的目标是介绍偏好调整和模型对齐的最新方法,增强研究人员和从业人员对这一领域的理解。我们希望鼓励这一领域的进一步参与和创新。摘要:Preference tuning is a crucial process for aligning deep generative models with human preferences. This survey offers a thorough overview of recent advancements in preference tuning and the integration of human feedback. The paper is organized into three main sections: 1) introduction and preliminaries: an introduction to reinforcement learning frameworks, preference tuning tasks, models, and datasets across various modalities: language, speech, and vision, as well as different policy approaches, 2) in-depth examination of each preference tuning approach: a detailed analysis of the methods used in preference tuning, and 3) applications, discussion, and future directions: an exploration of the applications of preference tuning in downstream tasks, including evaluation methods for different modalities, and an outlook on future research directions. Our objective is to present the latest methodologies in preference tuning and model alignment, enhancing the understanding of this field for researchers and practitioners. We hope to encourage further engagement and innovation in this area.
【25】 Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
标题: 增强、删除和交换:改善LLM字幕的多样性,以实现高效的音乐文本表示学习
作者:Ilaria Manco,Justin Salamon,Oriol Nieto
备注:To appear in the Proceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:音频文本对比模型已经成为音乐表征学习的一种强有力的方法。尽管他们的经验的成功,但是,鲜为人知的是通过这个框架学习的音乐文本表示的质量的关键设计选择的影响。在这项工作中,我们在有限的数据和计算预算的约束下暴露了这些设计选择,并沿着三个轴建立了基于经验观察的对其影响的更坚实的理解:基本编码器的选择,训练数据的策展水平,以及文本增强的使用。我们发现,在资源受限的情况下,数据策展是音乐文本对比训练的最重要因素。基于这一认识,我们引入了两种新技术,增强视图丢弃和文本交换,它们增加了训练中文本输入的多样性和连续性。通过我们的实验,我们证明了这些方法可以有效地提高不同预训练机制、模型架构和下游数据分布的性能,而不会产生更高的计算成本或需要额外的训练数据。摘要:Audio-text contrastive models have become a powerful approach in music representation learning. Despite their empirical success, however, little is known about the influence of key design choices on the quality of music-text representations learnt through this framework. In this work, we expose these design choices within the constraints of limited data and computation budgets, and establish a more solid understanding of their impact grounded in empirical observations along three axes: the choice of base encoders, the level of curation in training data, and the use of text augmentation. We find that data curation is the single most important factor for music-text contrastive training in resource-constrained scenarios. Motivated by this insight, we introduce two novel techniques, Augmented View Dropout and TextSwap, which increase the diversity and descriptiveness of text inputs seen in training. Through our experiments we demonstrate that these are effective at boosting performance across different pre-training regimes, model architectures, and downstream data distributions, without incurring higher computational costs or requiring additional training data.
【26】 Evaluation of pretrained language models on music understanding
标题: 预训练语言模型对音乐理解的评估
作者:Yannis Vasilakis,Rachel Bittner,Johan Pauwels
链接:点击下载PDF文件
摘要:音乐文本多模态系统已经使新的方法,音乐信息研究(MIR)的应用,如音频到文本和文本到音频检索,基于文本的歌曲生成,和音乐字幕。尽管报告的成功,很少有人投入到评估大语言模型(LLM)的音乐知识。在本文中,我们证明了LLM遭受1)提示敏感性,2)无法模拟否定(例如“没有吉他的摇滚歌曲”),以及3)对特定单词的敏感性。我们将这些属性量化为基于三元组的准确性,评估在分层本体中对标签的相对相似性进行建模的能力。我们利用Audioset本体来生成由流派和乐器子树的锚、正(相关)标签和负(不太相关)标签组成的三元组。我们评估了基于三元组的音乐知识的六个通用的基于变压器的模型。通过这种方法获得的三联体需要过滤,因为有些三联体难以判断,因此对于评价目的而言相对缺乏信息。尽管报告的准确性相对较高,但所有六种模型都存在明显的不一致性,这表明现成的LLM在使用前需要适应音乐。摘要:Music-text multimodal systems have enabled new approaches to Music Information Research (MIR) applications such as audio-to-text and text-to-audio retrieval, text-based song generation, and music captioning. Despite the reported success, little effort has been put into evaluating the musical knowledge of Large Language Models (LLM). In this paper, we demonstrate that LLMs suffer from 1) prompt sensitivity, 2) inability to model negation (e.g. 'rock song without guitar'), and 3) sensitivity towards the presence of specific words. We quantified these properties as a triplet-based accuracy, evaluating the ability to model the relative similarity of labels in a hierarchical ontology. We leveraged the Audioset ontology to generate triplets consisting of an anchor, a positive (relevant) label, and a negative (less relevant) label for the genre and instruments sub-tree. We evaluated the triplet-based musical knowledge for six general-purpose Transformer-based models. The triplets obtained through this methodology required filtering, as some were difficult to judge and therefore relatively uninformative for evaluation purposes. Despite the relatively high accuracy reported, inconsistencies are evident in all six models, suggesting that off-the-shelf LLMs need adaptation to music before use.
【27】 Machine listening in a neonatal intensive care unit
标题: 新生儿重症监护室中的机器监听
作者:Modan Tailleur,Vincent Lostanlen,Jean-Philippe Rivière,Pierre Aumond
Journal-ref:DCASE2024 Workshop, Nobutaka Ono; Noboru Harada; Yohei Kawaguchi, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
摘要:氧合器、报警装置和脚步声是医院中最常见的声源。检测它们对环境心理学具有科学价值,但也带来了自身的挑战:即隐私保护和有限的标记数据。在本文中,我们通过边缘计算和云计算的结合来解决这两个挑战。为了保护隐私,我们设计了一种声学传感器,它可以实时计算第三倍频程频谱图,而不是记录音频波形。为了实现样本高效的机器学习,我们通过频谱转码和标签空间自适应来重新利用预训练的音频神经网络(PANN)。在神经病学重症监护室(NICU)中的小规模研究证实,检测到的事件的时间序列与另一种测量方式一致:即,家长和医护人员的电子徽章。因此,本文论证了在医院病房中使用复音机收听的可行性,同时通过设计来保证隐私。摘要:Oxygenators, alarm devices, and footsteps are some of the most common sound sources in a hospital. Detecting them has scientific value for environmental psychology but comes with challenges of its own: namely, privacy preservation and limited labeled data. In this paper, we address these two challenges via a combination of edge computing and cloud computing. For privacy preservation, we have designed an acoustic sensor which computes third-octave spectrograms on the fly instead of recording audio waveforms. For sample-efficient machine learning, we have repurposed a pretrained audio neural network (PANN) via spectral transcoding and label space adaptation. A small-scale study in a neonatological intensive care unit (NICU) confirms that the time series of detected events align with another modality of measurement: i.e., electronic badges for parents and healthcare professionals. Hence, this paper demonstrates the feasibility of polyphonic machine listening in a hospital ward while guaranteeing privacy by design.
机器翻译,仅供参考
