本文经arXiv每日学术速递授权转载
【1】 Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
标题: 利用TTC合成数据高效开发关键词发现模型
作者:Hyun Jin Park,Dhruuv Agarwal,Neng Chen,Rentao Sun,Kurt Partridge,Justin Chen,Harry Zhang,Pai Zhu,Jacob Bartel,Kyle Kastner,Gary Wang,Andrew Rosenberg,Quan Wang
备注:to be published in a Workshop at Interspeech 2024, Synthetic Data's Transformative Role in Foundational Speech Models
链接:点击下载PDF文件
【2】 Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks
标题: 通过高保真生成对抗网络扩展语音带宽
作者:Mahmoud Salhab,Haidar Harmanani
链接:点击下载PDF文件
【3】 Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
标题: 使用具有交叉注意力的音频-视频Transformer融合的多模式情感识别
作者:Joe Dhanith P R,Shravan Venkatraman,Vigya Sharma,Santhosh Malarvannan
备注:38 Pages, 9 Tables, 12 Figures
链接:点击下载PDF文件
【4】 Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
标题: 使用自我监督语音模型提高NAM到语音合成的可理解度
作者:Neil Shah,Shirish Karande,Vineet Gandhi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【5】 SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
标题: FLIM:用于广义音频深度伪造检测的风格语言学不匹配模型
作者:Yi Zhu,Surya Koppisetti,Trang Tran,Gaurav Bharaj
链接:点击下载PDF文件
【6】 The formation of perceptual space in early phonetic acquisition: a cross-linguistic modeling approach
标题: 早期语音习得中感知空间的形成:跨语言建模方法
作者:Frank Lihui Tan,Youngah Do
备注:51 pages
链接:点击下载PDF文件
【7】 Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
标题: 通过基于原型的自适应增强隐形说话者的合成障碍语音识别
作者:Shiyao Wang,Shiwan Zhao,Jiaming Zhou,Aobo Kong,Yong Qin
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
【8】 Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
标题: 模型驱动的心率估计和基于音素图的心脏杂音检测
作者:Jingping Nie,Ran Liu,Behrooz Mahasseni,Erdrin Azemi,Vikramjit Mitra
备注:6 pages, 10 figures
链接:点击下载PDF文件
【9】 Simulation of Neural Responses to Classical Music Using Organoid Intelligence Methods
标题: 利用类器官智能方法模拟古典音乐的神经反应
作者:Daniel Szelogowski
备注:10 pages, 9 figures
链接:点击下载PDF文件
【10】 A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays
标题: 基于物理知识的神经网络的球形麦克风阵列空间上采样方法
作者:Federico Miotello,Ferdinando Terminiello,Mirco Pezzoli,Alberto Bernardini,Fabio Antonacci,Augusto Sarti
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
【11】 Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
标题: 分析无文本语音翻译的语音单元选择
作者:Jarod Duret,Yannick Estève,Titouan Parcollet
链接:点击下载PDF文件
标题: 基于物理知识的神经网络的球形麦克风阵列空间上采样方法
作者:Federico Miotello,Ferdinando Terminiello,Mirco Pezzoli,Alberto Bernardini,Fabio Antonacci,Augusto Sarti
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
【2】 Integrating Posture Control in Speech Motor Models: A Parallel-Structured Simulation Approach
标题: 语音运动模型中的姿势控制集成:并行结构的模拟方法
作者:Yadong Liu,Sidney Fels,Arian Shamei,Najeeb Khan,Bryan Gick
备注:11 pages, 3 figures
链接:点击下载PDF文件
【3】 VoxSim: A perceptual voice similarity dataset
标题: VoxSim:感知语音相似性数据集
作者:Junseok Ahn,Youkyum Kim,Yeunju Choi,Doyeop Kwak,Ji-Hoon Kim,Seongkyu Mun,Joon Son Chung
备注:INTERSPEECH 2024. The dataset is available from this https URL
链接:点击下载PDF文件
【4】 Matlab-based Epoch Extraction for Speaker Differentiation
标题: 基于Matlab的说话人区分的时代元提取
作者:Kunlun Li,Daniel Ferro,Xu Zhao,Abdul Jabbar Syed,Anil K Vuppala,Azeemuddin Syed
备注:8 pages, 11 figures, This paper is currently under review by the 9th ACMIEEE Symposium on Edge Computing (SEC 2024)
链接:点击下载PDF文件
【5】 Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
标题: 分析无文本语音翻译的语音单元选择
作者:Jarod Duret,Yannick Estève,Titouan Parcollet
链接:点击下载PDF文件
【6】 Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
标题: 利用TTC合成数据高效开发关键词发现模型
作者:Hyun Jin Park,Dhruuv Agarwal,Neng Chen,Rentao Sun,Kurt Partridge,Justin Chen,Harry Zhang,Pai Zhu,Jacob Bartel,Kyle Kastner,Gary Wang,Andrew Rosenberg,Quan Wang
备注:to be published in a Workshop at Interspeech 2024, Synthetic Data's Transformative Role in Foundational Speech Models
链接:点击下载PDF文件
【7】 Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks
标题: 通过高保真生成对抗网络扩展语音带宽
作者:Mahmoud Salhab,Haidar Harmanani
链接:点击下载PDF文件
【8】 Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
标题: 使用具有交叉注意力的音频-视频Transformer融合的多模式情感识别
作者:Joe Dhanith P R,Shravan Venkatraman,Vigya Sharma,Santhosh Malarvannan
备注:38 Pages, 9 Tables, 12 Figures
链接:点击下载PDF文件
【9】 Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
标题: 使用自我监督语音模型提高NAM到语音合成的可理解度
作者:Neil Shah,Shirish Karande,Vineet Gandhi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
【10】 SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
标题: FLIM:用于广义音频深度伪造检测的风格语言学不匹配模型
作者:Yi Zhu,Surya Koppisetti,Trang Tran,Gaurav Bharaj
链接:点击下载PDF文件
【11】 The formation of perceptual space in early phonetic acquisition: a cross-linguistic modeling approach
标题: 早期语音习得中感知空间的形成:跨语言建模方法
作者:Frank Lihui Tan,Youngah Do
备注:51 pages
链接:点击下载PDF文件
【12】 Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
标题: 通过基于原型的自适应增强隐形说话者的合成障碍语音识别
作者:Shiyao Wang,Shiwan Zhao,Jiaming Zhou,Aobo Kong,Yong Qin
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
【13】 Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
标题: 模型驱动的心率估计和基于音素图的心脏杂音检测
作者:Jingping Nie,Ran Liu,Behrooz Mahasseni,Erdrin Azemi,Vikramjit Mitra
备注:6 pages, 10 figures
链接:点击下载PDF文件
【14】 Simulation of Neural Responses to Classical Music Using Organoid Intelligence Methods
标题: 利用类器官智能方法模拟古典音乐的神经反应
作者:Daniel Szelogowski
备注:10 pages, 9 figures
链接:点击下载PDF文件
【15】 AMA-LSTM: Pioneering Robust and Fair Financial Audio Analysis for Stock Volatility Prediction
标题: AMA-LSTM:开创性的稳健、公平的金融音频分析,用于股票波动性预测
作者:Shengkun Wang,Taoran Ji,Jianfeng He,Mariam Almutairi,Dan Wang,Linhan Wang,Min Zhang,Chang-Tien Lu
链接:点击下载PDF文件
标题: 利用TTC合成数据高效开发关键词发现模型
作者:Hyun Jin Park,Dhruuv Agarwal,Neng Chen,Rentao Sun,Kurt Partridge,Justin Chen,Harry Zhang,Pai Zhu,Jacob Bartel,Kyle Kastner,Gary Wang,Andrew Rosenberg,Quan Wang
备注:to be published in a Workshop at Interspeech 2024, Synthetic Data's Transformative Role in Foundational Speech Models
链接:点击下载PDF文件
摘要:本文探讨了使用TTS合成训练数据的KWS(关键字定位)任务,同时最大限度地减少开发成本和时间。关键词识别模型需要大量的训练数据才能准确,而获得这些训练数据的成本可能很高。在目前的技术水平下,TTS模型可以生成大量的自然探测数据,这可以帮助减少KWS模型开发的成本和时间。尽管如此,TTS生成的数据与真实数据相比可能缺乏多样性。为了在有限的资源和当前TTS能力的约束下追求最大化KWS模型的准确性,我们探索了各种策略来混合TTS数据和真实人类语音数据,重点是最大限度地减少真实数据的使用和最大限度地提高TTS输出的多样性。我们的实验结果表明,相对少量的真实音频数据与扬声器的多样性(100扬声器,2k话语)和大量的TTS合成数据可以实现相当高的准确性(3倍误差率的基线),相比基线(训练与3.8M真正的积极话语)。摘要:This paper explores the use of TTS synthesized training data for KWS (keyword spotting) task while minimizing development cost and time. Keyword spotting models require a huge amount of training data to be accurate, and obtaining such training data can be costly. In the current state of the art, TTS models can generate large amounts of natural-sounding data, which can help reducing cost and time for KWS model development. Still, TTS generated data can be lacking diversity compared to real data. To pursue maximizing KWS model accuracy under the constraint of limited resources and current TTS capability, we explored various strategies to mix TTS data and real human speech data, with a focus on minimizing real data use and maximizing diversity of TTS output. Our experimental results indicate that relatively small amounts of real audio data with speaker diversity (100 speakers, 2k utterances) and large amounts of TTS synthesized data can achieve reasonably high accuracy (within 3x error rate of baseline), compared to the baseline (trained with 3.8M real positive utterances).
【2】 Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks
标题: 通过高保真生成对抗网络扩展语音带宽
作者:Mahmoud Salhab,Haidar Harmanani
链接:点击下载PDF文件
摘要:语音带宽扩展对于扩展低带宽语音信号的频率范围至关重要,从而提高数字应用中的音频质量、清晰度和可感知性。它的应用范围包括电话、压缩、文本到语音合成和语音识别。本文提出了一种使用高保真生成对抗网络的新方法,与级联系统不同,我们的系统在成对的窄带和宽带语音信号上进行端到端训练。我们的方法将各种带宽上采样率集成到一个统一的模型中,专门为语音带宽扩展应用而设计。我们的方法在各种带宽扩展因素中表现出强大的性能,包括那些在训练过程中没有遇到的,展示了zero-shot能力。据我们所知,这是第一个展示这种能力的作品。实验结果表明,我们的方法优于以前的端到端的方法,以及插值和传统的技术,展示了其在实际语音增强应用的有效性。摘要:Speech bandwidth expansion is crucial for expanding the frequency range of low-bandwidth speech signals, thereby improving audio quality, clarity and perceptibility in digital applications. Its applications span telephony, compression, text-to-speech synthesis, and speech recognition. This paper presents a novel approach using a high-fidelity generative adversarial network, unlike cascaded systems, our system is trained end-to-end on paired narrowband and wideband speech signals. Our method integrates various bandwidth upsampling ratios into a single unified model specifically designed for speech bandwidth expansion applications. Our approach exhibits robust performance across various bandwidth expansion factors, including those not encountered during training, demonstrating zero-shot capability. To the best of our knowledge, this is the first work to showcase this capability. The experimental results demonstrate that our method outperforms previous end-to-end approaches, as well as interpolation and traditional techniques, showcasing its effectiveness in practical speech enhancement applications.
【3】 Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
标题: 使用具有交叉注意力的音频-视频Transformer融合的多模式情感识别
作者:Joe Dhanith P R,Shravan Venkatraman,Vigya Sharma,Santhosh Malarvannan
备注:38 Pages, 9 Tables, 12 Figures
链接:点击下载PDF文件
摘要:理解情感是人类交流的一个基本方面。与依赖单个数据源(如语音或面部表情)的传统方法相比,整合音频和视频信号可以更全面地了解情绪状态。尽管它的潜力,多模态情感识别面临着重大的挑战,特别是在同步,特征提取和融合不同的数据源。为了解决这些问题,本文介绍了一种新的基于变压器的模型命名为音频-视频Transformer融合与交叉注意力(AVT-CA)。AVT-CA模型采用Transformer融合方法来有效地捕获和同步来自音频和视频输入的互连特征,从而解决同步问题。此外,AVT-CA中的交叉注意机制选择性地提取和强调关键特征,同时丢弃来自两种模态的不相关特征,从而解决特征提取和融合挑战。在CMU-MOSEI、RAVDESS和CREMA-D数据集上进行的大量实验分析证明了该模型的有效性。研究结果强调了AVT-CA在开发精确可靠的多模态情感识别系统中的重要性。摘要:Understanding emotions is a fundamental aspect of human communication. Integrating audio and video signals offers a more comprehensive understanding of emotional states compared to traditional methods that rely on a single data source, such as speech or facial expressions. Despite its potential, multimodal emotion recognition faces significant challenges, particularly in synchronization, feature extraction, and fusion of diverse data sources. To address these issues, this paper introduces a novel transformer-based model named Audio-Video Transformer Fusion with Cross Attention (AVT-CA). The AVT-CA model employs a transformer fusion approach to effectively capture and synchronize interlinked features from both audio and video inputs, thereby resolving synchronization problems. Additionally, the Cross Attention mechanism within AVT-CA selectively extracts and emphasizes critical features while discarding irrelevant ones from both modalities, addressing feature extraction and fusion challenges. Extensive experimental analysis conducted on the CMU-MOSEI, RAVDESS and CREMA-D datasets demonstrates the efficacy of the proposed model. The results underscore the importance of AVT-CA in developing precise and reliable multimodal emotion recognition systems for practical applications.
【4】 Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
标题: 使用自我监督语音模型提高NAM到语音合成的可理解度
作者:Neil Shah,Shirish Karande,Vineet Gandhi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的方法,利用自我监督和序列到序列(Seq2 Seq)学习技术,显着提高非可听杂音(NAM)到语音转换任务的清晰度。与明确记录地面实况语音的传统方法不同,我们的方法依赖于自我监督和语音到语音合成来模拟地面实况语音。尽管利用模拟语音,我们的方法超过了目前的最先进的(SOTA)的梅尔倒谱失真(MCD)度量的29.08%的改善。此外,我们提出了错误率,并证明了我们的模型的能力,以合成语音的新的声音的兴趣。此外,我们提出了一种方法来增强现有的CSTR NAM TIMIT Plus语料库,设置一个基准的字错误率(WER)为42.57%,以衡量合成语音的可懂度。语音样本可以在https: nam2speech.github.io NAM2Speech 上找到摘要:We propose a novel approach to significantly improve the intelligibility in the Non-Audible Murmur (NAM)-to-speech conversion task, leveraging self-supervision and sequence-to-sequence (Seq2Seq) learning techniques. Unlike conventional methods that explicitly record ground-truth speech, our methodology relies on self-supervision and speech-to-speech synthesis to simulate ground-truth speech. Despite utilizing simulated speech, our method surpasses the current state-of-the-art (SOTA) by 29.08% improvement in the Mel-Cepstral Distortion (MCD) metric. Additionally, we present error rates and demonstrate our model's proficiency to synthesize speech in novel voices of interest. Moreover, we present a methodology for augmenting the existing CSTR NAM TIMIT Plus corpus, setting a benchmark with a Word Error Rate (WER) of 42.57% to gauge the intelligibility of the synthesized speech. Speech samples can be found at https: nam2speech.github.io NAM2Speech
【5】 SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
标题: FLIM:用于广义音频深度伪造检测的风格语言学不匹配模型
作者:Yi Zhu,Surya Koppisetti,Trang Tran,Gaurav Bharaj
链接:点击下载PDF文件
摘要:音频deepfake检测(ADD)对于打击从生成AI模型合成的语音的滥用至关重要。现有的ADD模型存在泛化问题,域内和域外数据之间存在很大的性能差异。此外,现有模型的黑盒性质限制了它们在现实世界场景中的使用,在现实世界中,模型决策需要解释。为了缓解这些问题,我们引入了一个新的ADD模型,该模型显式地使用假语音中的StyleLInguistics Mismatch(SLIM)来将它们与真实语音分开。SLIM首先只对真实样本进行自我监督预训练,以学习真实类中的风格语言学依赖。然后将学习的特征与标准预训练的声学特征(例如,Wav2vec)学习一个分类器对真实类和假类的分类。当特征编码器被冻结时,SLIM在域外数据集上的性能优于基准方法,同时在域内数据上获得有竞争力的结果。SLIM学习的特征使我们能够量化样本中风格和语言内容之间的(错误)匹配,从而有助于解释模型决策。摘要:Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain data. Moreover, the black-box nature of existing models limits their use in real-world scenarios, where explanations are required for model decisions. To alleviate these issues, we introduce a new ADD model that explicitly uses the StyleLInguistics Mismatch (SLIM) in fake speech to separate them from real speech. SLIM first employs self-supervised pretraining on only real samples to learn the style-linguistics dependency in the real class. The learned features are then used in complement with standard pretrained acoustic features (e.g., Wav2vec) to learn a classifier on the real and fake classes. When the feature encoders are frozen, SLIM outperforms benchmark methods on out-of-domain datasets while achieving competitive results on in-domain data. The features learned by SLIM allow us to quantify the (mis)match between style and linguistic content in a sample, hence facilitating an explanation of the model decision.
【6】 The formation of perceptual space in early phonetic acquisition: a cross-linguistic modeling approach
标题: 早期语音习得中感知空间的形成:跨语言建模方法
作者:Frank Lihui Tan,Youngah Do
备注:51 pages
链接:点击下载PDF文件
摘要:本研究通过在两个关键方面推进前人的研究,探讨学习者如何在早期语音习得中组织感知空间。首先,它考察了学习的隐藏表征的形状以及它对语音类别进行分类的能力。其次,它探讨了训练模型对上下文无关的声学信息的影响,而不涉及上下文线索,对语音习得,密切模仿早期语言学习阶段。使用跨语言建模方法,自动编码器模型在英语和普通话上进行训练,并在母语和非母语条件下进行评估,遵循婴儿语言感知研究中使用的实验条件。结果表明,无监督的自下而上的培训上下文无关的声学信息导致可比的学习表征的感知空间之间的母语和非母语的条件下,英语和普通话,类似于早期阶段的婴儿普遍倾听。这些发现为我们理解早期语音习得过程中知觉空间的组织提供了新的视角,有助于我们理解语音范畴的形成和表征。摘要:This study investigates how learners organize perceptual space in early phonetic acquisition by advancing previous studies in two key aspects. Firstly, it examines the shape of the learned hidden representation as well as its ability to categorize phonetic categories. Secondly, it explores the impact of training models on context-free acoustic information, without involving contextual cues, on phonetic acquisition, closely mimicking the early language learning stage. Using a cross-linguistic modeling approach, autoencoder models are trained on English and Mandarin and evaluated in both native and non-native conditions, following experimental conditions used in infant language perception studies. The results demonstrate that unsupervised bottom-up training on context-free acoustic information leads to comparable learned representations of perceptual space between native and non-native conditions for both English and Mandarin, resembling the early stage of universal listening in infants. These findings provide insights into the organization of perceptual space during early phonetic acquisition and contribute to our understanding of the formation and representation of phonetic categories.
【7】 Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
标题: 通过基于原型的自适应增强隐形说话者的合成障碍语音识别
作者:Shiyao Wang,Shiwan Zhao,Jiaming Zhou,Aobo Kong,Yong Qin
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:构音障碍语音识别(DSR)提出了一个艰巨的挑战,由于固有的说话人间的变化,导致严重的性能下降时,应用DSR模型的新构音障碍的扬声器。传统的说话人自适应方法通常涉及微调模型为每个扬声器,但这种策略是成本高昂的,不方便残疾人用户,需要大量的数据收集。为了解决这个问题,我们引入了一个基于原型的方法,显着提高DSR性能看不见的构音障碍的扬声器没有额外的微调。我们的方法采用了一个用HuBERT训练的特征提取器来生成每个词的原型,这些原型封装了以前看不见的说话者的特征。这些原型是分类的基础。此外,我们结合了监督对比学习来改进特征提取。通过提高表示质量,我们进一步提高DSR性能,实现有效的个性化DSR。我们在https: github.com NKU-HLT PB-DSR上发布代码。摘要:Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional speaker adaptation methodologies typically involve fine-tuning models for each speaker, but this strategy is cost-prohibitive and inconvenient for disabled users, requiring substantial data collection. To address this issue, we introduce a prototype-based approach that markedly improves DSR performance for unseen dysarthric speakers without additional fine-tuning. Our method employs a feature extractor trained with HuBERT to produce per-word prototypes that encapsulate the characteristics of previously unseen speakers. These prototypes serve as the basis for classification. Additionally, we incorporate supervised contrastive learning to refine feature extraction. By enhancing representation quality, we further improve DSR performance, enabling effective personalized DSR. We release our code at https: github.com NKU-HLT PB-DSR.
【8】 Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
标题: 模型驱动的心率估计和基于音素图的心脏杂音检测
作者:Jingping Nie,Ran Liu,Behrooz Mahasseni,Erdrin Azemi,Vikramjit Mitra
备注:6 pages, 10 figures
链接:点击下载PDF文件
摘要:声学信号对于健康监测至关重要,特别是心音,它提供心率等基本数据并检测心脏异常,如杂音。本研究利用公开可用的心音图(PCG)数据集,使用模型驱动的方法来估计心率,并将性能最佳的模型扩展到多任务学习(MTL)框架,以同时进行心率估计和杂音检测。心率估计使用滑动窗口技术对心音片段进行推导,结合声学特征(Mel频谱图、倒谱系数、功率谱密度、均方根能量)进行分析。我们的研究结果表明,2D卷积神经网络(2dCNN)对于心率估计最有效,平均绝对误差(MAE)为1.312 bpm。我们系统地研究了不同特征组合的影响,发现利用所有四个特征会产生最好的结果。MTL模型( textbf{ texttt{2dCNN-MTL}})在杂音检测方面达到了95%以上的准确度,超越了现有模型,同时在心率估计方面保持了1.636 bpm的MAE,满足美国医疗器械促进协会(AAMI)的要求。摘要:Acoustic signals are crucial for health monitoring, particularly heart sounds which provide essential data like heart rate and detect cardiac anomalies such as murmurs. This study utilizes a publicly available phonocardiogram (PCG) dataset to estimate heart rate using model-driven methods and extends the best-performing model to a multi-task learning (MTL) framework for simultaneous heart rate estimation and murmur detection. Heart rate estimates are derived using a sliding window technique on heart sound snippets, analyzed with a combination of acoustic features (Mel spectrogram, cepstral coefficients, power spectral density, root mean square energy). Our findings indicate that a 2D convolutional neural network ( textbf{ texttt{2dCNN}}) is most effective for heart rate estimation, achieving a mean absolute error (MAE) of 1.312 bpm. We systematically investigate the impact of different feature combinations and find that utilizing all four features yields the best results. The MTL model ( textbf{ texttt{2dCNN-MTL}}) achieves accuracy over 95% in murmur detection, surpassing existing models, while maintaining an MAE of 1.636 bpm in heart rate estimation, satisfying the requirements stated by Association for the Advancement of Medical Instrumentation (AAMI).
【9】 Simulation of Neural Responses to Classical Music Using Organoid Intelligence Methods
标题: 利用类器官智能方法模拟古典音乐的神经反应
作者:Daniel Szelogowski
备注:10 pages, 9 figures
链接:点击下载PDF文件
摘要:音乐是一种复杂的听觉刺激,能够引起大脑活动的显着变化,影响认知过程,如记忆,注意力和情绪调节。然而,音乐诱导的认知过程的潜在机制在很大程度上仍然未知。类器官智能和深度学习模型显示出模拟和分析古典音乐的神经反应的希望,这是计算神经科学中尚未探索的领域。因此,我们提出了PyOrganoid库,这是一种创新的工具,可以促进类器官学习模型的模拟,将复杂的机器学习技术与生物启发的类器官模拟相结合。我们的研究重点是Pianoid模型的开发,这是一种“深度类器官学习”模型,利用双向LSTM网络来预测基于古典音乐录音音频特征的EEG响应。该模型证明了使用计算方法复制复杂神经过程的可行性,为音乐感知和认知提供了有价值的见解。同样,我们的研究结果强调了合成模型在神经科学研究中的实用性,并突出了PyOrganoid库作为推进神经科学和人工智能研究的多功能工具的潜力。摘要:Music is a complex auditory stimulus capable of eliciting significant changes in brain activity, influencing cognitive processes such as memory, attention, and emotional regulation. However, the underlying mechanisms of music-induced cognitive processes remain largely unknown. Organoid intelligence and deep learning models show promise for simulating and analyzing these neural responses to classical music, an area significantly unexplored in computational neuroscience. Hence, we present the PyOrganoid library, an innovative tool that facilitates the simulation of organoid learning models, integrating sophisticated machine learning techniques with biologically inspired organoid simulations. Our study features the development of the Pianoid model, a "deep organoid learning" model that utilizes a Bidirectional LSTM network to predict EEG responses based on audio features from classical music recordings. This model demonstrates the feasibility of using computational methods to replicate complex neural processes, providing valuable insights into music perception and cognition. Likewise, our findings emphasize the utility of synthetic models in neuroscience research and highlight the PyOrganoid library's potential as a versatile tool for advancing studies in neuroscience and artificial intelligence.
【10】 A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays
标题: 基于物理知识的神经网络的球形麦克风阵列空间上采样方法
作者:Federico Miotello,Ferdinando Terminiello,Mirco Pezzoli,Alberto Bernardini,Fabio Antonacci,Augusto Sarti
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
摘要:球形麦克风阵列是用于捕获声场的空间特性的方便工具。然而,实现优异的空间分辨率需要具有许多胶囊的阵列,因此导致昂贵的设备。为了解决这个问题,我们提出了一种空间上采样球形麦克风阵列与有限数量的胶囊的方法。我们的方法利用具有Rowdy激活函数的物理信息神经网络,利用物理约束从低阶设备开始提供高阶麦克风阵列信号。结果表明,在其应用范围内,我们的方法优于最先进的方法,基于信号处理的球形麦克风阵列上采样。摘要:Spherical microphone arrays are convenient tools for capturing the spatial characteristics of a sound field. However, achieving superior spatial resolution requires arrays with numerous capsules, consequently leading to expensive devices. To address this issue, we present a method for spatially upsampling spherical microphone arrays with a limited number of capsules. Our approach exploits a physics-informed neural network with Rowdy activation functions, leveraging physical constraints to provide high-order microphone array signals, starting from low-order devices. Results show that, within its domain of application, our approach outperforms a state of the art method based on signal processing for spherical microphone arrays upsampling.
【11】 Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
标题: 分析无文本语音翻译的语音单元选择
作者:Jarod Duret,Yannick Estève,Titouan Parcollet
链接:点击下载PDF文件
摘要:无文本语音到语音翻译系统的最新进展是通过采用自监督学习技术来驱动的。虽然大多数国家的最先进的系统采用类似的架构,将源语言语音转换成离散表示的目标语言的序列,选择这些目标语音单元的标准仍然是一个悬而未决的问题。这项工作探讨了选择过程中,通过研究的下游任务,如自动语音识别,语音合成,说话人识别,情感识别。有趣的是,我们的研究结果揭示了离散语音单元优化的差异:在再合成性能方面表现良好的单元不一定与提高翻译效率的单元相关。这种差异强调了目标特征选择的微妙复杂性及其对语音到语音翻译系统整体性能的影响。摘要:Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transform source language speech into sequences of discrete representations in the target language, the criteria for selecting these target speech units remains an open question. This work explores the selection process through a study of downstream tasks such as automatic speech recognition, speech synthesis, speaker recognition, and emotion recognition. Interestingly, our findings reveal a discrepancy in the optimization of discrete speech units: units that perform well in resynthesis performance do not necessarily correlate with those that enhance translation efficacy. This discrepancy underscores the nuanced complexity of target feature selection and its impact on the overall performance of speech-to-speech translation systems.
eess.AS音频处理
【1】 A Physics-Informed Neural Network-Based Approach for the Spatial Upsampling of Spherical Microphone Arrays标题: 基于物理知识的神经网络的球形麦克风阵列空间上采样方法
作者:Federico Miotello,Ferdinando Terminiello,Mirco Pezzoli,Alberto Bernardini,Fabio Antonacci,Augusto Sarti
备注:Accepted for publication at IWAENC 2024
链接:点击下载PDF文件
摘要:球形麦克风阵列是用于捕获声场的空间特性的方便工具。然而,实现优异的空间分辨率需要具有许多胶囊的阵列,因此导致昂贵的设备。为了解决这个问题,我们提出了一种空间上采样球形麦克风阵列与有限数量的胶囊的方法。我们的方法利用具有Rowdy激活函数的物理信息神经网络,利用物理约束从低阶设备开始提供高阶麦克风阵列信号。结果表明,在其应用范围内,我们的方法优于最先进的方法,基于信号处理的球形麦克风阵列上采样。摘要:Spherical microphone arrays are convenient tools for capturing the spatial characteristics of a sound field. However, achieving superior spatial resolution requires arrays with numerous capsules, consequently leading to expensive devices. To address this issue, we present a method for spatially upsampling spherical microphone arrays with a limited number of capsules. Our approach exploits a physics-informed neural network with Rowdy activation functions, leveraging physical constraints to provide high-order microphone array signals, starting from low-order devices. Results show that, within its domain of application, our approach outperforms a state of the art method based on signal processing for spherical microphone arrays upsampling.
【2】 Integrating Posture Control in Speech Motor Models: A Parallel-Structured Simulation Approach
标题: 语音运动模型中的姿势控制集成:并行结构的模拟方法
作者:Yadong Liu,Sidney Fels,Arian Shamei,Najeeb Khan,Bryan Gick
备注:11 pages, 3 figures
链接:点击下载PDF文件
摘要:重力是运动行为的一个重要方面,需要持续的肌肉激活来抵消重力。它在扰动下保持稳定,有助于保持身体平衡并实现运动执行。人们观察到全身姿势和言语姿势之间存在相似之处,例如涉及下巴、舌头和嘴唇的姿势,它们也表现出对扰动的弹性,并有助于平衡和运动。虽然姿势控制是人类运动和平衡的公认要素,特别是在更广泛的运动技能中,但它尚未充分纳入现有的言语运动控制模型中,这些模型通常集中于与特定言语运动相关联的手势或运动命令,忽视了姿势控制和重力的影响。在这里,我们介绍了一个模型,对齐语音姿势和运动,使用模拟来探索是否在这个框架内的语音姿势反映身体姿势控制的原则。我们的研究结果表明,类似于身体姿势,语音姿势也是强大的扰动,并在保持局部段平衡和提高语音生产中发挥了重要作用。摘要:Posture is an essential aspect of motor behavior, necessitating continuous muscle activation to counteract gravity. It remains stable under perturbation, aiding in maintaining bodily balance and enabling movement execution. Similarities have been observed between gross body postures and speech postures, such as those involving the jaw, tongue, and lips, which also exhibit resilience to perturbations and assist in equilibrium and movement. Although postural control is a recognized element of human movement and balance, particularly in broader motor skills, it has not been adequately incorporated into existing speech motor control models, which typically concentrate on the gestures or motor commands associated with specific speech movements, overlooking the influence of postural control and gravity. Here we introduce a model that aligns speech posture and movement, using simulations to explore whether speech posture within this framework mirrors the principles of bodily postural control. Our findings indicate that, akin to body posture, speech posture is also robust to perturbation and plays a significant role in maintaining local segment balance and enhancing speech production.
【3】 VoxSim: A perceptual voice similarity dataset
标题: VoxSim:感知语音相似性数据集
作者:Junseok Ahn,Youkyum Kim,Yeunju Choi,Doyeop Kwak,Ji-Hoon Kim,Seongkyu Mun,Joon Son Chung
备注:INTERSPEECH 2024. The dataset is available from this https URL
链接:点击下载PDF文件
摘要:本文介绍了VoxSim,感知语音相似性评级的数据集。最近的努力,以自动化语音合成技术的评估主要集中在预测平均意见评分的自然,离开说话人的声音相似性相对未开发的,由于缺乏广泛的训练数据。为了解决这个问题,我们从VoxCeleb数据集(一个广泛用于说话人识别的语音数据集)中生成了大约41k个话语对,并通过听力测试收集了近70k个说话人相似性分数。VoxSim为说话人相似性预测模型的开发和基准测试提供了宝贵的资源。我们提供了VoxSim测试集上说话人相似性预测模型的基线结果,并进一步证明了在我们的数据集上训练的模型可以推广到域外VCC 2018数据集。摘要:This paper introduces VoxSim, a dataset of perceptual voice similarity ratings. Recent efforts to automate the assessment of speech synthesis technologies have primarily focused on predicting mean opinion score of naturalness, leaving speaker voice similarity relatively unexplored due to a lack of extensive training data. To address this, we generate about 41k utterance pairs from the VoxCeleb dataset, a widely utilised speech dataset for speaker recognition, and collect nearly 70k speaker similarity scores through a listening test. VoxSim offers a valuable resource for the development and benchmarking of speaker similarity prediction models. We provide baseline results of speaker similarity prediction models on the VoxSim test set and further demonstrate that the model trained on our dataset generalises to the out-of-domain VCC2018 dataset.
【4】 Matlab-based Epoch Extraction for Speaker Differentiation
标题: 基于Matlab的说话人区分的时代元提取
作者:Kunlun Li,Daniel Ferro,Xu Zhao,Abdul Jabbar Syed,Anil K Vuppala,Azeemuddin Syed
备注:8 pages, 11 figures, This paper is currently under review by the 9th ACMIEEE Symposium on Edge Computing (SEC 2024)
链接:点击下载PDF文件
摘要:由于准确检测Epoch的位置对于分析语音信号是至关重要的,因此Epoch提取在近年来的语音分析研究中变得越来越流行。在多人对话中,声带系统兴奋的瞬间,特别是声门闭合时,Epoch在区分说话者方面起着重要的作用。然而,由于声道系统中的时变因素,Epoch的提取提出了挑战,这使得用于获得原始激励位置的去卷积更加复杂。本文将讨论用于历元提取的各种方法,包括零频滤波(ZFF)和零频谐振器(ZFR),并评估其优缺点。此外,还将比较每种方法的稳定性、准确性和可行性。评估将涉及基于Matlab的锁定算法,以及使用Raspberry pi进行扬声器区分的拟议硬件实现。该实验包括六个人说出短语“密西西比大学”,其中一个人作为参考或“锁定”扬声器。在与参考说话者相似的位置处发生的时期的数量将被计数为Delta,其中较大的Delta值指示较大的说话者相似性。实验结果表明,当说话人保持不变时,Delta的平均数量为7.5,而对于不同的说话人,Delta的平均数量分别减少到3,2,2和1,表示与参考说话人相比,相似位置处的epoch数量减少了约73%。摘要:Epoch extraction has become increasingly popular in recent years for speech analysis research because accurately detecting the location of the Epoch is crucial for analyzing speech signals. The Epoch, occurring at the instant of excitation in the vocal tract system, particularly during glottal closure, plays a significant role in differentiating speakers in multi-speaker conversations. However, the extraction of the Epoch poses a challenge due to the time-varying factors in the vocal tract system, which makes deconvolution for obtaining the original excitation location more complex. In this paper, various methods for Epoch extraction, including Zero Frequency Filtering (ZFF) and Zero Frequency Resonator (ZFR), will be discussed, and their pros and cons evaluated. In addition, the stability, accuracy, and feasibility of each method will be compared. The evaluation will involve a Matlab-based locking algorithm, and a proposed hardware implementation using Raspberry pi for speaker differentiation. The experiment includes six individuals uttering the phrase "The University of Mississippi," with one person acting as the reference or "lock" speaker. The number of epochs occurring at similar positions to the reference speaker will be counted as Delta, with larger Delta values indicating greater speaker similarity. Experimental results demonstrate that when the speaker remains the same, the average number of Delta is 7.5, while for different speakers, the average number of Delta decreases to 3, 2, 2, and 1, respectively, representing a decrease of approximately 73% in the number of epochs at similar positions compared to the reference speaker.
【5】 Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
标题: 分析无文本语音翻译的语音单元选择
作者:Jarod Duret,Yannick Estève,Titouan Parcollet
链接:点击下载PDF文件
摘要:无文本语音到语音翻译系统的最新进展是通过采用自监督学习技术来驱动的。虽然大多数国家的最先进的系统采用类似的架构,将源语言语音转换成离散表示的目标语言的序列,选择这些目标语音单元的标准仍然是一个悬而未决的问题。这项工作探讨了选择过程中,通过研究的下游任务,如自动语音识别,语音合成,说话人识别,情感识别。有趣的是,我们的研究结果揭示了离散语音单元优化的差异:在再合成性能方面表现良好的单元不一定与提高翻译效率的单元相关。这种差异强调了目标特征选择的微妙复杂性及其对语音到语音翻译系统整体性能的影响。摘要:Recent advancements in textless speech-to-speech translation systems have been driven by the adoption of self-supervised learning techniques. Although most state-of-the-art systems adopt a similar architecture to transform source language speech into sequences of discrete representations in the target language, the criteria for selecting these target speech units remains an open question. This work explores the selection process through a study of downstream tasks such as automatic speech recognition, speech synthesis, speaker recognition, and emotion recognition. Interestingly, our findings reveal a discrepancy in the optimization of discrete speech units: units that perform well in resynthesis performance do not necessarily correlate with those that enhance translation efficacy. This discrepancy underscores the nuanced complexity of target feature selection and its impact on the overall performance of speech-to-speech translation systems.
【6】 Utilizing TTS Synthesized Data for Efficient Development of Keyword Spotting Model
标题: 利用TTC合成数据高效开发关键词发现模型
作者:Hyun Jin Park,Dhruuv Agarwal,Neng Chen,Rentao Sun,Kurt Partridge,Justin Chen,Harry Zhang,Pai Zhu,Jacob Bartel,Kyle Kastner,Gary Wang,Andrew Rosenberg,Quan Wang
备注:to be published in a Workshop at Interspeech 2024, Synthetic Data's Transformative Role in Foundational Speech Models
链接:点击下载PDF文件
摘要:本文探讨了使用TTS合成训练数据的KWS(关键字定位)任务,同时最大限度地减少开发成本和时间。关键词识别模型需要大量的训练数据才能准确,而获得这些训练数据的成本可能很高。在目前的技术水平下,TTS模型可以生成大量的自然探测数据,这可以帮助减少KWS模型开发的成本和时间。尽管如此,TTS生成的数据与真实数据相比可能缺乏多样性。为了在有限的资源和当前TTS能力的约束下追求最大化KWS模型的准确性,我们探索了各种策略来混合TTS数据和真实人类语音数据,重点是最大限度地减少真实数据的使用和最大限度地提高TTS输出的多样性。我们的实验结果表明,相对少量的真实音频数据与扬声器的多样性(100扬声器,2k话语)和大量的TTS合成数据可以实现相当高的准确性(3倍误差率的基线),相比基线(训练与3.8M真正的积极话语)。摘要:This paper explores the use of TTS synthesized training data for KWS (keyword spotting) task while minimizing development cost and time. Keyword spotting models require a huge amount of training data to be accurate, and obtaining such training data can be costly. In the current state of the art, TTS models can generate large amounts of natural-sounding data, which can help reducing cost and time for KWS model development. Still, TTS generated data can be lacking diversity compared to real data. To pursue maximizing KWS model accuracy under the constraint of limited resources and current TTS capability, we explored various strategies to mix TTS data and real human speech data, with a focus on minimizing real data use and maximizing diversity of TTS output. Our experimental results indicate that relatively small amounts of real audio data with speaker diversity (100 speakers, 2k utterances) and large amounts of TTS synthesized data can achieve reasonably high accuracy (within 3x error rate of baseline), compared to the baseline (trained with 3.8M real positive utterances).
【7】 Speech Bandwidth Expansion Via High Fidelity Generative Adversarial Networks
标题: 通过高保真生成对抗网络扩展语音带宽
作者:Mahmoud Salhab,Haidar Harmanani
链接:点击下载PDF文件
摘要:语音带宽扩展对于扩展低带宽语音信号的频率范围至关重要,从而提高数字应用中的音频质量、清晰度和可感知性。它的应用范围包括电话、压缩、文本到语音合成和语音识别。本文提出了一种使用高保真生成对抗网络的新方法,与级联系统不同,我们的系统在成对的窄带和宽带语音信号上进行端到端训练。我们的方法将各种带宽上采样率集成到一个统一的模型中,专门为语音带宽扩展应用而设计。我们的方法在各种带宽扩展因素中表现出强大的性能,包括那些在训练过程中没有遇到的,展示了zero-shot能力。据我们所知,这是第一个展示这种能力的作品。实验结果表明,我们的方法优于以前的端到端的方法,以及插值和传统的技术,展示了其在实际语音增强应用的有效性。摘要:Speech bandwidth expansion is crucial for expanding the frequency range of low-bandwidth speech signals, thereby improving audio quality, clarity and perceptibility in digital applications. Its applications span telephony, compression, text-to-speech synthesis, and speech recognition. This paper presents a novel approach using a high-fidelity generative adversarial network, unlike cascaded systems, our system is trained end-to-end on paired narrowband and wideband speech signals. Our method integrates various bandwidth upsampling ratios into a single unified model specifically designed for speech bandwidth expansion applications. Our approach exhibits robust performance across various bandwidth expansion factors, including those not encountered during training, demonstrating zero-shot capability. To the best of our knowledge, this is the first work to showcase this capability. The experimental results demonstrate that our method outperforms previous end-to-end approaches, as well as interpolation and traditional techniques, showcasing its effectiveness in practical speech enhancement applications.
【8】 Multimodal Emotion Recognition using Audio-Video Transformer Fusion with Cross Attention
标题: 使用具有交叉注意力的音频-视频Transformer融合的多模式情感识别
作者:Joe Dhanith P R,Shravan Venkatraman,Vigya Sharma,Santhosh Malarvannan
备注:38 Pages, 9 Tables, 12 Figures
链接:点击下载PDF文件
摘要:理解情感是人类交流的一个基本方面。与依赖单个数据源(如语音或面部表情)的传统方法相比,整合音频和视频信号可以更全面地了解情绪状态。尽管它的潜力,多模态情感识别面临着重大的挑战,特别是在同步,特征提取和融合不同的数据源。为了解决这些问题,本文介绍了一种新的基于变压器的模型命名为音频-视频Transformer融合与交叉注意力(AVT-CA)。AVT-CA模型采用Transformer融合方法来有效地捕获和同步来自音频和视频输入的互连特征,从而解决同步问题。此外,AVT-CA中的交叉注意机制选择性地提取和强调关键特征,同时丢弃来自两种模态的不相关特征,从而解决特征提取和融合挑战。在CMU-MOSEI、RAVDESS和CREMA-D数据集上进行的大量实验分析证明了该模型的有效性。研究结果强调了AVT-CA在开发精确可靠的多模态情感识别系统中的重要性。摘要:Understanding emotions is a fundamental aspect of human communication. Integrating audio and video signals offers a more comprehensive understanding of emotional states compared to traditional methods that rely on a single data source, such as speech or facial expressions. Despite its potential, multimodal emotion recognition faces significant challenges, particularly in synchronization, feature extraction, and fusion of diverse data sources. To address these issues, this paper introduces a novel transformer-based model named Audio-Video Transformer Fusion with Cross Attention (AVT-CA). The AVT-CA model employs a transformer fusion approach to effectively capture and synchronize interlinked features from both audio and video inputs, thereby resolving synchronization problems. Additionally, the Cross Attention mechanism within AVT-CA selectively extracts and emphasizes critical features while discarding irrelevant ones from both modalities, addressing feature extraction and fusion challenges. Extensive experimental analysis conducted on the CMU-MOSEI, RAVDESS and CREMA-D datasets demonstrates the efficacy of the proposed model. The results underscore the importance of AVT-CA in developing precise and reliable multimodal emotion recognition systems for practical applications.
【9】 Towards Improving NAM-to-Speech Synthesis Intelligibility using Self-Supervised Speech Models
标题: 使用自我监督语音模型提高NAM到语音合成的可理解度
作者:Neil Shah,Shirish Karande,Vineet Gandhi
备注:Accepted at Interspeech 2024
链接:点击下载PDF文件
摘要:我们提出了一种新的方法来显着提高非可听杂音(NAM)到语音转换任务的可懂度,利用自我监督和序列到序列(Seq2Seq)学习技术。与明确记录地面实况语音的传统方法不同,我们的方法依赖于自我监督和语音到语音合成来模拟地面实况语音。尽管利用模拟语音,我们的方法超过了目前的最先进的(SOTA)的梅尔倒谱失真(MCD)度量的29.08%的改善。此外,我们提出了错误率,并证明了我们的模型的能力,以合成语音的新的声音的兴趣。此外,我们提出了一种方法来增强现有的CSTR NAM TIMIT Plus语料库,设置一个基准的字错误率(WER)为42.57%,以衡量合成语音的可懂度。语音样本可以在https: nam2speech.github.io NAM2Speech 上找到摘要:We propose a novel approach to significantly improve the intelligibility in the Non-Audible Murmur (NAM)-to-speech conversion task, leveraging self-supervision and sequence-to-sequence (Seq2Seq) learning techniques. Unlike conventional methods that explicitly record ground-truth speech, our methodology relies on self-supervision and speech-to-speech synthesis to simulate ground-truth speech. Despite utilizing simulated speech, our method surpasses the current state-of-the-art (SOTA) by 29.08% improvement in the Mel-Cepstral Distortion (MCD) metric. Additionally, we present error rates and demonstrate our model's proficiency to synthesize speech in novel voices of interest. Moreover, we present a methodology for augmenting the existing CSTR NAM TIMIT Plus corpus, setting a benchmark with a Word Error Rate (WER) of 42.57% to gauge the intelligibility of the synthesized speech. Speech samples can be found at https: nam2speech.github.io NAM2Speech
【10】 SLIM: Style-Linguistics Mismatch Model for Generalized Audio Deepfake Detection
标题: FLIM:用于广义音频深度伪造检测的风格语言学不匹配模型
作者:Yi Zhu,Surya Koppisetti,Trang Tran,Gaurav Bharaj
链接:点击下载PDF文件
摘要:音频deepfake检测(ADD)对于打击从生成AI模型合成的语音的滥用至关重要。现有的ADD模型存在泛化问题,域内数据和域外数据之间存在很大的性能差异。此外,现有模型的黑盒性质限制了它们在现实世界场景中的使用,在现实世界中,模型决策需要解释。为了缓解这些问题,我们引入了一个新的ADD模型,该模型显式地使用假语音中的StyleLInguistics Mismatch(SLIM)来将它们与真实语音分开。SLIM首先只对真实样本进行自我监督预训练,以学习真实类中的风格语言学依赖。然后将学习的特征与标准预训练的声学特征(例如,Wav2vec)学习一个分类器对真实类和假类的分类。当特征编码器被冻结时,SLIM在域外数据集上的性能优于基准方法,同时在域内数据上获得有竞争力的结果。SLIM学习的特征使我们能够量化样本中风格和语言内容之间的(错误)匹配,从而有助于解释模型决策。摘要:Audio deepfake detection (ADD) is crucial to combat the misuse of speech synthesized from generative AI models. Existing ADD models suffer from generalization issues, with a large performance discrepancy between in-domain and out-of-domain data. Moreover, the black-box nature of existing models limits their use in real-world scenarios, where explanations are required for model decisions. To alleviate these issues, we introduce a new ADD model that explicitly uses the StyleLInguistics Mismatch (SLIM) in fake speech to separate them from real speech. SLIM first employs self-supervised pretraining on only real samples to learn the style-linguistics dependency in the real class. The learned features are then used in complement with standard pretrained acoustic features (e.g., Wav2vec) to learn a classifier on the real and fake classes. When the feature encoders are frozen, SLIM outperforms benchmark methods on out-of-domain datasets while achieving competitive results on in-domain data. The features learned by SLIM allow us to quantify the (mis)match between style and linguistic content in a sample, hence facilitating an explanation of the model decision.
【11】 The formation of perceptual space in early phonetic acquisition: a cross-linguistic modeling approach
标题: 早期语音习得中感知空间的形成:跨语言建模方法
作者:Frank Lihui Tan,Youngah Do
备注:51 pages
链接:点击下载PDF文件
摘要:本研究通过在两个关键方面推进前人的研究,探讨学习者如何在早期语音习得中组织感知空间。首先,它考察了学习的隐藏表征的形状以及它对语音类别进行分类的能力。其次,它探讨了训练模型对上下文无关的声学信息的影响,而不涉及上下文线索,对语音习得,密切模仿早期语言学习阶段。使用跨语言建模方法,自动编码器模型在英语和普通话上进行训练,并在母语和非母语条件下进行评估,遵循婴儿语言感知研究中使用的实验条件。结果表明,无监督的自下而上的培训上下文无关的声学信息导致可比的学习表征的感知空间之间的母语和非母语的条件下,英语和普通话,类似于早期阶段的婴儿普遍倾听。这些发现为我们理解早期语音习得过程中知觉空间的组织提供了新的视角,有助于我们理解语音范畴的形成和表征。摘要:This study investigates how learners organize perceptual space in early phonetic acquisition by advancing previous studies in two key aspects. Firstly, it examines the shape of the learned hidden representation as well as its ability to categorize phonetic categories. Secondly, it explores the impact of training models on context-free acoustic information, without involving contextual cues, on phonetic acquisition, closely mimicking the early language learning stage. Using a cross-linguistic modeling approach, autoencoder models are trained on English and Mandarin and evaluated in both native and non-native conditions, following experimental conditions used in infant language perception studies. The results demonstrate that unsupervised bottom-up training on context-free acoustic information leads to comparable learned representations of perceptual space between native and non-native conditions for both English and Mandarin, resembling the early stage of universal listening in infants. These findings provide insights into the organization of perceptual space during early phonetic acquisition and contribute to our understanding of the formation and representation of phonetic categories.
【12】 Enhancing Dysarthric Speech Recognition for Unseen Speakers via Prototype-Based Adaptation
标题: 通过基于原型的自适应增强隐形说话者的合成障碍语音识别
作者:Shiyao Wang,Shiwan Zhao,Jiaming Zhou,Aobo Kong,Yong Qin
Journal-ref:INTERSPEECH 2024
链接:点击下载PDF文件
摘要:构音障碍语音识别(DSR)提出了一个艰巨的挑战,由于固有的说话人间的变化,导致严重的性能下降时,应用DSR模型的新构音障碍的扬声器。传统的说话人自适应方法通常涉及微调模型为每个扬声器,但这种策略是成本高昂的,不方便残疾人用户,需要大量的数据收集。为了解决这个问题,我们引入了一个基于原型的方法,显着提高DSR性能看不见的构音障碍的扬声器没有额外的微调。我们的方法采用了一个用HuBERT训练的特征提取器来生成每个词的原型,这些原型封装了以前看不见的说话者的特征。这些原型是分类的基础。此外,我们结合了监督对比学习来改进特征提取。通过提高表示质量,我们进一步提高DSR性能,实现有效的个性化DSR。我们在https: github.com NKU-HLT PB-DSR上发布代码。摘要:Dysarthric speech recognition (DSR) presents a formidable challenge due to inherent inter-speaker variability, leading to severe performance degradation when applying DSR models to new dysarthric speakers. Traditional speaker adaptation methodologies typically involve fine-tuning models for each speaker, but this strategy is cost-prohibitive and inconvenient for disabled users, requiring substantial data collection. To address this issue, we introduce a prototype-based approach that markedly improves DSR performance for unseen dysarthric speakers without additional fine-tuning. Our method employs a feature extractor trained with HuBERT to produce per-word prototypes that encapsulate the characteristics of previously unseen speakers. These prototypes serve as the basis for classification. Additionally, we incorporate supervised contrastive learning to refine feature extraction. By enhancing representation quality, we further improve DSR performance, enabling effective personalized DSR. We release our code at https: github.com NKU-HLT PB-DSR.
【13】 Model-driven Heart Rate Estimation and Heart Murmur Detection based on Phonocardiogram
标题: 模型驱动的心率估计和基于音素图的心脏杂音检测
作者:Jingping Nie,Ran Liu,Behrooz Mahasseni,Erdrin Azemi,Vikramjit Mitra
备注:6 pages, 10 figures
链接:点击下载PDF文件
摘要:声学信号对于健康监测至关重要,特别是心音,它提供心率等基本数据并检测心脏异常,如杂音。本研究利用公开可用的心音图(PCG)数据集,使用模型驱动的方法来估计心率,并将性能最佳的模型扩展到多任务学习(MTL)框架,以同时进行心率估计和杂音检测。心率估计使用滑动窗口技术对心音片段进行推导,结合声学特征(Mel频谱图、倒谱系数、功率谱密度、均方根能量)进行分析。我们的研究结果表明,2D卷积神经网络(2dCNN)对于心率估计最有效,平均绝对误差(MAE)为1.312 bpm。我们系统地研究了不同特征组合的影响,发现利用所有四个特征会产生最好的结果。MTL模型( textbf{ texttt{2dCNN-MTL}})在杂音检测方面达到了95%以上的准确度,超越了现有模型,同时在心率估计方面保持了1.636 bpm的MAE,满足美国医疗器械促进协会(AAMI)的要求。摘要:Acoustic signals are crucial for health monitoring, particularly heart sounds which provide essential data like heart rate and detect cardiac anomalies such as murmurs. This study utilizes a publicly available phonocardiogram (PCG) dataset to estimate heart rate using model-driven methods and extends the best-performing model to a multi-task learning (MTL) framework for simultaneous heart rate estimation and murmur detection. Heart rate estimates are derived using a sliding window technique on heart sound snippets, analyzed with a combination of acoustic features (Mel spectrogram, cepstral coefficients, power spectral density, root mean square energy). Our findings indicate that a 2D convolutional neural network ( textbf{ texttt{2dCNN}}) is most effective for heart rate estimation, achieving a mean absolute error (MAE) of 1.312 bpm. We systematically investigate the impact of different feature combinations and find that utilizing all four features yields the best results. The MTL model ( textbf{ texttt{2dCNN-MTL}}) achieves accuracy over 95% in murmur detection, surpassing existing models, while maintaining an MAE of 1.636 bpm in heart rate estimation, satisfying the requirements stated by Association for the Advancement of Medical Instrumentation (AAMI).
【14】 Simulation of Neural Responses to Classical Music Using Organoid Intelligence Methods
标题: 利用类器官智能方法模拟古典音乐的神经反应
作者:Daniel Szelogowski
备注:10 pages, 9 figures
链接:点击下载PDF文件
摘要:音乐是一种复杂的听觉刺激,能够引起大脑活动的显着变化,影响认知过程,如记忆,注意力和情绪调节。然而,音乐诱导的认知过程的潜在机制在很大程度上仍然未知。类器官智能和深度学习模型显示出模拟和分析古典音乐的神经反应的希望,这是计算神经科学中尚未探索的领域。因此,我们提出了PyOrganoid库,这是一种创新的工具,可以促进类器官学习模型的模拟,将复杂的机器学习技术与生物启发的类器官模拟相结合。我们的研究重点是Pianoid模型的开发,这是一种“深度类器官学习”模型,利用双向LSTM网络来预测基于古典音乐录音音频特征的EEG响应。该模型证明了使用计算方法复制复杂神经过程的可行性,为音乐感知和认知提供了有价值的见解。同样,我们的研究结果强调了合成模型在神经科学研究中的实用性,并突出了PyOrganoid库作为推进神经科学和人工智能研究的多功能工具的潜力。摘要:Music is a complex auditory stimulus capable of eliciting significant changes in brain activity, influencing cognitive processes such as memory, attention, and emotional regulation. However, the underlying mechanisms of music-induced cognitive processes remain largely unknown. Organoid intelligence and deep learning models show promise for simulating and analyzing these neural responses to classical music, an area significantly unexplored in computational neuroscience. Hence, we present the PyOrganoid library, an innovative tool that facilitates the simulation of organoid learning models, integrating sophisticated machine learning techniques with biologically inspired organoid simulations. Our study features the development of the Pianoid model, a "deep organoid learning" model that utilizes a Bidirectional LSTM network to predict EEG responses based on audio features from classical music recordings. This model demonstrates the feasibility of using computational methods to replicate complex neural processes, providing valuable insights into music perception and cognition. Likewise, our findings emphasize the utility of synthetic models in neuroscience research and highlight the PyOrganoid library's potential as a versatile tool for advancing studies in neuroscience and artificial intelligence.
【15】 AMA-LSTM: Pioneering Robust and Fair Financial Audio Analysis for Stock Volatility Prediction
标题: AMA-LSTM:开创性的稳健、公平的金融音频分析,用于股票波动性预测
作者:Shengkun Wang,Taoran Ji,Jianfeng He,Mariam Almutairi,Dan Wang,Linhan Wang,Min Zhang,Chang-Tien Lu
链接:点击下载PDF文件
摘要:股票波动预测是金融业的一项重要工作。多模态方法的最新进展,它集成了文本和听觉数据,已经证明了这一领域的显着改善,如收益电话(收益电话是公开的,通常涉及上市公司的管理团队和有关各方讨论公司的收益)。然而,这些多模式方法面临两个缺点。首先,由于它们吸收了来自股票市场的随机信息,它们往往无法产生可靠的模型并过度拟合数据。此外,使用多模态模型来预测股票波动受到性别偏见的影响,并且缺乏有效的方法来消除这种偏见。为了解决上述问题,我们使用对抗训练来生成扰动,通过在输入空间周围创建抵抗随机信息的区域来模拟固有的随机性和偏差,以提高模型的鲁棒性和公平性。我们在两个真实世界的金融音频数据集上的综合实验表明,该方法的性能超过了当前最先进的解决方案。这证实了对抗性训练在减少股票波动预测任务的随机性和偏差方面的价值。摘要:Stock volatility prediction is an important task in the financial industry. Recent advancements in multimodal methodologies, which integrate both textual and auditory data, have demonstrated significant improvements in this domain, such as earnings calls (Earnings calls are public available and often involve the management team of a public company and interested parties to discuss the company's earnings). However, these multimodal methods have faced two drawbacks. First, they often fail to yield reliable models and overfit the data due to their absorption of stochastic information from the stock market. Moreover, using multimodal models to predict stock volatility suffers from gender bias and lacks an efficient way to eliminate such bias. To address these aforementioned problems, we use adversarial training to generate perturbations that simulate the inherent stochasticity and bias, by creating areas resistant to random information around the input space to improve model robustness and fairness. Our comprehensive experiments on two real-world financial audio datasets reveal that this method exceeds the performance of current state-of-the-art solution. This confirms the value of adversarial training in reducing stochasticity and bias for stock volatility prediction tasks.
机器翻译,仅供参考
