本文经arXiv每日学术速递授权转载
【1】 CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval
链接:https://arxiv.org/abs/2412.13071
备注:accepted at ECIR 2025
摘要:本研究介绍CLASP(对比存储-语音预训练),一种多语言,多模态表示量身定制的音频文本信息检索。CLASP利用了口语内容和文本数据之间的协同作用。在训练过程中,我们使用了新引入的语音文本数据集,该数据集包含从小说到宗教的15个不同类别。CLASP的音频组件将音频频谱图与预训练的自监督语音模型集成在一起,而其语言编码对应部分则采用了一个在100多种语言上进行预训练的句子编码器。这种统一的轻量级模型弥合了各种模式和语言之间的差距,提高了处理和检索多语言和多模式数据的有效性。我们对多种语言的评估表明,CLASP在HITS@1,MRR和meanR指标中建立了新的基准,在特定场景中优于传统的基于ASR的检索方法。
摘要:This study introduces CLASP (Contrastive Language-Speech Pretraining), amultilingual, multimodal representation tailored for audio-text informationretrieval. CLASP leverages the synergy between spoken content and textual data.During training, we utilize our newly introduced speech-text dataset, whichencompasses 15 diverse categories ranging from fiction to religion. CLASP'saudio component integrates audio spectrograms with a pre-trainedself-supervised speech model, while its language encoding counterpart employs asentence encoder pre-trained on over 100 languages. This unified lightweightmodel bridges the gap between various modalities and languages, enhancing itseffectiveness in handling and retrieving multilingual and multimodal data. Ourevaluations across multiple languages demonstrate that CLASP establishes newbenchmarks in HITS@1, MRR, and meanR metrics, outperforming traditionalASR-based retrieval approaches in specific scenarios.
标题:多模式大型语言模型的模式不一致连续学习
链接:https://arxiv.org/abs/2412.13050
摘要:在本文中,我们介绍了模态不一致的持续学习(MICL),这是多模态大型语言模型(MLLM)的一种新的持续学习场景,涉及具有不一致模态(图像,音频或视频)和不同任务类型(字幕或问答)的任务。与现有的仅视觉或模态增量设置不同,MICI结合了模态和任务类型的转变,两者都会导致灾难性的遗忘。为了解决这些挑战,我们提出了MoInCL,它采用了一个伪目标生成模块,以减轻遗忘所造成的任务类型的变化,在以前看到的方式。它还结合了基于教学的知识蒸馏,以保留模型在引入新模式时处理先前学习的模式的能力。我们基准MICL使用共六个任务,并进行实验,以验证我们提出的MoInCL的有效性。实验结果突出了MoInCL的优越性,显示出比代表性和最先进的持续学习基线有显着改进。
摘要:In this paper, we introduce Modality-Inconsistent Continual Learning (MICL),a new continual learning scenario for Multimodal Large Language Models (MLLMs)that involves tasks with inconsistent modalities (image, audio, or video) andvarying task types (captioning or question-answering). Unlike existingvision-only or modality-incremental settings, MICL combines modality and tasktype shifts, both of which drive catastrophic forgetting. To address thesechallenges, we propose MoInCL, which employs a Pseudo Targets Generation Moduleto mitigate forgetting caused by task type shifts in previously seenmodalities. It also incorporates Instruction-based Knowledge Distillation topreserve the model's ability to handle previously learned modalities when newones are introduced. We benchmark MICL using a total of six tasks and conductexperiments to validate the effectiveness of our proposed MoInCL. Theexperimental results highlight the superiority of MoInCL, showing significantimprovements over representative and state-of-the-art continual learningbaselines.
标题:TAME:用于增强无人机轨迹估计和分类的基于时间音频的曼巴
链接:https://arxiv.org/abs/2412.13037
摘要:紧凑型无人机的日益普及给公共安全带来了重大风险,而传统的无人机检测系统往往体积庞大且成本高昂。为了解决这些挑战,我们提出了TAME,用于增强无人机轨迹估计和分类的基于时间音频的曼巴。这种创新的反无人机检测模型利用并行选择性状态空间模型来同时捕获和学习音频的时间和频谱特征,从而有效地分析声音的传播。为了进一步增强时间特征,我们引入了一个时间特征增强模块,它使用残差交叉注意将光谱特征集成到时间数据中。这种增强的时间信息,然后用于精确的3D轨迹估计和分类。我们的模型在MMUAD基准测试中树立了新的性能标准,展示了卓越的准确性和有效性。代码和训练模型在GitHub \url{https://github.com/AmazingDay1/TAME}上公开提供。
摘要:The increasing prevalence of compact UAVs has introduced significant risks topublic safety, while traditional drone detection systems are often bulky andcostly. To address these challenges, we present TAME, the Temporal Audio-basedMamba for Enhanced Drone Trajectory Estimation and Classification. Thisinnovative anti-UAV detection model leverages a parallel selective state-spacemodel to simultaneously capture and learn both the temporal and spectralfeatures of audio, effectively analyzing propagation of sound. To furtherenhance temporal features, we introduce a Temporal Feature Enhancement Module,which integrates spectral features into temporal data using residualcross-attention. This enhanced temporal information is then employed forprecise 3D trajectory estimation and classification. Our model sets a newstandard of performance on the MMUAD benchmarks, demonstrating superioraccuracy and effectiveness. The code and trained models are publicly availableon GitHub \url{https://github.com/AmazingDay1/TAME}.
标题:CAMEL:代码切换语音识别的交叉注意增强型专家混合和语言偏见
链接:https://arxiv.org/abs/2412.12760
备注:Accepted by ICASSP 2025. 5 pages, 2 figures
摘要:语码转换自动语音识别(ASR)旨在准确地转录包含两种或两种以上语言的语音。为了更好地捕获特定于语言的语音表示并解决代码切换ASR中的语言混乱,通常采用专家混合(MoE)架构和额外的语言日记化(LD)解码器。然而,大多数研究仍然停留在简单的操作,如加权求和或拼接融合特定语言的语音表示,留下了重大的机会,探索整合语言偏见信息的增强。在本文中,我们介绍了CAMEL,一个基于交叉注意的MoE和语言偏见的方法,用于语码转换ASR。具体来说,在每个MoE层之后,我们将特定语言的语音表示与交叉注意融合,利用其强大的上下文建模能力。此外,我们设计了一个基于源注意力的机制,将语言信息从LD解码器输出到文本嵌入。实验结果表明,该方法在SEAME、ASRU200和ASRU700+LibriSpeech460汉英语码转换ASR数据集上均取得了较好的性能。
摘要:Code-switching automatic speech recognition (ASR) aims to transcribe speechthat contains two or more languages accurately. To better capturelanguage-specific speech representations and address language confusion incode-switching ASR, the mixture-of-experts (MoE) architecture and an additionallanguage diarization (LD) decoder are commonly employed. However, mostresearches remain stagnant in simple operations like weighted summation orconcatenation to fuse language-specific speech representations, leavingsignificant opportunities to explore the enhancement of integrating languagebias information. In this paper, we introduce CAMEL, a cross-attention-basedMoE and language bias approach for code-switching ASR. Specifically, after eachMoE layer, we fuse language-specific speech representations withcross-attention, leveraging its strong contextual modeling abilities.Additionally, we design a source attention-based mechanism to incorporate thelanguage information from the LD decoder output into text embeddings.Experimental results demonstrate that our approach achieves state-of-the-artperformance on the SEAME, ASRU200, and ASRU700+LibriSpeech460 Mandarin-Englishcode-switching ASR datasets.
标题:基于音频阵列的3D无人机轨迹估计LiDART伪标记
链接:https://arxiv.org/abs/2412.12698
摘要:随着小型无人机(UAV)变得越来越普遍,人们越来越担心它们对公共安全和隐私的影响,这凸显了对先进跟踪和轨迹估计解决方案的需求。为此,本文提出了一种新的框架,利用音频阵列的三维无人机轨迹估计。我们的方法结合了一个自监督学习模型,从音频数据转换为梅尔频谱图,通过编码器进行分析,以提取关键的时间和频谱信息。同时,无人机轨迹估计使用激光雷达点云通过无监督的方法。这些基于激光雷达的估计充当伪标签,使得能够在不需要标记数据的情况下训练音频感知网络。在这种架构中,基于LiDAR的系统作为教师网络运行,指导作为学生网络的音频感知网络。一旦经过训练,该模型可以仅使用音频信号独立预测3D轨迹,而无需在部署期间使用LiDAR数据或外部地面实况。为了进一步提高精度,我们应用高斯过程建模改进时空跟踪。我们的方法在MMAUD数据集上提供了顶级性能,使用自监督学习技术建立了轨迹估计的新基准,而不依赖于地面实况注释。
摘要:As small unmanned aerial vehicles (UAVs) become increasingly prevalent, thereis growing concern regarding their impact on public safety and privacy,highlighting the need for advanced tracking and trajectory estimationsolutions. In response, this paper introduces a novel framework that utilizesaudio array for 3D UAV trajectory estimation. Our approach incorporates aself-supervised learning model, starting with the conversion of audio data intomel-spectrograms, which are analyzed through an encoder to extract crucialtemporal and spectral information. Simultaneously, UAV trajectories areestimated using LiDAR point clouds via unsupervised methods. These LiDAR-basedestimations act as pseudo labels, enabling the training of an Audio PerceptionNetwork without requiring labeled data. In this architecture, the LiDAR-basedsystem operates as the Teacher Network, guiding the Audio Perception Network,which serves as the Student Network. Once trained, the model can independentlypredict 3D trajectories using only audio signals, with no need for LiDAR dataor external ground truth during deployment. To further enhance precision, weapply Gaussian Process modeling for improved spatiotemporal tracking. Ourmethod delivers top-tier performance on the MMAUD dataset, establishing a newbenchmark in trajectory estimation using self-supervised learning techniqueswithout reliance on ground truth annotations.
标题:音素级特征差异:检测复杂语音Deepfakes的关键
链接:https://arxiv.org/abs/2412.12619
摘要:文本到语音和语音转换技术的最新进展使得能够创建高度令人信服的合成语音。虽然这些创新提供了许多实际好处,但当被恶意滥用时,它们也会带来重大的安全挑战。因此,迫切需要对这些合成语音信号进行检测。音素特征为deepfake检测提供了强大的语音表示。然而,以前的基于音素的检测方法通常集中在特定的音素,忽略了整个音素序列的时间不一致。在本文中,我们开发了一种新的机制,通过识别音素级语音特征的不一致性来检测语音deepfake。我们设计了一个自适应音素池技术,提取样本特定的音素级功能,从帧级语音数据。通过将这种技术应用于由预先训练的音频模型在以前看不见的deepfake数据集上提取的特征,我们证明了与真实语音相比,deepfake样本通常表现出音素级别的不一致性。为了进一步提高检测准确性,我们提出了一个deepfake检测器,它使用图形注意力网络来建模音素级特征的时间依赖性。此外,我们引入了随机音素替换增强技术,以增加训练过程中的特征多样性。在四个基准数据集上进行的大量实验表明,我们的方法比现有的最先进的检测方法具有更好的性能。
摘要:Recent advancements in text-to-speech and speech conversion technologies haveenabled the creation of highly convincing synthetic speech. While theseinnovations offer numerous practical benefits, they also cause significantsecurity challenges when maliciously misused. Therefore, there is an urgentneed to detect these synthetic speech signals. Phoneme features provide apowerful speech representation for deepfake detection. However, previousphoneme-based detection approaches typically focused on specific phonemes,overlooking temporal inconsistencies across the entire phoneme sequence. Inthis paper, we develop a new mechanism for detecting speech deepfakes byidentifying the inconsistencies of phoneme-level speech features. We design anadaptive phoneme pooling technique that extracts sample-specific phoneme-levelfeatures from frame-level speech data. By applying this technique to featuresextracted by pre-trained audio models on previously unseen deepfake datasets,we demonstrate that deepfake samples often exhibit phoneme-levelinconsistencies when compared to genuine speech. To further enhance detectionaccuracy, we propose a deepfake detector that uses a graph attention network tomodel the temporal dependencies of phoneme-level features. Additionally, weintroduce a random phoneme substitution augmentation technique to increasefeature diversity during training. Extensive experiments on four benchmarkdatasets demonstrate the superior performance of our method over existingstate-of-the-art detection methods.
标题:Libri2Vox数据集:采用不同说话者条件和合成数据的目标说话者提取
链接:https://arxiv.org/abs/2412.12512
摘要:目标说话人提取(TSE)在语音处理应用中是必不可少的,特别是在复杂的声学环境中。目前的TSE系统面临着有限的数据多样性和在现实世界条件下缺乏鲁棒性的挑战,主要是因为它们是在具有有限的说话者可变性和不切实际的噪声分布的人工混合数据集上训练的。为了解决这些挑战,我们提出了Libri2Vox,这是一个新的数据集,它将LibriTTS数据集的干净目标语音与嘈杂VoxCeleb2数据集的干扰语音相结合,在现实的嘈杂条件下提供了大量不同的扬声器。我们还使用最先进的语音生成模型生成的合成扬声器来增强Libri2Vox,以增强扬声器的多样性。此外,为了进一步提高合并合成数据的有效性,实施课程学习以逐步训练具有增加的难度水平的TSE模型。在多个TSE架构上进行的广泛实验显示出不同程度的改进,SpeakerBeam表现出最大的增益:与基线训练相比,Libri2Talker测试集上的信号失真比(SDR)提高了1.39 dB。在这些结果的基础上,我们通过使用Conformer架构的基于扬声器相似性的课程学习方法进一步增强了性能,与传统的随机采样方法相比,实现了额外的0.78 dB的改进,其中数据样本是从整个数据集中随机选择的。这些结果表明,在构建强大的TSE系统中,各种真实世界数据、合成扬声器增强和结构化训练策略具有互补优势。
摘要:Target speaker extraction (TSE) is essential in speech processingapplications, particularly in scenarios with complex acoustic environments.Current TSE systems face challenges in limited data diversity and a lack ofrobustness in real-world conditions, primarily because they are trained onartificially mixed datasets with limited speaker variability and unrealisticnoise profiles. To address these challenges, we propose Libri2Vox, a newdataset that combines clean target speech from the LibriTTS dataset withinterference speech from the noisy VoxCeleb2 dataset, providing a large anddiverse set of speakers under realistic noisy conditions. We also augmentLibri2Vox with synthetic speakers generated using state-of-the-art speechgenerative models to enhance speaker diversity. Additionally, to furtherimprove the effectiveness of incorporating synthetic data, curriculum learningis implemented to progressively train TSE models with increasing levels ofdifficulty. Extensive experiments across multiple TSE architectures revealvarying degrees of improvement, with SpeakerBeam demonstrating the mostsubstantial gains: a 1.39 dB improvement in signal-to-distortion ratio (SDR) onthe Libri2Talker test set compared to baseline training. Building upon theseresults, we further enhanced performance through our speaker similarity-basedcurriculum learning approach with the Conformer architecture, achieving anadditional 0.78 dB improvement over conventional random sampling methods inwhich data samples are randomly selected from the entire dataset. These resultsdemonstrate the complementary benefits of diverse real-world data, syntheticspeaker augmentation, and structured training strategies in building robust TSEsystems.
标题:语音合成中情感渲染的分层控制
链接:https://arxiv.org/abs/2412.12498
备注:Submitted to IEEE Transactions
摘要:情感文本到语音合成(TTS)的目的是从输入文本生成逼真的情感语音。然而,定量控制多层次情感渲染仍然具有挑战性。在本文中,我们提出了一个基于扩散的情感TTS框架,并采用了一种新颖的情感强度建模方法,以促进对音素、单词和话语级别的情感渲染的细粒度控制。我们引入了一个分层的情感分布(ED)提取器,捕获一个可量化的ED嵌入在不同的语音段水平。此外,我们探索各种声学特征,并评估其对情绪强度建模的影响。在TTS训练期间,分层ED嵌入有效地捕获来自参考音频的情感强度的变化,并将其与语言和说话者信息相关联。TTS模型不仅在推理过程中产生情感语音,而且定量地控制语音成分上的情感渲染。客观和主观的评价表明,我们的框架在语音质量,情感表达和层次情感控制方面的有效性。
摘要:Emotional text-to-speech synthesis (TTS) aims to generate realistic emotionalspeech from input text. However, quantitatively controlling multi-level emotionrendering remains challenging. In this paper, we propose a diffusion-basedemotional TTS framework with a novel approach for emotion intensity modeling tofacilitate fine-grained control over emotion rendering at the phoneme, word,and utterance levels. We introduce a hierarchical emotion distribution (ED)extractor that captures a quantifiable ED embedding across different speechsegment levels. Additionally, we explore various acoustic features and assesstheir impact on emotion intensity modeling. During TTS training, thehierarchical ED embedding effectively captures the variance in emotionintensity from the reference audio and correlates it with linguistic andspeaker information. The TTS model not only generates emotional speech duringinference, but also quantitatively controls the emotion rendering over thespeech constituents. Both objective and subjective evaluations demonstrate theeffectiveness of our framework in terms of speech quality, emotionalexpressiveness, and hierarchical emotion control.
标题:四类昆虫的健全分类
链接:https://arxiv.org/abs/2412.12395
备注:The manuscript is in submission
摘要:这个项目的目标是对四种不同的昆虫声音进行分类:蟋蟀、甲虫、白蚁和蟋蟀。该项目的一个应用是害虫控制,以监测和保护我们的生态系统。我们的项目利用数据增强,包括音高变换和速度变化,以提高模型的泛化能力。本项目将结合MFCC特征测试决策树、随机森林、SVM RBF、XGBoost和k-NN模型的性能。这个项目的一个潜在的新颖性是,使用了各种数据增强技术,并与原始声音一起创建了6个数据。该数据集由这四种昆虫的声音记录组成。该项目旨在实现高的分类精度,并减少过度拟合问题。
摘要:The goal of this project is to classify four different insect sounds: cicada,beetle, termite, and cricket. One application of this project is for pestcontrol to monitor and protect our ecosystem. Our project leverages dataaugmentation, including pitch shifting and speed changing, to improve modelgeneralization. This project will test the performance of Decision Tree, RandomForest, SVM RBF, XGBoost, and k-NN models, combined with MFCC feature. Apotential novelty of this project is that various data augmentation techniquesare used and created 6 data along with the original sound. The dataset consistsof the sound recordings of these four insects. This project aims to achieve ahigh classification accuracy and to reduce the over-fitting problem.
标题:用于鲁棒脑机接口的想象语音状态分类
链接:https://arxiv.org/abs/2412.12215
摘要:本研究考察了传统机器学习分类器与深度学习模型在使用脑电图数据检测想象语音方面的有效性。具体来说,我们评估了传统的机器学习技术,如CSP-SVM和LDA-SVM分类器,以及深度学习架构,如EEGNet,ShallowConvNet和DeepConvNet。机器学习分类器表现出显着较低的精度和召回率,表明有限的特征提取能力和想象的语音和空闲状态之间的泛化能力差。相比之下,深度学习模型,特别是EEGNet,达到了0.7080的最高准确度和0.6718的F1得分,证明了它们在自动特征提取和表示学习方面的增强能力,这对于捕获复杂的神经生理模式至关重要。这些发现突出了传统机器学习方法在脑机接口(BCI)应用中的局限性,并倡导采用深度学习方法来实现对检测想象语音的更精确和可靠的分类。这一基础研究有助于开发基于想象语音的BCI系统。
摘要:This study examines the effectiveness of traditional machine learningclassifiers versus deep learning models for detecting the imagined speech usingelectroencephalogram data. Specifically, we evaluated conventional machinelearning techniques such as CSP-SVM and LDA-SVM classifiers alongside deeplearning architectures such as EEGNet, ShallowConvNet, and DeepConvNet. Machinelearning classifiers exhibited significantly lower precision and recall,indicating limited feature extraction capabilities and poor generalizationbetween imagined speech and idle states. In contrast, deep learning models,particularly EEGNet, achieved the highest accuracy of 0.7080 and an F1 score of0.6718, demonstrating their enhanced ability in automatic feature extractionand representation learning, essential for capturing complex neurophysiologicalpatterns. These findings highlight the limitations of conventional machinelearning approaches in brain-computer interface (BCI) applications and advocatefor adopting deep learning methodologies to achieve more precise and reliableclassification of detecting imagined speech. This foundational researchcontributes to the development of imagined speech-based BCI systems.
标题:多语言背景下语音生物标志物分析和发音障碍言语的自动严重程度分类
链接:https://arxiv.org/abs/2412.12111
备注:SNU Doctoral thesis
摘要:构音障碍是一种运动性言语障碍,严重影响语音质量、发音和韵律,导致言语清晰度降低和生活质量下降。准确的评估是有效治疗的关键,但传统的感知评估受到主观性和资源强度的限制。为了减轻这些限制,已经提出了构音障碍语音自动评估方法来支持临床医生的决策。虽然这些方法已经显示出有希望的结果,但大多数研究都集中在单语环境中。然而,多语言方法对于解决构音障碍的全球负担并确保公平获得准确诊断是必要的。本论文提出一种新的多语言构音障碍严重程度分类方法,借由分析三种语言:英语、韩语、及泰米尔语。
摘要:Dysarthria, a motor speech disorder, severely impacts voice quality,pronunciation, and prosody, leading to diminished speech intelligibility andreduced quality of life. Accurate assessment is crucial for effectivetreatment, but traditional perceptual assessments are limited by theirsubjectivity and resource intensity. To mitigate the limitations, automaticdysarthric speech assessment methods have been proposed to support clinicianson their decision-making. While these methods have shown promising results,most research has focused on monolingual environments. However, multilingualapproaches are necessary to address the global burden of dysarthria and ensureequitable access to accurate diagnosis. This thesis proposes a novelmultilingual dysarthria severity classification method, by analyzing threelanguages: English, Korean, and Tamil.
标题:跨层识别一致性增强流媒体关键词发现
链接:https://arxiv.org/abs/2412.12635
备注:Submitted to ICASSP2025
摘要:连接主义时态分类(CTC)是一种非自回归的训练准则,被广泛应用于在线关键词识别(KWS)。然而,现有的基于CTC的KWS解码策略要么依赖于自动语音识别(ASR),其由于其在声学空间上的广泛搜索而没有关键字特定的优化而执行次优,要么依赖于KWS特定的解码图,其实现和维护复杂。在这项工作中,我们提出了一个流解码算法增强跨层歧视一致性(CDC),量身定制的CTC为基础的KWS。具体来说,我们引入了一个精简而有效的解码算法,能够检测到在任何任意位置的关键字的开始。此外,我们利用跨层的判别一致性信息来更好地区分正报警样本和假报警样本。我们在干净和有噪声的Hey Snips数据集上的实验表明,所提出的流解码策略优于基于ASR和基于图的KWS基线。CDC增强的解码进一步提高了性能,与基于图形的KWS基线相比,平均绝对召回率提高了6.8%,未命中率相对降低了46.3%,误报率非常低,为每小时0.05。
摘要:Connectionist Temporal Classification (CTC), a non-autoregressive trainingcriterion, is widely used in online keyword spotting (KWS). However, existingCTC-based KWS decoding strategies either rely on Automatic Speech Recognition(ASR), which performs suboptimally due to its broad search over the acousticspace without keyword-specific optimization, or on KWS-specific decodinggraphs, which are complex to implement and maintain. In this work, we propose astreaming decoding algorithm enhanced by Cross-layer Discrimination Consistency(CDC), tailored for CTC-based KWS. Specifically, we introduce a streamlined yeteffective decoding algorithm capable of detecting the start of the keyword atany arbitrary position. Furthermore, we leverage discrimination consistencyinformation across layers to better differentiate between positive and falsealarm samples. Our experiments on both clean and noisy Hey Snips datasets showthat the proposed streaming decoding strategy outperforms ASR-based andgraph-based KWS baselines. The CDC-boosted decoding further improvesperformance, yielding an average absolute recall improvement of 6.8% and a46.3% relative reduction in the miss rate compared to the graph-based KWSbaseline, with a very low false alarm rate of 0.05 per hour.
标题:NTC-KWS:用于强大关键字发现的噪音感知CTC
链接:https://arxiv.org/abs/2412.12614
备注:Submitted to ICASSP 2025
摘要:近年来,人们对设计小占用空间但有效的基于连接主义时间分类的关键字定位(CTC-KWS)系统的兴趣越来越大。它们通常部署在低资源计算平台上,其中模型大小和计算能力的限制在复杂的声学场景下产生瓶颈。这样的约束通常会导致关键字和背景噪声之间的过度拟合和混淆,从而导致高误报。为了解决这些问题,我们提出了一个噪声感知的CTC为基础的KWS(NTC-KWS)框架,旨在提高模型的鲁棒性在嘈杂的环境中,特别是在极低的信噪比。我们的方法引入了两个额外的噪声建模弧的训练和解码过程的基础上加权有限状态转换器(WFST)图:自循环弧,以解决噪声插入错误和旁路弧处理掩蔽和干扰造成的过度噪声。在干净和嘈杂的Hey Snips上进行的实验表明,NTC-KWS在各种声学条件下的性能优于最先进的(SOTA)端到端系统和CTC-KWS基线,在低SNR场景中表现尤其出色。
摘要:In recent years, there has been a growing interest in designingsmall-footprint yet effective Connectionist Temporal Classification basedkeyword spotting (CTC-KWS) systems. They are typically deployed on low-resourcecomputing platforms, where limitations on model size and computational capacitycreate bottlenecks under complicated acoustic scenarios. Such constraints oftenresult in overfitting and confusion between keywords and background noise,leading to high false alarms. To address these issues, we propose a noise-awareCTC-based KWS (NTC-KWS) framework designed to enhance model robustness in noisyenvironments, particularly under extremely low signal-to-noise ratios. Ourapproach introduces two additional noise-modeling wildcard arcs into thetraining and decoding processes based on weighted finite state transducer(WFST) graphs: self-loop arcs to address noise insertion errors and bypass arcsto handle masking and interference caused by excessive noise. Experiments onclean and noisy Hey Snips show that NTC-KWS outperforms state-of-the-art (SOTA)end-to-end systems and CTC-KWS baselines across various acoustic conditions,with particularly strong performance in low SNR scenarios.
标题:跨层识别一致性增强流媒体关键词发现
链接:https://arxiv.org/abs/2412.12635
备注:Submitted to ICASSP2025
摘要:连接主义时态分类(CTC)是一种非自回归的训练准则,被广泛应用于在线关键词识别(KWS)。然而,现有的基于CTC的KWS解码策略要么依赖于自动语音识别(ASR),其由于其在声学空间上的广泛搜索而没有关键字特定的优化而执行次优,要么依赖于KWS特定的解码图,其实现和维护复杂。在这项工作中,我们提出了一个流解码算法增强跨层歧视一致性(CDC),量身定制的CTC为基础的KWS。具体来说,我们引入了一个精简而有效的解码算法,能够检测到在任何任意位置的关键字的开始。此外,我们利用跨层的判别一致性信息来更好地区分正报警样本和假报警样本。我们在干净和有噪声的Hey Snips数据集上的实验表明,所提出的流解码策略优于基于ASR和基于图的KWS基线。CDC增强的解码进一步提高了性能,与基于图形的KWS基线相比,平均绝对召回率提高了6.8%,未命中率相对降低了46.3%,误报率非常低,为每小时0.05。
摘要:Connectionist Temporal Classification (CTC), a non-autoregressive trainingcriterion, is widely used in online keyword spotting (KWS). However, existingCTC-based KWS decoding strategies either rely on Automatic Speech Recognition(ASR), which performs suboptimally due to its broad search over the acousticspace without keyword-specific optimization, or on KWS-specific decodinggraphs, which are complex to implement and maintain. In this work, we propose astreaming decoding algorithm enhanced by Cross-layer Discrimination Consistency(CDC), tailored for CTC-based KWS. Specifically, we introduce a streamlined yeteffective decoding algorithm capable of detecting the start of the keyword atany arbitrary position. Furthermore, we leverage discrimination consistencyinformation across layers to better differentiate between positive and falsealarm samples. Our experiments on both clean and noisy Hey Snips datasets showthat the proposed streaming decoding strategy outperforms ASR-based andgraph-based KWS baselines. The CDC-boosted decoding further improvesperformance, yielding an average absolute recall improvement of 6.8% and a46.3% relative reduction in the miss rate compared to the graph-based KWSbaseline, with a very low false alarm rate of 0.05 per hour.
标题:NTC-KWS:用于强大关键字发现的噪音感知CTC
链接:https://arxiv.org/abs/2412.12614
备注:Submitted to ICASSP 2025
摘要:近年来,人们对设计小占用空间但有效的基于连接主义时间分类的关键字定位(CTC-KWS)系统的兴趣越来越大。它们通常部署在低资源计算平台上,其中模型大小和计算能力的限制在复杂的声学场景下产生瓶颈。这样的约束通常会导致关键字和背景噪声之间的过度拟合和混淆,从而导致高误报。为了解决这些问题,我们提出了一个噪声感知的CTC为基础的KWS(NTC-KWS)框架,旨在提高模型的鲁棒性在嘈杂的环境中,特别是在极低的信噪比。我们的方法引入了两个额外的噪声建模弧的训练和解码过程的基础上加权有限状态转换器(WFST)图:自循环弧,以解决噪声插入错误和旁路弧处理掩蔽和干扰造成的过度噪声。在干净和嘈杂的Hey Snips上进行的实验表明,NTC-KWS在各种声学条件下的性能优于最先进的(SOTA)端到端系统和CTC-KWS基线,在低SNR场景中表现尤其出色。
摘要:In recent years, there has been a growing interest in designingsmall-footprint yet effective Connectionist Temporal Classification basedkeyword spotting (CTC-KWS) systems. They are typically deployed on low-resourcecomputing platforms, where limitations on model size and computational capacitycreate bottlenecks under complicated acoustic scenarios. Such constraints oftenresult in overfitting and confusion between keywords and background noise,leading to high false alarms. To address these issues, we propose a noise-awareCTC-based KWS (NTC-KWS) framework designed to enhance model robustness in noisyenvironments, particularly under extremely low signal-to-noise ratios. Ourapproach introduces two additional noise-modeling wildcard arcs into thetraining and decoding processes based on weighted finite state transducer(WFST) graphs: self-loop arcs to address noise insertion errors and bypass arcsto handle masking and interference caused by excessive noise. Experiments onclean and noisy Hey Snips show that NTC-KWS outperforms state-of-the-art (SOTA)end-to-end systems and CTC-KWS baselines across various acoustic conditions,with particularly strong performance in low SNR scenarios.
标题:CLASP:用于多语言多模式信息检索的对比语音预训练
链接:https://arxiv.org/abs/2412.13071
备注:accepted at ECIR 2025
摘要:本研究介绍CLASP(对比存储-语音预训练),一种多语言,多模态表示量身定制的音频文本信息检索。CLASP利用了口语内容和文本数据之间的协同作用。在训练过程中,我们利用新引入的语音文本数据集,该数据集包含从小说到宗教等15个不同类别。CLASP的音频组件将音频频谱图与预训练的自监督语音模型集成在一起,而其语言编码对应部分则采用了一个在100多种语言上进行预训练的句子编码器。这种统一的轻量级模型弥合了各种模式和语言之间的差距,提高了处理和检索多语言和多模式数据的有效性。我们对多种语言的评估表明,CLASP在HITS@1,MRR和meanR指标中建立了新的基准,在特定场景中优于传统的基于ASR的检索方法。
摘要:This study introduces CLASP (Contrastive Language-Speech Pretraining), amultilingual, multimodal representation tailored for audio-text informationretrieval. CLASP leverages the synergy between spoken content and textual data.During training, we utilize our newly introduced speech-text dataset, whichencompasses 15 diverse categories ranging from fiction to religion. CLASP'saudio component integrates audio spectrograms with a pre-trainedself-supervised speech model, while its language encoding counterpart employs asentence encoder pre-trained on over 100 languages. This unified lightweightmodel bridges the gap between various modalities and languages, enhancing itseffectiveness in handling and retrieving multilingual and multimodal data. Ourevaluations across multiple languages demonstrate that CLASP establishes newbenchmarks in HITS@1, MRR, and meanR metrics, outperforming traditionalASR-based retrieval approaches in specific scenarios.
标题:多模式大型语言模型的模式不一致连续学习
链接:https://arxiv.org/abs/2412.13050
摘要:在本文中,我们介绍了模态不一致持续学习(MICL),这是一种新的多模态大型语言模型(MLLM)持续学习场景,涉及具有不一致模态(图像、音频或视频)和不同任务类型(字幕或问题)的任务。回答)。与现有的仅视觉或模态增量设置不同,MICI结合了模态和任务类型的转变,两者都会导致灾难性的遗忘。为了解决这些挑战,我们提出了MoInCL,它采用了一个伪目标生成模块,以减轻遗忘所造成的任务类型的变化,在以前看到的方式。它还结合了基于教学的知识蒸馏,以保留模型在引入新模式时处理先前学习的模式的能力。我们基准MICL使用共六个任务,并进行实验,以验证我们提出的MoInCL的有效性。实验结果突出了MoInCL的优越性,显示出比代表性和最先进的持续学习基线有显着改进。
摘要:In this paper, we introduce Modality-Inconsistent Continual Learning (MICL),a new continual learning scenario for Multimodal Large Language Models (MLLMs)that involves tasks with inconsistent modalities (image, audio, or video) andvarying task types (captioning or question-answering). Unlike existingvision-only or modality-incremental settings, MICL combines modality and tasktype shifts, both of which drive catastrophic forgetting. To address thesechallenges, we propose MoInCL, which employs a Pseudo Targets Generation Moduleto mitigate forgetting caused by task type shifts in previously seenmodalities. It also incorporates Instruction-based Knowledge Distillation topreserve the model's ability to handle previously learned modalities when newones are introduced. We benchmark MICL using a total of six tasks and conductexperiments to validate the effectiveness of our proposed MoInCL. Theexperimental results highlight the superiority of MoInCL, showing significantimprovements over representative and state-of-the-art continual learningbaselines.
标题:TAME:用于增强无人机轨迹估计和分类的基于时间音频的曼巴
链接:https://arxiv.org/abs/2412.13037
摘要:紧凑型无人机的日益普及给公共安全带来了重大风险,而传统的无人机检测系统往往体积庞大且成本高昂。为了解决这些挑战,我们提出了TAME,用于增强无人机轨迹估计和分类的基于时间音频的曼巴。这种创新的反无人机检测模型利用并行选择性状态空间模型来同时捕获和学习音频的时间和频谱特征,从而有效地分析声音的传播。为了进一步增强时间特征,我们引入了一个时间特征增强模块,它使用残差交叉注意将光谱特征集成到时间数据中。这种增强的时间信息,然后用于精确的3D轨迹估计和分类。我们的模型在MMUAD基准测试中树立了新的性能标准,展示了卓越的准确性和有效性。代码和训练模型在GitHub \url{https://github.com/AmazingDay1/TAME}上公开提供。
摘要:The increasing prevalence of compact UAVs has introduced significant risks topublic safety, while traditional drone detection systems are often bulky andcostly. To address these challenges, we present TAME, the Temporal Audio-basedMamba for Enhanced Drone Trajectory Estimation and Classification. Thisinnovative anti-UAV detection model leverages a parallel selective state-spacemodel to simultaneously capture and learn both the temporal and spectralfeatures of audio, effectively analyzing propagation of sound. To furtherenhance temporal features, we introduce a Temporal Feature Enhancement Module,which integrates spectral features into temporal data using residualcross-attention. This enhanced temporal information is then employed forprecise 3D trajectory estimation and classification. Our model sets a newstandard of performance on the MMUAD benchmarks, demonstrating superioraccuracy and effectiveness. The code and trained models are publicly availableon GitHub \url{https://github.com/AmazingDay1/TAME}.
标题:CAMEL:代码切换语音识别的交叉注意增强型专家混合和语言偏见
链接:https://arxiv.org/abs/2412.12760
备注:Accepted by ICASSP 2025. 5 pages, 2 figures
摘要:语码转换自动语音识别(ASR)旨在准确地转录包含两种或两种以上语言的语音。为了更好地捕获语言特定的语音表示并解决代码切换ASR中的语言混淆,通常采用专家混合(MoE)架构和附加的语言日记化(LD)解码器。然而,大多数研究仍然停留在简单的操作,如加权求和或拼接融合特定语言的语音表示,留下了重大的机会,探索整合语言偏见信息的增强。在本文中,我们介绍了CAMEL,一个基于交叉注意的MoE和语言偏见的方法,用于语码转换ASR。具体来说,在每个MoE层之后,我们将特定语言的语音表示与交叉注意融合,利用其强大的上下文建模能力。此外,我们设计了一个基于源注意力的机制,将语言信息从LD解码器输出到文本嵌入。实验结果表明,该方法在SEAME、ASRU200和ASRU700+LibriSpeech460汉英语码转换ASR数据集上均取得了较好的性能。
摘要:Code-switching automatic speech recognition (ASR) aims to transcribe speechthat contains two or more languages accurately. To better capturelanguage-specific speech representations and address language confusion incode-switching ASR, the mixture-of-experts (MoE) architecture and an additionallanguage diarization (LD) decoder are commonly employed. However, mostresearches remain stagnant in simple operations like weighted summation orconcatenation to fuse language-specific speech representations, leavingsignificant opportunities to explore the enhancement of integrating languagebias information. In this paper, we introduce CAMEL, a cross-attention-basedMoE and language bias approach for code-switching ASR. Specifically, after eachMoE layer, we fuse language-specific speech representations withcross-attention, leveraging its strong contextual modeling abilities.Additionally, we design a source attention-based mechanism to incorporate thelanguage information from the LD decoder output into text embeddings.Experimental results demonstrate that our approach achieves state-of-the-artperformance on the SEAME, ASRU200, and ASRU700+LibriSpeech460 Mandarin-Englishcode-switching ASR datasets.
标题:基于音频阵列的3D无人机轨迹估计LiDART伪标记
链接:https://arxiv.org/abs/2412.12698
摘要:随着小型无人机(UAV)变得越来越普遍,人们越来越担心它们对公共安全和隐私的影响,这凸显了对先进跟踪和轨迹估计解决方案的需求。为此,本文提出了一种新的框架,利用音频阵列的三维无人机轨迹估计。我们的方法结合了一个自监督学习模型,从音频数据转换为梅尔频谱图,通过编码器进行分析,以提取关键的时间和频谱信息。同时,无人机轨迹估计使用激光雷达点云通过无监督的方法。这些基于LiDAR的估计充当伪标签,使得能够在不需要标记数据的情况下训练音频感知网络。在这种架构中,基于LiDAR的系统作为教师网络运行,指导作为学生网络的音频感知网络。一旦经过训练,该模型可以仅使用音频信号独立预测3D轨迹,而无需在部署期间使用LiDAR数据或外部地面实况。为了进一步提高精度,我们应用高斯过程建模改进时空跟踪。我们的方法在MMAUD数据集上提供了顶级性能,使用自监督学习技术建立了轨迹估计的新基准,而不依赖于地面实况注释。
摘要:As small unmanned aerial vehicles (UAVs) become increasingly prevalent, thereis growing concern regarding their impact on public safety and privacy,highlighting the need for advanced tracking and trajectory estimationsolutions. In response, this paper introduces a novel framework that utilizesaudio array for 3D UAV trajectory estimation. Our approach incorporates aself-supervised learning model, starting with the conversion of audio data intomel-spectrograms, which are analyzed through an encoder to extract crucialtemporal and spectral information. Simultaneously, UAV trajectories areestimated using LiDAR point clouds via unsupervised methods. These LiDAR-basedestimations act as pseudo labels, enabling the training of an Audio PerceptionNetwork without requiring labeled data. In this architecture, the LiDAR-basedsystem operates as the Teacher Network, guiding the Audio Perception Network,which serves as the Student Network. Once trained, the model can independentlypredict 3D trajectories using only audio signals, with no need for LiDAR dataor external ground truth during deployment. To further enhance precision, weapply Gaussian Process modeling for improved spatiotemporal tracking. Ourmethod delivers top-tier performance on the MMAUD dataset, establishing a newbenchmark in trajectory estimation using self-supervised learning techniqueswithout reliance on ground truth annotations.
标题:音素级特征差异:检测复杂语音Deepfakes的关键
链接:https://arxiv.org/abs/2412.12619
摘要:文本到语音和语音转换技术的最新进展使得能够创建高度令人信服的合成语音。虽然这些创新提供了许多实际好处,但当被恶意滥用时,它们也会带来重大的安全挑战。因此,迫切需要对这些合成语音信号进行检测。音素特征为deepfake检测提供了强大的语音表示。然而,以前的基于音素的检测方法通常集中在特定的音素,忽略了整个音素序列的时间不一致。在本文中,我们开发了一种新的机制,通过识别音素级语音特征的不一致性来检测语音deepfake。我们设计了一个自适应音素池技术,提取样本特定的音素级功能,从帧级语音数据。通过将这种技术应用于由预先训练的音频模型在以前看不见的deepfake数据集上提取的特征,我们证明了与真实语音相比,deepfake样本通常表现出音素级别的不一致性。为了进一步提高检测准确性,我们提出了一个deepfake检测器,它使用图形注意力网络来建模音素级特征的时间依赖性。此外,我们引入了随机音素替换增强技术,以增加训练过程中的特征多样性。在四个基准数据集上进行的大量实验表明,我们的方法比现有的最先进的检测方法具有更好的性能。
摘要:Recent advancements in text-to-speech and speech conversion technologies haveenabled the creation of highly convincing synthetic speech. While theseinnovations offer numerous practical benefits, they also cause significantsecurity challenges when maliciously misused. Therefore, there is an urgentneed to detect these synthetic speech signals. Phoneme features provide apowerful speech representation for deepfake detection. However, previousphoneme-based detection approaches typically focused on specific phonemes,overlooking temporal inconsistencies across the entire phoneme sequence. Inthis paper, we develop a new mechanism for detecting speech deepfakes byidentifying the inconsistencies of phoneme-level speech features. We design anadaptive phoneme pooling technique that extracts sample-specific phoneme-levelfeatures from frame-level speech data. By applying this technique to featuresextracted by pre-trained audio models on previously unseen deepfake datasets,we demonstrate that deepfake samples often exhibit phoneme-levelinconsistencies when compared to genuine speech. To further enhance detectionaccuracy, we propose a deepfake detector that uses a graph attention network tomodel the temporal dependencies of phoneme-level features. Additionally, weintroduce a random phoneme substitution augmentation technique to increasefeature diversity during training. Extensive experiments on four benchmarkdatasets demonstrate the superior performance of our method over existingstate-of-the-art detection methods.
标题:Libri2Vox数据集:采用不同说话者条件和合成数据的目标说话者提取
链接:https://arxiv.org/abs/2412.12512
摘要:目标说话人提取(TSE)在语音处理应用中是必不可少的,特别是在复杂的声学环境中。目前的TSE系统面临着有限的数据多样性和在现实世界条件下缺乏鲁棒性的挑战,主要是因为它们是在具有有限的说话者可变性和不切实际的噪声分布的人工混合数据集上训练的。为了解决这些挑战,我们提出了Libri2Vox,这是一个新的数据集,它将LibriTTS数据集的干净目标语音与嘈杂VoxCeleb2数据集的干扰语音相结合,在现实的嘈杂条件下提供了大量不同的扬声器。我们还使用最先进的语音生成模型生成的合成扬声器来增强Libri2Vox,以增强扬声器的多样性。此外,为了进一步提高合并合成数据的有效性,实施课程学习以逐步训练具有增加的难度水平的TSE模型。在多个TSE架构上进行的广泛实验显示出不同程度的改进,SpeakerBeam表现出最大的增益:与基线训练相比,Libri2Talker测试集上的信号失真比(SDR)提高了1.39 dB。在这些结果的基础上,我们通过使用Conformer架构的基于扬声器相似性的课程学习方法进一步增强了性能,与传统的随机采样方法相比,实现了额外的0.78 dB的改进,其中数据样本是从整个数据集中随机选择的。这些结果表明,在构建强大的TSE系统中,各种真实世界数据、合成扬声器增强和结构化训练策略具有互补优势。
摘要:Target speaker extraction (TSE) is essential in speech processingapplications, particularly in scenarios with complex acoustic environments.Current TSE systems face challenges in limited data diversity and a lack ofrobustness in real-world conditions, primarily because they are trained onartificially mixed datasets with limited speaker variability and unrealisticnoise profiles. To address these challenges, we propose Libri2Vox, a newdataset that combines clean target speech from the LibriTTS dataset withinterference speech from the noisy VoxCeleb2 dataset, providing a large anddiverse set of speakers under realistic noisy conditions. We also augmentLibri2Vox with synthetic speakers generated using state-of-the-art speechgenerative models to enhance speaker diversity. Additionally, to furtherimprove the effectiveness of incorporating synthetic data, curriculum learningis implemented to progressively train TSE models with increasing levels ofdifficulty. Extensive experiments across multiple TSE architectures revealvarying degrees of improvement, with SpeakerBeam demonstrating the mostsubstantial gains: a 1.39 dB improvement in signal-to-distortion ratio (SDR) onthe Libri2Talker test set compared to baseline training. Building upon theseresults, we further enhanced performance through our speaker similarity-basedcurriculum learning approach with the Conformer architecture, achieving anadditional 0.78 dB improvement over conventional random sampling methods inwhich data samples are randomly selected from the entire dataset. These resultsdemonstrate the complementary benefits of diverse real-world data, syntheticspeaker augmentation, and structured training strategies in building robust TSEsystems.
标题:语音合成中情感渲染的分层控制
链接:https://arxiv.org/abs/2412.12498
备注:Submitted to IEEE Transactions
摘要:情感文本到语音合成(TTS)的目的是从输入文本生成逼真的情感语音。然而,定量控制多层次情感渲染仍然具有挑战性。在本文中,我们提出了一个基于扩散的情感TTS框架与情感强度建模的新方法,以促进细粒度的控制情绪渲染在音素,单词和话语水平。我们引入了一个分层的情感分布(ED)提取器,捕获一个可量化的ED嵌入在不同的语音段水平。此外,我们探索各种声学特征,并评估其对情绪强度建模的影响。在TTS训练期间,分层ED嵌入有效地捕获来自参考音频的情感强度的变化,并将其与语言和说话者信息相关联。TTS模型不仅在推理过程中产生情感语音,而且定量地控制语音成分上的情感渲染。客观和主观的评价表明,我们的框架在语音质量,情感表达和层次情感控制方面的有效性。
摘要:Emotional text-to-speech synthesis (TTS) aims to generate realistic emotionalspeech from input text. However, quantitatively controlling multi-level emotionrendering remains challenging. In this paper, we propose a diffusion-basedemotional TTS framework with a novel approach for emotion intensity modeling tofacilitate fine-grained control over emotion rendering at the phoneme, word,and utterance levels. We introduce a hierarchical emotion distribution (ED)extractor that captures a quantifiable ED embedding across different speechsegment levels. Additionally, we explore various acoustic features and assesstheir impact on emotion intensity modeling. During TTS training, thehierarchical ED embedding effectively captures the variance in emotionintensity from the reference audio and correlates it with linguistic andspeaker information. The TTS model not only generates emotional speech duringinference, but also quantitatively controls the emotion rendering over thespeech constituents. Both objective and subjective evaluations demonstrate theeffectiveness of our framework in terms of speech quality, emotionalexpressiveness, and hierarchical emotion control.
标题:四类昆虫的健全分类
链接:https://arxiv.org/abs/2412.12395
备注:The manuscript is in submission
摘要:这个项目的目标是对四种不同的昆虫声音进行分类:蟋蟀、甲虫、白蚁和蟋蟀。该项目的一个应用是害虫控制,以监测和保护我们的生态系统。我们的项目利用数据增强,包括音高变换和速度变化,以提高模型的泛化能力。该项目将结合MFCC功能测试决策树、随机森林、SVM RBF、XGBoost和k-NN模型的性能。这个项目的一个潜在的新颖性是,使用了各种数据增强技术,并与原始声音一起创建了6个数据。该数据集由这四种昆虫的声音记录组成。该项目旨在实现高的分类精度,并减少过度拟合问题。
摘要:The goal of this project is to classify four different insect sounds: cicada,beetle, termite, and cricket. One application of this project is for pestcontrol to monitor and protect our ecosystem. Our project leverages dataaugmentation, including pitch shifting and speed changing, to improve modelgeneralization. This project will test the performance of Decision Tree, RandomForest, SVM RBF, XGBoost, and k-NN models, combined with MFCC feature. Apotential novelty of this project is that various data augmentation techniquesare used and created 6 data along with the original sound. The dataset consistsof the sound recordings of these four insects. This project aims to achieve ahigh classification accuracy and to reduce the over-fitting problem.
标题:用于鲁棒脑机接口的想象语音状态分类
链接:https://arxiv.org/abs/2412.12215
摘要:本研究考察了传统机器学习分类器与深度学习模型在使用脑电图数据检测想象语音方面的有效性。具体来说,我们评估了传统的机器学习技术,如CSP-SVM和LDA-SVM分类器,以及深度学习架构,如EEGNet,ShallowConvNet和DeepConvNet。机器学习分类器表现出显着较低的精度和召回率,表明有限的特征提取能力和想象的语音和空闲状态之间的泛化能力差。相比之下,深度学习模型,特别是EEGNet,达到了0.7080的最高准确度和0.6718的F1得分,证明了它们在自动特征提取和表示学习方面的增强能力,这对于捕获复杂的神经生理模式至关重要。这些发现突出了传统机器学习方法在脑机接口(BCI)应用中的局限性,并倡导采用深度学习方法来实现对检测想象语音的更精确和可靠的分类。这一基础研究有助于开发基于想象语音的BCI系统。
摘要:This study examines the effectiveness of traditional machine learningclassifiers versus deep learning models for detecting the imagined speech usingelectroencephalogram data. Specifically, we evaluated conventional machinelearning techniques such as CSP-SVM and LDA-SVM classifiers alongside deeplearning architectures such as EEGNet, ShallowConvNet, and DeepConvNet. Machinelearning classifiers exhibited significantly lower precision and recall,indicating limited feature extraction capabilities and poor generalizationbetween imagined speech and idle states. In contrast, deep learning models,particularly EEGNet, achieved the highest accuracy of 0.7080 and an F1 score of0.6718, demonstrating their enhanced ability in automatic feature extractionand representation learning, essential for capturing complex neurophysiologicalpatterns. These findings highlight the limitations of conventional machinelearning approaches in brain-computer interface (BCI) applications and advocatefor adopting deep learning methodologies to achieve more precise and reliableclassification of detecting imagined speech. This foundational researchcontributes to the development of imagined speech-based BCI systems.
标题:Greek 2 MathTex:用于LaTeX方程生成的希腊语音到文本框架
链接:https://arxiv.org/abs/2412.12167
备注:4 pages, 2 figures, SETN2024: 13th EETN Conference on Artificial Intelligence
摘要:在绝大多数学术和科学领域,LaTeX已经成为排版复杂数学方程和公式的事实标准。然而,LaTeX复杂的语法和类似代码的外观为残疾人以及不熟悉编码约定的人带来了可访问性障碍。在本文中,我们提出了一种新的解决方案,通过专门为希腊语设计的一种新的语音到LaTeX方程系统的发展,这一挑战。我们提出了一个端到端的系统,利用自动语音识别(ASR)和自然语言处理(NLP)技术的力量,使用户能够口头口述自然语言中的数学表达式和方程,随后转换为LaTeX格式。我们提出了我们的系统的架构和设计原则,突出了关键组件,如ASR引擎,基于LLM的迭代驱动的方程生成机制,以及在整个开发过程中采用的自定义评估指标的应用。我们已经使我们的系统开源,并可在https://github.com/magcil/greek-speech-to-math。
摘要:In the vast majority of the academic and scientific domains, LaTeX hasestablished itself as the de facto standard for typesetting complexmathematical equations and formulae. However, LaTeX's complex syntax andcode-like appearance present accessibility barriers for individuals withdisabilities, as well as those unfamiliar with coding conventions. In thispaper, we present a novel solution to this challenge through the development ofa novel speech-to-LaTeX equations system specifically designed for the Greeklanguage. We propose an end-to-end system that harnesses the power of AutomaticSpeech Recognition (ASR) and Natural Language Processing (NLP) techniques toenable users to verbally dictate mathematical expressions and equations innatural language, which are subsequently converted into LaTeX format. Wepresent the architecture and design principles of our system, highlighting keycomponents such as the ASR engine, the LLM-based prompt-driven equationsgeneration mechanism, as well as the application of a custom evaluation metricemployed throughout the development process. We have made our system opensource and available at https://github.com/magcil/greek-speech-to-math.
标题:多语言背景下语音生物标志物分析和发音障碍言语的自动严重程度分类
链接:https://arxiv.org/abs/2412.12111
备注:SNU Doctoral thesis
摘要:构音障碍是一种运动性言语障碍,严重影响语音质量、发音和韵律,导致言语清晰度降低和生活质量下降。准确的评估是有效治疗的关键,但传统的感知评估受到主观性和资源强度的限制。为了减轻这些限制,已经提出了构音障碍语音自动评估方法来支持临床医生的决策。虽然这些方法已经显示出有希望的结果,但大多数研究都集中在单语环境中。然而,多语言方法对于解决构音障碍的全球负担并确保公平获得准确诊断是必要的。本论文提出一种新的多语言构音障碍严重程度分类方法,借由分析三种语言:英语、韩语、及泰米尔语。
摘要:Dysarthria, a motor speech disorder, severely impacts voice quality,pronunciation, and prosody, leading to diminished speech intelligibility andreduced quality of life. Accurate assessment is crucial for effectivetreatment, but traditional perceptual assessments are limited by theirsubjectivity and resource intensity. To mitigate the limitations, automaticdysarthric speech assessment methods have been proposed to support clinicianson their decision-making. While these methods have shown promising results,most research has focused on monolingual environments. However, multilingualapproaches are necessary to address the global burden of dysarthria and ensureequitable access to accurate diagnosis. This thesis proposes a novelmultilingual dysarthria severity classification method, by analyzing threelanguages: English, Korean, and Tamil.
