今日论文合集:cs.SD语音10篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音
【1】Local deployment of large-scale music AI models on commodity hardware
标题:在商品硬件上本地部署大型音乐AI模型
链接:https://arxiv.org/abs/2411.09625
作者:Xun Zhou,  Charlie Ruan,  Zihe Zhao,  Tianqi Chen,  Chris Donahue
备注:2 pages
摘要:我们提出了MIDInfinite,这是一个能够在商品硬件上本地使用大规模生成AI模型生成符号音乐的Web应用程序。创建此演示涉及将Anticipatory Music Transformer(一个在Lakh数据集上预训练的大型语言模型(LLM))移植到机器学习编译(MLC)框架。一旦模型被移植,MLC就可以在各种运行时(包括C++、移动设备和浏览器)上进行推理。我们设想MLC有潜力弥合日益强大的音乐AI模型与音乐软件开发人员更熟悉的技术之间的差距。作为概念证明,我们构建了一个Web应用程序,允许用户在浏览器中生成无休止的多工具流,无论是从头开始还是根据提示。在商用硬件(M3 Macbook Pro)上,我们的演示可以每秒生成51个音符,这比实时播放快72.9%的代,并增加到86.3%,有2秒的前期缓冲。
摘要:We present the MIDInfinite, a web application capable of generating symbolicmusic using a large-scale generative AI model locally on commodity hardware.Creating this demo involved porting the Anticipatory Music Transformer, a largelanguage model (LLM) pre-trained on the Lakh MIDI dataset, to the MachineLearning Compilation (MLC) framework. Once the model is ported, MLC facilitatesinference on a variety of runtimes including C++, mobile, and the browser. Weenvision that MLC has the potential to bridge the gap between the landscape ofincreasingly capable music AI models and technology more familiar to musicsoftware developers. As a proof of concept, we build a web application thatallows users to generate endless streams of multi-instrumental MIDI in thebrowser, either from scratch or conditioned on a prompt. On commodity hardware(an M3 Macbook Pro), our demo can generate 51 notes per second, which is fasterthan real-time playback for 72.9% of generations, and increases to 86.3% with 2seconds of upfront buffering.

【2】 ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics  over Acoustic Foundation Models
标题:ParaLBench:计算副语言学优于声学基础模型的大规模基准
链接:https://arxiv.org/abs/2411.09349
作者:Zixing Zhang,  Weixiang Xu,  Zhongren Dong,  Kanglin Wang,  Yimeng Wu,  Jing Peng,  Runming Wang,  Dong-Yan Huang
摘要:计算语言学(Comparal)旨在开发算法和模型来自动检测,分析和解释来自语音交流的非语言信息,例如。例如,在一个实施例中,情绪、健康状况、年龄和性别。尽管其发展迅速,但它严重依赖于特定语言任务的精心设计的模型。因此,CompParal模型的异质性和多样性在很大程度上阻碍了CompParal模型的实际实现。近年来,随着自监督学习的声学基础模型的出现,开发能够有效感知大量非语言信息的通用模型已经成为语音处理中的一个活跃话题。然而,它缺乏一个统一的评价框架,无法进行公平和一致的业绩比较。为了弥合这一差距,我们进行了一个大规模的基准测试,即ParaLBench,它集中于标准化的评估过程中的各种语言任务,包括情感计算的关键方面,如情感识别和情感维度预测,在不同的声学基础模型。该基准测试包含10个数据集,包含13个不同的语言任务,涵盖短期,中期和长期特征。每项任务都在统一的评估框架下对14个声学基础模型进行评估,从而可以进行公正的方法比较,并为Comparal社区提供有依据的参考。根据ParaLBench的见解,我们还指出了潜在的研究方向,即,跨语料库的泛化能力,以推动未来的Comparal研究。与本研究相关的代码将可用于促进后续研究人员这项工作的透明度和可复制性。
摘要:Computational paralinguistics (ComParal) aims to develop algorithms andmodels to automatically detect, analyze, and interpret non-verbal informationfrom speech communication, e. g., emotion, health state, age, and gender.Despite its rapid progress, it heavily depends on sophisticatedly designedmodels given specific paralinguistic tasks. Thus, the heterogeneity anddiversity of ComParal models largely prevent the realistic implementation ofComParal models. Recently, with the advent of acoustic foundation modelsbecause of self-supervised learning, developing more generic models that canefficiently perceive a plethora of paralinguistic information has become anactive topic in speech processing. However, it lacks a unified evaluationframework for a fair and consistent performance comparison. To bridge this gap,we conduct a large-scale benchmark, namely ParaLBench, which concentrates onstandardizing the evaluation process of diverse paralinguistic tasks, includingcritical aspects of affective computing such as emotion recognition and emotiondimensions prediction, over different acoustic foundation models. Thisbenchmark contains ten datasets with thirteen distinct paralinguistic tasks,covering short-, medium- and long-term characteristics. Each task is carriedout on 14 acoustic foundation models under a unified evaluation framework,which allows for an unbiased methodological comparison and offers a groundedreference for the ComParal community. Based on the insights gained fromParaLBench, we also point out potential research directions, i.e., thecross-corpus generalizability, to propel ComParal research in the future. Thecode associated with this study will be available to foster the transparencyand replicability of this work for succeeding researchers.

【3】 Re-Parameterization of Lightweight Transformer for On-Device Speech  Emotion Recognition
标题:用于设备上语音情感识别的轻型Transformer重新参数化
链接:https://arxiv.org/abs/2411.09339
作者:Zixing Zhang,  Zhongren Dong,  Weixiang Xu,  Jing Han
摘要:随着机器学习模型在边缘或物联网(IoT)设备上的实施越来越多,在资源受限的IoT设备上部署高级模型仍然具有挑战性。Transformer模型是目前占主导地位的神经架构,在广泛的领域取得了巨大的成功,但其复杂性阻碍了其在计算能力和存储大小有限的物联网设备上的部署。虽然已经探索了许多模型压缩方法,但它们经常遭受臭名昭著的性能下降。为了解决这个问题,我们引入了一种新的方法,即Transformer重新参数化,以提高轻量级Transformer模型的性能。它包括两个过程:训练阶段中的高阶分解(HRF)过程和推理阶段中的去高阶分解(deHRF)过程。在前一个过程中,我们在轻量级Transformer的前馈网络(FFN)之前插入一个额外的线性层。据推测,插入的HRF层可以增强模型的学习能力。在后面的过程中,辅助HRF层将与随后的FFN层合并到一个线性层中,从而恢复轻量级模型的原始结构。为了检验所提出的方法的有效性,我们对三种广泛使用的Transformer变体进行了评估,即,ConvTransformer、Conformer和SpeechFormer网络在IEMOCAP、M3 ED和DAIC-WOZ数据集上的语音情感识别应用。实验结果表明,我们提出的方法一致地提高了轻量级Transformers的性能,甚至使它们与大型模型相当。所提出的重新参数化方法使高级Transformer模型能够部署在资源受限的IoT设备上。
摘要:With the increasing implementation of machine learning models on edge orInternet-of-Things (IoT) devices, deploying advanced models onresource-constrained IoT devices remains challenging. Transformer models, acurrently dominant neural architecture, have achieved great success in broaddomains but their complexity hinders its deployment on IoT devices with limitedcomputation capability and storage size. Although many model compressionapproaches have been explored, they often suffer from notorious performancedegradation. To address this issue, we introduce a new method, namelyTransformer Re-parameterization, to boost the performance of lightweightTransformer models. It consists of two processes: the High-Rank Factorization(HRF) process in the training stage and the deHigh-Rank Factorization (deHRF)process in the inference stage. In the former process, we insert an additionallinear layer before the Feed-Forward Network (FFN) of the lightweightTransformer. It is supposed that the inserted HRF layers can enhance the modellearning capability. In the later process, the auxiliary HRF layer will bemerged together with the following FFN layer into one linear layer and thusrecover the original structure of the lightweight model. To examine theeffectiveness of the proposed method, we evaluate it on three widely usedTransformer variants, i.e., ConvTransformer, Conformer, and SpeechFormernetworks, in the application of speech emotion recognition on the IEMOCAP, M3EDand DAIC-WOZ datasets. Experimental results show that our proposed methodconsistently improves the performance of lightweight Transformers, even makingthem comparable to large models. The proposed re-parameterization approachenables advanced Transformer models to be deployed on resource-constrained IoTdevices.

【4】 EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble  Diffusion Models
标题:基于脑电波的语音解码:一种使用多核集合扩散模型的新方法
链接:https://arxiv.org/abs/2411.09302
作者:Soowon Kim,  Ha-Na Jo,  Eunyeong Ko
摘要:在这项研究中,我们提出了一个基于脑电图的公开语音分类的集成学习框架,利用不同卷积核大小的去噪扩散概率模型。该集成包括三个模型,内核大小为51,101和201,有效地捕捉信号中固有的多尺度时间特征。这种方法通过适应神经信号丰富的时间复杂性来提高语音解码的鲁棒性和准确性。集成模型与条件自动编码器结合使用,条件自动编码器可以优化重构信号并最大化下游分类任务的有用信息。结果表明,所提出的集成为基础的方法显着优于个人的模型和现有的国家的最先进的技术。这些发现表明了集成方法在推进大脑信号解码方面的潜力,为非语言通信应用提供了新的可能性,特别是在旨在帮助有语言障碍的个人的脑机接口系统中。
摘要:In this study, we propose an ensemble learning framework forelectroencephalogram-based overt speech classification, leveraging denoisingdiffusion probabilistic models with varying convolutional kernel sizes. Theensemble comprises three models with kernel sizes of 51, 101, and 201,effectively capturing multi-scale temporal features inherent in signals. Thisapproach improves the robustness and accuracy of speech decoding byaccommodating the rich temporal complexity of neural signals. The ensemblemodels work in conjunction with conditional autoencoders that refine thereconstructed signals and maximize the useful information for downstreamclassification tasks. The results indicate that the proposed ensemble-basedapproach significantly outperforms individual models and existingstate-of-the-art techniques. These findings demonstrate the potential ofensemble methods in advancing brain signal decoding, offering new possibilitiesfor non-verbal communication applications, particularly in brain-computerinterface systems aimed at aiding individuals with speech impairments.

【5】 Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech  from EEG Signals
标题:从脑电信号中感知、口语和想象语音的统一神经解码
链接:https://arxiv.org/abs/2411.09243
作者:Jung-Sun Lee,  Ha-Na Jo,  Seo-Hyun Lee
摘要:大脑信号伴随着与人类行为和心理意象相关的各种信息,这使得它们对于解释和理解人类意图至关重要。脑机接口技术利用这种大脑活动来生成控制环境的外部命令,为瘫痪或闭锁综合征患者提供关键优势。在脑-机接口领域,脑-语音研究已经引起了人们的关注,重点是从大脑信号直接合成可听语音。目前大多数研究使用侵入性技术从大脑活动中解码语音,并强调口语语音数据。然而,人类表达各种语音状态,通过非侵入性方法区分这些状态仍然是一项重要而具有挑战性的任务。这项研究调查了深度学习模型在基于非侵入性的神经信号解码中的有效性,重点是区分不同的语音范式,包括多个频带上的感知、公开、耳语和想象语音。与其他模型相比,利用空间常规神经网络模块的模型表现出优越的性能,特别是在伽马波段。此外,theta频段的想象语音,深度学习也表现出很强的效果,与其他语音范式相比,表现出统计学上的显著差异。
摘要:Brain signals accompany various information relevant to human actions andmental imagery, making them crucial to interpreting and understanding humanintentions. Brain-computer interface technology leverages this brain activityto generate external commands for controlling the environment, offeringcritical advantages to individuals with paralysis or locked-in syndrome. Withinthe brain-computer interface domain, brain-to-speech research has gainedattention, focusing on the direct synthesis of audible speech from brainsignals. Most current studies decode speech from brain activity using invasivetechniques and emphasize spoken speech data. However, humans express variousspeech states, and distinguishing these states through non-invasive approachesremains a significant yet challenging task. This research investigated theeffectiveness of deep learning models for non-invasive-based neural signaldecoding, with an emphasis on distinguishing between different speechparadigms, including perceived, overt, whispered, and imagined speech, acrossmultiple frequency bands. The model utilizing the spatial conventional neuralnetwork module demonstrated superior performance compared to other models,especially in the gamma band. Additionally, imagined speech in the thetafrequency band, where deep learning also showed strong effects, exhibitedstatistically significant differences compared to the other speech paradigms.

【6】 Improvement and Implementation of a Speech Emotion Recognition Model  Based on Dual-Layer LSTM
标题:基于双层LSTM的语音情感识别模型的改进与实现
链接:https://arxiv.org/abs/2411.09189
作者:Xiaoran Yang,  Shuhan Yu,  Wenxi Xu
摘要:本文在现有语音情感识别模型的基础上,通过增加一个额外的LSTM层来提高从音频数据中识别情感的准确性和处理效率。通过双层LSTM网络捕获音频序列内的长期依赖关系,该模型可以更准确地识别和分类复杂的情感模式。在RAVDESS数据集上进行的实验验证了这种方法,表明与单层LSTM相比,修改后的双层LSTM模型将准确率提高了2%,同时显着降低了识别延迟,从而提高了实时性能。这些结果表明,双层LSTM架构非常适合处理具有长期依赖性的情感特征,为语音情感识别系统提供了可行的优化。该研究为智能客户服务、情感分析和人机交互等领域的实际应用提供了参考。
摘要:This paper builds upon an existing speech emotion recognition model by addingan additional LSTM layer to improve the accuracy and processing efficiency ofemotion recognition from audio data. By capturing the long-term dependencieswithin audio sequences through a dual-layer LSTM network, the model canrecognize and classify complex emotional patterns more accurately. Experimentsconducted on the RAVDESS dataset validated this approach, showing that themodified dual layer LSTM model improves accuracy by 2% compared to thesingle-layer LSTM while significantly reducing recognition latency, therebyenhancing real-time performance. These results indicate that the dual-layerLSTM architecture is highly suitable for handling emotional features withlong-term dependencies, providing a viable optimization for speech emotionrecognition systems. This research provides a reference for practicalapplications in fields like intelligent customer service, sentiment analysisand human-computer interaction.

【7】 Robust AI-Synthesized Speech Detection Using Feature Decomposition  Learning and Synthesizer Feature Augmentation
标题:使用特征分解学习和合成器特征增强的稳健AI合成语音检测
链接:https://arxiv.org/abs/2411.09167
作者:Kuiyuan Zhang,  Zhongyun Hua,  Yushu Zhang,  Yifang Guo,  Tao Xiang
摘要:人工智能合成语音,也称为deepfake语音,最近由于语音合成和语音转换技术的快速发展而引起了人们的极大关注。以前的工作通常依赖于区分合成器伪影来识别deepfake语音。然而,过度依赖于这些特定的合成器伪像可能导致在寻址由看不见的合成器创建的语音信号时不令人满意的性能。在本文中,我们提出了一种鲁棒的deepfake语音检测方法,该方法采用特征分解来学习独立于合成器的内容特征作为检测的补充。具体来说,我们提出了一个双流特征分解学习策略,使用合成器流和内容流分解学习的语音表示。合成器流专门通过合成器标签的监督训练来学习合成器功能。同时,内容流专注于学习独立于合成器的内容特征,通过基于伪标签的监督学习方法实现。该方法随机变换语音以生成用于训练的速度和压缩标签。此外,我们采用对抗性学习技术来减少内容流中与合成器相关的组件。最终的分类是通过连接合成器和内容特征来确定的。为了增强模型对不同合成器特征的鲁棒性,我们进一步提出了一种合成器特征增强策略,该策略随机混合真实和虚假音频特征内的特征风格,并随机将合成器特征与内容特征混洗。该策略有效地增强了特征的多样性,模拟了更多的特征组合。
摘要:AI-synthesized speech, also known as deepfake speech, has recently raisedsignificant concerns due to the rapid advancement of speech synthesis andspeech conversion techniques. Previous works often rely on distinguishingsynthesizer artifacts to identify deepfake speech. However, excessive relianceon these specific synthesizer artifacts may result in unsatisfactoryperformance when addressing speech signals created by unseen synthesizers. Inthis paper, we propose a robust deepfake speech detection method that employsfeature decomposition to learn synthesizer-independent content features ascomplementary for detection. Specifically, we propose a dual-stream featuredecomposition learning strategy that decomposes the learned speechrepresentation using a synthesizer stream and a content stream. The synthesizerstream specializes in learning synthesizer features through supervised trainingwith synthesizer labels. Meanwhile, the content stream focuses on learningsynthesizer-independent content features, enabled by a pseudo-labeling-basedsupervised learning method. This method randomly transforms speech to generatespeed and compression labels for training. Additionally, we employ anadversarial learning technique to reduce the synthesizer-related components inthe content stream. The final classification is determined by concatenating thesynthesizer and content features. To enhance the model's robustness todifferent synthesizer characteristics, we further propose a synthesizer featureaugmentation strategy that randomly blends the characteristic styles withinreal and fake audio features and randomly shuffles the synthesizer featureswith the content features. This strategy effectively enhances the featurediversity and simulates more feature combinations.

【8】 Language Models for Music Medicine Generation
标题:音乐医学一代的语言模型
链接:https://arxiv.org/abs/2411.09080
作者:Emmanouil Nikolakakis,  Joann Ching,  Emmanouil Karystinaios,  Gabrielle Sipin,  Gerhard Widmer,  Razvan Marinescu
备注:Late-Breaking / Demo Session Extended Abstract, ISMIR 2024 Conference
摘要:近年来,音乐疗法已被证明可以提供与情绪健康相关的多种健康益处。反过来,保持健康的情绪状态已被证明对接受治疗的患者有效,例如帕金森病患者或患有压力和焦虑的患者。我们建议微调MusicGen,一个音乐生成Transformer模型,以创建简短的音乐片段,帮助患者从消极的情绪状态过渡到期望的情绪状态。使用低秩分解微调的MTG-Jamendo数据集与情感标签,我们生成30秒的剪辑,坚持iso原则,引导患者通过中间状态的效价唤醒环。使用音乐情感识别模型来评估所生成的音乐,以确保与预期的情感对齐。通过连接这些片段,我们制作了一个15分钟的“音乐医学”,类似于音乐治疗。我们的方法是第一个利用语言模型生成音乐医学的模型。最终,输出的目的是作为一个临时救济之间的音乐治疗会议与董事会认证的治疗师。
摘要:Music therapy has been shown in recent years to provide multiple healthbenefits related to emotional wellness. In turn, maintaining a healthyemotional state has proven to be effective for patients undergoing treatment,such as Parkinson's patients or patients suffering from stress and anxiety. Wepropose fine-tuning MusicGen, a music-generating transformer model, to createshort musical clips that assist patients in transitioning from negative todesired emotional states. Using low-rank decomposition fine-tuning on theMTG-Jamendo Dataset with emotion tags, we generate 30-second clips that adhereto the iso principle, guiding patients through intermediate states in thevalence-arousal circumplex. The generated music is evaluated using a musicemotion recognition model to ensure alignment with intended emotions. Byconcatenating these clips, we produce a 15-minute "music medicine" resembling amusic therapy session. Our approach is the first model to leverage LanguageModels to generate music medicine. Ultimately, the output is intended to beused as a temporary relief between music therapy sessions with aboard-certified therapist.

【9】 Multilingual Standalone Trustworthy Voice-Based Social Network for  Disaster Situations
标题:针对灾难情况的多语言独立可信的基于语音的社交网络
链接:https://arxiv.org/abs/2411.08889
作者:Majid Behravan,  Elham Mohammadrezaei,  Mohamed Azab,  Denis Gracanin
备注:Accepted for publication in IEEE UEMCON 2024, to appear in December 2024. 7 pages, 3 figures
摘要:在发生灾害的情况下,有效的沟通至关重要,但语言障碍往往阻碍及时和准确的信息传播,加剧了脆弱性,使应对工作复杂化。本文提出了一种新颖的,多语言的,基于语音的社交网络,专门设计来解决这些挑战。该系统将先进的人工智能(AI)与区块链技术相结合,以实现跨多种语言的安全、异步语音通信。该应用程序独立于外部服务器运行,通过本地网络离线运行,即使在受损环境中也能确保可靠性。主要功能包括人工智能驱动的语音消息实时翻译,确保无缝的跨语言通信,以及支持区块链的存储,用于所有交互的安全,不可变的记录,保护消息的完整性。该系统专为跨平台使用而设计,可在从移动电话到台式机的各种设备上提供一致的性能,使其在各种灾难情况下具有高度适应性。评估指标表明,语音识别和翻译的准确性高,延迟低,用户满意度,验证了系统在危机期间加强沟通的有效性。该解决方案代表了灾难通信领域的重大进步,弥合了语言差距,以支持更具包容性和更高效的应急响应。
摘要:In disaster scenarios, effective communication is crucial, yet languagebarriers often hinder timely and accurate information dissemination,exacerbating vulnerabilities and complicating response efforts. This paperpresents a novel, multilingual, voice-based social network specificallydesigned to address these challenges. The proposed system integrates advancedartificial intelligence (AI) with blockchain technology to enable secure,asynchronous voice communication across multiple languages. The applicationoperates independently of external servers, ensuring reliability even incompromised environments by functioning offline through local networks. Keyfeatures include AI-driven real-time translation of voice messages, ensuringseamless cross-linguistic communication, and blockchain-enabled storage forsecure, immutable records of all interactions, safeguarding message integrity.Designed for cross-platform use, the system offers consistent performanceacross devices, from mobile phones to desktops, making it highly adaptable indiverse disaster situations. Evaluation metrics demonstrate high accuracy inspeech recognition and translation, low latency, and user satisfaction,validating the system's effectiveness in enhancing communication during crises.This solution represents a significant advancement in disaster communication,bridging language gaps to support more inclusive and efficient emergencyresponse.

【10】 Enhancing Lie Detection Accuracy: A Comparative Study of Classic ML,  CNN, and GCN Models using Audio-Visual Features
标题:提高谎言检测准确性:使用视听特征的经典ML、CNN和GCN模型的比较研究
链接:https://arxiv.org/abs/2411.08885
作者:Abdelrahman Abdelwahab,  Abdelrahman Abdelwahab,  Ayaan Vaswani,  Advait Bharathulwar,  Arnav Kommaraju
备注:11 pages, 18 figures
摘要:测谎仪测试的不准确性往往会导致错误的定罪,错误的信息和偏见,所有这些都对法律和政治体系产生重大影响。最近,分析面部微表情已成为检测欺骗的方法,然而,目前的模型还没有达到高精度和泛化能力。本研究的目的是帮助解决这些问题。本研究中使用的独特的多模态Transformer架构通过使用听觉输入、视觉面部微表情和手动转录的手势注释来改进先前的方法,从而更接近可靠的非侵入性测谎模型。分别使用Vision Transformer和OpenSmile模型提取视觉和听觉特征,然后将其与参与者微表情和手势的transmittance连接起来。使用这些经过处理和连接的特征,训练了各种模型来分类谎言和真相。CNN Conv1D多模态模型的平均准确率为95.4%。然而,仍然需要进一步的研究来创建更高质量的数据集,甚至更广泛的模型,用于更多样化的应用。
摘要:Inaccuracies in polygraph tests often lead to wrongful convictions, falseinformation, and bias, all of which have significant consequences for bothlegal and political systems. Recently, analyzing facial micro-expressions hasemerged as a method for detecting deception; however, current models have notreached high accuracy and generalizability. The purpose of this study is to aidin remedying these problems. The unique multimodal transformer architectureused in this study improves upon previous approaches by using auditory inputs,visual facial micro-expressions, and manually transcribed gesture annotations,moving closer to a reliable non-invasive lie detection model. Visual andauditory features were extracted using the Vision Transformer and OpenSmilemodels respectively, which were then concatenated with the transcriptions ofparticipants micro-expressions and gestures. Various models were trained forthe classification of lies and truths using these processed and concatenatedfeatures. The CNN Conv1D multimodal model achieved an average accuracy of95.4%. However, further research is still required to create higher-qualitydatasets and even more generalized models for more diverse applications.

eess.AS音频处理

【1】 An End-To-End Stuttering Detection Method Based On Conformer And BILSTM
标题:基于Conformer和BILSTM的端到端口吃检测方法
链接:https://arxiv.org/abs/2411.09479
作者:Xiaokang Liu,  Changqing Xu,  Yudong Yang,  Lan Wang,  Nan Yan
备注:7 pages, 3 figures, 7 tables
摘要:口吃是一种神经发育性言语障碍,其特征是常见的言语症状,如停顿、感叹、重复和延长。语言病理学家通常通过观察这些症状来评估口吃的类型和严重程度。存在许多有效的端到端方法用于口吃检测,但通常被忽视的挑战是此过程中涉及的任务之间的不确定关系。采用合适的多任务策略可以提高口吃检测的性能。本文提出了一种新的口吃事件检测模型,旨在帮助语音语言病理学家评估口吃的类型和严重程度。首先,Conformer模型从口吃语音中提取声学特征,然后通过长短期记忆(LSTM)网络来捕获上下文信息。最后,我们探讨了口吃的多任务学习,并提出了一个有效的多任务策略。实验结果表明,我们的模型优于目前国家的最先进的口吃检测方法。在基于AS-70数据集[1]的2024年口吃语音挑战赛中,与基线方法相比,我们的模型将平均F1分数提高了24.8%,并获得了第一名。在此基础上,我们分别对LSTM和多任务学习策略进行了相关的广泛实验。结果表明,与基线方法相比,我们提出的方法将平均F1分数提高了39.8%。
摘要:Stuttering is a neurodevelopmental speech disorder characterized by commonspeech symptoms such as pauses, exclamations, repetition, and prolongation.Speech-language pathologists typically assess the type and severity ofstuttering by observing these symptoms. Many effective end-to-end methods existfor stuttering detection, but a commonly overlooked challenge is the uncertainrelationship between tasks involved in this process. Using a suitablemulti-task strategy could improve stuttering detection performance. This paperpresents a novel stuttering event detection model designed to helpspeech-language pathologists assess both the type and severity of stuttering.First, the Conformer model extracts acoustic features from stuttered speech,followed by a Long Short-Term Memory (LSTM) network to capture contextualinformation. Finally, we explore multi-task learning for stuttering and proposean effective multi-task strategy. Experimental results show that our modeloutperforms current state-of-the-art methods for stuttering detection. In theSLT 2024 Stuttering Speech Challenge based on the AS-70 dataset [1], our modelimproved the mean F1 score by 24.8% compared to the baseline method andachieved first place. On this basis, we conducted relevant extensiveexperiments on LSTM and multi-task learning strategies respectively. Theresults show that our proposed method improved the mean F1 score by 39.8%compared to the baseline method.

【2】 Transferable Adversarial Attacks against ASR
标题:针对ASC的可转移对抗攻击
链接:https://arxiv.org/abs/2411.09220
作者:Xiaoxue Gao,  Zexin Li,  Yiming Chen,  Cong Liu,  Haizhou Li
备注:IEEE SPL
摘要:鉴于自动语音识别(ASR)的广泛研究和实际应用,确保ASR模型对微小输入扰动的鲁棒性成为在实时场景中保持其有效性的关键考虑因素。以前对ASR模型鲁棒性的探索主要围绕着在完全访问ASR模型的情况下评估白盒设置的准确性。然而,完整的ASR模型细节在现实世界的应用中通常不可用。因此,评估黑盒ASR模型的鲁棒性对于全面理解ASR模型的弹性至关重要。在这方面,我们深入研究了尖端ASR模型中实际黑盒攻击的脆弱性,并提出采用两种先进的基于时域的可转移攻击以及我们的可区分特征提取器。我们还提出了一种语音感知的梯度优化方法(SAGO)的ASR,它迫使误译与最小的影响,通过语音活动检测规则和语音感知的梯度为导向的优化器对人类的不可感知性。我们全面的实验结果显示,性能增强的基线方法相比,在两个数据库上的五个模型。
摘要:Given the extensive research and real-world applications of automatic speechrecognition (ASR), ensuring the robustness of ASR models against minor inputperturbations becomes a crucial consideration for maintaining theireffectiveness in real-time scenarios. Previous explorations into ASR modelrobustness have predominantly revolved around evaluating accuracy on white-boxsettings with full access to ASR models. Nevertheless, full ASR model detailsare often not available in real-world applications. Therefore, evaluating therobustness of black-box ASR models is essential for a comprehensiveunderstanding of ASR model resilience. In this regard, we thoroughly study thevulnerability of practical black-box attacks in cutting-edge ASR models andpropose to employ two advanced time-domain-based transferable attacks alongsideour differentiable feature extractor. We also propose a speech-aware gradientoptimization approach (SAGO) for ASR, which forces mistranscription withminimal impact on human imperceptibility through voice activity detection ruleand a speech-aware gradient-oriented optimizer. Our comprehensive experimentalresults reveal performance enhancements compared to baseline approaches acrossfive models on two databases.

【3】 Local deployment of large-scale music AI models on commodity hardware
标题:在商品硬件上本地部署大型音乐AI模型
链接:https://arxiv.org/abs/2411.09625
作者:Xun Zhou,  Charlie Ruan,  Zihe Zhao,  Tianqi Chen,  Chris Donahue
备注:2 pages
摘要:我们提出了MIDInfinite,这是一个能够在商品硬件上本地使用大规模生成AI模型生成符号音乐的Web应用程序。创建此演示涉及将Anticipatory Music Transformer(一个在Lakh数据集上预训练的大型语言模型(LLM))移植到机器学习编译(MLC)框架。一旦模型被移植,MLC就可以在各种运行时(包括C++、移动设备和浏览器)上进行推理。我们设想MLC有潜力弥合日益强大的音乐AI模型与音乐软件开发人员更熟悉的技术之间的差距。作为概念验证,我们构建了一个网络应用程序,允许用户在浏览器中从头开始或根据提示生成源源不断的多乐器MIDI。在商用硬件(M3 Macbook Pro)上,我们的演示可以每秒生成51个音符,这比实时播放快72.9%的代,并增加到86.3%,有2秒的前期缓冲。
摘要:We present the MIDInfinite, a web application capable of generating symbolicmusic using a large-scale generative AI model locally on commodity hardware.Creating this demo involved porting the Anticipatory Music Transformer, a largelanguage model (LLM) pre-trained on the Lakh MIDI dataset, to the MachineLearning Compilation (MLC) framework. Once the model is ported, MLC facilitatesinference on a variety of runtimes including C++, mobile, and the browser. Weenvision that MLC has the potential to bridge the gap between the landscape ofincreasingly capable music AI models and technology more familiar to musicsoftware developers. As a proof of concept, we build a web application thatallows users to generate endless streams of multi-instrumental MIDI in thebrowser, either from scratch or conditioned on a prompt. On commodity hardware(an M3 Macbook Pro), our demo can generate 51 notes per second, which is fasterthan real-time playback for 72.9% of generations, and increases to 86.3% with 2seconds of upfront buffering.

【4】 ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics  over Acoustic Foundation Models
标题:ParaLBench:计算副语言学优于声学基础模型的大规模基准
链接:https://arxiv.org/abs/2411.09349
作者:Zixing Zhang,  Weixiang Xu,  Zhongren Dong,  Kanglin Wang,  Yimeng Wu,  Jing Peng,  Runming Wang,  Dong-Yan Huang
摘要:计算语言学(Comparal)旨在开发算法和模型来自动检测,分析和解释来自语音交流的非语言信息,例如。例如,在一个实施例中,情绪、健康状况、年龄和性别。尽管其发展迅速,但它严重依赖于特定语言任务的精心设计的模型。因此,CompParal模型的异质性和多样性在很大程度上阻碍了CompParal模型的实际实现。近年来,随着自监督学习的声学基础模型的出现,开发能够有效感知大量非语言信息的通用模型已经成为语音处理中的一个活跃话题。然而,它缺乏一个统一的评价框架,无法进行公平和一致的业绩比较。为了弥合这一差距,我们进行了一个大规模的基准测试,即ParaLBench,它集中于标准化的评估过程中的各种语言任务,包括情感计算的关键方面,如情感识别和情感维度预测,在不同的声学基础模型。该基准测试包含10个数据集,包含13个不同的语言任务,涵盖短期,中期和长期特征。每项任务都在统一的评估框架下对14个声学基础模型进行评估,从而可以进行公正的方法比较,并为Comparal社区提供有依据的参考。根据ParaLBench的见解,我们还指出了潜在的研究方向,即,跨语料库的泛化能力,以推动未来的Comparal研究。与本研究相关的代码将可用于促进后续研究人员这项工作的透明度和可复制性。
摘要:Computational paralinguistics (ComParal) aims to develop algorithms andmodels to automatically detect, analyze, and interpret non-verbal informationfrom speech communication, e. g., emotion, health state, age, and gender.Despite its rapid progress, it heavily depends on sophisticatedly designedmodels given specific paralinguistic tasks. Thus, the heterogeneity anddiversity of ComParal models largely prevent the realistic implementation ofComParal models. Recently, with the advent of acoustic foundation modelsbecause of self-supervised learning, developing more generic models that canefficiently perceive a plethora of paralinguistic information has become anactive topic in speech processing. However, it lacks a unified evaluationframework for a fair and consistent performance comparison. To bridge this gap,we conduct a large-scale benchmark, namely ParaLBench, which concentrates onstandardizing the evaluation process of diverse paralinguistic tasks, includingcritical aspects of affective computing such as emotion recognition and emotiondimensions prediction, over different acoustic foundation models. Thisbenchmark contains ten datasets with thirteen distinct paralinguistic tasks,covering short-, medium- and long-term characteristics. Each task is carriedout on 14 acoustic foundation models under a unified evaluation framework,which allows for an unbiased methodological comparison and offers a groundedreference for the ComParal community. Based on the insights gained fromParaLBench, we also point out potential research directions, i.e., thecross-corpus generalizability, to propel ComParal research in the future. Thecode associated with this study will be available to foster the transparencyand replicability of this work for succeeding researchers.

【5】 Re-Parameterization of Lightweight Transformer for On-Device Speech  Emotion Recognition
标题:用于设备上语音情感识别的轻型Transformer重新参数化
链接:https://arxiv.org/abs/2411.09339
作者:Zixing Zhang,  Zhongren Dong,  Weixiang Xu,  Jing Han
摘要:随着机器学习模型在边缘或物联网(IoT)设备上的实施越来越多,在资源受限的IoT设备上部署高级模型仍然具有挑战性。Transformer模型是目前占主导地位的神经架构,在广泛的领域取得了巨大的成功,但其复杂性阻碍了其在计算能力和存储大小有限的物联网设备上的部署。虽然已经探索了许多模型压缩方法,但它们经常遭受臭名昭著的性能下降。为了解决这个问题,我们引入了一种新的方法,即Transformer重新参数化,以提高轻量级Transformer模型的性能。它包括两个过程:训练阶段中的高阶分解(HRF)过程和推理阶段中的去高阶分解(deHRF)过程。在前一个过程中,我们在轻量级Transformer的前馈网络(FFN)之前插入一个额外的线性层。据推测,插入的HRF层可以增强模型的学习能力。在后面的过程中,辅助HRF层将与随后的FFN层合并到一个线性层中,从而恢复轻量级模型的原始结构。为了检验所提出的方法的有效性,我们对三种广泛使用的Transformer变体进行了评估,即,ConvTransformer、Conformer和SpeechFormer网络在IEMOCAP、M3 ED和DAIC-WOZ数据集上的语音情感识别应用。实验结果表明,我们提出的方法一致地提高了轻量级Transformers的性能,甚至使它们与大型模型相当。所提出的重新参数化方法使高级Transformer模型能够部署在资源受限的IoT设备上。
摘要:With the increasing implementation of machine learning models on edge orInternet-of-Things (IoT) devices, deploying advanced models onresource-constrained IoT devices remains challenging. Transformer models, acurrently dominant neural architecture, have achieved great success in broaddomains but their complexity hinders its deployment on IoT devices with limitedcomputation capability and storage size. Although many model compressionapproaches have been explored, they often suffer from notorious performancedegradation. To address this issue, we introduce a new method, namelyTransformer Re-parameterization, to boost the performance of lightweightTransformer models. It consists of two processes: the High-Rank Factorization(HRF) process in the training stage and the deHigh-Rank Factorization (deHRF)process in the inference stage. In the former process, we insert an additionallinear layer before the Feed-Forward Network (FFN) of the lightweightTransformer. It is supposed that the inserted HRF layers can enhance the modellearning capability. In the later process, the auxiliary HRF layer will bemerged together with the following FFN layer into one linear layer and thusrecover the original structure of the lightweight model. To examine theeffectiveness of the proposed method, we evaluate it on three widely usedTransformer variants, i.e., ConvTransformer, Conformer, and SpeechFormernetworks, in the application of speech emotion recognition on the IEMOCAP, M3EDand DAIC-WOZ datasets. Experimental results show that our proposed methodconsistently improves the performance of lightweight Transformers, even makingthem comparable to large models. The proposed re-parameterization approachenables advanced Transformer models to be deployed on resource-constrained IoTdevices.

【6】 EEG-Based Speech Decoding: A Novel Approach Using Multi-Kernel Ensemble  Diffusion Models
标题:基于脑电波的语音解码:一种使用多核集合扩散模型的新方法
链接:https://arxiv.org/abs/2411.09302
作者:Soowon Kim,  Ha-Na Jo,  Eunyeong Ko
摘要:在这项研究中,我们提出了一个基于脑电图的公开语音分类的集成学习框架,利用不同卷积核大小的去噪扩散概率模型。该集成包括三个模型,内核大小为51,101和201,有效地捕捉信号中固有的多尺度时间特征。这种方法通过适应神经信号丰富的时间复杂性来提高语音解码的鲁棒性和准确性。集成模型与条件自动编码器结合使用,条件自动编码器可以优化重构信号并最大化下游分类任务的有用信息。结果表明,所提出的集成为基础的方法显着优于个人的模型和现有的国家的最先进的技术。这些发现表明了集成方法在推进大脑信号解码方面的潜力,为非语言通信应用提供了新的可能性,特别是在旨在帮助有语言障碍的个人的脑机接口系统中。
摘要:In this study, we propose an ensemble learning framework forelectroencephalogram-based overt speech classification, leveraging denoisingdiffusion probabilistic models with varying convolutional kernel sizes. Theensemble comprises three models with kernel sizes of 51, 101, and 201,effectively capturing multi-scale temporal features inherent in signals. Thisapproach improves the robustness and accuracy of speech decoding byaccommodating the rich temporal complexity of neural signals. The ensemblemodels work in conjunction with conditional autoencoders that refine thereconstructed signals and maximize the useful information for downstreamclassification tasks. The results indicate that the proposed ensemble-basedapproach significantly outperforms individual models and existingstate-of-the-art techniques. These findings demonstrate the potential ofensemble methods in advancing brain signal decoding, offering new possibilitiesfor non-verbal communication applications, particularly in brain-computerinterface systems aimed at aiding individuals with speech impairments.

【7】 Towards Unified Neural Decoding of Perceived, Spoken and Imagined Speech  from EEG Signals
标题:从脑电信号中感知、口语和想象语音的统一神经解码
链接:https://arxiv.org/abs/2411.09243
作者:Jung-Sun Lee,  Ha-Na Jo,  Seo-Hyun Lee
摘要:大脑信号伴随着与人类行为和心理意象相关的各种信息,这使得它们对于解释和理解人类意图至关重要。脑机接口技术利用这种大脑活动来生成控制环境的外部命令,为瘫痪或闭锁综合征患者提供关键优势。在脑-机接口领域,脑-语音研究已经引起了人们的关注,重点是从大脑信号直接合成可听语音。目前大多数研究使用侵入性技术从大脑活动中解码语音,并强调口语语音数据。然而,人类表达各种语音状态,通过非侵入性方法区分这些状态仍然是一项重要而具有挑战性的任务。这项研究调查了深度学习模型在基于非侵入性的神经信号解码中的有效性,重点是区分不同的语音范式,包括多个频带上的感知、公开、耳语和想象语音。与其他模型相比,利用空间常规神经网络模块的模型表现出优越的性能,特别是在伽马波段。此外,深度学习也表现出强烈效果的theta频带中的想象语音与其他语音范式相比表现出统计上显着的差异。
摘要:Brain signals accompany various information relevant to human actions andmental imagery, making them crucial to interpreting and understanding humanintentions. Brain-computer interface technology leverages this brain activityto generate external commands for controlling the environment, offeringcritical advantages to individuals with paralysis or locked-in syndrome. Withinthe brain-computer interface domain, brain-to-speech research has gainedattention, focusing on the direct synthesis of audible speech from brainsignals. Most current studies decode speech from brain activity using invasivetechniques and emphasize spoken speech data. However, humans express variousspeech states, and distinguishing these states through non-invasive approachesremains a significant yet challenging task. This research investigated theeffectiveness of deep learning models for non-invasive-based neural signaldecoding, with an emphasis on distinguishing between different speechparadigms, including perceived, overt, whispered, and imagined speech, acrossmultiple frequency bands. The model utilizing the spatial conventional neuralnetwork module demonstrated superior performance compared to other models,especially in the gamma band. Additionally, imagined speech in the thetafrequency band, where deep learning also showed strong effects, exhibitedstatistically significant differences compared to the other speech paradigms.

【8】 Improvement and Implementation of a Speech Emotion Recognition Model  Based on Dual-Layer LSTM
标题:基于双层LSTM的语音情感识别模型的改进与实现
链接:https://arxiv.org/abs/2411.09189
作者:Xiaoran Yang,  Shuhan Yu,  Wenxi Xu
摘要:本文在现有语音情感识别模型的基础上,通过增加一个额外的LSTM层来提高从音频数据中识别情感的准确性和处理效率。通过双层LSTM网络捕获音频序列中的长期依赖关系,该模型可以更准确地识别和分类复杂的情感模式。在RAVDESS数据集上进行的实验验证了这种方法,表明与单层LSTM相比,修改后的双层LSTM模型将准确率提高了2%,同时显着降低了识别延迟,从而提高了实时性能。这些结果表明,双层LSTM架构非常适合处理具有长期依赖性的情感特征,为语音情感识别系统提供了可行的优化。该研究为智能客户服务、情感分析和人机交互等领域的实际应用提供了参考。
摘要:This paper builds upon an existing speech emotion recognition model by addingan additional LSTM layer to improve the accuracy and processing efficiency ofemotion recognition from audio data. By capturing the long-term dependencieswithin audio sequences through a dual-layer LSTM network, the model canrecognize and classify complex emotional patterns more accurately. Experimentsconducted on the RAVDESS dataset validated this approach, showing that themodified dual layer LSTM model improves accuracy by 2% compared to thesingle-layer LSTM while significantly reducing recognition latency, therebyenhancing real-time performance. These results indicate that the dual-layerLSTM architecture is highly suitable for handling emotional features withlong-term dependencies, providing a viable optimization for speech emotionrecognition systems. This research provides a reference for practicalapplications in fields like intelligent customer service, sentiment analysisand human-computer interaction.

【9】 Robust AI-Synthesized Speech Detection Using Feature Decomposition  Learning and Synthesizer Feature Augmentation
标题:使用特征分解学习和合成器特征增强的稳健AI合成语音检测
链接:https://arxiv.org/abs/2411.09167
作者:Kuiyuan Zhang,  Zhongyun Hua,  Yushu Zhang,  Yifang Guo,  Tao Xiang
摘要:人工智能合成语音,也称为deepfake语音,最近由于语音合成和语音转换技术的快速发展而引起了人们的极大关注。以前的工作通常依赖于区分合成器伪影来识别deepfake语音。然而,过度依赖于这些特定的合成器伪像可能导致在寻址由看不见的合成器创建的语音信号时不令人满意的性能。在本文中,我们提出了一种鲁棒的deepfake语音检测方法,该方法采用特征分解来学习独立于合成器的内容特征作为检测的补充。具体来说,我们提出了一个双流特征分解学习策略,使用合成器流和内容流分解学习的语音表示。合成器流专门通过合成器标签的监督训练来学习合成器功能。同时,内容流专注于学习独立于合成器的内容特征,通过基于伪标签的监督学习方法实现。该方法随机变换语音以生成用于训练的速度和压缩标签。此外,我们采用对抗性学习技术来减少内容流中与合成器相关的组件。最终的分类是通过连接合成器和内容特征来确定的。为了增强模型对不同合成器特征的鲁棒性,我们进一步提出了一种合成器特征增强策略,该策略随机混合真实和虚假音频特征内的特征风格,并随机将合成器特征与内容特征混洗。该策略有效地增强了特征的多样性,模拟了更多的特征组合。
摘要:AI-synthesized speech, also known as deepfake speech, has recently raisedsignificant concerns due to the rapid advancement of speech synthesis andspeech conversion techniques. Previous works often rely on distinguishingsynthesizer artifacts to identify deepfake speech. However, excessive relianceon these specific synthesizer artifacts may result in unsatisfactoryperformance when addressing speech signals created by unseen synthesizers. Inthis paper, we propose a robust deepfake speech detection method that employsfeature decomposition to learn synthesizer-independent content features ascomplementary for detection. Specifically, we propose a dual-stream featuredecomposition learning strategy that decomposes the learned speechrepresentation using a synthesizer stream and a content stream. The synthesizerstream specializes in learning synthesizer features through supervised trainingwith synthesizer labels. Meanwhile, the content stream focuses on learningsynthesizer-independent content features, enabled by a pseudo-labeling-basedsupervised learning method. This method randomly transforms speech to generatespeed and compression labels for training. Additionally, we employ anadversarial learning technique to reduce the synthesizer-related components inthe content stream. The final classification is determined by concatenating thesynthesizer and content features. To enhance the model's robustness todifferent synthesizer characteristics, we further propose a synthesizer featureaugmentation strategy that randomly blends the characteristic styles withinreal and fake audio features and randomly shuffles the synthesizer featureswith the content features. This strategy effectively enhances the featurediversity and simulates more feature combinations.

【10】 Language Models for Music Medicine Generation
标题:音乐医学一代的语言模型
链接:https://arxiv.org/abs/2411.09080
作者:Emmanouil Nikolakakis,  Joann Ching,  Emmanouil Karystinaios,  Gabrielle Sipin,  Gerhard Widmer,  Razvan Marinescu
备注:Late-Breaking / Demo Session Extended Abstract, ISMIR 2024 Conference
摘要:None
摘要:Music therapy has been shown in recent years to provide multiple healthbenefits related to emotional wellness. In turn, maintaining a healthyemotional state has proven to be effective for patients undergoing treatment,such as Parkinson's patients or patients suffering from stress and anxiety. Wepropose fine-tuning MusicGen, a music-generating transformer model, to createshort musical clips that assist patients in transitioning from negative todesired emotional states. Using low-rank decomposition fine-tuning on theMTG-Jamendo Dataset with emotion tags, we generate 30-second clips that adhereto the iso principle, guiding patients through intermediate states in thevalence-arousal circumplex. The generated music is evaluated using a musicemotion recognition model to ensure alignment with intended emotions. Byconcatenating these clips, we produce a 15-minute "music medicine" resembling amusic therapy session. Our approach is the first model to leverage LanguageModels to generate music medicine. Ultimately, the output is intended to beused as a temporary relief between music therapy sessions with aboard-certified therapist.

【11】 Multilingual Standalone Trustworthy Voice-Based Social Network for  Disaster Situations
标题:针对灾难情况的多语言独立可信的基于语音的社交网络
链接:https://arxiv.org/abs/2411.08889
作者:Majid Behravan,  Elham Mohammadrezaei,  Mohamed Azab,  Denis Gracanin
备注:Accepted for publication in IEEE UEMCON 2024, to appear in December 2024. 7 pages, 3 figures
摘要:在发生灾害的情况下,有效的沟通至关重要,但语言障碍往往阻碍及时和准确的信息传播,加剧了脆弱性,使应对工作复杂化。本文提出了一种新颖的,多语言的,基于语音的社交网络,专门设计来解决这些挑战。该系统将先进的人工智能(AI)与区块链技术相结合,以实现跨多种语言的安全、异步语音通信。该应用程序独立于外部服务器运行,通过本地网络离线运行,即使在受损环境中也能确保可靠性。主要功能包括人工智能驱动的语音消息实时翻译,确保无缝的跨语言通信,以及支持区块链的存储,用于所有交互的安全,不可变的记录,保护消息的完整性。该系统专为跨平台使用而设计,可在从移动电话到台式机的各种设备上提供一致的性能,使其在各种灾难情况下具有高度适应性。评估指标表明,语音识别和翻译的准确性高,延迟低,用户满意度,验证了系统在危机期间加强沟通的有效性。这一解决方案代表了灾害通信的重大进步,弥合了语言差距,以支持更具包容性和更有效的应急响应。
摘要:In disaster scenarios, effective communication is crucial, yet languagebarriers often hinder timely and accurate information dissemination,exacerbating vulnerabilities and complicating response efforts. This paperpresents a novel, multilingual, voice-based social network specificallydesigned to address these challenges. The proposed system integrates advancedartificial intelligence (AI) with blockchain technology to enable secure,asynchronous voice communication across multiple languages. The applicationoperates independently of external servers, ensuring reliability even incompromised environments by functioning offline through local networks. Keyfeatures include AI-driven real-time translation of voice messages, ensuringseamless cross-linguistic communication, and blockchain-enabled storage forsecure, immutable records of all interactions, safeguarding message integrity.Designed for cross-platform use, the system offers consistent performanceacross devices, from mobile phones to desktops, making it highly adaptable indiverse disaster situations. Evaluation metrics demonstrate high accuracy inspeech recognition and translation, low latency, and user satisfaction,validating the system's effectiveness in enhancing communication during crises.This solution represents a significant advancement in disaster communication,bridging language gaps to support more inclusive and efficient emergencyresponse.

【12】 Enhancing Lie Detection Accuracy: A Comparative Study of Classic ML,  CNN, and GCN Models using Audio-Visual Features
标题:提高谎言检测准确性:使用视听特征的经典ML、CNN和GCN模型的比较研究
链接:https://arxiv.org/abs/2411.08885
作者:Abdelrahman Abdelwahab,  Abdelrahman Abdelwahab,  Ayaan Vaswani,  Advait Bharathulwar,  Arnav Kommaraju
备注:11 pages, 18 figures
摘要:测谎仪测试的不准确通常会导致错误定罪、虚假信息和偏见,所有这些都会对法律和政治制度产生重大后果。最近,分析面部微表情已成为检测欺骗的方法,然而,目前的模型还没有达到高精度和泛化能力。本研究的目的是帮助解决这些问题。本研究中使用的独特的多模态Transformer架构通过使用听觉输入、视觉面部微表情和手动转录的手势注释来改进先前的方法,从而更接近可靠的非侵入性测谎模型。分别使用Vision Transformer和OpenSmile模型提取视觉和听觉特征,然后将其与参与者微表情和手势的transmittance连接起来。使用这些经过处理和连接的特征来训练各种模型,以分类谎言和真相。CNN Conv1D多模态模型的平均准确率为95.4%。然而,仍然需要进一步的研究来创建更高质量的数据集,甚至更广泛的模型,用于更多样化的应用。
摘要:Inaccuracies in polygraph tests often lead to wrongful convictions, falseinformation, and bias, all of which have significant consequences for bothlegal and political systems. Recently, analyzing facial micro-expressions hasemerged as a method for detecting deception; however, current models have notreached high accuracy and generalizability. The purpose of this study is to aidin remedying these problems. The unique multimodal transformer architectureused in this study improves upon previous approaches by using auditory inputs,visual facial micro-expressions, and manually transcribed gesture annotations,moving closer to a reliable non-invasive lie detection model. Visual andauditory features were extracted using the Vision Transformer and OpenSmilemodels respectively, which were then concatenated with the transcriptions ofparticipants micro-expressions and gestures. Various models were trained forthe classification of lies and truths using these processed and concatenatedfeatures. The CNN Conv1D multimodal model achieved an average accuracy of95.4%. However, further research is still required to create higher-qualitydatasets and even more generalized models for more diverse applications.

机器翻译由腾讯交互翻译提供,仅供参考