今日论文合集:cs.SD语音7篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】TinyDéjàVu: Smaller Memory Footprint & Faster Inference on Sensor Data Streams with Always-On Microcontrollers
标题:TinyDéjàVu:使用始终在线的微控制器更小的内存占用和更快的传感器数据流推理
链接:https://arxiv.org/abs/2512.09786

作者:Zhaolan Huang,Emmanuel Baccelli
摘要:人们越来越希望始终在线的传感器能够搭载各种微型神经网络,并不断对它们感测到的数据的时间序列进行推理。为了在电池上操作时满足寿命和能耗要求,这种硬件使用具有微小存储器预算的微控制器(MCU),例如,128kB RAM在这种情况下,优化跨神经网络层的数据流变得至关重要。在本文中,我们介绍了TinyDéjàVu,这是我们设计的一个新框架和新算法,用于在典型的微控制器硬件上使用各种微型ML模型进行传感器数据时间序列推断,从而大大减少所需的RAM占用量。我们将TinyDéjàVu的实现作为开源发布,并在硬件上执行可复制的基准测试。我们表明,TinyDéjàVu可以节省超过60%的RAM使用,并消除高达90%的重叠滑动窗口输入的冗余计算。
摘要:Always-on sensors are increasingly expected to embark a variety of tiny neural networks and to continuously perform inference on time-series of the data they sense. In order to fit lifetime and energy consumption requirements when operating on battery, such hardware uses microcontrollers (MCUs) with tiny memory budget e.g., 128kB of RAM. In this context, optimizing data flows across neural network layers becomes crucial. In this paper, we introduce TinyDéjàVu, a new framework and novel algorithms we designed to drastically reduce the RAM footprint required by inference using various tiny ML models for sensor data time-series on typical microcontroller hardware. We publish the implementation of TinyDéjàVu as open source, and we perform reproducible benchmarks on hardware. We show that TinyDéjàVu can save more than 60% of RAM usage and eliminate up to 90% of redundant compute on overlapping sliding window inputs.


【2】DMP-TTS: Disentangled multi-modal Prompting for Controllable Text-to-Speech with Chained Guidance
标题:DMP-TTC:具有链引导的可控文本到语音的分解多模式预处理
链接:https://arxiv.org/abs/2512.09504

作者:Kang Yin,Chunyu Qiang,Sirui Zhao,Xiaopeng Wang,Yuzhe Liang,Pengfei Cai,Tong Xu,Chen Zhang,Enhong Chen
摘要:可控制的文本到语音(TTS)系统面临着实现独立操纵扬声器音色和说话风格,往往遭受这些属性之间的纠缠的重大挑战。我们提出了DMP-TTS,一个潜在的扩散Transformer(DiT)框架,明确的解纠缠和多模态提示。基于CLAP的风格编码器(Style-CLAP)在共享空间中对齐来自参考音频和描述性文本的提示,并通过对比学习加上对风格属性的多任务监督进行训练。为了在推理过程中进行细粒度控制,我们引入了使用分层条件丢弃训练的链式无分类器指导(cCFG),可以独立调整内容,音色和风格指导强度。此外,我们还使用Representation Alignment(REPA)将声学语义特征从预训练的Whisper模型中提取到中间DiT表示中,从而稳定训练并加速收敛。实验表明,DMP-TTS提供了更强的风格可控性比开源基线,同时保持竞争力的可理解性和自然性。代码和演示将在https://y61329697.github.io/DMP-TTS/上提供。
摘要:Controllable text-to-speech (TTS) systems face significant challenges in achieving independent manipulation of speaker timbre and speaking style, often suffering from entanglement between these attributes. We present DMP-TTS, a latent Diffusion Transformer (DiT) framework with explicit disentanglement and multi-modal prompting. A CLAP-based style encoder (Style-CLAP) aligns cues from reference audio and descriptive text in a shared space and is trained with contrastive learning plus multi-task supervision on style attributes. For fine-grained control during inference, we introduce chained classifier-free guidance (cCFG) trained with hierarchical condition dropout, enabling independent adjustment of content, timbre, and style guidance strengths. Additionally, we employ Representation Alignment (REPA) to distill acoustic-semantic features from a pretrained Whisper model into intermediate DiT representations, stabilizing training and accelerating convergence. Experiments show that DMP-TTS delivers stronger style controllability than open-source baselines while maintaining competitive intelligibility and naturalness. Code and demos will be available at https://y61329697.github.io/DMP-TTS/.


【3】UniLS: End-to-End Audio-Driven Avatars for Unified Listening and Speaking
标题:UniLS:端到端音频驱动化身,实现统一的听力和口语
链接:https://arxiv.org/abs/2512.09327

作者:Xuangeng Chu,Ruicong Liu,Yifei Huang,Yun Liu,Yichen Peng,Bo Zheng
摘要:生成逼真的会话化身不仅需要建模孤立的扬声器,而且需要建模说话和倾听的动态交互。然而,对听者进行建模是非常具有挑战性的:直接音频驱动的训练失败了,产生僵硬的静态听觉动作。这种失败源于一个基本的不平衡:说话者的运动强烈地由语音音频驱动,而听者的运动主要遵循内部运动先验,并且仅由外部语音松散地引导。这一挑战使得大多数方法都专注于只说话的生成。联合生成的唯一先前尝试依赖于额外的说话者的运动来产生听者。这种设计不是端到端的,从而阻碍了实时适用性。为了解决这个限制,我们提出了UniLS,第一个端到端的框架,用于生成统一的说听表达,仅由双轨音频驱动。我们的方法引入了一种新的两阶段训练范式。阶段1首先通过训练无音频自回归生成器来学习内部运动,捕获自然面部运动的自发动态。然后,阶段2引入双声道音频,微调发生器以基于外部语音提示来调制所学习的运动先验。广泛的评估表明,UniLS达到了最先进的说话准确性。更重要的是,它提供了高达44.1%的听力指标的改善,产生更多样化和自然的听力表达。这有效地缓解了刚度问题,并为交互式数字人提供了实用的高保真音频驱动解决方案。
摘要:Generating lifelike conversational avatars requires modeling not just isolated speakers, but the dynamic, reciprocal interaction of speaking and listening. However, modeling the listener is exceptionally challenging: direct audio-driven training fails, producing stiff, static listening motions. This failure stems from a fundamental imbalance: the speaker's motion is strongly driven by speech audio, while the listener's motion primarily follows an internal motion prior and is only loosely guided by external speech. This challenge has led most methods to focus on speak-only generation. The only prior attempt at joint generation relies on extra speaker's motion to produce the listener. This design is not end-to-end, thereby hindering the real-time applicability. To address this limitation, we present UniLS, the first end-to-end framework for generating unified speak-listen expressions, driven by only dual-track audio. Our method introduces a novel two-stage training paradigm. Stage 1 first learns the internal motion prior by training an audio-free autoregressive generator, capturing the spontaneous dynamics of natural facial motion. Stage 2 then introduces the dual-track audio, fine-tuning the generator to modulate the learned motion prior based on external speech cues. Extensive evaluations show UniLS achieves state-of-the-art speaking accuracy. More importantly, it delivers up to 44.1\% improvement in listening metrics, generating significantly more diverse and natural listening expressions. This effectively mitigates the stiffness problem and provides a practical, high-fidelity audio-driven solution for interactive digital humans.


【4】VABench: A Comprehensive Benchmark for Audio-Video Generation
标题:VABench:音频视频生成的全面基准
链接:https://arxiv.org/abs/2512.09299

作者:Daili Hua,Xizhi Wang,Bohan Zeng,Xinyi Huang,Hao Liang,Junbo Niu,Xinlong Chen,Quanqing Xu,Wentao Zhang
备注:24 pages, 25 figures
摘要:视频生成方面的最新进展非常显著,使模型能够生成具有同步音频的视觉上引人注目的视频。虽然现有的视频生成基准提供了全面的视觉质量指标,但它们缺乏令人信服的音视频生成评估,特别是对于旨在生成同步音视频输出的模型。为了解决这一差距,我们引入VABench,一个全面的和多维的基准框架,旨在系统地评估同步音视频生成的能力。VABench包含三种主要任务类型:文本到音频视频(T2 AV),图像到音频视频(I2 AV)和立体声音频视频生成。它还建立了两个主要的评价模块,涵盖15个方面。这些维度专门评估成对相似性(文本-视频,文本-音频,视频-音频),音频-视频同步,唇语一致性以及精心策划的音频和视频问答(QA)对等。此外,VABench涵盖了七个主要内容类别:动物,人类声音,音乐,环境声音,同步物理声音,复杂场景和虚拟世界。我们对评估结果进行了系统的分析和可视化,旨在为评估具有同步音频功能的视频生成模型建立新的标准,并促进该领域的全面进步。
摘要:Recent advances in video generation have been remarkable, enabling models to produce visually compelling videos with synchronized audio. While existing video generation benchmarks provide comprehensive metrics for visual quality, they lack convincing evaluations for audio-video generation, especially for models aiming to generate synchronized audio-video outputs. To address this gap, we introduce VABench, a comprehensive and multi-dimensional benchmark framework designed to systematically evaluate the capabilities of synchronous audio-video generation. VABench encompasses three primary task types: text-to-audio-video (T2AV), image-to-audio-video (I2AV), and stereo audio-video generation. It further establishes two major evaluation modules covering 15 dimensions. These dimensions specifically assess pairwise similarities (text-video, text-audio, video-audio), audio-video synchronization, lip-speech consistency, and carefully curated audio and video question-answering (QA) pairs, among others. Furthermore, VABench covers seven major content categories: animals, human sounds, music, environmental sounds, synchronous physical sounds, complex scenes, and virtual worlds. We provide a systematic analysis and visualization of the evaluation results, aiming to establish a new standard for assessing video generation models with synchronous audio capabilities and to promote the comprehensive advancement of the field.


【5】Who Speaks What from Afar: Eavesdropping In-Person Conversations via mmWave Sensing
标题:谁在远方说了什么:通过mmWave Sensing进行令人惊叹的面对面对话
链接:https://arxiv.org/abs/2512.09285

作者:Shaoying Wang,Hansong Zhou,Yukun Yuan,Xiaonan Zhang
摘要:多参与者会议发生在各个领域,例如商业谈判和医疗咨询,在此期间,经常讨论商业秘密,商业策略和患者状况等敏感信息。先前的研究表明,在房间外使用毫米波雷达的攻击者可以通过检测物体上微小的语音引起的振动来偷听会议内容。然而,这些窃听攻击无法区分多人参与的会议中哪些语音内容来自哪个人,从而导致潜在的误解和决策失误。在本文中,我们回答了“谁在说什么”的问题。通过利用无处不在的对象引入的空间多样性,我们提出了一个攻击系统,使攻击者能够远程窃听在人的对话,而不需要先验知识,如身份,参与者的数量,或座位安排。由于面对面会议的参与者通常坐在不同的位置,他们的讲话会在附近的物体上引起不同的振动模式。为了利用这一点,我们设计了一种噪声鲁棒的无监督方法,通过检测频域中的语音引起的振动差异来区分参与者。同时,探索了一种基于深度学习的框架来组合来自对象的信号以增强语音质量。我们通过大量的实验验证了对语音分类和信号增强的概念验证攻击。实验结果表明,我们的攻击可以实现语音分类的准确率高达0.99 $与多个与会者在一个会议室。与此同时,我们的攻击在所有真实场景中都表现出一致的语音质量增强,包括雷达和物体之间的不同距离。
摘要:Multi-participant meetings occur across various domains, such as business negotiations and medical consultations, during which sensitive information like trade secrets, business strategies, and patient conditions is often discussed. Previous research has demonstrated that attackers with mmWave radars outside the room can overhear meeting content by detecting minute speech-induced vibrations on objects. However, these eavesdropping attacks cannot differentiate which speech content comes from which person in a multi-participant meeting, leading to potential misunderstandings and poor decision-making. In this paper, we answer the question ``who speaks what''. By leveraging the spatial diversity introduced by ubiquitous objects, we propose an attack system that enables attackers to remotely eavesdrop on in-person conversations without requiring prior knowledge, such as identities, the number of participants, or seating arrangements. Since participants in in-person meetings are typically seated at different locations, their speech induces distinct vibration patterns on nearby objects. To exploit this, we design a noise-robust unsupervised approach for distinguishing participants by detecting speech-induced vibration differences in the frequency domain. Meanwhile, a deep learning-based framework is explored to combine signals from objects for speech quality enhancement. We validate the proof-of-concept attack on speech classification and signal enhancement through extensive experiments. The experimental results show that our attack can achieve the speech classification accuracy of up to $0.99$ with several participants in a meeting room. Meanwhile, our attack demonstrates consistent speech quality enhancement across all real-world scenarios, including different distances between the radar and the objects.


【6】ORCA: Open-ended Response Correctness Assessment for Audio Question Answering
标题:ORCA:音频问题回答的开放式响应正确性评估
链接:https://arxiv.org/abs/2512.09066

作者:Šimon Sedláček,Sara Barahona,Bolaji Yusuf,Laura Herrera-Alarcón,Santosh Kesiraju,Cecilia Bolaños,Alicia Lozano-Diez,Sathvik Udupa,Fernando López,Allison Ferner,Ramani Duraiswami,Jan Černocký
摘要:评估来自大型音频语言模型(LALM)的开放式响应是具有挑战性的,因为由于多个有效解释,部分正确性和主观判断,人类注释者通常真正不同意答案的正确性。传统的指标只报告平均分数,无法捕捉这种不确定性。我们提出了ORCA(开放式响应正确性评估),一个框架,模型的变化,在人类的判断使用贝塔分布预测预期的正确性和不确定性。我们的三阶段注释框架将人类判断与结构化反馈和迭代细化相结合,以同时管理训练数据并提高基准测试质量。我们在两个音频QA基准测试中从15个LALM收集了3,580个问答对中的11,721个注释,实现了0.82(Krippendorff的alpha)的注释者间一致性。ORCA与平均人类判断的斯皮尔曼相关性达到0.91,匹配或优于LLM判断基线,同时提供不确定性估计,并需要显著减少计算。我们发布我们的模型,代码和策展数据集。
摘要:Evaluating open-ended responses from large audio language models (LALMs) is challenging because human annotators often genuinely disagree on answer correctness due to multiple valid interpretations, partial correctness, and subjective judgment. Traditional metrics reporting only mean scores fail to capture this uncertainty. We present ORCA (Open-ended Response Correctness Assessment), a framework that models the variability in human judgments using Beta distributions to predict both expected correctness and uncertainty. Our three-stage annotation framework combines human judgment with structured feedback and iterative refinement to simultaneously curate training data and improve benchmark quality. We collected 11,721 annotations across 3,580 question-answer pairs from 15 LALMs on two audio QA benchmarks, achieving inter-annotator agreement of 0.82 (Krippendorff's alpha). ORCA achieves 0.91 Spearman correlation with mean human judgments, matching or outperforming LLM-judge baselines while providing uncertainty estimates and requiring significantly less compute. We release our models, code, and curated dataset.


【7】Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
标题:通过集成噪音检测架构增强自动语音识别
链接:https://arxiv.org/abs/2512.08973

作者:Karamvir Singh
备注:5 figures
摘要:这项研究提出了一种新的方法,以提高自动语音识别系统的集成噪声检测能力直接到识别架构。基于wav2vec2框架,所提出的方法采用了一个专用的噪声识别模块,同时与语音转录。使用公开可用的语音和环境音频数据集的实验验证表明,转录质量和噪声歧视的显着改善。增强的系统在字错误率,字符错误率,和噪声检测精度相比,传统的架构实现了卓越的性能。结果表明,联合优化的转录和噪声分类目标产生更可靠的语音识别在具有挑战性的声学条件。
摘要:This research presents a novel approach to enhancing automatic speech recognition systems by integrating noise detection capabilities directly into the recognition architecture. Building upon the wav2vec2 framework, the proposed method incorporates a dedicated noise identification module that operates concurrently with speech transcription. Experimental validation using publicly available speech and environmental audio datasets demonstrates substantial improvements in transcription quality and noise discrimination. The enhanced system achieves superior performance in word error rate, character error rate, and noise detection accuracy compared to conventional architectures. Results indicate that joint optimization of transcription and noise classification objectives yields more reliable speech recognition in challenging acoustic conditions.


eess.AS音频处理


【1】Robust Speech Activity Detection in the Presence of Singing Voice
标题:存在歌声时的鲁棒语音活动检测
链接:https://arxiv.org/abs/2512.09713

作者:Philipp Grundhuber,Mhd Modar Halimeh,Martin Strauß,Emanuël A. P. Habets
备注:This paper has been published in: 2025 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA)
摘要:语音活动检测(SAD)系统经常将歌唱误分类为语音,导致对话增强和自动语音识别等应用的性能下降。我们介绍了歌唱鲁棒语音活动检测(SR-SAD),一个神经网络,旨在鲁棒地检测语音存在的歌唱。我们的主要贡献是:i)使用语音和歌唱样本的受控比率来提高辨别力的训练策略,ii)在减少推理运行时间的同时保持鲁棒性能的计算高效模型,以及iii)定制为评估混合语音歌唱场景中的SAD鲁棒性的新评估度量。在跨越多种音乐类型的具有挑战性的数据集上进行的实验表明,SR-SAD在拒绝唱歌的同时保持了较高的语音检测精度(AUC = 0.919)。通过明确学习区分语音和歌唱,SR-SAD在混合语音歌唱场景中实现了更可靠的SAD。
摘要:Speech Activity Detection (SAD) systems often misclassify singing as speech, leading to degraded performance in applications such as dialogue enhancement and automatic speech recognition. We introduce Singing-Robust Speech Activity Detection ( SR-SAD ), a neural network designed to robustly detect speech in the presence of singing. Our key contributions are: i) a training strategy using controlled ratios of speech and singing samples to improve discrimination, ii) a computationally efficient model that maintains robust performance while reducing inference runtime, and iii) a new evaluation metric tailored to assess SAD robustness in mixed speech-singing scenarios. Experiments on a challenging dataset spanning multiple musical genres show that SR-SAD maintains high speech detection accuracy (AUC = 0.919) while rejecting singing. By explicitly learning to distinguish between speech and singing, SR-SAD enables more reliable SAD in mixed speech-singing scenarios.


【2】Human perception of audio deepfakes: the role of language and speaking style
标题:人类对深度伪造音频的感知:语言和说话风格的作用
链接:https://arxiv.org/abs/2512.09221

作者:Eugenia San Segundo,Aurora López-Jareño,Xin Wang,Junichi Yamagishi
备注:Submitted to Speech Communication
摘要:音频深度伪造已经达到了一定的现实水平,使得区分人类和人工声音变得越来越困难,这带来了身份盗窃或虚假信息传播等风险。尽管存在这些担忧,但关于人类识别深度伪造的能力的研究仍然有限,大多数研究都集中在英语上,很少有人探索听众感知决定背后的原因。本研究通过一项感知实验解决了这一差距,在该实验中,54名听众(28名母语为西班牙语的人和26名母语为日语的人)将声音分类为自然或合成,并证明他们的选择是合理的。该实验包括80个刺激(50%人工),根据三个变量组织:语言(西班牙语/日语),演讲风格(有声读物/采访)和熟悉的声音(熟悉/不熟悉)。我们的目标是研究这些变量如何影响检测,并定性分析背后的推理听众的感知决定。结果表明,平均准确率为59.11%,具有更高的真实样本的性能。对声音自然度的判断依赖于语言和非语言线索的结合。比较日本和西班牙的听众,我们的定性分析进一步揭示了共享的线索和显着的跨语言的差异,听众如何概念化的“人性”的讲话。总体而言,参与者主要依赖于超音段和更高层次或语言外的特征-如语调,节奏,流畅性,停顿,速度,呼吸和笑声-而不是音段特征。这些发现强调了人类感知策略在区分自然语音和人工语音方面的复杂性,并与先前强调韵律和自发语音典型现象(如不流利)的重要性的研究部分一致。
摘要:Audio deepfakes have reached a level of realism that makes it increasingly difficult to distinguish between human and artificial voices, which poses risks such as identity theft or spread of disinformation. Despite these concerns, research on humans' ability to identify deepfakes is limited, with most studies focusing on English and very few exploring the reasons behind listeners' perceptual decisions. This study addresses this gap through a perceptual experiment in which 54 listeners (28 native Spanish speakers and 26 native Japanese speakers) classified voices as natural or synthetic, and justified their choices. The experiment included 80 stimuli (50% artificial), organized according to three variables: language (Spanish/Japanese), speech style (audiobooks/interviews), and familiarity with the voice (familiar/unfamiliar). The goal was to examine how these variables influence detection and to analyze qualitatively the reasoning behind listeners' perceptual decisions. Results indicate an average accuracy of 59.11%, with higher performance on authentic samples. Judgments of vocal naturalness rely on a combination of linguistic and non-linguistic cues. Comparing Japanese and Spanish listeners, our qualitative analysis further reveals both shared cues and notable cross-linguistic differences in how listeners conceptualize the "humanness" of speech. Overall, participants relied primarily on suprasegmental and higher-level or extralinguistic characteristics - such as intonation, rhythm, fluency, pauses, speed, breathing, and laughter - over segmental features. These findings underscore the complexity of human perceptual strategies in distinguishing natural from artificial speech and align partly with prior research emphasizing the importance of prosody and phenomena typical of spontaneous speech, such as disfluencies.


【3】LG Uplus System with Multi-Speaker IDs and Discriminator-based Sub-Judges for the WildSpoof Challenge
标题:LG Uplus系统采用多扬声器ID和基于鉴别器的子法官,用于WildSpoof挑战
链接:https://arxiv.org/abs/2512.09000

作者:Jinyoung Park,Won Jang,Jiwoong Park
备注:3 pages, 2 figures, 2 tables
摘要:本文介绍了我们提交的WildSpoof挑战赛轨道2,其重点是欺骗感知扬声器验证(SASV)在高质量的文本到语音(TTS)攻击的存在。我们采用ResNet-221骨干,研究了两种说话人标记策略,namyDual-Speaker ID和Multi-Speaker ID,以显式地扩大嵌入空间中真实语音和生成语音之间的裕度。此外,我们提出了基于判别器的子判断系统,该系统重用HiFi-GAN和BigVGAN判别器的内部特征,通过多查询多头注意统计池(MQMHA)进行聚合。在SpoofCeleb语料库上的实验结果表明,该系统能够有效地提高不可知检测代价函数(a-DCF)。
摘要:This paper describes our submission to the WildSpoof Challenge Track 2, which focuses on spoof-aware speaker verification (SASV) in the presence of high-quality text-to-speech (TTS) attacks. We adopt a ResNet-221 back-bone and study two speaker-labeling strategies, namelyDual-Speaker IDs and Multi-Speaker IDs, to explicitly enlarge the margin between bona fide and generated speech in the embedding space. In addition, we propose discriminator-based sub-judge systems that reuse internal features from HiFi-GAN and BigVGAN discriminators, aggregated via multi-query multi-head attentive statistics pooling(MQMHA). Experimental results on the SpoofCeleb corpus show that our system design is effective in improving agnostic detection cost function (a-DCF).


【4】TinyDéjàVu: Smaller Memory Footprint & Faster Inference on Sensor Data Streams with Always-On Microcontrollers
标题:TinyDéjàVu:使用始终在线的微控制器更小的内存占用和更快的传感器数据流推理
链接:https://arxiv.org/abs/2512.09786

作者:Zhaolan Huang,Emmanuel Baccelli
摘要:人们越来越希望始终在线的传感器能够搭载各种微型神经网络,并不断对它们感测到的数据的时间序列进行推理。为了在电池上操作时满足寿命和能耗要求,这种硬件使用具有微小存储器预算的微控制器(MCU),例如,128kB RAM在这种情况下,优化跨神经网络层的数据流变得至关重要。在本文中,我们介绍了TinyDéjàVu,这是我们设计的一个新框架和新算法,用于在典型的微控制器硬件上使用各种微型ML模型进行传感器数据时间序列推断,从而大大减少所需的RAM占用量。我们将TinyDéjàVu的实现作为开源发布,并在硬件上执行可复制的基准测试。我们表明,TinyDéjàVu可以节省超过60%的RAM使用,并消除高达90%的重叠滑动窗口输入的冗余计算。
摘要:Always-on sensors are increasingly expected to embark a variety of tiny neural networks and to continuously perform inference on time-series of the data they sense. In order to fit lifetime and energy consumption requirements when operating on battery, such hardware uses microcontrollers (MCUs) with tiny memory budget e.g., 128kB of RAM. In this context, optimizing data flows across neural network layers becomes crucial. In this paper, we introduce TinyDéjàVu, a new framework and novel algorithms we designed to drastically reduce the RAM footprint required by inference using various tiny ML models for sensor data time-series on typical microcontroller hardware. We publish the implementation of TinyDéjàVu as open source, and we perform reproducible benchmarks on hardware. We show that TinyDéjàVu can save more than 60% of RAM usage and eliminate up to 90% of redundant compute on overlapping sliding window inputs.


【5】Enhancing Automatic Speech Recognition Through Integrated Noise Detection Architecture
标题:通过集成噪音检测架构增强自动语音识别
链接:https://arxiv.org/abs/2512.08973

作者:Karamvir Singh
备注:5 figures
摘要:这项研究提出了一种新的方法,以提高自动语音识别系统的集成噪声检测能力直接到识别架构。基于wav2vec2框架,所提出的方法采用了一个专用的噪声识别模块,同时与语音转录。使用公开可用的语音和环境音频数据集的实验验证表明,转录质量和噪声歧视的显着改善。增强的系统在字错误率,字符错误率,和噪声检测精度相比,传统的架构实现了卓越的性能。结果表明,联合优化的转录和噪声分类目标产生更可靠的语音识别在具有挑战性的声学条件。
摘要:This research presents a novel approach to enhancing automatic speech recognition systems by integrating noise detection capabilities directly into the recognition architecture. Building upon the wav2vec2 framework, the proposed method incorporates a dedicated noise identification module that operates concurrently with speech transcription. Experimental validation using publicly available speech and environmental audio datasets demonstrates substantial improvements in transcription quality and noise discrimination. The enhanced system achieves superior performance in word error rate, character error rate, and noise detection accuracy compared to conventional architectures. Results indicate that joint optimization of transcription and noise classification objectives yields more reliable speech recognition in challenging acoustic conditions.


机器翻译由腾讯交互翻译提供,仅供参考