微信公众号:arXiv_Daily
cs.SD语音
标题:MiniMind-O技术报告:开放式小规模语音原生全方位模型
链接:https://arxiv.org/abs/2605.03937
备注:17 pages. Code, checkpoints, and training data are available at https://github.com/jingyaogong/minimind-o
摘要:MiniMind-O是一个基于MiniMind语言模型的开放的0. 1B级omni模型。它接受文本、语音和图像输入,并返回文本和流式语音。该版本包括模型代码,检查点和用于文本到音频,图像到文本和音频到音频训练的主要Parquet训练数据集,使完整的交互循环直接可检查。该模型使用完整的MiniMind主干作为思想者,并使用由MiniMind块制成的独立四层Talker。Frozen SenseVoice-Small和SigLIP 2编码器提供语音和图像特征,这些特征由轻量级MLP投影仪映射并注入模态占位符位置。Talker读取中间层Thinker状态以及自回归八层Mimi代码缓冲区。扬声器控制由专用扬声器令牌、右对齐参考编解码器提示和预先计算的CAM++扬声器嵌入处理,因此语音调节仍然是音频代码上下文的一部分,而不是单独的TTS模块。对于768维Talker,密集和MoE变体在Thinker-Talker一致性评价中达到0.0897和0.0900的平均CER,总体语音克隆相似性为0.5995和0.5937。除了报告一个工作系统之外,本文还确定了小型omni模型的三个关键设计选择:中间层语义桥接,发布的多模态序列格式和参数高效的八码本接口。
摘要:MiniMind-O is an open 0.1B-scale omni model built on the MiniMind language model. It accepts text, speech, and image inputs, and returns both text and streaming speech. The release includes model code, checkpoints, and the main Parquet training datasets for text-to-audio, image-to-text, and audio-to-audio training, making the complete interaction loop directly inspectable. The model uses a full MiniMind backbone as the Thinker and an independent four-layer Talker made from MiniMind blocks. Frozen SenseVoice-Small and SigLIP2 encoders provide speech and image features, which are mapped by lightweight MLP projectors and injected at modality-placeholder positions. The Talker reads a middle-layer Thinker state together with an autoregressive eight-layer Mimi-code buffer. Speaker control is handled by a dedicated speaker token, right-aligned reference codec prompts, and precomputed CAM++ speaker embeddings, so voice conditioning remains part of the audio-code context rather than a separate TTS module. With a 768-dimensional Talker, the dense and MoE variants reach average CERs of 0.0897 and 0.0900 in Thinker--Talker consistency evaluation, with overall voice-cloning similarities of 0.5995 and 0.5937. Beyond reporting a working system, the paper identifies three scale-critical design choices for small omni models: middle-layer semantic bridging, a released multimodal sequence format, and a parameter-efficient eight-codebook interface.
【2】Towards Open World Sound Event Detection
标题:迈向开放世界声音事件检测链接:https://arxiv.org/abs/2605.03934
备注:32 pages, 3 figures. Submitted to Signal Processing (Elsevier)
摘要:声音事件检测(SED)在音频理解中起着至关重要的作用,应用于监控、智慧城市、医疗保健和多媒体索引。然而,传统的SED系统在封闭世界的假设下操作,限制了它们在经常出现新的声学事件的真实世界环境中的有效性。受开放世界学习在计算机视觉中的成功启发,我们引入了开放世界声音事件检测(OW-SED)范式,其中模型必须检测已知事件,识别未知事件,并逐步从中学习。为了解决OW-SED的独特挑战,如重叠和模糊的事件,我们提出了一个1D可变形的架构,利用可变形的注意力,自适应地集中在突出的时间区域。此外,我们设计了一种新的开放世界可变形声音事件检测Transformer(WOOT)框架,结合特征解纠缠来分离类特定和类不可知的表示,以及一对多的匹配策略和多样性损失,以提高表示的多样性。实验结果表明,我们的方法实现了略优于现有的领先技术相比,在封闭世界的设置和显着改善现有的基线在开放世界的情况下,性能。
摘要:Sound Event Detection (SED) plays a vital role in audio understanding, with applications in surveillance, smart cities, healthcare, and multimedia indexing. However, conventional SED systems operate under a closed-world assumption, limiting their effectiveness in real-world environments where novel acoustic events frequently emerge. Inspired by the success of open-world learning in computer vision, we introduce the Open-World Sound Event Detection (OW-SED) paradigm, where models must detect known events, identify unseen ones, and incrementally learn from them. To tackle the unique challenges of OW-SED, such as overlapping and ambiguous events, we propose a 1D Deformable architecture that leverages deformable attention to adaptively focus on salient temporal regions. Furthermore, we design a novel Open-World Deformable Sound Event Detection Transformer (WOOT) framework incorporating feature disentanglement to separate class-specific and class-agnostic representations, together with a one-to-many matching strategy and a diversity loss to enhance representation diversity. Experimental results demonstrate that our method achieves marginally superior performance compared to existing leading techniques in closed-world settings and significantly improves over existing baselines in open-world scenarios.
【3】PHALAR: Phasors for Learned Musical Audio Representations
标题:PHALAR:用于学习音乐音频表示的相位器链接:https://arxiv.org/abs/2605.03929
摘要:词干检索是将缺失的词干与给定的音频子混合进行匹配的任务,是目前受到丢弃时间信息的模型限制的关键挑战。我们介绍PHALAR,一个对比框架,实现了相对准确性的增加高达$\approximately 70\%$的国家的最先进的,而需要$<50\%$的参数和7 $\times $的训练加速。通过利用学习频谱池层和复值头,PHALAR强制执行音高等变和相位等变偏置。PHALAR在MoisesDB、Slakh和ChocoChorales之间建立了新的最先进的检索技术,与人类一致性判断的相关性显著高于语义基线。最后,zero-shot节拍跟踪和线性和弦探测证实,PHALAR捕捉强大的音乐结构超出检索任务。
摘要:Stem retrieval, the task of matching missing stems to a given audio submix, is a key challenge currently limited by models that discard temporal information. We introduce PHALAR, a contrastive framework achieving a relative accuracy increase of up to $\approx 70\%$ over the state-of-the-art while requiring $<50\%$ of the parameters and a 7$\times$ training speedup. By utilizing a Learned Spectral Pooling layer and a complex-valued head, PHALAR enforces pitch-equivariant and phase-equivariant biases. PHALAR establishes new retrieval state-of-the-art across MoisesDB, Slakh, and ChocoChorales, correlating significantly higher with human coherence judgment than semantic baselines. Finally, zero-shot beat tracking and linear chord probing confirm that PHALAR captures robust musical structures beyond the retrieval task.
【4】Ecologically-Constrained Task Arithmetic for Multi-Taxa Bioacoustic Classifiers Without Shared Data
标题:无共享数据的多分类群生物声学分类器的动态约束任务算法链接:https://arxiv.org/abs/2605.03914
摘要:生物声学的训练数据分散在分类群、地区和机构中。将所有这些集中起来往往是不可行的。我们表明,独立微调BEAT编码器可以组成一个统一的661种分类器,通过任务向量算法,而无需共享数据。我们发现,生物声学任务向量是近正交(余弦0.01-0.09)。它们的分离与光谱分布距离密切相关,这一梯度与声学生态位假设一致。这种几何结构使得简单的平均化成为最佳,而符号冲突方法将准确性降低了一到六个百分点。组成也造成了一个不对称的差距:物种丰富的群体失去了相对于联合培训的准确性,而代表性不足的类群增益,重新分配有利于公平的生物多样性监测。我们验证了所有分类对的线性模式连接,展示了zero-shot转移到新的区域,并确定域否定作为组合失败的边界条件。这些结果使生物声学的合作范例,其中机构共享只组装多分类群分类器的任务向量,保护数据隐私。
摘要:Training data for bioacoustics is scattered across taxa, regions, and institutions. Centralizing it all is often infeasible. We show that independently fine-tuned BEATs encoders can be composed into a unified 661-species classifier via task vector arithmetic without sharing data. We find that bioacoustic task vectors are near-orthogonal (cosine 0.01-0.09). Their separation aligns closely with spectral distribution distance, a gradient consistent with the acoustic niche hypothesis. This geometry makes simple averaging optimal while sign-conflict methods reduce accuracy by one to six percentage points. Composition also creates an asymmetric gap: species-rich groups lose accuracy relative to joint training while underrepresented taxa gain, a redistribution useful for equitable biodiversity monitoring. We verify linear mode connectivity across all taxonomic pairs, demonstrate zero-shot transfer to new regions, and identify domain negation as a boundary condition where composition fails. These results enable a collaborative paradigm for bioacoustics where institutions share only task vectors to assemble multi-taxa classifiers, preserving data privacy.
【5】AfriVox-v2: A Domain-Verticalized Benchmark for In-the-Wild African Speech Recognition
标题:AfriVox-v2:野外非洲语音识别的领域垂直基准链接:https://arxiv.org/abs/2605.03590
摘要:最近的大型语言模型(LLM)显示出强大的语音识别和高资源语言的翻译能力。然而,非洲语言在基准中的代表性仍然严重不足,限制了它们在低资源环境中的实际使用。虽然早期的基准测试非洲语言和口音,但它们缺乏详尽的真实世界噪音和粒度域评估。我们提出了AfriVox-v2,一个全面的基准,旨在测试语音模型在现实的非洲部署条件下。AfriVox-v2为所有支持的语言引入了“野生”无脚本音频。我们还引入了严格的领域垂直化,评估政府、金融、卫生和农业等十个行业的模型准确性,并对数字和命名实体进行有针对性的测试。最后,我们对新一代语音模型进行了基准测试,包括Sahara-v2,Gemini 3 Flash和Omnilingual CTC模型。我们的研究结果揭示了现代语音模型在专门的、嘈杂的非洲环境中的真正泛化差距,并为开发人员构建本地化语音AI提供了可靠的蓝图。
摘要:Recent large language models (LLMs) show strong speech recognition and translation capabilities for high-resource languages. However, African languages remain dramatically underrepresented in benchmarks, limiting their practical use in low-resource settings. While early benchmarks tested African languages and accents, they lacked exhaustive real-world noise and granular domain evaluations. We present AfriVox-v2, a comprehensive benchmark designed to test speech models under realistic African deployment conditions. AfriVox-v2 introduces "in the wild" unscripted audio for all supported languages. We also introduce strict domain verticalization, evaluating model accuracy across ten sectors including government, finance, health, and agriculture and conducting targeted tests on numbers and named entities. Finally, we benchmark a new generation of speech models, including Sahara-v2, Gemini 3 Flash, and the Omnilingual CTC models. Our results expose the true generalization gap of modern speech models in specialized, noisy African contexts and provide a reliable blueprint for developers building localized voice AI.
【6】Cosmodoit: A Python Package for Adaptive, Efficient Pipelining of Feature Extraction from Performed Music
标题:Cosmodoit:一个用于自适应、高效地从表演音乐中提取特征的Python包链接:https://arxiv.org/abs/2605.03541
备注:6 pages, 1 figure
摘要:表演音乐的计算分析是音乐信息研究的一个关键组成部分,因为表演塑造了我们听到的大部分音乐。音乐表演分析研究由表演者引入的声学变化以及这些变化如何反映音乐的解释和结构。虽然存在许多算法和工具用于执行性能与分数对齐以及符号或音频特征提取等任务,但它们分布在不同的编程语言和数据格式中,使得它们难以有效组合。为了解决这个问题,我们提出了Cosmodoit,一种新颖的Python包,旨在简化从表演音乐中提取特征。Cosmodoit在一个模块化、灵活的管道中集成了性能到分数的对齐与符号和音频特征提取,该管道支持选择性处理、依赖性感知计算和增量更新。其可扩展设计减少了重复工作,最大限度地减少了错误,并实现了高效的大规模处理。通过容纳以多种语言实现的算法,并允许参数调整以实现一致的特征提取,Cosmodoit为音乐性能分析的研究和开发提供了一个多功能和实用的工具。
摘要:Computational analysis of performed music is a key component of music information research, as performance shapes much of the music we hear. Music performance analysis studies the acoustic variations introduced by performers and how these variations reflect musical interpretation and structure. Although many algorithms and tools exist for tasks such as performance-to-score alignment and symbolic or audio feature extraction, they are spread across different programming languages and data formats, making them difficult to combine efficiently. To address this problem, we present Cosmodoit, a novel Python package designed to streamline feature extraction from performed music. Cosmodoit integrates performance-to-score alignment with symbolic and audio feature extraction in a modular, flexible pipeline that supports selective processing, dependency-aware computation, and incremental updates. Its extensible design reduces duplicated work, minimizes errors, and enables efficient large-scale processing. By accommodating algorithms implemented in multiple languages and allowing parameter tuning for consistent feature extraction, Cosmodoit provides a versatile and practical tool for both research and development in music performance analysis.
【7】Deepfake Audio Detection Using Self-supervised Fusion Representations
标题:使用自监督融合表示的Deepfake音频检测链接:https://arxiv.org/abs/2605.03420
摘要:本文描述了2026年环境感知语音和声音Deepfake检测挑战赛(ESDD 2)的提交内容,该挑战赛使用CompSpoofV 2数据集解决了组件级的deepfake检测,其中语音和环境声音可以独立操作。为了应对这一挑战,提出了一种双分支深度伪造检测框架,用于从输入音频中联合建模语音和环境上下文表示。两个预训练的模型,XLS-R的语音和BEAT的环境声音,用于提取互补的上下文表示。引入匹配头,通过统计归一化和表示交互来建模表示差异,从而实现对原始类的估计。同时,多头交叉注意使语音和环境组件之间的有效信息交换。细化的表示用残余连接和层归一化进行处理,并传递到AASIST分类器以预测基于语音和基于环境的欺骗概率。该模型输出原始、语音和环境预测。在测试集上,该系统实现了70.20%的F1分数和16.54%的环境EER,优于基线系统。
摘要:This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds may be independently manipulated. To address this challenge, a dual-branch deepfake detection framework is proposed to jointly model speech and environmental contextual representations from input audio. Two pretrained models, XLS-R for speech and BEATs for environmental sound, are used to extract complementary contextual representations. A Matching Head is introduced to model representation differences through statistical normalization and representation interaction, enabling estimation of the original class. In parallel, multi-head cross-attention enables effective information exchange between speech and environmental components. The refined representations are processed with residual connections and layer normalization, and passed to an AASIST classifier to predict speech-based and environment-based spoofing probabilities. The model outputs original, speech, and environment predictions. On the test set, the proposed system achieves an F1-score of 70.20% and an environmental EER of 16.54%, outperforming the baseline system.
【8】Smart Passive Acoustic Monitoring: Embedding a Classifier on AudioMoth Microcontroller
标题:智能被动声学监测:在AudioMoth微控制器上嵌入分类器链接:https://arxiv.org/abs/2605.03412
备注:3 pages, 1 table, 2 figures. Video associated
摘要:被动声监测(PAM)是一种有效的、非侵入性的生态系统监测方法。通常,自主记录器允许采集大量的生物声学数据集,然后进行分析。然而,功耗和数据存储都是稀缺的,并限制了采集活动的持续时间。为了解决这个问题,我们提出了一个智能PAM系统,它允许通过直接嵌入到AudioMoth微控制器的分类器的音景的现场分析。具体来说,我们提出了一个优化但简单的1D卷积神经网络(1D-CNN)来对原始音频进行分类。该模型专注于Scopoli Shearwater海鸟(濒危物种)的特定呼叫,并在真实世界的数据集上进行训练,分类准确率为91%(平衡准确率为89%)。我们还提出了一个过程来优化模型,以适应AudioMoth的严重资源限制,实现了10 kB的RAM内存占用和20 ms的推理时间。最后,我们提供了一个关于模型优化和导出策略的开源教程,该教程可用于嵌入超出我们研究范围的模型。我们修改后的AudioMoth固件增加了两个功能:(F1)在检测到目标物种时选择性地记录数据,(F2)实时记录连续分类结果。这项工作旨在促进智能传感器的概念,提高生物声学监测活动的效率和可扩展性。
摘要:Passive Acoustic Monitoring (PAM) is an efficient and non-invasive method for surveying ecosystems at a reduced cost. Typically, autonomous recorders allow the acquisition of vast bioacoustic datasets which are then analyzed. However, power consumption and data storage are both scarce and limit the duration of acquisition campaigns. To address this issue, we propose a smart PAM system which allows the in-situ analysis of the soundscape by embedding a classifier directly onto an AudioMoth microcontroller. Specifically, we propose an optimized yet simple 1D Convolutional Neural Network (1D-CNN) to classify the raw audio. The model focuses on the specific call of Scopoli Shearwater seabirds (endangered species) and is trained on a real-world dataset with a classification accuracy of 91\% (balanced accuracy of 89\%). We also propose a process to optimize the model to fit the severe resource constraints of the AudioMoth, achieving a \~10kB RAM memory footprint and 20ms inference time. Finally, we present an open-source tutorial of our model optimization and export strategy which can be used for embedding models beyond the scope of our study. Our modified version of the AudioMoth firmware adds two functions: (F1) which selectively records data when the target species has been detected and (F2) which logs the continuous classification results in real time. This work intends to facilitate the conception of intelligent sensors, enhancing the efficiency and scalability of bioacoustic monitoring campaigns.
【9】APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music
标题:APEX:人工智能生成音乐的大规模多任务美学受欢迎程度预测链接:https://arxiv.org/abs/2605.03395
摘要:音乐流行度预测吸引了越来越多的研究兴趣,与艺术家,平台和推荐系统相关。然而,人工智能生成的音乐平台的爆炸性崛起创造了一个全新的、基本上未被探索的景观,每天都有大量歌曲被制作和消费,而没有传统的艺术家声誉或标签支持。在这种追求中,关键但尚未探索的是美学质量。我们提出了APEX,这是第一个用于人工智能生成音乐的大规模多任务学习框架,在Suno和Udio的超过211 k首歌曲(10 k小时的音频)上进行了训练,该框架联合预测基于成就的流行信号-流和喜欢分数-以及从MERT(一种自我监督的音乐理解模型)提取的冻结音频嵌入的五个感知美学质量维度。美学质量和流行度捕捉音乐的互补方面,共同证明是有价值的:在对音乐竞技场数据集的分布外评估中,包括训练期间未见过的11个生成音乐系统中的成对人类偏好战斗,包括美学特征始终改善偏好预测,展示了跨生成架构的学习表示的强大泛化。
摘要:Music popularity prediction has attracted growing research interest, with relevance to artists, platforms, and recommendation systems. However, the explosive rise of AI-generated music platforms has created an entirely new and largely unexplored landscape, where a surge of songs is produced and consumed daily without the traditional markers of artist reputation or label backing. Key, yet unexplored in this pursuit is aesthetic quality. We propose APEX, the first large-scale multi-task learning framework for AI-generated music, trained on over 211k songs (10k hours of audio) from Suno and Udio, that jointly predicts engagement-based popularity signals - streams and likes scores - alongside five perceptual aesthetic quality dimensions from frozen audio embeddings extracted from MERT, a self-supervised music understanding model. Aesthetic quality and popularity capture complementary aspects of music that together prove valuable: in an out-of-distribution evaluation on the Music Arena dataset, comprising pairwise human preference battles across eleven generative music systems unseen during training, including aesthetic features consistently improves preference prediction, demonstrating strong generalisation of the learned representations across generative architectures.
【10】DECKER: Domain-invariant Embedding for Cross-Keyboard Extraction and Recognition
标题:DECKER:用于跨键盘提取和识别的域不变嵌入链接:https://arxiv.org/abs/2605.03384
备注:Accepted to AsiaCCS'26
摘要:键盘上的声学侧信道攻击(ASCA)构成了重大的安全风险,因为可以从键入声学中推断出干扰,从而泄露敏感信息。之前的ASCA研究受到小规模数据集的限制,用户,键盘和环境的多样性有限,限制了对设备,麦克风和噪声条件的分析。我们介绍了HEAR,一个数据集,旨在研究ASCA沿三个轴:键盘泛化,噪声适应和用户偏见。HEAR包含53名参与者使用37个笔记本电脑键盘的录音,这些录音在三种真实的设置中收集:(1)外部麦克风捕获,(2)无网络噪音的设备麦克风捕获,以及(3)基于VoIP的流媒体捕获。这可以实现跨用户、键盘和环境的受控评估。在HEAR上,我们建立了ASCA基准测试,涵盖单模式和多模式设置中原始音频和声谱图的传统特征和预训练表示。我们提出了DECKER,一个具有四个阶段的域不变的非线性推理框架:(1)键盘签名归一化以减少设备着色,(2)域对抗解纠缠以抑制键盘身份,(3)监督跨键盘对比对齐以加强键的一致性,以及(4)声学风格随机化以合成看不见的键盘响应。我们进一步探索使用基于LLM的后处理层通过语言上下文来细化序列的推理级别。HEAR上的结果显示,DECKER在强基线上提高了识别能力,特别是在跨键盘和跨用户设置中,并从语言模型校正中获得了进一步的收益。这些发现强调了ASCA在不同的用户,设备和嘈杂的环境中仍然有效,强调了其实际的安全风险。
摘要:Acoustic side-channel attacks (ASCA) on keyboards pose a significant security risk, as keystrokes can be inferred from typing acoustics, revealing sensitive information. Prior ASCA studies are limited by small-scale datasets with restricted diversity in users, keyboards, and environments, constraining analysis across devices, microphones, and noise conditions. We introduce HEAR, a dataset designed to study ASCA along three axes: keyboard generalization, noise adaptation, and user bias. HEAR contains recordings from 53 participants using 37 laptop keyboards, collected in three realistic settings: (1) external microphone capture, (2) device microphone capture without network noise, and (3) VoIP-based streaming capture. This enables controlled evaluation across users, keyboards, and environments. On HEAR, we establish an ASCA benchmark spanning conventional features and pre-trained representations from raw audio and spectrograms in unimodal and multimodal settings. We propose DECKER, a domain-invariant keystroke inference framework with four stages: (1) Keyboard Signature Normalization to reduce device coloration, (2) domain-adversarial disentanglement to suppress keyboard identity, (3) supervised cross-keyboard contrastive alignment to enforce key consistency, and (4) Acoustic Style Randomization to synthesize unseen keyboard responses. We further explore sentence-level inference using an LLM-based post-processing layer to refine keystroke sequences via linguistic context. Results on HEAR show DECKER improves keystroke identification over strong baselines, particularly in cross-keyboard and cross-user settings, with further gains from language-model rectification. These findings highlight that ASCA remains effective across diverse users, devices, and noisy environments, underscoring its practical security risk.
【11】Contrastive Regularization for Accent-Robust ASR
标题:口音鲁棒的ASB的对比正规化链接:https://arxiv.org/abs/2605.03297
摘要:基于自监督声学预训练和CTC微调的ASR系统在母语上实现了强大的性能,但对口音变化仍然敏感。我们调查监督对比学习(SupCon)作为一个轻量级的,口音不变的辅助目标CTC微调。话语级对比损失规则化编码器表示,而无需架构修改或明确的口音监督。在L2-ARCTIC基准测试上的实验表明,多个预训练编码器的WER降低一致,在unseen-accent评估下相对降低高达25 - 29%。使用内转录余弦分散的分析表明,SupCon促进更紧凑,更稳定的代表几何重音变化。总的来说,SupCon提供了一个有效的和模型无关的正则化策略,以提高口音的鲁棒性。
摘要:ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability. We investigate supervised contrastive learning (SupCon) as a lightweight, accent-invariant auxiliary objective for CTC fine-tuning. An utterance-level contrastive loss regularizes encoder representations without architectural modification or explicit accent supervision. Experiments on the L2-ARCTIC benchmark show consistent WER reductions across multiple pretrained encoders, with up to 25 -- 29\% relative reduction under unseen-accent evaluation. Analysis using within-transcript cosine dispersion indicates that SupCon promotes more compact and stable representation geometry under accent variability. Overall, SupCon provides an effective and model-agnostic regularization strategy for improving accent robustness.
【12】Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
标题:使用自我监督嵌入在情感状况中进行音素级Deepfake检测链接:https://arxiv.org/abs/2605.03079
备注:6 pages, 2 figures, submitted to IEEE SMC 2026
摘要:情感语音转换(EVC)的最新进展已经能够生成富有表现力的合成语音,这引起了音频深度伪造检测的新关注。现有的方法将语音视为同质信号,并在很大程度上忽略了其内部的语音结构,限制了它们在情绪条件下的可解释性。在这项工作中,我们提出了一个音素级的框架来分析情感操纵的合成语音使用真实的和EVC生成的语音匹配的情感条件下与共享的成绩单,音素对齐的TextGrids,和基于WavLM的嵌入。我们的研究结果表明,音素行为不同类别,复杂的元音和摩擦音表现出较高的分歧,而简单的音素保持更稳定。具有较大分布差异的音素也被发现更容易被检测到,在多种情绪和合成系统中保持一致。这些研究结果表明,音素级分析是一种有效的和可解释的方法来检测情绪操纵的合成语音。
摘要:Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook its internal phonetic structure, limiting their interpretability in emotionally conditioned settings. In this work, we propose a phoneme-level framework to analyze emotionally manipulated synthetic speech using real and EVC-generated speech under matched emotional conditions with shared transcripts, phoneme-aligned TextGrids, and WavLM-based embeddings. Our results show that phoneme behavior varies across categories, with complex vowels and fricatives exhibiting higher divergence while simpler phonemes remain more stable. Phonemes with larger distributional differences are also found to be more easily detected, consistently across multiple emotions and synthesis systems. These findings demonstrate that phoneme-level analysis is an effective and interpretable approach for detecting emotionally manipulated synthetic speech.
【13】The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail
标题:TTS-STT Flywheel:合成超密音频缩小了商业和开源系统失败的印度ASB差距链接:https://arxiv.org/abs/2605.03073
备注:8 pages, 2 figures. Companion to arXiv:2604.25441 (Praxy Voice TTS), arXiv:2604.25476 (PSP), arXiv:2605.00777 (LASE)
摘要:利基领域的印度语ASR --数字字符串、货币金额、地址、品牌名称、英语/印度语代码混合--在开源SOTA和商业系统中都没有得到充分的服务。在合成的实体密集泰卢固语测试集(由合成系统支持)上,vasista 22/whisper-telugu-large-v2(开放SOTA)实现了安全命中率(EHR)0.027和Deepgram Nova-3(商业)0.16。我们通过一个独立的TTS STT飞轮缩小了这一差距<->:一个开源的Indic TTS管道以<$50的边际成本合成了~ 22,000个实体密集的Indic-English代码混合话语,并且在vasista 22之上的LoRA微调在保持测试上达到了EHR 0.473(17倍于开放SOTA,3倍于商业),其中在FLEURS-Te上阅读散文回归限制为+6.6ppWER。跨语言:beta-Hi 0.337(7 x vs vasista 22)和beta-Ta 0.543(22 x vs vasista 22,22 x vs Deepgram);在印地语中,Deepgram具有大量的实体覆盖率,飞轮的表现低于商业。所有三个测试模型都低于预先注册的EHR目标(Te为0.75,Hi/Ta为0.65);我们诚实地报告。一个本地人记录的健全检查(n=20泰卢固语)确认转移到真实的语音(beta-Te EHR 0.516对本地人和合成器上的0.473)。EDSA隔离消融(仅FLEURS-Te上的LoRA)在相同的保持时间上产生EHR 0.020,将~100%的增益归因于EDSA语料库。我们还报告了一个语言条件发现:vanilla Whisper-large-v3具有泰卢固语特定的脚本崩溃(SFR 0.46-0.71),每种语言LoRA纠正(SFR 0.81-0.97),但该配方在印地语和泰米尔语上是禁忌的,其中vanilla SFR >= 0.98。代码,holdouts,预测,EDSA语料库和实体字典都是开源的。
摘要:Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and commercial systems. On a synthesised entity-dense Telugu test set (held-out by synthesis system), vasista22/whisper-telugu-large-v2 (open SOTA) achieves Entity-Hit-Rate (EHR) 0.027 and Deepgram Nova-3 (commercial) 0.16. We close this gap with a self-contained TTS<->STT flywheel: an open-source Indic TTS pipeline synthesises ~22,000 entity-dense Indic-English code-mix utterances at <$50 marginal cost, and a LoRA fine-tune on top of vasista22 achieves EHR 0.473 on the held-out test (17x over open SOTA, 3x over commercial), with read-prose regression bounded to +6.6 pp WER on FLEURS-Te. Cross-language: beta-Hi 0.337 (7x vs vasista22) and beta-Ta 0.543 (22x vs vasista22, 22x vs Deepgram); on Hindi where Deepgram has substantial entity coverage, the flywheel underperforms commercial. All three beta models fall below pre-registered EHR targets (0.75 for Te, 0.65 for Hi/Ta); we report honestly. A native-human-recorded sanity check (n=20 Telugu) confirms transfer to real speech (beta-Te EHR 0.516 on native vs 0.473 on synth). An EDSA-isolation ablation (LoRA on FLEURS-Te alone) yields EHR 0.020 on the same held-out, attributing ~100% of the gain to the EDSA corpus. We additionally report a language-conditional finding: vanilla Whisper-large-v3 has Telugu-specific Script Collapse (SFR 0.46-0.71) that a per-language LoRA corrects (SFR 0.81-0.97), but the recipe is contraindicated on Hindi and Tamil where vanilla SFR >= 0.98. Code, holdouts, predictions, EDSA corpus, and entity dictionaries are released open-source.
【14】Mixed-Precision Information Bottlenecks for On-Device Trait-State Disentanglement in Bipolar Agitation Detection
标题:双极搅动检测中设备上特征状态解纠缠的混合精度信息瓶颈链接:https://arxiv.org/abs/2605.03039
摘要:通过声音生物标志物连续监测双相情感障碍激动需要在资源受限的边缘设备上将稳定的说话者特征与波动的情感状态分离开来。我们介绍MP-IB,第一个框架来处理混合精度量化作为临床特征状态分离的信息瓶颈。核心观点是数字精度本身控制容量:FP 16特征头(1,024位)编码说话者身份,而INT 4状态头(128位)捕获激动,产生8倍信息不对称,无需对抗训练。我们增加了动态精度调度和多尺度时间融合。关于Bridge 2AI-Voice(N=833,4个会话/参与者,严格的说话者独立CV),MP-IB达到rho = 0.117(95% CI:[0.089,0.145],p=0.003 vs. chance),优于具有域内SSL延续的94 M参数WavLM适配器(rho = -0.042)、β VAE解开(rho = 0.089)和手工韵律(rho = 0.031)的绝对值为2.8- 15.9点。Zero-shot转移到CREMA-D实现AUC=0.817。同一性泄漏被抑制至接近随机(EER=0.42,MIA-AUC=0.52)。端到端延迟为23.4 ms,占用空间为617 KB,可对低于20美元的设备进行实时监控。
摘要:Continuous monitoring of bipolar disorder agitation via voice biomarkers requires disentangling stable speaker traits from volatile affective states on resource-constrained edge devices. We introduce MP-IB, the first framework to treat mixed-precision quantization as an information bottleneck for clinical trait-state separation. The core insight is that numerical precision itself controls capacity: an FP16 trait head (1,024 bits) encodes speaker identity, while an INT4 state head (128 bits) captures agitation, yielding 8x information asymmetry without adversarial training. We augment this with Dynamic Precision Scheduling and Multi-Scale Temporal Fusion. On Bridge2AI-Voice (N=833, 4 sessions/participant, strict speaker-independent CV), MP-IB achieves rho = 0.117 (95\% CI: [0.089, 0.145], p=0.003 vs. chance), outperforming 94M-parameter WavLM-Adapter with in-domain SSL continuation (rho = -0.042), beta VAE disentanglement (rho = 0.089), and hand-crafted prosody (rho = 0.031) by 2.8--15.9 points absolute. Zero-shot transfer to CREMA-D achieves AUC=0.817. Identity leakage is suppressed to near-random (EER=0.42, MIA-AUC=0.52). End-to-end latency is 23.4 ms with a 617 KB footprint, enabling real-time monitoring on sub 20 dollar devices.
【15】AsymK-Talker: Real-Time and Long-Horizon Talking Head Generation via Asymmetric Kernel Distillation
标题:AsymK-Talker:通过不对称核蒸馏产生实时和长视野说话的头部链接:https://arxiv.org/abs/2605.02948
摘要:扩散模型的最新进展显着提高了音频驱动的说话头生成的视觉保真度。然而,现有的方法受到三个关键的限制:因果效率低下,阻碍实时推理,时间相干条件不兼容,并逐步漂移在长时间生成,共同阻碍他们的部署在实时应用。为了克服这些挑战,我们介绍了AsymK-Talker,一种新的扩散蒸馏方法,专为实时和长期的谈话头生成。AsymK-Talker包括三个关键组件:(1)Kernel-Conditioned Loop Generation(KCLG),一种因果的、分块的生成范式,利用运动核来实现时间一致的传播;(2)时间参考编码(TRE),将静态身份参考转换为时间感知的潜在表示,以增强视听同步;以及(3)非对称核蒸馏(AKD),教师-学生蒸馏框架,其中教师模型对用于监督的地面实况运动核进行调节,而学生学习从所生成的核生成,从而确保在扩展的生成序列期间的鲁棒性。AsymK-Talker在视觉保真度和嘴唇同步指标上都取得了令人满意的结果。
摘要:Recent advances in diffusion models have markedly enhanced the visual fidelity of audio-driven talking head generation. Nevertheless, existing methods are constrained by three critical limitations: causal inefficiency that impedes real-time inference, incompatibility with temporally coherent conditioning, and progressive drift over long-horizon generation, collectively hindering their deployment in real-time applications. To overcome these challenges, we introduce AsymK-Talker, a novel diffusion-distillation method designed for real-time and long-horizon talking head generation. AsymK-Talker comprises three key components: (1) Kernel-Conditioned Loop Generation (KCLG), a causal, chunk-wise generation paradigm that leverages motion kernels to enable temporally consistent propagation; (2) Temporal Reference Encoding (TRE), which converts a static identity reference into a time-aware latent representation to enhance audio-visual synchronization; and (3) Asymmetric Kernel Distillation (AKD), a teacher-student distillation framework wherein the teacher model conditions on ground-truth motion kernels for supervision, while the student learns to generate from generated kernels, thereby ensuring robustness during extended generation sequences. AsymK-Talker achieves promising results on both visual fidelity and lip synchronization metrics.
【16】Keyword spotting using convolutional neural network for speech recognition in Hindi
标题:使用卷积神经网络进行印地语语音识别的关键词定位链接:https://arxiv.org/abs/2605.02928
备注:Published in 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT)
摘要:在这项研究中,我们研究了关键字定位(KWS)在印地语语音识别领域的应用,利用包括40,000个音频样本的数据集。在44 kHz的采样率和每个样本平均持续时间为1.9秒的情况下,我们专注于开发一个针对用户特定查询的高效设备上KWS系统。利用卷积神经网络(CNN)进行分类,我们采用特征工程技术将原始音频记录转换为Mel频率倒谱系数(MFCC)作为我们网络的输入。我们的实验包括各种CNN架构,探索其在识别连续语音流中的预定义关键字的功效。通过严格的评估,我们基于CNN的方法实现了91.79%的准确率,表现出良好的性能,同时确保了印地语语音识别的计算效率和用户特定的定制。
摘要:In this study, we investigate the application of keyword spotting (KWS) in the domain of Hindi speech recognition, utilizing a dataset comprising 40,000 audio samples. With a sampling rate of 44 kHz and an average duration of 1.9 seconds per sample, we focus on developing an efficient on-device KWS system tailored for user-specific queries. Leveraging Convolutional Neural Networks (CNNs) for classification, we employ feature engineering techniques to convert raw audio recordings into Mel Frequency Cepstral Coefficients (MFCCs) as an input for our network. Our experiments encompass various CNN architectures, exploring their efficacy in identifying predefined keywords within the continuous speech stream. Our CNN-based approach achieves a commendable accuracy rate of 91.79% through rigorous evaluation, demonstrating promising performance while ensuring computational efficiency and user-specific customization in Hindi speech recognition.
标题:评估噪音和语音增强对语音编解码器可理解度的影响
链接:https://arxiv.org/abs/2605.03776
备注:submitted to Interspeech 2026
摘要:保持语音清晰度是通信中对语音编解码器的最低要求。最近,非常低比特率的神经编解码器已经获得了替代经典编解码器的兴趣,加强了评估在现实场景中是否保留可懂度的需要。在本文中,我们评估的清晰度和听力努力的经典和神经语音编解码器在干净和嘈杂的条件。此外,我们在编码之前评估语音增强(SE)的影响,模拟可能的音频处理管道。结果表明,经典编解码器比神经编解码器具有更好的抗噪声能力。此外,SE可以导致编解码器的显著可懂度和收听努力的改善,否则会受到噪声的负面影响。当可懂度饱和时,倾听努力揭示了细微的差异。最后,基于自动语音识别的客观可懂度与每个条件平均的主观可懂度分数高度相关。
摘要:Preserving speech intelligibility is a minimum requirement for speech codecs in communication. Recently, very low-bitrate neural codecs have gained interest for replacing classical codecs, reinforcing the need to evaluate whether intelligibility is preserved in realistic scenarios. In this paper, we evaluate the intelligibility and listening effort of classical and neural speech codecs in clean and noisy conditions. Further, we assess the impact of speech enhancement (SE) before coding, simulating a possible audio processing pipeline. The results show that classical codecs are more noise robust than neural codecs. Further, SE can lead to significant intelligibility and listening effort improvements for codecs otherwise negatively affected by noise. Listening effort reveals nuanced differences when intelligibility is saturated. Lastly, objective intelligibility based on automatic speech recognition is highly correlated with subjective intelligibility scores averaged per condition.
【2】MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
标题:MiniMind-O技术报告:开放式小规模语音原生全方位模型链接:https://arxiv.org/abs/2605.03937
备注:17 pages. Code, checkpoints, and training data are available at https://github.com/jingyaogong/minimind-o
摘要:MiniMind-O是一个基于MiniMind语言模型的开放的0. 1B级omni模型。它接受文本、语音和图像输入,并返回文本和流式语音。该版本包括模型代码,检查点和用于文本到音频,图像到文本和音频到音频训练的主要Parquet训练数据集,使完整的交互循环直接可检查。该模型使用完整的MiniMind主干作为思想者,并使用由MiniMind块制成的独立四层Talker。Frozen SenseVoice-Small和SigLIP 2编码器提供语音和图像特征,这些特征由轻量级MLP投影仪映射并注入模态占位符位置。Talker读取中间层Thinker状态以及自回归八层Mimi代码缓冲区。扬声器控制由专用扬声器令牌、右对齐参考编解码器提示和预先计算的CAM++扬声器嵌入处理,因此语音调节仍然是音频代码上下文的一部分,而不是单独的TTS模块。对于768维Talker,密集和MoE变体在Thinker-Talker一致性评价中达到0.0897和0.0900的平均CER,总体语音克隆相似性为0.5995和0.5937。除了报告一个工作系统之外,本文还确定了小型omni模型的三个关键设计选择:中间层语义桥接,发布的多模态序列格式和参数高效的八码本接口。
摘要:MiniMind-O is an open 0.1B-scale omni model built on the MiniMind language model. It accepts text, speech, and image inputs, and returns both text and streaming speech. The release includes model code, checkpoints, and the main Parquet training datasets for text-to-audio, image-to-text, and audio-to-audio training, making the complete interaction loop directly inspectable. The model uses a full MiniMind backbone as the Thinker and an independent four-layer Talker made from MiniMind blocks. Frozen SenseVoice-Small and SigLIP2 encoders provide speech and image features, which are mapped by lightweight MLP projectors and injected at modality-placeholder positions. The Talker reads a middle-layer Thinker state together with an autoregressive eight-layer Mimi-code buffer. Speaker control is handled by a dedicated speaker token, right-aligned reference codec prompts, and precomputed CAM++ speaker embeddings, so voice conditioning remains part of the audio-code context rather than a separate TTS module. With a 768-dimensional Talker, the dense and MoE variants reach average CERs of 0.0897 and 0.0900 in Thinker--Talker consistency evaluation, with overall voice-cloning similarities of 0.5995 and 0.5937. Beyond reporting a working system, the paper identifies three scale-critical design choices for small omni models: middle-layer semantic bridging, a released multimodal sequence format, and a parameter-efficient eight-codebook interface.
【3】Phoneme-Level Deepfake Detection Across Emotional Conditions Using Self-Supervised Embeddings
标题:使用自我监督嵌入在情感状况中进行音素级Deepfake检测链接:https://arxiv.org/abs/2605.03079
备注:6 pages, 2 figures, submitted to IEEE SMC 2026
摘要:情感语音转换(EVC)的最新进展已经能够生成富有表现力的合成语音,这引起了音频深度伪造检测的新关注。现有的方法将语音视为同质信号,并在很大程度上忽略了其内部的语音结构,限制了它们在情绪条件下的可解释性。在这项工作中,我们提出了一个音素级的框架来分析情感操纵的合成语音使用真实的和EVC生成的语音匹配的情感条件下与共享的成绩单,音素对齐的TextGrids,和基于WavLM的嵌入。我们的研究结果表明,音素行为不同类别,复杂的元音和摩擦音表现出较高的分歧,而简单的音素保持更稳定。具有较大分布差异的音素也被发现更容易被检测到,在多种情绪和合成系统中保持一致。这些研究结果表明,音素级分析是一种有效的和可解释的方法来检测情绪操纵的合成语音。
摘要:Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook its internal phonetic structure, limiting their interpretability in emotionally conditioned settings. In this work, we propose a phoneme-level framework to analyze emotionally manipulated synthetic speech using real and EVC-generated speech under matched emotional conditions with shared transcripts, phoneme-aligned TextGrids, and WavLM-based embeddings. Our results show that phoneme behavior varies across categories, with complex vowels and fricatives exhibiting higher divergence while simpler phonemes remain more stable. Phonemes with larger distributional differences are also found to be more easily detected, consistently across multiple emotions and synthesis systems. These findings demonstrate that phoneme-level analysis is an effective and interpretable approach for detecting emotionally manipulated synthetic speech.
机器翻译由腾讯交互翻译提供,仅供参考
