微信公众号:arXiv_Daily
cs.SD语音
标题:无崩溃的私人语音分类:稳定的DP训练和离线蒸馏
链接:https://arxiv.org/abs/2605.02718
摘要:We study example-level private supervised speech classification under a practical release constraint: training may access privileged side information, but the released model must be audio-only. This setting is important because speech systems can often exploit richer side information during development, whereas deployment and release require a lightweight unimodal model with auditable privacy guarantees. Using DP-SGD on the private dataset $D_{\text{priv}}$, we identify a strong-privacy failure mode ($ε\le 1$) on imbalanced tasks, where training may collapse to a near single-class predictor, a phenomenon that overall accuracy can obscure. We therefore emphasize Macro-F1, balanced accuracy, and a simple collapse diagnostic. This failure is especially problematic in our release setting because a collapsed private teacher cannot provide useful supervision for the downstream audio-only student. To address this setting under strong privacy, we propose a two-stage protocol: (i) train a (possibly multimodal) DP teacher on $D_{\text{priv}}$, and (ii) distill an audio-only student on a fixed, recording-disjoint auxiliary dataset $D_{\text{aux}}$ using one-shot offline teacher probability outputs, releasing only the student. The DP guarantee applies only to $D_{\text{priv}}$; we make no DP claim for $D_{\text{aux}}$, and privacy of the released student with respect to $D_{\text{priv}}$ follows by post-processing. We frame this setting as involving four coupled bottlenecks: speech-induced optimization instability under DP-SGD, minority-class erosion under clipping and noise, teacher over-reliance on privileged modalities unavailable at deployment, and train--deploy modality mismatch. We address them with a DP-stabilizing acoustic front-end (DSAF), minibatch-adaptive bounded loss reweighting (AW-DP), privileged-modality dropout, and offline teacher-to-student distillation.
摘要:We study example-level private supervised speech classification under a practical release constraint: training may access privileged side information, but the released model must be audio-only. This setting is important because speech systems can often exploit richer side information during development, whereas deployment and release require a lightweight unimodal model with auditable privacy guarantees. Using DP-SGD on the private dataset $D_{\text{priv}}$, we identify a strong-privacy failure mode ($ε\le 1$) on imbalanced tasks, where training may collapse to a near single-class predictor, a phenomenon that overall accuracy can obscure. We therefore emphasize Macro-F1, balanced accuracy, and a simple collapse diagnostic. This failure is especially problematic in our release setting because a collapsed private teacher cannot provide useful supervision for the downstream audio-only student. To address this setting under strong privacy, we propose a two-stage protocol: (i) train a (possibly multimodal) DP teacher on $D_{\text{priv}}$, and (ii) distill an audio-only student on a fixed, recording-disjoint auxiliary dataset $D_{\text{aux}}$ using one-shot offline teacher probability outputs, releasing only the student. The DP guarantee applies only to $D_{\text{priv}}$; we make no DP claim for $D_{\text{aux}}$, and privacy of the released student with respect to $D_{\text{priv}}$ follows by post-processing. We frame this setting as involving four coupled bottlenecks: speech-induced optimization instability under DP-SGD, minority-class erosion under clipping and noise, teacher over-reliance on privileged modalities unavailable at deployment, and train--deploy modality mismatch. We address them with a DP-stabilizing acoustic front-end (DSAF), minibatch-adaptive bounded loss reweighting (AW-DP), privileged-modality dropout, and offline teacher-to-student distillation.
【2】Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation
标题:Tibetan-TTC:具有大模型适应的低资源西藏语音合成链接:https://arxiv.org/abs/2605.02496
摘要:藏语文语转换长期以来受到语音资源匮乏、方言差异大、书面语与口语语音映射复杂等问题的挑战。为了解决这些问题,这项工作提出了,据我们所知,第一个大模型为基础的藏语TTS系统在行业内,建立在一个大的语音合成模型由星辰AGI实验室开发。该系统集成了数据质量增强,面向藏语的文本表示和标记器自适应,以及跨语言自适应训练低资源藏语语音合成。实验结果表明,该系统能够在低资源条件下生成稳定、自然、清晰的藏语语音。在主观评价中,音节级和基于BPE的系统的MOS得分分别达到4.28和4.35,发音准确率分别达到97.6%和96.6%,优于外部商业藏语TTS接口。这些结果表明,将大模型骨干与面向藏语的文本表示自适应和跨语言自适应训练相结合,可以实现高可用性的低资源藏语语音合成,也为未来统一的多方言藏语语音合成提供了技术基础。
摘要:Tibetan text-to-speech (TTS) has long been challenged by scarce speech resources, significant dialectal variation, and the complex mapping between written text and spoken pronunciation. To address these issues, this work presents, to the best of our knowledge, the first large-model-based Tibetan TTS system in the industry, built upon a large speech synthesis model developed by Xingchen AGI Lab. The proposed system integrates data quality enhancement, Tibetan-oriented text representation and tokenizer adaptation, and cross-lingual adaptive training for low-resource Tibetan speech synthesis. Experimental results show that the system can generate stable, natural, and intelligible Tibetan speech under low-resource conditions. In subjective evaluation, the MOS scores of the syllable-level and BPE-based systems reach 4.28 and 4.35, while their pronunciation accuracies reach 97.6% and 96.6%, respectively, outperforming an external commercial Tibetan TTS interface. These results demonstrate that combining a large-model backbone with Tibetan-oriented text representation adaptation and cross-lingual adaptive training enables highly usable low-resource Tibetan speech synthesis, and also provides a technical foundation for future unified multi-dialect Tibetan speech synthesis.
【3】Toward Fine-Grained Speech Inpainting Forensics:A Dataset, Method, and Metric for Multi-Region Tampering Localization
标题:迈向细粒度语音修复取证:用于多区域篡改定位的数据集、方法和指标链接:https://arxiv.org/abs/2605.02223
摘要:语音克隆和文本到语音合成的最新进展使得部分语音操纵-对手在保留说话者身份的同时替换话语中的几个单词以改变其含义-成为越来越现实的威胁。现有的音频deepfake检测基准专注于话语级二进制分类或单区域篡改,在检测和定位多个修复片段(其计数先验未知)方面存在关键差距。我们通过三个贡献来解决这一差距。首先,我们介绍了MIST(多区域修复语音篡改),这是一个跨6种语言的大规模多语言数据集,每个话语有1-3个独立修复的单词级片段,通过LLM引导的语义替换和神经语音克隆生成,其中虚假内容仅占每个话语的2-7%。其次,我们提出了ISA(迭代分段分析),一个骨干不可知的框架,执行粗到细的滑动窗口分类与间隙容忍区域的建议和边界细化恢复所有篡改的区域,而无需事先知道他们的计数。第三,我们定义了SF1@tau,这是一个基于时间IoU匹配的分段级F1度量,它联合评估区域计数精度和定位精度。Zero-shot评估显示,现有的深度伪造检测器仍然无法解决单词粒度的部分修复:在完全合成的语音上训练的话语级分类器将接近零的伪造概率分配给MIST话语,其中只有2-7%的内容被操纵。ISA在这种具有挑战性的环境中始终优于非迭代基线,并且数据集,代码和评估工具包都是公开发布的。
摘要:Recent advances in voice cloning and text-to-speech synthesis have made partial speech manipulation - where an adversary replaces a few words within an utterance to alter its meaning while preserving the speaker's identity - an increasingly realistic threat. Existing audio deepfake detection benchmarks focus on utterance-level binary classification or single-region tampering, leaving a critical gap in detecting and localizing multiple inpainted segments whose count is unknown a priori. We address this gap with three contributions. First, we introduce MIST (Multiregion Inpainting Speech Tampering), a large-scale multilingual dataset spanning 6 languages with 1-3 independently inpainted word-level segments per utterance, generated via LLM-guided semantic replacement and neural voice cloning, with fake content constituting only 2-7% of each utterance. Second, we propose ISA (Iterative Segment Analysis), a backbone-agnostic framework that performs coarse-to-fine sliding-window classification with gap-tolerant region proposal and boundary refinement to recover all tampered regions without prior knowledge of their count. Third, we define SF1@tau, a segment-level F1 metric based on temporal IoU matching that jointly evaluates region count accuracy and localization precision. Zero-shot evaluation reveals that partial inpainting at word granularity remains unsolved by existing deepfake detectors: utterance-level classifiers trained on fully synthesized speech assign near zero fake probability to MIST utterances where only 2-7% of content is manipulated. ISA consistently outperforms non-iterative baselines in this challenging setting, and the dataset, code, and evaluation toolkit are publicly released.
【4】RenCon 2025: Revival of the Expressive Performance Rendering Competition
标题:RenCon 2025:表现力表演渲染大赛的复兴链接:https://arxiv.org/abs/2605.02059
备注:Accepted at NIME 2026
摘要:本文介绍了RenCon 2025的全面文档,RenCon 2025是在韩国大田举行的ISMIR 2025上举行的表现力渲染比赛的复兴。比赛吸引了来自国际研究团体的9个参赛作品,代表了表达性钢琴演奏渲染的不同方法。两阶段评估结构包括初步在线评估和会议现场实时呈现。我们分析了竞争的形式,参与者的人口统计数据,系统性能,并为未来的迭代经验教训。研究结果表明,在表现力渲染能力的显着进步,同时突出了在实现人类水平的音乐表达仍然存在的挑战。
摘要:This paper presents a comprehensive documentation of RenCon 2025, the revival of the expressive performance rendering competition which took place at ISMIR 2025 in Daejeon, Korea. The competition attracted 9 entries from international research groups, representing diverse approaches to expressive piano performance rendering. The two-phase assessment structure comprised a preliminary online evaluation and live real-time rendering at the conference. We analyze the competition format, participant demographics, system performance, and lessons learned for future iterations. The results demonstrate significant advances in expressive rendering capabilities while highlighting remaining challenges in achieving human-level musical expression.
【5】Spoken Language Identification with Pre-trained Models and Margin Loss
标题:使用预训练模型和保证金损失的口语识别链接:https://arxiv.org/abs/2605.01905
备注:Technical report for the TidyLang 2026 Challenge. Accepted at Odyssey 2026
摘要:针对TidyLang Challenge 2026中提出的说话人控制的口语识别任务,本文提出了一种基于预训练模型和基于边缘损失的语言识别方法。该方法采用预训练的ECAPA-TDNN作为特征编码器,并结合基于边缘的损失来增强语言表示的区分能力,从而提高类间可分性并减少非语言因素(如说话人特征)的干扰。在Tidy-X数据集上的实验结果表明,该方法在语言识别任务上达到了85.95%的宏观准确率和90.96%的微观准确率,在验证任务上达到了17.08%的等误率。与官方基线相比,宏观精度提高了45.7%,微观精度提高了15.2%,EER降低了约50.8%,证明了该方法的有效性。该代码将在https://github.com/PunkMale/TidyLang2026上发布。
摘要:For the speaker-controlled spoken language identification task proposed in the TidyLang Challenge 2026, this paper proposes a language identification method based on pre-trained models and margin-based losses. The proposed method adopts a pre-trained ECAPA-TDNN as the feature encoder and incorporates margin-based losses to enhance the discriminative ability of language representations, thereby improving inter-class separability and reducing the interference of non-linguistic factors such as speaker characteristics. Experimental results on the Tidy-X dataset show that the proposed method achieves 85.95% macro accuracy and 90.96% micro accuracy on the language identification task and 17.08% equal error rate (EER) on the verification task. Compared with the official baseline, the macro accuracy improves by 45.7%, the micro accuracy improves by 15.2%, and the EER is reduced by approximately 50.8%, demonstrating the effectiveness of the proposed method. The code will be released at https://github.com/PunkMale/TidyLang2026.
【6】TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation
标题:TMD-Bench:音乐与舞蹈共生的多层次评估范式链接:https://arxiv.org/abs/2605.01809
摘要:统一的视听生成正在迅速获得工业和创造性的相关性,使虚拟制作和互动媒体的应用成为可能。然而,当从一般的音频-视频合成移动到音乐-舞蹈联合生成时,任务变得更加困难:音乐节奏,措辞和口音必须以精细的时间分辨率驱动舞蹈动作,并且这种节奏耦合不被当前评估实践中使用的单峰度量或通用视听一致性分数捕获。我们介绍了TMD-Bench,一个文本驱动的音乐舞蹈联合生成的基准,评估系统的单峰生成质量,指令遵守和跨模态的节奏对齐。该基准将可计算的物理指标与感知的多模态判断相结合,并由一个精心策划的节奏对齐的音乐舞蹈数据集和一个用于结构化音乐语义的细粒度音乐字幕支持。TMD-Bench进一步揭示了(i)现代商业视听模型,如Veo 3和Sora 2,产生高质量的音乐和视频,而节奏耦合仍然不太一致,并留下改进的空间,(ii)我们在节奏对齐数据上训练的统一基线RhyJAM实现了有竞争力的节拍级同步,同时保持了有竞争力的单峰保真度。这为构建下一代音乐舞蹈模型提供了前景,这些模型明确优化了节奏和动力学连贯性。
摘要:Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation, the task becomes substantially harder: musical rhythm, phrasing, and accents must drive choreographic motion at fine temporal resolution, and such rhythmic coupling is not captured by unimodal metrics or generic audiovisual consistency scores used in current evaluation practice. We introduce TMD-Bench, a benchmark for text-driven music-dance co-generation that assesses systems across unimodal generation quality, instruction adherence, and cross-modal rhythmic alignment. The benchmark integrates computable physical metrics with perceptual multimodal judgments, and is supported by a curated rhythm-aligned music-dance dataset and a fine-grained Music Captioner for structured music semantics. TMD-Bench further reveals that (i) modern commercial audio-visual models, such as Veo 3 and Sora 2, produce high-quality music and video, while rhythmic coupling remains less consistently optimized and leaves room for improvement, and (ii) our unified baseline RhyJAM trained on rhythm-aligned data achieves competitive beat-level synchronization while maintaining competitive unimodal fidelity. This presents prospects for building next-generation music-dance models that explicitly optimize rhythmic and kinetic coherence.
【7】Khala: Scaling Acoustic Token Language Models Toward High-Fidelity Music Generation
标题:Khala:将声学代币语言模型扩展为高保真音乐生成链接:https://arxiv.org/abs/2605.01790
摘要:高质量音乐生成中的一个常见设计模式是在不同的表示空间中处理结构和保真度:生成器首先对高级结构进行建模,然后是基于扩散或神经的解码阶段,重建精细细节。在这项工作中,我们探索另一种观点:两者都可以在一个单一的深层声学令牌层次结构中逐步建模。为了研究这一点,我们建立了一个64层的残差矢量量化(RVQ)的声学表示,并提出了一个两阶段的粗到精的生成框架。骨干模型首先生成完整音轨的粗声学标记,然后超分辨率模型在相同的声学标记空间内完成更精细的标记。超分辨率阶段以全轨迹规模工作,并在一段时间内并行运行时逐层细化令牌,从而形成固定的62步推理过程。为了共同改善歌词对齐和细节重建,我们进一步引入了混合注意力训练:对齐目标使用因果注意力,而逐层细化使用完全注意力。一个关键的发现是,文本-语音对齐可以出现在纯声学令牌语言建模,而不需要一个单独的语义令牌阶段。此外,从训练的骨干初始化超分辨率模型显着提高了收敛性和最终质量。两者合计,我们的研究结果表明,高品质的音乐生成可以有效地追求不分离的结构和保真度到异构的表示空间。相反,两者都可以在统一的声学令牌层次结构中逐步建模,指向更简单,更统一的高质量音乐生成路径。
摘要:A common design pattern in high-quality music generation is to handle structure and fidelity in different representation spaces: a generator first models high-level structure, followed by diffusion-based or neural decoding stages that reconstruct fine details. In this work, we explore an alternative view: both may be progressively modeled within a single deep acoustic-token hierarchy. To study this, we build a 64-layer residual vector quantization (RVQ) acoustic representation and propose a two-stage coarse-to-fine generation framework. A backbone model first generates coarse acoustic tokens for the full track, and a super-resolution model then completes finer tokens within the same acoustic token space. The super-resolution stage works at full-track scale and refines tokens layer by layer while running in parallel over time, leading to a fixed 62-step inference process. To jointly improve lyric alignment and fine-detail reconstruction, we further introduce hybrid-attention training: the alignment objective uses causal attention, while layer-wise refinement uses full attention. A key finding is that text--vocal alignment can emerge within pure acoustic-token language modeling, without requiring a separate semantic token stage. Moreover, initializing the super-resolution model from the trained backbone significantly improves convergence and final quality. Taken together, our results suggest that high-quality music generation can be effectively pursued without separating structure and fidelity into heterogeneous representation spaces. Instead, both can be progressively modeled within a unified acoustic-token hierarchy, pointing toward a simpler and more unified path to high-quality music generation.
【8】Delayed Commitment for Representation Readiness in Stage-wise Audio-Visual Learning
标题:阶段视听学习中的表示准备延迟承诺链接:https://arxiv.org/abs/2605.01673
摘要:逐阶段视听编码器跨层传播融合的中间状态,使得稍后表示的形成取决于早期融合状态的准备情况。强大的本地视听协议提供了有用的对应证据,但融合状态还需要足够的跨层和跨模态支持,才能可靠地指导后续融合。本文研究了这个问题,通过传播意识的表示准备和制定过早的知觉承诺作为一个准备不足的问题,其中本地可扩展性,传播的影响,和支持不足共同出现在一个中间阶段。我们提出了延迟感知承诺网络(DPC-Net),一个编码器级的框架,估计一个可观察到的准备不足的代理,本地化的干预敏感的瓶颈,并应用支持感知校正与跨层和跨模态的证据。DPC-Net保留了特定于任务的头、损失、解码模块和评估协议,使其通过编码器端干预适用于不同的视听任务。视听语音分离,视听事件定位,视听语音识别的实验表明,在重建,定位和识别制度的一致改善。对组件贡献、选择标准、反事实干预和就绪轨迹的进一步分析支持就绪引导的瓶颈校正的有效性。
摘要:Stage-wise audio-visual encoders propagate fused intermediate states across layers, making the formation of later representations depend on the readiness of earlier fusion states. Strong local audio-visual agreement provides useful correspondence evidence, yet a fused state also needs sufficient cross-layer and cross-modal support before it can reliably guide later fusion. This paper studies this issue through propagation-aware representation readiness and formulates premature perceptual commitment as a readiness-deficiency problem, where local plausibility, propagation influence, and support insufficiency jointly appear at an intermediate stage. We propose the Delayed Perceptual Commitment Network (DPC-Net), an encoder-level framework that estimates an observable readiness-deficiency surrogate, localizes the intervention-sensitive bottleneck, and applies support-aware correction with cross-layer and cross-modal evidence. DPC-Net preserves task-specific heads, losses, decoding modules, and evaluation protocols, making it applicable to different audio-visual tasks through encoder-side intervention. Experiments on audio-visual speech separation, audio-visual event localization, and audio-visual speech recognition show consistent improvements across reconstruction, localization, and recognition regimes. Further analyses on component contribution, selection criteria, counterfactual intervention, and readiness trajectories support the effectiveness of readiness-guided bottleneck correction.
【9】MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech
标题:MelShield:用于人工智能生成的合成语音起源归因的鲁棒Mel-域音频水印链接:https://arxiv.org/abs/2605.01515
备注:Accepted by ACISP 2026
摘要:在本文中,我们提出了MelShield,这是一个强大的,代内的,关键的音频水印框架,它将可识别的信号嵌入到AI生成的音频中,以实现版权保护和可靠的归属。具体而言,MelShield在生成过程中在Mel频谱图域中操作,以Mel条件管道中的中间声学表示为目标,用于文本到语音(TTS)生成。其核心思想是将中间梅尔频谱图视为宿主信号,并在波形合成之前通过分布在精心选择的时间-频率区域的低能量、键控扩频扰动来嵌入短的二进制有效载荷。通过在声码器推理之前执行水印,MelShield保持了Mel条件TTS架构的即插即用,并且不需要修改或重新训练底层TTS生成声码器,例如DiffWave和HiFi-GAN。此外,多用户键控构造实现了可扩展的用户特定属性,而键控验证机制限制了未经授权的解码,从而降低了大规模提取器探测和对抗性分析的风险。在DiffWave和HiFi-GAN上的大量实验表明,MelShield实现了可靠的水印提取,接近100\%的比特精度,即使在信号失真的情况下,例如,压缩和加性噪声,同时保持高感知音频质量。
摘要:In this paper, we propose MelShield, a robust, in-generation, keyed audio watermarking framework that embeds identifiable signals into AI-generated audio for copyright protection and reliable attribution. Specifically, MelShield operates in the Mel-spectrogram domain during the generation process, targeting intermediate acoustic representations in Mel-conditioned pipelines for text-to-speech (TTS) generation. The core idea is to treat the intermediate Mel-spectrogram as the host signal and embed a short binary payload via low-energy, keyed spread-spectrum perturbations distributed across carefully selected time-frequency regions prior to waveform synthesis. By performing watermarking before vocoder inference, MelShield remains plug-and-play for Mel-conditioned TTS architectures and does not require modification or retraining of the underlying TTS generation vocoder, such as DiffWave and HiFi-GAN. Moreover, the multi-user keyed construction enables scalable user-specific attribution, while the keyed verification mechanism limits unauthorized decoding, thereby reducing the risk of large-scale extractor probing and adversarial analysis. Extensive experiments on DiffWave and HiFi-GAN demonstrate that MelShield achieves reliable watermark extraction, approaching 100\% bit accuracy, even under signal distortions, e.g., compression and additive noise, while preserving high perceptual audio quality.
【10】MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention
标题:MindMelody:一个用于个性化音乐干预的闭环脑电驱动系统链接:https://arxiv.org/abs/2605.01235
摘要:随着全球心理健康状况负担的不断加重,基于音乐的干预措施作为一种非侵入性、具有成本效益的情绪调节和心理压力缓解方式引起了人们的极大关注。然而,目前的数字音乐服务依赖于静态偏好,无法适应用户的瞬时心理状态。此外,直接映射脑电图(EEG)的音乐生成仍然具有挑战性,由于严重的配对数据稀缺和缺乏可解释性。为了解决这些限制,我们提出了MindMelody,一个功能齐全的,闭环实时系统,用于EEG驱动的个性化音乐干预。MindMelody引入了一个以情感为中介的语义桥。具体来说,混合变换器-GNN首先将实时EEG信号解码为全局Valence-Arousal状态和局部时间影响轨迹。然后,这些状态被馈送到配备检索增强生成(RAG)的大型语言模型(LLM)中,以制定结构化的干预计划。随后,一种新型的分层EEG控制器将全局影响前缀和局部时间指导注入到预先训练的音乐骨干中,从而实现细粒度可控的音频合成。至关重要的是,该系统采用了一个连续的反馈回路,根据用户不断变化的EEG动态更新生成参数。大量的实验表明,MindMelody提高了控制坚持和情绪调整,并在短期的听力设置中获得更高的感知帮助,这表明它有希望作为一个自适应的情感感知音乐生成框架。
摘要:Driven by the escalating global burden of mental health conditions, music-based interventions have attracted significant attention as a non-invasive, cost-effective modality for emotion regulation and psychological stress relief. However, current digital music services rely on static preferences and fail to adapt to users' instantaneous psychological states. Furthermore, directly mapping electroencephalography (EEG) to music generation remains challenging due to severe paired-data scarcity and a lack of interpretability. To address these limitations, we propose MindMelody, a fully functional, closed-loop real-time system for EEG-driven personalized music intervention. MindMelody introduces an emotion-mediated semantic bridge. Specifically, a hybrid Transformer-GNN first decodes real-time EEG signals into global Valence-Arousal states and local temporal affect trajectories. These states are then fed into a Retrieval-Augmented Generation (RAG)-equipped Large Language Model (LLM) to formulate structured intervention plans. Subsequently, a novel Hierarchical EEG Controller injects global affect prefixes and local temporal guidance into a pretrained music backbone, enabling fine-grained controllable audio synthesis. Crucially, the system incorporates a continuous feedback loop that updates generation parameters on the fly based on the user's evolving EEG dynamics. Extensive experiments show that MindMelody improves control adherence and emotional alignment, and receives higher perceived helpfulness in a short-term listening setting, suggesting its promise as an adaptive affect-aware music generation framework.
【11】Multimodal Confidence Modeling in Audio-Visual Quality Assessment
标题:视听质量评估中的多峰置信度建模链接:https://arxiv.org/abs/2605.01219
备注:Accepted at ICIP 2026, 6 pages, 4 figures, no supplementary material
摘要:视听质量评估(AVQA)对于流媒体、电话会议和沉浸式媒体至关重要。在现实的流场景中,失真通常是不对称的,其中一种模态可能严重退化,而另一种模态保持干净。尽管如此,大多数当代AVQA指标将音频和视频视为同样可靠,导致置信度不敏感的融合强调不可靠的信号。本文提出了MCM-AVQA,一个多模态的信心感知AVQA框架,明确估计特定模态的信心,并将其注入到一个专用的视听混合器的跨模态注意。视听混合器利用帧级、置信度引导的通道注意力进行门融合,调制模态之间的特征交互,使得高置信度流占主导地位,同时抑制不可靠的输入,保留时间退化模式。多头视觉置信度估计器将帧级伪影概率转换为时间平滑的剪辑级视觉置信度分数,而音频置信度模块从语音质量线索中获得置信度,而不需要干净的参考。多个AVQA基准测试的实验表明,MCM-AVQA,特别是其置信度引导的视听混合器,提高了与人类平均意见分数的相关性,并在现实世界的非对称视听失真下产生更多可解释的行为。
摘要:Audio-visual quality assessment (AVQA) is essential for streaming, teleconferencing, and immersive media. In realistic streaming scenarios, distortions are often asymmetric, where one modality may be severely degraded while the other remains clean. Still, most contemporary AVQA metrics treat audio and video as equally reliable, causing confidence-unaware fusion to emphasize unreliable signals. This paper proposes MCM-AVQA, a multimodal confidence-aware AVQA framework that explicitly estimates modality-specific confidence and injects it into a dedicated audio-visual mixer for cross-modal attention. The Audio-Visual Mixer utilizes frame-level, confidence-guided channel attention to gate fusion, modulating feature interaction between modalities so that high-confidence streams dominate while unreliable inputs are suppressed, preserving temporal degradation patterns. A multi-head visual confidence estimator turns frame-level artifact probabilities into temporally smoothed, clip-level visual confidence scores, while an audio confidence module derives confidence from speech-quality cues without requiring a clean reference. Experiments on multiple AVQA benchmarks show that MCM-AVQA, and specifically its confidence-guided Audio-Visual Mixer, improve correlation with human mean opinion scores and yield more interpretable behavior under real-world asymmetric audio-visual distortions.
【12】MG-Former: A Transformer-Based Framework for Music-Driven 3D Conducting Gesture Generation
标题:MG-Former:基于变形器的框架,用于音乐驱动的3D指挥手势生成链接:https://arxiv.org/abs/2605.01197
摘要:从音乐中生成富有表现力的指挥手势是一个具有挑战性的跨模态运动合成问题:输出必须遵循长距离音乐结构,保持节拍级同步,并保持合理的细粒度3D人类表演。现有的传导运动研究通常受到稀疏姿态表示、小规模数据或评估协议的限制,这些协议不直接测量音乐和手势是否相互对齐。本文介绍了TransConductor,一个基于transformer的框架,用于音乐驱动的指挥手势生成。我们引入了ConductorMotion,这是一个SPL参数数据构建管道,可以从指挥视频中恢复详细的身体运动,并形成针对专业指挥手势的数据集。给定从音频和初始姿势提取的声学描述符,TransConductor使用跨时间音乐编码器和跨时间传导姿势解码器来自回归地预测SMPL姿势参数。为了更好地评估艺术对应关系,我们进一步构建了一个基于检索的评估模型,该模型将音乐和手势嵌入到共享空间中,并产生FID,模态距离,多模态距离和多样性度量。实验表明,TransConductor优于舞蹈生成和传导生成基线,而消融验证了Transformer骨干和拟议的对准损失的好处。
摘要:Generating expressive conducting gestures from music is a challenging cross-modal motion synthesis problem: the output must follow long-range musical structure, preserve beat-level synchronization, and remain plausible as a fine-grained 3D human performance. Existing conducting-motion studies are often limited by sparse pose representations, small-scale data, or evaluation protocols that do not directly measure whether music and gesture are mutually aligned. This paper presents TransConductor, a Transformer-based framework for music-driven conducting gesture generation. We introduce ConductorMotion, a SMPL-parameter data construction pipeline that recovers detailed body motion from conducting videos and forms a dataset targeted at professional conducting gestures. Given acoustic descriptors extracted from audio and an initial pose, TransConductor uses a Trans-Temporal Music Encoder and a Trans-Temporal Conducting Gesture Decoder to autoregressively predict SMPL pose parameters. To better assess artistic correspondence, we further build a retrieval-based evaluation model that embeds music and gestures into a shared space and yields FID, modality distance, multi-modality distance, and diversity metrics. Experiments show that TransConductor outperforms dance-generation and conducting-generation baselines, while ablations verify the benefits of the Transformer backbone and the proposed alignment loss.
【13】Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy
标题:虚拟言语治疗师:临床医生在环人工智能言语治疗代理,用于个性化和监督治疗链接:https://arxiv.org/abs/2605.01101
备注:Under Review
摘要:本文开发了虚拟语音治疗师(VST),这是一个基于智能代理的平台,可以简化口吃评估,并通过自动化和自适应AI驱动的工作流程提供定制的治疗计划。VST集成了最先进的基于深度学习的口吃分类和多智能体大语言模型(LLM)推理,以支持基于证据的临床决策。VST从患者语音样本的采集和特征提取开始,然后对口吃类型进行鲁棒分类。在这些输出的基础上,VST启动了一个代理推理过程,在这个过程中,专门的LLM代理自动生成,批判和迭代地完善个性化的治疗计划。一个专门的评论家代理评估所有生成的治疗计划,以确保临床安全性,方法的合理性,并与同行评审的证据和既定的专业指南。结果输出是一个全面的,患者特定的治疗草案,旨在供临床医生审查。通过验证临床医生反馈,系统然后产生适合于患者递送的最终治疗计划,从而维持临床医生在环范例。专家言语治疗师的实验评估证实,VST始终产生高质量的,以证据为基础的治疗建议。这些研究结果表明,该系统有可能增加临床工作流程,减轻临床医生的负担,并改善言语障碍患者的治疗效果。所提出的系统的交互式用户界面可在https://vocametrix.com/ai/stuttering-therapy-planning-agent在线获得,便于实时口吃评估和个性化治疗计划。
摘要:This paper develops Virtual Speech Therapist (VST), an intelligent agent-based platform that streamlines stuttering assessment and delivers customized therapy planning through automated and adaptive AI-driven workflows. VST integrates state-of-the-art deep learning-based stuttering classification, and multi-agent large language model (LLM) reasoning to support evidence-based clinical decision-making. The VST begins with the acquisition and feature extraction of patient speech samples, followed by robust classification of stuttering types. Building on these outputs, VST initiates an agentic reasoning process in which specialized LLM agents autonomously generate, critique, and iteratively refine individualized therapy plans. A dedicated critic agent evaluates all generated therapy plans to ensure clinical safety, methodological soundness, and alignment with peer-reviewed evidence and established professional guidelines. The resulting output is a comprehensive, patient-specific therapy draft intended for clinician review. Incorporating clinician feedback, the system then produces a finalized therapy plan suitable for patient delivery, thereby maintaining a clinician-in-the-loop paradigm. Experimental evaluation by expert speech therapists confirms that VST consistently generates high-quality, evidence-based therapy recommendations. These findings demonstrate the system's potential to augment clinical workflows, reduce clinician burden, and improve therapeutic outcomes for individuals with speech impairments. An interactive user interface for the proposed system is available online at: https://vocametrix.com/ai/stuttering-therapy-planning-agent , facilitating real-time stuttering assessment and personalized therapy planning.
【14】MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
标题:MedMosaic:多元化医疗音频的令人惊叹的大规模基准链接:https://arxiv.org/abs/2605.00969
备注:Accepted at ICML 2026. 12 pages main text, 35 pages appendix, 5 figures, 7 tables
摘要:我们提出了MedMosaic,一个医疗音频问答数据集,旨在在现实的临床约束下对语言和音频推理模型进行基准测试。由于隐私法规和领域专业知识带来的高注释成本,医疗音频数据难以收集。因此,现有的基准往往不能充分代表复杂的医疗音频场景。为了应对这些挑战,MedMosaic提供了各种各样的医疗音频类型,包括与条件相关的生理声音,精心构建的合成语音,以模仿带有人工制品的语音,以及真实的长短临床对话,以模拟不同的上下文长度。该数据集还包含46,701个问答对,涵盖多项选择、顺序多轮和开放式问答等类别,可以系统地评估多跳推理和答案生成能力。对13个音频和多模态推理模型进行基准测试表明,对于所有评估的系统来说,推理仍然具有挑战性,不同问题类型之间的性能差异很大。特别是,即使是像Gemini-2.5-pro这样的最先进的模型也只能达到大约68.1%的准确率。这些发现强调了医学推理的持续局限性,并强调了对更强大的,特定领域的多模态推理模型的需求。
摘要:We present MedMosaic, a medical audio question-answering dataset designed to benchmark language and audio reasoning models under realistic clinical constraints. Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address these challenges, MedMosaic features a diverse range of medical audio types, including condition-related physiological sounds, carefully constructed synthetic voices to mimic speech with artifacts as well as real short and long length clinical conversations to model varying context lengths. The dataset also features a total of 46,701 question-answer pairs, spanning categories such as multiple-choice, sequential multi-turn, and open-ended question-answers, enabling systematic evaluation of multi-hop reasoning and answer generation capabilities. Benchmarking 13 audio and multimodal reasoning models reveals that reasoning remains challenging for all evaluated systems, with substantial performance variation across question types. In particular, even state-of-the-art model like Gemini-2.5-pro can only achieve 68.1% accuracy approximately. These findings underscore persistent limitations in medical reasoning and highlight the need for more robust, domain-specific multimodal reasoning models.
【15】Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
标题:迈向公平语音技术:语音人工智能偏见和公平性的全面调查链接:https://arxiv.org/abs/2605.01597
备注:32 pages, work in progress
摘要:语音技术被部署在高风险的环境中,但公平性问题仍然分散在任务和学科中。现有的调查要么采用一般的机器学习视角,忽视语音特定的属性,要么专注于单个任务,忽略了整个语音领域共享的失败模式。综合了400多项跨越生成和感知任务以及新兴语音语言模型的研究,该调查提出了一个统一的框架,将正式的公平性定义与评估,诊断和缓解联系起来。我们正式七个公平的定义,适应语音模态和组织领域的概念演变,通过三种范式:鲁棒性,代表性和治理。然后,我们地面评估指标的数学核心,这些定义,并提供了一个决策树的指标选择。我们沿着语音处理管道诊断偏见来源,将特定于语音的机制(如通道偏见)作为人口统计学代理和情感标签中的注释主观性。我们将缓解策略系统化,分为四个干预阶段,将每个阶段映射到诊断的来源。最后,我们确定了开放的挑战,并提出了未来的研究方向。
摘要:Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-learning perspective that overlooks speech-specific properties or focus on a single task, missing failure patterns shared across the speech domain. Synthesizing over 400 studies spanning generation and perception tasks and emerging speech-language models, this survey presents a unified framework that links formal fairness definitions to evaluation, diagnosis, and mitigation. We formalize seven fairness definitions adapted to the speech modality and organize the field's conceptual evolution through three paradigms: Robustness, Representation, and Governance. We then ground evaluation metrics in the mathematical cores of these definitions and offer a decision tree for metric selection. We diagnose bias sources along the speech processing pipeline, surfacing speech-specific mechanisms such as channel bias as a demographic proxy and annotation subjectivity in emotion labels. We systematize mitigation strategies across four intervention stages, mapping each to the diagnosed sources. Finally, we identify open challenges and propose directions for future research.
【16】How Well Can We Decode Vowels from Auditory EEG -- A Rigorous Cross-Subject Benchmark with Honest Assessment
标题:从听觉脑电中解码元音的效果如何--具有诚实评估的严格跨学科基准链接:https://arxiv.org/abs/2605.00865
备注:31 pages, 11 figures; includes supplementary material (14 pages, additional figures and analyses)
摘要:基于EEG的音素解码对于脑机接口是有希望的,但是许多先前的研究依赖于受试者内评估、小群组或弱泄漏控制。我们提出了一个可重复的跨学科基准五类元音解码(a,e,i,o,u)从听觉脑电图使用OpenNeuro ds006104(16名受试者,61通道,256 Hz)。在严格的leave one subject out评估下,仅训练归一化和显式防泄漏检查,我们比较了经典机器学习,深度学习和黎曼方法的14个管道。最好的全特征模型(XGBoost)达到24.5%的准确率(机会20%),而LightGBM的差分熵特征在特征特定分析中达到25.5%。在多重比较校正之后,强成对模型优势受到限制。在这种低信号状态下,经典方法与深度模型具有竞争力。其他分析(消融、成对元音、受试者内CV、ERP、时间概括和电极重要性)表明元音信息是真实的,但较弱,主要由早期短暂的听觉反应携带。我们发布代码和评估脚本,以实现完全可重复性。
摘要:EEG based phoneme decoding is promising for brain computer interfaces, but many prior studies rely on within subject evaluation, small cohorts, or weak leakage control. We present a reproducible cross subject benchmark for five class vowel decoding (a, e, i, o, u) from auditory EEG using OpenNeuro ds006104 (16 subjects, 61 channels, 256 Hz). Under strict leave one subject out evaluation with training only normalization and explicit anti leakage checks, we compare 14 pipelines from classical machine learning, deep learning, and Riemannian methods. The best full feature model (XGBoost) reaches 24.5 percent accuracy (chance 20 percent), while differential entropy features with LightGBM reach 25.5 percent in feature specific analysis. After multiple comparison correction, strong pairwise model advantages are limited. Classical methods are competitive with deep models in this low signal regime. Additional analyses (ablation, pairwise vowels, within subject CV, ERP, temporal generalization, and electrode importance) indicate that vowel information is real but weak and mainly carried by early transient auditory responses. We release code and evaluation scripts for full reproducibility.
标题:通过因子划分嵌入的多轴语音相似性
链接:https://arxiv.org/abs/2605.02804
备注:7 pages, accepted at Odyssey 2026
摘要:语音编码多个同时属性-语言内容,说话者身份,方言,性别-传统的单向量嵌入合并。 我们提出了一个因子划分的嵌入框架,将每个话语映射到一个单一的向量,其子空间对应于不同的变化轴。 一个共享的声学编码器为每轴线性投影头提供信息,每个投影头都是通过从专业教师或共享标签对的对比目标中提取信息进行训练的。 由此产生的嵌入支持属性条件检索:相似性计算为每个轴余弦分数的带符号加权和,允许检索联合考虑说了什么以及如何-或显式抑制一个属性以显示另一个属性。 我们评估跨语料库检索共享哈佛句子提示语料库,表明有符号的轴加权可以抑制相同的扬声器偏见和表面语义匹配的话语在整个记录条件。
摘要:Speech encodes multiple simultaneous attributes--linguistic content, speaker identity, dialect, gender--that conventional single-vector embeddings conflate. We present a factor-partitioned embedding framework that maps each utterance into a single vector whose subspaces correspond to distinct axes of variation. A shared acoustic encoder feeds per-axis linear projection heads, each trained via distillation from a specialist teacher or a contrastive objective over shared-label pairs. The resulting embeddings support attribute-conditioned retrieval: similarity is computed as a signed weighted sum over per-axis cosine scores, allowing retrieval that jointly considers what was said and how --or explicitly suppresses one attribute to surface another. We evaluate on cross-corpus retrieval over corpora sharing the Harvard sentence prompts, demonstrating that signed axis weighting can suppress same-speaker bias and surface semantically matched utterances across recording conditions.
【2】Dimensionality-Aware Anomaly Detection in Learned Representations of Self-Supervised Speech Models
标题:自监督语音模型学习表示中的感知异常检测链接:https://arxiv.org/abs/2605.02715
备注:Submitted to Interspeech 2026
摘要:自监督语音模型(S3 M)实现了强大的下游性能,但在自然和对抗性扰动下,它们的学习表示仍然很难理解。先前的研究依赖于表示相似性或全局维度,对局部几何变化的可见性有限。我们要问:扰动如何使局部几何形状变形,这些变化是否跟踪下游自动语音识别(ASR)的退化?为了解决这个问题,我们提出了GRIDS,一个框架,使用本地固有的分层表示(LID)在WavLM和wav 2 vec 2.0。我们发现,LID增加所有低信噪比(SNR)的扰动和发散在高SNR:良性噪声收敛到干净的配置文件,而敌对的输入保留早期层LID海拔。我们发现LID升高与WER增加同时发生,并且逐层LID特征能够进行异常检测(AUROC 0.78-1.00),为S3 M中的无转录监测打开了大门。
摘要:Self-supervised speech models (S3Ms) achieve strong downstream performance, yet their learned representations remain poorly understood under natural and adversarial perturbations. Prior studies rely on representation similarity or global dimensionality, offering limited visibility into local geometric changes. We ask: how do perturbations deform local geometry, and do these shifts track downstream automatic speech recognition (ASR) degradation? To address this, we present GRIDS, a framework using Local Intrinsic Dimensionality (LID) across layer-wise representations in WavLM and wav2vec 2.0. We find that LID increases for all low signal-to noise ratio (SNR) perturbations and diverges at high SNR: benign noise converges toward the clean profile, while adversarial inputs retain early-layer LID elevation. We show LID elevation co-occurs with increased WER, and that layer-wise LID features enable anomaly detection (AUROC 0.78-1.00), opening the door to transcript-free monitoring in S3Ms.
【3】Neck-Learn: Attention-Based Multiple Instance Learning and Ensemble Framework for Ecological Momentary Assessment
标题:Neck-Learn:基于注意力的多实例学习和生态瞬间评估集成框架链接:https://arxiv.org/abs/2605.02700
摘要:发声功能亢进(VH)是一种普遍的语音障碍,尽管有大量的日常语音数据,但其动态检测仍然具有挑战性。先前的方法捕获长达一周的颈部表面加速度计记录,但将它们折叠成固定长度的主体级特征向量,丢弃编码细微差别的发声特征交互的日内时间动态。我们引入了一种新的混合架构,将基于CNN的多实例学习(MIL)框架与基于CNN的多实例学习(MIL)框架结合在一起,该框架可以在每天的时间动态中保留和学习。在保留的测试集上,我们的模型超过了挑战基线(AUC:0.82 PVH,0.77 NPVH),PVH的AUC为0.879(Rank 5),NPVH的AUC为0.848(Rank 3),同时还提供了对这两种病理的临床相关信息的见解。
摘要:Vocal hyperfunction (VH) is a prevalent voice disorder whose ambulatory detection remains challenging despite extensive daily voice data. Prior approaches capture week-long neck-surface accelerometer recordings but collapse them into fixed-length subject-level feature vectors, discarding within-day temporal dynamics encoding nuanced voicing feature interactions. We introduce a novel hybrid architecture combining gradient-boosted trees on day-level distributional features with a CNN-based multiple instance learning (MIL) framework that preserves and learns from from temporal dynamics throughout each day. On the held-out test set, our model exceeds the challenge baselines (AUC: 0.82 PVH, 0.77 NPVH), achieving AUCs of 0.879 for PVH (Rank 5) and 0.848 for NPVH (Rank 3), while also providing insights into clinically relevant information about both pathologies.
【4】Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI
标题:迈向公平语音技术:语音人工智能偏见和公平性的全面调查链接:https://arxiv.org/abs/2605.01597
备注:32 pages, work in progress
摘要:语音技术被部署在高风险的环境中,但公平性问题仍然分散在任务和学科中。现有的调查要么采用一般的机器学习视角,忽视语音特定的属性,要么专注于单个任务,忽略了整个语音领域共享的失败模式。综合了400多项跨越生成和感知任务以及新兴语音语言模型的研究,该调查提出了一个统一的框架,将正式的公平性定义与评估,诊断和缓解联系起来。我们正式七个公平的定义,适应语音模态和组织领域的概念演变,通过三种范式:鲁棒性,代表性和治理。然后,我们地面评估指标的数学核心,这些定义,并提供了一个决策树的指标选择。我们沿着语音处理管道诊断偏见来源,将特定于语音的机制(如通道偏见)作为人口统计学代理和情感标签中的注释主观性。我们将缓解策略系统化,分为四个干预阶段,将每个阶段映射到诊断的来源。最后,我们确定了开放的挑战,并提出了未来的研究方向。
摘要:Speech technologies are deployed in high-stakes settings, yet fairness concerns remain fragmented across tasks and disciplines. Existing surveys either adopt a general machine-learning perspective that overlooks speech-specific properties or focus on a single task, missing failure patterns shared across the speech domain. Synthesizing over 400 studies spanning generation and perception tasks and emerging speech-language models, this survey presents a unified framework that links formal fairness definitions to evaluation, diagnosis, and mitigation. We formalize seven fairness definitions adapted to the speech modality and organize the field's conceptual evolution through three paradigms: Robustness, Representation, and Governance. We then ground evaluation metrics in the mathematical cores of these definitions and offer a decision tree for metric selection. We diagnose bias sources along the speech processing pipeline, surfacing speech-specific mechanisms such as channel bias as a demographic proxy and annotation subjectivity in emotion labels. We systematize mitigation strategies across four intervention stages, mapping each to the diagnosed sources. Finally, we identify open challenges and propose directions for future research.
【5】Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment
标题:文本到语音系统的语音映射:基于度量的语音质量评估方法链接:https://arxiv.org/abs/2605.00861
摘要:本研究探讨语音映射作为评估框架的文本到语音(TTS)合成质量。本研究分析了六种TTS模式,包括历史和最近的。这些指标是波峰因子、频谱平衡和倒谱峰值突出度(CPP)。我们研究了6种有影响力的TTS模型:Merlin、Tacotron 2、Transformer TTS、FastSpeech 2、Glow-TTS和VITS。结果表明,语音范围作为模型能力的主要指标,与VITS显示最大的范围之间的测试模型。辉光TTS表现出优越的性能,在软发声,更高的频谱平衡,尽管有限的声音范围。结果表明,在7-8 dB之间的CPPs值表示自然的语音质量,而当CPPs超过10 dB时,语音听起来倾向于机器人。这些发现强调了语音映射的必要性,以评估声乐的努力,并捕捉TTS系统如何处理语音动态和表现力。
摘要:This study investigates voice mapping as an evaluation framework for text-to-speech (TTS) synthesis quality. The study analyzes six TTS models, including historical and recent ones. The metrics are crest factor, spectrum balance, and cepstral peak prominence (CPPs). We investigated 6 influential TTS models: Merlin, Tacotron 2, Transformer TTS, FastSpeech 2, Glow-TTS, and VITS. The results demonstrate that voice range serves as a primary indicator of model capability, with VITS showing the largest range among tested models. Glow-TTS exhibited superior performance in soft phonation, indicated by higher spectrum balance, despite limited voice range. The results showed that the CPPs values between 7-8 dB indicate natural voice quality, while with CPPs exceeding 10 dB, the speech tends to sound robotic. These findings underscore the need for voice mapping to evaluate vocal effort, and capture how TTS systems handle voice dynamic and expressiveness.
【6】When Audio-Language Models Fail to Leverage Multimodal Context for Dysarthric Speech Recognition
标题:当音频语言模型未能利用多模式上下文进行结构障碍语音识别时链接:https://arxiv.org/abs/2605.02782
摘要:自动语音识别(ASR)系统在构音障碍和其他非典型语音上仍然很脆弱。最近的音频语言模型提高了通过在推理时调节额外的临床背景来提高性能的可能性,但目前还不清楚这些模型是否可以利用这些信息。我们引入了一个建立在语音无障碍项目(SAP)数据集上的基准测试,该数据集测试诊断标签、临床医生获得的语音评级和逐渐丰富的临床描述是否能提高构音障碍语音的转录准确性。通过对9个模型的匹配比较,我们发现当前的模型没有有意义地使用这种背景:诊断信息和临床详细提示产生的改进可以忽略不计,并且经常降低单词错误率。我们补充提示分析与上下文相关的微调,表明LoRA适应与临床提示格式的混合物实现了WER为0.066,比冻结基线相对减少52%,同时保留性能时,上下文不可用。亚组分析显示,唐氏综合征和轻度严重的扬声器显着收益。这些结果澄清了当前模型的不足之处,并为衡量更具包容性的ASR的进展提供了测试平台。
摘要:Automatic speech recognition (ASR) systems remain brittle on dysarthric and other atypical speech. Recent audio-language models raise the possibility of improving performance by conditioning on additional clinical context at inference time, but it is unclear whether these models can make use of such information. We introduce a benchmark built on the Speech Accessibility Project (SAP) dataset that tests whether diagnosis labels, clinician-derived speech ratings, and progressively richer clinical descriptions improve transcription accuracy for dysarthric speech. Across matched comparisons on nine models, we find that current models do not meaningfully use this context: diagnosis-informed and clinically detailed prompts yield negligible improvements and often degrade word error rate. We complement the prompting analysis with context-dependent fine-tuning, showing that LoRA adaptation with a mixture of clinical prompt formats achieves a WER of 0.066, a 52% relative reduction over the frozen baseline, while preserving performance when context is unavailable. Subgroup analyses reveal significant gains for Down syndrome and mild-severity speakers. These results clarify where current models fall short and provide a testbed for measuring progress toward more inclusive ASR.
【7】Mitigating Multimodal LLMs Hallucinations via Relevance Propagation at Inference Time
标题:通过推理时的相关传播减轻多模态LLM幻觉链接:https://arxiv.org/abs/2605.01766
摘要:多模态大型语言模型(MLLM)彻底改变了人工智能的格局,在处理复杂的视觉和音频语言任务方面展示了令人印象深刻的能力。然而,一个关键的挑战仍然存在:这些模型经常遭受幻觉,产生与所提供的感知输入不同的输出。这种倾向源于推理过程中模态利用的内在不平衡,其中文本标记的主导地位破坏了感知输入的潜力。因此,该模型经常以牺牲有根据的证据为代价诉诸文本语言先验。为了解决这个问题,我们提出了学习推理时间模态增强(LIME),这是一个免训练的框架,旨在通过显式增强解码过程中的模态使用来支持多模态接地。LIME利用分层相关传播(LRP)来量化令牌级贡献,并定义了一个基于相关性的目标,以促进对感知输入的依赖。这个目标是通过对模型的键值表示进行推理时更新来实现的,而不需要修改模型参数或需要额外的训练数据。我们在视觉和音频领域的多个多模态基准测试中评估了LIME,在保持生成质量的同时,证明了幻觉和增强接地的一致减少。进一步的分析表明,LIME增加了情态贡献,并产生更多的本地化和语义对齐的相关模式。
摘要:Multimodal large language models (MLLMs) have revolutionized the landscape of AI, demonstrating impressive capabilities in tackling complex vision and audio-language tasks. However, a critical challenge remains: these models often suffer from hallucinations, generating outputs that diverge from the provided perceptual inputs. This tendency stems from an inherent imbalance in modality utilization during inference, where the dominance of textual tokens undermines the potential of perceptual inputs. As a result, the model frequently resorts to textual language priors at the expense of grounded evidence. To tackle this issue, we propose Learning Inference-time Modality Enhancement (LIME), a training-free framework designed to bolster multimodal grounding by explicitly enhancing modality usage during decoding. LIME leverages Layer-wise Relevance Propagation (LRP) to quantify token-level contributions and defines a relevance-based objective that promotes increased reliance on perceptual inputs. This objective is enforced through inference-time updates to the model's key-value representations, without modifying model parameters or requiring additional training data. We evaluate LIME across multiple multimodal benchmarks in both vision and audio domains, demonstrating consistent reductions in hallucinations and enhanced grounding while preserving generation quality. Further analysis shows that LIME increases modality contribution and produces more localized and semantically aligned relevance patterns.
【8】Virtual Speech Therapist: A Clinician-in-the-Loop AI Speech Therapy Agent for Personalized and Supervised Therapy
标题:虚拟言语治疗师:临床医生在环人工智能言语治疗代理,用于个性化和监督治疗链接:https://arxiv.org/abs/2605.01101
备注:Under Review
摘要:本文开发了虚拟语音治疗师(VST),这是一个基于智能代理的平台,可以简化口吃评估,并通过自动化和自适应AI驱动的工作流程提供定制的治疗计划。VST集成了最先进的基于深度学习的口吃分类和多智能体大语言模型(LLM)推理,以支持基于证据的临床决策。VST从患者语音样本的采集和特征提取开始,然后对口吃类型进行鲁棒分类。在这些输出的基础上,VST启动了一个代理推理过程,在这个过程中,专门的LLM代理自动生成,批判和迭代地完善个性化的治疗计划。一个专门的评论家代理评估所有生成的治疗计划,以确保临床安全性,方法的合理性,并与同行评审的证据和既定的专业指南。结果输出是一个全面的,患者特定的治疗草案,旨在供临床医生审查。通过验证临床医生反馈,系统然后产生适合于患者递送的最终治疗计划,从而维持临床医生在环范例。专家言语治疗师的实验评估证实,VST始终产生高质量的,以证据为基础的治疗建议。这些研究结果表明,该系统有可能增加临床工作流程,减轻临床医生的负担,并改善言语障碍患者的治疗效果。所提出的系统的交互式用户界面可在https://vocametrix.com/ai/stuttering-therapy-planning-agent在线获得,便于实时口吃评估和个性化治疗计划。
摘要:This paper develops Virtual Speech Therapist (VST), an intelligent agent-based platform that streamlines stuttering assessment and delivers customized therapy planning through automated and adaptive AI-driven workflows. VST integrates state-of-the-art deep learning-based stuttering classification, and multi-agent large language model (LLM) reasoning to support evidence-based clinical decision-making. The VST begins with the acquisition and feature extraction of patient speech samples, followed by robust classification of stuttering types. Building on these outputs, VST initiates an agentic reasoning process in which specialized LLM agents autonomously generate, critique, and iteratively refine individualized therapy plans. A dedicated critic agent evaluates all generated therapy plans to ensure clinical safety, methodological soundness, and alignment with peer-reviewed evidence and established professional guidelines. The resulting output is a comprehensive, patient-specific therapy draft intended for clinician review. Incorporating clinician feedback, the system then produces a finalized therapy plan suitable for patient delivery, thereby maintaining a clinician-in-the-loop paradigm. Experimental evaluation by expert speech therapists confirms that VST consistently generates high-quality, evidence-based therapy recommendations. These findings demonstrate the system's potential to augment clinical workflows, reduce clinician burden, and improve therapeutic outcomes for individuals with speech impairments. An interactive user interface for the proposed system is available online at: https://vocametrix.com/ai/stuttering-therapy-planning-agent , facilitating real-time stuttering assessment and personalized therapy planning.
机器翻译由腾讯交互翻译提供,仅供参考
