今日论文合集:cs.SD语音18篇,eess.AS音频处理13篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Hidden in Plain Tokens: Simply Robust, Gradient-Free Watermark for Synthetic Audio
标题:隐藏在纯令牌中:用于合成音频的简单稳健、无干扰水印
链接:https://arxiv.org/abs/2605.25967
作者:Georgios Milis,Yubin Qin,Yihan Wu,Heng Huang
备注:Accepted to ICML 2026
摘要:随着政策赶上生成AI的能力,水印是内容来源工作的核心。自回归模型的推理时间水印由于离散化不一致而不适合连续模态。现有的方法克服了这一点,微调的模态标记,无效的水印的训练自由的优势。在这项工作中,出于离散化的词汇冗余,我们提出了一个优雅的解决方案,强大的和鲁棒的合成音频水印。我们从理论上分析了令牌错误对水印检测的影响,并有效地减轻他们使用减少通过社区检测获得的词汇。彻底的实验表明,我们的无梯度方法可以提高几个数量级的可检测性,同时还实现了内置的音频修改的鲁棒性。从广义上讲,我们发现了一个新的国家的最先进的令牌级水印在多媒体中,这仅仅是由于离散表示学习的性质。
摘要:As policy catches up with the capabilities of generative AI, watermarking is central to content provenance efforts. Inference-time watermarks for autoregressive models are unfit for continuous modalities due to discretization inconsistencies. Existing methods overcome this by finetuning the modality tokenizers, nullifying the watermark's training-free advantage. In this work, motivated by the vocabulary redundancy of discretization, we propose an elegant solution for powerful and robust watermarking of synthetic audio. We theoretically analyze the impact of token errors on watermark detection, and effectively mitigate them using a reduced vocabulary obtained via community detection. Thorough experiments showcase that our gradient-free method can boost detectability by several orders of magnitude, while also achieving built-in robustness to audio modifications. Broadly, we discover a new state-of-the-art for token-level watermarks in multimedia, which simply arises from the nature of discrete representation learning.


【2】Continual Speaker Identity Unlearning with Minimal Interference

标题:以最小的干扰持续消除说话者身份
链接:https://arxiv.org/abs/2605.25962
作者:Jinju Kim,Yunsung Kang,Gyeong-Moon Park,Jong Hwan Ko
备注:preprint
摘要:机器非学习从预先训练的模型中删除指定的概念或知识。最近的工作已经扩展了这种模式的说话人身份遗忘在zero-shot文本到语音(TTS),选择性地擦除模型的能力,复制一个扬声器的声音的任务。然而,现有的方法悄悄地假设所有的遗忘请求同时到达;这是一个不切实际的假设,因为隐私动机的删除会随着时间的推移而顺序到达。我们表明,这种假设打破了最先进的方法:忘记每个新的扬声器完全恢复以前未学习的扬声器,重新引入非常隐私的风险忘记是为了消除。我们提出了累积ORTHOOPHOTOSTIC身份抑制(CORTIS),第一个框架,连续扬声器的身份unlearning在ST-TTS,不需要访问以前未学习的扬声器数据。CORTIS结合了基于Fisher信息的参数掩蔽,其将更新定位到说话者相关权重,并对先前的非学习更新所跨越的子空间进行正交投影。使用VoiceBox,CORTIS可以忘记每个被请求的说话者,同时保持以前未学习的说话者在长时间的请求序列中被遗忘,大大优于先前方法的顺序应用。该演示可在https://cumulativeortis.github.io/上获得。
摘要:Machine unlearning removes designated concepts or knowledge from pre-trained models. Recent work has extended this paradigm to speaker identity unlearning in zero-shot text-to-speech (ZS-TTS), the task of selectively erasing a model's ability to replicate a speaker's voice. Existing methods, however, quietly assume all unlearning requests arrive at once; an unrealistic assumption, since privacy-motivated removals arrive sequentially over time. We show this assumption breaks state-of-the-art methods: unlearning each new speaker fully revives previously unlearned speakers, reintroducing the very privacy risk unlearning was meant to eliminate. We present Cumulative ORThogonal Identity Suppression (CORTIS), the first framework for continual speaker identity unlearning in ZS-TTS that requires no access to previously-unlearned speaker data. CORTIS combines Fisher-information-based parameter masking, which localizes updates to speaker-relevant weights, with orthogonal projection against subspaces spanned by prior unlearning updates. With VoiceBox, CORTIS unlearns each requested speaker while keeping previously unlearned speakers forgotten across long request sequences, substantially outperforming sequential application of prior methods. The demo is available at https://cumulativeortis.github.io/ .


【3】Score-Agnostic Structure Analysis in Large-Scale Performance Datasets

标题:大规模绩效数据集中的得分不可知结构分析
链接:https://arxiv.org/abs/2605.25951
作者:Patricia Hu,Silvan Peter,Gerhard Widmer
备注:published at the Music Encoding Conference (MEC) 2026
摘要:近年来,由于自动音乐转录(AMT)的进步,已经发布了几个自动转录的钢琴独奏音乐的大规模数据集。虽然这些数据集无疑为性能研究提供了广泛的材料,但它们的质量差异很大。   在古典音乐的情况下,表演往往不仅在表现方面,如速度,但也在他们的乐谱结构解释(包括重复模式和版本特定的变体)不同。为了有意义地使用大规模转录数据集进行性能研究,必须根据其潜在的结构实现对同一片段的转录进行分组,以支持有效的比较。   我们通过应用序列对序列比对,然后进行分层聚类来解决这个问题:我们为给定片段的所有transmits对创建成对比对,并使用比对成本和执行序列长度的(不)相似性来解决结构错配作为分组特征。我们提出这种方法作为自动评估缺乏地面实况分数和/或音频的大规模转录数据集的第一步,将评估标准从基于事实的准确性转变为音乐连贯性和可扩展性。   我们展示了我们的分数不可知论的方法,从最近发表的大规模转录钢琴演奏数据集的88个组成的约1,500个转录。
摘要:In recent years, thanks to advances in automatic music transcription (AMT), several large-scale datasets of automatically transcribed piano solo music have been released. While these datasets undoubtedly offer extensive material for performance studies, they vary substantially in quality.   In the case of classical music, performances often differ not only in expressive aspects such as tempo, but also in their structural interpretation of the score (including repeat patterns and edition-specific variants). To meaningfully use large-scale transcribed datasets for performance research, transcriptions of the same piece must be grouped according to their underlying structural realisation to support valid comparison.   We address this by applying sequence-to-sequence alignment followed by hierarchical clustering: we create pairwise alignments for all pairs of transcriptions of a given piece, and use the alignment cost and (dis)similarity of performed sequence lengths to resolve structural mismatches as features for grouping. We propose this approach as a first step towards automatically evaluating large-scale transcribed datasets that lack ground-truth score and/or audio, shifting the evaluation criterion from truth-based accuracy to musical coherence and plausibility.   We demonstrate our score-agnostic approach on around 1,500 transcriptions of 88 compositions from a recently published large-scale transcribed piano performance dataset.


【4】CosyEdit2: Speech-Editing-Oriented Reinforcement Learning Unlocks Better Zero-Shot TTS

标题:CosyEditor 2:面向语音编辑的强化学习解锁更好的Zero-ShotTTC
链接:https://arxiv.org/abs/2605.25930
作者:Junyang Chen,Yuhang Jia,Hui Wang,Jiaming Zhou,Yongchang Gan,Yong Qin
摘要:语音编辑和zero-shot文本到语音(TTS)共享以语音提示为条件的类似生成基础,然而语音编辑要求与周围未编辑内容的更严格的局部声学一致性。虽然先前的工作已经表明,监督微调(SFT)使TTS模型获得功能编辑能力,这种方法仍然从根本上受到不完美的配对编辑数据和粗粒度优化信号的阻碍。为了解决这些局限性,我们提出了CosyEdit 2,这是一种基于两阶段后训练框架的语音编辑模型,该框架从监督编辑初始化到面向编辑的组相对策略优化(GRPO)。大量的实验表明,CosyEdit 2不仅大大提高了语音编辑性能,而且还解锁了更好的zero-shot TTS能力,揭示了这两个任务之间更深层次的相互关系。音频样本可在https://cjy1018.github.io/CosyEdit2上获得。
摘要:Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consistency with surrounding unedited content. While prior work has shown that Supervised Fine-Tuning (SFT) enables TTS models to acquire functional editing capability, this approach remains fundamentally bottlenecked by imperfect paired editing data and coarse-grained optimization signals. To address these limitations, we propose CosyEdit2, a speech editing model built on a two-stage post-training framework that progresses from supervised editing initialization to editing-oriented Group Relative Policy Optimization (GRPO) over target-speech-free data. Extensive experiments demonstrate that CosyEdit2 not only substantially advances speech editing performance, but also unlocks better zero-shot TTS capability, revealing a deeper mutual relationship between the two tasks. Audio samples are available at https://cjy1018.github.io/CosyEdit2.


【5】Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

标题:塔卡出席KSBA-2026任务2:阿拉伯语语音变音化的常规微调
链接:https://arxiv.org/abs/2605.25928
作者:Meshal Alamr,Hassan Alqaeri,Abdullah Aldahlawi
备注:4 pages, 1 figure. Published in Proceedings of OSACT7 (LREC 2026). Winning system for KSAA-2026 Task 2 on Arabic Speech Diacritization
摘要:我们描述了KSAA-2026阿拉伯语语音听写与自动变音的共享任务的任务2的获奖系统。该任务需要从语音音频和未变音的转录本中生成完全变音的阿拉伯语文本,只有2,327个训练样本,并且不允许使用外部数据。我们的系统微调CATT-Whisper,这是一种字符级多模态模型,将预训练的CATT文本编码器与冻结的Whisper语音编码器相结合。我们方法的关键是训练正则化:R-Drop一致性正则化,Optuna优化的高权重衰减超参数和焦点损失。在推理时,我们在softmax概率水平上使用Monte Carlo Dropout对四个模型检查点的200次随机向前传递进行平均。该系统在主要排行榜指标上实现了23.26%的WER(带有案例结尾,包括无变音符号位置),在所有参与者中排名第一。
摘要:We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires producing fully diacritized Arabic text from speech audio and undiacritized transcripts, with only 2,327 training samples available and no external data permitted. Our system fine-tunes CATT-Whisper, a character-level multimodal model combining a pretrained CATT text encoder with a frozen Whisper speech encoder. The key to our approach is training regularization: R-Drop consistency regularization, Optuna-optimized hyperparameters with high weight decay, and Focal Loss. At inference, we average 200 stochastic forward passes across four model checkpoints using Monte Carlo Dropout at the softmax probability level. The system achieves 23.26% WER on the primary leaderboard metric (with case endings, including no-diacritic positions), placing 1st among all participants.


【6】A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

标题:通过语言和声学表示学习进行痴呆症检测的多模式框架
链接:https://arxiv.org/abs/2605.25540
作者:Loukas Ilias,Dimitris Askounis
摘要:阿尔茨海默病(AD)是一种进行性神经退行性疾病,是痴呆的主要原因,影响记忆、推理、交流和日常功能。早期诊断尤其重要,因为及时干预可能有助于减缓认知能力下降并改善患者护理。最近的研究表明,自发言语包含有价值的语言和声学生物标志物与痴呆症。然而,现有的方法往往依赖于独立训练的模态特定的模型,特征级联策略,集成方法,或基于注意力的融合机制,没有显式地最大化语音和转录表示之间的依赖性。在这项工作中,我们提出了一个用于自动痴呆检测的多模式深度学习框架,该框架以端到端的可训练方式联合利用语音和转录信息。具体来说,语音记录被分成10秒的片段,并通过预先训练的HuBERT模型来提取上下文化的声学表示。为了更好地捕获信息丰富的时间语音特征,采用注意统计池来聚合帧级声学嵌入。对于文本模态,使用预训练的BERT模型对转录进行编码,其中[CLS]令牌表示用作语言嵌入。随后使用基于注意力的音频-文本融合(AT-融合)机制将声学和文本表示相结合。此外,我们引入了一个MINE目标,以最大限度地提高模态之间的互信息,提高多模态表示对齐。融合的多模态表示最终用于痴呆症分类。在公开的ADReSS Challenge和PROCESS-2数据集上进行的实验证明了所提出的基于语音的痴呆评估方法的有效性和鲁棒性。
摘要:Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, and daily functioning. Early diagnosis is particularly important, as timely intervention may help slow cognitive decline and improve patient care. Recent studies have demonstrated that spontaneous speech contains valuable linguistic and acoustic biomarkers associated with dementia. However, existing approaches often rely on independently trained modality-specific models, feature concatenation strategies, ensemble methods, or attention-based fusion mechanisms that do not explicitly maximize the dependency between speech and transcript representations. In this work, we propose a multimodal deep learning framework for automatic dementia detection that jointly exploits speech and transcript information in an end-to-end trainable manner. Specifically, speech recordings are divided into 10-second segments and passed through a pre-trained HuBERT model to extract contextualized acoustic representations. To better capture informative temporal speech characteristics, attentive statistics pooling is employed to aggregate frame-level acoustic embeddings. For the textual modality, transcripts are encoded using a pre-trained BERT model, where the [CLS] token representation is used as the linguistic embedding. The acoustic and textual representations are subsequently combined using an attention-based Audio-Text Fusion (AT-Fusion) mechanism. In addition, we introduce a MINE objective to maximize the mutual information between modalities and improve multimodal representation alignment. The fused multimodal representation is finally used for dementia classification. Experiments conducted on the publicly available ADReSS Challenge and PROCESS-2 dataset demonstrate the effectiveness and robustness of the proposed approach for speech-based dementia assessment.


【7】Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

标题:从语音中Zero-Shot帕金森病检测:比较大型音频和语言模型
链接:https://arxiv.org/abs/2605.24806
作者:Muhammad Ashad Kabir,Sirajam Munira
备注:6 pages
摘要:大型音频和语言模型最近已经在各个领域展示了zero-shot推理能力。然而,目前尚不清楚音频输入的形式,无论是从语音中提取的手工声学特征还是原始音频波形本身,如何影响不同语言的帕金森病(PD)检测性能。在这项研究中,我们系统地比较了两种输入模式的zero-shot PD检测:(i)手工制作的声学特征提取的语音记录分析的通用LLM,和(ii)直接波形输入分析的音频模型。在四种语言的PD语音数据集上的实验表明,性能在输入方式、语音任务和语言之间存在差异。手工制作的声学特征在低资源语言中提供更稳定的性能(例如,孟加拉语),而音频输入产生依赖于语音的增益。这些发现突出了输入模态对从语音中检测出zero-shot PD的影响。
摘要:Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted from speech or the raw audio waveform itself, affects performance for Parkinson's disease (PD) detection across different languages. In this study, we systematically compare two input modalities for zero-shot PD detection: (i) handcrafted acoustic features extracted from speech recordings analyzed by a general-purpose LLM, and (ii) direct waveform input analyzed by audio-capable models. Experiments on PD speech datasets in four languages show that performance varies across input modalities, speech tasks, and languages. Handcrafted acoustic features provide more stable performance in a low-resource language (e.g., Bengali), whereas audio input yields dataset-dependent gains. These findings highlight the impact of input modality on zero-shot PD detection from speech.


【8】Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care

标题:心理健康护理临床决策支持的感知言语特征探索
链接:https://arxiv.org/abs/2605.24678
作者:Vassilis Lyberatos,Edmund G. Dervakos,Eleni Adamidi,Athanasios Voulodimos,Giorgos Stamou
备注:Accepted to CLPsych 2026, part of ACL 2026
摘要:语音和语言技术通过客观和可解释的线索为支持心理健康评估提供了宝贵的机会。我们提出了一个系统的基于特征的分析框架,利用感知接地声学和语言特征,包括韵律,音质,语义连贯性,句法结构和讽刺。使用统计分析和可解释的机器学习(XGBoost with SHAP and LIME),我们研究了语音特征与抑郁、焦虑和ADHD的有效症状指标之间的关联。在对照基准数据集(StressID、DAIC-WOZ、Androids、EATD)和真实世界的临床数据集上进行评估,该框架揭示了症状严重程度和声音不规则之间稳定和一致的关系(例如,闪光,抖动),词汇句法模式,和情感基调。对所有数据集进行的消融研究进一步确定了信息量最大的特征组。这项工作探索了一种透明的和临床上可解释的方法,以语音为基础的心理健康分析。
摘要:Speech and language technologies offer valuable opportunities for supporting mental health assessment through objective and interpretable cues. We present a systematic feature-based analysis framework leveraging perceptually grounded acoustic and linguistic characteristics, including prosody, vocal quality, semantic coherence, syntactic structure, and sarcasm. Using statistical analysis and interpretable machine learning (XGBoost with SHAP and LIME), we examine associations between speech features and validated symptom measures of depression, anxiety, and ADHD. Evaluated on both controlled benchmark datasets (StressID, DAIC-WOZ, Androids, EATD) and a real-world clinical dataset, the framework reveals stable and consistent relationships between symptom severity and vocal irregularities (e.g., shimmer, jitter), lexical-syntactic patterns, and affective tone. An ablation study conducted across all datasets further identifies the most informative feature groups. This work explores a transparent and clinically interpretable approach to speech-based mental health analysis.


【9】AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

标题:AVBench:音频视频生成模型的人性化和自动化评估基准
链接:https://arxiv.org/abs/2605.24652
作者:Jialiang Yang,Bin Xia,Ruihang Chu,Dingdong Wang,Wanke Xia,Zhun Mou,Tianyang Zhong,Yiting Zhao,Wenming Yang
摘要:音频-视频(AV)生成的快速发展已经实现了具有同步声音的高保真合成,特别是对于涉及语音和交互的人类相关场景。然而,对AV生成的评估仍处于早期阶段,只有少数与人类相关的场景的粗粒度基准,并且依赖于使用通用多模态LLM的有限预设评估,导致对模型能力的评估不准确。为了解决这些问题,我们引入了AVBench,这是一个为以人为中心的AV生成量身定制的全自动基准测试。AVBench基于两个关键设计,以实现全面和准确的评估:(i)以人为本和细粒度的指标。AVBench集成了十个针对以人为中心的真实场景设计的评估维度,涵盖视觉质量、音频质量和跨模态的多级一致性。这些实用的指标捕捉了现有基准经常忽略的与人相关的细节。(ii)通过偏好学习的专业评估者。为了解决缺乏专业训练数据的问题,我们通过将真实世界的视频转换为具有受控扰动的不同训练对来构建大规模监督。在对这个高质量的数据集进行微调之后,评估器学会了可靠地检测微妙的跨模态不一致。至关重要的是,AVBench不是产生离散的文本判断,而是从模型对二元决策的预测置信度中获得连续的评估分数。这种概率评分机制能够实现比传统VQA风格评估更可靠的评估,并与人类判断密切相关。总的来说,AVBench为AV生成提供了自动化评估,展示了强大的数据过滤潜力,并作为从人类反馈强化学习(RLHF)的可区分奖励信号。
摘要:Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios involving speech and interactions. Yet evaluation for AV generation remains at an early stage, with only a few coarse-grained benchmarks for human-related scenarios and relying on limited preset evaluations with generic multimodal LLMs, leading to inaccurate assessments of model capabilities. To address these issues, we introduce AVBench, a fully automated benchmark tailored for human-centric AV generation. AVBench is built on two key designs for comprehensive and accurate evaluation: (i) Human-centric and fine-grained metrics. AVBench integrates ten evaluation dimensions designed for human-centered real-world scenarios, covering visual quality, audio quality, and multi-level consistency across modalities. These practical metrics capture human-related details that existing benchmarks often overlook. (ii) Specialized evaluators via preference learning. To address the lack of specialized training data, we construct large-scale supervision by transforming real-world videos into diverse training pairs with controlled perturbations. After fine-tuning on this high-quality dataset, the evaluators learn to reliably detect subtle cross-modal inconsistencies. Crucially, instead of producing discrete textual judgment, AVBench derives continuous evaluation scores from the model's prediction confidence on binary decisions. This probabilistic scoring mechanism enables a more reliable assessment than traditional VQA-style evaluation and aligns closely with human judgment. Taken together, AVBench offers automated evaluation for AV generation, demonstrates strong potential for data filtering, and serves as a differentiable reward signal for Reinforcement Learning from Human Feedback (RLHF).


【10】Rubato: Transcribing Piano Music with Timestamps

标题:鲁巴托:用时间戳抄写钢琴音乐
链接:https://arxiv.org/abs/2605.24291
作者:Nazif Can Tamer,Victoria Ebert,Guang Yang,Noah A. Smith
备注:18 pages, 7 figures, 5 tables
摘要:我们考虑将音乐录音转换为人类可读的带有时间戳注释的乐谱。这样的输出可以让听众清楚地看到rubato(时间表达演奏),学习者可以根据书面音乐诊断合奏精度和时机选择,音乐学学者可以比较同一作品的录音中的表演风格。我们介绍了(1)一个名为Rubato的无条件编码器-解码器模型,它被训练为输出(2)一个新的复调音乐文本表示,名为InterMo,我们设计它是为了与序列到序列训练兼容。我们的实验表明,Rubato产生的时间戳钢琴乐谱从音频具有更高的记谱精度比现有的最好的方法,这是基于级联。我们发现,即使级联被给予地面实况广播而不是音频,Rubato也表现得更好,这表明现有方法的天花板主要是代表性的,而不是声学的。此外,由于Rubato是在几个相关的任务(有提示)上训练的,因此它在相关但更简单的任务(如音符接地和节拍/强拍检测)上与最好的单任务系统竞争或优于它们。演示可在https://nctamer.github.io/rubato-transcription上获得。
摘要:We consider the conversion of musical recordings into human-readable sheet music annotated with timestamps. Such output lets a listener clearly visualize rubato (temporally expressive playing), a learner diagnose ensemble precision and timing choices against the written music, and a musicology scholar compare performance styles across recordings of the same work. We introduce (1) a prompt-conditioned encoder-decoder model, named Rubato, trained to output (2) a new textual representation for polyphonic music, named InterMo, which we designed for compatibility with sequence-to-sequence training. Our experiments demonstrate that Rubato produces timestamped piano sheet music from audio with higher notational accuracy than the best existing approaches, which are based on cascades. We find that even if the cascade is given ground-truth MIDI instead of audio, Rubato performs better, suggesting that the ceiling of existing approaches is primarily representational, not acoustic. Further, because Rubato is trained on several related tasks (with prompts), it competes with or outperforms the best single-task systems on related but simpler tasks like MIDI note grounding and beat/downbeat detection. A demo is available at https://nctamer.github.io/rubato-transcription .


【11】Music Transcription with (Almost) No Supervision

标题:(几乎)没有监督的音乐抄写
链接:https://arxiv.org/abs/2605.24193
作者:Saebyeol Shin,Chao Wan,Zhenzhen Liu,Justin Lovelace,Daniel C. Lin,Kilian Q. Weinberger,John Thickstun
摘要:竞争性音乐转录模型需要大量成对的音频乐谱数据,由于收集成本、对齐困难和版权限制,这些数据是稀缺的。与此同时,大量未配对的录音和符号乐谱可以免费获得,但没有使用。我们采用了一个周期一致的翻译框架,其中少量的配对数据作为最小的锚点,释放了未配对池的全部潜力。我们发现:未配对的数据产生令人惊讶的大增益,特别是在有限的监督下;未配对的音频比未配对的分数贡献更多;在训练期间合并来自新乐器的未标记的音频改进了该乐器的转录,而无需任何配对监督。总之,这些结果表明,缩放未配对数据为标记数据仍然稀缺的仪器提供了一条通往高质量转录的实用途径。
摘要:Competitive music transcription models require large amounts of paired audio-score data, which is scarce due to collection costs, alignment difficulty, and copyright restrictions. Meanwhile, vast quantities of unpaired audio recordings and symbolic scores are freely available but have gone unused. We adopt a cycle-consistent translation framework in which a small amount of paired data acts as a minimal anchor, unlocking the full potential of the unpaired pool. We find that: unpaired data yields surprisingly large gains, especially under limited supervision; unpaired audio contributes more than unpaired scores; incorporating unlabeled audio from a new instrument during training improves transcription for that instrument without any paired supervision. Together, these results suggest that scaling unpaired data offers a practical path toward high-quality transcription for instruments where labeled data remains scarce.


【12】PiAnnotate: A Web Annotation Tool for Piano Fingering, with a Diagnostic Probe

标题:PiAnnotate:一个用于钢琴指法的网络注释工具,带有诊断探针
链接:https://arxiv.org/abs/2605.23982
作者:Joonhyung Bae,Kirak Kim,Hyeyoon Cho,Sein Lee,Yoon-Seok Choi,Hyeon Hur,Gyubin Lee,Akira Maezawa,Jonghwa Park,Jaebum Park,Juhan Nam
摘要:钢琴指法塑造了一个段落如何被演奏,然而在演奏之后很难给它贴上标签。注释者必须决定哪个手指产生每个音符,同时协调乐谱、计时、视频和手部动作。我们提出了PiAnnotate,一个基于Web的管道,用于将专家指法注释添加到FurElise性能数据集。该工具汇集了钢琴滚动视图,性能视频和3D MANO手网,以便评审员可以在音乐和物理环境中检查每个任务。PiAnnotate不是只存储最终答案,而是保留配对的基于规则和人工编辑的指法轨迹。这些成对的轨迹通过显示几何规则在哪里足够、专家在哪里干预以及标签如何在审查过程中发生变化,使注释历史可审计。作为最后的诊断,我们在成对的轨道上训练一个小的Transformer探测器。探测器在保留片段的规则基线上进行了改进,同时对更改已经正确的标签保持保守,这表明编辑的标签包含可学习的结构,而不仅仅是孤立的修复。
摘要:Piano fingering shapes how a passage can be played, yet it is difficult to label after a performance. An annotator must decide which finger produced each note while reconciling the score, timing, video, and hand motion. We present PiAnnotate, a web-based pipeline for adding expert fingering annotations to the FurElise performance dataset. The tool brings together a piano-roll view, performance video, and a 3D MANO hand mesh so that reviewers can inspect each assignment in musical and physical context. Rather than storing only the final answer, PiAnnotate keeps paired rule-based and human-edited fingering tracks. These paired tracks make the annotation history auditable by showing where a geometric rule was sufficient, where experts intervened, and how labels changed across review passes. As a final diagnostic, we train a small Transformer probe on the paired tracks. The probe improves on the rule baseline on held-out pieces while remaining conservative about changing labels that were already correct, suggesting that the edited labels contain learnable structure rather than only isolated fixes.


【13】A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

标题:临床访谈抑郁检测基准的多探头审计
链接:https://arxiv.org/abs/2605.23977
作者:Takehiro Ishikawa,Jon Duke
摘要:本文通过DAIC/E-DAIC,CMDC,ANDROIDS,MODMA和PDCH的四个互补探针审计临床访谈抑郁症检测中的基准评估。首先,我们重新评估E-DAIC下严格的主题不相交留一主题交叉验证。一个轻量级的混合文本加LLM分数模型达到了macro-F1 = 0.723-据我们所知,这是该协议下报告的最高值-提供了一个保守的不依赖于特权官方坚持的参考点。其次,我们测试E-DAIC官方分裂是否支持细粒度的排行榜排名,通过扫描96个模型配置跨模态捆绑,池化策略和学习器。开发侧交叉验证和官方测试排名仅适度一致:最佳交叉验证配置在官方测试中排名第20,官方测试获胜者在交叉验证中排名第41,前3名重叠为零,只有32.3%的主题引导中明显获胜者排名第1。第三,我们在外部验证强大的公共CMDC和ANDROIDS基线,以实现接近上限的域内性能。零镜头迁移到外部语料库是相当弱的。最后,我们使用由基于SRDS的注释器定义的成对的高密度与低密度访谈切片对E-DAIC文本和音频模型进行压力测试。文本分数在高密度切片上急剧上升,而音频分数几乎保持不变;文本减去音频的差距在所有五个种子中都是正的。
摘要:This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS, MODMA, and PDCH. First, we re-evaluate E-DAIC under strict subject-disjoint leave-one-subject-out cross-validation. A lightweight hybrid text-plus-LLM-score model reaches macro-F1 = 0.723 - the highest reported under this protocol, to our knowledge - providing a conservative out-of-fold reference point that does not depend on the privileged official holdout. Second, we test whether the E-DAIC official split supports fine-grained leaderboard rankings by sweeping 96 model configurations across modality bundles, pooling strategies, and learners. Development-side cross-validation and official-test rankings align only moderately: the best cross-validation configuration ranks twentieth on the official test, the official-test winner ranks forty-first by cross-validation, top-3 overlap is zero, and the apparent winner is rank-1 in only 32.3% of subject bootstraps. Third, we externally validate strong public CMDC and ANDROIDS baselines that achieve near-ceiling in-domain performance. Zero-shot transfer to external corpora is substantially weaker. Finally, we stress-test E-DAIC text and audio models using paired symptom-dense versus symptom-light interview slices defined by an SRDS-based annotator. Text scores rise sharply on symptom-dense slices, whereas audio scores remain nearly flat; the text-minus-audio gap is positive across all five seeds.


【14】Direct Preference Optimization for English-Mandarin Code-Switching Speech Recognition in Audio LLMs

标题:音频LLM中英语-普通话代码转换语音识别的直接偏好优化
链接:https://arxiv.org/abs/2605.23975
作者:Trung Nguyen Quang,Cheng Yi Lewis Won,Minh Duc Pham,Yingxu He,Shuo Sun,Ai Ti Aw
摘要:音频大语言模型(音频LLM)表现出系统性的故障,尽管强大的多语言能力,在转录代码切换语音。针对英语-普通话,我们确定了三种失败模式:语言省略,以言代录,和幻觉。我们应用直接偏好优化(DPO)来对齐模型,构建偏好对,其中选择的响应保留混合语言内容,而拒绝的响应模仿失败模式。在10万对(570小时)上训练三个音频LLM,我们观察到一致的行为转变:模型学习保留语言成分,而不是在提示转录时进行翻译。这种一致性使MER降低高达89.6%(分布内)和20.0%(分布外)。我们的研究结果表明,DPO可以有效地诱导正确的语码转换转录行为的多语种音频LLM。
摘要:Audio large language models (Audio LLMs) exhibit systematic failures in transcribing code-switching speech despite strong multilingual capabilities. Focusing on English-Mandarin, we identify three failure modes: language omission, translation-instead-of-transcription, and hallucination. We apply Direct Preference Optimization (DPO) to align models, constructing preference pairs in which chosen responses preserve mixed-language content while rejected responses mimic failure patterns. Training three Audio LLMs on 100K pairs (570 hours), we observe consistent behavioral shifts: models learn to preserve language composition rather than translating when prompted for transcription. This alignment yields MER reductions up to 89.6% (in-distribution) and 20.0% (out-of-distribution). Our findings suggest DPO can effectively elicit correct code-switching transcription behavior from multilingual Audio LLMs.


【15】EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

标题:EchoDistill:用于稳健音频LLM的校准噪音清洁自蒸馏
链接:https://arxiv.org/abs/2605.23954
作者:Liang Lin,Chunxi Luo,Kaiwen Luo,Jie Zhang,Jin Wang,Yuanhe Zhang,Cai Yuchen,Qiankun Li,Gongli Xi,Zhenhong Zhou,Kun Wang,Junhao Dong
摘要:音频大语言模型(ALLM)非常容易受到真实世界噪声的影响,这通常会导致严重的语义漂移和幻觉。现有的鲁棒性方法主要依赖于波形级声学增强、应答级监督或噪声表示的内部抑制。为了解决这些问题,我们提出了echodistill,一个基于自蒸馏的自蒸馏框架。EchoDistill利用冻结的干净音频教师为推理时间嘈杂音频学生提供语义参考。具体来说,学生在噪声条件下对候选响应进行采样,以暴露其测试时的行为。然后,这些轨迹通过组相对策略优化(GRPO)进行优化,其中与教师的令牌级一致性作为奖励奖金。通过将嘈杂的学生的候选响应与干净的语义证据对齐,并应用音频感知奖励成形,我们的方法鼓励正确且真正声学接地的推理轨迹。Echodistill显著提高了复杂噪声下音频LLM的语义可靠性和任务性能,而不会引入任何额外的推理成本。大量的实验表明:(1)在强噪声条件下,与最强基线相比,回声蒸馏法平均提高了4.18%。(II)Qwen-Omni上的消融结果进一步表明,echodistill在Acc中比仅GRPO变体平均提高了3.02\%$\uparrow$,在Noisy中提高了3.89\%$\uparrow$,在GSR中提高了4.53\%$\uparrow$。我们的代码可在https://anonymous.4open.science/r/echodistill-10DE上获得。
摘要:Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing robustness methods primarily rely on waveform-level acoustic enhancement, answer-level supervision, or the internal suppression of noise representations. To address these issues, we propose echodistill, an alignment-based noisy-to-clean self-distillation framework. Echodistill leverages a frozen clean-audio teacher to provide semantic references for an inference-time noisy-audio student. Specifically, the student samples candidate responses under noisy conditions to expose its test-time behavior. These trajectories are then optimized via group-relative policy optimization (GRPO), where the token-level consistency with the teacher acts as a reward bonus. By aligning the noisy student's candidate responses with clean semantic evidence, and applying audio-aware reward shaping, our method encourages reasoning trajectories that are both correct and genuinely acoustically grounded. Echodistill significantly improves the semantic reliability and task performance of Audio LLMs under complex noise, without introducing any additional inference costs. Extensive experiments show that: (I) Compared with the strongest baseline, echodistill achieves average improvements of 4.18\%$\uparrow$ in GSR under strong noise. (II) Ablation results on Qwen-Omni further show that echodistill improves over the GRPO-only variant by 3.02\%$\uparrow$ in Acc, 3.89\%$\uparrow$ in Noisy, and 4.53\%$\uparrow$ in GSR on average. Our codes are available at https://anonymous.4open.science/r/echodistill-10DE.


【16】Raon-Speech Technical Report

标题:Raon-Speech技术报告
链接:https://arxiv.org/abs/2605.23912
作者:Beomsoo Kim,Changho Choi,Dohyun Kim,Dongki Lee,Ethan Ewer,Eunchong Kim,Gyeongman Kim,Haechan Kim,Hyeonghwan Kim,Inkyu Park,Jihun Yun,Jihwan Moon,Jiyun Kim,Joonghyun Bae,Junhyuck Kim,Minkyu Kim,Sehun Lee,Seungjun Chung,Sungwoo Cho,Dongmin Park,Dongwon Kim,Hara Kang,Jonghyun Lee,Keon Lee,Kangwook Lee,Jaewoong Cho
摘要:我们提出了Raon-Speech,一个用于英语和韩语语音理解,回答和生成的高性能9 B参数语音语言模型(SpeechLM),以及Raon-SpeechChat,一个用于自然实时对话的高性能全双工扩展。Raon-Speech成功地将预先训练的LLM转换为SpeechLM,该LLM既能理解又能生成语音,同时保留了强大的文本功能。它在138万小时的高度策划的英语和韩语语音和文本数据集上进行训练,包括以下训练阶段:(1)语音模块对齐,(2)端到端SpeechLM预训练和知识蒸馏,以及(3)基于多任务偏好优化的后训练。在42个英语和韩语语音和文本基准测试中,Raon-Speech在我们与八个类似规模的最近音频基础模型(包括Qwen2.5-Omni和Fun-Audio-Chat)的比较中,建立了以语音为中心的任务的最强整体配置文件,同时保持了强大的文本问答性能。在此基础上,Raon-SpeechChat通过对119 K小时的时间对齐的真实和合成对话数据进行持续训练,实现自然的全双工对话。它通过三个互补的训练阶段进行:(1)因果编码器自适应,(2)全双工预训练,(3)语音和角色控制的全双工微调。在多个全双工基准测试中,Raon-SpeechChat在FDB v1.0所涵盖的话轮转换和中断敏感行为方面表现出了最明显的优势,并且在更广泛的全双工评估套件中仍然具有竞争力。我们开源了所有模型检查点、训练和推理管道以及交互式演示。
摘要:We present Raon-Speech, a top-performing 9B-parameter speech language model (SpeechLM) for English and Korean speech understanding, answering, and generation, and Raon-SpeechChat, a high-performing full-duplex extension for natural real-time conversation. Raon-Speech successfully transforms a pre-trained LLM into a SpeechLM that both understands and generates speech while preserving strong text capabilities. It trains on 1.38M hours of highly curated English and Korean speech and text datasets with the following training stages: (1) speech modules alignment, (2) end-to-end SpeechLM pre-training with knowledge distillation, and (3) multi-task preference optimization-based post-training. Across 42 English and Korean speech and text benchmarks, Raon-Speech establishes the strongest overall profile on speech-centric tasks in our comparison against eight similarly sized recent audio foundation models, including Qwen2.5-Omni and Fun-Audio-Chat, while preserving strong text question answering performance. Building upon it, Raon-SpeechChat enables natural full-duplex conversation by continual training on 119K hours of time-aligned real and synthetic dialogue data. It proceeds through three complementary training stages: (1) causal encoder adaptation, (2) full-duplex pre-training, (3) full-duplex fine-tuning for voice and role-control. On multiple full-duplex benchmarks, Raon-SpeechChat shows its clearest strengths on the turn-taking and interruption-sensitive behaviors covered by FDB v1.0, and remains competitive across the broader full-duplex evaluation suite. We open-source all model checkpoints, the training and inference pipeline, and an interactive demo.


【17】Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

标题:重新思考语音和音频的持续学习:以代表为中心的分类学和开放问题
链接:https://arxiv.org/abs/2605.24863
作者:Yang Xiao,Siyi Wang,Eun-Jung Holden,Ting Dang
备注:4 pages, 1 figure, working in process
摘要:语音和音频系统工作在固有的非平稳环境中,但在这一领域的持续学习(CL)的研究,特别是在基础模型时代,仍然是支离破碎的,未能占耦合,几何敏感的性质的声学表示。现代语音基础模型在高度纠缠的连续表示上运行,这些表示在共享的潜在空间内联合编码语言,扬声器和非语言因素。因此,CL从根本上讲是关于保留和发展共享的表示结构,而不是保留孤立的任务知识。在这项工作中,我们重新CL语音代表为中心的角度来看,并介绍了一种新的分类法,组织CL根据底层的代表几何形状如何在非平稳的声学条件下演变。我们进一步确定了当前CL假设和语音基础模型行为之间的关键不匹配,最后概述了一组开放的挑战和未来的研究方向。
摘要:Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions.


【18】Time Segmented Beamforming via Dynamic Programming: Theory and Implementation

标题:通过动态规划进行时分束形成:理论与实现
链接:https://arxiv.org/abs/2605.24825
作者:Manan Mittal,Ryan M. Corey,Diego Cuji,John R. Buck,Andrew C. Singer
备注:16 pages, 17 figures, Beamforming New Approach Regret Bounds
摘要:在具有时变干扰的动态声学环境中,有效的波束形成需要识别随时间的静止区域。Capon波束形成器是一种白化匹配滤波器,理论上依赖于瞬时系综协方差矩阵。实际的实现依赖于批量Capon(或样本矩阵反演),它通过对快照块进行平均来估计样本协方差矩阵(SCM)。这种实用的方法隐含地假设批处理窗口内的数据是稳定的,并且可以被相干地组合。在非平稳设置中,在固定或过长的窗口上求平均的批量方法失败,因为移动的波束形成器模糊SCM并降低波束形成器的调零能力。为了解决这个问题,本文介绍了时间分段无失真响应波束形成器。受分段最小二乘法的启发,该方法将分段多项式拟合到数据,同时惩罚过度分割以防止过拟合,该框架通过结合数据驱动的时间分割来扩展实际的Capon波束形成。该公式最大限度地减少输出功率,同时动态地适应SCM估计窗口的本地平稳性,提供了一个原则性的方法来跟踪时变干扰。
摘要:In dynamic acoustic environments with time-varying interferers, effective beamforming requires identifying stationary regions over time. The Capon beamformer, a whitened matched filter constrained to maintain unity gain in the desired direction, theoretically relies on the instantaneous ensemble covariance matrix. Practical implementations rely on the batch Capon (or Sample Matrix Inversion), which estimates the sample covariance matrix (SCM) by averaging over a block of snapshots. This practical approach implicitly assumes that the data within the batch window is stationary and can be coherently combined. In non-stationary settings, a batch approach that averages over fixed or excessively long windows fails, as moving interferers smear the SCM and degrade the beamformer's nulling capabilities. To address this, this paper introduces a temporally segmented distortionless response beamformer. Inspired by the segmented least squares method, which fits piecewise polynomials to data while penalizing excessive segmentation to prevent overfitting, the framework extends practical Capon beamforming by incorporating data-driven temporal segmentation. This formulation minimizes output power while dynamically adapting the SCM estimation windows to local stationarity, offering a principled approach to tracking time-varying interferers.


eess.AS音频处理


【1】Ultra-Low-Bitrate Mel-Spectrogram-based Neural Speech Coding with Flow-Matching-based Refinement and Vocoding-driven Reconstruction
标题:基于超低比特率Mel谱图的神经语音编码、基于流匹配的细化和声码驱动的重建
链接:https://arxiv.org/abs/2605.25669
作者:Hui-Peng Du,Yang Ai,Xiao-Hang Jiang,Yuan Tian,Zhen-Hua Ling
备注:Published at IEEE/ACM Transactions on Audio, Speech, and Language Processing
摘要:超低比特率语音编码对于带宽受限的通信和深度压缩是关键的,然而由于明显的信息丢失和量化不稳定性,在这种极端的比特预算下保持自然度和说话者身份仍然具有挑战性。为此,我们提出了FMelCodec,一个超低比特率的神经语音编解码器在梅尔频谱域,铸造作为一个三阶段的编码-细化-重建(CRR)的框架,可以在低至250 bps。在CRR框架中,前端梅尔频谱图编码阶段采用具有单个1024条目VQ码本的高度积极的640 x压缩/解压缩编码器-解码器结构,再加上重新分配未充分使用的码字以防止码本崩溃并保持码本多样性的在线聚类策略。随后的基于条件流匹配(CFM)的梅尔频谱图细化阶段利用轻量级速度场估计器和基于CFM的求解器来细化由前一解码器产生的编解码器降级的梅尔频谱图,并采用支持较少迭代推理步骤的自一致性训练方案以减少计算开销。最后,声码驱动的波形重建阶段采用HiFi-GAN声码器从细化的梅尔频谱图忠实地重建波形。在两个采样率的数据集上进行的实验表明,在16 kHz的250 bps和48 kHz的750 bps的超低比特率约束下,客观和主观评价一致地表明,FMelCodec实现了更高的语音重建质量和说话人相似性,同时降低了计算和模型复杂度。
摘要:Ultra-low-bitrate speech coding is pivotal for bandwidth-constrained communication and deep compression, yet maintaining naturalness and speaker identity at such extreme bit budgets remains challenging due to pronounced information loss and quantization instability. To this end, we propose FMelCodec, an ultra-low-bitrate neural speech codec in the mel-spectrogram domain, cast as a three-stage coding-refinement-reconstruction (CRR) framework that can operate at as low as 250 bps. In the CRR framework, the front-end mel-spectrogram coding stage employs a highly aggressive 640x compression/decompression encoder-decoder structure with a single 1024-entry VQ codebook, coupled with an online clustering strategy that reassigns underused codewords to prevent codebook collapse and preserve codebook diversity. The subsequent conditional flow matching (CFM)-based mel-spectrogram refinement stage leverages a lightweight velocity-field estimator and CFM-based solver to refine the codec-degraded mel-spectrogram produced by the preceding decoder, and adopts a self-consistency training scheme that supports fewer iterative inference steps for the purpose of reducing computational overhead. Finally, the vocoding-driven waveform reconstruction stage employs a HiFi-GAN vocoder to faithfully reconstruct waveform from the refined mel-spectrogram. Experiments conducted on two datasets spanning two sampling rates show that, under ultra-low-bitrate constraints of 250 bps for 16 kHz and 750 bps for 48 kHz, both objective and subjective evaluations consistently demonstrate that FMelCodec achieves higher speech reconstruction quality and speaker similarity, while incurring lower computational and model complexity.


【2】Decoding Stimulus Reconstruction-Based Auditory Attention Robustly in Unbalanced EEG Datasets

标题:不平衡脑电数据集中基于刺激重建的听觉注意力解码
链接:https://arxiv.org/abs/2605.25605
作者:Yuanming Zhang,Yayun Liang,Zhibin Lin,Jing Lu
摘要:在过去的十年中,许多研究已经应用深度神经网络(DNN)通过刺激重建从脑电图(EEG)信号中解码听觉注意力(AAD)。然而,数据集平衡对基于刺激重构的AAD解码性能的影响仍未被探索。在这项研究中,三个公开可用的EEG-AAD数据集- KUL,DTU和NJU cEEGrid -被用来构建平衡和不平衡的实验条件。我们假设并证明,基于刺激重建的DNN解码器往往会在不平衡数据集上产生高估的解码性能。为了解决这个问题,我们提出了一个留一配对信封(LOPEO)的交叉验证协议。实验结果表明,LOPEO算法有效地防止了不平衡数据集上的解码精度膨胀。虽然平衡数据集通常是实验设计中的首选,但LOPEO为已经发表的不平衡数据集提供了一个原则性的评估框架,填补了该领域的一个重要空白。
摘要:In the past decade, numerous studies have applied deep neural networks (DNNs) to decode auditory attention (AAD) from Electroencephalogram (EEG) signals via stimulus reconstruction. However, the influence of dataset balance on the decoding performance of stimulus reconstruction-based AAD remains unexplored. In this study, three publicly available EEG-AAD datasets - KUL, DTU, and NJU cEEGrid - are used to construct both balanced and unbalanced experimental conditions. We hypothesize and demonstrate that stimulus reconstruction-based DNN decoders tend to produce overestimated decoding performance on unbalanced datasets. To address this issue, we propose a leave-one-paired-envelope-out (LOPEO) cross-validation protocol. Experimental results confirm that LOPEO effectively prevents inflated decoding accuracy on unbalanced datasets. While balanced datasets are generally preferred in experimental design, LOPEO provides a principled evaluation framework for unbalanced datasets that have already been published, filling an important gap in the field.


【3】cSTMM: A Unified Complex Spherical Student's $t$ Mixture Model for Directional Statistics in Mask-Based Blind Speech Separation

标题:cSTMM:基于口罩的盲语音分离中用于方向统计的统一复球形学生的$t$混合模型
链接:https://arxiv.org/abs/2605.25512
作者:Nobutaka Ito
摘要:基于掩码的盲语音分离(BSS)通过使用空间信息对多通道观测进行聚类来估计源方向的时频(TF)掩码。方向统计方法集群归一化多通道观测复杂的单位球,而不明确提取相位和电平差功能的基础上的平面波或球面波的假设。然而,以前的研究主要是比较了少数单独定义的方向统计混合模型,而更广泛的分布家庭将使一个更系统的研究密度分布如何影响分离性能。   我们提出了复球面学生t混合模型(cSTMM),一个方向的混合模型,连接复杂的角中心高斯混合模型(cACGMM),复杂的宾汉混合模型(cBMM),和复杂的沃森混合模型(cWMM)通过自由度参数$ν$。我们还推导出一个广义的最小化最大化(MM)为基础的参数估计过程。对无噪声LibriSpeech混合物与测量的房间脉冲响应进行的无重启评估表明,在所有声学条件下,单个开发选择的值$v ^\ast=1$实现了比cACGMM等效设置$v =M$更高的测试集平均信号失真比改善(SDRi),平均条件增益为0.25dB。实验还数值验证了所提出的配方数值恢复cACGMM,cBMM,cWMM的情况下。
摘要:Mask-based blind speech separation (BSS) estimates source-wise time-frequency (TF) masks by clustering multichannel observations using spatial information. The directional statistical approach clusters normalized multichannel observations on the complex unit sphere, without explicitly extracting phase and level difference features based on the plane-wave or spherical-wave assumptions. However, prior studies have mostly compared a small number of separately defined directional statistical mixture models, whereas a broader distribution family would enable a more systematic study of how density profiles affect separation performance.   We propose the complex spherical Student's t mixture model (cSTMM), a directional mixture model that connects the complex angular central Gaussian mixture model (cACGMM), complex Bingham mixture model (cBMM), and complex Watson mixture model (cWMM) through the degrees-of-freedom parameter $ν$. We also derive a generalized minorization-maximization (MM) based procedure for parameter estimation. A no-restart evaluation on noise-free LibriSpeech mixtures reverberated with measured room impulse responses shows that a single development-selected value $ν^\ast=1$ achieved higher test-set mean signal-to-distortion ratio improvements (SDRi) than the cACGMM-equivalent setting $ν=M$ in all acoustic conditions, with an average condition-wise gain of 0.25dB. The experiments also numerically verify that the proposed formulation numerically recovers the cACGMM, cBMM, and cWMM cases.


【4】WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

标题:WaveNeXt 2:基于ConvNeXtt的快速神经声码器,具有GAN和扩散模型的残余去噪和子建模
链接:https://arxiv.org/abs/2605.25506
作者:Wangzixi Zhou,Takuma Okamoto,Yamato Ohtani,Sakriani Sakti,Hisashi Kawai
备注:ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
摘要:大多数神经声码器限于一种类型:GAN或基于扩散的。虽然Vocos和WaveNeXt等最先进的模型使用强大的基于ConvNeXt的生成器,但它们仅用于GAN框架,并且在多扬声器设置中性能有限。此外,尽管扩散模型的训练速度比GANs快,但CPU推理速度很慢。在本文中,我们介绍了WaveNeXt 2,一个统一的基于ConvNeXt的框架兼容的GAN和扩散声码器。其核心创新是残差去噪和子模型,其中每个子模型逐步细化波形。在多说话人数据集上的实验结果证明了我们方法的有效性:(1)GAN-WaveNeXt 2比HiFi-GAN和WaveFit快得多,(2)Diff-WaveNeXt 2与4步FastDiff相比,也提供了更快的推理和有竞争力的合成质量。Diff-WaveNeXt 2的训练效率非常高,仅需32小时即可完成训练,非常适合资源受限的应用。
摘要:Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have only been used in GAN frameworks and have limited performance in multi-speaker settings. Moreover, diffusion models, despite training faster than GANs, have slow CPU inference. In this paper, we introduce WaveNeXt 2, a unified ConvNeXt-based framework compatible with both GAN and diffusion vocoders. Its core innovation is residual denoising and sub-modeling, where each sub-model progressively refines the waveform. Experimental results in the multi-speaker dataset demonstrate the effectiveness of our approach: (1) GAN-WaveNeXt 2 is much faster than HiFi-GAN and WaveFit, and (2) Diff-WaveNeXt 2 also delivers much faster inference and competitive synthesis quality compared with FastDiff with 4 steps. The Diff-WaveNeXt 2 is very training-efficient, training in only 32 hours, making it ideal for resource-constrained applications.


【5】Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

标题:迈向具有细粒度非言语表达控制的自然情感文本转语音系统
链接:https://arxiv.org/abs/2605.25504
作者:Wangzixi Zhou,Bagus Tris Atmaja,Sakriani Sakti
备注:2025 28th Conference of the Oriental COCOSDA International Committee for the Co-ordination and Standardisation of Speech Databases and Assessment Techniques (O-COCOSDA)
摘要:虽然目前的情感文本到语音(TTS)模型已经成功地控制了言语韵律,但它们往往忽略了非言语发声(NV),这是真实人类情感的关键。尽管最近出现了一些非语言数据集,但它们通常缺乏高质量、细粒度的注释,这限制了模型精确控制NV生成的能力。为了解决这个问题,我们提出了一种新的方法细粒度的非语言表达合成。我们从EARS语料库中挑选和重新处理女性NV话语,开发一种新的注释方案,使用标签对NV类型、频率和持续时间进行编码,并建立一个情感TTS基准来证明其有效性。我们的评估表明,虽然我们的NV方法在感知自然性方面有轻微的权衡,但它显着提高了表达力(eMOS 4.20)和情感识别准确率(78.8%)。针对特定情绪的分析进一步表明,NV线索对高唤醒情绪非常有效,如快乐(82.5%)和恐惧(82.7%),几乎完美地传达悲伤(98.3%)。
摘要:While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic human emotion. Although some non-verbal datasets have recently emerged, they often lack high-quality, fine-grained annotations, which restricts a model's ability to precisely control NV generation. To address this limitation, we propose a novel approach for fine-grained non-verbal expression synthesis. We curate and reprocess female NV utterances from the EARS corpus, develop a new annotation scheme using tags to encode NV types, frequencies, and durations, and build an emotional TTS benchmark to demonstrate its effectiveness. Our evaluation shows that while our NV approach leads to minor trade-offs in perceived naturalness, it significantly improves expressiveness (eMOS 4.20) and emotional recognition accuracy (78.8%). Emotion-specific analysis further reveals that NV cues are highly effective for high-arousal emotions like happy (82.5%) and fear (82.7%), and almost perfectly convey sadness (98.3%).


【6】Subspace Track-before-Detect for Passive Multi-Target Tracking with Unknown Emitted Signals

标题:未知发射信号被动多目标跟踪的子空间检测前跟踪
链接:https://arxiv.org/abs/2605.25498
作者:Nobutaka Ito,Yoshiaki Bando
摘要:被动多目标跟踪(MTT)的目的是从噪声传感器数据中推断多个目标的运动状态,其中来自未知目标发射信号的贡献被叠加。检测前跟踪(TBD)方法通过直接对原始传感器数据进行操作而不依赖于先前的检测阶段来提高对噪声的鲁棒性。然而,许多现有的TBD方法假设,每个目标的传感器数据的贡献是完全由其运动状态。这种假设限制了它们对被动MTT的适用性,其中每个目标的贡献取决于其运动状态和未知的发射信号。我们提出了子空间TBD,被动多目标TBD方法的基础上,从复杂的宾汉分布,不需要明确的建模或估计的未知发射信号的可能性。在粒子滤波(PF)框架中,每个多目标假设被映射到一个低维子空间,该子空间由与假设目标状态相对应的导向向量构成。然后使用似然性来评估归一化的多通道传感器数据与该子空间的对准。仿真实验表明,该方法可以在信噪比为-10dB的情况下跟踪两个发射未知信号的运动目标,而传统的TBD基线会产生较大的跟踪误差。
摘要:Passive multi-target tracking (MTT) aims to infer the kinematic states of multiple targets from noisy sensor data in which contributions from unknown target-emitted signals are superposed. Track-before-detect (TBD) methods improve robustness to noise by operating directly on raw sensor data without relying on a preceding detection stage. However, many existing TBD methods assume that each target's contribution to the sensor data is determined solely by its kinematic state. This assumption limits their applicability to passive MTT, where each target's contribution depends on both its kinematic state and the unknown emitted signal. We propose subspace TBD, a passive multi-target TBD method based on a likelihood derived from the complex Bingham distribution that does not require explicit modeling or estimation of the unknown emitted signals. In a particle filter (PF) framework, each multi-target hypothesis is mapped to a low-dimensional subspace spanned by the steering vectors corresponding to the hypothesized target states. The likelihood is then used to evaluate the alignment of the normalized multichannel sensor data with this subspace. Preliminary experiments with simulated acoustic measurements and a given target activity pattern show that the proposed method can track two moving targets emitting unknown signals at a signal-to-noise ratio (SNR) of -10dB, whereas a conventional TBD baseline yields substantially larger tracking errors.


【7】Rethinking Continual Learning for Speech and Audio: A Representation-Centric Taxonomy and Open Problems

标题:重新思考语音和音频的持续学习:以代表为中心的分类学和开放问题
链接:https://arxiv.org/abs/2605.24863
作者:Yang Xiao,Siyi Wang,Eun-Jung Holden,Ting Dang
备注:4 pages, 1 figure, working in process
摘要:语音和音频系统工作在固有的非平稳环境中,但在这一领域的持续学习(CL)的研究,特别是在基础模型时代,仍然是支离破碎的,未能占耦合,几何敏感的性质的声学表示。现代语音基础模型在高度纠缠的连续表示上运行,这些表示在共享的潜在空间内联合编码语言,扬声器和非语言因素。因此,CL从根本上讲是关于保留和发展共享的表示结构,而不是保留孤立的任务知识。在这项工作中,我们重新CL语音代表为中心的角度来看,并介绍了一种新的分类法,组织CL根据底层的代表几何形状如何在非平稳的声学条件下演变。我们进一步确定了当前CL假设和语音基础模型行为之间的关键不匹配,最后概述了一组开放的挑战和未来的研究方向。
摘要:Speech and audio systems operate in inherently non-stationary environments, yet continual learning (CL) research in this domain, especially in the foundation model era, remains fragmented that fail to account for the coupled, geometry-sensitive nature of acoustic representations. Modern speech foundation models operate over highly entangled, continuous representations that jointly encode linguistic, speaker, and paralinguistic factors within a shared latent space. CL is therefore fundamentally about preserving and evolving shared representation structure rather than retaining isolated task knowledge. In this work, we revisit CL for speech from a representation-centered perspective, and introduce a new taxonomy that organizes CL according to how underlying representation geometry evolves under non-stationary acoustic conditions. We further identify key mismatches between current CL assumptions and speech foundation model behavior, and finally outline a set of open challenges and future research directions.


【8】Time Segmented Beamforming via Dynamic Programming: Theory and Implementation

标题:通过动态规划进行时分束形成:理论与实现
链接:https://arxiv.org/abs/2605.24825
作者:Manan Mittal,Ryan M. Corey,Diego Cuji,John R. Buck,Andrew C. Singer
备注:16 pages, 17 figures, Beamforming New Approach Regret Bounds
摘要:在具有时变干扰的动态声学环境中,有效的波束形成需要识别随时间的静止区域。Capon波束形成器是一种白化匹配滤波器,理论上依赖于瞬时系综协方差矩阵。实际的实现依赖于批量Capon(或样本矩阵反演),它通过对快照块进行平均来估计样本协方差矩阵(SCM)。这种实用的方法隐含地假设批处理窗口内的数据是稳定的,并且可以被相干地组合。在非平稳设置中,在固定或过长的窗口上求平均的批量方法失败,因为移动的波束形成器模糊SCM并降低波束形成器的调零能力。为了解决这个问题,本文介绍了时间分段无失真响应波束形成器。受分段最小二乘法的启发,该方法将分段多项式拟合到数据,同时惩罚过度分割以防止过拟合,该框架通过结合数据驱动的时间分割来扩展实际的Capon波束形成。该公式最大限度地减少输出功率,同时动态地适应SCM估计窗口的本地平稳性,提供了一个原则性的方法来跟踪时变干扰。
摘要:In dynamic acoustic environments with time-varying interferers, effective beamforming requires identifying stationary regions over time. The Capon beamformer, a whitened matched filter constrained to maintain unity gain in the desired direction, theoretically relies on the instantaneous ensemble covariance matrix. Practical implementations rely on the batch Capon (or Sample Matrix Inversion), which estimates the sample covariance matrix (SCM) by averaging over a block of snapshots. This practical approach implicitly assumes that the data within the batch window is stationary and can be coherently combined. In non-stationary settings, a batch approach that averages over fixed or excessively long windows fails, as moving interferers smear the SCM and degrade the beamformer's nulling capabilities. To address this, this paper introduces a temporally segmented distortionless response beamformer. Inspired by the segmented least squares method, which fits piecewise polynomials to data while penalizing excessive segmentation to prevent overfitting, the framework extends practical Capon beamforming by incorporating data-driven temporal segmentation. This formulation minimizes output power while dynamically adapting the SCM estimation windows to local stationarity, offering a principled approach to tracking time-varying interferers.


【9】FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations

标题:FC-TTC:Zero-Shot文本到语音转换中的风格和音色控制,语音分解
链接:https://arxiv.org/abs/2605.24618
作者:Yoonhyung Lee,Hyunsin Park,Jinhwan Park,Jinkyu Lee
备注:Accepted to ACL 2026 (Main Conference). 20 pages, 8 figures, 7 tables. Demo page: https://qualcomm-ai-research.github.io/fc-tts
摘要:zero-shot文语转换(TTS)技术的最新进展使得能够在说话风格和说话者音色方面准确地模仿参考语音。然而,从单独的参考文献中实现对这些方面的分离控制仍然是一项具有挑战性的任务。几项研究已经提出了将语音分解为可解释的属性(例如,音色,韵律和内容),提供了一个有前途的基础TTS属性控制从单独的参考。然而,如何有效地将这种表示集成到TTS系统中,以实现独立和精确的控制仍然有待探索。在本文中,我们提出了FC-TTS,一个zero-shot的TTS框架,使分离控制的风格和音色上的两个不同的参考话语的条件。与现有的系统继承了这些预先训练的解纠缠表示的限制不同,FC-TTS引入了关键的设计策略,包括架构选择,训练框架和辅助训练目标,这些策略提高了属性分离和双参考控制的可靠性。实验表明,FC-TTS实现了高保真合成和竞争性的zero-shot自然度,同时独特地支持风格和音色的一致和独立操纵。音频样本可在https://qualcomm-ai-research.github.io/fc-tts上获得
摘要:Recent advances in zero-shot text-to-speech (TTS) have enabled accurate imitation of reference speech in terms of both speaking style and speaker timbre. However, achieving disentangled control over these aspects from separate references remains a challenging task. Several studies have proposed disentangled speech representations that decompose speech into interpretable attributes (e.g., timbre, prosody, and content), providing a promising foundation for TTS with attribute control from separate references. Yet, how to effectively integrate such representations into TTS systems to achieve independent and precise control remains underexplored. In this paper, we present FC-TTS, a zero-shot TTS framework that enables disentangled control of style and timbre by conditioning on two distinct reference utterances. Unlike existing systems that inherit limitations from those pre-trained disentangled representations, FC-TTS introduces key design strategies, including architectural choices, training framework, and auxiliary training objectives, which improve the reliability of attribute separation and dual-reference control. Experiments show that FC-TTS achieves high-fidelity synthesis and competitive zero-shot naturalness, while uniquely supporting consistent and independent manipulation of style and timbre. Audio samples are available at https://qualcomm-ai-research.github.io/fc-tts


【10】Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization

标题:塔卡出席KSBA-2026任务2:阿拉伯语语音变音化的常规微调
链接:https://arxiv.org/abs/2605.25928
作者:Meshal Alamr,Hassan Alqaeri,Abdullah Aldahlawi
备注:4 pages, 1 figure. Published in Proceedings of OSACT7 (LREC 2026). Winning system for KSAA-2026 Task 2 on Arabic Speech Diacritization
摘要:我们描述了KSAA-2026阿拉伯语语音听写与自动变音的共享任务的任务2的获奖系统。该任务需要从语音音频和未变音的转录本中生成完全变音的阿拉伯语文本,只有2,327个训练样本,并且不允许使用外部数据。我们的系统微调CATT-Whisper,这是一种字符级多模态模型,将预训练的CATT文本编码器与冻结的Whisper语音编码器相结合。我们方法的关键是训练正则化:R-Drop一致性正则化,Optuna优化的高权重衰减超参数和焦点损失。在推理时,我们在softmax概率水平上使用Monte Carlo Dropout对四个模型检查点的200次随机向前传递进行平均。该系统在主要排行榜指标上实现了23.26%的WER(带有案例结尾,包括无变音符号位置),在所有参与者中排名第一。
摘要:We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires producing fully diacritized Arabic text from speech audio and undiacritized transcripts, with only 2,327 training samples available and no external data permitted. Our system fine-tunes CATT-Whisper, a character-level multimodal model combining a pretrained CATT text encoder with a frozen Whisper speech encoder. The key to our approach is training regularization: R-Drop consistency regularization, Optuna-optimized hyperparameters with high weight decay, and Focal Loss. At inference, we average 200 stochastic forward passes across four model checkpoints using Monte Carlo Dropout at the softmax probability level. The system achieves 23.26% WER on the primary leaderboard metric (with case endings, including no-diacritic positions), placing 1st among all participants.


【11】Proactive for Uncertainty: Cause-Aware Error Diagnosis and Interactive Clarification for Spoken Dialogue Systems

标题:主动应对不确定性:口语对话系统的原因感知错误诊断和交互式澄清
链接:https://arxiv.org/abs/2605.25404
作者:Yizhou Peng,Ziyang Ma,Changsong Liu,Yi-Wen Chao,Xie Chen,Eng Siong Chng
摘要:级联自动语音识别-大语言模型(ASR-LLM)流水线在工业口语对话系统(SDS)中仍然很受欢迎,主要是因为它们的解耦设计确保了感知可验证性。然而,级联系统遭受错误传播,因为转录失败不可避免地级联到后续组件,从而降低最终的交互质量。虽然ASR置信度分数为不可靠的输入提供了一个简单的过滤器,但这种方法从根本上是有限的,因为它通常无法检测删除错误或区分声学(无法听清楚)和语言(无法理解)的不匹配,这两者都需要有针对性的恢复策略。在本文中,我们提出了一个原因感知的错误恢复范式,从根本上重新思考SDS的鲁棒性。与传统的置信度过滤不同,我们引入了一套小型的精确聚焦检测器,这些检测器利用深度ASR潜在表示将标记级错误分解为感知、理解和删除失败。这种细粒度的诊断智能使LLM能够编排有针对性的多轮澄清策略,有效地将模糊信号转化为无缝的用户交互。实验结果验证了我们方法的精确度,与基线相比,域转移错误的召回率增加了一倍多(57.96%与23.66%)。至关重要的是,这种诊断精度可以使WER减少30%,并在不同口音,失真和领域的下游任务上提高17%。
摘要:Cascaded Automatic Speech Recognition -- Large Language Model (ASR-LLM) pipelines remain popular for industrial Spoken Dialogue Systems (SDS), primarily because their decoupled design ensures perceptual verifiability. However, cascaded systems suffer from error propagation, as transcription failures inevitably cascade to subsequent components, thereby degrading the final interaction quality. Although ASR confidence scores offer a simple filter for unreliable inputs, this approach is fundamentally limited because it typically fails to detect deletion errors or to distinguish between acoustic (inability to hear clearly) and linguistic (inability to understand) mismatches, both of which require targeted recovery strategies. In this paper, we propose a cause-aware error recovery paradigm that fundamentally rethinks robustness in SDS. Unlike traditional confidence filtering, we introduce a suite of small precision-focused detectors that exploit deep ASR latent representations to disentangle token-level errors into perception, comprehension, and deletion failures. This fine-grained diagnostic intelligence empowers the LLM to orchestrate targeted, multi-turn clarification strategies, effectively transforming ambiguous signals into seamless user interactions. Experimental results validate the precision of our approach, which more than doubles the recall on domain-shift errors (57.96% vs. 23.66%) compared to baselines. Crucially, this diagnostic precision yields up to a 30% reduction in WER and a 17% improvement on the downstream task across diverse accents, distortions, and domains.


【12】Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

标题:从语音中Zero-Shot帕金森病检测:比较大型音频和语言模型
链接:https://arxiv.org/abs/2605.24806
作者:Muhammad Ashad Kabir,Sirajam Munira
备注:6 pages
摘要:大型音频和语言模型最近已经在各个领域展示了zero-shot推理能力。然而,目前尚不清楚音频输入的形式,无论是从语音中提取的手工声学特征还是原始音频波形本身,如何影响不同语言的帕金森病(PD)检测性能。在这项研究中,我们系统地比较了两种输入模式的zero-shot PD检测:(i)手工制作的声学特征提取的语音记录分析的通用LLM,和(ii)直接波形输入分析的音频模型。在四种语言的PD语音数据集上的实验表明,性能在输入方式、语音任务和语言之间存在差异。手工制作的声学特征在低资源语言中提供更稳定的性能(例如,孟加拉语),而音频输入产生依赖于语音的增益。这些发现突出了输入模态对从语音中检测出zero-shot PD的影响。
摘要:Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted from speech or the raw audio waveform itself, affects performance for Parkinson's disease (PD) detection across different languages. In this study, we systematically compare two input modalities for zero-shot PD detection: (i) handcrafted acoustic features extracted from speech recordings analyzed by a general-purpose LLM, and (ii) direct waveform input analyzed by audio-capable models. Experiments on PD speech datasets in four languages show that performance varies across input modalities, speech tasks, and languages. Handcrafted acoustic features provide more stable performance in a low-resource language (e.g., Bengali), whereas audio input yields dataset-dependent gains. These findings highlight the impact of input modality on zero-shot PD detection from speech.


【13】A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

标题:临床访谈抑郁检测基准的多探头审计
链接:https://arxiv.org/abs/2605.23977
作者:Takehiro Ishikawa,Jon Duke
摘要:本文通过DAIC/E-DAIC,CMDC,ANDROIDS,MODMA和PDCH的四个互补探针审计临床访谈抑郁症检测中的基准评估。首先,我们重新评估E-DAIC下严格的主题不相交留一主题交叉验证。一个轻量级的混合文本加LLM分数模型达到了macro-F1 = 0.723 -据我们所知,这是该协议下报告的最高值-提供了一个保守的不依赖于特权官方坚持的参考点。其次,我们测试E-DAIC官方分裂是否支持细粒度的排行榜排名,通过扫描96个模型配置跨模态捆绑,池化策略和学习器。开发侧交叉验证和官方测试排名仅适度一致:最佳交叉验证配置在官方测试中排名第20,官方测试获胜者在交叉验证中排名第41,前3名重叠为零,只有32.3%的主题引导中明显获胜者排名第1。第三,我们在外部验证强大的公共CMDC和ANDROIDS基线,以实现接近上限的域内性能。Zero-shot迁移到外部语料库是相当弱的。最后,我们使用由基于SRDS的注释器定义的成对的高密度与低密度访谈切片对E-DAIC文本和音频模型进行压力测试。文本分数在高密度切片上急剧上升,而音频分数几乎保持不变;文本减去音频的差距在所有五个种子中都是正的。
摘要:This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS, MODMA, and PDCH. First, we re-evaluate E-DAIC under strict subject-disjoint leave-one-subject-out cross-validation. A lightweight hybrid text-plus-LLM-score model reaches macro-F1 = 0.723 - the highest reported under this protocol, to our knowledge - providing a conservative out-of-fold reference point that does not depend on the privileged official holdout. Second, we test whether the E-DAIC official split supports fine-grained leaderboard rankings by sweeping 96 model configurations across modality bundles, pooling strategies, and learners. Development-side cross-validation and official-test rankings align only moderately: the best cross-validation configuration ranks twentieth on the official test, the official-test winner ranks forty-first by cross-validation, top-3 overlap is zero, and the apparent winner is rank-1 in only 32.3% of subject bootstraps. Third, we externally validate strong public CMDC and ANDROIDS baselines that achieve near-ceiling in-domain performance. Zero-shot transfer to external corpora is substantially weaker. Finally, we stress-test E-DAIC text and audio models using paired symptom-dense versus symptom-light interview slices defined by an SRDS-based annotator. Text scores rise sharply on symptom-dense slices, whereas audio scores remain nearly flat; the text-minus-audio gap is positive across all five seeds.


机器翻译由腾讯交互翻译提供,仅供参考