今日论文合集:cs.SD语音8篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】A Toolkit for Detecting Spurious Correlations in Speech Datasets
标题:检测语音数据集中杂散相关性的工具包
链接:https://arxiv.org/abs/2604.26676
作者:Lara Gauder,Pablo Riera,Andrea Slachevsky,Gonzalo Forno,Adolfo M. García,Luciana Ferrer
摘要:我们介绍了一个工具包,用于发现语音数据集的记录特征和目标类之间的虚假相关性。由于异构记录条件,可能会出现虚假相关性,这是健康相关数据集的常见情况。当存在于训练和测试数据中时,这些相关性会导致系统性能的高估-这是一种危险的情况,特别是在需要系统满足最低性能要求的高风险应用中。我们的工具包实现了一个诊断方法的基础上检测的目标类只使用音频中的非语音区域。在这项任务中优于机会的表现表明,可以从非语音区域提取关于目标类的信息,标记虚假相关的存在。该工具包可供研究使用。
摘要:We introduce a toolkit for uncovering spurious correlations between recording characteristics and target class in speech datasets. Spurious correlations may arise due to heterogeneous recording conditions, a common scenario for health-related datasets. When present both in the training and test data, these correlations result in an overestimation of the system performance -- a dangerous situation, specially in high-stakes application where systems are required to satisfy minimum performance requirements. Our toolkit implements a diagnostic method based on the detection of the target class using only the non-speech regions in the audio. Better than chance performance at this task indicates that information about the target class can be extracted from the non-speech regions, flagging the presence of spurious correlations. The toolkit is publicly available for research use.


【2】Full band denoising of room impulse response in the wavelet domain with dictionary learning

标题:基于字典学习的子波域房间脉冲响应全带去噪
链接:https://arxiv.org/abs/2604.26669
作者:Théophile Dupré,Romain Couderc,Miguel Moleron,Axel Coulon,Rémy Bruno,Arnaud Laborie
摘要:传统的小波域房间脉冲响应去噪方法依赖于对细节系数进行阈值化,这不适合低频情况。在这项工作中,我们引入了一个基于小波的后处理算法,扩展去噪近似系数的稀疏字典学习与随时间变化的误差容限。该方法利用指数衰减包络模型,根据局部信噪比来调整重建精度。与基线方法相比,该方法显著改善了合成和测量的房间脉冲响应的低频去噪,从而更准确地估计声学参数,例如衰减时间。
摘要:Conventional wavelet-domain methods for room impulse response denoising rely on thresholding detail coefficients, which is unsuited for low frequencies. In this work, we introduce a wavelet-based post-processing algorithm that extends denoising to approximation coefficients by means of sparse dictionary learning with a time-varying error tolerance. The proposed method leverages an exponential decay envelope model to adapt reconstruction accuracy according to the local signal-to-noise ratio. This approach significantly improves low-frequency denoising of synthetic and measured room impulse responses compared to the baseline method, leading to more accurate estimation of acoustic parameters such as decay time.


【3】Diffusion Reconstruction towards Generalizable Audio Deepfake Detection

标题:面向可推广音频Deepfake检测的扩散重建
链接:https://arxiv.org/abs/2604.26465
作者:Bo Cheng,Songjun Cao,Xiaoming Zhang,Jie Chen,Long Ma,Fei Chen
备注:5 pages, this paper was submitted to Interspeech2026 for review
摘要:在生成模型的快速演变驱动下,实现针对不可见攻击的鲁棒泛化仍然是音频深度伪造检测(ADD)中的一个挑战。为了解决这个问题,我们提出了一个以硬样本分类为中心的框架。其核心思想是,一个能够区分具有挑战性的硬样本的模型本身就具备有效处理简单情况的能力。我们研究了多种重建范例,确定了基于扩散的方法作为生成硬样本的最佳方法。此外,我们利用多层特征聚合并引入正则化辅助对比学习(RACL)目标来增强泛化能力。实验证明了我们的方法具有优越的泛化能力,与基线相比,我们的最佳模型实现了平均等错误率(EER)的显着降低。
摘要:Achieving robust generalization against unseen attacks remains a challenge in Audio Deepfake Detection (ADD), driven by the rapid evolution of generative models. To address this, we propose a framework centered on hard sample classification. The core idea is that a model capable of distinguishing challenging hard samples is inherently equipped to handle simpler cases effectively. We investigate multiple reconstruction paradigms, identifying the diffusion-based method as optimal for generating hard samples. Furthermore, we leverage multi-layer feature aggregation and introduce a Regularization-Assisted Contrastive Learning (RACL) objective to enhance generalizability. Experiments demonstrate the superior generalization of our approach, with our best model achieving a significant reduction in the average Equal Error Rate (EER) compared to the baseline.


【4】EmoTransCap: Dataset and Pipeline for Emotion Transition-Aware Speech Captioning in Discourses

标题:SYS Transcap:话语中情感转变感知语音字幕的数据集和管道
链接:https://arxiv.org/abs/2604.26417
作者:Shuhao Xu,Yifan Hu,Jingjing Wu,Zhihao Du,Zheng Lian,Rui Liu
备注:15 pages, 5 figures, including appendix
摘要:情感感知和自适应表达是人与智能体交互的基本能力。虽然语音情感字幕(SEC)的最新进展,提高了细粒度的情感建模,现有的系统仍然局限于静态的,孤立的句子内的单一情感表征,忽略了动态的情感转换在话语水平。为了解决这一差距,我们提出了情感转换感知语音字幕(EMOTION transCap),一个范例,集成了时间的情感动态与话语级的语音描述。为了构建一个情感过渡丰富的数据集,同时实现可扩展的扩展,我们设计了一个自动化的管道来创建数据集。这是第一个明确设计用于捕捉话语级情感转换的大规模数据集。为了生成语义丰富的描述,我们将声学属性和时间线索从话语级语音。我们的多任务情感转换识别(MTETR)模型执行联合情感转换检测和日记。利用LLM的语义分析能力,我们产生了两个注释版本:描述性和解释性。这些数据和注释为提高情感感知和情感表达提供了宝贵的资源。该数据集支持捕捉情感转变的语音字幕,促进了时间动态和细粒度的情感理解。我们还介绍了一个可控的,过渡意识的情感语音合成系统在话语水平,提高拟人化的情感表达能力和支持情感智能会话代理。
摘要:Emotion perception and adaptive expression are fundamental capabilities in human-agent interaction. While recent advances in speech emotion captioning (SEC) have improved fine-grained emotional modeling, existing systems remain limited to static, single-emotion characterization within isolated sentences, neglecting dynamic emotional transitions at the discourse level. To address this gap, we propose Emotion Transition-Aware Speech Captioning (EmoTransCap), a paradigm that integrates temporal emotion dynamics with discourse-level speech description. To construct a dataset rich in emotion transitions while enabling scalable expansion, we design an automated pipeline for dataset creation. This is the first large-scale dataset explicitly designed to capture discourse-level emotion transitions. To generate semantically rich descriptions, we incorporate acoustic attributes and temporal cues from discourse-level speech. Our Multi-Task Emotion Transition Recognition (MTETR) model performs joint emotion transition detection and diarization. Leveraging the semantic analysis capabilities of LLMs, we produce two annotation versions: descriptive and instruction-oriented. These data and annotations offer a valuable resource for advancing emotion perception and emotional expressiveness. The dataset enables speech captions that capture emotional transitions, facilitating temporal-dynamic and fine-grained emotion understanding. We also introduce a controllable, transition-aware emotional speech synthesis system at the discourse level, enhancing anthropomorphic emotional expressiveness and supporting emotionally intelligent conversational agents.


【5】Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech

标题:基于复发的非线性声乐动力学作为从对话言语中检测抑郁症的数字生物标志物
链接:https://arxiv.org/abs/2604.26242
作者:Himadri S Samanta
备注:12 pages, 5 figures
摘要:抑郁症的数字生物标志物在很大程度上依赖于静态声学描述符,汇总汇总统计数据或传统的机器学习表示。这样的方法可能会错过嵌入在会话语音动态的非线性时间组织。我们假设抑郁症与发声状态轨迹中的复发结构改变有关,反映了发声系统如何随时间重新访问声学状态的变化。使用DAIC-WOZ语料库的抑郁子集与142个标记的参与者,我们建模帧级COVAREP轨迹作为非线性动力系统,并从74个声乐通道推导出基于递归的生物标志物。Logistic回归与特征选择和分层交叉验证评估分类性能。基于复发的生物标志物实现了0.689的平均交叉验证AUC,超过静态声学基线、熵动力学特征、赫斯特指数特征、确定性特征和Lyapunov样不稳定性代理。排列检验表明具有统计学显著性,$p=0.004$。合并的交叉验证预测结果为AUC 0.665,95% bootstrap置信区间为[0.568,0.758]。这些研究结果表明,抑郁症的特点可能是改变复发结构的对话声乐动力学和支持非线性状态空间分析作为一个有前途的方向数字精神生物标志物。
摘要:Digital biomarkers for depression have largely relied on static acoustic descriptors, pooled summary statistics, or conventional machine learning representations. Such approaches may miss nonlinear temporal organization embedded in conversational vocal dynamics. We hypothesized that depression is associated with altered recurrence structure in vocal state trajectories, reflecting changes in how the vocal system revisits acoustic states over time. Using the depression subset of the DAIC-WOZ corpus with 142 labeled participants, we modeled frame-level COVAREP trajectories as nonlinear dynamical systems and derived recurrence-based biomarkers from 74 vocal channels. Logistic regression with feature selection and stratified cross-validation evaluated classification performance. Recurrence-based biomarkers achieved a mean cross-validated AUC of 0.689, exceeding static acoustic baselines, entropy-dynamics features, Hurst exponent features, determinism features, and Lyapunov-like instability proxies. Permutation testing indicated statistical significance with $p=0.004$. Pooled cross-validated predictions yielded AUC 0.665 with a 95\% bootstrap confidence interval of [0.568, 0.758]. These findings suggest that depression may be characterized by altered recurrence structure in conversational vocal dynamics and support nonlinear state-space analysis as a promising direction for digital psychiatric biomarkers.


【6】Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model

标题:使用MFCC特征和基于LSTM的深度学习模型的语音情感识别
链接:https://arxiv.org/abs/2604.25938
作者:Adelekun Oluwademilade,Ademola Adedamola,Abiola Abdulhakeem,Akinpelu Azeezat,Eraiyetan Israel,Omotosho Oluwadunsin,Ibenye Ikechukwu,Ayuba Muhammad,Olusanya Olamide,Kamorudeen Amuda
摘要:语音情感识别(Speech Emotion Recognition,SER)是利用机器来检测人类基于语音的情感状态,在自然的人机交互中越来越重要。言语是一种非常有价值的信息来源,因为情绪会改变言语的模式;音调,能量甚至时机。尽管如此,SER并不是一件容易的事情,因为说话者并不是恒定的,而且录音时的情况也会有所不同,声音之间的相似性也会有所不同。在这项工作中,作者介绍了一个语音情感识别系统依赖于梅尔频率倒谱系数和长短期记忆(LSTM)神经网络,作为特征提取方法。对多伦多情感语音集(TESS)语音信号进行预处理,并将其转换为MFCC特征,以了解时间方面的重要方面。然后将得到的特征引入LSTM模型,该模型能够学习序列音频数据的长期特征。经过训练的模型在数据集中出现的几个情感类上进行测量。如实验结果所示,所提出的MFCC-LSTM方法成功地捕获了语音中的情感模式,并在所有选定的情感分类中提供了高度真实的分类。本研究提出一个以Mel频率倒谱系数(MFCC)为特征的语音情感识别系统和一个深度学习LSTM分类器。具有RBF核的支持向量机(SVM)作为经典基线,实现了98%的准确度,并验证了所提出的LSTM模型,实现了99%的准确度。总的来说,可以确认基于LSTM的架构可以用于解决语音情感识别的任务。所提出的系统的实际应用可能是虚拟助理和心理健康监测。
摘要:Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in natural human-computer interaction. Speech is a very valuable source of information, as emotions modify the patterns of speech; pitch, energy and even timing. Nonetheless, SER is not an easy task because speakers are not constant, and situations vary when recording and the sound similarity between specific feelings. In this work, the author introduces a speech emotion recognition system relying on the Mel-Frequency Cepstral Coefficient and Long Short-Term Memory (LSTM) neural network, as a feature extraction method. The Toronto Emotional Speech Set (TESS) speech signal was pre-processed, and transformed into MFCC features to understand the important aspects in terms of time. The resultant features were then introduced to LSTM model, which is able to learn long term features of sequential audio data. The trained model was measured over several emotion classes occurring in the dataset. As seen in the results of experiments, the proposed MFCC-LSTM approach succeeds in capturing the patterns of emotions in speech and provides highly realistic classifications in all the chosen emotion classifications. This study presents a speech emotion recognition system using Mel-Frequency Cepstral Coefficients (MFCCs) as features and a deep learning LSTM classifier. A Support Vector Machine (SVM) with an RBF kernel served as a classical baseline, achieving 98% accuracy, against which the proposed LSTM model, achieving 99% accuracy, was validated. Overall, it is possible to confirm that LSTM-based architectures can be used to address the task of speech emotion recognition. Actual applications of the proposed system may be virtual assistants and mental health surveillance.


【7】DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

标题:DiffAnon:基于扩散的韵律控制语音识别
链接:https://arxiv.org/abs/2604.26281
作者:Ismail Rasim Ulgen,Zexin Cai,Nicholas Andrews,Philipp Koehn,Berrak Sisman
备注:Submitted to Interspeech 2026
摘要:保留或不保留韵律是语音匿名化的核心问题。韵律传达意义和情感,但与说话人身份紧密相连。现有的方法要么放弃韵律隐私或缺乏一个原则性的机制来控制效用隐私的权衡,在固定的设计点操作。我们提出了DiffAnon,一种基于扩散的匿名化方法,具有无分类器指导(CFG),提供了对韵律保留的明确的,连续的推理时间控制。DiffAnon在RVQ编解码器的语义嵌入上细化声学细节,在单个模型内实现匿名化强度和韵律保真度之间的平滑插值。据我们所知,这是第一个语音匿名框架,提供结构化的,可插值的推理时间韵律控制。实验证明结构化的权衡行为,实现强大的效用,同时保持竞争力的隐私在可控的操作点。
摘要:To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy or lack a principled mechanism to control the utility-privacy trade-off, operating at fixed design points. We propose DiffAnon, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation. DiffAnon refines acoustic detail over semantic embeddings of an RVQ codec, enabling smooth interpolation between anonymization strength and prosodic fidelity within a single model. To the best of our knowledge, it is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control. Experiments demonstrate structured trade-off behavior, achieving strong utility while maintaining competitive privacy across controllable operating points.


【8】SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

标题:SongBench:歌曲质量评估的细粒度多方面基准
链接:https://arxiv.org/abs/2604.25937
作者:Dapeng Wu,Shun Lei,Wei Tan,Guangzheng Li,Yunzhe Wang,Huaicheng Zhang,Lishi Zuo,Zhiyong Wu
摘要:文本到歌曲生成的最新进展使现实的音乐内容制作成为可能,但现有的评估基准缺乏捕捉多维美学细微差别的专业粒度。在本文中,我们提出了SongBench,一个专门的框架,细粒度的歌曲评估在七个关键方面:声乐,乐器,旋律,结构,安排,混合,和音乐性。利用这个框架,我们构建了一个专家注释的数据库,包括11,717个来自最先进的模型的样本,由音乐专业人士标记。大量的实验结果表明,SongBench实现了与专家评分的高度相关性。通过揭示当前最先进模型中的细粒度性能差距,SongBench可作为诊断基准,引导开发更专业和音乐连贯的歌曲生成。
摘要:Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose SongBench, a specialized framework for fine-grained song assessment across seven key dimensions: Vocal, Instrument, Melody, Structure, Arrangement, Mixing, and Musicality. Utilizing this framework, we construct an expert-annotated database comprising 11,717 samples from state-of-the-art models, labeled by music professionals. Extensive experimental results demonstrate that SongBench achieves high correlation with expert ratings. By revealing fine-grained performance gaps in current state-of-the-art models, SongBench serves as a diagnostic benchmark to steer the development toward more professional and musically coherent song generation.


eess.AS音频处理


【1】The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
标题:假共振:语音生成评估中情感嵌入相似性的批判性检查
链接:https://arxiv.org/abs/2604.26347
作者:Yun-Shao Tsai,Yi-Cheng Lin,Huang-Cheng Chou,Tzu-Wen Hsu,Yun-Man Hsu,Chun Wei Chen,Shrikanth Narayanan,Hung-yi Lee
备注:Submitted to Interspeech 2026
摘要:情感表现力的客观度量对于语音生成是至关重要的,特别是在需要情感韵律转移的表达性合成和语音转换中。为了量化这一点,该领域广泛依赖于参考样本和生成样本之间的情感相似性。这种方法计算来自emotion2vec等编码器的嵌入的余弦相似性,假设它们捕获情感线索,尽管语言和说话者的变化。我们通过受控对抗任务和人类对齐测试来挑战这一假设。尽管分类精度高,这些潜在空间是不适合的zero-shot相似性评价。表征的局限性导致语言和说话者的干扰掩盖了情感特征,降低了辨别能力。因此,该指标与人类感知不一致。这种听觉上的脆弱性表明,它会奖励声音模仿,而不是真正的情感合成。
摘要:Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on emotion similarity between reference and generated samples. This approach computes cosine similarity of embeddings from encoders like emotion2vec, assuming they capture affective cues despite linguistic and speaker variations. We challenge this assumption through controlled adversarial tasks and human alignment tests. Despite high classification accuracy, these latent spaces are unsuitable for zero-shot similarity evaluation. Representational limitations cause linguistic and speaker interference to overshadow emotional features, degrading discriminative ability. Consequently, the metric misaligns with human perception. This acoustic vulnerability reveals it rewards acoustic mimicry over genuine emotional synthesis.


【2】Dual-LoRA: Parameter-Efficient Adversarial Disentanglement for Cross-Lingual Speaker Verification

标题:Dual-LoRA:跨语言说话人验证的参数高效对抗解纠缠
链接:https://arxiv.org/abs/2604.26327
作者:Qituan Shangguan,Junhao Du,Kunyang Peng,Feng Xue,Hui Zhang,Xinsheng Wang,Kai Yu,Shuai Wang
备注:Submitted to Interspeech 2026; 5 pages
摘要:跨语言说话人确认存在严重的语言-说话人纠缠问题。这导致了最困难情况下的系统退化:正确接受来自不同语言的同一说话者的话语,同时拒绝来自共享同一语言的不同说话者的话语。标准的对抗性解纠缠降低了说话者的辨别能力;盲辨别器无意中惩罚了仅仅与语言相关的说话者辨别特征。为了解决这个问题,我们提出了Dual-LoRA,将可训练的任务分解LoRA适配器注入到冻结的预训练骨干中。我们的核心创新是一个基于锚定的Advertisement:通过将Advertisement与一个明确的语言分支联系起来,对抗梯度针对真实的语言线索,而不是任意的相关性,保留了说话者的基本特征。在TidyVoice基准测试中,我们的系统实现了0.91%的验证EER,并在官方挑战赛中获得第三名。
摘要:Cross-lingual speaker verification suffers from severe language-speaker entanglement. This causes systematic degradation in the hardest scenario: correctly accepting utterances from the same speaker across different languages while rejecting those from different speakers sharing the same language. Standard adversarial disentanglement degrades speaker discriminability; blind discriminators inadvertently penalize speaker-discriminative traits that merely correlate with language. To address this, we propose Dual-LoRA, injecting trainable task-factorized LoRA adapters into a frozen pre-trained backbone. Our core innovation is a Language-Anchored Adversary: by grounding the discriminator with an explicit language branch, adversarial gradients target true linguistic cues rather than arbitrary correlations, preserving essential speaker characteristics. Evaluated on the TidyVoice benchmark, our system achieves a 0.91% validation EER and achieves 3rd place in the official challenge.


【3】SPG-Codec: Exploring the Role and Boundaries of Semantic Priors in Ultra-Low-Bitrate Neural Speech Coding

标题:SPG-Codec:探索超低比特率神经语音编码中语义先验的作用和边界
链接:https://arxiv.org/abs/2604.26296
作者:Mingyu Zhao,Zijian Lin,Kun Wei,Zhiyong Wu
备注:6 pages, 6 figures, accepted to ICME 2026
摘要:传统的神经语音编解码器在超低比特率下遭受严重的可懂度退化,其中瓶颈从声学失真转变为语义丢失。为了解决这个问题,本文进行了系统的调查的作用和基本的限制整合冻结的语义先验-特别是休伯特和耳语-到神经语音编码。我们介绍并定量验证了一种新的语义退休现象:虽然语义约束降低了字错误率(WER)高达10%,相对于1.5 kbps的,他们的好处迅速减少超过6 kbps,表明一个实际的容量边界。我们进一步揭示了不同先验类型之间的明显权衡:声学丰富先验(HuBERT)更好地保留了韵律和音色细节,而高级语言先验(Whisper)有效地抑制了嘈杂环境中的语音幻觉(将幻觉率降低了26%),并大大缩小了未见过的扬声器的泛化差距。基于这些发现,我们提出了一种比特率感知的调节策略,动态调整先验强度,以优化语义一致性和感知自然性之间的权衡。广泛的实验评估证实,与现有基线相比,我们的方法实现了有竞争力的清晰度和噪声鲁棒性,为超低比特率生成语音编码提供了一条原则性途径。
摘要:Conventional neural speech codecs suffer from severe intelligibility degradation at ultra-low bitrates, where the bottleneck transitions from acoustic distortion to semantic loss. To address this issue, this paper conducts a systematic investigation into the role and fundamental limits of integrating frozen semantic priors -- specifically HuBERT and Whisper -- into neural speech coding. We introduce and quantitatively validate a novel Semantic Retirement phenomenon: while semantic constraints reduce the Word Error Rate (WER) by up to ~10% relatively at 1.5 kbps, their benefits rapidly diminish beyond 6 kbps, indicating a practical capacity boundary. We further uncover a clear trade-off between different prior types: acoustic-rich priors (HuBERT) better preserve prosodic and timbral details, whereas high-level linguistic priors (Whisper) effectively suppress phonetic hallucinations in noisy environments (reducing hallucination rates by 26 percent) and substantially narrow the generalization gap for unseen speakers. Building on these findings, we propose a bitrate-aware regulation strategy that dynamically adjusts prior strength to optimize the trade-off between semantic consistency and perceptual naturalness. Extensive experimental evaluations confirm that our approach achieves competitive intelligibility and noise robustness compared to existing baselines, offering a principled pathway toward ultra-low-bitrate generative speech coding.


【4】DiffAnon: Diffusion-based Prosody Control for Voice Anonymization

标题:DiffAnon:基于扩散的韵律控制语音识别
链接:https://arxiv.org/abs/2604.26281
作者:Ismail Rasim Ulgen,Zexin Cai,Nicholas Andrews,Philipp Koehn,Berrak Sisman
备注:Submitted to Interspeech 2026
摘要:保留或不保留韵律是语音匿名化的核心问题。韵律传达意义和情感,但与说话人身份紧密相连。现有的方法要么放弃韵律隐私或缺乏一个原则性的机制来控制效用隐私的权衡,在固定的设计点操作。我们提出了DiffAnon,一种基于扩散的匿名化方法,具有无分类器指导(CFG),提供了对韵律保留的明确的,连续的推理时间控制。DiffAnon在RVQ编解码器的语义嵌入上细化声学细节,在单个模型内实现匿名化强度和韵律保真度之间的平滑插值。据我们所知,这是第一个语音匿名框架,提供结构化的,可插值的推理时间韵律控制。实验证明结构化的权衡行为,实现强大的效用,同时保持竞争力的隐私在可控的操作点。
摘要:To preserve or not to preserve prosody is a central question in voice anonymization. Prosody conveys meaning and affect, yet is tightly coupled with speaker identity. Existing methods either discard prosody for privacy or lack a principled mechanism to control the utility-privacy trade-off, operating at fixed design points. We propose DiffAnon, a diffusion-based anonymization method with classifier-free guidance (CFG) that provides explicit, continuous inference-time control over prosody preservation. DiffAnon refines acoustic detail over semantic embeddings of an RVQ codec, enabling smooth interpolation between anonymization strength and prosodic fidelity within a single model. To the best of our knowledge, it is the first voice anonymization framework to provide structured, interpolatable inference-time prosody control. Experiments demonstrate structured trade-off behavior, achieving strong utility while maintaining competitive privacy across controllable operating points.


【5】One Voice, Many Tongues: Cross-Lingual Voice Cloning for Scientific Speech

标题:一种声音,多种语言:科学言语的跨语言语音克隆
链接:https://arxiv.org/abs/2604.26136
作者:Amanuel Gizachew Abebe,Yasmin Moslem
备注:IWSLT 2026
摘要:在用不同语言生成语音的同时保持说话者的语音身份仍然是口语技术的基本挑战,特别是在科学交流等专业领域。在本文中,我们通过我们的系统提交给国际会议口语翻译(IWITH 2026),跨语言语音克隆共享任务来应对这一挑战。首先,我们评估了几个国家的最先进的语音克隆模型的阿拉伯语,中文和法语的科学文本的跨语言语音生成。然后,基于OmniVoice基础模型构建了语音克隆系统。我们采用数据增强,通过多模型集成蒸馏从ACL 60/60语料库。我们研究了使用这种合成数据进行微调的效果,在保持说话者相似性的同时,展示了跨语言可懂度(WER和CER)的一致改善。
摘要:Preserving a speaker's voice identity while generating speech in a different language remains a fundamental challenge in spoken language technology, particularly in specialized domains such as scientific communication. In this paper, we address this challenge through our system submission to the International Conference on Spoken Language Translation (IWSLT 2026), the Cross-Lingual Voice Cloning shared task. First, we evaluate several state-of-the-art voice cloning models for cross-lingual speech generation of scientific texts in Arabic, Chinese, and French. Then, we build voice cloning systems based on the OmniVoice foundation model. We employ data augmentation via multi-model ensemble distillation from the ACL 60/60 corpus. We investigate the effect of using this synthetic data for fine-tuning, demonstrating consistent improvements in intelligibility (WER and CER) across languages while preserving speaker similarity.


【6】Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection

标题:Deepfake音频检测的监督对比学习中的相似性选择和负缩放
链接:https://arxiv.org/abs/2604.26057
作者:Jaskirat Sudan,Hashim Ali,Surya Subramani,Hafiz Malik
摘要:监督对比学习(SupCon)被广泛用于形状表示,但对音频深度伪造检测的针对性研究有限。现有的工作通常将对比术语与更广泛的管道相结合;然而,对SupCon本身的关注缺失。在这项工作中,我们对wav 2 vec 2 XLS-R(300 M)进行了一项对照研究,该研究改变了(i)SupCon中的相似性(余弦与来自超球面角的角度相似性)和(ii)使用热启动全局跨批次队列的负缩放。阶段1使用SupCon微调编码器和投影头;阶段2冻结它们并使用BCE训练线性分类器。在ASVspoof 2019 LA上进行训练,并在ASV 19 eval加上ITW和ASVspoof 2021 DF/LA上进行评估,具有延迟队列的余弦SupCon实现了最佳的ITW EER(8.29%)和合并EER(4.44),而角度相似性在没有排队否定(ITW 8.70)的情况下表现强劲,表明对大型否定集的依赖减少。
摘要:Supervised contrastive learning (SupCon) is widely used to shape representations, but has seen limited targeted study for audio deepfake detection. Existing work typically combines contrastive terms with broader pipelines; however, the focus on SupCon itself is missing. In this work, we run a controlled study on wav2vec2 XLS-R (300M) that varies (i) similarity in SupCon (cosine vs angular similarity derived from the hyperspherical angle) and (ii) negative scaling using a warm-started global cross-batch queue. Stage 1 fine-tunes the encoder and projection head with SupCon; Stage 2 freezes them and trains a linear classifier with BCE. Trained on ASVspoof 2019 LA and evaluated on ASV19 eval plus ITW and ASVspoof 2021 DF/LA, Cosine SupCon with a delayed queue achieves the best ITW EER (8.29%) and pooled EER (4.44), while angular similarity performs strongly without queued negatives (ITW 8.70), indicating reduced reliance on large negative sets.


【7】SongBench: A Fine-Grained Multi-Aspect Benchmark for Song Quality Assessment

标题:SongBench:歌曲质量评估的细粒度多方面基准
链接:https://arxiv.org/abs/2604.25937
作者:Dapeng Wu,Shun Lei,Wei Tan,Guangzheng Li,Yunzhe Wang,Huaicheng Zhang,Lishi Zuo,Zhiyong Wu
摘要:文本到歌曲生成的最新进展使现实的音乐内容制作成为可能,但现有的评估基准缺乏捕捉多维美学细微差别的专业粒度。在本文中,我们提出了SongBench,一个专门的框架,细粒度的歌曲评估在七个关键方面:声乐,乐器,旋律,结构,安排,混合,和音乐性。利用这个框架,我们构建了一个专家注释的数据库,包括11,717个来自最先进的模型的样本,由音乐专业人士标记。大量的实验结果表明,SongBench实现了与专家评分的高度相关性。通过揭示当前最先进模型中的细粒度性能差距,SongBench可作为诊断基准,引导开发更专业和音乐连贯的歌曲生成。
摘要:Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose SongBench, a specialized framework for fine-grained song assessment across seven key dimensions: Vocal, Instrument, Melody, Structure, Arrangement, Mixing, and Musicality. Utilizing this framework, we construct an expert-annotated database comprising 11,717 samples from state-of-the-art models, labeled by music professionals. Extensive experimental results demonstrate that SongBench achieves high correlation with expert ratings. By revealing fine-grained performance gaps in current state-of-the-art models, SongBench serves as a diagnostic benchmark to steer the development toward more professional and musically coherent song generation.


【8】Recurrence-Based Nonlinear Vocal Dynamics as Digital Biomarkers for Depression Detection from Conversational Speech

标题:基于复发的非线性声乐动力学作为从对话言语中检测抑郁症的数字生物标志物
链接:https://arxiv.org/abs/2604.26242
作者:Himadri S Samanta
备注:12 pages, 5 figures
摘要:抑郁症的数字生物标志物在很大程度上依赖于静态声学描述符,汇总汇总统计数据或传统的机器学习表示。这种方法可能会错过嵌入对话声音动态中的非线性时间组织。我们假设抑郁症与发声状态轨迹中的复发结构改变有关,反映了发声系统如何随时间重新访问声学状态的变化。使用DAIC-WOZ语料库的抑郁子集与142个标记的参与者,我们建模帧级COVAREP轨迹作为非线性动力系统,并从74个声乐通道推导出基于递归的生物标志物。Logistic回归与特征选择和分层交叉验证评估分类性能。基于复发的生物标志物实现了0.689的平均交叉验证AUC,超过静态声学基线、熵动力学特征、赫斯特指数特征、确定性特征和Lyapunov样不稳定性代理。排列检验表明具有统计学显著性,$p=0.004$。合并的交叉验证预测结果为AUC 0.665,95% bootstrap置信区间为[0.568,0.758]。这些研究结果表明,抑郁症的特点可能是改变复发结构的对话声乐动力学和支持非线性状态空间分析作为一个有前途的方向数字精神生物标志物。
摘要:Digital biomarkers for depression have largely relied on static acoustic descriptors, pooled summary statistics, or conventional machine learning representations. Such approaches may miss nonlinear temporal organization embedded in conversational vocal dynamics. We hypothesized that depression is associated with altered recurrence structure in vocal state trajectories, reflecting changes in how the vocal system revisits acoustic states over time. Using the depression subset of the DAIC-WOZ corpus with 142 labeled participants, we modeled frame-level COVAREP trajectories as nonlinear dynamical systems and derived recurrence-based biomarkers from 74 vocal channels. Logistic regression with feature selection and stratified cross-validation evaluated classification performance. Recurrence-based biomarkers achieved a mean cross-validated AUC of 0.689, exceeding static acoustic baselines, entropy-dynamics features, Hurst exponent features, determinism features, and Lyapunov-like instability proxies. Permutation testing indicated statistical significance with $p=0.004$. Pooled cross-validated predictions yielded AUC 0.665 with a 95\% bootstrap confidence interval of [0.568, 0.758]. These findings suggest that depression may be characterized by altered recurrence structure in conversational vocal dynamics and support nonlinear state-space analysis as a promising direction for digital psychiatric biomarkers.


【9】Speech Emotion Recognition Using MFCC Features and LSTM-Based Deep Learning Model

标题:使用MFCC特征和基于LSTM的深度学习模型的语音情感识别
链接:https://arxiv.org/abs/2604.25938
作者:Adelekun Oluwademilade,Ademola Adedamola,Abiola Abdulhakeem,Akinpelu Azeezat,Eraiyetan Israel,Omotosho Oluwadunsin,Ibenye Ikechukwu,Ayuba Muhammad,Olusanya Olamide,Kamorudeen Amuda
摘要:语音情感识别(Speech Emotion Recognition,SER)是利用机器来检测人类基于语音的情感状态,在自然的人机交互中越来越重要。言语是一种非常有价值的信息来源,因为情绪会改变言语的模式;音调,能量甚至时机。尽管如此,SER并不是一件容易的事情,因为说话者并不是恒定的,而且录音时的情况也会有所不同,声音之间的相似性也会有所不同。在这项工作中,作者介绍了一个语音情感识别系统依赖于梅尔频率倒谱系数和长短期记忆(LSTM)神经网络,作为特征提取方法。对多伦多情感语音集(TESS)语音信号进行预处理,并将其转换为MFCC特征,以了解时间方面的重要方面。然后将得到的特征引入LSTM模型,该模型能够学习序列音频数据的长期特征。经过训练的模型在数据集中出现的几个情感类上进行测量。如实验结果所示,所提出的MFCC-LSTM方法成功地捕获了语音中的情感模式,并在所有选定的情感分类中提供了高度真实的分类。本研究提出了一种使用梅尔倒谱系数(MFCC)作为特征和深度学习LSTM分类器的语音情感识别系统。具有RBF核的支持向量机(SVM)作为经典基线,实现了98%的准确度,并验证了所提出的LSTM模型,实现了99%的准确度。总的来说,可以确认基于LSTM的架构可以用于解决语音情感识别的任务。所提出的系统的实际应用可能是虚拟助理和心理健康监测。
摘要:Speech Emotion Recognition (SER) is the use of machines to detect the emotional state of humans based on the speech, which is gaining importance in natural human-computer interaction. Speech is a very valuable source of information, as emotions modify the patterns of speech; pitch, energy and even timing. Nonetheless, SER is not an easy task because speakers are not constant, and situations vary when recording and the sound similarity between specific feelings. In this work, the author introduces a speech emotion recognition system relying on the Mel-Frequency Cepstral Coefficient and Long Short-Term Memory (LSTM) neural network, as a feature extraction method. The Toronto Emotional Speech Set (TESS) speech signal was pre-processed, and transformed into MFCC features to understand the important aspects in terms of time. The resultant features were then introduced to LSTM model, which is able to learn long term features of sequential audio data. The trained model was measured over several emotion classes occurring in the dataset. As seen in the results of experiments, the proposed MFCC-LSTM approach succeeds in capturing the patterns of emotions in speech and provides highly realistic classifications in all the chosen emotion classifications. This study presents a speech emotion recognition system using Mel-Frequency Cepstral Coefficients (MFCCs) as features and a deep learning LSTM classifier. A Support Vector Machine (SVM) with an RBF kernel served as a classical baseline, achieving 98% accuracy, against which the proposed LSTM model, achieving 99% accuracy, was validated. Overall, it is possible to confirm that LSTM-based architectures can be used to address the task of speech emotion recognition. Actual applications of the proposed system may be virtual assistants and mental health surveillance.


机器翻译由腾讯交互翻译提供,仅供参考