今日论文合集:CS.SD语音与音频 | 共 7 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 语音合成与声音生成 1 篇
3. 语音翻译与语音语言模型 1 篇
4. 安全、隐私与深度伪造音频 2 篇
5. 其他/综合语音音频 2 篇
1. 语音识别与关键词检测 | 1 篇
1. CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
CircleMatch:基于循环时间统计的原型匹配用于微型关键词唤醒
AI 总结:CircleMatch通过原型匹配与循环时间统计,以约1k-7k参数的微型模型在多个关键词唤醒数据集上实现高精度,并展现出平移等变性与时间压缩适应性。
链接:https://arxiv.org/abs/2609.20070
机构:Shanghai Normal University(上海师范大学)
作者:Jiajun Sun, Zhe Gao
英文摘要:Keyword spotting (KWS), the task of identifying predefined words in speech, is a core capability of voice-enabled devices. Achieving high KWS accuracy under tight parameter budgets across different vocabulary sizes remains challenging. We present CircleMatch, a matching framework enabling KWS with very few parameters. Its encoder independently compresses frequency bands and fuses them into frame features. These features are then matched against learned class-specific prototypes to produce temporal response curves. Parameter-free circular aggregation encodes time as angles and summarizes response distributions and relative timing for classification. We develop four tiny variants, Circle-D4, Circle-D8, Circle-D16, and Circle-D32, ranging from approximately 1k to 7k parameters in the 12-class setting. Experiments with multiple random seeds on Speech Commands v1/v2 and the English and Spanish Micro subsets of the Multilingual Spoken Words Corpus demonstrate competitive accuracy with tiny models. Our qualitative analysis further suggests approximate shift equivariance of prototype responses and adaptation to temporal compression. Code and model weights are available at this https URL.
2. 语音合成与声音生成 | 1 篇
2. Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis
面向直播语音合成的多维韵律评判
链接:https://arxiv.org/abs/2609.20124
机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); TaoLive-AIGC Team Taobao & Tmall Group of Alibaba(阿里巴巴淘宝天猫集团淘Live-AIGC团队)
作者:Zifan Guan, Longyu Lu, Junan Zhang, Zhizheng Wu, Meiguang Jin, Junfeng Ma
英文摘要:Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective pairwise evaluator distilled from Gemini into Qwen3-Omni. However, we identify a critical flaw in standard multi-dimensional evaluation: verdict coupling. The judge tends to lazily align all individual dimension scores with its overall preference, collapsing a rich multi-dimensional rubric into a single preference bit. To resolve this, we further propose Decoupled-Live-ProsodyJudge (D-LPJ). D-LPJ eliminates the overall verdict target to prevent blind following, masks uncertain pair-dimensions during Supervised Fine-Tuning(SFT), and introduces a novel span-local GRPO strategy that applies normalized advantages strictly to their corresponding rationale spans. Evaluated on highly curated human-annotated test sets, 10 sample balanced-order LPJ achieves higher point accuracy than a single Gemini call, while D-LPJ successfully produces independent,decoupled dimension judgments. Furthermore, in a Best-of-8 TTS candidate selection tournament, the LPJ-selected utterance falls within the human top-3 in 85.29% of high-confidence cases, demonstrating its efficacy for fine-grained TTS preference optimization.
3. 语音翻译与语音语言模型 | 1 篇
3. Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study
音频语言模型中的音乐幻觉:一种分层形式化与实证研究
AI 总结:本研究首次对音频语言模型的音乐幻觉进行分层多范式实证分析,提出MuseDiag诊断框架,并验证两种免训练缓解方法在不同范式下的效果差异。
链接:https://arxiv.org/abs/2609.20195
机构:Institute of Information Engineering, CAS(中国科学院信息工程研究所); School of Cyber Security, UCAS(中国科学院大学网络空间安全学院); Central Conservatory of Music(中央音乐学院); University of Electronic Science and Technology of China(电子科技大学)
作者:Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu
英文摘要:Audio-language models increasingly generate confident music descriptions that are unsupported by the input audio. We present, to our knowledge, the first music-specific, layer-wise, multi-paradigm empirical study of hallucination in audio-language models and formulate it as a hierarchical perceptual grounding failure across five layers: sound events, temporal properties, tonal attributes, style, and emotion. We introduce MuseDiag, a multi-paradigm diagnostic framework with contradiction-based verification, and evaluate nine models (four open-source and five closed-source). We find that (1) vocal misperception is a universal weakness across all nine models, tonal perception is a major axis of architectural differentiation, and Audio-Flamingo-3 remains the stable leader while substantial reordering below it reveals paradigm-specific vulnerability profiles; (2) affirmative bias, generation-mode effects, and layer-specific perceptual limitations are each empirically associated with the observed patterns, with convergent evidence from multiple analyses rather than strict causal attribution; and (3) our two training-free mitigation methods, Audio-Dependency-Aware Decoding for Music (ADD-M) and Taxonomy-Guided Perceptual Anchoring (TPA), can reduce hallucination in probing, but their gains vary by model and often do not carry over to free-form generation, showing that music hallucination mitigation must be evaluated across paradigms.
4. 安全、隐私与深度伪造音频 | 2 篇
4. CoRELoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection
CoRELoop:用于音频深度伪造检测的参数高效受控循环精化
AI 总结:针对音频深度伪造检测泛化难题,提出CoReLoop方法,在冻结的SSL检测器上通过轻量级循环精化模块和低秩适配器实现参数高效精化,在14个跨域测试集上将EER从4.85%降至3.74%。
链接:https://arxiv.org/abs/2609.19818
机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); Amphion Technology Co., Ltd.(安飞昂科技有限公司)
作者:Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu
英文摘要:Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlling state updates, and aligning refined outputs with the frozen classifier. By training only lightweight refinement modules and loop-specific low-rank adapters on the original data, CoReLoop enables additional refinement while preserving the detector's original first-pass prediction. On 14 cross-domain test sets, the 24-layer model reduces pooled equal error rate (EER) from 4.85% to 3.74% with two passes, with approximately 10M trainable parameters out of 598M. To selectively apply this refinement, an optional halting head chooses the depth for each utterance, achieving 3.73% pooled EER with an average of 1.18 passes.
5. Robust Workflow Generation via Adversarial Learning for Audio Deepfake Detection
通过对抗学习实现鲁棒工作流生成的音频深度伪造检测
AI 总结:针对音频深度伪造检测在真实扰动下泛化性差的问题,提出ROGUE框架,通过双智能体对抗学习动态生成检测工作流,显著提升鲁棒性和泛化能力。
链接:https://arxiv.org/abs/2609.20063
机构:Fordham University(福特汉姆大学); IBM Research(IBM研究院)
作者:Xiang Li, Pin-Yu Chen, Wenqi Wei
英文摘要:The rapid advancement of speech synthesis and voice conversion technologies has made audio deepfakes increasingly realistic, posing serious security risks in practical applications. While existing detection methods achieve strong performance under controlled conditions, they often fail to generalize under real-world perturbations and corruptions. In this paper, we propose ROGUE, a framework that dynamically constructs robust detection workflows by orchestrating multiple detection tools. ROGUE formulates workflow generation as a sequential decision-making problem and introduces a dual-agent paradigm, where a perturbation agent generates audio perturbations and a policy agent learns to select and execute detection tools under perturbed conditions. Through adversarial learning, ROGUE enables perturbation-aware tool selection, adaptive execution strategies, and improved robustness to distribution shifts. Extensive experiments across multiple datasets and real-world corruptions demonstrate that ROGUE consistently outperforms strong baselines in both robustness and generalization. Our results highlight the effectiveness of adversarially optimized workflow generation for building reliable audio deepfake detection systems in real-world deployment settings.
5. 其他/综合语音音频 | 2 篇
6. A State-Space Model of Figured-Bass Realization: Local Constraints, Coupled Voices, and Polynomial-Time Solvability
数字低音实现的状态空间模型:局部约束、耦合声部与多项式时间可解性
AI 总结:本文为数字低音实现建立显式数学模型,将约束表示为谓词,证明固定声部下可行性与最小成本实现为多项式时间问题,并通过实例展示合法性、优化及贪婪选择的失败。
链接:https://arxiv.org/abs/2609.19397
机构:National Taiwan Normal University(国立台湾师范大学)
作者:Evan Unit Lim
英文摘要:Figured-bass realization can be described as a sequence of choices constrained both within each sonority and between successive sonorities. This paper gives an explicit mathematical model of a restricted, examination-style four-part realization problem. Pitch spelling, range, chord membership, doubling, omission, spacing, crossing, overlap, melodic motion, consecutive perfect intervals, and selected resolution requirements are expressed as predicates. We distinguish hard constraints from optional preference costs. Four labeled notes are represented visually as the vertices of a quadrilateral and computationally as one ordered voicing state. Legal progressions become paths through a layered graph. We prove that feasibility and minimum-cost realization are polynomial-time problems for a fixed number of voices with explicit finite note domains and fixed local rules. For fixed ranges, a fixed note alphabet, and adjacent-event rules, the number of graph operations is linear in the number of events. Worked two-, four-, and eight-beat examples illustrate legality, optimization, and the failure of a greedy choice. The result concerns the stated formal model; it is not a claim that every musical judgment is captured by local predicates.
7. A Cross-Lingual Acoustic Disease-Alignment Framework for Respiratory Health Assessment from Spontaneous Speech
跨语言声学疾病对齐框架:基于自发语音的呼吸健康评估
AI 总结:提出跨语言疾病对齐框架CL-DAF,从自发语音中识别跨语言一致的声学特征,提升呼吸疾病(如COPD)跨语言检测性能,为多语言临床语音模型奠定基础。
链接:https://arxiv.org/abs/2609.19398
机构:University of Maryland, College Park(马里兰大学学院公园分校); Sylhet MAG Osmani Medical College Hospital(锡尔赫特MAG奥斯马尼医学院医院); Dr. M R Khan Shishu Hospital & Institute of Child Health(M R 汗儿童医院与儿童健康研究所); Line Reflection Ltd.(Line Reflection 有限公司)
作者:Roksana Khanom, Raghib Asfak Tasnim, Bodrun Nahar Bithi, Shafia Shirin Supty, Saiful Islam Raju, Ashok Agrawala, Nirupam Roy
英文摘要:Spontaneous speech offers a scalable, noninvasive signal for respiratory health assessment, yet interpretable models that generalize across languages remain challenging because disease-related acoustic changes are confounded by language-specific phonetic variation. We present CL-DAF, a Cross-Lingual Disease-Alignment Framework that identifies acoustic dimensions whose disease effects remain consistent across languages. Using 201 English and 75 newly collected Bangla speakers, we construct a common 272-dimensional acoustic representation and quantify disease alignment using signed rank-biserial effects and the Language Invariance Score. We first show that spontaneous Bangla speech separates COPD from controls (AUC 0.85); however, 133 features reverse their disease direction across languages and the full representation transfers poorly (AUC 0.49 from Bangla to English). CL-DAF isolates 26 disease-aligned features that raise AUCs to 0.825 and 0.722 from English to Bangla and Bangla to English, respectively. These findings provide a foundation for multilingual clinical speech models emphasizing pathology over language-dependent variation.
