今日论文合集:CS.SD语音与音频 | 共 9 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 语音合成与声音生成 3 篇
3. 语音翻译与语音语言模型 1 篇
4. 多模态音频与视听学习 1 篇
5. 数据集、基准与评测 1 篇
6. 其他/综合语音音频 2 篇
1. 语音识别与关键词检测 | 1 篇
1. SongCraft: Unified Song Generation and Editing with Reconstructive Learning
SongCraft:基于重建学习的统一歌曲生成与编辑
AI 总结:SongCraft通过重建预训练统一歌曲生成与编辑,利用条件重建实现细粒度属性编辑,并引入音素对齐、节拍条件等提升质量,达到最低词错误率。
链接:https://arxiv.org/abs/2609.16315
机构:Meta AI
作者:Haohe Liu, Varun Nagaraja, Gael Le Lan, Xinhao Mei, Zhaoheng Ni, Vikas Chandra, Abdelrahman Mohamed, Yangyang Shi
英文摘要:Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in which a model is trained to reconstruct audio from varying numbers of interpretable conditions. With conditions such as text and lyrics, the model learns to generate diverse songs. With dense conditions specifying fine-grained music attributes, the model learns to reconstruct the target and enables editing by modifying any single attribute while keeping others fixed. This leads to SongCraft, a latent flow matching based model trained for both generation and fine-grained editing. To improve song generation quality, we further introduce word-level phoneme alignment that improves pronunciation learning and accelerates convergence, beat conditioning that improves general musicality, and representation alignment on VAE latent space that produces semantically meaningful latents for improved generation quality. Experiments show that SongCraft achieves the lowest word error rate among evaluated song generation baselines while maintaining competitive audio quality. We further show that a single model can support editing of lyrics, vocal melody, beats, and singer identity, and we also study the trade-off between reconstruction quality and editability.
2. 语音合成与声音生成 | 3 篇
2. Taming Long-form Text-to-Speech
驯服长文本语音合成
AI 总结:针对长文本TTS性能退化问题,提出LACI推理方法,通过近实时错误检测与回滚重生成,显著提升长文本WER和语音克隆可靠性。
链接:https://arxiv.org/abs/2609.16989
机构:Argmax, Inc.(Argmax公司)
作者:Rongxiang Wang, Berkin Durmus, Aysegul Orhon, Eduardo Pacheco, Atila Orhon
英文摘要:Long-form text-to-speech (TTS) enables multi-turn conversations with consistent prosody and higher quality voice cloning from longer reference audio. Recent open-weights autoregressive TTS models such as Qwen3-TTS and VoxCPM2 attain state-of-the-art word error rate (WER) and speaker similarity (SIM) on short-form prompts but significantly deteriorate when used with long-form prompts. We propose Localized Attention-Constrained Inference (LACI), an inference-only method to detect TTS errors in near real-time, roll back to the error onset and regenerate with temporary guardrails, adding negligible computational overhead. Using LACI, we improve worst-of-N WER across 10 RNG seeds for Qwen3-TTS-0.6B from 35.2% to 3.4% on prompts longer than 1500 words, even surpassing its short-form reliability of 5.4\% on prompts with fewer than 500 words. To demonstrate the efficacy of LACI on voice cloning reliability, we propose a sliding-window version of the SIM metric that we call wSIM. wSIM exposes several novel failure patterns that are not captured by SIM. LACI improves worst-of-N wSIM from 0.01 to 0.47 on 120 seconds of reference audio while reducing the rate of catastrophic generations with WER above 30% from 26% to below 1%
3. Self-Distilled Pronunciation and Accent Control for Neural Text-to-Speech
自蒸馏发音与口音控制用于神经文本到语音
AI 总结:提出自蒸馏方法,用冻结主干自身输出作为教师,训练带口音生僻词阅读,无需词典或外部编辑,在多个TTS模型上显著提升口音准确率且不损害自然度。
链接:https://arxiv.org/abs/2609.17234
作者:Shuhei Kato
英文摘要:Text-to-speech that reads raw text has no lexicon: a rare word is read as guessed. Remedies train a reading-and-accent channel on recorded speech or edit words one at a time from exemplars. We do neither. The frozen backbone reads a sentence containing a common word it already says correctly, and its own output then serves as the teacher for the same sentence, with that word replaced by a tagged, accented reading; this training pair is the whole idea. On Sarashina2.2-TTS, screened raters at Fleiss' kappa = 0.85 hear the prescribed accent on 0.89 of unseen words against 0.57 for kana, which cannot express one; kana wins no pair; naturalness is not measurably hurt. Moved untuned to autoregressive, diffusion, and encoder-decoder backbones, it transfers reading, 0.25 to 0.47 above no edit on 319 words, and on CosyVoice 2 accent on two words in three, but not on Irodori; the paper locates why.
4. LACE: Layer-Wise Compression for Dynamic Frame Rate Codecs
LACE:用于动态帧率编解码器的分层压缩
AI 总结:LACE提出了一种层自适应动态帧率音频编解码器,通过每层独立压缩和联合对齐机制,在重建任务上取得更优的率-质量权衡,并提升TTS推理效率。
链接:https://arxiv.org/abs/2609.17509
机构:Carnegie Mellon University(卡内基梅隆大学)
作者:Thanapat Trachu, Samuele Cornell, William Chen, Shinji Watanabe
英文摘要:Neural audio codecs are a key component in speech language modeling. However, their high frame rates lead to long sequence lengths, increasing computational costs. Dynamic frame rate codecs mitigate this by reducing the effective frame rate using a compression step to merge multiple frames together. However, most prior methods either operate on single-codebook codecs or apply a single compression step before multi-layer quantization. This forces all quantization layers to share the same segmentation boundaries, despite the residual embeddings at different quantization layers exhibiting different rates of change over time. We propose LACE (Layer-Adaptive Codec Encoding), a dynamic frame rate codec that applies an independent compression step at each quantization layer, enabling layer-specific segmentation boundaries. To use LACE tokens in downstream text-to-speech (TTS), we further introduce union alignment and boundary anchor mechanisms to make durations consistent across layers while preserving compression benefits. Experiments on LibriTTS show that LACE offers a better rate-quality tradeoff than prior dynamic frame rate methods on the reconstruction task and improves TTS inference efficiency while maintaining competitive synthesis quality. Our code is released as part of the ESPnet3 codec recipe.
3. 语音翻译与语音语言模型 | 1 篇
5. CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
CLASH:口语讽刺检测中词汇与韵律依赖的反事实审计
AI 总结:CLASH提出双语反事实诊断框架,审计口语讽刺检测中词汇与韵律的依赖,揭示模型主要依赖词汇线索,区分声学敏感性与讽刺判别。
链接:https://arxiv.org/abs/2609.16582
机构:Imperial College London(伦敦帝国理工学院); New York University(纽约大学); Technical University of Munich(慕尼黑工业大学); Tencent Inc.(腾讯公司); Wuhan University(武汉大学); Nankai University(南开大学)
作者:Qiyang Sun, Xudong Li, Yupei Li, Jiabin Xue, Yuhang Dai, Jiaming Li, Bjorn W. Schuller
英文摘要:Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each utterance under original, lexical-preserving, prosody-preserving, and approximately neutralised conditions. We evaluate handcrafted acoustic-feature systems, self-supervised learning (SSL) probes, and large audio language models (LALMs) on CMMA and MUStARD. For target-only Qwen3-Omni, lexical-preserving speech retains a 0.135--0.148 AUROC advantage over prosody-preserving speech after duration balancing, with cluster-bootstrap intervals above zero; alternative lexical resynthesis preserves this advantage. Acoustic interventions shift scores without consistently improving discrimination or changing binary predictions under the evaluated conditions. Context and interaction estimates vary across corpora. These findings distinguish acoustic sensitivity from sarcasm discrimination while exposing duration, identity, and transformation effects.
4. 多模态音频与视听学习 | 1 篇
6. Audio-Visual Turn-taking Prediction in Cocktail Party Scenarios
鸡尾酒会场景中的音视频轮流说话预测
AI 总结:本研究评估了音视频轮流说话预测模型在鸡尾酒会嘈杂场景下的泛化与适应能力,发现性能显著下降(加权F1最高降38%),微调可改善但效果因模态和数据量而异,并公开了代码与标签。
链接:https://arxiv.org/abs/2609.17056
机构:Trinity College Dublin(都柏林圣三一大学)
作者:Long-Vu Hoang, Naomi Harte
英文摘要:Current predictive turn-taking models (PTTMs) achieve strong performance on benchmarks with controlled acoustic conditions and clean audio signals. Their generalisation to conversations with overlapping speech and background interference remains underexplored. In this research, we evaluate audio-visual PTTMs trained with clean data on a challenging cocktail-party testbed derived from the AVCocktail dataset, and analyse their adaptation behaviour to this new domain. Experimental results show consistent performance degradation across audio and visual modalities under noisy conditions, with up to 38% relative drop in weighted F1. Fine-tuning on the new domain improves robustness, but gains vary across modalities and depend on the size of the available pre-training data. These findings provide insights into the different generalisation and adaptation capabilities of the audio and visual modalities, and indicate the need for robust modelling strategies to adapt to the complexities of human interactions in noise. All code and turn labels are made publicly available to facilitate further research.
5. 数据集、基准与评测 | 1 篇
7. MUUNRiver-Bench: Diagnosing Relation-Dependent Music Retrieval with Multimodal Instructions
MUUNRiver-Bench:基于多模态指令的关系依赖音乐检索诊断基准
AI 总结:MUUNRiver-Bench通过多模态指令诊断关系依赖的音乐检索,涵盖13个流派、3,440首曲目和七项任务,揭示声学与文本编码器的互补偏差及融合方案的局限。
链接:https://arxiv.org/abs/2609.16090
机构:Central Conservatory of Music(中央音乐学院); Tsinghua University(清华大学)
作者:Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao, Xiaobing Li, Maosong Sun
英文摘要:Music retrieval is relation-dependent: given a reference track, a listener may seek its style with a new theme, a cover, or a comparable voice, and these intents demand contradictory rankings. We present MUUNRiver-Bench, a diagnostic benchmark whose reference-audio queries use natural-language instructions to define relevance. A pipeline combining expert genre priors, LLM-generated prompts and lyrics, synthesis, and expert review yields 3,440 tracks spanning 13 genres and 116 sub-genres, and seven tasks: similar-music, style-preserving lyric-rewriting, lyric-preserving style-rewriting, cover, vocal-timbre, isolated-vocal, and segment retrieval. Across six models in eight configurations, task-wise rank reversals reveal complementary biases: acoustic encoders favour local identity, whereas text-aligned encoders favour semantic relations. Frozen encoders diagnose default similarity preferences; instruction-aware and audio-text fusion systems provide exploratory tests of textual conditioning, with neither simple fusion scheme consistently improving its backbone
6. 其他/综合语音音频 | 2 篇
8. Structure Across Voices: Comparing acoustic-event type accumulation and sequence dependence across four vocal repertoires using frozen audio encoders
跨声音结构:使用冻结音频编码器比较四种声音库的声学事件类型累积与序列依赖性
AI 总结:本研究用冻结音频编码器比较四种声音库,发现声学事件累积与序列依赖性差异取决于测量属性与时间尺度,而非单一等级。
链接:https://arxiv.org/abs/2609.16612
作者:Mudit Sinha, Sanika Chavan
英文摘要:Vocal repertoires can differ in acoustic-event type accumulation and temporal organization, yet direct comparison is difficult because corpora use different native events and unequal amounts of sequence. We compare sperm whale codas, human speech phones, Bengalese finch syllables, and common marmoset calls using the same frozen-audio-encoder procedure while matching event count and local sequence opportunity. Whale shows the fastest type accumulation; Finch shows the strongest immediate dependence and repeated-subsequence recurrence. Physically interpretable acoustics recover complementary parts of this profile, continuous analyses without clustering support broad Whale acoustic coverage, and source- and position-preserving nulls retain both Finch order effects. Extending predictive context shifts the comparison toward Whale. Thus repertoire differences depend on the acoustic property and temporal scale measured rather than forming a single hierarchy.
9. SpiroPhonia: Non-Invasive Respiratory Health Assessment from Spontaneous Speech
SpiroPhonia:基于自发语音的无创呼吸健康评估
AI 总结:本文提出SpiroPhonia框架,利用自发语音通过机器学习评估呼吸健康,在201人数据集上达到78%准确率,证明日常语音可编码呼吸生物标志物,支持无创连续监测。
链接:https://arxiv.org/abs/2609.17350
机构:University of Maryland, College Park(马里兰大学学院公园分校); DR. M R Khan Shishu Hospital & Institute of Child Health(M R 汗儿童医院与儿童健康研究所)
作者:Roksana Khanom, Shafia Supty, Nirupam Roy, Ashok Agrawala
英文摘要:Chronic Obstructive Pulmonary Disease (COPD) remains a major global health challenge, emphasizing the need for accessible and non-invasive detection. Since speech production is fundamentally linked to respiratory physiology, its disruptions can serve as indirect indicators of pulmonary impairment. This study introduces SpiroPhonia, a machine learning framework that leverages spontaneous speech for respiratory health assessment. We evaluated SpiroPhonia on a new dataset of 201 speakers (102 with COPD, 99 healthy controls). By integrating statistical analysis with recursive feature selection, we identified a compact set of discriminative speech markers. Our best model achieved 78% accuracy, 80% F1-score, and 87% AUC. This performance on spontaneous speech is competitive with methods using controlled laboratory recordings. Findings demonstrate that everyday speech encodes robust respiratory biomarkers, paving the way for continuous health monitoring via voice-enabled technologies.
