今日论文合集:CS.SD语音与音频 | 共 6 篇。
本文经arXiv每日学术速递授权转载微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. MMAG: A Multi-Control Mixed Audio Generation Benchmark

MMAG:多控制混合音频生成基准

AI 总结:本研究针对现有音频生成基准的不足,推出MMAG混合音频生成基准,构建含4000个标注音频的数据集,提出系统评估协议,测试发现各类模型在多控制能力间存在性能权衡,为该领域提供综合基准。

链接:https://arxiv.org/abs/2608.06900

机构:Shanghai Jiao Tong University(上海交通大学); Shanghai AI Lab(上海人工智能实验室)

作者:Zihao Zheng, Xuenan Xu, Jiahao Mei, Yixuan Li, Minghao Lv, Wen Wu, Chao Zhang, Mengyue Wu

英文摘要:Recent audio generation systems have progressed from single-modality synthesis to generating complex acoustic scenes containing speech, music, and sound effects. Therefore, evaluating these models requires assessing multiple interacting capabilities, including semantic fidelity, speaker consistency, and temporal control, yet existing benchmarks focus on isolated domains or coarse-grained descriptions. To address this gap, we introduce the Multi-control Mixed Audio Generation (MMAG) benchmark. MMAG contains approximately 4,000 manually verified audio clips with rich annotations covering speech content, speaker identity, music attributes, sound events, and temporal relationships, together with dedicated subsets for voice cloning and timestamp-conditioned generation. We further propose a systematic evaluation protocol that measures acoustic fidelity, speech quality, semantic alignment, and temporal accuracy. Benchmarking representative agentic orchestrators, unified audio-visual generation models, and native mixed-audio generators reveals substantial performance trade-offs across these capabilities, with no existing model performing consistently well. Our results highlight the remaining challenges of controllable mixed audio generation and establish MMAG as a comprehensive benchmark for future research.

2. Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement

云端增强型低计算量多通道语音增强

AI 总结:针对可穿戴设备语音增强的计算约束瓶颈,提出融合延迟服务器输出、分层特征增强及协作多通道维纳滤波的云端协作框架,以低额外开销显著提升边缘模型性能。

链接:https://arxiv.org/abs/2608.07423

机构:University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校); Meta Reality Labs Research(Meta现实实验室)

作者:Xulin Fan, Juan Azcarreta, Ashutosh Pandey, Jesus Alvarez, Ke Tan, Jacob Donley, Ritwik Giri, Buye Xu

英文摘要:Low-latency, low-compute speech enhancement is essential for wearable devices with real-time communication requirements, but strict computational constraints significantly limit on-device performance. Knowledge Boosting has been proposed as an effective approach to improve edge model performance by leveraging a more capable server-side model, but performance gains for speech enhancement have been limited. We propose a collaborative framework incorporating three techniques: (1) delayed server output as additional input, (2) layerwise feature boosting that transfers intermediate server representations to guide edge inference, and (3) collaborative multichannel Wiener filtering, which fuses weighted covariance matrices estimated from both server and edge models for improved beamforming. Experimental results demonstrate that the proposed collaborative framework significantly outperforms the edge-only baseline with minimal additional computational overhead.

3. MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering

MI-MIDI:基于探测、透镜与调控的文本生成MIDI模型的可机理解释

AI 总结:本研究针对两款文本生成MIDI模型,采用探测、透镜等方法开展可机理解释,揭示架构对音乐结构的塑造作用,提出双取向调控协议并提供音乐概念追踪工具包。

链接:https://arxiv.org/abs/2608.06638

作者:Jakub Poćwiardowski, Mateusz Modrzejewski

英文摘要:Mechanistic interpretability of music generation has concentrated on audio models, leaving symbolic models largely unexplored. We analyze two public text-to-MIDI systems of contrasting design: the purpose-built encoder--decoder text2midi and MIDI-LLM, a Llama~3.2~1B model extended with MIDI tokens using linear probing, the logit and tuned lenses, activation patching and difference-in-means steering. Across these methods, we recover musically meaningful structure and show how architecture shapes its formation and control. Pitch, instrumentation, harmony and texture are linearly decodable in both models. text2midi refines predictions gradually across depth, whereas MIDI-LLM works largely in its inherited textual basis before a sharp late rotation into the musical vocabulary; patching identifies a matching late attenuation of prompt-driven instrument transfer. Steering produces bidirectional changes in register and polyphony in both systems, and in tempo/energy in MIDI-LLM. Our two-orientation protocol isolates directional control and shows that all-layer interventions are robust in text2midi but accumulate disruptively in MIDI-LLM. Together, the results provide a practical toolkit for tracing and controlling musical concepts in symbolic generators. Audio examples are available on a demo website.

4. Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

用于文本引导语音修复与编辑的多编解码器离散扩散模型

AI 总结:该研究提出基于分层编解码器令牌的离散扩散框架SIEDD,其核心架构HiCoDD结合音素级约束等技术,在RealEdit基准中实现最优语音编辑性能,且优于自回归基线,显著提升上下文保留的语音重建与编辑效果。

链接:https://arxiv.org/abs/2608.06424

机构:Ben-Gurion University of the Negev(内盖夫本-古里安大学); University of Haifa(海法大学)

作者:Iftach Shoham, Tali Dror, Oren Gal, Haim Permuter, Gilad Katz, Eliya Nachmani

英文摘要:Speech recordings often contain missing, corrupted, or incorrect regions that must be reconstructed or modified without re-synthesizing the entire utterance. Speech inpainting restores missing segments, whereas speech editing replaces spoken content according to an edited transcript. Both tasks require the generated speech to express the intended words while remaining consistent with the surrounding speaker identity, prosody, timing, and recording conditions. Discrete diffusion is particularly well suited to these tasks because it can iteratively refine masked tokens while jointly conditioning on both left and right acoustic context. We introduce SIEDD, a discrete diffusion framework for text-guided speech inpainting and editing over hierarchical codec tokens. Its core architecture, HiCoDD, follows the RVQ generation order by representing previously generated codebooks as clean, committed acoustic context and applying diffusion only to the current refinement codebook. This separation enables leakage-free joint training while matching sequential coarse-to-fine inference. The model further combines phoneme-level conditioning, span-localized classifier-free guidance, and duration prediction to support both fixed-duration inpainting and variable-duration text edits. On the RealEdit benchmark, SIEDD achieves the best overall speech-editing performance among the evaluated methods. It also outperforms the evaluated autoregressive baselines across all speech-inpainting settings, on both single and multiple gaps. These results demonstrate that explicitly modeling the codec hierarchy substantially improves context-preserving speech reconstruction and editing. See our full code at this https URL.

5. Frame-Level Pansori Mode Classification with Complementary Audio Representations

基于互补音频表示的帧级盘索里调式分类

AI 总结:本研究针对韩国传统声乐盘索里,采用46小时帧级标注与4种互补音频表示,构建调式分类模型,验证其学习调式特征而非记忆曲目,揭示了源分离对线索的影响及预训练的局限性。

链接:https://arxiv.org/abs/2608.06633

作者:Sangheon Park, Seonguk Ju, Suin Chung, Danbinaerin Han, Dasaem Jeong

英文摘要:Pansori is a traditional Korean vocal genre whose mode system (jo) is defined not by scale alone but by the entanglement of pitch collection, microtonal ornament (sigimsae), and vocal timbre. In this study, we introduce a 46-hour frame-level pansori mode annotation, expert-labeled across all five canonical batang, and evaluate four complementary input representations (mel spectrogram, F0 contour, MIDI piano roll, and a multi-cultural SSL encoder) under two split strategies designed to detect shortcut learning. Across the three well-represented modes, performance degrades by only 2.1--3.6 points of F1 when entire works are held out, indicating that the models learn mode-relevant features rather than memorizing repertoire. Per-class results further show that source separation removes the percussion cue on which changjo depends, and that generic multi-cultural pre-training fails specifically on the Ujo--Gyemyeonjo distinction. Qualitative analysis of cross-modal disagreement recovers musicologically documented phenomena and agrees with published score-based analyses of modern changjak pansori.

6. From Prompting to Describing: A Cross-Cultural Study of Language for AI-Generated Music

从提示到描述:针对AI生成音乐语言的跨文化研究

AI 总结:该研究对比了TTM系统提示与音乐描述的语言差异,构建了音乐提示词汇分类法,发现提示与感知存在结构不对称,且不同文化背景的描述特征存在差异,引发对现有TTM系统文化适应性的思考。

链接:https://arxiv.org/abs/2608.06634

作者:Sangheon Park, Claire Arthur

英文摘要:Text-to-music (TTM) generation systems allow users to create music through natural language prompts, yet it is unclear whether the descriptive language used to prompt aligns with descriptive language used to summarize or describe heard music. We pair 200 real-world Udio prompts with their generated audio and free-form descriptions collected from English- (n = 70) and Korean-speaking (n = 78) listeners, and contribute a human-derived taxonomy of musical prompting vocabulary grounded in real user data. Using this framework, alongside word- and vector-level analyses, we find a consistent structural asymmetry: prompts are dominated by Genre and Story/Narrative language. Genre terms propagate most reliably from prompt to perception, while narrative-heavy prompts are the strongest predictor of semantic misalignment. A preliminary cross-cultural comparison further suggests that description profiles vary across listener populations along narrative, functional, and affective dimensions, raising questions about whether current TTM systems, trained on aggregated English-centric corpora, can accommodate the full diversity of how people naturally express musical ideas.