今日论文合集:CS.SD语音与音频 | 共 7 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音识别与关键词检测 1 篇

2. 语音合成与声音生成 1 篇

3. 音乐信息检索与音乐生成 1 篇

4. 安全、隐私与深度伪造音频 1 篇

5. 其他/综合语音音频 3 篇

1. 语音识别与关键词检测 | 1 篇

1. Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding

倾听,勿抄袭:内化音频基础的支架上下文以实现稳健的全模型语音理解

AI 总结:研究针对全模型在嘈杂重叠语音中准确率下降问题,提出音频基础的支架上下文(AGSC)方法,通过三步构建线索,经测试优化后用于训练,降低了无线索错误率,还制定联合任务,内化后几乎不增加推理开销。

链接:https://arxiv.org/abs/2607.21943

作者:Pengfei Zhang, Biao Tian, Tianxin Xie, Minghao Yang, Xiangang Li, Li Liu

英文摘要:Omni models transcribe clean, single-speaker speech well, but their accuracy drops sharply when speakers overlap and the scene is noisy, exactly where knowing who said what matters most. A natural fix is a short scene description. We show why this is risky: answer-bearing text lets the model copy instead of listen, so the score rises although nothing has been heard; a silent test exposes this shortcut at once. We call this failure mode perception bypass and address it with Audio-Grounded Scaffold Context (AGSC). AGSC links three steps: first, we build clues from audio to guide listening without giving the answer; second, answer-overlap and silence tests probe them for leakage and audio dependence; finally, those clues scaffold training but vanish at test time, yielding no-clue capability. Across three heterogeneous Omni models, training on AGSC lowers no-clue capped mean permutation word error rate (mpWER) on overlapping, noisy speech from 25%-71% to 9%-15%. For streaming control, we formulate a joint GDPO task in which the model learns when to use a clue and how to produce a speaker-attributed transcript from separately normalized format, gate, and transcript rewards. After internalization, AGSC adds almost no inference overhead.

2. 语音合成与声音生成 | 1 篇

2. SoundscapeAgent: Agentic Soundscape Construction for Controllable Synthesis and Scalable Audio-Language Supervision

SoundscapeAgent:用于可控合成和可扩展音频-语言监督的智能音景构建

AI 总结:提出智能音景构建框架SoundscapeAgent,基于大语言模型的智能体将用户意图转化为场景计划,经多步骤实现可控音频合成与可扩展音频-语言数据构建,在生成性能及下游音频推理方面表现出色。

链接:https://arxiv.org/abs/2607.21857

机构:Tencent(腾讯)

作者:Hao Zhang, Yiwen Zhao, Yixuan Zhang, Yiwen Shao, Steve Yves

英文摘要:We present an agentic soundscape construction framework for controllable compositional audio generation that makes explicit the scene planning, source selection, temporal layout, and rendering steps typically handled implicitly by single-shot text-to-audio models. An LLM-based agent converts user intent into an executable scene plan, acquires assets through retrieval and on-demand generation, renders controllable multi-event mixtures, and exports aligned scene metadata. The framework also supports human-in-the-loop interaction through user-guided tool selection and editable scene plans. Together, these components provide an inspectable and reusable approach to controllable soundscape synthesis and scalable audio-language data construction. Listener studies and objective metrics demonstrate competitive generation performance against text-to-audio baselines, while models trained with agent-generated data consistently outperform real-only baselines in downstream audio reasoning. Code, demos, and listening-test results are available at this https URL.

3. 音乐信息检索与音乐生成 | 1 篇

3. Music-JEPA: Learning a World Model of Sound from Action

Music-JEPA:从动作中学习声音的世界模型

AI 总结:研究提出Music-JEPA,将音乐构建为动作条件系统,用JEPA学习钢琴声音世界模型,在离线状态下用配对数据训练。实验表明模型能捕捉音乐动作与声音关系,其表示支持下游任务并能通过规划实现钢琴转录。

链接:https://arxiv.org/abs/2607.22000

作者:Ziyu Wang, Kun Fang, Yann LeCun

英文摘要:Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remains unclear how such frameworks can naturally support the formation of a world model for music. In this work, we propose to learn a world model of piano sound using JEPA by framing music as an action-conditioned system: the audio is treated as the state, and the pianoroll as the instrument action. Given a current audio state and an action, the model predicts the resulting future audio state, mirroring how humans learn musical sound through interaction. The model is trained in a fully offline setting using paired audio-pianoroll data, without environment interaction. Experiments show that the learned model captures the relationships between musical actions and their resulting sound. The resulting representations support downstream tasks, including beat tracking, composer identification, and key estimation, and enable piano transcription via planning, by searching for actions that best explain a target sound.

4. 安全、隐私与深度伪造音频 | 1 篇

4. Probing Speaker Identity Sensitivity in Audio Deepfake Detectors

探究音频深度伪造检测器中的说话者身份敏感性

AI 总结:研究音频深度伪造检测器对说话者身份的敏感性问题,提出身份敏感性分数(ISS),该方法无需真实标签,通过计算检测器分数变化量化身份敏感程度,能有效预测错误分类,为基于说话者的故障分析提供实用诊断。

链接:https://arxiv.org/abs/2607.21820

机构:Michigan State University(密歇根州立大学)

作者:Daniyal Kabir Dar, Arun Ross

英文摘要:Audio deepfake detectors are trained to distinguish genuine speech from synthetic speech and often perform well on standard benchmarks. Yet the same detector that achieves less than 1% error on one dataset can see its error rate increase twentyfold when evaluated on a different dataset. We argue that one contributing factor is speaker-identity reliance: standard training corpora correlate speaker identity with the genuine/synthetic label, allowing detectors to partially rely on speaker-related cues rather than synthesis artifacts alone. We propose the Identity Sensitivity Score (ISS), a per-utterance diagnostic that quantifies how much a detector's output changes across different speaker identity contexts. ISS requires no ground-truth labels at inference time and can be computed from the detector score and a pool of reference speaker examples. Across two detectors and two datasets, incorrectly classified utterances have ISS scores 29 to 52 times higher than correctly classified utterances, and ISS alone predicts misclassification with area-under-curve (AUC) up to 0.954. To test whether ISS actually captures identity-sensitive behavior rather than serving only as a proxy for prediction confidence, we apply voice conversion to 500 utterances and measure the resulting detector-score shift. Utterances flagged as identity-sensitive by ISS respond 19 to 30 times more strongly to this manipulation than utterances flagged as stable. These results position ISS as a practical inference-time diagnostic for speaker-dependent failure analysis in audio deepfake detection.

5. 其他/综合语音音频 | 3 篇

5. CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following

CODA:用于基于图像的实时乐谱跟随的级联在线不连续感知对齐

AI 总结:本文针对实时乐谱跟随难题,提出CODA系统。它利用乐谱级联结构增强预测一致性,通过静音驱动中断模式实现不连续恢复。在多模态乐谱数据集钢琴基准测试中,CODA在实时吞吐量下取得了领先的跟踪精度和不连续恢复性能。

链接:https://arxiv.org/abs/2607.21899

作者:Yining Yang, Ruogu Chen, Jie Han

英文摘要:Real-time score following from sheet images remains chal- lenging because the model must process streaming au- dio while resolving highly repetitive visual patterns un- der strict latency constraints. Recent image-based meth- ods have attempted to use multi-resolution prediction by simultaneously predicting the positions of the active sys- tem, bar, and note. However, their predictions across these different levels of notation are independent, which makes the predictions unstable and introduces unnecessary ex- tra search space for bar- and note-level predictions. Most existing methods also lack mechanisms to recover from score discontinuities, such as repeats, da capo (D.C.), or coda jumps. This paper proposes CODA, to the best of our knowledge, the first real-time score following system that addresses both gaps. CODA explicitly exploits the cascaded structure of music scores: it first selects the ac- tive system, then the active bar within it, and finally the active note within the selected bar. This enforces pre- diction consistency across resolutions. A silence-driven break mode enables recovery from arbitrary score discon- tinuities without requiring knowledge of the repeat struc- ture. Evaluated on the Multimodal Sheet Music Dataset (MSMD) piano benchmarks, CODA achieves state-of-the- art tracking accuracy and discontinuity-recovery perfor- mance under real-time throughput. Code is available at this https URL.

6. MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection

MemNMF:基于线性预测编码频谱的内存增强非负矩阵分解用于异常声音检测

AI 总结:研究针对基于自动编码器的异常声音检测问题,提出MemNMF方法,该方法基于线性预测编码频谱,通过初始化内存模块并进行注意力加权组合来重建输入,实验表明其能改进自动编码器基线,在复杂条件下鲁棒性强。

链接:https://arxiv.org/abs/2607.22086

机构:Institute of Science Tokyo(东京理科大学)

作者:Phurich Saengthong, Takahiro Shinozaki

英文摘要:Autoencoder-based anomalous sound detection is attractive for machine condition monitoring because it can be trained using only normal recordings and yields an interpretable anomaly score from reconstruction error. Most prior work uses spectrogram autoencoders, but reconstructing detailed time--frequency patterns is sensitive to noise and transients, and models can reconstruct some anomalous inputs well, weakening normal--anomaly separation. We propose MemNMF, a constrained reconstruction method that operates on the Linear Predictive Coding spectrum, a compact estimate of the spectral envelope. MemNMF initializes a memory module from an NMF dictionary learned on normal LPC spectra and reconstructs each input as an attention-weighted combination of prototypical normal spectral patterns. Experiments on MIMII and DCASE 2020 Task 2 across multiple machine types and operating conditions show that LPC-spectrum inputs improve a standard autoencoder baseline and that MemNMF yields further gains, with especially strong robustness under noisy, non-stationary settings.

7. Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition

Reflector:用于基于样本的作曲的排列感知和声检索

AI 总结:研究针对编曲发展中样本检索问题,提出Reflector交互式音频工作站,通过固定音级类预言机及相关技术跟踪和声组合、调整检索,能揭示和声关系,管道本地运行且免费开源,其嵌入保留判断并覆盖整个库。

链接:https://arxiv.org/abs/2607.22413

作者:Austin Rockman

英文摘要:Sample retrieval tools can help composers find harmonically compatible material, but querying from a fixed reference sample becomes less informative as arrangements evolve and the harmonic context shifts with each musical decision. We present Reflector, an interactive audio workstation that tracks harmonic combinations as they accumulate on the composer's timeline and adapts retrieval as the arrangement develops. The system is organized around a fixed interval-class oracle: a hand-designed table of weights that scores how pitch-class content combines between sources. An encoder trained entirely on synthetic audio learns to approximate the oracle in a 128-dimensional embedding space, where dot products stand in for compatibility scores at interactive speed. As the composer arranges material on a multi-track timeline, a sweep-line analysis discovers co-sounding regions, computes oracle-weighted centroids, and retrieves against the composite harmonic identity of the session as it evolves. Session centroids projected into a navigable 3-D space reveal structural harmonic relations across the composer's body of work. This paper is a systems account: we give the design rationale for each architectural decision, characterize Reflector's behavior through intrinsic measurements on a working sample library, and describe the implementation. The characterization yields a central finding: the learned embedding preserves the kernel's pairwise judgments while covering the whole library, something the kernel cannot do when used directly as a retrieval rule, because the embedding's normalized geometry cannot express the degenerate solutions that direct scoring favors. The entire pipeline runs locally with no copyrighted training data. Reflector is free, and the training pipeline is open source.