微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 音乐信息检索与音乐生成 1 篇
2. 语音翻译与语音语言模型 1 篇
3. 其他/综合语音音频 1 篇
1. 音乐信息检索与音乐生成 | 1 篇
1. Finding the noise: Zero-shot AI Music Detection
寻找噪音:零样本人工智能音乐检测
AI 总结:研究无人知晓输入样本生成模型时的人工智能音乐检测,提出结合伪像提取法、非负矩阵分解及分类聚类方法,用于区分真实与合成音乐及零样本多类识别,在相关任务中性能优异,可监测大规模目录。
链接:https://arxiv.org/abs/2607.25530
作者:Darius Afchar, Romain Hennequin
英文摘要:We present a novel method for AI-generated music detection in scenarios where the models that generated the input samples are unknown to the detector (e.g., from a newly released service). Since 2023, there has been a multiplication of user-friendly AI-music generation services (e.g., Suno, Udio), along with regular updates and new features. There is thus a need to address synthetic content detection in an unsupervised way to adapt to this rapidly changing context. This angle has not been much studied in music yet. We propose to study two tasks. First, discriminating between real and synthetic music. This may be approached in a one-class manner, namely, using some baseline real music and trying to determine what falls outside. Second, zero-shot multi-class identification, which is more similar to an unsupervised clustering task on a mix of real and various AI-music generations, where the goal is to create coherent, high-purity clusters. We propose a combination of a previously proposed artifact-extraction method, on top of which we apply non-negative matrix factorization and simple classification and clustering methods. We achieve excellent performance on both tasks, showing that the proposed methods may be used to monitor large-scale catalogs that may receive AI-generated samples from various newly released generative models.
2. 语音翻译与语音语言模型 | 1 篇
2. From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
从语义到读出:时间音频接地微调后音频令牌的机制理解
AI 总结:研究大型音频语言模型中音频令牌在微调后的机制,通过四种分析研究其分层语义等,发现微调前已有潜在证据,微调后解码器更易访问事件信息,支持从语义到读出的解释,有助于解码器连接时间输出。
链接:https://arxiv.org/abs/2607.25355
作者:Yujian Ma, Jinqiu Sang, Ruizhe Li, Jiaao Yu, Ang Li
英文摘要:Large audio-language models (LALMs) convey acoustic evidence to language decoders through native audio tokens, yet the internal roles of these tokens remain poorly understood. Using temporal audio grounding as a diagnostic setting, we examine how language-model fine-tuning affects the layerwise semantics, decoder accessibility, and temporal output alignment of native audio-token states through four complementary analyses: query-conditioned token semantics, calibrated token readout, temporal-window probes, and residual-delta erasure during generation. Alongside substantial improvements in temporal localization, semantic analysis of Qwen2.5-Omni shows that latent evidence for queried events is already present before fine-tuning and that the audio tokens most strongly aligned with the queried event appear at similar temporal positions before and after fine-tuning. After fine-tuning, event-related information in audio tokens becomes more accessible to the decoder, especially in early and middle layers, and a cross-checkpoint control shows that this improvement arises primarily from decoder adaptation. Temporal probes show that the base checkpoint already contains recoverable information about annotated windows and that fine-tuning mainly improves alignment with each checkpoint's own predicted temporal support. Residual-delta erasure further shows that removing audio-token updates within predicted windows harms timestamp generation more than removing the same number of randomly selected updates. The same broad improvements in decoder readability and prediction alignment also appear in Qwen2-Audio. Together, these results support a semantics-to-readout account in which grounding fine-tuning helps the decoder read existing event evidence and connect it more reliably to temporal outputs.
3. 其他/综合语音音频 | 1 篇
3. GraphIDyOM: A graph-native Python reimplementation of IDyOM for musical expectation modelling
GraphIDyOM:用于音乐期望建模的IDyOM的原生Python重新实现
AI 总结:研究针对IDyOM难以与Python集成及内存结构不易处理的问题,提出GraphIDyOM进行重新实现,保留其架构,验证了实现效果并与其他实现做了对比,展示了新实现支持多种分析及应用,为研究音乐期望提供平台。
链接:https://arxiv.org/abs/2607.25787
机构:Institute for Interdisciplinary Studies on Artificial Intelligence (IRIDIA), Université Libre de Bruxelles(布鲁塞尔自由大学跨学科人工智能研究所(IRIDIA))
作者:Lluc Bono Rosselló
英文摘要:The Information Dynamics of Music model (IDyOM) has played a central role in computational accounts of musical expectation by providing event-by-event estimates of uncertainty and surprise from symbolic musical sequences. However, its reference implementation is difficult to integrate with contemporary Python workflows, and its internal memory structures are not easily accessible for inspection or modification. We introduce GraphIDyOM, a graph-native Python reimplementation of IDyOM that represents long-term and short-term predictive memories as explicit graph objects while preserving the model's variable-order, multiple-viewpoint architecture. GraphIDyOM returns event-wise information content and entropy, exposes internal memory structures for analysis and export, and supports access through a local server. We validate the implementation against the original Lisp IDyOM across single, projected, and multiple-viewpoint configurations, and benchmark its coverage and computational performance against a recent reimplementation. We then demonstrate how the explicit memory representation supports network analysis of learned memories, projection of expectation values onto musical networks, recency-sensitive memory retrieval, and interactive applications. GraphIDyOM therefore provides both a faithful and accessible reimplementation of a widely used model and a platform for studying musical expectation through memory, topology, and interaction.
