微信公众号:arXiv_Daily
cs.SD语音
标题:流行音乐和爵士乐混音比例对流派适应性和弦一代的实证研究
链接:https://arxiv.org/abs/2605.04998
备注:3 figures, 5 tables. Companion HuggingFace models: https://huggingface.co/PearlLeeStudio
摘要:和弦进行生成在实际中很重要,但研究不足。大多数大型符号音乐系统的目标是旋律,多轨道安排,或音频合成,和弦模型往往被降级为较大管道内的调节组件。本文将和弦生成作为一个独立的任务,并解决了一个问题,每当这样的模型是适应跨流派:有多少旧域的数据必须保留在微调过程中获得一个新的域,而不会忘记旧的?我研究爵士乐微调开始从流行预训练25 M参数音乐Transformer(84.24%的前1和弦准确率举行了流行测试集)。可用的爵士乐语料库比流行音乐语料库小一个数量级,因此每次微调运行都使用所有1,513个爵士乐训练序列。扫描变量是与之混合的pop“排练”数据的量,取值为{0,1 K,2.5K,5 K,10 K}。每一个微调模型增益7至9点的爵士乐顶1。流行音乐的准确性崩溃了2.14点下,只有爵士乐微调,恢复到基线约2.5K排练样本(1.65倍的爵士乐音量),并饱和超过这一点。一个补充的观察:度量最佳运行(F3,2.5K混合)并不总是感知首选的。倾向流行音乐(10 K)和倾向爵士乐(1 K)的端点具有更坚定的风格特征,作者在非正式聆听中更经常选择这些风格特征作为最终输出。我讨论了这对音乐共同创作工具的建议,但没有提出感知要求,因为没有进行正式的听力研究。所有六个检查站都在https://huggingface.co/PearlLeeStudio的HuggingFace Hub上发布。
摘要:Chord progression generation is practically important but understudied. Most large-scale symbolic music systems target melody, multi-track arrangement, or audio synthesis, and chord-only models tend to be relegated to conditioning components inside larger pipelines. This paper treats chord generation as a standalone task and addresses a question that arises whenever such a model is adapted across genres: how much old-domain data must be retained during fine-tuning to acquire a new domain without forgetting the old? I study jazz fine-tuning starting from a pop-pretrained 25M-parameter Music Transformer (84.24% top-1 chord accuracy on a held-out pop test set). The available jazz corpus is an order of magnitude smaller than the pop corpus, so every fine-tune run uses all 1,513 jazz training sequences. The swept variable is the volume of pop "rehearsal" data mixed alongside, taking values in {0, 1K, 2.5K, 5K, 10K}. Every fine-tuned model gains 7 to 9 points of jazz top-1. Pop accuracy collapses by 2.14 points under jazz-only fine-tuning, recovers to baseline at approximately 2.5K rehearsal samples (1.65x the jazz volume), and saturates beyond that point. A complementary observation: the metric-best run (F3, 2.5K mix) is not always the perceptually preferred one. The pop-leaning (10K) and jazz-leaning (1K) endpoints carry more committed stylistic identities that the author more often selects as finished output in informal listening. I discuss what this suggests for music co-creation tools but make no perceptual claim, since no formal listening study has been conducted. All six checkpoints are released on the HuggingFace Hub at https://huggingface.co/PearlLeeStudio.
【2】Hearing the Ocean: Bio-inspired Gammatone-CNN framework for Robust Underwater Acoustic Target Classification
标题:聆听海洋:基于生物启发的Gammatone-CNN框架,用于稳健的水下声学目标分类链接:https://arxiv.org/abs/2605.04839
摘要:提出了一种基于生物启发的水声目标鲁棒识别方法。最新的最先进的方法往往无法解决密集的低频谐波结构的船舶推进信号在高噪声条件下,这是解决所提出的框架,使用生物启发伽玛通滤波器组,模仿耳蜗的非线性频率选择性。通过根据等效矩形带宽(ERB)尺度分布滤波器,该框架实现了发动机辐射音调的高保真表示,同时有效地抑制了各向同性环境干扰。由此产生的卷积图特征由一个轻量级的、定制设计的卷积神经网络(CNN)处理,该网络利用大的感受野来整合频谱-时间连续性。在VTUAD数据集上的实验结果表明,最先进的分类准确率为98.41%,分别优于连续小波变换和Mel频率倒谱系数基线的3.5%和7.7%。此外,该框架的推理延迟仅为0.77 ms,Cohen Kappa得分为0.971,验证了其在自主低功耗声纳硬件上实时部署的有效性。
摘要:This study presents a bio inspired signal processing framework for robust Underwater Acoustic Target Recognition (UATR). The latest state of the art methods often fail to resolve dense low frequency harmonic structures in vessel propulsion signals under high noise conditions, which is addressed by the proposed framework using a biologically inspired Gammatone filter bank that emulates the cochlea nonlinear frequency selectivity. By distributing filters according to the Equivalent Rectangular Bandwidth (ERB) scale, the framework achieves a high fidelity representation of engine radiated tonals while effectively suppressing isotropic ambient interference. The resulting Cochleagram features are processed by a lightweight, custom designed Convolutional Neural Network (CNN) that leverages large receptive fields to integrate spectral-temporal continuities. Experimental results on the VTUAD dataset demonstrate a state of the art classification accuracy of 98.41%, outperforming Continuous Wavelet Transform and Mel Frequency Cepstral Coefficients baselines by 3.5% and 7.7% respectively. Furthermore, the framework achieves an inference latency of only 0.77 ms and a 0.971 Cohen Kappa score, validating its efficacy for real time deployment on autonomous, low-power sonar hardware.
【3】Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization
标题:稀疏令牌就足够了:通过令牌感知梯度优化越狱音频语言模型链接:https://arxiv.org/abs/2605.04700
摘要:对音频语言模型(ALM)的越狱攻击优化音频扰动以引发不安全的生成,并且它们通常在整个优化过程中密集地更新整个波形。在这项工作中,我们调查这种密集优化的必要性,通过分析结构的令牌对齐的梯度在ALMs。我们发现,梯度能量是高度不均匀的音频令牌,这表明只有一个小的令牌对齐的音频区域的子集占主导地位的优化信号。受此观察的启发,我们提出了令牌感知梯度优化(TAGO),它通过仅保留与具有高梯度能量的音频令牌对齐的波形梯度来实现稀疏越狱优化,同时在每次迭代时掩蔽剩余的梯度。在三个ALM中,TAGO的表现优于基线,并且大量的稀疏化保持了强大的攻击成功率(例如,在Qwen 3-Omni上,$\mathrm{ASR}_{l}$保持在86%,令牌保留率为0.25,而完全令牌保留率为87%)。这些结果表明,密集的波形更新在很大程度上是冗余的,我们主张未来的音频越狱和安全对齐研究应该进一步利用这种异构的令牌级梯度结构。
摘要:Jailbreak attacks on audio language models (ALMs) optimize audio perturbations to elicit unsafe generations, and they typically update the entire waveform densely throughout optimization. In this work, we investigate the necessity of such dense optimization by analyzing the structure of token-aligned gradients in ALMs. We find that gradient energy is highly non-uniform across audio tokens, indicating that only a small subset of token-aligned audio regions dominates the optimization signal. Motivated by this observation, we propose Token-Aware Gradient Optimization (TAGO), which enables sparse jailbreak optimization by retaining only waveform gradients aligned with audio tokens that have high gradient energy, while masking the remaining gradients at each iteration. Across three ALMs, TAGO outperforms baselines, and substantial sparsification preserves strong attack success rates (e.g. on Qwen3-Omni, $\mathrm{ASR}_{l}$ remains at 86% with a token retention ratio of 0.25, compared to 87% with full token retention). These results demonstrate that dense waveform updates are largely redundant, and we advocate that future audio jailbreak and safety alignment research should further leverage this heterogeneous token-level gradient structure.
【4】VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
标题:VocalParse:利用大型音频语言模型实现统一和可扩展的歌唱声音转录链接:https://arxiv.org/abs/2605.04613
摘要:高质量的歌唱标注是现代歌唱语音合成(SVS)系统的基础。然而,由于需要大量的劳动力和音乐专业知识,通过手动标记来大规模地获得这些注释是不现实的,使得自动注释非常必要。尽管它们的实用性,当前的自动转录系统面临着重大的挑战:它们往往依赖于复杂的多级管道,努力恢复文本注释对齐,并表现出对分布外(OOD)歌唱数据的泛化能力差。为了缓解这些问题,我们提出了VocalParse,一个统一的歌唱语音转录(SVT)模型建立在一个大型音频语言模型(LALM)。具体来说,我们的新贡献是引入一个交错的提示制定,共同模型歌词,旋律,和单词音符对应,产生一个生成的序列,直接映射到一个结构化的乐谱。此外,我们提出了一个思想链(CoT)风格的提示策略,首先解码歌词作为语义支架,显着减轻上下文中断的问题,同时保留交错生成的结构优势。实验表明,VocalParse在多个歌唱数据集上实现了最先进的SVT性能。源代码和检查点可以在https://github.com/pymaster17/VocalParse上找到。
摘要:High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.
【5】Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)
标题:对LLM进行大规模声音嵌入基准(MSEB)的基准链接:https://arxiv.org/abs/2605.04556
摘要:Massive Sound Embedding Benchmark(MSEB)已成为评估音频模型功能广度的标准。虽然最初的基线集中在专门的编码器上,但向“音频原生”大型语言模型(LLM)的转变表明了一种新的范式,即单一的多模态主干可以取代复杂的特定任务管道。本文提供了一个严格的经验评估领先的法学硕士-包括双子座和GPT家庭的成员-在八个核心MSEB能力,以评估其功效和音频文本的奇偶性。我们的研究结果表明,虽然一个显着的模态差距持续有关的性能和鲁棒性,经验证据的“最佳”建模方法仍然是不确定的。最终,听觉和级联架构之间的选择在很大程度上取决于特定的用例需求以及关于延迟,成本和推理深度的基本假设。
摘要:The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines focused on specialized encoders, the shift toward "audio-native" Large Language Models (LLMs) suggests a new paradigm where a single multimodal backbone may replace complex, task-specific pipelines. This paper provides a rigorous empirical evaluation of leading LLMs - including members from the Gemini and GPT families - across the eight core MSEB capabilities to assess their efficacy and audio-text parity. Our results indicate that while a significant modality gap persists regarding performance and robustness, the empirical evidence for an "optimal" modeling approach remains inconclusive. Ultimately, the choice between audionative and cascaded architectures depends heavily on specific use-case requirements and the underlying assumptions regarding latency, cost, and reasoning depth.
【6】Stage-adaptive audio diffusion modeling
标题:阶段自适应音频扩散建模链接:https://arxiv.org/abs/2605.04547
摘要:基于扩散的音频生成和恢复的最新进展大大提高了异构调节机制的性能,包括文本调节音频生成和音频调节超分辨率。然而,训练音频扩散模型在计算上仍然是昂贵的,并且大多数现有的流水线仍然依赖于静态优化配方,该静态优化配方在整个学习过程中将训练信号的相对重要性视为固定的。在这项工作中,我们认为,效率低下的一个主要来源在于语义采集和面向生成的细化之间的不断发展的平衡。早期的训练更加强调获得条件一致的语义结构和粗略的全局组织,而后期的训练越来越强调时间一致性,感知保真度和细节细化。为了表征这种不断变化的平衡,我们引入了一个基于进度的制度变量,该变量来自SSL空间差异的训练时间斜率,它可以测量训练过程中的语义进展。基于这个信号,我们开发了三个互补的阶段感知机制:早期语义引导的衰减SSL指导,由政权变量驱动的自适应时间步长采样,以及从参数空间中的收敛分组组织激活的结构感知正则化。我们评估这些机制的文本条件音频生成和音频条件超分辨率。在这两种设置中,所提出的阶段感知策略改善了收敛行为,并在标准静态基线上获得了主要生成和频谱重建指标的收益。这些结果支持了这样的观点,即有效的音频扩散训练可以受益于将外部指导,内部组织和优化强调作为依赖于阶段的组件,而不是固定的成分。
摘要:Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution. However, training audio diffusion models remains computationally expensive, and most existing pipelines still rely on static optimization recipes that treat the relative importance of training signals as fixed throughout learning. In this work, we argue that a major source of inefficiency lies in the evolving balance between semantic acquisition and generation-oriented refinement. Early training places stronger emphasis on acquiring condition-aligned semantic structure and coarse global organization, whereas later training increasingly emphasizes temporal consistency, perceptual fidelity, and fine-detail refinement. To characterize this evolving balance, we introduce a progress-based regime variable derived from the training-time slope of an SSL-space discrepancy, which measures semantic progress during training. Based on this signal, we develop three complementary stage-aware mechanisms: decayed SSL guidance for early semantic bootstrapping, self-adaptive timestep sampling driven by the regime variable, and structure-aware regularization activated from convergent grouped organization in parameter space. We evaluate these mechanisms on text-conditioned audio generation and audio-conditioned super-resolution. Across both settings, the proposed stage-aware strategies improve convergence behavior and yield gains on the primary generation and spectral reconstruction metrics over standard static baselines. These results support the view that efficient audio diffusion training can benefit from treating external guidance, internal organization, and optimization emphasis as stage-dependent components rather than fixed ingredients.
【7】Adaptive Diagonal Loading for Norm Constrained Beamforming
标题:规范约束束成形的自适应对角加载链接:https://arxiv.org/abs/2605.04342
备注:5 pages, 5 figures
摘要:可靠的自适应波束形成对于在高动态声学环境中工作的大型麦克风阵列至关重要。在以快速移动的说话者和说话者为特征的场景中,用于估计空间相关矩阵的可用样本支持通常是快照不足的。这种缺陷,加上阵列的缺陷,降低了白噪声增益(WNG),导致严重的目标信号抵消。为了保证波束形成的稳定性和鲁棒性,我们提出了一种新的自适应对角加载方法,保证WNG严格保持在指定的范围内。通过利用Kantorovich不等式,我们将所需的WNG映射到相关矩阵条件数的严格上界。此外,我们提出了三种自适应加载水平的估计技术,从基于跟踪的边界到精确的特征值分解,提供可扩展的计算复杂度$\mathcal{O}(M)$,$\mathcal{O}(M^2)$和$\mathcal{O}(M^3)$。我们的方法在快速变化的干扰下表现出高度稳定的波束形成。
摘要:Reliable adaptive beamforming is critical for large microphone arrays operating in highly dynamic acoustic environments. In scenarios characterized by fast-moving talkers and interferers, the available sample support for estimating the spatial correlation matrix is often snapshot-deficient. This deficiency, coupled with array imperfections, degrades the White Noise Gain (WNG), leading to severe target signal cancellation. To ensure stable and robust beamforming, we propose a novel adaptive diagonal loading method that guarantees the WNG remains strictly within specified bounds. By leveraging the Kantorovich inequality, we map the desired WNG to a strict upper bound on the condition number of the correlation matrix. Furthermore, we present three estimation techniques for the adaptive loading level, ranging from trace-based bounding to exact eigenvalue decomposition, offering scalable computational complexities of $\mathcal{O}(M)$, $\mathcal{O}(M^2)$, and $\mathcal{O}(M^3)$. Our approach demonstrates highly stable beamforming under fast-changing interference.
【8】JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
标题:JASTIN:通过自然语言指令调整LLM以进行Zero-Shot音频和语音评估链接:https://arxiv.org/abs/2605.04505
摘要:生成音频模型的快速发展已经超过了强大的评估方法的发展。现有的客观指标和通用多模态大型语言模型(MLLM)往往与领域泛化,zero-shot能力和教学灵活性的斗争。为了解决这些瓶颈,我们提出了JASTIN,一个可推广的,推理驱动的音频评估框架,制定音频评估作为一个自我指导的推理任务。JASTIN通过可训练的音频适配器将冻结的高性能音频编码器与微调的LLM主干连接起来。为了确保鲁棒的zero-shot泛化,我们引入了一个全面的指令后的数据准备管道,将多源,多任务,多校准,多描述数据。实验结果表明,JASTIN实现了与人类主观评级的最先进的Pearson和Spearman相关性。它在语音,声音,音乐和域外评估任务方面始终优于一般MLLM,而无需进行特定任务的再培训。
摘要:The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language models (MLLMs) often struggle with domain generalization, zero-shot capabilities, and instructional flexibility. To address these bottlenecks, we propose JASTIN, a generalizable, instruction-driven audio evaluation framework that formulates audio assessment as a self-instructed reasoning task. JASTIN bridges a frozen high-performance audio encoder with a fine-tuned LLM backbone via a trainable audio adapter. To ensure robust zero-shot generalization, we introduce a comprehensive instruction following data preparation pipeline, incorporating Multi-Source, Multi-Task, Multi-Calibration, and Multi-Description data. Experimental results demonstrate that JASTIN achieves state-of-the-art Pearson and Spearman correlations with human subjective ratings. It consistently outperforms general MLLMs across speech, sound, music, and out-of-domain evaluation tasks without the need for task-specific retraining.
标题:空间放大镜:用于多通道语音增强的空间上采样
链接:https://arxiv.org/abs/2605.04749
备注:5 pages, 2 figures, 4 tables
摘要:虽然多通道语音增强算法的空间方向性随着麦克风的数量而改善,但是将大型捕获阵列适配到现实世界的边缘设备中通常受到物理约束的限制。为了克服这一限制,我们提出了空间放大器,一个神经网络,旨在从一组有限的真实麦克风(RM)测量生成虚拟麦克风(VM)信号。此外,我们引入了空间音频表示学习(SARL)框架,该框架利用估计的VM信号和特征来调节下游语音增强系统。实验结果表明,所提出的框架优于现有的空间上采样基线在各种语音提取系统,包括端到端的多通道语音增强和神经波束形成。所提出的方法几乎恢复了所有麦克风可用时所达到的预言性能。
摘要:While the spatial directivity of multichannel speech enhancement algorithms improves with the number of microphones, fitting large capture arrays into real-world edge devices is typically limited by physical constraints. To overcome this limitation, we propose Spatial-Magnifier, a neural network designed to generate virtual microphone (VM) signals from a limited set of real microphone (RM) measurements. Moreover, we introduce the Spatial Audio Representation Learning (SARL) framework, which leverages estimated VM signals and features to condition a downstream speech enhancement system. Experimental results demonstrate that the proposed framework outperforms existing spatial upsampling baselines across various speech extraction systems, including end-to-end multichannel speech enhancement and neural beamforming. The proposed method nearly recovers the oracle performance achieved when all microphones are available.
【2】JASTIN: Aligning LLMs for Zero-Shot Audio and Speech Evaluation via Natural Language Instructions
标题:JASTIN:通过自然语言指令调整LLM以进行Zero-Shot音频和语音评估链接:https://arxiv.org/abs/2605.04505
摘要:生成音频模型的快速发展已经超过了强大的评估方法的发展。现有的客观指标和通用多模态大型语言模型(MLLM)往往与领域泛化,zero-shot能力和教学灵活性的斗争。为了解决这些瓶颈,我们提出了JASTIN,一个可推广的,推理驱动的音频评估框架,制定音频评估作为一个自我指导的推理任务。JASTIN通过可训练的音频适配器将冻结的高性能音频编码器与微调的LLM主干连接起来。为了确保鲁棒的zero-shot泛化,我们引入了一个全面的指令后的数据准备管道,将多源,多任务,多校准,多描述数据。实验结果表明,JASTIN实现了最先进的皮尔逊和斯皮尔曼相关性与人类的主观评级。它在语音,声音,音乐和域外评估任务方面始终优于一般MLLM,而无需进行特定任务的再培训。
摘要:The rapid advancement of generative audio models has outpaced the development of robust evaluation methodologies. Existing objective metrics and general multimodal large language models (MLLMs) often struggle with domain generalization, zero-shot capabilities, and instructional flexibility. To address these bottlenecks, we propose JASTIN, a generalizable, instruction-driven audio evaluation framework that formulates audio assessment as a self-instructed reasoning task. JASTIN bridges a frozen high-performance audio encoder with a fine-tuned LLM backbone via a trainable audio adapter. To ensure robust zero-shot generalization, we introduce a comprehensive instruction following data preparation pipeline, incorporating Multi-Source, Multi-Task, Multi-Calibration, and Multi-Description data. Experimental results demonstrate that JASTIN achieves state-of-the-art Pearson and Spearman correlations with human subjective ratings. It consistently outperforms general MLLMs across speech, sound, music, and out-of-domain evaluation tasks without the need for task-specific retraining.
机器翻译由腾讯交互翻译提供,仅供参考
