微信公众号:arXiv_Daily
cs.SD语音
【1】AcousTools: A `Full-Stack', Python-Based, Acoustic Holography Library
标题:AcoustTools:一个“全栈”、基于Python的声学全息图书馆
链接:https://arxiv.org/abs/2511.07336
备注:14 Pages, 7 Figures, 2 Tables, To be submitted to APL Computational Physics
摘要:声全息是一个新兴的领域,其中半空中的超声波是控制和操纵的新颖和令人兴奋的应用。这些应用包括半空触觉、体积显示、非接触式制造,甚至化学和生物医学应用,如药物输送。为了开发这些应用,需要一个软件框架来预测声学行为并模拟所产生的效果,例如施加的力或散射模式。已经有各种各样的软件库和平台试图填补这个角色,但还没有一个单一的软件作为一个“全栈”的解决方案。我们将此全栈定义为从抽象到物理化的过程,从设置开始,建模声学传播,换能器相位检索,声场分析和声学全息硬件本身的控制。现有方法不能满足这些类别中的一个或多个。为了解决这个问题,我们提出了Acoustools,一个基于Python的声学全息库,旨在支持全套声学全息应用程序,我们展示了Acoustools满足全栈要求的每一步的能力。Acoustools有可能成为声全息的标准代码库,其独特的完整功能套件封装在一种众所周知易于使用的语言中,Acoustools将提高研究人员开发新颖应用程序以及准确审查的能力。其他人的工作。除了软件之外,全栈也将对研究人员有用-通过了解它们在堆栈中的位置,提供一种查看和比较方法的方法。
摘要:Acoustic Holography is an emerging field where mid-air ultrasound is controlled and manipulated for novel and exciting applications. These range from mid-air haptics, volumetric displays, contactless fabrication, and even chemical and biomedical applications such as drug delivery. To develop these applications, a software framework to predict acoustic behaviour and simulating resulting effects, such as applied forces or scattering patterns is desirable. There have been various software libraries and platforms that attempt to fill this role, but there is yet to be a single piece of software that acts as a 'full-stack' solution. We define this full-stack as the process from abstraction to physicalisation starting with setup, modelling acoustic propagation, transducer phase retrieval, sound field analysis, and control of the acoustic holographic hardware itself. Existing methods fail to fulfil one or more of these categories. To address this, we present AcousTools, a Python-based acoustic holography library, designed to support the full suite of acoustic holographic applications and we show AcousTools's ability to meet each step of the full-stack's requirements. AcousTools has the potential to become the standard code library for acoustic holography, with the uniquely complete suite of features wrapped in a language that is known to be easy to use, AcousTools will increase the ability for researchers to develop novel applications as well as accurately review other's work. The full-stack, aside from software, will also be useful for researchers - providing a way to view and compare methodologies by understanding where they fit into the stack.
【2】Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
标题:用Transformer生成钢琴音乐:规模、数据和时间的比较研究
链接:https://arxiv.org/abs/2511.07268
备注:NeurIPS 2025 Workshop on AI for Music
摘要:虽然近年来已经提出了各种Transformers的符号音乐生成,仍然很少有全面的研究,具体的设计选择如何影响生成的音乐的质量。在这项工作中,我们系统地比较了不同的数据集,模型架构,模型大小和符号钢琴音乐生成任务的训练策略。为了支持模型的开发和评估,我们研究了一系列定量指标,并分析了它们与通过听力研究收集的人类判断的相关性。我们表现最好的模型是一个950 M参数的Transformer,它在来自不同流派的80K音频文件上训练,产生的输出在图灵风格的听力调查中通常被评为人类组成。
摘要:Although a variety of transformers have been proposed for symbolic music generation in recent years, there is still little comprehensive study on how specific design choices affect the quality of the generated music. In this work, we systematically compare different datasets, model architectures, model sizes, and training strategies for the task of symbolic piano music generation. To support model development and evaluation, we examine a range of quantitative metrics and analyze how well they correlate with human judgment collected through listening studies. Our best-performing model, a 950M-parameter transformer trained on 80K MIDI files from diverse genres, produces outputs that are often rated as human-composed in a Turing-style listening survey.
【3】Twenty-Five Years of MIR Research: Achievements, Practices, Evaluations, and Future Challenges
标题:MIR研究二十五年:成就、实践、评估和未来挑战
链接:https://arxiv.org/abs/2511.07205
摘要:在本文中,我们跟踪音乐信息检索(MIR)在过去25年的演变。虽然MIR收集了与音乐信息学相关的各种研究,但其中很大一部分集中在音乐数据的信号处理技术上,与IEEE音频和声学信号处理技术委员会建立了密切的关系。本文沿着与音乐分析、处理和生成相关的三个EDICS反映了MIR的主要研究成果。然后,我们回顾了一组成功的做法,燃料的MIR研究的快速发展。一种做法是年度研究基准,音乐信息检索评估交换,参与者在一组研究任务上竞争。另一种做法是追求可复制和开放的研究。积极参与行业研究和产品是实现巨大社会影响和激励年轻一代学生加入该领域的另一个关键因素。最后但并非最不重要的是,对多样性,公平性和包容性的承诺确保MIR成为一个充满活力和开放的社区,各种想法,方法和职业道路相互碰撞。最后,我们提供了MIR将不得不面对的未来挑战。
摘要:In this paper, we trace the evolution of Music Information Retrieval (MIR) over the past 25 years. While MIR gathers all kinds of research related to music informatics, a large part of it focuses on signal processing techniques for music data, fostering a close relationship with the IEEE Audio and Acoustic Signal Processing Technical Commitee. In this paper, we reflect the main research achievements of MIR along the three EDICS related to music analysis, processing and generation. We then review a set of successful practices that fuel the rapid development of MIR research. One practice is the annual research benchmark, the Music Information Retrieval Evaluation eXchange, where participants compete on a set of research tasks. Another practice is the pursuit of reproducible and open research. The active engagement with industry research and products is another key factor for achieving large societal impacts and motivating younger generations of students to join the field. Last but not the least, the commitment to diversity, equity and inclusion ensures MIR to be a vibrant and open community where various ideas, methodologies, and career pathways collide. We finish by providing future challenges MIR will have to face.
【4】Generating Novel and Realistic Speakers for Voice Conversion
标题:生成新颖且真实的扬声器以进行语音转换
链接:https://arxiv.org/abs/2511.07135
摘要:语音转换模型修改音色,同时保留非语言特征,支持配音和身份保护等应用。然而,大多数VC系统需要访问目标话语,从而在目标数据不可用或用户希望转换为全新、不可见的声音时限制了它们的使用。为了解决这个问题,我们引入了一个轻量级的方法SpeakerVAE来生成新的VC扬声器。我们的方法使用一个深层次的变分自动编码器来模拟扬声器音色空间。通过从训练好的模型中采样,我们在VC管道中生成了用于语音合成的新的说话人表示。所提出的方法是一个灵活的插件模块兼容各种VC模型,没有共同训练或微调的基础VC系统。我们使用最先进的VC模型评估了我们的方法:FACodec和CosyVoice2。结果表明,我们的方法成功地产生了新的,看不见的扬声器与质量相媲美的训练扬声器。
摘要:Voice conversion models modify timbre while preserving paralinguistic features, enabling applications like dubbing and identity protection. However, most VC systems require access to target utterances, limiting their use when target data is unavailable or when users desire conversion to entirely novel, unseen voices. To address this, we introduce a lightweight method SpeakerVAE to generate novel speakers for VC. Our approach uses a deep hierarchical variational autoencoder to model the speaker timbre space. By sampling from the trained model, we generate novel speaker representations for voice synthesis in a VC pipeline. The proposed method is a flexible plug-in module compatible with various VC models, without co-training or fine-tuning of the base VC system. We evaluated our approach with state-of-the-art VC models: FACodec and CosyVoice2. The results demonstrate that our method successfully generates novel, unseen speakers with quality comparable to that of the training speakers.
【5】BridgeVoC: Revitalizing Neural Vocoder from a Restoration Perspective
标题:BridgeVoC:从恢复角度重振神经声码器
链接:https://arxiv.org/abs/2511.07116
备注:18 pages, 16 figures
摘要:本文通过音频恢复的镜头重新审视了神经声码器任务,并提出了一种新的扩散声码器称为BridgeVoC。具体来说,通过秩分析,我们比较梅尔频谱的秩特性与其他常见的声学退化因素,并铸造声码器任务作为一个特殊的情况下,音频恢复,其中的距离空间频谱(RSS)替代目标频谱作为退化的输入。在此基础上,我们介绍了薛定谔桥框架的扩散建模,它定义的RSS和目标光谱的随机生成轨迹的双端点。此外,为了充分利用时频域中子带的分层先验,我们精心设计了一种新的子带感知卷积扩散网络作为数据预测器,其中子带按照不均匀策略进行划分,并采用卷积式注意力模块,采用大内核进行有效的T-F上下文建模。为了实现单步推理,我们提出了一个全向蒸馏损失,以促进有效的信息从教师模型到学生模型的传输,并通过结合目标相关和双射一致性损失的性能得到改善。在各种基准测试和分布外数据集上进行了全面的实验。定量和定性的结果表明,同时享有更少的参数,更低的计算成本,和竞争力的推理速度,建议BridgeVoC产量最先进的性能比现有的先进的GAN,DDPM和流量匹配为基础的基线,只有4个采样步骤。并且通过单步推理仍然可以获得一致的优越性。
摘要:This paper revisits the neural vocoder task through the lens of audio restoration and propose a novel diffusion vocoder called BridgeVoC. Specifically, by rank analysis, we compare the rank characteristics of Mel-spectrum with other common acoustic degradation factors, and cast the vocoder task as a specialized case of audio restoration, where the range-space spectral (RSS) surrogate of the target spectrum acts as the degraded input. Based on that, we introduce the Schrodinger bridge framework for diffusion modeling, which defines the RSS and target spectrum as dual endpoints of the stochastic generation trajectory. Further, to fully utilize the hierarchical prior of subbands in the time-frequency (T-F) domain, we elaborately devise a novel subband-aware convolutional diffusion network as the data predictor, where subbands are divided following an uneven strategy, and convolutional-style attention module is employed with large kernels for efficient T-F contextual modeling. To enable single-step inference, we propose an omnidirectional distillation loss to facilitate effective information transfer from the teacher model to the student model, and the performance is improved by combining target-related and bijective consistency losses. Comprehensive experiments are conducted on various benchmarks and out-of-distribution datasets. Quantitative and qualitative results show that while enjoying fewer parameters, lower computational cost, and competitive inference speed, the proposed BridgeVoC yields stateof-the-art performance over existing advanced GAN-, DDPMand flow-matching-based baselines with only 4 sampling steps. And consistent superiority is still achieved with single-step inference.
【6】E2E-VGuard: Adversarial Prevention for Production LLM-based End-To-End Speech Synthesis
标题:E2 E-VGuard:基于生产LLM的端到端语音合成的对抗预防
链接:https://arxiv.org/abs/2511.07099
备注:Accepted to NeurIPS 2025
摘要:语音合成技术的最新进展丰富了我们的日常生活,高质量和人性化的音频在现实世界的应用中被广泛采用。然而,像语音克隆欺诈这样的恶意利用会带来严重的安全风险。现有的防御技术难以解决基于大语言模型(LLM)的语音合成问题.虽然之前的研究考虑了对微调合成器的保护,但它们假设手动注释的转录本。考虑到人工注释的劳动强度,利用自动语音识别(ASR)来生成转录本的端到端(E2 E)系统变得越来越普遍,例如,通过商业API进行语音克隆。因此,这种E2 E语音合成也需要新的安全机制。为了应对这些挑战,我们提出了E2 E-VGuard,这是一个针对两种新兴威胁的主动防御框架:(1)基于生产LLM的语音合成,以及(2)由ASR驱动的E2 E场景引起的新型攻击。具体来说,我们使用带有特征提取器的编码器集成来保护音色,而针对ASR的对抗性示例会破坏发音。此外,我们将心理声学模型,以确保扰动不可感知。为了进行全面的评估,我们在中文和英文数据集上测试了16个开源合成器和3个商业API,证实了E2 E-VGuard在音色和发音保护方面的有效性。还进行了实际部署验证。我们的代码和演示页面可以在https://wxzyd123.github.io/e2e-vguard/上找到。
摘要:Recent advancements in speech synthesis technology have enriched our daily lives, with high-quality and human-like audio widely adopted across real-world applications. However, malicious exploitation like voice-cloning fraud poses severe security risks. Existing defense techniques struggle to address the production large language model (LLM)-based speech synthesis. While previous studies have considered the protection for fine-tuning synthesizers, they assume manually annotated transcripts. Given the labor intensity of manual annotation, end-to-end (E2E) systems leveraging automatic speech recognition (ASR) to generate transcripts are becoming increasingly prevalent, e.g., voice cloning via commercial APIs. Therefore, this E2E speech synthesis also requires new security mechanisms. To tackle these challenges, we propose E2E-VGuard, a proactive defense framework for two emerging threats: (1) production LLM-based speech synthesis, and (2) the novel attack arising from ASR-driven E2E scenarios. Specifically, we employ the encoder ensemble with a feature extractor to protect timbre, while ASR-targeted adversarial examples disrupt pronunciation. Moreover, we incorporate the psychoacoustic model to ensure perturbative imperceptibility. For a comprehensive evaluation, we test 16 open-source synthesizers and 3 commercial APIs across Chinese and English datasets, confirming E2E-VGuard's effectiveness in timbre and pronunciation protection. Real-world deployment validation is also conducted. Our code and demo page are available at https://wxzyd123.github.io/e2e-vguard/.
【7】Metric Analysis for Spatial Semantic Segmentation of Sound Scenes
标题:声音场景空间语义分割的度量分析
链接:https://arxiv.org/abs/2511.07075
备注:5 pages; content+bibliography
摘要:声音场景的空间语义分割(S5)包括从多通道音频混合物联合执行音频源分离和声音事件分类。为了评估S5系统,可以考虑两个单独的指标,即一个用于源分离,另一个用于声音事件分类,但是这种方法使得比较S5系统具有挑战性。因此,联合类感知的信号失真比(CA-SDR)的度量被提出来评估S5系统。在这项工作中,我们首先比较了CA-SDR与经典SDR的情况下,只有分类错误。然后,我们分析的情况下,度量可能不允许适当的比较系统。为了解决这个问题,我们提出了一个修改后的版本的CA-SDR,首先专注于类不可知的SDR,然后占错误标记的来源。我们还分析了这两个指标的性能下的交叉污染分离的音频源之间。最后,我们提出了第一组惩罚,试图使度量更能反映标签和分离错误。
摘要:Spatial semantic segmentation of sound scenes (S5) consists of jointly performing audio source separation and sound event classification from a multichannel audio mixture. To evaluate S5 systems, one can consider two individual metrics, i.e., one for source separation and another for sound event classification, but this approach makes it challenging to compare S5 systems. Thus, a joint class-aware signal-to-distortion ratio (CA-SDR) metric was proposed to evaluate S5 systems. In this work, we first compare the CA-SDR with the classical SDR on scenarios with only classification errors. We then analyze the cases where the metric might not allow proper comparison of the systems. To address this problem, we propose a modified version of the CA-SDR which first focuses on class-agnostic SDR and then accounts for the wrongly labeled sources. We also analyze the performance of the two metrics under cross-contamination between separated audio sources. Finally, we propose a first set of penalties in an attempt to make the metric more reflective of the labeling and separation errors.
【8】CLiFT-ASR: A Cross-Lingual Fine-Tuning Framework for Low-Resource Taiwanese Hokkien Speech Recognition
标题:CLiFT-ASB:用于低资源台湾闽南语语音识别的跨语言微调框架
链接:https://arxiv.org/abs/2511.06860
备注:Accepted for an oral presentation at the 37th Conference on Computational Linguistics and Speech Processing (ROCLING 2025)
摘要:自动语音识别(ASR)的低资源的语言,如闽南语是困难的,由于缺乏注释的数据。然而,直接对汉字拼音进行微调往往无法捕捉到详细的语音和音调线索,而只对罗马化进行训练则缺乏词汇和句法覆盖。此外,以前的研究很少探索阶段性的策略,整合这两种注释类型。为了解决这个差距,我们提出了CLiFT-ASR,一个跨语言微调框架,建立在普通话休伯特模型,并逐步适应台湾闽南语。该框架采用了两个阶段的过程中,它首先学习声学和音调表示从语音Tai-lo注释,然后捕获的词汇和句法从汉字音译。这种渐进的适应使得语音和正字法结构之间能够有效地对齐。在TAT-MOE语料库上的实验表明,与强基线相比,CLiFT-ASR的字符错误率(CER)相对降低了24.88%.结果表明,CLiFT-ASR为台湾闽南语ASR提供了一个有效的和参数高效的解决方案,它有可能使其他低资源的语言场景受益。
摘要:Automatic speech recognition (ASR) for low-resource languages such as Taiwanese Hokkien is difficult due to the scarcity of annotated data. However, direct fine-tuning on Han-character transcriptions often fails to capture detailed phonetic and tonal cues, while training only on romanization lacks lexical and syntactic coverage. In addition, prior studies have rarely explored staged strategies that integrate both annotation types. To address this gap, we present CLiFT-ASR, a cross-lingual fine-tuning framework that builds on Mandarin HuBERT models and progressively adapts them to Taiwanese Hokkien. The framework employs a two-stage process in which it first learns acoustic and tonal representations from phonetic Tai-lo annotations and then captures vocabulary and syntax from Han-character transcriptions. This progressive adaptation enables effective alignment between speech sounds and orthographic structures. Experiments on the TAT-MOE corpus demonstrate that CLiFT-ASR achieves a 24.88\% relative reduction in character error rate (CER) compared with strong baselines. The results indicate that CLiFT-ASR provides an effective and parameter-efficient solution for Taiwanese Hokkien ASR and that it has potential to benefit other low-resource language scenarios.
【9】SAR-LM: Symbolic Audio Reasoning with Large Language Models
标题:SAR-LM:使用大型语言模型的符号音频推理
链接:https://arxiv.org/abs/2511.06483
摘要:大型语言模型(LLM)在文本和视觉方面取得了进步,但它们对音频的推理仍然有限。大多数现有的方法依赖于密集的音频嵌入,这是难以解释的,往往失败的结构化推理任务。在最近的基准测试(如MMAU)中引入的基于标题的方法通过将音频转换为文本来提高性能,但仍然依赖于密集的嵌入作为输入,当模型失败时几乎没有洞察力。 我们提出了SAR-LM,一个符号音频推理管道,建立在这个基于字幕的范例,通过将音频转换为结构化的,人类可读的功能,跨越语音,声音事件和音乐。这些符号输入支持推理和透明的错误分析,使我们能够将故障追溯到特定的功能。在MMAU、MMAR和OmniBench三个基准测试中,SAR-LM取得了有竞争力的结果,同时优先考虑可解释性作为其主要贡献。
摘要:Large language models (LLMs) have advanced in text and vision, but their reasoning on audio remains limited. Most existing methods rely on dense audio embeddings, which are difficult to interpret and often fail on structured reasoning tasks. Caption-based approaches, introduced in recent benchmarks such as MMAU, improve performance by translating audio into text, yet still depend on dense embeddings as input, offering little insight when models fail. We present SAR-LM, a symbolic audio reasoning pipeline that builds on this caption-based paradigm by converting audio into structured, human-readable features across speech, sound events, and music. These symbolic inputs support both reasoning and transparent error analysis, enabling us to trace failures to specific features. Across three benchmarks, MMAU, MMAR, and OmniBench, SAR-LM achieves competitive results, while prioritizing interpretability as its primary contribution.
【10】EchoMark: Perceptual Acoustic Environment Transfer with Watermark-Embedded Room Impulse Response
标题:EchoMark:具有嵌入水印的房间脉冲响应的感知声学环境传输
链接:https://arxiv.org/abs/2511.06458
摘要:声学环境匹配(AEM)是将干净的音频传输到目标声学环境中的任务,从而实现音频配音和听觉沉浸式虚拟现实(VR)等引人入胜的应用。直接从混响语音中恢复相似的房间脉冲响应(RIR)提供了更容易和灵活的AEM解决方案。然而,如果被恶意用户滥用,这种能力也会引入任意“重新定位”的漏洞,例如促进高级语音欺骗攻击或破坏记录证据的真实性。为了解决这个问题,我们提出了EchoMark,这是第一个基于深度学习的AEM框架,可以生成具有嵌入水印的感知相似RIR。我们的设计通过在潜在域中操作来解决可变RIR特性(例如不同的持续时间和能量衰减)所带来的挑战。通过联合优化RIR重建的感知损失和水印检测损失的模型,EchoMark实现了高质量的环境传输和可靠的水印恢复。在不同数据集上的实验验证了EchoMark实现的室内声学参数匹配性能可与最先进的RIR估计器FiNS相媲美。此外,高平均意见分数(MOS)的4.22出5,水印检测准确率超过99%,误码率(BER)低于0.3%,共同证明了EchoMark在保持感知质量,同时确保可靠的水印嵌入的有效性。
摘要:Acoustic Environment Matching (AEM) is the task of transferring clean audio into a target acoustic environment, enabling engaging applications such as audio dubbing and auditory immersive virtual reality (VR). Recovering similar room impulse response (RIR) directly from reverberant speech offers more accessible and flexible AEM solution. However, this capability also introduces vulnerabilities of arbitrary ``relocation" if misused by malicious user, such as facilitating advanced voice spoofing attacks or undermining the authenticity of recorded evidence. To address this issue, we propose EchoMark, the first deep learning-based AEM framework that generates perceptually similar RIRs with embedded watermark. Our design tackle the challenges posed by variable RIR characteristics, such as different durations and energy decays, by operating in the latent domain. By jointly optimizing the model with a perceptual loss for RIR reconstruction and a loss for watermark detection, EchoMark achieves both high-quality environment transfer and reliable watermark recovery. Experiments on diverse datasets validate that EchoMark achieves room acoustic parameter matching performance comparable to FiNS, the state-of-the-art RIR estimator. Furthermore, a high Mean Opinion Score (MOS) of 4.22 out of 5, watermark detection accuracy exceeding 99\%, and bit error rates (BER) below 0.3\% collectively demonstrate the effectiveness of EchoMark in preserving perceptual quality while ensuring reliable watermark embedding.
【11】MT-HuBERT: Self-Supervised Mix-Training for Few-Shot Keyword Spotting in Mixed Speech
标题:MT-HuBERT:自监督混合训练,用于混合语音中的Few-Shot关键词发现
链接:https://arxiv.org/abs/2511.06296
摘要:Few-Shot关键字定位的目的是用非常有限的标记样本检测以前未见过的关键字。一个预训练和适应范式通常采用这项任务。虽然在干净的条件下有效,但大多数现有的方法都难以进行混合关键字识别-在单个话语中检测多个重叠的关键字-这是现实世界应用程序所必需的功能。我们之前提出了一种基于混合训练(MT)的预训练方法来解决混合关键字检测问题,并证明了其效率。然而,这种方法是完全监督的,无法利用大量未标记的数据。为此,我们提出了混合训练HuBERT(MT-HuBERT),这是一种自监督学习(SSL)预训练框架,在预训练期间实现MT标准。MT-HuBERT以自我监督的方式预测来自上下文线索的每个组成信号的干净声学单元,而不是预测混合语音的组成模式。在Google Speech Commands(GSC v2)语料库上进行的实验表明,我们提出的MT-HuBERT在混合和干净条件下,在Few-Shot KWS任务中的性能始终优于几个最先进的基线。
摘要:Few-shot keyword spotting aims to detect previously unseen keywords with very limited labeled samples. A pre-training and adaptation paradigm is typically adopted for this task. While effective in clean conditions, most existing approaches struggle with mixed keyword spotting--detecting multiple overlapping keywords within a single utterance--a capability essential for real-world applications. We have previously proposed a pre-training approach based on Mix-Training (MT) to tackle the mixed keyword detection problem and demonstrated its efficiency. However, this approach is fully supervised, unable to utilize vast unlabeled data. To this end, we propose Mix-Training HuBERT (MT-HuBERT), a self-supervised learning (SSL) pre-training framework that implements the MT criterion during pre-training. MT-HuBERT predicts, in a self-supervised manner, the clean acoustic units of each constituent signal from contextual cues, in contrast to predicting compositional patterns of mixed speech. Experiments conducted on the Google Speech Commands (GSC v2) corpus demonstrate that our proposed MT-HuBERT consistently outperforms several state-of-the-art baselines in few-shot KWS tasks under both mixed and clean conditions.
【12】ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
标题:ELEGANCE:用于视听目标语音提取的高效LLM制导
链接:https://arxiv.org/abs/2511.06288
摘要:视听目标说话人提取(AV-TSE)模型主要依赖于来自目标说话人的视觉线索。然而,人类也利用语言学知识,如句法约束,下一个词的预测,和会话的先验知识,提取目标语音。受这一观察的启发,我们提出了ELEGANCE,这是一个新的框架,它通过三种不同的指导策略将来自大型语言模型(LLM)的语言知识整合到AV-TSE模型中:输出语言约束,中间语言预测和输入语言先验。在两个AV-TSE主干上使用RoBERTa、Qwen 3 -0.6B和Qwen 3 - 4 B进行的综合实验证明了我们方法的有效性。在具有挑战性的情况下,包括视觉线索受损,看不见的语言,目标扬声器开关,增加干扰扬声器,和域外测试集,观察到显着的改善。演示页面:https://alexwxwu.github.io/ELEGANCE/。
摘要:Audio-visual target speaker extraction (AV-TSE) models primarily rely on visual cues from the target speaker. However, humans also leverage linguistic knowledge, such as syntactic constraints, next word prediction, and prior knowledge of conversation, to extract target speech. Inspired by this observation, we propose ELEGANCE, a novel framework that incorporates linguistic knowledge from large language models (LLMs) into AV-TSE models through three distinct guidance strategies: output linguistic constraints, intermediate linguistic prediction, and input linguistic prior. Comprehensive experiments with RoBERTa, Qwen3-0.6B, and Qwen3-4B on two AV-TSE backbones demon- strate the effectiveness of our approach. Significant improvements are observed in challenging scenarios, including visual cue impaired, unseen languages, target speaker switches, increased interfering speakers, and out-of-domain test set. Demo page: https://alexwxwu.github.io/ELEGANCE/.
【13】We Can Hear You with mmWave Radar! An End-to-End Eavesdropping System
标题:我们可以通过毫米波雷达听到您的声音!端到端的发射器投放系统
链接:https://arxiv.org/abs/2511.06205
摘要:随着语音技术的兴起,扬声器播放变得越来越普遍,对语音隐私构成了越来越大的风险。传统的窃听方法通常需要侵入式访问或视线,限制了其实用性。在本文中,我们提出了mmSpeech,一个端到端的基于mmWave的窃听系统,它仅从扬声器播放引起的振动信号中重建可理解的语音,即使是通过墙壁,也没有扬声器的先验知识。为了实现这一目标,我们揭示了振动材料和雷达采样率的最佳组合,用于使用窄带毫米波信号捕获高质量的振动。然后,我们设计了一个深度神经网络,从估计的噪声频谱图中重建可理解的语音。为了进一步支持下游语音理解,我们引入了合成训练管道,并选择性地微调预训练ASR模型的编码器。我们实现了毫米波商用毫米波雷达的mmSpeech,并通过大量的实验验证其性能。结果表明,mmSpeech实现了最先进的语音质量,并且在看不见的说话者和各种条件下都能很好地推广。
摘要:With the rise of voice-enabled technologies, loudspeaker playback has become widespread, posing increasing risks to speech privacy. Traditional eavesdropping methods often require invasive access or line-of-sight, limiting their practicality. In this paper, we present mmSpeech, an end-to-end mmWave-based eavesdropping system that reconstructs intelligible speech solely from vibration signals induced by loudspeaker playback, even through walls and without prior knowledge of the speaker. To achieve this, we reveal an optimal combination of vibrating material and radar sampling rate for capturing high- quality vibrations using narrowband mmWave signals. We then design a deep neural network that reconstructs intelligible speech from the estimated noisy spectrograms. To further support downstream speech understanding, we introduce a synthetic training pipeline and selectively fine-tune the encoder of a pre-trained ASR model. We implement mmSpeech with a commercial mmWave radar and validate its performance through extensive experiments. Results show that mmSpeech achieves state-of-the-art speech quality and generalizes well across unseen speakers and various conditions.
【14】Who Gets Heard? Rethinking Fairness in AI for Music Systems
标题:谁会被听到?重新思考音乐系统人工智能的公平性
链接:https://arxiv.org/abs/2511.05953
备注:7 pages, Accepted at NeurIPS'25 workshop on AI for Music
摘要:近年来,音乐研究界已经研究了音乐AI模型的风险,特别是生成AI模型,引起了对版权,深度伪造和透明度的担忧。在我们的工作中,我们对音乐AI系统(音乐AI系统)中的文化和流派偏见提出了担忧,这些偏见影响了包括创作者、发行商和听众在内的利益相关者,从而影响了音乐AI的表现。这些偏见可能会歪曲边缘化的传统,特别是来自全球南方的传统,产生不真实的输出(例如,扭曲的ragas),这降低了创作者对这些系统的信任。这些危害有可能强化偏见,限制创造力,并导致文化抹除。为了解决这个问题,我们在音乐人工智能系统的数据集,模型和接口级别提供建议。
摘要:In recent years, the music research community has examined risks of AI models for music, with generative AI models in particular, raised concerns about copyright, deepfakes, and transparency. In our work, we raise concerns about cultural and genre biases in AI for music systems (music-AI systems) which affect stakeholders including creators, distributors, and listeners shaping representation in AI for music. These biases can misrepresent marginalized traditions, especially from the Global South, producing inauthentic outputs (e.g., distorted ragas) that reduces creators' trust on these systems. Such harms risk reinforcing biases, limiting creativity, and contributing to cultural erasure. To address this, we offer recommendations at dataset, model and interface level in music-AI systems.
【15】Loud-loss: A Perceptually Motivated Loss Function for Speech Enhancement Based on Equal-Loudness Contours
标题:响度损失:一种基于音量-响度轮廓的语音增强的感知动机损失函数
链接:https://arxiv.org/abs/2511.05945
摘要:均方误差(MSE)是语音增强中普遍存在的损失函数,但其问题是误差不能反映听觉感知质量。这是因为MSE导致模型过度强调具有高能量的低频分量,导致感知上重要的高频信息的建模不足。为了克服这个限制,我们提出了一个基于心理声学原理的感知加权损失函数。具体地说,它利用等响轮廓为重建误差分配频率相关的权重,从而以与人类听觉灵敏度一致的方式惩罚偏差。建议的损失是模型无关的和灵活的,表现出很强的通用性。在VoiceBank+DEMAND数据集上的实验表明,在GTCRN模型中用我们的损失替换MSE将WB-PESQ分数从2.17提高到2.93-感知质量的显着改善。
摘要:The mean squared error (MSE) is a ubiquitous loss function for speech enhancement, but its problem is that the error cannot reflect the auditory perception quality. This is because MSE causes models to over-emphasize low-frequency components which has high energy, leading to the inadequate modeling of perceptually important high-frequency information. To overcome this limitation, we propose a perceptually-weighted loss function grounded in psychoacoustic principles. Specifically, it leverages equal-loudness contours to assign frequency-dependent weights to the reconstruction error, thereby penalizing deviations in a way aligning with human auditory sensitivity. The proposed loss is model-agnostic and flexible, demonstrating strong generality. Experiments on the VoiceBank+DEMAND dataset show that replacing MSE with our loss in a GTCRN model elevates the WB-PESQ score from 2.17 to 2.93-a significant improvement in perceptual quality.
【16】TalkSketch: Multimodal Generative AI for Real-time Sketch Ideation with Speech
标题:TalkSketch:多模式生成人工智能,用于通过语音进行实时草图构思
链接:https://arxiv.org/abs/2511.05817
备注:Accepted at AAAI 2026 Workshop on Creative AI for Live Interactive Performances (CLIP). To be published in Springer CCIS series
摘要:草图是一种广泛使用的媒介,用于生成和探索早期设计概念。虽然生成式AI(GenAI)聊天机器人越来越多地用于创意生成,但设计师往往难以制作有效的提示,并且很难仅通过文本来表达不断发展的视觉概念。在形成性研究(N=6)中,我们研究了设计师在构思过程中如何使用GenAI,揭示了基于文本的提示会扰乱创意流程。为了解决这些问题,我们开发了TalkSketch,这是一个嵌入式多模态AI草图系统,它集成了手绘和实时语音输入。TalkSketch旨在通过在草图绘制过程中捕获口头描述并生成上下文感知的AI响应来支持更流畅的构思过程。我们的工作突出了GenAI工具参与设计过程本身而不是专注于输出的潜力。
摘要:Sketching is a widely used medium for generating and exploring early-stage design concepts. While generative AI (GenAI) chatbots are increasingly used for idea generation, designers often struggle to craft effective prompts and find it difficult to express evolving visual concepts through text alone. In the formative study (N=6), we examined how designers use GenAI during ideation, revealing that text-based prompting disrupts creative flow. To address these issues, we developed TalkSketch, an embedded multimodal AI sketching system that integrates freehand drawing with real-time speech input. TalkSketch aims to support a more fluid ideation process through capturing verbal descriptions during sketching and generating context-aware AI responses. Our work highlights the potential of GenAI tools to engage the design process itself rather than focusing on output.
【17】Persian Musical Instruments Classification Using Polyphonic Data Augmentation
标题:使用复音数据增强的波斯乐器分类
链接:https://arxiv.org/abs/2511.05717
备注:9 pages, 2 figures, 4 tables
摘要:乐器分类是音乐信息检索和音乐生成系统的基础。然而,对非西方传统的研究,特别是波斯音乐,仍然有限。我们通过引入一个新的孤立记录数据集来解决这一差距,该数据集涵盖了七种传统的波斯乐器,两种常见但最初非波斯乐器(即,小提琴,钢琴)和声乐。我们提出了一个文化上知情的数据增强策略,从单声道样本产生现实的复调混合物。使用带有分类头的MERT模型(带有大规模自监督训练的音乐理解模型),我们使用通过手动标记传统歌曲片段获得的分布外数据来评估我们的方法。在现实世界的复调波斯音乐,所提出的方法产生了最好的ROC AUC(0.795),突出了音调和时间连贯性的互补优势。这些结果证明了基于文化的增强对波斯乐器识别的有效性,并为文化包容性的MIR和多样化的音乐生成系统提供了基础。
摘要:Musical instrument classification is essential for music information retrieval (MIR) and generative music systems. However, research on non-Western traditions, particularly Persian music, remains limited. We address this gap by introducing a new dataset of isolated recordings covering seven traditional Persian instruments, two common but originally non-Persian instruments (i.e., violin, piano), and vocals. We propose a culturally informed data augmentation strategy that generates realistic polyphonic mixtures from monophonic samples. Using the MERT model (Music undERstanding with large-scale self-supervised Training) with a classification head, we evaluate our approach with out-of-distribution data which was obtained by manually labeling segments of traditional songs. On real-world polyphonic Persian music, the proposed method yielded the best ROC-AUC (0.795), highlighting complementary benefits of tonal and temporal coherence. These results demonstrate the effectiveness of culturally grounded augmentation for robust Persian instrument recognition and provide a foundation for culturally inclusive MIR and diverse music generation systems.
【18】Factual and Musical Evaluation Metrics for Music Language Models
标题:音乐语言模型的事实性与音乐性评价
链接:https://arxiv.org/abs/2511.05550
备注:18 pages; first submission
摘要:音乐语言模型(Music LM)与视觉语言模型一样,利用多模态表示来回答关于音乐音频记录的自然语言查询。虽然据报道,音乐LM正在改进,但我们发现目前的评估无法捕捉他们的答案是否正确。具体来说,对于我们研究的所有音乐LM,广泛使用的评估指标,如BLEU,METEOR和BERTScore,除了模型响应的语言流畅性之外,无法衡量任何东西。为了衡量音乐LM的真实性能,我们提出了(1)一个更好的通用评估指标,适用于音乐领域的音乐LM和(2)一个事实的评估框架,以量化音乐LM的响应的正确性。我们的框架是不可知的问答模型的模态,可以推广到量化性能在其他开放式问答域。我们在实验中使用开放数据集,并将在发布时发布所有代码。
摘要:Music language models (Music LMs), like vision language models, leverage multimodal representations to answer natural language queries about musical audio recordings. Although Music LMs are reportedly improving, we find that current evaluations fail to capture whether their answers are correct. Specifically, for all Music LMs that we examine, widely-used evaluation metrics such as BLEU, METEOR, and BERTScore fail to measure anything beyond linguistic fluency of the model's responses. To measure the true performance of Music LMs, we propose (1) a better general-purpose evaluation metric for Music LMs adapted to the music domain and (2) a factual evaluation framework to quantify the correctness of a Music LM's responses. Our framework is agnostic to the modality of the question-answering model and could be generalized to quantify performance in other open-ended question-answering domains. We use open datasets in our experiments and will release all code on publication.
【19】Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
标题:Ming-UniAudio:通过统一表示进行联合理解、生成和编辑的语音LLM
链接:https://arxiv.org/abs/2511.05516
备注:32 pages, 8 figures
摘要:现有的语音模型受到理解和生成任务对令牌表示的竞争要求的影响。这种表示上的差异阻止了语音语言模型执行基于简化的自由形式编辑。为了解决这一挑战,我们引入了一个新的框架,统一的语音理解,生成和编辑。我们的统一模型的核心是一个统一的连续语音标记器MingTok-Audio,第一个连续标记器,有效地集成语义和声学特征,这使得它适合理解和生成任务。在此基础上,我们开发了语音语言模型Ming-UniAudio,实现了生成和理解能力之间的平衡。Ming-UniAudio在ContextASR基准测试的12项指标中,有8项创下了最新的SOTA记录。值得注意的是,对于中文语音克隆,它实现了0.95的高度竞争性的Seed-TTS-WER。利用这个基础模型,我们进一步训练了一个专用的语音编辑模型Ming-UniAudio-Edit,这是第一个语音语言模型,可以实现仅由自然语言指令指导的通用、自由形式的语音编辑,在没有时间戳条件的情况下处理语义和声学修改。为了严格评估编辑能力,并为未来的研究奠定基础,我们引入了Ming-Freeform-Audio-Edit,这是第一个为基于发音的自由形式语音编辑量身定制的综合基准,具有多种场景和评估维度,涵盖语义正确性,声学质量和指令对齐。我们开源了连续音频标记器、统一基础模型和基于自由格式的编辑模型,以促进统一音频理解、生成和操作的开发。
摘要:Existing speech models suffer from competing requirements on token representations by understanding and generation tasks. This discrepancy in representation prevents speech language models from performing instruction-based free-form editing. To solve this challenge, we introduce a novel framework that unifies speech understanding, generation, and editing. The core of our unified model is a unified continuous speech tokenizer MingTok-Audio, the first continuous tokenizer to effectively integrate semantic and acoustic features, which makes it suitable for both understanding and generation tasks. Based on this unified continuous audio tokenizer, we developed the speech language model Ming-UniAudio, which achieved a balance between generation and understanding capabilities. Ming-UniAudio sets new state-of-the-art (SOTA) records on 8 out of 12 metrics on the ContextASR benchmark. Notably, for Chinese voice cloning, it achieves a highly competitive Seed-TTS-WER of 0.95. Leveraging this foundational model, we further trained a dedicated speech editing model Ming-UniAudio-Edit, the first speech language model that enables universal, free-form speech editing guided solely by natural language instructions, handling both semantic and acoustic modifications without timestamp condition. To rigorously assess the editing capability and establish a foundation for future research, we introduce Ming-Freeform-Audio-Edit, the first comprehensive benchmark tailored for instruction-based free-form speech editing, featuring diverse scenarios and evaluation dimensions spanning semantic correctness, acoustic quality, and instruction alignment. We open-sourced the continuous audio tokenizer, the unified foundational model, and the free-form instruction-based editing model to facilitate the development of unified audio understanding, generation, and manipulation.
【20】Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
标题:Omni-AVSR:采用大型语言模型实现统一多模式语音识别
链接:https://arxiv.org/abs/2511.07253
备注:Project website: this https URL
摘要:大型语言模型(LLM)最近在多种模态的语音识别中取得了令人印象深刻的结果,包括听觉语音识别(ASR),视觉语音识别(VSR)和视听语音识别(AVSR)。尽管取得了这一进展,但目前基于LLM的方法通常独立地处理每个任务,训练单独的模型,提高计算和部署资源的使用,同时失去潜在的跨任务协同作用。它们还依赖于固定速率的令牌压缩,这限制了平衡准确性和效率的灵活性。这些限制突出了对统一框架的需求,该框架可以支持ASR,VSR和AVSR,同时启用弹性推理。为此,我们提出了Omni-AVSR,这是一个统一的视听LLM,它将高效的多粒度训练与参数高效的自适应相结合。具体来说,我们适应matryoshka表示学习范式,有效地训练在多个音频和视觉粒度,减少其固有的训练资源的使用。此外,我们探索了三种基于LoRA的策略,以适应骨干LLM,平衡共享和特定任务的专业化。在LRS 2和LRS 3上的实验表明,Omni-AVSR实现了与最先进的基线相当或更高的精度,同时以更低的训练和部署资源使用来训练单个模型。该模型在声学噪声下也保持鲁棒性,并且我们分析了其缩放行为,因为LLM大小增加,提供了对性能和效率之间权衡的见解。
摘要:Large language models (LLMs) have recently achieved impressive results in speech recognition across multiple modalities, including Auditory Speech Recognition (ASR), Visual Speech Recognition (VSR), and Audio-Visual Speech Recognition (AVSR). Despite this progress, current LLM-based approaches typically address each task independently, training separate models that raise computational and deployment resource use while missing potential cross-task synergies. They also rely on fixed-rate token compression, which restricts flexibility in balancing accuracy with efficiency. These limitations highlight the need for a unified framework that can support ASR, VSR, and AVSR while enabling elastic inference. To this end, we present Omni-AVSR, a unified audio-visual LLM that combines efficient multi-granularity training with parameter-efficient adaptation. Specifically, we adapt the matryoshka representation learning paradigm to efficiently train across multiple audio and visual granularities, reducing its inherent training resource use. Furthermore, we explore three LoRA-based strategies for adapting the backbone LLM, balancing shared and task-specific specialization. Experiments on LRS2 and LRS3 show that Omni-AVSR achieves comparable or superior accuracy to state-of-the-art baselines while training a single model at substantially lower training and deployment resource use. The model also remains robust under acoustic noise, and we analyze its scaling behavior as LLM size increases, providing insights into the trade-off between performance and efficiency.
【1】Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
标题:Omni-AVSR:采用大型语言模型实现统一多模式语音识别
链接:https://arxiv.org/abs/2511.07253
备注:Project website: this https URL
摘要:大型语言模型(LLM)最近在多种模态的语音识别中取得了令人印象深刻的结果,包括听觉语音识别(ASR),视觉语音识别(VSR)和视听语音识别(AVSR)。尽管取得了这一进展,但目前基于LLM的方法通常独立地处理每个任务,训练单独的模型,提高计算和部署资源的使用,同时失去潜在的跨任务协同作用。它们还依赖于固定速率的令牌压缩,这限制了平衡准确性和效率的灵活性。这些限制突出了对统一框架的需求,该框架可以支持ASR,VSR和AVSR,同时启用弹性推理。为此,我们提出了Omni-AVSR,这是一个统一的视听LLM,它将高效的多粒度训练与参数高效的自适应相结合。具体来说,我们适应matryoshka表示学习范式,有效地训练在多个音频和视觉粒度,减少其固有的训练资源的使用。此外,我们探索了三种基于LoRA的策略,以适应骨干LLM,平衡共享和特定任务的专业化。在LRS2和LRS3上的实验表明,Omni-AVSR实现了与最先进的基线相当或更高的精度,同时以更低的训练和部署资源使用来训练单个模型。该模型在声学噪声下也保持鲁棒性,并且我们分析了其缩放行为,因为LLM大小增加,提供了对性能和效率之间权衡的见解。
【2】Neural Directional Filtering Using a Compact Microphone Array
标题:使用紧凑型麦克风阵列的神经定向过滤
链接:https://arxiv.org/abs/2511.07185
摘要:在许多音频应用中,使用紧凑型麦克风阵列的具有期望方向性图案的波束形成是必不可少的。使用传统波束形成器可实现的方向图取决于麦克风的数量和阵列孔径。一般来说,它们的有效性降低紧凑的阵列。为了克服这些限制,我们提出了一种神经方向滤波(NDF)方法,该方法利用深度神经网络来实现具有预定义方向性模式的声音捕获。NDF根据麦克风阵列信号计算单通道复掩模,然后将其应用于参考麦克风以产生近似具有所需方向性图案的虚拟定向麦克风的输出。我们介绍了训练策略,并提出了依赖于数据的指标来评估方向性模式和方向性因子。我们表明,所提出的方法:i)实现了频率不变的方向性模式,甚至高于空间混叠频率,ii)可以近似不同的和高阶的模式,iii)可以在不同的方向上引导模式,和iv)推广到看不见的条件。最后,实验比较表明优于传统的波束形成和参数化方法的性能。
【3】SPUR: A Plug-and-Play Framework for Integrating Spatial Audio Understanding and Reasoning into Large Audio-Language Models
标题:SPPUR:一个即插即用框架,用于将空间音频理解和推理集成到大型音频语言模型中
链接:https://arxiv.org/abs/2511.06606
备注:Project: this https URL
摘要:空间感知是听觉智能的核心,能够准确理解真实世界的声学场景,并提高人类对周围世界的感知。虽然最近的大型音频语言模型(LALM)在复杂的音频上表现出很强的推理能力,但大多数模型都是在单声道输入上操作的,并且缺乏捕获空间线索(如方向、海拔和距离)的能力。我们介绍SPUR,一个轻量级的,插件的方法,装备LALM的空间感知,通过最小的架构变化。SPUR包括:(i)一阶高保真度立体声(FOA)编码器,将(W、X、Y、Z)通道映射到旋转感知、以听众为中心的空间特征,并通过多模式适配器集成到目标LAM中;和(ii)SPUR-Set,一个空间QA数据集,将开源FOA记录与受控模拟相结合,强调相对方向、海拔、距离和重叠,以进行监督空间推理。在SPUR-Set上对我们的模型进行微调,可以持续改善空间QA和多说话者属性,同时保持一般的音频理解。SPUR提供了一个简单的方法,将单声道LALM转换为空间感知模型。广泛的消融验证了我们的方法的有效性。
【4】IDMap: A Pseudo-Speaker Generator Framework Based on Speaker Identity Index to Vector Mapping
标题:IDMap:一种基于说话者身份索引到载体映射的伪说话者生成器框架
链接:https://arxiv.org/abs/2511.06246
摘要:语音生成框架将语音分解为内容、说话人和韵律,通过将原始说话人嵌入向量替换为伪说话人嵌入向量来实现语音匿名化。在这个框架中,伪说话人生成形成了一个基本的挑战。现有的伪说话人生成方法在识别出的伪说话人的唯一性方面存在一定的局限性,从而限制了其在语音隐私保护方面的有效性。此外,现有的基于模型的方法遭受沉重的计算成本。特别是在产生大量伪说话人的大规模场景中,唯一性和计算效率的限制变得更加明显。为此,本文提出了一种伪说话人生成框架,该框架在前馈结构中建立了从说话人身份索引到说话人向量的映射,称为IDMap。具体来说,该框架被指定为两个模型:IDMap-MLP和IDMap-Diff。在小型和大型评估数据集上进行了实验。LibriSpeech数据集上的小规模评估验证了所提出的IDMap框架在增强伪说话者的唯一性方面的有效性,从而提高了语音隐私保护,同时降低了计算成本。对MLS和Common Voice数据集的大规模评估进一步证明了IDMap框架在语音隐私保护能力的稳定性方面的优越性,因为伪说话者的数量增加了。音频样本和开源代码可以在https://github.com/VoicePrivacy/IDMap上找到。
【5】BSCodec: A Band-Split Neural Codec for High-Quality Universal Audio Reconstruction
标题:BSCodec:用于高质量通用音频重建的带分裂神经编解码器
链接:https://arxiv.org/abs/2511.06150
摘要:神经音频编解码器最近已经能够以高压缩率进行高保真重建,特别是对于语音。然而,语音和非语音音频表现出根本不同的频谱特性:语音能量集中在音高谐波(80-400 Hz)周围的窄带中,而非语音音频需要在整个频谱上忠实再现,特别是保留定义音色和纹理的较高频率。这就带来了一个挑战:语音优化的神经编解码器在音乐或声音上会受到影响。整体地处理全频谱是次优的:频带具有根据内容类型的极大不同的信息密度和感知重要性,然而全频带方法在频率上应用统一的容量而不考虑这些声学结构。为了解决这一差距,我们提出了BSCodec(带分割编解码器),一种新的神经音频编解码器架构,将频谱维度分割成单独的频带,并独立压缩每个频带。实验结果表明,BSCodec在语音、音乐和声音的相同组合数据集上训练时,在语音域中保持有竞争力的质量的同时,在声音和音乐的基线上实现了更好的重建。下游基准测试任务进一步证实了BSCodec在下游应用中的强大潜力。
【6】Generating Piano Music with Transformers: A Comparative Study of Scale, Data, and Metrics
标题:用Transformer生成钢琴音乐:规模、数据和时间的比较研究
链接:https://arxiv.org/abs/2511.07268
备注:NeurIPS 2025 Workshop on AI for Music
摘要:虽然近年来已经提出了各种Transformers的符号音乐生成,仍然很少有全面的研究,具体的设计选择如何影响生成的音乐的质量。在这项工作中,我们系统地比较了不同的数据集,模型架构,模型大小和符号钢琴音乐生成任务的训练策略。为了支持模型的开发和评估,我们研究了一系列定量指标,并分析了它们与通过听力研究收集的人类判断的相关性。我们表现最好的模型是一个950 M参数的Transformer,它在来自不同流派的80K音频文件上训练,产生的输出在图灵风格的听力调查中通常被评为人类组成。
【7】Conditional Diffusion as Latent Constraints for Controllable Symbolic Music Generation
标题:条件扩散作为可控符号音乐生成的潜在约束
链接:https://arxiv.org/abs/2511.07156
备注:None
摘要:潜在扩散模型的最新进展已经证明了高维时间序列数据合成的最新性能,同时通过调节和指导提供灵活的控制。然而,现有的方法主要依赖于音乐上下文或自然语言作为与生成过程交互的主要模态,这对于寻求对特定音乐属性进行精确的类似推子的控制的专家用户来说可能不是理想的。在这项工作中,我们探讨了应用去噪扩散过程作为即插即用的无条件符号音乐生成模型的潜在约束。我们专注于一个框架,利用一个小的条件扩散模型库操作的隐式概率先验的潜伏期冻结无条件骨干。虽然以前的研究已经探索了特定领域的用例,但据我们所知,这项工作是第一次证明这种方法在各种音乐属性中的多功能性,例如音符密度,音高范围,轮廓和节奏复杂性。我们的实验表明,扩散驱动的约束优于传统的属性正则化和其他潜在的约束架构,实现目标和生成的属性之间的相关性显着更强,同时保持高的感知质量和多样性。
【8】Generating Novel and Realistic Speakers for Voice Conversion
标题:生成新颖且真实的扬声器以进行语音转换
链接:https://arxiv.org/abs/2511.07135
摘要:语音转换模型修改音色,同时保留非语言特征,支持配音和身份保护等应用。然而,大多数VC系统需要访问目标话语,从而在目标数据不可用或用户希望转换为全新、不可见的声音时限制了它们的使用。为了解决这个问题,我们引入了一个轻量级的方法SpeakerVAE来生成新的VC扬声器。我们的方法使用一个深层次的变分自动编码器来模拟扬声器音色空间。通过从训练好的模型中采样,我们在VC管道中生成了用于语音合成的新的说话人表示。所提出的方法是一个灵活的插件模块兼容各种VC模型,没有共同训练或微调的基础VC系统。我们使用最先进的VC模型评估了我们的方法:FACodec和CosyVoice2。结果表明,我们的方法成功地产生了新的,看不见的扬声器与质量相媲美的训练扬声器。
【9】On the Joint Minimization of Regularization Loss Functions in Deep Variational Bayesian Methods for Attribute-Controlled Symbolic Music Generation
标题:属性控制符号音乐生成的深度变分Bayesian方法中正规化损失函数的联合最小化
链接:https://arxiv.org/abs/2511.07118
备注:IEEE Catalog No.: CFP2540S-ART ISBN: 978-9-46-459362-4
摘要:显式潜变量模型为数据综合提供了一个灵活而强大的框架,使生成因子的控制操作成为可能。通过从可以进一步约束的易处理的概率密度函数中提取潜在变量,这些模型通过导航其潜在空间来实现对输出空间的连续和语义丰富的探索。结构化潜在表示通常通过正则化损失函数的联合最小化来获得。在变分信息瓶颈模型中,重构损失和Kullback-Leibler散度(KLD)通常与辅助属性正则化(AR)损失线性组合。然而,平衡KLD和AR是一个非常微妙的问题。当KLD优于AR时,生成模型往往缺乏可控性;当AR优于KLD时,随机编码器被鼓励违反标准的正态先验。我们探索这种权衡的背景下,符号音乐生成明确控制连续的音乐属性。我们表明,现有的方法努力共同最小化两个正则化目标,而合适的属性变换可以帮助实现目标潜在维度的可控性和正则化。
【10】MedVoiceBias: A Controlled Study of Audio LLM Behavior in Clinical Decision-Making
标题:MedVoiceBias:临床决策中音频LLM行为的对照研究
链接:https://arxiv.org/abs/2511.06592
摘要:随着大型语言模型从基于文本的界面过渡到临床环境中的音频交互,它们可能会通过音频中的非语言线索引入新的漏洞。我们在170个临床病例中评估了这些模型,每个病例都从36个不同的声音轮廓中合成语音,这些声音轮廓跨越了年龄,性别和情绪的变化。我们的研究结果揭示了严重的模态偏差:与相同的基于文本的输入相比,音频输入的手术建议变化高达35%,其中一个模型提供的建议减少了80%。进一步的分析发现,年轻人和老年人的声音之间存在高达12%的年龄差异,尽管有思想链的提示,但大多数模型仍然存在这种差异。虽然外显推理成功地消除了性别偏见,但由于识别性能差,没有检测到情感的影响。这些结果表明,音频LLM容易根据患者的声音特征而不是医学证据做出临床决策,这一缺陷有可能使医疗保健差异永久化。我们的结论是,偏见意识的架构是必不可少的,迫切需要这些模型的临床部署之前。
【11】EchoMark: Perceptual Acoustic Environment Transfer with Watermark-Embedded Room Impulse Response
标题:EchoMark:具有嵌入水印的房间脉冲响应的感知声学环境传输
链接:https://arxiv.org/abs/2511.06458
摘要:声学环境匹配(AEM)是将干净的音频传输到目标声学环境中的任务,从而实现音频配音和听觉沉浸式虚拟现实(VR)等引人入胜的应用。直接从混响语音中恢复相似的房间脉冲响应(RIR)提供了更容易和灵活的AEM解决方案。然而,如果被恶意用户滥用,这种能力也会引入任意“重新定位”的漏洞,例如促进高级语音欺骗攻击或破坏记录证据的真实性。为了解决这个问题,我们提出了EchoMark,这是第一个基于深度学习的AEM框架,可以生成具有嵌入水印的感知相似RIR。我们的设计通过在潜在域中操作来解决可变RIR特性(例如不同的持续时间和能量衰减)所带来的挑战。通过联合优化RIR重建的感知损失和水印检测损失的模型,EchoMark实现了高质量的环境传输和可靠的水印恢复。在不同数据集上的实验验证了EchoMark实现的室内声学参数匹配性能可与最先进的RIR估计器FiNS相媲美。此外,高平均意见分数(MOS)的4.22出5,水印检测准确率超过99%,误码率(BER)低于0.3%,共同证明了EchoMark在保持感知质量,同时确保可靠的水印嵌入的有效性。
【12】ELEGANCE: Efficient LLM Guidance for Audio-Visual Target Speech Extraction
标题:ELEGANCE:用于视听目标语音提取的高效LLM制导
链接:https://arxiv.org/abs/2511.06288
摘要:视听目标说话人提取(AV-TSE)模型主要依赖于来自目标说话人的视觉线索。然而,人类也利用语言学知识,如句法约束,下一个词的预测,和会话的先验知识,提取目标语音。受这一观察的启发,我们提出了ELEGANCE,这是一个新的框架,它通过三种不同的指导策略将来自大型语言模型(LLM)的语言知识整合到AV-TSE模型中:输出语言约束,中间语言预测和输入语言先验。在两个AV-TSE主干上使用RoBERTa、Qwen 3 -0.6B和Qwen 3 - 4 B进行的综合实验证明了我们方法的有效性。在具有挑战性的情况下,包括视觉线索受损,看不见的语言,目标扬声器开关,增加干扰扬声器,和域外测试集,观察到显着的改善。演示页面:https://alexwxwu.github.io/ELEGANCE/。
【13】Who Gets Heard? Rethinking Fairness in AI for Music Systems
标题:谁会被听到?重新思考音乐系统人工智能的公平性
链接:https://arxiv.org/abs/2511.05953
备注:7 pages, Accepted at NeurIPS'25 workshop on AI for Music
摘要:近年来,音乐研究界已经研究了音乐AI模型的风险,特别是生成AI模型,引起了对版权,深度伪造和透明度的担忧。在我们的工作中,我们对音乐AI系统(音乐AI系统)中的文化和流派偏见提出了担忧,这些偏见影响了包括创作者、发行商和听众在内的利益相关者,从而影响了音乐AI的表现。这些偏见可能会歪曲边缘化的传统,特别是来自全球南方的传统,产生不真实的输出(例如,扭曲的ragas),这降低了创作者对这些系统的信任。这些危害有可能强化偏见,限制创造力,并导致文化抹除。为了解决这个问题,我们在音乐人工智能系统的数据集,模型和接口级别提供建议。
【14】Ming-UniAudio: Speech LLM for Joint Understanding, Generation and Editing with Unified Representation
标题:Ming-UniAudio:通过统一表示进行联合理解、生成和编辑的语音LLM
链接:https://arxiv.org/abs/2511.05516
备注:32 pages, 8 figures
摘要:现有的语音模型受到理解和生成任务对令牌表示的竞争要求的影响。这种表示上的差异阻止了语音语言模型执行基于简化的自由形式编辑。为了解决这一挑战,我们引入了一个新的框架,统一的语音理解,生成和编辑。我们的统一模型的核心是一个统一的连续语音标记器MingTok-Audio,第一个连续标记器,有效地集成语义和声学特征,这使得它适合理解和生成任务。在此基础上,我们开发了语音语言模型Ming-UniAudio,实现了生成和理解能力之间的平衡。Ming-UniAudio在ContextASR基准测试的12项指标中,有8项创下了最新的SOTA记录。值得注意的是,对于中文语音克隆,它实现了0.95的高度竞争性的Seed-TTS-WER。利用这个基础模型,我们进一步训练了一个专用的语音编辑模型Ming-UniAudio-Edit,这是第一个语音语言模型,可以实现仅由自然语言指令指导的通用、自由形式的语音编辑,在没有时间戳条件的情况下处理语义和声学修改。为了严格评估编辑能力,并为未来的研究奠定基础,我们引入了Ming-Freeform-Audio-Edit,这是第一个为基于发音的自由形式语音编辑量身定制的综合基准,具有多种场景和评估维度,涵盖语义正确性,声学质量和指令对齐。我们开源了连续音频标记器、统一基础模型和基于自由格式的编辑模型,以促进统一音频理解、生成和操作的开发。
