今日论文合集:cs.SD语音9篇,eess.AS音频处理2篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
标题:NaijaS2ST:资源匮乏的尼日利亚语言语音到语音翻译的多口音基准
链接:https://arxiv.org/abs/2604.16287

作者:Marie Maltais,Yejin Jeon,Min Ma,Shamsuddeen Hassan Muhammad,Idris Abdulmumin,Maryam Ibrahim Mukhtar,Daud Abolade,Joel Okepefi,Johnson Sewedo,David Ifeoluwa Adelani
备注:Preprint
摘要:低资源语言的语音翻译仍然受到高质量,多样化并行语音数据稀缺的根本限制,这一挑战在非洲语言环境中尤为突出。为了解决这个问题,我们引入了NaijaS2ST,这是一个并行语音翻译数据集,涵盖了伊博语、豪萨语、约克夏语和尼日利亚洋泾浜语,并与英语配对。该数据集包括每种语言大约50小时的语音,并捕捉了说话者和口音的大量变化,反映了现实的多语言和多口音条件。通过NaijaS2ST,我们在双向翻译设置中对级联、端到端(E2E)和基于AudioLLM的方法进行了全面的基准测试。我们的研究结果表明,具有Few-Shot示例的音频LLM比在微调数据上训练的级联和端到端方法更有效地用于语音到文本翻译。然而,对于语音到语音翻译,级联和音频LLM范例产生相当的性能,这表明在为这种设置开发有针对性的特定任务模型方面仍有相当大的改进空间。通过提供高质量的数据集和系统的基准,我们希望NaijaS2ST将成为推进低资源多语言语音翻译研究的坚实基础。
摘要:Speech translation for low-resource languages remains fundamentally limited by the scarcity of high-quality, diverse parallel speech data, a challenge that is especially pronounced in African linguistic contexts. To address this, we introduce NaijaS2ST, a parallel speech translation dataset spanning Igbo, Hausa, Yorùbá, and Nigerian Pidgin paired with English. The dataset comprises approximately 50 hours of speech per language and captures substantial variation in speakers and accents, reflecting realistic multilingual and multi-accent conditions. With NaijaS2ST, we conduct a comprehensive benchmark of cascaded, end-to-end (E2E), and AudioLLM-based approaches across bidirectional translation settings. Our results show that audio LLMs with few-shot examples are more effective for speech-to-text translation than cascaded and end-to-end methods trained on fine-tuned data. However, for speech-to-speech translation, the cascaded and audio LLM paradigms yield comparable performance, indicating that there is still considerable room for improvement in developing targeted, task-specific models for this setting. By providing both a high-quality dataset and a systematic benchmark, we hope that NaijaS2ST will serve as a strong foundation for advancing research in low-resource, multilingual speech translation.


【2】ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics
标题:ArtifactNet:通过法医残留物理检测人工智能生成的音乐
链接:https://arxiv.org/abs/2604.16254

作者:Heewon Oh
备注:9 pages, 7 figures, 9 tables
摘要:我们提出了ArtifactNet,一个轻量级的框架,通过将问题重新定义为法医物理学来检测AI生成的音乐-提取和分析神经音频编解码器不可避免地在生成的音频上留下的物理伪影。有界掩码UNet(ArtifactUNet,3.6M参数)从幅度谱图中提取编解码器残差,然后通过HPSS将其分解为7通道取证特征,用于通过紧凑的CNN(0.4M参数;总共4.0M)进行分类。我们引入了ArtifactBench,这是一个多生成器评估基准,包括6,183个跟踪(来自22个生成器的4,383个AI和来自6个不同来源的1,800个Real)。每个轨道都标记有基准原点,以进行公平的zero-shot评估。在看不见的测试分区(n= 2,263)上,ArtifactNet实现了F1 = 0.9829,FPR = 1.49%,而CLAM(F1 = 0.7576,FPR = 69.26%)和SpecTTTra(F1 = 0.7713,FPR = 19.43%)在相同条件下使用已发布的检查点进行了评估。编解码器感知训练(4路WAV/MP3/AAC/Opus增强)进一步将跨编解码器概率漂移降低了83%(Delta = 0.95 -> 0.16),解决了主要的编解码器不变性故障模式。这些结果建立了法医物理学-直接提取编解码器级伪影-作为AI音乐检测的一种比表征学习更通用和参数效率更高的范例,使用的参数比CLAM少49倍,比SpecTTTra少4.8倍。
摘要:We present ArtifactNet, a lightweight framework that detects AI-generated music by reframing the problem as forensic physics -- extracting and analyzing the physical artifacts that neural audio codecs inevitably imprint on generated audio. A bounded-mask UNet (ArtifactUNet, 3.6M parameters) extracts codec residuals from magnitude spectrograms, which are then decomposed via HPSS into 7-channel forensic features for classification by a compact CNN (0.4M parameters; 4.0M total). We introduce ArtifactBench, a multi-generator evaluation benchmark comprising 6,183 tracks (4,383 AI from 22 generators and 1,800 real from 6 diverse sources). Each track is tagged with bench_origin for fair zero-shot evaluation. On the unseen test partition (n=2,263), ArtifactNet achieves F1 = 0.9829 with FPR = 1.49%, compared to CLAM (F1 = 0.7576, FPR = 69.26%) and SpecTTTra (F1 = 0.7713, FPR = 19.43%) evaluated under identical conditions with published checkpoints. Codec-aware training (4-way WAV/MP3/AAC/Opus augmentation) further reduces cross-codec probability drift by 83% (Delta = 0.95 -> 0.16), resolving the primary codec-invariance failure mode. These results establish forensic physics -- direct extraction of codec-level artifacts -- as a more generalizable and parameter-efficient paradigm for AI music detection than representation learning, using 49x fewer parameters than CLAM and 4.8x fewer than SpecTTTra.


【3】NVBench: A Benchmark for Speech Synthesis with Non-Verbal Vocalizations
标题:NVBench:非言语发声语音合成的基准
链接:https://arxiv.org/abs/2604.16211

作者:Liumeng Xue,Weizhen Bian,Jiahao Pan,Wenxuan Wang,Yilin Ren,Boyi Kang,Jingbin Hu,Ziyang Ma,Shuai Wang,Xinyuan Qian,Hung-yi Lee,Yike Guo
摘要:像笑、叹息和哭泣这样的非语言发声(NVV)对于人类语音来说是必不可少的,但是标准化的评估仍然局限于联合评估系统是否可以生成预期的NVV,正确地放置它们,并在不损害语音的情况下保持它们的突出性。我们提出了非语言发声基准(NVBench),一个双语(英语/中文)基准,评估语音合成与NVV。NVBench将统一的45种类型分类与策划的双语数据集配对,并引入了一个多轴协议,将一般语音自然度和质量与NVV特定的可控性,放置和显著性分开。我们基准测试15 TTS系统使用客观指标,听力测试,并基于LLM的多评价。结果表明,NVV的可控性往往从质量,而低信噪比的口头线索和长时间的情感NVV仍然是持久的瓶颈。NVBench能够在统一的标准化框架下对不同的控制接口进行公平的跨系统比较。
摘要:Non-verbal vocalizations (NVVs) like laugh, sigh, and sob are essential for human-like speech, yet standardized evaluation remains limited in jointly assessing whether systems can generate the intended NVVs, place them correctly, and keep them salient without harming speech. We present Non-verbal Vocalization Benchmark (NVBench), a bilingual (English/Chinese) benchmark that evaluates speech synthesis with NVVs. NVBench pairs a unified 45-type taxonomy with a curated bilingual dataset and introduces a multi-axis protocol that separates general speech naturalness and quality from NVV-specific controllability, placement, and salience. We benchmark 15 TTS systems using objective metrics, listening tests, and an LLM-based multi-rater evaluation. Results reveal that NVVs controllability often decouples from quality, while low-SNR oral cues and long-duration affective NVVs remain persistent bottlenecks. NVBench enables fair cross-system comparison across diverse control interfaces under a unified, standardized framework.


【4】AST: Adaptive, Seamless, and Training-Free Precise Speech Editing
标题:AST:自适应、无缝、免训练的精确语音编辑
链接:https://arxiv.org/abs/2604.16056

作者:Sihan Lv,Yechen Jin,Zhen Li,Jintao Chen,Jinshan Zhang,Ying Li,Jianwei Yin,Meng Xi
摘要:基于文本的语音编辑旨在修改特定片段,同时保留说话者身份和声学背景。现有的方法依赖于特定于任务的训练,这会产生很高的数据成本,并且在未经编辑的区域中难以实现时间保真度。与此同时,调整文本到语音(TTS)模型通常面临编辑质量和一致性之间的权衡。为了解决这些问题,我们提出了AST,一个自适应的,无缝的,和训练免费的精确语音编辑框架。利用预训练的自回归TTS模型,AST引入了潜在重组,以选择性地将保留的源片段与新合成的目标片段缝合在一起。此外,AST扩展了这种潜在的操作,以实现特定语音片段的精确风格编辑。为了防止在这些编辑边界出现伪影,该框架采用了自适应弱事实指导(AWFG)。AWFG动态地调制一个梅尔空间引导信号,只在必要时执行结构约束,而不破坏生成流形。为了填补公开访问基准的空白,我们引入了LibriSpeech-Edit,一个新的更大的语音编辑数据集。由于现有的指标评价时间一致性差,在未经编辑的区域,我们提出了字级动态时间规整(WDTW)。大量的实验表明,AST解决了不需要额外训练的可扩展性质量权衡。与之前时间上最一致的基线相比,AST提高了一致性,同时将单词错误率降低了近70%。此外,将AST应用于基础TTS模型将WDTW减少了27%,实现了最先进的说话人保留和时间保真度。
摘要:Text-based speech editing aims to modify specific segments while preserving speaker identity and acoustic context. Existing methods rely on task-specific training, which incurs high data costs and struggles with temporal fidelity in unedited regions. Meanwhile, adapting Text-to-Speech (TTS) models often faces a trade-off between editing quality and consistency. To address these issues, we propose AST, an Adaptive, Seamless, and Training-free precise speech editing framework. Leveraging a pre-trained autoregressive TTS model, AST introduces Latent Recomposition to selectively stitch preserved source segments with newly synthesized targets. Furthermore, AST extends this latent manipulation to enable precise style editing for specific speech segments. To prevent artifacts at these edit boundaries, the framework incorporates Adaptive Weak Fact Guidance (AWFG). AWFG dynamically modulates a mel-space guidance signal, enforcing structural constraints only where necessary without disrupting the generative manifold. To fill the gap of publicly accessible benchmarks, we introduce LibriSpeech-Edit, a new and larger speech editing dataset. As existing metrics poorly evaluate temporal consistency in unedited regions, we propose Word-level Dynamic Time Warping (WDTW). Extensive experiments demonstrate that AST resolves the controllability-quality trade-off without extra training. Compared to the previous most temporally consistent baseline, AST improves consistency while reducing Word Error Rate by nearly 70%. Moreover, applying AST to a foundation TTS model reduces WDTW by 27%, achieving state-of-the-art speaker preservation and temporal fidelity.


【5】Breakout-picker: Reducing false positives in deep learning-based borehole breakout characterization from acoustic image logs
标题:突破点选择器:减少来自声学图像日志的基于深度学习的井眼突破特征中的误报
链接:https://arxiv.org/abs/2604.16011

作者:Guangyu Wang,Xiaodong Ma,Xinming Wu
摘要:钻孔崩落是钻孔壁上的应力诱导剥落,其在声学成像测井中可识别为具有近对称方位角、低声学振幅和增加的钻孔半径的成对区域。准确的崩落特征对于地应力分析至关重要。近年来,深度学习已经被引入到自动化耗时和劳动密集型的分拣过程中。然而,现有的方法往往遭受误分类的非突破功能,导致高误报率。为了解决这一限制,本研究开发了一个深度学习框架,称为Breakout-picker,特别关注减少自动突破特征中的误报。Breakout-picker通过两种策略减少误报。首先,突破选择器的训练包含非突破特征的负样本,包括天然裂缝、键槽和测井伪影。它们具有与崩落相似的特征,例如低声幅或局部扩大的井眼半径。这些负训练样本使突破选择器能够更好地区分真正的突破和类似的非突破特征。第二,通过方位角对称性标准进一步验证由突破拾取器识别的候选突破,从而排除不表现出突破方位角的近对称特性的检测。使用来自不同地区的三个声波图像测井数据集来评估突破拾取器的性能。结果表明,突破采摘优于其他自动方法,具有更高的准确性和显着降低误报率。通过减少误报,Breakout-picker增强了从声学图像测井中自动描述崩落特征的可靠性,这反过来又有利于基于钻孔崩落的原地应力分析。
摘要:Borehole breakouts are stress-induced spalling on the borehole wall, which are identifiable in acoustic image logs as paired zones with near-symmetry azimuths, low acoustic amplitudes, and increased borehole radius. Accurate breakout characterization is crucial for in-situ stress analysis. In recent years, deep learning has been introduced to automate the time-consuming and labor-intensive breakout picking process. However, existing approaches often suffer from misclassification of non-breakout features, leading to high false positive rates. To address this limitation, this study develops a deep learning framework, termed Breakout-picker, with a specific focus on reducing false positives in automatic breakout characterization. Breakout-picker reduces false positives through two strategies. First, the training of Breakout-picker incorporates negative samples of non-breakout features, including natural fractures, keyseats, and logging artifacts. They share similar characteristics with breakouts, such as low acoustic amplitude or locally enlarged borehole radius. These negative training samples enables Breakout-picker to better discriminate true breakouts and similar non-breakout features. Second, candidate breakouts identified by Breakout-picker are further validated by azimuthal symmetry criteria, whereby detections that do not exhibit the near-symmetry characteristics of breakout azimuth are excluded. The performance of Breakout-picker is evaluated using three acoustic image log datasets from different regions. The results demonstrate that Breakout-picker outperforms other automatic methods with higher accuracy and substantially lower false positive rates. By reducing false positives, Breakout-picker enhances the reliability of automatic breakout characterization from acoustic image logs, which in turn benefits in-situ stress analysis based on borehole breakouts.


【6】Hierarchical Codec Diffusion for Video-to-Speech Generation
标题:用于视频到语音生成的分层编解码器扩散
链接:https://arxiv.org/abs/2604.15923

作者:Jiaxin Ye,Gaoxiang Cong,Chenhui Wang,Xin-Cheng Wen,Zhaoyang Li,Boyuan Cao,Hongming Shan
备注:CVPR 2026
摘要:视频到语音(VTS)生成的目的是从一个无声的视频合成语音没有听觉信号。然而,现有的VTS方法忽略了语音的层次性,其跨越粗说话者感知语义到细粒度韵律细节。这种疏忽阻碍了在属性匹配期间在特定层次级别上视觉和语音特征之间的直接对齐。在本文中,利用层次结构的残差矢量量化(RVQ)为基础的编解码器,我们提出HiCoDiT,一种新的分层编解码器扩散Transformer,利用固有的层次结构的离散语音令牌,以实现强大的视听对齐。具体而言,由于较低级别的令牌编码粗糙的说话者感知语义和较高级别的令牌捕获细粒度的韵律,HiCoDiT采用低级别和高级别的块来生成不同级别的令牌。低级别块的条件唇同步运动和面部身份捕捉说话人意识的内容,而高级别块使用面部表情来调制韵律动态。最后,为了实现更有效的从粗到细的调节,我们提出了一种双尺度自适应实例层归一化,通过通道归一化和局部韵律动态通过时间归一化联合捕获全局声乐风格。大量实验表明,HiCoDiT在保真度和表现力方面优于基线,突出了VTS离散建模的潜力。代码和语音演示都可以在https://github.com/Jiaxin-Ye/HiCoDiT上找到。
摘要:Video-to-Speech (VTS) generation aims to synthesize speech from a silent video without auditory signals. However, existing VTS methods disregard the hierarchical nature of speech, which spans coarse speaker-aware semantics to fine-grained prosodic details. This oversight hinders direct alignment between visual and speech features at specific hierarchical levels during property matching. In this paper, leveraging the hierarchical structure of Residual Vector Quantization (RVQ)-based codec, we propose HiCoDiT, a novel Hierarchical Codec Diffusion Transformer that exploits the inherent hierarchy of discrete speech tokens to achieve strong audio-visual alignment. Specifically, since lower-level tokens encode coarse speaker-aware semantics and higher-level tokens capture fine-grained prosody, HiCoDiT employs low-level and high-level blocks to generate tokens at different levels. The low-level blocks condition on lip-synchronized motion and facial identity to capture speaker-aware content, while the high-level blocks use facial expression to modulate prosodic dynamics. Finally, to enable more effective coarse-to-fine conditioning, we propose a dual-scale adaptive instance layer normalization that jointly captures global vocal style through channel-wise normalization and local prosody dynamics through temporal-wise normalization. Extensive experiments demonstrate that HiCoDiT outperforms baselines in fidelity and expressiveness, highlighting the potential of discrete modelling for VTS. The code and speech demo are both available at https://github.com/Jiaxin-Ye/HiCoDiT.


【7】TinyMU: A Compact Audio-Language Model for Music Understanding
标题:TinyMU:音乐理解的紧凑音频语言模型
链接:https://arxiv.org/abs/2604.15849

作者:Xiquan Li,Aurian Quelennec,Slim Essid
备注:ICASSP 2026
摘要:音乐理解和推理是音乐信息研究领域的核心挑战,其应用范围从检索和推荐到音乐代理和虚拟助理。最近的大型音频语言模型(LALM)在回答与音乐相关的问题方面取得了显着的进展。然而,其庞大的规模,通常是数十亿个参数,导致昂贵的训练,缓慢的推理和边缘设备上的有限部署。在这项工作中,我们提出了TinyMU,一个轻量级的(229 M)音乐语言模型(MLM),实现性能媲美更大的LALM,同时保持高效和紧凑。为了训练TinyMU,我们引入了MusicSkills-3.5M,这是一个精心策划的,以音乐为基础的问答数据集,拥有350万个样本。跨越多项选择,二进制和开放式格式,该数据集提供了对不同音乐概念的细粒度监督。对于其架构,TinyMU利用MATPAC++,SOTA自监督音频编码器进行细粒度特征提取。搭配轻量级线性投影仪,它可以有效地将音频嵌入与语言模型对齐。通过广泛的评估,我们表明TinyMU在基本的音乐理解和复杂的推理方面都表现出色。值得注意的是,在MuChoMusic基准测试中,它实现了SOTA LALM性能的82%,尽管它比SOTA LALM小35倍,突出了小型MLMs在有限的计算预算下的潜力。
摘要:Music understanding and reasoning are central challenges in the Music Information Research field, with applications ranging from retrieval and recommendation to music agents and virtual assistants. Recent Large Audio-Language Models (LALMs) have shown remarkable progress in answering music-related questions by following user instructions. However, their massive scale, often billions of parameters, results in expensive training, slow inference, and limited deployability on edge devices. In this work, we present TinyMU, a lightweight (229M) Music-Language Model (MLM) that achieves performance comparable to much larger LALMs while remaining efficient and compact. To train TinyMU, we introduce MusicSkills-3.5M, a carefully curated, music-grounded question-answering dataset with 3.5M samples. Spanning multiple-choice, binary, and open-ended formats, this dataset provides fine-grained supervision across diverse musical concepts. For its architecture, TinyMU leverages MATPAC++, the SOTA self-supervised audio encoder for fine-grained feature extraction. Paired with a lightweight linear projector, it efficiently aligns audio embeddings with the language model. Through extensive evaluation, we show that TinyMU performs strongly in both basic music understanding and complex reasoning. Notably, on the MuChoMusic benchmark, it achieves 82\% of SOTA LALM's performance despite being 35x smaller, highlighting the potential of small MLMs under constrained computational budgets.


【8】VoxMind: An End-to-End Agentic Spoken Dialogue System
标题:VoxMind:一个端到端的远程口语对话系统
链接:https://arxiv.org/abs/2604.15710

作者:Tianle Liang,Yifu Chen,Shengpeng Ji,Yijun Chen,Zhiyang Jia,Jingyu Lu,Fan Zhuo,Xueyi Pu,Yangzhuo Li,Zhou Zhao
备注:Accepted to ACL 2026 Main Conference.Code and data available at https://github.com/MM-Speech/VoxMind
摘要:最近的端到端口语对话模型实现了自然交互。然而,随着用户需求变得越来越复杂,仅仅依赖于会话能力的模型往往难以应对。因此,增强代理能力至关重要:通过启用工具使用,这些模型可以扩展其知识边界并更好地解决现实世界的任务。然而,现有的研究主要集中在核心感知和生成,相对有限的探索,这种工具增强的扩展。为了弥合这一差距,我们提出了VoxMind,一个集成的框架,旨在为端到端的口语对话模型提供全面的代理能力。利用我们精心策划的470小时AgentChat数据集,我们引入了“先思考后发言”机制,使模型能够将结构化推理内化为规划和响应生成的关键先决条件。此外,为了缓解大规模工具集成所造成的延迟瓶颈,我们提出了一个多Agent动态工具管理体系结构。通过异步委托检索任务的辅助代理对齐的主要模型的推理轨迹,该系统有效地从工具集大小的推理延迟。实验结果证实,VoxMind实现了代理性能的显着改善:与强基线相比,任务完成率从34.88%提高到74.57%,在口语代理任务上优于Gemini-2.5-Pro,同时保持一般会话质量。源代码和相关数据可在https://github.com/MM-Speech/VoxMind上公开获取。
摘要:Recent end-to-end spoken dialogue models enable natural interaction. However, as user demands become increasingly complex, models that rely solely on conversational abilities often struggle to cope. Incorporating agentic capabilities is therefore essential: by enabling tool use, these models can extend their knowledge boundaries and better solve real-world tasks. Yet, existing research has largely concentrated on core perception and generation, with comparatively limited exploration of such tool-augmented extensions. To bridge this gap, we present VoxMind, an integrated framework designed to equip end-to-end spoken dialogue models with comprehensive agentic abilities. Leveraging our curated 470-hour AgentChat dataset, we incorporate a "Think-before-Speak" mechanism, enabling the model to internalize structured reasoning as a critical prerequisite for planning and response generation. Furthermore, to mitigate latency bottlenecks caused by large-scale tool integration, we propose a Multi-Agent Dynamic Tool Management architecture. By asynchronously delegating retrieval tasks to an auxiliary agent aligned with the main model's reasoning trajectory, this system effectively decouples inference latency from toolset size. Experimental results confirm that VoxMind achieves significant improvements in agent performance: compared with strong baselines, the task completion rate increases from 34.88% to 74.57%, outperforming Gemini-2.5-Pro on spoken agent tasks while preserving general conversational quality. The source code and associated data are publicly available at https://github.com/MM-Speech/VoxMind.


【9】Temporal Contrastive Decoding: A Training-Free Method for Large Audio-Language Models
标题:时间对比解码:大型音频语言模型的免训练方法
链接:https://arxiv.org/abs/2604.15383

作者:Yanda Li,Yuhan Liu,Zirui Song,Yunchao Wei,Martin Takáč,Salem Lahlou
备注:ACL 2026 Findings
摘要:大型音频语言模型(LALM)概括了语音,声音和音乐,但统一的解码器可以表现出时间平滑偏差:瞬态声学线索可能未被充分利用,有利于时间平滑的上下文,更好地支持语言先验,导致不太具体的音频接地输出。我们提出了时间对比解码(TCD),一种用于统一LALM的免训练解码方法,可以在推理时减轻这种影响。TCD通过平滑输入波形并重新编码来构建时间模糊的慢径视图,然后将下一个令牌logits与原始视图和慢径视图进行对比。对比信号被应用为限制于小候选集的令牌级logit更新。自归一化稳定性分数设置模糊窗口和更新尺度,并且基于不确定性和音频依赖的步进式门仅在需要时激活更新。在MMAU和AIR-Bench上的实验表明,强统一LALM得到了一致的改进。我们进一步进行消融和架构适用性研究,以分析关键组件的贡献以及TCD在大型音频语言模型设计中的表现。
摘要:Large audio-language models (LALMs) generalize across speech, sound, and music, but unified decoders can exhibit a \emph{temporal smoothing bias}: transient acoustic cues may be underutilized in favor of temporally smooth context that is better supported by language priors, leading to less specific audio-grounded outputs. We propose \emph{Temporal Contrastive Decoding} (TCD), a training-free decoding method for unified LALMs that mitigates this effect at inference time. TCD constructs a temporally blurred slow-path view by smoothing the input waveform and re-encoding it, then contrasts next-token logits from the original and slow-path views. The contrastive signal is applied as a token-level logit update restricted to a small candidate set. A self-normalized stability score sets the blur window and update scale, and a step-wise gate based on uncertainty and audio reliance activates the update only when needed. Experiments on MMAU and AIR-Bench show consistent improvements on strong unified LALMs. We further conduct ablations and an architectural applicability study to analyze the contributions of key components and how TCD behaves across large audio-language model designs.


eess.AS音频处理


【1】ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics
标题:ArtifactNet:通过法医残留物理检测人工智能生成的音乐
链接:https://arxiv.org/abs/2604.16254

作者:Heewon Oh
备注:9 pages, 7 figures, 9 tables
摘要:我们提出了ArtifactNet,一个轻量级的框架,通过将问题重新定义为法医物理学来检测AI生成的音乐-提取和分析神经音频编解码器不可避免地在生成的音频上留下的物理伪影。有界掩码UNet(ArtifactUNet,3.6M参数)从幅度谱图中提取编解码器残差,然后通过HPSS将其分解为7通道取证特征,用于通过紧凑的CNN(0.4M参数;总共4.0M)进行分类。我们引入了ArtifactBench,这是一个多生成器评估基准,包括6,183个跟踪(来自22个生成器的4,383个AI和来自6个不同来源的1,800个Real)。每个轨道都标记有基准原点,以进行公平的zero-shot评估。在看不见的测试分区(n= 2,263)上,ArtifactNet实现了F1 = 0.9829,FPR = 1.49%,而CLAM(F1 = 0.7576,FPR = 69.26%)和SpecTTTra(F1 = 0.7713,FPR = 19.43%)在相同条件下使用已发布的检查点进行了评估。编解码器感知训练(4路WAV/MP3/AAC/Opus增强)进一步将跨编解码器概率漂移降低了83%(Delta = 0.95 -> 0.16),解决了主要的编解码器不变性故障模式。这些结果建立了法医物理学-直接提取编解码器级伪影-作为AI音乐检测的一种比表征学习更通用和参数效率更高的范例,使用的参数比CLAM少49倍,比SpecTTTra少4.8倍。
摘要:We present ArtifactNet, a lightweight framework that detects AI-generated music by reframing the problem as forensic physics -- extracting and analyzing the physical artifacts that neural audio codecs inevitably imprint on generated audio. A bounded-mask UNet (ArtifactUNet, 3.6M parameters) extracts codec residuals from magnitude spectrograms, which are then decomposed via HPSS into 7-channel forensic features for classification by a compact CNN (0.4M parameters; 4.0M total). We introduce ArtifactBench, a multi-generator evaluation benchmark comprising 6,183 tracks (4,383 AI from 22 generators and 1,800 real from 6 diverse sources). Each track is tagged with bench_origin for fair zero-shot evaluation. On the unseen test partition (n=2,263), ArtifactNet achieves F1 = 0.9829 with FPR = 1.49%, compared to CLAM (F1 = 0.7576, FPR = 69.26%) and SpecTTTra (F1 = 0.7713, FPR = 19.43%) evaluated under identical conditions with published checkpoints. Codec-aware training (4-way WAV/MP3/AAC/Opus augmentation) further reduces cross-codec probability drift by 83% (Delta = 0.95 -> 0.16), resolving the primary codec-invariance failure mode. These results establish forensic physics -- direct extraction of codec-level artifacts -- as a more generalizable and parameter-efficient paradigm for AI music detection than representation learning, using 49x fewer parameters than CLAM and 4.8x fewer than SpecTTTra.


【2】Qwen3.5-Omni Technical Report
标题:Qwen 3.5-Omni技术报告
链接:https://arxiv.org/abs/2604.15804

作者:Qwen Team
摘要:在这项工作中,我们提出了Qwen3.5-Omni,在Qwen-Omni模型家族的最新进展。Qwen3.5-Omni代表了其前身的重大发展,可扩展到数千亿个参数,并支持256 k上下文长度。通过利用由异构文本-视觉对和超过1亿小时的视听内容组成的庞大数据集,该模型展示了强大的全模态功能。Qwen3.5-Omni-plus在215个音频和视听理解、推理和交互子任务和基准测试中实现了SOTA结果,在关键音频任务中超过了Gemini-3.1 Pro,在综合视听理解方面与之相当。在架构上,Qwen3.5-Omni为Thinker和Talker采用了混合注意力专家混合(MoE)框架,实现了高效的长序列推理。该模型促进了复杂的交互,支持超过10小时的音频理解和400秒的720 P视频(1 FPS)。为了解决流语音合成中固有的不稳定性和不自然性,通常由文本和语音标记器之间的编码效率差异引起,我们引入ARIA。ARIA动态对齐文本和语音单元,显著增强会话语音的稳定性和韵律,同时将延迟影响降至最低。此外,Qwen3.5-Omni扩展了语言边界,支持跨10种语言的多语言理解和语音生成,具有类似人类的情感细微差别。最后,Qwen3.5-Omni具有卓越的视听基础能力,可生成具有精确时间同步和自动场景分割的脚本级结构化字幕。值得注意的是,我们观察到全模态模型中出现了一种新的能力:直接根据视听指令进行编码,我们称之为视听氛围编码。
摘要:In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor, Qwen3.5-Omni scales to hundreds of billions of parameters and supports a 256k context length. By leveraging a massive dataset comprising heterogeneous text-vision pairs and over 100 million hours of audio-visual content, the model demonstrates robust omni-modality capabilities. Qwen3.5-Omni-plus achieves SOTA results across 215 audio and audio-visual understanding, reasoning, and interaction subtasks and benchmarks, surpassing Gemini-3.1 Pro in key audio tasks and matching it in comprehensive audio-visual understanding. Architecturally, Qwen3.5-Omni employs a Hybrid Attention Mixture-of-Experts (MoE) framework for both Thinker and Talker, enabling efficient long-sequence inference. The model facilitates sophisticated interaction, supporting over 10 hours of audio understanding and 400 seconds of 720P video (at 1 FPS). To address the inherent instability and unnaturalness in streaming speech synthesis, often caused by encoding efficiency discrepancies between text and speech tokenizers, we introduce ARIA. ARIA dynamically aligns text and speech units, significantly enhancing the stability and prosody of conversational speech with minimal latency impact. Furthermore, Qwen3.5-Omni expands linguistic boundaries, supporting multilingual understanding and speech generation across 10 languages with human-like emotional nuance. Finally, Qwen3.5-Omni exhibits superior audio-visual grounding capabilities, generating script-level structured captions with precise temporal synchronization and automated scene segmentation. Remarkably, we observed the emergence of a new capability in omnimodal models: directly performing coding based on audio-visual instructions, which we call Audio-Visual Vibe Coding.


机器翻译由腾讯交互翻译提供,仅供参考