今日论文合集:CS.SD语音与音频 | 共 8 篇


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音合成与声音生成 1 篇

2. 语音增强、降噪与音频修复 2 篇

3. 音乐信息检索与音乐生成 2 篇

4. 数据集、基准与评测 2 篇

5. 其他/综合语音音频 1 篇

1. 语音合成与声音生成 | 1 篇

1. ZONOS2 Technical Report

ZONOS2 技术报告

AI 总结:提出ZONOS2 8B TTS模型,通过MoE架构、6M小时训练数据和简化后训练配方,在自然度、韵律和声音克隆保真度上达到最先进水平。

链接:https://arxiv.org/abs/2606.24320

机构:Zyphra

作者:Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close, Beren Millidge

英文摘要:We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity. We improve upon Zonos-v0.1 across scale, data, and training recipe. We scale the model from 1.6B to 8B total parameters (900M active) with a novel mixture-of-experts (MoE) backbone, improving inference latency and throughput. We expand our training corpus from 200K to over 6M hours using a new data processing pipeline, and we simplify our post-training and conditioning recipes to improve naturalness and voice cloning fidelity. We evaluate ZONOS2 8B on quality, speaker similarity, WER, and ZTTS1-Eval, our novel TTS benchmark, where it performs competitively with state-of-the-art systems while maintaining good streaming latency. We release our model weights and example inference code under an Apache 2.0 license on GitHub and Hugging Face.

2. 语音增强、降噪与音频修复 | 2 篇

2. Neuromorphic Speech Enhancement with Dual-Branch Spiking Neural Networks

基于双分支脉冲神经网络的神经形态语音增强

AI 总结:提出双分支脉冲神经网络GSU-DBNet,通过门控脉冲单元同时建模幅度和复数频谱,在仅394K参数下达到3.04 PESQ,优于现有SNN方法且参数仅为ANN模型的4.5%-10.6%。

链接:https://arxiv.org/abs/2606.23761

机构:School of Communication Engineering, Hangzhou Dianzi University(杭州电子科技大学通信工程学院)

作者:Taiyu Meng, Wenbin Jiang, Haoyi Zhang, Yuhan Zhou, Haibing Yin

英文摘要:Spiking neural network (SNN)-based neuromorphic speech enhancement has emerged as a promising paradigm due to its energy efficiency, yet it still underperforms classical artificial neural network (ANN)-based approaches owing to binary activations and the lack of well-designed network architectures. To overcome this limitation, we propose a novel dual-branch spiking neural network architecture equipped with a gated spiking unit (GSU), termed GSU-DBNet. Specifically, GSU-DBNet simultaneously models the speech magnitude spectrum and complex spectrum, predicting the corresponding magnitude and complex spectral masks. Meanwhile, a dual-path GSU module is adopted to exploit temporal and frequency information for enhanced spatiotemporal feature representation. Experiments on a popular benchmark dataset show that GSU-DBNet achieves a PESQ score of 3.04 with only 394K parameters, outperforming existing SNN-based methods while using only 4.5%--10.6% of the parameters of representative ANN-based models.

3. Beyond U-Net: A Latent-Representation-Aligned Skip-Free Backbone for Flow-Matching Speech Enhancement

超越U-Net:用于流匹配语音增强的无跳跃连接潜在表示对齐骨干网络

AI 总结:提出一种无跳跃连接的编码器-解码器骨干网络,通过潜在表示对齐(LRA)引导流匹配语音增强,避免U-Net跳跃连接传递噪声相关特征,仅需五次函数评估即可提升PESQ和感知质量。

链接:https://arxiv.org/abs/2606.24745

机构:Department of Information Engineering, Electronics and Telecommunications(信息工程、电子与电信系)

作者:Wangyi Pu, Michele Scarpiniti

英文摘要:Generative models, particularly diffusion and score-based approaches, have recently achieved strong performance in speech enhancement, but their iterative sampling process limits real-time deployment. Flow Matching offers an efficient alternative by transporting noisy speech toward clean speech through an ordinary differential equation with few function evaluations. In this work, we propose a skip-free encoder-decoder backbone for flow-matching speech enhancement, guided by Latent Representation Alignment (LRA). Instead of relying on U-Net skip connections, which may transfer noise-correlated low-level features to the decoder, the proposed model aligns its bottleneck and decoder representations with clean latent features extracted from a frozen Descript Audio Codec encoder-decoder without quantization. This codec-aligned supervision promotes compact clean-speech representations while preserving efficient few-step inference. Experiments on WSJ0-CHiME3 and VoiceBank-DEMAND show improved PESQ and perceptual quality, especially on VoiceBank-DEMAND, using only five function evaluations.

3. 音乐信息检索与音乐生成 | 2 篇

4. Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment

使用指令微调和反馈驱动对齐将 MusicLLM 与情感对齐

AI 总结:本文研究如何通过指令微调和反馈驱动对齐策略,使音乐大语言模型(MusicLLM)具备情感回归能力,并证明反馈驱动对齐能显著提升唤醒度和效价预测性能,同时保持音乐问答能力。

链接:https://arxiv.org/abs/2606.24123

机构:LY Corporation

作者:Takuya Hasumi, Welly Naptali

英文摘要:This paper investigates whether music large language models (MusicLLMs) can be aligned for emotion regression. While MusicLLMs have shown strong performance in music information retrieval tasks, their ability to predict arousal and valence scores remains limited, since emotion regression has not been an explicit training objective. To examine whether MusicLLMs can be aligned with emotion, we train MusicLLMs on emotion regression and compare two strategies: instruction tuning and feedback-driven alignment. Our experiments show that task-aware instruction tuning enables MusicLLMs to predict emotion levels to some extent, although the accuracy remains limited. Applying feedback-driven alignment with a verifiable numerical reward substantially improves performance on both arousal and valence over instruction tuning alone. We further show that our approach improves emotion regression performance while maintaining MusicQA capability.

5. Real-Time Interactive Music Generation via Data-Free Streaming Consistency Distillation

实时交互式音乐生成:基于无数据流式一致性蒸馏

AI 总结:提出无数据流式一致性蒸馏框架,将文本到音乐模型转化为低延迟、可实时交互的乐器,支持动态输入引导音乐轨迹。

链接:https://arxiv.org/abs/2606.24307

机构:ZhuoLab(卓实验室)

作者:Baisen Wang, Chenxi Bao, Qisong Han

英文摘要:Interactive music and live performance relies on real-time human expression, but modern generative music AI remains largely absent from this domain due to its prohibitive inference latency and offline rendering paradigm. To provide pioneer musicians with a novel medium for interactive composition, we should fundamentally change these static models into dynamic, playable instruments. In this paper, we propose a framework that bridges this gap. To achieve the low latency required for live interaction without sacrificing structural coherence, we formulate distillation within a streaming autoregressive latent space. Our approach gets rid of the need for expensive paired audio-latent datasets by utilizing prompt-only inputs to synthesize teacher-guided, chunk-wise trajectories on the fly. Because live instruments require high acoustic fidelity, we introduce music-aware consistency objectives, which combine latent, spectral, and temporal-difference losses, to preserve crucial qualities like timbre, transients, and rhythmic stability during accelerated single-step streaming generation. Implemented via parameter-efficient adaptation, our distillation reduces generation steps to achieve a low real-time factor. Crucially, by operating as a continuous autoregressive stream, the system can seamlessly assimilate dynamic human inputs on the fly, allowing users to instantly steer the musical trajectory without interrupting the audio flow. Ultimately, this work recontextualizes generative text-to-music models not as passive prompt-and-wait systems, but as responsive instruments, opening new frontiers for live human-AI musical co-creation.

4. 数据集、基准与评测 | 2 篇

6. VieSpeaker: A Large-Scale Vietnamese Speaker Recognition Dataset Beyond Visual Dependency

VieSpeaker:超越视觉依赖的大规模越南语说话人识别数据集

AI 总结:提出无需面部线索的数据集构建流程,利用文本元数据和大型语言模型推理说话人身份,构建包含4715名说话人约902小时语音的越南语数据集VieSpeaker,实验证明其提升模型鲁棒性和泛化能力。

链接:https://arxiv.org/abs/2606.24066

机构:Hanoi University of Science and Technology(河内科技大学)

作者:Viet Hoang Pham, Tran Trung Nguyen, Bao Thu Ho, Phuong Tuan Dat, Thi Thu Trang Nguyen

英文摘要:Speaker recognition has advanced rapidly with large-scale training datasets, yet Vietnamese remains under-resourced, with existing corpora limited in scale and acoustic diversity. Most large-scale datasets rely on facial cues to link speech with speaker identities, restricting data collection to recordings where speakers appear on camera. We propose a face-independent dataset construction pipeline and introduce VieSpeaker, a large-scale Vietnamese speaker recognition dataset. Our approach leverages textual metadata and large language model reasoning to infer speaker identities from transcripts and contextual information. VieSpeaker contains approximately 902 hours of speech from 4,715 speakers. Experiments show that models trained on VieSpeaker achieve improved robustness and generalization compared to existing Vietnamese datasets. This work demonstrates the feasibility of face-independent dataset construction and provides a new direction for building large-scale speech resources.

7. ParaPairAudioBench: Paralinguistic Pairwise Audio Benchmark for LALM-as-a-Judge

ParaPairAudioBench: 面向LALM-as-a-Judge的副语言成对音频基准

AI 总结:提出ParaPairAudioBench,包含5,175对音频,覆盖风格、语速、重音、年龄和性别五个副语言维度,用于评估大型音频语言模型作为评判者的可靠性,发现其与人类判断差距大且校准失败严重。

链接:https://arxiv.org/abs/2606.24648

机构:Hongik University(弘益大学); Seoul National University(首尔大学); NAVER Cloud(NAVER云); KAIST(韩国科学技术院)

作者:Jisu Jeon, Seungyeon Jwa, Joosung Lee, Jinhyeon Kim, Woojin Chung, Hwiyeol Jo, Jeonghoon Kim, Jonghyun Choi, Soyoon Kim

英文摘要:Large Audio-Language Models (LALMs) have been widely used as judge models for the automatic evaluation of generated speech. However, prior approaches predominantly focus on holistic naturalness, leaving fine-grained paralinguistic distinctions underexplored. We introduce ParaPairAudioBench, a pairwise benchmark of 5,175 audio pairs across five paralinguistic dimensions: Style, Rate, Emphasis, Age, and Gender. Our experiments show that current LALM judges still lag behind human judgments by 32%p on average and exhibit severe calibration failures, particularly in Tie cases where the correct decision is to abstain. To further analyze lexical versus acoustic reliance, the benchmark includes both same-transcript and cross-transcript conditions. ParaPairAudioBench enables multi-dimensional, calibration-aware assessment of the reliability of LALM-as-a-Judge for paralinguistic speech evaluation.

5. 其他/综合语音音频 | 1 篇

8. Statistical validation and full-sphere extension of a Bayesian model for human static sound localisation

人类静态声源定位贝叶斯模型的统计验证与全空间扩展

AI 总结:提出贝叶斯声源定位模型的显式似然函数,通过参数恢复和行为数据拟合验证其可靠性,并比较四种HRTF模板插值方法,发现全空间覆盖和高频保真度是关键。

链接:https://arxiv.org/abs/2606.24367

机构:Dyson School of Design Engineering, Imperial College London(帝国理工学院戴森设计工程学院); Audio Communication Group, Technische Universität Berlin(柏林工业大学音频通信组); Department of Industrial Systems Technology and Management, University of Padova(帕多瓦大学工业系统技术与管理系)

作者:Roberto Barumerli, Fabian Brinkmann, Emanuele Zanoni, Anton Hoyer, Lorenzo Picinali, Michele Geronazzo

英文摘要:Auditory models are central tools for studying spatial hearing, yet their validation typically relies on heuristic performance metrics rather than principled statistical methods. We present two contributions building on a Bayesian sound localisation model that jointly infers sound direction from noisy perceptual features and individual head-related transfer functions (HRTFs). First, we derive an explicit likelihood function and validate it through parameter recovery on simulated data and fitting to behavioural responses from 33 participants, demonstrating that the framework reliably identifies individual sensorimotor and spectral parameters. Second, we use this framework to compare four HRTF template interpolation methods, showing that full-sphere spatial coverage and high-frequency spectral fidelity are the primary determinants of template quality, while the specific interpolation algorithm is secondary. Together, these results show that standard model-based statistical methods can address both fundamental questions in spatial hearing and applied problems such as perceptual HRTF evaluation. An open-source Python implementation is released alongside this work.