今日论文合集:CS.SD语音与音频 | 共 11 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition
基于发音特征分解的多任务非标准音素识别
AI 总结:该研究提出分层多任务学习结合交叉注意力融合模块与半监督学习的方法,在L2-ARCTIC数据集上显著提升非标准音素识别性能,为病理语音等非标准语音的鲁棒可解释识别提供新策略。
链接:https://arxiv.org/abs/2608.22273
机构:Kaliber AI(卡利伯人工智能公司)
作者:Sophia Riaz, Haoze Zheng, Amos Roche, Miyu Zhang, Anamika Ragu, Salvatore Penachio, Kaustav Mukherjee, Aneesh Jonelagadda
英文摘要:Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and limited availability of labeled clinical speech data. Existing phoneme recognition systems are typically trained on canonical speech and treat phonemes as atomic categorical labels, limiting their ability to detect structured articulatory errors common in speech disorders and accents. In this work, we introduce a linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. We implement this formulation using a hierarchical multi-task learning architecture in which task-specific articulatory feature heads learn feature-level representations that are subsequently integrated through a cross-attention-based fusion module to produce phoneme predictions. To address the scarcity and noise of pathological speech labels, we combine this framework with semi-supervised learning via Momentum Pseudo-Labeling (MPL) and propose a cascaded training strategy that progressively introduces articulatory feature tasks while employing staged unfreezing of a pretrained speech encoder. Experiments on L2-ARCTIC, used as a proxy for pathological speech variation, show that the proposed approach achieves substantial improvements in phoneme recognition performance compared to strong baseline architectures, while yielding interpretable error patterns aligned with phonological feature structure. These results suggest that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition in non-canonical speech, and motivate future validation on clinically diagnosed pathological speech datasets.

2. sanoTTS: The Smallest Real-Time Neural TTS on a General-Purpose Microcontroller
sanoTTS:通用微控制器上最小的实时神经TTS系统
AI 总结:本文提出sanoTTS,为通用微控制器上无需神经加速器的最小实时神经TTS,基于Piper/VITS教师模型蒸馏,规模小、速度快,在ESP32系列上实现实时语音生成,同时评估了其性能与局限。
链接:https://arxiv.org/abs/2608.21378
作者:Ashish Thapa (Ampixa Labs)
英文摘要:This paper describes an audited neural text-to-speech stack that runs from phoneme IDs to 22.05-kHz PCM on general-purpose microcontrollers. Its deployed graph has 567,008 parameters, and its two int8 blobs occupy 679,832 bytes. On an ESP32-S3, the complete duration-acoustic-inverse-STFT path generates 4.54 s of speech in 1.02 s (0.22x real time) without a neural accelerator. The same portable C core runs offline at 5.72x real time on an FPU-less ESP32-C3. To our knowledge, this is the smallest complete phoneme-to-waveform neural TTS graph demonstrated in real time on a general-purpose microcontroller without a neural accelerator. We derive the students from the conditional-VAE objective of their Piper/VITS teachers and state the duration, latent-interface, waveform, adversarial, and joint-distillation losses used in training. The size and speed come with an audible cost: on unseen text, the embedded stack distilled from en_US-kristin-medium scores 2.54 SCOREQ and 2.80 UTMOS, compared with 4.68 and 4.42 for its teacher. A separate English quality package uses the stronger en_US-amy-medium teacher. Its 1,454,284-parameter Pareto point scores 4.13 SCOREQ and 4.10 UTMOS; a 1,834,380-parameter variant scores 4.16 SCOREQ. A controlled capacity study with Kristin identifies the decoder, rather than the output representation, as the main constraint. Two evaluation failures also affected the work: a narrow, templated test set overstated one early student's SCOREQ by 1.35, and aggregate quality predictors missed a sibilant failure that was evident in listening and in a phoneme-resolved spectral probe. Checksums cover the reported model blobs, runtime ports, and golden vectors.

3. AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS
AudioNoisePrints:基于流匹配TTS空间相关性的无模型音频水印
AI 总结:该研究提出无训练的AudioNoisePrints音频水印流水线,利用流匹配TTS的噪声-输出相关性实现水印,在强增强下优于AudioSeal,可适配多种TTS模型。
链接:https://arxiv.org/abs/2608.22186
机构:National Research Council Canada(加拿大国家研究委员会); University of British Columbia(不列颠哥伦比亚大学)
作者:Timothy Tin-Long, Jian Zhu, Aidan Pine, Mengzhe Geng
英文摘要:We present AudioNoisePrints, a training-free watermarking pipeline for flow matching and diffusion TTS models, which requires minimal extra computation during inference and does not require retraining the TTS model or reducing the generation quality. We exploited the fact that there are strong correlations between the initial Gaussian noises and the generated outputs in diffusion and flow matching models, such that a simple cosine correlation between the initial noise and the generated output can be used to perform watermaking. Moreover, we train a lightweight detector on top for more aggressive augmentations. Our method outperforms AudioSeal, a strong baseline for audio watermarking under strong augmentations. We experimented on F5TTS and other TTS and vocoder models, and concluded that they all exhibit similar spatial correlation properties, suggesting our watermarking scheme can be used for more flow-matching TTS models and even vocoders in the future.

4. LipsAM: Lipschitz-continuous Neural Networks for Convergent Plug-and-Play Audio Signal Recovery
LipsAM:用于收敛即插即用音频信号恢复的利普希茨连续神经网络
AI 总结:本文针对现有理论框架无法适配音频处理常用DNN的问题,提出利普希茨连续的LipsAM架构,开发其利普希茨常数评估框架,并将其应用于即插即用音频恢复,实现了可收敛的语音去混响算法。
链接:https://arxiv.org/abs/2608.23038
机构:Tokyo University of Agriculture and Technology(东京农工大学)
作者:Kazuki Matsumoto, Ren Uchida, Natsuki Yoshino, Kohei Yatabe
英文摘要:The Lipschitz continuity of deep neural networks (DNNs) is essential for establishing theoretical guarantees regarding their behavior. From both theoretical and practical perspectives, various methods have been proposed to construct Lipschitz-continuous architectures and control their Lipschitz constants. However, several DNN architectures common in audio signal processing fall outside the scope of existing theoretical frameworks, hindering the development of Lipschitz-continuous models in acoustic applications. In particular, despite their widespread adoption, DNNs that separately process the magnitude and phase of complex-valued signals cannot be Lipschitz continuous under existing frameworks. In this paper, to address this limitation, we establish a theoretical foundation for constructing amplitude modifiers (AMs), a class of DNN architectures that operate solely on the magnitude of a complex-valued input, with provable Lipschitz continuity. Specifically, we derive a necessary and sufficient condition for an AM to be Lipschitz continuous and propose LipsAMs (Lipschitz-continuous AMs) corresponding to common architectures for audio signals, including time-frequency masking. Furthermore, we develop an efficient framework for evaluating their Lipschitz constants and analytically derive these constants for some of the proposed architectures. As an application, we propose CoReM-LipsAM (Controlled Residual Maps via LipsAM) for plug-and-play (PnP) audio signal recovery, integrating a DNN as a data-driven prior within a model-based signal processing algorithm. The convergence of the obtained PnP algorithm is structurally guaranteed by the CoReM-LipsAM architecture and empirically validated through speech dereverberation experiments.

5. MusPyExpress: Extending MusPy with Enhanced Expression Text Support
MusPyExpress:为MusPy扩展增强的表情文本支持功能
AI 总结:该研究提出MusPyExpress扩展MusPy库,支持提取表情文本,解析PDMX数据集展示相关数据丰富性,并开展三类利用该信息的生成任务以推动符号音乐建模。
链接:https://arxiv.org/abs/2608.21678
机构:University of California, San Diego(加利福尼亚大学圣迭戈分校); University of Michigan, Ann Arbor(密歇根大学安娜堡分校)
作者:Phillip Long, Hao-Wen Dong, Julian McAuley, Zachary Novack
英文摘要:Current work in modeling symbolic music primarily relies on representations extracted from MIDI-like data. While such formats allow for modeling symbolic music as sequences of notes, they omit the large space of symbolic annotations common in western sheet music broadly known as expression text, such as tempo or dynamics, which specify time- and velocity-dependent controls on the musical composition and performance. To alleviate this gap, we present MusPyExpress, an extension to the popular symbolic music processing library MusPy that enables the extraction of expression text along with symbolic music for downstream modeling. Utilizing this extension, we parse the PDMX dataset to illustrate the wealth of expression text available in MusicXML datasets. Additionally, we introduce multiple generative tasks, including joint expression-note generation, expression-conditioned music generation, and expression tagging, that take advantage of this additional notational information.

6. MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
MRMAD:用于评估大型音频语言模型中声学降级感知的多轮多音频基准
AI 总结:该研究提出MRMAD多轮多音频基准,评估18种LALMs的音频降级感知,发现现有模型难可靠推理降级,为构建鲁棒LALMs提供基础。
链接:https://arxiv.org/abs/2608.22236
机构:Northeastern University(东北大学); Bose Corporation(博士公司); Stony Brook University(石溪大学)
作者:Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang, Andrew C. Singer, Xue Lin, Chuan-Che Huang, Shuo Zhang

7. Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering
面向推理的后训练与推理时LoRA重缩放:针对音频相关问答任务
AI 总结:该研究针对音频相关问答任务,提出面向推理的LoRA后训练与推理时重缩放方法,在Qwen和MOSS-Audio模型上验证了有效性,提交系统在挑战赛中获总体第三、轻量级系统第二。
链接:https://arxiv.org/abs/2608.23092
作者:Weiteng Hu, Yin Cao, Jun Yang

8. AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection
AT-ADD:面向鲁棒全类型音频深度伪造检测的基准与挑战赛
AI 总结:本文提出AT-ADD基准与挑战赛,设双赛道分别评估鲁棒语音深度伪造检测与全类型音频深度伪造检测,官方基线及获胜系统取得对应性能,同时揭示泛化关键及未解决问题。
链接:https://arxiv.org/abs/2608.23437
机构:Communication University of China(中国传媒大学); Ant Group(蚂蚁集团); Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所); Beijing Institute of Technology(北京理工大学); Shanghai Jiao Tong University(上海交通大学)
作者:Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo, Tao Wang, Changhao Zhang, Jian Liu, Weiqiang Wang, Ruibo Fu, Xiaopeng Wang, Hengyan Huang, Xiaoying Huang, Long Ye, Guangtao Zhai

9. Vibrato Matching for Modulation Control and Blending in Sound Mixtures
声音混合中用于调制控制与融合的颤音匹配
AI 总结:本研究提出颤音匹配算法,可抑制目标信号颤音并转移源信号颤音,用于声音混合的颤音控制与声源融合,且匹配颤音会降低声源分离算法性能。
链接:https://arxiv.org/abs/2608.22057
作者:Jeremy Hyrkas


10. FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation
FlowSep 2:用于语言查询音频源分离的自监督流匹配方法
AI 总结:本研究提出FlowSep2,一种结合Self-Flow与Diffusion Transformer的文本条件流匹配生成模型,用于语言查询音频源分离,在多个基准上达到SOTA性能,可有效分离重叠声源。
链接:https://arxiv.org/abs/2608.22111
机构:School of Computer Science and Electronic Engineering, University of Surrey(萨里大学计算机科学与电子工程学院); Meta Superintelligence Labs(元宇宙超级智能实验室); Department of Informatics, King’s College London(伦敦国王学院信息学系)
作者:Yi Yuan, Xubo Liu, Haohe Liu, Xiyuan Kang, Mark D. Plumbley, Wenwu Wang

11. Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings
从非侵入式脑记录中解码感知语音的跨被试泛化
AI 总结:该研究提出CPSD框架,结合对比学习、PESA模块等,在三类感知语音数据集上跨被试解码性能优于基线,提升了Top-10准确率。
链接:https://arxiv.org/abs/2608.22420
机构:School of Intelligence Science and Technology, Peking University(北京大学智能科学与技术学院); Academy for Advanced Interdisciplinary Studies, Peking University(北京大学前沿交叉学科研究院); National Key Laboratory of General Artificial Intelligence(通用人工智能国家重点实验室)
作者:Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen