微信公众号:arXiv_Daily
cs.SD语音
【1】HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
标题:HQ-SRC:在低资源场景下实现高质量的Zero-Shot歌唱声音转换
链接:https://arxiv.org/abs/2511.08496
备注:Accepted by AAAI 2026 main technical track
【2】Uncertainty Calibration of Multi-Label Bird Sound Classifiers
标题:多标签鸟声分类器的不确定度校准
链接:https://arxiv.org/abs/2511.08261
备注:Under review at ICAART 2026
【3】Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models
标题:Melodia:在扩散模型中以注意力探索为指导的免训练音乐编辑
链接:https://arxiv.org/abs/2511.08252
备注:AAAI 2026
【4】DOA Estimation with Lightweight Network on LLM-Aided Simulated Acoustic Scenes
标题:LLM辅助模拟声场景下轻量级网络的波达方向估计
链接:https://arxiv.org/abs/2511.08012
【5】Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
标题:利用发音兴奋信息和关节运动学的语音情感识别
链接:https://arxiv.org/abs/2511.07955
【6】SpeechJudge: Towards Human-Level Judgment for Speech Naturalness
标题:SpeechJudge:迈向人类水平的言语自然性判断
链接:https://arxiv.org/abs/2511.07931
备注:Project Page: this https URL
【7】SpikCommander: A High-performance Spiking Transformer with Multi-view Learning for Efficient Speech Command Recognition
标题:SpikCommander:一款高性能Spiking Transformer,具有多视图学习功能,用于高效的语音命令识别
链接:https://arxiv.org/abs/2511.07883
备注:Accepted by The Fortieth AAAI Conference on Artificial Intelligence (AAAI 2026)
【8】SynTTS-Commands: A Public Dataset for On-Device KWS via TTS-Synthesized Multilingual Speech
标题:SynTTS-Commands:通过TTS-合成多语言语音的设备上KWS的公共数据集
链接:https://arxiv.org/abs/2511.07821
【9】Speech Separation for Hearing-Impaired Children in the Classroom
标题:课堂上听力障碍儿童的言语分离
链接:https://arxiv.org/abs/2511.07677
备注:13 pages
【10】Enabling Automatic Self-Talk Detection via Earables
标题:通过Earables启用自动自言自语检测
链接:https://arxiv.org/abs/2511.07493
【11】Quantizing Whisper-small: How design choices affect ASR performance
标题:量化Whisper-small:设计选择如何影响ASB性能
链接:https://arxiv.org/abs/2511.08093
备注:Submitted to ICASSP 2026
【12】Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
标题:作为规则化的修剪:ASB中敏感性意识的一次修剪
链接:https://arxiv.org/abs/2511.08092
备注:Submitted to ICASSP 2026
【13】Automatic Music Mixing using a Generative Model of Effect Embeddings
标题:使用效果嵌入生成模型的自动音乐混音
链接:https://arxiv.org/abs/2511.08040
备注:submitted to IEEE ICASSP 2026
【1】Unifying Model and Layer Fusion for Speech Foundation Models
标题:语音基础模型的统一模型和层融合
链接:https://arxiv.org/abs/2511.08389
备注:Accepted by IEEE ASRU 2025
【2】Quantizing Whisper-small: How design choices affect ASR performance
标题:量化Whisper-small:设计选择如何影响ASB性能
链接:https://arxiv.org/abs/2511.08093
备注:Submitted to ICASSP 2026
【3】Pruning as Regularization: Sensitivity-Aware One-Shot Pruning in ASR
标题:作为规则化的修剪:ASB中敏感性意识的一次修剪
链接:https://arxiv.org/abs/2511.08092
备注:Submitted to ICASSP 2026
【4】Automatic Music Mixing using a Generative Model of Effect Embeddings
标题:使用效果嵌入生成模型的自动音乐混音
链接:https://arxiv.org/abs/2511.08040
备注:submitted to IEEE ICASSP 2026
【5】HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
标题:HQ-SRC:在低资源场景下实现高质量的Zero-Shot歌唱声音转换
链接:https://arxiv.org/abs/2511.08496
备注:Accepted by AAAI 2026 main technical track
【6】Melodia: Training-Free Music Editing Guided by Attention Probing in Diffusion Models
标题:Melodia:在扩散模型中以注意力探索为指导的免训练音乐编辑
链接:https://arxiv.org/abs/2511.08252
备注:AAAI 2026
机器翻译由腾讯交互翻译提供,仅供参考
