今日论文合集:CS.SD语音与音频 | 共 10 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
感知耳语的大语言模型:用于鲁棒耳语语音识别的自监督不确定性学习
AI 总结:本文提出Whisper-Aware LLM框架,通过自监督学习量化声学信号缺陷并结合置信融合解码机制,在AISHELL6-Whisper数据集上将耳语语音识别的CER相对降低17%,幻觉率降至4.5%,达到最优性能。
链接:https://arxiv.org/abs/2608.10836
机构:Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
作者:Gaopeng Xu, Zhenyu Wang, Zheng Xue, Yinfeng Xia, Haitao Yao
英文摘要:The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to this uncertainty. Our model develops an intrinsic self-awareness by learning to quantify the physical deficiencies of acoustic signals through targeted self-supervised tasks. This learned uncertainty is then operationalized via a novel Confidence-Fused Decoding mechanism, which provides both high-level instructions and frame-level attention modulation to the LLM decoder. Our experiments confirm the effectiveness of this approach. The model sets a new state-of-the-art on whispered speech with a 17% relative CER reduction on AISHELL6-Whisper. At the same time, it directly addresses the reliability trade-off, with hallucination rates dropping from over 25% to 4.5%.

2. Training Set Synthesis for Bioacoustic Denoising: A Case Study With Mice
用生物声学去噪的训练集合成:以小鼠为例
AI 总结:针对生物声学去噪的训练集合成方法,开发了脊线引导损失的U-Net去噪模型,以小鼠超声发声为案例,提升了脊线跟踪与USV分类性能,可推广至其他生物声学信号。
链接:https://arxiv.org/abs/2608.10054
作者:Reyhaneh Abbasi, Peter Balazs, Vincent Lostanlen, Clara Hollomey, Dustin J. Penn, Sarah M. Zala, Nicki Holighaus
英文摘要:Bioacoustic recordings are often degraded by ambient noise, which complicates the analysis of weak or noise-overlapped vocalizations. Convolutional neural networks, particularly U-Net architectures, have shown a strong denoising performance in speech and music processing. However, their direct application to bioacoustic signals is limited by the scarcity of clean training data. To address this issue, we propose a training set synthesis approach and develop a supervised denoising model that predicts a complex ratio mask in the time-frequency domain. The model leverages ridges, or frequency contours, that represent the fundamental frequency together with one or more harmonic partial components of vocalizations. These ridges are used both for the synthesis of training sets and to design a loss function that assigns higher weights to the ridge regions (ridge-guided loss function). This weighting step helps the network better preserve vocalization details during denoising. As a case study, we evaluate our approach using ultrasonic vocalizations (USVs) recordings of house mice, which are widely studied in behavioral biology and neuroscience. In actual field recordings, the proposed method enhances fundamental and harmonic partial ridge tracking compared to our previous signal-processing approach. In addition, a classifier trained on denoised data improves USV classification on out-of-sample, noisy recordings from wild and domesticated mice compared to classifiers trained on noisy recordings. Our proposed method also substantially improves the scale-invariant signal-to-distortion ratio on synthetic testing data across a wide range of input signal-to-noise ratios. Although we focus on USVs, the proposed approach should be broadly applicable to other bioacoustic signals with trackable ridges, and thus enables ridgebased training set synthesis and denoising.

3. DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning
DINO-A:将自蒸馏视觉Transformer适配至通用音频表示学习
AI 总结:本研究提出DINO-A,将自蒸馏视觉Transformer适配至通用音频表示学习,通过替换输入模态和增强方式实现,在多音频数据集上验证了其性能,发现了补丁分辨率、骨干选择对任务的影响及与BYOL-A v2的性能差异原因。
链接:https://arxiv.org/abs/2608.10659
机构:Warsaw University of Technology(华沙理工大学)
作者:Tomasz Radzikowski, Mateusz Modrzejewski, Przemysław Rokita
英文摘要:We present DINO-A, an adaptation of self-distillation from vision to general audio representation learning. While DINO has become a canonical method in self-supervised vision and prior audio work has explored latent prediction (BYOL-A) and masked modeling (Audio-MAE, BEATs), no prior work has brought canonical DINO to general audio classification in the way BYOL-A brought BYOL. DINO-A retains DINO's multi-crop, EMA teacher, and high-dimensional projection, replacing only the input modality and augmentations with log-mel spectrograms and the BYOL-A v2 augmentation block. We pretrain three backbones, two Vision Transformers with 8x8 and 16x16 patches and a convolutional encoder, on FSD50K and evaluate them with linear probing on ESC-50, Speech Commands v2, UrbanSound8K, and GTZAN. Three findings characterize the resulting representations. Patch resolution within the Vision Transformer family has consistent effect on representation quality, with smaller patches winning across all four tasks. The choice between Vision Transformer and convolutional backbone interacts with task type: convolutional networks lead on speech while Vision Transformers lead on environmental sounds and music. Under identical pretraining and evaluation conditions, DINO-A and BYOL-A v2 differ by 11.96 percentage points on average, and we trace this difference to two mechanisms: the interaction between DINO's high-dimensional projection space and FSD50K's limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2 does not. The high-dimensional projection space, central to DINO's success in vision, becomes a liability at FSD50K scale.

4. Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
永不停止说话:针对端到端语音语言模型的拒绝服务攻击
AI 总结:本研究针对端到端语音语言模型提出基于扰动的拒绝服务攻击,通过优化声学扰动抑制EOS生成以延长解码,在三类开源模型上验证了其攻击有效性及安全风险。
链接:https://arxiv.org/abs/2608.10405
作者:Shuozhe Cheng, Kunlan Xiang, Mingxuan Li, Ji Zhang, Dongxiao Liu, Wenbo Jiang
英文摘要:Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) attacks target text-only LLMs, end-to-end (E2E) speech LLMs are rapidly emerging. Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inputs and therefore cannot be directly transferred to continuous speech inputs. Moreover, prior studies on speech model security mainly focus on ASR or TTS systems, leaving the DoS vulnerability of E2E speech LLMs largely unexplored. To address this gap, we propose the perturbation-based DoS attack targeting E2E speech models. Instead of inducing long outputs through prompt manipulation, our method optimizes imperceptible acoustic perturbations to directly influence the model's autoregressive generation process while preserving the original input length. Specifically, we formulate the attack as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alignment loss. To further improve stealthiness, we employ voice activity detection (VAD) to inject perturbations only into voiced regions. Extensive experiments on three open-source E2E speech LLMs demonstrate that our method achieves stable attack success rate while significantly increasing generation length and GPU resource consumption, revealing security risks in modern ALLMs.

5. VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation
VoxSumm:用于联合摘要与翻译的多语言长篇口语新闻语料库
AI 总结:本研究针对长文档摘要与多语言语音翻译的研究缺口,提出联合语音摘要与翻译任务,构建首个多语言跨语言基准VoxSumm,并评估相关模型表现,为多语言语音处理系统提供支撑。
链接:https://arxiv.org/abs/2608.10359
机构:Mila - Quebec AI Institute(米拉-魁北克人工智能研究所); McGill University(麦吉尔大学); Google DeepMind(谷歌DeepMind); Canada CIFAR AI Chair(加拿大CIFAR人工智能主席项目)
作者:Yejin Jeon, Marie Maltais, Virginia Ceccatelli, Min Ma, David Ifeoluwa Adelani
英文摘要:As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.

6. DIY e-HandPan: A new DIY Low-Cost Handpan Interface based on Arduino and ESP32 Microcontrollers
DIY电子手碟:一种基于Arduino和ESP32微控制器的新型DIY低成本手碟界面
AI 总结:该研究提出一种基于Arduino和ESP32的开源低成本可定制DIY电子手碟界面,含双色LED引导系统,适用于音乐表演、教育与研究,相关资源开源以促进社区开发。
链接:https://arxiv.org/abs/2608.10185
机构:IBISC, University of Évry Paris-Saclay(埃夫里巴黎-萨克雷大学IBISC研究所)
作者:Benoit Collin, Dominique Fourer, Eric Genotelle
英文摘要:We present DIY e-HandPan, a new open-source, low-cost and customizable handpan audio and MIDI protocol interface designed for musical performance, education and research. The proposed hardware is built from inexpensive electronic components and recycled materials using widely available fabrication techniques, making it accessible to makers, educators and researchers. The instrument can be implemented on two distinct MicroController Unit (MCU): Arduino or ESP32. The microcontroller captures strike velocity to provide expressive musical performance comparable to that of an acoustic handpan. In addition to real-time audio and MIDI generation, DIY e-HandPan integrates a bi-color LED guidance system capable of displaying musical sequences from MIDI files, providing an effective learning aid for beginners and educational activities. The modular architecture allows users to easily customize the number of notes, hardware configuration and embedded software according to specific applications. We present the complete hardware design, firmware and assembly instructions and we discuss the design choices and limitations, to evaluate the system in representative educational and musical performance scenarios. All design files, source code and documentation are released under an open-source license to improve the reproducibility and encourage further developments by the open hardware community.

7. Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning
超越干参考:通过对比距离学习学习相对音频效果表征
AI 总结:针对现有音频效果建模依赖干参考的问题,提出无干参考的对比学习框架RelFx,通过双分支暹罗编码器等实现相对效果表征,在Fx风格迁移任务上优于现有方法
链接:https://arxiv.org/abs/2608.10573
作者:Xinlu Liu, Huibin Lin, Weixing Wei, Zhenhai Yan
英文摘要:Audio effects (Fx) representation learning plays a key role in intelligent music production, including automatic mixing and Fx style transfer. Existing methods typically rely on dry or nearly dry references for effect modeling, yet truly unprocessed audio is rarely available in practice, as real recordings inevitably reflect the microphone, room acoustics, and preceding signal processing. Instead of pursuing absolute effect encodings, we argue that the relative effect distance between audio signals is more meaningful for real-world music production. Motivated by this, we propose RelFx, a contrastive learning framework that learns relative effect transformations from general audio collections without requiring dry references during representation training. Our approach uses a dual-branch Siamese encoder equipped with cross-attention and differential gating fusion to infer the shared effect transformation from a reference clip and an effect-processed, content-related clip. We further propose an antisymmetric fusion variant for bidirectional effect encoding, such that swapping the input order directly produces a nearly sign-reversed embedding, a property not explored in earlier work. Moreover, our dry-reference-free formulation eliminates the reliance on dry multitrack datasets and enables training on effect-bearing audio. Experiments on Fx style transfer demonstrate state-of-the-art performance under the standard Fx-Encoder++ MUSDB18 evaluation protocol, consistently outperforming existing approaches across all four instrument categories.

8. DuplexWorld: Can voice agents help you get through the day?
DuplexWorld:语音智能体能帮你度过一天吗?
AI 总结:DuplexWorld针对现有语音智能体评估基准的不足,构建含六大领域的156个场景开展评估,发现现有最优语音智能体在多维度仍有较大改进空间,并分析了相关性能与失败模式。
链接:https://arxiv.org/abs/2608.10716
机构:Centific Global Solutions Inc.(森蒂菲克全球解决方案公司); University of Maryland(马里兰大学)
作者:Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji, Dinesh Manocha
英文摘要:Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.

9. Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis
基于VQ-VAE的音高轮廓分词及其在韩国传统音乐分析中的应用
AI 总结:该研究针对连续音高运动的音乐传统,用VQ-VAE从无标注音频学习音高轮廓分词,在韩国传统音乐分析中验证了分词可作为语料库级分析单元的有效性。
链接:https://arxiv.org/abs/2608.10979
作者:Seonguk Ju, Seola Cho, Sooin Chung, Danbinaerin Han, Dasaem Jeong
英文摘要:Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learning a vocabulary of local pitch-contour patterns directly from unlabeled audio, using a VQ-VAE that quantizes fixed-length contour segments into a finite codebook. To make the learned tokens stable across segmentation positions and small variations in timing and pitch range, we train the model with a reconstruction objective evaluated under the best alignment among a set of candidate temporal and pitch-domain transformations. Applied to Korean traditional music, the learned tokens recover information about expert-defined sigimsae categories without supervision, and in pansori individual tokens align with the two principal modes, Gyemyeonjo and Ujo, supporting their use as units for corpus-level analysis of contour-centric traditions.

10. Measuring Cross-Cultural Style Diffusion Through Era Classification: US and Korean Popular Music
通过时代分类测量跨文化风格传播:美国与韩国流行音乐
AI 总结:本研究提出时代分类框架,通过CNN分类器量化美韩流行音乐的跨文化风格传播,发现1960-80年代韩流滞后美流4-5年,90年代后偏差缩小,该框架可推广至其他榜单文化对。
链接:https://arxiv.org/abs/2608.10980
作者:Dasol Lee, Minhee Lee, Seonguk Ju, Daewoong Kim, Harin Lee, Dasaem Jeong
英文摘要:Popular music circulates globally while being locally reinterpreted, yet this process of cross-cultural style diffusion has rarely been quantified. We propose an era-classification framework for measuring temporal alignment between chart cultures. CNN classifiers trained from scratch on Billboard Hot 100 audio are applied to Korean Melon chart songs. Korean chart songs from the 1960s through the 1980s are consistently inferred as belonging to earlier Billboard eras, by a median of about four to five years, while the same models remain unbiased on held-out Billboard audio. The offset then halves at the 1990s, to roughly two to three years, and holds there through the 2000s. Reverse inference shows a complementary narrowing, and the pattern holds across architectures and seeds. We interpret these results as reflecting how globally circulating pop styles were locally adopted and progressively synchronized. The framework can be applied to other pairs of chart cultures beyond the US-Korea case examined here.