今日论文合集:CS.SD语音与音频 | 共 12 篇


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音识别与关键词检测 2 篇

2. 语音合成与声音生成 1 篇

3. 音频事件检测与场景理解 2 篇

4. 音乐信息检索与音乐生成 2 篇

5. 语音翻译与语音语言模型 2 篇

6. 安全、隐私与深度伪造音频 2 篇

7. 其他/综合语音音频 1 篇

1. 语音识别与关键词检测 | 2 篇

1. Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition

通过在线精炼提升BEST-RQ伪标签质量用于自动语音识别

AI 总结:提出三种在线精炼方法改进BEST-RQ的伪标签生成,在LibriSpeech上实现12%的词错误率相对降低。

链接:https://arxiv.org/abs/2606.30671

机构:Machine Learning and Human Language Technology Group, RWTH Aachen University(亚琛工业大学机器学习与人类语言技术组); Apptek GmbH(Apptek 有限公司)

作者:Jingjing Xu, Zijian Yang, Mohammad Zeineldeen, Eugen Beck, Ralf Schlueter, Hermann Ney

英文摘要:BEST-RQ is a simple and effective self-supervised training method for speech representation learning that performs well on automatic speech recognition (ASR) tasks. It generates pseudolabels using a fixed online quantization scheme, which simplifies training but provides weaker supervision than HuBERT-style models that iteratively refine pseudo-labels. In this work, we improve online pseudo-label generation while preserving simplicity. We propose three modifications: replacing the quantizer's linear projection with Principal Component Analysis (PCA), updating the codebook via iterative codebook refinement, and introducing an additional codebook updated via codebook distillation. We pre-train on the LibriSpeech 960-hour dataset and fine-tune using 100 hours of supervised LibriSpeech data. With all three modifications enabled, we achieve a 12% relative reduction in word error rate (WER) on the LibriSpeech test-other set, improving from 10.1% to 8.8%.

2. BEST-RQ-2: Contextualize-Then-Predict, a Two-Step Approach for Self-Supervised Audio Representations

BEST-RQ-2:先上下文化再预测——自监督音频表示的两步方法

AI 总结:提出BEST-RQ-2,通过两步式上下文化-预测预训练方案,使用ViT编码器和轻量预测器,在保持推理计算不变下提升跨域迁移性能。

链接:https://arxiv.org/abs/2606.30700

作者:Ludovic K. Tuncay (IRIT-SAMoVA), Etienne Labbé (IRIT-SAMoVA), Thomas Pellegrini (IRIT-SAMoVA)

英文摘要:Self-supervised learning enables audio representations that transfer across domains and tasks. We present BEST-RQ-2, an evolution of BEST-RQ that retains frozen randomprojection-based discrete targets while introducing a two-step contextualize-then-predict pretraining scheme. A ViT context encoder processes only the unmasked spectrogram regions, and a lightweight predictor infers targets for the masked regions; the predictor is discarded after pretraining. Replacing the original Conformer encoder with a ViT shifts performance across domains, slightly reducing speech performance while improving music and environmental sounds, with comparable average scores. The main improvement comes from decomposing masked prediction into separate contextualization and prediction stages. On the X-ARES and XARES-LLM benchmarks, BEST-RQ-2 consistently outperforms one-stage baselines in overall transfer while keeping inference compute unchanged. Code and model checkpoints are publicly available.

2. 语音合成与声音生成 | 1 篇

3. UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

UniSAE:通过离散音素后验图建模实现说话人、情感和低级内容的统一语音属性编辑

AI 总结:提出UniSAE统一框架,通过离散音素后验图(DPPG)表示语音内容,支持从子音素到词级的说话人、情感和内容编辑,并采用扩散解码器生成语音。

链接:https://arxiv.org/abs/2606.31128

作者:Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang, Kun Qian, Yike Guo, Wei Xue

英文摘要:Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE introduces a Discrete Phonetic PosteriorGram (DPPG) representation that factorizes speech content into discrete tokens encoding phoneme identity, pronunciation variants, and duration, enabling direct phoneme- and sub-phoneme-level editing. For higher-level modifications, an autoregressive content transformer predicts edited DPPG sequences for word-level content editing. The edited sequences are rendered into speech by a diffusion-based acoustic decoder, conditioned on disentangled speaker and emotion representations. Experimental results demonstrate that the proposed unified framework supports precise speaker and emotion control, content editing at multiple granularities, and joint modification of all three attributes within a single framework.

3. 音频事件检测与场景理解 | 2 篇

4. SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

SwiftAudio: 用于一步式文本到音频扩散生成的数据高效纯文本蒸馏

AI 总结:提出SwiftAudio框架,通过仅使用文本描述从预训练扩散教师模型进行无音频蒸馏,实现一步式文本到音频生成,在AudioCaps和Clotho上达到最优性能。

链接:https://arxiv.org/abs/2606.31259

机构:Posts and Telecommunications Institute of Technology(邮电技术学院)

作者:Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran

英文摘要:Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising. Existing one-step approaches alleviate this issue but still rely on paired text--audio data during distillation. To address these limitations, we propose SwiftAudio, a one-step TTA framework that performs audio-free distillation from a pretrained diffusion teacher using only text captions. Specifically, we adapt Variational Score Distillation (VSD) to the audio domain and introduce a temporal smoothness regularization objective to encourage coherent latent audio representations. This design enables the student model to inherit the teacher's generative prior without requiring paired audio supervision and allows effective training with only approximately 45K captions. Experiments on AudioCaps and Clotho demonstrate that SwiftAudio achieves state-of-the-art performance among strict one-step methods and substantially narrows the gap to multi-step diffusion systems. Project page: this https URL

5. ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

ZEBRA:面向音频语言模型中基类到新类泛化的零样本熵正则化提示学习

AI 总结:提出ZEBRA框架,通过融合零样本logits与提示学习logits并采用自熵正则化,解决音频语言模型中提示学习导致的新类性能下降问题,实现基类到新类的泛化提升。

链接:https://arxiv.org/abs/2606.31587

机构:Mohamed Bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)

作者:Asif Hanif, Mohammad Yaqub

英文摘要:Audio-Language Models (ALMs) achieve strong zero-shot performance by aligning audio with textual class descriptions. Although prompt learning improves accuracy on base classes through few-shot supervised adaptation, we observe a critical trade-off: it often degrades performance on novel classes, sometimes falling below zero-shot accuracy. This exposes a base-to-novel generalization gap in prompt learning for ALMs. To address this issue, we propose \textbf{ZEBRA} (Zero-shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization), a plug-and-play framework that fuses zero-shot logits with prompt-learning logits, and employs self-entropy regularization to reduce overfitting to base classes. Experiments across multiple audio classification datasets show that ZEBRA consistently improves novel-class performance while maintaining strong base accuracy, significantly reducing the base-to-novel gap compared to standard prompt learning. The code is available at: this https URL.

4. 音乐信息检索与音乐生成 | 2 篇

6. Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

超越二元乐器问答:探究音乐音频-语言模型中的乐器接地

AI 总结:本文通过多轴诊断基准测试发现,音乐音频-语言模型在二元乐器问答中的高准确率常掩盖选项位置偏差、易混淆乐器错误和时间响应偏差等问题,表明乐器接地评估需采用多维度基准而非单一准确率。

链接:https://arxiv.org/abs/2606.31338

作者:Yujun Lee, Joonhyeok Shin, Hyoeun Kim, Kyuhong Shim

英文摘要:Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts. In this paper, we introduce an OpenMIC-derived diagnostic benchmark sequence for instrument grounding in music audio-language models, extending binary instrument-presence QA to genre-prior-reduced examples, confusable instrument discrimination, longer audio context, and temporal localization. Across these settings, high binary QA accuracy often fails to predict model behavior: models can exhibit option-position bias, confusable-instrument errors, and temporal response bias. These results suggest that instrument grounding should be evaluated with multi-axis diagnostic benchmarks rather than a single aggregate accuracy.

7. Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

Dilemmadata:异构罗马数字数据集的可互操作性研究

AI 总结:提出Dilemmadata,通过共享逐音符模式协调两个异构罗马数字语料库,解决标注、表示、工具链和策展四类困境,生成最大同质语料库(1621首作品),并保留84首双重分析作品以支持比较研究。

链接:https://arxiv.org/abs/2606.31595

作者:Johannes Hentschel, Emmanouil Karystinaios, Gerhard Widmer, Markus Neuwirth

英文摘要:In recent years, there has been growing effort to annotate and collect large-scale corpora of Roman numeral analyses in support of data-driven studies in tonal harmony. We introduce dilemmadata, the first resource to reconcile two major collections, the AugmentedNet Dataset (AN) and the Distant Listening Corpus (DLC), making them interoperable through a shared note-wise TSV schema. The reconciliation confronts four families of dilemmata: annotation-standard (the two encode the same musical fact differently in terms of vocabulary size, syntax, conventions for chord extensions, inventory of special chord functions), representational (what counts as a row, and which information survives the conversion), toolchain (incompatible Python ecosystems built around music21 vs. ms3+dimcat), and curatorial (which pieces to include, exclude, or retain twice). We resolve each by deliberately transforming, augmenting, and omitting information, formalising the mismatches, preserving musical semantics, and flagging transformations that may subtly affect annotation fidelity. Consistency checks and qualitative inspections offer a preliminary assessment of post-conversion validity and a basis for critiquing the theoretical assumptions embedded in each original standard. After removing duplicates and merging the two collections, the resulting dilemmadata (1,621 pieces and aprox. 2.8 M note-wise annotations) is the largest homogeneous Roman-numeral corpus currently available, albeit far from perfect. Crucially, we retain 84 pieces common to both corpora under each of their original analyses, yielding a shared reference set in which two equally legitimate analytical traditions can be compared note-for-note over identical musical material. Released on Zenodo, dilemmadata supports interoperability, comparative harmonization modeling, and future refinement of Roman-numeral encoding standards.

5. 语音翻译与语音语言模型 | 2 篇

8. ALM2Vec: Learning Audio Embeddings for Universal Audio Retrieval with Large Audio-Language Models

ALM2Vec: 利用大型音频语言模型学习通用音频检索的音频嵌入

AI 总结:提出ALM2Vec框架,利用预训练大型音频语言模型学习统一嵌入空间,实现跨域、任务感知的音频检索,支持指令感知检索,在标准基准上表现优异。

链接:https://arxiv.org/abs/2606.30682

作者:Fengjie Lu, Chenang Jiang, Jiarui Hai, Helin Wang, Aaron Yee

英文摘要:Recent advances in language--audio retrieval have been largely driven by contrastive dual-encoder architectures that align audio and text in a shared embedding space. While effective, existing retrieval embeddings are primarily optimized for audio--caption matching, limiting their ability to support diverse retrieval objectives and controllable retrieval behaviors. We present ALM2Vec, a universal audio embedding framework derived from pretrained large audio--language models (LALMs). By transferring the audio understanding, instruction-following, and reasoning capabilities acquired through large-scale multimodal training, ALM2Vec learns a unified embedding space for retrieval across audio domains and task types. Beyond conventional text--audio retrieval, ALM2Vec incorporates natural-language instructions into the embedding process, enabling instruction-aware retrieval for scenarios such as audio question answering and aspect-conditioned retrieval. Experimental results show that ALM2Vec achieves competitive performance on standard audio and speech retrieval benchmarks while exhibiting promising compositional and controllable retrieval capabilities, highlighting its potential as a unified audio embedding model for retrieval across domains, tasks, and user intents.

9. FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

FlexiSLM:一种动态可控帧率的语音语言模型

AI 总结:提出FlexiSLM,首个支持动态可控帧率的语音语言模型,利用动态帧率表示在高质量点超越固定帧率7B模型,并在6.25 Hz时推理速度减半且保持高质量。

链接:https://arxiv.org/abs/2606.31247

机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); ByteDance(字节跳动)

作者:Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang, Xu Li, Yufan Lin, Junwen Qiu, Jun Zhang, Lu Lu, Haizhou Li, Zhizheng Wu

英文摘要:Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering no flexibility to trade off quality for speed at inference time. Recent audio tokenizer research has proposed dynamic frame rate speech coding, which exploits this non-uniformity and enables two new capabilities: very low average frame rates and frame rate controllability. However, this technique has not yet been applied to SLMs. We introduce Flexible Spoken Language Model (FlexiSLM), the first SLM that supports dynamic and controllable frame rates on both speech input and output. Using dynamic frame rate representations, FlexiSLM outperforms fixed-frame-rate 7B models including Qwen2.5-Omni and Kimi-Audio at its high-quality operating points. We further verify that FlexiSLM can be accurately steered down to 4.0 Hz; at 6.25 Hz, it roughly halves inference time relative to 12.5 Hz while retaining strong speech-to-speech quality. Audio samples are available at this https URL.

6. 安全、隐私与深度伪造音频 | 2 篇

10. Probing-Guided Layer Selection from Self-Supervised Speech Models for Generalizable Audio Deepfake Detection

基于探针引导的自监督语音模型层选择用于泛化音频深度伪造检测

AI 总结:提出一种模型无关的两阶段方法,通过XGBoost探针评估各层跨域判别力,选择信息层进行融合,在保持冻结骨干的同时显著提升跨域泛化性能。

链接:https://arxiv.org/abs/2606.30791

机构:Michigan Technological University(密歇根理工大学)

作者:Marjan Beheshti, Majid Rostami, Bo Chen

英文摘要:Audio deepfake detection systems often fail to generalize across domains because they rely on features tied to specific attacks or recording conditions. Self-supervised speech models offer rich multi-layer representations, yet existing approaches either use a single layer or fuse all layers indiscriminately, and only reveal layer importance after training. We propose a model-agnostic, two-stage methodology that identifies informative depth zones before any task-specific model is trained. In the first stage, lightweight XGBoost probes evaluate each transformer layer's cross-domain discriminative power, producing a layer ranking. In the second stage, a compact neural classifier fuses only the selected layers through per-layer attention pooling and a shared bottleneck projection, while the backbone remains frozen. Applied across three backbones, the probing reveals two key findings. First, informative layers cluster in depth zones rather than at uniquely optimal positions: within-zone substitutions fall within multi-seed noise, while zone violations degrade performance by up to 5x. Second, the probing produces backbone-specific selections rather than a fixed layer recipe. On XLS-R-300M, four probing-selected layers with 1.34M trainable parameters achieve 4.94 +/- 0.32% equal error rate on In-The-Wild and 5.07% cross-domain average over four shared datasets, a 28% relative improvement over the best prior frozen-backbone result (Xiao and Vu, 2025) using all 25 layers with identical training data.

11. Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

攻击UTMOS:探测语音质量评估模型的鲁棒性

AI 总结:通过分数保持和质量保持两种攻击方式,在原始波形、mel谱图+HiFi-GAN声码器和EnCodec潜在空间三个输入空间上测试UTMOS的鲁棒性,发现分数保持攻击有效,而EnCodec潜在空间最有利于质量保持攻击。

链接:https://arxiv.org/abs/2606.31105

机构:Nagoya University(名古屋大学)

作者:Wen-Chin Huang, Tomoki Toda

英文摘要:UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to probe its robustness. Starting from high-quality speech samples, we optimize the input in two directions: a score-preserving attack, which degrades perceived quality while maintaining the predicted score, and a quality-preserving attack, which lowers the predicted score while maintaining perceived quality. We consider three input spaces: raw waveform, mel spectrogram with a HiFi-GAN vocoder, and the latent space of EnCodec, a neural audio codec. Experimental results show that score-preserving attacks are effective against UTMOS. Although perfect quality-preserving attacks are more difficult, optimization in the EnCodec latent space provides the best chance of success. These results reveal failure modes of UTMOS and highlight the importance of robustness analysis for DNN-based SQA metrics.

7. 其他/综合语音音频 | 1 篇

12. ASR-Agnostic Multimodal Spectrotemporal Modeling for Early Dementia Detection

ASR无关的多模态谱时建模用于早期痴呆检测

AI 总结:提出一种不依赖语音识别的框架,直接从梅尔频谱图中提取谱时位移场作为数字生物标志物,结合CNN-ConvGRU声学嵌入和交叉注意力融合,在三种语言语料库上验证了多模态融合的语料依赖性。

链接:https://arxiv.org/abs/2606.30646

作者:Chukwuemeka Ugwu, Oluwafemi Richard Oyeleke

英文摘要:Speech recruits the same executive, attentional, and working memory processes underlying instrumental activities of daily living, or IADLs, providing a non-invasive proxy for cognitive assessment. Yet most speech-based dementia detection systems depend on transcription, discard within-recording temporal structure, and are validated on a single English corpus with known recording artifacts. We propose an ASR-agnostic framework operating directly on Mel spectrograms. Our key contribution is extracting spectrotemporal displacement fields from consecutive spectrogram frames, capturing shifting spectral energy patterns as digital biomarkers of cognitive decline. These features are fused with CNN-ConvGRU acoustic embeddings via a learned cross-attention mechanism and aggregated using a Transformer encoder with learnable query pooling. A composite temporal loss enforces smoothness and contrastive coherence across segments. We train independent models on English DementiaBank, Slovak EWA-DB, and Spanish Ivanova corpora, using clinical elicitation protocols taxing IADL-relevant cognitive domains. The Slovak model achieves 83.9% accuracy, and Spanish achieves, while the English baseline yields 53.2%, confirming known artifacts. Cross-lingual ablation studies reveal distinct fusion regimes: removing cross-attention collapses Spanish performance to 53.7%, below unimodal models, while the Slovak audio encoder alone outperforms the full model, 93.7% vs. 83.9%, and all English configurations remain near chance. Thus, multimodal fusion's value is corpus-dependent: essential when signal is distributed across modalities, counterproductive when one dominates, and irrelevant when no signal exists. Auxiliary temporal losses converge to language-invariant values, indicating cross-lingual architectural stability.