今日论文合集:cs.SD语音9篇,eess.AS音频处理5篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
标题:HPSO:现实世界口语理解中人类感知的基准
链接:https://arxiv.org/abs/2511.23178

作者:Chen Li,Peiji Yang,Yicheng Zhong,Jianxing Yu,Zhisheng Wang,Zihao Gou,Wenqing Chen,Jian Yin
备注:Accepted by AAAI 2026
摘要:语音大语言模型(Speech LLM)的最新进展已经导致了语音理解任务(如自动语音识别(ASR)和语音情感识别(SER))的巨大进步。然而,这些模型是否能够实现人类水平的听觉感知,特别是在理解现实世界口语中潜在意图和隐含情感的能力方面,仍然有待探索。为此,我们引入了人类水平的感知口语理解(HPSU),一个新的基准充分评估人类水平的感知和理解能力的语音LLM。HPSU包含超过20,000个经过专家验证的英语和汉语口语理解样本。它建立了一个全面的评估框架,涵盖了一系列的任务,从基本的说话人属性识别到复杂的推理的潜在意图和隐含的情绪。为了解决现实场景中数据稀缺和人工标注成本高的问题,我们开发了一个半自动标注过程。该过程融合了音频、文本和视觉信息,以实现精确的语音理解和标记,从而提高注释效率和质量。我们系统地评估各种开源和专有的语音LLM。结果表明,即使是表现最好的模型,在理解真正的口语交互方面,仍然远远低于人类的能力。因此,HPSU将有助于指导语音LLM向人类水平的感知和认知的发展。
摘要:Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of their ability to comprehend latent intentions and implicit emotions in real-world spoken language, remains underexplored. To this end, we introduce the Human-level Perception in Spoken Speech Understanding (HPSU), a new benchmark for fully evaluating the human-level perceptual and understanding capabilities of Speech LLMs. HPSU comprises over 20,000 expert-validated spoken language understanding samples in English and Chinese. It establishes a comprehensive evaluation framework by encompassing a spectrum of tasks, ranging from basic speaker attribute recognition to complex inference of latent intentions and implicit emotions. To address the issues of data scarcity and high cost of manual annotation in real-world scenarios, we developed a semi-automatic annotation process. This process fuses audio, textual, and visual information to enable precise speech understanding and labeling, thus enhancing both annotation efficiency and quality. We systematically evaluate various open-source and proprietary Speech LLMs. The results demonstrate that even top-performing models still fall considerably short of human capabilities in understanding genuine spoken interactions. Consequently, HPSU will be useful for guiding the development of Speech LLMs toward human-level perception and cognition.


【2】Adapting Neural Audio Codecs to EEG
标题:使神经音频编解码器适应脑电
链接:https://arxiv.org/abs/2511.23142

作者:Ard Kastrati,Luca Lanzendörfer,Riccardo Rigoni,John Staib Matilla,Roger Wattenhofer
备注:Foundation Models for the Brain and Body (BrainBodyFM@NeurIPS)
摘要:EEG和音频本质上是不同的模态,在采样率、通道结构和尺度上不同。然而,我们表明,预训练的神经音频编解码器可以作为EEG压缩的有效起点,前提是数据经过预处理以适合编解码器的输入约束。使用DAC,最先进的神经音频编解码器作为我们的基础,我们证明了原始EEG可以映射到编解码器的基于步幅的帧,使音频预训练的编码器-解码器的直接重用。即使没有修改,这种设置也会产生稳定的EEG重建,与从头开始训练相比,对EEG数据的微调进一步提高了保真度和泛化能力。我们系统地探索压缩质量的权衡,通过改变残差码本深度,码本(词汇)的大小,和输入采样率。为了捕获电极之间的空间依赖性,我们提出了DAC-MC,这是一种具有基于注意力的跨通道聚合和通道特定解码的多通道扩展,同时保留了音频预训练的初始化。对TUH异常和癫痫数据集的评估表明,自适应编解码器保留了临床相关信息,如基于频谱的重建损失和下游分类准确性所反映的。
摘要:EEG and audio are inherently distinct modalities, differing in sampling rate, channel structure, and scale. Yet, we show that pretrained neural audio codecs can serve as effective starting points for EEG compression, provided that the data are preprocessed to be suitable to the codec's input constraints. Using DAC, a state-of-the-art neural audio codec as our base, we demonstrate that raw EEG can be mapped into the codec's stride-based framing, enabling direct reuse of the audio-pretrained encoder-decoder. Even without modification, this setup yields stable EEG reconstructions, and fine-tuning on EEG data further improves fidelity and generalization compared to training from scratch. We systematically explore compression-quality trade-offs by varying residual codebook depth, codebook (vocabulary) size, and input sampling rate. To capture spatial dependencies across electrodes, we propose DAC-MC, a multi-channel extension with attention-based cross-channel aggregation and channel-specific decoding, while retaining the audio-pretrained initialization. Evaluations on the TUH Abnormal and Epilepsy datasets show that the adapted codecs preserve clinically relevant information, as reflected in spectrogram-based reconstruction loss and downstream classification accuracy.


【3】Probabilistic Fusion and Calibration of Neural Speaker Diarization Models
标题:神经说话人模型的概率融合与校正
链接:https://arxiv.org/abs/2511.22696

作者:Juan Ignacio Alvarez-Trejos,Sergio A. Balanya,Daniel Ramos,Alicia Lozano-Diez
摘要:端到端神经日志化(EEND)系统产生帧级概率说话者活动估计,然而由于评估主要集中在日志化错误率(DER)上,因此这些置信度分数的可靠性和校准在很大程度上被忽略。当融合多个日志系统时,DOVER-Lap仍然是唯一建立的方法,在细分市场层面上进行艰难的决策。我们建议使用连续的概率输出,这使得更复杂的校准和融合技术,可以利用模型的不确定性和不同架构的互补优势。本文提出了第一个全面的框架,在概率水平上校准和融合EEND模型。我们调查两个输出公式(多标签和幂集表示)和它们的影响,校准和融合的有效性。通过对CallHome双扬声器基准测试的大量实验,我们证明了适当的校准即使对于单个模型也可以提供实质性的改进(相对DER降低高达19%),在某些情况下可以减轻域自适应的缺失。我们发现,在幂集空间中的联合校准始终优于独立的每个扬声器校准,融合,然后校准排序一般优于融合前校准单个模型,而只需要一个单一的组合模型的校准。我们的最佳配置在DER方面优于DOVER-Lap,同时提供下游应用所需的可靠置信度估计。这项工作提出了概率级融合的EEND系统的最佳实践,并证明了利用软输出硬决策的优势。
摘要:End-to-End Neural Diarization (EEND) systems produce frame-level probabilistic speaker activity estimates, yet since evaluation focuses primarily on Diarization Error Rate (DER), the reliability and calibration of these confidence scores have been largely neglected. When fusing multiple diarization systems, DOVER-Lap remains the only established approach, operating at the segment level with hard decisions. We propose working with continuous probability outputs, which enables more sophisticated calibration and fusion techniques that can leverage model uncertainty and complementary strengths across different architectures. This paper presents the first comprehensive framework for calibrating and fusing EEND models at the probability level. We investigate two output formulations (multilabel and powerset representations) and their impact on calibration and fusion effectiveness. Through extensive experiments on the CallHome two-speaker benchmark, we demonstrate that proper calibration provides substantial improvements even for individual models (up to 19% relative DER reduction), in some cases mitigating the absence of domain adaptation. We reveal that joint calibration in powerset space consistently outperforms independent per-speaker calibration, and that the Fuse-then-Calibrate ordering generally outperforms calibrating individual models before fusion while requiring calibration of only a single combined model. Our best configuration outperforms DOVER-Lap in terms of DER while providing reliable confidence estimates essential for downstream applications. This work proposes best practices for probability-level fusion of EEND systems and demonstrates the advantages of leveraging soft outputs over hard decisions.


【4】PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
标题:PURE Codec:语音编解码器学习的剩余信息的渐进展开
链接:https://arxiv.org/abs/2511.22687

作者:Jiatong Shi,Haoran Wang,William Chen,Chenda Li,Wangyou Zhang,Jinchuan Tian,Shinji Watanabe
备注:Accepted by ASRU2025
摘要:神经语音编解码器在低比特率压缩方面取得了很好的性能,但残差矢量量化(RVQ)通常存在训练不稳定和分解无效的问题,限制了重建质量和效率。我们提出了PURE编解码器(残差熵的渐进展开),这是一种新的框架,它使用预先训练的语音增强模型来指导多级量化。第一个量化阶段重建低熵、去噪的语音嵌入,而后续阶段对残留的高熵分量进行编码。该设计显著提高了训练稳定性。实验表明,PURE在重建和下游基于语音语言模型的文本到语音转换中始终优于传统的基于RVQ的编解码器,特别是在嘈杂的训练条件下。
摘要:Neural speech codecs have achieved strong performance in low-bitrate compression, but residual vector quantization (RVQ) often suffers from unstable training and ineffective decomposition, limiting reconstruction quality and efficiency. We propose PURE Codec (Progressive Unfolding of Residual Entropy), a novel framework that guides multi-stage quantization using a pre-trained speech enhancement model. The first quantization stage reconstructs low-entropy, denoised speech embeddings, while subsequent stages encode residual high-entropy components. This design improves training stability significantly. Experiments demonstrate that PURE consistently outperforms conventional RVQ-based codecs in reconstruction and downstream speech language model-based text-to-speech, particularly under noisy training conditions.


【5】Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
标题:基于LLM的端到端口语对话状态跟踪的联合语音和文本训练
链接:https://arxiv.org/abs/2511.22503

作者:Katia Vendrame,Bolaji Yusuf,Santosh Kesiraju,Šimon Sedláček,Oldřich Plchot,Jan Černocký
备注:submitted to ICASSP 2026
摘要:端到端的口语对话状态跟踪(DST)由于必须处理语音输入和数据稀缺而变得困难。在最近的工作中,已经提出了将语音基础编码器和大型语言模型相结合,以减轻这种困难。虽然这种方法已被证明可以产生强大的口语DST模型,在现实的多回合DST中实现最先进的性能,但它很难跨领域进行概括,并且需要为每个感兴趣的领域提供注释的口语DST训练数据。然而,为每个目标领域收集此类数据既昂贵又困难。注意到文本DST数据更容易获得各个领域,在这项工作中,我们建议联合训练可用的口语DST数据和来自其他领域的书面文本数据,作为实现跨领域泛化的一种方式。我们进行的实验表明,我们提出的方法获得良好的跨域DST性能,而不依赖于口语训练数据的目标领域的有效性。
摘要:End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large language models has been proposed in recent work as to alleviate some of this difficulty. Although this approach has been shown to result in strong spoken DST models, achieving state-of-the-art performance in realistic multi-turn DST, it struggles to generalize across domains and requires annotated spoken DST training data for each domain of interest. However, collecting such data for every target domain is both costly and difficult. Noting that textual DST data is more easily obtained for various domains, in this work, we propose jointly training on available spoken DST data and written textual data from other domains as a way to achieve cross-domain generalization. We conduct experiments which show the efficacy of our proposed method for getting good cross-domain DST performance without relying on spoken training data from the target domains.


【6】GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
标题:GLA-Grad++:一种改进的Griffin-Lim引导的语音合成扩散模型
链接:https://arxiv.org/abs/2511.22293

作者:Teysir Baoueb,Xiaoyu Bie,Mathieu Fontaine,Gaël Richard
摘要:扩散模型的最新进展使其成为语音合成的强大生成框架,在音频质量和稳定性方面有了实质性的改进。然而,它们在以梅尔频谱图为条件的声码器中的有效性仍然受到限制,特别是当条件偏离训练分布时。最近提出的GLA-Grad模型引入了对WaveGrad声码器的相位感知扩展,该扩展将Griffin-Lim算法(GLA)集成到反向过程中,以减少生成的信号和调节梅尔频谱图之间的不一致。在本文中,我们进一步改善GLA-Grad通过创新的选择,在如何应用校正。特别是,我们计算的校正项只有一次,与一个单一的应用程序的GLA,以加速生成过程。实验结果表明,我们的方法始终优于基线模型,特别是在域外的情况下。
摘要:Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders conditioned on mel spectrograms remains constrained, particularly when the conditioning diverges from the training distribution. The recently proposed GLA-Grad model introduced a phase-aware extension to the WaveGrad vocoder that integrated the Griffin-Lim algorithm (GLA) into the reverse process to reduce inconsistencies between generated signals and conditioning mel spectrogram. In this paper, we further improve GLA-Grad through an innovative choice in how to apply the correction. Particularly, we compute the correction term only once, with a single application of GLA, to accelerate the generation process. Experimental results demonstrate that our method consistently outperforms the baseline models, particularly in out-of-domain scenarios.


【7】Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
标题:利用深生成模型推进海洋生物声学:南方常驻虎鲸检测的混合增强策略
链接:https://arxiv.org/abs/2511.21872

作者:Bruno Padovese,Fabio Frazao,Michael Dowd,Ruth Joy
备注:16 pages, 6 Figures, 2 Tables, submitted to Marine Mammal Science as part of a special issue on Machine Learning and Artificial Intelligence in Marine Mammal Research
摘要:海洋哺乳动物发声的自动检测和分类对于保护和管理工作至关重要,但受到有限的注释数据集和现实世界海洋环境声学复杂性的阻碍。数据扩充已被证明是一种有效的策略,通过增加数据集的多样性和提高模型的泛化能力,而不需要额外的现场数据来解决这一限制。然而,迄今为止使用的大多数增强技术都依赖于有效但相对简单的转换,这就留下了一个问题,即深度生成模型是否可以提供额外的好处。在这项研究中,我们评估了深度生成在海洋哺乳动物呼叫检测中的数据增强潜力,包括:变分自编码器,生成对抗网络和去噪扩散概率模型。使用南方居民虎鲸(Orcinus虎鲸)发声从两个长期的水听器部署在萨利希海,我们比较这些方法对传统的增强方法,如时间转移和发声掩蔽。虽然所有生成方法相对于基线都提高了分类性能,但基于扩散的增强产生了最高的召回率(0.87)和总体F1分数(0.75)。将基于生成的合成与传统方法相结合的混合策略实现了最佳的整体性能,F1得分为0.81。我们希望这项研究鼓励进一步探索深层生成模型作为补充增强策略,以推进对受威胁海洋哺乳动物种群的声学监测。
摘要:Automated detection and classification of marine mammals vocalizations is critical for conservation and management efforts but is hindered by limited annotated datasets and the acoustic complexity of real-world marine environments. Data augmentation has proven to be an effective strategy to address this limitation by increasing dataset diversity and improving model generalization without requiring additional field data. However, most augmentation techniques used to date rely on effective but relatively simple transformations, leaving open the question of whether deep generative models can provide additional benefits. In this study, we evaluate the potential of deep generative for data augmentation in marine mammal call detection including: Variational Autoencoders, Generative Adversarial Networks, and Denoising Diffusion Probabilistic Models. Using Southern Resident Killer Whale (Orcinus orca) vocalizations from two long-term hydrophone deployments in the Salish Sea, we compare these approaches against traditional augmentation methods such as time-shifting and vocalization masking. While all generative approaches improved classification performance relative to the baseline, diffusion-based augmentation yielded the highest recall (0.87) and overall F1-score (0.75). A hybrid strategy combining generative-based synthesis with traditional methods achieved the best overall performance with an F1-score of 0.81. We hope this study encourages further exploration of deep generative models as complementary augmentation strategies to advance acoustic monitoring of threatened marine mammal populations.


【8】3MDiT: Unified Tri-Modal Diffusion Transformer for Text-Driven Synchronized Audio-Video Generation
标题:3MDiT:用于文本驱动同步音视频生成的统一三模扩散Transformer
链接:https://arxiv.org/abs/2511.21780

作者:Yaoru Li,Heyu Si,Federico Landi,Pilar Oplustil Gallegos,Ioannis Koutsoumpas,O. Ricardo Cortez Vazquez,Ruiju Fu,Qi Guo,Xin Jin,Shunyu Liu,Mingli Song
摘要:文本到视频(T2 V)扩散模型最近取得了令人印象深刻的视觉质量,但大多数系统仍然生成无声的剪辑,并将音频视为次要问题。现有的音频-视频生成流水线通常将任务分解为级联阶段,这些级联阶段跨模态累积误差,并在单独的目标下进行训练。最近的联合音频-视频生成器缓解了这个问题,但通常依赖于具有临时跨模态桥和静态单次文本调节的双塔架构,这使得很难重用T2 V主干,也很难理解音频、视频和语言如何随着时间的推移进行交互。为了解决这些挑战,我们提出了3 MDiT,一个统一的三模式扩散Transformer文本驱动的同步音频视频生成。我们的框架模型视频,音频和文本作为共同发展的流:一个同构的音频分支反映了T2 V的骨干,三模态全向块执行功能级融合的三种形式,和一个可选的动态文本调节机制更新的文本表示作为音频和视频证据共同发展。该设计支持两种机制:从头开始对音频视频数据进行训练,以及在不修改其主干的情况下正交调整预训练的T2 V模型。实验表明,我们的方法生成高质量的视频和逼真的音频,同时不断提高音视频同步和三模态对齐在一系列的定量指标。
摘要:Text-to-video (T2V) diffusion models have recently achieved impressive visual quality, yet most systems still generate silent clips and treat audio as a secondary concern. Existing audio-video generation pipelines typically decompose the task into cascaded stages, which accumulate errors across modalities and are trained under separate objectives. Recent joint audio-video generators alleviate this issue but often rely on dual-tower architectures with ad-hoc cross-modal bridges and static, single-shot text conditioning, making it difficult to both reuse T2V backbones and to reason about how audio, video and language interact over time. To address these challenges, we propose 3MDiT, a unified tri-modal diffusion transformer for text-driven synchronized audio-video generation. Our framework models video, audio and text as jointly evolving streams: an isomorphic audio branch mirrors a T2V backbone, tri-modal omni-blocks perform feature-level fusion across the three modalities, and an optional dynamic text conditioning mechanism updates the text representation as audio and video evidence co-evolve. The design supports two regimes: training from scratch on audio-video data, and orthogonally adapting a pretrained T2V model without modifying its backbone. Experiments show that our approach generates high-quality videos and realistic audio while consistently improving audio-video synchronization and tri-modal alignment across a range of quantitative metrics.


【9】On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
标题:预训练基于wav 2vec 2的模型的跨语言可移植性
链接:https://arxiv.org/abs/2511.21704

作者:Jonatas Grosman,Cassio Almeida,Guilherme Schardong,Hélio Lopes
摘要:使用大型预训练模型提供的表示已成为在各种任务中实现最先进结果的主要策略。最近提出的一个大型预训练模型wav2vec 2.0对其他几个在语音数据上预训练大型模型的工作具有开创性意义。许多模型使用与wav2vec 2.0相同的架构进行预训练,并在各种语音相关任务中获得最先进的技术。先前的工作已经证明,在这些基于wav2vec2的模型的预训练期间使用的数据可以影响模型在下游任务中的性能,并且在使用这些模型之前应该考虑到这一点。然而,很少有人提出进一步研究这些预训练模型的知识转移如何在不同的语言中表现,即使目标语言与模型预训练期间使用的语言不同。我们的工作旨在研究这些基于wav2vec2的模型的跨语言可移植性。我们使用15个大型预训练模型对18种语言的语音识别任务进行了几次微调实验。我们的实验结果表明,在这些模型的预训练过程中使用的数据大小对最终性能的影响不如多样性那么重要。我们注意到,在评估的模型中,印欧语言的性能优于非印欧语言。我们使用单语模型观察到了积极的跨语言知识转移,这在我们使用的所有语言中都很明显,但当预训练期间使用的语言与下游任务语言更相似时,这种情况更加明显。有了这些发现,我们的目标是帮助科学界利用现有的基于wav2vec2的预训练模型,并促进新模型的预训练。
摘要:Using representations provided by a large pre-trained model has become the primary strategy for achieving state-of-the-art results in a wide range of tasks. A recently proposed large pre-trained model, wav2vec 2.0, was seminal for several other works on pre-training large models on speech data. Many models are being pre-trained using the same architecture as wav2vec 2.0 and are getting state-of-the-art in various speech-related tasks. Previous work has demonstrated that the data used during the pre-training of these wav2vec2-based models can impact the model's performance in downstream tasks, and this should be taken into consideration before utilizing these models. However, few works have proposed investigating further how the transfer knowledge of these pre-trained models behaves in different languages, even when the target language differs from the one used during the model's pre-training. Our work aims to investigate the cross-lingual transferability of these wav2vec2-based models. We performed several fine-tuning experiments on the speech recognition task in 18 languages using 15 large pre-trained models. The results of our experiments showed us that the size of data used during the pre-training of these models is not as important to the final performance as the diversity. We noticed that the performance of Indo-European languages is superior to non-Indo-European languages in the evaluated models. We have observed a positive cross-lingual transfer of knowledge using monolingual models, which was evident in all the languages we used, but more pronounced when the language used during pre-training was more similar to the downstream task language. With these findings, we aim to assist the scientific community in utilizing existing wav2vec2-based pre-trained models, as well as facilitate the pre-training of new ones.


eess.AS音频处理


【1】Group-Aware Partial Model Merging for Children's Automatic Speech Recognition
标题:儿童自动语音识别的群体感知部分模型合并
链接:https://arxiv.org/abs/2511.23098

作者:Thomas Rolland,Alberto Abad
备注:IEEE ASRU 2025 Workshop AI4CSL
摘要:儿童的自动语音识别(ASR)仍然具有挑战性,主要是由于大的声学可变性和有限的训练数据。虽然成人预训练模型的监督微调已经显示出希望,但它往往无法捕捉儿童群体特定的特征变化。为了解决这个问题,我们引入了组感知PARtial模型合并(GRAPAM),这是一种参数高效的方法,它结合了无监督聚类,部分微调和模型合并。我们的方法首先根据声学相似性对儿童的数据进行分组,从而使成人预先训练的模型适应儿童。每组用于部分微调成人预训练模型,并在参数级别合并所得模型。在MyST儿童语音语料库上进行的实验表明,GRAPAM在使用相同数据量的情况下实现了6%的字错误率(WER)的相对改善,在训练较少参数的情况下优于完全微调。这些结果突出了模型合并作为儿童ASR的可扩展和有效的策略的承诺。
摘要:Automatic Speech Recognition (ASR) for children remains challenging, primarily due to large acoustic variability and limited availability of training data. While supervised fine-tuning of adult pre-trained models has shown promise, it often fails to capture group-specific characteristics variations among children. To address this, we introduce GRoup-Aware PARtial model Merging (GRAPAM), a parameter-efficient approach that combines unsupervised clustering, partial fine-tuning, and model merging. Our approach adapts adult-pre-trained models to children by first grouping the children's data based on acoustic similarity. Each group is used to partially fine-tune an adult pre-trained model, and the resulting models are merged at the parameter level. Experiments conducted on the MyST children's speech corpus indicate that GRAPAM achieves a relative improvement of 6% of Word Error Rate (WER), using the same amount of data, outperforming full fine-tuning while training fewer parameters. These results highlight the promise of model merging as a scalable and effective strategy for children's ASR.


【2】PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
标题:PURE Codec:语音编解码器学习的剩余信息的渐进展开
链接:https://arxiv.org/abs/2511.22687

作者:Jiatong Shi,Haoran Wang,William Chen,Chenda Li,Wangyou Zhang,Jinchuan Tian,Shinji Watanabe
备注:Accepted by ASRU2025
摘要:神经语音编解码器在低比特率压缩方面取得了很好的性能,但残差矢量量化(RVQ)通常存在训练不稳定和分解无效的问题,限制了重建质量和效率。我们提出了PURE编解码器(残差熵的渐进展开),这是一种新的框架,它使用预先训练的语音增强模型来指导多级量化。第一个量化阶段重建低熵,去噪语音嵌入,而后续阶段编码残留的高熵成分。该设计显著提高了训练稳定性。实验表明,PURE在重建和下游基于语音语言模型的文本到语音转换中始终优于传统的基于RVQ的编解码器,特别是在嘈杂的训练条件下。
摘要:Neural speech codecs have achieved strong performance in low-bitrate compression, but residual vector quantization (RVQ) often suffers from unstable training and ineffective decomposition, limiting reconstruction quality and efficiency. We propose PURE Codec (Progressive Unfolding of Residual Entropy), a novel framework that guides multi-stage quantization using a pre-trained speech enhancement model. The first quantization stage reconstructs low-entropy, denoised speech embeddings, while subsequent stages encode residual high-entropy components. This design improves training stability significantly. Experiments demonstrate that PURE consistently outperforms conventional RVQ-based codecs in reconstruction and downstream speech language model-based text-to-speech, particularly under noisy training conditions.


【3】Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
标题:基于LLM的端到端口语对话状态跟踪的联合语音和文本训练
链接:https://arxiv.org/abs/2511.22503

作者:Katia Vendrame,Bolaji Yusuf,Santosh Kesiraju,Šimon Sedláček,Oldřich Plchot,Jan Černocký
备注:submitted to ICASSP 2026
摘要:端到端的口语对话状态跟踪(DST)由于必须处理语音输入和数据稀缺而变得困难。在最近的工作中,已经提出了将语音基础编码器和大型语言模型相结合,以减轻这种困难。虽然这种方法已被证明可以产生强大的口语DST模型,在现实的多回合DST中实现最先进的性能,但它很难跨领域进行概括,并且需要为每个感兴趣的领域提供注释的口语DST训练数据。然而,为每个目标领域收集此类数据既昂贵又困难。注意到文本DST数据更容易获得各个领域,在这项工作中,我们建议联合训练可用的口语DST数据和来自其他领域的书面文本数据,作为实现跨领域泛化的一种方式。我们进行的实验表明,我们提出的方法获得良好的跨域DST性能,而不依赖于口语训练数据的目标领域的有效性。
摘要:End-to-end spoken dialogue state tracking (DST) is made difficult by the tandem of having to handle speech input and data scarcity. Combining speech foundation encoders and large language models has been proposed in recent work as to alleviate some of this difficulty. Although this approach has been shown to result in strong spoken DST models, achieving state-of-the-art performance in realistic multi-turn DST, it struggles to generalize across domains and requires annotated spoken DST training data for each domain of interest. However, collecting such data for every target domain is both costly and difficult. Noting that textual DST data is more easily obtained for various domains, in this work, we propose jointly training on available spoken DST data and written textual data from other domains as a way to achieve cross-domain generalization. We conduct experiments which show the efficacy of our proposed method for getting good cross-domain DST performance without relying on spoken training data from the target domains.


【4】GLA-Grad++: An Improved Griffin-Lim Guided Diffusion Model for Speech Synthesis
标题:GLA-Grad++:一种改进的Griffin-Lim引导的语音合成扩散模型
链接:https://arxiv.org/abs/2511.22293

作者:Teysir Baoueb,Xiaoyu Bie,Mathieu Fontaine,Gaël Richard
摘要:扩散模型的最新进展使其成为语音合成的强大生成框架,在音频质量和稳定性方面有了实质性的改进。然而,它们在以梅尔频谱图为条件的声码器中的有效性仍然受到限制,特别是当条件偏离训练分布时。最近提出的GLA-Grad模型引入了对WaveGrad声码器的相位感知扩展,该扩展将Griffin-Lim算法(GLA)集成到反向过程中,以减少生成的信号和调节梅尔频谱图之间的不一致。在本文中,我们进一步改善GLA-Grad通过创新的选择,在如何应用校正。特别是,我们计算的校正项只有一次,与一个单一的应用程序的GLA,以加速生成过程。实验结果表明,我们的方法始终优于基线模型,特别是在域外的情况下。
摘要:Recent advances in diffusion models have positioned them as powerful generative frameworks for speech synthesis, demonstrating substantial improvements in audio quality and stability. Nevertheless, their effectiveness in vocoders conditioned on mel spectrograms remains constrained, particularly when the conditioning diverges from the training distribution. The recently proposed GLA-Grad model introduced a phase-aware extension to the WaveGrad vocoder that integrated the Griffin-Lim algorithm (GLA) into the reverse process to reduce inconsistencies between generated signals and conditioning mel spectrogram. In this paper, we further improve GLA-Grad through an innovative choice in how to apply the correction. Particularly, we compute the correction term only once, with a single application of GLA, to accelerate the generation process. Experimental results demonstrate that our method consistently outperforms the baseline models, particularly in out-of-domain scenarios.


【5】Advancing Marine Bioacoustics with Deep Generative Models: A Hybrid Augmentation Strategy for Southern Resident Killer Whale Detection
标题:利用深生成模型推进海洋生物声学:南方常驻虎鲸检测的混合增强策略
链接:https://arxiv.org/abs/2511.21872

作者:Bruno Padovese,Fabio Frazao,Michael Dowd,Ruth Joy
备注:16 pages, 6 Figures, 2 Tables, submitted to Marine Mammal Science as part of a special issue on Machine Learning and Artificial Intelligence in Marine Mammal Research
摘要:海洋哺乳动物发声的自动检测和分类对于保护和管理工作至关重要,但受到有限的注释数据集和现实世界海洋环境声学复杂性的阻碍。数据扩充已被证明是一种有效的策略,通过增加数据集的多样性和提高模型的泛化能力,而不需要额外的现场数据来解决这一限制。然而,迄今为止使用的大多数增强技术都依赖于有效但相对简单的转换,这就留下了一个问题,即深度生成模型是否可以提供额外的好处。在这项研究中,我们评估了深度生成在海洋哺乳动物呼叫检测中的数据增强潜力,包括:变分自编码器,生成对抗网络和去噪扩散概率模型。使用南方居民虎鲸(Orcinus虎鲸)发声从两个长期的水听器部署在萨利希海,我们比较这些方法对传统的增强方法,如时间转移和发声掩蔽。虽然所有生成方法相对于基线都提高了分类性能,但基于扩散的增强产生了最高的召回率(0.87)和总体F1分数(0.75)。将基于生成的合成与传统方法相结合的混合策略实现了最佳的整体性能,F1得分为0.81。我们希望这项研究鼓励进一步探索深层生成模型作为补充增强策略,以推进对受威胁海洋哺乳动物种群的声学监测。
摘要:Automated detection and classification of marine mammals vocalizations is critical for conservation and management efforts but is hindered by limited annotated datasets and the acoustic complexity of real-world marine environments. Data augmentation has proven to be an effective strategy to address this limitation by increasing dataset diversity and improving model generalization without requiring additional field data. However, most augmentation techniques used to date rely on effective but relatively simple transformations, leaving open the question of whether deep generative models can provide additional benefits. In this study, we evaluate the potential of deep generative for data augmentation in marine mammal call detection including: Variational Autoencoders, Generative Adversarial Networks, and Denoising Diffusion Probabilistic Models. Using Southern Resident Killer Whale (Orcinus orca) vocalizations from two long-term hydrophone deployments in the Salish Sea, we compare these approaches against traditional augmentation methods such as time-shifting and vocalization masking. While all generative approaches improved classification performance relative to the baseline, diffusion-based augmentation yielded the highest recall (0.87) and overall F1-score (0.75). A hybrid strategy combining generative-based synthesis with traditional methods achieved the best overall performance with an F1-score of 0.81. We hope this study encourages further exploration of deep generative models as complementary augmentation strategies to advance acoustic monitoring of threatened marine mammal populations.


机器翻译由腾讯交互翻译提供,仅供参考