微信公众号:arXiv_Daily
cs.SD语音
【1】DIFFA: Large Language Diffusion Models Can Listen and Understand
标题:DIFFA:大型语言扩散模型可以听和理解
链接:http://arxiv.org/pdf/2507.18452v1
摘要:大型语言模型(LLM)的最新进展显示出跨文本和多模态领域的卓越能力。同时,基于扩散的语言模型已经成为自回归范式的一个有前途的替代方案,提供了改进的可控性,双向上下文建模和鲁棒的生成。然而,它们在音频模态中的应用仍然没有得到充分探索。在这项工作中,我们介绍了 textbf{DIFFA},第一个基于扩散的大型音频语言模型,旨在执行口语理解。DIFFA将冻结扩散语言模型与轻量级双适配器架构集成在一起,该架构连接了语音理解和自然语言推理。我们采用了两个阶段的训练管道:首先,通过ASR目标对齐语义表示;然后,通过提示LLM自动生成的合成音频-字幕对学习注意力跟随能力。尽管只接受了960小时的ASR和127小时的合成指令数据的训练,但DIFFA在包括MMSU、MMAU和VoiceBench在内的主要基准测试中表现出了竞争力,优于几个自回归开源基线。我们的研究结果揭示了基于扩散的语言模型在高效和可扩展的音频理解方面的潜力,为语音驱动的AI开辟了新的方向。我们的代码将在https: github.com NKU-HLT DIFFA.git上提供。
摘要:Recent advances in Large language models (LLMs) have shown remarkable capabilities across textual and multimodal domains. In parallel, diffusion-based language models have emerged as a promising alternative to the autoregressive paradigm, offering improved controllability, bidirectional context modeling, and robust generation. However, their application to the audio modality remains underexplored. In this work, we introduce textbf{DIFFA}, the first diffusion-based Large Audio-Language Model designed to perform spoken language understanding. DIFFA integrates a frozen diffusion language model with a lightweight dual-adapter architecture that bridges speech understanding and natural language reasoning. We employ a two-stage training pipeline: first, aligning semantic representations via an ASR objective; then, learning instruction-following abilities through synthetic audio-caption pairs automatically generated by prompting LLMs. Despite being trained on only 960 hours of ASR and 127 hours of synthetic instruction data, DIFFA demonstrates competitive performance on major benchmarks, including MMSU, MMAU, and VoiceBench, outperforming several autoregressive open-source baselines. Our results reveal the potential of diffusion-based language models for efficient and scalable audio understanding, opening a new direction for speech-driven AI. Our code will be available at https: github.com NKU-HLT DIFFA.git.
【2】Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
标题:微型还不够小:通过混合知识提炼高质量、低资源的面部动画模型
链接:https://arxiv.org/abs/2507.18352
备注:Accepted to ACM Transactions on Graphics 2025 (SIGGRAPH journal track)
摘要:为语音驱动的3D面部动画训练高质量、鲁棒的机器学习模型需要大量不同的高质量音频动画对数据集。为了克服这种数据集的缺乏,最近的工作引入了大型的预训练语音编码器,这些编码器对输入音频的变化具有鲁棒性,因此,使面部动画模型能够跨扬声器,音频质量和语言进行概括。然而,由此产生的面部动画模型是非常大的,并且只适合在专用机器上进行离线推理。在这项工作中,我们将探讨在游戏开发的背景下,设备上的实时面部动画模型。我们克服了缺乏大型数据集,使用混合知识蒸馏与伪标签。给定一个大的音频数据集,我们采用高性能的教师模型来训练非常小的学生模型。与预先训练的语音编码器相比,我们的学生模型仅由卷积层和全连接层组成,不需要注意上下文或定期更新。在我们的实验中,我们证明,我们可以减少内存占用高达3.4 MB,并要求未来的音频上下文高达81毫秒,同时保持高品质的动画。这为设备上的推理铺平了道路,这是迈向逼真的模型驱动数字角色的重要一步。
摘要:The training of high-quality, robust machine learning models for speech-driven 3D facial animation requires a large, diverse dataset of high-quality audio-animation pairs. To overcome the lack of such a dataset, recent work has introduced large pre-trained speech encoders that are robust to variations in the input audio and, therefore, enable the facial animation model to generalize across speakers, audio quality, and languages. However, the resulting facial animation models are prohibitively large and lend themselves only to offline inference on a dedicated machine. In this work, we explore on-device, real-time facial animation models in the context of game development. We overcome the lack of large datasets by using hybrid knowledge distillation with pseudo-labeling. Given a large audio dataset, we employ a high-performing teacher model to train very small student models. In contrast to the pre-trained speech encoders, our student models only consist of convolutional and fully-connected layers, removing the need for attention context or recurrent updates. In our experiments, we demonstrate that we can reduce the memory footprint to up to 3.4 MB and required future audio context to up to 81 ms while maintaining high-quality animations. This paves the way for on-device inference, an important step towards realistic, model-driven digital characters.
【3】Improving Bird Classification with Primary Color Additives
标题:用原色添加剂改善鸟类分类
链接:https://arxiv.org/abs/2507.18334
备注:5 pages (Accepted to Interspeech 2025)
摘要:我们解决的问题,鸟类的分类使用他们的歌曲录音,一个具有挑战性的任务,由于环境噪声,重叠发声,和丢失的标签。现有的模型与低信噪比或多物种记录的斗争。我们假设,鸟类可以通过可视化他们的音高模式,速度和重复,统称为图案进行分类。应用于光谱图图像的深度学习模型有所帮助,但不同物种之间的相似图案会导致混淆。为了减轻这种情况,我们使用原色添加剂将频率信息嵌入到频谱图中。这增强了物种区分并提高了分类准确性。我们的实验表明,所提出的方法在没有着色的模型上实现了统计学上的显著增益,并超过了BirdCLEF 2024的获胜者,将F1提高了7.3%,ROC-AUC提高了6.2%,CMAP提高了6.6%。这些结果证明了通过彩色化结合频率信息的有效性。
摘要:We address the problem of classifying bird species using their song recordings, a challenging task due to environmental noise, overlapping vocalizations, and missing labels. Existing models struggle with low-SNR or multi-species recordings. We hypothesize that birds can be classified by visualizing their pitch pattern, speed, and repetition, collectively called motifs. Deep learning models applied to spectrogram images help, but similar motifs across species cause confusion. To mitigate this, we embed frequency information into spectrograms using primary color additives. This enhances species distinction and improves classification accuracy. Our experiments show that the proposed approach achieves statistically significant gains over models without colorization and surpasses the BirdCLEF 2024 winner, improving F1 by 7.3%, ROC-AUC by 6.2%, and CMAP by 6.6%. These results demonstrate the effectiveness of incorporating frequency information via colorization.
【4】GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
标题:GOAT-LAM:具有副语言和说话者特征意识的口语模型
链接:https://arxiv.org/abs/2507.18119
摘要:端到端口语模型(SLM)的最新进展显着提高了人工智能系统参与自然口语交互的能力。然而,大多数现有的模型只把语音作为语言内容的载体,往往忽略了丰富的语言和说话者特征线索嵌入在人类语音,如方言,年龄,情感和非语音发声。在这项工作中,我们介绍了GOAT-SLM,一种新的口语模型,具有语言学和说话者特征意识,旨在扩展口语建模超越文本语义。GOAT-SLM采用双模态头部架构,将语言建模与声学实现相结合,实现了强大的语言理解,同时支持表达性和自适应语音生成。为了提高模型的效率和通用性,我们提出了一种模块化的分阶段训练策略,该策略使用大规模语音文本语料库逐步调整语言,语言学和说话人特征信息。TELEVAL上的实验结果表明,GOAT-SLM在语义和非语义任务上实现了良好的平衡性能,在处理情感,方言变化和年龄敏感的交互方面优于现有的开源模型。这项工作强调了建模超越语言内容的重要性,并推动了更自然,适应性和社会意识的口语系统的发展。
摘要:Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing models treat speech merely as a vehicle for linguistic content, often overlooking the rich paralinguistic and speaker characteristic cues embedded in human speech, such as dialect, age, emotion, and non-speech vocalizations. In this work, we introduce GOAT-SLM, a novel spoken language model with paralinguistic and speaker characteristic awareness, designed to extend spoken language modeling beyond text semantics. GOAT-SLM adopts a dual-modality head architecture that decouples linguistic modeling from acoustic realization, enabling robust language understanding while supporting expressive and adaptive speech generation. To enhance model efficiency and versatility, we propose a modular, staged training strategy that progressively aligns linguistic, paralinguistic, and speaker characteristic information using large-scale speech-text corpora. Experimental results on TELEVAL, a multi-dimensional evaluation benchmark, demonstrate that GOAT-SLM achieves well-balanced performance across both semantic and non-semantic tasks, and outperforms existing open-source models in handling emotion, dialectal variation, and age-sensitive interactions. This work highlights the importance of modeling beyond linguistic content and advances the development of more natural, adaptive, and socially aware spoken language systems.
【5】TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
标题:TELEVAL:汉语交互场景中口语模型的动态基准
链接:https://arxiv.org/abs/2507.18061
摘要:近年来,口语模型(SLM)取得了快速的进展,并开发了许多用于评估其性能的基准。然而,大多数现有的基准主要集中在评估SLM是否可以执行与大型语言模型(LLM)处理的复杂任务相当的复杂任务,通常无法与用户在现实世界的会话场景中自然交互的方式保持一致。在本文中,我们提出了TELEVAL,一个动态的基准,专门设计来评估SLM的有效性,在现实的中国互动设置的会话代理。TELEVAL定义了三个评估维度:显式语义、副语言和隐式语义以及系统能力。它采用与现实世界使用一致的对话格式,并分别评估文本和音频输出。TELEVAL特别关注模型从用户语音中提取隐含线索并在没有额外指令的情况下做出适当反应的能力。我们的实验表明,尽管最近的进展,现有的SLM仍然有相当大的改进空间,在自然的会话任务。我们希望TELEVAL可以作为一个以用户为中心的评估框架,直接反映用户体验,并有助于开发更有能力的对话导向的SLM。
摘要:Spoken language models (SLMs) have seen rapid progress in recent years, along with the development of numerous benchmarks for evaluating their performance. However, most existing benchmarks primarily focus on evaluating whether SLMs can perform complex tasks comparable to those tackled by large language models (LLMs), often failing to align with how users naturally interact in real-world conversational scenarios. In this paper, we propose TELEVAL, a dynamic benchmark specifically designed to evaluate SLMs' effectiveness as conversational agents in realistic Chinese interactive settings. TELEVAL defines three evaluation dimensions: Explicit Semantics, Paralinguistic and Implicit Semantics, and System Abilities. It adopts a dialogue format consistent with real-world usage and evaluates text and audio outputs separately. TELEVAL particularly focuses on the model's ability to extract implicit cues from user speech and respond appropriately without additional instructions. Our experiments demonstrate that despite recent progress, existing SLMs still have considerable room for improvement in natural conversational tasks. We hope that TELEVAL can serve as a user-centered evaluation framework that directly reflects the user experience and contributes to the development of more capable dialogue-oriented SLMs.
【6】The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
标题:MLC-SL 2025挑战赛中用于多语言对话语音识别和语音规模化的TEA-ASLP系统
链接:https://arxiv.org/abs/2507.18051
备注:Interspeech 2025 workshop
摘要:本文介绍了TEA-ASLP提交给MLC-SLM 2025挑战赛的系统,在任务I中解决了多语言会话自动语音识别(ASR),在任务II中解决了语音日志化ASR。对于任务I,我们通过集成已知的语言识别和多语言MoE LoRA结构来增强Ideal-LLM模型,同时使用CTC预测的令牌作为提示来改进自回归生成。该模型基于大约18万小时的多语言ASR数据进行训练。在任务II中,我们用一个更合适的纯英语版本替换了基线的英汉说话者日记模型。与基线语音语言模型相比,我们的方法实现了30.8%的单词错误率(WER)降低,导致任务I中的最终WER为9.60%,任务II中的时间约束最小排列WER为17.49%,在各自的挑战任务中获得第一名和第二名。
摘要:This paper presents the TEA-ASLP's system submitted to the MLC-SLM 2025 Challenge, addressing multilingual conversational automatic speech recognition (ASR) in Task I and speech diarization ASR in Task II. For Task I, we enhance Ideal-LLM model by integrating known language identification and a multilingual MOE LoRA structure, along with using CTC-predicted tokens as prompts to improve autoregressive generation. The model is trained on approximately 180k hours of multilingual ASR data. In Task II, we replace the baseline English-Chinese speaker diarization model with a more suitable English-only version. Our approach achieves a 30.8% reduction in word error rate (WER) compared to the baseline speech language model, resulting in a final WER of 9.60% in Task I and a time-constrained minimum-permutation WER of 17.49% in Task II, earning first and second place in the respective challenge tasks.
【7】Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
标题:用于声音事件定位、检测和距离估计的具有共享权重和注意力机制的Resnet-conformer网络
链接:https://arxiv.org/abs/2507.17941
备注:This paper has been submitted as a technical report outlining our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 and can be found in DCASE2024 technical reports
摘要:本技术报告概述了我们对2024年声学场景和事件的检测和分类(DCASE)任务3A的方法,重点是声音事件定位和检测(SELD)。SELD通过估计声音事件定位和检测提供有价值的见解,帮助各种机器认知任务,如环境推理,导航和其他声音定位相关的应用。今年的挑战赛对使用纯音频(音轨A)或视听(音轨B)输入的真实声音场景的注释录音进行评估。今年的一个显著变化是采用了距离估计,并相应调整了评价指标,以进行全面评估。我们提交的是挑战赛的任务A,重点是音频轨道。我们的方法利用对数梅尔光谱图,强度矢量,并采用多个数据增强。我们提出了一种基于EINV2的网络架构[1],取得了改进的结果:在开发数据集[2,3]的测试集上,F分数为40.2%,角度误差(DOA)为17.7度,相对距离误差(RDE)为0.32。
摘要:This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating sound event localization and detection, aiding in various machine cognition tasks such as environmental inference, navigation, and other sound localization-related applications. This year's challenge evaluates models using either audio-only (Track A) or audiovisual (Track B) inputs on annotated recordings of real sound scenes. A notable change this year is the introduction of distance estimation, with evaluation metrics adjusted accordingly for a comprehensive assessment. Our submission is for Task A of the Challenge, which focuses on the audio-only track. Our approach utilizes log-mel spectrograms, intensity vectors, and employs multiple data augmentations. We proposed an EINV2-based [1] network architecture, achieving improved results: an F-score of 40.2%, Angular Error (DOA) of 17.7 degrees, and Relative Distance Error (RDE) of 0.32 on the test set of the Development Dataset [2 ,3].
【8】Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
标题:鲍勃的五彩纸屑:音乐和视频生成中的语音同步攻击
链接:https://arxiv.org/abs/2507.17937
摘要:歌词到歌曲(LS 2)生成模型承诺从文本进行端到端音乐合成,但它们在训练数据记忆方面的脆弱性仍然没有得到充分研究。我们介绍了对抗性音素替换(APT),这是一种新颖的攻击,在这种攻击中,歌词在语义上被改变,同时通过谐音替换(例如,阿姆的著名的“妈妈的意大利面”$\rightarrow$“鲍勃的五彩纸屑”)。尽管存在这些扭曲,但我们发现了一种强大的子词汇记忆形式:像SUNO和YuE这样的模型重新生成的输出与已知的训练内容惊人地相似,在音频域指标(包括CLAP,AudioJudge和CoverID)之间实现了高度相似性。此漏洞在多种语言和流派中持续存在。更令人惊讶的是,我们发现,在文本到视频模型中,仅音素改变的歌词就可以触发视觉记忆。当提示与语音修改歌词从失去自己,Veo 3重建视觉元素从原来的音乐视频-包括字符外观和场景组成-尽管没有视觉线索的提示。我们称这种现象为语音-视觉回流。总之,这些研究结果暴露了一个关键的脆弱性,在成绩单条件下的多模态生成:语音提示单独可以解锁记忆的视听内容,提出了迫切的问题,在现代生成系统的版权,安全和内容出处。示例生成可以在我们的演示页面(jrohsc.github.io/music_attack/)上找到。
摘要:Lyrics-to-Song (LS2) generation models promise end-to-end music synthesis from text, yet their vulnerability to training data memorization remains underexplored. We introduce Adversarial PhoneTic Prompting (APT), a novel attack where lyrics are semantically altered while preserving their acoustic structure through homophonic substitutions (e.g., Eminem's famous "mom's spaghetti" $\rightarrow$ "Bob's confetti"). Despite these distortions, we uncover a powerful form of sub-lexical memorization: models like SUNO and YuE regenerate outputs strikingly similar to known training content, achieving high similarity across audio-domain metrics, including CLAP, AudioJudge, and CoverID. This vulnerability persists across multiple languages and genres. More surprisingly, we discover that phoneme-altered lyrics alone can trigger visual memorization in text-to-video models. When prompted with phonetically modified lyrics from Lose Yourself, Veo 3 reconstructs visual elements from the original music video -- including character appearance and scene composition -- despite no visual cues in the prompt. We term this phenomenon phonetic-to-visual regurgitation. Together, these findings expose a critical vulnerability in transcript-conditioned multimodal generation: phonetic prompting alone can unlock memorized audiovisual content, raising urgent questions about copyright, safety, and content provenance in modern generative systems. Example generations are available on our demo page (jrohsc.github.io/music_attack/).
【9】Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
标题:基于可解释性的语音预训练模型的说话人去纠缠
链接:https://arxiv.org/abs/2507.17851
备注:20 pages, 9 figures, 2 tables
摘要:语音预训练模型包含跨不同层的特定于任务的信息,但解耦内容和音色信息仍然具有挑战性,因为删除特定于说话者的信息通常会导致内容丢失。目前的研究缺乏直接的指标来量化模型编码中的音色残留,依赖于通过下游任务的间接评估。本文通过在语音预训练模型中基于可解释性的说话人解纠缠来解决这些挑战。我们定量评估模型嵌入的音色残留,并使用解释性表示提高扬声器解纠缠。我们的贡献包括:(1)InterpTRQE-SptME基准-一个使用可解释性的音色残留识别框架。该基准将内容嵌入与音色嵌入连接起来进行说话人分类,然后应用Gradient SHAP Explainer量化音色残差。我们评估了七个语音预训练模型的变化。(2)InterpTF-SptME方法-使用SHAP噪声和SHAP裁剪技术的基于可解释性的音色过滤方法。这种与模型无关的方法转换中间编码以在保留内容的同时去除音色。使用HuBERT LARGE在VCTK数据集上进行的实验证明了成功的内容保留和显着的扬声器解纠缠优化。结果表明,SHAP噪声方法可以将音色残留从18.05%降低到接近0%,同时保持内容完整性,有助于增强与内容相关的语音处理任务的性能,并防止音色隐私泄露。
摘要:Speech pretrained models contain task-specific information across different layers, but decoupling content and timbre information remains challenging as removing speaker-specific information often causes content loss. Current research lacks direct metrics to quantify timbre residual in model encodings, relying on indirect evaluation through downstream tasks. This paper addresses these challenges through interpretability-based speaker disentanglement in speech pretraining models. We quantitatively evaluate timbre residual in model embeddings and improve speaker disentanglement using interpretive representations. Our contributions include: (1) InterpTRQE-SptME Benchmark - a timbre residual recognition framework using interpretability. The benchmark concatenates content embeddings with timbre embeddings for speaker classification, then applies Gradient SHAP Explainer to quantify timbre residual. We evaluate seven speech pretraining model variations. (2) InterpTF-SptME method - an interpretability-based timbre filtering approach using SHAP Noise and SHAP Cropping techniques. This model-agnostic method transforms intermediate encodings to remove timbre while preserving content. Experiments on VCTK dataset with HuBERT LARGE demonstrate successful content preservation and significant speaker disentanglement optimization. Results show the SHAP Noise method can reduce timbre residual from 18.05% to near 0% while maintaining content integrity, contributing to enhanced performance in content-related speech processing tasks and preventing timbre privacy leakage.
【10】Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
标题:流媒体排序器:基于扬声器缓存的在线扬声器拨号,并通过到达时间排序
链接:https://arxiv.org/abs/2507.18446
备注:Accepted to Interspeech 2025
摘要:本文提出了一个流扩展的Sortformer扬声器日记框架,其关键属性是输出扬声器的到达时间排序。所提出的方法采用到达顺序扬声器缓存(AOSC)存储帧级的声学嵌入先前观察到的扬声器。与传统的说话者跟踪缓冲区不同,AOSC通过与其到达时间顺序相对应的说话者索引对嵌入进行排序,并通过基于模型的过去预测选择具有最高分数的帧来动态更新。值得注意的是,每个扬声器存储的嵌入数量由更新机制动态确定,确保高效的缓存利用率和精确的扬声器跟踪。在基准数据集上的实验证实了我们方法的有效性和灵活性,即使在低延迟设置中也是如此。这些结果确立了Streaming Sortformer作为实时多说话者跟踪的鲁棒解决方案和流式多说话者语音处理的基础。
摘要:This paper presents a streaming extension for the Sortformer speaker diarization framework, whose key property is the arrival-time ordering of output speakers. The proposed approach employs an Arrival-Order Speaker Cache (AOSC) to store frame-level acoustic embeddings of previously observed speakers. Unlike conventional speaker-tracing buffers, AOSC orders embeddings by speaker index corresponding to their arrival time order, and is dynamically updated by selecting frames with the highest scores based on the model's past predictions. Notably, the number of stored embeddings per speaker is determined dynamically by the update mechanism, ensuring efficient cache utilization and precise speaker tracking. Experiments on benchmark datasets confirm the effectiveness and flexibility of our approach, even in low-latency setups. These results establish Streaming Sortformer as a robust solution for real-time multi-speaker tracking and a foundation for streaming multi-talker speech processing.
【11】Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
标题:双路径多通道线性预测过滤器和多规范束形成的语音增强
链接:https://arxiv.org/abs/2507.18350
备注:Paper accepted by Interspeech 2025
摘要:本文提出了一种基于双通道多通道线性预测(MCLP)滤波器的语音增强方法 和多范数波束形成。具体而言,MCLP部分 所提出的方法在两种情况下都设计有双路滤波器, 时间和频率维度。对于波束成形部分,我们 最小化麦克风阵列输出的功率,以及 去噪信号的L1范数,同时保留来自目标方向的源信号。一种有效的方法来选择 提出了双径滤波器的预测阶数, 对于具有不同混响时间(T60)值的信号具有鲁棒性,并且可以应用于其他基于MCLP的方法。评估表明,我们提出的方法优于 语音增强的基线方法,特别是在高 混响场景
摘要:In this paper, we propose a speech enhancement method us ing dual-path Multi-Channel Linear Prediction (MCLP) filters and multi-norm beamforming. Specifically, the MCLP part in the proposed method is designed with dual-path filters in both time and frequency dimensions. For the beamforming part, we minimize the power of the microphone array output as well as the l1 norm of the denoised signals while preserving source sig nals from the target directions. An efficient method to select the prediction orders in the dual-path filters is also proposed, which is robust for signals with different reverberation time (T60) val ues and can be applied to other MCLP-based methods. Eval uations demonstrate that our proposed method outperforms the baseline methods for speech enhancement, particularly in high reverberation scenarios.
【12】SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
标题:SpecASB:通过推测解码加速基于LLM的自动语音识别
链接:https://arxiv.org/abs/2507.18181
摘要:基于大语言模型(LLM)的自动语音识别(ASR)因其识别准确率高、对多方言的支持增强等优点,近年来受到了广泛关注。然而,LLM的高解码延迟对实时ASR要求提出了挑战。虽然推测解码已经被探索以获得更好的解码效率,但是它们通常忽略ASR任务的关键特征并且实现有限的加速。为了进一步减少实时ASR延迟,在本文中,我们提出了一种新的投机解码框架专门为ASR,称为SpecASR。SpecASR是基于我们的核心观察而开发的,即ASR解码是音频调节的,这导致小型和大型ASR模型之间的高输出对齐,即使在中间解码步骤中存在输出失配。因此,SpecASR具有自适应草案序列生成过程,该过程动态修改草案序列长度以最大化令牌接受长度。SpecASR进一步提出了一种草稿序列回收策略,该策略重用先前生成的草稿序列以减少草稿ASR模型延迟。此外,还提出了一种两遍稀疏令牌树生成算法,以平衡草稿和目标ASR模型的延迟。通过大量的实验结果,我们证明SpecASR分别比基线自回归解码和推测解码实现了3.04x-3.79x和1.25x-1.84x的加速,而没有任何识别准确性的损失。
摘要:Large language model (LLM)-based automatic speech recognition (ASR) has recently attracted a lot of attention due to its high recognition accuracy and enhanced multi-dialect support. However, the high decoding latency of LLMs challenges the real-time ASR requirements. Although speculative decoding has been explored for better decoding efficiency, they usually ignore the key characteristics of the ASR task and achieve limited speedup. To further reduce the real-time ASR latency, in this paper, we propose a novel speculative decoding framework specialized for ASR, dubbed SpecASR. SpecASR is developed based on our core observation that ASR decoding is audio-conditioned, which results in high output alignment between small and large ASR models, even given output mismatches in intermediate decoding steps. Therefore, SpecASR features an adaptive draft sequence generation process that dynamically modifies the draft sequence length to maximize the token acceptance length. SpecASR further proposes a draft sequence recycling strategy that reuses the previously generated draft sequence to reduce the draft ASR model latency. Moreover, a two-pass sparse token tree generation algorithm is also proposed to balance the latency of draft and target ASR models. With extensive experimental results, we demonstrate SpecASR achieves 3.04x-3.79x and 1.25x-1.84x speedup over the baseline autoregressive decoding and speculative decoding, respectively, without any loss in recognition accuracy.
【13】Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
标题:远距离对话语音识别的最新趋势:CHiME-7和8 DSVR挑战回顾
链接:https://arxiv.org/abs/2507.18161
摘要:CHiME-7和8远程语音识别(DASR)挑战集中在多通道,可推广,联合自动语音识别(ASR)和对话语音的日记化。有9个团队提交了32个不同的系统,这些挑战有助于该领域的最先进的研究。本文概述了挑战的设计,评估指标,数据集和基线系统,同时分析参与者提交的关键趋势。从该分析中可以看出:1)大多数参与者使用端到端(e2 e)ASR系统,而混合系统在以前的CHiME挑战中很普遍。这种转变主要是由于强大的大规模预训练模型的可用性,这降低了e2 e-ASR的数据负担。2)尽管最近在神经语音分离和增强(SSE)方面取得了进展,但所有团队仍然严重依赖于引导源分离,这表明当前的神经SSE技术仍然无法可靠地处理复杂的场景和不同的记录设置。3)所有最好的系统都通过目标说话人日记化技术来采用日记化细化。因此,在第一次日记化过程中准确的说话者计数对于避免复合错误至关重要,CHiME-8 DASR参与者特别关注这一部分。4)通过会议摘要进行的下游评估与转录质量的相关性很弱,因为大语言模型在处理错误方面具有显着的有效性。在NOTSOFAR-1场景中,即使具有超过50\%时间约束最小置换WER的系统也可以与最有效的系统(约11\%)大致相当。5)尽管最近取得了进展,但在具有挑战性的声学环境中准确转录自发语音仍然很困难,即使在使用计算密集型系统集成时也是如此。
摘要:The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse systems, these challenges have contributed to state-of-the-art research in the field. This paper outlines the challenges' design, evaluation metrics, datasets, and baseline systems while analyzing key trends from participant submissions. From this analysis it emerges that: 1) Most participants use end-to-end (e2e) ASR systems, whereas hybrid systems were prevalent in previous CHiME challenges. This transition is mainly due to the availability of robust large-scale pre-trained models, which lowers the data burden for e2e-ASR. 2) Despite recent advances in neural speech separation and enhancement (SSE), all teams still heavily rely on guided source separation, suggesting that current neural SSE techniques are still unable to reliably deal with complex scenarios and different recording setups. 3) All best systems employ diarization refinement via target-speaker diarization techniques. Accurate speaker counting in the first diarization pass is thus crucial to avoid compounding errors and CHiME-8 DASR participants especially focused on this part. 4) Downstream evaluation via meeting summarization can correlate weakly with transcription quality due to the remarkable effectiveness of large-language models in handling errors. On the NOTSOFAR-1 scenario, even systems with over 50\% time-constrained minimum permutation WER can perform roughly on par with the most effective ones (around 11\%). 5) Despite recent progress, accurately transcribing spontaneous speech in challenging acoustic environments remains difficult, even when using computationally intensive system ensembles.
【14】A Concept-based approach to Voice Disorder Detection
标题:基于概念的语音障碍检测方法
链接:https://arxiv.org/abs/2507.17799
摘要:语音障碍影响了很大一部分人口,使用自动化,非侵入性技术诊断语音障碍的能力将代表医疗保健的重大进步,提高患者的生活质量。最近的研究表明,人工智能模型,特别是深度神经网络(DNN),可以有效地解决这一任务。然而,由于它们的复杂性,这些模型的决策过程通常保持不透明,限制了它们在临床环境中的可信度。本文研究了一种基于可解释AI(XAI)的替代方法,该领域旨在通过提供不同形式的解释来提高DNN的可解释性。具体来说,这项工作侧重于基于概念的模型,如概念瓶颈模型(CBM)和概念嵌入模型(CEM),以及它们如何实现与传统深度学习方法相当的性能,同时提供更透明和可解释的决策框架。
摘要:Voice disorders affect a significant portion of the population, and the ability to diagnose them using automated, non-invasive techniques would represent a substantial advancement in healthcare, improving the quality of life of patients. Recent studies have demonstrated that artificial intelligence models, particularly Deep Neural Networks (DNNs), can effectively address this task. However, due to their complexity, the decision-making process of such models often remain opaque, limiting their trustworthiness in clinical contexts. This paper investigates an alternative approach based on Explainable AI (XAI), a field that aims to improve the interpretability of DNNs by providing different forms of explanations. Specifically, this works focuses on concept-based models such as Concept Bottleneck Model (CBM) and Concept Embedding Model (CEM) and how they can achieve performance comparable to traditional deep learning methods, while offering a more transparent and interpretable decision framework.
【1】Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
标题:流媒体排序器:基于扬声器缓存的在线扬声器拨号,并通过到达时间排序
链接:https://arxiv.org/abs/2507.18446
备注:Accepted to Interspeech 2025
摘要:本文提出了一个流扩展的Sortformer扬声器日记框架,其关键属性是输出扬声器的到达时间排序。所提出的方法采用到达顺序扬声器缓存(AOSC)存储帧级的声学嵌入先前观察到的扬声器。与传统的说话者跟踪缓冲区不同,AOSC通过与其到达时间顺序相对应的说话者索引对嵌入进行排序,并通过基于模型的过去预测选择具有最高分数的帧来动态更新。值得注意的是,每个扬声器存储的嵌入数量由更新机制动态确定,确保高效的缓存利用率和精确的扬声器跟踪。在基准数据集上的实验证实了我们方法的有效性和灵活性,即使在低延迟设置中也是如此。这些结果确立了Streaming Sortformer作为实时多说话者跟踪的鲁棒解决方案和流式多说话者语音处理的基础。
摘要:This paper presents a streaming extension for the Sortformer speaker diarization framework, whose key property is the arrival-time ordering of output speakers. The proposed approach employs an Arrival-Order Speaker Cache (AOSC) to store frame-level acoustic embeddings of previously observed speakers. Unlike conventional speaker-tracing buffers, AOSC orders embeddings by speaker index corresponding to their arrival time order, and is dynamically updated by selecting frames with the highest scores based on the model's past predictions. Notably, the number of stored embeddings per speaker is determined dynamically by the update mechanism, ensuring efficient cache utilization and precise speaker tracking. Experiments on benchmark datasets confirm the effectiveness and flexibility of our approach, even in low-latency setups. These results establish Streaming Sortformer as a robust solution for real-time multi-speaker tracking and a foundation for streaming multi-talker speech processing.
【2】Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
标题:双路径多通道线性预测过滤器和多规范束形成的语音增强
链接:https://arxiv.org/abs/2507.18350
备注:Paper accepted by Interspeech 2025
摘要:本文提出了一种基于双通道多通道线性预测(MCLP)滤波器的语音增强方法 和多范数波束形成。具体而言,MCLP部分 所提出的方法在两种情况下都设计有双路滤波器, 时间和频率维度。对于波束成形部分,我们 最小化麦克风阵列输出的功率,以及 去噪信号的L1范数,同时保留来自目标方向的源信号。一种有效的方法来选择 提出了双径滤波器的预测阶数, 对于具有不同混响时间(T60)值的信号具有鲁棒性,并且可以应用于其他基于MCLP的方法。评估表明,我们提出的方法优于 语音增强的基线方法,特别是在高 混响场景
摘要:In this paper, we propose a speech enhancement method us ing dual-path Multi-Channel Linear Prediction (MCLP) filters and multi-norm beamforming. Specifically, the MCLP part in the proposed method is designed with dual-path filters in both time and frequency dimensions. For the beamforming part, we minimize the power of the microphone array output as well as the l1 norm of the denoised signals while preserving source sig nals from the target directions. An efficient method to select the prediction orders in the dual-path filters is also proposed, which is robust for signals with different reverberation time (T60) val ues and can be applied to other MCLP-based methods. Eval uations demonstrate that our proposed method outperforms the baseline methods for speech enhancement, particularly in high reverberation scenarios.
【3】SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding
标题:SpecASB:通过推测解码加速基于LLM的自动语音识别
链接:https://arxiv.org/abs/2507.18181
摘要:基于大语言模型(LLM)的自动语音识别(ASR)因其识别准确率高、对多方言的支持增强等优点,近年来受到了广泛关注。然而,LLM的高解码延迟对实时ASR要求提出了挑战。虽然推测解码已经被探索以获得更好的解码效率,但是它们通常忽略ASR任务的关键特征并且实现有限的加速。为了进一步减少实时ASR延迟,在本文中,我们提出了一种新的投机解码框架专门为ASR,称为SpecASR。SpecASR是基于我们的核心观察而开发的,即ASR解码是音频调节的,这导致小型和大型ASR模型之间的高输出对齐,即使在中间解码步骤中存在输出失配。因此,SpecASR具有自适应草案序列生成过程,该过程动态修改草案序列长度以最大化令牌接受长度。SpecASR进一步提出了一种草稿序列回收策略,该策略重用先前生成的草稿序列以减少草稿ASR模型延迟。此外,还提出了一种两遍稀疏令牌树生成算法,以平衡草稿和目标ASR模型的延迟。通过大量的实验结果,我们证明SpecASR分别比基线自回归解码和推测解码实现了3.04x-3.79x和1.25x-1.84x的加速,而没有任何识别准确性的损失。
摘要:Large language model (LLM)-based automatic speech recognition (ASR) has recently attracted a lot of attention due to its high recognition accuracy and enhanced multi-dialect support. However, the high decoding latency of LLMs challenges the real-time ASR requirements. Although speculative decoding has been explored for better decoding efficiency, they usually ignore the key characteristics of the ASR task and achieve limited speedup. To further reduce the real-time ASR latency, in this paper, we propose a novel speculative decoding framework specialized for ASR, dubbed SpecASR. SpecASR is developed based on our core observation that ASR decoding is audio-conditioned, which results in high output alignment between small and large ASR models, even given output mismatches in intermediate decoding steps. Therefore, SpecASR features an adaptive draft sequence generation process that dynamically modifies the draft sequence length to maximize the token acceptance length. SpecASR further proposes a draft sequence recycling strategy that reuses the previously generated draft sequence to reduce the draft ASR model latency. Moreover, a two-pass sparse token tree generation algorithm is also proposed to balance the latency of draft and target ASR models. With extensive experimental results, we demonstrate SpecASR achieves 3.04x-3.79x and 1.25x-1.84x speedup over the baseline autoregressive decoding and speculative decoding, respectively, without any loss in recognition accuracy.
【4】Recent Trends in Distant Conversational Speech Recognition: A Review of CHiME-7 and 8 DASR Challenges
标题:远距离对话语音识别的最新趋势:CHiME-7和8 DSVR挑战回顾
链接:https://arxiv.org/abs/2507.18161
摘要:CHiME-7和8远程语音识别(DASR)挑战集中在多通道,可推广,联合自动语音识别(ASR)和对话语音的日记化。有9个团队提交了32个不同的系统,这些挑战有助于该领域的最先进的研究。本文概述了挑战的设计,评估指标,数据集和基线系统,同时分析参与者提交的关键趋势。从该分析中可以看出:1)大多数参与者使用端到端(e2 e)ASR系统,而混合系统在以前的CHiME挑战中很普遍。这种转变主要是由于强大的大规模预训练模型的可用性,这降低了e2 e-ASR的数据负担。2)尽管最近在神经语音分离和增强(SSE)方面取得了进展,但所有团队仍然严重依赖于引导源分离,这表明当前的神经SSE技术仍然无法可靠地处理复杂的场景和不同的记录设置。3)所有最好的系统都通过目标说话人日记化技术来采用日记化细化。因此,在第一次日记化过程中准确的说话者计数对于避免复合错误至关重要,CHiME-8 DASR参与者特别关注这一部分。4)通过会议摘要进行的下游评估与转录质量的相关性很弱,因为大语言模型在处理错误方面具有显着的有效性。在NOTSOFAR-1场景中,即使具有超过50\%时间约束最小置换WER的系统也可以与最有效的系统(约11\%)大致相当。5)尽管最近取得了进展,但在具有挑战性的声学环境中准确转录自发语音仍然很困难,即使在使用计算密集型系统集成时也是如此。
摘要:The CHiME-7 and 8 distant speech recognition (DASR) challenges focus on multi-channel, generalizable, joint automatic speech recognition (ASR) and diarization of conversational speech. With participation from 9 teams submitting 32 diverse systems, these challenges have contributed to state-of-the-art research in the field. This paper outlines the challenges' design, evaluation metrics, datasets, and baseline systems while analyzing key trends from participant submissions. From this analysis it emerges that: 1) Most participants use end-to-end (e2e) ASR systems, whereas hybrid systems were prevalent in previous CHiME challenges. This transition is mainly due to the availability of robust large-scale pre-trained models, which lowers the data burden for e2e-ASR. 2) Despite recent advances in neural speech separation and enhancement (SSE), all teams still heavily rely on guided source separation, suggesting that current neural SSE techniques are still unable to reliably deal with complex scenarios and different recording setups. 3) All best systems employ diarization refinement via target-speaker diarization techniques. Accurate speaker counting in the first diarization pass is thus crucial to avoid compounding errors and CHiME-8 DASR participants especially focused on this part. 4) Downstream evaluation via meeting summarization can correlate weakly with transcription quality due to the remarkable effectiveness of large-language models in handling errors. On the NOTSOFAR-1 scenario, even systems with over 50\% time-constrained minimum permutation WER can perform roughly on par with the most effective ones (around 11\%). 5) Despite recent progress, accurately transcribing spontaneous speech in challenging acoustic environments remains difficult, even when using computationally intensive system ensembles.
【5】A Concept-based approach to Voice Disorder Detection
标题:基于概念的语音障碍检测方法
链接:https://arxiv.org/abs/2507.17799
摘要:语音障碍影响了很大一部分人口,使用自动化,非侵入性技术诊断语音障碍的能力将代表医疗保健的重大进步,提高患者的生活质量。最近的研究表明,人工智能模型,特别是深度神经网络(DNN),可以有效地解决这一任务。然而,由于它们的复杂性,这些模型的决策过程通常保持不透明,限制了它们在临床环境中的可信度。本文研究了一种基于可解释AI(XAI)的替代方法,该领域旨在通过提供不同形式的解释来提高DNN的可解释性。具体来说,这项工作侧重于基于概念的模型,如概念瓶颈模型(CBM)和概念嵌入模型(CEM),以及它们如何实现与传统深度学习方法相当的性能,同时提供更透明和可解释的决策框架。
摘要:Voice disorders affect a significant portion of the population, and the ability to diagnose them using automated, non-invasive techniques would represent a substantial advancement in healthcare, improving the quality of life of patients. Recent studies have demonstrated that artificial intelligence models, particularly Deep Neural Networks (DNNs), can effectively address this task. However, due to their complexity, the decision-making process of such models often remain opaque, limiting their trustworthiness in clinical contexts. This paper investigates an alternative approach based on Explainable AI (XAI), a field that aims to improve the interpretability of DNNs by providing different forms of explanations. Specifically, this works focuses on concept-based models such as Concept Bottleneck Model (CBM) and Concept Embedding Model (CEM) and how they can achieve performance comparable to traditional deep learning methods, while offering a more transparent and interpretable decision framework.
【6】ASR-Guided Speaker-Role Diarization and Diarization-Guided ASR Decoding
标题:ASB引导的扬声器角色分区和分区引导的ASB解码
链接:https://arxiv.org/abs/2507.17765
备注:Interspeech 2025 Submission
摘要:从应用的角度来看,说话者角色日记(RD),如医生与患者,主人与客人等,通常比传统的说话者日记(SD)更有用,后者分配通用标签,如说话者-1,说话者-2等。在联合自动语音识别(ASR)+ SD(谁说了什么?)的上下文中,最近的端到端模型采用与ASR换能器同步的辅助SD换能器来预测每个单词的说话者。在本文中,我们将这一框架扩展到RD,并做出了三个关键贡献:(1)我们通过强制对齐和交叉熵损失而不是RNNT损失来简化训练,(2)我们表明单词预测和角色预测需要不同数量的预测器上下文,导致单独的任务特定预测器,与现有的共享预测器模型不同,以及(3)我们提出了一种利用RD后验活动来影响ASR解码并减少小词删除错误的方法。
摘要:From an application standpoint, speaker-role diarization (RD), such as doctor vs. patient, host vs. guest, etc. is often more useful than traditional speaker diarization (SD), which assigns generic labels like speaker-1, speaker-2 etc. In the context of joint automatic speech recognition (ASR) + SD (who spoke what?), recent end-to-end models employ an auxiliary SD transducer, synchronized with the ASR transducer, to predict speakers per word. In this paper, we extend this framework to RD with three key contributions: (1) we simplify the training via forced alignment and cross-entropy loss instead of RNNT loss, (2) we show that word prediction and role prediction require different amounts of predictor's context, leading to separate task-specific predictors, unlike existing shared-predictor models, and (3) we propose a way to leverage RD posterior activity to influence ASR decoding and reduce small-word deletion errors.
【7】Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
标题:微型还不够小:通过混合知识提炼高质量、低资源的面部动画模型
链接:https://arxiv.org/abs/2507.18352
备注:Accepted to ACM Transactions on Graphics 2025 (SIGGRAPH journal track)
摘要:为语音驱动的3D面部动画训练高质量、鲁棒的机器学习模型需要大量不同的高质量音频动画对数据集。为了克服这种数据集的缺乏,最近的工作引入了大型的预训练语音编码器,这些编码器对输入音频的变化具有鲁棒性,因此,使面部动画模型能够跨扬声器,音频质量和语言进行概括。然而,由此产生的面部动画模型是非常大的,并且只适合在专用机器上进行离线推理。在这项工作中,我们将探讨在游戏开发的背景下,设备上的实时面部动画模型。我们克服了缺乏大型数据集,使用混合知识蒸馏与伪标签。给定一个大的音频数据集,我们采用高性能的教师模型来训练非常小的学生模型。与预先训练的语音编码器相比,我们的学生模型仅由卷积层和全连接层组成,不需要注意上下文或定期更新。在我们的实验中,我们证明,我们可以减少内存占用高达3.4 MB,并要求未来的音频上下文高达81毫秒,同时保持高品质的动画。这为设备上的推理铺平了道路,这是迈向逼真的模型驱动数字角色的重要一步。
摘要:The training of high-quality, robust machine learning models for speech-driven 3D facial animation requires a large, diverse dataset of high-quality audio-animation pairs. To overcome the lack of such a dataset, recent work has introduced large pre-trained speech encoders that are robust to variations in the input audio and, therefore, enable the facial animation model to generalize across speakers, audio quality, and languages. However, the resulting facial animation models are prohibitively large and lend themselves only to offline inference on a dedicated machine. In this work, we explore on-device, real-time facial animation models in the context of game development. We overcome the lack of large datasets by using hybrid knowledge distillation with pseudo-labeling. Given a large audio dataset, we employ a high-performing teacher model to train very small student models. In contrast to the pre-trained speech encoders, our student models only consist of convolutional and fully-connected layers, removing the need for attention context or recurrent updates. In our experiments, we demonstrate that we can reduce the memory footprint to up to 3.4 MB and required future audio context to up to 81 ms while maintaining high-quality animations. This paves the way for on-device inference, an important step towards realistic, model-driven digital characters.
【8】Improving Bird Classification with Primary Color Additives
标题:用原色添加剂改善鸟类分类
链接:https://arxiv.org/abs/2507.18334
备注:5 pages (Accepted to Interspeech 2025)
摘要:我们解决的问题,鸟类的分类使用他们的歌曲录音,一个具有挑战性的任务,由于环境噪声,重叠发声,和丢失的标签。现有的模型与低信噪比或多物种记录的斗争。我们假设,鸟类可以通过可视化他们的音高模式,速度和重复,统称为图案进行分类。应用于光谱图图像的深度学习模型有所帮助,但不同物种之间的相似图案会导致混淆。为了减轻这种情况,我们使用原色添加剂将频率信息嵌入到频谱图中。这增强了物种区分并提高了分类准确性。我们的实验表明,所提出的方法在没有着色的模型上实现了统计学上的显著增益,并超过了BirdCLEF 2024的获胜者,将F1提高了7.3%,ROC-AUC提高了6.2%,CMAP提高了6.6%。这些结果证明了通过彩色化结合频率信息的有效性。
摘要:We address the problem of classifying bird species using their song recordings, a challenging task due to environmental noise, overlapping vocalizations, and missing labels. Existing models struggle with low-SNR or multi-species recordings. We hypothesize that birds can be classified by visualizing their pitch pattern, speed, and repetition, collectively called motifs. Deep learning models applied to spectrogram images help, but similar motifs across species cause confusion. To mitigate this, we embed frequency information into spectrograms using primary color additives. This enhances species distinction and improves classification accuracy. Our experiments show that the proposed approach achieves statistically significant gains over models without colorization and surpasses the BirdCLEF 2024 winner, improving F1 by 7.3%, ROC-AUC by 6.2%, and CMAP by 6.6%. These results demonstrate the effectiveness of incorporating frequency information via colorization.
【9】GOAT-SLM: A Spoken Language Model with Paralinguistic and Speaker Characteristic Awareness
标题:GOAT-LAM:具有副语言和说话者特征意识的口语模型
链接:https://arxiv.org/abs/2507.18119
摘要:端到端口语模型(SLM)的最新进展显着提高了人工智能系统参与自然口语交互的能力。然而,大多数现有的模型只把语音作为语言内容的载体,往往忽略了丰富的语言和说话者特征线索嵌入在人类语音,如方言,年龄,情感和非语音发声。在这项工作中,我们介绍了GOAT-SLM,一种新的口语模型,具有语言学和说话者特征意识,旨在扩展口语建模超越文本语义。GOAT-SLM采用双模态头部架构,将语言建模与声学实现相结合,实现了强大的语言理解,同时支持表达性和自适应语音生成。为了提高模型的效率和通用性,我们提出了一种模块化的分阶段训练策略,该策略使用大规模语音文本语料库逐步调整语言,语言学和说话人特征信息。TELEVAL上的实验结果表明,GOAT-SLM在语义和非语义任务上实现了良好的平衡性能,在处理情感,方言变化和年龄敏感的交互方面优于现有的开源模型。这项工作强调了建模超越语言内容的重要性,并推动了更自然,适应性和社会意识的口语系统的发展。
摘要:Recent advances in end-to-end spoken language models (SLMs) have significantly improved the ability of AI systems to engage in natural spoken interactions. However, most existing models treat speech merely as a vehicle for linguistic content, often overlooking the rich paralinguistic and speaker characteristic cues embedded in human speech, such as dialect, age, emotion, and non-speech vocalizations. In this work, we introduce GOAT-SLM, a novel spoken language model with paralinguistic and speaker characteristic awareness, designed to extend spoken language modeling beyond text semantics. GOAT-SLM adopts a dual-modality head architecture that decouples linguistic modeling from acoustic realization, enabling robust language understanding while supporting expressive and adaptive speech generation. To enhance model efficiency and versatility, we propose a modular, staged training strategy that progressively aligns linguistic, paralinguistic, and speaker characteristic information using large-scale speech-text corpora. Experimental results on TELEVAL, a multi-dimensional evaluation benchmark, demonstrate that GOAT-SLM achieves well-balanced performance across both semantic and non-semantic tasks, and outperforms existing open-source models in handling emotion, dialectal variation, and age-sensitive interactions. This work highlights the importance of modeling beyond linguistic content and advances the development of more natural, adaptive, and socially aware spoken language systems.
【10】TELEVAL: A Dynamic Benchmark Designed for Spoken Language Models in Chinese Interactive Scenarios
标题:TELEVAL:汉语交互场景中口语模型的动态基准
链接:https://arxiv.org/abs/2507.18061
摘要:近年来,口语模型(SLM)取得了快速的进展,并开发了许多用于评估其性能的基准。然而,大多数现有的基准主要集中在评估SLM是否可以执行与大型语言模型(LLM)处理的复杂任务相当的复杂任务,通常无法与用户在现实世界的会话场景中自然交互的方式保持一致。在本文中,我们提出了TELEVAL,一个动态的基准,专门设计来评估SLM的有效性,在现实的中国互动设置的会话代理。TELEVAL定义了三个评估维度:显式语义、副语言和隐式语义以及系统能力。它采用与现实世界使用一致的对话格式,并分别评估文本和音频输出。TELEVAL特别关注模型从用户语音中提取隐含线索并在没有额外指令的情况下做出适当反应的能力。我们的实验表明,尽管最近的进展,现有的SLM仍然有相当大的改进空间,在自然的会话任务。我们希望TELEVAL可以作为一个以用户为中心的评估框架,直接反映用户体验,并有助于开发更有能力的对话导向的SLM。
摘要:Spoken language models (SLMs) have seen rapid progress in recent years, along with the development of numerous benchmarks for evaluating their performance. However, most existing benchmarks primarily focus on evaluating whether SLMs can perform complex tasks comparable to those tackled by large language models (LLMs), often failing to align with how users naturally interact in real-world conversational scenarios. In this paper, we propose TELEVAL, a dynamic benchmark specifically designed to evaluate SLMs' effectiveness as conversational agents in realistic Chinese interactive settings. TELEVAL defines three evaluation dimensions: Explicit Semantics, Paralinguistic and Implicit Semantics, and System Abilities. It adopts a dialogue format consistent with real-world usage and evaluates text and audio outputs separately. TELEVAL particularly focuses on the model's ability to extract implicit cues from user speech and respond appropriately without additional instructions. Our experiments demonstrate that despite recent progress, existing SLMs still have considerable room for improvement in natural conversational tasks. We hope that TELEVAL can serve as a user-centered evaluation framework that directly reflects the user experience and contributes to the development of more capable dialogue-oriented SLMs.
【11】The TEA-ASLP System for Multilingual Conversational Speech Recognition and Speech Diarization in MLC-SLM 2025 Challenge
标题:MLC-SL 2025挑战赛中用于多语言对话语音识别和语音规模化的TEA-ASLP系统
链接:https://arxiv.org/abs/2507.18051
备注:Interspeech 2025 workshop
摘要:本文介绍了TEA-ASLP提交给MLC-SLM 2025挑战赛的系统,在任务I中解决了多语言会话自动语音识别(ASR),在任务II中解决了语音日志化ASR。对于任务I,我们通过集成已知的语言识别和多语言MoE LoRA结构来增强Ideal-LLM模型,同时使用CTC预测的令牌作为提示来改进自回归生成。该模型基于大约18万小时的多语言ASR数据进行训练。在任务II中,我们用一个更合适的纯英语版本替换了基线的英汉说话者日记模型。与基线语音语言模型相比,我们的方法实现了30.8%的单词错误率(WER)降低,导致任务I中的最终WER为9.60%,任务II中的时间约束最小排列WER为17.49%,在各自的挑战任务中获得第一名和第二名。
摘要:This paper presents the TEA-ASLP's system submitted to the MLC-SLM 2025 Challenge, addressing multilingual conversational automatic speech recognition (ASR) in Task I and speech diarization ASR in Task II. For Task I, we enhance Ideal-LLM model by integrating known language identification and a multilingual MOE LoRA structure, along with using CTC-predicted tokens as prompts to improve autoregressive generation. The model is trained on approximately 180k hours of multilingual ASR data. In Task II, we replace the baseline English-Chinese speaker diarization model with a more suitable English-only version. Our approach achieves a 30.8% reduction in word error rate (WER) compared to the baseline speech language model, resulting in a final WER of 9.60% in Task I and a time-constrained minimum-permutation WER of 17.49% in Task II, earning first and second place in the respective challenge tasks.
【12】Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
标题:用于声音事件定位、检测和距离估计的具有共享权重和注意力机制的Resnet-conformer网络
链接:https://arxiv.org/abs/2507.17941
备注:This paper has been submitted as a technical report outlining our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024 and can be found in DCASE2024 technical reports
摘要:本技术报告概述了我们对2024年声学场景和事件的检测和分类(DCASE)任务3A的方法,重点是声音事件定位和检测(SELD)。SELD通过估计声音事件定位和检测提供有价值的见解,帮助各种机器认知任务,如环境推理,导航和其他声音定位相关的应用。今年的挑战赛对使用纯音频(音轨A)或视听(音轨B)输入的真实声音场景的注释录音进行评估。今年的一个显著变化是采用了距离估计,并相应调整了评价指标,以进行全面评估。我们提交的是挑战赛的任务A,重点是音频轨道。我们的方法利用对数梅尔光谱图,强度矢量,并采用多个数据增强。我们提出了一种基于EINV2的网络架构[1],取得了改进的结果:在开发数据集[2,3]的测试集上,F分数为40.2%,角度误差(DOA)为17.7度,相对距离误差(RDE)为0.32。
摘要:This technical report outlines our approach to Task 3A of the Detection and Classification of Acoustic Scenes and Events (DCASE) 2024, focusing on Sound Event Localization and Detection (SELD). SELD provides valuable insights by estimating sound event localization and detection, aiding in various machine cognition tasks such as environmental inference, navigation, and other sound localization-related applications. This year's challenge evaluates models using either audio-only (Track A) or audiovisual (Track B) inputs on annotated recordings of real sound scenes. A notable change this year is the introduction of distance estimation, with evaluation metrics adjusted accordingly for a comprehensive assessment. Our submission is for Task A of the Challenge, which focuses on the audio-only track. Our approach utilizes log-mel spectrograms, intensity vectors, and employs multiple data augmentations. We proposed an EINV2-based [1] network architecture, achieving improved results: an F-score of 40.2%, Angular Error (DOA) of 17.7 degrees, and Relative Distance Error (RDE) of 0.32 on the test set of the Development Dataset [2 ,3].
【13】Bob's Confetti: Phonetic Memorization Attacks in Music and Video Generation
标题:鲍勃的五彩纸屑:音乐和视频生成中的语音同步攻击
链接:https://arxiv.org/abs/2507.17937
摘要:歌词到歌曲(LS 2)生成模型承诺从文本进行端到端音乐合成,但它们在训练数据记忆方面的脆弱性仍然没有得到充分研究。我们介绍了对抗性音素替换(APT),这是一种新颖的攻击,在这种攻击中,歌词在语义上被改变,同时通过谐音替换(例如,阿姆的著名的“妈妈的意大利面”$\rightarrow$“鲍勃的五彩纸屑”)。尽管存在这些扭曲,但我们发现了一种强大的子词汇记忆形式:像SUNO和YuE这样的模型重新生成的输出与已知的训练内容惊人地相似,在音频域指标(包括CLAP,AudioJudge和CoverID)之间实现了高度相似性。此漏洞在多种语言和流派中持续存在。更令人惊讶的是,我们发现,在文本到视频模型中,仅音素改变的歌词就可以触发视觉记忆。当提示与语音修改歌词从失去自己,Veo 3重建视觉元素从原来的音乐视频-包括字符外观和场景组成-尽管没有视觉线索的提示。我们称这种现象为语音-视觉回流。总之,这些研究结果暴露了一个关键的脆弱性,在成绩单条件下的多模态生成:语音提示单独可以解锁记忆的视听内容,提出了迫切的问题,在现代生成系统的版权,安全和内容出处。示例生成可以在我们的演示页面(jrohsc.github.io/music_attack/)上找到。
摘要:Lyrics-to-Song (LS2) generation models promise end-to-end music synthesis from text, yet their vulnerability to training data memorization remains underexplored. We introduce Adversarial PhoneTic Prompting (APT), a novel attack where lyrics are semantically altered while preserving their acoustic structure through homophonic substitutions (e.g., Eminem's famous "mom's spaghetti" $\rightarrow$ "Bob's confetti"). Despite these distortions, we uncover a powerful form of sub-lexical memorization: models like SUNO and YuE regenerate outputs strikingly similar to known training content, achieving high similarity across audio-domain metrics, including CLAP, AudioJudge, and CoverID. This vulnerability persists across multiple languages and genres. More surprisingly, we discover that phoneme-altered lyrics alone can trigger visual memorization in text-to-video models. When prompted with phonetically modified lyrics from Lose Yourself, Veo 3 reconstructs visual elements from the original music video -- including character appearance and scene composition -- despite no visual cues in the prompt. We term this phenomenon phonetic-to-visual regurgitation. Together, these findings expose a critical vulnerability in transcript-conditioned multimodal generation: phonetic prompting alone can unlock memorized audiovisual content, raising urgent questions about copyright, safety, and content provenance in modern generative systems. Example generations are available on our demo page (jrohsc.github.io/music_attack/).
【14】One Whisper to Grade Them All
标题:一个耳语就可以给他们全部评分
链接:https://arxiv.org/abs/2507.17918
备注:Accepted to SLaTE 2025 workshop
摘要:我们为多部分第二语言测试的整体自动口语评估(ASA)提供了一种高效的端到端方法,该方法是为2025年Speak & Improve Challenge开发的。我们的系统的主要新颖性是能够处理所有四个口头回应与一个耳语小编码器,通过一个轻量级的聚合器结合所有信息,并预测最终得分。这种架构消除了对转录和每部分模型的需要,减少了推理时间,并使ASA适用于大规模计算机辅助语言学习系统。 我们的系统实现了0.384的均方根误差(RMSE),优于基于文本的基线(0.44),同时使用最多168 M参数(约70%的耳语小)。此外,我们提出了一个数据采样策略,允许该模型训练上只有44.8%的发言人在语料库中,仍然达到0.383 RMSE,表现出改进的性能不平衡类和强大的数据效率。
摘要:We present an efficient end-to-end approach for holistic Automatic Speaking Assessment (ASA) of multi-part second-language tests, developed for the 2025 Speak & Improve Challenge. Our system's main novelty is the ability to process all four spoken responses with a single Whisper-small encoder, combine all information via a lightweight aggregator, and predict the final score. This architecture removes the need for transcription and per-part models, cuts inference time, and makes ASA practical for large-scale Computer-Assisted Language Learning systems. Our system achieved a Root Mean Squared Error (RMSE) of 0.384, outperforming the text-based baseline (0.44) while using at most 168M parameters (about 70% of Whisper-small). Furthermore, we propose a data sampling strategy, allowing the model to train on only 44.8% of the speakers in the corpus and still reach 0.383 RMSE, demonstrating improved performance on imbalanced classes and strong data efficiency.
【15】Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability
标题:基于可解释性的语音预训练模型的说话人去纠缠
链接:https://arxiv.org/abs/2507.17851
备注:20 pages, 9 figures, 2 tables
摘要:语音预训练模型包含跨不同层的特定于任务的信息,但解耦内容和音色信息仍然具有挑战性,因为删除特定于说话者的信息通常会导致内容丢失。目前的研究缺乏直接的指标来量化模型编码中的音色残留,依赖于通过下游任务的间接评估。本文通过在语音预训练模型中基于可解释性的说话人解纠缠来解决这些挑战。我们定量评估模型嵌入的音色残留,并使用解释性表示提高扬声器解纠缠。我们的贡献包括:(1)InterpTRQE-SptME基准-一个使用可解释性的音色残留识别框架。该基准将内容嵌入与音色嵌入连接起来进行说话人分类,然后应用Gradient SHAP Explainer量化音色残差。我们评估了七个语音预训练模型的变化。(2)InterpTF-SptME方法-使用SHAP噪声和SHAP裁剪技术的基于可解释性的音色过滤方法。这种与模型无关的方法转换中间编码以在保留内容的同时去除音色。使用HuBERT LARGE在VCTK数据集上进行的实验证明了成功的内容保留和显着的扬声器解纠缠优化。结果表明,SHAP噪声方法可以将音色残留从18.05%降低到接近0%,同时保持内容完整性,有助于增强与内容相关的语音处理任务的性能,并防止音色隐私泄露。
摘要:Speech pretrained models contain task-specific information across different layers, but decoupling content and timbre information remains challenging as removing speaker-specific information often causes content loss. Current research lacks direct metrics to quantify timbre residual in model encodings, relying on indirect evaluation through downstream tasks. This paper addresses these challenges through interpretability-based speaker disentanglement in speech pretraining models. We quantitatively evaluate timbre residual in model embeddings and improve speaker disentanglement using interpretive representations. Our contributions include: (1) InterpTRQE-SptME Benchmark - a timbre residual recognition framework using interpretability. The benchmark concatenates content embeddings with timbre embeddings for speaker classification, then applies Gradient SHAP Explainer to quantify timbre residual. We evaluate seven speech pretraining model variations. (2) InterpTF-SptME method - an interpretability-based timbre filtering approach using SHAP Noise and SHAP Cropping techniques. This model-agnostic method transforms intermediate encodings to remove timbre while preserving content. Experiments on VCTK dataset with HuBERT LARGE demonstrate successful content preservation and significant speaker disentanglement optimization. Results show the SHAP Noise method can reduce timbre residual from 18.05% to near 0% while maintaining content integrity, contributing to enhanced performance in content-related speech processing tasks and preventing timbre privacy leakage.
机器翻译由腾讯交互翻译提供,仅供参考
