微信公众号:arXiv_Daily
cs.SD语音
【1】Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
标题:Audio 2Face-3D:适合数字化身的音频驱动逼真面部动画
链接:https://arxiv.org/abs/2508.16401
摘要:Audio-driven facial animation presents an effective solution for animating digital avatars. In this paper, we detail the technical aspects of NVIDIA Audio2Face-3D, including data acquisition, network architecture, retargeting methodology, evaluation metrics, and use cases. Audio2Face-3D system enables real-time interaction between human users and interactive avatars, facilitating facial animation authoring for game characters. To assist digital avatar creators and game developers in generating realistic facial animations, we have open-sourced Audio2Face-3D networks, SDK, training framework, and example dataset.
【2】Vevo2: Bridging Controllable Speech and Singing Voice Generation via Unified Prosody Learning
标题:Vevo 2:通过统一韵律学习弥合可控语音和歌唱声音生成
链接:https://arxiv.org/abs/2508.16332
备注:We will release code and model checkpoints at this https URL
摘要:可控的人类声音生成,特别是对于像唱歌这样的表达领域,仍然是一个重大的挑战。本文介绍了Vevo 2,这是一个用于可控语音和歌声生成的统一框架。为了解决带注释的歌唱数据稀缺等问题并实现灵活的可控性,Vevo 2引入了两个音频标记器:(1)一个无音乐符号的韵律分词器,它从讲话、歌唱甚至乐器声音中捕捉韵律和旋律,以及(2)一个低帧率的韵律分词器,(12.5 Hz)内容风格标记器,其对语音和歌唱的语言内容、韵律和风格进行编码,同时实现音色分离。Vevo2包括一个自回归(AR)内容风格建模阶段,旨在实现对文本、韵律和风格的可控性,以及一个允许音色控制的流匹配声学建模阶段。特别地,在AR模型的预训练期间,我们提出了外显和内隐韵律学习策略来桥接语音和歌声。此外,为了进一步增强AR模型跟踪文本和韵律的能力,我们设计了一个多目标后训练任务,该任务集成了可理解性和韵律相似性对齐。实验结果表明,Vevo2中的统一建模对语音和歌声的生成都有很好的效果。此外,Vevo2在语音和歌唱的各种合成、转换和编辑任务中的有效性进一步证明了其强大的泛化能力和多功能性。音频样本可在https://versasinger.github.io/上获得。
摘要:Controllable human voice generation, particularly for expressive domains like singing, remains a significant challenge. This paper introduces Vevo2, a unified framework for controllable speech and singing voice generation. To tackle issues like the scarcity of annotated singing data and to enable flexible controllability, Vevo2 introduces two audio tokenizers: (1) a music-notation-free prosody tokenizer that captures prosody and melody from speech, singing, and even instrumental sounds, and (2) a low-frame-rate (12.5 Hz) content-style tokenizer that encodes linguistic content, prosody, and style for both speech and singing, while enabling timbre disentanglement. Vevo2 consists of an auto-regressive (AR) content-style modeling stage, which aims to enable controllability over text, prosody, and style, as well as a flow-matching acoustic modeling stage that allows for timbre control. Particularly, during pre-training of the AR model, we propose both explicit and implicit prosody learning strategies to bridge speech and singing voice. Moreover, to further enhance the AR model's ability to follow text and prosody, we design a multi-objective post-training task that integrates both intelligibility and prosody similarity alignment. Experimental results show that the unified modeling in Vevo2 brings mutual benefits to both speech and singing voice generation. Additionally, Vevo2's effectiveness across a wide range of synthesis, conversion, and editing tasks for both speech and singing further demonstrates its strong generalization ability and versatility. Audio samples are are available at https://versasinger.github.io/.
【3】Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
标题:亲眼所见:用于表达性语音生成的描述感知视听语言建模
链接:https://arxiv.org/abs/2508.16188
备注:EMNLP 2025 (Findings)
摘要:我们提出了一个视听语言模型(AVLM)的表达性语音生成集成到一个预先训练的表达性语音模型的全脸视觉线索。我们在预训练期间探索多个视觉编码器和多模态融合策略,以确定最有效的集成方法。随后对情感识别和表达性对话任务的微调产生了比仅语音基线(例如,+5 F1在情绪识别)。AVLM强调了表达性视觉信息在指导语音生成方面的价值,并为端到端多模态会话系统提供了基础。
摘要:We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during pre-training to identify the most effective integration approach. Subsequent fine-tuning on emotion recognition and expressive dialogue tasks yields substantial gains over speech-only baselines (e.g., +5 F1 in emotion recognition). AVLM highlights the value of expressive visual information in guiding speech generation and offers a foundation for end-to-end multimodal conversational systems.
【4】Head-Related Transfer Function Individualization Using Anthropometric Features and Spatially Independent Latent Representation
标题:使用人体测量特征和空间独立潜在表示的头部相关传递函数个性化
链接:https://arxiv.org/abs/2508.16176
备注:Accepted to IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:提出了一种根据人体测量参数进行头部相关传递函数(HRTF)个体化的方法。由于测量成本高,许多HRTF数据集中包含的受试者数量有限,包含人体测量参数的受试者数量甚至更少。因此,基于深度神经网络(DNN)的HRTF个性化是一项具有挑战性的任务。我们提出了一种HRTF个性化的方法,使用HRTF幅度的潜在表示通过自动编码器条件声源的位置,这使得有可能结合多个HRTF数据集与不同的测量源位置,并使网络训练易于处理,通过减少参数的数量来估计从人体测量参数。实验评估表明,高的估计精度是通过该方法实现的,相比目前基于DNN的方法。
摘要:A method for head-related transfer function (HRTF) individualization from the subject's anthropometric parameters is proposed. Due to the high cost of measurement, the number of subjects included in many HRTF datasets is limited, and the number of those that include anthropometric parameters is even smaller. Therefore, HRTF individualization based on deep neural networks (DNNs) is a challenging task. We propose a HRTF individualization method using the latent representation of HRTF magnitude obtained through an autoencoder conditioned on sound source positions, which makes it possible to combine multiple HRTF datasets with different measured source positions, and makes the network training tractable by reducing the number of parameters to be estimated from anthropometric parameters. Experimental evaluation shows that high estimation accuracy is achieved by the proposed method, compared to current DNN-based methods.
【5】QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
标题:基于差分相对属性学习的语音音色属性检测
链接:https://arxiv.org/abs/2508.15931
备注:Accepted by National Conference on Man-Machine Speech Communication, NCMMSC'2025
摘要:语音音色属性检测在语音生成任务的细粒度音色建模中起着关键作用。然而,由于音色描述符固有的主观性质和现有数据集中严重的标签不平衡,它仍然具有挑战性。在这项工作中,我们提出了Qvestival,一种新的成对比较框架的基础上差分注意,旨在提高建模的感知音色属性。为了解决VCTK-RVA数据集中的标签不平衡问题,我们引入了一种基于图的数据增强策略,该策略构建了一个有向无环图,并采用不相交集联合技术来自动挖掘具有有效属性比较的未观察话语对。我们的框架利用来自预训练的FACodec的扬声器嵌入,并结合了相对音色偏移感知差分注意力模块。该模块通过差分去噪和对比度放大机制显式地对成对话语之间的属性特定对比进行建模。在VCTK-RVA基准测试上的实验结果表明,Qvendor在多个音色描述符上实现了实质性的改进,特别是在跨说话人泛化场景中有显著的增益。
摘要:Voice Timbre Attribute Detection (vTAD) plays a pivotal role in fine-grained timbre modeling for speech generation tasks. However, it remains challenging due to the inherently subjective nature of timbre descriptors and the severe label imbalance in existing datasets. In this work, we present QvTAD, a novel pairwise comparison framework based on differential attention, designed to enhance the modeling of perceptual timbre attributes. To address the label imbalance in the VCTK-RVA dataset, we introduce a graph-based data augmentation strategy that constructs a Directed Acyclic Graph and employs Disjoint-Set Union techniques to automatically mine unobserved utterance pairs with valid attribute comparisons. Our framework leverages speaker embeddings from a pretrained FACodec, and incorporates a Relative Timbre Shift-Aware Differential Attention module. This module explicitly models attribute-specific contrasts between paired utterances via differential denoising and contrast amplification mechanisms. Experimental results on the VCTK-RVA benchmark demonstrate that QvTAD achieves substantial improvements across multiple timbre descriptors, with particularly notable gains in cross-speaker generalization scenarios.
【6】Beyond Transcription: Mechanistic Interpretability in ASR
标题:超越转录:ASB中的机械解释性
链接:https://arxiv.org/abs/2508.15882
摘要:可解释性方法最近得到了极大的关注,特别是在大型语言模型的背景下,使人们能够深入了解语言表示,错误检测和模型行为,如幻觉和重复。然而,这些技术在自动语音识别(ASR)中仍然没有得到充分的探索,尽管它们有可能提高ASR系统的性能和可解释性。在这项工作中,我们适应并系统地应用已建立的可解释性方法,如logit透镜,线性探测和激活修补,以研究声学和语义信息如何在ASR系统中跨层演变。我们的实验揭示了以前未知的内部动态,包括负责重复幻觉和语义偏见编码深内声学表示的特定编码器-解码器的相互作用。这些见解证明了将可解释性技术扩展和应用于语音识别的好处,为未来提高模型透明度和鲁棒性的研究开辟了有前途的方向。
摘要:Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error detection, and model behaviors such as hallucinations and repetitions. However, these techniques remain underexplored in automatic speech recognition (ASR), despite their potential to advance both the performance and interpretability of ASR systems. In this work, we adapt and systematically apply established interpretability methods such as logit lens, linear probing, and activation patching, to examine how acoustic and semantic information evolves across layers in ASR systems. Our experiments reveal previously unknown internal dynamics, including specific encoder-decoder interactions responsible for repetition hallucinations and semantic biases encoded deep within acoustic representations. These insights demonstrate the benefits of extending and applying interpretability techniques to speech recognition, opening promising directions for future research on improving model transparency and robustness.
【7】MGSC: A Multi-granularity Consistency Framework for Robust End-to-end Asr
标题:MGSC:一种多粒度的健壮端到端Asr一致性框架
链接:https://arxiv.org/abs/2508.15853
备注:12 pages, 5figures
摘要:端到端的ASR模型,尽管在基准测试中取得了成功,但在嘈杂的环境中往往会产生灾难性的语义错误。我们将这种脆弱性归因于流行的“直接映射”目标,该目标仅惩罚最终输出错误,同时使模型的内部计算过程不受约束。为了解决这个问题,我们引入了多粒度软一致性(MGSC)框架,一个模型不可知的,即插即用的模块,通过同时正则化宏层次的句子语义和微观层次的令牌对齐来执行内部的自我一致性。关键是,我们的工作是第一个发现这两个一致性粒度之间的强大协同作用:它们的联合优化产生的鲁棒性增益显着超过它们各自的贡献之和。在公共数据集上,MGSC在不同的噪声条件下将平均字符错误率降低了8.7%,主要是通过防止严重的意义改变错误。我们的工作表明,加强内部一致性是构建更强大和值得信赖的AI的关键一步。
摘要:End-to-end ASR models, despite their success on benchmarks, often pro-duce catastrophic semantic errors in noisy environments. We attribute this fragility to the prevailing 'direct mapping' objective, which solely penalizes final output errors while leaving the model's internal computational pro-cess unconstrained. To address this, we introduce the Multi-Granularity Soft Consistency (MGSC) framework, a model-agnostic, plug-and-play module that enforces internal self-consistency by simultaneously regulariz-ing macro-level sentence semantics and micro-level token alignment. Cru-cially, our work is the first to uncover a powerful synergy between these two consistency granularities: their joint optimization yields robustness gains that significantly surpass the sum of their individual contributions. On a public dataset, MGSC reduces the average Character Error Rate by a relative 8.7% across diverse noise conditions, primarily by preventing se-vere meaning-altering mistakes. Our work demonstrates that enforcing in-ternal consistency is a crucial step towards building more robust and trust-worthy AI.
【1】Hybrid Pruning: In-Situ Compression of Self-Supervised Speech Models for Speaker Verification and Anti-Spoofing
标题:混合修剪:自监督语音模型的现场压缩,用于说话人验证和反欺骗
链接:https://arxiv.org/abs/2508.16232
摘要:虽然像WavLM这样的大规模自监督学习(SSL)模型在语音处理方面已经达到了最先进的性能,但它们的巨大规模阻碍了在资源受限设备上的部署。虽然结构化修剪是模型压缩的关键技术,但现有方法通常将其与特定于任务的微调分开。这种多阶段方法努力创建针对不同下游任务量身定制的最佳架构。在这项工作中,我们引入了一个统一的框架,将结构化修剪到下游的微调过程。我们的框架统一了这些步骤,在单个阶段中联合优化任务性能和模型稀疏性。这允许模型学习专门用于最终任务的压缩架构,消除了对复杂的多级管道和知识蒸馏的需要。我们的修剪模型在大规模数据集上实现了高达70%的参数减少,性能下降可以忽略不计,在Vox 1-O,-E和-H上分别实现了0.7%,0.8%和1.6%的相等错误率。此外,我们的方法还证明了在低资源场景中改进的泛化能力,减少了过拟合并在ASVspoof 5上实现了最先进的3.7\% EER。
摘要:Although large-scale self-supervised learning (SSL) models like WavLM have achieved state-of-the-art performance in speech processing, their significant size impedes deployment on resource-constrained devices. While structured pruning is a key technique for model compression, existing methods typically separate it from task-specific fine-tuning. This multi-stage approach struggles to create optimal architectures tailored for diverse downstream tasks. In this work, we introduce a unified framework that integrates structured pruning into the downstream fine-tuning process. Our framework unifies these steps, jointly optimizing for task performance and model sparsity in a single stage. This allows the model to learn a compressed architecture specifically for the end task, eliminating the need for complex multi-stage pipelines and knowledge distillation. Our pruned models achieve up to a 70\% parameter reduction with negligible performance degradation on large-scale datasets, achieving equal error rates of 0.7\%, 0.8\%, and 1.6\% on Vox1-O, -E, and -H, respectively. Furthermore, our approach demonstrates improved generalization in low-resource scenarios, reducing overfitting and achieving a state-of-the-art 3.7\% EER on ASVspoof5.
【2】Robust Residual Finite Scalar Quantization for Neural Compression
标题:神经压缩的鲁棒剩余有限量量化
链接:https://arxiv.org/abs/2508.15860
备注:11 pages, 7 figures
摘要:有限标量量化(FSQ)已成为神经压缩中矢量量化(VQ)的一种有前途的替代方案,提供简化的训练和改进的稳定性。然而,FSQ在残差量化框架中的天真应用受到\textbf{残差幅度衰减问题}的困扰,其中后续FSQ层接收逐渐变弱的信号,严重限制了它们的有效性。我们提出了\textbf{Robust Residual Finite Scalar Quantization(RFSQ)},这是一个通用框架,通过两种新的调节策略解决了这一基本限制:可学习的缩放因子和可逆层归一化。我们的方法保持了FSQ的简单性,同时实现了有效的多级残差量化。在ImageNet上进行的综合实验表明,RFSQ变体的性能明显优于VQ-EMA、FSQ和LFQ等强基线,感知损失改善高达45%,L1重建错误减少28.7%。所提出的LayerNorm策略在不同的配置中表现出最一致的改进,将RFSQ确立为神经压缩的优越量化方法。
摘要:Finite Scalar Quantization (FSQ) has emerged as a promising alternative to Vector Quantization (VQ) in neural compression, offering simplified training and improved stability. However, naive application of FSQ in residual quantization frameworks suffers from the \textbf{residual magnitude decay problem}, where subsequent FSQ layers receive progressively weaker signals, severely limiting their effectiveness. We propose \textbf{Robust Residual Finite Scalar Quantization (RFSQ)}, a general framework that addresses this fundamental limitation through two novel conditioning strategies: learnable scaling factors and invertible layer normalization. Our approach maintains the simplicity of FSQ while enabling effective multi-stage residual quantization. Comprehensive experiments on ImageNet demonstrate that RFSQ variants significantly outperform strong baselines including VQ-EMA, FSQ, and LFQ, achieving up to 45\% improvement in perceptual loss and 28.7\% reduction in L1 reconstruction error. The proposed LayerNorm strategy shows the most consistent improvements across different configurations, establishing RFSQ as a superior quantization method for neural compression.
【3】A XAI-based Framework for Frequency Subband Characterization of Cough Spectrograms in Chronic Respiratory Disease
标题:基于XAI的慢性呼吸道疾病咳嗽谱图子带特征框架
链接:https://arxiv.org/abs/2508.16237
摘要:本文提出了一种基于可解释人工智能(XAI)的框架,用于与慢性呼吸系统疾病相关的咳嗽声的频谱分析,特别关注慢性阻塞性肺疾病(COPD)。卷积神经网络(CNN)在咳嗽信号的时间-频率表示上进行训练,并且遮挡图用于识别频谱图内的诊断相关区域。这些突出显示的区域随后被分解为五个频率子带,从而实现有针对性的光谱特征提取和分析。结果表明,频谱模式不同的子带和疾病组,揭示了整个频谱的互补和补偿趋势。值得注意的是,该方法基于可解释的光谱标志物将COPD与其他呼吸疾病区分开,并且将慢性患者组与非慢性患者组区分开。这些发现提供了对咳嗽声学的潜在病理生理学特征的深入了解,并证明了频率分辨、XAI增强分析对生物医学信号解释和转化呼吸疾病诊断的价值。
摘要:This paper presents an explainable artificial intelligence (XAI)-based framework for the spectral analysis of cough sounds associated with chronic respiratory diseases, with a particular focus on Chronic Obstructive Pulmonary Disease (COPD). A Convolutional Neural Network (CNN) is trained on time-frequency representations of cough signals, and occlusion maps are used to identify diagnostically relevant regions within the spectrograms. These highlighted areas are subsequently decomposed into five frequency subbands, enabling targeted spectral feature extraction and analysis. The results reveal that spectral patterns differ across subbands and disease groups, uncovering complementary and compensatory trends across the frequency spectrum. Noteworthy, the approach distinguishes COPD from other respiratory conditions, and chronic from non-chronic patient groups, based on interpretable spectral markers. These findings provide insight into the underlying pathophysiological characteristics of cough acoustics and demonstrate the value of frequency-resolved, XAI-enhanced analysis for biomedical signal interpretation and translational respiratory disease diagnostics.
【4】Head-Related Transfer Function Individualization Using Anthropometric Features and Spatially Independent Latent Representation
标题:使用人体测量特征和空间独立潜在表示的头部相关传递函数个性化
链接:https://arxiv.org/abs/2508.16176
备注:Accepted to IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA) 2025
摘要:提出了一种根据人体测量参数进行头部相关传递函数(HRTF)个体化的方法。由于测量成本高,许多HRTF数据集中包含的受试者数量有限,包含人体测量参数的受试者数量甚至更少。因此,基于深度神经网络(DNN)的HRTF个性化是一项具有挑战性的任务。我们提出了一种HRTF个性化的方法,使用HRTF幅度的潜在表示通过自动编码器条件声源的位置,这使得有可能结合多个HRTF数据集与不同的测量源位置,并使网络训练易于处理,通过减少参数的数量来估计从人体测量参数。实验评估表明,高的估计精度是通过该方法实现的,相比目前基于DNN的方法。
摘要:A method for head-related transfer function (HRTF) individualization from the subject's anthropometric parameters is proposed. Due to the high cost of measurement, the number of subjects included in many HRTF datasets is limited, and the number of those that include anthropometric parameters is even smaller. Therefore, HRTF individualization based on deep neural networks (DNNs) is a challenging task. We propose a HRTF individualization method using the latent representation of HRTF magnitude obtained through an autoencoder conditioned on sound source positions, which makes it possible to combine multiple HRTF datasets with different measured source positions, and makes the network training tractable by reducing the number of parameters to be estimated from anthropometric parameters. Experimental evaluation shows that high estimation accuracy is achieved by the proposed method, compared to current DNN-based methods.
【5】QvTAD: Differential Relative Attribute Learning for Voice Timbre Attribute Detection
标题:基于差分相对属性学习的语音音色属性检测
链接:https://arxiv.org/abs/2508.15931
备注:Accepted by National Conference on Man-Machine Speech Communication, NCMMSC'2025
摘要:语音音色属性检测在语音生成任务的细粒度音色建模中起着关键作用。然而,由于音色描述符固有的主观性质和现有数据集中严重的标签不平衡,它仍然具有挑战性。在这项工作中,我们提出了Qvestival,一种新的成对比较框架的基础上差分注意,旨在提高建模的感知音色属性。为了解决VCTK-RVA数据集中的标签不平衡问题,我们引入了一种基于图的数据增强策略,该策略构建了一个有向无环图,并采用不相交集联合技术来自动挖掘具有有效属性比较的未观察话语对。我们的框架利用来自预训练的FACodec的扬声器嵌入,并结合了相对音色偏移感知差分注意力模块。该模块通过差分去噪和对比度放大机制显式地对成对话语之间的属性特定对比进行建模。在VCTK-RVA基准测试上的实验结果表明,Qvendor在多个音色描述符上实现了实质性的改进,特别是在跨说话人泛化场景中有显著的增益。
摘要:Voice Timbre Attribute Detection (vTAD) plays a pivotal role in fine-grained timbre modeling for speech generation tasks. However, it remains challenging due to the inherently subjective nature of timbre descriptors and the severe label imbalance in existing datasets. In this work, we present QvTAD, a novel pairwise comparison framework based on differential attention, designed to enhance the modeling of perceptual timbre attributes. To address the label imbalance in the VCTK-RVA dataset, we introduce a graph-based data augmentation strategy that constructs a Directed Acyclic Graph and employs Disjoint-Set Union techniques to automatically mine unobserved utterance pairs with valid attribute comparisons. Our framework leverages speaker embeddings from a pretrained FACodec, and incorporates a Relative Timbre Shift-Aware Differential Attention module. This module explicitly models attribute-specific contrasts between paired utterances via differential denoising and contrast amplification mechanisms. Experimental results on the VCTK-RVA benchmark demonstrate that QvTAD achieves substantial improvements across multiple timbre descriptors, with particularly notable gains in cross-speaker generalization scenarios.
【6】Beyond Transcription: Mechanistic Interpretability in ASR
标题:超越转录:ASB中的机械解释性
链接:https://arxiv.org/abs/2508.15882
摘要:可解释性方法最近得到了极大的关注,特别是在大型语言模型的背景下,使人们能够深入了解语言表示,错误检测和模型行为,如幻觉和重复。然而,这些技术在自动语音识别(ASR)中仍然没有得到充分的探索,尽管它们有可能提高ASR系统的性能和可解释性。在这项工作中,我们适应并系统地应用已建立的可解释性方法,如logit透镜,线性探测和激活修补,以研究声学和语义信息如何在ASR系统中跨层演变。我们的实验揭示了以前未知的内部动态,包括负责重复幻觉和语义偏见编码深内声学表示的特定编码器-解码器的相互作用。这些见解证明了将可解释性技术扩展和应用于语音识别的好处,为未来提高模型透明度和鲁棒性的研究开辟了有前途的方向。
摘要:Interpretability methods have recently gained significant attention, particularly in the context of large language models, enabling insights into linguistic representations, error detection, and model behaviors such as hallucinations and repetitions. However, these techniques remain underexplored in automatic speech recognition (ASR), despite their potential to advance both the performance and interpretability of ASR systems. In this work, we adapt and systematically apply established interpretability methods such as logit lens, linear probing, and activation patching, to examine how acoustic and semantic information evolves across layers in ASR systems. Our experiments reveal previously unknown internal dynamics, including specific encoder-decoder interactions responsible for repetition hallucinations and semantic biases encoded deep within acoustic representations. These insights demonstrate the benefits of extending and applying interpretability techniques to speech recognition, opening promising directions for future research on improving model transparency and robustness.
【7】Mini-Omni-Reasoner: Token-Level Thinking-in-Speaking in Large Speech Models
标题:迷你全方位推理:大型语音模型中的代币级口语思维
链接:https://arxiv.org/abs/2508.15827
备注:Technical report; Work in progress. Project page: this https URL
摘要:推理对于有效的沟通和决策至关重要。虽然LLM和MLLM的最新进展表明,将显式推理显着提高理解和概括,LSM的推理仍然处于初级阶段。早期的努力试图将“先思后说”的范式从语篇模式转移到言语模式。然而,这种顺序公式化引入了显著的延迟,因为口语响应被延迟直到推理完全完成,从而损害了实时交互和通信效率。为了解决这个问题,我们提出了迷你全方位推理,一个框架,使推理的语音通过一个新的“思维在说话”的制定。Mini-Omni-Reasoner不是在产生任何口头输出之前完成推理,而是在令牌级别将无声推理令牌与口头响应令牌交织在一起。这种设计允许连续语音生成,同时嵌入结构化内部推理,利用模型的高频令牌处理能力。虽然是交错的,但局部语义对齐是强制执行的,以确保每个响应令牌都由其先前的推理通知。为了支持这个框架,我们引入了Spoken-Math-Problems-3 M,这是一个为交错推理和响应量身定制的大规模数据集。该数据集确保语言标记始终遵循相关的推理内容,从而能够准确有效地学习语音耦合推理。Mini-Omni-Reasoner建立在层次化的思想者-说话者架构之上,提供流畅但逻辑上有基础的口语回答,保持自然和精确。在Spoken-MQA基准测试中,它在算术推理方面获得了+19.1%的增益,在上下文理解方面获得了+6.4%的增益,输出更短,解码延迟为零。
摘要:Reasoning is essential for effective communication and decision-making. While recent advances in LLMs and MLLMs have shown that incorporating explicit reasoning significantly improves understanding and generalization, reasoning in LSMs remains in a nascent stage. Early efforts attempt to transfer the "Thinking-before-Speaking" paradigm from textual models to speech. However, this sequential formulation introduces notable latency, as spoken responses are delayed until reasoning is fully completed, impairing real-time interaction and communication efficiency. To address this, we propose Mini-Omni-Reasoner, a framework that enables reasoning within speech via a novel "Thinking-in-Speaking" formulation. Rather than completing reasoning before producing any verbal output, Mini-Omni-Reasoner interleaves silent reasoning tokens with spoken response tokens at the token level. This design allows continuous speech generation while embedding structured internal reasoning, leveraging the model's high-frequency token processing capability. Although interleaved, local semantic alignment is enforced to ensure that each response token is informed by its preceding reasoning. To support this framework, we introduce Spoken-Math-Problems-3M, a large-scale dataset tailored for interleaved reasoning and response. The dataset ensures that verbal tokens consistently follow relevant reasoning content, enabling accurate and efficient learning of speech-coupled reasoning. Built on a hierarchical Thinker-Talker architecture, Mini-Omni-Reasoner delivers fluent yet logically grounded spoken responses, maintaining both naturalness and precision. On the Spoken-MQA benchmark, it achieves a +19.1% gain in arithmetic reasoning and +6.4% in contextual understanding, with shorter outputs and zero decoding latency.
机器翻译由腾讯交互翻译提供,仅供参考
