微信公众号:arXiv_Daily
cs.SD语音
【1】Steer-MoE: Efficient Audio-Language Alignment with a Mixture-of-Experts Steering Module
标题:Steer-MoE:通过专家混合指导模块实现高效的音频语言协调
链接:https://arxiv.org/abs/2510.13558
备注:5 pages, 1 figures. Code is available at: this https URL. Submitted to ICASSP 2026
摘要:对齐预训练的音频编码器和大型语言模型(LLM)为构建强大的多模态代理提供了一条有前途的参数高效路径。然而,现有的方法通常需要昂贵的全模型微调或依赖于静态适配器,可能缺乏表达能力。从柏拉图表征假说的灵感,我们介绍SteerMoE,一种新颖的和模块化的框架,音频语言对齐。SteerMoE冻结音频编码器和LLM解码器,仅训练集成在编码器层中的轻量级转向模块。该模块使用混合专家(MoE)路由器来动态地选择和应用学习的导向向量,逐步地将连续音频表示转换为LLM可理解的空间。通过完全在连续嵌入空间中操作,我们的方法不需要修改LLM的词汇表,并保留其先进的推理和代理能力。我们通过ASR,音频理解和定性函数调用任务的实验证明,SteerMoE实现了强大的性能,同时保持高度模块化和计算效率,为开发复杂的音频语言系统提供了一个强大的新范式。
摘要:Aligning pretrained audio encoders and Large Language Models (LLMs) offers a promising, parameter-efficient path to building powerful multimodal agents. However, existing methods often require costly full-model finetuning or rely on static adapters that may lack expressive power. Drawing inspiration from the Platonic Representation Hypothesis, we introduce SteerMoE, a novel and modular framework for audio-language alignment. SteerMoE freezes both the audio encoder and the LLM decoder, training only a lightweight steering module integrated within the encoder's layers. This module uses a Mixture-of-Experts (MoE) router to dynamically select and apply learned steering vectors, progressively transforming continuous audio representations into a space comprehensible to the LLM. By operating entirely in the continuous embedding space, our approach requires no modifications to the LLM's vocabulary and preserves its advanced reasoning and agentic capabilities. We demonstrate through experiments on ASR, audio understanding, and a qualitative function-calling task that SteerMoE achieves strong performance while remaining highly modular and computationally efficient, offering a robust new paradigm for developing sophisticated audio-language systems.
【2】UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
标题:UniMoE-Audio:具有动态容量MoE的统一语音和音乐生成
链接:https://arxiv.org/abs/2510.13344
摘要:统一的多模态模型的最新进展表明,一个明显的趋势,全面的内容生成。然而,听觉领域仍然是一个重大的挑战,音乐和语音往往是孤立发展的,阻碍了通用音频合成的进展。这种分离源于固有的任务冲突和严重的数据不平衡,这阻碍了真正统一的音频生成模型的发展。为了解决这一挑战,我们提出了UniMoE音频,一个统一的语音和音乐生成模型在一个新的动态容量混合专家(MoE)的框架。在架构上,UniMoE-Audio引入了Top-P路由策略用于动态专家数量分配,以及混合专家设计,包括用于特定领域知识的路由专家,用于领域不可知特征的共享专家,以及用于自适应计算跳过的空专家。为了解决数据不平衡的问题,我们引入了一个三阶段的训练课程:1)独立专家训练利用原始数据集将特定领域的知识不受干扰地灌输给每个“原型专家”; 2)MoE集成和预热将这些专家纳入UniMoE-Audio架构,使用平衡数据集的子集预热门模块和共享专家;协同联合训练在完全平衡的数据集上对整个模型进行端到端的训练,促进增强的跨域协同。大量的实验表明,UniMoE-Audio不仅在主要的语音和音乐生成基准上实现了最先进的性能,而且还展示了卓越的协同学习,减轻了通常在朴素联合训练中看到的性能下降。我们的研究结果突出了专业MoE架构和策划培训策略在推进通用音频生成领域的巨大潜力。主页:https://mukioxun.github.io/Uni-MoE-site/home.html
摘要:Recent advances in unified multimodal models indicate a clear trend towards comprehensive content generation. However, the auditory domain remains a significant challenge, with music and speech often developed in isolation, hindering progress towards universal audio synthesis. This separation stems from inherent task conflicts and severe data imbalances, which impede the development of a truly unified audio generation model. To address this challenge, we propose UniMoE-Audio, a unified speech and music generation model within a novel Dynamic-Capacity Mixture-of-Experts (MoE) framework. Architecturally, UniMoE-Audio introduces a Top-P routing strategy for dynamic expert number allocation, and a hybrid expert design comprising routed experts for domain-specific knowledge, shared experts for domain-agnostic features, and null experts for adaptive computation skipping. To tackle data imbalance, we introduce a three-stage training curriculum: 1) Independent Specialist Training leverages original datasets to instill domain-specific knowledge into each "proto-expert" without interference; 2) MoE Integration and Warmup incorporates these specialists into the UniMoE-Audio architecture, warming up the gate module and shared expert using a subset of balanced dataset; and 3) Synergistic Joint Training trains the entire model end-to-end on the fully balanced dataset, fostering enhanced cross-domain synergy. Extensive experiments show that UniMoE-Audio not only achieves state-of-the-art performance on major speech and music generation benchmarks, but also demonstrates superior synergistic learning, mitigating the performance degradation typically seen in naive joint training. Our findings highlight the substantial potential of specialized MoE architecture and curated training strategies in advancing the field of universal audio generation. Homepage: https://mukioxun.github.io/Uni-MoE-site/home.html
【3】MotionBeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding
标题:MotionBeat:通过同步对比学习和Bar等变接触感知编码的运动对齐音乐表示
链接:https://arxiv.org/abs/2510.13244
备注:5 pages, 1 figure. demo page: this https URL
摘要:音乐既是一种听觉现象,也是一种身体现象,与人体运动密切相关,并通过舞蹈自然表达。然而,大多数现有的音频表示忽略了这一具体的维度,限制了它们捕捉驱动运动的节奏和结构线索的能力。我们提出了MotionBeat,一个用于运动对齐音乐表示学习的框架。MotionBeat使用两个新提出的目标进行训练:增强的对比损失(ECL),一种增强的InfoNCE公式,具有节奏感知和节拍抖动负性,以实现细粒度的节奏辨别,以及结构节奏对齐损失(SRAL),通过将音乐重音与相应的运动事件对齐来确保节奏一致性。在架构上,MotionBeat引入了条等变相位旋转来捕捉周期性的节奏模式和接触引导的注意力,以强调与音乐口音同步的运动事件。实验表明,MotionBeat在音乐到舞蹈生成方面优于最先进的音频编码器,并有效地转移到节拍跟踪,音乐标记,流派和乐器分类,情感识别和视听检索。我们的项目演示页面:https://motionbeat2025.github.io/。
摘要:Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect this embodied dimension, limiting their ability to capture rhythmic and structural cues that drive movement. We propose MotionBeat, a framework for motion-aligned music representation learning. MotionBeat is trained with two newly proposed objectives: the Embodied Contrastive Loss (ECL), an enhanced InfoNCE formulation with tempo-aware and beat-jitter negatives to achieve fine-grained rhythmic discrimination, and the Structural Rhythm Alignment Loss (SRAL), which ensures rhythm consistency by aligning music accents with corresponding motion events. Architecturally, MotionBeat introduces bar-equivariant phase rotations to capture cyclic rhythmic patterns and contact-guided attention to emphasize motion events synchronized with musical accents. Experiments show that MotionBeat outperforms state-of-the-art audio encoders in music-to-dance generation and transfers effectively to beat tracking, music tagging, genre and instrument classification, emotion recognition, and audio-visual retrieval. Our project demo page: https://motionbeat2025.github.io/.
【4】VCTR: A Transformer-Based Model for Non-parallel Voice Conversion
标题:VCTF:一种基于转换器的非并行语音转换模型
链接:https://arxiv.org/abs/2510.12964
摘要:非并行语音转换的目标是在没有成对训练数据的情况下将语音从源域转换到目标域。循环一致生成对抗网络(CycleGAN)和变分自编码器(VAE)已用于此任务,但这些模型存在训练困难和结果不满意的问题。后来,引入了对比语音转换(CVC),利用基于对比学习的方法来解决这些问题。然而,这些方法使用基于CNN的生成器,它可以捕获局部语义,但缺乏捕获全局语义所需的长期依赖关系的能力。在本文中,我们提出了VCTR,这是一种有效的非并行语音转换方法,它利用了混合感知块(HPB)和双修剪自注意力(DPSA)以及基于对比学习的对抗方法。代码可以在https://github.com/Maharnab-Saikia/VCTR中找到。
摘要:Non-parallel voice conversion aims to convert voice from a source domain to a target domain without paired training data. Cycle-Consistent Generative Adversarial Networks (CycleGAN) and Variational Autoencoders (VAE) have been used for this task, but these models suffer from difficult training and unsatisfactory results. Later, Contrastive Voice Conversion (CVC) was introduced, utilizing a contrastive learning-based approach to address these issues. However, these methods use CNN-based generators, which can capture local semantics but lacks the ability to capture long-range dependencies necessary for global semantics. In this paper, we propose VCTR, an efficient method for non-parallel voice conversion that leverages the Hybrid Perception Block (HPB) and Dual Pruned Self-Attention (DPSA) along with a contrastive learning-based adversarial approach. The code can be found in https://github.com/Maharnab-Saikia/VCTR.
【5】A Critical Review of the Need for Knowledge-Centric Evaluation of Quranic Recitation
标题:对《古兰经》背诵以知识为中心评估的必要性的批判性评论
链接:https://arxiv.org/abs/2510.12858
备注:33 pages
摘要:古兰经背诵(Tajweed)的神圣实践,受到精确的语音,韵律和神学规则的支配,在现代面临着重大的教学挑战。虽然数字技术承诺前所未有的教育机会,但用于背诵评估的自动化工具未能实现广泛采用或教学效果。这篇文献综述调查了这一关键差距,对过去二十年来开发的学术研究,网络平台和商业应用程序进行了全面分析。我们的综合揭示了一个根本的错位,在现行的方法,重新调整自动语音识别(ASR)架构,优先词汇识别定性声学评估和困扰的数据依赖性,人口统计学偏见,并无法提供诊断有用的反馈。批评这些数据驱动的范式,我们认为一个基本的范式转向以知识为中心的计算框架。利用古兰经文本的不可变性和Tajweed精确定义的规则,我们建议一个强大的评估器必须围绕基于规范规则和衔接点(Makhraj)的预期声学建模进行构建,而不是依赖于从不完美和有偏见的数据集中学习到的统计模式。这篇评论的结论是,自动化古兰经评估的未来在于将深层语言知识与先进的音频分析相结合的混合系统,提供了一条通往强大,公平和教学合理的工具的道路,可以忠实地支持世界各地的学习者。
摘要:The sacred practice of Quranic recitation (Tajweed), governed by precise phonetic, prosodic, and theological rules, faces significant pedagogical challenges in the modern era. While digital technologies promise unprecedented access to education, automated tools for recitation evaluation have failed to achieve widespread adoption or pedagogical efficacy. This literature review investigates this critical gap, conducting a comprehensive analysis of academic research, web platforms, and commercial applications developed over the past two decades. Our synthesis reveals a fundamental misalignment in prevailing approaches that repurpose Automatic Speech Recognition (ASR) architectures, which prioritize lexical recognition over qualitative acoustic assessment and are plagued by data dependency, demographic biases, and an inability to provide diagnostically useful feedback. Critiquing these data--driven paradigms, we argue for a foundational paradigm shift towards a knowledge-centric computational framework. Capitalizing on the immutable nature of the Quranic text and the precisely defined rules of Tajweed, we propose that a robust evaluator must be architected around anticipatory acoustic modeling based on canonical rules and articulation points (Makhraj), rather than relying on statistical patterns learned from imperfect and biased datasets. This review concludes that the future of automated Quranic evaluation lies in hybrid systems that integrate deep linguistic knowledge with advanced audio analysis, offering a path toward robust, equitable, and pedagogically sound tools that can faithfully support learners worldwide.
【6】Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
标题:自适应向量转向:一种免训练、分层干预,用于缓解大型音频和多模式模型中的幻觉
链接:https://arxiv.org/abs/2510.12851
备注:Note: This preprint is a version of the paper submitted to ICASSP 2026. The author list here includes contributors who provided additional supervision and guidance. The official ICASSP submission may differ slightly in author composition
摘要:大型音频语言模型和多模态大型语言模型在音频问题分类(AQA)、音频字幕和自动语音识别(ASR)等任务中表现出强大的能力。然而,越来越多的证据表明,这些模型可以对音频内容产生幻觉。为了解决这个问题,我们探测模型的内部状态,并提出自适应矢量转向(AVS),一种更好地在音频内容中生成的方法。我们还确定了输出正确性和内部表示之间的强相关性。实验表明,在两个模型和两个基准一致的性能增益。在音频幻觉QA数据集上,我们的方法将Gemma的F1分数从0.550提高到0.619,Qwen从0.626提高到0.632。此外,我们的方法将Qwen对MMAU的准确度从0.548提高到0.592,相对提高了8%。据我们所知,这是第一个应用矢量转向来减轻音频中的幻觉的工作。
摘要:Large Audio-Language Models and Multi-Modal Large Language Models have demonstrated strong capabilities in tasks such as Audio Question Answering (AQA), Audio Captioning, and Automatic Speech Recognition (ASR). However, there is growing evidence that these models can hallucinate about the content of the audio. To address this issue, we probe the models' internal states and propose Adaptive Vector Steering (AVS), a method that better grounds generation in audio content. We also identify a strong correlation between output correctness and internal representations. Experiments show consistent performance gains across two models and two benchmarks. On the Audio Hallucination QA dataset, our method boosts the F1-score of Gemma from 0.550 to 0.619 and Qwen from 0.626 to 0.632. Furthermore, our method increases the accuracy of Qwen on MMAU from 0.548 to 0.592, marking an 8% relative increase. To the best of our knowledge, this is the first work to apply vector steering to mitigate hallucination in audio.
【7】Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
标题:Gelina:通过交织令牌预测统一语音和手势合成
链接:https://arxiv.org/abs/2510.12834
备注:5 pages
摘要:人类交流是多模态的,语音和手势紧密耦合,但大多数用于生成语音和手势的计算方法按顺序合成它们,削弱了同步和韵律对齐。我们介绍Gelina,一个统一的框架,联合合成语音和语音手势从文本中使用交错令牌序列在离散自回归骨干,与特定于模态的解码器。Gelina支持多说话者和多风格克隆,并支持从语音输入进行仅手势合成。主观和客观的评价表明竞争力的语音质量和改进的手势生成超过单峰基线。
摘要:Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines.
【8】Production and Manufacturing of 3D Printed Acoustic Guitars
标题:3D打印原声吉他的生产制造
链接:https://arxiv.org/abs/2510.12823
摘要:这项研究调查了使用3D打印生产经济实惠的功能性原声吉他的可行性,重点是生产具有适当音调性能的结构设计。该研究与William Schiesser合作进行,使用古典吉他模型,选择其较低的弦张力,以评估由聚乳酸(PLA)制成的3D打印原型的音调特征。由于Prusa Mark 4打印机的构造板尺寸限制,吉他主体被分成多个部分,以压配合公差和最小的氰基丙烯酸酯粘合剂连接。Fusion 360中的CAD建模确保了压配连接和整体装配的尺寸精度。组装完成后,吉他被用尼龙弦串起来,并使用Audacity软件进行测试,将记录的频率和音符与标准参考值进行比较。结果表明,在较低的字符串频率,可能是由印刷中使用的材料选择造成的大的偏差。所有琴弦都达到了精确的音高,尽管通过调音存在频率差异,这表明尽管存在不可避免的挑战,但PLA和现代制造方法可以生产出价格合理、可演奏的原声吉他。进一步的研究可能会调查替代塑料的优良频率匹配。这种方法具有巨大的潜力,可以扩大获得优质乐器的机会,同时减少对濒危音木的依赖,从而鼓励可持续的乐器生产和增加音乐参与。这也为获得乐器仍然是一项挑战的弱势社区创造了机会。 关键词:制琴师,光固化成型,3D打印,吉他制作
摘要:This research investigates the feasibility of producing affordable, functional acoustic guitars using 3D printing, with a focus on producing structural designs with proper tonal performance. Conducted in collaboration with William Schiesser, the study uses a classical guitar model, chosen for its lower string tension, to evaluate the tonal characteristics of a 3D-printed prototype made from polylactic acid (PLA). Due to the build plate size constraints of the Prusa Mark 4 printer, the guitar body was divided into multiple sections joined with press-fit tolerances and minimal cyanoacrylate adhesive. CAD modeling in Fusion 360 ensured dimensional accuracy in press-fit connections and the overall assembly. Following assembly, the guitar was strung with nylon strings and tested using Audacity software to compare recorded frequencies and notes with standard reference values. Results showed large deviations in lower string frequencies, likely caused by the material choice utilized in printing. Accurate pitches were reached with all strings despite frequency differences through tuning, demonstrating that PLA and modern manufacturing methods can produce affordable, playable acoustic guitars despite inevitable challenges. Further research may investigate alternative plastics for superior frequency matching. This approach holds significant potential for expanding access to quality instruments while reducing reliance on endangered tonewoods, thereby encouraging both sustainable instrument production and increased musical participation. This also creates opportunities for disadvantaged communities where access to musical instruments remains a challenge. Keywords: Luthiery, Stereolithography, 3D-Print, Guitar Making
【9】Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
标题:超越离散类别:用于宠物发声分析的多任务价-唤醒建模
链接:https://arxiv.org/abs/2510.12819
备注:24 pages, 6 figures, 4 tables. First continuous VA framework for pet vocalization analysis with 42,553 samples
摘要:传统的宠物情感识别发声,基于离散分类,与模糊性和捕捉强度变化的斗争。我们提出了一个连续的效价唤醒(VA)模型,表示在一个二维空间的情绪。我们的方法使用自动VA标签生成算法,能够对42,553个宠物发声样本进行大规模注释。多任务学习框架将VA回归与辅助任务(情感、体型、性别)联合训练,以通过改进特征学习来增强预测。我们的Audio Transformer模型实现了r = 0.9024和r = 0.7155的验证Valence Pearson相关性,有效地解决了“领土”和“快乐”等离散类别之间的混淆。“这项工作引入了第一个用于宠物发声分析的连续VA框架,为人类与宠物的互动,兽医诊断和行为训练提供了更具表现力的表示。该方法显示出在消费产品中部署的强大潜力,如AI宠物情感翻译器。
摘要:Traditional pet emotion recognition from vocalizations, based on discrete classification, struggles with ambiguity and capturing intensity variations. We propose a continuous Valence-Arousal (VA) model that represents emotions in a two-dimensional space. Our method uses an automatic VA label generation algorithm, enabling large-scale annotation of 42,553 pet vocalization samples. A multi-task learning framework jointly trains VA regression with auxiliary tasks (emotion, body size, gender) to enhance prediction by improving feature learning. Our Audio Transformer model achieves a validation Valence Pearson correlation of r = 0.9024 and an Arousal r = 0.7155, effectively resolving confusion between discrete categories like "territorial" and "happy." This work introduces the first continuous VA framework for pet vocalization analysis, offering a more expressive representation for human-pet interaction, veterinary diagnostics, and behavioral training. The approach shows strong potential for deployment in consumer products like AI pet emotion translators.
【10】Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
标题:多模式LLM中演讲者参考的TTC的连续令牌扩散
链接:https://arxiv.org/abs/2510.12995
摘要:多模态大型语言模型(MLLM)中的统一架构在单个框架内处理不同任务方面表现出了希望。在文本到语音(TTS)的任务,目前MLLM为基础的方法依赖于离散的令牌表示,忽视了语音的固有连续性,并可能导致损失的细粒度的声学信息。在这项工作中,我们调查的TTS MLLM范式使用连续的语音表示。我们设计了一个双头架构,并实现了两个互补的训练策略,一个强大的模型。(1)在MLLM上增加了一个产生连续语音表示的扩散头,它是帧级的,并且是严格自回归的。(2)原始语言模型头被保留以保持多任务能力并控制语音合成的开始和结束。(3)采用掩蔽训练来解决自回归解码中的暴露偏差。(4)为了稳定优化,我们提出了一个两阶段的方案,其中LM在第二阶段被冻结,确保扩散头从固定的输入分布中学习。LibriSpeech(PC)测试清洁的评估表明,我们的方法实现了最先进的自回归性能,WER为1.95%,说话人相似度为0.54,UTMOS为4.00。两阶段训练比一阶段训练基线产生46%的相对WER降低。这些结果突出了自回归建模与连续令牌扩散相结合的有效性,由两阶段训练过程支持。
摘要:Unified architectures in multimodal large language models (MLLM) have shown promise in handling diverse tasks within a single framework. In the text-to-speech (TTS) task, current MLLM-based approaches rely on discrete token representations, which disregard the inherently continuous nature of speech and can lead to loss of fine-grained acoustic information.In this work, we investigate the TTS within the MLLM paradigm using continuous speech representations. We design a dual-head architecture and implement two complementary training strategies for a robust model. (1) A diffusion head generating continuous speech representations is added on the MLLM, which is on frame-level and strictly autoregressive. (2) The original language model head is retained to preserve multitask capability and to control the start and end of speech synthesis. (3) Masked training is employed to address exposure bias in autoregressive decoding. (4) To stabilize optimization, we propose a two-stage scheme where the LM is frozen in the second stage, ensuring the diffusion head learns from a fixed input distribution. Evaluations on LibriSpeech(PC) test-clean show that our approach achieves state-of-the-art autoregressive performance, with a WER of 1.95%, speaker similarity of 0.54, and UTMOS of 4.00. The two-stage training yields a 46% relative WER reduction over the one-stage training baseline. These results highlight the effectiveness of combining autoregressive modeling with continuous-token diffusion, supported by a two-stage training procedure.
【11】HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
标题:HyWA:超网络重量自适应个性化语音活动检测
链接:https://arxiv.org/abs/2510.12947
备注:Mahsa Ghazvini Nejad and Hamed Jafarzadeh Asl contributed equally to this work
摘要:个性化语音活动检测(PVAD)系统通过结合来自登记话语的说话者嵌入而仅响应于特定目标说话者而激活。与现有的方法,需要架构的变化,如电影层,我们的方法采用了超网络来修改一个标准的语音活动检测(VAD)模型中的几个选定的层的权重。这使得扬声器调节无需改变VAD架构,允许相同的VAD模型通过仅更新一小部分层来适应不同的扬声器。我们提出了HyWA-PVAD,超网络的权重自适应方法,并评估它对多个基线条件技术。我们的比较显示PVAD性能的持续改进。HyWA还通过保留核心VAD架构为部署提供了实际优势。我们的新方法在两个方面改进了当前的条件技术:i)提高了平均精度,ii)通过重用相同的VAD架构简化了部署。
摘要:Personalized Voice Activity Detection (PVAD) systems activate only in response to a specific target speaker by incorporating speaker embeddings from enrollment utterances. Unlike existing methods that require architectural changes, such as FiLM layers, our approach employs a hypernetwork to modify the weights of a few selected layers within a standard voice activity detection (VAD) model. This enables speaker conditioning without changing the VAD architecture, allowing the same VAD model to adapt to different speakers by updating only a small subset of the layers. We propose HyWA-PVAD, a hypernetwork weight adaptation method, and evaluate it against multiple baseline conditioning techniques. Our comparison shows consistent improvements in PVAD performance. HyWA also offers practical advantages for deployment by preserving the core VAD architecture. Our new approach improves the current conditioning techniques in two ways: i) increases the mean average precision, ii) simplifies deployment by reusing the same VAD architecture.
【12】Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
标题:现代的自动语音识别:架构、训练和评估
链接:https://arxiv.org/abs/2510.12827
摘要:在过去十年中,自动语音识别(ASR)在深度学习的推动下经历了深刻的变革。本调查全面概述了ASR的现代时代,描绘了其从传统混合系统(如高斯混合模型-隐马尔可夫模型(GMM-Hyndom)和深度神经网络-Hyndom(DNN-Hyndom))到现在占主导地位的端到端神经架构的演变。我们系统地回顾了基本的端到端范例:连接主义时间分类(CTC),基于注意力的编码器-解码器模型和递归神经网络转换器(RNN-T),它们为完全集成的语音到文本系统奠定了基础。然后,我们详细介绍了随后的架构转向Transformer和Conformer模型,利用自我注意力来捕获高计算效率的长期依赖关系。本次调查的一个中心主题是培训模式的平行革命。我们研究了从完全监督学习(通过SpecAugment等技术增强)到自我监督学习(SSL)的兴起以及wav 2 vec 2.0等基础模型的进展,这些模型大大减少了对转录数据的依赖。此外,我们分析了像Whisper这样的大规模弱监督模型的影响,这些模型通过大量数据多样性实现了前所未有的鲁棒性。该文件还涵盖了生态系统的基本组成部分,包括关键数据集和基准(例如,LibriSpeech、Switchboard、CHiME)、标准评估度量(例如,字错误率),以及现实世界部署的关键考虑因素,如流式推理,设备效率以及公平性和鲁棒性的道德要求。最后,我们概述了开放的挑战和未来的研究方向。
摘要:Automatic Speech Recognition (ASR) has undergone a profound transformation over the past decade, driven by advances in deep learning. This survey provides a comprehensive overview of the modern era of ASR, charting its evolution from traditional hybrid systems, such as Gaussian Mixture Model-Hidden Markov Models (GMM-HMMs) and Deep Neural Network-HMMs (DNN-HMMs), to the now-dominant end-to-end neural architectures. We systematically review the foundational end-to-end paradigms: Connectionist Temporal Classification (CTC), attention-based encoder-decoder models, and the Recurrent Neural Network Transducer (RNN-T), which established the groundwork for fully integrated speech-to-text systems. We then detail the subsequent architectural shift towards Transformer and Conformer models, which leverage self-attention to capture long-range dependencies with high computational efficiency. A central theme of this survey is the parallel revolution in training paradigms. We examine the progression from fully supervised learning, augmented by techniques like SpecAugment, to the rise of self-supervised learning (SSL) with foundation models such as wav2vec 2.0, which drastically reduce the reliance on transcribed data. Furthermore, we analyze the impact of largescale, weakly supervised models like Whisper, which achieve unprecedented robustness through massive data diversity. The paper also covers essential ecosystem components, including key datasets and benchmarks (e.g., LibriSpeech, Switchboard, CHiME), standard evaluation metrics (e.g., Word Error Rate), and critical considerations for real-world deployment, such as streaming inference, on-device efficiency, and the ethical imperatives of fairness and robustness. We conclude by outlining open challenges and future research directions.
【1】Towards Multimodal Query-Based Spatial Audio Source Extraction
标题:基于多模式查询的空间音频源提取
链接:https://arxiv.org/abs/2510.13308
备注:Submitted to ICASSP 2026
摘要:基于查询的音频源提取试图从以查询为条件的混合中恢复目标源。现有的方法在很大程度上局限于单声道音频,留下的空间信息在多声道录音未充分利用。我们引入了一个基于查询的空间音频源提取框架,用于从一阶立体混响(FOA)混合信号中恢复干目标信号。我们的方法接受音频提示或文本提示作为条件输入,从而实现灵活的端到端提取。我们所提出的模型的核心在于一个三轴Transformer,联合建模的时间,频率和空间通道的依赖性。该模型使用对比语言音频预训练(CLAP)嵌入,通过特征线性调制(FILM)实现统一的音频文本调节。为了消除昂贵的注释并提高泛化能力,我们提出了一种无标签的数据管道,可以动态生成空间混合和相应的训练目标。高分离质量的实验结果证明了多峰调节和三轴建模的有效性。这项工作为沉浸式应用中的高保真空间音频分离建立了一个新的范例。
摘要:Query-based audio source extraction seeks to recover a target source from a mixture conditioned on a query. Existing approaches are largely confined to single-channel audio, leaving the spatial information in multi-channel recordings underexploited. We introduce a query-based spatial audio source extraction framework for recovering dry target signals from first-order ambisonics (FOA) mixtures. Our method accepts either an audio prompt or a text prompt as condition input, enabling flexible end-to-end extraction. The core of our proposed model lies in a tri-axial Transformer that jointly models temporal, frequency, and spatial channel dependencies. The model uses contrastive language-audio pretraining (CLAP) embeddings to enable unified audio-text conditioning via feature-wise linear modulation (FiLM). To eliminate costly annotations and improve generalization, we propose a label-free data pipeline that dynamically generates spatial mixtures and corresponding targets for training. The result of our experiment with high separation quality demonstrates the efficacy of multimodal conditioning and tri-axial modeling. This work establishes a new paradigm for high-fidelity spatial audio separation in immersive applications.
【2】Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
标题:两个头比一个头好:双重假设下的视听语音错误纠正
链接:https://arxiv.org/abs/2510.13281
备注:Preprint work
摘要:本文介绍了一种新的生成纠错框架的视听语音识别(AVSR)的原因,在特定模态的证据直接在语言空间。我们的框架,DualHyp,授权一个大型语言模型(LLM)组成独立的N-最好的假设,从单独的自动语音识别(ASR)和视觉语音识别(VSR)模型。为了最大限度地提高DualHyp的有效性,我们进一步引入了RelPrompt,这是一种噪声感知指导机制,为LLM提供模态接地提示。RelPrompt提供了每个模态流的时间可靠性,引导模型在ASR和VSR假设之间动态切换焦点,以进行准确的校正。在各种腐败情况下,我们的框架在LRS 2基准测试中获得了高达57.7%的错误率增益,而单流GER方法仅获得10%的增益。为了促进我们的DualHyp框架内的研究,我们在https://github.com/sungnyun/dualhyp上发布了代码和包括ASR和VSR假设的数据集。
摘要:This paper introduces a new paradigm for generative error correction (GER) framework in audio-visual speech recognition (AVSR) that reasons over modality-specific evidences directly in the language space. Our framework, DualHyp, empowers a large language model (LLM) to compose independent N-best hypotheses from separate automatic speech recognition (ASR) and visual speech recognition (VSR) models. To maximize the effectiveness of DualHyp, we further introduce RelPrompt, a noise-aware guidance mechanism that provides modality-grounded prompts to the LLM. RelPrompt offers the temporal reliability of each modality stream, guiding the model to dynamically switch its focus between ASR and VSR hypotheses for an accurate correction. Under various corruption scenarios, our framework attains up to 57.7% error rate gain on the LRS2 benchmark over standard ASR baseline, contrary to single-stream GER approaches that achieve only 10% gain. To facilitate research within our DualHyp framework, we release the code and the dataset comprising ASR and VSR hypotheses at https://github.com/sungnyun/dualhyp.
【3】Acoustic Teleportation via Disentangled Neural Audio Codec Representations
标题:通过解开神经音频编解码器表示的声学隐形传输
链接:https://arxiv.org/abs/2510.13221
摘要:本文提出了一种方法,通过从神经音频编解码器表示的声学环境特性中分离出语音内容来实现声学隐形传态。声学隐形传态在语音记录之间传输房间特征,同时保留内容和说话者身份。我们建立在以前的工作使用EnCodec架构,实现了实质性的客观质量改进与非侵入性的ScoreQ得分为3.03,相比2.44为以前的方法。我们的训练策略包括五个任务:干净的重建,混响重建,去混响,和两个变体的声学隐形传态。我们证明了声学嵌入的时间下采样会显着降低性能,即使是2x下采样也会导致质量在统计上显着降低。学习的声学嵌入表现出与RT 60的强相关性。使用t-SNE聚类分析证明了有效的解纠缠,其中声学嵌入按房间聚类,而语音嵌入按扬声器聚类。
摘要:This paper presents an approach for acoustic teleportation by disentangling speech content from acoustic environment characteristics in neural audio codec representations. Acoustic teleportation transfers room characteristics between speech recordings while preserving content and speaker identity. We build upon previous work using the EnCodec architecture, achieving substantial objective quality improvements with non-intrusive ScoreQ scores of 3.03, compared to 2.44 for prior methods. Our training strategy incorporates five tasks: clean reconstruction, reverberated reconstruction, dereverberation, and two variants of acoustic teleportation. We demonstrate that temporal downsampling of the acoustic embedding significantly degrades performance, with even 2x downsampling resulting in a statistically significant reduction in quality. The learned acoustic embeddings exhibit strong correlations with RT60. Effective disentanglement is demonstrated using t-SNE clustering analysis, where acoustic embeddings cluster by room while speech embeddings cluster by speaker.
【4】Continuous-Token Diffusion for Speaker-Referenced TTS in Multimodal LLMs
标题:多模式LLM中演讲者参考的TTC的连续令牌扩散
链接:https://arxiv.org/abs/2510.12995
摘要:多模态大型语言模型(MLLM)中的统一架构在单个框架内处理不同任务方面表现出了希望。在文本到语音(TTS)的任务,目前MLLM为基础的方法依赖于离散的令牌表示,忽视了语音的固有连续性,并可能导致损失的细粒度的声学信息。在这项工作中,我们调查的TTS MLLM范式使用连续的语音表示。我们设计了一个双头架构,并实现了两个互补的训练策略,一个强大的模型。(1)在MLLM上增加了一个产生连续语音表示的扩散头,它是帧级的,并且是严格自回归的。(2)原始语言模型头被保留以保持多任务能力并控制语音合成的开始和结束。(3)采用掩蔽训练来解决自回归解码中的暴露偏差。(4)为了稳定优化,我们提出了一个两阶段的方案,其中LM在第二阶段被冻结,确保扩散头从固定的输入分布中学习。LibriSpeech(PC)测试清洁的评估表明,我们的方法实现了最先进的自回归性能,WER为1.95%,说话人相似度为0.54,UTMOS为4.00。两阶段训练比一阶段训练基线产生46%的相对WER降低。这些结果突出了自回归建模与连续令牌扩散相结合的有效性,由两阶段训练过程支持。
摘要:Unified architectures in multimodal large language models (MLLM) have shown promise in handling diverse tasks within a single framework. In the text-to-speech (TTS) task, current MLLM-based approaches rely on discrete token representations, which disregard the inherently continuous nature of speech and can lead to loss of fine-grained acoustic information.In this work, we investigate the TTS within the MLLM paradigm using continuous speech representations. We design a dual-head architecture and implement two complementary training strategies for a robust model. (1) A diffusion head generating continuous speech representations is added on the MLLM, which is on frame-level and strictly autoregressive. (2) The original language model head is retained to preserve multitask capability and to control the start and end of speech synthesis. (3) Masked training is employed to address exposure bias in autoregressive decoding. (4) To stabilize optimization, we propose a two-stage scheme where the LM is frozen in the second stage, ensuring the diffusion head learns from a fixed input distribution. Evaluations on LibriSpeech(PC) test-clean show that our approach achieves state-of-the-art autoregressive performance, with a WER of 1.95%, speaker similarity of 0.54, and UTMOS of 4.00. The two-stage training yields a 46% relative WER reduction over the one-stage training baseline. These results highlight the effectiveness of combining autoregressive modeling with continuous-token diffusion, supported by a two-stage training procedure.
【5】HyWA: Hypernetwork Weight Adapting Personalized Voice Activity Detection
标题:HyWA:超网络重量自适应个性化语音活动检测
链接:https://arxiv.org/abs/2510.12947
备注:Mahsa Ghazvini Nejad and Hamed Jafarzadeh Asl contributed equally to this work
摘要:个性化语音活动检测(PVAD)系统通过结合来自登记话语的说话者嵌入而仅响应于特定目标说话者而激活。与现有的方法,需要架构的变化,如电影层,我们的方法采用了超网络来修改一个标准的语音活动检测(VAD)模型中的几个选定的层的权重。这使得扬声器调节无需改变VAD架构,允许相同的VAD模型通过仅更新一小部分层来适应不同的扬声器。我们提出了HyWA-PVAD,超网络的权重自适应方法,并评估它对多个基线条件技术。我们的比较显示PVAD性能的持续改进。HyWA还通过保留核心VAD架构为部署提供了实际优势。我们的新方法在两个方面改进了当前的条件技术:i)提高了平均精度,ii)通过重用相同的VAD架构简化了部署。
摘要:Personalized Voice Activity Detection (PVAD) systems activate only in response to a specific target speaker by incorporating speaker embeddings from enrollment utterances. Unlike existing methods that require architectural changes, such as FiLM layers, our approach employs a hypernetwork to modify the weights of a few selected layers within a standard voice activity detection (VAD) model. This enables speaker conditioning without changing the VAD architecture, allowing the same VAD model to adapt to different speakers by updating only a small subset of the layers. We propose HyWA-PVAD, a hypernetwork weight adaptation method, and evaluate it against multiple baseline conditioning techniques. Our comparison shows consistent improvements in PVAD performance. HyWA also offers practical advantages for deployment by preserving the core VAD architecture. Our new approach improves the current conditioning techniques in two ways: i) increases the mean average precision, ii) simplifies deployment by reusing the same VAD architecture.
【6】Automatic Speech Recognition in the Modern Era: Architectures, Training, and Evaluation
标题:现代的自动语音识别:架构、训练和评估
链接:https://arxiv.org/abs/2510.12827
摘要:在过去十年中,自动语音识别(ASR)在深度学习的推动下经历了深刻的变革。本调查全面概述了ASR的现代时代,描绘了其从传统混合系统(如高斯混合模型-隐马尔可夫模型(GMM-Hyndom)和深度神经网络-Hyndom(DNN-Hyndom))到现在占主导地位的端到端神经架构的演变。我们系统地回顾了基本的端到端范例:连接主义时间分类(CTC),基于注意力的编码器-解码器模型和递归神经网络转换器(RNN-T),它们为完全集成的语音到文本系统奠定了基础。然后,我们详细介绍了随后的架构转向Transformer和Conformer模型,利用自我注意力来捕获高计算效率的长期依赖关系。本次调查的一个中心主题是培训模式的平行革命。我们研究了从完全监督学习(通过SpecAugment等技术增强)到自我监督学习(SSL)的兴起以及wav 2 vec 2.0等基础模型的进展,这些模型大大减少了对转录数据的依赖。此外,我们分析了像Whisper这样的大规模弱监督模型的影响,这些模型通过大量数据多样性实现了前所未有的鲁棒性。该文件还涵盖了生态系统的基本组成部分,包括关键数据集和基准(例如,LibriSpeech、Switchboard、CHiME)、标准评估度量(例如,字错误率),以及现实世界部署的关键考虑因素,如流式推理、设备上的效率以及公平性和鲁棒性的道德要求。最后,我们概述了开放的挑战和未来的研究方向。
摘要:Automatic Speech Recognition (ASR) has undergone a profound transformation over the past decade, driven by advances in deep learning. This survey provides a comprehensive overview of the modern era of ASR, charting its evolution from traditional hybrid systems, such as Gaussian Mixture Model-Hidden Markov Models (GMM-HMMs) and Deep Neural Network-HMMs (DNN-HMMs), to the now-dominant end-to-end neural architectures. We systematically review the foundational end-to-end paradigms: Connectionist Temporal Classification (CTC), attention-based encoder-decoder models, and the Recurrent Neural Network Transducer (RNN-T), which established the groundwork for fully integrated speech-to-text systems. We then detail the subsequent architectural shift towards Transformer and Conformer models, which leverage self-attention to capture long-range dependencies with high computational efficiency. A central theme of this survey is the parallel revolution in training paradigms. We examine the progression from fully supervised learning, augmented by techniques like SpecAugment, to the rise of self-supervised learning (SSL) with foundation models such as wav2vec 2.0, which drastically reduce the reliance on transcribed data. Furthermore, we analyze the impact of largescale, weakly supervised models like Whisper, which achieve unprecedented robustness through massive data diversity. The paper also covers essential ecosystem components, including key datasets and benchmarks (e.g., LibriSpeech, Switchboard, CHiME), standard evaluation metrics (e.g., Word Error Rate), and critical considerations for real-world deployment, such as streaming inference, on-device efficiency, and the ethical imperatives of fairness and robustness. We conclude by outlining open challenges and future research directions.
【7】Closing the Gap Between Text and Speech Understanding in LLMs
标题:缩小法学硕士中文本和言语理解之间的差距
链接:https://arxiv.org/abs/2510.13632
摘要:大型语言模型(LLM)可以被适配以将其文本能力扩展到语音输入。然而,这些语音适应LLM在语言理解任务上始终表现不佳,甚至是级联管道。我们将这种不足称为文本-语音理解差距:当语音适应LLM处理语音输入时,相对于原始基于文本的LLM处理等效文本时,观察到的性能下降。最近缩小这一差距的方法要么依赖于文本语料库的大规模语音合成,这是昂贵的,严重依赖于合成数据,或大规模的专有语音数据集,这是不可复制的。因此,仍然需要更有效的数据替代方案来缩小文本-语音理解差距。在这项工作中,我们分析了由两个因素驱动的差距:(i)在适应过程中忘记文本功能,以及(ii)语音和文本之间的跨模态不一致。基于这种分析,我们引入了SALAD-通过主动选择和跨模态蒸馏进行学习的样本有效对齐-它将跨模态蒸馏与目标合成数据相结合,以改善对齐,同时减轻遗忘。应用于3B和7 B LLM,SALAD在知识,语言理解和推理的广泛领域基准中具有强大的开放权重模型,同时在公共语料库中训练数量级更少的语音数据,从而实现了具有竞争力的性能。
摘要:Large Language Models (LLMs) can be adapted to extend their text capabilities to speech inputs. However, these speech-adapted LLMs consistently underperform their text-based counterparts--and even cascaded pipelines--on language understanding tasks. We term this shortfall the text-speech understanding gap: the performance drop observed when a speech-adapted LLM processes spoken inputs relative to when the original text-based LLM processes the equivalent text. Recent approaches to narrowing this gap either rely on large-scale speech synthesis of text corpora, which is costly and heavily dependent on synthetic data, or on large-scale proprietary speech datasets, which are not reproducible. As a result, there remains a need for more data-efficient alternatives for closing the text-speech understanding gap. In this work, we analyze the gap as driven by two factors: (i) forgetting of text capabilities during adaptation, and (ii) cross-modal misalignment between speech and text. Based on this analysis, we introduce SALAD--Sample-efficient Alignment with Learning through Active selection and cross-modal Distillation--which combines cross-modal distillation with targeted synthetic data to improve alignment while mitigating forgetting. Applied to 3B and 7B LLMs, SALAD achieves competitive performance with a strong open-weight model across broad-domain benchmarks in knowledge, language understanding, and reasoning, while training on over an order of magnitude less speech data from public corpora.
【8】Adaptive vector steering: A training-free, layer-wise intervention for hallucination mitigation in large audio and multimodal models
标题:自适应向量转向:一种免训练、分层干预,用于缓解大型音频和多模式模型中的幻觉
链接:https://arxiv.org/abs/2510.12851
备注:Note: This preprint is a version of the paper submitted to ICASSP 2026. The author list here includes contributors who provided additional supervision and guidance. The official ICASSP submission may differ slightly in author composition
摘要:大型音频语言模型和多模态大型语言模型在音频问题分类(AQA)、音频字幕和自动语音识别(ASR)等任务中表现出强大的能力。然而,越来越多的证据表明,这些模型可以对音频内容产生幻觉。为了解决这个问题,我们探测模型的内部状态,并提出自适应矢量转向(AVS),一种更好地在音频内容中生成的方法。我们还确定了输出正确性和内部表示之间的强相关性。实验表明,在两个模型和两个基准一致的性能增益。在音频幻觉QA数据集上,我们的方法将Gemma的F1分数从0.550提高到0.619,Qwen从0.626提高到0.632。此外,我们的方法将Qwen对MMAU的准确度从0.548提高到0.592,相对提高了8%。据我们所知,这是第一个应用矢量转向来减轻音频中的幻觉的工作。
摘要:Large Audio-Language Models and Multi-Modal Large Language Models have demonstrated strong capabilities in tasks such as Audio Question Answering (AQA), Audio Captioning, and Automatic Speech Recognition (ASR). However, there is growing evidence that these models can hallucinate about the content of the audio. To address this issue, we probe the models' internal states and propose Adaptive Vector Steering (AVS), a method that better grounds generation in audio content. We also identify a strong correlation between output correctness and internal representations. Experiments show consistent performance gains across two models and two benchmarks. On the Audio Hallucination QA dataset, our method boosts the F1-score of Gemma from 0.550 to 0.619 and Qwen from 0.626 to 0.632. Furthermore, our method increases the accuracy of Qwen on MMAU from 0.548 to 0.592, marking an 8% relative increase. To the best of our knowledge, this is the first work to apply vector steering to mitigate hallucination in audio.
【9】Gelina: Unified Speech and Gesture Synthesis via Interleaved Token Prediction
标题:Gelina:通过交织令牌预测统一语音和手势合成
链接:https://arxiv.org/abs/2510.12834
备注:5 pages
摘要:人类交流是多模态的,语音和手势紧密耦合,但大多数用于生成语音和手势的计算方法按顺序合成它们,削弱了同步和韵律对齐。我们介绍Gelina,一个统一的框架,联合合成语音和语音手势从文本中使用交错令牌序列在离散自回归骨干,与特定于模态的解码器。Gelina支持多说话者和多风格克隆,并支持从语音输入进行仅手势合成。主观和客观的评价表明竞争力的语音质量和改进的手势生成超过单峰基线。
摘要:Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a unified framework that jointly synthesizes speech and co-speech gestures from text using interleaved token sequences in a discrete autoregressive backbone, with modality-specific decoders. Gelina supports multi-speaker and multi-style cloning and enables gesture-only synthesis from speech inputs. Subjective and objective evaluations demonstrate competitive speech quality and improved gesture generation over unimodal baselines.
【10】Production and Manufacturing of 3D Printed Acoustic Guitars
标题:3D打印原声吉他的生产制造
链接:https://arxiv.org/abs/2510.12823
摘要:这项研究调查了使用3D打印生产经济实惠的功能性原声吉他的可行性,重点是生产具有适当音调性能的结构设计。该研究与William Schiesser合作进行,使用古典吉他模型,选择其较低的弦张力,以评估由聚乳酸(PLA)制成的3D打印原型的音调特征。由于Prusa Mark 4打印机的构造板尺寸限制,吉他主体被分成多个部分,以压配合公差和最小的氰基丙烯酸酯粘合剂连接。Fusion 360中的CAD建模确保了压配连接和整体装配的尺寸精度。组装完成后,吉他被用尼龙弦串起来,并使用Audacity软件进行测试,将记录的频率和音符与标准参考值进行比较。结果表明,在较低的字符串频率,可能是由印刷中使用的材料选择造成的大的偏差。所有琴弦都达到了精确的音高,尽管通过调音存在频率差异,这表明尽管存在不可避免的挑战,但PLA和现代制造方法可以生产出价格合理、可演奏的原声吉他。进一步的研究可能会调查替代塑料的优良频率匹配。这种方法具有巨大的潜力,可以扩大获得优质乐器的机会,同时减少对濒危音木的依赖,从而鼓励可持续的乐器生产和增加音乐参与。这也为获得乐器仍然是一项挑战的弱势社区创造了机会。 关键词:制琴师,光固化成型,3D打印,吉他制作
摘要:This research investigates the feasibility of producing affordable, functional acoustic guitars using 3D printing, with a focus on producing structural designs with proper tonal performance. Conducted in collaboration with William Schiesser, the study uses a classical guitar model, chosen for its lower string tension, to evaluate the tonal characteristics of a 3D-printed prototype made from polylactic acid (PLA). Due to the build plate size constraints of the Prusa Mark 4 printer, the guitar body was divided into multiple sections joined with press-fit tolerances and minimal cyanoacrylate adhesive. CAD modeling in Fusion 360 ensured dimensional accuracy in press-fit connections and the overall assembly. Following assembly, the guitar was strung with nylon strings and tested using Audacity software to compare recorded frequencies and notes with standard reference values. Results showed large deviations in lower string frequencies, likely caused by the material choice utilized in printing. Accurate pitches were reached with all strings despite frequency differences through tuning, demonstrating that PLA and modern manufacturing methods can produce affordable, playable acoustic guitars despite inevitable challenges. Further research may investigate alternative plastics for superior frequency matching. This approach holds significant potential for expanding access to quality instruments while reducing reliance on endangered tonewoods, thereby encouraging both sustainable instrument production and increased musical participation. This also creates opportunities for disadvantaged communities where access to musical instruments remains a challenge. Keywords: Luthiery, Stereolithography, 3D-Print, Guitar Making
【11】Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
标题:超越离散类别:用于宠物发声分析的多任务价-唤醒建模
链接:https://arxiv.org/abs/2510.12819
备注:24 pages, 6 figures, 4 tables. First continuous VA framework for pet vocalization analysis with 42,553 samples
摘要:传统的宠物情感识别发声,基于离散分类,与模糊性和捕捉强度变化的斗争。我们提出了一个连续的效价唤醒(VA)模型,表示在一个二维空间的情绪。我们的方法使用自动VA标签生成算法,能够对42,553个宠物发声样本进行大规模注释。多任务学习框架联合训练VA回归与辅助任务(情感,身体大小,性别),以通过改进特征学习来增强预测。我们的Audio Transformer模型实现了r = 0.9024和r = 0.7155的验证Valence Pearson相关性,有效地解决了“领土”和“快乐”等离散类别之间的混淆。“这项工作引入了第一个用于宠物发声分析的连续VA框架,为人类与宠物的互动,兽医诊断和行为训练提供了更具表现力的表示。该方法显示出在消费产品中部署的强大潜力,如AI宠物情感翻译器。
摘要:Traditional pet emotion recognition from vocalizations, based on discrete classification, struggles with ambiguity and capturing intensity variations. We propose a continuous Valence-Arousal (VA) model that represents emotions in a two-dimensional space. Our method uses an automatic VA label generation algorithm, enabling large-scale annotation of 42,553 pet vocalization samples. A multi-task learning framework jointly trains VA regression with auxiliary tasks (emotion, body size, gender) to enhance prediction by improving feature learning. Our Audio Transformer model achieves a validation Valence Pearson correlation of r = 0.9024 and an Arousal r = 0.7155, effectively resolving confusion between discrete categories like "territorial" and "happy." This work introduces the first continuous VA framework for pet vocalization analysis, offering a more expressive representation for human-pet interaction, veterinary diagnostics, and behavioral training. The approach shows strong potential for deployment in consumer products like AI pet emotion translators.
机器翻译由腾讯交互翻译提供,仅供参考
