今日论文合集:cs.SD语音6篇,eess.AS音频处理8篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily

cs.SD语音
【1】HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
标题:HARP:大规模更高级立体声房间脉冲响应数据集
链接:https://arxiv.org/abs/2411.14207
作者:Shivam Saini,  Jürgen Peissig
备注:Submitted to ICASSP 2025 Workshop Dataset and code to be uploaded at: this https URL
摘要:本文介绍了一个使用图像源方法创建的7阶高保真立体声房间脉冲响应(HOA-RIR)数据集。通过采用高阶Ambisonics,我们的数据集可以实现精确的空间音频再现,这是逼真沉浸式音频应用的关键要求。利用虚拟仿真,我们提出了一种基于叠加原理的独特麦克风配置,旨在优化声场覆盖范围,同时解决传统麦克风阵列的局限性。所提出的64麦克风配置允许我们直接在球谐域中捕获RIR。该数据集具有广泛的房间配置,包括房间几何形状,吸声材料和源-接收器距离的变化。同时提供了模拟设置的详细描述,以实现准确的再现。该数据集是研究空间音频的研究人员的重要资源,特别是在涉及机器学习的应用中,以改善室内声学建模和声场合成。它还提供了非常高的空间分辨率和逼真度,对于声源定位、混响预测和沉浸式声音再现等任务至关重要。
摘要:This contribution introduces a dataset of 7th-order Ambisonic Room ImpulseResponses (HOA-RIRs), created using the Image Source Method. By employinghigher-order Ambisonics, our dataset enables precise spatial audioreproduction, a critical requirement for realistic immersive audioapplications. Leveraging the virtual simulation, we present a unique microphoneconfiguration, based on the superposition principle, designed to optimize soundfield coverage while addressing the limitations of traditional microphonearrays. The presented 64-microphone configuration allows us to capture RIRsdirectly in the Spherical Harmonics domain. The dataset features a wide rangeof room configurations, encompassing variations in room geometry, acousticabsorption materials, and source-receiver distances. A detailed description ofthe simulation setup is provided alongside for an accurate reproduction. Thedataset serves as a vital resource for researchers working on spatial audio,particularly in applications involving machine learning to improve roomacoustics modeling and sound field synthesis. It further provides a very highlevel of spatial resolution and realism crucial for tasks such as sourcelocalization, reverberation prediction, and immersive sound reproduction.

【2】 X-CrossNet: A complex spectral mapping approach to target speaker  extraction with cross attention speaker embedding fusion
标题:X-CrossNet:一种复杂的频谱映射方法,通过交叉注意说话人嵌入融合来提取目标说话人
链接:https://arxiv.org/abs/2411.13811
作者:Chang Sun,  Bo Qin
摘要:目标说话人提取(TSE)是一种利用与目标说话人相关联的辅助特征从混合语音中分离出目标说话人的语音的技术。这种方法解决了鸡尾酒会的问题,通常被认为是更有前途的实际应用比传统的语音分离方法。虽然这一领域的学术研究在公共数据集上取得了很高的准确性和评估分数,但大多数模型在现实世界的噪声或混响条件下表现出显着降低的性能。为了解决这个问题,我们提出了一种新的TSE模型,X-CrossNet,它利用CrossNet作为其骨干。CrossNet是一种语音分离网络,专门针对具有挑战性的噪声和混响环境进行了优化,在这些条件下实现了扬声器分离等任务的最先进性能。此外,为了增强网络捕获和利用目标说话人辅助特征的能力,我们将交叉注意机制集成到每个CrossNet块内的全局多头自注意(GMHSA)模块中。这有助于目标说话者特征与混合语音特征的更有效整合。实验结果表明,该方法对WSJ 0 - 2 mix和WHAMR!数据集,表现出强大的鲁棒性和稳定性。
摘要:Target speaker extraction (TSE) is a technique for isolating a targetspeaker's voice from mixed speech using auxiliary features associated with thetarget speaker. This approach addresses the cocktail party problem and isgenerally considered more promising for practical applications thanconventional speech separation methods. Although academic research in this areahas achieved high accuracy and evaluation scores on public datasets, mostmodels exhibit significantly reduced performance in real-world noisy orreverberant conditions. To address this limitation, we propose a novel TSEmodel, X-CrossNet, which leverages CrossNet as its backbone. CrossNet is aspeech separation network specifically optimized for challenging noisy andreverberant environments, achieving state-of-the-art performance in tasks suchas speaker separation under these conditions. Additionally, to enhance thenetwork's ability to capture and utilize auxiliary features of the targetspeaker, we integrate a Cross-Attention mechanism into the global multi-headself-attention (GMHSA) module within each CrossNet block. This facilitates moreeffective integration of target speaker features with mixed speech features.Experimental results show that our method performs superior separation on theWSJ0-2mix and WHAMR! datasets, demonstrating strong robustness and stability.

【3】 Tiny-Align: Bridging Automatic Speech Recognition and Large Language  Model on the Edge
标题:Tiny-Align:在边缘架起自动语音识别和大型语言模型
链接:https://arxiv.org/abs/2411.13766
作者:Ruiyang Qin,  Dancheng Liu,  Gelei Xu,  Zheyu Yan,  Chenhui Xu,  Yuting Hu,  X. Sharon Hu,  Jinjun Xiong,  Yiyu Shi
备注:7 pages, 8 figures
摘要:大型语言模型(LLM)和自动语音识别(ASR)的组合,当部署在边缘设备(称为边缘ASR-LLM)上时,可以作为一个强大的个性化助手,为用户提供基于音频的交互。与基于文本的交互相比,边缘ASR-LLM允许可访问和自然的音频交互。不幸的是,现有的ASR-LLM模型主要是在高性能计算环境中训练的,并产生大量的模型权重,使得它们难以部署在边缘设备上。更重要的是,为了更好地满足用户的个性化需求,ASR-LLM必须能够从每个不同的用户那里学习,因为音频输入通常包含高度个性化的特征,需要个性化的设备上培训。由于单独微调ASR或LLM通常会因特定于模态的限制而导致次优结果,因此端到端培训可确保音频功能和语言理解的无缝集成(跨模态对齐),最终实现更加个性化和高效的适配边缘设备。然而,由于现有方法的复杂训练要求和大量计算需求,ASR音频和LLM之间的跨模态对齐在边缘设备上可能具有挑战性。在这项工作中,我们提出了一个资源高效的跨模态对齐框架,该框架将边缘设备上的ASR和LLM连接起来,以处理个性化的音频输入。我们的框架能够在资源受限的设备(如NVIDIA Jetson Orin(8 GB RAM))上实现高效的ASR-LLM对齐,实现50倍的训练时间加速,同时将对齐质量提高50%以上。据我们所知,这是研究资源受限边缘设备上的有效ASR-LLM对齐的第一项工作。
摘要:The combination of Large Language Models (LLM) and Automatic SpeechRecognition (ASR), when deployed on edge devices (called edge ASR-LLM), canserve as a powerful personalized assistant to enable audio-based interactionfor users. Compared to text-based interaction, edge ASR-LLM allows accessibleand natural audio interactions. Unfortunately, existing ASR-LLM models aremainly trained in high-performance computing environments and producesubstantial model weights, making them difficult to deploy on edge devices.More importantly, to better serve users' personalized needs, the ASR-LLM mustbe able to learn from each distinct user, given that audio input often containshighly personalized characteristics that necessitate personalized on-devicetraining. Since individually fine-tuning the ASR or LLM often leads tosuboptimal results due to modality-specific limitations, end-to-end trainingensures seamless integration of audio features and language understanding(cross-modal alignment), ultimately enabling a more personalized and efficientadaptation on edge devices. However, due to the complex training requirementsand substantial computational demands of existing approaches, cross-modalalignment between ASR audio and LLM can be challenging on edge devices. In thiswork, we propose a resource-efficient cross-modal alignment framework thatbridges ASR and LLMs on edge devices to handle personalized audio input. Ourframework enables efficient ASR-LLM alignment on resource-constrained deviceslike NVIDIA Jetson Orin (8GB RAM), achieving 50x training time speedup whileimproving the alignment quality by more than 50\%. To the best of ourknowledge, this is the first work to study efficient ASR-LLM alignment onresource-constrained edge devices.

【4】 FabuLight-ASD: Unveiling Speech Activity via Body Language
标题:FabuLight-ASD:通过肢体语言揭示言语活动
链接:https://arxiv.org/abs/2411.13674
作者:Hugo Carneiro,  Stefan Wermter
备注:23 pages, 8 figures, 3 tables, accepted for publication in Neural Computing and Applications
摘要:多模态环境中的主动说话人检测(ASD)对于从视频会议到人机交互的各种应用至关重要。本文介绍了FabuLight-ASD,一种先进的ASD模型,它集成了面部,音频和身体姿势信息,以提高检测精度和鲁棒性。我们的模型建立在现有的Light-ASD框架上,通过合并人体姿势数据,通过骨架图表示,最大限度地减少了计算开销。使用怀尔德主动说话者检测(WASD)数据集,可靠的面部和身体边界框注释而闻名,我们证明了FabuLight-ASD在现实世界中的有效性。FabuLight-ASD的总体平均精度(mAP)达到94.3%,优于Light-ASD,后者在各种具有挑战性的场景中的总体mAP为93.7%。身体姿势信息的结合显示出特别有利的影响,在具有语音障碍、面部遮挡和人类语音背景噪声的场景中观察到mAP的显著改善。此外,效率分析表明,参数计数(27.3%)和乘法累加运算(高达2.4%)只有适度的增加,强调了模型的效率和可行性。这些发现验证了FabuLight-ASD通过整合身体姿势数据增强ASD性能的功效。FabuLight-ASD的代码和模型重量可在https://github.com/knowledgetechnologyuhh/FabuLight-ASD上获得。
摘要:Active speaker detection (ASD) in multimodal environments is crucial forvarious applications, from video conferencing to human-robot interaction. Thispaper introduces FabuLight-ASD, an advanced ASD model that integrates facial,audio, and body pose information to enhance detection accuracy and robustness.Our model builds upon the existing Light-ASD framework by incorporating humanpose data, represented through skeleton graphs, which minimises computationaloverhead. Using the Wilder Active Speaker Detection (WASD) dataset, renownedfor reliable face and body bounding box annotations, we demonstrateFabuLight-ASD's effectiveness in real-world scenarios. Achieving an overallmean average precision (mAP) of 94.3%, FabuLight-ASD outperforms Light-ASD,which has an overall mAP of 93.7% across various challenging scenarios. Theincorporation of body pose information shows a particularly advantageousimpact, with notable improvements in mAP observed in scenarios with speechimpairment, face occlusion, and human voice background noise. Furthermore,efficiency analysis indicates only a modest increase in parameter count (27.3%)and multiply-accumulate operations (up to 2.4%), underscoring the model'sefficiency and feasibility. These findings validate the efficacy ofFabuLight-ASD in enhancing ASD performance through the integration of body posedata. FabuLight-ASD's code and model weights are available athttps://github.com/knowledgetechnologyuhh/FabuLight-ASD.

【5】 A Novel Speech Analysis and Correction Tool for Arabic-Speaking Children
标题:针对阿拉伯语儿童的新型言语分析和纠正工具
链接:https://arxiv.org/abs/2411.13592
作者:Lamia Berriche,  Maha Driss,  Areej Ahmed Almuntashri,  Asma Mufreh Lghabi,  Heba Saleh Almudhi,  Munerah Abdul-Aziz Almansour
摘要:本文介绍了一个新的应用程序名为ArPA的阿拉伯孩子谁有发音困难。我们的应用程序包括两个关键组件:诊断模块和治疗模块。诊断过程包括捕获儿童的语音信号,预处理,并使用不同的机器学习分类器(如K-最近邻(KNN),支持向量机(SVM)和决策树)以及深度神经网络分类器(如ResNet 18)进行分析。治疗模块提供了引人注目的游戏化界面,其中每个正确发音的字母都可以获得更高的化身级别,为孩子的发音改善提供积极的强化。两个数据集用于实验评估:一个来自儿童保育中心,另一个包括阿拉伯字母发音录音。我们的工作使用了一种新的语音识别技术,使用Melspectrogram和MFCC图像。结果表明,ResNet 18分类器在语音到图像转换数据上有效地识别阿拉伯语语音中的错误发音,准确率为99.015%,Mel-Spectrogram图像优于ResNet 18 MFCC图像。
摘要:This paper introduces a new application named ArPA for Arabic kids who havetrouble with pronunciation. Our application comprises two key components: thediagnostic module and the therapeutic module. The diagnostic process involvescapturing the child's speech signal, preprocessing, and analyzing it usingdifferent machine learning classifiers like K-Nearest Neighbors (KNN), SupportVector Machine (SVM), and Decision Trees as well as deep neural networkclassifiers like ResNet18. The therapeutic module offers eye-catching gamifiedinterfaces in which each correctly spoken letter earns a higher avatar level,providing positive reinforcement for the child's pronunciation improvement. Twodatasets were used for experimental evaluation: one from a childcare centre andthe other including Arabic alphabet pronunciation recordings. Our work uses anovel technique for speech recognition using Melspectrogram and MFCC images.The results show that the ResNet18 classifier on speech-to-image converted dataeffectively identifies mispronunciations in Arabic speech with an accuracy of99.015\% with Mel-Spectrogram images outperforming ResNet18 with MFCC images.

【6】 WavChat: A Survey of Spoken Dialogue Models
标题:WavChat:口语对话模式调查
链接:https://arxiv.org/abs/2411.13577
作者:Shengpeng Ji,  Yifu Chen,  Minghui Fang,  Jialong Zuo,  Jingyu Lu,  Hanting Wang,  Ziyue Jiang,  Long Zhou,  Shujie Liu,  Xize Cheng,  Xiaoda Yang,  Zehan Wang,  Qian Yang,  Jian Li,  Yidi Jiang,  Jingzhen He,  Yunfei Chu,  Jin Xu,  Zhou Zhao
备注:60 papes, working in progress
摘要:口语对话模型的最新进展,例如GPT-4 o系统,在语音领域引起了极大的关注。与包括语音识别(ASR)、大语言模型(LLM)和文本到语音(TTS)的传统三层级联口语对话模型相比,现代口语对话模型表现出更高的智能。这些高级口语对话模型不仅能够理解音频、音乐和其他语音相关特征,还能捕捉语音中的风格和音色特征。此外,它们能够以低延迟生成高质量的多轮语音响应,通过同时听和说功能实现实时交互。尽管口语对话系统取得了进展,但缺乏系统地组织和分析这些系统及其基础技术的全面调查。为了解决这个问题,我们首先按时间顺序编译了现有的口语对话系统,并将其分类为级联和端到端范式。然后,我们深入概述了口语对话模型的核心技术,包括语音表示、训练范式、流、双工和交互功能等方面。每一节都讨论了这些技术的局限性,并概述了未来研究的考虑因素。此外,我们提出了一个彻底的审查相关的数据集,评估指标,并从培训和评估口语对话系统的角度基准。我们希望这项调查将有助于推进口语对话系统领域的学术研究和工业应用。有关材料可在https://github.com/jishengpeng/WavChat上查阅。
摘要:Recent advancements in spoken dialogue models, exemplified by systems likeGPT-4o, have captured significant attention in the speech domain. Compared totraditional three-tier cascaded spoken dialogue models that comprise speechrecognition (ASR), large language models (LLMs), and text-to-speech (TTS),modern spoken dialogue models exhibit greater intelligence. These advancedspoken dialogue models not only comprehend audio, music, and otherspeech-related features, but also capture stylistic and timbral characteristicsin speech. Moreover, they generate high-quality, multi-turn speech responseswith low latency, enabling real-time interaction through simultaneous listeningand speaking capability. Despite the progress in spoken dialogue systems, thereis a lack of comprehensive surveys that systematically organize and analyzethese systems and the underlying technologies. To address this, we have firstcompiled existing spoken dialogue systems in the chronological order andcategorized them into the cascaded and end-to-end paradigms. We then provide anin-depth overview of the core technologies in spoken dialogue models, coveringaspects such as speech representation, training paradigm, streaming, duplex,and interaction capabilities. Each section discusses the limitations of thesetechnologies and outlines considerations for future research. Additionally, wepresent a thorough review of relevant datasets, evaluation metrics, andbenchmarks from the perspectives of training and evaluating spoken dialoguesystems. We hope this survey will contribute to advancing both academicresearch and industrial applications in the field of spoken dialogue systems.The related material is available at https://github.com/jishengpeng/WavChat.

eess.AS音频处理

【1】 MVANet: Multi-Stage Video Attention Network for Sound Event Localization  and Detection with Source Distance Estimation
标题:MVANet:通过源距离估计进行声音事件定位和检测的多阶段视频注意力网络
链接:https://arxiv.org/abs/2411.14153
作者:Hengyi Hong,  Qing Wang,  Jun Du,  Ruoyu Wei,  Mingqi Cai,  Xin Fang
摘要:声源距离估计的声事件定位和检测(3D SELD)不仅涉及识别声音类别及其到达方向(DOA),而且还涉及预测声源的距离,旨在提供关于声音位置的完整信息。本文提出了一种用于视听(AV)3D SELD的多级视频注意力网络(MVANet)。多阶段音频特征用于自适应地捕获视频中声源的空间信息。我们提出了一种新颖的输出表示,通过计算真实的笛卡尔坐标将波达方向与声源距离结合起来,以解决声学场景和事件检测和分类(DCASE)2024挑战赛中新引入的源距离估计(SDE)任务。我们还采用了各种有效的数据增强和预训练方法。STARSS23数据集上的实验结果证明了我们提出的MVANet的有效性。通过集成上述技术,我们的系统优于我们在DCASE 2024挑战赛的AV 3D SELD任务中使用的排名第一的方法,而无需模型集成。该守则将在今后公开提供。
摘要:Sound event localization and detection with source distance estimation (3DSELD) involves not only identifying the sound category and itsdirection-of-arrival (DOA) but also predicting the source's distance, aiming toprovide full information about the sound position. This paper proposes amulti-stage video attention network (MVANet) for audio-visual (AV) 3D SELD.Multi-stage audio features are used to adaptively capture the spatialinformation of sound sources in videos. We propose a novel outputrepresentation that combines the DOA with distance of sound sources bycalculating the real Cartesian coordinates to address the newly introducedsource distance estimation (SDE) task in the Detection and Classification ofAcoustic Scenes and Events (DCASE) 2024 Challenge. We also employ a variety ofeffective data augmentation and pre-training methods. Experimental results onthe STARSS23 dataset have proven the effectiveness of our proposed MVANet. Byintegrating the aforementioned techniques, our system outperforms thetop-ranked method we used in the AV 3D SELD task of the DCASE 2024 Challengewithout model ensemble. The code will be made publicly available in the future.

【2】 BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken  Term Detection
标题:BEST-SD:用于口语检测的双向Mamba增强语音令牌化
链接:https://arxiv.org/abs/2411.14100
作者:Anup Singh,  Kris Demuynck,  Vipul Arora
备注:Submitted to ICASSP 2025
摘要:口语词检测(STD)往往受到依赖帧级特征和计算密集的基于DTW的模板匹配的阻碍,限制了其实用性。为了解决这些挑战,我们提出了一种新的方法,将语音编码成离散的,说话者不可知的语义令牌。这有助于使用基于文本的搜索算法进行快速检索,并有效地处理词汇表外的术语。我们的方法侧重于在同一术语的不同话语中生成一致的令牌序列。我们还提出了一个双向的状态空间建模内的曼巴编码器,在一个自我监督的学习框架中训练,学习上下文帧级的功能,进一步编码成离散的令牌。我们的分析表明,我们的语音令牌表现出更大的扬声器不变性比现有的tokenizer,使他们更适合STD任务。LibriSpeech和TIMIT数据库的实证评估表明,我们的方法优于现有的STD基线,同时更有效。
摘要:Spoken term detection (STD) is often hindered by reliance on frame-levelfeatures and the computationally intensive DTW-based template matching,limiting its practicality. To address these challenges, we propose a novelapproach that encodes speech into discrete, speaker-agnostic semantic tokens.This facilitates fast retrieval using text-based search algorithms andeffectively handles out-of-vocabulary terms. Our approach focuses on generatingconsistent token sequences across varying utterances of the same term. We alsopropose a bidirectional state space modeling within the Mamba encoder, trainedin a self-supervised learning framework, to learn contextual frame-levelfeatures that are further encoded into discrete tokens. Our analysis shows thatour speech tokens exhibit greater speaker invariance than those from existingtokenizers, making them more suitable for STD tasks. Empirical evaluation onLibriSpeech and TIMIT databases indicates that our method outperforms existingSTD baselines while being more efficient.

【3】 Single-Model Attribution for Spoofed Speech via Vocoder Fingerprints in  an Open-World Setting
标题:开放世界环境中通过声码器指纹实现欺骗语音的单模型归因
链接:https://arxiv.org/abs/2411.14013
作者:Matías Pizarro,  Mike Laszkiewicz,  Dorothea Kolossa,  Asja Fischer
摘要:随着语音生成技术的进步,滥用伪造语音信号的潜在威胁也在增加。解决这些威胁的一种方法是将信号归因于其源生成模型。在这项工作中,我们是第一个在开放世界环境中解决单模型归属任务的人,也就是说,我们的目标是识别来自未知来源的欺骗语音信号是否来自特定的声码器。我们表明,音频信号和它们的低通滤波或EnCodec滤波版本之间的标准化平均残差可以作为强大的声码器指纹。该方法仅需要来自目标声码器的数据,并允许简单但高度准确的基于距离的模型归因。我们在LJSpeech和JSUT上证明了其有效性,在大多数环境中平均AUROC超过99%。随附的鲁棒性研究表明,它在一定程度上对噪声水平也有弹性。
摘要:As speech generation technology advances, so do the potential threats ofmisusing spoofed speech signals. One way to address these threats is byattributing the signals to their source generative model. In this work, we arethe first to tackle the single-model attribution task in an open-world setting,that is, we aim at identifying whether spoofed speech signals from unknownsources originate from a specific vocoder. We show that the standardizedaverage residual between audio signals and their low-pass filtered or EnCodecfiltered versions can serve as powerful vocoder fingerprints. The approach onlyrequires data from the target vocoder and allows for simple but highly accuratedistance-based model attribution. We demonstrate its effectiveness on LJSpeechand JSUT, achieving an average AUROC of over 99% in most settings. Theaccompanying robustness study shows that it is also resilient to noise levelsup to a certain degree.

【4】 Sequence-to-Sequence Neural Diarization with Automatic Speaker Detection  and Representation
标题:具有自动说话人检测和表示的序列到序列神经扩展
链接:https://arxiv.org/abs/2411.13849
作者:Ming Cheng,  Yuke Lin,  Ming Li
摘要:本文提出了一种新的序列到序列神经日记(SSND)框架,执行在线和离线的发言人日记。它是从我们以前的目标说话人语音活动检测系统的序列到序列的架构,然后通过解决两个关键问题演变成一个新的日记范式。1)扬声器检测:该方法可以利用不完全给定的说话人嵌入来发现未知说话人,并预测音频信号中的目标语音活动。它不需要事先登记发言人的事先日记系统。2)发言人代表:该方法可以采用预测的语音活动作为参考信息,同时从音频信号中提取说话人嵌入。在整个日志化网络中联合学习说话人嵌入的表示空间,而不使用额外的说话人嵌入模型。在推理过程中,SSND框架可以按块处理长音频记录。检测模块利用先前获得的说话者嵌入缓冲器来预测每个即将到来的音频块的登记的和未知的说话者的语音活动。接下来,根据表示模块的预测来更新说话者嵌入缓冲器。假设一个新的扬声器可能会出现在一个小的块移位,我们的模型迭代地预测每个块的结果,并提取后续块的目标嵌入,直到信号结束。最后,最后一个说话者嵌入缓冲区可以对整个音频进行重新评分,从而实现作为离线系统的高度准确的日志化性能。(......)
摘要:This paper proposes a novel Sequence-to-Sequence Neural Diarization (SSND)framework to perform online and offline speaker diarization. It is developedfrom the sequence-to-sequence architecture of our previous target-speaker voiceactivity detection system and then evolves into a new diarization paradigm byaddressing two critical problems. 1) Speaker Detection: The proposed approachcan utilize incompletely given speaker embeddings to discover the unknownspeaker and predict the target voice activities in the audio signal. It doesnot require a prior diarization system for speaker enrollment in advance. 2)Speaker Representation: The proposed approach can adopt the predicted voiceactivities as reference information to extract speaker embeddings from theaudio signal simultaneously. The representation space of speaker embedding isjointly learned within the whole diarization network without using an extraspeaker embedding model. During inference, the SSND framework can process longaudio recordings blockwise. The detection module utilizes the previouslyobtained speaker-embedding buffer to predict both enrolled and unknownspeakers' voice activities for each coming audio block. Next, thespeaker-embedding buffer is updated according to the predictions of therepresentation module. Assuming that up to one new speaker may appear in asmall block shift, our model iteratively predicts the results of each block andextracts target embeddings for the subsequent blocks until the signal ends.Finally, the last speaker-embedding buffer can re-score the entire audio,achieving highly accurate diarization performance as an offline system.(......)

【5】 WavChat: A Survey of Spoken Dialogue Models
标题:WavChat:口语对话模式调查
链接:https://arxiv.org/abs/2411.13577
作者:Shengpeng Ji,  Yifu Chen,  Minghui Fang,  Jialong Zuo,  Jingyu Lu,  Hanting Wang,  Ziyue Jiang,  Long Zhou,  Shujie Liu,  Xize Cheng,  Xiaoda Yang,  Zehan Wang,  Qian Yang,  Jian Li,  Yidi Jiang,  Jingzhen He,  Yunfei Chu,  Jin Xu,  Zhou Zhao
备注:60 papes, working in progress
摘要:口语对话模型的最新进展,例如GPT-4 o系统,在语音领域引起了极大的关注。与包括语音识别(ASR)、大语言模型(LLM)和文本到语音(TTS)的传统三层级联口语对话模型相比,现代口语对话模型表现出更高的智能。这些高级口语对话模型不仅能够理解音频、音乐和其他语音相关特征,还能捕捉语音中的风格和音色特征。此外,它们能够以低延迟生成高质量的多轮语音响应,通过同时听和说功能实现实时交互。尽管口语对话系统取得了进展,但缺乏系统地组织和分析这些系统及其基础技术的全面调查。为了解决这个问题,我们首先按时间顺序编译了现有的口语对话系统,并将其分类为级联和端到端范式。然后,我们深入概述了口语对话模型的核心技术,包括语音表示、训练范式、流、双工和交互功能等方面。每一节都讨论了这些技术的局限性,并概述了未来研究的考虑因素。此外,我们提出了一个彻底的审查相关的数据集,评估指标,并从培训和评估口语对话系统的角度基准。我们希望这项调查将有助于推进口语对话系统领域的学术研究和工业应用。有关材料可在https://github.com/jishengpeng/WavChat上查阅。
摘要:Recent advancements in spoken dialogue models, exemplified by systems likeGPT-4o, have captured significant attention in the speech domain. Compared totraditional three-tier cascaded spoken dialogue models that comprise speechrecognition (ASR), large language models (LLMs), and text-to-speech (TTS),modern spoken dialogue models exhibit greater intelligence. These advancedspoken dialogue models not only comprehend audio, music, and otherspeech-related features, but also capture stylistic and timbral characteristicsin speech. Moreover, they generate high-quality, multi-turn speech responseswith low latency, enabling real-time interaction through simultaneous listeningand speaking capability. Despite the progress in spoken dialogue systems, thereis a lack of comprehensive surveys that systematically organize and analyzethese systems and the underlying technologies. To address this, we have firstcompiled existing spoken dialogue systems in the chronological order andcategorized them into the cascaded and end-to-end paradigms. We then provide anin-depth overview of the core technologies in spoken dialogue models, coveringaspects such as speech representation, training paradigm, streaming, duplex,and interaction capabilities. Each section discusses the limitations of thesetechnologies and outlines considerations for future research. Additionally, wepresent a thorough review of relevant datasets, evaluation metrics, andbenchmarks from the perspectives of training and evaluating spoken dialoguesystems. We hope this survey will contribute to advancing both academicresearch and industrial applications in the field of spoken dialogue systems.The related material is available at https://github.com/jishengpeng/WavChat.

【6】 HARP: A Large-Scale Higher-Order Ambisonic Room Impulse Response Dataset
标题:HARP:大规模更高级立体声房间脉冲响应数据集
链接:https://arxiv.org/abs/2411.14207
作者:Shivam Saini,  Jürgen Peissig
备注:Submitted to ICASSP 2025 Workshop Dataset and code to be uploaded at: this https URL
摘要:本文介绍了一个使用图像源方法创建的7阶高保真立体声房间脉冲响应(HOA-RIR)数据集。通过采用高阶Ambisonics,我们的数据集可以实现精确的空间音频再现,这是逼真沉浸式音频应用的关键要求。利用虚拟仿真,我们提出了一种基于叠加原理的独特麦克风配置,旨在优化声场覆盖范围,同时解决传统麦克风阵列的局限性。所提出的64麦克风配置允许我们直接在球谐域中捕获RIR。该数据集具有广泛的房间配置,包括房间几何形状,吸声材料和源-接收器距离的变化。同时提供了模拟设置的详细描述,以实现准确的再现。该数据集是研究空间音频的研究人员的重要资源,特别是在涉及机器学习的应用中,以改善室内声学建模和声场合成。它还提供了非常高的空间分辨率和逼真度,对于声源定位、混响预测和沉浸式声音再现等任务至关重要。
摘要:This contribution introduces a dataset of 7th-order Ambisonic Room ImpulseResponses (HOA-RIRs), created using the Image Source Method. By employinghigher-order Ambisonics, our dataset enables precise spatial audioreproduction, a critical requirement for realistic immersive audioapplications. Leveraging the virtual simulation, we present a unique microphoneconfiguration, based on the superposition principle, designed to optimize soundfield coverage while addressing the limitations of traditional microphonearrays. The presented 64-microphone configuration allows us to capture RIRsdirectly in the Spherical Harmonics domain. The dataset features a wide rangeof room configurations, encompassing variations in room geometry, acousticabsorption materials, and source-receiver distances. A detailed description ofthe simulation setup is provided alongside for an accurate reproduction. Thedataset serves as a vital resource for researchers working on spatial audio,particularly in applications involving machine learning to improve roomacoustics modeling and sound field synthesis. It further provides a very highlevel of spatial resolution and realism crucial for tasks such as sourcelocalization, reverberation prediction, and immersive sound reproduction.

【7】 X-CrossNet: A complex spectral mapping approach to target speaker  extraction with cross attention speaker embedding fusion
标题:X-CrossNet:一种复杂的频谱映射方法,通过交叉注意说话人嵌入融合来提取目标说话人
链接:https://arxiv.org/abs/2411.13811
作者:Chang Sun,  Bo Qin
摘要:目标说话人提取(TSE)是一种利用与目标说话人相关联的辅助特征从混合语音中分离出目标说话人的语音的技术。这种方法解决了鸡尾酒会的问题,通常被认为是更有前途的实际应用比传统的语音分离方法。虽然这一领域的学术研究在公共数据集上取得了很高的准确性和评估分数,但大多数模型在现实世界的噪声或混响条件下表现出显着降低的性能。为了解决这个问题,我们提出了一种新的TSE模型,X-CrossNet,它利用CrossNet作为其骨干。CrossNet是一种语音分离网络,专门针对具有挑战性的噪声和混响环境进行了优化,在这些条件下实现了扬声器分离等任务的最先进性能。此外,为了增强网络捕获和利用目标说话人辅助特征的能力,我们将交叉注意机制集成到每个CrossNet块内的全局多头自注意(GMHSA)模块中。这有利于目标说话者特征与混合语音特征的更有效集成。实验结果表明,该方法对WSJ 0 - 2 mix和WHAMR!数据集,表现出强大的鲁棒性和稳定性。
摘要:Target speaker extraction (TSE) is a technique for isolating a targetspeaker's voice from mixed speech using auxiliary features associated with thetarget speaker. This approach addresses the cocktail party problem and isgenerally considered more promising for practical applications thanconventional speech separation methods. Although academic research in this areahas achieved high accuracy and evaluation scores on public datasets, mostmodels exhibit significantly reduced performance in real-world noisy orreverberant conditions. To address this limitation, we propose a novel TSEmodel, X-CrossNet, which leverages CrossNet as its backbone. CrossNet is aspeech separation network specifically optimized for challenging noisy andreverberant environments, achieving state-of-the-art performance in tasks suchas speaker separation under these conditions. Additionally, to enhance thenetwork's ability to capture and utilize auxiliary features of the targetspeaker, we integrate a Cross-Attention mechanism into the global multi-headself-attention (GMHSA) module within each CrossNet block. This facilitates moreeffective integration of target speaker features with mixed speech features.Experimental results show that our method performs superior separation on theWSJ0-2mix and WHAMR! datasets, demonstrating strong robustness and stability.

【8】 Tiny-Align: Bridging Automatic Speech Recognition and Large Language  Model on the Edge
标题:Tiny-Align:在边缘架起自动语音识别和大型语言模型
链接:https://arxiv.org/abs/2411.13766
作者:Ruiyang Qin,  Dancheng Liu,  Gelei Xu,  Zheyu Yan,  Chenhui Xu,  Yuting Hu,  X. Sharon Hu,  Jinjun Xiong,  Yiyu Shi
备注:7 pages, 8 figures
摘要:大型语言模型(LLM)和自动语音识别(ASR)的组合,当部署在边缘设备(称为边缘ASR-LLM)上时,可以作为一个强大的个性化助手,为用户提供基于音频的交互。与基于文本的交互相比,边缘ASR-LLM允许可访问和自然的音频交互。不幸的是,现有的ASR-LLM模型主要是在高性能计算环境中训练的,并产生大量的模型权重,使得它们难以部署在边缘设备上。更重要的是,为了更好地满足用户的个性化需求,ASR-LLM必须能够从每个不同的用户那里学习,因为音频输入通常包含高度个性化的特征,需要个性化的设备上培训。由于单独微调ASR或LLM通常会由于特定模态的限制而导致次优结果,因此端到端训练可确保音频功能和语言理解(跨模态对齐)的无缝集成,最终实现边缘设备上更个性化和更有效的适应。然而,由于现有方法的复杂训练要求和大量计算需求,ASR音频和LLM之间的跨模态对齐在边缘设备上可能具有挑战性。在这项工作中,我们提出了一个资源高效的跨模态对齐框架,该框架将边缘设备上的ASR和LLM连接起来,以处理个性化的音频输入。我们的框架能够在资源受限的设备(如NVIDIA Jetson Orin(8 GB RAM))上实现高效的ASR-LLM对齐,实现50倍的训练时间加速,同时将对齐质量提高50%以上。据我们所知,这是研究资源受限边缘设备上的有效ASR-LLM对齐的第一项工作。
摘要:The combination of Large Language Models (LLM) and Automatic SpeechRecognition (ASR), when deployed on edge devices (called edge ASR-LLM), canserve as a powerful personalized assistant to enable audio-based interactionfor users. Compared to text-based interaction, edge ASR-LLM allows accessibleand natural audio interactions. Unfortunately, existing ASR-LLM models aremainly trained in high-performance computing environments and producesubstantial model weights, making them difficult to deploy on edge devices.More importantly, to better serve users' personalized needs, the ASR-LLM mustbe able to learn from each distinct user, given that audio input often containshighly personalized characteristics that necessitate personalized on-devicetraining. Since individually fine-tuning the ASR or LLM often leads tosuboptimal results due to modality-specific limitations, end-to-end trainingensures seamless integration of audio features and language understanding(cross-modal alignment), ultimately enabling a more personalized and efficientadaptation on edge devices. However, due to the complex training requirementsand substantial computational demands of existing approaches, cross-modalalignment between ASR audio and LLM can be challenging on edge devices. In thiswork, we propose a resource-efficient cross-modal alignment framework thatbridges ASR and LLMs on edge devices to handle personalized audio input. Ourframework enables efficient ASR-LLM alignment on resource-constrained deviceslike NVIDIA Jetson Orin (8GB RAM), achieving 50x training time speedup whileimproving the alignment quality by more than 50\%. To the best of ourknowledge, this is the first work to study efficient ASR-LLM alignment onresource-constrained edge devices.

机器翻译由腾讯交互翻译提供,仅供参考