本文经arXiv每日学术速递授权转载
【1】CAFE A Novel Code switching Dataset for Algerian Dialect French and English
标题:CAFE阿尔及利亚方言法语和英语的新型代码切换数据集
链接:https://arxiv.org/abs/2411.13424
作者:Houssam Eddine-Othman Lachemat, Akli Abbas, Nourredine Oukas, Yassine El Kheir, Samia Haboussi, Absar Showdhury Shammur
备注:24 pages, submitted to tallip
摘要:本文介绍并公开发布(接受后提供数据下载链接)CAFE -阿尔及利亚方言,法语和英语之间的第一个代码转换数据集。CAFE语音数据是独特的:(a)其自发的说话风格在体内人与人的对话捕捉现象,如代码转换和重叠的讲话,(b)解决了北非阿拉伯语方言的独特语言挑战;(c)CAFE捕捉不同社会语言学背景下阿尔及利亚各地的方言变化。CAFE数据包含大约37小时的语音,其中2小时36分钟的子集CAFE-small通过人工注释发布,包括语音分割,转录,代码转换点的明确注释,重叠语音以及其他事件,如噪音和笑声等。其余约34.58小时含有假标签transcription。除了数据发布之外,该论文还强调了使用最先进的自动语音识别(ASR)模型(如Whisper large-v2,3和AdvertingWhisper)来处理此类内容的挑战。接下来,我们使用上述Whisper模型对CAFE数据进行基准测试,并展示了精心设计的数据处理管道和先进的解码技术如何提高ASR性能,混合错误率(MER)为0.310,字符错误率(CER)为0.329,字错误率(WER)为0.538。
摘要:The paper introduces and publicly releases (Data download link availableafter acceptance) CAFE -- the first Code-switching dataset between Algeriandialect, French, and english languages. The CAFE speech data is unique for (a)its spontaneous speaking style in vivo human-human conversation capturingphenomena like code-switching and overlapping speech, (b) addresses distinctlinguistic challenges in North African Arabic dialect; (c) the CAFE capturesdialectal variations from various parts of Algeria within differentsociolinguistic contexts. CAFE data contains approximately 37 hours of speech,with a subset, CAFE-small, of 2 hours and 36 minutes released with manual humanannotation including speech segmentation, transcription, explicit annotation ofcode-switching points, overlapping speech, and other events such as noises, andlaughter among others. The rest approximately 34.58 hours contain pseudo labeltranscriptions. In addition to the data release, the paper also highlighted thechallenges of using state-of-the-art Automatic Speech Recognition (ASR) modelssuch as Whisper large-v2,3 and PromptingWhisper to handle such content.Following, we benchmark CAFE data with the aforementioned Whisper models andshow how well-designed data processing pipelines and advanced decodingtechniques can improve the ASR performance in terms of Mixed Error Rate (MER)of 0.310, Character Error Rate (CER) of 0.329 and Word Error Rate (WER) of0.538.
标题:I2 TTC:具有空间感知的图像指示沉浸式文本到语音合成
链接:https://arxiv.org/abs/2411.13314
备注:5pages,4figures
摘要:控制语音合成的风格和特征对于使输出适应特定上下文和用户需求至关重要。以前的文语转换(TTS)工作主要集中在产生自然发音的语音的技术方面,如语调,节奏和清晰度。然而,他们忽略了一个事实,即越来越重视合成语音的空间感知,这可能会在游戏和虚拟现实中提供身临其境的体验。为了解决这个问题,在本文中,我们提出了一种新的多模态TTS方法,即图像指示沉浸式文本到语音合成(I2 TTS)。具体来说,我们引入了一个场景提示编码器,将视觉场景提示直接集成到合成管道中,以控制语音生成过程。此外,我们提出了一种混响分类和细化技术,可以调整合成的梅尔谱图以增强沉浸式体验,确保所涉及的混响条件与场景准确匹配。实验结果表明,我们的模型实现了高质量的场景和空间匹配,而不影响语音的自然性,标志着在上下文感知语音合成领域的一个重大进步。项目演示页面:https://spatialTTS.github.io/索引术语-语音合成,场景提示,空间感知
摘要:Controlling the style and characteristics of speech synthesis is crucial foradapting the output to specific contexts and user requirements. PreviousText-to-speech (TTS) works have focused primarily on the technical aspects ofproducing natural-sounding speech, such as intonation, rhythm, and clarity.However, they overlook the fact that there is a growing emphasis on spatialperception of synthesized speech, which may provide immersive experience ingaming and virtual reality. To solve this issue, in this paper, we present anovel multi-modal TTS approach, namely Image-indicated Immersive Text-to-speechSynthesis (I2TTS). Specifically, we introduce a scene prompt encoder thatintegrates visual scene prompts directly into the synthesis pipeline to controlthe speech generation process. Additionally, we propose a reverberationclassification and refinement technique that adjusts the synthesizedmel-spectrogram to enhance the immersive experience, ensuring that the involvedreverberation condition matches the scene accurately. Experimental resultsdemonstrate that our model achieves high-quality scene and spatial matchingwithout compromising speech naturalness, marking a significant advancement inthe field of context-aware speech synthesis. Project demo page:https://spatialTTS.github.io/ Index Terms-Speech synthesis, scene prompt,spatial perception
标题:使用乐高积木和Raspberry Pi构建音乐
链接:https://arxiv.org/abs/2411.13224
备注:21 pages
摘要:在本文中,一个系统,以建立一个直观的和可访问的方式,与乐高积木的音乐。该系统利用技术提供的新的强大和廉价的可能性,以新的方式制造旧的东西。Raspberry Pi用于控制系统并运行必要的算法,定制的乐高积木用于构建旋律,定制电子设计,软件件和3D打印部件完成所使用的项目。该系统的设计是模块化的,它允许创建旋律与和弦和打击乐或只是旋律或作为一个节拍箱或旋律箱执行。与系统的主要交互是使用乐高积木。测试证明了它的多功能性和易用性,以及它在儿童和成人音乐学习中的有用性。
摘要:In this paper, a system to build music in an intuitive and accessible way,with Lego bricks, is presented. The system makes use of the new powerful andcheap possibilities that technology offers for making old things in a new way.The Raspberry Pi is used to control the system and run the necessaryalgorithms, customized Lego bricks are used for building melodies, customelectronic designs, software pieces and 3D printed parts complete the itemsemployed. The system designed is modular, it allows creating melodies withchords and percussion or just melodies or perform as a beatbox or a melody box.The main interaction with the system is made using Lego-type building blocks.Tests have demonstrated its versatility and ease of use, as well as itsusefulness in music learning for both children and adults.
标题:实时说话肖像合成的音频特征提取比较分析
链接:https://arxiv.org/abs/2411.13209
备注:16 pages, 6 figures, 3 tables. submitted to MDPI journal in as Big Data and Cognitive Computing
摘要:本文研究了用于面试官培训的实时讲话头生成的集成,重点是克服音频特征提取(AFE)中的挑战,这通常会在实时应用中引入延迟并限制响应能力。为了解决这些问题,我们提出并实现了一个完全集成的系统,用Open AI的Whisper取代传统的AFE模型,利用其编码器优化处理并提高整体系统效率。我们对三个不同数据集的两个开源实时模型的评估表明,Whisper不仅加快了处理速度,而且还提高了渲染质量的特定方面,从而实现了更逼真和响应更快的通话头交互。这些进步使该系统成为沉浸式、交互式培训应用程序的更有效工具,扩大了人工智能驱动的化身在面试官培训中的潜力。
摘要:This paper examines the integration of real-time talking-head generation forinterviewer training, focusing on overcoming challenges in Audio FeatureExtraction (AFE), which often introduces latency and limits responsiveness inreal-time applications. To address these issues, we propose and implement afully integrated system that replaces conventional AFE models with Open AI'sWhisper, leveraging its encoder to optimize processing and improve overallsystem efficiency. Our evaluation of two open-source real-time models acrossthree different datasets shows that Whisper not only accelerates processing butalso improves specific aspects of rendering quality, resulting in morerealistic and responsive talking-head interactions. These advancements make thesystem a more effective tool for immersive, interactive training applications,expanding the potential of AI-driven avatars in interviewer training.
标题:SONNET:通过利用模拟音频增强时间延迟估计
链接:https://arxiv.org/abs/2411.13179
摘要:时间延迟估计或到达时间差估计是多个定位应用(例如,定位、到达方向和自校准)的关键组成部分。任务是估计信号到达两个不同传感器之间的时间差。对于音频传感器模态,大多数当前系统都基于经典方法,例如广义互相关相位变换(GCC-PHAT)方法。在本文中,我们证明了基于学习的方法,即使是基于合成数据,也可以在新的真实世界数据上显着优于GCC-PHAT。为了克服任务缺乏地面真实数据的问题,我们在一个足够大和多样的模拟数据集上训练我们的模型,并捕捉现实世界问题的相关特征。我们提供了我们的训练模型SONNET(时移模拟优化神经网络估计),它可实时运行,并可用于许多真实数据应用程序的新数据,即无需重新训练。我们进一步证明,与经典方法相比,使用我们的模型时,自校准的下游任务的性能大大提高。
摘要:Time delay estimation or Time-Difference-Of-Arrival estimates is a criticalcomponent for multiple localization applications such as multilateration,direction of arrival, and self-calibration. The task is to estimate the timedifference between a signal arriving at two different sensors. For the audiosensor modality, most current systems are based on classical methods such asthe Generalized Cross-Correlation Phase Transform (GCC-PHAT) method. In thispaper we demonstrate that learning based methods can, even based on syntheticdata, significantly outperform GCC-PHAT on novel real world data. To overcomethe lack of data with ground truth for the task, we train our model on asimulated dataset which is sufficiently large and varied, and that captures therelevant characteristics of the real world problem. We provide our trainedmodel, SONNET (Simulation Optimized Neural Network Estimator of Timeshifts),which is runnable in real-time and works on novel data out of the box for manyreal data applications, i.e. without re-training. We further demonstrategreatly improved performance on the downstream task of self-calibration whenusing our model compared to classical methods.
标题:Hard-Synth:使用Zero-Shot TTC和LLM合成用于ASC的各种硬样本
链接:https://arxiv.org/abs/2411.13159
摘要:文语转换(TTS)模型已被广泛采用,以增强使用纯文本语料库的自动语音识别(ASR)系统,从而降低标记真实语音数据的成本。现有的研究主要利用额外的文本数据和预定义的语音风格支持的TTS模型。在本文中,我们提出了Hard-Synth,一种新的ASR数据增强方法,利用大型语言模型(LLM)和先进的zero-shot TTS。我们的方法采用LLM通过重写生成不同的域内文本,而不依赖于额外的文本数据。而不是使用预定义的语音风格,我们引入了一个硬提示选择方法与zero-shot TTS克隆语音风格,ASR模型发现具有挑战性的认识。实验结果表明,Hard-Synth显著增强了Conformer模型,在LibriSpeech dev/test-other子集上实现了6.5%/4.4%的相对单词错误率(WER)降低。此外,我们还证明了Hard-Synth是数据高效的,能够减少ASR中的偏差。
摘要:Text-to-speech (TTS) models have been widely adopted to enhance automaticspeech recognition (ASR) systems using text-only corpora, thereby reducing thecost of labeling real speech data. Existing research primarily utilizesadditional text data and predefined speech styles supported by TTS models. Inthis paper, we propose Hard-Synth, a novel ASR data augmentation method thatleverages large language models (LLMs) and advanced zero-shot TTS. Our approachemploys LLMs to generate diverse in-domain text through rewriting, withoutrelying on additional text data. Rather than using predefined speech styles, weintroduce a hard prompt selection method with zero-shot TTS to clone speechstyles that the ASR model finds challenging to recognize. Experimentsdemonstrate that Hard-Synth significantly enhances the Conformer model,achieving relative word error rate (WER) reductions of 6.5\%/4.4\% onLibriSpeech dev/test-other subsets. Additionally, we show that Hard-Synth isdata-efficient and capable of reducing bias in ASR.
标题:ESARM:通过自动排名演示的奖励模型进行3D情感演讲到动画
链接:https://arxiv.org/abs/2411.13089
备注:Accepted by the 26th IEEE International Conference on High Performance Computing and Communications (HPCC2024)
摘要:本文提出了一种新颖的3D语音转动画(STA)生成框架,旨在解决现有模型在生成多样化且情感共鸣的动画方面的缺点。目前的STA模型通常生成缺乏情感深度和多样性的动画,无法与人类的期望保持一致。为了克服这些限制,我们引入了一种新的STA模型加上奖励模型。这种组合使得能够通过交叉耦合训练方法在音频条件下解耦情感和内容。此外,我们开发了一种训练方法,该方法利用生成的面部动画的自动质量评估来指导强化学习过程。这种方法鼓励STA模型探索更广泛的可能性,从而产生多样化和情感表达的优质面部动画。我们在基准数据集上进行了广泛的实证实验,结果验证了我们提出的框架在生成高质量,情感丰富的3D动画,更好地与人类的喜好。
摘要:This paper proposes a novel 3D speech-to-animation (STA) generation frameworkdesigned to address the shortcomings of existing models in producing diverseand emotionally resonant animations. Current STA models often generateanimations that lack emotional depth and variety, failing to align with humanexpectations. To overcome these limitations, we introduce a novel STA modelcoupled with a reward model. This combination enables the decoupling of emotionand content under audio conditions through a cross-coupling training approach.Additionally, we develop a training methodology that leverages automaticquality evaluation of generated facial animations to guide the reinforcementlearning process. This methodology encourages the STA model to explore abroader range of possibilities, resulting in the generation of diverse andemotionally expressive facial animations of superior quality. We conductextensive empirical experiments on a benchmark dataset, and the resultsvalidate the effectiveness of our proposed framework in generatinghigh-quality, emotionally rich 3D animations that are better aligned with humanpreferences.
标题:基于能量的特征和双LSTM神经网络用于基于脑电波的音乐和语音分类
链接:https://arxiv.org/abs/2411.13217
备注:12 pages
摘要:人类大脑以多种方式接收刺激;其中,音频构成了大脑有关通信,娱乐,警告等相关刺激的重要来源。在这种情况下,本手稿的目的是推进大脑对不同类型音乐和不同性质声音的反应分类:语音和音乐。为此,我们设计了两个不同的实验,分别从不同音乐类型的歌曲和不同语言的句子中获取被试的脑电信号。有了这个,提出了一种新的方案来表征大脑信号进行分类;该方案是基于在不同EEG通道测量的能量与bi-LSTM神经网络的使用之间的关系构建的特征矩阵。利用所获得的数据,进行关于语音和音乐之间的基于EEG的分类、不同的音乐流派以及受试者是否喜欢所听的歌曲的评估。实验结果表明,该方案具有良好的性能.对二进制音频类型的分类取得了98.66%的成功率。在4种音乐流派之间的多类分类中,准确率达到61.59%,音乐品味的二元分类结果上升到96.96%。
摘要:The human brain receives stimuli in multiple ways; among them, audioconstitutes an important source of relevant stimuli for the brain regardingcommunication, amusement, warning, etc. In this context, the aim of thismanuscript is to advance in the classification of brain responses to music ofdiverse genres and to sounds of different nature: speech and music. For thispurpose, two different experiments have been designed to acquiere EEG signalsfrom subjects listening to songs of different musical genres and sentences invarious languages. With this, a novel scheme is proposed to characterize brainsignals for their classification; this scheme is based on the construction of afeature matrix built on relations between energy measured at the different EEGchannels and the usage of a bi-LSTM neural network. With the data obtained,evaluations regarding EEG-based classification between speech and music,different musical genres, and whether the subject likes the song listened to ornot are carried out. The experiments unveil satisfactory performance to theproposed scheme. The results obtained for binary audio type classificationattain 98.66% of success. In multi-class classification between 4 musicalgenres, the accuracy attained is 61.59%, and results for binary classificationof musical taste rise to 96.96%.
标题:基于能量的特征和双LSTM神经网络用于基于脑电波的音乐和语音分类
链接:https://arxiv.org/abs/2411.13217
备注:12 pages
摘要:人类大脑以多种方式接收刺激;其中,音频构成了大脑有关通信,娱乐,警告等相关刺激的重要来源。在这种情况下,本手稿的目的是推进大脑对不同类型音乐和不同性质声音的反应分类:语音和音乐。为此,我们设计了两个不同的实验,分别从不同音乐类型的歌曲和不同语言的句子中获取脑电信号。有了这个,提出了一种新的方案来表征大脑信号进行分类;该方案是基于在不同EEG通道测量的能量与bi-LSTM神经网络的使用之间的关系构建的特征矩阵。利用所获得的数据,进行关于语音和音乐之间的基于EEG的分类、不同的音乐流派以及受试者是否喜欢所听的歌曲的评估。实验结果表明,该方案具有良好的性能.对二进制音频类型的分类取得了98.66%的成功率。在4种音乐流派之间的多类分类中,准确率达到61.59%,音乐品味的二元分类结果上升到96.96%。
摘要:The human brain receives stimuli in multiple ways; among them, audioconstitutes an important source of relevant stimuli for the brain regardingcommunication, amusement, warning, etc. In this context, the aim of thismanuscript is to advance in the classification of brain responses to music ofdiverse genres and to sounds of different nature: speech and music. For thispurpose, two different experiments have been designed to acquiere EEG signalsfrom subjects listening to songs of different musical genres and sentences invarious languages. With this, a novel scheme is proposed to characterize brainsignals for their classification; this scheme is based on the construction of afeature matrix built on relations between energy measured at the different EEGchannels and the usage of a bi-LSTM neural network. With the data obtained,evaluations regarding EEG-based classification between speech and music,different musical genres, and whether the subject likes the song listened to ornot are carried out. The experiments unveil satisfactory performance to theproposed scheme. The results obtained for binary audio type classificationattain 98.66% of success. In multi-class classification between 4 musicalgenres, the accuracy attained is 61.59%, and results for binary classificationof musical taste rise to 96.96%.
标题:用于声音事件定位和检测的类增量学习
链接:https://arxiv.org/abs/2411.12830
摘要:本文研究了类增量学习(CIL)的声音事件定位和检测(SELD)任务的可行性。该方法具有增量学习器,可以独立学习新的声音类,同时保留旧类的知识。通过基于均方误差的蒸馏损失来实现连续学习,以最小化后续学习器之间的输出差异。实验在TAU-NIGENS Spatial Sound Events 2021数据集上进行,该数据集包括12种不同的声音类别,并证明了所提出的方法的有效性。我们从学习8个类开始,并在下一阶段介绍4个新类。在增量阶段之后,系统在学习的类的完整集合上进行评估。结果表明,对于这个现实的数据集,我们提出的方法成功地保持了所有指标的基线性能。
摘要:This paper investigates the feasibility of class-incremental learning (CIL)for Sound Event Localization and Detection (SELD) tasks. The method features anincremental learner that can learn new sound classes independently whilepreserving knowledge of old classes. The continual learning is achieved througha mean square error-based distillation loss to minimize output discrepanciesbetween subsequent learners. The experiments are conducted on the TAU-NIGENSSpatial Sound Events 2021 dataset, which includes 12 different sound classesand demonstrate the efficacy of proposed method. We begin by learning 8 classesand introduce the 4 new classes at next stage. After the incremental phase, thesystem is evaluated on the full set of learned classes. Results show that, forthis realistic dataset, our proposed method successfully maintains baselineperformance across all metrics.
标题:CAFE阿尔及利亚方言法语和英语的新型代码切换数据集
链接:https://arxiv.org/abs/2411.13424
备注:24 pages, submitted to tallip
摘要:本文介绍并公开发布(接受后提供数据下载链接)CAFE -阿尔及利亚方言,法语和英语之间的第一个代码转换数据集。CAFE语音数据是独特的:(a)其自发的说话风格在体内人与人的对话捕捉现象,如代码转换和重叠的讲话,(b)解决了北非阿拉伯语方言的独特语言挑战;(c)CAFE捕捉不同社会语言学背景下阿尔及利亚各地的方言变化。CAFE数据包含大约37小时的语音,其中2小时36分钟的子集CAFE-small通过人工注释发布,包括语音分割,转录,代码转换点的明确注释,重叠语音以及其他事件,如噪音和笑声等。其余约34.58小时含有假标签transcription。除了数据发布之外,该论文还强调了使用最先进的自动语音识别(ASR)模型(例如Whisper large-v2,3和PromptingWhisper)来处理此类内容所面临的挑战。接下来,我们使用上述Whisper模型对CAFE数据进行基准测试,并展示了精心设计的数据处理管道和先进的解码技术如何提高ASR性能,混合错误率(MER)为0.310,字符错误率(CER)为0.329,字错误率(WER)为0.538。
摘要:The paper introduces and publicly releases (Data download link availableafter acceptance) CAFE -- the first Code-switching dataset between Algeriandialect, French, and english languages. The CAFE speech data is unique for (a)its spontaneous speaking style in vivo human-human conversation capturingphenomena like code-switching and overlapping speech, (b) addresses distinctlinguistic challenges in North African Arabic dialect; (c) the CAFE capturesdialectal variations from various parts of Algeria within differentsociolinguistic contexts. CAFE data contains approximately 37 hours of speech,with a subset, CAFE-small, of 2 hours and 36 minutes released with manual humanannotation including speech segmentation, transcription, explicit annotation ofcode-switching points, overlapping speech, and other events such as noises, andlaughter among others. The rest approximately 34.58 hours contain pseudo labeltranscriptions. In addition to the data release, the paper also highlighted thechallenges of using state-of-the-art Automatic Speech Recognition (ASR) modelssuch as Whisper large-v2,3 and PromptingWhisper to handle such content.Following, we benchmark CAFE data with the aforementioned Whisper models andshow how well-designed data processing pipelines and advanced decodingtechniques can improve the ASR performance in terms of Mixed Error Rate (MER)of 0.310, Character Error Rate (CER) of 0.329 and Word Error Rate (WER) of0.538.
标题:I2 TTC:具有空间感知的图像指示沉浸式文本到语音合成
链接:https://arxiv.org/abs/2411.13314
备注:5pages,4figures
摘要:控制语音合成的风格和特征对于使输出适应特定上下文和用户需求至关重要。以前的文语转换(TTS)工作主要集中在产生自然发音的语音的技术方面,如语调,节奏和清晰度。然而,他们忽略了一个事实,即越来越重视合成语音的空间感知,这可能会在游戏和虚拟现实中提供身临其境的体验。为了解决这个问题,在本文中,我们提出了一种新的多模态TTS方法,即图像指示沉浸式文本到语音合成(I2 TTS)。具体来说,我们引入了一个场景提示编码器,将视觉场景提示直接集成到合成管道中,以控制语音生成过程。此外,我们提出了一种混响分类和细化技术,可以调整合成的梅尔谱图以增强沉浸式体验,确保所涉及的混响条件与场景准确匹配。实验结果表明,我们的模型实现了高质量的场景和空间匹配,而不影响语音的自然性,标志着在上下文感知语音合成领域的重大进步。项目演示页面:https://spatialTTS.github.io/索引术语-语音合成,场景提示,空间感知
摘要:Controlling the style and characteristics of speech synthesis is crucial foradapting the output to specific contexts and user requirements. PreviousText-to-speech (TTS) works have focused primarily on the technical aspects ofproducing natural-sounding speech, such as intonation, rhythm, and clarity.However, they overlook the fact that there is a growing emphasis on spatialperception of synthesized speech, which may provide immersive experience ingaming and virtual reality. To solve this issue, in this paper, we present anovel multi-modal TTS approach, namely Image-indicated Immersive Text-to-speechSynthesis (I2TTS). Specifically, we introduce a scene prompt encoder thatintegrates visual scene prompts directly into the synthesis pipeline to controlthe speech generation process. Additionally, we propose a reverberationclassification and refinement technique that adjusts the synthesizedmel-spectrogram to enhance the immersive experience, ensuring that the involvedreverberation condition matches the scene accurately. Experimental resultsdemonstrate that our model achieves high-quality scene and spatial matchingwithout compromising speech naturalness, marking a significant advancement inthe field of context-aware speech synthesis. Project demo page:https://spatialTTS.github.io/ Index Terms-Speech synthesis, scene prompt,spatial perception
标题:使用乐高积木和Raspberry Pi构建音乐
链接:https://arxiv.org/abs/2411.13224
备注:21 pages
摘要:在本文中,一个系统,以建立一个直观的和可访问的方式,与乐高积木的音乐。该系统利用技术提供的新的强大和廉价的可能性,以新的方式制造旧的东西。Raspberry Pi用于控制系统并运行必要的算法,定制的乐高积木用于构建旋律,定制电子设计,软件件和3D打印部件完成所使用的项目。该系统的设计是模块化的,它允许创建旋律与和弦和打击乐或只是旋律或作为一个节拍箱或旋律箱执行。与系统的主要交互是使用乐高积木。测试证明了它的多功能性和易用性,以及它在儿童和成人音乐学习中的有用性。
摘要:In this paper, a system to build music in an intuitive and accessible way,with Lego bricks, is presented. The system makes use of the new powerful andcheap possibilities that technology offers for making old things in a new way.The Raspberry Pi is used to control the system and run the necessaryalgorithms, customized Lego bricks are used for building melodies, customelectronic designs, software pieces and 3D printed parts complete the itemsemployed. The system designed is modular, it allows creating melodies withchords and percussion or just melodies or perform as a beatbox or a melody box.The main interaction with the system is made using Lego-type building blocks.Tests have demonstrated its versatility and ease of use, as well as itsusefulness in music learning for both children and adults.
标题:实时说话肖像合成的音频特征提取比较分析
链接:https://arxiv.org/abs/2411.13209
备注:16 pages, 6 figures, 3 tables. submitted to MDPI journal in as Big Data and Cognitive Computing
摘要:本文研究了用于面试官培训的实时讲话头生成的集成,重点是克服音频特征提取(AFE)中的挑战,这通常会在实时应用中引入延迟并限制响应能力。为了解决这些问题,我们提出并实现了一个完全集成的系统,用Open AI的Whisper取代传统的AFE模型,利用其编码器优化处理并提高整体系统效率。我们对三个不同数据集的两个开源实时模型的评估表明,Whisper不仅加快了处理速度,而且还提高了渲染质量的特定方面,从而实现了更逼真和响应更快的通话头交互。这些进步使该系统成为沉浸式交互式培训应用程序的更有效工具,扩大了人工智能驱动的化身在面试官培训中的潜力。
摘要:This paper examines the integration of real-time talking-head generation forinterviewer training, focusing on overcoming challenges in Audio FeatureExtraction (AFE), which often introduces latency and limits responsiveness inreal-time applications. To address these issues, we propose and implement afully integrated system that replaces conventional AFE models with Open AI'sWhisper, leveraging its encoder to optimize processing and improve overallsystem efficiency. Our evaluation of two open-source real-time models acrossthree different datasets shows that Whisper not only accelerates processing butalso improves specific aspects of rendering quality, resulting in morerealistic and responsive talking-head interactions. These advancements make thesystem a more effective tool for immersive, interactive training applications,expanding the potential of AI-driven avatars in interviewer training.
标题:SONNET:通过利用模拟音频增强时间延迟估计
链接:https://arxiv.org/abs/2411.13179
摘要:时间延迟估计或到达时间差估计是多个定位应用(例如,定位、到达方向和自校准)的关键组成部分。任务是估计信号到达两个不同传感器之间的时间差。对于音频传感器模态,大多数当前系统都基于经典方法,例如广义互相关相位变换(GCC-PHAT)方法。在本文中,我们证明了基于学习的方法,即使是基于合成数据,也可以在新的真实世界数据上显着优于GCC-PHAT。为了克服任务缺乏地面真实数据的问题,我们在一个足够大和多样的模拟数据集上训练我们的模型,并捕捉现实世界问题的相关特征。我们提供了我们的训练模型SONNET(时移模拟优化神经网络估计),它可实时运行,并可用于许多真实数据应用程序的新数据,即无需重新训练。我们进一步证明,与经典方法相比,使用我们的模型时,自校准的下游任务的性能大大提高。
摘要:Time delay estimation or Time-Difference-Of-Arrival estimates is a criticalcomponent for multiple localization applications such as multilateration,direction of arrival, and self-calibration. The task is to estimate the timedifference between a signal arriving at two different sensors. For the audiosensor modality, most current systems are based on classical methods such asthe Generalized Cross-Correlation Phase Transform (GCC-PHAT) method. In thispaper we demonstrate that learning based methods can, even based on syntheticdata, significantly outperform GCC-PHAT on novel real world data. To overcomethe lack of data with ground truth for the task, we train our model on asimulated dataset which is sufficiently large and varied, and that captures therelevant characteristics of the real world problem. We provide our trainedmodel, SONNET (Simulation Optimized Neural Network Estimator of Timeshifts),which is runnable in real-time and works on novel data out of the box for manyreal data applications, i.e. without re-training. We further demonstrategreatly improved performance on the downstream task of self-calibration whenusing our model compared to classical methods.
标题:Hard-Synth:使用Zero-Shot TTC和LLM合成用于ASC的各种硬样本
链接:https://arxiv.org/abs/2411.13159
摘要:文语转换(TTS)模型已被广泛采用,以增强使用纯文本语料库的自动语音识别(ASR)系统,从而降低标记真实语音数据的成本。现有的研究主要利用额外的文本数据和预定义的语音风格支持的TTS模型。在本文中,我们提出了Hard-Synth,一种新的ASR数据增强方法,利用大型语言模型(LLM)和先进的zero-shot TTS。我们的方法采用LLM通过重写生成不同的域内文本,而不依赖于额外的文本数据。而不是使用预定义的语音风格,我们引入了一个硬提示选择方法与zero-shot TTS克隆语音风格,ASR模型发现具有挑战性的认识。实验结果表明,Hard-Synth显著增强了Conformer模型,在LibriSpeech dev/test-other子集上实现了6.5%/4.4%的相对单词错误率(WER)降低。此外,我们还证明了Hard-Synth是数据高效的,能够减少ASR中的偏差。
摘要:Text-to-speech (TTS) models have been widely adopted to enhance automaticspeech recognition (ASR) systems using text-only corpora, thereby reducing thecost of labeling real speech data. Existing research primarily utilizesadditional text data and predefined speech styles supported by TTS models. Inthis paper, we propose Hard-Synth, a novel ASR data augmentation method thatleverages large language models (LLMs) and advanced zero-shot TTS. Our approachemploys LLMs to generate diverse in-domain text through rewriting, withoutrelying on additional text data. Rather than using predefined speech styles, weintroduce a hard prompt selection method with zero-shot TTS to clone speechstyles that the ASR model finds challenging to recognize. Experimentsdemonstrate that Hard-Synth significantly enhances the Conformer model,achieving relative word error rate (WER) reductions of 6.5\%/4.4\% onLibriSpeech dev/test-other subsets. Additionally, we show that Hard-Synth isdata-efficient and capable of reducing bias in ASR.
标题:ESARM:通过自动排名演示的奖励模型进行3D情感演讲到动画
链接:https://arxiv.org/abs/2411.13089
备注:Accepted by the 26th IEEE International Conference on High Performance Computing and Communications (HPCC2024)
摘要:本文提出了一种新颖的3D语音转动画(STA)生成框架,旨在解决现有模型在生成多样化且情感共鸣的动画方面的缺点。目前的STA模型通常生成缺乏情感深度和多样性的动画,无法与人类的期望保持一致。为了克服这些限制,我们引入了一种新的STA模型加上奖励模型。这种组合使得能够通过交叉耦合训练方法在音频条件下解耦情感和内容。此外,我们开发了一种训练方法,该方法利用生成的面部动画的自动质量评估来指导强化学习过程。这种方法鼓励STA模型探索更广泛的可能性,从而产生多样化和情感表达的优质面部动画。我们在基准数据集上进行了广泛的实证实验,结果验证了我们提出的框架在生成高质量,情感丰富的3D动画,更好地与人类的喜好。
摘要:This paper proposes a novel 3D speech-to-animation (STA) generation frameworkdesigned to address the shortcomings of existing models in producing diverseand emotionally resonant animations. Current STA models often generateanimations that lack emotional depth and variety, failing to align with humanexpectations. To overcome these limitations, we introduce a novel STA modelcoupled with a reward model. This combination enables the decoupling of emotionand content under audio conditions through a cross-coupling training approach.Additionally, we develop a training methodology that leverages automaticquality evaluation of generated facial animations to guide the reinforcementlearning process. This methodology encourages the STA model to explore abroader range of possibilities, resulting in the generation of diverse andemotionally expressive facial animations of superior quality. We conductextensive empirical experiments on a benchmark dataset, and the resultsvalidate the effectiveness of our proposed framework in generatinghigh-quality, emotionally rich 3D animations that are better aligned with humanpreferences.
