今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音

【1】Evaluating ASR Confidence Scores for Automated Error Detection in  User-Assisted Correction Interfaces

标题:评估用于用户辅助纠正界面中自动错误检测的ASB置信度分数
链接:https://arxiv.org/abs/2503.15124
作者:Korbinian Kuhn,  Verena Kersken,  Gottfried Zimmermann
备注:7 pages, 1 figure, to be published in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA '25)
摘要:尽管自动语音识别(ASR)取得了进展,但转录错误仍然存在,需要手动纠正。表示ASR结果的确定性的置信度分数可以帮助用户识别和纠正错误。本研究通过对端到端ASR模型的全面分析和一项有36名参与者的用户研究,评估了错误检测置信度分数的可靠性。结果表明,虽然置信度得分与转录准确性相关,但它们的错误检测性能是有限的。分类器经常错过错误或产生许多误报,破坏了它们的实际效用。基于信心的错误检测既没有提高纠错效率,也被认为是有益的参与者。这些发现突出了置信度分数的局限性,以及需要更复杂的方法来改善用户交互和ASR结果的可解释性。
摘要:Despite advances in Automatic Speech Recognition (ASR), transcription errorspersist and require manual correction. Confidence scores, which indicate thecertainty of ASR results, could assist users in identifying and correctingerrors. This study evaluates the reliability of confidence scores for errordetection through a comprehensive analysis of end-to-end ASR models and a userstudy with 36 participants. The results show that while confidence scorescorrelate with transcription accuracy, their error detection performance islimited. Classifiers frequently miss errors or generate many false positives,undermining their practical utility. Confidence-based error detection neitherimproved correction efficiency nor was perceived as helpful by participants.These findings highlight the limitations of confidence scores and the need formore sophisticated approaches to improve user interaction and explainability ofASR results.

【2】 Communication Access Real-Time Translation Through Collaborative  Correction of Automatic Speech Recognition
标题:通过自动语音识别的协同纠正实现通信访问实时翻译
链接:https://arxiv.org/abs/2503.15120
作者:Korbinian Kuhn,  Verena Kersken,  Gottfried Zimmermann
备注:8 pages, 2 figures, to be published in Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (CHI EA '25)
摘要:通信接入实时翻译(CART)是聋人和重听人(DHH)的基本无障碍服务,但训练有素的人员的成本和稀缺性限制了其可用性。虽然自动语音识别(ASR)提供了一种廉价且可扩展的替代方案,但转录错误可能会导致严重的可访问性问题。由非专业人员进行的ASR实时校正提出了一种解决这些限制的未充分探索的CART工作流程。我们进行了一项有75名参与者的用户研究,以评估此工作流程的可行性和效率。作为补充,我们与25名DHH人员举行了焦点小组,以确定可接受的准确性水平和影响实时字幕可访问性的因素。结果表明,协作编辑可以提高转录准确性的程度,DHH用户积极评价它的可理解性。焦点小组还表明,人类为改善字幕所做的努力受到高度重视,支持半自动化方法作为独立ASR和传统CART服务的替代方案。
摘要:Communication access real-time translation (CART) is an essentialaccessibility service for d/Deaf and hard of hearing (DHH) individuals, but thecost and scarcity of trained personnel limit its availability. While AutomaticSpeech Recognition (ASR) offers a cheap and scalable alternative, transcriptionerrors can lead to serious accessibility issues. Real-time correction of ASR bynon-professionals presents an under-explored CART workflow that addresses theselimitations. We conducted a user study with 75 participants to evaluate thefeasibility and efficiency of this workflow. Complementary, we held focusgroups with 25 DHH individuals to identify acceptable accuracy levels andfactors affecting the accessibility of real-time captioning. Results suggestthat collaborative editing can improve transcription accuracy to the extentthat DHH users rate it positively regarding understandability. Focus groupsalso showed that human effort to improve captioning is highly valued,supporting a semi-automated approach as an alternative to stand-alone ASR andtraditional CART services.

【3】 InsectSet459: an open dataset of insect sounds for bioacoustic machine  learning
标题:InsectSet 459:用于生物声学机器学习的昆虫声音开放数据集
链接:https://arxiv.org/abs/2503.15074
作者:Marius Faiß,  Burooj Ghani,  Dan Stowell
摘要:自动识别昆虫的声音可以帮助我们了解世界各地不断变化的生物多样性趋势,但即使是深度学习也很难识别昆虫的声音。我们提出了一个新的数据集,包括26399音频文件,从459种直翅目和蝉。这是第一个大规模的昆虫声音数据集,可以很容易地应用于开发新的深度学习方法。它的录音是由各种录音机使用不同的采样率,以捕捉昆虫产生的极其广泛的频率范围。我们使用两种最先进的深度学习分类器进行性能基准测试,表现出良好的性能,但在声学昆虫分类方面也有很大的改进空间。该数据集可以作为实现昆虫监测工作流程的现实测试案例,并作为开发可以处理高度可变频率和/或采样率的音频表示方法的挑战性基础。
摘要:Automatic recognition of insect sound could help us understand changingbiodiversity trends around the world -- but insect sounds are challenging torecognize even for deep learning. We present a new dataset comprised of 26399audio files, from 459 species of Orthoptera and Cicadidae. It is the firstlarge-scale dataset of insect sound that is easily applicable for developingnovel deep-learning methods. Its recordings were made with a variety of audiorecorders using varying sample rates to capture the extremely broad range offrequencies that insects produce. We benchmark performance with twostate-of-the-art deep learning classifiers, demonstrating good performance butalso significant room for improvement in acoustic insect classification. Thisdataset can serve as a realistic test case for implementing insect monitoringworkflows, and as a challenging basis for the development of audiorepresentation methods that can handle highly variable frequencies and/orsample rates.

【4】 Shushing! Let's Imagine an Authentic Speech from the Silent Video
标题:嘘!让我们想象一下无声视频中的真实演讲
链接:https://arxiv.org/abs/2503.14928
作者:Jiaxin Ye,  Hongming Shan
备注:Project Page: this https URL
摘要:视觉引导的语音生成旨在从面部外观或嘴唇运动产生真实的语音,而不依赖于听觉信号,为电影制作中的配音和帮助患有失声症的个人等应用提供了巨大的潜力。尽管最近的进展,现有的方法努力实现统一的跨模态对齐语义,音色,和情感韵律的视觉线索,促使我们提出一致的视频到语音(CV2S)作为一个扩展的任务,以提高跨模态的一致性。为了应对新出现的挑战,我们引入了ImaginTalk,这是一种新颖的跨模态扩散框架,仅使用视觉输入生成忠实的语音,在离散空间内操作。具体来说,我们提出了一个离散唇对齐器,预测离散语音令牌从唇视频捕捉语义信息,而错误检测器识别未对齐的令牌,随后通过掩蔽语言建模与BERT细化。为了进一步提高所生成的语音的表现力,我们开发了一个风格扩散Transformer配备了一个脸式适配器,自适应定制的身份和韵律动态跨通道和时间维度,同时确保同步与唇感知语义功能。大量的实验表明,与最先进的基线相比,ImaginTalk可以生成具有更准确的语义细节和更强的音色和情感表现力的高保真语音。演示显示在我们的项目页面:www.example.com。
摘要:Vision-guided speech generation aims to produce authentic speech from facialappearance or lip motions without relying on auditory signals, offeringsignificant potential for applications such as dubbing in filmmaking andassisting individuals with aphonia. Despite recent progress, existing methodsstruggle to achieve unified cross-modal alignment across semantics, timbre, andemotional prosody from visual cues, prompting us to propose ConsistentVideo-to-Speech (CV2S) as an extended task to enhance cross-modal consistency.To tackle emerging challenges, we introduce ImaginTalk, a novel cross-modaldiffusion framework that generates faithful speech using only visual input,operating within a discrete space. Specifically, we propose a discrete lipaligner that predicts discrete speech tokens from lip videos to capturesemantic information, while an error detector identifies misaligned tokens,which are subsequently refined through masked language modeling with BERT. Tofurther enhance the expressiveness of the generated speech, we develop a stylediffusion transformer equipped with a face-style adapter that adaptivelycustomizes identity and prosody dynamics across both the channel and temporaldimensions while ensuring synchronization with lip-aware semantic features.Extensive experiments demonstrate that ImaginTalk can generate high-fidelityspeech with more accurate semantic details and greater expressiveness in timbreand emotion compared to state-of-the-art baselines. Demos are shown at ourproject page: https://imagintalk.github.io.

【5】 PANDORA: Diffusion Policy Learning for Dexterous Robotic Piano Playing
标题:PANDORA:灵巧机器人钢琴演奏的扩散政策学习
链接:https://arxiv.org/abs/2503.14545
作者:Yanjia Huang,  Renjie Li,  Zhengzhong Tu
摘要:我们提出了PANDORA,一种新的基于扩散的政策学习框架,专门为灵巧的机器人钢琴性能。我们的方法采用了一个有条件的U-Net架构,增强了基于FiLM的全局调节,迭代去噪噪声动作序列平滑,高维轨迹。为了实现精确的关键执行加上富有表现力的音乐表现,我们设计了一个复合奖励函数,该函数集成了特定于任务的准确性,音频保真度和来自大型语言模型(LLM)Oracle的高级语义反馈。LLM oracle评估音乐表现力和风格上的细微差别,从而实现动态的、针对具体手型的奖励调整。通过残余逆运动学细化策略进一步增强,PANDORA在ROBOPIANIST环境中实现了最先进的性能,在精度和表现力方面都显着优于基线。消融研究验证了基于扩散的去噪和LLM驱动的语义反馈在增强机器人音乐才能方面的关键贡献。视频网址:https://taco-group.github.io/PANDORA
摘要:We present PANDORA, a novel diffusion-based policy learning frameworkdesigned specifically for dexterous robotic piano performance. Our approachemploys a conditional U-Net architecture enhanced with FiLM-based globalconditioning, which iteratively denoises noisy action sequences into smooth,high-dimensional trajectories. To achieve precise key execution coupled withexpressive musical performance, we design a composite reward function thatintegrates task-specific accuracy, audio fidelity, and high-level semanticfeedback from a large language model (LLM) oracle. The LLM oracle assessesmusical expressiveness and stylistic nuances, enabling dynamic, hand-specificreward adjustments. Further augmented by a residual inverse-kinematicsrefinement policy, PANDORA achieves state-of-the-art performance in theROBOPIANIST environment, significantly outperforming baselines in bothprecision and expressiveness. Ablation studies validate the criticalcontributions of diffusion-based denoising and LLM-driven semantic feedback inenhancing robotic musicianship. Videos available at:https://taco-group.github.io/PANDORA

【6】 Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
标题:索拉:迈向倾听声学背景的以演讲为导向的法学硕士
链接:https://arxiv.org/abs/2503.15338
作者:Junyi Ao,  Dekun Chen,  Xiaohai Tian,  Wenjie Feng,  Jun Zhang,  Lu Lu,  Yuxuan Wang,  Haizhou Li,  Zhizheng Wu
摘要:大型语言模型(LLM)最近表现出了非凡的能力,不仅可以处理文本,还可以处理语音和音频等多模态输入。然而,大多数现有的模型主要集中在使用文本指令分析输入信号,忽略了语音指令和音频混合并作为模型输入的场景。为了解决这些挑战,我们引入Solla,这是一种新颖的框架,旨在理解基于语音的问题并同时听到声学背景。Solla集成了一个音频标记模块来有效地识别和表示音频事件,以及一个ASR辅助预测方法来提高对口语内容的理解。为了严格评估Solla和其他公开可用的模型,我们提出了一个名为SA-Eval的新基准数据集,它包括三个任务:音频事件分类,音频字幕和音频问答。SA-Eval具有各种说话风格的多样化语音教学,包括两个难度级别,简单和困难,以捕捉真实世界的声学条件范围。实验结果表明,Solla在简单和困难的测试集上的表现与基线模型相当或优于基线模型,强调了其在联合理解语音和音频方面的有效性。
摘要:Large Language Models (LLMs) have recently shown remarkable ability toprocess not only text but also multimodal inputs such as speech and audio.However, most existing models primarily focus on analyzing input signals usingtext instructions, overlooking scenarios in which speech instructions and audioare mixed and serve as inputs to the model. To address these challenges, weintroduce Solla, a novel framework designed to understand speech-basedquestions and hear the acoustic context concurrently. Solla incorporates anaudio tagging module to effectively identify and represent audio events, aswell as an ASR-assisted prediction method to improve comprehension of spokencontent. To rigorously evaluate Solla and other publicly available models, wepropose a new benchmark dataset called SA-Eval, which includes three tasks:audio event classification, audio captioning, and audio question answering.SA-Eval has diverse speech instruction with various speaking styles,encompassing two difficulty levels, easy and hard, to capture the range ofreal-world acoustic conditions. Experimental results show that Solla performson par with or outperforms baseline models on both the easy and hard test sets,underscoring its effectiveness in jointly understanding speech and audio.

eess.AS音频处理

【1】 Solla: Towards a Speech-Oriented LLM That Hears Acoustic Context
标题:索拉:迈向倾听声学背景的以演讲为导向的法学硕士
链接:https://arxiv.org/abs/2503.15338
作者:Junyi Ao,  Dekun Chen,  Xiaohai Tian,  Wenjie Feng,  Jun Zhang,  Lu Lu,  Yuxuan Wang,  Haizhou Li,  Zhizheng Wu
摘要:大型语言模型(LLM)最近表现出了非凡的能力,不仅可以处理文本,还可以处理语音和音频等多模态输入。然而,大多数现有的模型主要集中在使用文本指令分析输入信号,忽略了语音指令和音频混合并作为模型输入的场景。为了解决这些挑战,我们引入Solla,这是一种新颖的框架,旨在理解基于语音的问题并同时听到声学背景。Solla集成了一个音频标记模块来有效地识别和表示音频事件,以及一个ASR辅助预测方法来提高对口语内容的理解。为了严格评估Solla和其他公开可用的模型,我们提出了一个名为SA-Eval的新基准数据集,它包括三个任务:音频事件分类,音频字幕和音频问答。SA-Eval具有各种说话风格的多样化语音教学,包括两个难度级别,简单和困难,以捕捉真实世界的声学条件范围。实验结果表明,Solla在简单和困难的测试集上的表现与基线模型相当或优于基线模型,强调了其在联合理解语音和音频方面的有效性。
摘要:Large Language Models (LLMs) have recently shown remarkable ability toprocess not only text but also multimodal inputs such as speech and audio.However, most existing models primarily focus on analyzing input signals usingtext instructions, overlooking scenarios in which speech instructions and audioare mixed and serve as inputs to the model. To address these challenges, weintroduce Solla, a novel framework designed to understand speech-basedquestions and hear the acoustic context concurrently. Solla incorporates anaudio tagging module to effectively identify and represent audio events, aswell as an ASR-assisted prediction method to improve comprehension of spokencontent. To rigorously evaluate Solla and other publicly available models, wepropose a new benchmark dataset called SA-Eval, which includes three tasks:audio event classification, audio captioning, and audio question answering.SA-Eval has diverse speech instruction with various speaking styles,encompassing two difficulty levels, easy and hard, to capture the range ofreal-world acoustic conditions. Experimental results show that Solla performson par with or outperforms baseline models on both the easy and hard test sets,underscoring its effectiveness in jointly understanding speech and audio.

【2】 Gridless Chirp Parameter Retrieval via Constrained Two-Dimensional  Atomic Norm Minimization
标题:通过约束二维原子规范最小化的无网格Chirp参数检索
链接:https://arxiv.org/abs/2503.15164
作者:Dehui Yang,  Feng Xi
摘要:本文研究了从线性线性调频信号的混合信号中估计线性调频信号参数的基本问题。与大多数以前的方法,解决问题的离散化的参数空间,然后估计啁啾参数,我们提出了一个无网格的方法,通过重新制定的逆问题作为一个约束的二维原子范数最小化结构化测量。这种重构使得能够直接估计连续值参数而无需离散化,从而解决了基础失配的问题。采用近似半定规划(SDP)求解所提出的凸规划。此外,构造了一个对偶多项式来证明原子分解的最优性。数值模拟表明,精确的恢复啁啾参数是可以实现的,使用所提出的原子范数最小化。
摘要:This paper is concerned with the fundamental problem of estimating chirpparameters from a mixture of linear chirp signals. Unlike most previousmethods, which solve the problem by discretizing the parameter space and thenestimating the chirp parameters, we propose a gridless approach byreformulating the inverse problem as a constrained two-dimensional atomic normminimization from structured measurements. This reformulation enables thedirect estimation of continuous-valued parameters without discretization,thereby resolving the issue of basis mismatch. An approximate semidefiniteprogramming (SDP) is employed to solve the proposed convex program.Additionally, a dual polynomial is constructed to certify the optimality of theatomic decomposition. Numerical simulations demonstrate that exact recovery ofchirp parameters is achievable using the proposed atomic norm minimization.

【3】 Analysis and Extension of Noisy-target Training for Unsupervised Target  Signal Enhancement
标题:无监督目标信号增强的噪音目标训练分析与扩展
链接:https://arxiv.org/abs/2503.14854
作者:Takuya Fujimura,  Tomoki Toda
摘要:基于深度神经网络的目标信号增强(TSE)通常使用干净的目标信号以监督的方式进行训练。然而,收集干净的目标信号是昂贵的,并且这样的信号并不总是可用的。因此,期望开发一种不依赖于干净目标信号的无监督方法。在无监督TSE方法的各种研究中,噪声目标训练(NyTT)已被确立为基本方法。NyTT在典型的监督训练中简单地将干净的目标信号替换为有噪声的信号,并且已经通过实验证明可以实现TSE。尽管它的有效性和简单性,其机制和详细的行为仍然不清楚。在本文中,推进NyTT,从而,无监督的方法作为一个整体,我们从不同的角度分析NyTT。我们通过实验证明了NyTT的机制,理想的条件,以及在少量干净的目标信号可用的情况下利用噪声信号的有效性。此外,我们提出了一个改进版本的NyTT的基础上,它的属性,并探讨其功能的去混响和declipping任务,超出了去噪任务。
摘要:Deep neural network-based target signal enhancement (TSE) is usually trainedin a supervised manner using clean target signals. However, collecting cleantarget signals is costly and such signals are not always available. Thus, it isdesirable to develop an unsupervised method that does not rely on clean targetsignals. Among various studies on unsupervised TSE methods, Noisy-targetTraining (NyTT) has been established as a fundamental method. NyTT simplyreplaces clean target signals with noisy ones in the typical supervisedtraining, and it has been experimentally shown to achieve TSE. Despite itseffectiveness and simplicity, its mechanism and detailed behavior are stillunclear. In this paper, to advance NyTT and, thus, unsupervised methods as awhole, we analyze NyTT from various perspectives. We experimentally demonstratethe mechanism of NyTT, the desirable conditions, and the effectiveness ofutilizing noisy signals in situations where a small number of clean targetsignals are available. Furthermore, we propose an improved version of NyTTbased on its properties and explore its capabilities in the dereverberation anddeclipping tasks, beyond the denoising task.

【4】 InsectSet459: an open dataset of insect sounds for bioacoustic machine  learning
标题:InsectSet 459:用于生物声学机器学习的昆虫声音开放数据集
链接:https://arxiv.org/abs/2503.15074
作者:Marius Faiß,  Burooj Ghani,  Dan Stowell
摘要:自动识别昆虫的声音可以帮助我们了解世界各地不断变化的生物多样性趋势,但即使是深度学习也很难识别昆虫的声音。我们提出了一个新的数据集,包括26399音频文件,从459种直翅目和蝉。这是第一个大规模的昆虫声音数据集,可以很容易地应用于开发新的深度学习方法。它的录音是由各种录音机使用不同的采样率,以捕捉昆虫产生的极其广泛的频率范围。我们使用两种最先进的深度学习分类器进行性能基准测试,表现出良好的性能,但在声学昆虫分类方面也有很大的改进空间。该数据集可以作为实现昆虫监测工作流程的现实测试案例,并作为开发可以处理高度可变频率和/或采样率的音频表示方法的挑战性基础。
摘要:Automatic recognition of insect sound could help us understand changingbiodiversity trends around the world -- but insect sounds are challenging torecognize even for deep learning. We present a new dataset comprised of 26399audio files, from 459 species of Orthoptera and Cicadidae. It is the firstlarge-scale dataset of insect sound that is easily applicable for developingnovel deep-learning methods. Its recordings were made with a variety of audiorecorders using varying sample rates to capture the extremely broad range offrequencies that insects produce. We benchmark performance with twostate-of-the-art deep learning classifiers, demonstrating good performance butalso significant room for improvement in acoustic insect classification. Thisdataset can serve as a realistic test case for implementing insect monitoringworkflows, and as a challenging basis for the development of audiorepresentation methods that can handle highly variable frequencies and/orsample rates.

【5】 Shushing! Let's Imagine an Authentic Speech from the Silent Video
标题:嘘!让我们想象一下无声视频中的真实演讲
链接:https://arxiv.org/abs/2503.14928
作者:Jiaxin Ye,  Hongming Shan
备注:Project Page: this https URL
摘要:视觉引导的语音生成旨在从面部外观或嘴唇运动产生真实的语音,而不依赖于听觉信号,为电影制作中的配音和帮助患有失声症的个人等应用提供了巨大的潜力。尽管最近的进展,现有的方法努力实现统一的跨模态对齐语义,音色,和情感韵律的视觉线索,促使我们提出一致的视频到语音(CV2S)作为一个扩展的任务,以提高跨模态的一致性。为了应对新出现的挑战,我们引入了ImaginTalk,这是一种新颖的跨模态扩散框架,仅使用视觉输入生成忠实的语音,在离散空间内操作。具体来说,我们提出了一个离散唇对齐器,预测离散语音令牌从唇视频捕捉语义信息,而错误检测器识别未对齐的令牌,随后通过掩蔽语言建模与BERT细化。为了进一步提高所生成的语音的表现力,我们开发了一个风格扩散Transformer配备了一个脸式适配器,自适应定制的身份和韵律动态跨通道和时间维度,同时确保同步与唇感知语义功能。大量的实验表明,与最先进的基线相比,ImaginTalk可以生成具有更准确的语义细节和更强的音色和情感表现力的高保真语音。演示显示在我们的项目页面:https://imagintalk.github.io。
摘要:Vision-guided speech generation aims to produce authentic speech from facialappearance or lip motions without relying on auditory signals, offeringsignificant potential for applications such as dubbing in filmmaking andassisting individuals with aphonia. Despite recent progress, existing methodsstruggle to achieve unified cross-modal alignment across semantics, timbre, andemotional prosody from visual cues, prompting us to propose ConsistentVideo-to-Speech (CV2S) as an extended task to enhance cross-modal consistency.To tackle emerging challenges, we introduce ImaginTalk, a novel cross-modaldiffusion framework that generates faithful speech using only visual input,operating within a discrete space. Specifically, we propose a discrete lipaligner that predicts discrete speech tokens from lip videos to capturesemantic information, while an error detector identifies misaligned tokens,which are subsequently refined through masked language modeling with BERT. Tofurther enhance the expressiveness of the generated speech, we develop a stylediffusion transformer equipped with a face-style adapter that adaptivelycustomizes identity and prosody dynamics across both the channel and temporaldimensions while ensuring synchronization with lip-aware semantic features.Extensive experiments demonstrate that ImaginTalk can generate high-fidelityspeech with more accurate semantic details and greater expressiveness in timbreand emotion compared to state-of-the-art baselines. Demos are shown at ourproject page: https://imagintalk.github.io.

【6】 PANDORA: Diffusion Policy Learning for Dexterous Robotic Piano Playing
标题:PANDORA:灵巧机器人钢琴演奏的扩散政策学习
链接:https://arxiv.org/abs/2503.14545
作者:Yanjia Huang,  Renjie Li,  Zhengzhong Tu
摘要:我们提出了PANDORA,一种新的基于扩散的政策学习框架,专门为灵巧的机器人钢琴性能。我们的方法采用了一个有条件的U-Net架构,增强了基于FiLM的全局调节,迭代去噪噪声动作序列平滑,高维轨迹。为了实现精确的关键执行加上富有表现力的音乐表现,我们设计了一个复合奖励函数,该函数集成了特定于任务的准确性,音频保真度和来自大型语言模型(LLM)Oracle的高级语义反馈。LLM oracle评估音乐表现力和风格上的细微差别,从而实现动态的、针对具体手型的奖励调整。通过残余逆运动学细化策略进一步增强,PANDORA在ROBOPIANIST环境中实现了最先进的性能,在精度和表现力方面都显着优于基线。消融研究验证了基于扩散的去噪和LLM驱动的语义反馈在增强机器人音乐才能方面的关键贡献。视频网址:https://taco-group.github.io/PANDORA
摘要:We present PANDORA, a novel diffusion-based policy learning frameworkdesigned specifically for dexterous robotic piano performance. Our approachemploys a conditional U-Net architecture enhanced with FiLM-based globalconditioning, which iteratively denoises noisy action sequences into smooth,high-dimensional trajectories. To achieve precise key execution coupled withexpressive musical performance, we design a composite reward function thatintegrates task-specific accuracy, audio fidelity, and high-level semanticfeedback from a large language model (LLM) oracle. The LLM oracle assessesmusical expressiveness and stylistic nuances, enabling dynamic, hand-specificreward adjustments. Further augmented by a residual inverse-kinematicsrefinement policy, PANDORA achieves state-of-the-art performance in theROBOPIANIST environment, significantly outperforming baselines in bothprecision and expressiveness. Ablation studies validate the criticalcontributions of diffusion-based denoising and LLM-driven semantic feedback inenhancing robotic musicianship. Videos available at:https://taco-group.github.io/PANDORA

机器翻译由腾讯交互翻译提供,仅供参考