今日论文合集:cs.SD语音6篇,eess.AS音频处理6篇。本文经arXiv每日学术速递授权转载
【1】RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization
标题:RealMAN:用于动态语音增强和定位的实时记录和注释麦克风阵列数据集作者:Bing Yang,Changsheng Quan,Yabo Wang,Pengyu Wang,Yujie Yang,Ying Fang,Nian Shao,Hui Bu,Xin Xu,Xiaofei Li摘要:基于深度学习的多通道语音增强和源定位系统的训练在很大程度上依赖于房间脉冲响应和多通道扩散噪声的模拟,因为缺乏大规模的真实记录数据集。然而,模拟和真实世界的数据之间的声学失配可能会降低模型的性能时,应用在现实世界的场景。为了弥合这一模拟与真实的差距,本文提出了一个新的相对大规模的实时记录和注释麦克风阵列语音和噪声(RealMAN)数据集。所提出的数据集在两个方面具有价值:1)在真实场景中对语音增强和定位算法进行基准测试; 2)提供大量真实世界的训练数据,以潜在地提高真实世界应用的性能。具体而言,使用具有高保真麦克风的32通道阵列进行记录。扬声器用于播放源语音信号。在32个不同的场景中记录了总共83小时的语音信号(静态扬声器48小时和移动扬声器35小时),并且在31个不同的场景中记录了144小时的背景噪声。语音和噪声记录场景覆盖各种常见的室内、室外、半室外和交通环境,这使得通用语音增强和源定位网络的训练成为可能。为了获得特定于任务的注释,扬声器的方位角通过自动检测扬声器用全向鱼眼相机进行注释。直接路径信号被设置为用于语音增强的目标干净语音,该目标干净语音是通过利用估计的直接路径传播滤波器对源语音信号进行滤波而获得的。摘要:The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded datasets. However, the acoustic mismatch between simulated and real-world data could degrade the model performance when applying in real-world scenarios. To bridge this simulation-to-real gap, this paper presents a new relatively large-scale Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset. The proposed dataset is valuable in two aspects: 1) benchmarking speech enhancement and localization algorithms in real scenarios; 2) offering a substantial amount of real-world training data for potentially improving the performance of real-world applications. Specifically, a 32-channel array with high-fidelity microphones is used for recording. A loudspeaker is used for playing source speech signals. A total of 83-hour speech signals (48 hours for static speaker and 35 hours for moving speaker) are recorded in 32 different scenes, and 144 hours of background noise are recorded in 31 different scenes. Both speech and noise recording scenes cover various common indoor, outdoor, semi-outdoor and transportation environments, which enables the training of general-purpose speech enhancement and source localization networks. To obtain the task-specific annotations, the azimuth angle of the loudspeaker is annotated with an omni-direction fisheye camera by automatically detecting the loudspeaker. The direct-path signal is set as the target clean speech for speech enhancement, which is obtained by filtering the source speech signal with an estimated direct-path propagation filter.
【2】 BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5标题:BESTow:高效且可流化的语音语言模型,在GPT和T5中具有两种最佳效果作者:Zhehuai Chen,He Huang,Oleksii Hrinchuk,Krishna C. Puvvada,Nithin Rao Koluguri,Piotr Żelasko,Jagadeesh Balam,Boris Ginsburg摘要:将语音理解能力扩展到预训练的大型语言模型中已成为一个重要的研究方向(SpeechLLM)。以前的架构可以分类为:i)GPT风格,将语音提示作为LLM输入序列(如仅解码器模型)前置到文本提示; ii)T5风格,将语音交叉注意力引入预训练LLM的每一层。我们提出BESTOW架构,将BEST功能从两个世界到一个单一的模型,是高效的,具有强大的多任务能力。此外,对于任何一种风格都没有明确的流解决方案,特别是考虑到该解决方案应该推广到语音多任务。我们重新制定流SpeechLLM作为一个读写策略的问题,并结合离线和流的研究与BESTOW架构。因此,我们展示了第一个开源的SpeechLLM解决方案,该解决方案可以同时实现大规模(超越ASR)的流和多任务。这种可流式传输的解决方案在各种语音任务(ASR,AST,SQA,看不见的DynamicSuperb)上实现了非常强大的性能。它是端到端可优化的,具有较低的训练/推理成本,并展示了LLM知识到语音的可移植性。摘要:Incorporating speech understanding capabilities into pretrained large-language models has become a vital research direction (SpeechLLM). The previous architectures can be categorized as: i) GPT-style, prepend speech prompts to the text prompts as a sequence of LLM inputs like a decoder-only model; ii) T5-style, introduce speech cross-attention to each layer of the pretrained LLMs. We propose BESTOW architecture to bring the BESt features from TwO Worlds into a single model that is highly efficient and has strong multitask capabilities. Moreover, there is no clear streaming solution for either style, especially considering the solution should generalize to speech multitask. We reformulate streamable SpeechLLM as a read-write policy problem and unifies the offline and streaming research with BESTOW architecture. Hence we demonstrate the first open-source SpeechLLM solution that enables Streaming and Multitask at scale (beyond ASR) at the same time. This streamable solution achieves very strong performance on a wide range of speech tasks (ASR, AST, SQA, unseen DynamicSuperb). It is end-to-end optimizable, with lower training/inference cost, and demonstrates LLM knowledge transferability to speech.【3】 SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR标题:SAL:端到端ASB LoRA专家的扬声器自适应混合作者:Qiuming Zhao,Guangzhi Sun,Chao Zhang,Mingxing Xu,Thomas Fang Zheng备注:5 pages, accepted by Interspeech 2024. arXiv admin note: substantial text overlap with arXiv:2309.09136摘要:混合专家模型(Mixture-of-Experts,MoE)在许多任务中取得了很好的效果。然而,传统的MoE模型通常非常大,使得它们难以部署在资源受限的边缘设备上。在本文中,我们提出了一种新的说话人自适应混合LoRA专家(SAML)的方法,它使用低秩自适应(LoRA)模块作为专家,以减少MoE的可训练参数的数量。具体而言,SAML应用于量化和个性化的端到端自动语音识别模型,它结合了测试时说话人自适应,以提高在特定于说话人的场景中严重压缩模型的性能。在LibriSpeech和TED-LIUM 3语料库上进行了实验。值得注意的是,与原始的全精度模型相比,在模型大小减少7倍的情况下,量化Whisper模型和基于Conformer的基于注意力的编码器-解码器ASR模型分别实现了29.1%和31.1%的相对字错误率降低。摘要:Mixture-of-experts (MoE) models have achieved excellent results in many tasks. However, conventional MoE models are often very large, making them challenging to deploy on resource-constrained edge devices. In this paper, we propose a novel speaker adaptive mixture of LoRA experts (SAML) approach, which uses low-rank adaptation (LoRA) modules as experts to reduce the number of trainable parameters in MoE. Specifically, SAML is applied to the quantised and personalised end-to-end automatic speech recognition models, which combines test-time speaker adaptation to improve the performance of heavily compressed models in speaker-specific scenarios. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size, 29.1% and 31.1% relative word error rate reductions were achieved on the quantised Whisper model and Conformer-based attention-based encoder-decoder ASR model respectively, comparing to the original full precision models.【4】 Less is More: Accurate Speech Recognition & Translation without Web-Scale Data标题:少即是多:无需网络规模数据即可准确语音识别和翻译作者:Krishna C. Puvvada,Piotr Żelasko,He Huang,Oleksii Hrinchuk,Nithin Rao Koluguri,Kunal Dhawan,Somshubra Majumdar,Elena Rastorgueva,Zhehuai Chen,Vitaly Lavrukhin,Jagadeesh Balam,Boris Ginsburg备注:Accepted at Interspeech-2024摘要:语音识别和翻译的最新进展依赖于数十万小时的互联网语音数据。我们认为,国家的最先进的准确性可以达到不依赖于网络规模的数据。Canary -多语言ASR和语音翻译模型,在英语,法语,西班牙语和德语上优于当前最先进的模型- Whisper,OWSM和无障碍M4 T,同时在比这些模型少一个数量级的数据上训练。三个关键因素使这种数据高效模型成为可能:(1)基于FastConformer的注意力编码器-解码器架构(2)对机器翻译生成的合成数据进行训练,以及(3)高级训练技术:数据平衡,动态数据混合,动态桶和噪声鲁棒微调。模型、权重和训练代码将是开源的。摘要:Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data. We argue that state-of-the art accuracy can be reached without relying on web-scale data. Canary - multilingual ASR and speech translation model, outperforms current state-of-the-art models - Whisper, OWSM, and Seamless-M4T on English, French, Spanish, and German languages, while being trained on an order of magnitude less data than these models. Three key factors enables such data-efficient model: (1) a FastConformer-based attention encoder-decoder architecture (2) training on synthetic data generated with machine translation and (3) advanced training techniques: data-balancing, dynamic data blending, dynamic bucketing and noise-robust fine-tuning. The model, weights, and training code will be open-sourced.
【5】 Network Bending of Diffusion Models for Audio-Visual Generation作者:Luke Dzwonczyk,Carmine Emanuele Cella,David Ban备注:8 pages, 5 figures, to be published in the proceedings of the 27th International Conference on Digital Audio Effects (DAFx24), for additional image and video examples see this https URL摘要:在本文中,我们介绍了创建工具的第一步,该工具使艺术家能够使用预先训练的生成机器学习模型创建音乐可视化。首先,我们研究了网络弯曲的应用,即在生成网络的层内应用变换的过程,通过利用一系列逐点、逐张量和形态学运算符,将其应用于图像生成扩散模型。我们确定了一些视觉效果,从不同的运营商,包括一些不容易重建与标准的图像编辑工具。我们发现,这个过程允许连续的,细粒度的图像生成控制,这可能有助于创造性的应用程序。接下来,我们通过将音频特征作为参数传递给网络弯曲算子,使用稳定扩散生成音乐反应视频。最后,我们评论某些变换,从根本上改变图像和学习更多关于稳定扩散的潜在空间的可能性,基于这些变换。摘要:In this paper we present the first steps towards the creation of a tool which enables artists to create music visualizations using pre-trained, generative, machine learning models. First, we investigate the application of network bending, the process of applying transforms within the layers of a generative network, to image generation diffusion models by utilizing a range of point-wise, tensor-wise, and morphological operators. We identify a number of visual effects that result from various operators, including some that are not easily recreated with standard image editing tools. We find that this process allows for continuous, fine-grain control of image generation which can be helpful for creative applications. Next, we generate music-reactive videos using Stable Diffusion by passing audio features as parameters to network bending operators. Finally, we comment on certain transforms which radically shift the image and the possibilities of learning more about the latent space of Stable Diffusion based on these transforms.【6】 ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data标题:ManiWave:从野外视听数据学习机器人操纵作者:Zeyi Liu,Cheng Chi,Eric Cousineau,Naveen Kuppuswamy,Benjamin Burchfiel,Shuran Song摘要:音频信号通过接触为机器人交互和对象属性提供了丰富的信息。这些信息可以令人惊讶地简化接触丰富的机器人操作技能的学习,特别是当视觉信息本身模糊或不完整时。然而,在机器人操作中使用的音频数据已被限制到通过将麦克风连接到机器人或对象收集的遥控演示,这大大限制了其在机器人学习管道中的使用。在这项工作中,我们引入ManiWAV:一个“耳朵在手”的数据收集设备,收集在野外的人类演示同步音频和视觉反馈,和相应的政策接口,直接从演示学习机器人操作政策。我们展示了我们的系统的能力,通过四个接触丰富的操作任务,需要被动地感知接触事件和模式,或主动地感知物体表面材料和状态。此外,我们表明,我们的系统可以推广到看不见的野生环境,从不同的野生人类示范学习。项目网址:https://mani-wav.github.io/摘要:Audio signals provide rich information for the robot interaction and object properties through contact. These information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio data in robot manipulation has been constrained to teleoperated demonstrations collected by either attaching a microphone to the robot or object, which significantly limits its usage in robot learning pipelines. In this work, we introduce ManiWAV: an 'ear-in-hand' data collection device to collect in-the-wild human demonstrations with synchronous audio and visual feedback, and a corresponding policy interface to learn robot manipulation policy directly from the demonstrations. We demonstrate the capabilities of our system through four contact-rich manipulation tasks that require either passively sensing the contact events and modes, or actively sensing the object surface materials and states. In addition, we show that our system can generalize to unseen in-the-wild environments, by learning from diverse in-the-wild human demonstrations. Project website: https://mani-wav.github.io/【1】 RealMAN: A Real-Recorded and Annotated Microphone Array Dataset for Dynamic Speech Enhancement and Localization标题:RealMAN:用于动态语音增强和定位的实时记录和注释麦克风阵列数据集作者:Bing Yang,Changsheng Quan,Yabo Wang,Pengyu Wang,Yujie Yang,Ying Fang,Nian Shao,Hui Bu,Xin Xu,Xiaofei Li摘要:基于深度学习的多通道语音增强和源定位系统的训练在很大程度上依赖于房间脉冲响应和多通道扩散噪声的模拟,因为缺乏大规模的真实记录数据集。然而,模拟和真实世界的数据之间的声学失配可能会降低模型的性能时,应用在现实世界的场景。为了弥合这一模拟与真实的差距,本文提出了一个新的相对大规模的实时记录和注释麦克风阵列语音和噪声(RealMAN)数据集。所提出的数据集在两个方面具有价值:1)在真实场景中对语音增强和定位算法进行基准测试; 2)提供大量真实世界的训练数据,以潜在地提高真实世界应用的性能。具体而言,使用具有高保真麦克风的32通道阵列进行记录。扬声器用于播放源语音信号。在32个不同的场景中记录了总共83小时的语音信号(静态扬声器48小时和移动扬声器35小时),并且在31个不同的场景中记录了144小时的背景噪声。语音和噪声记录场景覆盖各种常见的室内、室外、半室外和交通环境,这使得通用语音增强和源定位网络的训练成为可能。为了获得特定于任务的注释,扬声器的方位角通过自动检测扬声器用全向鱼眼相机进行注释。直接路径信号被设置为用于语音增强的目标干净语音,该目标干净语音是通过利用估计的直接路径传播滤波器对源语音信号进行滤波而获得的。摘要:The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded datasets. However, the acoustic mismatch between simulated and real-world data could degrade the model performance when applying in real-world scenarios. To bridge this simulation-to-real gap, this paper presents a new relatively large-scale Real-recorded and annotated Microphone Array speech&Noise (RealMAN) dataset. The proposed dataset is valuable in two aspects: 1) benchmarking speech enhancement and localization algorithms in real scenarios; 2) offering a substantial amount of real-world training data for potentially improving the performance of real-world applications. Specifically, a 32-channel array with high-fidelity microphones is used for recording. A loudspeaker is used for playing source speech signals. A total of 83-hour speech signals (48 hours for static speaker and 35 hours for moving speaker) are recorded in 32 different scenes, and 144 hours of background noise are recorded in 31 different scenes. Both speech and noise recording scenes cover various common indoor, outdoor, semi-outdoor and transportation environments, which enables the training of general-purpose speech enhancement and source localization networks. To obtain the task-specific annotations, the azimuth angle of the loudspeaker is annotated with an omni-direction fisheye camera by automatically detecting the loudspeaker. The direct-path signal is set as the target clean speech for speech enhancement, which is obtained by filtering the source speech signal with an estimated direct-path propagation filter.
【2】 BESTOW: Efficient and Streamable Speech Language Model with the Best of Two Worlds in GPT and T5标题:BESTow:高效且可流化的语音语言模型,在GPT和T5中具有两种最佳效果作者:Zhehuai Chen,He Huang,Oleksii Hrinchuk,Krishna C. Puvvada,Nithin Rao Koluguri,Piotr Żelasko,Jagadeesh Balam,Boris Ginsburg摘要:将语音理解能力扩展到预训练的大型语言模型中已成为一个重要的研究方向(SpeechLLM)。以前的架构可以分类为:i)GPT风格,将语音提示作为LLM输入序列(如仅解码器模型)前置到文本提示; ii)T5风格,将语音交叉注意力引入预训练LLM的每一层。我们提出BESTOW架构,将BEST功能从两个世界到一个单一的模型,是高效的,具有强大的多任务能力。此外,对于任何一种风格都没有明确的流解决方案,特别是考虑到该解决方案应该推广到语音多任务。我们重新制定流SpeechLLM作为一个读写策略的问题,并结合离线和流的研究与BESTOW架构。因此,我们展示了第一个开源的SpeechLLM解决方案,该解决方案可以同时实现大规模(超越ASR)的流和多任务。这种可流式传输的解决方案在各种语音任务(ASR,AST,SQA,看不见的DynamicSuperb)上实现了非常强大的性能。它是端到端可优化的,具有较低的训练/推理成本,并展示了LLM知识到语音的可移植性。摘要:Incorporating speech understanding capabilities into pretrained large-language models has become a vital research direction (SpeechLLM). The previous architectures can be categorized as: i) GPT-style, prepend speech prompts to the text prompts as a sequence of LLM inputs like a decoder-only model; ii) T5-style, introduce speech cross-attention to each layer of the pretrained LLMs. We propose BESTOW architecture to bring the BESt features from TwO Worlds into a single model that is highly efficient and has strong multitask capabilities. Moreover, there is no clear streaming solution for either style, especially considering the solution should generalize to speech multitask. We reformulate streamable SpeechLLM as a read-write policy problem and unifies the offline and streaming research with BESTOW architecture. Hence we demonstrate the first open-source SpeechLLM solution that enables Streaming and Multitask at scale (beyond ASR) at the same time. This streamable solution achieves very strong performance on a wide range of speech tasks (ASR, AST, SQA, unseen DynamicSuperb). It is end-to-end optimizable, with lower training/inference cost, and demonstrates LLM knowledge transferability to speech.
【3】 SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR标题:SAL:端到端ASB LoRA专家的扬声器自适应混合作者:Qiuming Zhao,Guangzhi Sun,Chao Zhang,Mingxing Xu,Thomas Fang Zheng备注:5 pages, accepted by Interspeech 2024. arXiv admin note: substantial text overlap with arXiv:2309.09136摘要:混合专家模型(Mixture-of-Experts,MoE)在许多任务中取得了很好的效果。然而,传统的MoE模型通常非常大,使得它们难以部署在资源受限的边缘设备上。在本文中,我们提出了一种新的说话人自适应混合LoRA专家(SAML)的方法,它使用低秩自适应(LoRA)模块作为专家,以减少MoE的可训练参数的数量。具体而言,SAML应用于量化和个性化的端到端自动语音识别模型,它结合了测试时说话人自适应,以提高在特定于说话人的场景中严重压缩模型的性能。在LibriSpeech和TED-LIUM 3语料库上进行了实验。值得注意的是,与原始的全精度模型相比,在模型大小减少7倍的情况下,量化Whisper模型和基于Conformer的基于注意力的编码器-解码器ASR模型分别实现了29.1%和31.1%的相对字错误率降低。摘要:Mixture-of-experts (MoE) models have achieved excellent results in many tasks. However, conventional MoE models are often very large, making them challenging to deploy on resource-constrained edge devices. In this paper, we propose a novel speaker adaptive mixture of LoRA experts (SAML) approach, which uses low-rank adaptation (LoRA) modules as experts to reduce the number of trainable parameters in MoE. Specifically, SAML is applied to the quantised and personalised end-to-end automatic speech recognition models, which combines test-time speaker adaptation to improve the performance of heavily compressed models in speaker-specific scenarios. Experiments have been performed on the LibriSpeech and the TED-LIUM 3 corpora. Remarkably, with a 7x reduction in model size, 29.1% and 31.1% relative word error rate reductions were achieved on the quantised Whisper model and Conformer-based attention-based encoder-decoder ASR model respectively, comparing to the original full precision models.
【4】 Less is More: Accurate Speech Recognition & Translation without Web-Scale Data标题:少即是多:无需网络规模数据即可准确语音识别和翻译作者:Krishna C. Puvvada,Piotr Żelasko,He Huang,Oleksii Hrinchuk,Nithin Rao Koluguri,Kunal Dhawan,Somshubra Majumdar,Elena Rastorgueva,Zhehuai Chen,Vitaly Lavrukhin,Jagadeesh Balam,Boris Ginsburg备注:Accepted at Interspeech-2024摘要:语音识别和翻译的最新进展依赖于数十万小时的互联网语音数据。我们认为,国家的最先进的准确性可以达到不依赖于网络规模的数据。Canary -多语言ASR和语音翻译模型,在英语,法语,西班牙语和德语上优于当前最先进的模型- Whisper,OWSM和无障碍M4 T,同时在比这些模型少一个数量级的数据上训练。三个关键因素使这种数据高效模型成为可能:(1)基于FastConformer的注意力编码器-解码器架构(2)对机器翻译生成的合成数据进行训练,以及(3)高级训练技术:数据平衡,动态数据混合,动态桶和噪声鲁棒微调。模型、权重和训练代码将是开源的。摘要:Recent advances in speech recognition and translation rely on hundreds of thousands of hours of Internet speech data. We argue that state-of-the art accuracy can be reached without relying on web-scale data. Canary - multilingual ASR and speech translation model, outperforms current state-of-the-art models - Whisper, OWSM, and Seamless-M4T on English, French, Spanish, and German languages, while being trained on an order of magnitude less data than these models. Three key factors enables such data-efficient model: (1) a FastConformer-based attention encoder-decoder architecture (2) training on synthetic data generated with machine translation and (3) advanced training techniques: data-balancing, dynamic data blending, dynamic bucketing and noise-robust fine-tuning. The model, weights, and training code will be open-sourced.【5】 Network Bending of Diffusion Models for Audio-Visual Generation作者:Luke Dzwonczyk,Carmine Emanuele Cella,David Ban备注:8 pages, 5 figures, to be published in the proceedings of the 27th International Conference on Digital Audio Effects (DAFx24), for additional image and video examples see this https URL摘要:在本文中,我们介绍了创建工具的第一步,该工具使艺术家能够使用预先训练的生成机器学习模型创建音乐可视化。首先,我们研究了网络弯曲的应用,即在生成网络的层内应用变换的过程,通过利用一系列逐点、逐张量和形态学运算符,将其应用于图像生成扩散模型。我们确定了一些视觉效果,从不同的运营商,包括一些不容易重建与标准的图像编辑工具。我们发现,这个过程允许连续的,细粒度的图像生成控制,这可能有助于创造性的应用程序。接下来,我们通过将音频特征作为参数传递给网络弯曲算子,使用稳定扩散生成音乐反应视频。最后,我们评论某些变换,从根本上改变图像和学习更多关于稳定扩散的潜在空间的可能性,基于这些变换。摘要:In this paper we present the first steps towards the creation of a tool which enables artists to create music visualizations using pre-trained, generative, machine learning models. First, we investigate the application of network bending, the process of applying transforms within the layers of a generative network, to image generation diffusion models by utilizing a range of point-wise, tensor-wise, and morphological operators. We identify a number of visual effects that result from various operators, including some that are not easily recreated with standard image editing tools. We find that this process allows for continuous, fine-grain control of image generation which can be helpful for creative applications. Next, we generate music-reactive videos using Stable Diffusion by passing audio features as parameters to network bending operators. Finally, we comment on certain transforms which radically shift the image and the possibilities of learning more about the latent space of Stable Diffusion based on these transforms.【6】 ManiWAV: Learning Robot Manipulation from In-the-Wild Audio-Visual Data标题:ManiWave:从野外视听数据学习机器人操纵作者:Zeyi Liu,Cheng Chi,Eric Cousineau,Naveen Kuppuswamy,Benjamin Burchfiel,Shuran Song摘要:音频信号通过接触为机器人交互和对象属性提供了丰富的信息。这些信息可以令人惊讶地简化接触丰富的机器人操作技能的学习,特别是当视觉信息本身模糊或不完整时。然而,在机器人操作中使用的音频数据已被限制到通过将麦克风连接到机器人或对象收集的遥控演示,这大大限制了其在机器人学习管道中的使用。在这项工作中,我们引入ManiWAV:一个“耳朵在手”的数据收集设备,收集在野外的人类演示同步音频和视觉反馈,和相应的政策接口,直接从演示学习机器人操作政策。我们展示了我们的系统的能力,通过四个接触丰富的操作任务,需要被动地感知接触事件和模式,或主动地感知物体表面材料和状态。此外,我们表明,我们的系统可以推广到看不见的野生环境,从不同的野生人类示范学习。项目网站:https://mani-wav.github.io/摘要:Audio signals provide rich information for the robot interaction and object properties through contact. These information can surprisingly ease the learning of contact-rich robot manipulation skills, especially when the visual information alone is ambiguous or incomplete. However, the usage of audio data in robot manipulation has been constrained to teleoperated demonstrations collected by either attaching a microphone to the robot or object, which significantly limits its usage in robot learning pipelines. In this work, we introduce ManiWAV: an 'ear-in-hand' data collection device to collect in-the-wild human demonstrations with synchronous audio and visual feedback, and a corresponding policy interface to learn robot manipulation policy directly from the demonstrations. We demonstrate the capabilities of our system through four contact-rich manipulation tasks that require either passively sensing the contact events and modes, or actively sensing the object surface materials and states. In addition, we show that our system can generalize to unseen in-the-wild environments, by learning from diverse in-the-wild human demonstrations. Project website: https://mani-wav.github.io/