今天跟大家分享一篇语音相关的论文合集:cs.SD语音8篇,eess.AS音频处理11篇。
cs.SD语音
【1】 Finding Fallen Objects Via Asynchronous Audio-Visual Integration
标题:基于异步音视频融合的坠落物体定位
链接:https://arxiv.org/abs/2207.03483
作者:Chuang Gan,Yi Gu,Siyuan Zhou,Jeremy Schwartz,Seth Alter,James Traer,Dan Gutfreund,Joshua B. Tenenbaum,Josh McDermott,Antonio Torralba备注:CVPR 2022. Project page: this http URL摘要:物体的外观和声音提供了其物理特性的互补反射。在许多情况下,视觉和听觉的提示是异步到达的,但必须进行整合,就像我们听到物体掉在地板上,然后必须找到它一样。在本文中,我们介绍了一种用于研究三维虚拟环境中的多模式对象定位的环境。一个物体掉在房间的某个地方。配备摄像机和麦克风的嵌入式机器人代理必须通过将音频和视觉信号与基本物理知识相结合来确定掉落了什么物体以及掉落在哪里。为了研究这个问题,我们生成了一个大规模数据集——坠落物体数据集,其中包括64个房间中30个物理物体类别的8000个实例。该数据集使用ThreeDWorld平台,该平台可以在真实照片环境中模拟基于物理的冲击声音和对象之间的复杂物理交互。作为解决这一挑战的第一步,我们基于模仿学习、强化学习和模块化规划开发了一组具体的agent基线,并对这一新任务的挑战进行了深入分析。摘要:The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in which to study multi-modal object localization in 3D virtual environments. An object is dropped somewhere in a room. An embodied robot agent, equipped with a camera and microphone, must determine what object has been dropped -- and where -- by combining audio and visual signals with knowledge of the underlying physics. To study this problem, we have generated a large-scale dataset -- the Fallen Objects dataset -- that includes 8000 instances of 30 physical object categories in 64 rooms. The dataset uses the ThreeDWorld platform which can simulate physics-based impact sounds and complex physical interactions between objects in a photorealistic setting. As a first step toward addressing this challenge, we develop a set of embodied agent baselines, based on imitation learning, reinforcement learning, and modular planning, and perform an in-depth analysis of the challenge of this new task.
【2】 Self-Supervised Learning of Music-Dance Representation through Explicit-Implicit Rhythm Synchronization
标题:基于显性-隐性节奏同步的乐舞表征自监督学习
链接:https://arxiv.org/abs/2207.03190
作者:Jiashuo Yu,Junfu Pu,Ying Cheng,Rui Feng,Ying Shan摘要:虽然视听表示已被证明适用于许多下游任务,但舞蹈视频的表示更为具体,总是伴随着具有复杂听觉内容的音乐,仍然具有挑战性,尚未进行研究。考虑到舞者的学员动作和音乐节奏之间的内在一致性,我们引入了一种新的音乐舞蹈表征学习框架MuDaR,以显式和隐式方式实现音乐和舞蹈节奏的同步。具体来说,我们根据视觉外观和受音乐节奏分析启发的动作线索推导出舞蹈节奏。然后,视觉节奏与音乐对应物在时间上对齐,音乐对应物由声音强度的幅度提取。同时,通过对比学习,我们开发了音频和视频流中隐含的节奏的内隐连贯性。该模型通过预测视听对之间的时间一致性来学习联合嵌入。音乐舞蹈表示,以及检测音频和视觉节奏的能力,可以进一步应用于三个下游任务:(a)舞蹈分类,(b)音乐舞蹈检索,和(c)音乐舞蹈重定位。大量实验表明,我们提出的框架在很大程度上优于其他自监督方法。摘要:Although audio-visual representation has been proved to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of dancer and music rhythm, we introduce MuDaR, a novel Music-Dance Representation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin.
【3】 Visual-Assisted Sound Source Depth Estimation in the Wild
标题:野外视觉辅助声源深度估计
链接:https://arxiv.org/abs/2207.03074
备注:13 pages;in submission;摘要:深度估计可以实现多种3D应用,例如机器人、自动驾驶和虚拟现实。尽管在这一领域开展了大量工作,但如何实现准确、低成本、高分辨率和大范围深度估计仍然是个未知数。受闪电爆炸现象(即看到闪电后听到雷声)的启发,本文开发了第一个视听深度估计框架FBDepth。它利用光和声音的飞行时间(ToF)之间的差异来推断声源深度。FBDepth是第一个将视频和音频与语义特征和空间提示结合起来进行距离估计的方法。它首先对齐视频轨迹和音频轨迹之间的对应关系,以粗粒度定位目标对象和目标声音。基于对运动物体轨迹的观察,FBDepth提出了在声音产生之前和之后估计光流交点的方法,以及时定位视频事件。FBDepth为最终深度估计提供视频事件和音频片段的估计时间戳。我们用一部手机收集了3000多个视频片段,其中有20个不同的物体,价格高达5000万美元。与基于RGB的方法相比,FBDepth将绝对相对误差(AbsRel)降低了55%。摘要:Depth estimation enables a wide variety of 3D applications, such as robotics, autonomous driving, and virtual reality. Despite significant work in this area, it remains open how to enable accurate, low-cost, high-resolution, and large-range depth estimation. Inspired by the flash-to-bang phenomenon (\ie hearing the thunder after seeing the lightning), this paper develops FBDepth, the first audio-visual depth estimation framework. It takes the difference between the time-of-flight (ToF) of the light and the sound to infer the sound source depth. FBDepth is the first to incorporate video and audio with both semantic features and spatial hints for range estimation. It first aligns correspondence between the video track and audio track to locate the target object and target sound in a coarse granularity. Based on the observation of moving objects' trajectories, FBDepth proposes to estimate the intersection of optical flow before and after the sound production to locate video events in time. FBDepth feeds the estimated timestamp of the video event and the audio clip for the final depth estimation. We use a mobile phone to collect 3000+ video clips with 20 different objects at up to $50m$. FBDepth decreases the Absolute Relative error (AbsRel) by 55\% compared to RGB-based methods.
【4】 Cross-Scale Vector Quantization for Scalable Neural Speech Coding
标题:用于可伸缩神经语音编码的跨尺度矢量量化
链接:https://arxiv.org/abs/2207.03067
作者:Xue Jiang,Xiulian Peng,Huaying Xue,Yuan Zhang,Yan Lu备注:INTERSPEECH 2022(Accepted)摘要:比特率可扩展性是实时通信中音频编码的理想特性。现有的神经音频编解码器通常在训练期间强制执行特定的比特率,因此需要针对每个目标比特率训练不同的模型,这增加了发送方和接收方的内存占用,并且通常需要转码来支持多个接收器。在本文中,我们介绍了一种跨尺度可扩展矢量量化方案(CSVQ),在该方案中,多尺度特征通过逐步特征融合和细化进行渐进编码。通过这种方式,如果仅接收到比特流的一部分,则重构粗电平信号,并且随着更多比特可用,逐渐提高质量。提出的CSVQ方案可以灵活地应用于具有镜像自动编码器结构的任何神经音频编码网络,以实现比特率可扩展性。主观结果表明,该方案在可扩展性方面优于经典残差矢量量化(RVQ)。此外,提出的3 kbps的CSVQ优于9 kbps的Opus和3kbps的Lyra,并且可以随着比特率的增加提供优雅的质量提升。摘要:Bitrate scalability is a desirable feature for audio coding in real-time communications. Existing neural audio codecs usually enforce a specific bitrate during training, so different models need to be trained for each target bitrate, which increases the memory footprint at the sender and the receiver side and transcoding is often needed to support multiple receivers. In this paper, we introduce a cross-scale scalable vector quantization scheme (CSVQ), in which multi-scale features are encoded progressively with stepwise feature fusion and refinement. In this way, a coarse-level signal is reconstructed if only a portion of the bitstream is received, and progressively improves the quality as more bits are available. The proposed CSVQ scheme can be flexibly applied to any neural audio coding network with a mirrored auto-encoder structure to achieve bitrate scalability. Subjective results show that the proposed scheme outperforms the classical residual VQ (RVQ) with scalability. Moreover, the proposed CSVQ at 3 kbps outperforms Opus at 9 kbps and Lyra at 3kbps and it could provide a graceful quality boost with bitrate increase.
【5】 Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding
标题:Branchform:为语音识别和理解捕获局部和全局上下文的并行MLP-注意体系结构
链接:https://arxiv.org/abs/2207.02971
作者:Yifan Peng,Siddharth Dalmia,Ian Lane,Shinji Watanabe摘要:事实证明,整合器在许多语音处理任务中是有效的。它结合了使用卷积提取局部依赖和使用自注意力提取全局依赖的优点。受此启发,我们提出了一种更灵活、可解释和可定制的编码器替代方案Branchformer,它具有并行分支,用于建模端到端语音处理中的各种范围依存关系。在每个编码器层中,一个分支使用自注意力或其变体来捕获长程依赖关系,而另一个分支使用带卷积选通(cgMLP)的MLP模块来提取局部关系。我们在几个语音识别和口语理解基准上进行了实验。结果表明,我们的模型优于Transformer和cgMLP。它还与Conformer取得的最先进成果相匹配或优于Conformer取得的最先进成果。此外,由于双分支架构,我们展示了各种减少计算的策略,包括在单个训练模型中具有可变推理复杂度的能力。为合并分支而学习的权重表示如何在不同层中利用局部和全局依赖关系,这有利于模型设计。摘要:Conformer has proven to be effective in many speech processing tasks. It combines the benefits of extracting local dependencies using convolutions and global dependencies using self-attention. Inspired by this, we propose a more flexible, interpretable and customizable encoder alternative, Branchformer, with parallel branches for modeling various ranged dependencies in end-to-end speech processing. In each encoder layer, one branch employs self-attention or its variant to capture long-range dependencies, while the other branch utilizes an MLP module with convolutional gating (cgMLP) to extract local relationships. We conduct experiments on several speech recognition and spoken language understanding benchmarks. Results show that our model outperforms both Transformer and cgMLP. It also matches with or outperforms state-of-the-art results achieved by Conformer. Furthermore, we show various strategies to reduce computation thanks to the two-branch architecture, including the ability to have variable inference complexity in a single trained model. The weights learned for merging branches indicate how local and global dependencies are utilized in different layers, which benefits model designing.
【6】 Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation
标题:言语情感:模型表征、多任务学习与知识提炼研究
链接:https://arxiv.org/abs/2207.03334
作者:Vikramjit Mitra,Hsiang-Yun Sherry Chien,Vasudha Kowtha,Joseph Yitan Cheng,Erdrin Azemi备注:5 pages, 3 figures, Interspeech 2022摘要:在过去的几年里,从语音信号中估计维度情绪,例如激活、价和优势,已经得到了广泛的探索。虽然从语音中准确估计激活度和优势度似乎是可能的,但对价态的准确估计仍然具有挑战性。以往的研究表明,词汇信息的使用可以提高配价估计性能。词汇信息可以从预先训练的声学模型中获得,其中学习的表示可以改进语音的价估计。我们研究了使用预训练的模型表示来改进声学语音信号的价估计。我们还探索了表征的融合,以改善所有三个情绪维度的情绪估计:激活、配价和优势。此外,我们还研究了是否可以将预训练模型中的表示提取到使用低级特征训练的模型中,从而得到参数数量较少的模型。我们表明,与标准声学特征基线(mel滤波器组能量)相比,预训练模型嵌入的融合导致价估计的协方差相关系数CCC相对提高79%,而从预训练模型嵌入到低维表示的蒸馏产生了相对提高12%。在两个评估集上观察到了这样的性能增益,这表明我们提出的架构在这些评估集上得到了推广。我们在两个MSP播客评估集上报告了最新的“无文本”纯声学维度情感估计值$CCC$。摘要:Estimating dimensional emotions, such as activation, valence and dominance, from acoustic speech signals has been widely explored over the past few years. While accurate estimation of activation and dominance from speech seem to be possible, the same for valence remains challenging. Previous research has shown that the use of lexical information can improve valence estimation performance. Lexical information can be obtained from pre-trained acoustic models, where the learned representations can improve valence estimation from speech. We investigate the use of pre-trained model representations to improve valence estimation from acoustic speech signal. We also explore fusion of representations to improve emotion estimation across all three emotion dimensions: activation, valence and dominance. Additionally, we investigate if representations from pre-trained models can be distilled into models trained with low-level features, resulting in models with a less number of parameters. We show that fusion of pre-trained model embeddings result in a 79% relative improvement in concordance correlation coefficient CCC on valence estimation compared to standard acoustic feature baseline (mel-filterbank energies), while distillation from pre-trained model embeddings to lower-dimensional representations yielded a relative 12% improvement. Such performance gains were observed over two evaluation sets, indicating that our proposed architecture generalizes across those evaluation sets. We report new state-of-the-art "text-free" acoustic-only dimensional emotion estimation $CCC$ values on two MSP-Podcast evaluation sets.
【7】 NESC: Robust Neural End-2-End Speech Coding with GANs
标题:NESC:基于GANS的稳健神经端双端语音编码
链接:https://arxiv.org/abs/2207.03282
作者:Nicola Pia,Kishan Gupta,Srikanth Korse,Markus Multrus,Guillaume Fuchs备注:Paper accepted to Interspeech 2022 Please check our demo at: this https URL摘要:神经网络已被证明是解决极低比特率语音编码问题的强大工具。然而,设计一种能够在现实条件下稳健运行的神经编码器仍然是一个重大挑战。因此,我们提出了一种稳健、可扩展的端到端神经语音编解码器,用于3kbps的高质量宽带语音编码。编码器使用了一种新的架构配置,它依赖于我们提出的双路径连接神经网络(DPCRNN)层,而解码器架构基于我们之前的工作流式风格。我们对纯净和噪声语音的主观听力测试表明,NESC对看不见的条件和信号扰动特别鲁棒。摘要:Neural networks have proven to be a formidable tool to tackle the problem of speech coding at very low bit rates. However, the design of a neural coder that can be operated robustly under real-world conditions remains a major challenge. Therefore, we present Neural End-2-End Speech Codec (NESC) a robust, scalable end-to-end neural speech codec for high-quality wideband speech coding at 3 kbps. The encoder uses a new architecture configuration, which relies on our proposed Dual-PathConvRNN (DPCRNN) layer, while the decoder architecture is based on our previous work Streamwise-StyleMelGAN. Our subjective listening tests on clean and noisy speech show that NESC is particularly robust to unseen conditions and signal perturbations.
【8】 End-to-end Speech-to-Punctuated-Text Recognition
标题:端到端语音到标点符号文本识别
链接:https://arxiv.org/abs/2207.03169
作者:Jumon Nozaki,Tatsuya Kawahara,Kenkichi Ishizuka,Taiichi Hashimoto备注:Accepted to INTERSPEECH2022摘要:传统的自动语音识别系统不会产生标点符号,标点符号对语音识别结果的可读性很重要。后续的自然语言处理任务(如机器翻译)也需要它们。标点预测模型在语音识别结果中插入标点符号作为后处理,已经有很多工作。然而,这些研究没有利用声学信息进行标点预测,并且直接受到语音识别错误的影响。在本研究中,我们提出了一种端到端模型,该模型将语音作为输入并输出标点文本。该模型有望在使用声学信息时针对语音识别错误稳健地预测标点。我们还建议使用中间层和非同步文本的输出合并辅助损耗来训练模型。通过实验,我们比较了该模型与级联系统的性能。该模型在不牺牲语音识别错误率的情况下,实现了比级联系统更高的标点预测精度。研究还表明,使用中间输出对非定时文本进行多任务学习是有效的。此外,与级联系统相比,该模型仅具有约1/7的参数。摘要:Conventional automatic speech recognition systems do not produce punctuation marks which are important for the readability of the speech recognition results. They are also needed for subsequent natural language processing tasks such as machine translation. There have been a lot of works on punctuation prediction models that insert punctuation marks into speech recognition results as post-processing. However, these studies do not utilize acoustic information for punctuation prediction and are directly affected by speech recognition errors. In this study, we propose an end-to-end model that takes speech as input and outputs punctuated texts. This model is expected to predict punctuation robustly against speech recognition errors while using acoustic information. We also propose to incorporate an auxiliary loss to train the model using the output of the intermediate layer and unpunctuated texts. Through experiments, we compare the performance of the proposed model to that of a cascaded system. The proposed model achieves higher punctuation prediction accuracy than the cascaded system without sacrificing the speech recognition error rate. It is also demonstrated that the multi-task learning using the intermediate output against the unpunctuated text is effective. Moreover, the proposed model has only about 1/7th of the parameters compared to the cascaded system.
【1】 Speech Emotion: Investigating Model Representations, Multi-Task Learning and Knowledge Distillation
标题:言语情感:模型表征、多任务学习与知识提炼研究
链接:https://arxiv.org/abs/2207.03334
* 与cs.SD语音【6】为同一篇
作者:Vikramjit Mitra,Hsiang-Yun Sherry Chien,Vasudha Kowtha,Joseph Yitan Cheng,Erdrin Azemi备注:5 pages, 3 figures, Interspeech 2022摘要:在过去的几年里,从语音信号中估计维度情绪,例如激活、价和优势,已经得到了广泛的探索。虽然从语音中准确估计激活度和优势度似乎是可能的,但对价态的准确估计仍然具有挑战性。以往的研究表明,词汇信息的使用可以提高配价估计性能。词汇信息可以从预先训练的声学模型中获得,其中学习的表示可以改进语音的价估计。我们研究了使用预训练的模型表示来改进声学语音信号的价估计。我们还探索了表征的融合,以改善所有三个情绪维度的情绪估计:激活、配价和优势。此外,我们还研究了是否可以将预训练模型中的表示提取到使用低级特征训练的模型中,从而得到参数数量较少的模型。我们表明,与标准声学特征基线(mel滤波器组能量)相比,预训练模型嵌入的融合导致价估计的协方差相关系数CCC相对提高79%,而从预训练模型嵌入到低维表示的蒸馏产生了相对提高12%。在两个评估集上观察到了这样的性能增益,这表明我们提出的架构在这些评估集上得到了推广。我们在两个MSP播客评估集上报告了最新的“无文本”纯声学维度情感估计值$CCC$。摘要:Estimating dimensional emotions, such as activation, valence and dominance, from acoustic speech signals has been widely explored over the past few years. While accurate estimation of activation and dominance from speech seem to be possible, the same for valence remains challenging. Previous research has shown that the use of lexical information can improve valence estimation performance. Lexical information can be obtained from pre-trained acoustic models, where the learned representations can improve valence estimation from speech. We investigate the use of pre-trained model representations to improve valence estimation from acoustic speech signal. We also explore fusion of representations to improve emotion estimation across all three emotion dimensions: activation, valence and dominance. Additionally, we investigate if representations from pre-trained models can be distilled into models trained with low-level features, resulting in models with a less number of parameters. We show that fusion of pre-trained model embeddings result in a 79% relative improvement in concordance correlation coefficient CCC on valence estimation compared to standard acoustic feature baseline (mel-filterbank energies), while distillation from pre-trained model embeddings to lower-dimensional representations yielded a relative 12% improvement. Such performance gains were observed over two evaluation sets, indicating that our proposed architecture generalizes across those evaluation sets. We report new state-of-the-art "text-free" acoustic-only dimensional emotion estimation $CCC$ values on two MSP-Podcast evaluation sets.
【2】 Low-resource Low-footprint Wake-word Detection using Knowledge Distillation
标题:基于知识蒸馏的低资源低占用尾音检测
链接:https://arxiv.org/abs/2207.03331
作者:Arindam Ghosh,Mark Fuhs,Deblin Bagchi,Bahman Farahani,Monika Woszczyna备注:Accepted to INTERSPEECH 2022摘要:随着虚拟助手变得更加多样化和专业化,对应用程序或品牌特定唤醒词的需求也越来越大。然而,通常用于训练尾词检测器的尾词特定数据集的创建成本很高。在本文中,我们探讨了两种利用声学建模数据进行大词汇量语音识别的技术,以改进专门构建的尾词检测器:转移学习和知识提取。我们还探讨了这些技术如何与时间同步训练目标交互以提高检测延迟。在开源的“Hey-Snips”数据集和更具挑战性的内部远场数据集上进行了实验。使用电话同步目标和从大型声学模型中提取知识,我们能够提高两个数据集的跨数据集大小的准确性,同时减少延迟。摘要:As virtual assistants have become more diverse and specialized, so has the demand for application or brand-specific wake words. However, the wake-word-specific datasets typically used to train wake-word detectors are costly to create. In this paper, we explore two techniques to leverage acoustic modeling data for large-vocabulary speech recognition to improve a purpose-built wake-word detector: transfer learning and knowledge distillation. We also explore how these techniques interact with time-synchronous training targets to improve detection latency. Experiments are presented on the open-source "Hey Snips" dataset and a more challenging in-house far-field dataset. Using phone-synchronous targets and knowledge distillation from a large acoustic model, we are able to improve accuracy across dataset sizes for both datasets while reducing latency.
【3】 NESC: Robust Neural End-2-End Speech Coding with GANs
标题:NESC:基于GANS的稳健神经端双端语音编码
链接:https://arxiv.org/abs/2207.03282
* 与cs.SD语音【7】为同一篇
作者:Nicola Pia,Kishan Gupta,Srikanth Korse,Markus Multrus,Guillaume Fuchs备注:Paper accepted to Interspeech 2022 Please check our demo at: this https URL摘要:神经网络已被证明是解决极低比特率语音编码问题的强大工具。然而,设计一种能够在现实条件下稳健运行的神经编码器仍然是一个重大挑战。因此,我们提出了一种稳健、可扩展的端到端神经语音编解码器,用于3kbps的高质量宽带语音编码。编码器使用了一种新的架构配置,它依赖于我们提出的双路径连接神经网络(DPCRNN)层,而解码器架构基于我们之前的工作流式风格。我们对纯净和噪声语音的主观听力测试表明,NESC对看不见的条件和信号扰动特别鲁棒。摘要:Neural networks have proven to be a formidable tool to tackle the problem of speech coding at very low bit rates. However, the design of a neural coder that can be operated robustly under real-world conditions remains a major challenge. Therefore, we present Neural End-2-End Speech Codec (NESC) a robust, scalable end-to-end neural speech codec for high-quality wideband speech coding at 3 kbps. The encoder uses a new architecture configuration, which relies on our proposed Dual-PathConvRNN (DPCRNN) layer, while the decoder architecture is based on our previous work Streamwise-StyleMelGAN. Our subjective listening tests on clean and noisy speech show that NESC is particularly robust to unseen conditions and signal perturbations.
【4】 End-to-end Speech-to-Punctuated-Text Recognition
标题:端到端语音到标点符号文本识别
链接:https://arxiv.org/abs/2207.03169
* 与cs.SD语音【8】为同一篇
作者:Jumon Nozaki,Tatsuya Kawahara,Kenkichi Ishizuka,Taiichi Hashimoto备注:Accepted to INTERSPEECH2022摘要:传统的自动语音识别系统不会产生标点符号,标点符号对语音识别结果的可读性很重要。后续的自然语言处理任务(如机器翻译)也需要它们。标点预测模型在语音识别结果中插入标点符号作为后处理,已经有很多工作。然而,这些研究没有利用声学信息进行标点预测,并且直接受到语音识别错误的影响。在本研究中,我们提出了一种端到端模型,该模型将语音作为输入并输出标点文本。该模型有望在使用声学信息时针对语音识别错误稳健地预测标点。我们还建议使用中间层和非同步文本的输出合并辅助损耗来训练模型。通过实验,我们比较了该模型与级联系统的性能。该模型在不牺牲语音识别错误率的情况下,实现了比级联系统更高的标点预测精度。研究还表明,使用中间输出对非定时文本进行多任务学习是有效的。此外,与级联系统相比,该模型仅具有约1/7的参数。摘要:Conventional automatic speech recognition systems do not produce punctuation marks which are important for the readability of the speech recognition results. They are also needed for subsequent natural language processing tasks such as machine translation. There have been a lot of works on punctuation prediction models that insert punctuation marks into speech recognition results as post-processing. However, these studies do not utilize acoustic information for punctuation prediction and are directly affected by speech recognition errors. In this study, we propose an end-to-end model that takes speech as input and outputs punctuated texts. This model is expected to predict punctuation robustly against speech recognition errors while using acoustic information. We also propose to incorporate an auxiliary loss to train the model using the output of the intermediate layer and unpunctuated texts. Through experiments, we compare the performance of the proposed model to that of a cascaded system. The proposed model achieves higher punctuation prediction accuracy than the cascaded system without sacrificing the speech recognition error rate. It is also demonstrated that the multi-task learning using the intermediate output against the unpunctuated text is effective. Moreover, the proposed model has only about 1/7th of the parameters compared to the cascaded system.
【5】 Finding Fallen Objects Via Asynchronous Audio-Visual Integration
标题:基于异步音视频融合的坠落物体定位
链接:https://arxiv.org/abs/2207.03483
* 与cs.SD语音【1】为同一篇
作者:Chuang Gan,Yi Gu,Siyuan Zhou,Jeremy Schwartz,Seth Alter,James Traer,Dan Gutfreund,Joshua B. Tenenbaum,Josh McDermott,Antonio Torralba备注:CVPR 2022. Project page: this http URL摘要:物体的外观和声音提供了其物理特性的互补反射。在许多情况下,视觉和听觉的提示是异步到达的,但必须进行整合,就像我们听到物体掉在地板上,然后必须找到它一样。在本文中,我们介绍了一种用于研究三维虚拟环境中的多模式对象定位的环境。一个物体掉在房间的某个地方。配备摄像机和麦克风的嵌入式机器人代理必须通过将音频和视觉信号与基本物理知识相结合来确定掉落了什么物体以及掉落在哪里。为了研究这个问题,我们生成了一个大规模数据集——坠落物体数据集,其中包括64个房间中30个物理物体类别的8000个实例。该数据集使用ThreeDWorld平台,该平台可以在真实照片环境中模拟基于物理的冲击声音和对象之间的复杂物理交互。作为解决这一挑战的第一步,我们基于模仿学习、强化学习和模块化规划开发了一组具体的agent基线,并对这一新任务的挑战进行了深入分析。摘要:The way an object looks and sounds provide complementary reflections of its physical properties. In many settings cues from vision and audition arrive asynchronously but must be integrated, as when we hear an object dropped on the floor and then must find it. In this paper, we introduce a setting in which to study multi-modal object localization in 3D virtual environments. An object is dropped somewhere in a room. An embodied robot agent, equipped with a camera and microphone, must determine what object has been dropped -- and where -- by combining audio and visual signals with knowledge of the underlying physics. To study this problem, we have generated a large-scale dataset -- the Fallen Objects dataset -- that includes 8000 instances of 30 physical object categories in 64 rooms. The dataset uses the ThreeDWorld platform which can simulate physics-based impact sounds and complex physical interactions between objects in a photorealistic setting. As a first step toward addressing this challenge, we develop a set of embodied agent baselines, based on imitation learning, reinforcement learning, and modular planning, and perform an in-depth analysis of the challenge of this new task.
【6】 Non-Linear Pairwise Language Mappings for Low-Resource Multilingual Acoustic Model Fusion
标题:用于低资源多语言声学模型融合的非线性成对语言映射
链接:https://arxiv.org/abs/2207.03391
作者:Muhammad Umar Farooq,Darshan Adiga Haniya Narayana,Thomas Hain备注:Accepted for Interspeech 2022摘要:多语言语音识别作为一种有效的方法来弥补低资源语言的数据不足,已引起人们的广泛关注。端到端(e2e)建模优于传统的混合系统,主要是因为没有词典要求。然而,在有限的数据场景中,混合DNN HMM仍然优于e2e模型。此外,许多语言的字音位(G2P)和文本到IPA音译的公共训练模型缓解了手动创建词典的问题。在本文中,提出了一种新的混合DNN-HMM声学模型融合方法,用于低资源语言的多语言设置。针对目标语言语音信号,将不同单语声学模型的后验分布融合在一起。为每个源-目标语言对训练一个单独的回归神经网络,以将后验值从源声学模型转换为目标语言。与ASR训练相比,这些网络需要的数据非常有限。与多语言和单语基线相比,后融合的相对增益分别为14.65%和6.5%。跨语言模型融合表明,在不使用语言相关ASR的后验值的情况下,可以获得类似的结果。摘要:Multilingual speech recognition has drawn significant attention as an effective way to compensate data scarcity for low-resource languages. End-to-end (e2e) modelling is preferred over conventional hybrid systems, mainly because of no lexicon requirement. However, hybrid DNN-HMMs still outperform e2e models in limited data scenarios. Furthermore, the problem of manual lexicon creation has been alleviated by publicly available trained models of grapheme-to-phoneme (G2P) and text to IPA transliteration for a lot of languages. In this paper, a novel approach of hybrid DNN-HMM acoustic models fusion is proposed in a multilingual setup for the low-resource languages. Posterior distributions from different monolingual acoustic models, against a target language speech signal, are fused together. A separate regression neural network is trained for each source-target language pair to transform posteriors from source acoustic model to the target language. These networks require very limited data as compared to the ASR training. Posterior fusion yields a relative gain of 14.65% and 6.5% when compared with multilingual and monolingual baselines respectively. Cross-lingual model fusion shows that the comparable results can be achieved without using posteriors from the language dependent ASR.
【7】 Investigating the Impact of Cross-lingual Acoustic-Phonetic Similarities on Multilingual Speech Recognition
标题:跨语种声学语音相似性对多语种语音识别的影响研究
链接:https://arxiv.org/abs/2207.03390
作者:Muhammad Umar Farooq,Thomas Hain备注:Accepted for Interspeech 2022摘要:多语言自动语音识别(ASR)系统主要受益于低资源语言,但与单语相对应的几种语言的性能下降。有限的研究集中在理解多语言语音识别设置中的语言行为。本文提出了一种新的数据驱动方法来研究跨语言声学语音相似性。该技术测量不同单语声学模型的后验分布与目标语音信号之间的相似性。深度神经网络训练为映射网络,将不同声学模型的分布转换为直接可比形式。分析发现,语言贴近度不能通过重叠音素集的数量来真正估计。对拟议映射网络的熵分析表明,重叠较少的语言更易于跨语言迁移,因此在多语言环境中更有利。最后,利用提出的后验变换方法融合目标语言的单语模型。与单语对应词相比,相对提高了约8%。摘要:Multilingual automatic speech recognition (ASR) systems mostly benefit low resource languages but suffer degradation in performance across several languages relative to their monolingual counterparts. Limited studies have focused on understanding the languages behaviour in the multilingual speech recognition setups. In this paper, a novel data-driven approach is proposed to investigate the cross-lingual acoustic-phonetic similarities. This technique measures the similarities between posterior distributions from various monolingual acoustic models against a target speech signal. Deep neural networks are trained as mapping networks to transform the distributions from different acoustic models into a directly comparable form. The analysis observes that the languages closeness can not be truly estimated by the volume of overlapping phonemes set. Entropy analysis of the proposed mapping networks exhibits that a language with lesser overlap can be more amenable to cross-lingual transfer, and hence more beneficial in the multilingual setup. Finally, the proposed posterior transformation approach is leveraged to fuse monolingual models for a target language. A relative improvement of ~8% over monolingual counterpart is achieved.
【8】 Self-Supervised Learning of Music-Dance Representation through Explicit-Implicit Rhythm Synchronization
标题:基于显性-隐性节奏同步的乐舞表征自监督学习
链接:https://arxiv.org/abs/2207.03190
* 与cs.SD语音【2】为同一篇
作者:Jiashuo Yu,Junfu Pu,Ying Cheng,Rui Feng,Ying Shan摘要:虽然视听表示已被证明适用于许多下游任务,但舞蹈视频的表示更为具体,总是伴随着具有复杂听觉内容的音乐,仍然具有挑战性,尚未进行研究。考虑到舞者的学员动作和音乐节奏之间的内在一致性,我们引入了一种新的音乐舞蹈表征学习框架MuDaR,以显式和隐式方式实现音乐和舞蹈节奏的同步。具体来说,我们根据视觉外观和受音乐节奏分析启发的动作线索推导出舞蹈节奏。然后,视觉节奏与音乐对应物在时间上对齐,音乐对应物由声音强度的幅度提取。同时,通过对比学习,我们开发了音频和视频流中隐含的节奏的内隐连贯性。该模型通过预测视听对之间的时间一致性来学习联合嵌入。音乐舞蹈表示,以及检测音频和视觉节奏的能力,可以进一步应用于三个下游任务:(a)舞蹈分类,(b)音乐舞蹈检索,和(c)音乐舞蹈重定位。大量实验表明,我们提出的框架在很大程度上优于其他自监督方法。摘要:Although audio-visual representation has been proved to be applicable in many downstream tasks, the representation of dancing videos, which is more specific and always accompanied by music with complex auditory contents, remains challenging and uninvestigated. Considering the intrinsic alignment between the cadent movement of dancer and music rhythm, we introduce MuDaR, a novel Music-Dance Representation learning framework to perform the synchronization of music and dance rhythms both in explicit and implicit ways. Specifically, we derive the dance rhythms based on visual appearance and motion cues inspired by the music rhythm analysis. Then the visual rhythms are temporally aligned with the music counterparts, which are extracted by the amplitude of sound intensity. Meanwhile, we exploit the implicit coherence of rhythms implied in audio and visual streams by contrastive learning. The model learns the joint embedding by predicting the temporal consistency between audio-visual pairs. The music-dance representation, together with the capability of detecting audio and visual rhythms, can further be applied to three downstream tasks: (a) dance classification, (b) music-dance retrieval, and (c) music-dance retargeting. Extensive experiments demonstrate that our proposed framework outperforms other self-supervised methods by a large margin.
【9】 Visual-Assisted Sound Source Depth Estimation in the Wild
标题:野外视觉辅助声源深度估计
链接:https://arxiv.org/abs/2207.03074
* 与cs.SD语音【3】为同一篇
作者:Wei Sun,Lili Qiu备注:13 pages;in submission;摘要:深度估计可以实现多种3D应用,例如机器人、自动驾驶和虚拟现实。尽管在这一领域开展了大量工作,但如何实现准确、低成本、高分辨率和大范围深度估计仍然是个未知数。受闪电爆炸现象(即看到闪电后听到雷声)的启发,本文开发了第一个视听深度估计框架FBDepth。它利用光和声音的飞行时间(ToF)之间的差异来推断声源深度。FBDepth是第一个将视频和音频与语义特征和空间提示结合起来进行距离估计的方法。它首先对齐视频轨迹和音频轨迹之间的对应关系,以粗粒度定位目标对象和目标声音。基于对运动物体轨迹的观察,FBDepth提出了在声音产生之前和之后估计光流交点的方法,以及时定位视频事件。FBDepth为最终深度估计提供视频事件和音频片段的估计时间戳。我们用一部手机收集了3000多个视频片段,其中有20个不同的物体,价格高达5000万美元。与基于RGB的方法相比,FBDepth将绝对相对误差(AbsRel)降低了55%。摘要:Depth estimation enables a wide variety of 3D applications, such as robotics, autonomous driving, and virtual reality. Despite significant work in this area, it remains open how to enable accurate, low-cost, high-resolution, and large-range depth estimation. Inspired by the flash-to-bang phenomenon (\ie hearing the thunder after seeing the lightning), this paper develops FBDepth, the first audio-visual depth estimation framework. It takes the difference between the time-of-flight (ToF) of the light and the sound to infer the sound source depth. FBDepth is the first to incorporate video and audio with both semantic features and spatial hints for range estimation. It first aligns correspondence between the video track and audio track to locate the target object and target sound in a coarse granularity. Based on the observation of moving objects' trajectories, FBDepth proposes to estimate the intersection of optical flow before and after the sound production to locate video events in time. FBDepth feeds the estimated timestamp of the video event and the audio clip for the final depth estimation. We use a mobile phone to collect 3000+ video clips with 20 different objects at up to $50m$. FBDepth decreases the Absolute Relative error (AbsRel) by 55\% compared to RGB-based methods.
【10】 Cross-Scale Vector Quantization for Scalable Neural Speech Coding
标题:用于可伸缩神经语音编码的跨尺度矢量量化
链接:https://arxiv.org/abs/2207.03067
* 与cs.SD语音【4】为同一篇
作者:Xue Jiang,Xiulian Peng,Huaying Xue,Yuan Zhang,Yan Lu备注:INTERSPEECH 2022(Accepted)摘要:比特率可扩展性是实时通信中音频编码的理想特性。现有的神经音频编解码器通常在训练期间强制执行特定的比特率,因此需要针对每个目标比特率训练不同的模型,这增加了发送方和接收方的内存占用,并且通常需要转码来支持多个接收器。在本文中,我们介绍了一种跨尺度可扩展矢量量化方案(CSVQ),在该方案中,多尺度特征通过逐步特征融合和细化进行渐进编码。通过这种方式,如果仅接收到比特流的一部分,则重构粗电平信号,并且随着更多比特可用,逐渐提高质量。提出的CSVQ方案可以灵活地应用于具有镜像自动编码器结构的任何神经音频编码网络,以实现比特率可扩展性。主观结果表明,该方案在可扩展性方面优于经典残差矢量量化(RVQ)。此外,提出的3 kbps的CSVQ优于9 kbps的Opus和3kbps的Lyra,并且可以随着比特率的增加提供优雅的质量提升。摘要:Bitrate scalability is a desirable feature for audio coding in real-time communications. Existing neural audio codecs usually enforce a specific bitrate during training, so different models need to be trained for each target bitrate, which increases the memory footprint at the sender and the receiver side and transcoding is often needed to support multiple receivers. In this paper, we introduce a cross-scale scalable vector quantization scheme (CSVQ), in which multi-scale features are encoded progressively with stepwise feature fusion and refinement. In this way, a coarse-level signal is reconstructed if only a portion of the bitstream is received, and progressively improves the quality as more bits are available. The proposed CSVQ scheme can be flexibly applied to any neural audio coding network with a mirrored auto-encoder structure to achieve bitrate scalability. Subjective results show that the proposed scheme outperforms the classical residual VQ (RVQ) with scalability. Moreover, the proposed CSVQ at 3 kbps outperforms Opus at 9 kbps and Lyra at 3kbps and it could provide a graceful quality boost with bitrate increase.
【11】 Branchformer: Parallel MLP-Attention Architectures to Capture Local and Global Context for Speech Recognition and Understanding
标题:Branchform:为语音识别和理解捕获局部和全局上下文的并行MLP-注意体系结构
链接:https://arxiv.org/abs/2207.02971
* 与cs.SD语音【5】为同一篇
作者:Yifan Peng,Siddharth Dalmia,Ian Lane,Shinji Watanabe摘要:事实证明,整合器在许多语音处理任务中是有效的。它结合了使用卷积提取局部依赖和使用自注意力提取全局依赖的优点。受此启发,我们提出了一种更灵活、可解释和可定制的编码器替代方案Branchformer,它具有并行分支,用于建模端到端语音处理中的各种范围依存关系。在每个编码器层中,一个分支使用自注意力或其变体来捕获长程依赖关系,而另一个分支使用带卷积选通(cgMLP)的MLP模块来提取局部关系。我们在几个语音识别和口语理解基准上进行了实验。结果表明,我们的模型优于Transformer和cgMLP。它还与Conformer取得的最先进成果相匹配或优于Conformer取得的最先进成果。此外,由于双分支架构,我们展示了各种减少计算的策略,包括在单个训练模型中具有可变推理复杂度的能力。为合并分支而学习的权重表示如何在不同层中利用局部和全局依赖关系,这有利于模型设计。摘要:Conformer has proven to be effective in many speech processing tasks. It combines the benefits of extracting local dependencies using convolutions and global dependencies using self-attention. Inspired by this, we propose a more flexible, interpretable and customizable encoder alternative, Branchformer, with parallel branches for modeling various ranged dependencies in end-to-end speech processing. In each encoder layer, one branch employs self-attention or its variant to capture long-range dependencies, while the other branch utilizes an MLP module with convolutional gating (cgMLP) to extract local relationships. We conduct experiments on several speech recognition and spoken language understanding benchmarks. Results show that our model outperforms both Transformer and cgMLP. It also matches with or outperforms state-of-the-art results achieved by Conformer. Furthermore, we show various strategies to reduce computation thanks to the two-branch architecture, including the ability to have variable inference complexity in a single trained model. The weights learned for merging branches indicate how local and global dependencies are utilized in different layers, which benefits model designing.机器翻译,仅供参考