今日论文合集:cs.SD语音4篇,eess.AS音频处理5篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 On the Condition Monitoring of Bolted Joints through Acoustic Emission and Deep Transfer Learning: Generalization, Ordinal Loss and Super-Convergence
标题: 通过声发射和深度传输学习监测螺栓接头的状态:概括、有序损失和超收敛
作者:Emmanuel Ramasso,Rafael de O. Teloli,Romain Marcel
链接:点击下载PDF文件
摘要:本文研究了使用基于卷积神经网络(CNN)的深度迁移学习来使用声发射监测螺栓连接的状况。螺栓结构是许多机械系统中的关键部件,监测其状态的能力对于有效的结构健康监测至关重要。我们使用ORION-AE基准评估了我们的方法的性能,ORION-AE基准是由两个由三个螺栓连接的薄梁组成的结构,其中采用高噪声声发射测量来检测螺栓施加的拧紧扭矩的变化。该结构中使用的数据来自于使用连续小波变换将声发射数据流转换为图像,并利用预训练的CNN进行特征提取和去噪。我们的实验比较了单传感器与多传感器融合估计拧紧水平(松动)的螺栓,并评估了使用原始与预过滤数据的性能。我们特别关注基于CNN的迁移学习在不同测量活动中的泛化能力,并且我们研究了顺序损失函数,以便在接近地面事实时不那么严重地惩罚不正确的预测,从而鼓励错误分类错误出现在相邻的类别中。还研究了网络结构和学习率的变化,并得到了超收敛性,即,用不同的网络在几次迭代中实现了高的分类精度。此外,结果表明,基于CNN的迁移学习的泛化能力监测螺栓结构的声发射与训练过程中所需的先验信息的变化量。摘要:This paper investigates the use of deep transfer learning based on convolutional neural networks (CNNs) to monitor the condition of bolted joints using acoustic emissions. Bolted structures are critical components in many mechanical systems, and the ability to monitor their condition status is crucial for effective structural health monitoring. We evaluated the performance of our methodology using the ORION-AE benchmark, a structure composed of two thin beams connected by three bolts, where highly noisy acoustic emission measurements were taken to detect changes in the applied tightening torque of the bolts. The data used from this structure is derived from the transformation of acoustic emission data streams into images using continuous wavelet transform, and leveraging pretrained CNNs for feature extraction and denoising. Our experiments compared single-sensor versus multiple-sensor fusion for estimating the tightening level (loosening) of bolts and evaluated the use of raw versus prefiltered data on the performance. We particularly focused on the generalization capabilities of CNN-based transfer learning across different measurement campaigns and we studied ordinal loss functions to penalize incorrect predictions less severely when close to the ground truth, thereby encouraging misclassification errors to be in adjacent classes. Network configurations as well as learning rate schedulers are also investigated, and super-convergence is obtained, i.e., high classification accuracy is achieved in a few number of iterations with different networks. Furthermore, results demonstrate the generalization capabilities of CNN-based transfer learning for monitoring bolted structures by acoustic emission with varying amounts of prior information required during training.

【2】 Effects of Dataset Sampling Rate for Noise Cancellation through Deep Learning
标题: 数据集采样率对深度学习降噪的影响
作者:Brandon Colelough,Andrew Zheng
备注:16 pages, 8 pictures, 3 tables
链接:点击下载PDF文件
摘要:背景:有源噪声消除已经是几十年来的研究主题。传统的技术,如快速傅立叶变换,在某些情况下有局限性。这项研究探索了深度神经网络(DNN)作为一种更好的替代方案的使用。目的:该研究旨在确定训练数据中的采样率对在移动设备的处理约束下运行的轻量级、高效DNN的影响。方法:我们选择了ConvTasNET网络,因为它在语音分离和增强方面的效率已经得到了证明。ConvTasNET是在WHAM!,LibriMix和MS-2023 DNS挑战。以8 kHz、16 kHz和48 kHz的速率对数据集进行采样,以分析采样速率对噪声消除效率和有效性的影响。该模型在2023年的英特尔酷睿i7处理器上进行了测试,评估了网络在过滤背景噪音的同时产生清晰音频的能力。结果:在更高的采样率(48 kHz)下训练的模型提供了更好的总谐波失真(THD)和生成神经语音编解码器质量预测(WARP-Q)值的评估指标,表明音频质量得到了改善。然而,注意到一种折衷,即较高采样率的处理时间较长。结论:Conv-TasNET网络在以48 kHz等更高速率采样的数据集上进行训练,为移动设备提供了一个强大的解决方案,通过语音分离和增强实现噪声消除。未来的工作包括进一步优化模型的效率,并在移动设备上进行测试。摘要:Background: Active noise cancellation has been a subject of research for decades. Traditional techniques, like the Fast Fourier Transform, have limitations in certain scenarios. This research explores the use of deep neural networks (DNNs) as a superior alternative. Objective: The study aims to determine the effect sampling rate within training data has on lightweight, efficient DNNs that operate within the processing constraints of mobile devices. Methods: We chose the ConvTasNET network for its proven efficiency in speech separation and enhancement. ConvTasNET was trained on datasets such as WHAM!, LibriMix, and the MS-2023 DNS Challenge. The datasets were sampled at rates of 8kHz, 16kHz, and 48kHz to analyze the effect of sampling rate on noise cancellation efficiency and effectiveness. The model was tested on a core-i7 Intel processor from 2023, assessing the network's ability to produce clear audio while filtering out background noise. Results: Models trained at higher sampling rates (48kHz) provided much better evaluation metrics against Total Harmonic Distortion (THD) and Quality Prediction For Generative Neural Speech Codecs (WARP-Q) values, indicating improved audio quality. However, a trade-off was noted with the processing time being longer for higher sampling rates. Conclusions: The Conv-TasNET network, trained on datasets sampled at higher rates like 48kHz, offers a robust solution for mobile devices in achieving noise cancellation through speech separation and enhancement. Future work involves optimizing the model's efficiency further and testing on mobile devices.

【3】 SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
标题: 无限期ExpressiveLM:具有思想链的表达性言语翻译的语音语言模型
作者:Hongyu Gong,Bandhav Veluri
链接:点击下载PDF文件
摘要:表达性语音到语音翻译(S2ST)是无缝通信中的一个重要研究课题,其重点是在翻译的语音中保留语义和说话人的声音风格。早期的工作合成说话人风格对齐语音,以便直接学习从语音到目标语音谱图的映射。在不依赖风格对齐数据的情况下,最近的研究利用语言建模(LM)的进步,并在语义和声学标记上构建级联LM。这项工作提出了一个单一的语音语言模型,表达S2ST无障碍ExpressiveLM。我们将复杂的源到目标的语音映射分解为中间生成步骤与思想链提示。该模型首先引导目标语义内容的翻译,然后将说话人风格转换为多流声学单元。在西班牙语到英语和匈牙利语到英语的翻译上进行了评估,在语义质量和风格转换方面,无偏表达式LM都优于级联LM,同时实现了更好的参数效率。摘要:Expressive speech-to-speech translation (S2ST) is a key research topic in seamless communication, which focuses on the preservation of semantics and speaker vocal style in translated speech. Early works synthesized speaker style aligned speech in order to directly learn the mapping from speech to target speech spectrogram. Without reliance on style aligned data, recent studies leverage the advances of language modeling (LM) and build cascaded LMs on semantic and acoustic tokens. This work proposes SeamlessExpressiveLM, a single speech language model for expressive S2ST. We decompose the complex source-to-target speech mapping into intermediate generation steps with chain-of-thought prompting. The model is first guided to translate target semantic content and then transfer the speaker style to multi-stream acoustic units. Evaluated on Spanish-to-English and Hungarian-to-English translations, SeamlessExpressiveLM outperforms cascaded LMs in both semantic quality and style transfer, meanwhile achieving better parameter efficiency.

【4】 Cross-Talk Reduction
标题: 串扰抑制
作者:Zhong-Qiu Wang,Anurag Kumar,Shinji Watanabe
备注:in International Joint Conference on Artificial Intelligence (IJCAI), 2024
链接:点击下载PDF文件
摘要:当记录远场多说话者混合时,每个说话者可以佩戴近距离通话麦克风,以便可以同时记录近距离通话混合。虽然每个近距离说话混合具有佩戴者的高信噪比(SNR),但是其具有非常有限的应用范围,因为其还包含其他说话者的显著串扰语音并且不够干净。在这种情况下,我们提出了一种新的任务称为串扰减少(CTR),旨在减少串扰语音,和一种新的解决方案命名为CTRnet,它是基于无监督或弱监督神经语音分离。在无监督的CTRnet中,近距离说话和远场混合物被堆叠作为DNN的输入,以估计每个说话者的近距离说话语音。它以无监督的、有区别的方式进行训练,使得每个扬声器的DNN估计可以被线性过滤,以消除在其他麦克风处捕获的扬声器的串扰语音。在弱监督的CTRnet中,我们假设在训练过程中每个说话者的活动时间戳的可用性,并利用它们来改进无监督的CTRnet的训练。在一个模拟的两个扬声器CTR任务和一个真实记录的会话语音分离和识别任务的评估结果表明CTRnet的有效性和潜力。摘要:While far-field multi-talker mixtures are recorded, each speaker can wear a close-talk microphone so that close-talk mixtures can be recorded at the same time. Although each close-talk mixture has a high signal-to-noise ratio (SNR) of the wearer, it has a very limited range of applications, as it also contains significant cross-talk speech by other speakers and is not clean enough. In this context, we propose a novel task named cross-talk reduction (CTR) which aims at reducing cross-talk speech, and a novel solution named CTRnet which is based on unsupervised or weakly-supervised neural speech separation. In unsupervised CTRnet, close-talk and far-field mixtures are stacked as input for a DNN to estimate the close-talk speech of each speaker. It is trained in an unsupervised, discriminative way such that the DNN estimate for each speaker can be linearly filtered to cancel out the speaker's cross-talk speech captured at other microphones. In weakly-supervised CTRnet, we assume the availability of each speaker's activity timestamps during training, and leverage them to improve the training of unsupervised CTRnet. Evaluation results on a simulated two-speaker CTR task and on a real-recorded conversational speech separation and recognition task show the effectiveness and potential of CTRnet.


eess.AS音频处理
【1】 Very Low Complexity Speech Synthesis Using Framewise Autoregressive GAN (FARGAN) with Pitch Prediction
标题: 使用带音调预测的逐帧自回归GAN(FARGAN)进行极低复杂度语音合成
作者:Jean-Marc Valin,Ahmed Mustafa,Jan Büthe
备注:5 pages
链接:点击下载PDF文件
摘要:神经声码器现在被广泛用于语音处理应用中。在许多应用中,声码器可能是最复杂的组件,因此找到较低复杂度的算法可以带来显着的实际好处。在这项工作中,我们提出了FARGAN,一个自回归声码器,利用长期的音高预测,在小的子帧合成高质量的语音,而不需要教师强迫。实验结果表明,与现有的低复杂度声码器相比,本文提出的600~MFLOPS FARGAN声码器可以实现更高的质量和更低的复杂度。质量甚至与现有的更高复杂度的声码器相匹配。摘要:Neural vocoders are now being used in a wide range of speech processing applications. In many of those applications, the vocoder can be the most complex component, so finding lower complexity algorithms can lead to significant practical benefits. In this work, we propose FARGAN, an autoregressive vocoder that takes advantage of long-term pitch prediction to synthesize high-quality speech in small subframes, without the need for teacher-forcing. Experimental results show that the proposed 600~MFLOPS FARGAN vocoder can achieve both higher quality and lower complexity than existing low-complexity vocoders. The quality even matches that of existing higher-complexity vocoders.

【2】 Cross-Talk Reduction
标题: 串扰抑制
作者:Zhong-Qiu Wang,Anurag Kumar,Shinji Watanabe
备注:in International Joint Conference on Artificial Intelligence (IJCAI), 2024
链接:点击下载PDF文件
摘要:当记录远场多说话者混合时,每个说话者可以佩戴近距离通话麦克风,以便可以同时记录近距离通话混合。虽然每个近距离说话混合具有佩戴者的高信噪比(SNR),但是其具有非常有限的应用范围,因为其还包含其他说话者的显著串扰语音并且不够干净。在这种情况下,我们提出了一种新的任务称为串扰减少(CTR),旨在减少串扰语音,和一种新的解决方案命名为CTRnet,它是基于无监督或弱监督神经语音分离。在无监督的CTRnet中,近距离说话和远场混合物被堆叠作为DNN的输入,以估计每个说话者的近距离说话语音。它以无监督的、有区别的方式进行训练,使得每个扬声器的DNN估计可以被线性过滤,以消除在其他麦克风处捕获的扬声器的串扰语音。在弱监督的CTRnet中,我们假设在训练过程中每个说话者的活动时间戳的可用性,并利用它们来改进无监督的CTRnet的训练。在一个模拟的两个扬声器CTR任务和一个真实记录的会话语音分离和识别任务的评估结果表明CTRnet的有效性和潜力。摘要:While far-field multi-talker mixtures are recorded, each speaker can wear a close-talk microphone so that close-talk mixtures can be recorded at the same time. Although each close-talk mixture has a high signal-to-noise ratio (SNR) of the wearer, it has a very limited range of applications, as it also contains significant cross-talk speech by other speakers and is not clean enough. In this context, we propose a novel task named cross-talk reduction (CTR) which aims at reducing cross-talk speech, and a novel solution named CTRnet which is based on unsupervised or weakly-supervised neural speech separation. In unsupervised CTRnet, close-talk and far-field mixtures are stacked as input for a DNN to estimate the close-talk speech of each speaker. It is trained in an unsupervised, discriminative way such that the DNN estimate for each speaker can be linearly filtered to cancel out the speaker's cross-talk speech captured at other microphones. In weakly-supervised CTRnet, we assume the availability of each speaker's activity timestamps during training, and leverage them to improve the training of unsupervised CTRnet. Evaluation results on a simulated two-speaker CTR task and on a real-recorded conversational speech separation and recognition task show the effectiveness and potential of CTRnet.

【3】 On the Condition Monitoring of Bolted Joints through Acoustic Emission and Deep Transfer Learning: Generalization, Ordinal Loss and Super-Convergence
标题: 通过声发射和深度传输学习监测螺栓接头的状态:概括、有序损失和超收敛
作者:Emmanuel Ramasso,Rafael de O. Teloli,Romain Marcel
链接:点击下载PDF文件
摘要:本文研究了使用基于卷积神经网络(CNN)的深度迁移学习来使用声发射监测螺栓连接的状况。螺栓结构是许多机械系统中的关键部件,监测其状态的能力对于有效的结构健康监测至关重要。我们使用ORION-AE基准评估了我们的方法的性能,ORION-AE基准是由两个由三个螺栓连接的薄梁组成的结构,其中采用高噪声声发射测量来检测螺栓施加的拧紧扭矩的变化。该结构中使用的数据来自于使用连续小波变换将声发射数据流转换为图像,并利用预训练的CNN进行特征提取和去噪。我们的实验比较了单传感器与多传感器融合估计拧紧水平(松动)的螺栓,并评估了使用原始与预过滤数据的性能。我们特别关注基于CNN的迁移学习在不同测量活动中的泛化能力,并且我们研究了顺序损失函数,以便在接近地面事实时不那么严重地惩罚不正确的预测,从而鼓励错误分类错误出现在相邻的类别中。还研究了网络结构和学习率的变化,并得到了超收敛性,即,用不同的网络在几次迭代中实现了高的分类精度。此外,结果表明,基于CNN的迁移学习的泛化能力监测螺栓结构的声发射与训练过程中所需的先验信息的变化量。摘要:This paper investigates the use of deep transfer learning based on convolutional neural networks (CNNs) to monitor the condition of bolted joints using acoustic emissions. Bolted structures are critical components in many mechanical systems, and the ability to monitor their condition status is crucial for effective structural health monitoring. We evaluated the performance of our methodology using the ORION-AE benchmark, a structure composed of two thin beams connected by three bolts, where highly noisy acoustic emission measurements were taken to detect changes in the applied tightening torque of the bolts. The data used from this structure is derived from the transformation of acoustic emission data streams into images using continuous wavelet transform, and leveraging pretrained CNNs for feature extraction and denoising. Our experiments compared single-sensor versus multiple-sensor fusion for estimating the tightening level (loosening) of bolts and evaluated the use of raw versus prefiltered data on the performance. We particularly focused on the generalization capabilities of CNN-based transfer learning across different measurement campaigns and we studied ordinal loss functions to penalize incorrect predictions less severely when close to the ground truth, thereby encouraging misclassification errors to be in adjacent classes. Network configurations as well as learning rate schedulers are also investigated, and super-convergence is obtained, i.e., high classification accuracy is achieved in a few number of iterations with different networks. Furthermore, results demonstrate the generalization capabilities of CNN-based transfer learning for monitoring bolted structures by acoustic emission with varying amounts of prior information required during training.

【4】 Effects of Dataset Sampling Rate for Noise Cancellation through Deep Learning
标题: 数据集采样率对深度学习降噪的影响
作者:Brandon Colelough,Andrew Zheng
备注:16 pages, 8 pictures, 3 tables
链接:点击下载PDF文件
摘要:背景:有源噪声消除已经是几十年来的研究主题。传统的技术,如快速傅立叶变换,在某些情况下有局限性。这项研究探索了深度神经网络(DNN)作为一种更好的替代方案的使用。目的:该研究旨在确定训练数据中的采样率对在移动设备的处理约束下运行的轻量级、高效DNN的影响。方法:我们选择了ConvTasNET网络,因为它在语音分离和增强方面的效率已经得到了证明。ConvTasNET是在WHAM!,LibriMix和MS-2023 DNS挑战。以8 kHz、16 kHz和48 kHz的速率对数据集进行采样,以分析采样速率对噪声消除效率和有效性的影响。该模型在2023年的英特尔酷睿i7处理器上进行了测试,评估了网络在过滤背景噪音的同时产生清晰音频的能力。结果如下:在更高的采样率(48 kHz)下训练的模型提供了更好的总谐波失真(THD)和生成神经语音编解码器质量预测(WARP-Q)值的评估指标,表明音频质量得到了改善。然而,注意到一种折衷,即较高采样率的处理时间较长。结论:Conv-TasNET网络在以48 kHz等更高速率采样的数据集上进行训练,为移动设备提供了一个强大的解决方案,通过语音分离和增强实现噪声消除。未来的工作包括进一步优化模型的效率,并在移动设备上进行测试。摘要:Background: Active noise cancellation has been a subject of research for decades. Traditional techniques, like the Fast Fourier Transform, have limitations in certain scenarios. This research explores the use of deep neural networks (DNNs) as a superior alternative. Objective: The study aims to determine the effect sampling rate within training data has on lightweight, efficient DNNs that operate within the processing constraints of mobile devices. Methods: We chose the ConvTasNET network for its proven efficiency in speech separation and enhancement. ConvTasNET was trained on datasets such as WHAM!, LibriMix, and the MS-2023 DNS Challenge. The datasets were sampled at rates of 8kHz, 16kHz, and 48kHz to analyze the effect of sampling rate on noise cancellation efficiency and effectiveness. The model was tested on a core-i7 Intel processor from 2023, assessing the network's ability to produce clear audio while filtering out background noise. Results: Models trained at higher sampling rates (48kHz) provided much better evaluation metrics against Total Harmonic Distortion (THD) and Quality Prediction For Generative Neural Speech Codecs (WARP-Q) values, indicating improved audio quality. However, a trade-off was noted with the processing time being longer for higher sampling rates. Conclusions: The Conv-TasNET network, trained on datasets sampled at higher rates like 48kHz, offers a robust solution for mobile devices in achieving noise cancellation through speech separation and enhancement. Future work involves optimizing the model's efficiency further and testing on mobile devices.

【5】 SeamlessExpressiveLM: Speech Language Model for Expressive Speech-to-Speech Translation with Chain-of-Thought
标题: 无限期ExpressiveLM:具有思想链的表达性言语翻译的语音语言模型
作者:Hongyu Gong,Bandhav Veluri
链接:点击下载PDF文件
摘要:表达性语音到语音翻译(S2ST)是无缝通信中的一个重要研究课题,其重点是在翻译的语音中保留语义和说话人的声音风格。早期的工作合成说话人风格对齐语音,以便直接学习从语音到目标语音谱图的映射。在不依赖风格对齐数据的情况下,最近的研究利用语言建模(LM)的进步,并在语义和声学标记上构建级联LM。这项工作提出了一个单一的语音语言模型,表达S2ST无障碍ExpressiveLM。我们将复杂的源到目标的语音映射分解为中间生成步骤与思想链提示。该模型首先引导目标语义内容的翻译,然后将说话人风格转换为多流声学单元。在西班牙语到英语和匈牙利语到英语的翻译上进行了评估,在语义质量和风格转换方面,无偏表达式LM都优于级联LM,同时实现了更好的参数效率。摘要:Expressive speech-to-speech translation (S2ST) is a key research topic in seamless communication, which focuses on the preservation of semantics and speaker vocal style in translated speech. Early works synthesized speaker style aligned speech in order to directly learn the mapping from speech to target speech spectrogram. Without reliance on style aligned data, recent studies leverage the advances of language modeling (LM) and build cascaded LMs on semantic and acoustic tokens. This work proposes SeamlessExpressiveLM, a single speech language model for expressive S2ST. We decompose the complex source-to-target speech mapping into intermediate generation steps with chain-of-thought prompting. The model is first guided to translate target semantic content and then transfer the speaker style to multi-stream acoustic units. Evaluated on Spanish-to-English and Hungarian-to-English translations, SeamlessExpressiveLM outperforms cascaded LMs in both semantic quality and style transfer, meanwhile achieving better parameter efficiency.


机器翻译,仅供参考