今日论文合集:cs.SD语音10篇,eess.AS音频处理11篇。

本文经arXiv每日学术速递授权转载


cs.SD语音

【1】Robust Dual-Modal Speech Keyword Spotting for XR Headsets
标题:面向XR耳机的稳健双模式语音关键词检测
链接:https://arxiv.org/abs/2401.14978
作者:Zhuojiang Cai,Yuhan Ma,Feng Lu
备注:Accepted to IEEE VR 2024
摘要:虽然语音交互在延展实境(XR)领域内得到广泛应用,但是传统的有声语音关键词发现系统继续应对巨大的挑战,包括在嘈杂环境中的次优性能、在需要静音的情况下的不切实际性以及当其他人在附近说话时对无意激活的敏感性。然而,这些挑战可以通过具有成本效益的语音和嘴唇运动信息的融合来克服。因此,我们提出了一种新的人声回声双模态关键字定位系统设计的XR耳机。我们设计了两种不同的模态融合方法,并进行实验来测试系统在不同场景下的性能。结果表明,我们的双模态系统不仅始终优于其单模态同行,在典型和嘈杂的环境中表现出更高的精度,但也擅长准确地识别无声的话语。此外,我们已经成功地将该系统应用于实时演示,取得了可喜的成果。该代码可在https://github.com/caizhuojiang/VE-KWS上获得。
摘要:While speech interaction finds widespread utility within the Extended Reality (XR) domain, conventional vocal speech keyword spotting systems continue to grapple with formidable challenges, including suboptimal performance in noisy environments, impracticality in situations requiring silence, and susceptibility to inadvertent activations when others speak nearby. These challenges, however, can potentially be surmounted through the cost-effective fusion of voice and lip movement information. Consequently, we propose a novel vocal-echoic dual-modal keyword spotting system designed for XR headsets. We devise two different modal fusion approches and conduct experiments to test the system's performance across diverse scenarios. The results show that our dual-modal system not only consistently outperforms its single-modal counterparts, demonstrating higher precision in both typical and noisy environments, but also excels in accurately identifying silent utterances. Furthermore, we have successfully applied the system in real-time demonstrations, achieving promising results. The code is available at https://github.com/caizhuojiang/VE-KWS.


【2】 Comparison of parameters of vowel sounds of russian and english  languages
标题:俄语和英语元音参数的比较
链接:https://arxiv.org/abs/2401.14890
作者:V. I. Fedoseev,A. A. Konev,A. Yu. Yakimuk
备注:7 pages, 1 figures, 3 tables
摘要:在多语言语音识别系统中,经常会出现这样一种情况,即事先不知道语言,但信号已经被接收并正在处理。对于这种情况,需要某种通用模型,它将能够响应语音差异,并根据它们正确地识别所需语言中的语音。为了建立这样一个模型,有必要设置语音参数的值,然后比较相似的声音,建立显着差异。
摘要:In multilingual speech recognition systems, a situation can often arise when the language is not known in advance, but the signal has already been received and is being processed. For such cases, some generalized model is needed that will be able to respond to phonetic differences and, depending on them, correctly recog-nize speech in the desired language. To build such a model, it is necessary to set the values of phonetic parameters, and then compare similar sounds, establishing significant differences.

【3】 Expressivity-aware Music Performance Retrieval using Mid-level  Perceptual Features and Emotion Word Embeddings
标题:基于中层感知特征和情感词嵌入的表现力感知音乐表演检索
链接:https://arxiv.org/abs/2401.14826
作者:Shreyan Chowdhury,Gerhard Widmer
备注:Presented at FIRE 2023 (Forum for Information Retrieval Evaluation) conference, Goa, India
摘要:本文探讨了跨模态音乐检索的一个具体子任务。我们认为,检索一个性能或再现的音乐作品的基础上,其风格的描述,表达的特点,或从一组不同的性能相同的作品的情感的微妙的任务。我们观察到,一个通用的跨模态系统训练学习一个共同的文本音频嵌入空间不会产生最佳的结果,为这项任务。通过引入两个变化-文本编码器和音频编码器各一个-我们在钢琴演奏和相关自由文本描述的数据集上展示了改进的性能。在文本方面,我们使用情感丰富的词嵌入(EWE),在音频方面,我们提取中级感知特征,而不是通用的音频嵌入。我们的研究结果突出了从音乐中学习的中级感知特征和从情感标记文本中学习的情感丰富的词嵌入在跨模态设置中捕获音乐表达的有效性。此外,我们的可解释的中级功能提供了一个路线,在检索和下游推荐过程中引入可解释性。
摘要:This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a set of different performances of the same piece. We observe that a general purpose cross-modal system trained to learn a common text-audio embedding space does not yield optimal results for this task. By introducing two changes -- one each to the text encoder and the audio encoder -- we demonstrate improved performance on a dataset of piano performances and associated free-text descriptions. On the text side, we use emotion-enriched word embeddings (EWE) and on the audio side, we extract mid-level perceptual features instead of generic audio embeddings. Our results highlight the effectiveness of mid-level perceptual features learnt from music and emotion enriched word embeddings learnt from emotion-labelled text in capturing musical expression in a cross-modal setting. Additionally, our interpretable mid-level features provide a route for introducing explainability in the retrieval and downstream recommendation processes.

【4】 Turn-taking and Backchannel Prediction with Acoustic and Large Language  Model Fusion
标题:基于声学和大语言模型融合的话轮转换和回声预测
链接:https://arxiv.org/abs/2401.14717
作者:Jinhan Wang,Long Chen,Aparna Khare,Anirudh Raju,Pranav Dheram,Di He,Minhua Wu,Andreas Stolcke,Venkatesh Ravichandran
备注:To appear in IEEE ICASSP 2024
摘要:我们提出了一种方法,通过融合神经声学模型与大语言模型(LLM),连续预测口语对话中的话轮转换和反向通道位置。在Switchboard人-人对话数据集上的实验表明,我们的方法始终优于单一模态的基线模型。我们还开发了一种新的多任务指令微调策略,以进一步受益于LLM编码的知识,用于理解任务和会话上下文,从而实现额外的改进。我们的方法展示了LLM和声学模型相结合的潜力,可以在人类和支持语音的AI代理之间进行更自然和对话式的交互。
摘要:We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms the baseline models with single modality. We also develop a novel multi-task instruction fine-tuning strategy to further benefit from LLM-encoded knowledge for understanding the tasks and conversational contexts, leading to additional improvements. Our approach demonstrates the potential of combined LLMs and acoustic models for a more natural and conversational interaction between humans and speech-enabled AI agents.


【5】 UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit  Normalization
标题:UNIT-DSR:基于语音单元归一化的动态语音重建系统
链接:https://arxiv.org/abs/2401.14664
作者:Yuejiao Wang,Xixin Wu,Disong Wang,Lingwei Meng,Helen Meng
备注:Accepted to ICASSP 2024
摘要:构音障碍语音重建(DSR)系统旨在自动将构音障碍语音转换为正常发音语音。这项技术简化了与受神经运动障碍影响的说话者的沟通,并增强了他们的社会包容性。与基于GAN(生成对抗网络)的方法相比,基于NED(神经编码器-解码器)的系统显着提高了重建语音的可理解性,但该方法仍然受到级联管道和内容编码器辅助任务导致的训练效率低下的限制,这反过来又可能影响重建的质量。受自监督语音表示学习和离散语音单元的启发,我们提出了一个Unit-DSR系统,该系统利用HuBERT强大的域自适应能力来提高训练效率,并利用语音单元来约束离散语言空间中的构音障碍内容恢复。与NED方法相比,Unit-DSR系统仅由语音单元归一化器和Unit HiFi-GAN声码器组成,这是相当简单的,没有级联的子模块或辅助任务。UASpeech语料库上的结果表明,Unit-DSR在内容恢复方面优于竞争基线,与原始构音障碍语音相比,相对平均单词错误率降低了28.2%,并且对速度扰动和噪声具有鲁棒性。
摘要:Dysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected by the neuromotor disorder and enhances their social inclusion. NED-based (Neural Encoder-Decoder) systems have significantly improved the intelligibility of the reconstructed speech as compared with GAN-based (Generative Adversarial Network) approaches, but the approach is still limited by training inefficiency caused by the cascaded pipeline and auxiliary tasks of the content encoder, which may in turn affect the quality of reconstruction. Inspired by self-supervised speech representation learning and discrete speech units, we propose a Unit-DSR system, which harnesses the powerful domain-adaptation capacity of HuBERT for training efficiency improvement and utilizes speech units to constrain the dysarthric content restoration in a discrete linguistic space. Compared with NED approaches, the Unit-DSR system only consists of a speech unit normalizer and a Unit HiFi-GAN vocoder, which is considerably simpler without cascaded sub-modules or auxiliary tasks. Results on the UASpeech corpus indicate that Unit-DSR outperforms competitive baselines in terms of content restoration, reaching a 28.2% relative average word error rate reduction when compared to original dysarthric speech, and shows robustness against speed perturbation and noise.

【6】 Exploring Musical Roots: Applying Audio Embeddings to Empower Influence  Attribution for a Generative Music Model
标题:探索音乐根源:应用音频嵌入增强生成性音乐模型的影响力归因
链接:https://arxiv.org/abs/2401.14542
作者:Julia Barnett,Hugo Flores Garcia,Bryan Pardo
备注:14 pages + references. Under conference review
摘要:每个艺术家都有一个创作过程,从以前的艺术家和他们的作品中汲取灵感。今天,“灵感”已经被生成音乐模型自动化。这些模型的黑箱性质掩盖了影响其创造性产出的作品的身份。因此,用户可能会无意中盗用、误用或复制现有艺术家的作品。我们建立了一种可复制的方法,以系统地识别类似的音乐音频片段,这种方法有助于理解训练数据的属性。我们的方法的一个关键方面是利用一个有效的音乐音频相似性度量。我们比较了应用CLMR和CLAP嵌入相似性测量的效果,在一组用于训练VampNet的500万个音频片段,最近的开源生成音乐模型。我们用人类听力研究验证了这种方法。我们还探索了音频示例的修改(例如,音调偏移、时间拉伸、背景噪声)对相似性测量有影响。这项工作是将自动影响归因纳入生成建模的基础,这有望让模型创建者和用户从无知的挪用转向知情的创建。本文附带的音频样本可在https://tinyurl.com/exploring-musical-roots上获得。
摘要:Every artist has a creative process that draws inspiration from previous artists and their works. Today, "inspiration" has been automated by generative music models. The black box nature of these models obscures the identity of the works that influence their creative output. As a result, users may inadvertently appropriate, misuse, or copy existing artists' works. We establish a replicable methodology to systematically identify similar pieces of music audio in a manner that is useful for understanding training data attribution. A key aspect of our approach is to harness an effective music audio similarity measure. We compare the effect of applying CLMR and CLAP embeddings to similarity measurement in a set of 5 million audio clips used to train VampNet, a recent open source generative music model. We validate this approach with a human listening study. We also explore the effect that modifications of an audio example (e.g., pitch shifting, time stretching, background noise) have on similarity measurements. This work is foundational to incorporating automated influence attribution into generative modeling, which promises to let model creators and users move from ignorant appropriation to informed creation. Audio samples that accompany this paper are available at https://tinyurl.com/exploring-musical-roots.

【7】 ICASSP 2024 Speech Signal Improvement Challenge
标题:ICASSP 2024语音信号改善挑战赛
链接:https://arxiv.org/abs/2401.14444
作者:Nicolae Catalin Ristea,Ando Saabas,Ross Cutler,Babak Naderi,Sebastian Braun,Solomiya Branets摘要:ICASSP 2024语音信号改善大挑战赛旨在促进提高通信系统中语音信号质量领域的研究。这标志着我们的第二个挑战,建立在上一届ICASSP 2023大挑战赛的成功基础上。我们通过引入数据集合成器来增强竞争,使所有参赛团队能够从更高的基线开始,这是我们扩展的P.804测试的客观指标,2023测试集的成绩单,我们还添加了单词准确性(WAcc)作为指标。我们评估了13个系统的实时跟踪和11个系统的非实时跟踪使用主观的P.804和客观的字的准确性指标。
摘要:The ICASSP 2024 Speech Signal Improvement Grand Challenge is intended to stimulate research in the area of improving the speech signal quality in communication systems. This marks our second challenge, building upon the success from the previous ICASSP 2023 Grand Challenge. We enhance the competition by introducing a dataset synthesizer, enabling all participating teams to start at a higher baseline, an objective metric for our extended P.804 tests, transcripts for the 2023 test set, and we add Word Accuracy (WAcc) as a metric. We evaluate a total of 13 systems in the real-time track and 11 systems in the non-real-time track using both subjective P.804 and objective Word Accuracy metrics.

【8】 Spatial Analysis and Synthesis Methods: Subjective and Objective  Evaluations Using Various Microphone Arrays in the Auralization of a Critical  Listening Room
标题:空间分析和综合方法:在关键听音室可听化中使用不同麦克风阵列的主客观评价
链接:https://arxiv.org/abs/2401.15023
作者:Alan Pawlak,Hyunkook Lee,Aki Mäkivirta,Thomas Lund
备注:13 pages, 6 figures
摘要:参数声场合成方法,如空间分解方法(SDM)和高阶空间脉冲响应绘制(HO-SIRR),被广泛用于声场的分析和可听化。本文研究了各种声场合成方法的性能在可听化的一个关键的听音室的背景下。考虑了以下因素对感知空间和音色保真度的影响:渲染框架、到达方向(DOA)估计方法、麦克风阵列结构以及使用具有SDM的专用中心参考麦克风。听力测试将合成声场与参考双耳渲染条件进行比较。测量几个声学参数,以了解方法之间的客观差异。高质量的压力麦克风提高了SDM框架的音色保真度。此外,SDM和HO-SIRR在空间保真度方面表现出相似性。SDM配置之间的性能变化受到DOA估计方法和麦克风阵列结构的影响。双耳SDM(BSDM)呈现显示影响声音质量的时间伪影。
摘要:Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies the performances of various sound field synthesis methods in the context of the auralization of a critical listening room. The influence on the perceived spatial and timbral fidelity of the following factors is considered: the rendering framework, direction of arrival (DOA) estimation method, microphone array structure, and use of a dedicated center reference microphone with SDM. Listening tests compare the synthesized sound fields to a reference binaural rendering condition. Several acoustic parameters are measured to gain insights into objective differences between methods. A high-quality pressure microphone improves the SDM framework's timbral fidelity. Additionally, SDM and HO-SIRR show similarities in spatial fidelity. Performance variation between SDM configurations is influenced by the DOA estimation method and microphone array construction. The binaural SDM (BSDM) presentations display temporal artifacts impacting sound quality.

【9】 Enhancement of a Text-Independent Speaker Verification System by using  Feature Combination and Parallel-Structure Classifiers
标题:基于特征组合和并行结构分类器的文本无关说话人确认系统改进
链接:https://arxiv.org/abs/2401.15018
作者:Kerlos Atia Abdalmalak,Ascensión Gallardo-Antol'in
备注:None
摘要:说话人确认系统主要包括两个阶段:特征提取和分类。在本文中,我们探讨这两个模块的目的是提高性能的说话人确认系统在嘈杂的条件下。一方面,选择最合适的声学特征是进行鲁棒说话人确认的关键因素。在所提出的系统中使用的声学参数是:梅尔频率倒谱系数(MFCC)、它们的一阶和二阶导数(Δ和Δ-Δ)、巴克频率倒谱系数(BFCC)、感知线性预测(PLP)和相对谱变换-感知线性预测(RASTA-PLP)。在本文中,一个完整的比较不同的组合,以前的功能进行了讨论。另一方面,传统的支持向量机(SVM)分类器的主要缺点是使用通用的传统核函数来计算数据点之间的距离。然而,支持向量机的核函数对其性能有很大的影响。在这项工作中,我们提出了两个基于SVM的分类器与不同的核函数的组合:线性核和高斯径向基函数(RBF)核与逻辑回归(LR)分类器。该组合是通过一个并行结构的方法,其中考虑不同的投票规则,采取最终的决定。结果表明,无论是在干净的语音或存在噪声的SV系统的性能显着改善,通过使用组合特征与组合分类器。最后,为了提高系统在嘈杂的环境中,包括多频带噪声去除技术作为预处理阶段提出。
摘要:Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under noisy conditions. On the one hand, the choice of the most appropriate acoustic features is a crucial factor for performing robust speaker verification. The acoustic parameters used in the proposed system are: Mel Frequency Cepstral Coefficients (MFCC), their first and second derivatives (Deltas and Delta- Deltas), Bark Frequency Cepstral Coefficients (BFCC), Perceptual Linear Predictive (PLP), and Relative Spectral Transform - Perceptual Linear Predictive (RASTA-PLP). In this paper, a complete comparison of different combinations of the previous features is discussed. On the other hand, the major weakness of a conventional Support Vector Machine (SVM) classifier is the use of generic traditional kernel functions to compute the distances among data points. However, the kernel function of an SVM has great influence on its performance. In this work, we propose the combination of two SVM-based classifiers with different kernel functions: Linear kernel and Gaussian Radial Basis Function (RBF) kernel with a Logistic Regression (LR) classifier. The combination is carried out by means of a parallel structure approach, in which different voting rules to take the final decision are considered. Results show that significant improvement in the performance of the SV system is achieved by using the combined features with the combined classifiers either with clean speech or in the presence of noise. Finally, to enhance the system more in noisy environments, the inclusion of the multiband noise removal technique as a preprocessing stage is proposed.


【10】 Acoustic characterization of speech rhythm: going beyond metrics with  recurrent neural networks
标题:语音节奏的声学表征:超越递归神经网络的度量
链接:https://arxiv.org/abs/2401.14416
作者:François Deloche,Laurent Bonnasse-Gahot,Judit Gervain
备注:15 pages, 7 figures
摘要:长期以来,语言一直是根据其感知的节奏属性来描述的。相关的类型学在心理语言学中很有意义,因为它们部分预测了新生儿区分语言的能力,并提供了成年听众如何处理非母语的见解。尽管相对成功的节奏度量在支持语言节奏类的存在,定量研究尚未捕捉到完整的复杂性与语音节奏的时间间隔。我们认为,深度学习提供了一种强大的模式识别方法,以推进语音节奏的声学基础的表征。为了探索这一假设,我们在21种语言的大型语音记录数据库上训练了一个中型递归神经网络进行语言识别任务。该网络可以访问幅度包络和标识有声片段的变量,假设该信号将不好地传达语音信息,但保留韵律特征。在40%的案例中,该网络能够识别10秒录音的语言,并且在三分之二的案例中,该语言位于前三名。可视化方法表明,从网络激活建立的表示与语音节奏类型学是一致的,虽然由此产生的地图比重音和音节定时语言之间的两个单独的集群更复杂。我们通过识别网络激活与已知语音节奏度量之间的相关性来进一步分析该模型。这些发现说明了深度学习工具通过识别和探索语言相关的声学特征空间来促进我们对语音节奏的理解的潜力。
摘要:Languages have long been described according to their perceived rhythmic attributes. The associated typologies are of interest in psycholinguistics as they partly predict newborns' abilities to discriminate between languages and provide insights into how adult listeners process non-native languages. Despite the relative success of rhythm metrics in supporting the existence of linguistic rhythmic classes, quantitative studies have yet to capture the full complexity of temporal regularities associated with speech rhythm. We argue that deep learning offers a powerful pattern-recognition approach to advance the characterization of the acoustic bases of speech rhythm. To explore this hypothesis, we trained a medium-sized recurrent neural network on a language identification task over a large database of speech recordings in 21 languages. The network had access to the amplitude envelopes and a variable identifying the voiced segments, assuming that this signal would poorly convey phonetic information but preserve prosodic features. The network was able to identify the language of 10-second recordings in 40% of the cases, and the language was in the top-3 guesses in two-thirds of the cases. Visualization methods show that representations built from the network activations are consistent with speech rhythm typologies, although the resulting maps are more complex than two separated clusters between stress and syllable-timed languages. We further analyzed the model by identifying correlations between network activations and known speech rhythm metrics. The findings illustrate the potential of deep learning tools to advance our understanding of speech rhythm through the identification and exploration of linguistically relevant acoustic feature spaces.


eess.AS音频处理
【1】 Spatial Analysis and Synthesis Methods: Subjective and Objective  Evaluations Using Various Microphone Arrays in the Auralization of a Critical  Listening Room
标题:空间分析和综合方法:在关键听音室可听化中使用不同麦克风阵列的主客观评价
链接:https://arxiv.org/abs/2401.15023
作者:Alan Pawlak,Hyunkook Lee,Aki Mäkivirta,Thomas Lund
备注:13 pages, 6 figures
摘要:参数声场合成方法,如空间分解方法(SDM)和高阶空间脉冲响应绘制(HO-SIRR),被广泛用于声场的分析和可听化。本文研究了各种声场合成方法的性能在可听化的一个关键的听音室的背景下。考虑了以下因素对感知空间和音色保真度的影响:渲染框架、到达方向(DOA)估计方法、麦克风阵列结构以及使用具有SDM的专用中心参考麦克风。听力测试将合成声场与参考双耳渲染条件进行比较。测量几个声学参数,以了解方法之间的客观差异。高质量的压力麦克风提高了SDM框架的音色保真度。此外,SDM和HO-SIRR在空间保真度方面表现出相似性。SDM配置之间的性能变化受到DOA估计方法和麦克风阵列结构的影响。双耳SDM(BSDM)呈现显示影响声音质量的时间伪影。
摘要:Parametric sound field synthesis methods, such as the Spatial Decomposition Method (SDM) and Higher-Order Spatial Impulse Response Rendering (HO-SIRR), are widely used for the analysis and auralization of sound fields. This paper studies the performances of various sound field synthesis methods in the context of the auralization of a critical listening room. The influence on the perceived spatial and timbral fidelity of the following factors is considered: the rendering framework, direction of arrival (DOA) estimation method, microphone array structure, and use of a dedicated center reference microphone with SDM. Listening tests compare the synthesized sound fields to a reference binaural rendering condition. Several acoustic parameters are measured to gain insights into objective differences between methods. A high-quality pressure microphone improves the SDM framework's timbral fidelity. Additionally, SDM and HO-SIRR show similarities in spatial fidelity. Performance variation between SDM configurations is influenced by the DOA estimation method and microphone array construction. The binaural SDM (BSDM) presentations display temporal artifacts impacting sound quality.


【2】 Enhancement of a Text-Independent Speaker Verification System by using  Feature Combination and Parallel-Structure Classifiers
标题:基于特征组合和并行结构分类器的文本无关说话人确认系统改进
链接:https://arxiv.org/abs/2401.15018
作者:Kerlos Atia Abdalmalak,Ascensión Gallardo-Antol'in
备注:None
摘要:说话人确认系统主要包括两个阶段:特征提取和分类。在本文中,我们探讨这两个模块的目的是提高性能的说话人确认系统在嘈杂的条件下。一方面,选择最合适的声学特征是进行鲁棒说话人确认的关键因素。在所提出的系统中使用的声学参数是:梅尔频率倒谱系数(MFCC)、它们的一阶和二阶导数(Δ和Δ-Δ)、巴克频率倒谱系数(BFCC)、感知线性预测(PLP)和相对谱变换-感知线性预测(RASTA-PLP)。在本文中,一个完整的比较不同的组合,以前的功能进行了讨论。另一方面,传统的支持向量机(SVM)分类器的主要缺点是使用通用的传统核函数来计算数据点之间的距离。然而,支持向量机的核函数对其性能有很大的影响。在这项工作中,我们提出了两个基于SVM的分类器与不同的核函数的组合:线性核和高斯径向基函数(RBF)核与逻辑回归(LR)分类器。该组合是通过一个并行结构的方法,其中考虑不同的投票规则,采取最终的决定。结果表明,无论是在干净的语音或存在噪声的SV系统的性能显着改善,通过使用组合特征与组合分类器。最后,为了提高系统在嘈杂的环境中,包括多频带噪声去除技术作为预处理阶段提出。
摘要:Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under noisy conditions. On the one hand, the choice of the most appropriate acoustic features is a crucial factor for performing robust speaker verification. The acoustic parameters used in the proposed system are: Mel Frequency Cepstral Coefficients (MFCC), their first and second derivatives (Deltas and Delta- Deltas), Bark Frequency Cepstral Coefficients (BFCC), Perceptual Linear Predictive (PLP), and Relative Spectral Transform - Perceptual Linear Predictive (RASTA-PLP). In this paper, a complete comparison of different combinations of the previous features is discussed. On the other hand, the major weakness of a conventional Support Vector Machine (SVM) classifier is the use of generic traditional kernel functions to compute the distances among data points. However, the kernel function of an SVM has great influence on its performance. In this work, we propose the combination of two SVM-based classifiers with different kernel functions: Linear kernel and Gaussian Radial Basis Function (RBF) kernel with a Logistic Regression (LR) classifier. The combination is carried out by means of a parallel structure approach, in which different voting rules to take the final decision are considered. Results show that significant improvement in the performance of the SV system is achieved by using the combined features with the combined classifiers either with clean speech or in the presence of noise. Finally, to enhance the system more in noisy environments, the inclusion of the multiband noise removal technique as a preprocessing stage is proposed.

【3】 Acoustic characterization of speech rhythm: going beyond metrics with  recurrent neural networks
标题:语音节奏的声学表征:用递归神经网络超越度量
链接:https://arxiv.org/abs/2401.14416
作者:François Deloche,Laurent Bonnasse-Gahot,Judit Gervain
备注:15 pages, 7 figures
摘要:长期以来,语言一直是根据其感知的节奏属性来描述的。相关的类型学在心理语言学中很有意义,因为它们部分预测了新生儿区分语言的能力,并提供了成年听众如何处理非母语的见解。尽管相对成功的节奏度量在支持语言节奏类的存在,定量研究尚未捕捉到完整的复杂性与语音节奏的时间间隔。我们认为,深度学习提供了一种强大的模式识别方法,以推进语音节奏的声学基础的表征。为了探索这一假设,我们在21种语言的大型语音记录数据库上训练了一个中型递归神经网络进行语言识别任务。该网络可以访问幅度包络和标识有声片段的变量,假设该信号将不好地传达语音信息,但保留韵律特征。在40%的案例中,该网络能够识别10秒录音的语言,并且在三分之二的案例中,该语言位于前三名。可视化方法表明,从网络激活建立的表示与语音节奏类型学是一致的,虽然由此产生的地图比重音和音节定时语言之间的两个单独的集群更复杂。我们通过识别网络激活与已知语音节奏度量之间的相关性来进一步分析该模型。这些发现说明了深度学习工具通过识别和探索语言相关的声学特征空间来促进我们对语音节奏的理解的潜力。
摘要:Languages have long been described according to their perceived rhythmic attributes. The associated typologies are of interest in psycholinguistics as they partly predict newborns' abilities to discriminate between languages and provide insights into how adult listeners process non-native languages. Despite the relative success of rhythm metrics in supporting the existence of linguistic rhythmic classes, quantitative studies have yet to capture the full complexity of temporal regularities associated with speech rhythm. We argue that deep learning offers a powerful pattern-recognition approach to advance the characterization of the acoustic bases of speech rhythm. To explore this hypothesis, we trained a medium-sized recurrent neural network on a language identification task over a large database of speech recordings in 21 languages. The network had access to the amplitude envelopes and a variable identifying the voiced segments, assuming that this signal would poorly convey phonetic information but preserve prosodic features. The network was able to identify the language of 10-second recordings in 40% of the cases, and the language was in the top-3 guesses in two-thirds of the cases. Visualization methods show that representations built from the network activations are consistent with speech rhythm typologies, although the resulting maps are more complex than two separated clusters between stress and syllable-timed languages. We further analyzed the model by identifying correlations between network activations and known speech rhythm metrics. The findings illustrate the potential of deep learning tools to advance our understanding of speech rhythm through the identification and exploration of linguistically relevant acoustic feature spaces.

【4】 Revisiting proximity effect using broadband signals
标题:使用宽带信号重温邻近效应
链接:https://arxiv.org/abs/2401.14410
作者:Laurent Millot,Mohammed Elliq,Manuel Lopes,Gérard Pelé,Dominique Lambert
备注:None
摘要:主要研究邻近效应的实验。粉红噪声和音乐被用作刺激,组合吉他放大器作为源来测试几个麦克风:全向和定向。我们绘制轴内水平和光谱平衡作为x的函数,到源的距离。全向麦克风的邻近效应。轴内水平曲线表明,1/x定律似乎不太有效。频谱平衡的演变取决于麦克风,而且对刺激:更大的下降与粉红噪声的低频;更大的增加与音乐的其他频率。对于一个裸扬声器,我们发现类似的轴内电平曲线下和以上的截止频率,并提出了一个解释。聆听均衡的音乐录音将有助于证明被测麦克风的邻近效应。论文7106在维也纳音频工程学会第122届会议上发表,2007年
摘要:Experiments studying mainly proximity effect are presented. Pink noise and music were used as stimuli and a combo guitar amplifier as source to test several microphones: omnidirectional and directional. We plot in-axis levels and spectral balances as functions of x, the distance to the source. Proximity effect was found for omnidirectional microphones. In-axis level curves show that 1/x law seems poorly valid. Spectral balance evolutions depend on microphones and moreover on stimuli: bigger decreases of low frequencies with pink noise; larger increases of other frequencies with music. For a naked loudspeaker, we found similar in-axis level curves under and above the cut-off frequency and propose an explanation. Listening equalized music recordings will help to demonstrate proximity effect for tested microphones.Paper 7106 presented at the 122th Convention of the Audio Engineering Society, Wien, 2007


【5】 Robust Dual-Modal Speech Keyword Spotting for XR Headsets
标题:面向XR耳机的稳健双模式语音关键词检测
链接:https://arxiv.org/abs/2401.14978
作者:Zhuojiang Cai,Yuhan Ma,Feng Lu
备注:Accepted to IEEE VR 2024
摘要:虽然语音交互在延展实境(XR)领域内得到广泛应用,但是传统的有声语音关键词发现系统继续应对巨大的挑战,包括在嘈杂环境中的次优性能、在需要静音的情况下的不切实际性以及当其他人在附近说话时对无意激活的敏感性。然而,这些挑战可以通过具有成本效益的语音和嘴唇运动信息的融合来克服。因此,我们提出了一种新的人声回声双模态关键字定位系统设计的XR耳机。我们设计了两种不同的模态融合方法,并进行实验来测试系统在不同场景下的性能。结果表明,我们的双模态系统不仅始终优于其单模态同行,在典型和嘈杂的环境中表现出更高的精度,但也擅长准确地识别无声的话语。此外,我们已经成功地将该系统应用于实时演示,取得了可喜的成果。该代码可在https://github.com/caizhuojiang/VE-KWS上获得。
摘要:While speech interaction finds widespread utility within the Extended Reality (XR) domain, conventional vocal speech keyword spotting systems continue to grapple with formidable challenges, including suboptimal performance in noisy environments, impracticality in situations requiring silence, and susceptibility to inadvertent activations when others speak nearby. These challenges, however, can potentially be surmounted through the cost-effective fusion of voice and lip movement information. Consequently, we propose a novel vocal-echoic dual-modal keyword spotting system designed for XR headsets. We devise two different modal fusion approches and conduct experiments to test the system's performance across diverse scenarios. The results show that our dual-modal system not only consistently outperforms its single-modal counterparts, demonstrating higher precision in both typical and noisy environments, but also excels in accurately identifying silent utterances. Furthermore, we have successfully applied the system in real-time demonstrations, achieving promising results. The code is available at https://github.com/caizhuojiang/VE-KWS.

【6】 Comparison of parameters of vowel sounds of russian and english  languages
标题:俄语和英语元音参数的比较
链接:https://arxiv.org/abs/2401.14890
作者:V. I. Fedoseev,A. A. Konev,A. Yu. Yakimuk
备注:7 pages, 1 figures, 3 tables
摘要:在多语言语音识别系统中,经常会出现这样一种情况,即事先不知道语言,但信号已经被接收并正在处理。对于这种情况,需要某种通用模型,它将能够响应语音差异,并根据它们正确地识别所需语言中的语音。为了建立这样一个模型,有必要设置语音参数的值,然后比较相似的声音,建立显着差异。
摘要:In multilingual speech recognition systems, a situation can often arise when the language is not known in advance, but the signal has already been received and is being processed. For such cases, some generalized model is needed that will be able to respond to phonetic differences and, depending on them, correctly recog-nize speech in the desired language. To build such a model, it is necessary to set the values of phonetic parameters, and then compare similar sounds, establishing significant differences.

【7】 Expressivity-aware Music Performance Retrieval using Mid-level  Perceptual Features and Emotion Word Embeddings
标题:基于中层感知特征和情感词嵌入的表现力感知音乐表演检索
链接:https://arxiv.org/abs/2401.14826
作者:Shreyan Chowdhury,Gerhard Widmer
备注:Presented at FIRE 2023 (Forum for Information Retrieval Evaluation) conference, Goa, India
摘要:本文探讨了跨模态音乐检索的一个具体子任务。我们认为,检索一个性能或再现的音乐作品的基础上,其风格的描述,表达的特点,或从一组不同的性能相同的作品的情感的微妙的任务。我们观察到,一个通用的跨模态系统训练学习一个共同的文本音频嵌入空间不会产生最佳的结果,为这项任务。通过引入两个变化-文本编码器和音频编码器各一个-我们在钢琴演奏和相关自由文本描述的数据集上展示了改进的性能。在文本方面,我们使用情感丰富的词嵌入(EWE),在音频方面,我们提取中级感知特征,而不是通用的音频嵌入。我们的研究结果突出了从音乐中学习的中级感知特征和从情感标记文本中学习的情感丰富的词嵌入在跨模态设置中捕获音乐表达的有效性。此外,我们的可解释的中级功能提供了一个路线,在检索和下游推荐过程中引入可解释性。
摘要:This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a set of different performances of the same piece. We observe that a general purpose cross-modal system trained to learn a common text-audio embedding space does not yield optimal results for this task. By introducing two changes -- one each to the text encoder and the audio encoder -- we demonstrate improved performance on a dataset of piano performances and associated free-text descriptions. On the text side, we use emotion-enriched word embeddings (EWE) and on the audio side, we extract mid-level perceptual features instead of generic audio embeddings. Our results highlight the effectiveness of mid-level perceptual features learnt from music and emotion enriched word embeddings learnt from emotion-labelled text in capturing musical expression in a cross-modal setting. Additionally, our interpretable mid-level features provide a route for introducing explainability in the retrieval and downstream recommendation processes.


【8】 Turn-taking and Backchannel Prediction with Acoustic and Large Language  Model Fusion
标题:基于声学和大语言模型融合的话轮转换和回声预测
链接:https://arxiv.org/abs/2401.14717
作者:Jinhan Wang,Long Chen,Aparna Khare,Anirudh Raju,Pranav Dheram,Di He,Minhua Wu,Andreas Stolcke,Venkatesh Ravichandran
备注:To appear in IEEE ICASSP 2024
摘要:我们提出了一种方法,通过融合神经声学模型与大语言模型(LLM),连续预测口语对话中的话轮转换和反向通道位置。在Switchboard人-人对话数据集上的实验表明,我们的方法始终优于单一模态的基线模型。我们还开发了一种新的多任务指令微调策略,以进一步受益于LLM编码的知识,用于理解任务和会话上下文,从而实现额外的改进。我们的方法展示了LLM和声学模型相结合的潜力,可以在人类和支持语音的AI代理之间进行更自然和对话式的交互。
摘要:We propose an approach for continuous prediction of turn-taking and backchanneling locations in spoken dialogue by fusing a neural acoustic model with a large language model (LLM). Experiments on the Switchboard human-human conversation dataset demonstrate that our approach consistently outperforms the baseline models with single modality. We also develop a novel multi-task instruction fine-tuning strategy to further benefit from LLM-encoded knowledge for understanding the tasks and conversational contexts, leading to additional improvements. Our approach demonstrates the potential of combined LLMs and acoustic models for a more natural and conversational interaction between humans and speech-enabled AI agents.

【9】 UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit  Normalization
标题:UNIT-DSR:基于语音单元归一化的动态语音重建系统
链接:https://arxiv.org/abs/2401.14664
作者:Yuejiao Wang,Xixin Wu,Disong Wang,Lingwei Meng,Helen Meng
备注:Accepted to ICASSP 2024
摘要:构音障碍语音重建(DSR)系统旨在自动将构音障碍语音转换为正常发音语音。这项技术简化了与受神经运动障碍影响的说话者的沟通,并增强了他们的社会包容性。与基于GAN(生成对抗网络)的方法相比,基于NED(神经编码器-解码器)的系统显着提高了重建语音的可理解性,但该方法仍然受到级联管道和内容编码器辅助任务导致的训练效率低下的限制,这反过来又可能影响重建的质量。受自监督语音表示学习和离散语音单元的启发,我们提出了一个Unit-DSR系统,该系统利用HuBERT强大的域自适应能力来提高训练效率,并利用语音单元来约束离散语言空间中的构音障碍内容恢复。与NED方法相比,Unit-DSR系统仅由语音单元归一化器和Unit HiFi-GAN声码器组成,这是相当简单的,没有级联的子模块或辅助任务。UASpeech语料库上的结果表明,Unit-DSR在内容恢复方面优于竞争基线,与原始构音障碍语音相比,相对平均单词错误率降低了28.2%,并且对速度扰动和噪声具有鲁棒性。
摘要:Dysarthric speech reconstruction (DSR) systems aim to automatically convert dysarthric speech into normal-sounding speech. The technology eases communication with speakers affected by the neuromotor disorder and enhances their social inclusion. NED-based (Neural Encoder-Decoder) systems have significantly improved the intelligibility of the reconstructed speech as compared with GAN-based (Generative Adversarial Network) approaches, but the approach is still limited by training inefficiency caused by the cascaded pipeline and auxiliary tasks of the content encoder, which may in turn affect the quality of reconstruction. Inspired by self-supervised speech representation learning and discrete speech units, we propose a Unit-DSR system, which harnesses the powerful domain-adaptation capacity of HuBERT for training efficiency improvement and utilizes speech units to constrain the dysarthric content restoration in a discrete linguistic space. Compared with NED approaches, the Unit-DSR system only consists of a speech unit normalizer and a Unit HiFi-GAN vocoder, which is considerably simpler without cascaded sub-modules or auxiliary tasks. Results on the UASpeech corpus indicate that Unit-DSR outperforms competitive baselines in terms of content restoration, reaching a 28.2% relative average word error rate reduction when compared to original dysarthric speech, and shows robustness against speed perturbation and noise.

【10】 Exploring Musical Roots: Applying Audio Embeddings to Empower Influence  Attribution for a Generative Music Model
标题:探索音乐根源:应用音频嵌入增强生成性音乐模型的影响力归因
链接:https://arxiv.org/abs/2401.14542
作者:Julia Barnett,Hugo Flores Garcia,Bryan Pardo
备注:14 pages + references. Under conference review
摘要:每个艺术家都有一个创作过程,从以前的艺术家和他们的作品中汲取灵感。今天,“灵感”已经被生成音乐模型自动化。这些模型的黑箱性质掩盖了影响其创造性产出的作品的身份。因此,用户可能会无意中盗用、误用或复制现有艺术家的作品。我们建立了一种可复制的方法,以系统地识别类似的音乐音频片段,这种方法有助于理解训练数据的属性。我们的方法的一个关键方面是利用一个有效的音乐音频相似性度量。我们比较了应用CLMR和CLAP嵌入相似性测量的效果,在一组用于训练VampNet的500万个音频片段,最近的开源生成音乐模型。我们用人类听力研究验证了这种方法。我们还探索了音频示例的修改(例如,音调偏移、时间拉伸、背景噪声)对相似性测量有影响。这项工作是将自动影响归因纳入生成建模的基础,这有望让模型创建者和用户从无知的挪用转向知情的创建。本文附带的音频样本可在https://tinyurl.com/exploring-musical-roots上获得。
摘要:Every artist has a creative process that draws inspiration from previous artists and their works. Today, "inspiration" has been automated by generative music models. The black box nature of these models obscures the identity of the works that influence their creative output. As a result, users may inadvertently appropriate, misuse, or copy existing artists' works. We establish a replicable methodology to systematically identify similar pieces of music audio in a manner that is useful for understanding training data attribution. A key aspect of our approach is to harness an effective music audio similarity measure. We compare the effect of applying CLMR and CLAP embeddings to similarity measurement in a set of 5 million audio clips used to train VampNet, a recent open source generative music model. We validate this approach with a human listening study. We also explore the effect that modifications of an audio example (e.g., pitch shifting, time stretching, background noise) have on similarity measurements. This work is foundational to incorporating automated influence attribution into generative modeling, which promises to let model creators and users move from ignorant appropriation to informed creation. Audio samples that accompany this paper are available at https://tinyurl.com/exploring-musical-roots.

【11】 ICASSP 2024 Speech Signal Improvement Challenge
标题:ICASSP 2024语音信号改善挑战赛
链接:https://arxiv.org/abs/2401.14444
作者:Nicolae Catalin Ristea,Ando Saabas,Ross Cutler,Babak Naderi,Sebastian Braun,Solomiya Branets
摘要:ICASSP 2024语音信号改善大挑战赛旨在促进提高通信系统中语音信号质量领域的研究。这标志着我们的第二个挑战,建立在上一届ICASSP 2023大挑战赛的成功基础上。我们通过引入数据集合成器来增强竞争,使所有参赛团队能够从更高的基线开始,这是我们扩展的P.804测试的客观指标,2023测试集的成绩单,我们还添加了单词准确性(WAcc)作为指标。我们评估了13个系统的实时跟踪和11个系统的非实时跟踪使用主观的P.804和客观的字的准确性指标。
摘要:The ICASSP 2024 Speech Signal Improvement Grand Challenge is intended to stimulate research in the area of improving the speech signal quality in communication systems. This marks our second challenge, building upon the success from the previous ICASSP 2023 Grand Challenge. We enhance the competition by introducing a dataset synthesizer, enabling all participating teams to start at a higher baseline, an objective metric for our extended P.804 tests, transcripts for the 2023 test set, and we add Word Accuracy (WAcc) as a metric. We evaluate a total of 13 systems in the real-time track and 11 systems in the non-real-time track using both subjective P.804 and objective Word Accuracy metrics.
机器翻译由腾讯交互翻译提供,仅供参考