今天跟大家分享一篇语音相关的论文合集:cs.SD语音6篇,eess.AS音频处理6篇。
【1】 Extract fundamental frequency based on CNN combined with PYIN
标题:基于CNN结合PYIN的基频提取
链接:https://arxiv.org/abs/2208.08354
作者:Ruowei Xing,Shengchen Li摘要:本文主要研究多基频信号的提取(多个F0)基于PYIN,用于提取基频的算法(F0)的单声道音乐,并训练卷积神经网络(CNN)模型,其中产生输入信号的音调显著函数以估计倍数F0。本文讨论了这两种算法的实现及其各自的优缺点,分析了这两种算法的不同性能,为了结合这两种算法的优点,本文采用了PYIN算法来补充从训练好的CNN模型中提取的F0,并根据提取的F0曲线的平坦程度来评价模型的性能。实验结果表明,该组合模型在单声道和多声道音乐中的F0提取效果均优于原算法。摘要:This paper refers to the extraction of multiple fundamental frequencies (multiple F0) based on PYIN, an algorithm for extracting the fundamental frequency (F0) of monophonic music, and a trained convolutional neural networks (CNN) model, where a pitch salience function of the input signal is produced to estimate the multiple F0. The implementation of these two algorithms and their corresponding advantages and disadvantages are discussed in this article. Analysing the different performance of these two methods, PYIN is applied to supplement the F0 extracted from the trained CNN model to combine the advantages of these two algorithms. For evaluation, four pieces played by two violins are used, and the performance of the models are evaluated accoring to the flatness of the F0 curve extracted. The result shows the combined model outperforms the original algorithms when extracting F0 from monophonic music and polyphonic music.
【2】 Domestic sound event detection by shift consistency mean-teacher training and adversarial domain adaptation
标题:通过移位一致性均值—教师训练和对抗域自适应的家庭声音事件检测
链接:https://arxiv.org/abs/2208.08131
作者:Fang-Ching Chen,Kuan-Dar Chen,Yi-Wen Liu摘要:半监督学习和领域自适应技术在家庭声音事件检测领域受到了越来越多的关注,这得益于大量未标记数据的可用性和相对容易生成合成强标记数据的优势.在先前的工作中,设计了几种半监督学习策略来提高均值—教师模型的性能,即这些策略包括移位一致性训练(shift consistency training(SCT),插值一致性训练(ICT)和伪标签。然而,对抗性领域适应在本研究中,我们通过实验发现ICT往往会将真实数据和合成数据在t-SNE图中的分布拉开,因此ICT被放弃,而SCT,在DCASE 2020 task 4数据集上,该系统的F1得分达到47.2%,比前人的研究结果提高了2.1个百分点.摘要:Semi-supervised learning and domain adaptation techniques have drawn increasing attention in the field of domestic sound event detection thanks to the availability of large amounts of unlabeled data and the relative ease to generate synthetic strongly-labeled data. In a previous work, several semi-supervised learning strategies were designed to boost the performance of a mean-teacher model. Namely, these strategies include shift consistency training (SCT), interpolation consistency training (ICT), and pseudo-labeling. However, adversarial domain adaptation (ADA) did not seem to improve the event detection accuracy further when we attempt to compensate for the domain gap between synthetic and real data. In this research, we empirically found that ICT tends to pull apart the distributions of synthetic and real data in t-SNE plots. Therefore, ICT is abandoned while SCT, in contrast, is applied to train both the student and the teacher models. With these modifications, the system successfully integrates with an ADA network, and we achieve 47.2% in the F1 score on the DCASE 2020 task 4 dataset, which is 2.1% higher than what was reported in the previous work.
【3】 A Hybrid SFANC-FxNLMS Algorithm for Active Noise Control based on Deep Learning
标题:基于深度学习的SFANC-FxNLMS混合有源噪声控制算法
链接:https://arxiv.org/abs/2208.08082
作者:Zhengding Luo,Dongyuan Shi,Woon-Seng Gan摘要:选择性固定滤波器有源噪声控制(SFANC)方法针对各种噪声类型选择最佳的预训练控制滤波器,可以获得较快的响应时间,但由于滤波器选择不准确和缺乏适应性,可能会导致较大的稳态误差.相比之下,滤波-X归一化最小均方方法(filtered-X normalized least mean square,SFANC)可以获得较好的稳态性能.(FxNLMS)算法通过自适应优化可以获得较低的稳态误差,但其收敛速度慢对动态噪声的衰减有不利影响,本文提出了一种SFANC-FxNLMS混合算法,克服了自适应算法收敛速度慢的缺点,并提供了比SFANC方法更好的降噪效果.设计了一种能自动为每一帧原始噪声选择最合适的控制滤波器的一维CNN神经网络,FxNLMS算法继续以采样率更新所选择的预训练控制滤波器的系数,实验结果表明,混合SFANC-FxNLMS算法具有响应速度快、降噪误差小、鲁棒性强等优点。摘要:The selective fixed-filter active noise control (SFANC) method selecting the best pre-trained control filters for various types of noise can achieve a fast response time. However, it may lead to large steady-state errors due to inaccurate filter selection and the lack of adaptability. In comparison, the filtered-X normalized least-mean-square (FxNLMS) algorithm can obtain lower steady-state errors through adaptive optimization. Nonetheless, its slow convergence has a detrimental effect on dynamic noise attenuation. Therefore, this paper proposes a hybrid SFANC-FxNLMS approach to overcome the adaptive algorithm's slow convergence and provide a better noise reduction level than the SFANC method. A lightweight one-dimensional convolutional neural network (1D CNN) is designed to automatically select the most suitable pre-trained control filter for each frame of the primary noise. Meanwhile, the FxNLMS algorithm continues to update the coefficients of the chosen pre-trained control filter at the sampling rate. Owing to the effective combination of the two algorithms, experimental results show that the hybrid SFANC-FxNLMS algorithm can achieve a rapid response time, a low noise reduction error, and a high degree of robustness.
【4】 The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines
标题:会话短句说话人对话(CSSD)任务:数据集、评估指标和基线
链接:https://arxiv.org/abs/2208.08042
作者:Gaofeng Cheng,Yifan Chen,Runyan Yang,Qingxuan Li,Zehui Yang,Lingxuan Ye,Pengyuan Zhang,Qingqing Zhang,Lei Xie,Yanmin Qian,Kong Aik Lee,Yonghong Yan备注:arXiv admin note: text overlap with arXiv:2203.16844摘要:会话场景是语音处理技术中最重要也是最具挑战性的场景之一,因为会话中的人们以随意的风格相互响应,检测会话中每个人的语音活动对于诸如自然语言处理、机器翻译人们把“谁在何时说话”的检测技术称为说话人日记化(SD)。传统上,日记化错误率(DER)一直被用作SD系统的标准评价指标,DER未能充分重视短会话短语,这些短语虽然短但在语义层面上很重要,而且在语音界还没有一个经过仔细和准确的人工标注的适合评估会话SD技术的测试数据集.本文提出了一种新的会话SD技术,我们设计并描述了会话短语说话人对话在数据集方面,尽管之前有开源的180小时会话式MagicData-RAMC数据集,我们为CSSD任务准备了一个20小时的会话语音测试数据集,我们设计了一种新的会话DER(CDER)评价指标,在话语层面上计算SD的准确率,在基线方面,我们采用了一种常用的方法:变分贝叶斯HMM x向量系统,作为CSSD任务的基线。我们的评估指标可在www.example.com上公开获得https://github.com/SpeechClub/CDER_Metric。摘要:The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each person in a conversation is vital to downstream tasks, like natural language processing, machine translation, etc. People refer to the detection technology of "who speak when" as speaker diarization (SD). Traditionally, diarization error rate (DER) has been used as the standard evaluation metric of SD systems for a long time. However, DER fails to give enough importance to short conversational phrases, which are short but important on the semantic level. Also, a carefully and accurately manually-annotated testing dataset suitable for evaluating the conversational SD technologies is still unavailable in the speech community. In this paper, we design and describe the Conversational Short-phrases Speaker Diarization (CSSD) task, which consists of training and testing datasets, evaluation metric and baselines. In the dataset aspect, despite the previously open-sourced 180-hour conversational MagicData-RAMC dataset, we prepare an individual 20-hour conversational speech test dataset with carefully and artificially verified speakers timestamps annotations for the CSSD task. In the metric aspect, we design the new conversational DER (CDER) evaluation metric, which calculates the SD accuracy at the utterance level. In the baseline aspect, we adopt a commonly used method: Variational Bayes HMM x-vector system, as the baseline of the CSSD task. Our evaluation metric is publicly available at https://github.com/SpeechClub/CDER_Metric.
【5】 Enhancing Audio Perception of Music By AI Picked Room Acoustics
标题:通过人工智能拾取房间声学增强音乐的音频感知
链接:https://arxiv.org/abs/2208.07994
作者:Prateek Verma,Jonathan Berger备注:24th International Congress on Acoustics, Gyeongju, South Korea摘要:我们听到的每个声音都是连续卷积运算的结果(例如,房间声学、麦克风特性、乐器本身的共振特性,更不用说声音再现系统的特性和限制)。在这项工作中,我们试图确定使用人工智能演奏特定乐曲的最佳房间。此外,我们使用室内声学作为增强给定声音的感知质量的一种方式。2历史上,会议室(特别是教堂和音乐厅)被设计成举办和服务于特定的音乐功能。在某些情况下,建筑的声学质量增强了在那里表演的音乐。我们试着模仿它,作为第一步,通过指定与产生特定音乐的增强的声音质量相关的房间脉冲响应,首先训练卷积结构,以接收音频样本,并以大约78%的准确度模仿专家对各种乐器系列和音符的评价,以获得感知质量。这为我们提供了一个针对任何音频样本的评分函数,该函数可以自动对音符的感知愉悦度进行评级。现在,通过模拟各种房间、材料等的大约60,000个合成脉冲响应的库,我们使用简单的卷积运算,转换声音,使其听起来像是在特定的房间里播放的一样。感知评估器用于对音乐声音进行排序,并产生播放声音的“最佳房间或音乐厅”。作为一种副产品,它还可以使用房间声学将质量差的声音转换为“好”的声音。摘要:Every sound that we hear is the result of successive convolutional operations (e.g. room acoustics, microphone characteristics, resonant properties of the instrument itself, not to mention characteristics and limitations of the sound reproduction system). In this work we seek to determine the best room in which to perform a particular piece using AI. Additionally, we use room acoustics as a way to enhance the perceptual qualities of a given sound. Historically, rooms (particularly Churches and concert halls) were designed to host and serve specific musical functions. In some cases the architectural acoustical qualities enhanced the music performed there. We try to mimic this, as a first step, by designating room impulse responses that would correlate to producing enhanced sound quality for particular music. A convolutional architecture is first trained to take in an audio sample and mimic the ratings of experts with about 78 % accuracy for various instrument families and notes for perceptual qualities. This gives us a scoring function for any audio sample which can rate the perceptual pleasantness of a note automatically. Now, via a library of about 60,000 synthetic impulse responses mimicking all kinds of room, materials, etc, we use a simple convolution operation, to transform the sound as if it was played in a particular room. The perceptual evaluator is used to rank the musical sounds, and yield the "best room or the concert hall" to play a sound. As a byproduct it can also use room acoustics to turn a poor quality sound into a "good" sound.
【6】 Disentangled Speaker Representation Learning via Mutual Information Minimization
标题:基于互信息最小化的解纠缠说话人表征学习
链接:https://arxiv.org/abs/2208.08012
作者:Sung Hwan Mun,Min Hyun Han,Minchan Kim,Dongjune Lee,Nam Soo Kim备注:7 pages, 4 figures, and 1 table摘要:针对说话人无关特征引起的域失配问题,提出一种基于互信息的显式解纠缠框架,从说话人无关特征中分离出说话人相关特征为了实现最小化说话人相关和说话人无关特征之间的MI的目标,我们采用了对比对数比上界该框架采用3级结构:首先在前端编码器中将输入语音编码为共享初始嵌入,然后在去耦合模块中将共享初始嵌入分解为说话人相关嵌入和说话人无关嵌入,最后在去耦合模块中将共享初始嵌入与说话人相关嵌入进行比较,在最后一个阶段,通过MI最小化进行解纠缠。(FFSVC2022)的实验结果表明,本文提出的框架是有效的,同时,我们使用VoxCeleb数据集对前端编码器进行预训练,然后使用FFSVC 2022数据集对解纠缠框架中的说话人嵌入模型进行微调。实验结果表明,在现有预训练模型上使用解纠缠框架进行微调是有效的,可以进一步提高性能。摘要:Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unravel speaker-relevant features from speaker-unrelated features via mutual information (MI) minimization. To achieve our goal of minimizing MI between speaker-related and speaker-unrelated features, we adopt a contrastive log-ratio upper bound (CLUB), which exploits the upper bound of MI. Our framework is constructed in a 3-stage structure. First, in the front-end encoder, input speech is encoded into shared initial embedding. Next, in the decoupling block, shared initial embedding is split into separate speaker-related and speaker-unrelated embeddings. Finally, disentanglement is conducted by MI minimization in the last stage. Experiments on Far-Field Speaker Verification Challenge 2022 (FFSVC2022) demonstrate that our proposed framework is effective for disentanglement. Also, to utilize domain-unknown datasets containing numerous speakers, we pre-trained the front-end encoder with VoxCeleb datasets. We then fine-tuned the speaker embedding model in the disentanglement framework with FFSVC 2022 dataset. The experimental results show that fine-tuning with a disentanglement framework on a existing pre-trained model is valid and can further improve performance.
【1】 A Hybrid SFANC-FxNLMS Algorithm for Active Noise Control based on Deep Learning
标题:基于深度学习的SFANC-FxNLMS混合有源噪声控制算法
链接:https://arxiv.org/abs/2208.08082
* 与cs.SD语音【3】为同一篇
作者:Zhengding Luo,Dongyuan Shi,Woon-Seng Gan摘要:选择性固定滤波器有源噪声控制(SFANC)方法针对各种噪声类型选择最佳的预训练控制滤波器,可以获得较快的响应时间,但由于滤波器选择不准确和缺乏适应性,可能会导致较大的稳态误差.相比之下,滤波-X归一化最小均方方法(filtered-X normalized least mean square,SFANC)可以获得较好的稳态性能.(FxNLMS)算法通过自适应优化可以获得较低的稳态误差,但其收敛速度慢对动态噪声的衰减有不利影响,本文提出了一种SFANC-FxNLMS混合算法,克服了自适应算法收敛速度慢的缺点,并提供了比SFANC方法更好的降噪效果.设计了一种能自动为每一帧原始噪声选择最合适的控制滤波器的一维CNN神经网络,FxNLMS算法继续以采样率更新所选择的预训练控制滤波器的系数,实验结果表明,混合SFANC-FxNLMS算法具有响应速度快、降噪误差小、鲁棒性强等优点。摘要:The selective fixed-filter active noise control (SFANC) method selecting the best pre-trained control filters for various types of noise can achieve a fast response time. However, it may lead to large steady-state errors due to inaccurate filter selection and the lack of adaptability. In comparison, the filtered-X normalized least-mean-square (FxNLMS) algorithm can obtain lower steady-state errors through adaptive optimization. Nonetheless, its slow convergence has a detrimental effect on dynamic noise attenuation. Therefore, this paper proposes a hybrid SFANC-FxNLMS approach to overcome the adaptive algorithm's slow convergence and provide a better noise reduction level than the SFANC method. A lightweight one-dimensional convolutional neural network (1D CNN) is designed to automatically select the most suitable pre-trained control filter for each frame of the primary noise. Meanwhile, the FxNLMS algorithm continues to update the coefficients of the chosen pre-trained control filter at the sampling rate. Owing to the effective combination of the two algorithms, experimental results show that the hybrid SFANC-FxNLMS algorithm can achieve a rapid response time, a low noise reduction error, and a high degree of robustness.
【2】 Disentangled Speaker Representation Learning via Mutual Information Minimization
标题:基于互信息最小化的解纠缠说话人表征学习
链接:https://arxiv.org/abs/2208.08012
* 与cs.SD语音【6】为同一篇
作者:Sung Hwan Mun,Min Hyun Han,Minchan Kim,Dongjune Lee,Nam Soo Kim备注:7 pages, 4 figures, and 1 table摘要:针对说话人无关特征引起的域失配问题,提出一种基于互信息的显式解纠缠框架,从说话人无关特征中分离出说话人相关特征为了实现最小化说话人相关和说话人无关特征之间的MI的目标,我们采用了对比对数比上界该框架采用3级结构:首先在前端编码器中将输入语音编码为共享初始嵌入,然后在去耦合模块中将共享初始嵌入分解为说话人相关嵌入和说话人无关嵌入,最后在去耦合模块中将共享初始嵌入与说话人相关嵌入进行比较,在最后一个阶段,通过MI最小化进行解纠缠。(FFSVC2022)的实验结果表明,本文提出的框架是有效的,同时,我们使用VoxCeleb数据集对前端编码器进行预训练,然后使用FFSVC 2022数据集对解纠缠框架中的说话人嵌入模型进行微调。实验结果表明,在现有预训练模型上使用解纠缠框架进行微调是有效的,可以进一步提高性能。摘要:Domain mismatch problem caused by speaker-unrelated feature has been a major topic in speaker recognition. In this paper, we propose an explicit disentanglement framework to unravel speaker-relevant features from speaker-unrelated features via mutual information (MI) minimization. To achieve our goal of minimizing MI between speaker-related and speaker-unrelated features, we adopt a contrastive log-ratio upper bound (CLUB), which exploits the upper bound of MI. Our framework is constructed in a 3-stage structure. First, in the front-end encoder, input speech is encoded into shared initial embedding. Next, in the decoupling block, shared initial embedding is split into separate speaker-related and speaker-unrelated embeddings. Finally, disentanglement is conducted by MI minimization in the last stage. Experiments on Far-Field Speaker Verification Challenge 2022 (FFSVC2022) demonstrate that our proposed framework is effective for disentanglement. Also, to utilize domain-unknown datasets containing numerous speakers, we pre-trained the front-end encoder with VoxCeleb datasets. We then fine-tuned the speaker embedding model in the disentanglement framework with FFSVC 2022 dataset. The experimental results show that fine-tuning with a disentanglement framework on a existing pre-trained model is valid and can further improve performance.
【3】 Extract fundamental frequency based on CNN combined with PYIN
标题:基于CNN结合PYIN的基频提取
链接:https://arxiv.org/abs/2208.08354
* 与cs.SD语音【1】为同一篇
作者:Ruowei Xing,Shengchen Li摘要:本文主要研究多基频信号的提取(多个F0)基于PYIN,用于提取基频的算法(F0)的单声道音乐,并训练卷积神经网络(CNN)模型,其中产生输入信号的音调显著函数以估计倍数F0。本文讨论了这两种算法的实现及其各自的优缺点,分析了这两种算法的不同性能,为了结合这两种算法的优点,本文采用了PYIN算法来补充从训练好的CNN模型中提取的F0,并根据提取的F0曲线的平坦程度来评价模型的性能。实验结果表明,该组合模型在单声道和多声道音乐中的F0提取效果均优于原算法。摘要:This paper refers to the extraction of multiple fundamental frequencies (multiple F0) based on PYIN, an algorithm for extracting the fundamental frequency (F0) of monophonic music, and a trained convolutional neural networks (CNN) model, where a pitch salience function of the input signal is produced to estimate the multiple F0. The implementation of these two algorithms and their corresponding advantages and disadvantages are discussed in this article. Analysing the different performance of these two methods, PYIN is applied to supplement the F0 extracted from the trained CNN model to combine the advantages of these two algorithms. For evaluation, four pieces played by two violins are used, and the performance of the models are evaluated accoring to the flatness of the F0 curve extracted. The result shows the combined model outperforms the original algorithms when extracting F0 from monophonic music and polyphonic music.
【4】 Domestic sound event detection by shift consistency mean-teacher training and adversarial domain adaptation
标题:通过移位一致性均值—教师训练和对抗域自适应的家庭声音事件检测
链接:https://arxiv.org/abs/2208.08131
* 与cs.SD语音【2】为同一篇
作者:Fang-Ching Chen,Kuan-Dar Chen,Yi-Wen Liu摘要:半监督学习和领域自适应技术在家庭声音事件检测领域受到了越来越多的关注,这得益于大量未标记数据的可用性和相对容易生成合成强标记数据的优势.在先前的工作中,设计了几种半监督学习策略来提高均值—教师模型的性能,即这些策略包括移位一致性训练(shift consistency training(SCT),插值一致性训练(ICT)和伪标签。然而,对抗性领域适应在本研究中,我们通过实验发现ICT往往会将真实数据和合成数据在t-SNE图中的分布拉开,因此ICT被放弃,而SCT,在DCASE 2020 task 4数据集上,该系统的F1得分达到47.2%,比前人的研究结果提高了2.1个百分点.摘要:Semi-supervised learning and domain adaptation techniques have drawn increasing attention in the field of domestic sound event detection thanks to the availability of large amounts of unlabeled data and the relative ease to generate synthetic strongly-labeled data. In a previous work, several semi-supervised learning strategies were designed to boost the performance of a mean-teacher model. Namely, these strategies include shift consistency training (SCT), interpolation consistency training (ICT), and pseudo-labeling. However, adversarial domain adaptation (ADA) did not seem to improve the event detection accuracy further when we attempt to compensate for the domain gap between synthetic and real data. In this research, we empirically found that ICT tends to pull apart the distributions of synthetic and real data in t-SNE plots. Therefore, ICT is abandoned while SCT, in contrast, is applied to train both the student and the teacher models. With these modifications, the system successfully integrates with an ADA network, and we achieve 47.2% in the F1 score on the DCASE 2020 task 4 dataset, which is 2.1% higher than what was reported in the previous work.
【5】 The Conversational Short-phrase Speaker Diarization (CSSD) Task: Dataset, Evaluation Metric and Baselines
标题:会话短句说话人对话(CSSD)任务:数据集、评估指标和基线
链接:https://arxiv.org/abs/2208.08042
* 与cs.SD语音【4】为同一篇
作者:Gaofeng Cheng,Yifan Chen,Runyan Yang,Qingxuan Li,Zehui Yang,Lingxuan Ye,Pengyuan Zhang,Qingqing Zhang,Lei Xie,Yanmin Qian,Kong Aik Lee,Yonghong Yan备注:arXiv admin note: text overlap with arXiv:2203.16844摘要:会话场景是语音处理技术中最重要也是最具挑战性的场景之一,因为会话中的人们以随意的风格相互响应,检测会话中每个人的语音活动对于诸如自然语言处理、机器翻译人们把“谁在何时说话”的检测技术称为说话人日记化(SD)。传统上,日记化错误率(DER)一直被用作SD系统的标准评价指标,DER未能充分重视短会话短语,这些短语虽然短但在语义层面上很重要,而且在语音界还没有一个经过仔细和准确的人工标注的适合评估会话SD技术的测试数据集.本文提出了一种新的会话SD技术,我们设计并描述了会话短语说话人对话在数据集方面,尽管之前有开源的180小时会话式MagicData-RAMC数据集,我们为CSSD任务准备了一个20小时的会话语音测试数据集,我们设计了一种新的会话DER(CDER)评价指标,在话语层面上计算SD的准确率,在基线方面,我们采用了一种常用的方法:变分贝叶斯HMM x向量系统,作为CSSD任务的基线。我们的评估指标可在www.example.com上公开获得https://github.com/SpeechClub/CDER_Metric。摘要:The conversation scenario is one of the most important and most challenging scenarios for speech processing technologies because people in conversation respond to each other in a casual style. Detecting the speech activities of each person in a conversation is vital to downstream tasks, like natural language processing, machine translation, etc. People refer to the detection technology of "who speak when" as speaker diarization (SD). Traditionally, diarization error rate (DER) has been used as the standard evaluation metric of SD systems for a long time. However, DER fails to give enough importance to short conversational phrases, which are short but important on the semantic level. Also, a carefully and accurately manually-annotated testing dataset suitable for evaluating the conversational SD technologies is still unavailable in the speech community. In this paper, we design and describe the Conversational Short-phrases Speaker Diarization (CSSD) task, which consists of training and testing datasets, evaluation metric and baselines. In the dataset aspect, despite the previously open-sourced 180-hour conversational MagicData-RAMC dataset, we prepare an individual 20-hour conversational speech test dataset with carefully and artificially verified speakers timestamps annotations for the CSSD task. In the metric aspect, we design the new conversational DER (CDER) evaluation metric, which calculates the SD accuracy at the utterance level. In the baseline aspect, we adopt a commonly used method: Variational Bayes HMM x-vector system, as the baseline of the CSSD task. Our evaluation metric is publicly available at https://github.com/SpeechClub/CDER_Metric.
【6】 Enhancing Audio Perception of Music By AI Picked Room Acoustics
标题:通过人工智能拾取房间声学增强音乐的音频感知
链接:https://arxiv.org/abs/2208.07994
* 与cs.SD语音【5】为同一篇
作者:Prateek Verma,Jonathan Berger备注:24th International Congress on Acoustics, Gyeongju, South Korea摘要:我们听到的每个声音都是连续卷积运算的结果(例如,房间声学、麦克风特性、乐器本身的共振特性,更不用说声音再现系统的特性和限制)。在这项工作中,我们试图确定使用人工智能演奏特定乐曲的最佳房间。此外,我们使用室内声学作为增强给定声音的感知质量的一种方式。2历史上,会议室(特别是教堂和音乐厅)被设计成举办和服务于特定的音乐功能。在某些情况下,建筑的声学质量增强了在那里表演的音乐。我们试着模仿它,作为第一步,通过指定与产生特定音乐的增强的声音质量相关的房间脉冲响应,首先训练卷积结构,以接收音频样本,并以大约78%的准确度模仿专家对各种乐器系列和音符的评价,以获得感知质量。这为我们提供了一个针对任何音频样本的评分函数,该函数可以自动对音符的感知愉悦度进行评级。现在,通过模拟各种房间、材料等的大约60,000个合成脉冲响应的库,我们使用简单的卷积运算,转换声音,使其听起来像是在特定的房间里播放的一样。感知评估器用于对音乐声音进行排序,并产生播放声音的“最佳房间或音乐厅”。作为一种副产品,它还可以使用房间声学将质量差的声音转换为“好”的声音。摘要:Every sound that we hear is the result of successive convolutional operations (e.g. room acoustics, microphone characteristics, resonant properties of the instrument itself, not to mention characteristics and limitations of the sound reproduction system). In this work we seek to determine the best room in which to perform a particular piece using AI. Additionally, we use room acoustics as a way to enhance the perceptual qualities of a given sound. Historically, rooms (particularly Churches and concert halls) were designed to host and serve specific musical functions. In some cases the architectural acoustical qualities enhanced the music performed there. We try to mimic this, as a first step, by designating room impulse responses that would correlate to producing enhanced sound quality for particular music. A convolutional architecture is first trained to take in an audio sample and mimic the ratings of experts with about 78 % accuracy for various instrument families and notes for perceptual qualities. This gives us a scoring function for any audio sample which can rate the perceptual pleasantness of a note automatically. Now, via a library of about 60,000 synthetic impulse responses mimicking all kinds of room, materials, etc, we use a simple convolution operation, to transform the sound as if it was played in a particular room. The perceptual evaluator is used to rank the musical sounds, and yield the "best room or the concert hall" to play a sound. As a byproduct it can also use room acoustics to turn a poor quality sound into a "good" sound.
机器翻译,仅供参考