今日论文合集:cs.SD语音6篇,eess.AS音频处理7篇。本文经arXiv每日学术速递授权转载
【1】HiddenSpeaker: Generate Imperceptible Unlearnable Audios for Speaker Verification System
标题:HiddenSpeaker:为说话人验证系统生成难以感知的、不可学习的音频作者:Zhisheng Zhang,Pengyang Huang摘要:近年来,深度神经网络的显著进步带来了巨大的便利。然而,一个高效的模型的训练过程需要大量的样本,这带来了巨大的潜在威胁,如未经授权的利用与隐私泄露。作为回应,我们提出了一个名为HiddenSpeaker的框架,在训练语音样本中嵌入不可感知的扰动,并使其对于基于深度学习的说话人验证系统不可学习,该系统采用大规模说话人进行有效训练。HiddenSpeaker采用一种简化的误差最小化方法,称为单级误差最小化(SLEM),以产生特定的和有效的扰动。此外,采用混合目标函数进行人类感知优化,确保扰动与人类听众无法区分。我们在说话人验证领域的多个最先进的(SOTA)模型上进行了广泛的实验,以评估HiddenSpeaker。我们的研究结果表明,HiddenSpeaker不仅用不可学习的样本欺骗模型,而且还增强了扰动的不可感知性,显示出在不同模型之间的强可移植性。摘要:In recent years, the remarkable advancements in deep neural networks have brought tremendous convenience. However, the training process of a highly effective model necessitates a substantial quantity of samples, which brings huge potential threats, like unauthorized exploitation with privacy leakage. In response, we propose a framework named HiddenSpeaker, embedding imperceptible perturbations within the training speech samples and rendering them unlearnable for deep-learning-based speaker verification systems that employ large-scale speakers for efficient training. The HiddenSpeaker utilizes a simplified error-minimizing method named Single-Level Error-Minimizing (SLEM) to generate specific and effective perturbations. Additionally, a hybrid objective function is employed for human perceptual optimization, ensuring the perturbation is indistinguishable from human listeners. We conduct extensive experiments on multiple state-of-the-art (SOTA) models in the speaker verification domain to evaluate HiddenSpeaker. Our results demonstrate that HiddenSpeaker not only deceives the model with unlearnable samples but also enhances the imperceptibility of the perturbations, showcasing strong transferability across different models.
【2】 SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation标题:SoundLoCD:用于文本到声音生成的高效条件离散对比潜在扩散模型作者:Xinlei Niu,Jing Zhang,Christian Walder,Charles Patrick Martin摘要:我们提出了SoundLoCD,一种新的文本到声音生成框架,它采用了基于LoRA的条件离散对比潜在扩散模型。与最近的大规模声音生成模型不同,我们的模型可以在有限的计算资源下有效地训练。对比学习策略的整合进一步增强了文本条件和生成的输出之间的联系,从而产生连贯和高保真的性能。我们的实验表明,SoundLoCD优于基线,大大减少了计算资源。一项全面的消融研究进一步验证了SoundLoCD中每个组件的贡献。演示页面:\url {https://XinleiNIU.github.io/demo—SoundLoCD/}。摘要:We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently trained under limited computational resources. The integration of a contrastive learning strategy further enhances the connection between text conditions and the generated outputs, resulting in coherent and high-fidelity performance. Our experiments demonstrate that SoundLoCD outperforms the baseline with greatly reduced computational resources. A comprehensive ablation study further validates the contribution of each component within SoundLoCD. Demo page: \url{https://XinleiNIU.github.io/demo-SoundLoCD/}.
【3】 Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition作者:Zijin Gu,Tatiana Likhomanenko,He Bai,Erik McDermott,Ronan Collobert,Navdeep Jaitly摘要:长期以来,语言模型(LM)一直被用来改善自动语音识别(ASR)系统的结果,但它们没有意识到ASR系统所犯的错误。纠错模型被设计用于修复ASR错误,然而,它们与传统的LM相比几乎没有改进,主要是由于缺乏监督训练数据。在本文中,我们提出了去噪LM(DLM),这是一个$\textit{scaled}$误差校正模型,使用大量的合成数据进行训练,大大超过了先前的尝试,同时实现了新的最先进的ASR性能。我们使用文本到语音(TTS)系统来合成音频,将其输入ASR系统以产生噪声假设,然后将其与原始文本配对以训练DLM。DLM有几个$\textit{key ingredients}$:(i)放大模型和数据;(ii)使用多扬声器TTS系统;(iii)多种噪声增强策略的组合;以及(iv)新的解码技术。借助Transformer-CTC ASR,DLM在Librispeech上实现了$\textit{test-clean}$1.5%的单词错误率(WER)和$\textit{test-other}$3.3%的WER,据我们所知,这是在不使用外部音频数据的情况下报告的最佳数字,甚至与使用外部音频数据的自监督方法相匹配。此外,单个DLM适用于不同的ASR,并且大大超过了传统的基于LM的波束搜索重新评分的性能。这些结果表明,适当调查的误差校正模型有可能取代传统的LM,持有一个新的水平的准确性在ASR系统的关键。摘要:Language models (LMs) have long been used to improve results of automatic speech recognition (ASR) systems, but they are unaware of the errors that ASR systems make. Error correction models are designed to fix ASR errors, however, they showed little improvement over traditional LMs mainly due to the lack of supervised training data. In this paper, we present Denoising LM (DLM), which is a $\textit{scaled}$ error correction model trained with vast amounts of synthetic data, significantly exceeding prior attempts meanwhile achieving new state-of-the-art ASR performance. We use text-to-speech (TTS) systems to synthesize audio, which is fed into an ASR system to produce noisy hypotheses, which are then paired with the original texts to train the DLM. DLM has several $\textit{key ingredients}$: (i) up-scaled model and data; (ii) usage of multi-speaker TTS systems; (iii) combination of multiple noise augmentation strategies; and (iv) new decoding techniques. With a Transformer-CTC ASR, DLM achieves 1.5% word error rate (WER) on $\textit{test-clean}$ and 3.3% WER on $\textit{test-other}$ on Librispeech, which to our knowledge are the best reported numbers in the setting where no external audio data are used and even match self-supervised methods which use external audio data. Furthermore, a single DLM is applicable to different ASRs, and greatly surpassing the performance of conventional LM based beam-search rescoring. These results indicate that properly investigated error correction models have the potential to replace conventional LMs, holding the key to a new level of accuracy in ASR systems.【4】 The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation摘要:白噪声信号可以访问任何可能的值配置,尽管在统计上在许多样本上趋于均匀的频谱分布,并且极不可能产生可理解的声音。但有多不可能?基于在真实音乐音频信号中观察到的一些必要特征,如最接近的运动和过零率,分析了白噪声在不同持续时间内产生类音乐信号的概率。根据数学结果,音乐作为信号的稀有性被认为是整体性的。这项研究的适用性不仅表明音乐具有珍贵的稀有价值,而且对音乐大小相对于音频信号空间整体大小的检查提供了信息,以告知新一代算法音乐系统(现在通常直接建立在音频信号生成上,并且可能通过诸如扩散的机器学习过程与白噪声相关)。估计上限的稀有音乐的各种物理和音乐空间的大小进行比较,以更好地理解结果的大小(双关语)。这项研究的基础是这样一个问题:“还有多少音乐存在?”以及“机器学习过程实际上可以达到多少音乐?”'.摘要:A white noise signal can access any possible configuration of values, though statistically over many samples tends to a uniform spectral distribution, and is highly unlikely to produce intelligible sound. But how unlikely? The probability that white noise generates a music-like signal over different durations is analyzed, based on some necessary features observed in real music audio signals such as mostly proximate movement and zero crossing rate. Given the mathematical results, the rarity of music as a signal is considered overall. The applicability of this study is not just to show that music has a precious rarity value, but that examination of the size of music relative to the overall size of audio signal space provides information to inform new generations of algorithmic music system (which are now often founded on audio signal generation directly, and may relate to white noise via such machine learning processes as diffusion). Estimated upper bounds on the rarity of music to the size of various physical and musical spaces are compared, to better understand the magnitude of the results (pun intended). Underlying the research are the questions `how much music is still out there?' and `how much music could a machine learning process actually reach?'.【5】 Music Genre Classification: Training an AI model摘要:音乐流派分类是利用机器学习模型和技术来处理音频信号的领域,其中应用范围从内容推荐系统到音乐推荐系统。在这项研究中,我探索了各种机器学习算法,用于音乐流派分类,使用从音频信号中提取的特征。这些系统分别是多层感知器(从头开始构建),k最近邻(也是从头开始构建),卷积神经网络,最后是随机森林模型。为了处理音频信号,执行诸如短时傅立叶变换的特征提取方法和梅尔倒谱系数(MFCC)的提取。通过这项广泛的研究,我的目标是评估体裁分类的机器学习模型的鲁棒性,并比较它们的结果。摘要:Music genre classification is an area that utilizes machine learning models and techniques for the processing of audio signals, in which applications range from content recommendation systems to music recommendation systems. In this research I explore various machine learning algorithms for the purpose of music genre classification, using features extracted from audio signals.The systems are namely, a Multilayer Perceptron (built from scratch), a k-Nearest Neighbours (also built from scratch), a Convolutional Neural Network and lastly a Random Forest wide model. In order to process the audio signals, feature extraction methods such as Short-Time Fourier Transform, and the extraction of Mel Cepstral Coefficients (MFCCs), is performed. Through this extensive research, I aim to asses the robustness of machine learning models for genre classification, and to compare their results.【6】 Acoustical Features as Knee Health Biomarkers: A Critical Analysis标题:声学特征作为膝关节健康生物标志物:批判性分析作者:Christodoulos Kechris,Jerome Thevenot,Tomas Teijeiro,Vincent A. Stadelmann,Nicola A. Maffiuletti,David Atienza摘要:声学膝关节健康评估长期以来一直承诺替代临床可用的医学成像工具,但这种模式尚未在医疗实践中采用。该领域目前由处理声学特征的机器学习模型领导,这些模型具有很好的诊断性能。然而,这些方法忽略了音频信号的复杂多源性质和起作用的潜在机制。通过解决这个关键的差距,本文介绍了一种新的因果框架,用于验证膝盖的声学功能。我们认为,目前用于声学膝关节诊断的机器学习方法缺乏所需的保证,因此不能用于将声学特征分类为生物标志物。我们的框架建立了一套必要的理论保证,以验证这一主张。我们将我们的方法应用到三个真实世界的实验中,研究人员的期望,实验协议和可穿戴传感器的影响。这项调查揭示了潜在的问题,如潜在的捷径学习和性能膨胀。这项研究是在声学膝关节健康评估领域的第一个独立的结果再现研究。最后,我们从我们的研究结果中得出了可操作的见解,为在未来的研究中克服这些关键限制提供了有价值的指导。摘要:Acoustical knee health assessment has long promised an alternative to clinically available medical imaging tools, but this modality has yet to be adopted in medical practice. The field is currently led by machine learning models processing acoustical features, which have presented promising diagnostic performances. However, these methods overlook the intricate multi-source nature of audio signals and the underlying mechanisms at play. By addressing this critical gap, the present paper introduces a novel causal framework for validating knee acoustical features. We argue that current machine learning methodologies for acoustical knee diagnosis lack the required assurances and thus cannot be used to classify acoustic features as biomarkers. Our framework establishes a set of essential theoretical guarantees necessary to validate this claim. We apply our methodology to three real-world experiments investigating the effect of researchers' expectations, the experimental protocol and the wearable employed sensor. This investigation reveals latent issues such as underlying shortcut learning and performance inflation. This study is the first independent result reproduction study in the field of acoustical knee health evaluation. We conclude with actionable insights from our findings, offering valuable guidance to navigate these crucial limitations in future research.【1】 Real-Time and Accurate: Zero-shot High-Fidelity Singing Voice Conversion with Multi-Condition Flow Synthesis标题:实时准确:采用多条件流合成的Zero-Shot高保真歌唱声音转换链接:https://arxiv.org/abs/2405.15093作者:Hui Li,Hongyu Wang,Zhijin Chen,Bohan Sun,Bo Li摘要:演唱声音转换是将原演唱声音转换为除内容外的目标演唱声音。目前,基于流的模型可以完成语音转换的任务,但他们很难有效地提取潜在变量,在更富有节奏和情感表达的歌唱语音转换任务,同时还面临着语音处理效率低的问题。在本文中,我们提出了一个高保真的基于流的模型的基础上,多解耦的特征约束,提高了捕获的声音细节,通过集成多个编码器。我们还使用iSTFT通过替换声码器的某些层来提高语音处理的速度。我们从多个维度将合成的歌声与其他模型进行了比较,我们提出的模型与当前最先进的模型高度一致,演示可在\url{https://lazycat1119.github.io/RASVC-demo/}摘要:Singing voice conversion is to convert the source sing voice into the target sing voice except for the content. Currently, flow-based models can complete the task of voice conversion, but they struggle to effectively extract latent variables in the more rhythmically rich and emotionally expressive task of singing voice conversion, while also facing issues with low efficiency in speech processing. In this paper, we propose a high-fidelity flow-based model based on multi-decoupling feature constraints, which enhances the capture of vocal details by integrating multiple encoders. We also use iSTFT to enhance the speed of speech processing by replacing some layers of the Vocoder. We compare the synthesized singing voice with other models from multiple dimensions, and our proposed model is highly consistent with the current state-of-the-art, with the demo which is available at \url{https://lazycat1119.github.io/RASVC-demo/}
【2】 Acoustical Features as Knee Health Biomarkers: A Critical Analysis标题:声学特征作为膝关节健康生物标志物:批判性分析作者:Christodoulos Kechris,Jerome Thevenot,Tomas Teijeiro,Vincent A. Stadelmann,Nicola A. Maffiuletti,David Atienza摘要:声学膝关节健康评估长期以来一直承诺替代临床可用的医学成像工具,但这种模式尚未在医疗实践中采用。该领域目前由处理声学特征的机器学习模型领导,这些模型具有很好的诊断性能。然而,这些方法忽略了音频信号的复杂多源性质和起作用的潜在机制。通过解决这个关键的差距,本文介绍了一种新的因果框架,用于验证膝盖的声学功能。我们认为,目前用于声学膝关节诊断的机器学习方法缺乏所需的保证,因此不能用于将声学特征分类为生物标志物。我们的框架建立了一套必要的理论保证,以验证这一主张。我们将我们的方法应用到三个真实世界的实验中,研究人员的期望,实验协议和可穿戴传感器的影响。这项调查揭示了潜在的问题,如潜在的捷径学习和性能膨胀。这项研究是在声学膝关节健康评估领域的第一个独立的结果再现研究。最后,我们从我们的研究结果中得出了可操作的见解,为在未来的研究中克服这些关键限制提供了有价值的指导。摘要:Acoustical knee health assessment has long promised an alternative to clinically available medical imaging tools, but this modality has yet to be adopted in medical practice. The field is currently led by machine learning models processing acoustical features, which have presented promising diagnostic performances. However, these methods overlook the intricate multi-source nature of audio signals and the underlying mechanisms at play. By addressing this critical gap, the present paper introduces a novel causal framework for validating knee acoustical features. We argue that current machine learning methodologies for acoustical knee diagnosis lack the required assurances and thus cannot be used to classify acoustic features as biomarkers. Our framework establishes a set of essential theoretical guarantees necessary to validate this claim. We apply our methodology to three real-world experiments investigating the effect of researchers' expectations, the experimental protocol and the wearable employed sensor. This investigation reveals latent issues such as underlying shortcut learning and performance inflation. This study is the first independent result reproduction study in the field of acoustical knee health evaluation. We conclude with actionable insights from our findings, offering valuable guidance to navigate these crucial limitations in future research.【3】 HiddenSpeaker: Generate Imperceptible Unlearnable Audios for Speaker Verification System标题:HiddenSpeaker:为说话人验证系统生成难以感知的、不可学习的音频作者:Zhisheng Zhang,Pengyang Huang摘要:近年来,深度神经网络的显著进步带来了巨大的便利。然而,一个高效的模型的训练过程需要大量的样本,这带来了巨大的潜在威胁,如未经授权的利用与隐私泄露。作为回应,我们提出了一个名为HiddenSpeaker的框架,在训练语音样本中嵌入不可感知的扰动,并使其对于基于深度学习的说话人验证系统不可学习,该系统采用大规模说话人进行有效训练。HiddenSpeaker采用一种简化的误差最小化方法,称为单级误差最小化(SLEM),以产生特定的和有效的扰动。此外,采用混合目标函数进行人类感知优化,确保扰动与人类听众无法区分。我们在说话人验证领域的多个最先进的(SOTA)模型上进行了广泛的实验,以评估HiddenSpeaker。我们的研究结果表明,HiddenSpeaker不仅用不可学习的样本欺骗模型,而且还增强了扰动的不可感知性,显示出在不同模型之间的强可移植性。摘要:In recent years, the remarkable advancements in deep neural networks have brought tremendous convenience. However, the training process of a highly effective model necessitates a substantial quantity of samples, which brings huge potential threats, like unauthorized exploitation with privacy leakage. In response, we propose a framework named HiddenSpeaker, embedding imperceptible perturbations within the training speech samples and rendering them unlearnable for deep-learning-based speaker verification systems that employ large-scale speakers for efficient training. The HiddenSpeaker utilizes a simplified error-minimizing method named Single-Level Error-Minimizing (SLEM) to generate specific and effective perturbations. Additionally, a hybrid objective function is employed for human perceptual optimization, ensuring the perturbation is indistinguishable from human listeners. We conduct extensive experiments on multiple state-of-the-art (SOTA) models in the speaker verification domain to evaluate HiddenSpeaker. Our results demonstrate that HiddenSpeaker not only deceives the model with unlearnable samples but also enhances the imperceptibility of the perturbations, showcasing strong transferability across different models.
【4】 SoundLoCD: An Efficient Conditional Discrete Contrastive Latent Diffusion Model for Text-to-Sound Generation标题:SoundLoCD:用于文本到声音生成的高效条件离散对比潜在扩散模型作者:Xinlei Niu,Jing Zhang,Christian Walder,Charles Patrick Martin摘要:我们提出了SoundLoCD,一种新的文本到声音生成框架,它采用了基于LoRA的条件离散对比潜在扩散模型。与最近的大规模声音生成模型不同,我们的模型可以在有限的计算资源下有效地训练。对比学习策略的整合进一步增强了文本条件和生成的输出之间的联系,从而产生连贯和高保真的性能。我们的实验表明,SoundLoCD优于基线,大大减少了计算资源。一项全面的消融研究进一步验证了SoundLoCD中每个组件的贡献。演示页面:\url {https://XinleiNIU.github.io/demo—SoundLoCD/}。摘要:We present SoundLoCD, a novel text-to-sound generation framework, which incorporates a LoRA-based conditional discrete contrastive latent diffusion model. Unlike recent large-scale sound generation models, our model can be efficiently trained under limited computational resources. The integration of a contrastive learning strategy further enhances the connection between text conditions and the generated outputs, resulting in coherent and high-fidelity performance. Our experiments demonstrate that SoundLoCD outperforms the baseline with greatly reduced computational resources. A comprehensive ablation study further validates the contribution of each component within SoundLoCD. Demo page: \url{https://XinleiNIU.github.io/demo-SoundLoCD/}.【5】 Denoising LM: Pushing the Limits of Error Correction Models for Speech Recognition作者:Zijin Gu,Tatiana Likhomanenko,He Bai,Erik McDermott,Ronan Collobert,Navdeep Jaitly摘要:长期以来,语言模型(LM)一直被用来改善自动语音识别(ASR)系统的结果,但它们没有意识到ASR系统所犯的错误。纠错模型被设计用于修复ASR错误,然而,它们与传统的LM相比几乎没有改进,主要是由于缺乏监督训练数据。在本文中,我们提出了去噪LM(DLM),这是一个$\textit{scaled}$误差校正模型,使用大量的合成数据进行训练,大大超过了先前的尝试,同时实现了新的最先进的ASR性能。我们使用文本到语音(TTS)系统来合成音频,将其输入ASR系统以产生噪声假设,然后将其与原始文本配对以训练DLM。DLM有几个$\textit{key ingredients}$:(i)放大模型和数据;(ii)使用多扬声器TTS系统;(iii)多种噪声增强策略的组合;以及(iv)新的解码技术。借助Transformer-CTC ASR,DLM在Librispeech上实现了$\textit{test-clean}$1.5%的单词错误率(WER)和$\textit{test-other}$3.3%的WER,据我们所知,这是在不使用外部音频数据的情况下报告的最佳数字,甚至与使用外部音频数据的自监督方法相匹配。此外,单个DLM适用于不同的ASR,并且大大超过了传统的基于LM的波束搜索重新评分的性能。这些结果表明,适当调查的误差校正模型有可能取代传统的LM,持有一个新的水平的准确性在ASR系统的关键。摘要:Language models (LMs) have long been used to improve results of automatic speech recognition (ASR) systems, but they are unaware of the errors that ASR systems make. Error correction models are designed to fix ASR errors, however, they showed little improvement over traditional LMs mainly due to the lack of supervised training data. In this paper, we present Denoising LM (DLM), which is a $\textit{scaled}$ error correction model trained with vast amounts of synthetic data, significantly exceeding prior attempts meanwhile achieving new state-of-the-art ASR performance. We use text-to-speech (TTS) systems to synthesize audio, which is fed into an ASR system to produce noisy hypotheses, which are then paired with the original texts to train the DLM. DLM has several $\textit{key ingredients}$: (i) up-scaled model and data; (ii) usage of multi-speaker TTS systems; (iii) combination of multiple noise augmentation strategies; and (iv) new decoding techniques. With a Transformer-CTC ASR, DLM achieves 1.5% word error rate (WER) on $\textit{test-clean}$ and 3.3% WER on $\textit{test-other}$ on Librispeech, which to our knowledge are the best reported numbers in the setting where no external audio data are used and even match self-supervised methods which use external audio data. Furthermore, a single DLM is applicable to different ASRs, and greatly surpassing the performance of conventional LM based beam-search rescoring. These results indicate that properly investigated error correction models have the potential to replace conventional LMs, holding the key to a new level of accuracy in ASR systems.
【6】 The Rarity of Musical Audio Signals Within the Space of Possible Audio Generation摘要:白噪声信号可以访问任何可能的值配置,尽管在统计上在许多样本上趋于均匀的频谱分布,并且极不可能产生可理解的声音。但有多不可能?基于在真实音乐音频信号中观察到的一些必要特征,如最接近的运动和过零率,分析了白噪声在不同持续时间内产生类音乐信号的概率。根据数学结果,音乐作为信号的稀有性被认为是整体性的。这项研究的适用性不仅表明音乐具有珍贵的稀有价值,而且对音乐大小相对于音频信号空间整体大小的检查提供了信息,以告知新一代算法音乐系统(现在通常直接建立在音频信号生成上,并且可能通过诸如扩散的机器学习过程与白噪声相关)。估计上限的稀有音乐的各种物理和音乐空间的大小进行比较,以更好地理解结果的大小(双关语)。这项研究的基础是这样一个问题:“还有多少音乐存在?”以及“机器学习过程实际上可以达到多少音乐?”'.摘要:A white noise signal can access any possible configuration of values, though statistically over many samples tends to a uniform spectral distribution, and is highly unlikely to produce intelligible sound. But how unlikely? The probability that white noise generates a music-like signal over different durations is analyzed, based on some necessary features observed in real music audio signals such as mostly proximate movement and zero crossing rate. Given the mathematical results, the rarity of music as a signal is considered overall. The applicability of this study is not just to show that music has a precious rarity value, but that examination of the size of music relative to the overall size of audio signal space provides information to inform new generations of algorithmic music system (which are now often founded on audio signal generation directly, and may relate to white noise via such machine learning processes as diffusion). Estimated upper bounds on the rarity of music to the size of various physical and musical spaces are compared, to better understand the magnitude of the results (pun intended). Underlying the research are the questions `how much music is still out there?' and `how much music could a machine learning process actually reach?'.
【7】 Music Genre Classification: Training an AI model摘要:音乐流派分类是利用机器学习模型和技术来处理音频信号的领域,其中应用范围从内容推荐系统到音乐推荐系统。在这项研究中,我探索了各种机器学习算法,用于音乐流派分类,使用从音频信号中提取的特征。这些系统分别是多层感知器(从头开始构建),k最近邻(也是从头开始构建),卷积神经网络,最后是随机森林模型。为了处理音频信号,执行诸如短时傅立叶变换的特征提取方法和梅尔倒谱系数(MFCC)的提取。通过这项广泛的研究,我的目标是评估体裁分类的机器学习模型的鲁棒性,并比较它们的结果。摘要:Music genre classification is an area that utilizes machine learning models and techniques for the processing of audio signals, in which applications range from content recommendation systems to music recommendation systems. In this research I explore various machine learning algorithms for the purpose of music genre classification, using features extracted from audio signals.The systems are namely, a Multilayer Perceptron (built from scratch), a k-Nearest Neighbours (also built from scratch), a Convolutional Neural Network and lastly a Random Forest wide model. In order to process the audio signals, feature extraction methods such as Short-Time Fourier Transform, and the extraction of Mel Cepstral Coefficients (MFCCs), is performed. Through this extensive research, I aim to asses the robustness of machine learning models for genre classification, and to compare their results.