本文经arXiv每日学术速递授权转载
【1】 Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
标题: 通过原始音频上的音乐相似性来评估音乐生成中的数据复制
作者:Roser Batlle-Roca,Wei-Hisang Liao,Xavier Serra,Yuki Mitsufuji,Emilia Gómez
备注:Accepted at ISMIR 2024
链接:点击下载PDF文件
【2】 Stable Audio Open
标题: 稳定的音频打开
作者:Zach Evans,Julian D. Parker,CJ Carr,Zack Zukowski,Josiah Taylor,Jordi Pons
备注:Demo: this https URL Weights: this https URL Code: this https URL arXiv admin note: text overlap with arXiv:2404.10301
链接:点击下载PDF文件
【3】 Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
标题: 使用来自大型语言模型的声音属性知识增强Zero-Shot音频分类
作者:Xuenan Xu,Pingyue Zhang,Ming Yan,Ji Zhang,Mengyue Wu
备注:Interspeech 2024
链接:点击下载PDF文件
【4】 Efficient Audio Captioning with Encoder-Level Knowledge Distillation
标题: 具有编码器级知识提炼的高效音频字幕
作者:Xuenan Xu,Haohe Liu,Mengyue Wu,Wenwu Wang,Mark D. Plumbley
备注:Interspeech 2024
链接:点击下载PDF文件
【5】 Guitar Chord Diagram Suggestion for Western Popular Music
标题: 吉他和弦图对西方流行音乐的建议
作者:Alexandre d'Hooge,Louis Bigo,Ken Déguernel,Nicolas Martin
Journal-ref:Sound and Music Computing Conference, Jul 2024, Porto, Portugal
链接:点击下载PDF文件
【6】 Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
标题: 盲文语音生成器:基于CLIP和Fastspeech 2联合微调的音频生成
作者:Chun Xu,En-Wei Sun
链接:点击下载PDF文件
【7】 Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
标题: Rasa:在低资源环境中为印度语言构建表达性语音合成系统
作者:Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024. First two authors listed contributed equally
链接:点击下载PDF文件
【8】 Topology-Independent GEVD-Based Distributed Adaptive Node-Specific Signal Estimation in Ad-Hoc Wireless Acoustic Sensor Networks
标题: 自组织无线声学传感器网络中基于独立的GEVD的分布式自适应特定节点信号估计
作者:Paul Didier,Toon van Waterschoot,Marc Moonen
备注:Presented in the 2024 32nd European Signal Processing Conference (EUSIPCO)
链接:点击下载PDF文件
【9】 GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
标题: GE 2 E-AC:口音分类的广义端到端损失训练
作者:Chihiro Watanabe,Hirokazu Kameoka
链接:点击下载PDF文件
【10】 MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
标题: MSceneSpeech:用于表达性语音合成的多场景语音数据集
作者:Qian Yang,Jialong Zuo,Zhe Su,Ziyue Jiang,Mingze Li,Zhou Zhao,Feiyang Chen,Zhefeng Wang,Baoxing Huai
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【11】 Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
标题: 用于发音障碍和老年人语音识别的自我监督ASB模型和功能
作者:Shujie Hu,Xurong Xie,Mengzhe Geng,Zengrui Jin,Jiajun Deng,Guinan Li,Yi Wang,Mingyu Cui,Tianzi Wang,Helen Meng,Xunying Liu
备注:IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
标题: PolySinger:从英语到日语的歌唱声到歌唱声翻译
作者:Silas Antonisen,Iván López-Espejo
备注:This paper was accepted at ISMIR 2024
链接:点击下载PDF文件
【2】 Topology-Independent GEVD-Based Distributed Adaptive Node-Specific Signal Estimation in Ad-Hoc Wireless Acoustic Sensor Networks
标题: 自组织无线声学传感器网络中基于独立的GEVD的分布式自适应特定节点信号估计
作者:Paul Didier,Toon van Waterschoot,Marc Moonen
备注:Presented in the 2024 32nd European Signal Processing Conference (EUSIPCO)
链接:点击下载PDF文件
【3】 Wideband Relative Transfer Function (RTF) Estimation Exploiting Frequency Correlations
标题: 利用频率相关性的宽带相对传递函数(RTI)估计
作者:Giovanni Bologni,Richard C. Hendriks,Richard Heusdens
备注:Under review at IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
【4】 GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
标题: GE 2 E-AC:口音分类的广义端到端损失训练
作者:Chihiro Watanabe,Hirokazu Kameoka
链接:点击下载PDF文件
【5】 MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
标题: MSceneSpeech:用于表达性语音合成的多场景语音数据集
作者:Qian Yang,Jialong Zuo,Zhe Su,Ziyue Jiang,Mingze Li,Zhou Zhao,Feiyang Chen,Zhefeng Wang,Baoxing Huai
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
【6】 Improving Robustness and Clinical Applicability of Respiratory Sound Classification via Audio Enhancement
标题: 通过音频增强提高呼吸音分类的稳健性和临床适用性
作者:Jing-Tong Tzeng,Jeng-Lin Li,Huan-Yu Chen,Chun-Hsiang Huang,Chi-Hsin Chen,Cheng-Yi Fan,Edward Pei-Chuan Huang,Chi-Chun Lee
备注:The following article has been submitted to The Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this https URL
链接:点击下载PDF文件
【7】 Semi-Supervised Contrastive Learning of Musical Representations
标题: 音乐表象的半监督对比学习
作者:Julien Guinot,Elio Quinton,György Fazekas
备注:Accepted to be published at the Proceedings of the 25th International Society for Music Information Retrieval Conference 2024, Includes non-proceedings appendix
链接:点击下载PDF文件
【8】 Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
标题: 用于发音障碍和老年人语音识别的自我监督ASB模型和功能
作者:Shujie Hu,Xurong Xie,Mengzhe Geng,Zengrui Jin,Jiajun Deng,Guinan Li,Yi Wang,Mingyu Cui,Tianzi Wang,Helen Meng,Xunying Liu
备注:IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
【9】 Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
标题: 通过原始音频上的音乐相似性来评估音乐生成中的数据复制
作者:Roser Batlle-Roca,Wei-Hisang Liao,Xavier Serra,Yuki Mitsufuji,Emilia Gómez
备注:Accepted at ISMIR 2024
链接:点击下载PDF文件
【10】 Stable Audio Open
标题: 稳定的音频打开
作者:Zach Evans,Julian D. Parker,CJ Carr,Zack Zukowski,Josiah Taylor,Jordi Pons
备注:Demo: this https URL Weights: this https URL Code: this https URL arXiv admin note: text overlap with arXiv:2404.10301
链接:点击下载PDF文件
【11】 Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
标题: 使用来自大型语言模型的声音属性知识增强Zero-Shot音频分类
作者:Xuenan Xu,Pingyue Zhang,Ming Yan,Ji Zhang,Mengyue Wu
备注:Interspeech 2024
链接:点击下载PDF文件
【12】 Efficient Audio Captioning with Encoder-Level Knowledge Distillation
标题: 具有编码器级知识提炼的高效音频字幕
作者:Xuenan Xu,Haohe Liu,Mengyue Wu,Wenwu Wang,Mark D. Plumbley
备注:Interspeech 2024
链接:点击下载PDF文件
【13】 CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation Units
标题: CoVoSwitch:基于语调单位的合成代码切换文本的机器翻译
作者:Yeeun Kang
备注:Accepted to ACL 2024 Student Research Workshop (ACL-SRW 2024)
链接:点击下载PDF文件
【14】 Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
标题: 盲文语音生成器:基于CLIP和Fastspeech 2联合微调的音频生成
作者:Chun Xu,En-Wei Sun
链接:点击下载PDF文件
【15】 Automatic Classification of News Subjects in Broadcast News: Application to a Gender Bias Representation Analysis
标题: 广播新闻中新闻主题的自动分类:性别偏见表示分析的应用
作者:Valentin Pelloin,Lena Dodson,Émile Chapuis,Nicolas Hervé,David Doukhan
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【16】 Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
标题: Rasa:在低资源环境中为印度语言构建表达性语音合成系统
作者:Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024. First two authors listed contributed equally
链接:点击下载PDF文件
标题: 通过原始音频上的音乐相似性来评估音乐生成中的数据复制
作者:Roser Batlle-Roca,Wei-Hisang Liao,Xavier Serra,Yuki Mitsufuji,Emilia Gómez
备注:Accepted at ISMIR 2024
链接:点击下载PDF文件
摘要:音乐生成的最新进展引发了人们对人工智能在创意音乐过程中的影响、当前商业模式以及与知识产权管理相关的影响的多种担忧。一个相关的挑战是人工智能生成的音乐中训练集的潜在复制和剽窃,这可能导致数据滥用和侵犯知识产权。为了解决这个问题,我们提出了音乐复制评估(MiRA)工具:一个独立于模型的开放式评估方法,基于不同的音频音乐相似性指标来评估训练集的数据复制。我们评估五个指标,以确定准确的复制的能力,通过进行控制复制实验,在不同的音乐流派的基础上合成的样本。我们的研究结果表明,所提出的方法可以估计准确的数据复制的比例高于10%。通过引入MiRA工具,我们打算鼓励研究人员,开发人员和用户对音乐生成模型进行关于数据复制的开放评估,强调生成AI在音乐领域的伦理,社会,法律和经济后果的重要性。摘要:Recent advancements in music generation are raising multiple concerns about the implications of AI in creative music processes, current business models and impacts related to intellectual property management. A relevant challenge is the potential replication and plagiarism of the training set in AI-generated music, which could lead to misuse of data and intellectual property rights violations. To tackle this issue, we present the Music Replication Assessment (MiRA) tool: a model-independent open evaluation method based on diverse audio music similarity metrics to assess data replication of the training set. We evaluate the ability of five metrics to identify exact replication, by conducting a controlled replication experiment in different music genres based on synthetic samples. Our results show that the proposed methodology can estimate exact data replication with a proportion higher than 10%. By introducing the MiRA tool, we intend to encourage the open evaluation of music generative models by researchers, developers and users concerning data replication, highlighting the importance of ethical, social, legal and economic consequences of generative AI in the music domain.
【2】 Stable Audio Open
标题: 稳定的音频打开
作者:Zach Evans,Julian D. Parker,CJ Carr,Zack Zukowski,Josiah Taylor,Jordi Pons
备注:Demo: this https URL Weights: this https URL Code: this https URL arXiv admin note: text overlap with arXiv:2404.10301
链接:点击下载PDF文件
摘要:开放生成模型对社区至关重要,允许微调并在呈现新模型时作为基线。然而,大多数当前的文本到音频模型都是私有的,艺术家和研究人员无法访问。在这里,我们描述了一个新的开放权重文本到音频模型的架构和训练过程,该模型使用知识共享数据进行训练。我们的评估表明,该模型的性能是具有竞争力的国家的最先进的各种指标。值得注意的是,报告的FDopenl3结果(测量世代的真实性)展示了其在44.1kHz下高质量立体声合成的潜力。摘要:Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.
【3】 Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
标题: 使用来自大型语言模型的声音属性知识增强Zero-Shot音频分类
作者:Xuenan Xu,Pingyue Zhang,Ming Yan,Ji Zhang,Mengyue Wu
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:Zero-shot音频分类旨在识别和分类模型在训练期间从未见过的声音类别。本文提出了一种新的方法,zero-shot音频分类自动生成的声音属性描述。我们提出了一个声音属性列表,并利用大型语言模型的领域知识来生成详细的属性描述每个类。与以前主要依赖于类标签或简单描述的作品相比,我们的方法侧重于多维先天听觉属性,捕捉声音类的不同特征。此外,我们采用了对比学习的方法,以提高从文本标签的zero-shot学习。我们在VGGSound和AudioSet上验证了我们的方法的有效性 footnote{代码可在 url{https: www.github.com wsntxxn AttrEnhZsAc}.}获得。我们的研究结果表明,在zero-shot分类精度的实质性改善。烧蚀结果显示出强大的性能增强,无论模型架构。摘要:Zero-shot audio classification aims to recognize and classify a sound class that the model has never seen during training. This paper presents a novel approach for zero-shot audio classification using automatically generated sound attribute descriptions. We propose a list of sound attributes and leverage large language model's domain knowledge to generate detailed attribute descriptions for each class. In contrast to previous works that primarily relied on class labels or simple descriptions, our method focuses on multi-dimensional innate auditory attributes, capturing different characteristics of sound classes. Additionally, we incorporate a contrastive learning approach to enhance zero-shot learning from textual labels. We validate the effectiveness of our method on VGGSound and AudioSet footnote{The code is available at url{https: www.github.com wsntxxn AttrEnhZsAc}.}. Our results demonstrate a substantial improvement in zero-shot classification accuracy. Ablation results show robust performance enhancement, regardless of the model architecture.
【4】 Efficient Audio Captioning with Encoder-Level Knowledge Distillation
标题: 具有编码器级知识提炼的高效音频字幕
作者:Xuenan Xu,Haohe Liu,Mengyue Wu,Wenwu Wang,Mark D. Plumbley
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:最近的模型在自动音频字幕(AAC)方面取得了显着的改进。然而,这些模型随着其性能的增强而变得越来越大。在这项工作中,我们提出了一个知识蒸馏(KD)的AAC框架。我们的分析表明,在基于编码器-解码器的AAC模型中,与解码器相比,将知识提取到编码器中更有效。为此,除了标准的监督损失和序列级KD损失之外,我们还将编码器级KD损失纳入训练。我们研究了两种编码器级KD方法,分别基于均方误差(MSE)损失和对比损失。实验结果表明,对比KD比MSE KD更强大,在数据稀缺的情况下表现出优越的性能。通过在KD框架中利用纯音频数据进行训练,我们的学生模型实现了具有竞争力的性能,推理速度提高了19倍 footnote{在线演示可在 url{https: huggingface.co spaces wsntxxn efficient_audio_captioning}}上获得。摘要:Significant improvement has been achieved in automated audio captioning (AAC) with recent models. However, these models have become increasingly large as their performance is enhanced. In this work, we propose a knowledge distillation (KD) framework for AAC. Our analysis shows that in the encoder-decoder based AAC models, it is more effective to distill knowledge into the encoder as compared with the decoder. To this end, we incorporate encoder-level KD loss into training, in addition to the standard supervised loss and sequence-level KD loss. We investigate two encoder-level KD methods, based on mean squared error (MSE) loss and contrastive loss, respectively. Experimental results demonstrate that contrastive KD is more robust than MSE KD, exhibiting superior performance in data-scarce situations. By leveraging audio-only data into training in the KD framework, our student model achieves competitive performance, with an inference speed that is 19 times faster footnote{An online demo is available at url{https: huggingface.co spaces wsntxxn efficient_audio_captioning}}.
【5】 Guitar Chord Diagram Suggestion for Western Popular Music
标题: 吉他和弦图对西方流行音乐的建议
作者:Alexandre d'Hooge,Louis Bigo,Ken Déguernel,Nicolas Martin
Journal-ref:Sound and Music Computing Conference, Jul 2024, Porto, Portugal
链接:点击下载PDF文件
摘要:和弦图是吉他手用来显示在指板上哪里以及如何演奏和弦的。他们是有用的初学者学习和弦或分享所需的手的位置演奏一首歌。然而,吉他学习工具上呈现的图表通常是从现有的数据库中选择的,很少代表表演者使用的实际位置。在本文中,我们提出了一个工具,建议和弦图的和弦标签,在对DadaGP和mySongBook数据集进行统计分析的基础上,我们发现有些和弦图是过度的一些和弦可以用20多种不同的方式演奏。我们认为,考虑到上下文可以提高多样性,和弦图建议的质量,并将此方法与仅考虑当前和弦标签的模型进行比较。我们表明,添加先前的上下文可以将此任务的F1分数提高高达27%并减少了模型建议标准开放和弦的倾向。我们还定义了和弦图上下文中的纹理概念,并通过各种度量显示我们的模型改进与上图一致。摘要:Chord diagrams are used by guitar players to show where and how to play a chord on the fretboard. They are useful to beginners learning chords or for sharing the hand positions required to play a song.However, the diagrams presented on guitar learning toolsare usually selected from an existing databaseand rarely represent the actual positions used by performers.In this paper, we propose a tool which suggests a chord diagram for achord label,taking into account the diagram of the previous chord.Based on statistical analysis of the DadaGP and mySongBook datasets, we show that some chord diagrams are over-represented in western popular musicand that some chords can be played in more than 20 different ways.We argue that taking context into account can improve the variety and the quality of chord diagram suggestion, and compare this approach with a model taking only the current chord label into account.We show that adding previous context improves the F1-score on this task by up to 27% and reduces the propensity of the model to suggest standard open chords.We also define the notion of texture in the context of chord diagrams andshow through a variety of metrics that our model improves textureconsistencywith the previous diagram.
【6】 Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
标题: 盲文语音生成器:基于CLIP和Fastspeech 2联合微调的音频生成
作者:Chun Xu,En-Wei Sun
链接:点击下载PDF文件
摘要:越来越多的中国人受到不同程度的视觉障碍的困扰,这使得视野中的单个图像或视频帧与表达相同信息的音频之间的模态转换成为研究热点。OCR+Vocoder和Im 2 Wav等深度学习技术可以以自我监督的方式实现英语音频合成或图像到声音的匹配。然而,用于培训的音频数据有限,英语对于不同教育水平的视障者来说并不普遍。因此,为了解决数据量和语言适用性问题,提高视障人士的阅读效率,构建了一套基于中文语境的图像语音转换框架CLIP-KNN-Fastspeech 2。该框架集成了多个基本模型,采用独立预训练和联合微调的策略。首先,分别在MUGE和Baker两个公共数据集上对中文CLIP和Fastspeech 2文语转换模型进行了预训练,并验证了其收敛性。随后,使用自建的盲文图像数据集进行联合微调。在VGGSound、Flickr 8 k、ImageHear等多个公共数据集和自建盲文数据集BIT-DP上的实验结果表明,该模型在BLEU 4、FAD(Fr 'echet Audio Distance)、WER(Word Error Ratio)等客观指标上都有提高,甚至推理速度也有所提高。这验证了所构建的模型在有限的数据下仍具有合成高质量语音的能力,也证明了融合多个基本模型的联合训练策略的有效性。摘要:An increasing number of Chinese people are troubled by different degrees of visual impairment, which has made the modal conversion between a single image or video frame in the visual field and the audio expressing the same information a research hotspot. Deep learning technologies such as OCR+Vocoder and Im2Wav enable English audio synthesis or image-to-sound matching in a self-supervised manner. However, the audio data used for training is limited and English is not universal for visually impaired people with different educational levels. Therefore, for the sake of solving the problems of data volume and language applicability to improve the reading efficiency of visually impaired people, a set of image-to-speech framework CLIP-KNN-Fastspeech2 based on the Chinese context was constructed. The framework integrates multiple basic models and adopts the strategy of independent pre-training and joint fine-tuning. First, the Chinese CLIP and Fastspeech2 text-to-speech models were pre-trained on two public datasets, MUGE and Baker, respectively, and their convergence was verified. Subsequently, joint fine-tuning was performed using a self-built Braille image dataset. Experimental results on multiple public datasets such as VGGSound, Flickr8k, ImageHear, and the self-built Braille dataset BIT-DP show that the model has improved objective indicators such as BLEU4,FAD(Fr 'echet Audio Distance), WER(Word Error Ratio), and even inference speed. This verifies that the constructed model still has the ability to synthesize high-quality speech under limited data, and also proves the effectiveness of the joint training strategy that integrates multiple basic models.
【7】 Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
标题: Rasa:在低资源环境中为印度语言构建表达性语音合成系统
作者:Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024. First two authors listed contributed equally
链接:点击下载PDF文件
摘要:我们发布了Rasa,这是第一个针对任何印度语言的多语言表达性TTS数据集,其中包含10小时的中性语音和1-3小时的表达性语音,用于涵盖3种语言的6种Ekman情绪:阿萨姆语,孟加拉语和泰米尔语。我们的消融研究表明,仅1小时的中性和30分钟的表达数据就可以产生一个公平的系统,如MUSHRA评分所示。将中性数据增加到10小时,使用最少的表达性数据,显著增强了表达能力。这为资源受限的语言提供了一个实用的方法,优先考虑容易获得的中性数据以及少量的表达性数据。我们的音节平衡的数据和汇集的情绪,以提高表现力的重要性。我们还强调了产生特定情绪的挑战,例如,恐惧和惊讶。摘要:We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the 6 Ekman emotions covering 3 languages: Assamese, Bengali, & Tamil. Our ablation studies reveal that just 1 hour of neutral and 30 minutes of expressive data can yield a Fair system as indicated by MUSHRA scores. Increasing neutral data to 10 hours, with minimal expressive data, significantly enhances expressiveness. This offers a practical recipe for resource-constrained languages, prioritizing easily obtainable neutral data alongside smaller amounts of expressive data. We show the importance of syllabically balanced data and pooling emotions to enhance expressiveness. We also highlight challenges in generating specific emotions, e.g., fear and surprise.
【8】 Topology-Independent GEVD-Based Distributed Adaptive Node-Specific Signal Estimation in Ad-Hoc Wireless Acoustic Sensor Networks
标题: 自组织无线声学传感器网络中基于独立的GEVD的分布式自适应特定节点信号估计
作者:Paul Didier,Toon van Waterschoot,Marc Moonen
备注:Presented in the 2024 32nd European Signal Processing Conference (EUSIPCO)
链接:点击下载PDF文件
摘要:提出了一种基于低秩近似的拓扑无关分布式自适应节点特定信号估计(TI-DANSE)算法,并将其应用于ad-hoc无线声传感器网络中。该TI-GEVD-DANSE算法以及原始TI-DANSE算法表现出非严格收敛性,这可能导致随时间推移的数值不稳定性,特别是在精确空间协方差矩阵的估计具有挑战性的情况下。提出了一种自适应滤波器系数归一化策略来缓解这个问题,并使TI-(GEVD-)DANSE的性能稳定。该方法在包括动态声学场景的数值模拟中得到了验证,证明了额外归一化的重要性。摘要:A low-rank approximation-based version of the topology-independent distributed adaptive node-specific signal estimation (TI-DANSE) algorithm is introduced, using a generalized eigenvalue decomposition (GEVD) for application in ad-hoc wireless acoustic sensor networks. This TI-GEVD-DANSE algorithm as well as the original TI-DANSE algorithm exhibit a non-strict convergence, which can lead to numerical instability over time, particularly in scenarios where the estimation of accurate spatial covariance matrices is challenging. An adaptive filter coefficient normalization strategy is proposed to mitigate this issue and enable the stable performance of TI-(GEVD-)DANSE. The method is validated in numerical simulations including dynamic acoustic scenarios, demonstrating the importance of the additional normalization.
【9】 GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
标题: GE 2 E-AC:口音分类的广义端到端损失训练
作者:Chihiro Watanabe,Hirokazu Kameoka
链接:点击下载PDF文件
摘要:口音分类或AC是一项预测输入话语的口音类型的任务,它可以用作带口音语音识别和口音转换的初步步骤。现有的研究通常通过训练神经网络模型来实现这种分类,以最小化预测口音标签的分类误差,该预测口音标签可以作为模型输出来获得。由于在这种方法中,我们仅从训练期间分类损失的角度优化整个模型,因此模型可能会学习从不相关的特征(例如单个说话者身份)预测口音类型,这些特征在测试期间没有信息。为了解决这个问题,我们提出了一个GE 2 E-AC,其中我们训练一个模型来提取口音嵌入或输入话语的AE,使得相同口音类的AE更接近,而不是直接最小化分类损失。我们通过实验证明了所提出的GE 2 E-AC的有效性,与传统的基于交叉熵的损失训练的基线模型相比。摘要:Accent classification or AC is a task to predict the accent type of an input utterance, and it can be used as a preliminary step toward accented speech recognition and accent conversion. Existing studies have often achieved such classification by training a neural network model to minimize the classification error of the predicted accent label, which can be obtained as a model output. Since we optimize the entire model only from the perspective of classification loss during training time in this approach, the model might learn to predict the accent type from irrelevant features, such as individual speaker identity, which are not informative during test time. To address this problem, we propose a GE2E-AC, in which we train a model to extract accent embedding or AE of an input utterance such that the AEs of the same accent class get closer, instead of directly minimizing the classification loss. We experimentally show the effectiveness of the proposed GE2E-AC, compared to the baseline model trained with the conventional cross-entropy-based loss.
【10】 MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
标题: MSceneSpeech:用于表达性语音合成的多场景语音数据集
作者:Qian Yang,Jialong Zuo,Zhe Su,Ziyue Jiang,Mingze Li,Zhou Zhao,Feiyang Chen,Zhefeng Wang,Baoxing Huai
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们介绍了一个开源的高质量的普通话TTS数据集MSceneSpeech(多场景语音数据集),旨在为表达性语音合成提供资源。MSceneSpeech包括根据日常生活场景执行和记录的大量音频记录和文本。每个场景包括多个扬声器和各种韵律风格,使其适合需要多扬声器风格和韵律建模的语音合成。我们已经建立了一个强大的基线,通过提示机制,可以有效地合成语音的特点是用户特定的音色和场景特定的韵律与任意文本输入。开源MSceneSpeech Dataset和我们基线的音频样本可在https: speechai-demo.github.io MSceneSpeech 上获得。摘要:We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio recordings and texts performed and recorded according to daily life scenarios. Each scenario includes multiple speakers and a diverse range of prosodic styles, making it suitable for speech synthesis that entails multi-speaker style and prosody modeling. We have established a robust baseline, through the prompting mechanism, that can effectively synthesize speech characterized by both user-specific timbre and scene-specific prosody with arbitrary text input. The open source MSceneSpeech Dataset and audio samples of our baseline are available at https: speechai-demo.github.io MSceneSpeech .
【11】 Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
标题: 用于发音障碍和老年人语音识别的自我监督ASB模型和功能
作者:Shujie Hu,Xurong Xie,Mengzhe Geng,Zengrui Jin,Jiajun Deng,Guinan Li,Yi Wang,Mingyu Cui,Tianzi Wang,Helen Meng,Xunying Liu
备注:IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
摘要:基于自监督学习(SSL)的语音基础模型已被应用于广泛的ASR任务。然而,他们的应用程序构音障碍和老年人的语音通过数据密集型参数微调面临着在域数据稀缺和不匹配。为此,本文探索了一系列方法,将领域微调的SSL预训练模型及其功能集成到TDNN和Conformer ASR系统中,用于构音障碍和老年人语音识别。其中包括:a)标准声学前端和域微调SSL语音表示之间的输入特征融合; b)单独使用标准声学特征单独训练的TDNN系统和具有附加域微调SSL特征的TDNN系统之间的帧级联合解码;以及c)涉及要使用域微调预训练ASR模型重新评分的TDNN Conformer系统输出的多遍解码。此外,微调SSL语音功能中使用的声学发音(A2 A)的反演,以构建多模态ASR系统。实验在四个任务上进行:英语UASpeech和TORGO构音障碍言语语料库;以及英语DementiaBank Pitt和粤语JCCOCC MoCA老年言语数据集。通过集成域适应的HuBERT、wav 2 vec 2-conformer或多语言XLSR模型及其功能构建的TDNN系统始终优于独立的微调SSL预训练模型。这些系统在四项任务上分别产生了统计学显著的WER或CER减少,分别为6.53%、1.90%、2.04%和7.97%(相对减少24.10%、23.84%、10.14%和31.39%)。使用DementiaBank Pitt老年人语音识别输出也获得了阿尔茨海默病检测准确性的一致改善。摘要:Self-supervised learning (SSL) based speech foundation models have been applied to a wide range of ASR tasks. However, their application to dysarthric and elderly speech via data-intensive parameter fine-tuning is confronted by in-domain data scarcity and mismatch. To this end, this paper explores a series of approaches to integrate domain fine-tuned SSL pre-trained models and their features into TDNN and Conformer ASR systems for dysarthric and elderly speech recognition. These include: a) input feature fusion between standard acoustic frontends and domain fine-tuned SSL speech representations; b) frame-level joint decoding between TDNN systems separately trained using standard acoustic features alone and those with additional domain fine-tuned SSL features; and c) multi-pass decoding involving the TDNN Conformer system outputs to be rescored using domain fine-tuned pre-trained ASR models. In addition, fine-tuned SSL speech features are used in acoustic-to-articulatory (A2A) inversion to construct multi-modal ASR systems. Experiments are conducted on four tasks: the English UASpeech and TORGO dysarthric speech corpora; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets. The TDNN systems constructed by integrating domain-adapted HuBERT, wav2vec2-conformer or multi-lingual XLSR models and their features consistently outperform the standalone fine-tuned SSL pre-trained models. These systems produced statistically significant WER or CER reductions of 6.53%, 1.90%, 2.04% and 7.97% absolute (24.10%, 23.84%, 10.14% and 31.39% relative) on the four tasks respectively. Consistent improvements in Alzheimer's Disease detection accuracy are also obtained using the DementiaBank Pitt elderly speech recognition outputs.
eess.AS音频处理
【1】 PolySinger: Singing-Voice to Singing-Voice Translation from English to Japanese标题: PolySinger:从英语到日语的歌唱声到歌唱声翻译
作者:Silas Antonisen,Iván López-Espejo
备注:This paper was accepted at ISMIR 2024
链接:点击下载PDF文件
摘要:语音领域在几个自然语言处理(NLP)任务的聚光灯下占主导地位,而唱歌领域仍然较少探索。NLP的顶点是语音到语音翻译(S2 ST)任务,指的是人类语音的翻译和合成。S2 ST和可能的适应歌唱领域,我们称之为歌唱声音到歌唱声音翻译(SV 2SVT)之间的差距,正变得突出,因为前者的进展越来越快,而后者则处于停滞状态。尽管人们对多语言歌曲创作和歌曲翻译的关注有限,但歌唱声音合成系统正在克服多语言合成的障碍。本文试图确定成功的SV 2SVT需要什么,并提出PolySinger( textbf{Poly}glot textbf{Singer}):第一个SV 2SVT系统,执行从英语到日语的歌词翻译。一个级联的方法,提出了建立一个框架,具有高度的控制,可以潜在地减少SV 2SVT和S2 ST之间的差距。PolySinger的性能通过日语母语者的平均意见得分测试进行评估。结果和深入的讨论,与测试对象提出了一个坚实的基础SV 2SVT,但必须克服的几个缺点,这是讨论了未来的SV 2SVT。摘要:The speech domain prevails in the spotlight for several natural language processing (NLP) tasks while the singing domain remains less explored. The culmination of NLP is the speech-to-speech translation (S2ST) task, referring to translation and synthesis of human speech. A disparity between S2ST and the possible adaptation to the singing domain, which we describe as singing-voice to singing-voice translation (SV2SVT), is becoming prominent as the former is progressing ever faster, while the latter is at a standstill. Singing-voice synthesis systems are overcoming the barrier of multi-lingual synthesis, despite limited attention has been paid to multi-lingual songwriting and song translation. This paper endeavors to determine what is required for successful SV2SVT and proposes PolySinger ( textbf{Poly}glot textbf{Singer}): the first system for SV2SVT, performing lyrics translation from English to Japanese. A cascaded approach is proposed to establish a framework with a high degree of control which can potentially diminish the disparity between SV2SVT and S2ST. The performance of PolySinger is evaluated by a mean opinion score test with native Japanese speakers. Results and in-depth discussions with test subjects suggest a solid foundation for SV2SVT, but several shortcomings must be overcome, which are discussed for the future of SV2SVT.
【2】 Topology-Independent GEVD-Based Distributed Adaptive Node-Specific Signal Estimation in Ad-Hoc Wireless Acoustic Sensor Networks
标题: 自组织无线声学传感器网络中基于独立的GEVD的分布式自适应特定节点信号估计
作者:Paul Didier,Toon van Waterschoot,Marc Moonen
备注:Presented in the 2024 32nd European Signal Processing Conference (EUSIPCO)
链接:点击下载PDF文件
摘要:提出了一种基于低秩近似的拓扑无关分布式自适应节点特定信号估计(TI-DANSE)算法,并将其应用于ad-hoc无线声传感器网络中。该TI-GEVD-DANSE算法以及原始TI-DANSE算法表现出非严格收敛性,这可能导致随时间推移的数值不稳定性,特别是在精确空间协方差矩阵的估计具有挑战性的情况下。提出了一种自适应滤波器系数归一化策略来缓解这个问题,并使TI-(GEVD-)DANSE的性能稳定。该方法在包括动态声学场景的数值模拟中得到了验证,证明了额外归一化的重要性。摘要:A low-rank approximation-based version of the topology-independent distributed adaptive node-specific signal estimation (TI-DANSE) algorithm is introduced, using a generalized eigenvalue decomposition (GEVD) for application in ad-hoc wireless acoustic sensor networks. This TI-GEVD-DANSE algorithm as well as the original TI-DANSE algorithm exhibit a non-strict convergence, which can lead to numerical instability over time, particularly in scenarios where the estimation of accurate spatial covariance matrices is challenging. An adaptive filter coefficient normalization strategy is proposed to mitigate this issue and enable the stable performance of TI-(GEVD-)DANSE. The method is validated in numerical simulations including dynamic acoustic scenarios, demonstrating the importance of the additional normalization.
【3】 Wideband Relative Transfer Function (RTF) Estimation Exploiting Frequency Correlations
标题: 利用频率相关性的宽带相对传递函数(RTI)估计
作者:Giovanni Bologni,Richard C. Hendriks,Richard Heusdens
备注:Under review at IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
摘要:本文主要讨论波束成形应用中的相对传递函数(RTF)估计。虽然传统的方法假设频谱是不相关的,但由于诸如多普勒效应的自然现象、诸如时域加窗的人工操作或如在语音中观察到的信号的非平稳性质,在实际场景中经常违反该假设。为了解决这个问题,我们提出了一个RTF估计技术,利用光谱和空间相关性,通过子空间分析。为了克服估计真实数据的二阶谱统计量的挑战,我们采用了最初在发动机故障检测的背景下提出的相位调整估计器。此外,我们推导出Cram 'er-Rao界(CRBs)的RTF估计任务,提供理论上的见解,可实现的估计精度。边界表明,信道估计可以更准确地执行,如果噪声或目标呈现谱相关性。真实和合成数据上的实验表明,当目标表现出谱相关性时,我们的技术优于窄带最大似然估计。虽然所提出的算法的精度通常接近的界限,有一些改进的空间,特别是当噪声信号具有高频谱相关性。虽然信道估计的应用是多种多样的,我们证明了在语音阵列处理的背景下的方法。摘要:This article focuses on estimating relative transfer functions (RTFs) for beamforming applications. While traditional methods assume that spectra are uncorrelated, this assumption is often violated in practical scenarios due to natural phenomena such as the Doppler effect, artificial manipulations like time-domain windowing, or the non-stationary nature of the signals, as observed in speech. To address this, we propose an RTF estimation technique that leverages spectral and spatial correlations through subspace analysis. To overcome the challenge of estimating second-order spectral statistics for real data, we employ a phase-adjusted estimator originally proposed in the context of engine fault detection. Additionally, we derive Cram 'er--Rao bounds (CRBs) for the RTF estimation task, providing theoretical insights into the achievable estimation accuracy. The bounds show that channel estimation can be performed more accurately if the noise or the target presents spectral correlations. Experiments on real and synthetic data show that our technique outperforms the narrowband maximum-likelihood estimator when the target exhibits spectral correlations. Although the accuracy of the proposed algorithm is generally close to the bound, there is some room for improvement, especially when noise signals with high spectral correlation are present. While the applications of channel estimation are diverse, we demonstrate the method in the context of array processing for speech.
【4】 GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
标题: GE 2 E-AC:口音分类的广义端到端损失训练
作者:Chihiro Watanabe,Hirokazu Kameoka
链接:点击下载PDF文件
摘要:口音分类或AC是一项预测输入话语的口音类型的任务,它可以用作带口音语音识别和口音转换的初步步骤。现有的研究通常通过训练神经网络模型来实现这种分类,以最小化预测口音标签的分类误差,该预测口音标签可以作为模型输出来获得。由于在这种方法中,我们仅从训练期间分类损失的角度优化整个模型,因此模型可能会学习从不相关的特征(例如单个说话者身份)预测口音类型,这些特征在测试期间没有信息。为了解决这个问题,我们提出了一个GE 2 E-AC,其中我们训练一个模型来提取口音嵌入或输入话语的AE,使得相同口音类的AE更接近,而不是直接最小化分类损失。我们通过实验证明了所提出的GE 2 E-AC的有效性,与传统的基于交叉熵的损失训练的基线模型相比。摘要:Accent classification or AC is a task to predict the accent type of an input utterance, and it can be used as a preliminary step toward accented speech recognition and accent conversion. Existing studies have often achieved such classification by training a neural network model to minimize the classification error of the predicted accent label, which can be obtained as a model output. Since we optimize the entire model only from the perspective of classification loss during training time in this approach, the model might learn to predict the accent type from irrelevant features, such as individual speaker identity, which are not informative during test time. To address this problem, we propose a GE2E-AC, in which we train a model to extract accent embedding or AE of an input utterance such that the AEs of the same accent class get closer, instead of directly minimizing the classification loss. We experimentally show the effectiveness of the proposed GE2E-AC, compared to the baseline model trained with the conventional cross-entropy-based loss.
【5】 MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
标题: MSceneSpeech:用于表达性语音合成的多场景语音数据集
作者:Qian Yang,Jialong Zuo,Zhe Su,Ziyue Jiang,Mingze Li,Zhou Zhao,Feiyang Chen,Zhefeng Wang,Baoxing Huai
备注:Accepted by INTERSPEECH 2024
链接:点击下载PDF文件
摘要:我们介绍了一个开源的高质量的普通话TTS数据集MSceneSpeech(多场景语音数据集),旨在为表达性语音合成提供资源。MSceneSpeech包括根据日常生活场景执行和记录的大量音频记录和文本。每个场景包括多个扬声器和各种韵律风格,使其适合需要多扬声器风格和韵律建模的语音合成。我们通过提示机制建立了一个强大的基线,可以通过任意文本输入有效地合成具有用户特定音色和场景特定韵律特征的语音。开源MSceneSpeech Dataset和我们基线的音频样本可在https: speechai-demo.github.io MSceneSpeech 上获得。摘要:We introduce an open source high-quality Mandarin TTS dataset MSceneSpeech (Multiple Scene Speech Dataset), which is intended to provide resources for expressive speech synthesis. MSceneSpeech comprises numerous audio recordings and texts performed and recorded according to daily life scenarios. Each scenario includes multiple speakers and a diverse range of prosodic styles, making it suitable for speech synthesis that entails multi-speaker style and prosody modeling. We have established a robust baseline, through the prompting mechanism, that can effectively synthesize speech characterized by both user-specific timbre and scene-specific prosody with arbitrary text input. The open source MSceneSpeech Dataset and audio samples of our baseline are available at https: speechai-demo.github.io MSceneSpeech .
【6】 Improving Robustness and Clinical Applicability of Respiratory Sound Classification via Audio Enhancement
标题: 通过音频增强提高呼吸音分类的稳健性和临床适用性
作者:Jing-Tong Tzeng,Jeng-Lin Li,Huan-Yu Chen,Chun-Hsiang Huang,Chi-Hsin Chen,Cheng-Yi Fan,Edward Pei-Chuan Huang,Chi-Chun Lee
备注:The following article has been submitted to The Journal of the Acoustical Society of America (JASA). After it is published, it will be found at this https URL
链接:点击下载PDF文件
摘要:深度学习技术在呼吸音自动分类方面显示出了很好的效果。然而,在现实世界的嘈杂条件下准确区分这些声音对临床部署提出了挑战。此外,仅用背景噪声预测信号可能会破坏用户对系统的信任。在这项研究中,我们提出了一个音频增强(AE)管道作为呼吸声分类之前的预处理步骤,旨在提高在嘈杂环境中的性能。使用不同的音频增强模型结构进行了多个实验,与噪声注入数据增强的基线方法相比,证明了改进的分类性能。具体而言,AE管道的集成导致ICBHI呼吸音数据集上的ICBHI分类评分增加了2.59%,并且在多类噪声场景中,我们最近收集的Formosa Archive of Breath Sounds(FABS)提高了2.51%。此外,医生验证研究评估了我们的系统的临床效用。定量分析显示,与原始噪声记录相比,我们的系统在模型辅助诊断过程中提高了效率,诊断信心和信任。集成增强音频的工作流程使诊断灵敏度提高了11.61%,并促进了高置信度的诊断。我们的研究结果表明,将音频增强算法显着提高了鲁棒性和临床实用性。摘要:Deep learning techniques have shown promising results in the automatic classification of respiratory sounds. However, accurately distinguishing these sounds in real-world noisy conditions poses challenges for clinical deployment. Additionally, predicting signals with only background noise could undermine user trust in the system. In this study, we propose an audio enhancement (AE) pipeline as a pre-processing step before respiratory sound classification, aiming to improve performance in noisy environments. Multiple experiments were conducted using different audio enhancement model structures, demonstrating improved classification performance compared to the baseline method of noise injection data augmentation. Specifically, the integration of the AE pipeline resulted in a 2.59% increase in the ICBHI classification score on the ICBHI respiratory sound dataset and a 2.51% improvement on our recently collected Formosa Archive of Breath Sounds (FABS) in multi-class noisy scenarios. Furthermore, a physician validation study assessed the clinical utility of our system. Quantitative analysis revealed enhancements in efficiency, diagnostic confidence, and trust during model-assisted diagnosis with our system compared to raw noisy recordings. Workflows integrating enhanced audio led to an 11.61% increase in diagnostic sensitivity and facilitated high-confidence diagnoses. Our findings demonstrate that incorporating an audio enhancement algorithm significantly enhances robustness and clinical utility.
【7】 Semi-Supervised Contrastive Learning of Musical Representations
标题: 音乐表象的半监督对比学习
作者:Julien Guinot,Elio Quinton,György Fazekas
备注:Accepted to be published at the Proceedings of the 25th International Society for Music Information Retrieval Conference 2024, Includes non-proceedings appendix
链接:点击下载PDF文件
摘要:尽管对比学习在音乐信息检索中取得了成功,但对比自我监督的内在模糊性提出了挑战。仅仅依靠增强链和自我监督的正采样策略可能会导致预训练目标无法捕获下游任务的关键音乐信息。我们介绍了半监督对比学习(SemiSupCon),这是一种在音乐表征的对比学习中利用音乐信息标记数据(监督信号)的简单方法。我们的方法通过在比以前的方法更简单的框架中结合监督和自监督对比目标,将音乐相关的监督信号引入自监督对比学习。该框架提高了下游的性能和鲁棒性音频损坏的下游MIR任务的范围与适量的标记数据。我们的方法能够通过选择标记数据来塑造学习到的相似性度量,这些标记数据(1)为表示注入音乐领域知识,(2)以最小的一般下游性能损失提高域外性能。我们在与音乐相关但又不太相似的任务上表现出了很强的迁移学习性能,比如音高和基调估计。此外,我们的方法显示了自动标记的性能优于自我监督方法,预训练中仅包含5%的可用标签。摘要:Despite the success of contrastive learning in Music Information Retrieval, the inherent ambiguity of contrastive self-supervision presents a challenge. Relying solely on augmentation chains and self-supervised positive sampling strategies can lead to a pretraining objective that does not capture key musical information for downstream tasks. We introduce semi-supervised contrastive learning (SemiSupCon), a simple method for leveraging musically informed labeled data (supervision signals) in the contrastive learning of musical representations. Our approach introduces musically relevant supervision signals into self-supervised contrastive learning by combining supervised and self-supervised contrastive objectives in a simpler framework than previous approaches. This framework improves downstream performance and robustness to audio corruptions on a range of downstream MIR tasks with moderate amounts of labeled data. Our approach enables shaping the learned similarity metric through the choice of labeled data that (1) infuses the representations with musical domain knowledge and (2) improves out-of-domain performance with minimal general downstream performance loss. We show strong transfer learning performance on musically related yet not trivially similar tasks - such as pitch and key estimation. Additionally, our approach shows performance improvement on automatic tagging over self-supervised approaches with only 5 % of available labels included in pretraining.
【8】 Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition
标题: 用于发音障碍和老年人语音识别的自我监督ASB模型和功能
作者:Shujie Hu,Xurong Xie,Mengzhe Geng,Zengrui Jin,Jiajun Deng,Guinan Li,Yi Wang,Mingyu Cui,Tianzi Wang,Helen Meng,Xunying Liu
备注:IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
摘要:基于自监督学习(SSL)的语音基础模型已被应用于广泛的ASR任务。然而,他们的应用程序构音障碍和老年人的语音通过数据密集型参数微调面临着在域数据稀缺和不匹配。为此,本文探索了一系列方法,将领域微调的SSL预训练模型及其功能集成到TDNN和Conformer ASR系统中,用于构音障碍和老年人语音识别。其中包括:a)标准声学前端和域微调SSL语音表示之间的输入特征融合; b)单独使用标准声学特征单独训练的TDNN系统和具有附加域微调SSL特征的TDNN系统之间的帧级联合解码;以及c)涉及要使用域微调预训练ASR模型重新评分的TDNN Conformer系统输出的多遍解码。此外,微调SSL语音功能中使用的声学发音(A2 A)的反演,以构建多模态ASR系统。实验在四个任务上进行:英语UASpeech和TORGO构音障碍言语语料库;以及英语DementiaBank Pitt和粤语JCCOCC MoCA老年言语数据集。通过集成域适应的HuBERT、wav 2 vec 2-conformer或多语言XLSR模型及其功能构建的TDNN系统始终优于独立的微调SSL预训练模型。这些系统在四项任务上分别产生了统计学显著的WER或CER减少,分别为6.53%、1.90%、2.04%和7.97%(相对减少24.10%、23.84%、10.14%和31.39%)。使用DementiaBank Pitt老年人语音识别输出也获得了阿尔茨海默病检测准确性的一致改善。摘要:Self-supervised learning (SSL) based speech foundation models have been applied to a wide range of ASR tasks. However, their application to dysarthric and elderly speech via data-intensive parameter fine-tuning is confronted by in-domain data scarcity and mismatch. To this end, this paper explores a series of approaches to integrate domain fine-tuned SSL pre-trained models and their features into TDNN and Conformer ASR systems for dysarthric and elderly speech recognition. These include: a) input feature fusion between standard acoustic frontends and domain fine-tuned SSL speech representations; b) frame-level joint decoding between TDNN systems separately trained using standard acoustic features alone and those with additional domain fine-tuned SSL features; and c) multi-pass decoding involving the TDNN Conformer system outputs to be rescored using domain fine-tuned pre-trained ASR models. In addition, fine-tuned SSL speech features are used in acoustic-to-articulatory (A2A) inversion to construct multi-modal ASR systems. Experiments are conducted on four tasks: the English UASpeech and TORGO dysarthric speech corpora; and the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech datasets. The TDNN systems constructed by integrating domain-adapted HuBERT, wav2vec2-conformer or multi-lingual XLSR models and their features consistently outperform the standalone fine-tuned SSL pre-trained models. These systems produced statistically significant WER or CER reductions of 6.53%, 1.90%, 2.04% and 7.97% absolute (24.10%, 23.84%, 10.14% and 31.39% relative) on the four tasks respectively. Consistent improvements in Alzheimer's Disease detection accuracy are also obtained using the DementiaBank Pitt elderly speech recognition outputs.
【9】 Towards Assessing Data Replication in Music Generation with Music Similarity Metrics on Raw Audio
标题: 通过原始音频上的音乐相似性来评估音乐生成中的数据复制
作者:Roser Batlle-Roca,Wei-Hisang Liao,Xavier Serra,Yuki Mitsufuji,Emilia Gómez
备注:Accepted at ISMIR 2024
链接:点击下载PDF文件
摘要:音乐生成的最新进展引发了人们对人工智能在创意音乐过程中的影响、当前商业模式以及与知识产权管理相关的影响的多种担忧。一个相关的挑战是人工智能生成的音乐中训练集的潜在复制和剽窃,这可能导致数据滥用和侵犯知识产权。为了解决这个问题,我们提出了音乐复制评估(MiRA)工具:一个独立于模型的开放式评估方法,基于不同的音频音乐相似性指标来评估训练集的数据复制。我们评估五个指标,以确定准确的复制的能力,通过进行控制复制实验,在不同的音乐流派的基础上合成的样本。我们的研究结果表明,所提出的方法可以估计准确的数据复制的比例高于10%。通过引入MiRA工具,我们打算鼓励研究人员,开发人员和用户对音乐生成模型进行关于数据复制的开放评估,强调生成AI在音乐领域的伦理,社会,法律和经济后果的重要性。摘要:Recent advancements in music generation are raising multiple concerns about the implications of AI in creative music processes, current business models and impacts related to intellectual property management. A relevant challenge is the potential replication and plagiarism of the training set in AI-generated music, which could lead to misuse of data and intellectual property rights violations. To tackle this issue, we present the Music Replication Assessment (MiRA) tool: a model-independent open evaluation method based on diverse audio music similarity metrics to assess data replication of the training set. We evaluate the ability of five metrics to identify exact replication, by conducting a controlled replication experiment in different music genres based on synthetic samples. Our results show that the proposed methodology can estimate exact data replication with a proportion higher than 10%. By introducing the MiRA tool, we intend to encourage the open evaluation of music generative models by researchers, developers and users concerning data replication, highlighting the importance of ethical, social, legal and economic consequences of generative AI in the music domain.
【10】 Stable Audio Open
标题: 稳定的音频打开
作者:Zach Evans,Julian D. Parker,CJ Carr,Zack Zukowski,Josiah Taylor,Jordi Pons
备注:Demo: this https URL Weights: this https URL Code: this https URL arXiv admin note: text overlap with arXiv:2404.10301
链接:点击下载PDF文件
摘要:开放生成模型对社区至关重要,允许微调并在呈现新模型时作为基线。然而,大多数当前的文本到音频模型都是私有的,艺术家和研究人员无法访问。在这里,我们描述了一个新的开放权重文本到音频模型的架构和训练过程,该模型使用知识共享数据进行训练。我们的评估表明,该模型的性能是具有竞争力的国家的最先进的各种指标。值得注意的是,报告的FDopenl3结果(测量世代的真实性)展示了其在44.1kHz下高质量立体声合成的潜力。摘要:Open generative models are vitally important for the community, allowing for fine-tunes and serving as baselines when presenting new models. However, most current text-to-audio models are private and not accessible for artists and researchers to build upon. Here we describe the architecture and training process of a new open-weights text-to-audio model trained with Creative Commons data. Our evaluation shows that the model's performance is competitive with the state-of-the-art across various metrics. Notably, the reported FDopenl3 results (measuring the realism of the generations) showcase its potential for high-quality stereo sound synthesis at 44.1kHz.
【11】 Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
标题: 使用来自大型语言模型的声音属性知识增强Zero-Shot音频分类
作者:Xuenan Xu,Pingyue Zhang,Ming Yan,Ji Zhang,Mengyue Wu
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:Zero-shot音频分类旨在识别和分类模型在训练期间从未见过的声音类别。本文提出了一种新的方法,zero-shot音频分类自动生成的声音属性描述。我们提出了一个声音属性列表,并利用大型语言模型的领域知识来生成详细的属性描述每个类。与以前主要依赖于类标签或简单描述的作品相比,我们的方法侧重于多维先天听觉属性,捕捉声音类的不同特征。此外,我们采用对比学习方法来增强文本标签的zero-shot学习。我们在VGGSound和AudioSet上验证了我们的方法的有效性 footnote{该代码可在 url{https: www.github.com wsntxxn AttrEnhZsAc}.}获得。我们的研究结果表明,在zero-shot分类精度的实质性改善。烧蚀结果显示出强大的性能增强,无论模型架构。摘要:Zero-shot audio classification aims to recognize and classify a sound class that the model has never seen during training. This paper presents a novel approach for zero-shot audio classification using automatically generated sound attribute descriptions. We propose a list of sound attributes and leverage large language model's domain knowledge to generate detailed attribute descriptions for each class. In contrast to previous works that primarily relied on class labels or simple descriptions, our method focuses on multi-dimensional innate auditory attributes, capturing different characteristics of sound classes. Additionally, we incorporate a contrastive learning approach to enhance zero-shot learning from textual labels. We validate the effectiveness of our method on VGGSound and AudioSet footnote{The code is available at url{https: www.github.com wsntxxn AttrEnhZsAc}.}. Our results demonstrate a substantial improvement in zero-shot classification accuracy. Ablation results show robust performance enhancement, regardless of the model architecture.
【12】 Efficient Audio Captioning with Encoder-Level Knowledge Distillation
标题: 具有编码器级知识提炼的高效音频字幕
作者:Xuenan Xu,Haohe Liu,Mengyue Wu,Wenwu Wang,Mark D. Plumbley
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:最近的模型在自动音频字幕(AAC)方面取得了显着的改进。然而,这些模型随着其性能的增强而变得越来越大。在这项工作中,我们提出了一个知识蒸馏(KD)的AAC框架。我们的分析表明,在基于编码器-解码器的AAC模型中,与解码器相比,将知识提取到编码器中更有效。为此,除了标准的监督损失和序列级KD损失之外,我们还将编码器级KD损失纳入训练。我们研究了两种编码器级KD方法,分别基于均方误差(MSE)损失和对比损失。实验结果表明,对比KD比MSE KD更强大,在数据稀缺的情况下表现出优越的性能。通过在KD框架中利用纯音频数据进行训练,我们的学生模型实现了具有竞争力的性能,推理速度提高了19倍 footnote{在线演示可在 url{https: huggingface.co spaces wsntxxn efficient_audio_captioning}}上获得。摘要:Significant improvement has been achieved in automated audio captioning (AAC) with recent models. However, these models have become increasingly large as their performance is enhanced. In this work, we propose a knowledge distillation (KD) framework for AAC. Our analysis shows that in the encoder-decoder based AAC models, it is more effective to distill knowledge into the encoder as compared with the decoder. To this end, we incorporate encoder-level KD loss into training, in addition to the standard supervised loss and sequence-level KD loss. We investigate two encoder-level KD methods, based on mean squared error (MSE) loss and contrastive loss, respectively. Experimental results demonstrate that contrastive KD is more robust than MSE KD, exhibiting superior performance in data-scarce situations. By leveraging audio-only data into training in the KD framework, our student model achieves competitive performance, with an inference speed that is 19 times faster footnote{An online demo is available at url{https: huggingface.co spaces wsntxxn efficient_audio_captioning}}.
【13】 CoVoSwitch: Machine Translation of Synthetic Code-Switched Text Based on Intonation Units
标题: CoVoSwitch:基于语调单位的合成代码切换文本的机器翻译
作者:Yeeun Kang
备注:Accepted to ACL 2024 Student Research Workshop (ACL-SRW 2024)
链接:点击下载PDF文件
摘要:多语言语码转换研究经常受到可用数据集的缺乏和语言偏见的阻碍。为了扩展语言表示,我们通过使用语音到文本翻译数据集CoVoST 2替换通过PSST检测到的语调单元来合成代码切换数据,PSST是一种从OpenAI的Whisper微调的语音分割模型。使用我们的数据集CoVoSwitch,跨越13种语言,我们评估了两种多语言翻译模型M2M-100 418 M和NLLB-200 600 M的代码切换翻译性能。我们发现,包括代码转换单元的结果在更高的翻译性能比单语设置和模型更好地在代码转换翻译成英语比非英语。此外,低资源语言在翻译成英语时从语码转换单元的整合中获益最多,但在翻译成非英语时获益较少。翻译成低资源语言的性能甚至比原始代码切换输入更差。我们发现,系统擅长复制英语令牌,但与非英语令牌的斗争,在单语设置中的脱靶问题也相关的代码切换设置,模型幻觉在代码切换翻译中引入的话缺席的两个原始的源句子。CoVoSwitch和代码可在https: github.com sophiayk20 covoswitch上获得。摘要:Multilingual code-switching research is often hindered by the lack and linguistically biased status of available datasets. To expand language representation, we synthesize code-switching data by replacing intonation units detected through PSST, a speech segmentation model fine-tuned from OpenAI's Whisper, using a speech-to-text translation dataset, CoVoST 2. With our dataset, CoVoSwitch, spanning 13 languages, we evaluate the code-switching translation performance of two multilingual translation models, M2M-100 418M and NLLB-200 600M. We reveal that the inclusion of code-switching units results in higher translation performance than monolingual settings and that models are better at code-switching translation into English than non-English. Further, low-resource languages gain most from integration of code-switched units when translating into English but much less when translating into non-English. Translations into low-resource languages also perform worse than even raw code-switched inputs. We find that systems excel at copying English tokens but struggle with non-English tokens, that the off-target problem in monolingual settings is also relevant in code-switching settings, and that models hallucinate in code-switching translation by introducing words absent in both of the original source sentences. CoVoSwitch and code are available at https: github.com sophiayk20 covoswitch.
【14】 Braille-to-Speech Generator: Audio Generation Based on Joint Fine-Tuning of CLIP and Fastspeech2
标题: 盲文语音生成器:基于CLIP和Fastspeech 2联合微调的音频生成
作者:Chun Xu,En-Wei Sun
链接:点击下载PDF文件
摘要:越来越多的中国人受到不同程度的视觉障碍的困扰,这使得视野中的单个图像或视频帧与表达相同信息的音频之间的模态转换成为研究热点。OCR+Vocoder和Im 2 Wav等深度学习技术可以以自我监督的方式实现英语音频合成或图像到声音的匹配。然而,用于培训的音频数据有限,英语对于不同教育水平的视障者来说并不普遍。因此,为了解决数据量和语言适用性问题,提高视障人群的阅读效率,构建了一套基于中文语境的图像转语音框架CLIP-KNN-Fastspeech 2。该框架集成了多个基本模型,采用独立预训练和联合微调的策略。首先,分别在MUGE和Baker两个公共数据集上对中文CLIP和Fastspeech 2文语转换模型进行了预训练,并验证了其收敛性。随后,使用自建的盲文图像数据集进行联合微调。在VGGSound、Flickr 8 k、ImageHear等多个公共数据集和自建盲文数据集BIT-DP上的实验结果表明,该模型在BLEU 4、FAD(Fr 'echet Audio Distance)、WER(Word Error Ratio)等客观指标上都有提高,甚至推理速度也有所提高。这验证了所构建的模型在有限的数据下仍具有合成高质量语音的能力,也证明了融合多个基本模型的联合训练策略的有效性。摘要:An increasing number of Chinese people are troubled by different degrees of visual impairment, which has made the modal conversion between a single image or video frame in the visual field and the audio expressing the same information a research hotspot. Deep learning technologies such as OCR+Vocoder and Im2Wav enable English audio synthesis or image-to-sound matching in a self-supervised manner. However, the audio data used for training is limited and English is not universal for visually impaired people with different educational levels. Therefore, for the sake of solving the problems of data volume and language applicability to improve the reading efficiency of visually impaired people, a set of image-to-speech framework CLIP-KNN-Fastspeech2 based on the Chinese context was constructed. The framework integrates multiple basic models and adopts the strategy of independent pre-training and joint fine-tuning. First, the Chinese CLIP and Fastspeech2 text-to-speech models were pre-trained on two public datasets, MUGE and Baker, respectively, and their convergence was verified. Subsequently, joint fine-tuning was performed using a self-built Braille image dataset. Experimental results on multiple public datasets such as VGGSound, Flickr8k, ImageHear, and the self-built Braille dataset BIT-DP show that the model has improved objective indicators such as BLEU4,FAD(Fr 'echet Audio Distance), WER(Word Error Ratio), and even inference speed. This verifies that the constructed model still has the ability to synthesize high-quality speech under limited data, and also proves the effectiveness of the joint training strategy that integrates multiple basic models.
【15】 Automatic Classification of News Subjects in Broadcast News: Application to a Gender Bias Representation Analysis
标题: 广播新闻中新闻主题的自动分类:性别偏见表示分析的应用
作者:Valentin Pelloin,Lena Dodson,Émile Chapuis,Nicolas Hervé,David Doukhan
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文介绍了一个计算框架,旨在描绘性别分布偏见的主题所涵盖的法国电视和广播新闻。我们转录了一个11.7k小时的数据集,于2023年在21个法国频道播出。一个大的语言模型(LLM)被用于在Few-Shot会话模式,以获得这些transmits的主题分类。使用生成的LLM注释,我们探索了一个专门的较小的分类模型的微调,以减少计算成本。为了评估这些模型的性能,我们构建并注释了804个对话的数据集。该数据集免费提供用于研究目的。我们发现,妇女在体育、政治和冲突等科目中的代表性明显不足。相反,在天气、商业和健康等话题上,女性的发言时间超过了所有学科的总体平均水平。我们还观察到私营和公共服务渠道之间的代表性差异。摘要:This paper introduces a computational framework designed to delineate gender distribution biases in topics covered by French TV and radio news. We transcribe a dataset of 11.7k hours, broadcasted in 2023 on 21 French channels. A Large Language Model (LLM) is used in few-shot conversation mode to obtain a topic classification on those transcriptions. Using the generated LLM annotations, we explore the finetuning of a specialized smaller classification model, to reduce the computational cost. To evaluate the performances of these models, we construct and annotate a dataset of 804 dialogues. This dataset is made available free of charge for research purposes. We show that women are notably underrepresented in subjects such as sports, politics and conflicts. Conversely, on topics such as weather, commercials and health, women have more speaking time than their overall average across all subjects. We also observe representations differences between private and public service channels.
【16】 Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
标题: Rasa:在低资源环境中为印度语言构建表达性语音合成系统
作者:Praveen Srinivasa Varadhan,Ashwin Sankar,Giri Raju,Mitesh M. Khapra
备注:Accepted at INTERSPEECH 2024. First two authors listed contributed equally
链接:点击下载PDF文件
摘要:我们发布了Rasa,这是第一个针对任何印度语言的多语言表达性TTS数据集,其中包含10小时的中性语音和1-3小时的表达性语音,用于涵盖3种语言的6种Ekman情绪:阿萨姆语,孟加拉语和泰米尔语。我们的消融研究表明,仅1小时的中性和30分钟的表达数据就可以产生一个公平的系统,如MUSHRA评分所示。将中性数据增加到10小时,使用最少的表达性数据,显著增强了表达能力。这为资源受限的语言提供了一个实用的方法,优先考虑容易获得的中性数据以及少量的表达性数据。我们的音节平衡的数据和汇集的情绪,以提高表现力的重要性。我们还强调了产生特定情绪的挑战,例如,恐惧和惊讶。摘要:We release Rasa, the first multilingual expressive TTS dataset for any Indian language, which contains 10 hours of neutral speech and 1-3 hours of expressive speech for each of the 6 Ekman emotions covering 3 languages: Assamese, Bengali, & Tamil. Our ablation studies reveal that just 1 hour of neutral and 30 minutes of expressive data can yield a Fair system as indicated by MUSHRA scores. Increasing neutral data to 10 hours, with minimal expressive data, significantly enhances expressiveness. This offers a practical recipe for resource-constrained languages, prioritizing easily obtainable neutral data alongside smaller amounts of expressive data. We show the importance of syllabically balanced data and pooling emotions to enhance expressiveness. We also highlight challenges in generating specific emotions, e.g., fear and surprise.
机器翻译,仅供参考
