本文经arXiv每日学术速递授权转载
【1】 Taming Data and Transformers for Audio Generation
标题: 驯服数据和Transformer以生成音频
作者:Moayed Haji-Ali,Willi Menapace,Aliaksandr Siarohin,Guha Balakrishnan,Sergey Tulyakov,Vicente Ordonez
备注:Project Webpage: this https URL
链接:点击下载PDF文件
【2】 Subtractive Training for Music Stem Insertion using Latent Diffusion Models
标题: 使用潜在扩散模型进行音乐干插入的减法训练
作者:Ivan Villa-Renteria,Mason L. Wang,Zachary Shah,Zhe Li,Soohyun Kim,Neelesh Ramachandran,Mert Pilanci
链接:点击下载PDF文件
【3】 Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems
标题: 黑匣子自动语音识别系统的零查询对抗攻击
作者:Zheng Fang,Tao Wang,Lingchen Zhao,Shenyi Zhang,Bowen Li,Yunjie Ge,Qi Li,Chao Shen,Qian Wang
备注:To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024
链接:点击下载PDF文件
【4】 Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
标题: ASA在VC和TTC模型中持续时间预测器改进后的语音识别中的应用
作者:Borodin Kirill Nikolayevich,Kudryavtsev Vasiliy Dmitrievich,Mkrtchian Grach Maratovich,Gorodnichev Mikhail Genadievich,Korzh Dmitrii Sergeevich
链接:点击下载PDF文件
【5】 Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
标题: 利用前端自适应网络增强的ASB对数据包丢失的鲁棒性
作者:Yehoshua Dissen,Shiry Yonash,Israel Cohen,Joseph Keshet
备注:Accepted for publication at INTERSPEECH 2024
链接:点击下载PDF文件
【6】 Factor-Conditioned Speaking-Style Captioning
标题: 因素条件说话风格字幕
作者:Atsushi Ando,Takafumi Moriya,Shota Horiguchi,Ryo Masumura
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【7】 Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
标题: 具有离散语音单元的仅流解码器自动语音识别:试点研究
作者:Peikun Chen,Sining Sun,Changhao Shan,Qing Yang,Lei Xie
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
【8】 A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
标题: 一种用于超越四个词干的音乐源分离的与词干无关的单解码器系统
作者:Karn N. Watcharasupat,Alexander Lerch
备注:Submitted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【9】 Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
标题: 使用纵向语音Transformer自动预测肌萎缩侧索硬化进展
作者:Liming Wang,Yuan Gong,Nauman Dawalatabad,Marco Vilela,Katerina Placek,Brian Tracey,Yishu Gong,Alan Premasiri,Fernando Vieira,James Glass
链接:点击下载PDF文件
【10】 Towards Deep Active Learning in Avian Bioacoustics
标题: 鸟类生物声学中的深度主动学习
作者:Lukas Rauch,Denis Huseljic,Moritz Wirth,Jens Decke,Bernhard Sick,Christoph Scholz
备注:preprint, under review IAL@ECML-PKDD24
链接:点击下载PDF文件
标题: 传统还是创新:现代ASB强制对齐方法的比较
作者:Rotem Rousso,Eyal Cohen,Joseph Keshet,Eleanor Chodroff
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
【2】 DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
标题: DEX-TTC:基于扩散的表达性文本到语音,具有时间变化性风格建模
作者:Hyun Joon Park,Jin Sob Kim,Wooseok Shin,Sung Won Han
备注:Preprint
链接:点击下载PDF文件
【3】 Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
标题: 应用LLM重新筛选临时对话的N-最佳ASB假设:领域适应和上下文结转的影响
作者:Atsunori Ogawa,Naoyuki Kamo,Kohei Matsuura,Takanori Ashihara,Takafumi Moriya,Takatomo Kano,Naohiro Tawara,Marc Delcroix
备注:5 pages
链接:点击下载PDF文件
【4】 DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
标题: DeSTA:通过描述性语音-文本对齐增强语音语言模型
作者:Ke-Han Lu,Zhehuai Chen,Szu-Wei Fu,He Huang,Boris Ginsburg,Yu-Chiang Frank Wang,Hung-yi Lee
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【5】 WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
标题: WavRx:一种疾病不可知、可概括且保护隐私的语音健康诊断模型
作者:Yi Zhu,Tiago Falk
备注:Under review; Model script available at this https URL
链接:点击下载PDF文件
【6】 Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
标题: 扬声器未嵌入:长形式神经扩张的免嵌入方法
作者:Xiang Li,Vivek Govindan,Rohit Paturi,Sundararajan Srinivasan
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【7】 Taming Data and Transformers for Audio Generation
标题: 驯服数据和Transformer以生成音频
作者:Moayed Haji-Ali,Willi Menapace,Aliaksandr Siarohin,Guha Balakrishnan,Sergey Tulyakov,Vicente Ordonez
备注:Project Webpage: this https URL
链接:点击下载PDF文件
【8】 Subtractive Training for Music Stem Insertion using Latent Diffusion Models
标题: 使用潜在扩散模型进行音乐干插入的减法训练
作者:Ivan Villa-Renteria,Mason L. Wang,Zachary Shah,Zhe Li,Soohyun Kim,Neelesh Ramachandran,Mert Pilanci
链接:点击下载PDF文件
【9】 Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems
标题: 黑匣子自动语音识别系统的零查询对抗攻击
作者:Zheng Fang,Tao Wang,Lingchen Zhao,Shenyi Zhang,Bowen Li,Yunjie Ge,Qi Li,Chao Shen,Qian Wang
备注:To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024
链接:点击下载PDF文件
【10】 Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
标题: ASA在VC和TTC模型中持续时间预测器改进后的语音识别中的应用
作者:Borodin Kirill Nikolayevich,Kudryavtsev Vasiliy Dmitrievich,Mkrtchian Grach Maratovich,Gorodnichev Mikhail Genadievich,Korzh Dmitrii Sergeevich
链接:点击下载PDF文件
【11】 Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
标题: 利用前端自适应网络增强的ASB对数据包丢失的鲁棒性
作者:Yehoshua Dissen,Shiry Yonash,Israel Cohen,Joseph Keshet
备注:Accepted for publication at INTERSPEECH 2024
链接:点击下载PDF文件
【12】 Factor-Conditioned Speaking-Style Captioning
标题: 因素条件说话风格字幕
作者:Atsushi Ando,Takafumi Moriya,Shota Horiguchi,Ryo Masumura
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
【13】 Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
标题: 具有离散语音单元的仅流解码器自动语音识别:试点研究
作者:Peikun Chen,Sining Sun,Changhao Shan,Qing Yang,Lei Xie
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
【14】 A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
标题: 一种用于超越四个词干的音乐源分离的与词干无关的单解码器系统
作者:Karn N. Watcharasupat,Alexander Lerch
备注:Submitted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
【15】 Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
标题: 使用纵向语音Transformer自动预测肌萎缩侧索硬化进展
作者:Liming Wang,Yuan Gong,Nauman Dawalatabad,Marco Vilela,Katerina Placek,Brian Tracey,Yishu Gong,Alan Premasiri,Fernando Vieira,James Glass
链接:点击下载PDF文件
【16】 Towards Deep Active Learning in Avian Bioacoustics
标题: 鸟类生物声学中的深度主动学习
作者:Lukas Rauch,Denis Huseljic,Moritz Wirth,Jens Decke,Bernhard Sick,Christoph Scholz
备注:preprint, under review IAL@ECML-PKDD24
链接:点击下载PDF文件
标题: 驯服数据和Transformer以生成音频
作者:Moayed Haji-Ali,Willi Menapace,Aliaksandr Siarohin,Guha Balakrishnan,Sergey Tulyakov,Vicente Ordonez
备注:Project Webpage: this https URL
链接:点击下载PDF文件
摘要:生成环境声音和效果是一个具有挑战性的问题,由于数据稀缺和字幕质量往往不足,因此很难采用大规模的生成模型来完成任务。在这项工作中,我们通过引入两个新的模型来解决这个问题。首先,我们提出了AutoCap,一个高质量和高效的自动音频字幕模型。我们表明,通过利用元数据与音频模态,我们可以大大提高字幕的质量。AutoCap的CIDER得分达到83.2,比最佳字幕模型提高了3.2%,推理速度提高了四倍。然后,我们使用AutoCap为现有数据集的片段添加字幕,获得了761,000个具有高质量字幕的音频片段,形成了最大的可用音频文本数据集。其次,我们提出了GenAu,这是一种可扩展的基于变换器的音频生成架构,我们可以扩展到1.25 B参数,并使用我们的新数据集进行训练。与最先进的音频生成器相比,GenAu在FAD评分中获得了15.7%的显着改善,IS评分为22.7%,CLAP评分为13.5%,这表明与以前的作品相比,生成的音频质量显着提高。这表明,数据的质量往往与其数量一样重要。此外,由于AutoCap是全自动的,因此可以将新的音频样本添加到训练数据集中,从而解锁用于音频合成的更大生成模型的训练。摘要:Generating ambient sounds and effects is a challenging problem due to data scarcity and often insufficient caption quality, making it difficult to employ large-scale generative models for the task. In this work, we tackle the problem by introducing two new models. First, we propose AutoCap, a high-quality and efficient automatic audio captioning model. We show that by leveraging metadata available with the audio modality, we can substantially improve the quality of captions. AutoCap reaches CIDEr score of 83.2, marking a 3.2% improvement from the best available captioning model at four times faster inference speed. We then use AutoCap to caption clips from existing datasets, obtaining 761,000 audio clips with high-quality captions, forming the largest available audio-text dataset. Second, we propose GenAu, a scalable transformer-based audio generation architecture that we scale up to 1.25B parameters and train with our new dataset. When compared to state-of-the-art audio generators, GenAu obtains significant improvements of 15.7% in FAD score, 22.7% in IS, and 13.5% in CLAP score, indicating significantly improved quality of generated audio compared to previous works. This shows that the quality of data is often as important as its quantity. Besides, since AutoCap is fully automatic, new audio samples can be added to the training dataset, unlocking the training of even larger generative models for audio synthesis.
【2】 Subtractive Training for Music Stem Insertion using Latent Diffusion Models
标题: 使用潜在扩散模型进行音乐干插入的减法训练
作者:Ivan Villa-Renteria,Mason L. Wang,Zachary Shah,Zhe Li,Soohyun Kim,Neelesh Ramachandran,Mert Pilanci
链接:点击下载PDF文件
摘要:我们提出了减法训练,一个简单而新颖的方法来合成个别乐器干给定其他文书的背景。该方法将完整音乐混合的数据集与1)缺少特定干的数据集的变体和2)LLM生成的描述如何重新引入缺失干的指令配对。然后,我们微调预训练的文本到音频扩散模型,以在现有词干和文本指令的指导下生成缺失的乐器词干。我们的结果证明了减法训练在创建与现有曲目无缝融合的真实鼓干方面的功效。我们还表明,我们可以使用文本指令来控制插入的干的节奏,动态和流派方面的生成,使我们能够修改一个完整的歌曲中的单一乐器的风格,同时保持其余的乐器相同。最后,我们将这种技术扩展到了非完整的格式,成功地为不完整的安排生成了兼容的低音,鼓和吉他部分。摘要:We present Subtractive Training, a simple and novel method for synthesizing individual musical instrument stems given other instruments as context. This method pairs a dataset of complete music mixes with 1) a variant of the dataset lacking a specific stem, and 2) LLM-generated instructions describing how the missing stem should be reintroduced. We then fine-tune a pretrained text-to-audio diffusion model to generate the missing instrument stem, guided by both the existing stems and the text instruction. Our results demonstrate Subtractive Training's efficacy in creating authentic drum stems that seamlessly blend with the existing tracks. We also show that we can use the text instruction to control the generation of the inserted stem in terms of rhythm, dynamics, and genre, allowing us to modify the style of a single instrument in a full song while keeping the remaining instruments the same. Lastly, we extend this technique to MIDI formats, successfully generating compatible bass, drum, and guitar parts for incomplete arrangements.
【3】 Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems
标题: 黑匣子自动语音识别系统的零查询对抗攻击
作者:Zheng Fang,Tao Wang,Lingchen Zhao,Shenyi Zhang,Bowen Li,Yunjie Ge,Qi Li,Chao Shen,Qian Wang
备注:To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024
链接:点击下载PDF文件
摘要:近年来,人们对ASR系统的脆弱性进行了广泛的研究,揭示了黑盒对抗性示例攻击对现实世界的ASR系统构成了重大威胁。然而,大多数现有的黑盒攻击依赖于对目标ASR的查询,这在不允许查询时是不切实际的。在本文中,我们提出了ZQ攻击,在零查询黑盒设置的ASR系统的基于转移的对抗性攻击。通过对现代ASR技术的全面回顾和分类,我们首先精心选择不同类型的代理ASR来生成对抗性示例。在此之后,ZQ-Attack使用缩放的目标命令音频来消除对抗干扰,使其相对难以察觉,同时保持有效性。随后,为了实现对抗性扰动的高可转移性,我们提出了一种顺序集成优化算法,该算法利用来自其他模型的协作信息,迭代优化每个代理模型上的对抗性扰动。我们进行了大量的实验来评估ZQ攻击。在离线环境下,ZQ-Attack在4个在线语音识别服务上实现了100%的攻击成功率(SRoA),平均信噪比(SNR)为21.91dB,在16个开源ASR上实现了100%的平均SRoA和19.67dB的SNR。对于商业智能语音控制设备,ZQ-Attack在空中设置中也实现了100%的SRoA,平均SNR为15.77dB。摘要:In recent years, extensive research has been conducted on the vulnerability of ASR systems, revealing that black-box adversarial example attacks pose significant threats to real-world ASR systems. However, most existing black-box attacks rely on queries to the target ASRs, which is impractical when queries are not permitted. In this paper, we propose ZQ-Attack, a transfer-based adversarial attack on ASR systems in the zero-query black-box setting. Through a comprehensive review and categorization of modern ASR technologies, we first meticulously select surrogate ASRs of diverse types to generate adversarial examples. Following this, ZQ-Attack initializes the adversarial perturbation with a scaled target command audio, rendering it relatively imperceptible while maintaining effectiveness. Subsequently, to achieve high transferability of adversarial perturbations, we propose a sequential ensemble optimization algorithm, which iteratively optimizes the adversarial perturbation on each surrogate model, leveraging collaborative information from other models. We conduct extensive experiments to evaluate ZQ-Attack. In the over-the-line setting, ZQ-Attack achieves a 100% success rate of attack (SRoA) with an average signal-to-noise ratio (SNR) of 21.91dB on 4 online speech recognition services, and attains an average SRoA of 100% and SNR of 19.67dB on 16 open-source ASRs. For commercial intelligent voice control devices, ZQ-Attack also achieves a 100% SRoA with an average SNR of 15.77dB in the over-the-air setting.
【4】 Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
标题: ASA在VC和TTC模型中持续时间预测器改进后的语音识别中的应用
作者:Borodin Kirill Nikolayevich,Kudryavtsev Vasiliy Dmitrievich,Mkrtchian Grach Maratovich,Gorodnichev Mikhail Genadievich,Korzh Dmitrii Sergeevich
链接:点击下载PDF文件
摘要:生物识别安全领域中最关键的组成部分之一是基于说话人声音的自动说话人验证系统。可以单独使用ASV或与其他AI模型结合使用。在当代,神经网络的质量和数量都在呈指数级增长。同时,有越来越多的系统旨在通过使用语音转换和文本到语音模型来操纵数据。语音生物识别伪造领域受到许多挑战的帮助,包括SSTC,ASVSpoof和SingFake。 本文介绍了一个自动说话人确认系统。我们的模型的主要目标是从目标扬声器的音频中提取嵌入,以获得有关他的声音的重要特征,如音高,能量和音素的持续时间的信息。这些信息用于我们的多语音TTS管道,目前正在开发中。然而,在SSTC挑战中使用该模型来验证其语音经历了语音转换的用户,其EER为20.669。摘要:One of the most crucial components in the field of biometric security is the automatic speaker verification system, which is based on the speaker's voice. It is possible to utilise ASVs in isolation or in conjunction with other AI models. In the contemporary era, the quality and quantity of neural networks are increasing exponentially. Concurrently, there is a growing number of systems that aim to manipulate data through the use of voice conversion and text-to-speech models. The field of voice biometrics forgery is aided by a number of challenges, including SSTC, ASVSpoof, and SingFake. This paper presents a system for automatic speaker verification. The primary objective of our model is the extraction of embeddings from the target speaker's audio in order to obtain information about important characteristics of his voice, such as pitch, energy, and the duration of phonemes. This information is used in our multivoice TTS pipeline, which is currently under development. However, this model was employed within the SSTC challenge to verify users whose voice had undergone voice conversion, where it demonstrated an EER of 20.669.
【5】 Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
标题: 利用前端自适应网络增强的ASB对数据包丢失的鲁棒性
作者:Yehoshua Dissen,Shiry Yonash,Israel Cohen,Joseph Keshet
备注:Accepted for publication at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在自动语音识别(ASR)领域,噪声环境中的鲁棒性仍然是一个重大挑战。最近的ASR模型,如Whisper,已经显示出了希望,但它们在噪声条件下的功效可以进一步增强。本研究的重点是从数据包丢失中恢复,以改善ASR模型的字错误率(WER)。我们建议使用连接到冻结ASR模型的前端自适应网络。自适应网络被训练为通过最小化ASR模型的标准以及增强损失函数来修改损坏的输入频谱。我们的实验表明,根据Whisper标准训练的自适应网络在丢包情况下显著降低了跨域和语言的单词错误率。这种改进是在对Whisper模型的基本性能影响最小的情况下实现的,强调了我们的方法在具有挑战性的声学环境中增强ASR模型的实用性和潜力。摘要:In the realm of automatic speech recognition (ASR), robustness in noisy environments remains a significant challenge. Recent ASR models, such as Whisper, have shown promise, but their efficacy in noisy conditions can be further enhanced. This study is focused on recovering from packet loss to improve the word error rate (WER) of ASR models. We propose using a front-end adaptation network connected to a frozen ASR model. The adaptation network is trained to modify the corrupted input spectrum by minimizing the criteria of the ASR model in addition to an enhancement loss function. Our experiments demonstrate that the adaptation network, trained on Whisper's criteria, notably reduces word error rates across domains and languages in packet-loss scenarios. This improvement is achieved with minimal affect to Whisper model's foundational performance, underscoring our method's practicality and potential in enhancing ASR models in challenging acoustic environments.
【6】 Factor-Conditioned Speaking-Style Captioning
标题: 因素条件说话风格字幕
作者:Atsushi Ando,Takafumi Moriya,Shota Horiguchi,Ryo Masumura
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的说话风格字幕生成方法,产生不同的描述,同时准确地预测说话风格的信息。传统的学习准则直接使用原始字幕,其中不仅包含说话风格的因素条款,但也语法的话,这干扰学习说话风格的信息。为了解决这个问题,我们引入了因子条件字幕(FCC),它首先输出一个表示说话风格因子的短语(例如,性别、音高等),然后生成字幕以确保模型明确地学习说话风格因素。我们还提出了贪婪然后采样(GtS)解码,它首先预测说话风格的因素确定性,以保证语义的准确性,然后生成一个字幕的基础上因素条件采样,以确保多样性。实验表明,FCC优于原来的基于字幕的训练,并与GtS,它产生更多样化的字幕,同时保持风格预测性能。摘要:This paper presents a novel speaking-style captioning method that generates diverse descriptions while accurately predicting speaking-style information. Conventional learning criteria directly use original captions that contain not only speaking-style factor terms but also syntax words, which disturbs learning speaking-style information. To solve this problem, we introduce factor-conditioned captioning (FCC), which first outputs a phrase representing speaking-style factors (e.g., gender, pitch, etc.), and then generates a caption to ensure the model explicitly learns speaking-style factors. We also propose greedy-then-sampling (GtS) decoding, which first predicts speaking-style factors deterministically to guarantee semantic accuracy, and then generates a caption based on factor-conditioned sampling to ensure diversity. Experiments show that FCC outperforms the original caption-based training, and with GtS, it generates more diverse captions while keeping style prediction performance.
【7】 Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
标题: 具有离散语音单元的仅流解码器自动语音识别:试点研究
作者:Peikun Chen,Sining Sun,Changhao Shan,Qing Yang,Lei Xie
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:统一的语音文本模型,如SpeechGPT,VioLA和AudioPaLM,在各种语音相关的任务中表现出令人印象深刻的性能,特别是在自动语音识别(ASR)中。这些模型通常采用统一的方法来对离散的语音和文本令牌进行建模,然后训练仅解码器的Transformer。然而,它们都是为非流式ASR任务而设计的,其中在解码期间需要整个语音话语。因此,我们引入了一个专为流识别设计的仅解码器模型,该模型包含专用的边界令牌以促进流识别,并在训练阶段采用因果注意掩蔽。此外,我们引入右块注意和各种数据增强技术,以提高模型的上下文建模能力。在实现流式语音识别的同时,在AISHELL-1和-2数据集上的实验证明了我们的流式方法与非流式解码器的竞争性能。摘要:Unified speech-text models like SpeechGPT, VioLA, and AudioPaLM have shown impressive performance across various speech-related tasks, especially in Automatic Speech Recognition (ASR). These models typically adopt a unified method to model discrete speech and text tokens, followed by training a decoder-only transformer. However, they are all designed for non-streaming ASR tasks, where the entire speech utterance is needed during decoding. Hence, we introduce a decoder-only model exclusively designed for streaming recognition, incorporating a dedicated boundary token to facilitate streaming recognition and employing causal attention masking during the training phase. Furthermore, we introduce right-chunk attention and various data augmentation techniques to improve the model's contextual modeling abilities. While achieving streaming speech recognition, experiments on the AISHELL-1 and -2 datasets demonstrate the competitive performance of our streaming approach with non-streaming decoder-only counterparts.
【8】 A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
标题: 一种用于超越四个词干的音乐源分离的与词干无关的单解码器系统
作者:Karn N. Watcharasupat,Alexander Lerch
备注:Submitted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:尽管最近在音频源分离的多个子任务方面取得了重大进展,但很少有音乐源分离系统支持四干人声、鼓、低音和其他(VDBO)设置之外的分离。在极少数当前系统中,除了这种设置之外,还支持源分离,大多数系统继续依赖于只能支持固定的预定义词干集的不灵活的解码器设置。在这些不灵活的系统中增加茎支持相应地需要增加计算复杂性,使得这些系统的扩展在计算上对于长尾仪器是不可行的。在这项工作中,我们提出宴会,一个系统,允许源分离的多个干使用一个解码器。一个bandsplit源分离模型扩展到工作在一个基于查询的设置与乐器识别PaSST模型。在MoisesDB数据集上,Banquet在只有24.9 M可训练参数的情况下,在VDBO茎上接近明显更复杂的6茎混合Transformer Demucs的性能水平,并在吉他和钢琴上表现出色。基于查询的设置允许分离狭窄的乐器类别,如干净的原声吉他,并可以成功地应用于提取不太常见的茎,如芦苇和器官。可在https: github.com kwatcharasupat query-bandit上获得实现。摘要:Despite significant recent progress across multiple subtasks of audio source separation, few music source separation systems support separation beyond the four-stem vocals, drums, bass, and other (VDBO) setup. Of the very few current systems that support source separation beyond this setup, most continue to rely on an inflexible decoder setup that can only support a fixed pre-defined set of stems. Increasing stem support in these inflexible systems correspondingly requires increasing computational complexity, rendering extensions of these systems computationally infeasible for long-tail instruments. In this work, we propose Banquet, a system that allows source separation of multiple stems using just one decoder. A bandsplit source separation model is extended to work in a query-based setup in tandem with a music instrument recognition PaSST model. On the MoisesDB dataset, Banquet, at only 24.9 M trainable parameters, approached the performance level of the significantly more complex 6-stem Hybrid Transformer Demucs on VDBO stems and outperformed it on guitar and piano. The query-based setup allows for the separation of narrow instrument classes such as clean acoustic guitars, and can be successfully applied to the extraction of less common stems such as reeds and organs. Implementation is available at https: github.com kwatcharasupat query-bandit.
【9】 Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
标题: 使用纵向语音Transformer自动预测肌萎缩侧索硬化进展
作者:Liming Wang,Yuan Gong,Nauman Dawalatabad,Marco Vilela,Katerina Placek,Brian Tracey,Yishu Gong,Alan Premasiri,Fernando Vieira,James Glass
链接:点击下载PDF文件
摘要:肌萎缩侧索硬化症(ALS)疾病进展的自动预测提供了一个更有效和客观的替代方法比手动。我们提出了ALS纵向语音Transformer(ALST),一种基于神经网络的ALS疾病进展的自动预测器,从ALS患者的纵向语音记录。通过利用高质量的预训练语音特征和录音中的纵向信息,我们的最佳模型达到了91.0 % AUC,在ALS TDI数据集上比以前的最佳模型提高了5.6 %。仔细分析表明,ALST能够对ALS进展进行细粒度和可解释的预测,特别是用于区分罕见和严重病例。代码是公开的。摘要:Automatic prediction of amyotrophic lateral sclerosis (ALS) disease progression provides a more efficient and objective alternative than manual approaches. We propose ALS longitudinal speech transformer (ALST), a neural network-based automatic predictor of ALS disease progression from longitudinal speech recordings of ALS patients. By taking advantage of high-quality pretrained speech features and longitudinal information in the recordings, our best model achieves 91.0 % AUC, improving upon the previous best model by 5.6 % relative on the ALS TDI dataset. Careful analysis reveals that ALST is capable of fine-grained and interpretable predictions of ALS progression, especially for distinguishing between rarer and more severe cases. Code is publicly available.
【10】 Towards Deep Active Learning in Avian Bioacoustics
标题: 鸟类生物声学中的深度主动学习
作者:Lukas Rauch,Denis Huseljic,Moritz Wirth,Jens Decke,Bernhard Sick,Christoph Scholz
备注:preprint, under review IAL@ECML-PKDD24
链接:点击下载PDF文件
摘要:鸟类生物声学中的被动声学监测(PAM)能够以最小的自然栖息地破坏进行成本效益和广泛的数据收集。尽管在计算鸟类生物声学方面取得了进展,但深度学习模型在实际PAM场景中适应不同环境方面仍然面临挑战。这主要是由于注释的稀缺性,这需要人类专家的劳动密集型工作。主动学习(AL)通过查询最具信息量的实例进行标记,降低了注释成本,并加快了对不同场景的适应速度。本文概述了深入的人工智能方法,介绍了关键挑战,并进行了小规模的试点研究。摘要:Passive acoustic monitoring (PAM) in avian bioacoustics enables cost-effective and extensive data collection with minimal disruption to natural habitats. Despite advancements in computational avian bioacoustics, deep learning models continue to encounter challenges in adapting to diverse environments in practical PAM scenarios. This is primarily due to the scarcity of annotations, which requires labor-intensive efforts from human experts. Active learning (AL) reduces annotation cost and speed ups adaption to diverse scenarios by querying the most informative instances for labeling. This paper outlines a deep AL approach, introduces key challenges, and conducts a small-scale pilot study.
eess.AS音频处理
【1】 Tradition or Innovation: A Comparison of Modern ASR Methods for Forced Alignment标题: 传统还是创新:现代ASB强制对齐方法的比较
作者:Rotem Rousso,Eyal Cohen,Joseph Keshet,Eleanor Chodroff
Journal-ref:Interspeech 2024
链接:点击下载PDF文件
摘要:强制对齐是语音研究中的一个关键问题,它通过自动的时间对齐来实现语音信号与文本的对齐。尽管语音技术正朝着端到端架构发展,但FA仍然主要通过经典的GMM-HMM声学模型来实现。这项工作直接比较了领先的自动语音识别(ASR)方法WhisperX和大规模多语言语音识别(MMS)与基于Kaldi的GMM-HMM系统Montreal Forced Aligner(MFA)的对齐性能。在手动对齐的TIMIT和Buckeye数据集上评估性能,仅对WhisperX和MMS正确识别的单词进行比较。MFA的表现优于WhisperX和MMS,揭示了现代ASR系统的缺点。这些发现强调了在强制调整方面取得进展的必要性,并强调了将传统专业知识与现代创新相结合以促进进步的重要性。索引术语:强制对齐、音素对齐、单词对齐摘要:Forced alignment (FA) plays a key role in speech research through the automatic time alignment of speech signals with corresponding text transcriptions. Despite the move towards end-to-end architectures for speech technology, FA is still dominantly achieved through a classic GMM-HMM acoustic model. This work directly compares alignment performance from leading automatic speech recognition (ASR) methods, WhisperX and Massively Multilingual Speech Recognition (MMS), against a Kaldi-based GMM-HMM system, the Montreal Forced Aligner (MFA). Performance was assessed on the manually aligned TIMIT and Buckeye datasets, with comparisons conducted only on words correctly recognized by WhisperX and MMS. The MFA outperformed both WhisperX and MMS, revealing a shortcoming of modern ASR systems. These findings highlight the need for advancements in forced alignment and emphasize the importance of integrating traditional expertise with modern innovation to foster progress. Index Terms: forced alignment, phoneme alignment, word alignment
【2】 DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
标题: DEX-TTC:基于扩散的表达性文本到语音,具有时间变化性风格建模
作者:Hyun Joon Park,Jin Sob Kim,Wooseok Shin,Sung Won Han
备注:Preprint
链接:点击下载PDF文件
摘要:基于参考语音的表达型文语转换(TTS)技术在合成自然语音方面得到了广泛的研究,但在获得良好的表达风格和提高模型泛化能力方面存在局限性。在这项研究中,我们提出了基于扩散的表达TTS(DEX-TTS),一个声学模型设计的基于参考的语音合成增强风格的表示。基于一个通用的扩散TTS框架,DEX-TTS包括编码器和适配器来处理从参考语音中提取的风格。关键的创新包括区分风格的时不变和时变的类别,有效的风格提取,以及具有高泛化能力的编码器和适配器的设计。此外,我们引入重叠补丁和卷积频率补丁嵌入策略,以改善基于DiT的扩散网络的TTS。DEX-TTS在英语多说话人和情感多说话人数据集的客观和主观评价方面表现出色,而不依赖于预训练策略。最后,在单说话人数据集上的一般TTS的比较结果验证了我们的增强扩散骨干的有效性。演示可在这里。摘要:Expressive Text-to-Speech (TTS) using reference speech has been studied extensively to synthesize natural speech, but there are limitations to obtaining well-represented styles and improving model generalization ability. In this study, we present Diffusion-based EXpressive TTS (DEX-TTS), an acoustic model designed for reference-based speech synthesis with enhanced style representations. Based on a general diffusion TTS framework, DEX-TTS includes encoders and adapters to handle styles extracted from reference speech. Key innovations contain the differentiation of styles into time-invariant and time-variant categories for effective style extraction, as well as the design of encoders and adapters with high generalization ability. In addition, we introduce overlapping patchify and convolution-frequency patch embedding strategies to improve DiT-based diffusion networks for TTS. DEX-TTS yields outstanding performance in terms of objective and subjective evaluation in English multi-speaker and emotional multi-speaker datasets, without relying on pre-training strategies. Lastly, the comparison results for the general TTS on a single-speaker dataset verify the effectiveness of our enhanced diffusion backbone. Demos are available here.
【3】 Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
标题: 应用LLM重新筛选临时对话的N-最佳ASB假设:领域适应和上下文结转的影响
作者:Atsunori Ogawa,Naoyuki Kamo,Kohei Matsuura,Takanori Ashihara,Takafumi Moriya,Takatomo Kano,Naohiro Tawara,Marc Delcroix
备注:5 pages
链接:点击下载PDF文件
摘要:大语言模型(LLM)已成功应用于自动语音识别(ASR)假设的再评分。然而,他们的能力rescore ASR假设的随意交谈还没有得到充分的探讨。在这项研究中,我们通过使用Llama 2对CHiME-7远距离ASR(DASR)任务执行N-最佳ASR假设重新评分来揭示这一点。Llama 2是最具代表性的LLM之一,CHiME-7 DASR任务提供了多个参与者之间的随意对话的数据集。我们调查域适应的LLM和上下文结转时,执行N-最好的重新评分的影响。实验结果表明,即使没有域自适应,Llama 2也优于标准大小的域自适应Transformer-LM,特别是在使用长上下文时。域自适应缩短了Llama 2实现其最佳性能所需的上下文长度,即,它降低了Llama 2的计算成本。摘要:Large language models (LLMs) have been successfully applied for rescoring automatic speech recognition (ASR) hypotheses. However, their ability to rescore ASR hypotheses of casual conversations has not been sufficiently explored. In this study, we reveal it by performing N-best ASR hypotheses rescoring using Llama2 on the CHiME-7 distant ASR (DASR) task. Llama2 is one of the most representative LLMs, and the CHiME-7 DASR task provides datasets of casual conversations between multiple participants. We investigate the effects of domain adaptation of the LLM and context carry-over when performing N-best rescoring. Experimental results show that, even without domain adaptation, Llama2 outperforms a standard-size domain-adapted Transformer-LM, especially when using a long context. Domain adaptation shortens the context length needed with Llama2 to achieve its best performance, i.e., it reduces the computational cost of Llama2.
【4】 DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
标题: DeSTA:通过描述性语音-文本对齐增强语音语言模型
作者:Ke-Han Lu,Zhehuai Chen,Szu-Wei Fu,He Huang,Boris Ginsburg,Yu-Chiang Frank Wang,Hung-yi Lee
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:最近的语音语言模型(SLM)通常包含预训练的语音模型,以扩展大型语言模型(LLM)的功能。在本文中,我们提出了一种描述性的语音文本对齐方法,利用语音字幕来弥合语音和文本模态之间的差距,使SLM能够解释和生成全面的自然语言描述,从而促进理解语音中的语言和非语言特征的能力。增强所提出的方法,我们的模型表现出卓越的性能动态SUPERB基准,特别是在推广到看不见的任务。此外,我们发现,对齐的模型表现出zero-shot后续的能力,没有明确的语音指令调整。这些研究结果强调了通过整合丰富的描述性语音字幕来重塑预防性后续SLM的潜力。摘要:Recent speech language models (SLMs) typically incorporate pre-trained speech models to extend the capabilities from large language models (LLMs). In this paper, we propose a Descriptive Speech-Text Alignment approach that leverages speech captioning to bridge the gap between speech and text modalities, enabling SLMs to interpret and generate comprehensive natural language descriptions, thereby facilitating the capability to understand both linguistic and non-linguistic features in speech. Enhanced with the proposed approach, our model demonstrates superior performance on the Dynamic-SUPERB benchmark, particularly in generalizing to unseen tasks. Moreover, we discover that the aligned model exhibits a zero-shot instruction-following capability without explicit speech instruction tuning. These findings highlight the potential to reshape instruction-following SLMs by incorporating rich, descriptive speech captions.
【5】 WavRx: a Disease-Agnostic, Generalizable, and Privacy-Preserving Speech Health Diagnostic Model
标题: WavRx:一种疾病不可知、可概括且保护隐私的语音健康诊断模型
作者:Yi Zhu,Tiago Falk
备注:Under review; Model script available at this https URL
链接:点击下载PDF文件
摘要:众所周知,语音具有与健康相关的属性,已成为远程和长期健康监测的新场所。然而,现有的模型通常是针对特定类型的疾病量身定制的,并且已被证明缺乏跨数据集的通用性。此外,最近也有人担心健康嵌入会导致说话人身份泄露。为了减轻这些限制,我们提出了WavRx,一种语音健康诊断模型,它从通用语音表示中捕获呼吸和发音相关的动态。我们在六个病理语音数据集上进行的域内和跨域实验表明,WavRx是一种新的最先进的健康诊断模型。此外,我们表明,在WavRx健康嵌入中所涉及的说话人身份的数量显着减少,在训练过程中没有额外的指导。对该模型进行了深入的分析,从而提供了其改进的泛化能力和隐私保护能力的生理解释。摘要:Speech is known to carry health-related attributes, which has emerged as a novel venue for remote and long-term health monitoring. However, existing models are usually tailored for a specific type of disease, and have been shown to lack generalizability across datasets. Furthermore, concerns have been raised recently towards the leakage of speaker identity from health embeddings. To mitigate these limitations, we propose WavRx, a speech health diagnostics model that captures the respiration and articulation related dynamics from a universal speech representation. Our in-domain and cross-domain experiments on six pathological speech datasets demonstrate WavRx as a new state-of-the-art health diagnostic model. Furthermore, we show that the amount of speaker identity entailed in the WavRx health embeddings is significantly reduced without extra guidance during training. An in-depth analysis of the model was performed, thus providing physiological interpretation of its improved generalizability and privacy-preserving ability.
【6】 Speakers Unembedded: Embedding-free Approach to Long-form Neural Diarization
标题: 扬声器未嵌入:长形式神经扩张的免嵌入方法
作者:Xiang Li,Vivek Govindan,Rohit Paturi,Sundararajan Srinivasan
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:端到端神经日志化(EEND)模型比传统的基于嵌入的说话者日志化(SD)方法提供了显着的改进,但在推广到具有大量说话者的长格式音频方面存在不足。EEND向量聚类方法通过将本地EEND与来自本地窗口的说话者嵌入的全局聚类相结合来缓解这一点,但这需要在EEND模块旁边的额外的说话者嵌入框架。在本文中,我们提出了一个新的框架,适用于本地和全球的长形式的音频没有单独的扬声器嵌入的EEND。该方法在Callhome美式英语和RT 03-CTS数据集上分别实现了比传统的1遍EEND显著的相对DER降低13%和10%,并且在不需要额外的说话人嵌入的情况下,比EEND向量聚类有了边际改进。此外,我们还讨论了我们提出的框架的计算复杂性,并探讨了减少处理时间的策略。摘要:End-to-end neural diarization (EEND) models offer significant improvements over traditional embedding-based Speaker Diarization (SD) approaches but falls short on generalizing to long-form audio with large number of speakers. EEND-vector-clustering method mitigates this by combining local EEND with global clustering of speaker embeddings from local windows, but this requires an additional speaker embedding framework alongside the EEND module. In this paper, we propose a novel framework applying EEND both locally and globally for long-form audio without separate speaker embeddings. This approach achieves significant relative DER reduction of 13% and 10% over the conventional 1-pass EEND on Callhome American English and RT03-CTS datasets respectively and marginal improvements over EEND-vector-clustering without the need for additional speaker embeddings. Furthermore, we discuss the computational complexity of our proposed framework and explore strategies for reducing processing times.
【7】 Taming Data and Transformers for Audio Generation
标题: 驯服数据和Transformer以生成音频
作者:Moayed Haji-Ali,Willi Menapace,Aliaksandr Siarohin,Guha Balakrishnan,Sergey Tulyakov,Vicente Ordonez
备注:Project Webpage: this https URL
链接:点击下载PDF文件
摘要:生成环境声音和效果是一个具有挑战性的问题,由于数据稀缺和字幕质量往往不足,因此很难采用大规模的生成模型来完成任务。在这项工作中,我们通过引入两个新的模型来解决这个问题。首先,我们提出了AutoCap,一个高质量和高效的自动音频字幕模型。我们表明,通过利用元数据与音频模态,我们可以大大提高字幕的质量。AutoCap的CIDER得分达到83.2,比最佳字幕模型提高了3.2%,推理速度提高了四倍。然后,我们使用AutoCap为现有数据集的片段添加字幕,获得了761,000个具有高质量字幕的音频片段,形成了最大的可用音频文本数据集。其次,我们提出了GenAu,这是一种可扩展的基于变换器的音频生成架构,我们可以扩展到1.25 B参数,并使用我们的新数据集进行训练。与最先进的音频生成器相比,GenAu在FAD评分中获得了15.7%的显着改善,IS评分为22.7%,CLAP评分为13.5%,这表明与以前的作品相比,生成的音频质量显着提高。这表明,数据的质量往往与其数量一样重要。此外,由于AutoCap是全自动的,因此可以将新的音频样本添加到训练数据集中,从而解锁用于音频合成的更大生成模型的训练。摘要:Generating ambient sounds and effects is a challenging problem due to data scarcity and often insufficient caption quality, making it difficult to employ large-scale generative models for the task. In this work, we tackle the problem by introducing two new models. First, we propose AutoCap, a high-quality and efficient automatic audio captioning model. We show that by leveraging metadata available with the audio modality, we can substantially improve the quality of captions. AutoCap reaches CIDEr score of 83.2, marking a 3.2% improvement from the best available captioning model at four times faster inference speed. We then use AutoCap to caption clips from existing datasets, obtaining 761,000 audio clips with high-quality captions, forming the largest available audio-text dataset. Second, we propose GenAu, a scalable transformer-based audio generation architecture that we scale up to 1.25B parameters and train with our new dataset. When compared to state-of-the-art audio generators, GenAu obtains significant improvements of 15.7% in FAD score, 22.7% in IS, and 13.5% in CLAP score, indicating significantly improved quality of generated audio compared to previous works. This shows that the quality of data is often as important as its quantity. Besides, since AutoCap is fully automatic, new audio samples can be added to the training dataset, unlocking the training of even larger generative models for audio synthesis.
【8】 Subtractive Training for Music Stem Insertion using Latent Diffusion Models
标题: 使用潜在扩散模型进行音乐干插入的减法训练
作者:Ivan Villa-Renteria,Mason L. Wang,Zachary Shah,Zhe Li,Soohyun Kim,Neelesh Ramachandran,Mert Pilanci
链接:点击下载PDF文件
摘要:我们提出了减法训练,一个简单而新颖的方法来合成个别乐器干给定其他文书的背景。该方法将完整音乐混合的数据集与1)缺少特定干的数据集的变体和2)LLM生成的描述如何重新引入缺失干的指令配对。然后,我们微调预训练的文本到音频扩散模型,以在现有词干和文本指令的指导下生成缺失的乐器词干。我们的结果证明了减法训练在创建与现有曲目无缝融合的真实鼓干方面的功效。我们还表明,我们可以使用文本指令来控制插入的干的节奏,动态和流派方面的生成,使我们能够修改一个完整的歌曲中的单一乐器的风格,同时保持其余的乐器相同。最后,我们将这种技术扩展到了非完整的格式,成功地为不完整的安排生成了兼容的低音,鼓和吉他部分。摘要:We present Subtractive Training, a simple and novel method for synthesizing individual musical instrument stems given other instruments as context. This method pairs a dataset of complete music mixes with 1) a variant of the dataset lacking a specific stem, and 2) LLM-generated instructions describing how the missing stem should be reintroduced. We then fine-tune a pretrained text-to-audio diffusion model to generate the missing instrument stem, guided by both the existing stems and the text instruction. Our results demonstrate Subtractive Training's efficacy in creating authentic drum stems that seamlessly blend with the existing tracks. We also show that we can use the text instruction to control the generation of the inserted stem in terms of rhythm, dynamics, and genre, allowing us to modify the style of a single instrument in a full song while keeping the remaining instruments the same. Lastly, we extend this technique to MIDI formats, successfully generating compatible bass, drum, and guitar parts for incomplete arrangements.
【9】 Zero-Query Adversarial Attack on Black-box Automatic Speech Recognition Systems
标题: 黑匣子自动语音识别系统的零查询对抗攻击
作者:Zheng Fang,Tao Wang,Lingchen Zhao,Shenyi Zhang,Bowen Li,Yunjie Ge,Qi Li,Chao Shen,Qian Wang
备注:To appear in the Proceedings of The ACM Conference on Computer and Communications Security (CCS), 2024
链接:点击下载PDF文件
摘要:近年来,人们对ASR系统的脆弱性进行了广泛的研究,揭示了黑盒对抗性示例攻击对现实世界的ASR系统构成了重大威胁。然而,大多数现有的黑盒攻击依赖于对目标ASR的查询,这在不允许查询时是不切实际的。在本文中,我们提出了ZQ攻击,在零查询黑盒设置的ASR系统的基于转移的对抗性攻击。通过对现代ASR技术的全面回顾和分类,我们首先精心选择不同类型的代理ASR来生成对抗性示例。在此之后,ZQ-Attack使用缩放的目标命令音频来消除对抗干扰,使其相对难以察觉,同时保持有效性。随后,为了实现对抗性扰动的高可转移性,我们提出了一种顺序集成优化算法,该算法利用来自其他模型的协作信息,迭代优化每个代理模型上的对抗性扰动。我们进行了大量的实验来评估ZQ攻击。在离线环境下,ZQ-Attack在4个在线语音识别服务上实现了100%的攻击成功率(SRoA),平均信噪比(SNR)为21.91dB,在16个开源ASR上实现了100%的平均SRoA和19.67dB的SNR。对于商业智能语音控制设备,ZQ-Attack在空中设置中也实现了100%的SRoA,平均SNR为15.77dB。摘要:In recent years, extensive research has been conducted on the vulnerability of ASR systems, revealing that black-box adversarial example attacks pose significant threats to real-world ASR systems. However, most existing black-box attacks rely on queries to the target ASRs, which is impractical when queries are not permitted. In this paper, we propose ZQ-Attack, a transfer-based adversarial attack on ASR systems in the zero-query black-box setting. Through a comprehensive review and categorization of modern ASR technologies, we first meticulously select surrogate ASRs of diverse types to generate adversarial examples. Following this, ZQ-Attack initializes the adversarial perturbation with a scaled target command audio, rendering it relatively imperceptible while maintaining effectiveness. Subsequently, to achieve high transferability of adversarial perturbations, we propose a sequential ensemble optimization algorithm, which iteratively optimizes the adversarial perturbation on each surrogate model, leveraging collaborative information from other models. We conduct extensive experiments to evaluate ZQ-Attack. In the over-the-line setting, ZQ-Attack achieves a 100% success rate of attack (SRoA) with an average signal-to-noise ratio (SNR) of 21.91dB on 4 online speech recognition services, and attains an average SRoA of 100% and SNR of 19.67dB on 16 open-source ASRs. For commercial intelligent voice control devices, ZQ-Attack also achieves a 100% SRoA with an average SNR of 15.77dB in the over-the-air setting.
【10】 Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
标题: ASA在VC和TTC模型中持续时间预测器改进后的语音识别中的应用
作者:Borodin Kirill Nikolayevich,Kudryavtsev Vasiliy Dmitrievich,Mkrtchian Grach Maratovich,Gorodnichev Mikhail Genadievich,Korzh Dmitrii Sergeevich
链接:点击下载PDF文件
摘要:生物识别安全领域中最关键的组成部分之一是基于说话人声音的自动说话人验证系统。可以单独使用ASV或与其他AI模型结合使用。在当代,神经网络的质量和数量都在呈指数级增长。同时,有越来越多的系统旨在通过使用语音转换和文本到语音模型来操纵数据。语音生物识别伪造领域受到许多挑战的帮助,包括SSTC,ASVSpoof和SingFake。 本文介绍了一个自动说话人确认系统。我们的模型的主要目标是从目标扬声器的音频中提取嵌入,以获得有关他的声音的重要特征,如音高,能量和音素的持续时间的信息。这些信息用于我们的多语音TTS管道,目前正在开发中。然而,在SSTC挑战中使用该模型来验证其语音经历了语音转换的用户,其EER为20.669。摘要:One of the most crucial components in the field of biometric security is the automatic speaker verification system, which is based on the speaker's voice. It is possible to utilise ASVs in isolation or in conjunction with other AI models. In the contemporary era, the quality and quantity of neural networks are increasing exponentially. Concurrently, there is a growing number of systems that aim to manipulate data through the use of voice conversion and text-to-speech models. The field of voice biometrics forgery is aided by a number of challenges, including SSTC, ASVSpoof, and SingFake. This paper presents a system for automatic speaker verification. The primary objective of our model is the extraction of embeddings from the target speaker's audio in order to obtain information about important characteristics of his voice, such as pitch, energy, and the duration of phonemes. This information is used in our multivoice TTS pipeline, which is currently under development. However, this model was employed within the SSTC challenge to verify users whose voice had undergone voice conversion, where it demonstrated an EER of 20.669.
【11】 Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
标题: 利用前端自适应网络增强的ASB对数据包丢失的鲁棒性
作者:Yehoshua Dissen,Shiry Yonash,Israel Cohen,Joseph Keshet
备注:Accepted for publication at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在自动语音识别(ASR)领域,噪声环境中的鲁棒性仍然是一个重大挑战。最近的ASR模型,如Whisper,已经显示出了希望,但它们在噪声条件下的功效可以进一步增强。本研究的重点是从数据包丢失中恢复,以改善ASR模型的字错误率(WER)。我们建议使用连接到冻结ASR模型的前端自适应网络。自适应网络被训练为通过最小化ASR模型的标准以及增强损失函数来修改损坏的输入频谱。我们的实验表明,根据Whisper标准训练的自适应网络在丢包情况下显著降低了跨域和语言的单词错误率。这种改进是在对Whisper模型的基本性能影响最小的情况下实现的,强调了我们的方法在具有挑战性的声学环境中增强ASR模型的实用性和潜力。摘要:In the realm of automatic speech recognition (ASR), robustness in noisy environments remains a significant challenge. Recent ASR models, such as Whisper, have shown promise, but their efficacy in noisy conditions can be further enhanced. This study is focused on recovering from packet loss to improve the word error rate (WER) of ASR models. We propose using a front-end adaptation network connected to a frozen ASR model. The adaptation network is trained to modify the corrupted input spectrum by minimizing the criteria of the ASR model in addition to an enhancement loss function. Our experiments demonstrate that the adaptation network, trained on Whisper's criteria, notably reduces word error rates across domains and languages in packet-loss scenarios. This improvement is achieved with minimal affect to Whisper model's foundational performance, underscoring our method's practicality and potential in enhancing ASR models in challenging acoustic environments.
【12】 Factor-Conditioned Speaking-Style Captioning
标题: 因素条件说话风格字幕
作者:Atsushi Ando,Takafumi Moriya,Shota Horiguchi,Ryo Masumura
备注:Accepted to Interspeech 2024
链接:点击下载PDF文件
摘要:本文提出了一种新的说话风格字幕生成方法,产生不同的描述,同时准确地预测说话风格的信息。传统的学习准则直接使用原始字幕,其中不仅包含说话风格的因素条款,但也语法的话,这干扰学习说话风格的信息。为了解决这个问题,我们引入了因子条件字幕(FCC),它首先输出一个表示说话风格因子的短语(例如,性别、音高等),然后生成字幕以确保模型明确地学习说话风格因素。我们还提出了贪婪然后采样(GtS)解码,它首先预测说话风格的因素确定性,以保证语义的准确性,然后生成一个字幕的基础上因素条件采样,以确保多样性。实验表明,FCC优于原来的基于字幕的训练,并与GtS,它产生更多样化的字幕,同时保持风格预测性能。摘要:This paper presents a novel speaking-style captioning method that generates diverse descriptions while accurately predicting speaking-style information. Conventional learning criteria directly use original captions that contain not only speaking-style factor terms but also syntax words, which disturbs learning speaking-style information. To solve this problem, we introduce factor-conditioned captioning (FCC), which first outputs a phrase representing speaking-style factors (e.g., gender, pitch, etc.), and then generates a caption to ensure the model explicitly learns speaking-style factors. We also propose greedy-then-sampling (GtS) decoding, which first predicts speaking-style factors deterministically to guarantee semantic accuracy, and then generates a caption based on factor-conditioned sampling to ensure diversity. Experiments show that FCC outperforms the original caption-based training, and with GtS, it generates more diverse captions while keeping style prediction performance.
【13】 Streaming Decoder-Only Automatic Speech Recognition with Discrete Speech Units: A Pilot Study
标题: 具有离散语音单元的仅流解码器自动语音识别:试点研究
作者:Peikun Chen,Sining Sun,Changhao Shan,Qing Yang,Lei Xie
备注:Accepted for Interspeech 2024
链接:点击下载PDF文件
摘要:统一的语音文本模型,如SpeechGPT,VioLA和AudioPaLM,在各种语音相关的任务中表现出令人印象深刻的性能,特别是在自动语音识别(ASR)中。这些模型通常采用统一的方法来对离散的语音和文本令牌进行建模,然后训练仅解码器的Transformer。然而,它们都是为非流式ASR任务而设计的,其中在解码期间需要整个语音话语。因此,我们引入了一个专为流识别设计的仅解码器模型,该模型包含专用的边界令牌以促进流识别,并在训练阶段采用因果注意掩蔽。此外,我们引入右块注意和各种数据增强技术,以提高模型的上下文建模能力。在实现流式语音识别的同时,在AISHELL-1和-2数据集上的实验证明了我们的流式方法与非流式解码器的竞争性能。摘要:Unified speech-text models like SpeechGPT, VioLA, and AudioPaLM have shown impressive performance across various speech-related tasks, especially in Automatic Speech Recognition (ASR). These models typically adopt a unified method to model discrete speech and text tokens, followed by training a decoder-only transformer. However, they are all designed for non-streaming ASR tasks, where the entire speech utterance is needed during decoding. Hence, we introduce a decoder-only model exclusively designed for streaming recognition, incorporating a dedicated boundary token to facilitate streaming recognition and employing causal attention masking during the training phase. Furthermore, we introduce right-chunk attention and various data augmentation techniques to improve the model's contextual modeling abilities. While achieving streaming speech recognition, experiments on the AISHELL-1 and -2 datasets demonstrate the competitive performance of our streaming approach with non-streaming decoder-only counterparts.
【14】 A Stem-Agnostic Single-Decoder System for Music Source Separation Beyond Four Stems
标题: 一种用于超越四个词干的音乐源分离的与词干无关的单解码器系统
作者:Karn N. Watcharasupat,Alexander Lerch
备注:Submitted to the 25th International Society for Music Information Retrieval Conference (ISMIR 2024)
链接:点击下载PDF文件
摘要:尽管最近在音频源分离的多个子任务方面取得了重大进展,但很少有音乐源分离系统支持四干人声、鼓、低音和其他(VDBO)设置之外的分离。在极少数当前系统中,除了这种设置之外,还支持源分离,大多数系统继续依赖于只能支持固定的预定义词干集的不灵活的解码器设置。在这些不灵活的系统中增加茎支持相应地需要增加计算复杂性,使得这些系统的扩展在计算上对于长尾仪器是不可行的。在这项工作中,我们提出宴会,一个系统,允许源分离的多个干使用一个解码器。一个bandsplit源分离模型扩展到工作在一个基于查询的设置与乐器识别PaSST模型。在MoisesDB数据集上,Banquet在只有24.9 M可训练参数的情况下,在VDBO茎上接近明显更复杂的6茎混合Transformer Demucs的性能水平,并在吉他和钢琴上表现出色。基于查询的设置允许分离狭窄的乐器类别,如干净的原声吉他,并可以成功地应用于提取不太常见的茎,如芦苇和器官。可在https: github.com kwatcharasupat query-bandit上获得实现。摘要:Despite significant recent progress across multiple subtasks of audio source separation, few music source separation systems support separation beyond the four-stem vocals, drums, bass, and other (VDBO) setup. Of the very few current systems that support source separation beyond this setup, most continue to rely on an inflexible decoder setup that can only support a fixed pre-defined set of stems. Increasing stem support in these inflexible systems correspondingly requires increasing computational complexity, rendering extensions of these systems computationally infeasible for long-tail instruments. In this work, we propose Banquet, a system that allows source separation of multiple stems using just one decoder. A bandsplit source separation model is extended to work in a query-based setup in tandem with a music instrument recognition PaSST model. On the MoisesDB dataset, Banquet, at only 24.9 M trainable parameters, approached the performance level of the significantly more complex 6-stem Hybrid Transformer Demucs on VDBO stems and outperformed it on guitar and piano. The query-based setup allows for the separation of narrow instrument classes such as clean acoustic guitars, and can be successfully applied to the extraction of less common stems such as reeds and organs. Implementation is available at https: github.com kwatcharasupat query-bandit.
【15】 Automatic Prediction of Amyotrophic Lateral Sclerosis Progression using Longitudinal Speech Transformer
标题: 使用纵向语音Transformer自动预测肌萎缩侧索硬化进展
作者:Liming Wang,Yuan Gong,Nauman Dawalatabad,Marco Vilela,Katerina Placek,Brian Tracey,Yishu Gong,Alan Premasiri,Fernando Vieira,James Glass
链接:点击下载PDF文件
摘要:肌萎缩侧索硬化症(ALS)疾病进展的自动预测提供了一个更有效和客观的替代方法比手动。我们提出了ALS纵向语音Transformer(ALST),一种基于神经网络的ALS疾病进展的自动预测器,从ALS患者的纵向语音记录。通过利用高质量的预训练语音特征和录音中的纵向信息,我们的最佳模型达到了91.0 % AUC,在ALS TDI数据集上比以前的最佳模型提高了5.6 %。仔细分析表明,ALST能够对ALS进展进行细粒度和可解释的预测,特别是用于区分罕见和严重病例。代码是公开的。摘要:Automatic prediction of amyotrophic lateral sclerosis (ALS) disease progression provides a more efficient and objective alternative than manual approaches. We propose ALS longitudinal speech transformer (ALST), a neural network-based automatic predictor of ALS disease progression from longitudinal speech recordings of ALS patients. By taking advantage of high-quality pretrained speech features and longitudinal information in the recordings, our best model achieves 91.0 % AUC, improving upon the previous best model by 5.6 % relative on the ALS TDI dataset. Careful analysis reveals that ALST is capable of fine-grained and interpretable predictions of ALS progression, especially for distinguishing between rarer and more severe cases. Code is publicly available.
【16】 Towards Deep Active Learning in Avian Bioacoustics
标题: 鸟类生物声学中的深度主动学习
作者:Lukas Rauch,Denis Huseljic,Moritz Wirth,Jens Decke,Bernhard Sick,Christoph Scholz
备注:preprint, under review IAL@ECML-PKDD24
链接:点击下载PDF文件
摘要:鸟类生物声学中的被动声学监测(PAM)能够以最小的自然栖息地破坏进行成本效益和广泛的数据收集。尽管在计算鸟类生物声学方面取得了进展,但深度学习模型在实际PAM场景中适应不同环境方面仍然面临挑战。这主要是由于注释的稀缺性,这需要人类专家的劳动密集型工作。主动学习(AL)通过查询最具信息量的实例进行标记,降低了注释成本,并加快了对不同场景的适应速度。本文概述了深入的人工智能方法,介绍了关键挑战,并进行了小规模的试点研究。摘要:Passive acoustic monitoring (PAM) in avian bioacoustics enables cost-effective and extensive data collection with minimal disruption to natural habitats. Despite advancements in computational avian bioacoustics, deep learning models continue to encounter challenges in adapting to diverse environments in practical PAM scenarios. This is primarily due to the scarcity of annotations, which requires labor-intensive efforts from human experts. Active learning (AL) reduces annotation cost and speed ups adaption to diverse scenarios by querying the most informative instances for labeling. This paper outlines a deep AL approach, introduces key challenges, and conducts a small-scale pilot study.
机器翻译,仅供参考
