今日论文合集:cs.SD语音8篇,eess.AS音频处理12篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 Speech Editing -- a Summary
标题: 语音编辑--总结
作者:Tobias Kässmann,Yining Liu,Danni Liu
链接:点击下载PDF文件
摘要:随着视频制作和社交媒体的兴起,语音编辑对于创作者解决录音中的发音错误、漏词或口吃等问题至关重要。本文探讨了基于文本的语音编辑方法,修改音频通过文本转录,而无需手动波形编辑。这些方法通过改变梅尔频谱图来确保编辑的音频与原始音频无法区分。最近的进展,如上下文感知韵律校正和先进的注意力机制,提高了语音编辑质量。本文回顾了最先进的方法,比较了关键指标,并检查了广泛使用的数据集。其目的是突出正在进行的问题,并激发进一步的研究和创新的语音编辑。摘要:With the rise of video production and social media, speech editing has become crucial for creators to address issues like mispronunciations, missing words, or stuttering in audio recordings. This paper explores text-based speech editing methods that modify audio via text transcripts without manual waveform editing. These approaches ensure edited audio is indistinguishable from the original by altering the mel-spectrogram. Recent advancements, such as context-aware prosody correction and advanced attention mechanisms, have improved speech editing quality. This paper reviews state-of-the-art methods, compares key metrics, and examines widely used datasets. The aim is to highlight ongoing issues and inspire further research and innovation in speech editing.

【2】 Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
标题: 使用预训练的捷克SpeechT5模型的Zero-Shot与Few-Shot多扬声器TTC
作者:Jan Lehečka,Zdeněk Hanzlíček,Jindřich Matoušek,Daniel Tihelka
备注:Accepted to TSD2024
链接:点击下载PDF文件
摘要:在本文中,我们使用在大规模数据集上预训练的SpeechT5模型进行了实验。我们从头开始预训练基础模型,并在大规模鲁棒的多说话者文本到语音(TTS)任务中对其进行微调。我们在零镜头和Few-Shot场景中测试了模型功能。基于两个听力测试,我们评估了合成音频质量和合成语音与真实语音的相似性。我们的研究结果表明,SpeechT5模型可以生成一个合成语音的任何扬声器使用只有一分钟的目标扬声器的数据。我们成功地在捷克知名政治家和名人身上展示了我们合成声音的高质量和相似性。摘要:In this paper, we experimented with the SpeechT5 model pre-trained on large-scale datasets. We pre-trained the foundation model from scratch and fine-tuned it on a large-scale robust multi-speaker text-to-speech (TTS) task. We tested the model capabilities in a zero- and few-shot scenario. Based on two listening tests, we evaluated the synthetic audio quality and the similarity of how synthetic voices resemble real voices. Our results showed that the SpeechT5 model can generate a synthetic voice for any speaker using only one minute of the target speaker's data. We successfully demonstrated the high quality and similarity of our synthetic voices on publicly known Czech politicians and celebrities.

【3】 Collaboration Between Robots, Interfaces and Humans: Practice-Based and Audience Perspectives
标题: 机器人、界面和人类之间的协作:基于实践和受众的视角
作者:Anna Savery,Richard Savery
Journal-ref:International Computer Music Conference 2024, Seoul, South Korea
链接:点击下载PDF文件
摘要:本文提供了一个混合媒体实验音乐作品的分析,探讨了人类音乐互动与新开发的小提琴界面的整合,由一个即兴小提琴手,交互式视觉效果,机器人鼓手和即兴合成管弦乐队操纵。我们首先介绍了所涉及的系统的详细技术概述,包括每个组件的设计和功能。然后,我们进行基于实践的审查,审查作品的创作过程和艺术决策,重点关注其发展过程中遇到的挑战和突破。通过这种内省的分析,我们揭示了人类表演者和技术代理人之间的协作动态,揭示了将传统音乐表现力与人工智能和机器人技术相结合的复杂性。为了评估公众的接受程度和诠释观点,我们进行了一项在线调查,与不同的观众分享了表演的视频。从这次调查中收集的反馈提供了关于作品的可访问性,情感影响和感知艺术价值的宝贵观点。受访者的反应强调了将先进技术融入音乐表演的变革潜力,同时也强调了进一步探索和改进的领域。摘要:This paper provides an analysis of a mixed-media experimental musical work that explores the integration of human musical interaction with a newly developed interface for the violin, manipulated by an improvising violinist, interactive visuals, a robotic drummer and an improvised synthesised orchestra. We first present a detailed technical overview of the systems involved including the design and functionality of each component. We then conduct a practice-based review examining the creative processes and artistic decisions underpinning the work, focusing on the challenges and breakthroughs encountered during its development. Through this introspective analysis, we uncover insights into the collaborative dynamics between the human performer and technological agents, revealing the complexities of blending traditional musical expressiveness with artificial intelligence and robotics. To gauge public reception and interpretive perspectives, we conducted an online survey, sharing a video of the performance with a diverse audience. The feedback collected from this survey offers valuable viewpoints on the accessibility, emotional impact, and perceived artistic value of the work. Respondents' reactions underscore the transformative potential of integrating advanced technologies in musical performance, while also highlighting areas for further exploration and refinement.

【4】 Long-Term, Store-Front Robotics: Interactive Music for Robotic Arm, Caxixi and Frame Drums
标题: 长期店面机器人技术:机器人手臂、Caixi和框架鼓的互动音乐
作者:Richard Savery,Fouad Sukkar
Journal-ref:International Computer Music Conference 2024, Seoul, South Korea
链接:点击下载PDF文件
摘要:本文介绍了一种创新的探索,在商业零售环境中整合互动机器人音乐,特别是通过一个为期三周的店内安装具有UR3机械臂,定制框架鼓,和自适应音乐生成系统。该项目位于世界上最大的城市之一的一个突出的店面,旨在通过创造动态的,引人入胜的音乐互动来增强购物体验,以响应商店的环境音景。主要贡献包括工业机器人在艺术表达中的新颖应用,互动音乐的部署,以丰富零售氛围,以及在公共环境中长时间连续机器人操作的演示。系统可靠性、音乐输出的变化、互动环境中的安全性和品牌一致性等挑战都得到了解决,以确保安装的成功。该项目不仅展示了机器人音乐在零售空间中的技术可行性和艺术潜力,还提供了对这种集成的实际影响的见解,包括系统可靠性,人机交互的动态以及对商店运营的影响。这种探索为通过技术、音乐和互动艺术的交叉来增强消费者的零售体验开辟了新的途径,表明未来机器人音乐将为公共和商业空间做出有意义的贡献。摘要:This paper presents an innovative exploration into the integration of interactive robotic musicianship within a commercial retail environment, specifically through a three-week-long in-store installation featuring a UR3 robotic arm, custom-built frame drums, and an adaptive music generation system. Situated in a prominent storefront in one of the world's largest cities, this project aimed to enhance the shopping experience by creating dynamic, engaging musical interactions that respond to the store's ambient soundscape. Key contributions include the novel application of industrial robotics in artistic expression, the deployment of interactive music to enrich retail ambiance, and the demonstration of continuous robotic operation in a public setting over an extended period. Challenges such as system reliability, variation in musical output, safety in interactive contexts, and brand alignment were addressed to ensure the installation's success. The project not only showcased the technical feasibility and artistic potential of robotic musicianship in retail spaces but also offered insights into the practical implications of such integration, including system reliability, the dynamics of human-robot interaction, and the impact on store operations. This exploration opens new avenues for enhancing consumer retail experiences through the intersection of technology, music, and interactive art, suggesting a future where robotic musicianship contributes meaningfully to public and commercial spaces.

【5】 A Framework for AI assisted Musical Devices
标题: 人工智能辅助音乐设备框架
作者:Miguel Civit,Luis Munoz Saavedra,Francisco Jose Cuadrado,Maria J. Escalona
Journal-ref:IntechOpen (2023)
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个新的框架,研究和设计人工智能辅助音乐设备(AIME)。首先,我们提出了这些设备的分类,并说明它与一组场景和人物角色。后来,我们提出了一个通用的架构实现AIME,并提出了一些例子,从场景。我们表明,所提出的框架和体系结构是一个有效的工具,智能音乐设备的研究。摘要:In this paper we present a novel framework for the study and design of AI assisted musical devices (AIMEs). Initially, we present a taxonomy of these devices and illustrate it with a set of scenarios and personas. Later, we propose a generic architecture for the implementation of AIMEs and present some examples from the scenarios. We show that the proposed framework and architecture are a valid tool for the study of intelligent musical devices.

【6】 Towards better visualizations of urban sound environments: insights from interviews
标题: 实现更好的城市声音环境可视化:采访的见解
作者:Modan Tailleur,Pierre Aumond,Vincent Tourre,Mathieu Lagrange
Journal-ref:INTERNOISE 2024, Aug 2024, Nantes (France), France
链接:点击下载PDF文件
摘要:传统上,城市噪声地图和噪声可视化提供了城市噪声水平的宏观表示。然而,这些表示无法准确衡量与这些声音环境相关的声音感知,因为感知高度依赖于所涉及的声源。本文的目的是分析声源的代表性的需要,通过确定城市的利益相关者,这种代表性被认为是重要的。通过与各种城市利益相关者的口头采访,我们深入了解了当前的做法,现有工具的优缺点以及将声源纳入现有城市声环境表示的相关性。在这项研究中出现了三种不同的使用声源表示:1)工业和专业市民的噪声相关投诉,2)市民的声景观质量评估,3)城市规划师的指导。调查结果还揭示了使用可视化的不同观点,这应该使用适应目标受众的指标,并使数据的可访问性。摘要:Urban noise maps and noise visualizations traditionally provide macroscopic representations of noise levels across cities. However, those representations fail at accurately gauging the sound perception associated with these sound environments, as perception highly depends on the sound sources involved. This paper aims at analyzing the need for the representations of sound sources, by identifying the urban stakeholders for whom such representations are assumed to be of importance. Through spoken interviews with various urban stakeholders, we have gained insight into current practices, the strengths and weaknesses of existing tools and the relevance of incorporating sound sources into existing urban sound environment representations. Three distinct use of sound source representations emerged in this study: 1) noise-related complaints for industrials and specialized citizens, 2) soundscape quality assessment for citizens, and 3) guidance for urban planners. Findings also reveal diverse perspectives for the use of visualizations, which should use indicators adapted to the target audience, and enable data accessibility.

【7】 Reduction of Nonlinear Distortion in Condenser Microphones Using a Simple Post-Processing Technique
标题: 使用简单的后处理技术减少电容麦克风的非线性失真
作者:Petr Honzík,Antonin Novak
备注:10 pages, 9 figures
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一种新的方法,有效地减少非线性失真的单背板电容麦克风,即,大多数MEMS麦克风、录音室录音电容麦克风和实验室测量麦克风。这种简单的后处理技术可以很容易地集成在外部硬件上,例如模拟电路、微控制器、音频编解码器、DSP单元,或者在MEMS麦克风的情况下集成在ASIC芯片内。它显著降低了麦克风在其频率和动态范围内的失真。它依赖于一个单一的参数,它可以从麦克风的物理参数或本文中提出的一个简单的测量。该参数的最佳估计实现了最佳的失真减少,而过高估计它永远不会增加超过原始水平的失真。该技术在MEMS麦克风上进行了测试。我们的研究结果表明,谐波激励所提出的技术减少了约40分贝的二次谐波,导致显着减少总谐波失真(THD)。通过双音和多音实验证明了失真降低技术对于更复杂信号的效率,其中二阶互调产物至少降低了20 dB。摘要:In this paper, we introduce a novel approach for effectively reducing nonlinear distortion in single back-plate condenser microphones, i.e., most MEMS microphones, studio recording condenser microphones, and laboratory measurement microphones. This simple post-processing technique can be easily integrated on an external hardware such as an analog circuit, microcontroller, audio codec, DSP unit, or within the ASIC chip in a case of MEMS microphones. It significantly reduces microphone distortion across its frequency and dynamic range. It relies on a single parameter, which can be derived from either the microphone's physical parameters or a straightforward measurement presented in this paper. An optimal estimate of this parameter achieves the best distortion reduction, whereas overestimating it never increases distortion beyond the original level. The technique was tested on a MEMS microphone. Our findings indicate that for harmonic excitation the proposed technique reduces the second harmonic by approximately 40 dB, leading to a significant reduction in the Total Harmonic Distortion (THD). The efficiency of the distortion reduction technique for more complex signals is demonstrated through two-tone and multitone experiments, where second-order intermodulation products are reduced by at least 20 dB.

【8】 Automatic Detection and Annotation of Sperm Whale Codas
标题: 抹香鲸尾线的自动检测与注释
作者:Guy Gubnitsky,Yaly Mevorach,Shane Gero,David F. Gruber,Roee Diamant
链接:点击下载PDF文件
摘要:监测抹香鲸(Physeter macrocephalus)的一项关键技术是识别抹香鲸的通信信号,称为尾波。在本文中,我们提出了第一个自动尾波检测和注释。我们的检测器的主要创新是基于图的聚类,它利用了构成尾波的点击之间的预期相似性。结果表明,检测和准确的注释,在低信噪比,分离之间的尾声和回声定位点击,并从同时发射鲸鱼的尾声之间的歧视。使用这个自动注释器,深入了解抹香鲸通信的特点。结果包括新类型的尾波信号,分析不同鲸鱼之间和不同年份的尾波类型的分布,以及在尾波类型和尾波传输时间方面通信鲸鱼之间同步的证据。这些结果表明,这一鲸目物种的通信系统高度复杂。为了确保可追溯性,我们共享了尾波检测器的实现代码。摘要:A key technology in sperm whale (Physeter macrocephalus) monitoring is the identification of sperm whale communication signals, known as codas. In this paper we present the first automatic coda detector and annotator. The main innovation in our detector is graph-based clustering, which utilizes the expected similarity between the clicks that make up the coda. Results show detection and accurate annotation at low signal-to-noise ratios, separation between codas and echolocation clicks, and discrimination between codas from simultaneously emitting whales. Using this automatic annotator, insights into the characterization of sperm whale communication are presented. The results include new types of coda signals, analyzes of the distribution of coda types among different whales and for different years, and evidence for synchronization between communicating whales in terms of coda type and coda transmission time. These results indicate a high degree of complexity in the communication system of this cetacean species. To ensure traceability, we share the implementation code of our coda detector.


eess.AS音频处理
【1】 A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation
标题: 真实语音动画视听同步技术的全面回顾和分类
作者:Jose Geraldo Fernandes,Sinval Nascimento,Daniel Dominguete,André Oliveira,Lucas Rotsen,Gabriel Souza,Gabriel Lara,Mateus Vilela,Pedro Mapa,David Brochero,Hebert Costa,Frederico Coelho,Antônio P. Braga
链接:点击下载PDF文件
摘要:在许多应用程序中,同步音频与视觉效果至关重要,例如为电影或游戏创建图形动画,将电影音频翻译成不同的语言,以及开发Metaverse应用程序。本文探讨了从音频输入实现逼真的面部动画,突出生成和自适应模型的各种方法。为了解决模型训练成本、数据集可用性和音频数据中的无声时刻分布等挑战,它提出了创新的解决方案来增强性能和真实感。该研究还引入了一种新的分类法,根据后勤方面对视听同步方法进行分类,提高了虚拟助手,游戏和交互式数字媒体的功能。摘要:In many applications, synchronizing audio with visuals is crucial, such as in creating graphic animations for films or games, translating movie audio into different languages, and developing metaverse applications. This review explores various methodologies for achieving realistic facial animations from audio inputs, highlighting generative and adaptive models. Addressing challenges like model training costs, dataset availability, and silent moment distributions in audio data, it presents innovative solutions to enhance performance and realism. The research also introduces a new taxonomy to categorize audio-visual synchronization methods based on logistical aspects, advancing the capabilities of virtual assistants, gaming, and interactive digital media.

【2】 Explaining Spectrograms in Machine Learning: A Study on Neural Networks for Speech Classification
标题: 机器学习中解释频谱图:用于语音分类的神经网络研究
作者:Jesin James,Balamurali B. T.,Binu Abeysinghe,Junchen Liu
备注:5th International Conference on Artificial Intelligence and Speech Technology (AIST-2023), New Delhi, India
链接:点击下载PDF文件
摘要:本研究探讨了神经网络学习的判别模式,以实现准确的语音分类,特别关注元音分类任务。通过研究元音分类神经网络的激活和特征,我们可以深入了解网络在声谱图中“看到”了什么。通过使用类激活映射,我们确定的频率,有助于元音分类和比较这些发现与语言知识。在美国英语元音数据集上的实验展示了神经网络的可解释性,并为将其与清音语音区分开来时错误分类的原因及其特征提供了有价值的见解。这项研究不仅增强了我们对元音分类中潜在声学线索的理解,而且通过弥合神经网络中的抽象表征与既定语言学知识之间的差距,为改善语音识别提供了机会摘要:This study investigates discriminative patterns learned by neural networks for accurate speech classification, with a specific focus on vowel classification tasks. By examining the activations and features of neural networks for vowel classification, we gain insights into what the networks "see" in spectrograms. Through the use of class activation mapping, we identify the frequencies that contribute to vowel classification and compare these findings with linguistic knowledge. Experiments on a American English dataset of vowels showcases the explainability of neural networks and provides valuable insights into the causes of misclassifications and their characteristics when differentiating them from unvoiced speech. This study not only enhances our understanding of the underlying acoustic cues in vowel classification but also offers opportunities for improving speech recognition by bridging the gap between abstract representations in neural networks and established linguistic knowledge

【3】 Reduction of Nonlinear Distortion in Condenser Microphones Using a Simple Post-Processing Technique
标题: 使用简单的后处理技术减少电容麦克风的非线性失真
作者:Petr Honzík,Antonin Novak
备注:10 pages, 9 figures
链接:点击下载PDF文件
摘要:在本文中,我们介绍了一种新的方法,有效地减少非线性失真的单背板电容麦克风,即,大多数MEMS麦克风、录音室录音电容麦克风和实验室测量麦克风。这种简单的后处理技术可以很容易地集成在外部硬件上,例如模拟电路、微控制器、音频编解码器、DSP单元,或者在MEMS麦克风的情况下集成在ASIC芯片内。它显著降低了麦克风在其频率和动态范围内的失真。它依赖于一个单一的参数,它可以从麦克风的物理参数或本文中提出的一个简单的测量。该参数的最佳估计实现了最佳的失真减少,而过高估计它永远不会增加超过原始水平的失真。该技术在MEMS麦克风上进行了测试。我们的研究结果表明,谐波激励所提出的技术减少了约40分贝的二次谐波,导致显着减少总谐波失真(THD)。通过双音和多音实验证明了失真降低技术对于更复杂信号的效率,其中二阶互调产物至少降低了20 dB。摘要:In this paper, we introduce a novel approach for effectively reducing nonlinear distortion in single back-plate condenser microphones, i.e., most MEMS microphones, studio recording condenser microphones, and laboratory measurement microphones. This simple post-processing technique can be easily integrated on an external hardware such as an analog circuit, microcontroller, audio codec, DSP unit, or within the ASIC chip in a case of MEMS microphones. It significantly reduces microphone distortion across its frequency and dynamic range. It relies on a single parameter, which can be derived from either the microphone's physical parameters or a straightforward measurement presented in this paper. An optimal estimate of this parameter achieves the best distortion reduction, whereas overestimating it never increases distortion beyond the original level. The technique was tested on a MEMS microphone. Our findings indicate that for harmonic excitation the proposed technique reduces the second harmonic by approximately 40 dB, leading to a significant reduction in the Total Harmonic Distortion (THD). The efficiency of the distortion reduction technique for more complex signals is demonstrated through two-tone and multitone experiments, where second-order intermodulation products are reduced by at least 20 dB.

【4】 Automatic Detection and Annotation of Sperm Whale Codas
标题: 抹香鲸尾线的自动检测与注释
作者:Guy Gubnitsky,Yaly Mevorach,Shane Gero,David F. Gruber,Roee Diamant
链接:点击下载PDF文件
摘要:监测抹香鲸(Physeter macrocephalus)的一项关键技术是识别抹香鲸的通信信号,称为尾波。在本文中,我们提出了第一个自动尾波检测和注释。我们的检测器的主要创新是基于图的聚类,它利用了构成尾波的点击之间的预期相似性。结果表明,检测和准确的注释,在低信噪比,分离之间的尾声和回声定位点击,并从同时发射鲸鱼的尾声之间的歧视。使用这个自动注释器,深入了解抹香鲸通信的特点。结果包括新类型的尾波信号,分析不同鲸鱼之间和不同年份的尾波类型的分布,以及在尾波类型和尾波传输时间方面通信鲸鱼之间同步的证据。这些结果表明,这一鲸目物种的通信系统高度复杂。为了确保可追溯性,我们共享了尾波检测器的实现代码。摘要:A key technology in sperm whale (Physeter macrocephalus) monitoring is the identification of sperm whale communication signals, known as codas. In this paper we present the first automatic coda detector and annotator. The main innovation in our detector is graph-based clustering, which utilizes the expected similarity between the clicks that make up the coda. Results show detection and accurate annotation at low signal-to-noise ratios, separation between codas and echolocation clicks, and discrimination between codas from simultaneously emitting whales. Using this automatic annotator, insights into the characterization of sperm whale communication are presented. The results include new types of coda signals, analyzes of the distribution of coda types among different whales and for different years, and evidence for synchronization between communicating whales in terms of coda type and coda transmission time. These results indicate a high degree of complexity in the communication system of this cetacean species. To ensure traceability, we share the implementation code of our coda detector.

【5】 Uncertainty-Based Ensemble Learning For Speech Classification
标题: 基于不确定性的语音分类集合学习
作者:Bagus Tris Atmaja,Felix Burkhardt
备注:Submitted to OCOCOSDA 2024
链接:点击下载PDF文件
摘要:语音分类由于其广泛的应用,特别是在生理和心理状态分类方面的应用,引起了越来越多的关注。然而,由于语音信号的高度可变性,这些任务具有挑战性。当多个分类器被组合以提高性能时,包围学习已经显示出有希望的结果。随着硬件开发的最新进展,结合几种模型并不是深度学习研究和应用的限制。在本文中,我们提出了一种基于不确定性的集成学习语音分类方法。具体来说,我们在同一个分类器上训练一组基本特征,并量化其预测的不确定性。使用不确定性计算的变体组合预测以产生最终预测。可视化的不确定性及其集成学习的结果表明语音分类任务的潜在改进。所提出的方法优于单一的模型和传统的集成学习方法的加权精度或未加权的准确性。摘要:Speech classification has attracted increasing attention due to its wide applications, particularly in classifying physical and mental states. However, these tasks are challenging due to the high variability in speech signals. Ensemble learning has shown promising results when multiple classifiers are combined to improve performance. With recent advancements in hardware development, combining several models is not a limitation in deep learning research and applications. In this paper, we propose an uncertainty-based ensemble learning approach for speech classification. Specifically, we train a set of base features on the same classifier and quantify the uncertainty of their predictions. The predictions are combined using variants of uncertainty calculation to produce the final prediction. The visualization of the effect of uncertainty and its ensemble learning results show potential improvements in speech classification tasks. The proposed method outperforms single models and conventional ensemble learning methods in terms of unweighted accuracy or weighted accuracy.

【6】 Synth4Kws: Synthesized Speech for User Defined Keyword Spotting in Low Resource Environments
标题: Synh 4Kws:用于低资源环境中用户定义关键字定位的合成语音
作者:Pai Zhu,Dhruuv Agarwal,Jacob W. Bartel,Kurt Partridge,Hyun Jin Park,Quan Wang
备注:5 pages, 5 figures, 2 tables The paper is accepted in Interspeech SynData4GenAI 2024 Workshop - this https URL
链接:点击下载PDF文件
摘要:开发高质量的自定义关键字识别(KWS)模型的挑战之一是收集涵盖各种语言、短语和说话风格的训练数据的过程冗长而昂贵。我们介绍了Synth 4Kws-一个框架,利用文本到语音(TTS)合成数据的自定义KWS在不同的资源设置。在没有真实数据的情况下,我们发现增加TTS短语多样性和话语采样单调地提高了模型性能,如通过语音命令数据集的11 k话语的EER和AUC度量所评估的。在低资源设置中,以50 k真实话语为基线,我们发现使用最佳数量的TTS数据可以将EER提高30.1%,AUC提高46.7%。此外,我们将TTS数据与不同数量的真实数据混合,并对实现各种质量目标所需的真实数据进行插值。我们的实验是基于英语和单个单词的话语,但研究结果推广到i18 n语言和其他关键字类型。摘要:One of the challenges in developing a high quality custom keyword spotting (KWS) model is the lengthy and expensive process of collecting training data covering a wide range of languages, phrases and speaking styles. We introduce Synth4Kws - a framework to leverage Text to Speech (TTS) synthesized data for custom KWS in different resource settings. With no real data, we found increasing TTS phrase diversity and utterance sampling monotonically improves model performance, as evaluated by EER and AUC metrics over 11k utterances of the speech command dataset. In low resource settings, with 50k real utterances as a baseline, we found using optimal amounts of TTS data can improve EER by 30.1% and AUC by 46.7%. Furthermore, we mix TTS data with varying amounts of real data and interpolate the real data needed to achieve various quality targets. Our experiments are based on English and single word utterances but the findings generalize to i18n languages and other keyword types.

【7】 Speech Editing -- a Summary
标题: 语音编辑--总结
作者:Tobias Kässmann,Yining Liu,Danni Liu
链接:点击下载PDF文件
摘要:随着视频制作和社交媒体的兴起,语音编辑对于创作者解决录音中的发音错误、漏词或口吃等问题至关重要。本文探讨了基于文本的语音编辑方法,修改音频通过文本转录,而无需手动波形编辑。这些方法通过改变梅尔频谱图来确保编辑的音频与原始音频无法区分。最近的进展,如上下文感知韵律校正和先进的注意力机制,提高了语音编辑质量。本文回顾了最先进的方法,比较了关键指标,并检查了广泛使用的数据集。其目的是突出正在进行的问题,并激发进一步的研究和创新的语音编辑。摘要:With the rise of video production and social media, speech editing has become crucial for creators to address issues like mispronunciations, missing words, or stuttering in audio recordings. This paper explores text-based speech editing methods that modify audio via text transcripts without manual waveform editing. These approaches ensure edited audio is indistinguishable from the original by altering the mel-spectrogram. Recent advancements, such as context-aware prosody correction and advanced attention mechanisms, have improved speech editing quality. This paper reviews state-of-the-art methods, compares key metrics, and examines widely used datasets. The aim is to highlight ongoing issues and inspire further research and innovation in speech editing.

【8】 Zero-Shot vs. Few-Shot Multi-Speaker TTS Using Pre-trained Czech SpeechT5 Model
标题: 使用预训练的捷克SpeechT5模型的Zero-Shot与Few-Shot多扬声器TTC
作者:Jan Lehečka,Zdeněk Hanzlíček,Jindřich Matoušek,Daniel Tihelka
备注:Accepted to TSD2024
链接:点击下载PDF文件
摘要:在本文中,我们使用在大规模数据集上预训练的SpeechT5模型进行了实验。我们从头开始预训练基础模型,并在大规模鲁棒的多说话者文本到语音(TTS)任务中对其进行微调。我们在零镜头和Few-Shot场景中测试了模型功能。基于两个听力测试,我们评估了合成音频质量和合成语音与真实语音的相似性。我们的研究结果表明,SpeechT5模型可以生成一个合成语音的任何扬声器使用只有一分钟的目标扬声器的数据。我们成功地在捷克知名政治家和名人身上展示了我们合成声音的高质量和相似性。摘要:In this paper, we experimented with the SpeechT5 model pre-trained on large-scale datasets. We pre-trained the foundation model from scratch and fine-tuned it on a large-scale robust multi-speaker text-to-speech (TTS) task. We tested the model capabilities in a zero- and few-shot scenario. Based on two listening tests, we evaluated the synthetic audio quality and the similarity of how synthetic voices resemble real voices. Our results showed that the SpeechT5 model can generate a synthetic voice for any speaker using only one minute of the target speaker's data. We successfully demonstrated the high quality and similarity of our synthetic voices on publicly known Czech politicians and celebrities.

【9】 Collaboration Between Robots, Interfaces and Humans: Practice-Based and Audience Perspectives
标题: 机器人、界面和人类之间的协作:基于实践和受众的视角
作者:Anna Savery,Richard Savery
Journal-ref:International Computer Music Conference 2024, Seoul, South Korea
链接:点击下载PDF文件
摘要:本文提供了一个混合媒体实验音乐作品的分析,探讨了人类音乐互动与新开发的小提琴界面的整合,由一个即兴小提琴手,交互式视觉效果,机器人鼓手和即兴合成管弦乐队操纵。我们首先介绍了所涉及的系统的详细技术概述,包括每个组件的设计和功能。然后,我们进行基于实践的审查,审查作品的创作过程和艺术决策,重点关注其发展过程中遇到的挑战和突破。通过这种内省的分析,我们揭示了人类表演者和技术代理人之间的协作动态,揭示了将传统音乐表现力与人工智能和机器人技术相结合的复杂性。为了评估公众的接受程度和诠释观点,我们进行了一项在线调查,与不同的观众分享了表演的视频。从这次调查中收集的反馈提供了关于作品的可访问性,情感影响和感知艺术价值的宝贵观点。受访者的反应强调了将先进技术融入音乐表演的变革潜力,同时也强调了进一步探索和改进的领域。摘要:This paper provides an analysis of a mixed-media experimental musical work that explores the integration of human musical interaction with a newly developed interface for the violin, manipulated by an improvising violinist, interactive visuals, a robotic drummer and an improvised synthesised orchestra. We first present a detailed technical overview of the systems involved including the design and functionality of each component. We then conduct a practice-based review examining the creative processes and artistic decisions underpinning the work, focusing on the challenges and breakthroughs encountered during its development. Through this introspective analysis, we uncover insights into the collaborative dynamics between the human performer and technological agents, revealing the complexities of blending traditional musical expressiveness with artificial intelligence and robotics. To gauge public reception and interpretive perspectives, we conducted an online survey, sharing a video of the performance with a diverse audience. The feedback collected from this survey offers valuable viewpoints on the accessibility, emotional impact, and perceived artistic value of the work. Respondents' reactions underscore the transformative potential of integrating advanced technologies in musical performance, while also highlighting areas for further exploration and refinement.

【10】 Long-Term, Store-Front Robotics: Interactive Music for Robotic Arm, Caxixi and Frame Drums
标题: 长期店面机器人技术:机器人手臂、Caixi和框架鼓的互动音乐
作者:Richard Savery,Fouad Sukkar
Journal-ref:International Computer Music Conference 2024, Seoul, South Korea
链接:点击下载PDF文件
摘要:本文介绍了一种创新的探索,在商业零售环境中整合互动机器人音乐,特别是通过一个为期三周的店内安装具有UR3机械臂,定制框架鼓,和自适应音乐生成系统。该项目位于世界上最大的城市之一的一个突出的店面,旨在通过创造动态的,引人入胜的音乐互动来增强购物体验,以响应商店的环境音景。主要贡献包括工业机器人在艺术表达中的新颖应用,互动音乐的部署,以丰富零售氛围,以及在公共环境中长时间连续机器人操作的演示。系统可靠性、音乐输出的变化、互动环境中的安全性和品牌一致性等挑战都得到了解决,以确保安装的成功。该项目不仅展示了机器人音乐在零售空间中的技术可行性和艺术潜力,还提供了对这种集成的实际影响的见解,包括系统可靠性,人机交互的动态以及对商店运营的影响。这种探索为通过技术、音乐和互动艺术的交叉来增强消费者的零售体验开辟了新的途径,表明未来机器人音乐将为公共和商业空间做出有意义的贡献。摘要:This paper presents an innovative exploration into the integration of interactive robotic musicianship within a commercial retail environment, specifically through a three-week-long in-store installation featuring a UR3 robotic arm, custom-built frame drums, and an adaptive music generation system. Situated in a prominent storefront in one of the world's largest cities, this project aimed to enhance the shopping experience by creating dynamic, engaging musical interactions that respond to the store's ambient soundscape. Key contributions include the novel application of industrial robotics in artistic expression, the deployment of interactive music to enrich retail ambiance, and the demonstration of continuous robotic operation in a public setting over an extended period. Challenges such as system reliability, variation in musical output, safety in interactive contexts, and brand alignment were addressed to ensure the installation's success. The project not only showcased the technical feasibility and artistic potential of robotic musicianship in retail spaces but also offered insights into the practical implications of such integration, including system reliability, the dynamics of human-robot interaction, and the impact on store operations. This exploration opens new avenues for enhancing consumer retail experiences through the intersection of technology, music, and interactive art, suggesting a future where robotic musicianship contributes meaningfully to public and commercial spaces.

【11】 A Framework for AI assisted Musical Devices
标题: 人工智能辅助音乐设备框架
作者:Miguel Civit,Luis Munoz Saavedra,Francisco Jose Cuadrado,Maria J. Escalona
Journal-ref:IntechOpen (2023)
链接:点击下载PDF文件
摘要:在本文中,我们提出了一个新的框架,研究和设计人工智能辅助音乐设备(AIME)。首先,我们提出了这些设备的分类,并说明它与一组场景和人物角色。后来,我们提出了一个通用的架构实现AIME,并提出了一些例子,从场景。我们表明,所提出的框架和体系结构是一个有效的工具,智能音乐设备的研究。摘要:In this paper we present a novel framework for the study and design of AI assisted musical devices (AIMEs). Initially, we present a taxonomy of these devices and illustrate it with a set of scenarios and personas. Later, we propose a generic architecture for the implementation of AIMEs and present some examples from the scenarios. We show that the proposed framework and architecture are a valid tool for the study of intelligent musical devices.

【12】 Towards better visualizations of urban sound environments: insights from interviews
标题: 实现更好的城市声音环境可视化:采访的见解
作者:Modan Tailleur,Pierre Aumond,Vincent Tourre,Mathieu Lagrange
Journal-ref:INTERNOISE 2024, Aug 2024, Nantes (France), France
链接:点击下载PDF文件
摘要:传统上,城市噪声地图和噪声可视化提供了城市噪声水平的宏观表示。然而,这些表示无法准确地测量与这些声音环境相关联的声音感知,因为感知高度依赖于所涉及的声源。本文的目的是分析声源的代表性的需要,通过确定城市的利益相关者,这种代表性被认为是重要的。通过与各种城市利益相关者的口头采访,我们深入了解了当前的做法,现有工具的优缺点以及将声源纳入现有城市声环境表示的相关性。在这项研究中出现了三种不同的使用声源表示:1)工业和专业市民的噪声相关投诉,2)市民的声景观质量评估,3)城市规划师的指导。调查结果还揭示了使用可视化的不同观点,这应该使用适应目标受众的指标,并使数据的可访问性。摘要:Urban noise maps and noise visualizations traditionally provide macroscopic representations of noise levels across cities. However, those representations fail at accurately gauging the sound perception associated with these sound environments, as perception highly depends on the sound sources involved. This paper aims at analyzing the need for the representations of sound sources, by identifying the urban stakeholders for whom such representations are assumed to be of importance. Through spoken interviews with various urban stakeholders, we have gained insight into current practices, the strengths and weaknesses of existing tools and the relevance of incorporating sound sources into existing urban sound environment representations. Three distinct use of sound source representations emerged in this study: 1) noise-related complaints for industrials and specialized citizens, 2) soundscape quality assessment for citizens, and 3) guidance for urban planners. Findings also reveal diverse perspectives for the use of visualizations, which should use indicators adapted to the target audience, and enable data accessibility.


机器翻译,仅供参考