本文经arXiv每日学术速递授权转载
标题: 具有自动回归的视频时间对齐音频
作者:Ilpo Viertola,Vladimir Iashin,Esa Rahtu
备注:Submitted to ICASSP 2025. Project page this https URL
链接:点击下载PDF文件
摘要:我们介绍V-AURA,第一个自回归模型,以实现高的时间对齐和相关性的视频到音频的生成。V-AURA使用高帧率视觉特征提取器和跨模态视听特征融合策略来捕获细粒度的视觉运动事件并确保精确的时间对齐。此外,我们提出了VisualSound,一个具有高视听相关性的基准数据集。VisualSound基于VGGSound,VGGSound是一个视频数据集,由从YouTube提取的原始样本组成。在策展过程中,我们删除了听觉事件与视觉事件不一致的样本。V-AURA在时间对齐和语义相关性方面优于当前最先进的模型,同时保持相当的音频质量。代码、示例、VisualSound和模型可在https: v-aura.notion.site上获得摘要:We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at https: v-aura.notion.site
【2】 A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
标题: 健全的描述:探索提示模板和类描述以增强Zero-Shot音频分类
作者:Michel Olvera,Paraskevas Stamatiadis,Slim Essid
备注:DCASE 2024 - 9th Workshop on Detection and Classification of Acoustic Scenes and Events, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
摘要:通过对比学习训练的音频文本模型提供了一种实用的方法,可以通过自然语言提示执行音频分类,例如“这是一个声音”后跟类别名称。在这项工作中,我们探索替代提示模板zero-shot音频分类,展示了更高性能的选项的存在。首先,我们发现提示的格式会显着影响性能,因此简单地用正确格式的类标签提示模型,就可以与优化的提示模板甚至提示集成竞争。此外,我们还研究了以音频为中心的描述来补充类标签。通过利用大型语言模型,我们生成文本描述,优先考虑声音事件的声学特征,以消除类之间的歧义,而无需大量的提示工程。我们表明,提示与类描述导致国家的最先进的结果在zero-shot音频分类在主要的环境声音数据集。值得注意的是,这种方法不需要额外的训练,并且保持完全的zero-shot。摘要:Audio-text models trained via contrastive learning offer a practical approach to perform audio classification through natural language prompts, such as "this is a sound of" followed by category names. In this work, we explore alternative prompt templates for zero-shot audio classification, demonstrating the existence of higher-performing options. First, we find that the formatting of the prompts significantly affects performance so that simply prompting the models with properly formatted class labels performs competitively with optimized prompt templates and even prompt ensembling. Moreover, we look into complementing class labels by audio-centric descriptions. By leveraging large language models, we generate textual descriptions that prioritize acoustic features of sound events to disambiguate between classes, without extensive prompt engineering. We show that prompting with class descriptions leads to state-of-the-art results in zero-shot audio classification across major ambient sound datasets. Remarkably, this method requires no additional training and remains fully zero-shot.
【3】 EMMeTT: Efficient Multimodal Machine Translation Training
标题: EMMeTT:高效的多模式机器翻译训练
作者:Piotr Żelasko,Zhehuai Chen,Mengru Wang,Daniel Galvez,Oleksii Hrinchuk,Shuoyang Ding,Ke Hu,Jagadeesh Balam,Vitaly Lavrukhin,Boris Ginsburg
备注:4 pages, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:对基础语言模型的模态扩展的兴趣日益增加,这就需要讨论最有效、最高效的多模态训练方法。这项工作的重点是神经机器翻译(NMT),并提出了一个联合多模式的语音LLM训练制度,包括自动语音翻译(AST)。我们研究了两种不同的基础模型架构,解码器只有GPT和编码器-解码器T5,扩展与金丝雀-1B的语音编码器。为了处理联合多模式训练,我们提出了一种新的训练框架,称为EMMeTT。EMMeTT通过以下方式提高了训练效率:跨语言,数据集和模态的平衡采样;高效的顺序数据迭代;以及针对多模态数据的新型2D桶化方案,并辅以批量大小优化器(OOMptimizer)。我们表明,多模态训练始终有助于这两种架构。此外,使用EMMeTT训练的SALM-T5保留了原始的NMT能力,同时在FLORES和FLEURS的四种语言子集上优于AST基线。由此产生的多模态翻译模型同时产生强大的文本和语音翻译结果。摘要:A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal training regime of Speech-LLM to include automatic speech translation (AST). We investigate two different foundation model architectures, decoder-only GPT and encoder-decoder T5, extended with Canary-1B's speech encoder. To handle joint multimodal training, we propose a novel training framework called EMMeTT. EMMeTT improves training efficiency with the following: balanced sampling across languages, datasets, and modalities; efficient sequential data iteration; and a novel 2D bucketing scheme for multimodal data, complemented by a batch size optimizer (OOMptimizer). We show that a multimodal training consistently helps with both architectures. Moreover, SALM-T5 trained with EMMeTT retains the original NMT capability while outperforming AST baselines on four-language subsets of FLORES and FLEURS. The resultant Multimodal Translation Model produces strong text and speech translation results at the same time.
【4】 LM-assisted keyword biasing with Aho-Corasick algorithm for Transducer-based ASR
标题: 使用Aho-Corasick算法的LM辅助关键字偏置用于基于传感器的ASB
作者:Iuliia Thorbecke,Juan Zuluaga-Gomez,Esaú Villatoro-Tello,Andres Carofilis,Shashi Kumar,Petr Motlicek,Karthik Pandia,Aravind Ganapathiraju
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:尽管自动语音识别的端到端模型最近取得了成功,但识别特殊的罕见词汇和词汇表外的单词,以及快速的文本域适应,仍然具有挑战性。经常发生的是,对特殊实体的偏见导致整体性能的下降。我们提出了一个轻的飞行方法,以提高自动语音识别性能相结合的偏见列表命名的实体与单词级的n-gram语言模型与浅融合方法的基础上的Aho-Corasick字符串匹配算法。Aho-Corasick算法已被证明比其他方法更有效,并且允许快速上下文自适应。一个n-gram语言模型被引入作为一个图的失败和输出弧,弧的权重是从n-gram概率。当语言模型与单个上下文图中的偏置实体组合以照顾整体性能时,语言模型被用作对关键字偏置的附加支持。我们在4种语言、2个公共数据集和1个私有数据集上展示了我们的研究结果,包括命名实体和词汇表外实体的性能。我们实现了高达21.6%的相对改善,在一般的字错误率没有实际的差异,在逆实时因素。摘要:Despite the recent success of end-to-end models for automatic speech recognition, recognizing special rare and out-of-vocabulary words, as well as fast domain adaptation with text, are still challenging. It often happens that biasing to the special entities leads to a degradation in the overall performance. We propose a light on-the-fly method to improve automatic speech recognition performance by combining a bias list of named entities with a word-level n-gram language model with the shallow fusion approach based on the Aho-Corasick string matching algorithm. The Aho-Corasick algorithm has proved to be more efficient than other methods and allows fast context adaptation. An n-gram language model is introduced as a graph with fail and output arcs, where the arc weights are adapted from the n-gram probabilities. The language model is used as an additional support to keyword biasing when the language model is combined with bias entities in a single context graph to take care of the overall performance. We demonstrate our findings on 4 languages, 2 public and 1 private datasets including performance on named entities and out-of-vocabulary entities. We achieve up to 21.6% relative improvement in the general word error rate with no practical difference in the inverse real-time factor.
【5】 Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
作者:Matthew Caren,Kartik Chandra,Joshua B. Tenenbaum,Jonathan Ragan-Kelley,Karima Ma
Journal-ref:SIGGRAPH Asia 2024
链接:点击下载PDF文件
摘要:我们提出了一种自动产生类似人类的声音模仿的方法:相当于“素描”,但听觉而不是视觉表示。从人类声道的模拟模型开始,我们首先尝试通过调整模型的控制参数来生成声音模仿,以使合成的发声在感知上突出的听觉特征方面与目标声音相匹配。然后,为了更好地匹配人类的直觉,我们应用了认知理论的沟通,考虑到人类扬声器的原因,他们的听众的战略。最后,我们通过几个实验和用户研究表明,当我们将这种类型的交流推理添加到我们的方法中时,它比单独匹配听觉特征更符合人类的直觉。这一观察结果对计算机图形学中的描绘研究具有广泛的意义。摘要:We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try generating vocal imitations by tuning the model's control parameters to make the synthesized vocalization match the target sound in terms of perceptually-salient auditory features. Then, to better match human intuitions, we apply a cognitive theory of communication to take into account how human speakers reason strategically about their listeners. Finally, we show through several experiments and user studies that when we add this type of communicative reasoning to our method, it aligns with human intuitions better than matching auditory features alone does. This observation has broad implications for the study of depiction in computer graphics.
【6】 Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
标题: 通过Whisper知识提炼快速流媒体传感器ASO原型
作者:Iuliia Thorbecke,Juan Zuluaga-Gomez,Esaú Villatoro-Tello,Shashi Kumar,Pradeep Rangappa,Sergio Burdisso,Petr Motlicek,Karthik Pandia,Aravind Ganapathiraju
备注:Accepted to EMNLP Findings 2024
链接:点击下载PDF文件
摘要:在几乎没有监督数据的情况下训练自动语音识别(ASR)仍然是一个悬而未决的问题。在这项工作中,我们证明了流转换器-转换器(TT)模型可以在消费者和可访问的GPU中从头开始训练,并完全使用来自基础语音模型(FSM)的伪标记(PL)语音。这允许仅在一个阶段中训练鲁棒的ASR模型,并且与具有预训练和微调的两步方案相比,不需要大量数据和计算预算。我们对基于PL的流TT模型的不同方面进行了全面的消融,例如(1)n-gram LM的浅融合,(2)命名实体的上下文偏置,(3)低延迟流应用的分块解码,以及(4)TT整体性能作为FSM大小的函数。我们的研究结果表明,TT可以从头开始训练,没有监督的数据,即使是非常嘈杂的PL。我们从CommonVoice的6种语言上验证了所提出的框架,并提出了多种算法来过滤出幻觉PL。摘要:The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (TT) models can be trained from scratch in consumer and accessible GPUs in their entirety with pseudo-labeled (PL) speech from foundational speech models (FSM). This allows training a robust ASR model just in one stage and does not require large data and computational budget compared to the two-step scenario with pre-training and fine-tuning. We perform a comprehensive ablation on different aspects of PL-based streaming TT models such as the impact of (1) shallow fusion of n-gram LMs, (2) contextual biasing with named entities, (3) chunk-wise decoding for low-latency streaming applications, and (4) TT overall performance as the function of the FSM size. Our results demonstrate that TT can be trained from scratch without supervised data, even with very noisy PLs. We validate the proposed framework on 6 languages from CommonVoice and propose multiple heuristics to filter out hallucinated PLs.
【7】 DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
标题: 迪夫声音:可区分的模式声音渲染和反向渲染,用于多样化推理任务
作者:Xutong Jin,Chenxi Xu,Ruohan Gao,Jiajun Wu,Guoping Wang,Sheng Li
备注:12 pages, 10 figures. Published in Siggraph 2024. Project page: this https URL
链接:点击下载PDF文件
摘要:从真实世界的录音中准确估计和模拟物体的物理特性在视觉、图形学和机器人领域具有重要的实际意义。然而,在这些方向上的进展受到限制-由于音频的高采样率,先前的可微分刚体或软体模拟技术不能直接应用于模态声音合成,而先前的音频合成器通常不能完全模拟发声对象的准确物理特性。我们提出了DiffSound,一个微分的声音渲染框架,基于物理的模态声音合成,这是基于一个隐式的形状表示,一个新的高阶有限元分析模块,和微分音频合成器。由于整个管道的可微性,我们的框架可以解决广泛的逆问题,包括物理参数估计,几何形状推理和碰撞位置预测。实验结果表明,我们的方法的有效性,突出了它的能力,准确地再现目标声音在一个基于物理的方式。DiffSound是各种声音合成和分析应用程序的宝贵工具。摘要:Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been limited -- prior differentiable rigid or soft body simulation techniques cannot be directly applied to modal sound synthesis due to the high sampling rate of audio, while previous audio synthesizers often do not fully model the accurate physical properties of the sounding objects. We propose DiffSound, a differentiable sound rendering framework for physics-based modal sound synthesis, which is based on an implicit shape representation, a new high-order finite element analysis module, and a differentiable audio synthesizer. Our framework can solve a wide range of inverse problems thanks to the differentiability of the entire pipeline, including physical parameter estimation, geometric shape reasoning, and impact position prediction. Experimental results demonstrate the effectiveness of our approach, highlighting its ability to accurately reproduce the target sound in a physics-based manner. DiffSound serves as a valuable tool for various sound synthesis and analysis applications.
【8】 Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
标题: 音频编解码器增强用于语音合成的鲁棒协作水印
作者:Lauri Juvela,Xin Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:合成语音的自动检测变得越来越重要,因为当前的合成方法与人类语音几乎无法区分,并且广泛地为公众所用。音频水印和其他主动披露方法正在吸引研究活动,因为它们可以补充基于被动检测的传统深度伪造防御。在主动和被动检测中,鲁棒性是主要的兴趣。传统的音频水印特别容易受到音频编解码器应用程序的移除攻击。大多数生成的语音和音频内容发布到野外纯粹作为分发方法通过音频编解码器。我们最近提出了协同水印的方法,使生成的语音更容易检测到的噪声,但可区分的传输通道。本文将通道增强扩展到不可区分的传统音频编解码器和神经音频编解码器,并评估各种配置下编解码器比特率的可转移性和效果。结果表明,黑盒音频编解码器可以使用波形域直通估计器进行梯度逼近来可靠地增强协作水印。此外,该结果表明,使用神经音频编解码器的通道增强可以很好地转移到传统编解码器。听力测试表明,协同水印在8kbps的高比特率编解码器或DAC中引起的感知退化可以忽略不计。摘要:Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to the public. Audio watermarking and other active disclosure methods of are attracting research activity, as they can complement traditional deepfake defenses based on passive detection. In both active and passive detection, robustness is of major interest. Traditional audio watermarks are particularly susceptible to removal attacks by audio codec application. Most generated speech and audio content released into the wild passes through an audio codec purely as a distribution method. We recently proposed collaborative watermarking as method for making generated speech more easily detectable over a noisy but differentiable transmission channel. This paper extends the channel augmentation to work with non-differentiable traditional audio codecs and neural audio codecs and evaluates transferability and effect of codec bitrate over various configurations. The results show that collaborative watermarking can be reliably augmented by black-box audio codecs using a waveform-domain straight-through-estimator for gradient approximation. Furthermore, that results show that channel augmentation with a neural audio codec transfers well to traditional codecs. Listening tests demonstrate collaborative watermarking incurs negligible perceptual degradation with high bitrate codecs or DAC at 8kbps.
【9】 Beyond the binary: Limitations and possibilities of gender-related speech technology research
标题: 超越二元:性别相关语音技术研究的局限性和可能性
作者:Ariadna Sanchez,Alice Ross,Nina Markl
备注:Accepted at Spoken Language Technology (SLT) Workshop 2024
链接:点击下载PDF文件
摘要:本文回顾了2013年至2023年ISCA Interspeech出版物中107篇与言语和性别或性别相关的研究论文。我们注意到关于这一主题的工作的稀缺性,并发现术语,特别是 textit{gender}一词的使用方式不够明确,而且往往与社会科学中的流行观点不一致,即性别是社会构建的,是一个光谱,而不是二元类别。我们提请注意潜在的问题,这可能会导致已经被边缘化的群体,并提出一些问题,研究人员问自己时,进行工作的言论和性别。摘要:This paper presents a review of 107 research papers relating to speech and sex or gender in ISCA Interspeech publications between 2013 and 2023. We note the scarcity of work on this topic and find that terminology, particularly the word textit{gender}, is used in ways that are underspecified and often out of step with the prevailing view in social sciences that gender is socially constructed and is a spectrum as opposed to a binary category. We draw attention to the potential problems that this can cause for already marginalised groups, and suggest some questions for researchers to ask themselves when undertaking work on speech and gender.
【10】 Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
标题: 大型语言模型应了解拼音才能进行中文ASB错误纠正
作者:Yuang Li,Xiaosong Qiao,Xiaofeng Zhao,Huan Zhao,Wei Tang,Min Zhang,Hao Yang
链接:点击下载PDF文件
摘要:大型语言模型可以通过生成纠错来增强自动语音识别系统。在本文中,我们提出了拼音增强的GEC,它利用拼音,汉语普通话的语音表示,作为补充信息,以提高中国ASR纠错。我们的方法只利用合成误差进行训练,并在推理过程中采用最佳假设。此外,我们引入了一种多任务训练方法,涉及拼音和文本之间的转换任务,以对齐它们的特征空间。在Aishell-1和Common Voice数据集上的实验表明,我们的方法在纯文本输入的情况下始终优于GEC。更重要的是,我们从两个方面为PY-GEC和多任务训练的有效性提供了直观的解释:1)增加了对拼音特征的注意力权重; 2)拼音和文本隐藏状态之间的对齐特征空间。摘要:Large language models can enhance automatic speech recognition systems through generative error correction. In this paper, we propose Pinyin-enhanced GEC, which leverages Pinyi, the phonetic representation of Mandarin Chinese, as supplementary information to improve Chinese ASR error correction. Our approach only utilizes synthetic errors for training and employs the one-best hypothesis during inference. Additionally, we introduce a multitask training approach involving conversion tasks between Pinyin and text to align their feature spaces. Experiments on the Aishell-1 and the Common Voice datasets demonstrate that our approach consistently outperforms GEC with text-only input. More importantly, we provide intuitive explanations for the effectiveness of PY-GEC and multitask training from two aspects: 1) increased attention weight on Pinyin features; and 2) aligned feature space between Pinyin and text hidden states.
【11】 MuCodec: Ultra Low-Bitrate Music Codec
标题: MuCodec:超低比特率音乐编解码器
作者:Yaoxun Xu,Hangting Chen,Jianwei Yu,Wei Tan,Rongzhi Gu,Shun Lei,Zhiwei Lin,Zhiyong Wu
链接:点击下载PDF文件
摘要:音乐编解码器是音频编解码器研究的一个重要方面,超低比特率压缩对于音乐传输和生成具有重要意义。由于音乐背景的复杂性和人声的丰富性,仅仅依靠语义或声学信息建模无法有效地重建具有人声和背景的音乐。为了解决这个问题,我们提出了MuCodec,专门针对超低比特率的音乐压缩和重建任务。MuCodec采用MuEncoder提取语音和语义特征,用RVQ离散化,通过流匹配得到Mel-VAE特征。然后使用预先训练的MEL-VAE解码器和HiFi-GAN重建音乐。MuCodec能够以超低(0.35kbps)或高比特率(1.35kbps)重建高保真音乐,在主观和客观指标上都达到了迄今为止的最佳效果。代码和演示:https: xuyaoxun.github.io MuCodec_demo 。摘要:Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity of music backgrounds and the richness of vocals, solely relying on modeling semantic or acoustic information cannot effectively reconstruct music with both vocals and backgrounds. To address this issue, we propose MuCodec, specifically targeting music compression and reconstruction tasks at ultra low bitrates. MuCodec employs MuEncoder to extract both acoustic and semantic features, discretizes them with RVQ, and obtains Mel-VAE features via flow-matching. The music is then reconstructed using a pre-trained MEL-VAE decoder and HiFi-GAN. MuCodec can reconstruct high-fidelity music at ultra low (0.35kbps) or high bitrates (1.35kbps), achieving the best results to date in both subjective and objective metrics. Code and Demo: https: xuyaoxun.github.io MuCodec_demo .
【12】 Personalized Speech Recognition for Children with Test-Time Adaptation
标题: 具有测试时间适应的儿童个性化语音识别
作者:Zhonghao Shi,Harshvardhan Srivastava,Xuan Shi,Shrikanth Narayanan,Maja J. Matarić
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
摘要:准确的儿童自动语音识别(ASR)对于有效的实时儿童-AI交互至关重要,特别是在教育应用中。然而,由于数据域从成人转移到儿童,主要在成人数据上进行预训练的现成ASR模型往往对儿童语音的泛化能力很差。最近的研究发现,对儿童语音数据进行监督微调可以帮助弥合这种域转移,但对于现实世界的应用程序来说,人类注释可能是不切实际的,并且训练时的适应可能会忽略测试时发生的额外域转移。我们设计了一种新型的ASR管道,将无监督测试时自适应(TTA)方法应用于儿童语音识别,以便根据成人语音预训练的ASR模型可以在测试时连续适应每个儿童说话者,而无需进一步的人工注释。我们的研究结果表明,与TTA方法相适应的ASR模型显着优于未经调整的现成的ASR基线的平均和统计在个别儿童扬声器。我们的分析还发现,显着的数据域之间的儿童扬声器和每个儿童扬声器内的变化,这进一步激发了测试时间适应的需要。摘要:Accurate automatic speech recognition (ASR) for children is crucial for effective real-time child-AI interaction, especially in educational applications. However, off-the-shelf ASR models primarily pre-trained on adult data tend to generalize poorly to children's speech due to the data domain shift from adults to children. Recent studies have found that supervised fine-tuning on children's speech data can help bridge this domain shift, but human annotations may be impractical to obtain for real-world applications and adaptation at training time can overlook additional domain shifts occurring at test time. We devised a novel ASR pipeline to apply unsupervised test-time adaptation (TTA) methods for child speech recognition, so that ASR models pre-trained on adult speech can be continuously adapted to each child speaker at test time without further human annotations. Our results show that ASR models adapted with TTA methods significantly outperform the unadapted off-the-shelf ASR baselines both on average and statistically across individual child speakers. Our analysis also discovered significant data domain shifts both between child speakers and within each child speaker, which further motivates the need for test-time adaptation.
【13】 DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
标题: 迪夫编辑器:通过语义丰富和声学一致性增强语音编辑
作者:Yang Chen,Yuhang Jia,Shiwan Zhao,Ziyue Jiang,Haoran Li,Jiarong Kang,Yong Qin
链接:点击下载PDF文件
摘要:随着基于文本的语音编辑变得越来越普遍,对无限制的自由文本编辑的需求持续增长。然而,现有的语音编辑技术遇到了重大的挑战,特别是在处理域外(OOD)文本时保持可懂度和声学一致性。在本文中,我们介绍,DiffEditor,一种新的语音编辑模型,旨在提高性能,在OOD文本的情况下,通过语义丰富和声学的一致性。为了提高编辑后语音的清晰度,我们通过整合从预训练语言模型中提取的词嵌入来丰富音素嵌入的语义信息。此外,我们强调,帧间平滑属性是建模声学一致性的关键,因此,我们提出了一个一阶损失函数,以促进编辑边界处的平滑过渡,并提高编辑语音的整体流畅性。实验结果表明,我们的模型实现了国家的最先进的性能在域和面向对象的文本场景。摘要:As text-based speech editing becomes increasingly prevalent, the demand for unrestricted free-text editing continues to grow. However, existing speech editing techniques encounter significant challenges, particularly in maintaining intelligibility and acoustic consistency when dealing with out-of-domain (OOD) text. In this paper, we introduce, DiffEditor, a novel speech editing model designed to enhance performance in OOD text scenarios through semantic enrichment and acoustic consistency. To improve the intelligibility of the edited speech, we enrich the semantic information of phoneme embeddings by integrating word embeddings extracted from a pretrained language model. Furthermore, we emphasize that interframe smoothing properties are critical for modeling acoustic consistency, and thus we propose a first-order loss function to promote smoother transitions at editing boundaries and enhance the overall fluency of the edited speech. Experimental results demonstrate that our model achieves state-of-the-art performance in both in-domain and OOD text scenarios.
【14】 Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection
标题: 时间和令牌:端到端语音流畅性检测基准
作者:Xuanru Zhou,Jiachen Lian,Cheol Jun Cho,Jingwen Liu,Zongli Ye,Jinming Zhang,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Maria Luisa Gorno Tempini,Gopala Anumanchipalli
链接:点击下载PDF文件
摘要:语音不流畅建模是检测语音中的重复、阻塞、插入、替换和删除等不流畅现象的任务。最新的进展对待这个问题作为一个基于时间的对象检测问题。在这项工作中,我们重新审视这个问题,从一个新的角度:tokenizing dysfluencies和建模的检测问题作为一个基于令牌的自动语音识别(ASR)的问题。我们提出了基于规则的语音和文本不流畅模拟器,并开发了VCTK-token,然后开发了一个类似Whisper的seq 2seq架构,以建立一个新的基准测试。我们还系统地比较了我们提出的基于令牌的方法与基于时间的方法,并提出了一个统一的基准,以促进未来的研究工作。我们为更广泛的科学界开放这些资源。该项目的网页可在https: rorizzz.github.io 上找到摘要:Speech dysfluency modeling is a task to detect dysfluencies in speech, such as repetition, block, insertion, replacement, and deletion. Most recent advancements treat this problem as a time-based object detection problem. In this work, we revisit this problem from a new perspective: tokenizing dysfluencies and modeling the detection problem as a token-based automatic speech recognition (ASR) problem. We propose rule-based speech and text dysfluency simulators and develop VCTK-token, and then develop a Whisper-like seq2seq architecture to build a new benchmark with decent performance. We also systematically compare our proposed token-based methods with time-based methods, and propose a unified benchmark to facilitate future research endeavors. We open-source these resources for the broader scientific community. The project page is available at https: rorizzz.github.io
【15】 Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
标题: 神经方向过滤:使用小型麦克风阵列进行远场方向性控制
作者:Julian Wechsler,Srikanth Raj Chetupalli,Mhd Modar Halimeh,Oliver Thiergart,Emanuël A. P. Habets
备注:Presented at the International Workshop on Acoustic Signal Enhancement (IWAENC), 2024
链接:点击下载PDF文件
摘要:捕获具有特定方向性模式的音频信号在语音通信中是必不可少的。这项研究提出了一种基于深度神经网络(DNN)的定向滤波方法,减轻了对显式信号模型的需求。更具体地说,我们提出的方法使用DNN从麦克风阵列的信号中估计单通道复杂掩码。然后将该掩模应用于参考麦克风以呈现呈现呈现期望的方向性图案的信号。我们研究了训练数据集的组成及其对DNN在推理过程中实现的方向性的影响。使用相对较小的DNN,发现所提出的方法接近所需的方向性图案。此外,它允许使用少量麦克风实现高阶方向性图案,这对于线性和参数方向滤波来说是一项困难的任务。摘要:Capturing audio signals with specific directivity patterns is essential in speech communication. This study presents a deep neural network (DNN)-based approach to directional filtering, alleviating the need for explicit signal models. More specifically, our proposed method uses a DNN to estimate a single-channel complex mask from the signals of a microphone array. This mask is then applied to a reference microphone to render a signal that exhibits a desired directivity pattern. We investigate the training dataset composition and its effect on the directivity realized by the DNN during inference. Using a relatively small DNN, the proposed method is found to approximate the desired directivity pattern closely. Additionally, it allows for the realization of higher-order directivity patterns using a small number of microphones, which is a difficult task for linear and parametric directional filtering.
【16】 Exploring Text-Queried Sound Event Detection with Audio Source Separation
标题: 探索使用音频源分离的文本查询声音事件检测
作者:Han Yin,Jisheng Bai,Yang Xiao,Hui Wang,Siqi Zheng,Yafeng Chen,Rohan Kumar Das,Chong Deng,Jianfeng Chen
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:在声音事件检测(SED)中,重叠的声音事件构成了一个重大的挑战,因为某些事件可以很容易地被背景噪声或其他事件掩盖,导致检测性能差。为了解决这个问题,我们提出了文本查询SED(TQ-SED)框架。具体来说,我们首先预训练语言查询音频源分离(LASS)模型,以从输入音频中分离出与不同事件对应的音轨。然后,采用多个目标SED分支来检测各个事件。AudioSep是最先进的LASS模型,但由于其用于分离的纯卷积结构,在提取动态音频信息方面存在局限性。为了解决这个问题,我们将双路径递归神经网络块集成到模型中。我们将这种结构称为AudioSep-DP,它在DCASE 2024任务9中实现了语言查询音频源分离(客观单一模型轨道)的第一名。实验结果表明,TQ-SED能显著提高SED性能,与传统框架相比,F1分数提高了7.22%.此外,我们还设置了全面的实验来探索模型复杂性的影响。源代码和预训练模型在https: github.com apple-yinhan TQ-SED上发布。摘要:In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22 % on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https: github.com apple-yinhan TQ-SED.
【17】 LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
标题: LiSenNet:用于实时语音增强的轻量级子带和双路径建模
作者:Haoyin Yan,Jie Zhang,Cunhang Fan,Yeping Zhou,Peiqi Liu
备注:5 pages, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
链接:点击下载PDF文件
摘要:语音增强的目的是从被噪声污染的测量信号中提取出干净的波形,以提高语音质量和可懂度。虽然基于学习的方法可以比传统方法表现得更好,但计算复杂性和模型大小严重限制了在延迟敏感和低资源边缘设备上的部署。在这项工作中,我们提出了一个轻量级的SE网络(LiSenNet)的实时应用。我们设计了子带下采样和上采样模块以及双路径递归模块,分别用于捕获频带感知特征和时频模式。噪声检测器被开发用于检测噪声区域,以自适应地执行SE并节省计算成本。与最近的更高资源依赖的基线模型相比,所提出的LiSenNet可以实现具有竞争力的性能,仅需37k个参数(最先进模型的一半)和每秒56M的乘法累加(MAC)操作。摘要:Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the large computational complexity and model size heavily limit the deployment on latency-sensitive and low-resource edge devices. In this work, we propose a lightweight SE network (LiSenNet) for real-time applications. We design sub-band downsampling and upsampling blocks and a dual-path recurrent module to capture band-aware features and time-frequency patterns, respectively. A noise detector is developed to detect noisy regions in order to perform SE adaptively and save computational costs. Compared to recent higher-resource-dependent baseline models, the proposed LiSenNet can achieve a competitive performance with only 37k parameters (half of the state-of-the-art model) and 56M multiply-accumulate (MAC) operations per second.
【18】 Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
标题: 利用纯音频数据进行文本查询目标声音提取
作者:Kohei Saijo,Janek Ebbers,François G. Germain,Sameer Khurana,Gordon Wichern,Jonathan Le Roux
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:文本查询的目标声音提取(TSE)的目标是从一个混合的声音源指定的自然语言的标题。虽然最好能够访问大规模的文本-音频对来处理各种文本提示,但是有限数量的可用高质量文本-音频对阻碍了数据缩放。为此,这项工作探讨了如何利用纯音频数据,没有任何字幕的文本查询TSE任务,以潜在地扩大数据量。这样做的一种直接方式是使用联合音频-文本嵌入模型(诸如对比语言-音频预训练(CLAP)模型)作为查询编码器,并使用从地面实况音频获得的音频嵌入来训练TSE模型。然后,TSE模型可以在推理时通过切换到文本编码器来接受文本查询。虽然如果CLAP中的音频和文本嵌入空间很好地对齐,这种方法应该有效,但在实践中,嵌入具有特定于域的信息,导致TSE模型过拟合音频查询。我们研究了几种方法来避免过拟合,并表明简单的嵌入操作方法,如dropout可以有效地缓解这个问题。大量的实验表明,在训练过程中使用带有嵌入丢弃的纯音频数据与使用文本标题一样有效,并且纯音频数据可以有效地用于改进文本查询的TSE模型。摘要:The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text prompts, the limited number of available high-quality text-audio pairs hinders the data scaling. To this end, this work explores how to leverage audio-only data without any captions for the text-queried TSE task to potentially scale up the data amount. A straightforward way to do so is to use a joint audio-text embedding model, such as the contrastive language-audio pre-training (CLAP) model, as a query encoder and train a TSE model using audio embeddings obtained from the ground-truth audio. The TSE model can then accept text queries at inference time by switching to the text encoder. While this approach should work if the audio and text embedding spaces in CLAP were well aligned, in practice, the embeddings have domain-specific information that causes the TSE model to overfit to audio queries. We investigate several methods to avoid overfitting and show that simple embedding-manipulation methods such as dropout can effectively alleviate this issue. Extensive experiments demonstrate that using audio-only data with embedding dropout is as effective as using text captions during training, and audio-only data can be effectively leveraged to improve text-queried TSE models.
【19】 DiffSSD: A Diffusion-Based Dataset For Speech Forensics
标题: 迪夫SSD:用于语音取证的基于扩散的数据集
作者:Kratika Bhagtani,Amit Kumar Singh Yadav,Paolo Bestagini,Edward J. Delp
备注:Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2025
链接:点击下载PDF文件
摘要:基于扩散的语音生成器无处不在。这些方法可以生成非常高质量的合成语音,最近的几起事件报告了它们的恶意使用。为了对抗这种滥用,已经开发了合成语音检测器。这些检测器中的许多是在不包括基于扩散的合成器的数据集上训练的。在本文中,我们证明了在这样一个数据集ASVspoof2019上训练的现有检测器在检测最近基于扩散的合成器的合成语音方面表现不佳。我们提出了基于扩散的合成语音数据集(DiffSSD),该数据集由大约200小时的标记语音组成,包括由8个基于扩散的开源和2个商业生成器生成的合成语音。我们还研究了现有的合成语音检测器在闭集和开集的情况下的DiffSSD的性能。结果突出了该数据集在检测最近的开源和商业语音生成器生成的合成语音方面的重要性。摘要:Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do not include diffusion-based synthesizers. In this paper, we demonstrate that existing detectors trained on one such dataset, ASVspoof2019, do not perform well in detecting synthetic speech from recent diffusion-based synthesizers. We propose the Diffusion-Based Synthetic Speech Dataset (DiffSSD), a dataset consisting of about 200 hours of labeled speech, including synthetic speech generated by 8 diffusion-based open-source and 2 commercial generators. We also examine the performance of existing synthetic speech detectors on DiffSSD in both closed-set and open-set scenarios. The results highlight the importance of this dataset in detecting synthetic speech generated from recent open-source and commercial speech generators.
eess.AS音频处理
【1】 Time and Tokens: Benchmarking End-to-End Speech Dysfluency Detection标题: 时间和令牌:端到端语音流畅性检测基准
作者:Xuanru Zhou,Jiachen Lian,Cheol Jun Cho,Jingwen Liu,Zongli Ye,Jinming Zhang,Brittany Morin,David Baquirin,Jet Vonk,Zoe Ezzes,Zachary Miller,Maria Luisa Gorno Tempini,Gopala Anumanchipalli
链接:点击下载PDF文件
摘要:语音不流畅建模是检测语音中的重复、阻塞、插入、替换和删除等不流畅现象的任务。最新的进展对待这个问题作为一个基于时间的对象检测问题。在这项工作中,我们重新审视这个问题,从一个新的角度:tokenizing dysfluencies和建模的检测问题作为一个基于令牌的自动语音识别(ASR)的问题。我们提出了基于规则的语音和文本不流畅模拟器,并开发了VCTK-token,然后开发了一个类似Whisper的seq 2seq架构,以建立一个新的基准测试。我们还系统地比较了我们提出的基于令牌的方法与基于时间的方法,并提出了一个统一的基准,以促进未来的研究工作。我们为更广泛的科学界开放这些资源。该项目的网页可在https: rorizzz.github.io 上找到摘要:Speech dysfluency modeling is a task to detect dysfluencies in speech, such as repetition, block, insertion, replacement, and deletion. Most recent advancements treat this problem as a time-based object detection problem. In this work, we revisit this problem from a new perspective: tokenizing dysfluencies and modeling the detection problem as a token-based automatic speech recognition (ASR) problem. We propose rule-based speech and text dysfluency simulators and develop VCTK-token, and then develop a Whisper-like seq2seq architecture to build a new benchmark with decent performance. We also systematically compare our proposed token-based methods with time-based methods, and propose a unified benchmark to facilitate future research endeavors. We open-source these resources for the broader scientific community. The project page is available at https: rorizzz.github.io
【2】 Neural Directional Filtering: Far-Field Directivity Control With a Small Microphone Array
标题: 神经方向过滤:使用小型麦克风阵列进行远场方向性控制
作者:Julian Wechsler,Srikanth Raj Chetupalli,Mhd Modar Halimeh,Oliver Thiergart,Emanuël A. P. Habets
备注:Presented at the International Workshop on Acoustic Signal Enhancement (IWAENC), 2024
链接:点击下载PDF文件
摘要:捕获具有特定方向性模式的音频信号在语音通信中是必不可少的。这项研究提出了一种基于深度神经网络(DNN)的定向滤波方法,减轻了对显式信号模型的需求。更具体地说,我们提出的方法使用DNN从麦克风阵列的信号中估计单通道复杂掩码。然后将该掩模应用于参考麦克风以呈现呈现呈现期望的方向性图案的信号。我们研究了训练数据集的组成及其对DNN在推理过程中实现的方向性的影响。使用相对较小的DNN,发现所提出的方法接近所需的方向性图案。此外,它允许使用少量麦克风实现高阶方向性图案,这对于线性和参数方向滤波来说是一项困难的任务。摘要:Capturing audio signals with specific directivity patterns is essential in speech communication. This study presents a deep neural network (DNN)-based approach to directional filtering, alleviating the need for explicit signal models. More specifically, our proposed method uses a DNN to estimate a single-channel complex mask from the signals of a microphone array. This mask is then applied to a reference microphone to render a signal that exhibits a desired directivity pattern. We investigate the training dataset composition and its effect on the directivity realized by the DNN during inference. Using a relatively small DNN, the proposed method is found to approximate the desired directivity pattern closely. Additionally, it allows for the realization of higher-order directivity patterns using a small number of microphones, which is a difficult task for linear and parametric directional filtering.
【3】 Exploring Text-Queried Sound Event Detection with Audio Source Separation
标题: 探索使用音频源分离的文本查询声音事件检测
作者:Han Yin,Jisheng Bai,Yang Xiao,Hui Wang,Siqi Zheng,Yafeng Chen,Rohan Kumar Das,Chong Deng,Jianfeng Chen
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:在声音事件检测(SED)中,重叠的声音事件构成了一个重大的挑战,因为某些事件可以很容易地被背景噪声或其他事件掩盖,导致检测性能差。为了解决这个问题,我们提出了文本查询SED(TQ-SED)框架。具体来说,我们首先预训练语言查询音频源分离(LASS)模型,以从输入音频中分离出与不同事件对应的音轨。然后,采用多个目标SED分支来检测各个事件。AudioSep是最先进的LASS模型,但由于其用于分离的纯卷积结构,在提取动态音频信息方面存在局限性。为了解决这个问题,我们将双路径递归神经网络块集成到模型中。我们将这种结构称为AudioSep-DP,它在DCASE 2024任务9中实现了语言查询音频源分离(客观单一模型轨道)的第一名。实验结果表明,TQ-SED能显著提高SED性能,与传统框架相比,F1分数提高了7.22%.此外,我们还设置了全面的实验来探索模型复杂性的影响。源代码和预训练模型已在www.example.com上发布。摘要:In sound event detection (SED), overlapping sound events pose a significant challenge, as certain events can be easily masked by background noise or other events, resulting in poor detection performance. To address this issue, we propose the text-queried SED (TQ-SED) framework. Specifically, we first pre-train a language-queried audio source separation (LASS) model to separate the audio tracks corresponding to different events from the input audio. Then, multiple target SED branches are employed to detect individual events. AudioSep is a state-of-the-art LASS model, but has limitations in extracting dynamic audio information because of its pure convolutional structure for separation. To address this, we integrate a dual-path recurrent neural network block into the model. We refer to this structure as AudioSep-DP, which achieves the first place in DCASE 2024 Task 9 on language-queried audio source separation (objective single model track). Experimental results show that TQ-SED can significantly improve the SED performance, with an improvement of 7.22 % on F1 score over the conventional framework. Additionally, we setup comprehensive experiments to explore the impact of model complexity. The source code and pre-trained model are released at https: github.com apple-yinhan TQ-SED.
【4】 LiSenNet: Lightweight Sub-band and Dual-Path Modeling for Real-Time Speech Enhancement
标题: LiSenNet:用于实时语音增强的轻量级子带和双路径建模
作者:Haoyin Yan,Jie Zhang,Cunhang Fan,Yeping Zhou,Peiqi Liu
备注:5 pages, submitted to 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2025)
链接:点击下载PDF文件
摘要:语音增强的目的是从被噪声污染的测量信号中提取出干净的波形,以提高语音质量和可懂度。虽然基于学习的方法可以比传统方法表现得更好,但计算复杂性和模型大小严重限制了在延迟敏感和低资源边缘设备上的部署。在这项工作中,我们提出了一个轻量级的SE网络(LiSenNet)的实时应用。我们设计了子带下采样和上采样模块以及双路径递归模块,分别用于捕获频带感知特征和时频模式。噪声检测器被开发用于检测噪声区域,以自适应地执行SE并节省计算成本。与最近的更高资源依赖的基线模型相比,所提出的LiSenNet可以实现具有竞争力的性能,仅需37k个参数(最先进模型的一半)和每秒56M的乘法累加(MAC)操作。摘要:Speech enhancement (SE) aims to extract the clean waveform from noise-contaminated measurements to improve the speech quality and intelligibility. Although learning-based methods can perform much better than traditional counterparts, the large computational complexity and model size heavily limit the deployment on latency-sensitive and low-resource edge devices. In this work, we propose a lightweight SE network (LiSenNet) for real-time applications. We design sub-band downsampling and upsampling blocks and a dual-path recurrent module to capture band-aware features and time-frequency patterns, respectively. A noise detector is developed to detect noisy regions in order to perform SE adaptively and save computational costs. Compared to recent higher-resource-dependent baseline models, the proposed LiSenNet can achieve a competitive performance with only 37k parameters (half of the state-of-the-art model) and 56M multiply-accumulate (MAC) operations per second.
【5】 Leveraging Audio-Only Data for Text-Queried Target Sound Extraction
标题: 利用纯音频数据进行文本查询目标声音提取
作者:Kohei Saijo,Janek Ebbers,François G. Germain,Sameer Khurana,Gordon Wichern,Jonathan Le Roux
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:文本查询的目标声音提取(TSE)的目标是从一个混合的声音源指定的自然语言的标题。虽然最好能够访问大规模的文本-音频对来处理各种文本提示,但是有限数量的可用高质量文本-音频对阻碍了数据缩放。为此,这项工作探讨了如何利用纯音频数据,没有任何字幕的文本查询TSE任务,以潜在地扩大数据量。这样做的一种直接方式是使用联合音频-文本嵌入模型(诸如对比语言-音频预训练(CLAP)模型)作为查询编码器,并使用从地面实况音频获得的音频嵌入来训练TSE模型。然后,TSE模型可以在推理时通过切换到文本编码器来接受文本查询。虽然如果CLAP中的音频和文本嵌入空间很好地对齐,这种方法应该有效,但在实践中,嵌入具有特定于域的信息,导致TSE模型过拟合音频查询。我们研究了几种方法来避免过拟合,并表明简单的嵌入操作方法,如dropout可以有效地缓解这个问题。大量的实验表明,在训练过程中使用带有嵌入丢弃的纯音频数据与使用文本标题一样有效,并且纯音频数据可以有效地用于改进文本查询的TSE模型。摘要:The goal of text-queried target sound extraction (TSE) is to extract from a mixture a sound source specified with a natural-language caption. While it is preferable to have access to large-scale text-audio pairs to address a variety of text prompts, the limited number of available high-quality text-audio pairs hinders the data scaling. To this end, this work explores how to leverage audio-only data without any captions for the text-queried TSE task to potentially scale up the data amount. A straightforward way to do so is to use a joint audio-text embedding model, such as the contrastive language-audio pre-training (CLAP) model, as a query encoder and train a TSE model using audio embeddings obtained from the ground-truth audio. The TSE model can then accept text queries at inference time by switching to the text encoder. While this approach should work if the audio and text embedding spaces in CLAP were well aligned, in practice, the embeddings have domain-specific information that causes the TSE model to overfit to audio queries. We investigate several methods to avoid overfitting and show that simple embedding-manipulation methods such as dropout can effectively alleviate this issue. Extensive experiments demonstrate that using audio-only data with embedding dropout is as effective as using text captions during training, and audio-only data can be effectively leveraged to improve text-queried TSE models.
【6】 DiffSSD: A Diffusion-Based Dataset For Speech Forensics
标题: 迪夫SSD:用于语音取证的基于扩散的数据集
作者:Kratika Bhagtani,Amit Kumar Singh Yadav,Paolo Bestagini,Edward J. Delp
备注:Submitted to IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP) 2025
链接:点击下载PDF文件
摘要:基于扩散的语音生成器无处不在。这些方法可以生成非常高质量的合成语音,最近的几起事件报告了它们的恶意使用。为了对抗这种滥用,已经开发了合成语音检测器。这些检测器中的许多是在不包括基于扩散的合成器的数据集上训练的。在本文中,我们证明了在这样一个数据集ASVspoof2019上训练的现有检测器在检测最近基于扩散的合成器的合成语音方面表现不佳。我们提出了基于扩散的合成语音数据集(DiffSSD),该数据集由大约200小时的标记语音组成,包括由8个基于扩散的开源和2个商业生成器生成的合成语音。我们还研究了现有的合成语音检测器在闭集和开集的情况下的DiffSSD的性能。结果突出了该数据集在检测最近的开源和商业语音生成器生成的合成语音方面的重要性。摘要:Diffusion-based speech generators are ubiquitous. These methods can generate very high quality synthetic speech and several recent incidents report their malicious use. To counter such misuse, synthetic speech detectors have been developed. Many of these detectors are trained on datasets which do not include diffusion-based synthesizers. In this paper, we demonstrate that existing detectors trained on one such dataset, ASVspoof2019, do not perform well in detecting synthetic speech from recent diffusion-based synthesizers. We propose the Diffusion-Based Synthetic Speech Dataset (DiffSSD), a dataset consisting of about 200 hours of labeled speech, including synthetic speech generated by 8 diffusion-based open-source and 2 commercial generators. We also examine the performance of existing synthetic speech detectors on DiffSSD in both closed-set and open-set scenarios. The results highlight the importance of this dataset in detecting synthetic speech generated from recent open-source and commercial speech generators.
【7】 Temporally Aligned Audio for Video with Autoregression
标题: 具有自动回归的视频时间对齐音频
作者:Ilpo Viertola,Vladimir Iashin,Esa Rahtu
备注:Submitted to ICASSP 2025. Project page this https URL
链接:点击下载PDF文件
摘要:None摘要:We introduce V-AURA, the first autoregressive model to achieve high temporal alignment and relevance in video-to-audio generation. V-AURA uses a high-framerate visual feature extractor and a cross-modal audio-visual feature fusion strategy to capture fine-grained visual motion events and ensure precise temporal alignment. Additionally, we propose VisualSound, a benchmark dataset with high audio-visual relevance. VisualSound is based on VGGSound, a video dataset consisting of in-the-wild samples extracted from YouTube. During the curation, we remove samples where auditory events are not aligned with the visual ones. V-AURA outperforms current state-of-the-art models in temporal alignment and semantic relevance while maintaining comparable audio quality. Code, samples, VisualSound and models are available at https: v-aura.notion.site
【8】 A sound description: Exploring prompt templates and class descriptions to enhance zero-shot audio classification
标题: 健全的描述:探索提示模板和类描述以增强Zero-Shot音频分类
作者:Michel Olvera,Paraskevas Stamatiadis,Slim Essid
备注:DCASE 2024 - 9th Workshop on Detection and Classification of Acoustic Scenes and Events, Oct 2024, Tokyo, Japan
链接:点击下载PDF文件
摘要:通过对比学习训练的音频文本模型提供了一种实用的方法,可以通过自然语言提示执行音频分类,例如“这是一个声音”后跟类别名称。在这项工作中,我们探索替代提示模板zero-shot音频分类,展示了更高性能的选项的存在。首先,我们发现提示的格式会显着影响性能,因此简单地用正确格式的类标签提示模型,就可以与优化的提示模板甚至提示集成竞争。此外,我们还研究了以音频为中心的描述来补充类标签。通过利用大型语言模型,我们生成文本描述,优先考虑声音事件的声学特征,以消除类之间的歧义,而无需大量的提示工程。我们表明,提示与类描述导致国家的最先进的结果在zero-shot音频分类在主要的环境声音数据集。值得注意的是,这种方法不需要额外的训练,并且保持完全的zero-shot。摘要:Audio-text models trained via contrastive learning offer a practical approach to perform audio classification through natural language prompts, such as "this is a sound of" followed by category names. In this work, we explore alternative prompt templates for zero-shot audio classification, demonstrating the existence of higher-performing options. First, we find that the formatting of the prompts significantly affects performance so that simply prompting the models with properly formatted class labels performs competitively with optimized prompt templates and even prompt ensembling. Moreover, we look into complementing class labels by audio-centric descriptions. By leveraging large language models, we generate textual descriptions that prioritize acoustic features of sound events to disambiguate between classes, without extensive prompt engineering. We show that prompting with class descriptions leads to state-of-the-art results in zero-shot audio classification across major ambient sound datasets. Remarkably, this method requires no additional training and remains fully zero-shot.
【9】 EMMeTT: Efficient Multimodal Machine Translation Training
标题: EMMeTT:高效的多模式机器翻译训练
作者:Piotr Żelasko,Zhehuai Chen,Mengru Wang,Daniel Galvez,Oleksii Hrinchuk,Shuoyang Ding,Ke Hu,Jagadeesh Balam,Vitaly Lavrukhin,Boris Ginsburg
备注:4 pages, submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:对基础语言模型的模态扩展的兴趣日益增加,这就需要讨论最有效、最高效的多模态训练方法。这项工作的重点是神经机器翻译(NMT),并提出了一个联合多模式的语音LLM训练制度,包括自动语音翻译(AST)。我们研究了两种不同的基础模型架构,解码器只有GPT和编码器-解码器T5,扩展与金丝雀-1B的语音编码器。为了处理联合多模式训练,我们提出了一种新的训练框架,称为EMMeTT。EMMeTT通过以下方式提高了训练效率:跨语言,数据集和模态的平衡采样;高效的顺序数据迭代;以及针对多模态数据的新型2D桶化方案,并辅以批量大小优化器(OOMptimizer)。我们表明,多模态训练始终有助于这两种架构。此外,使用EMMeTT训练的SALM-T5保留了原始的NMT能力,同时在FLORES和FLEURS的四种语言子集上优于AST基线。由此产生的多模态翻译模型同时产生强大的文本和语音翻译结果。摘要:A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and proposes a joint multimodal training regime of Speech-LLM to include automatic speech translation (AST). We investigate two different foundation model architectures, decoder-only GPT and encoder-decoder T5, extended with Canary-1B's speech encoder. To handle joint multimodal training, we propose a novel training framework called EMMeTT. EMMeTT improves training efficiency with the following: balanced sampling across languages, datasets, and modalities; efficient sequential data iteration; and a novel 2D bucketing scheme for multimodal data, complemented by a batch size optimizer (OOMptimizer). We show that a multimodal training consistently helps with both architectures. Moreover, SALM-T5 trained with EMMeTT retains the original NMT capability while outperforming AST baselines on four-language subsets of FLORES and FLEURS. The resultant Multimodal Translation Model produces strong text and speech translation results at the same time.
【10】 LM-assisted keyword biasing with Aho-Corasick algorithm for Transducer-based ASR
标题: 使用Aho-Corasick算法的LM辅助关键字偏置用于基于传感器的ASB
作者:Iuliia Thorbecke,Juan Zuluaga-Gomez,Esaú Villatoro-Tello,Andres Carofilis,Shashi Kumar,Petr Motlicek,Karthik Pandia,Aravind Ganapathiraju
备注:Submitted to ICASSP2025
链接:点击下载PDF文件
摘要:尽管自动语音识别的端到端模型最近取得了成功,但识别特殊的罕见词汇和词汇表外的单词,以及快速的文本域适应,仍然具有挑战性。经常发生的是,对特殊实体的偏见导致整体性能的下降。我们提出了一个轻的飞行方法,以提高自动语音识别性能相结合的偏见列表命名的实体与单词级的n-gram语言模型与浅融合方法的基础上的Aho-Corasick字符串匹配算法。Aho-Corasick算法已被证明比其他方法更有效,并且允许快速上下文自适应。一个n-gram语言模型被引入作为一个图的失败和输出弧,弧的权重是从n-gram概率。当语言模型与单个上下文图中的偏置实体组合以照顾整体性能时,语言模型被用作对关键字偏置的附加支持。我们在4种语言、2个公共数据集和1个私有数据集上展示了我们的研究结果,包括命名实体和词汇表外实体的性能。我们实现了高达21.6%的相对改善,在一般的字错误率没有实际的差异,在逆实时因素。摘要:Despite the recent success of end-to-end models for automatic speech recognition, recognizing special rare and out-of-vocabulary words, as well as fast domain adaptation with text, are still challenging. It often happens that biasing to the special entities leads to a degradation in the overall performance. We propose a light on-the-fly method to improve automatic speech recognition performance by combining a bias list of named entities with a word-level n-gram language model with the shallow fusion approach based on the Aho-Corasick string matching algorithm. The Aho-Corasick algorithm has proved to be more efficient than other methods and allows fast context adaptation. An n-gram language model is introduced as a graph with fail and output arcs, where the arc weights are adapted from the n-gram probabilities. The language model is used as an additional support to keyword biasing when the language model is combined with bias entities in a single context graph to take care of the overall performance. We demonstrate our findings on 4 languages, 2 public and 1 private datasets including performance on named entities and out-of-vocabulary entities. We achieve up to 21.6% relative improvement in the general word error rate with no practical difference in the inverse real-time factor.
【11】 Sketching With Your Voice: "Non-Phonorealistic" Rendering of Sounds via Vocal Imitation
作者:Matthew Caren,Kartik Chandra,Joshua B. Tenenbaum,Jonathan Ragan-Kelley,Karima Ma
Journal-ref:SIGGRAPH Asia 2024
链接:点击下载PDF文件
摘要:我们提出了一种自动产生类似人类的声音模仿的方法:相当于“素描”,但听觉而不是视觉表示。从人类声道的模拟模型开始,我们首先尝试通过调整模型的控制参数来生成声音模仿,以使合成的发声在感知上突出的听觉特征方面与目标声音相匹配。然后,为了更好地匹配人类的直觉,我们应用了认知理论的沟通,考虑到人类扬声器的原因,他们的听众的战略。最后,我们通过几个实验和用户研究表明,当我们将这种类型的交流推理添加到我们的方法中时,它比单独匹配听觉特征更符合人类的直觉。这一观察结果对计算机图形学中的描绘研究具有广泛的意义。摘要:We present a method for automatically producing human-like vocal imitations of sounds: the equivalent of "sketching," but for auditory rather than visual representation. Starting with a simulated model of the human vocal tract, we first try generating vocal imitations by tuning the model's control parameters to make the synthesized vocalization match the target sound in terms of perceptually-salient auditory features. Then, to better match human intuitions, we apply a cognitive theory of communication to take into account how human speakers reason strategically about their listeners. Finally, we show through several experiments and user studies that when we add this type of communicative reasoning to our method, it aligns with human intuitions better than matching auditory features alone does. This observation has broad implications for the study of depiction in computer graphics.
【12】 Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
标题: 通过Whisper知识提炼快速流媒体传感器ASO原型
作者:Iuliia Thorbecke,Juan Zuluaga-Gomez,Esaú Villatoro-Tello,Shashi Kumar,Pradeep Rangappa,Sergio Burdisso,Petr Motlicek,Karthik Pandia,Aravind Ganapathiraju
备注:Accepted to EMNLP Findings 2024
链接:点击下载PDF文件
摘要:在几乎没有监督数据的情况下训练自动语音识别(ASR)仍然是一个悬而未决的问题。在这项工作中,我们证明了流转换器-转换器(TT)模型可以在消费者和可访问的GPU中从头开始训练,并完全使用来自基础语音模型(FSM)的伪标记(PL)语音。这允许仅在一个阶段中训练鲁棒的ASR模型,并且与具有预训练和微调的两步方案相比,不需要大量数据和计算预算。我们对基于PL的流TT模型的不同方面进行了全面的消融,例如(1)n-gram LM的浅融合,(2)命名实体的上下文偏置,(3)低延迟流应用的分块解码,以及(4)TT整体性能作为FSM大小的函数。我们的研究结果表明,TT可以从头开始训练,没有监督的数据,即使是非常嘈杂的PL。我们从CommonVoice的6种语言上验证了所提出的框架,并提出了多种算法来过滤出幻觉PL。摘要:The training of automatic speech recognition (ASR) with little to no supervised data remains an open question. In this work, we demonstrate that streaming Transformer-Transducer (TT) models can be trained from scratch in consumer and accessible GPUs in their entirety with pseudo-labeled (PL) speech from foundational speech models (FSM). This allows training a robust ASR model just in one stage and does not require large data and computational budget compared to the two-step scenario with pre-training and fine-tuning. We perform a comprehensive ablation on different aspects of PL-based streaming TT models such as the impact of (1) shallow fusion of n-gram LMs, (2) contextual biasing with named entities, (3) chunk-wise decoding for low-latency streaming applications, and (4) TT overall performance as the function of the FSM size. Our results demonstrate that TT can be trained from scratch without supervised data, even with very noisy PLs. We validate the proposed framework on 6 languages from CommonVoice and propose multiple heuristics to filter out hallucinated PLs.
【13】 DiffSound: Differentiable Modal Sound Rendering and Inverse Rendering for Diverse Inference Tasks
标题: 迪夫声音:可区分的模式声音渲染和反向渲染,用于多样化推理任务
作者:Xutong Jin,Chenxi Xu,Ruohan Gao,Jiajun Wu,Guoping Wang,Sheng Li
备注:12 pages, 10 figures. Published in Siggraph 2024. Project page: this https URL
链接:点击下载PDF文件
摘要:从真实世界的声音记录中准确地估计和模拟物体的物理特性在视觉、图形学和机器人领域具有重要的实际意义。然而,在这些方向上的进展受到限制-由于音频的高采样率,先前的可微分刚体或软体模拟技术不能直接应用于模态声音合成,而先前的音频合成器通常不能完全模拟发声对象的准确物理特性。我们提出了DiffSound,这是一个用于基于物理的模态声音合成的可微分声音渲染框架,它基于隐式形状表示、新的高阶有限元分析模块和可微分音频合成器。由于整个管道的可微性,我们的框架可以解决广泛的逆问题,包括物理参数估计,几何形状推理和碰撞位置预测。实验结果表明,我们的方法的有效性,突出了它的能力,准确地再现目标声音在一个基于物理的方式。DiffSound是各种声音合成和分析应用程序的宝贵工具。摘要:Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been limited -- prior differentiable rigid or soft body simulation techniques cannot be directly applied to modal sound synthesis due to the high sampling rate of audio, while previous audio synthesizers often do not fully model the accurate physical properties of the sounding objects. We propose DiffSound, a differentiable sound rendering framework for physics-based modal sound synthesis, which is based on an implicit shape representation, a new high-order finite element analysis module, and a differentiable audio synthesizer. Our framework can solve a wide range of inverse problems thanks to the differentiability of the entire pipeline, including physical parameter estimation, geometric shape reasoning, and impact position prediction. Experimental results demonstrate the effectiveness of our approach, highlighting its ability to accurately reproduce the target sound in a physics-based manner. DiffSound serves as a valuable tool for various sound synthesis and analysis applications.
【14】 Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
标题: 音频编解码器增强用于语音合成的鲁棒协作水印
作者:Lauri Juvela,Xin Wang
备注:Submitted to ICASSP 2025
链接:点击下载PDF文件
摘要:合成语音的自动检测变得越来越重要,因为当前的合成方法与人类语音几乎无法区分,并且广泛地为公众所用。音频水印和其他主动披露方法正在吸引研究活动,因为它们可以补充基于被动检测的传统深度伪造防御。在主动和被动检测中,鲁棒性是主要的兴趣。传统的音频水印特别容易受到音频编解码器应用程序的移除攻击。大多数生成的语音和音频内容发布到野外纯粹作为分发方法通过音频编解码器。我们最近提出了协同水印的方法,使生成的语音更容易检测到的噪声,但可区分的传输通道。本文将通道增强扩展到不可区分的传统音频编解码器和神经音频编解码器,并评估各种配置下编解码器比特率的可转移性和效果。结果表明,黑盒音频编解码器可以使用波形域直通估计器进行梯度逼近来可靠地增强协作水印。此外,该结果表明,使用神经音频编解码器的通道增强可以很好地转移到传统编解码器。听力测试表明,协同水印在8kbps的高比特率编解码器或DAC中引起的感知退化可以忽略不计。摘要:Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to the public. Audio watermarking and other active disclosure methods of are attracting research activity, as they can complement traditional deepfake defenses based on passive detection. In both active and passive detection, robustness is of major interest. Traditional audio watermarks are particularly susceptible to removal attacks by audio codec application. Most generated speech and audio content released into the wild passes through an audio codec purely as a distribution method. We recently proposed collaborative watermarking as method for making generated speech more easily detectable over a noisy but differentiable transmission channel. This paper extends the channel augmentation to work with non-differentiable traditional audio codecs and neural audio codecs and evaluates transferability and effect of codec bitrate over various configurations. The results show that collaborative watermarking can be reliably augmented by black-box audio codecs using a waveform-domain straight-through-estimator for gradient approximation. Furthermore, that results show that channel augmentation with a neural audio codec transfers well to traditional codecs. Listening tests demonstrate collaborative watermarking incurs negligible perceptual degradation with high bitrate codecs or DAC at 8kbps.
【15】 Beyond the binary: Limitations and possibilities of gender-related speech technology research
标题: 超越二元:性别相关语音技术研究的局限性和可能性
作者:Ariadna Sanchez,Alice Ross,Nina Markl
备注:Accepted at Spoken Language Technology (SLT) Workshop 2024
链接:点击下载PDF文件
摘要:本文回顾了2013年至2023年ISCA Interspeech出版物中107篇与言语和性别或性别相关的研究论文。我们注意到关于这一主题的工作的稀缺性,并发现术语,特别是 textit{gender}一词的使用方式不够明确,而且往往与社会科学中的流行观点不一致,即性别是社会构建的,是一个光谱,而不是二元类别。我们提请注意潜在的问题,这可能会导致已经被边缘化的群体,并提出一些问题,研究人员问自己时,进行工作的言论和性别。摘要:This paper presents a review of 107 research papers relating to speech and sex or gender in ISCA Interspeech publications between 2013 and 2023. We note the scarcity of work on this topic and find that terminology, particularly the word textit{gender}, is used in ways that are underspecified and often out of step with the prevailing view in social sciences that gender is socially constructed and is a spectrum as opposed to a binary category. We draw attention to the potential problems that this can cause for already marginalised groups, and suggest some questions for researchers to ask themselves when undertaking work on speech and gender.
【16】 Large Language Model Should Understand Pinyin for Chinese ASR Error Correction
标题: 大型语言模型应了解拼音才能进行中文ASB错误纠正
作者:Yuang Li,Xiaosong Qiao,Xiaofeng Zhao,Huan Zhao,Wei Tang,Min Zhang,Hao Yang
链接:点击下载PDF文件
摘要:大型语言模型可以通过生成纠错来增强自动语音识别系统。在本文中,我们提出了拼音增强的GEC,它利用拼音,汉语普通话的语音表示,作为补充信息,以提高中国ASR纠错。我们的方法只利用合成误差进行训练,并在推理过程中采用最佳假设。此外,我们引入了一种多任务训练方法,涉及拼音和文本之间的转换任务,以对齐它们的特征空间。在Aishell-1和Common Voice数据集上的实验表明,我们的方法在纯文本输入的情况下始终优于GEC。更重要的是,我们从两个方面为PY-GEC和多任务训练的有效性提供了直观的解释:1)增加了对拼音特征的注意力权重; 2)拼音和文本隐藏状态之间的对齐特征空间。摘要:Large language models can enhance automatic speech recognition systems through generative error correction. In this paper, we propose Pinyin-enhanced GEC, which leverages Pinyi, the phonetic representation of Mandarin Chinese, as supplementary information to improve Chinese ASR error correction. Our approach only utilizes synthetic errors for training and employs the one-best hypothesis during inference. Additionally, we introduce a multitask training approach involving conversion tasks between Pinyin and text to align their feature spaces. Experiments on the Aishell-1 and the Common Voice datasets demonstrate that our approach consistently outperforms GEC with text-only input. More importantly, we provide intuitive explanations for the effectiveness of PY-GEC and multitask training from two aspects: 1) increased attention weight on Pinyin features; and 2) aligned feature space between Pinyin and text hidden states.
【17】 MuCodec: Ultra Low-Bitrate Music Codec
标题: MuCodec:超低比特率音乐编解码器
作者:Yaoxun Xu,Hangting Chen,Jianwei Yu,Wei Tan,Rongzhi Gu,Shun Lei,Zhiwei Lin,Zhiyong Wu
链接:点击下载PDF文件
摘要:音乐编解码器是音频编解码器研究的一个重要方面,超低比特率压缩对于音乐传输和生成具有重要意义。由于音乐背景的复杂性和人声的丰富性,单纯依靠语义或声学信息建模不能有效地重建具有人声和背景的音乐。为了解决这个问题,我们提出了MuCodec,专门针对超低比特率的音乐压缩和重建任务。MuCodec采用MuEncoder提取语音和语义特征,用RVQ离散化,通过流匹配得到Mel-VAE特征。然后使用预先训练的MEL-VAE解码器和HiFi-GAN重建音乐。MuCodec能够以超低(0.35kbps)或高比特率(1.35kbps)重建高保真音乐,在主观和客观指标上都达到了迄今为止的最佳效果。代码和演示:https: xuyaoxun.github.io MuCodec_demo 。摘要:Music codecs are a vital aspect of audio codec research, and ultra low-bitrate compression holds significant importance for music transmission and generation. Due to the complexity of music backgrounds and the richness of vocals, solely relying on modeling semantic or acoustic information cannot effectively reconstruct music with both vocals and backgrounds. To address this issue, we propose MuCodec, specifically targeting music compression and reconstruction tasks at ultra low bitrates. MuCodec employs MuEncoder to extract both acoustic and semantic features, discretizes them with RVQ, and obtains Mel-VAE features via flow-matching. The music is then reconstructed using a pre-trained MEL-VAE decoder and HiFi-GAN. MuCodec can reconstruct high-fidelity music at ultra low (0.35kbps) or high bitrates (1.35kbps), achieving the best results to date in both subjective and objective metrics. Code and Demo: https: xuyaoxun.github.io MuCodec_demo .
【18】 Personalized Speech Recognition for Children with Test-Time Adaptation
标题: 具有测试时间适应的儿童个性化语音识别
作者:Zhonghao Shi,Harshvardhan Srivastava,Xuan Shi,Shrikanth Narayanan,Maja J. Matarić
备注:This work has been submitted to the IEEE for possible publication
链接:点击下载PDF文件
摘要:准确的儿童自动语音识别(ASR)对于有效的实时儿童-AI交互至关重要,特别是在教育应用中。然而,由于数据域从成人转移到儿童,主要在成人数据上进行预训练的现成ASR模型往往对儿童语音的泛化能力很差。最近的研究发现,对儿童语音数据进行监督微调可以帮助弥合这种域转移,但对于现实世界的应用程序来说,人类注释可能是不切实际的,并且训练时的适应可能会忽略测试时发生的额外域转移。我们设计了一种新型的ASR管道,将无监督测试时自适应(TTA)方法应用于儿童语音识别,以便根据成人语音预训练的ASR模型可以在测试时连续适应每个儿童说话者,而无需进一步的人工注释。我们的研究结果表明,与TTA方法相适应的ASR模型显着优于未经调整的现成的ASR基线的平均和统计在个别儿童扬声器。我们的分析还发现,显着的数据域之间的儿童扬声器和每个儿童扬声器内的变化,这进一步激发了测试时间适应的需要。摘要:Accurate automatic speech recognition (ASR) for children is crucial for effective real-time child-AI interaction, especially in educational applications. However, off-the-shelf ASR models primarily pre-trained on adult data tend to generalize poorly to children's speech due to the data domain shift from adults to children. Recent studies have found that supervised fine-tuning on children's speech data can help bridge this domain shift, but human annotations may be impractical to obtain for real-world applications and adaptation at training time can overlook additional domain shifts occurring at test time. We devised a novel ASR pipeline to apply unsupervised test-time adaptation (TTA) methods for child speech recognition, so that ASR models pre-trained on adult speech can be continuously adapted to each child speaker at test time without further human annotations. Our results show that ASR models adapted with TTA methods significantly outperform the unadapted off-the-shelf ASR baselines both on average and statistically across individual child speakers. Our analysis also discovered significant data domain shifts both between child speakers and within each child speaker, which further motivates the need for test-time adaptation.
【19】 DiffEditor: Enhancing Speech Editing with Semantic Enrichment and Acoustic Consistency
标题: 迪夫编辑器:通过语义丰富和声学一致性增强语音编辑
作者:Yang Chen,Yuhang Jia,Shiwan Zhao,Ziyue Jiang,Haoran Li,Jiarong Kang,Yong Qin
链接:点击下载PDF文件
摘要:随着基于文本的语音编辑变得越来越普遍,对无限制的自由文本编辑的需求持续增长。然而,现有的语音编辑技术遇到了重大的挑战,特别是在处理域外(OOD)文本时保持可懂度和声学一致性。在本文中,我们介绍,DiffEditor,一种新的语音编辑模型,旨在提高性能,在OOD文本的情况下,通过语义丰富和声学的一致性。为了提高编辑语音的可理解性,我们通过整合从预训练的语言模型中提取的单词嵌入来丰富音素嵌入的语义信息。此外,我们强调,帧间平滑属性是建模声学一致性的关键,因此,我们提出了一个一阶损失函数,以促进编辑边界处的平滑过渡,并提高编辑语音的整体流畅性。实验结果表明,我们的模型实现了国家的最先进的性能在域和面向对象的文本场景。摘要:As text-based speech editing becomes increasingly prevalent, the demand for unrestricted free-text editing continues to grow. However, existing speech editing techniques encounter significant challenges, particularly in maintaining intelligibility and acoustic consistency when dealing with out-of-domain (OOD) text. In this paper, we introduce, DiffEditor, a novel speech editing model designed to enhance performance in OOD text scenarios through semantic enrichment and acoustic consistency. To improve the intelligibility of the edited speech, we enrich the semantic information of phoneme embeddings by integrating word embeddings extracted from a pretrained language model. Furthermore, we emphasize that interframe smoothing properties are critical for modeling acoustic consistency, and thus we propose a first-order loss function to promote smoother transitions at editing boundaries and enhance the overall fluency of the edited speech. Experimental results demonstrate that our model achieves state-of-the-art performance in both in-domain and OOD text scenarios.
机器翻译,仅供参考
