今日论文合集:cs.SD语音13篇,eess.AS音频处理13篇。

本文经arXiv每日学术速递授权转载


cs.SD语音
【1】 RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
标题: RT-LA-VocE:实时低SNR视听语音增强
作者:Honglie Chen,Rodrigo Mira,Stavros Petridis,Maja Pantic
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们的目标是从实时视频流和嘈杂的音频流中逐帧生成干净的语音,而不依赖于未来的输入。为此,我们提出了RT-LA-VocE,它完全重新设计了LA-VocE,一个国家的最先进的非因果视听语音增强模型的每个组件,执行因果实时推理与40毫秒的输入帧。我们通过设计新的视觉和音频编码器来实现这一点,这些编码器仅依赖于过去的帧,用Emformer取代Transformer编码器,并设计新的因果神经声码器C-HiFi-GAN。在流行的AVSpeech数据集上,我们证明了我们的算法在所有实时场景中都达到了最先进的效果。更重要的是,每个组件都经过精心调整,以将算法延迟降至理论最低值(40 ms),同时保持每帧28.15 ms的低端到端处理延迟,从而以最小的延迟实现实时逐帧增强。摘要:In this paper, we aim to generate clean speech frame by frame from a live video stream and a noisy audio stream without relying on future inputs. To this end, we propose RT-LA-VocE, which completely re-designs every component of LA-VocE, a state-of-the-art non-causal audio-visual speech enhancement model, to perform causal real-time inference with a 40ms input frame. We do so by devising new visual and audio encoders that rely solely on past frames, replacing the Transformer encoder with the Emformer, and designing a new causal neural vocoder C-HiFi-GAN. On the popular AVSpeech dataset, we show that our algorithm achieves state-of-the-art results in all real-time scenarios. More importantly, each component is carefully tuned to minimize the algorithm latency to the theoretical minimum (40ms) while maintaining a low end-to-end processing latency of 28.15ms per frame, enabling real-time frame-by-frame enhancement with minimal delay.

【2】 SaMoye: Zero-shot Singing Voice Conversion Based on Feature Disentanglement and Synthesis
标题: SaMoye:基于特征解纠缠和合成的Zero-Shot歌唱声音转换
作者:Zihao Wang,Le Ma,Yan Liu,Kejun Zhang
备注:7 pages, 4 figures
链接:点击下载PDF文件
摘要:歌唱声音转换(SVC)旨在将给定音乐作品中的歌手的声音转换为另一个歌手,同时保持原始内容。我们提出了一个端到端的功能分解为基础的模型,我们命名为SaMoye,使zero-shot多对多的歌声转换。SaMoye将歌唱声音的特征分解为内容特征、音色特征和音高特征。使用基于GPT的模型来执行与歌词的音素的交叉预测,以增强内容特征。SaMoye可以通过将音色特征替换为目标歌手来生成具有转换语音的音乐。我们还建立了一个无与伦比的大规模数据集,以保证zero-shot性能。该数据集由1500k个纯歌唱声乐片段组成,其中包含至少10,000名歌手。摘要:Singing voice conversion (SVC) aims to convert a singer's voice in a given music piece to another singer while keeping the original content. We propose an end-to-end feature disentanglement-based model, which we named SaMoye, to enable zero-shot many-to-many singing voice conversion. SaMoye disentangles the features of the singing voice into content features, timbre features, and pitch features respectively. The content features are enhanced using a GPT-based model to perform cross-prediction with the phoneme of the lyrics. SaMoye can generate the music with converted voice by replacing the timbre features with the target singer. We also establish an unparalleled large-scale dataset to guarantee zero-shot performance. The dataset consists of 1500k pure singing vocal clips containing at least 10,000 singers.

【3】 Targeted Augmented Data for Audio Deepfake Detection
标题: 用于音频Deepfake检测的定向增强数据
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in EUSIPCO 2024
链接:点击下载PDF文件
摘要:高度可信的音频深度伪造生成器的可用性突出了设计强大的音频深度伪造检测器的必要性。现有的工作往往仅仅依赖于训练集中的真实和虚假数据,这可能导致过拟合,从而降低了对不可见操作的鲁棒性。为了增强音频deepfake检测器的泛化能力,我们提出了一种新的增强方法,用于生成针对模型决策边界的音频伪假。受对抗性攻击的启发,我们扰动原始真实数据来合成具有模糊预测概率的伪伪造。两个著名的架构上的综合实验表明,所提出的增强有助于提高这些架构的泛化能力。摘要:The availability of highly convincing audio deepfake generators highlights the need for designing robust audio deepfake detectors. Existing works often rely solely on real and fake data available in the training set, which may lead to overfitting, thereby reducing the robustness to unseen manipulations. To enhance the generalization capabilities of audio deepfake detectors, we propose a novel augmentation method for generating audio pseudo-fakes targeting the decision boundary of the model. Inspired by adversarial attacks, we perturb original real data to synthesize pseudo-fakes with ambiguous prediction probabilities. Comprehensive experiments on two well-known architectures demonstrate that the proposed augmentation contributes to improving the generalization capabilities of these architectures.

【4】 HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
标题: HebDB:希伯来语语音处理的弱监督数据集
作者:Arnon Turetzky,Or Tal,Yael Segal-Feldman,Yehoshua Dissen,Ella Zeldes,Amit Roth,Eyal Cohen,Yosi Shrem,Bronya R. Chernyak,Olga Seleznova,Joseph Keshet,Yossi Adi
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:我们提出了HebDB,一个弱监督的数据集,用于希伯来语的口语处理。HebDB提供大约2500小时的希伯来语自然和自发的演讲录音,包括各种各样的演讲者和主题。我们提供原始录音以及预处理,弱监督和过滤版本。HebDB的目标是进一步加强希伯来语口语处理工具的研究和开发。因此,我们还提供了两个用于自动语音识别(ASR)的基线系统:(i)自监督模型;以及(ii)完全监督模型。我们提出了这两种方法在HebDB上优化的性能,并将其与当前的多语言ASR替代方案进行了比较。结果表明,所提出的方法达到更好的结果比评估基线考虑类似的模型大小。数据集、代码和模型可在https: pages.cs.huji.ac.il adiyoss-lab HebDB 上公开获取。摘要:We present HebDB, a weakly supervised dataset for spoken language processing in the Hebrew language. HebDB offers roughly 2500 hours of natural and spontaneous speech recordings in the Hebrew language, consisting of a large variety of speakers and topics. We provide raw recordings together with a pre-processed, weakly supervised, and filtered version. The goal of HebDB is to further enhance research and development of spoken language processing tools for the Hebrew language. Hence, we additionally provide two baseline systems for Automatic Speech Recognition (ASR): (i) a self-supervised model; and (ii) a fully supervised model. We present the performance of these two methods optimized on HebDB and compare them to current multi-lingual ASR alternatives. Results suggest the proposed method reaches better results than the evaluated baselines considering similar model sizes. Dataset, code, and models are publicly available under https: pages.cs.huji.ac.il adiyoss-lab HebDB .

【5】 Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
标题: Beat-It:节拍同步多条件3D舞蹈生成
作者:Zikai Huang,Xuemiao Xu,Cheng Xu,Huaidong Zhang,Chenxi Zheng,Jing Qin,Shengfeng He
备注:ECCV 2024
链接:点击下载PDF文件
摘要:舞蹈作为一种艺术形式,从根本上讲取决于与音乐节拍的精确同步。然而,从音乐中实现美观的舞蹈序列是具有挑战性的,现有的方法往往在可控性和节拍对齐方面不足。为了解决这些缺点,本文介绍了Beat-It,一种新颖的框架,用于特定节拍,关键姿势引导的舞蹈生成。与之前的方法不同,Beat-It独特地集成了明确的节拍意识和关键姿势指导,有效地解决了两个主要问题:生成的舞蹈动作与音乐节拍的不一致,以及无法将关键姿势映射到特定节拍,这对于实际的编舞至关重要。我们的方法使用最近的节拍距离表示从音乐中解开节拍条件,并采用分层多条件融合机制。这种机制无缝地整合了关键姿势,节拍和音乐特征,减轻了条件冲突,并为舞蹈生成提供了丰富的多条件指导。此外,一个专门设计的节拍对齐损失确保生成的舞蹈动作保持与指定的节拍同步。大量的实验证实了Beat-It在节拍对齐和运动可控性方面优于现有的最先进的方法。摘要:Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging, with existing methods often falling short in controllability and beat alignment. To address these shortcomings, this paper introduces Beat-It, a novel framework for beat-specific, key pose-guided dance generation. Unlike prior approaches, Beat-It uniquely integrates explicit beat awareness and key pose guidance, effectively resolving two main issues: the misalignment of generated dance motions with musical beats, and the inability to map key poses to specific beats, critical for practical choreography. Our approach disentangles beat conditions from music using a nearest beat distance representation and employs a hierarchical multi-condition fusion mechanism. This mechanism seamlessly integrates key poses, beats, and music features, mitigating condition conflicts and offering rich, multi-conditioned guidance for dance generation. Additionally, a specially designed beat alignment loss ensures the generated dance movements remain in sync with the designated beats. Extensive experiments confirm Beat-It's superiority over existing state-of-the-art methods in terms of beat alignment and motion controllability.

【6】 Video-to-Audio Generation with Hidden Alignment
标题: 具有隐藏对齐的视频到音频生成
作者:Manjie Xu,Chenxing Li,Yong Ren,Rilin Chen,Yu Gu,Wei Liang,Dong Yu
备注:this https URL
链接:点击下载PDF文件
摘要:根据视频输入生成语义上和时间上对齐的音频内容已经成为研究人员的焦点,特别是在文本到视频生成方面取得显著突破之后。在这项工作中,我们的目标是提供对视频到音频生成范式的见解,重点关注三个关键方面:视觉编码器,辅助嵌入和数据增强技术。从建立在简单但令人惊讶的有效直觉基础上的基础模型VTA-LDM开始,我们通过消融研究探索各种视觉编码器和辅助嵌入。采用全面的评估管道,强调生成质量和视频-音频同步对齐,我们证明了我们的模型具有最先进的视频到音频生成能力。此外,我们提供了重要的见解不同的数据增强方法对提高生成框架的整体能力的影响。我们展示的可能性,以推进从语义和时间的角度生成同步音频的挑战。我们希望这些见解将成为开发更现实和准确的视听生成模型的垫脚石。摘要:Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to offer insights into the video-to-audio generation paradigm, focusing on three crucial aspects: vision encoders, auxiliary embeddings, and data augmentation techniques. Beginning with a foundational model VTA-LDM built on a simple yet surprisingly effective intuition, we explore various vision encoders and auxiliary embeddings through ablation studies. Employing a comprehensive evaluation pipeline that emphasizes generation quality and video-audio synchronization alignment, we demonstrate that our model exhibits state-of-the-art video-to-audio generation capabilities. Furthermore, we provide critical insights into the impact of different data augmentation methods on enhancing the generation framework's overall capacity. We showcase possibilities to advance the challenge of generating synchronized audio from semantic and temporal perspectives. We hope these insights will serve as a stepping stone toward developing more realistic and accurate audio-visual generation models.

【7】 Out-of-distribution generalisation in spoken language understanding
标题: 口语理解中的非分布概括
作者:Dejan Porjazovski,Anssi Moisio,Mikko Kurimo
备注:Accepted for INTERSPEECH 2024
链接:点击下载PDF文件
摘要:当测试数据意外地与训练数据不同时,它被称为分布外(OOD),这是机器学习的真实用例中的一个常见挑战。虽然OOD概括近年来获得了兴趣,很少有作品集中在OOD概括口语理解(SLU)任务。为了促进对这一主题的研究,我们介绍了一个修改版本的流行SLU数据集SLURP,具有数据分裂测试OOD概括在SLU任务。我们将修改后的数据集称为SLURP For OOD泛化,或SLURPFOOD。利用我们的OOD数据分割,我们发现端到端SLU模型的泛化能力有限。此外,通过采用模型的可解释性技术,我们揭示了导致模型泛化困难的因素。为了提高泛化能力,我们尝试了两种技术,这两种技术提高了一些但不是所有分裂的结果,强调了对新技术的需求。摘要:Test data is said to be out-of-distribution (OOD) when it unexpectedly differs from the training data, a common challenge in real-world use cases of machine learning. Although OOD generalisation has gained interest in recent years, few works have focused on OOD generalisation in spoken language understanding (SLU) tasks. To facilitate research on this topic, we introduce a modified version of the popular SLU dataset SLURP, featuring data splits for testing OOD generalisation in the SLU task. We call our modified dataset SLURP For OOD generalisation, or SLURPFOOD. Utilising our OOD data splits, we find end-to-end SLU models to have limited capacity for generalisation. Furthermore, by employing model interpretability techniques, we shed light on the factors contributing to the generalisation difficulties of the models. To improve the generalisation, we experiment with two techniques, which improve the results on some, but not all the splits, emphasising the need for new techniques.

【8】 STONE: Self-supervised Tonality Estimator
标题: STONE:自我监督的调性估计器
作者:Yuexuan Kong,Vincent Lostanlen,Gabriel Meseguer-Brocal,Stella Wong,Mathieu Lagrange,Romain Hennequin
链接:点击下载PDF文件
摘要:虽然深度神经网络可以估计音乐作品的基调,但它们的监督会导致大量的注释工作。针对这一缺点,我们提出了石头,第一个自我监督的音调估计。STONE背后的架构名为ChromaNet,是一个具有八度等效性的convnet,它输出12个结构化logit的密钥签名配置文件(KSP)。首先,我们训练ChromaNet来回归来自同一音轨的任何两个未标记的音乐摘录之间的人工音高置换,如在五分之一圈(CoF)内测量的交叉功率谱密度(CPSD)。我们观察到,这种自我监督的借口任务导致KSP与音调的关键签名。基于这一观察,我们扩展石头输出一个结构化的KSP的24个logits,并引入监督,以消除歧义的主要与次要的密钥共享相同的密钥签名。应用不同的监督量产生半监督和全监督音调估计器:即,半色调和超色调。我们评估这些估计FMAK,一个新的数据集的5489个真实世界的音乐录音与专家注释的24个主要和次要的关键。我们发现,Semi-TONE匹配Sup-TONE的分类精度与减少监督,并优于它与平等的监督。摘要:Although deep neural networks can estimate the key of a musical piece, their supervision incurs a massive annotation effort. Against this shortcoming, we present STONE, the first self-supervised tonality estimator. The architecture behind STONE, named ChromaNet, is a convnet with octave equivalence which outputs a key signature profile (KSP) of 12 structured logits. First, we train ChromaNet to regress artificial pitch transpositions between any two unlabeled musical excerpts from the same audio track, as measured as cross-power spectral density (CPSD) within the circle of fifths (CoF). We observe that this self-supervised pretext task leads KSP to correlate with tonal key signature. Based on this observation, we extend STONE to output a structured KSP of 24 logits, and introduce supervision so as to disambiguate major versus minor keys sharing the same key signature. Applying different amounts of supervision yields semi-supervised and fully supervised tonality estimators: i.e., Semi-TONEs and Sup-TONEs. We evaluate these estimators on FMAK, a new dataset of 5489 real-world musical recordings with expert annotation of 24 major and minor keys. We find that Semi-TONE matches the classification accuracy of Sup-TONE with reduced supervision and outperforms it with equal supervision.

【9】 SimuSOE: A Simulated Snoring Dataset for Obstructive Sleep Apnea-Hypopnea Syndrome Evaluation during Wakefulness
标题: SimuSOE:清醒期间评估阻塞性睡眠呼吸暂停-低通气综合征的模拟打鼾数据集
作者:Jie Lin,Xiuping Yang,Li Xiao,Xinhong Li,Weiyan Yi,Yuhong Yang,Weiping Tu,Xiong Chen
链接:点击下载PDF文件
摘要:阻塞性睡眠呼吸暂停低通气综合征(OSAHS)是一种常见的上呼吸道阻塞引起的慢性呼吸障碍。以前的研究通过基于睡眠打鼾或语音信号数据集训练的基于机器学习的系统来推进OSAHS评估。然而,构建用于训练精确且快速的OSAHS评估系统的数据集提出了挑战,因为1)收集睡眠打鼾是耗时的,以及2)语音信号在反映上气道阻塞方面受到限制。在本文中,我们提出了一个新的打鼾数据集的OSAHS评估,命名为SimuSOE,其中引入了一种新颖的和时间有效的打鼾收集方法来解决上述问题。特别是,我们采用模拟打鼾,这是一种由患者故意发出的打鼾,以取代自然打鼾。实验结果表明,模拟清醒状态下的鼾声信号可以作为OSAHS初筛的有效特征。摘要:Obstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) is a prevalent chronic breathing disorder caused by upper airway obstruction. Previous studies advanced OSAHS evaluation through machine learning-based systems trained on sleep snoring or speech signal datasets. However, constructing datasets for training a precise and rapid OSAHS evaluation system poses a challenge, since 1) it is time-consuming to collect sleep snores and 2) the speech signal is limited in reflecting upper airway obstruction. In this paper, we propose a new snoring dataset for OSAHS evaluation, named SimuSOE, in which a novel and time-effective snoring collection method is introduced for tackling the above problems. In particular, we adopt simulated snoring which is a type of snore intentionally emitted by patients to replace natural snoring. Experimental results indicate that the simulated snoring signal during wakefulness can serve as an effective feature in OSAHS preliminary screening.

【10】 Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology
标题: 性别后的言语:言语科学与技术下一步的跨女性视角
作者:Robin Netzorg,Alyssa Cote,Sumi Koshin,Klo Vivienne Garoute,Gopala Krishna Anumanchipalli
链接:点击下载PDF文件
摘要:作为语音修改的专家,跨女性性别肯定语音教师对语音有独特的观点,混淆了目前对说话者身份的理解。为了证明这一点,我们提出了多功能语音数据集(VVD),一个集合的三个扬声器修改他们的声音沿性别轴。的VVD说明,目前的方法在扬声器建模,基于性别的分类概念和静态的声乐纹理的理解,未能考虑声道的灵活性。利用公开的扬声器嵌入,我们表明,性别分类系统是高度敏感的语音修改,和扬声器验证系统无法识别语音来自同一扬声器的语音修改变得更加激烈。作为走向超越说话者身份的分类和静态概念的一条路径,我们建议建模的声音纹理,如音高,共鸣和重量的个人素质。摘要:As experts in voice modification, trans-feminine gender-affirming voice teachers have unique perspectives on voice that confound current understandings of speaker identity. To demonstrate this, we present the Versatile Voice Dataset (VVD), a collection of three speakers modifying their voices along gendered axes. The VVD illustrates that current approaches in speaker modeling, based on categorical notions of gender and a static understanding of vocal texture, fail to account for the flexibility of the vocal tract. Utilizing publicly-available speaker embeddings, we demonstrate that gender classification systems are highly sensitive to voice modification, and speaker verification systems fail to identify voices as coming from the same speaker as voice modification becomes more drastic. As one path towards moving beyond categorical and static notions of speaker identity, we propose modeling individual qualities of vocal texture such as pitch, resonance, and weight.

【11】 AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
标题: AVCAP:利用视听功能作为文本标记进行字幕
作者:Jongsuk Kim,Jiwon Shin,Junmo Kim
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,表征学习和语言模型的进步将自动字幕(AC)推向了新的高度,从而能够生成人类级别的描述。利用这些进步,我们提出了 textbf{AVCap},一个 textbf{A}udio- textbf{V}isual textbf{Cap}tioning框架,一个简单而强大的基线方法适用于视听字幕。AVCap利用视听特征作为文本标记,不仅在性能上,而且在模型的可扩展性和可伸缩性方面具有许多优势。AVCap围绕三个关键维度设计:探索最佳视听编码器架构,根据生成文本的特征调整预训练模型,以及调查字幕中模态融合的有效性。我们的方法在所有指标上都优于现有的视听字幕方法,代码可在https: github.com JongSuk1 AVCap上获得摘要:In recent years, advancements in representation learning and language models have propelled Automated Captioning (AC) to new heights, enabling the generation of human-level descriptions. Leveraging these advancements, we propose textbf{AVCap}, an textbf{A}udio- textbf{V}isual textbf{Cap}tioning framework, a simple yet powerful baseline approach applicable to audio-visual captioning. AVCap utilizes audio-visual features as text tokens, which has many advantages not only in performance but also in the extensibility and scalability of the model. AVCap is designed around three pivotal dimensions: the exploration of optimal audio-visual encoder architectures, the adaptation of pre-trained models according to the characteristics of generated text, and the investigation into the efficacy of modality fusion in captioning. Our method outperforms existing audio-visual captioning methods across all metrics and the code is available on https: github.com JongSuk1 AVCap

【12】 Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data
标题: 神经数据中的比例定律:使用175小时脑电数据的无创语音解码
作者:Motoshige Sato,Kenichi Tomeoka,Ilya Horiguchi,Kai Arulkumaran,Ryota Kanai,Shuntaro Sasai
链接:点击下载PDF文件
摘要:脑机接口(BCI)在帮助有语言障碍的人方面具有巨大的潜力。利用脑电图(EEG)解码语音是特别有前途的,由于其非侵入性的性质。然而,记录通常很短,并且EEG数据的高度可变性导致研究人员专注于具有几十个类别的分类任务。为了评估其在语音神经假体中的实用性,我们研究了开放词汇设置中EEG数据大小与解码准确性之间的关系。我们收集了大量的EEG数据,从一个单一的参与者(175小时),并进行了zero-shot语音段分类使用自监督表示学习。在整个数据集上训练的模型达到了48%的前1准确率和76%的前10准确率,同时减轻了肌电位伪影的影响。相反,当数据被限制在实践中使用的典型数量(10小时)时,前1名的准确率下降到2.5%,显示出显着的缩放效应。此外,随着训练数据量的增加,EEG潜在表征逐渐表现出更清晰的口语短语的时间结构。这表明解码器可以以数据驱动的方式识别语音片段,而无需显式测量单词识别。这项研究标志着基于脑电的语音脑机接口的实际实现迈出了重要的一步。摘要:Brain-computer interfaces (BCIs) hold great potential for aiding individuals with speech impairments. Utilizing electroencephalography (EEG) to decode speech is particularly promising due to its non-invasive nature. However, recordings are typically short, and the high variability in EEG data has led researchers to focus on classification tasks with a few dozen classes. To assess its practical applicability for speech neuroprostheses, we investigate the relationship between the size of EEG data and decoding accuracy in the open vocabulary setting. We collected extensive EEG data from a single participant (175 hours) and conducted zero-shot speech segment classification using self-supervised representation learning. The model trained on the entire dataset achieved a top-1 accuracy of 48 % and a top-10 accuracy of 76 %, while mitigating the effects of myopotential artifacts. Conversely, when the data was limited to the typical amount used in practice ($ sim$10 hours), the top-1 accuracy dropped to 2.5 %, revealing a significant scaling effect. Additionally, as the amount of training data increased, the EEG latent representation progressively exhibited clearer temporal structures of spoken phrases. This indicates that the decoder can recognize speech segments in a data-driven manner without explicit measurements of word recognition. This research marks a significant step towards the practical realization of EEG-based speech BCIs.

【13】 Remastering Divide and Remaster: A Cinematic Audio Source Separation Dataset with Multilingual Support
标题: 重制Divide和Remaster:具有多语言支持的电影音频源分离数据集
作者:Karn N. Watcharasupat,Chih-Wei Wu,Iroro Orife
备注:Submitted to the 5th IEEE International Symposium on the Internet of Sounds
链接:点击下载PDF文件
摘要:电影音频源分离(CASS)是音频源分离的一个相对较新的子任务,涉及到将混合物分离成对话,音乐和效果茎。到目前为止,CASS只有一个公开可用的数据集,即Divide and Remaster(DnR)数据集,目前版本为2。虽然DnR v2一直是CASS的一个非常有用的资源,但已经确定了几个改进领域,特别是通过在2023年声音分层挑战赛中的使用。在这项工作中,我们开发了DnR数据集的第3版,解决了与非对话干中的声音内容,响度分布,掌握过程和语言多样性有关的问题。特别地,DnR v3的对话词干包括来自多个家族的30多种语言的语音内容,包括但不限于日耳曼语、罗曼语、印度-雅利安语、德拉维语、马来-波利尼西亚语和班图语家族。使用Bandit模型的基准测试结果表明,对多语言数据的训练即使在数据可用性低的语言中也能产生显著的模型泛化能力。即使在数据可用性高的语言中,多语言模型的性能通常与在单语言CASS数据集上训练的专用模型相当或更好。摘要:Cinematic audio source separation (CASS) is a relatively new subtask of audio source separation, concerned with the separation of a mixture into the dialogue, music, and effects stems. To date, only one publicly available dataset exists for CASS, that is, the Divide and Remaster (DnR) dataset, which is currently at version 2. While DnR v2 has been an incredibly useful resource for CASS, several areas of improvement have been identified, particularly through its use in the 2023 Sound Demixing Challenge. In this work, we develop version 3 of the DnR dataset, addressing issues relating to vocal content in non-dialogue stems, loudness distributions, mastering process, and linguistic diversity. In particular, the dialogue stem of DnR v3 includes speech content from more than 30 languages from multiple families including but not limited to the Germanic, Romance, Indo-Aryan, Dravidian, Malayo-Polynesian, and Bantu families. Benchmark results using the Bandit model indicated that training on multilingual data yields significant generalizability to the model even in languages with low data availability. Even in languages with high data availability, the multilingual model often performs on par or better than dedicated models trained on monolingual CASS datasets.


eess.AS音频处理
【1】 AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning
标题: AVCAP:利用视听功能作为文本标记进行字幕
作者:Jongsuk Kim,Jiwon Shin,Junmo Kim
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:近年来,表征学习和语言模型的进步将自动字幕(AC)推向了新的高度,从而能够生成人类级别的描述。利用这些进步,我们提出了 textbf{AVCap},一个 textbf{A}udio- textbf{V}isual textbf{Cap}tioning框架,一个简单而强大的基线方法适用于视听字幕。AVCap利用视听特征作为文本标记,不仅在性能上,而且在模型的可扩展性和可伸缩性方面具有许多优势。AVCap围绕三个关键维度设计:探索最佳视听编码器架构,根据生成文本的特征调整预训练模型,以及调查字幕中模态融合的有效性。我们的方法在所有指标上都优于现有的视听字幕方法,代码可在https: github.com JongSuk1 AVCap上获得摘要:In recent years, advancements in representation learning and language models have propelled Automated Captioning (AC) to new heights, enabling the generation of human-level descriptions. Leveraging these advancements, we propose textbf{AVCap}, an textbf{A}udio- textbf{V}isual textbf{Cap}tioning framework, a simple yet powerful baseline approach applicable to audio-visual captioning. AVCap utilizes audio-visual features as text tokens, which has many advantages not only in performance but also in the extensibility and scalability of the model. AVCap is designed around three pivotal dimensions: the exploration of optimal audio-visual encoder architectures, the adaptation of pre-trained models according to the characteristics of generated text, and the investigation into the efficacy of modality fusion in captioning. Our method outperforms existing audio-visual captioning methods across all metrics and the code is available on https: github.com JongSuk1 AVCap

【2】 Remastering Divide and Remaster: A Cinematic Audio Source Separation Dataset with Multilingual Support
标题: 重制Divide和Remaster:具有多语言支持的电影音频源分离数据集
作者:Karn N. Watcharasupat,Chih-Wei Wu,Iroro Orife
备注:Submitted to the 5th IEEE International Symposium on the Internet of Sounds
链接:点击下载PDF文件
摘要:电影音频源分离(CASS)是音频源分离的一个相对较新的子任务,涉及到将混合物分离成对话,音乐和效果茎。到目前为止,CASS只有一个公开可用的数据集,即Divide and Remaster(DnR)数据集,目前版本为2。虽然DnR v2一直是CASS的一个非常有用的资源,但已经确定了几个改进领域,特别是通过在2023年声音分层挑战赛中的使用。在这项工作中,我们开发了DnR数据集的第3版,解决了与非对话干中的声音内容,响度分布,掌握过程和语言多样性有关的问题。特别地,DnR v3的对话词干包括来自多个家族的30多种语言的语音内容,包括但不限于日耳曼语、罗曼语、印度-雅利安语、德拉维语、马来-波利尼西亚语和班图语家族。使用Bandit模型的基准测试结果表明,对多语言数据的训练即使在数据可用性低的语言中也能产生显著的模型泛化能力。即使在数据可用性高的语言中,多语言模型的性能通常与在单语言CASS数据集上训练的专用模型相当或更好。摘要:Cinematic audio source separation (CASS) is a relatively new subtask of audio source separation, concerned with the separation of a mixture into the dialogue, music, and effects stems. To date, only one publicly available dataset exists for CASS, that is, the Divide and Remaster (DnR) dataset, which is currently at version 2. While DnR v2 has been an incredibly useful resource for CASS, several areas of improvement have been identified, particularly through its use in the 2023 Sound Demixing Challenge. In this work, we develop version 3 of the DnR dataset, addressing issues relating to vocal content in non-dialogue stems, loudness distributions, mastering process, and linguistic diversity. In particular, the dialogue stem of DnR v3 includes speech content from more than 30 languages from multiple families including but not limited to the Germanic, Romance, Indo-Aryan, Dravidian, Malayo-Polynesian, and Bantu families. Benchmark results using the Bandit model indicated that training on multilingual data yields significant generalizability to the model even in languages with low data availability. Even in languages with high data availability, the multilingual model often performs on par or better than dedicated models trained on monolingual CASS datasets.

【3】 RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
标题: RT-LA-VocE:实时低SNR视听语音增强
作者:Honglie Chen,Rodrigo Mira,Stavros Petridis,Maja Pantic
备注:Interspeech 2024
链接:点击下载PDF文件
摘要:在本文中,我们的目标是从实时视频流和嘈杂的音频流中逐帧生成干净的语音,而不依赖于未来的输入。为此,我们提出了RT-LA-VocE,它完全重新设计了LA-VocE,一个国家的最先进的非因果视听语音增强模型的每个组件,执行因果实时推理与40毫秒的输入帧。我们通过设计新的视觉和音频编码器来实现这一点,这些编码器仅依赖于过去的帧,用Emformer取代Transformer编码器,并设计新的因果神经声码器C-HiFi-GAN。在流行的AVSpeech数据集上,我们证明了我们的算法在所有实时场景中都达到了最先进的效果。更重要的是,每个组件都经过精心调整,以将算法延迟降至理论最低值(40 ms),同时保持每帧28.15 ms的低端到端处理延迟,从而以最小的延迟实现实时逐帧增强。摘要:In this paper, we aim to generate clean speech frame by frame from a live video stream and a noisy audio stream without relying on future inputs. To this end, we propose RT-LA-VocE, which completely re-designs every component of LA-VocE, a state-of-the-art non-causal audio-visual speech enhancement model, to perform causal real-time inference with a 40ms input frame. We do so by devising new visual and audio encoders that rely solely on past frames, replacing the Transformer encoder with the Emformer, and designing a new causal neural vocoder C-HiFi-GAN. On the popular AVSpeech dataset, we show that our algorithm achieves state-of-the-art results in all real-time scenarios. More importantly, each component is carefully tuned to minimize the algorithm latency to the theoretical minimum (40ms) while maintaining a low end-to-end processing latency of 28.15ms per frame, enabling real-time frame-by-frame enhancement with minimal delay.

【4】 SaMoye: Zero-shot Singing Voice Conversion Based on Feature Disentanglement and Synthesis
标题: SaMoye:基于特征解纠缠和合成的Zero-Shot歌唱声音转换
作者:Zihao Wang,Le Ma,Yan Liu,Kejun Zhang
备注:7 pages, 4 figures
链接:点击下载PDF文件
摘要:歌唱声音转换(SVC)旨在将给定音乐作品中的歌手的声音转换为另一个歌手,同时保持原始内容。我们提出了一个端到端的功能分解为基础的模型,我们命名为SaMoye,使zero-shot多对多的歌声转换。SaMoye将歌唱声音的特征分解为内容特征、音色特征和音高特征。使用基于GPT的模型来执行与歌词的音素的交叉预测,以增强内容特征。SaMoye可以通过将音色特征替换为目标歌手来生成具有转换语音的音乐。我们还建立了一个无与伦比的大规模数据集,以保证zero-shot性能。该数据集由1500k个纯歌唱声乐片段组成,其中包含至少10,000名歌手。摘要:Singing voice conversion (SVC) aims to convert a singer's voice in a given music piece to another singer while keeping the original content. We propose an end-to-end feature disentanglement-based model, which we named SaMoye, to enable zero-shot many-to-many singing voice conversion. SaMoye disentangles the features of the singing voice into content features, timbre features, and pitch features respectively. The content features are enhanced using a GPT-based model to perform cross-prediction with the phoneme of the lyrics. SaMoye can generate the music with converted voice by replacing the timbre features with the target singer. We also establish an unparalleled large-scale dataset to guarantee zero-shot performance. The dataset consists of 1500k pure singing vocal clips containing at least 10,000 singers.

【5】 Targeted Augmented Data for Audio Deepfake Detection
标题: 用于音频Deepfake检测的定向增强数据
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in EUSIPCO 2024
链接:点击下载PDF文件
摘要:高度可信的音频深度伪造生成器的可用性突出了设计强大的音频深度伪造检测器的必要性。现有的工作往往仅仅依赖于训练集中的真实和虚假数据,这可能导致过拟合,从而降低了对不可见操作的鲁棒性。为了增强音频deepfake检测器的泛化能力,我们提出了一种新的增强方法,用于生成针对模型决策边界的音频伪假。受对抗性攻击的启发,我们扰动原始真实数据来合成具有模糊预测概率的伪伪造。两个著名的架构上的综合实验表明,所提出的增强有助于提高这些架构的泛化能力。摘要:The availability of highly convincing audio deepfake generators highlights the need for designing robust audio deepfake detectors. Existing works often rely solely on real and fake data available in the training set, which may lead to overfitting, thereby reducing the robustness to unseen manipulations. To enhance the generalization capabilities of audio deepfake detectors, we propose a novel augmentation method for generating audio pseudo-fakes targeting the decision boundary of the model. Inspired by adversarial attacks, we perturb original real data to synthesize pseudo-fakes with ambiguous prediction probabilities. Comprehensive experiments on two well-known architectures demonstrate that the proposed augmentation contributes to improving the generalization capabilities of these architectures.

【6】 Scaling Law in Neural Data: Non-Invasive Speech Decoding with 175 Hours of EEG Data
标题: 神经数据中的比例定律:使用175小时脑电数据的无创语音解码
作者:Motoshige Sato,Kenichi Tomeoka,Ilya Horiguchi,Kai Arulkumaran,Ryota Kanai,Shuntaro Sasai
链接:点击下载PDF文件
摘要:脑机接口(BCI)在帮助有语言障碍的人方面具有巨大的潜力。利用脑电图(EEG)解码语音是特别有前途的,由于其非侵入性的性质。然而,记录通常很短,并且EEG数据的高度可变性导致研究人员专注于具有几十个类别的分类任务。为了评估其在语音神经假体中的实用性,我们研究了开放词汇设置中EEG数据大小与解码准确性之间的关系。我们收集了大量的EEG数据,从一个单一的参与者(175小时),并进行了zero-shot语音段分类使用自监督表示学习。在整个数据集上训练的模型达到了48%的前1准确率和76%的前10准确率,同时减轻了肌电位伪影的影响。相反,当数据被限制在实践中使用的典型数量(10小时)时,前1名的准确率下降到2.5%,显示出显着的缩放效应。此外,随着训练数据量的增加,EEG潜在表征逐渐表现出更清晰的口语短语的时间结构。这表明解码器可以以数据驱动的方式识别语音片段,而无需显式测量单词识别。这项研究标志着基于脑电的语音脑机接口的实际实现迈出了重要的一步。摘要:Brain-computer interfaces (BCIs) hold great potential for aiding individuals with speech impairments. Utilizing electroencephalography (EEG) to decode speech is particularly promising due to its non-invasive nature. However, recordings are typically short, and the high variability in EEG data has led researchers to focus on classification tasks with a few dozen classes. To assess its practical applicability for speech neuroprostheses, we investigate the relationship between the size of EEG data and decoding accuracy in the open vocabulary setting. We collected extensive EEG data from a single participant (175 hours) and conducted zero-shot speech segment classification using self-supervised representation learning. The model trained on the entire dataset achieved a top-1 accuracy of 48 % and a top-10 accuracy of 76 %, while mitigating the effects of myopotential artifacts. Conversely, when the data was limited to the typical amount used in practice ($ sim$10 hours), the top-1 accuracy dropped to 2.5 %, revealing a significant scaling effect. Additionally, as the amount of training data increased, the EEG latent representation progressively exhibited clearer temporal structures of spoken phrases. This indicates that the decoder can recognize speech segments in a data-driven manner without explicit measurements of word recognition. This research marks a significant step towards the practical realization of EEG-based speech BCIs.

【7】 HebDB: a Weakly Supervised Dataset for Hebrew Speech Processing
标题: HebDB:希伯来语语音处理的弱监督数据集
作者:Arnon Turetzky,Or Tal,Yael Segal-Feldman,Yehoshua Dissen,Ella Zeldes,Amit Roth,Eyal Cohen,Yosi Shrem,Bronya R. Chernyak,Olga Seleznova,Joseph Keshet,Yossi Adi
备注:Accepted at Interspeech2024
链接:点击下载PDF文件
摘要:我们提出了HebDB,一个弱监督的数据集,用于希伯来语的口语处理。HebDB提供大约2500小时的希伯来语自然和自发的演讲录音,包括各种各样的演讲者和主题。我们提供原始录音以及预处理,弱监督和过滤版本。HebDB的目标是进一步加强希伯来语口语处理工具的研究和开发。因此,我们还提供了两个用于自动语音识别(ASR)的基线系统:(i)自监督模型;以及(ii)完全监督模型。我们提出了这两种方法在HebDB上优化的性能,并将其与当前的多语言ASR替代方案进行了比较。结果表明,所提出的方法达到更好的结果比评估基线考虑类似的模型大小。数据集、代码和模型可在https: pages.cs.huji.ac.il adiyoss-lab HebDB 上公开获取。摘要:We present HebDB, a weakly supervised dataset for spoken language processing in the Hebrew language. HebDB offers roughly 2500 hours of natural and spontaneous speech recordings in the Hebrew language, consisting of a large variety of speakers and topics. We provide raw recordings together with a pre-processed, weakly supervised, and filtered version. The goal of HebDB is to further enhance research and development of spoken language processing tools for the Hebrew language. Hence, we additionally provide two baseline systems for Automatic Speech Recognition (ASR): (i) a self-supervised model; and (ii) a fully supervised model. We present the performance of these two methods optimized on HebDB and compare them to current multi-lingual ASR alternatives. Results suggest the proposed method reaches better results than the evaluated baselines considering similar model sizes. Dataset, code, and models are publicly available under https: pages.cs.huji.ac.il adiyoss-lab HebDB .

【8】 Beat-It: Beat-Synchronized Multi-Condition 3D Dance Generation
标题: Beat-It:节拍同步多条件3D舞蹈生成
作者:Zikai Huang,Xuemiao Xu,Cheng Xu,Huaidong Zhang,Chenxi Zheng,Jing Qin,Shengfeng He
备注:ECCV 2024
链接:点击下载PDF文件
摘要:舞蹈作为一种艺术形式,从根本上讲取决于与音乐节拍的精确同步。然而,从音乐中实现美观的舞蹈序列是具有挑战性的,现有的方法往往在可控性和节拍对齐方面不足。为了解决这些缺点,本文介绍了Beat-It,一种新颖的框架,用于特定节拍,关键姿势引导的舞蹈生成。与之前的方法不同,Beat-It独特地集成了明确的节拍意识和关键姿势指导,有效地解决了两个主要问题:生成的舞蹈动作与音乐节拍的不一致,以及无法将关键姿势映射到特定节拍,这对于实际的编舞至关重要。我们的方法使用最近的节拍距离表示从音乐中解开节拍条件,并采用分层多条件融合机制。这种机制无缝地整合了关键姿势,节拍和音乐特征,减轻了条件冲突,并为舞蹈生成提供了丰富的多条件指导。此外,一个专门设计的节拍对齐损失确保生成的舞蹈动作保持与指定的节拍同步。大量的实验证实了Beat-It在节拍对齐和运动可控性方面优于现有的最先进的方法。摘要:Dance, as an art form, fundamentally hinges on the precise synchronization with musical beats. However, achieving aesthetically pleasing dance sequences from music is challenging, with existing methods often falling short in controllability and beat alignment. To address these shortcomings, this paper introduces Beat-It, a novel framework for beat-specific, key pose-guided dance generation. Unlike prior approaches, Beat-It uniquely integrates explicit beat awareness and key pose guidance, effectively resolving two main issues: the misalignment of generated dance motions with musical beats, and the inability to map key poses to specific beats, critical for practical choreography. Our approach disentangles beat conditions from music using a nearest beat distance representation and employs a hierarchical multi-condition fusion mechanism. This mechanism seamlessly integrates key poses, beats, and music features, mitigating condition conflicts and offering rich, multi-conditioned guidance for dance generation. Additionally, a specially designed beat alignment loss ensures the generated dance movements remain in sync with the designated beats. Extensive experiments confirm Beat-It's superiority over existing state-of-the-art methods in terms of beat alignment and motion controllability.

【9】 Video-to-Audio Generation with Hidden Alignment
标题: 具有隐藏对齐的视频到音频生成
作者:Manjie Xu,Chenxing Li,Yong Ren,Rilin Chen,Yu Gu,Wei Liang,Dong Yu
备注:this https URL
链接:点击下载PDF文件
摘要:根据视频输入生成语义上和时间上对齐的音频内容已经成为研究人员的焦点,特别是在文本到视频生成方面取得显著突破之后。在这项工作中,我们的目标是提供对视频到音频生成范式的见解,重点关注三个关键方面:视觉编码器,辅助嵌入和数据增强技术。从建立在简单但令人惊讶的有效直觉基础上的基础模型VTA-LDM开始,我们通过消融研究探索各种视觉编码器和辅助嵌入。采用全面的评估管道,强调生成质量和视频-音频同步对齐,我们证明了我们的模型具有最先进的视频到音频生成能力。此外,我们提供了重要的见解不同的数据增强方法对提高生成框架的整体能力的影响。我们展示的可能性,以推进从语义和时间的角度生成同步音频的挑战。我们希望这些见解将成为开发更现实和准确的视听生成模型的垫脚石。摘要:Generating semantically and temporally aligned audio content in accordance with video input has become a focal point for researchers, particularly following the remarkable breakthrough in text-to-video generation. In this work, we aim to offer insights into the video-to-audio generation paradigm, focusing on three crucial aspects: vision encoders, auxiliary embeddings, and data augmentation techniques. Beginning with a foundational model VTA-LDM built on a simple yet surprisingly effective intuition, we explore various vision encoders and auxiliary embeddings through ablation studies. Employing a comprehensive evaluation pipeline that emphasizes generation quality and video-audio synchronization alignment, we demonstrate that our model exhibits state-of-the-art video-to-audio generation capabilities. Furthermore, we provide critical insights into the impact of different data augmentation methods on enhancing the generation framework's overall capacity. We showcase possibilities to advance the challenge of generating synchronized audio from semantic and temporal perspectives. We hope these insights will serve as a stepping stone toward developing more realistic and accurate audio-visual generation models.

【10】 Out-of-distribution generalisation in spoken language understanding
标题: 口语理解中的非分布概括
作者:Dejan Porjazovski,Anssi Moisio,Mikko Kurimo
备注:Accepted for INTERSPEECH 2024
链接:点击下载PDF文件
摘要:当测试数据意外地与训练数据不同时,它被称为分布外(OOD),这是机器学习的真实用例中的一个常见挑战。虽然OOD概括近年来获得了兴趣,很少有作品集中在OOD概括口语理解(SLU)任务。为了促进对这一主题的研究,我们介绍了一个修改版本的流行SLU数据集SLURP,具有数据分裂测试OOD概括在SLU任务。我们将修改后的数据集称为SLURP For OOD泛化,或SLURPFOOD。利用我们的OOD数据分割,我们发现端到端SLU模型的泛化能力有限。此外,通过采用模型的可解释性技术,我们揭示了导致模型泛化困难的因素。为了提高泛化能力,我们尝试了两种技术,这两种技术提高了一些但不是所有分裂的结果,强调了对新技术的需求。摘要:Test data is said to be out-of-distribution (OOD) when it unexpectedly differs from the training data, a common challenge in real-world use cases of machine learning. Although OOD generalisation has gained interest in recent years, few works have focused on OOD generalisation in spoken language understanding (SLU) tasks. To facilitate research on this topic, we introduce a modified version of the popular SLU dataset SLURP, featuring data splits for testing OOD generalisation in the SLU task. We call our modified dataset SLURP For OOD generalisation, or SLURPFOOD. Utilising our OOD data splits, we find end-to-end SLU models to have limited capacity for generalisation. Furthermore, by employing model interpretability techniques, we shed light on the factors contributing to the generalisation difficulties of the models. To improve the generalisation, we experiment with two techniques, which improve the results on some, but not all the splits, emphasising the need for new techniques.

【11】 STONE: Self-supervised Tonality Estimator
标题: STONE:自我监督的调性估计器
作者:Yuexuan Kong,Vincent Lostanlen,Gabriel Meseguer-Brocal,Stella Wong,Mathieu Lagrange,Romain Hennequin
链接:点击下载PDF文件
摘要:虽然深度神经网络可以估计音乐作品的基调,但它们的监督会导致大量的注释工作。针对这一缺点,我们提出了石头,第一个自我监督的音调估计。STONE背后的架构名为ChromaNet,是一个具有八度等效性的convnet,它输出12个结构化logit的密钥签名配置文件(KSP)。首先,我们训练ChromaNet来回归来自同一音轨的任何两个未标记的音乐摘录之间的人工音高置换,如在五分之一圈(CoF)内测量的交叉功率谱密度(CPSD)。我们观察到,这种自我监督的借口任务导致KSP与音调的关键签名。基于这一观察,我们扩展石头输出一个结构化的KSP的24个logits,并引入监督,以消除歧义的主要与次要的密钥共享相同的密钥签名。应用不同的监督量产生半监督和全监督音调估计器:即,半色调和超色调。我们评估这些估计FMAK,一个新的数据集的5489个真实世界的音乐录音与专家注释的24个主要和次要的关键。我们发现,Semi-TONE匹配Sup-TONE的分类精度与减少监督,并优于它与平等的监督。摘要:Although deep neural networks can estimate the key of a musical piece, their supervision incurs a massive annotation effort. Against this shortcoming, we present STONE, the first self-supervised tonality estimator. The architecture behind STONE, named ChromaNet, is a convnet with octave equivalence which outputs a key signature profile (KSP) of 12 structured logits. First, we train ChromaNet to regress artificial pitch transpositions between any two unlabeled musical excerpts from the same audio track, as measured as cross-power spectral density (CPSD) within the circle of fifths (CoF). We observe that this self-supervised pretext task leads KSP to correlate with tonal key signature. Based on this observation, we extend STONE to output a structured KSP of 24 logits, and introduce supervision so as to disambiguate major versus minor keys sharing the same key signature. Applying different amounts of supervision yields semi-supervised and fully supervised tonality estimators: i.e., Semi-TONEs and Sup-TONEs. We evaluate these estimators on FMAK, a new dataset of 5489 real-world musical recordings with expert annotation of 24 major and minor keys. We find that Semi-TONE matches the classification accuracy of Sup-TONE with reduced supervision and outperforms it with equal supervision.

【12】 SimuSOE: A Simulated Snoring Dataset for Obstructive Sleep Apnea-Hypopnea Syndrome Evaluation during Wakefulness
标题: SimuSOE:清醒期间评估阻塞性睡眠呼吸暂停-低通气综合征的模拟打鼾数据集
作者:Jie Lin,Xiuping Yang,Li Xiao,Xinhong Li,Weiyan Yi,Yuhong Yang,Weiping Tu,Xiong Chen
链接:点击下载PDF文件
摘要:阻塞性睡眠呼吸暂停低通气综合征(OSAHS)是一种常见的上呼吸道阻塞引起的慢性呼吸障碍。以前的研究通过基于睡眠打鼾或语音信号数据集训练的基于机器学习的系统来推进OSAHS评估。然而,构建用于训练精确且快速的OSAHS评估系统的数据集提出了挑战,因为1)收集睡眠打鼾是耗时的,以及2)语音信号在反映上气道阻塞方面受到限制。在本文中,我们提出了一个新的打鼾数据集的OSAHS评估,命名为SimuSOE,其中引入了一种新颖的和时间有效的打鼾收集方法来解决上述问题。特别是,我们采用模拟打鼾,这是一种由患者故意发出的打鼾,以取代自然打鼾。实验结果表明,模拟清醒状态下的鼾声信号可以作为OSAHS初筛的有效特征。摘要:Obstructive Sleep Apnea-Hypopnea Syndrome (OSAHS) is a prevalent chronic breathing disorder caused by upper airway obstruction. Previous studies advanced OSAHS evaluation through machine learning-based systems trained on sleep snoring or speech signal datasets. However, constructing datasets for training a precise and rapid OSAHS evaluation system poses a challenge, since 1) it is time-consuming to collect sleep snores and 2) the speech signal is limited in reflecting upper airway obstruction. In this paper, we propose a new snoring dataset for OSAHS evaluation, named SimuSOE, in which a novel and time-effective snoring collection method is introduced for tackling the above problems. In particular, we adopt simulated snoring which is a type of snore intentionally emitted by patients to replace natural snoring. Experimental results indicate that the simulated snoring signal during wakefulness can serve as an effective feature in OSAHS preliminary screening.

【13】 Speech After Gender: A Trans-Feminine Perspective on Next Steps for Speech Science and Technology
标题: 性别后的言语:言语科学与技术下一步的跨女性视角
作者:Robin Netzorg,Alyssa Cote,Sumi Koshin,Klo Vivienne Garoute,Gopala Krishna Anumanchipalli
链接:点击下载PDF文件
摘要:作为语音修改的专家,跨女性性别肯定语音教师对语音有独特的观点,混淆了目前对说话者身份的理解。为了证明这一点,我们提出了多功能语音数据集(VVD),一个集合的三个扬声器修改他们的声音沿性别轴。的VVD说明,目前的方法在扬声器建模,基于性别的分类概念和静态的声乐纹理的理解,未能考虑声道的灵活性。利用公开的扬声器嵌入,我们表明,性别分类系统是高度敏感的语音修改,和扬声器验证系统无法识别语音来自同一扬声器的语音修改变得更加激烈。作为走向超越说话者身份的分类和静态概念的一条路径,我们建议建模的声音纹理,如音高,共鸣和重量的个人素质。摘要:As experts in voice modification, trans-feminine gender-affirming voice teachers have unique perspectives on voice that confound current understandings of speaker identity. To demonstrate this, we present the Versatile Voice Dataset (VVD), a collection of three speakers modifying their voices along gendered axes. The VVD illustrates that current approaches in speaker modeling, based on categorical notions of gender and a static understanding of vocal texture, fail to account for the flexibility of the vocal tract. Utilizing publicly-available speaker embeddings, we demonstrate that gender classification systems are highly sensitive to voice modification, and speaker verification systems fail to identify voices as coming from the same speaker as voice modification becomes more drastic. As one path towards moving beyond categorical and static notions of speaker identity, we propose modeling individual qualities of vocal texture such as pitch, resonance, and weight.


机器翻译,仅供参考