本文经arXiv每日学术速递授权转载
【1】 Statistics-aware Audio-visual Deepfake Detector
标题: 统计感知视听Deepfake检测器
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in ICIP 2024
链接:点击下载PDF文件
【2】 Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
标题: 研究标签布局和训练标准对ASB性能和对齐质量的影响
作者:Tina Raissi,Christoph Lüscher,Simon Berger,Ralf Schlüter,Hermann Ney
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
【3】 MMSD-Net: Towards Multi-modal Stuttering Detection
标题: MMSD-Net:走向多模式口吃检测
作者:Liangyu Nie,Sudarsana Reddy Kadiri,Ruchit Agrawal
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【4】 A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
标题: 基于GLM仅使用母语语音库模拟外国口音的初步研究
作者:Kentaro Onda,Joonyong Park,Nobuaki Minematsu,Daisuke Saito
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
【5】 Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
标题: 超越二进制:使用生成性预训练变形器和端到端模型进行多类失语症检测
作者:Matthew Perez,Aneesha Sampath,Minxue Niu,Emily Mower Provost
链接:点击下载PDF文件
【6】 Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
标题: 用于多模式物理场景理解的解开声学场
作者:Jie Yin,Andrew Luo,Yilun Du,Anoop Cherian,Tim K. Marks,Jonathan Le Roux,Chuang Gan
链接:点击下载PDF文件
【7】 Knowledge boosting during low-latency inference
标题: 低延迟推理期间的知识提升
作者:Vidya Srinivas,Malek Itani,Tuochao Chen,Emre Sefik Eskimez,Takuya Yoshioka,Shyamnath Gollakota
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【8】 Exploring Gender-Specific Speech Patterns in Automatic Suicide Risk Assessment
标题: 在自动自杀风险评估中探索特定性别的言语模式
作者:Maurice Gerczuk,Shahin Amiriparian,Justina Lutz,Wolfgang Strube,Irina Papazova,Alkomiet Hasan,Björn W. Schuller
备注:accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【9】 Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
标题: 具有自我监督音频屏蔽自动编码器的通用声音分离
作者:Junqi Zhao,Xubo Liu,Jinzheng Zhao,Yi Yuan,Qiuqiang Kong,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
标题: Vibravox:用体导音频传感器捕获的法语语音数据集
作者:Julien Hauret,Malo Olivier,Thomas Joubaud,Christophe Langrenne,Sarah Poirée,Véronique Zimpfer,Éric Bavu
备注:19 pages, 15 figures
链接:点击下载PDF文件
【2】 Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
标题: 具有自我监督音频屏蔽自动编码器的通用声音分离
作者:Junqi Zhao,Xubo Liu,Jinzheng Zhao,Yi Yuan,Qiuqiang Kong,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
【3】 MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
标题: MUSA:通过序列解纠缠实现多语言说话者匿名化
作者:Jixun Yao,Qing Wang,Pengcheng Guo,Ziqian Ning,Yuguang Yang,Yu Pan,Lei Xie
备注:Submitted to TASLP
链接:点击下载PDF文件
【4】 The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
标题: VoicePrivacy 2022挑战:语音匿名化的进展和前景
作者:Michele Panariello,Natalia Tomashenko,Xin Wang,Xiaoxiao Miao,Pierre Champion,Hubert Nourtel,Massimiliano Todisco,Nicholas Evans,Emmanuel Vincent,Junichi Yamagishi
备注:Accepted at IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
【5】 VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
标题: VoxBlink 2:10万+说话人识别数据库和开放式说话人识别基准
作者:Yuke Lin,Ming Cheng,Fulin Zhang,Yingying Gao,Shilei Zhang,Ming Li
备注:Accepted By InterSpeech2024
链接:点击下载PDF文件
【6】 Team HYU ASML ROBOVOX SP Cup 2024 System Description
标题: HyU ASML ROBOVOX SP Cup 2024系统描述
作者:Jeong-Hwan Choi,Gaeun Kim,Hee-Jae Lee,Seyun Ahn,Hyun-Soo Kim,Joon-Hyuk Chang
备注:Technical report for IEEE Signal Processing Cup 2024, 9 pages
链接:点击下载PDF文件
【7】 Statistics-aware Audio-visual Deepfake Detector
标题: 统计感知视听Deepfake检测器
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in ICIP 2024
链接:点击下载PDF文件
【8】 Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
标题: 研究标签布局和训练标准对ASB性能和对齐质量的影响
作者:Tina Raissi,Christoph Lüscher,Simon Berger,Ralf Schlüter,Hermann Ney
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
【9】 MMSD-Net: Towards Multi-modal Stuttering Detection
标题: MMSD-Net:走向多模式口吃检测
作者:Liangyu Nie,Sudarsana Reddy Kadiri,Ruchit Agrawal
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【10】 A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
标题: 基于GLM仅使用母语语音库模拟外国口音的初步研究
作者:Kentaro Onda,Joonyong Park,Nobuaki Minematsu,Daisuke Saito
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
【11】 Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
标题: 超越二进制:使用生成性预训练变形器和端到端模型进行多类失语症检测
作者:Matthew Perez,Aneesha Sampath,Minxue Niu,Emily Mower Provost
链接:点击下载PDF文件
【12】 Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
标题: 用于多模式物理场景理解的解开声学场
作者:Jie Yin,Andrew Luo,Yilun Du,Anoop Cherian,Tim K. Marks,Jonathan Le Roux,Chuang Gan
链接:点击下载PDF文件
【13】 Target conversation extraction: Source separation using turn-taking dynamics
标题: 目标对话提取:使用轮流动力学进行源分离
作者:Tuochao Chen,Qirui Wang,Bohan Wu,Malek Itani,Emre Sefik Eskimez,Takuya Yoshioka,Shyamnath Gollakota
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【14】 Knowledge boosting during low-latency inference
标题: 低延迟推理期间的知识提升
作者:Vidya Srinivas,Malek Itani,Tuochao Chen,Emre Sefik Eskimez,Takuya Yoshioka,Shyamnath Gollakota
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
【15】 Exploring Gender-Specific Speech Patterns in Automatic Suicide Risk Assessment
标题: 在自动自杀风险评估中探索特定性别的言语模式
作者:Maurice Gerczuk,Shahin Amiriparian,Justina Lutz,Wolfgang Strube,Irina Papazova,Alkomiet Hasan,Björn W. Schuller
备注:accepted at INTERSPEECH 2024
链接:点击下载PDF文件
【16】 Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
标题: 级联流语音翻译中MT束搜索雷区的导航
作者:Rastislav Rabatin,Frank Seide,Ernie Chang
链接:点击下载PDF文件
标题: 统计感知视听Deepfake检测器
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in ICIP 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种增强的视听深度检测方法。最近的视听Deepfake检测方法主要评估音频和视觉特征之间的同步。虽然他们已经显示出有前途的结果,他们是基于孤立的特征距离的最大化 最小化,而不考虑特征统计。此外,它们依赖于繁琐的深度学习架构,并且严重依赖于经验固定的超参数。在此,为了克服这些限制,我们提出:(1)统计特征损失以增强模型的区分能力,而不是仅仅依赖于特征距离;(2)使用波形来描述音频作为基于频率的表示的替代;(3)伪造分数的后处理归一化;(4)使用更浅的网络来降低计算复杂度。在DFDC和FakeAVCeleb数据集上的实验证明了该方法的相关性。摘要:In this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results, they are based on the maximization minimization of isolated feature distances without considering feature statistics. Moreover, they rely on cumbersome deep learning architectures and are heavily dependent on empirically fixed hyperparameters. Herein, to overcome these limitations, we propose: (1) a statistical feature loss to enhance the discrimination capability of the model, instead of relying solely on feature distances; (2) using the waveform for describing the audio as a replacement of frequency-based representations; (3) a post-processing normalization of the fakeness score; (4) the use of shallower network for reducing the computational complexity. Experiments on the DFDC and FakeAVCeleb datasets demonstrate the relevance of the proposed method.
【2】 Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
标题: 研究标签布局和训练标准对ASB性能和对齐质量的影响
作者:Tina Raissi,Christoph Lüscher,Simon Berger,Ralf Schlüter,Hermann Ney
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
摘要:正在进行的自动语音识别(ASR)研究设想了端到端方法和经典模块化系统之间的明确划分。即使两种方法之间的高层次比较,其要求和(缺点)的优点,通常是解决,在类似的条件下,更密切的比较是不容易在文献中。在这项工作中,我们提出了一个比较集中的标签拓扑结构和训练标准。我们比较了两个歧视性的对齐模型与隐马尔可夫模型(HMM)和连接主义的时间分类拓扑结构,和两个一阶标签上下文ASR模型,分别利用因子HMM和严格单调递归神经网络换能器。我们使用不同的测量来评估对齐质量,并比较我们最好的系统的字错误率和实时因素。在LibriSpeech 960h和Switchboard 300h任务上进行了实验。摘要:The ongoing research scenario for automatic speech recognition (ASR) envisions a clear division between end-to-end approaches and classic modular systems. Even though a high-level comparison between the two approaches in terms of their requirements and (dis)advantages is commonly addressed, a closer comparison under similar conditions is not readily available in the literature. In this work, we present a comparison focused on the label topology and training criterion. We compare two discriminative alignment models with hidden Markov model (HMM) and connectionist temporal classification topology, and two first-order label context ASR models utilizing factored HMM and strictly monotonic recurrent neural network transducer, respectively. We use different measurements for the evaluation of the alignment quality, and compare word error rate and real time factor of our best systems. Experiments are conducted on the LibriSpeech 960h and Switchboard 300h tasks.
【3】 MMSD-Net: Towards Multi-modal Stuttering Detection
标题: MMSD-Net:走向多模式口吃检测
作者:Liangyu Nie,Sudarsana Reddy Kadiri,Ruchit Agrawal
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:口吃是一种常见的语言障碍,由言语产生的不规则中断引起,影响着全世界7000多万人。标准的自动语音处理工具不考虑语音疾病,因此当以口吃语音作为输入时不能生成有意义的结果。口吃的自动检测是构建高效的、上下文感知的语音处理系统的重要一步。虽然以前的方法探索了统计和神经方法来检测口吃,但所有这些方法本质上都是单峰的。本文介绍了MMSD-Net,第一个多模态神经网络框架口吃检测。实验和结果表明,将视觉信号显着辅助口吃检测,我们的模型产生了2-17%的改善,在F1分数比现有的国家的最先进的单峰方法。摘要:Stuttering is a common speech impediment that is caused by irregular disruptions in speech production, affecting over 70 million people across the world. Standard automatic speech processing tools do not take speech ailments into account and are thereby not able to generate meaningful results when presented with stuttered speech as input. The automatic detection of stuttering is an integral step towards building efficient, context-aware speech processing systems. While previous approaches explore both statistical and neural approaches for stuttering detection, all of these methods are uni-modal in nature. This paper presents MMSD-Net, the first multi-modal neural framework for stuttering detection. Experiments and results demonstrate that incorporating the visual signal significantly aids stuttering detection, and our model yields an improvement of 2-17% in the F1-score over existing state-of-the-art uni-modal approaches.
【4】 A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
标题: 基于GLM仅使用母语语音库模拟外国口音的初步研究
作者:Kentaro Onda,Joonyong Park,Nobuaki Minematsu,Daisuke Saito
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
摘要:本文提出了一种基于生成式口语模型(GSLM)的外语重读模拟方法。当一个人听一种外语的口语单词并重复它们时,重复的语音通常带有该听者的L1的口音。据说这是因为说出来的单词在心理上被表征为L1的一系列语音单位,而这些单位用于口头再现。我们模拟这个过程中输入语言A的语音到GSLM的语言B添加B的口音到输入语音。针对外国输入语音运行L1的ASR并将ASR结果给予L1的TTS的过程可以被视为该方法的简单实现。我们的实验结果表明,合成的口音的输出语音是非常自然的,相比,真实的样本A产生的扬声器的L1是B,和重音的程度是可控的。摘要:We propose a method of simulating the human process of foreign accentuation using Generative Spoken Language Model (GSLM) only with native speech corpora. When one listens to spoken words of a foreign language and repeats them, the repeated speech is often with the accent of that listener's L1. This is said to be because the spoken words are mentally represented as a sequence of phonological units of the L1, and those units are used for oral reproduction. We simulate this process by inputting speech of language A into GSLM of language B to add B's accent onto the input speech. The process of running ASR of the L1 for foreign input speech and giving the ASR result to TTS of the L1 can be viewed as a naive implementation of this approach. The results of our experiments show that the synthesized accent of the output speech is highly natural, compared to real samples of A generated by speakers whose L1 is B, and that the degree of accentuation is controllable.
【5】 Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
标题: 超越二进制:使用生成性预训练变形器和端到端模型进行多类失语症检测
作者:Matthew Perez,Aneesha Sampath,Minxue Niu,Emily Mower Provost
链接:点击下载PDF文件
摘要:失语症是一种语言障碍,可导致言语错误,称为paraphasias,涉及误用,替换或发明的话。自动错语检测可以通过促进临床评估和治疗计划选择来帮助失语症患者。然而,大多数自动错语检测工作只集中在二进制检测,这涉及到识别的存在或不存在的错语。多类错语检测代表了一个未探索的研究领域,其重点是识别多种类型的错语以及它们在给定的语音片段中发生的位置。我们提出了新的方法,使用生成预训练的Transformer(GPT),以识别从成绩单以及两个端到端的方法,重点是建模自动语音识别(ASR)和失语症分类为多个序列与单一序列。我们证明了一个单一的序列模型优于GPT基线多类错语检测。摘要:Aphasia is a language disorder that can lead to speech errors known as paraphasias, which involve the misuse, substitution, or invention of words. Automatic paraphasia detection can help those with Aphasia by facilitating clinical assessment and treatment planning options. However, most automatic paraphasia detection works have focused solely on binary detection, which involves recognizing only the presence or absence of a paraphasia. Multiclass paraphasia detection represents an unexplored area of research that focuses on identifying multiple types of paraphasias and where they occur in a given speech segment. We present novel approaches that use a generative pretrained transformer (GPT) to identify paraphasias from transcripts as well as two end-to-end approaches that focus on modeling both automatic speech recognition (ASR) and paraphasia classification as multiple sequences vs. a single sequence. We demonstrate that a single sequence model outperforms GPT baselines for multiclass paraphasia detection.
【6】 Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
标题: 用于多模式物理场景理解的解开声学场
作者:Jie Yin,Andrew Luo,Yilun Du,Anoop Cherian,Tim K. Marks,Jonathan Le Roux,Chuang Gan
链接:点击下载PDF文件
摘要:我们研究了多模态物理场景理解的问题,其中一个具体的代理需要通过推断对象属性,方向和距离的影响声源找到倒下的物体。以前的工作采用前馈神经网络直接从声音回归变量,导致泛化能力差和域适应问题。在本文中,我们说明,学习的解纠缠模型的声学形成,称为解纠缠声场(ESTA),捕捉声音的产生和传播过程中,使体现代理构建一个空间的不确定性地图的对象可能已经下降。我们证明,我们的分析合成框架可以联合推断声音属性明确分解和分解的潜在空间的解开模型。我们进一步表明,空间不确定性地图可以显着提高坠落物体的定位成功率,提出多个合理的探索位置。摘要:We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward neural networks to directly regress the variables from sound, leading to poor generalization and domain adaptation issues. In this paper, we illustrate that learning a disentangled model of acoustic formation, referred to as disentangled acoustic field (DAF), to capture the sound generation and propagation process, enables the embodied agent to construct a spatial uncertainty map over where the objects may have fallen. We demonstrate that our analysis-by-synthesis framework can jointly infer sound properties by explicitly decomposing and factorizing the latent space of the disentangled model. We further show that the spatial uncertainty map can significantly improve the success rate for the localization of fallen objects by proposing multiple plausible exploration locations.
【7】 Knowledge boosting during low-latency inference
标题: 低延迟推理期间的知识提升
作者:Vidya Srinivas,Malek Itani,Tuochao Chen,Emre Sefik Eskimez,Takuya Yoshioka,Shyamnath Gollakota
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:用于低延迟、流应用的模型可以从较大模型的知识容量中受益,但由于资源限制,边缘设备无法运行这些模型。一个可能的解决方案是在推理过程中将提示从远程运行的大型模型传输到设备上运行的小型模型。然而,这会导致通信延迟,打破实时要求,并且不能保证两个模型将同时对相同的数据进行操作。我们提出了知识提升,这是一种新技术,允许大型模型在推理过程中对延时输入进行操作,同时仍能提高小型模型的性能。使用处理8 ms块的流神经网络,我们评估了通信延迟高达6个块或48 ms的不同语音分离和增强任务。我们的结果显示,当小模型和大模型之间的性能差距很大时,可以获得更大的收益,这表明这是一种很有前途的方法。用于低延迟应用的大小模型协作。代码、数据集和音频示例可在https: knowledgeboosting.cs.washington.edu 上获取。摘要:Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. Code, dataset, and audio samples available at https: knowledgeboosting.cs.washington.edu .
【8】 Exploring Gender-Specific Speech Patterns in Automatic Suicide Risk Assessment
标题: 在自动自杀风险评估中探索特定性别的言语模式
作者:Maurice Gerczuk,Shahin Amiriparian,Justina Lutz,Wolfgang Strube,Irina Papazova,Alkomiet Hasan,Björn W. Schuller
备注:accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在急诊医学中,对有自杀风险的患者的及时干预往往受到延迟获得专门精神病护理的阻碍。为了弥补这一差距,我们引入了一种基于语音的自动自杀风险评估方法。我们的研究涉及一个新的数据集,包括20名阅读中性文本的患者的语音记录。我们提取了四个语音表示,包括可解释的和深层的功能。此外,我们探讨了基于性别的建模和短语水平的正常化的影响。通过应用性别排斥建模,从情绪微调wav2vec2.0模型中提取的特征可以用于区分高自杀风险和低自杀风险,平衡准确率为81%。最后,我们的分析揭示了性别之间的言语特征和自杀风险的关系的差异。在我们的数据集中,男性的自杀风险随着情绪激动而增加,而女性受试者的声音特征则相反。摘要:In emergency medicine, timely intervention for patients at risk of suicide is often hindered by delayed access to specialised psychiatric care. To bridge this gap, we introduce a speech-based approach for automatic suicide risk assessment. Our study involves a novel dataset comprising speech recordings of 20 patients who read neutral texts. We extract four speech representations encompassing interpretable and deep features. Further, we explore the impact of gender-based modelling and phrase-level normalisation. By applying gender-exclusive modelling, features extracted from an emotion fine-tuned wav2vec2.0 model can be utilised to discriminate high- from low- suicide risk with a balanced accuracy of 81%. Finally, our analysis reveals a discrepancy in the relationship of speech characteristics and suicide risk between female and male subjects. For men in our dataset, suicide risk increases together with agitation while voice characteristics of female subjects point the other way.
【9】 Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
标题: 具有自我监督音频屏蔽自动编码器的通用声音分离
作者:Junqi Zhao,Xubo Liu,Jinzheng Zhao,Yi Yuan,Qiuqiang Kong,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
摘要:通用声音分离(USS)是分离任意声源混合的任务。通常,通用分离模型使用标记数据以监督方式从头开始训练。自监督学习(SSL)是一种新兴的深度学习方法,它利用未标记的数据来获得与任务无关的表示,这可以使许多下游任务受益。在本文中,我们提出将一个自监督的预训练模型,即音频掩蔽自动编码器(A-MAE),集成到一个通用的声音分离系统,以提高其分离性能。我们采用两种策略来利用SSL嵌入:在微调期间冻结或更新A-MAE的参数。SSL嵌入与短时傅立叶变换(STFT)级联,作为分离模型的输入特征。我们在AudioSet数据集上对我们的方法进行了评估,实验结果表明,所提出的方法成功地提高了最先进的基于ResUNet的USS模型的分离性能。摘要:Universal sound separation (USS) is a task of separating mixtures of arbitrary sound sources. Typically, universal separation models are trained from scratch in a supervised manner, using labeled data. Self-supervised learning (SSL) is an emerging deep learning approach that leverages unlabeled data to obtain task-agnostic representations, which can benefit many downstream tasks. In this paper, we propose integrating a self-supervised pre-trained model, namely the audio masked autoencoder (A-MAE), into a universal sound separation system to enhance its separation performance. We employ two strategies to utilize SSL embeddings: freezing or updating the parameters of A-MAE during fine-tuning. The SSL embeddings are concatenated with the short-time Fourier transform (STFT) to serve as input features for the separation model. We evaluate our methods on the AudioSet dataset, and the experimental results indicate that the proposed methods successfully enhance the separation performance of a state-of-the-art ResUNet-based USS model.
eess.AS音频处理
【1】 Vibravox: A Dataset of French Speech Captured with Body-conduction Audio Sensors标题: Vibravox:用体导音频传感器捕获的法语语音数据集
作者:Julien Hauret,Malo Olivier,Thomas Joubaud,Christophe Langrenne,Sarah Poirée,Véronique Zimpfer,Éric Bavu
备注:19 pages, 15 figures
链接:点击下载PDF文件
摘要:Vibravox是一个符合《通用数据保护条例》(GDPR)的数据集,包含使用五种不同的体传导音频传感器录制的音频:两个入耳式麦克风,两个骨传导振动拾音器和一个喉麦克风。该数据集还包括来自用作参考的机载麦克风的音频数据。Vibravox语料库包含由188名参与者在由高阶立体混响3D空间化器施加的不同声学条件下记录的38小时的语音样本和生理声音。语料库中还包括对录音条件和语言转换的注释。我们进行了一系列的实验,对各种语音相关的任务,包括语音识别,语音增强和说话人确认。这些实验是使用最先进的模型进行的,以评估和比较它们对Vibravox数据集提供的不同音频传感器捕获的信号的性能,目的是更好地掌握它们的个体特征。摘要:Vibravox is a dataset compliant with the General Data Protection Regulation (GDPR) containing audio recordings using five different body-conduction audio sensors : two in-ear microphones, two bone conduction vibration pickups and a laryngophone. The data set also includes audio data from an airborne microphone used as a reference. The Vibravox corpus contains 38 hours of speech samples and physiological sounds recorded by 188 participants under different acoustic conditions imposed by an high order ambisonics 3D spatializer. Annotations about the recording conditions and linguistic transcriptions are also included in the corpus. We conducted a series of experiments on various speech-related tasks, including speech recognition, speech enhancement and speaker verification. These experiments were carried out using state-of-the-art models to evaluate and compare their performances on signals captured by the different audio sensors offered by the Vibravox dataset, with the aim of gaining a better grasp of their individual characteristics.
【2】 Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
标题: 具有自我监督音频屏蔽自动编码器的通用声音分离
作者:Junqi Zhao,Xubo Liu,Jinzheng Zhao,Yi Yuan,Qiuqiang Kong,Mark D. Plumbley,Wenwu Wang
链接:点击下载PDF文件
摘要:通用声音分离(USS)是分离任意声源混合的任务。通常,通用分离模型使用标记数据以监督方式从头开始训练。自监督学习(SSL)是一种新兴的深度学习方法,它利用未标记的数据来获得与任务无关的表示,这可以使许多下游任务受益。在本文中,我们提出将一个自监督的预训练模型,即音频掩蔽自动编码器(A-MAE),集成到一个通用的声音分离系统,以提高其分离性能。我们采用两种策略来利用SSL嵌入:在微调期间冻结或更新A-MAE的参数。SSL嵌入与短时傅立叶变换(STFT)级联,作为分离模型的输入特征。我们在AudioSet数据集上对我们的方法进行了评估,实验结果表明,所提出的方法成功地提高了最先进的基于ResUNet的USS模型的分离性能。摘要:Universal sound separation (USS) is a task of separating mixtures of arbitrary sound sources. Typically, universal separation models are trained from scratch in a supervised manner, using labeled data. Self-supervised learning (SSL) is an emerging deep learning approach that leverages unlabeled data to obtain task-agnostic representations, which can benefit many downstream tasks. In this paper, we propose integrating a self-supervised pre-trained model, namely the audio masked autoencoder (A-MAE), into a universal sound separation system to enhance its separation performance. We employ two strategies to utilize SSL embeddings: freezing or updating the parameters of A-MAE during fine-tuning. The SSL embeddings are concatenated with the short-time Fourier transform (STFT) to serve as input features for the separation model. We evaluate our methods on the AudioSet dataset, and the experimental results indicate that the proposed methods successfully enhance the separation performance of a state-of-the-art ResUNet-based USS model.
【3】 MUSA: Multi-lingual Speaker Anonymization via Serial Disentanglement
标题: MUSA:通过序列解纠缠实现多语言说话者匿名化
作者:Jixun Yao,Qing Wang,Pengcheng Guo,Ziqian Ning,Yuguang Yang,Yu Pan,Lei Xie
备注:Submitted to TASLP
链接:点击下载PDF文件
摘要:说话人匿名是一种有效的隐私保护方案,旨在隐藏说话人的身份,同时保留原始语音的语言内容和副语言信息。虽然大多数先前的研究只关注一种语言,但理想的说话人匿名系统应该能够处理多种语言。本文提出了一种多语言说话人自动化方法MUSA,该方法采用串行解纠缠策略来执行从全局时不变表示到时间时变表示的逐步解纠缠。通过使用语义提取和自监督说话人提取,串行解纠缠策略可以避免强归纳偏差,并在不同语言之间表现出优异的泛化性能。与此同时,我们提出了一种简单的匿名化策略,该策略采用零值空嵌入来模拟说话人身份隐藏过程,无需转换为伪说话人身份,从而降低了说话人匿名化过程的复杂性。在VoicePrivacy官方数据集和多语言数据集上的实验结果表明,MUSA可以有效保护说话者隐私,同时保留语言内容和准语言信息。摘要:Speaker anonymization is an effective privacy protection solution designed to conceal the speaker's identity while preserving the linguistic content and para-linguistic information of the original speech. While most prior studies focus solely on a single language, an ideal speaker anonymization system should be capable of handling multiple languages. This paper proposes MUSA, a Multi-lingual Speaker Anonymization approach that employs a serial disentanglement strategy to perform a step-by-step disentanglement from a global time-invariant representation to a temporal time-variant representation. By utilizing semantic distillation and self-supervised speaker distillation, the serial disentanglement strategy can avoid strong inductive biases and exhibit superior generalization performance across different languages. Meanwhile, we propose a straightforward anonymization strategy that employs empty embedding with zero values to simulate the speaker identity concealment process, eliminating the need for conversion to a pseudo-speaker identity and thereby reducing the complexity of speaker anonymization process. Experimental results on VoicePrivacy official datasets and multi-lingual datasets demonstrate that MUSA can effectively protect speaker privacy while preserving linguistic content and para-linguistic information.
【4】 The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
标题: VoicePrivacy 2022挑战:语音匿名化的进展和前景
作者:Michele Panariello,Natalia Tomashenko,Xin Wang,Xiaoxiao Miao,Pierre Champion,Hubert Nourtel,Massimiliano Todisco,Nicholas Evans,Emmanuel Vincent,Junichi Yamagishi
备注:Accepted at IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
摘要:VoicePrivacy Challenge旨在促进语音技术的语音匿名化解决方案的发展。在本文中,我们对2022年举行的第二届会议进行了系统的概述和分析。我们描述了用于系统开发和评估的语音匿名化任务和数据集,提出了用于评估的不同攻击模型,以及相关的客观和主观指标。我们描述了三个匿名化基线,提供了由挑战参与者开发的匿名化系统的摘要描述,并报告了所有人的客观和主观评估结果。此外,我们描述了后评估分析和公开文献中报告的相关工作的总结。结果表明,基于语音转换的解决方案更好地保留了效用,将自动语音识别与合成相结合的替代方案实现了更大的隐私,并且隐私效用权衡仍然是当前匿名化解决方案所固有的。最后,我们提出了我们的想法和优先事项,为未来的语音隐私挑战版本。摘要:The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datasets used for system development and evaluation, present the different attack models used for evaluation, and the associated objective and subjective metrics. We describe three anonymisation baselines, provide a summary description of the anonymisation systems developed by challenge participants, and report objective and subjective evaluation results for all. In addition, we describe post-evaluation analyses and a summary of related work reported in the open literature. Results show that solutions based on voice conversion better preserve utility, that an alternative which combines automatic speech recognition with synthesis achieves greater privacy, and that a privacy-utility trade-off remains inherent to current anonymisation solutions. Finally, we present our ideas and priorities for future VoicePrivacy Challenge editions.
【5】 VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark
标题: VoxBlink 2:10万+说话人识别数据库和开放式说话人识别基准
作者:Yuke Lin,Ming Cheng,Fulin Zhang,Yingying Gao,Shilei Zhang,Ming Li
备注:Accepted By InterSpeech2024
链接:点击下载PDF文件
摘要:在本文中,我们提供了一个大型的视听说话人识别数据集VoxBlink 2,其中包括大约10 M的话语和来自110 K+说话人的视频。该数据集代表了VoxBlink数据集的显着扩展,通过优化的数据收集管道涵盖了更广泛的扬声器和场景。之后,我们探讨了训练策略,数据规模和模型复杂度对说话人验证的影响,并最终在VoxCeleb 1-O测试集上建立了一个新的单模型最先进的EER为0.170%,minDCF为0.006%。这些显著的结果促使我们从一个新的具有挑战性的角度来探索说话人识别。我们提出了开放集扬声器识别任务,其目的是要么匹配一个探测话语与已知的画廊扬声器或将其归类为一个未知的查询。与此相关的任务,我们设计了具体的基准和评估协议。数据和模型资源可以在http: voxblink2.github.io中找到。摘要:In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This dataset represents a significant expansion over the VoxBlink dataset, encompassing a broader diversity of speakers and scenarios by the grace of an optimized data collection pipeline. Afterward, we explore the impact of training strategies, data scale, and model complexity on speaker verification and finally establish a new single-model state-of-the-art EER at 0.170% and minDCF at 0.006% on the VoxCeleb1-O test set. Such remarkable results motivate us to explore speaker recognition from a new challenging perspective. We raise the Open-Set Speaker-Identification task, which is designed to either match a probe utterance with a known gallery speaker or categorize it as an unknown query. Associated with this task, we design concrete benchmark and evaluation protocols. The data and model resources can be found in http: voxblink2.github.io.
【6】 Team HYU ASML ROBOVOX SP Cup 2024 System Description
标题: HyU ASML ROBOVOX SP Cup 2024系统描述
作者:Jeong-Hwan Choi,Gaeun Kim,Hee-Jae Lee,Seyun Ahn,Hyun-Soo Kim,Joon-Hyuk Chang
备注:Technical report for IEEE Signal Processing Cup 2024, 9 pages
链接:点击下载PDF文件
摘要:本报告介绍了HYU ASML团队提交IEEE信号处理杯2024(SP杯2024)。这个挑战,题为“ROBOVOX:远场扬声器识别的移动机器人,”重点是扬声器识别使用移动机器人在嘈杂和混响条件。我们的解决方案结合了深度残差神经网络和基于时延神经网络的说话人嵌入模型的结果。这些模型是在包括法语语音的不同数据集上训练的。为了应对具有高噪声、混响和短语音条件的挑战性评估环境,我们专注于说话人嵌入模型的数据增强和训练语音持续时间。我们提交的作品在SP Cup 2024公开排行榜上获得第二名,检测成本函数为0.5245,等错误率为6.46%。摘要:This report describes the submission of HYU ASML team to the IEEE Signal Processing Cup 2024 (SP Cup 2024). This challenge, titled "ROBOVOX: Far-Field Speaker Recognition by a Mobile Robot," focuses on speaker recognition using a mobile robot in noisy and reverberant conditions. Our solution combines the result of deep residual neural networks and time-delay neural network-based speaker embedding models. These models were trained on a diverse dataset that includes French speech. To account for the challenging evaluation environment characterized by high noise, reverberation, and short speech conditions, we focused on data augmentation and training speech duration for the speaker embedding model. Our submission achieved second place on the SP Cup 2024 public leaderboard, with a detection cost function of 0.5245 and an equal error rate of 6.46%.
【7】 Statistics-aware Audio-visual Deepfake Detector
标题: 统计感知视听Deepfake检测器
作者:Marcella Astrid,Enjie Ghorbel,Djamila Aouada
备注:Accepted in ICIP 2024
链接:点击下载PDF文件
摘要:在本文中,我们提出了一种增强的视听深度检测方法。最近的视听Deepfake检测方法主要评估音频和视觉特征之间的同步。虽然他们已经显示出有前途的结果,他们是基于孤立的特征距离的最大化 最小化,而不考虑特征统计。此外,它们依赖于繁琐的深度学习架构,并且严重依赖于经验固定的超参数。在此,为了克服这些限制,我们提出:(1)统计特征损失以增强模型的区分能力,而不是仅仅依赖于特征距离;(2)使用波形来描述音频作为基于频率的表示的替代;(3)伪造分数的后处理归一化;(4)使用更浅的网络来降低计算复杂度。在DFDC和FakeAVCeleb数据集上的实验证明了该方法的相关性。摘要:In this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results, they are based on the maximization minimization of isolated feature distances without considering feature statistics. Moreover, they rely on cumbersome deep learning architectures and are heavily dependent on empirically fixed hyperparameters. Herein, to overcome these limitations, we propose: (1) a statistical feature loss to enhance the discrimination capability of the model, instead of relying solely on feature distances; (2) using the waveform for describing the audio as a replacement of frequency-based representations; (3) a post-processing normalization of the fakeness score; (4) the use of shallower network for reducing the computational complexity. Experiments on the DFDC and FakeAVCeleb datasets demonstrate the relevance of the proposed method.
【8】 Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
标题: 研究标签布局和训练标准对ASB性能和对齐质量的影响
作者:Tina Raissi,Christoph Lüscher,Simon Berger,Ralf Schlüter,Hermann Ney
备注:Accepted for presentation at Interspeech 2024
链接:点击下载PDF文件
摘要:正在进行的自动语音识别(ASR)研究设想了端到端方法和经典模块化系统之间的明确划分。即使两种方法之间的高层次比较,其要求和(缺点)的优点,通常是解决,在类似的条件下,更密切的比较是不容易在文献中。在这项工作中,我们提出了一个比较集中的标签拓扑结构和训练标准。我们比较了两个歧视性的对齐模型与隐马尔可夫模型(HMM)和连接主义的时间分类拓扑结构,和两个一阶标签上下文ASR模型,分别利用因子HMM和严格单调递归神经网络换能器。我们使用不同的测量来评估对齐质量,并比较我们最好的系统的字错误率和实时因素。在LibriSpeech 960h和Switchboard 300h任务上进行了实验。摘要:The ongoing research scenario for automatic speech recognition (ASR) envisions a clear division between end-to-end approaches and classic modular systems. Even though a high-level comparison between the two approaches in terms of their requirements and (dis)advantages is commonly addressed, a closer comparison under similar conditions is not readily available in the literature. In this work, we present a comparison focused on the label topology and training criterion. We compare two discriminative alignment models with hidden Markov model (HMM) and connectionist temporal classification topology, and two first-order label context ASR models utilizing factored HMM and strictly monotonic recurrent neural network transducer, respectively. We use different measurements for the evaluation of the alignment quality, and compare word error rate and real time factor of our best systems. Experiments are conducted on the LibriSpeech 960h and Switchboard 300h tasks.
【9】 MMSD-Net: Towards Multi-modal Stuttering Detection
标题: MMSD-Net:走向多模式口吃检测
作者:Liangyu Nie,Sudarsana Reddy Kadiri,Ruchit Agrawal
备注:Accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:口吃是一种常见的语言障碍,由言语产生的不规则中断引起,影响着全世界7000多万人。标准的自动语音处理工具不考虑语音疾病,因此当以口吃语音作为输入时不能生成有意义的结果。口吃的自动检测是构建高效的、上下文感知的语音处理系统的重要一步。虽然以前的方法探索了统计和神经方法来检测口吃,但所有这些方法本质上都是单峰的。本文介绍了MMSD-Net,第一个多模态神经网络框架口吃检测。实验和结果表明,将视觉信号显着辅助口吃检测,我们的模型产生了2-17%的改善,在F1分数比现有的国家的最先进的单峰方法。摘要:Stuttering is a common speech impediment that is caused by irregular disruptions in speech production, affecting over 70 million people across the world. Standard automatic speech processing tools do not take speech ailments into account and are thereby not able to generate meaningful results when presented with stuttered speech as input. The automatic detection of stuttering is an integral step towards building efficient, context-aware speech processing systems. While previous approaches explore both statistical and neural approaches for stuttering detection, all of these methods are uni-modal in nature. This paper presents MMSD-Net, the first multi-modal neural framework for stuttering detection. Experiments and results demonstrate that incorporating the visual signal significantly aids stuttering detection, and our model yields an improvement of 2-17% in the F1-score over existing state-of-the-art uni-modal approaches.
【10】 A Pilot Study of GSLM-based Simulation of Foreign Accentuation Only Using Native Speech Corpora
标题: 基于GLM仅使用母语语音库模拟外国口音的初步研究
作者:Kentaro Onda,Joonyong Park,Nobuaki Minematsu,Daisuke Saito
备注:Accepted to INTERSPEECH2024
链接:点击下载PDF文件
摘要:本文提出了一种基于生成式口语模型(GSLM)的外语重读模拟方法。当一个人听一种外语的口语单词并重复它们时,重复的语音通常带有该听者的L1的口音。据说这是因为说出来的单词在心理上被表征为L1的一系列语音单位,而这些单位用于口头再现。我们模拟这个过程中输入语言A的语音到GSLM的语言B添加B的口音到输入语音。针对外国输入语音运行L1的ASR并将ASR结果给予L1的TTS的过程可以被视为该方法的简单实现。我们的实验结果表明,合成的口音的输出语音是非常自然的,相比,真实的样本A产生的扬声器的L1是B,和重音的程度是可控的。摘要:We propose a method of simulating the human process of foreign accentuation using Generative Spoken Language Model (GSLM) only with native speech corpora. When one listens to spoken words of a foreign language and repeats them, the repeated speech is often with the accent of that listener's L1. This is said to be because the spoken words are mentally represented as a sequence of phonological units of the L1, and those units are used for oral reproduction. We simulate this process by inputting speech of language A into GSLM of language B to add B's accent onto the input speech. The process of running ASR of the L1 for foreign input speech and giving the ASR result to TTS of the L1 can be viewed as a naive implementation of this approach. The results of our experiments show that the synthesized accent of the output speech is highly natural, compared to real samples of A generated by speakers whose L1 is B, and that the degree of accentuation is controllable.
【11】 Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
标题: 超越二进制:使用生成性预训练变形器和端到端模型进行多类失语症检测
作者:Matthew Perez,Aneesha Sampath,Minxue Niu,Emily Mower Provost
链接:点击下载PDF文件
摘要:失语症是一种语言障碍,可导致言语错误,称为paraphasias,涉及误用,替换或发明的话。自动错语检测可以通过促进临床评估和治疗计划选择来帮助失语症患者。然而,大多数自动错语检测工作只集中在二进制检测,这涉及到识别的存在或不存在的错语。多类错语检测代表了一个未探索的研究领域,其重点是识别多种类型的错语以及它们在给定的语音片段中发生的位置。我们提出了新的方法,使用生成预训练的Transformer(GPT),以识别从成绩单以及两个端到端的方法,重点是建模自动语音识别(ASR)和失语症分类为多个序列与单一序列。我们证明了一个单一的序列模型优于GPT基线多类错语检测。摘要:Aphasia is a language disorder that can lead to speech errors known as paraphasias, which involve the misuse, substitution, or invention of words. Automatic paraphasia detection can help those with Aphasia by facilitating clinical assessment and treatment planning options. However, most automatic paraphasia detection works have focused solely on binary detection, which involves recognizing only the presence or absence of a paraphasia. Multiclass paraphasia detection represents an unexplored area of research that focuses on identifying multiple types of paraphasias and where they occur in a given speech segment. We present novel approaches that use a generative pretrained transformer (GPT) to identify paraphasias from transcripts as well as two end-to-end approaches that focus on modeling both automatic speech recognition (ASR) and paraphasia classification as multiple sequences vs. a single sequence. We demonstrate that a single sequence model outperforms GPT baselines for multiclass paraphasia detection.
【12】 Disentangled Acoustic Fields For Multimodal Physical Scene Understanding
标题: 用于多模式物理场景理解的解开声学场
作者:Jie Yin,Andrew Luo,Yilun Du,Anoop Cherian,Tim K. Marks,Jonathan Le Roux,Chuang Gan
链接:点击下载PDF文件
摘要:我们研究了多模态物理场景理解的问题,其中一个具体的代理需要通过推断对象属性,方向和距离的影响声源找到倒下的物体。以前的工作采用前馈神经网络直接从声音回归变量,导致泛化能力差和域适应问题。在本文中,我们说明,学习的解纠缠模型的声学形成,称为解纠缠声场(ESTA),捕捉声音的产生和传播过程中,使体现代理构建一个空间的不确定性地图的对象可能已经下降。我们证明,我们的分析合成框架可以联合推断声音属性明确分解和分解的潜在空间的解开模型。我们进一步表明,空间不确定性地图可以显着提高坠落物体的定位成功率,提出多个合理的探索位置。摘要:We study the problem of multimodal physical scene understanding, where an embodied agent needs to find fallen objects by inferring object properties, direction, and distance of an impact sound source. Previous works adopt feed-forward neural networks to directly regress the variables from sound, leading to poor generalization and domain adaptation issues. In this paper, we illustrate that learning a disentangled model of acoustic formation, referred to as disentangled acoustic field (DAF), to capture the sound generation and propagation process, enables the embodied agent to construct a spatial uncertainty map over where the objects may have fallen. We demonstrate that our analysis-by-synthesis framework can jointly infer sound properties by explicitly decomposing and factorizing the latent space of the disentangled model. We further show that the spatial uncertainty map can significantly improve the success rate for the localization of fallen objects by proposing multiple plausible exploration locations.
【13】 Target conversation extraction: Source separation using turn-taking dynamics
标题: 目标对话提取:使用轮流动力学进行源分离
作者:Tuochao Chen,Qirui Wang,Bohan Wu,Malek Itani,Emre Sefik Eskimez,Takuya Yoshioka,Shyamnath Gollakota
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:在干扰扬声器和噪声中提取对话中参与者的语音提出了一个具有挑战性的问题。在本文中,我们介绍了新的任务目标会话提取,其目标是提取音频的目标会话的基础上,其参与者之一的说话人嵌入。为了实现这一目标,我们建议利用人类对话中固有的时间模式,特别是话轮转换动态,它独特地表征了参与对话的扬声器,并将其与干扰扬声器和噪声区分开来。使用神经网络,我们展示了我们的方法在英语和普通话会话数据集上的可行性。在干扰扬声器的存在下,我们的研究结果表明,8.19 dB的改善,信噪比为2扬声器的对话和7.92 dB的改善2-4扬声器的对话。代码,数据集可在https: github.com chentuochao Target-Conversation-Extraction上获得。摘要:Extracting the speech of participants in a conversation amidst interfering speakers and noise presents a challenging problem. In this paper, we introduce the novel task of target conversation extraction, where the goal is to extract the audio of a target conversation based on the speaker embedding of one of its participants. To accomplish this, we propose leveraging temporal patterns inherent in human conversations, particularly turn-taking dynamics, which uniquely characterize speakers engaged in conversation and distinguish them from interfering speakers and noise. Using neural networks, we show the feasibility of our approach on English and Mandarin conversation datasets. In the presence of interfering speakers, our results show an 8.19 dB improvement in signal-to-noise ratio for 2-speaker conversations and a 7.92 dB improvement for 2-4-speaker conversations. Code, dataset available at https: github.com chentuochao Target-Conversation-Extraction.
【14】 Knowledge boosting during low-latency inference
标题: 低延迟推理期间的知识提升
作者:Vidya Srinivas,Malek Itani,Tuochao Chen,Emre Sefik Eskimez,Takuya Yoshioka,Shyamnath Gollakota
备注:Accepted by Interspeech 2024
链接:点击下载PDF文件
摘要:None摘要:Models for low-latency, streaming applications could benefit from the knowledge capacity of larger models, but edge devices cannot run these models due to resource constraints. A possible solution is to transfer hints during inference from a large model running remotely to a small model running on-device. However, this incurs a communication delay that breaks real-time requirements and does not guarantee that both models will operate on the same data at the same time. We propose knowledge boosting, a novel technique that allows a large model to operate on time-delayed input during inference, while still boosting small model performance. Using a streaming neural network that processes 8 ms chunks, we evaluate different speech separation and enhancement tasks with communication delays of up to six chunks or 48 ms. Our results show larger gains where the performance gap between the small and large models is wide, demonstrating a promising method for large-small model collaboration for low-latency applications. Code, dataset, and audio samples available at https: knowledgeboosting.cs.washington.edu .
【15】 Exploring Gender-Specific Speech Patterns in Automatic Suicide Risk Assessment
标题: 在自动自杀风险评估中探索特定性别的言语模式
作者:Maurice Gerczuk,Shahin Amiriparian,Justina Lutz,Wolfgang Strube,Irina Papazova,Alkomiet Hasan,Björn W. Schuller
备注:accepted at INTERSPEECH 2024
链接:点击下载PDF文件
摘要:在急诊医学中,对有自杀风险的患者的及时干预往往受到延迟获得专门精神病护理的阻碍。为了弥补这一差距,我们引入了一种基于语音的自动自杀风险评估方法。我们的研究涉及一个新的数据集,包括20名阅读中性文本的患者的语音记录。我们提取了四个语音表示,包括可解释的和深层的功能。此外,我们探讨了基于性别的建模和短语水平的正常化的影响。通过应用性别排斥建模,从情绪微调wav2vec2.0模型中提取的特征可以用于区分高自杀风险和低自杀风险,平衡准确率为81%。最后,我们的分析揭示了性别之间的言语特征和自杀风险的关系的差异。在我们的数据集中,男性的自杀风险随着情绪激动而增加,而女性受试者的声音特征则相反。摘要:In emergency medicine, timely intervention for patients at risk of suicide is often hindered by delayed access to specialised psychiatric care. To bridge this gap, we introduce a speech-based approach for automatic suicide risk assessment. Our study involves a novel dataset comprising speech recordings of 20 patients who read neutral texts. We extract four speech representations encompassing interpretable and deep features. Further, we explore the impact of gender-based modelling and phrase-level normalisation. By applying gender-exclusive modelling, features extracted from an emotion fine-tuned wav2vec2.0 model can be utilised to discriminate high- from low- suicide risk with a balanced accuracy of 81%. Finally, our analysis reveals a discrepancy in the relationship of speech characteristics and suicide risk between female and male subjects. For men in our dataset, suicide risk increases together with agitation while voice characteristics of female subjects point the other way.
【16】 Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation
标题: 级联流语音翻译中MT束搜索雷区的导航
作者:Rastislav Rabatin,Frank Seide,Ernie Chang
链接:点击下载PDF文件
摘要:我们适应著名的波束搜索算法的机器翻译操作的级联实时语音翻译系统。由于四个关键挑战,这被证明比最初预期的更复杂:(1)实时处理来自ASR的不完整单词的中间和最终翻译,(2)以最小的用户感知延迟发出中间和最终翻译,(3)处理具有不等长度和不同模型状态的波束搜索假设,以及(4)处理句子边界。在同步机器翻译领域中的先前工作仅实现贪婪解码。我们提出了一个波束搜索实现,处理所有上述问题,通过挑战雷区提供指导。与贪婪搜索相比,我们的方法将BLEU分数提高了1分,与重复重新翻译输入的基线启发式相比,将CPU时间减少了40%,字符闪烁率减少了20+%。摘要:We adapt the well-known beam-search algorithm for machine translation to operate in a cascaded real-time speech translation system. This proved to be more complex than initially anticipated, due to four key challenges: (1) real-time processing of intermediate and final transcriptions with incomplete words from ASR, (2) emitting intermediate and final translations with minimal user perceived latency, (3) handling beam search hypotheses that have unequal length and different model state, and (4) handling sentence boundaries. Previous work in the field of simultaneous machine translation only implemented greedy decoding. We present a beam-search realization that handles all of the above, providing guidance through the minefield of challenges. Our approach increases the BLEU score by 1 point compared to greedy search, reduces the CPU time by up to 40% and character flicker rate by 20+% compared to a baseline heuristic that just retranslates input repeatedly.
机器翻译,仅供参考
