本文经arXiv每日学术速递授权转载
【1】 A Framework for Multimodal Medical Image Interaction
标题: 多模式医学图像交互框架
作者:Laura Schütz,Sasan Matinfar,Gideon Schafroth,Navid Navab,Merle Fairhurst,Arthur Wagner,Benedikt Wiestler,Ulrich Eck,Nassir Navab
备注:Accepted for publication in IEEE TVCG; presentation at IEEE ISMAR 2024
链接:点击下载PDF文件
【2】 Audio-Language Datasets of Scenes and Events: A Survey
标题: 场景和事件的音频语言数据集:调查
作者:Gijs Wijngaard,Elia Formisano,Michele Esposito,Michel Dumontier
链接:点击下载PDF文件
【3】 RespEar: Earable-Based Robust Respiratory Rate Monitoring
标题: RespEar:基于Early的稳健呼吸率监测
作者:Yang Liu,Kayla-Jade Butkow,Jake Stuchbury-Wass,Adam Pullin,Dong Ma,Cecilia Mascolo
链接:点击下载PDF文件
【4】 Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
标题: 通过将通道间和频段特征与双分支一致器集成来改善语音增强
作者:Jizhen Li,Xinmeng Xu,Weiping Tu,Yuhong Yang,Rong Zhu
链接:点击下载PDF文件
【5】 Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
标题: 用于动态合成障碍和老年说话者适应的同质说话者功能
作者:Mengzhe Geng,Xurong Xie,Jiajun Deng,Zengrui Jin,Guinan Li,Tianzi Wang,Shujie Hu,Zhaoqing Li,Helen Meng,Xunying Liu
备注:In submission to IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
【6】 Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
标题: DS@GT BirdCREF 2024的伪多标签鸟鸣分类的迁移学习
作者:Anthony Miyaguchi,Adrian Cheung,Murilo Gustineli,Ashley Kim
备注:Submitted and accepted into CLEF 2024 CEUR-WS proceedings
链接:点击下载PDF文件
【7】 Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
标题: 复和实球调和高斯系数及其在球阵处理和立体声合成中的应用
作者:Archontis Politis
链接:点击下载PDF文件
【8】 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
标题: 学习并不要忘记:向ASB基础模型添加新语言
作者:Mengjie Qian,Siyuan Tang,Rao Ma,Kate M. Knill,Mark J. F. Gales
链接:点击下载PDF文件
标题: 公平地听和说:集成大型语言模型的言语中语义性别偏见的研究
作者:Yi-Cheng Lin,Tzu-Quan Lin,Chih-Kai Yang,Ke-Han Lu,Wei-Chih Chen,Chun-Yi Kuan,Hung-yi Lee
链接:点击下载PDF文件
【2】 Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
标题: 复和实球调和高斯系数及其在球阵处理和立体声合成中的应用
作者:Archontis Politis
链接:点击下载PDF文件
【3】 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
标题: 学习并不要忘记:向ASB基础模型添加新语言
作者:Mengjie Qian,Siyuan Tang,Rao Ma,Kate M. Knill,Mark J. F. Gales
链接:点击下载PDF文件
【4】 XANE Background Acoustic Embeddings: Ablation and Clustering Analysis
标题: XANE背景声学嵌入:消融和聚集分析
作者:Dushyant Sharma,James Fosburgh,Sri Harsha Dumpala,Chandramouli Shama Sastri,Stanislav Yu. Kruchinin,Patrick A. Naylor
备注:arXiv admin note: substantial text overlap with arXiv:2406.05199
链接:点击下载PDF文件
【5】 A Framework for Multimodal Medical Image Interaction
标题: 多模式医学图像交互框架
作者:Laura Schütz,Sasan Matinfar,Gideon Schafroth,Navid Navab,Merle Fairhurst,Arthur Wagner,Benedikt Wiestler,Ulrich Eck,Nassir Navab
备注:Accepted for publication in IEEE TVCG; presentation at IEEE ISMAR 2024
链接:点击下载PDF文件
【6】 Audio-Language Datasets of Scenes and Events: A Survey
标题: 场景和事件的音频语言数据集:调查
作者:Gijs Wijngaard,Elia Formisano,Michele Esposito,Michel Dumontier
链接:点击下载PDF文件
【7】 RespEar: Earable-Based Robust Respiratory Rate Monitoring
标题: RespEar:基于Early的稳健呼吸率监测
作者:Yang Liu,Kayla-Jade Butkow,Jake Stuchbury-Wass,Adam Pullin,Dong Ma,Cecilia Mascolo
链接:点击下载PDF文件
【8】 Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
标题: 通过将通道间和频段特征与双分支一致器集成来改善语音增强
作者:Jizhen Li,Xinmeng Xu,Weiping Tu,Yuhong Yang,Rong Zhu
链接:点击下载PDF文件
【9】 Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
标题: 用于动态合成障碍和老年说话者适应的同质说话者功能
作者:Mengzhe Geng,Xurong Xie,Jiajun Deng,Zengrui Jin,Guinan Li,Tianzi Wang,Shujie Hu,Zhaoqing Li,Helen Meng,Xunying Liu
备注:In submission to IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
【10】 Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
标题: DS@GT BirdCREF 2024的伪多标签鸟鸣分类的迁移学习
作者:Anthony Miyaguchi,Adrian Cheung,Murilo Gustineli,Ashley Kim
备注:Submitted and accepted into CLEF 2024 CEUR-WS proceedings
链接:点击下载PDF文件
标题: 多模式医学图像交互框架
作者:Laura Schütz,Sasan Matinfar,Gideon Schafroth,Navid Navab,Merle Fairhurst,Arthur Wagner,Benedikt Wiestler,Ulrich Eck,Nassir Navab
备注:Accepted for publication in IEEE TVCG; presentation at IEEE ISMAR 2024
链接:点击下载PDF文件
摘要:医生依靠人体解剖学的图像,例如磁共振成像(MRI),在诊断和治疗期间定位患者的感兴趣区域。尽管医学成像技术取得了进步,但信息传递仍然是单峰的。这种视觉表现未能捕捉到与人体组织的真实、多感官互动的复杂性。然而,实时感知关于患者解剖结构和疾病的多模态信息对于医疗程序和患者结果的成功至关重要。我们引入了一个多模态医学图像交互(MMII)框架,让医学专家在三维空间中与人体组织进行动态的视听交互。在虚拟现实环境中,用户接收物理通知的视听反馈以改善解剖结构的空间感知。MMII使用基于模型的声音处理方法来产生来自组织的几何形状和物理特性的声音,从而消除了手工制作声音设计的需要。两个用户研究,涉及34个一般和9个临床专家进行了评估建议的互动框架的可学习性,可用性和准确性。我们的研究结果显示,视听对应的可学习性非常好,因为在研究过程中,正确联想的比率显著提高(p < 0.001)。与传统医学图像交互相比,MMII的脑肿瘤定位准确性更高(p < 0.05)。我们的研究结果证实了这种新框架的潜力,以加强与医学图像的互动,例如,在手术过程中,需要立即和精确的反馈。摘要:Medical doctors rely on images of the human anatomy, such as magnetic resonance imaging (MRI), to localize regions of interest in the patient during diagnosis and treatment. Despite advances in medical imaging technology, the information conveyance remains unimodal. This visual representation fails to capture the complexity of the real, multisensory interaction with human tissue. However, perceiving multimodal information about the patient's anatomy and disease in real-time is critical for the success of medical procedures and patient outcome. We introduce a Multimodal Medical Image Interaction (MMII) framework to allow medical experts a dynamic, audiovisual interaction with human tissue in three-dimensional space. In a virtual reality environment, the user receives physically informed audiovisual feedback to improve the spatial perception of anatomical structures. MMII uses a model-based sonification approach to generate sounds derived from the geometry and physical properties of tissue, thereby eliminating the need for hand-crafted sound design. Two user studies involving 34 general and nine clinical experts were conducted to evaluate the proposed interaction framework's learnability, usability, and accuracy. Our results showed excellent learnability of audiovisual correspondence as the rate of correct associations significantly improved (p < 0.001) over the course of the study. MMII resulted in superior brain tumor localization accuracy (p < 0.05) compared to conventional medical image interaction. Our findings substantiate the potential of this novel framework to enhance interaction with medical images, for example, during surgical procedures where immediate and precise feedback is needed.
【2】 Audio-Language Datasets of Scenes and Events: A Survey
标题: 场景和事件的音频语言数据集:调查
作者:Gijs Wijngaard,Elia Formisano,Michele Esposito,Michel Dumontier
链接:点击下载PDF文件
摘要:音频语言模型(ALM)处理声音,以提供对产生声音的事件和场景的语言描述。计算能力和数据集创建的最新进展导致了这一领域的重大进展。本文调查了用于训练音频语言模型的现有数据集,强调了最近使用大型,多样化数据集来增强模型性能的趋势。这些数据集的主要来源包括Freesound平台和AudioSet,它们为该领域的快速增长做出了贡献。虽然以前的调查主要涉及技术和培训细节,但本调查对各种数据集进行了分类和评估,并讨论了它们的起源,特征和用例。它还执行数据泄漏分析,以确保数据集的完整性并减轻数据集之间的偏差。本次调查是通过分析截至2023年12月(含)的研究论文进行的,不包含该时期之后的任何论文。摘要:Audio-language models (ALMs) process sounds to provide a linguistic description of sound-producing events and scenes. Recent advances in computing power and dataset creation have led to significant progress in this domain. This paper surveys existing datasets used for training audio-language models, emphasizing the recent trend towards using large, diverse datasets to enhance model performance. Key sources of these datasets include the Freesound platform and AudioSet that have contributed to the field's rapid growth. Although prior surveys primarily address techniques and training details, this survey categorizes and evaluates a wide array of datasets, addressing their origins, characteristics, and use cases. It also performs a data leak analysis to ensure dataset integrity and mitigate bias between datasets. This survey was conducted by analyzing research papers up to and including December 2023, and does not contain any papers after that period.
【3】 RespEar: Earable-Based Robust Respiratory Rate Monitoring
标题: RespEar:基于Early的稳健呼吸率监测
作者:Yang Liu,Kayla-Jade Butkow,Jake Stuchbury-Wass,Adam Pullin,Dong Ma,Cecilia Mascolo
链接:点击下载PDF文件
摘要:呼吸率(RR)监测对于了解身心健康和跟踪健康状况是不可或缺的。现有研究已经证明了在特定用户条件下(例如,同时保持静止,或者同时沉重地呼吸)。然而,在不同的日常生活和活动中进行准确,连续和非侵入性的RR监测仍然具有挑战性。在这项工作中,我们提出了RespEar,一个基于earable的系统,强大的RR监测。通过利用耳塞中入耳式麦克风的独特特性,RespEar能够使用呼吸性窦性心律失常(RSA)和运动性呼吸耦合(LRC),心血管活动,步态和呼吸之间的生理耦合,间接确定RR。这有效地解决了日常活动中几乎无法察觉的呼吸信号所带来的挑战。我们进一步提出了一套精心制作的信号处理方案,以提高RR估计的准确性和鲁棒性。通过从18名受试者的8项活动中收集的数据,RespEar测量RR,在久坐状态下的平均绝对误差(MAE)为1.48次呼吸 分钟(BPM),平均绝对百分比误差(MAPE)为9.12%,在活动状态下的MAE为2.28 BPM,MAPE为11.04%,这对于能够用单一模态概括各种条件的方法来说是前所未有的。摘要:Respiratory rate (RR) monitoring is integral to understanding physical and mental health and tracking fitness. Existing studies have demonstrated the feasibility of RR monitoring under specific user conditions (e.g., while remaining still, or while breathing heavily). Yet, performing accurate, continuous and non-obtrusive RR monitoring across diverse daily routines and activities remains challenging. In this work, we present RespEar, an earable-based system for robust RR monitoring. By leveraging the unique properties of in-ear microphones in earbuds, RespEar enables the use of Respiratory Sinus Arrhythmia (RSA) and Locomotor Respiratory Coupling (LRC), physiological couplings between cardiovascular activity, gait and respiration, to indirectly determine RR. This effectively addresses the challenges posed by the almost imperceptible breathing signals under daily activities. We further propose a suite of meticulously crafted signal processing schemes to improve RR estimation accuracy and robustness. With data collected from 18 subjects over 8 activities, RespEar measures RR with a mean absolute error (MAE) of 1.48 breaths per minutes (BPM) and a mean absolute percent error (MAPE) of 9.12% in sedentary conditions, and a MAE of 2.28 BPM and a MAPE of 11.04% in active conditions, respectively, which is unprecedented for a method capable of generalizing across conditions with a single modality.
【4】 Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
标题: 通过将通道间和频段特征与双分支一致器集成来改善语音增强
作者:Jizhen Li,Xinmeng Xu,Weiping Tu,Yuhong Yang,Rong Zhu
链接:点击下载PDF文件
摘要:基于卷积神经网络(CNN)和Transformer的语音增强方法能够有效地提取语音信号的时频信息。然而,语音特征的各个通道之间的相关性却没有得到很好的探索。理论上,由不同卷积核获得的语音特征的每个通道图包含具有不同尺度的信息,表现出强相关性。为了填补这一空白,我们提出了一种新的双分支架构命名为通道感知双分支构象(CADB-Conformer),有效地探索不同的通道之间的长范围的时间和频率的相关性,分别提取信道关系感知的时频信息。在DNS挑战2020数据集上进行的消融研究证明了通道特征利用的重要性,同时显示了通道关系感知T-F信息对语音增强的重要性。大量的实验也表明,该模型实现了优越的性能比最近的方法具有吸引力的计算成本。摘要:Recent speech enhancement methods based on convolutional neural networks (CNNs) and transformer have been demonstrated to efficaciously capture time-frequency (T-F) information on spectrogram. However, the correlation of each channels of speech features is failed to explore. Theoretically, each channel map of speech features obtained by different convolution kernels contains information with different scales demonstrating strong correlations. To fill this gap, we propose a novel dual-branch architecture named channel-aware dual-branch conformer (CADB-Conformer), which effectively explores the long range time and frequency correlations among different channels, respectively, to extract channel relation aware time-frequency information. Ablation studies conducted on DNS-Challenge 2020 dataset demonstrate the importance of channel feature leveraging while showing the significance of channel relation aware T-F information for speech enhancement. Extensive experiments also show that the proposed model achieves superior performance than recent methods with an attractive computational costs.
【5】 Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
标题: 用于动态合成障碍和老年说话者适应的同质说话者功能
作者:Mengzhe Geng,Xurong Xie,Jiajun Deng,Zengrui Jin,Guinan Li,Tianzi Wang,Shujie Hu,Zhaoqing Li,Helen Meng,Xunying Liu
备注:In submission to IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
摘要:将数据密集型自动语音识别(ASR)技术应用于构音障碍和老年成人语音,面临着与健康和非老年人语音的不匹配、数据稀缺和说话者水平的大变异性。为此,本文提出了两种新的数据高效的方法来学习同质构音障碍和老年人说话者级别的功能,用于DNN TDNN和Conformer ASR模型的快速,动态测试时间适应。这些措施包括:1)说话人级方差正则化谱基嵌入(VR-SBE)特征,其利用特殊的正则化项来加强自适应中说话人特征的同质性;以及2)基于特征的学习隐藏单元贡献(f-LHUC)变换,其以VR-SBE特征为条件。实验在两种语言的四个任务上进行:英语UASpeech和TORGO构音障碍语音数据集,英语DementiaBank Pitt和粤语JCCOCC MoCA老年人语音语料库。所提出的动态扬声器自适应技术始终优于基线iVector和xVector自适应,统计上显着的单词或字符错误率降低了5.32%的绝对值(相对值为18.57%),批处理模式LHUC扬声器自适应降低了2.24%的绝对值(相对值为9.20%),同时在自适应期间使用实时因素加速到xVector的33.6倍。所提出的适应技术的有效性在与当前ASR技术的比较中得到了证明,包括UASpeech上的SSL预训练系统,其中我们最好的系统产生了23.33%的最先进的WER。分析表明,VR-SBE特征和f-LHUC变换在测试时自适应中对说话人级数据量不敏感。T-SNE可视化显示,它们比基线iVectors,xVectors和批处理模式LHUC变换具有更强的说话者级别均匀性。摘要:The application of data-intensive automatic speech recognition (ASR) technologies to dysarthric and elderly adult speech is confronted by their mismatch against healthy and nonaged voices, data scarcity and large speaker-level variability. To this end, this paper proposes two novel data-efficient methods to learn homogeneous dysarthric and elderly speaker-level features for rapid, on-the-fly test-time adaptation of DNN TDNN and Conformer ASR models. These include: 1) speaker-level variance-regularized spectral basis embedding (VR-SBE) features that exploit a special regularization term to enforce homogeneity of speaker features in adaptation; and 2) feature-based learning hidden unit contributions (f-LHUC) transforms that are conditioned on VR-SBE features. Experiments are conducted on four tasks across two languages: the English UASpeech and TORGO dysarthric speech datasets, the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech corpora. The proposed on-the-fly speaker adaptation techniques consistently outperform baseline iVector and xVector adaptation by statistically significant word or character error rate reductions up to 5.32% absolute (18.57% relative) and batch-mode LHUC speaker adaptation by 2.24% absolute (9.20% relative), while operating with real-time factors speeding up to 33.6 times against xVectors during adaptation. The efficacy of the proposed adaptation techniques is demonstrated in a comparison against current ASR technologies including SSL pre-trained systems on UASpeech, where our best system produces a state-of-the-art WER of 23.33%. Analyses show VR-SBE features and f-LHUC transforms are insensitive to speaker-level data quantity in testtime adaptation. T-SNE visualization reveals they have stronger speaker-level homogeneity than baseline iVectors, xVectors and batch-mode LHUC transforms.
【6】 Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
标题: DS@GT BirdCREF 2024的伪多标签鸟鸣分类的迁移学习
作者:Anthony Miyaguchi,Adrian Cheung,Murilo Gustineli,Ashley Kim
备注:Submitted and accepted into CLEF 2024 CEUR-WS proceedings
链接:点击下载PDF文件
摘要:我们为DS@GT团队提供了关于迁移学习的工作笔记,其中包含BirdCLEF 2024竞赛的伪多标签鸟鸣分类,重点是在录制的音景中识别印度鸟类。我们的方法利用生产级模型,如Google Bird Vocalization Classifier,BirdNET和EnCodec,以解决竞争中的表示和标签挑战。我们探讨了今年版本的未标记的音景代表隐藏的测试集之间的分布变化,并提出了一个伪多标签分类策略,以利用未标记的数据。我们的最高赛后公开排行榜分数是0.63,使用BirdNET嵌入与鸟发声伪标签。我们的代码可在https: github.com dsgt-kaggle-clef birdclef-2024上获得摘要:We present working notes for the DS@GT team on transfer learning with pseudo multi-label birdcall classification for the BirdCLEF 2024 competition, focused on identifying Indian bird species in recorded soundscapes. Our approach utilizes production-grade models such as the Google Bird Vocalization Classifier, BirdNET, and EnCodec to address representation and labeling challenges in the competition. We explore the distributional shift between this year's edition of unlabeled soundscapes representative of the hidden test set and propose a pseudo multi-label classification strategy to leverage the unlabeled data. Our highest post-competition public leaderboard score is 0.63 using BirdNET embeddings with Bird Vocalization pseudo-labels. Our code is available at https: github.com dsgt-kaggle-clef birdclef-2024
【7】 Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
标题: 复和实球调和高斯系数及其在球阵处理和立体声合成中的应用
作者:Archontis Politis
链接:点击下载PDF文件
摘要:声场的方向表示的声学信号处理,包括源,接收器和散射体传递函数,通常在球谐域(SHD)中表示和建模。某些这样的建模操作或那些模型的应用涉及那些方向量的乘法,这些方向量也可以通过称为冈特系数的耦合系数在SHD中方便地表示。由于Gaunt系数的定义和符号在声学出版物中有所不同,因此这项工作基于复和实球谐函数(SH)的既定惯例以及用于定向带限球函数的球乘的方便矩阵形式来定义它们。此外,该报告还提供了真实SH的Gaunt系数的推导,这在文献中是缺失的,可以直接用于空间音频框架,如高保真度立体声。提供了Matlab代码,可以计算用户指定的SH阶的所有系数。最后,一些相关的声学处理的例子,从文献中,下面的矩阵形式主义的系数在报告中介绍。摘要:Acoustical signal processing of directional representations of sound fields, including source, receiver, and scatterer transfer functions, are often expressed and modeled in the spherical harmonic domain (SHD). Certain such modeling operations, or applications of those models, involve multiplications of those directional quantities, which can also be expressed conveniently in the SHD through coupling coefficients known as Gaunt coefficients. Since the definition and notation of Gaunt coefficients varies across acoustical publications, this work defines them based on established conventions of complex and real spherical harmonics (SHs) along with a convenient matrix form for spherical multiplication of directionally band-limited spherical functions. Additionally, the report provides a derivation of the Gaunt coefficients for real SHs, which has been missing from the literature and can be used directly in spatial audio frameworks such as Ambisonics. Matlab code is provided that can compute all coefficients up to user specified SH orders. Finally, a number of relevant acoustical processing examples from the literature are presented, following the matrix formalism of coefficients introduced in the report.
【8】 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
标题: 学习并不要忘记:向ASB基础模型添加新语言
作者:Mengjie Qian,Siyuan Tang,Rao Ma,Kate M. Knill,Mark J. F. Gales
链接:点击下载PDF文件
摘要:基础ASR模型通常支持多种语言,例如Whisper中的100种语言。然而,在集成一种额外的、通常资源较少的语言,同时保持原始语言集的性能方面的工作有限。微调虽然简单,但可能会降低原始集合的准确性。我们比较了三种方法,利用自适应参数:软语言代码调整,训练只有语言代码;软提示调整,训练前置令牌;和LoRA的一小部分额外的参数进行了优化。弹性权重合并(Elastic Weight Consolidation,EWC)提供了另一种折衷方案,有可能保持特定目标语言的性能。结果表明,直接微调产生的新语言的最佳性能,但降低现有的语言能力。EWC可以针对特定语言解决此问题。如果仅使用自适应参数,则保持语言能力,但以新语言的性能为代价。摘要:Foundation ASR models often support many languages, e.g. 100 languages in Whisper. However, there has been limited work on integrating an additional, typically low-resource, language, while maintaining performance on the original language set. Fine-tuning, while simple, may degrade the accuracy of the original set. We compare three approaches that exploit adaptation parameters: soft language code tuning, train only the language code; soft prompt tuning, train prepended tokens; and LoRA where a small set of additional parameters are optimised. Elastic Weight Consolidation (EWC) offers an alternative compromise with the potential to maintain performance in specific target languages. Results show that direct fine-tuning yields the best performance for the new language but degrades existing language capabilities. EWC can address this issue for specific languages. If only adaptation parameters are used, the language capabilities are maintained but at the cost of performance in the new language.
eess.AS音频处理
【1】 Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models标题: 公平地听和说:集成大型语言模型的言语中语义性别偏见的研究
作者:Yi-Cheng Lin,Tzu-Quan Lin,Chih-Kai Yang,Ke-Han Lu,Wei-Chih Chen,Chun-Yi Kuan,Hung-yi Lee
链接:点击下载PDF文件
摘要:语音集成大语言模型(SILLM)将大语言模型与语音感知相结合,以执行各种任务,例如情感识别到说话人验证,展示了通用的音频理解能力。然而,这些模型可能会放大训练数据中存在的偏见,可能导致边缘化群体有偏见地获取信息。这项工作介绍了一个策划的口语偏见评估工具包和相应的数据集。我们在四个语义相关的任务中评估了SILLM中的性别偏见:语音到文本翻译(STT),口语共指消解(SCR),口语句子延续(SSC)和口语问题回答(SQA)。我们的分析表明,偏见水平是依赖于语言和不同的评价方法。我们的研究结果强调了采用多种方法来全面评估SILLM中的偏差的必要性,为开发更公平的SILLM系统提供了见解。摘要:Speech Integrated Large Language Models (SILLMs) combine large language models with speech perception to perform diverse tasks, such as emotion recognition to speaker verification, demonstrating universal audio understanding capability. However, these models may amplify biases present in training data, potentially leading to biased access to information for marginalized groups. This work introduces a curated spoken bias evaluation toolkit and corresponding dataset. We evaluate gender bias in SILLMs across four semantic-related tasks: speech-to-text translation (STT), spoken coreference resolution (SCR), spoken sentence continuation (SSC), and spoken question answering (SQA). Our analysis reveals that bias levels are language-dependent and vary with different evaluation methods. Our findings emphasize the necessity of employing multiple approaches to comprehensively assess biases in SILLMs, providing insights for developing fairer SILLM systems.
【2】 Gaunt coefficients for complex and real spherical harmonics with applications to spherical array processing and Ambisonics
标题: 复和实球调和高斯系数及其在球阵处理和立体声合成中的应用
作者:Archontis Politis
链接:点击下载PDF文件
摘要:声场的方向表示的声学信号处理,包括源,接收器和散射体传递函数,通常在球谐域(SHD)中表示和建模。某些这样的建模操作或那些模型的应用涉及那些方向量的乘法,这些方向量也可以通过称为冈特系数的耦合系数在SHD中方便地表示。由于Gaunt系数的定义和符号在声学出版物中有所不同,因此这项工作基于复和实球谐函数(SH)的既定惯例以及用于定向带限球函数的球乘的方便矩阵形式来定义它们。此外,该报告还提供了真实SH的Gaunt系数的推导,这在文献中是缺失的,可以直接用于空间音频框架,如高保真度立体声。提供了Matlab代码,可以计算用户指定的SH阶的所有系数。最后,一些相关的声学处理的例子,从文献中,下面的矩阵形式主义的系数在报告中介绍。摘要:Acoustical signal processing of directional representations of sound fields, including source, receiver, and scatterer transfer functions, are often expressed and modeled in the spherical harmonic domain (SHD). Certain such modeling operations, or applications of those models, involve multiplications of those directional quantities, which can also be expressed conveniently in the SHD through coupling coefficients known as Gaunt coefficients. Since the definition and notation of Gaunt coefficients varies across acoustical publications, this work defines them based on established conventions of complex and real spherical harmonics (SHs) along with a convenient matrix form for spherical multiplication of directionally band-limited spherical functions. Additionally, the report provides a derivation of the Gaunt coefficients for real SHs, which has been missing from the literature and can be used directly in spatial audio frameworks such as Ambisonics. Matlab code is provided that can compute all coefficients up to user specified SH orders. Finally, a number of relevant acoustical processing examples from the literature are presented, following the matrix formalism of coefficients introduced in the report.
【3】 Learn and Don't Forget: Adding a New Language to ASR Foundation Models
标题: 学习并不要忘记:向ASB基础模型添加新语言
作者:Mengjie Qian,Siyuan Tang,Rao Ma,Kate M. Knill,Mark J. F. Gales
链接:点击下载PDF文件
摘要:基础ASR模型通常支持多种语言,例如Whisper中的100种语言。然而,在集成一种额外的、通常资源较少的语言,同时保持原始语言集的性能方面的工作有限。微调虽然简单,但可能会降低原始集合的准确性。我们比较了三种方法,利用自适应参数:软语言代码调整,训练只有语言代码;软提示调整,训练前置令牌;和LoRA的一小部分额外的参数进行了优化。弹性权重合并(Elastic Weight Consolidation,EWC)提供了另一种折衷方案,有可能保持特定目标语言的性能。结果表明,直接微调产生的新语言的最佳性能,但降低现有的语言能力。EWC可以针对特定语言解决此问题。如果仅使用自适应参数,则保持语言能力,但以新语言的性能为代价。摘要:Foundation ASR models often support many languages, e.g. 100 languages in Whisper. However, there has been limited work on integrating an additional, typically low-resource, language, while maintaining performance on the original language set. Fine-tuning, while simple, may degrade the accuracy of the original set. We compare three approaches that exploit adaptation parameters: soft language code tuning, train only the language code; soft prompt tuning, train prepended tokens; and LoRA where a small set of additional parameters are optimised. Elastic Weight Consolidation (EWC) offers an alternative compromise with the potential to maintain performance in specific target languages. Results show that direct fine-tuning yields the best performance for the new language but degrades existing language capabilities. EWC can address this issue for specific languages. If only adaptation parameters are used, the language capabilities are maintained but at the cost of performance in the new language.
【4】 XANE Background Acoustic Embeddings: Ablation and Clustering Analysis
标题: XANE背景声学嵌入:消融和聚集分析
作者:Dushyant Sharma,James Fosburgh,Sri Harsha Dumpala,Chandramouli Shama Sastri,Stanislav Yu. Kruchinin,Patrick A. Naylor
备注:arXiv admin note: substantial text overlap with arXiv:2406.05199
链接:点击下载PDF文件
摘要:我们探讨了最近提出的可解释的声学神经嵌入~(XANE)系统,该系统以非侵入的方式对语音信号的背景声学进行建模。XANE嵌入用于估计与信号的背景声学特性相关的特定参数,这使得嵌入可以根据这些参数来解释。我们对XANE系统进行消融研究,结果表明,联合估计所有声学参数具有总体积极影响。此外,我们通过对看不见的测试数据进行聚类实验来说明XANE嵌入的价值,并表明所提出的嵌入在三个不同的任务中实现了92%的平均F1分数,显著优于基于WavLM的信号嵌入,并且与说话人嵌入互补。摘要:We explore the recently proposed explainable acoustic neural embedding~(XANE) system that models the background acoustics of a speech signal in a non-intrusive manner. The XANE embeddings are used to estimate specific parameters related to the background acoustic properties of the signal which allows the embeddings to be explainable in terms of those parameters. We perform ablation studies on the XANE system and show that estimating all acoustic parameters jointly has an overall positive effect. Furthermore, we illustrate the value of XANE embeddings by performing clustering experiments on unseen test data and show that the proposed embeddings achieve a mean F1 score of 92 % for three different tasks, outperforming significantly the WavLM based signal embeddings and are complimentary to speaker embeddings.
【5】 A Framework for Multimodal Medical Image Interaction
标题: 多模式医学图像交互框架
作者:Laura Schütz,Sasan Matinfar,Gideon Schafroth,Navid Navab,Merle Fairhurst,Arthur Wagner,Benedikt Wiestler,Ulrich Eck,Nassir Navab
备注:Accepted for publication in IEEE TVCG; presentation at IEEE ISMAR 2024
链接:点击下载PDF文件
摘要:医生依靠人体解剖学的图像,例如磁共振成像(MRI),在诊断和治疗期间定位患者的感兴趣区域。尽管医学成像技术取得了进步,但信息传递仍然是单峰的。这种视觉表现未能捕捉到与人体组织的真实、多感官互动的复杂性。然而,实时感知关于患者解剖结构和疾病的多模态信息对于医疗程序和患者结果的成功至关重要。我们引入了一个多模态医学图像交互(MMII)框架,让医学专家在三维空间中与人体组织进行动态的视听交互。在虚拟现实环境中,用户接收物理通知的视听反馈以改善解剖结构的空间感知。MMII使用基于模型的声音处理方法来产生来自组织的几何形状和物理特性的声音,从而消除了手工制作声音设计的需要。两个用户研究,涉及34个一般和9个临床专家进行了评估建议的互动框架的可学习性,可用性和准确性。我们的研究结果显示,视听对应的可学习性非常好,因为在研究过程中,正确联想的比率显著提高(p < 0.001)。与传统医学图像交互相比,MMII的脑肿瘤定位准确性更高(p < 0.05)。我们的研究结果证实了这种新框架的潜力,以加强与医学图像的互动,例如,在手术过程中,需要立即和精确的反馈。摘要:Medical doctors rely on images of the human anatomy, such as magnetic resonance imaging (MRI), to localize regions of interest in the patient during diagnosis and treatment. Despite advances in medical imaging technology, the information conveyance remains unimodal. This visual representation fails to capture the complexity of the real, multisensory interaction with human tissue. However, perceiving multimodal information about the patient's anatomy and disease in real-time is critical for the success of medical procedures and patient outcome. We introduce a Multimodal Medical Image Interaction (MMII) framework to allow medical experts a dynamic, audiovisual interaction with human tissue in three-dimensional space. In a virtual reality environment, the user receives physically informed audiovisual feedback to improve the spatial perception of anatomical structures. MMII uses a model-based sonification approach to generate sounds derived from the geometry and physical properties of tissue, thereby eliminating the need for hand-crafted sound design. Two user studies involving 34 general and nine clinical experts were conducted to evaluate the proposed interaction framework's learnability, usability, and accuracy. Our results showed excellent learnability of audiovisual correspondence as the rate of correct associations significantly improved (p < 0.001) over the course of the study. MMII resulted in superior brain tumor localization accuracy (p < 0.05) compared to conventional medical image interaction. Our findings substantiate the potential of this novel framework to enhance interaction with medical images, for example, during surgical procedures where immediate and precise feedback is needed.
【6】 Audio-Language Datasets of Scenes and Events: A Survey
标题: 场景和事件的音频语言数据集:调查
作者:Gijs Wijngaard,Elia Formisano,Michele Esposito,Michel Dumontier
链接:点击下载PDF文件
摘要:音频语言模型(ALM)处理声音,以提供对产生声音的事件和场景的语言描述。计算能力和数据集创建的最新进展导致了这一领域的重大进展。本文调查了用于训练音频语言模型的现有数据集,强调了最近使用大型,多样化数据集来增强模型性能的趋势。这些数据集的主要来源包括Freesound平台和AudioSet,它们为该领域的快速增长做出了贡献。虽然以前的调查主要涉及技术和培训细节,但本调查对各种数据集进行了分类和评估,并讨论了它们的起源,特征和用例。它还执行数据泄漏分析,以确保数据集的完整性并减轻数据集之间的偏差。本次调查是通过分析截至2023年12月(含)的研究论文进行的,不包含该时期之后的任何论文。摘要:Audio-language models (ALMs) process sounds to provide a linguistic description of sound-producing events and scenes. Recent advances in computing power and dataset creation have led to significant progress in this domain. This paper surveys existing datasets used for training audio-language models, emphasizing the recent trend towards using large, diverse datasets to enhance model performance. Key sources of these datasets include the Freesound platform and AudioSet that have contributed to the field's rapid growth. Although prior surveys primarily address techniques and training details, this survey categorizes and evaluates a wide array of datasets, addressing their origins, characteristics, and use cases. It also performs a data leak analysis to ensure dataset integrity and mitigate bias between datasets. This survey was conducted by analyzing research papers up to and including December 2023, and does not contain any papers after that period.
【7】 RespEar: Earable-Based Robust Respiratory Rate Monitoring
标题: RespEar:基于Early的稳健呼吸率监测
作者:Yang Liu,Kayla-Jade Butkow,Jake Stuchbury-Wass,Adam Pullin,Dong Ma,Cecilia Mascolo
链接:点击下载PDF文件
摘要:呼吸率(RR)监测对于了解身心健康和跟踪健康状况是不可或缺的。现有研究已经证明了在特定用户条件下(例如,同时保持静止,或者同时沉重地呼吸)。然而,在不同的日常生活和活动中进行准确,连续和非侵入性的RR监测仍然具有挑战性。在这项工作中,我们提出了RespEar,一个基于earable的系统,强大的RR监测。通过利用耳塞中入耳式麦克风的独特特性,RespEar能够使用呼吸性窦性心律失常(RSA)和运动性呼吸耦合(LRC),心血管活动,步态和呼吸之间的生理耦合,间接确定RR。这有效地解决了日常活动中几乎无法察觉的呼吸信号所带来的挑战。我们进一步提出了一套精心制作的信号处理方案,以提高RR估计的准确性和鲁棒性。通过从18名受试者的8项活动中收集的数据,RespEar测量RR,在久坐状态下的平均绝对误差(MAE)为1.48次呼吸 分钟(BPM),平均绝对百分比误差(MAPE)为9.12%,在活动状态下的MAE为2.28 BPM,MAPE为11.04%,这对于能够用单一模态概括各种条件的方法来说是前所未有的。摘要:Respiratory rate (RR) monitoring is integral to understanding physical and mental health and tracking fitness. Existing studies have demonstrated the feasibility of RR monitoring under specific user conditions (e.g., while remaining still, or while breathing heavily). Yet, performing accurate, continuous and non-obtrusive RR monitoring across diverse daily routines and activities remains challenging. In this work, we present RespEar, an earable-based system for robust RR monitoring. By leveraging the unique properties of in-ear microphones in earbuds, RespEar enables the use of Respiratory Sinus Arrhythmia (RSA) and Locomotor Respiratory Coupling (LRC), physiological couplings between cardiovascular activity, gait and respiration, to indirectly determine RR. This effectively addresses the challenges posed by the almost imperceptible breathing signals under daily activities. We further propose a suite of meticulously crafted signal processing schemes to improve RR estimation accuracy and robustness. With data collected from 18 subjects over 8 activities, RespEar measures RR with a mean absolute error (MAE) of 1.48 breaths per minutes (BPM) and a mean absolute percent error (MAPE) of 9.12% in sedentary conditions, and a MAE of 2.28 BPM and a MAPE of 11.04% in active conditions, respectively, which is unprecedented for a method capable of generalizing across conditions with a single modality.
【8】 Improving Speech Enhancement by Integrating Inter-Channel and Band Features with Dual-branch Conformer
标题: 通过将通道间和频段特征与双分支一致器集成来改善语音增强
作者:Jizhen Li,Xinmeng Xu,Weiping Tu,Yuhong Yang,Rong Zhu
链接:点击下载PDF文件
摘要:基于卷积神经网络(CNN)和Transformer的语音增强方法能够有效地提取语音信号的时频信息。然而,语音特征的各个通道之间的相关性却没有得到很好的探索。理论上,由不同卷积核获得的语音特征的每个通道图包含具有不同尺度的信息,表现出强相关性。为了填补这一空白,我们提出了一种新的双分支架构命名为通道感知双分支构象(CADB-Conformer),有效地探索不同的通道之间的长范围的时间和频率的相关性,分别提取信道关系感知的时频信息。在DNS挑战2020数据集上进行的消融研究证明了通道特征利用的重要性,同时显示了通道关系感知T-F信息对语音增强的重要性。大量的实验也表明,该模型实现了优越的性能比最近的方法具有吸引力的计算成本。摘要:Recent speech enhancement methods based on convolutional neural networks (CNNs) and transformer have been demonstrated to efficaciously capture time-frequency (T-F) information on spectrogram. However, the correlation of each channels of speech features is failed to explore. Theoretically, each channel map of speech features obtained by different convolution kernels contains information with different scales demonstrating strong correlations. To fill this gap, we propose a novel dual-branch architecture named channel-aware dual-branch conformer (CADB-Conformer), which effectively explores the long range time and frequency correlations among different channels, respectively, to extract channel relation aware time-frequency information. Ablation studies conducted on DNS-Challenge 2020 dataset demonstrate the importance of channel feature leveraging while showing the significance of channel relation aware T-F information for speech enhancement. Extensive experiments also show that the proposed model achieves superior performance than recent methods with an attractive computational costs.
【9】 Homogeneous Speaker Features for On-the-Fly Dysarthric and Elderly Speaker Adaptation
标题: 用于动态合成障碍和老年说话者适应的同质说话者功能
作者:Mengzhe Geng,Xurong Xie,Jiajun Deng,Zengrui Jin,Guinan Li,Tianzi Wang,Shujie Hu,Zhaoqing Li,Helen Meng,Xunying Liu
备注:In submission to IEEEACM Transactions on Audio, Speech, and Language Processing
链接:点击下载PDF文件
摘要:将数据密集型自动语音识别(ASR)技术应用于构音障碍和老年成人语音,面临着与健康和非老年人语音的不匹配、数据稀缺和说话者水平的大变异性。为此,本文提出了两种新的数据高效的方法来学习同质构音障碍和老年人说话者级别的功能,用于DNN TDNN和Conformer ASR模型的快速,动态测试时间适应。这些措施包括:1)说话人级方差正则化谱基嵌入(VR-SBE)特征,其利用特殊的正则化项来加强自适应中说话人特征的同质性;以及2)基于特征的学习隐藏单元贡献(f-LHUC)变换,其以VR-SBE特征为条件。实验在两种语言的四个任务上进行:英语UASpeech和TORGO构音障碍语音数据集,英语DementiaBank Pitt和粤语JCCOCC MoCA老年人语音语料库。所提出的动态扬声器自适应技术始终优于基线iVector和xVector自适应,统计上显着的单词或字符错误率降低了5.32%的绝对值(相对值为18.57%),批处理模式LHUC扬声器自适应降低了2.24%的绝对值(相对值为9.20%),同时在自适应期间使用实时因素加速到xVector的33.6倍。所提出的适应技术的有效性在与当前ASR技术的比较中得到了证明,包括UASpeech上的SSL预训练系统,其中我们最好的系统产生了23.33%的最先进的WER。分析表明,VR-SBE特征和f-LHUC变换在测试时自适应中对说话人级数据量不敏感。T-SNE可视化显示,它们比基线iVectors,xVectors和批处理模式LHUC变换具有更强的说话者级别均匀性。摘要:The application of data-intensive automatic speech recognition (ASR) technologies to dysarthric and elderly adult speech is confronted by their mismatch against healthy and nonaged voices, data scarcity and large speaker-level variability. To this end, this paper proposes two novel data-efficient methods to learn homogeneous dysarthric and elderly speaker-level features for rapid, on-the-fly test-time adaptation of DNN TDNN and Conformer ASR models. These include: 1) speaker-level variance-regularized spectral basis embedding (VR-SBE) features that exploit a special regularization term to enforce homogeneity of speaker features in adaptation; and 2) feature-based learning hidden unit contributions (f-LHUC) transforms that are conditioned on VR-SBE features. Experiments are conducted on four tasks across two languages: the English UASpeech and TORGO dysarthric speech datasets, the English DementiaBank Pitt and Cantonese JCCOCC MoCA elderly speech corpora. The proposed on-the-fly speaker adaptation techniques consistently outperform baseline iVector and xVector adaptation by statistically significant word or character error rate reductions up to 5.32% absolute (18.57% relative) and batch-mode LHUC speaker adaptation by 2.24% absolute (9.20% relative), while operating with real-time factors speeding up to 33.6 times against xVectors during adaptation. The efficacy of the proposed adaptation techniques is demonstrated in a comparison against current ASR technologies including SSL pre-trained systems on UASpeech, where our best system produces a state-of-the-art WER of 23.33%. Analyses show VR-SBE features and f-LHUC transforms are insensitive to speaker-level data quantity in testtime adaptation. T-SNE visualization reveals they have stronger speaker-level homogeneity than baseline iVectors, xVectors and batch-mode LHUC transforms.
【10】 Transfer Learning with Pseudo Multi-Label Birdcall Classification for DS@GT BirdCLEF 2024
标题: DS@GT BirdCREF 2024的伪多标签鸟鸣分类的迁移学习
作者:Anthony Miyaguchi,Adrian Cheung,Murilo Gustineli,Ashley Kim
备注:Submitted and accepted into CLEF 2024 CEUR-WS proceedings
链接:点击下载PDF文件
摘要:我们为DS@GT团队提供了关于迁移学习的工作笔记,其中包含BirdCLEF 2024竞赛的伪多标签鸟鸣分类,重点是在录制的音景中识别印度鸟类。我们的方法利用生产级模型,如Google Bird Vocalization Classifier,BirdNET和EnCodec,以解决竞争中的表示和标签挑战。我们探讨了今年版本的未标记的音景代表隐藏的测试集之间的分布变化,并提出了一个伪多标签分类策略,以利用未标记的数据。我们的最高赛后公开排行榜分数是0.63,使用BirdNET嵌入与鸟发声伪标签。我们的代码可在https: github.com dsgt-kaggle-clef birdclef-2024上获得摘要:We present working notes for the DS@GT team on transfer learning with pseudo multi-label birdcall classification for the BirdCLEF 2024 competition, focused on identifying Indian bird species in recorded soundscapes. Our approach utilizes production-grade models such as the Google Bird Vocalization Classifier, BirdNET, and EnCodec to address representation and labeling challenges in the competition. We explore the distributional shift between this year's edition of unlabeled soundscapes representative of the hidden test set and propose a pseudo multi-label classification strategy to leverage the unlabeled data. Our highest post-competition public leaderboard score is 0.63 using BirdNET embeddings with Bird Vocalization pseudo-labels. Our code is available at https: github.com dsgt-kaggle-clef birdclef-2024
机器翻译,仅供参考
