【1】Analyzing Musical Characteristics of National Anthems in Relation to Global Indices标题:从全球指标看国歌的音乐特征链接:https://arxiv.org/abs/2404.03606作者:S M Rakib Hasan,Aakar Dhakal,Ms. Ayesha Siddiqua,Mohammad Mominur Rahman,Md Maidul Islam,Mohammed Arfat Raihan Chowdhury,S M Masfequier Rahman Swapno,SM Nuruzzaman Nobel摘要:音乐在塑造人们的心理和行为模式方面起着巨大的作用。本文采用计算音乐分析和统计相关分析的方法,研究了各国国歌与不同的全球指数之间的关系。我们分析国歌的音乐数据,以确定某些音乐特征是否与和平,幸福,自杀率,犯罪率等,以实现这一目标,我们收集了来自169个国家的国歌,并使用计算音乐分析技术来提取音高,节奏,节拍,和其他相关的音频功能。然后,我们将这些音乐特征与不同全球指数的数据进行比较,以确定是否存在显着的相关性。研究结果表明,国歌的音乐特征与我们所研究的指标之间可能存在相关性。我们的研究结果对音乐心理学和有兴趣促进社会福祉的政策制定者的影响进行了讨论。本文强调了音乐数据分析在社会研究中的潜力,并为音乐与社会指数之间的关系提供了一个新的视角。源代码和数据是开放的,可重复性和未来的研究工作。可在http://bit.ly/na_code上查阅。摘要:Music plays a huge part in shaping peoples' psychology and behavioral patterns. This paper investigates the connection between national anthems and different global indices with computational music analysis and statistical correlation analysis. We analyze national anthem musical data to determine whether certain musical characteristics are associated with peace, happiness, suicide rate, crime rate, etc. To achieve this, we collect national anthems from 169 countries and use computational music analysis techniques to extract pitch, tempo, beat, and other pertinent audio features. We then compare these musical characteristics with data on different global indices to ascertain whether a significant correlation exists. Our findings indicate that there may be a correlation between the musical characteristics of national anthems and the indices we investigated. The implications of our findings for music psychology and policymakers interested in promoting social well-being are discussed. This paper emphasizes the potential of musical data analysis in social research and offers a novel perspective on the relationship between music and social indices. The source code and data are made open-access for reproducibility and future research endeavors. It can be accessed at http://bit.ly/na_code. 【2】 M3TCM: Multi-modal Multi-task Context Model for Utterance Classification in Motivational Interviews标题:M3TCM:多模态多任务语境模型在动机性访谈中的话语分类链接:https://arxiv.org/abs/2404.03312作者:Sayed Muddashir Hossain,Jan Alexandersson,Philipp Müller备注:Accepted for publication at LREC-COLING'24摘要:动机访谈中准确的话语分类对于自动理解客户-治疗师互动的质量和动态至关重要,并且它可以作为调解此类互动的系统的关键输入。动机性面试表现出三个重要特征。首先,有两个不同的角色,即客户和治疗师。第二,它们往往充满了高度的情感,这可以通过文本和韵律来表达。最后,语境对任何给定的话语进行分类都是至关重要的。以前的作品没有充分地将所有这些特征纳入心理健康对话的话语分类方法。相比之下,我们提出了M3 TCM,一个多模态,多任务上下文模型的话语分类。我们的方法首次采用多任务学习来有效地模拟治疗师和客户行为的联合和个人组成部分。此外,M3 TCM集成了来自文本和语音模态以及会话上下文的信息。通过我们的新方法,我们在最近推出的AnnoMI数据集上的话语分类方面优于最先进的技术,客户端的相对改善为20%,治疗师的话语分类为15%。在广泛的消融研究中,我们量化了每种贡献所带来的改善。摘要:Accurate utterance classification in motivational interviews is crucial to automatically understand the quality and dynamics of client-therapist interaction, and it can serve as a key input for systems mediating such interactions. Motivational interviews exhibit three important characteristics. First, there are two distinct roles, namely client and therapist. Second, they are often highly emotionally charged, which can be expressed both in text and in prosody. Finally, context is of central importance to classify any given utterance. Previous works did not adequately incorporate all of these characteristics into utterance classification approaches for mental health dialogues. In contrast, we present M3TCM, a Multi-modal, Multi-task Context Model for utterance classification. Our approach for the first time employs multi-task learning to effectively model both joint and individual components of therapist and client behaviour. Furthermore, M3TCM integrates information from the text and speech modality as well as the conversation context. With our novel approach, we outperform the state of the art for utterance classification on the recently introduced AnnoMI dataset with a relative improvement of 20% for the client- and by 15% for therapist utterance classification. In extensive ablation studies, we quantify the improvement resulting from each contribution.
【3】 UniAV: Unified Audio-Visual Perception for Multi-Task Video Localization标题:UniAV:多任务视频本地化的统一视听感知链接:https://arxiv.org/abs/2404.03179作者:Tiantian Geng,Teng Wang,Yanfu Zhang,Jinming Duan,Weili Guan,Feng Zheng摘要:视频定位任务旨在对视频中的特定实例进行时间定位,包括时间动作定位(TAL)、声音事件检测(SED)和视听事件定位(AVEL)。现有的方法过度专注于每个任务,忽略了这些实例经常出现在同一视频中以形成完整的视频内容的事实。在这项工作中,我们提出了UniAV,一个统一的视听感知网络,首次实现了TAL,SED和AVEL任务的联合学习。UniAV可以利用特定任务数据集中的各种数据,使模型能够跨任务和模式学习和共享互利知识。为了解决数据集(大小/域/持续时间)和不同任务特征的巨大变化所带来的挑战,我们建议对所有视频的视觉和音频模式进行统一编码,以获得通用表示,同时还设计特定于任务的专家来捕获每个任务的独特知识。此外,我们通过使用预训练的文本编码器开发了一个统一的语言感知分类器,使模型能够通过在推理过程中简单地改变提示来灵活地检测各种类型的实例和以前看不见的实例。UniAV在参数较少的情况下大幅优于其单任务同行,与ActivityNet 1.3,DESED和UnAV-100基准测试中最先进的特定任务方法相比,UniAV的性能达到了同等或更高的水平。摘要:Video localization tasks aim to temporally locate specific instances in videos, including temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods over-specialize on each task, overlooking the fact that these instances often occur in the same video to form the complete video content. In this work, we present UniAV, a Unified Audio-Visual perception network, to achieve joint learning of TAL, SED and AVEL tasks for the first time. UniAV can leverage diverse data available in task-specific datasets, allowing the model to learn and share mutually beneficial knowledge across tasks and modalities. To tackle the challenges posed by substantial variations in datasets (size/domain/duration) and distinct task characteristics, we propose to uniformly encode visual and audio modalities of all videos to derive generic representations, while also designing task-specific experts to capture unique knowledge for each task. Besides, we develop a unified language-aware classifier by utilizing a pre-trained text encoder, enabling the model to flexibly detect various types of instances and previously unseen ones by simply changing prompts during inference. UniAV outperforms its single-task counterparts by a large margin with fewer parameters, achieving on-par or superior performances compared to state-of-the-art task-specific methods across ActivityNet 1.3, DESED and UnAV-100 benchmarks. 【4】 Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian标题:Mai Ho 'omāuna i ka'Ai:语言模型改善夏威夷语的自动语音识别链接:https://arxiv.org/abs/2404.03073作者:Kaavya Chaparala,Guido Zarrella,Bruce Torres Fischer,Larry Kimura,Oiwi Parker Jones摘要:在本文中,我们将大量的独立的文本数据整合到一个自动语音识别(ASR)的基础模型,耳语,以解决提高低资源的语言,夏威夷语的挑战。为了做到这一点,我们训练一个外部语言模型(LM)的夏威夷语文本约150万字。然后,我们使用LM对Whisper进行重新评分,并在手动策划的标记夏威夷数据测试集上计算单词错误率(WER)。作为基线,我们使用Whisper而不使用外部LM。实验结果表明,一个小的,但显着的改善WER时,ASR输出与夏威夷LM rescored。结果支持利用所有可用的数据在开发ASR系统的代表性不足的语言。摘要:In this paper we address the challenge of improving Automatic Speech Recognition (ASR) for a low-resource language, Hawaiian, by incorporating large amounts of independent text data into an ASR foundation model, Whisper. To do this, we train an external language model (LM) on ~1.5M words of Hawaiian text. We then use the LM to rescore Whisper and compute word error rates (WERs) on a manually curated test set of labeled Hawaiian data. As a baseline, we use Whisper without an external LM. Experimental results reveal a small but significant improvement in WER when ASR outputs are rescored with a Hawaiian LM. The results support leveraging all available data in the development of ASR systems for underrepresented languages.
【5】 Correlation and Spectral Density Functions in Mode-Stirred Reverberation - II. Spectral Moments, Sampling, Noise, EMI and Understirring标题:模式搅拌混响中的相关和谱密度函数—II。频谱矩、采样、噪声、EMI和欠搅拌链接:https://arxiv.org/abs/2404.03520作者:Luk R. Arnaut,John M. Ladbury备注:None摘要:在第一部分中,谱矩和峰度作为参数建立了动态混响场的相关函数和谱密度函数的解析模型。在这第二部分中,几个实际的限制影响的准确性估计这些参数从测量搅拌扫描数据进行了调查。对于采样场,有限差分和混叠的贡献进行了评估。有限差分的结果在负偏置依赖于,以领先的顺序,二次的采样时间间隔和搅拌带宽的产品。直接从采样搅拌扫描提取的矩的数值估计显示出良好的协议,通过自协方差方法获得的值。确定和实验验证的RMS振幅的数据抽取和噪声搅拌比的影响。此外,依赖于噪声的搅拌带宽比,EMI,和未搅拌的能量的特点。摘要:In part I, spectral moments and kurtosis were established as parameters in analytic models of correlation and spectral density functions for dynamic reverberation fields. In this part II, several practical limitations affecting the accuracy of estimating these parameters from measured stir sweep data are investigated. For sampled fields, the contributions of finite differencing and aliasing are evaluated. Finite differencing results in a negative bias that depends, to leading order, quadratically on the product of the sampling time interval and the stir bandwidth. Numerical estimates of moments extracted directly from sampled stir sweeps show good agreement with values obtained by an autocovariance method. The effects of data decimation and noise-to-stir ratios of RMS amplitudes are determined and experimentally verified. In addition, the dependencies on the noise-to-stir-bandwidth ratio, EMI, and unstirred energy are characterized. 【6】 Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation标题:使用分层相关传播解释端到端深度学习模型以进行语音源定位链接:https://arxiv.org/abs/2404.03436作者:Luca Comanducci,Fabio Antonacci,Augusto Sarti摘要:深度学习模型广泛应用于信号处理领域,但其内部工作过程通常被视为黑箱。在本文中,我们研究了使用可解释的人工智能(XAI)技术,以学习为基础的端到端的语音源定位模型。我们考虑分层相关传播(LRP)技术,其目的是确定输入的哪些部分对输出预测更重要。使用LRP,我们分析了两个国家的最先进的模型,不同的建筑复杂性,将麦克风获得的音频信号映射到源的carbohydrate坐标。具体来说,我们检查了与两个模型的输入特征相关的相关性,并发现两个网络都对麦克风信号进行去噪和去混响,以计算它们之间更准确的统计相关性,从而定位源。为了进一步证明这一事实,我们估计到达时间差(TDoAs)通过广义互相关相位变换(GCC-PHAT)使用麦克风信号和相关信号提取的两个网络,并表明,通过后者,我们得到更准确的时间延迟估计结果。摘要:Deep learning models are widely applied in the signal processing community, yet their inner working procedure is often treated as a black box. In this paper, we investigate the use of eXplainable Artificial Intelligence (XAI) techniques to learning-based end-to-end speech source localization models. We consider the Layer-wise Relevance Propagation (LRP) technique, which aims to determine which parts of the input are more important for the output prediction. Using LRP we analyze two state-of-the-art models, of differing architectural complexity that map audio signals acquired by the microphones to the cartesian coordinates of the source. Specifically, we inspect the relevance associated with the input features of the two models and discover that both networks denoise and de-reverberate the microphone signals to compute more accurate statistical correlations between them and consequently localize the sources. To further demonstrate this fact, we estimate the Time-Difference of Arrivals (TDoAs) via the Generalized Cross Correlation with Phase Transform (GCC-PHAT) using both microphone signals and relevance signals extracted from the two networks and show that through the latter we obtain more accurate time-delay estimation results. 【7】 RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis标题:RALL—E:基于思想链的鲁棒编解码语言建模链接:https://arxiv.org/abs/2404.03204作者:Detai Xin,Xu Tan,Kai Shen,Zeqian Ju,Dongchao Yang,Yuancheng Wang,Shinnosuke Takamichi,Hiroshi Saruwatari,Shujie Liu,Jinyu Li,Sheng Zhao摘要:我们提出了RALL-E,一个强大的语言建模方法的文本到语音(TTS)合成。虽然以前的工作基于大型语言模型(LLM)显示出令人印象深刻的性能上zero-shot TTS,这样的方法往往遭受较差的鲁棒性,如不稳定的韵律(怪异的音高和节奏/持续时间)和高的单词错误率(WER),由于自回归预测风格的语言模型。RALL-E背后的核心思想是思想链(CoT)提示,它将任务分解为更简单的步骤,以增强基于LLM的TTS的鲁棒性。为了实现这个想法,RALL-E首先预测输入文本的韵律特征(音高和持续时间),并使用它们作为中间条件来预测CoT风格的语音标记。其次,RALL-E利用预测的持续时间提示来指导Transformer中自我注意权重的计算,以强制模型在预测语音令牌时关注相应的音素和韵律特征。综合客观和主观评价的结果表明,与一个强大的基线方法VALL-E相比,RALL-E显著提高了zero-shot TTS的WER,分别从6.3美元(不重排序)和2.1美元(重排序)提高到2.8美元和1.0美元。此外,我们证明了RALL-E正确地合成的句子,是很难VALL-E和减少错误率从68\%$到4\%$。摘要:We present RALL-E, a robust language modeling method for text-to-speech (TTS) synthesis. While previous work based on large language models (LLMs) shows impressive performance on zero-shot TTS, such methods often suffer from poor robustness, such as unstable prosody (weird pitch and rhythm/duration) and a high word error rate (WER), due to the autoregressive prediction style of language models. The core idea behind RALL-E is chain-of-thought (CoT) prompting, which decomposes the task into simpler steps to enhance the robustness of LLM-based TTS. To accomplish this idea, RALL-E first predicts prosody features (pitch and duration) of the input text and uses them as intermediate conditions to predict speech tokens in a CoT style. Second, RALL-E utilizes the predicted duration prompt to guide the computing of self-attention weights in Transformer to enforce the model to focus on the corresponding phonemes and prosody features when predicting speech tokens. Results of comprehensive objective and subjective evaluations demonstrate that, compared to a powerful baseline method VALL-E, RALL-E significantly improves the WER of zero-shot TTS from $6.3\%$ (without reranking) and $2.1\%$ (with reranking) to $2.8\%$ and $1.0\%$, respectively. Furthermore, we demonstrate that RALL-E correctly synthesizes sentences that are hard for VALL-E and reduces the error rate from $68\%$ to $4\%$. eess.AS音频处理【1】 Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation标题:使用分层相关传播解释端到端深度学习模型以进行语音源定位链接:https://arxiv.org/abs/2404.03436作者:Luca Comanducci,Fabio Antonacci,Augusto Sarti摘要:深度学习模型广泛应用于信号处理领域,但其内部工作过程通常被视为黑箱。在本文中,我们研究了使用可解释的人工智能(XAI)技术,以学习为基础的端到端的语音源定位模型。我们考虑分层相关传播(LRP)技术,其目的是确定输入的哪些部分对输出预测更重要。使用LRP,我们分析了两个国家的最先进的模型,不同的建筑复杂性,将麦克风获得的音频信号映射到源的carbohydrate坐标。具体来说,我们检查了与两个模型的输入特征相关的相关性,并发现两个网络都对麦克风信号进行去噪和去混响,以计算它们之间更准确的统计相关性,从而定位源。为了进一步证明这一事实,我们估计到达时间差(TDoAs)通过广义互相关相位变换(GCC-PHAT)使用麦克风信号和相关信号提取的两个网络,并表明,通过后者,我们得到更准确的时间延迟估计结果。摘要:Deep learning models are widely applied in the signal processing community, yet their inner working procedure is often treated as a black box. In this paper, we investigate the use of eXplainable Artificial Intelligence (XAI) techniques to learning-based end-to-end speech source localization models. We consider the Layer-wise Relevance Propagation (LRP) technique, which aims to determine which parts of the input are more important for the output prediction. Using LRP we analyze two state-of-the-art models, of differing architectural complexity that map audio signals acquired by the microphones to the cartesian coordinates of the source. Specifically, we inspect the relevance associated with the input features of the two models and discover that both networks denoise and de-reverberate the microphone signals to compute more accurate statistical correlations between them and consequently localize the sources. To further demonstrate this fact, we estimate the Time-Difference of Arrivals (TDoAs) via the Generalized Cross Correlation with Phase Transform (GCC-PHAT) using both microphone signals and relevance signals extracted from the two networks and show that through the latter we obtain more accurate time-delay estimation results. 【2】 RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis标题:RALL—E:基于思想链的鲁棒编解码语言建模链接:https://arxiv.org/abs/2404.03204作者:Detai Xin,Xu Tan,Kai Shen,Zeqian Ju,Dongchao Yang,Yuancheng Wang,Shinnosuke Takamichi,Hiroshi Saruwatari,Shujie Liu,Jinyu Li,Sheng Zhao摘要:我们提出了RALL-E,一个强大的语言建模方法的文本到语音(TTS)合成。虽然以前的工作基于大型语言模型(LLM)显示出令人印象深刻的性能上zero-shot TTS,这样的方法往往遭受较差的鲁棒性,如不稳定的韵律(怪异的音高和节奏/持续时间)和高的单词错误率(WER),由于自回归预测风格的语言模型。RALL-E背后的核心思想是思想链(CoT)提示,它将任务分解为更简单的步骤,以增强基于LLM的TTS的鲁棒性。为了实现这个想法,RALL-E首先预测输入文本的韵律特征(音高和持续时间),并使用它们作为中间条件来预测CoT风格的语音标记。其次,RALL-E利用预测的持续时间提示来指导Transformer中自我注意权重的计算,以强制模型在预测语音令牌时关注相应的音素和韵律特征。综合客观和主观评价的结果表明,与一个强大的基线方法VALL-E相比,RALL-E显著提高了zero-shot TTS的WER,分别从6.3美元(不重排序)和2.1美元(重排序)提高到2.8美元和1.0美元。此外,我们证明了RALL-E正确地合成的句子,是很难VALL-E和减少错误率从68\%$到4\%$。摘要:We present RALL-E, a robust language modeling method for text-to-speech (TTS) synthesis. While previous work based on large language models (LLMs) shows impressive performance on zero-shot TTS, such methods often suffer from poor robustness, such as unstable prosody (weird pitch and rhythm/duration) and a high word error rate (WER), due to the autoregressive prediction style of language models. The core idea behind RALL-E is chain-of-thought (CoT) prompting, which decomposes the task into simpler steps to enhance the robustness of LLM-based TTS. To accomplish this idea, RALL-E first predicts prosody features (pitch and duration) of the input text and uses them as intermediate conditions to predict speech tokens in a CoT style. Second, RALL-E utilizes the predicted duration prompt to guide the computing of self-attention weights in Transformer to enforce the model to focus on the corresponding phonemes and prosody features when predicting speech tokens. Results of comprehensive objective and subjective evaluations demonstrate that, compared to a powerful baseline method VALL-E, RALL-E significantly improves the WER of zero-shot TTS from $6.3\%$ (without reranking) and $2.1\%$ (with reranking) to $2.8\%$ and $1.0\%$, respectively. Furthermore, we demonstrate that RALL-E correctly synthesizes sentences that are hard for VALL-E and reduces the error rate from $68\%$ to $4\%$. 【3】 Analyzing Musical Characteristics of National Anthems in Relation to Global Indices标题:从全球指标看国歌的音乐特征链接:https://arxiv.org/abs/2404.03606作者:S M Rakib Hasan,Aakar Dhakal,Ms. Ayesha Siddiqua,Mohammad Mominur Rahman,Md Maidul Islam,Mohammed Arfat Raihan Chowdhury,S M Masfequier Rahman Swapno,SM Nuruzzaman Nobel摘要:音乐在塑造人们的心理和行为模式方面起着巨大的作用。本文采用计算音乐分析和统计相关分析的方法,研究了各国国歌与不同的全球指数之间的关系。我们分析国歌的音乐数据,以确定某些音乐特征是否与和平,幸福,自杀率,犯罪率等,以实现这一目标,我们收集了来自169个国家的国歌,并使用计算音乐分析技术来提取音高,节奏,节拍,和其他相关的音频功能。然后,我们将这些音乐特征与不同全球指数的数据进行比较,以确定是否存在显着的相关性。研究结果表明,国歌的音乐特征与我们所研究的指标之间可能存在相关性。我们的研究结果对音乐心理学和有兴趣促进社会福祉的政策制定者的影响进行了讨论。本文强调了音乐数据分析在社会研究中的潜力,并为音乐与社会指数之间的关系提供了一个新的视角。源代码和数据是开放的,可重复性和未来的研究工作。可在http://bit.ly/na_code上查阅。摘要:Music plays a huge part in shaping peoples' psychology and behavioral patterns. This paper investigates the connection between national anthems and different global indices with computational music analysis and statistical correlation analysis. We analyze national anthem musical data to determine whether certain musical characteristics are associated with peace, happiness, suicide rate, crime rate, etc. To achieve this, we collect national anthems from 169 countries and use computational music analysis techniques to extract pitch, tempo, beat, and other pertinent audio features. We then compare these musical characteristics with data on different global indices to ascertain whether a significant correlation exists. Our findings indicate that there may be a correlation between the musical characteristics of national anthems and the indices we investigated. The implications of our findings for music psychology and policymakers interested in promoting social well-being are discussed. This paper emphasizes the potential of musical data analysis in social research and offers a novel perspective on the relationship between music and social indices. The source code and data are made open-access for reproducibility and future research endeavors. It can be accessed at http://bit.ly/na_code. 【4】 Correlation and Spectral Density Functions in Mode-Stirred Reverberation - II. Spectral Moments, Sampling, Noise, EMI and Understirring标题:模式搅拌混响中的相关和谱密度函数—II。频谱矩、采样、噪声、EMI和欠搅拌链接:https://arxiv.org/abs/2404.03520作者:Luk R. Arnaut,John M. Ladbury备注:None摘要:在第一部分中,谱矩和峰度作为参数建立了动态混响场的相关函数和谱密度函数的解析模型。在这第二部分中,几个实际的限制影响的准确性估计这些参数从测量搅拌扫描数据进行了调查。对于采样场,有限差分和混叠的贡献进行了评估。有限差分的结果在负偏置依赖于,以领先的顺序,二次的采样时间间隔和搅拌带宽的产品。直接从采样搅拌扫描提取的矩的数值估计显示出良好的协议,通过自协方差方法获得的值。确定和实验验证的RMS振幅的数据抽取和噪声搅拌比的影响。此外,依赖于噪声的搅拌带宽比,EMI,和未搅拌的能量的特点。摘要:In part I, spectral moments and kurtosis were established as parameters in analytic models of correlation and spectral density functions for dynamic reverberation fields. In this part II, several practical limitations affecting the accuracy of estimating these parameters from measured stir sweep data are investigated. For sampled fields, the contributions of finite differencing and aliasing are evaluated. Finite differencing results in a negative bias that depends, to leading order, quadratically on the product of the sampling time interval and the stir bandwidth. Numerical estimates of moments extracted directly from sampled stir sweeps show good agreement with values obtained by an autocovariance method. The effects of data decimation and noise-to-stir ratios of RMS amplitudes are determined and experimentally verified. In addition, the dependencies on the noise-to-stir-bandwidth ratio, EMI, and unstirred energy are characterized. 【5】 M3TCM: Multi-modal Multi-task Context Model for Utterance Classification in Motivational Interviews标题:M3TCM:多模态多任务语境模型在动机性访谈中的话语分类链接:https://arxiv.org/abs/2404.03312作者:Sayed Muddashir Hossain,Jan Alexandersson,Philipp Müller备注:Accepted for publication at LREC-COLING'24摘要:动机访谈中准确的话语分类对于自动理解客户-治疗师互动的质量和动态至关重要,并且它可以作为调解此类互动的系统的关键输入。动机性面试表现出三个重要特征。首先,有两个不同的角色,即客户和治疗师。第二,它们往往充满了高度的情感,这可以通过文本和韵律来表达。最后,语境对任何给定的话语进行分类都是至关重要的。以前的作品没有充分地将所有这些特征纳入心理健康对话的话语分类方法。相比之下,我们提出了M3 TCM,一个多模态,多任务上下文模型的话语分类。我们的方法首次采用多任务学习来有效地模拟治疗师和客户行为的联合和个人组成部分。此外,M3 TCM集成了来自文本和语音模态以及会话上下文的信息。通过我们的新方法,我们在最近推出的AnnoMI数据集上的话语分类方面优于最先进的技术,客户端的相对改善为20%,治疗师的话语分类为15%。在广泛的消融研究中,我们量化了每种贡献所带来的改善。摘要:Accurate utterance classification in motivational interviews is crucial to automatically understand the quality and dynamics of client-therapist interaction, and it can serve as a key input for systems mediating such interactions. Motivational interviews exhibit three important characteristics. First, there are two distinct roles, namely client and therapist. Second, they are often highly emotionally charged, which can be expressed both in text and in prosody. Finally, context is of central importance to classify any given utterance. Previous works did not adequately incorporate all of these characteristics into utterance classification approaches for mental health dialogues. In contrast, we present M3TCM, a Multi-modal, Multi-task Context Model for utterance classification. Our approach for the first time employs multi-task learning to effectively model both joint and individual components of therapist and client behaviour. Furthermore, M3TCM integrates information from the text and speech modality as well as the conversation context. With our novel approach, we outperform the state of the art for utterance classification on the recently introduced AnnoMI dataset with a relative improvement of 20% for the client- and by 15% for therapist utterance classification. In extensive ablation studies, we quantify the improvement resulting from each contribution.
【6】 UniAV: Unified Audio-Visual Perception for Multi-Task Video Localization标题:UniAV:多任务视频本地化的统一视听感知链接:https://arxiv.org/abs/2404.03179作者:Tiantian Geng,Teng Wang,Yanfu Zhang,Jinming Duan,Weili Guan,Feng Zheng摘要:视频定位任务旨在对视频中的特定实例进行时间定位,包括时间动作定位(TAL)、声音事件检测(SED)和视听事件定位(AVEL)。现有的方法过度专注于每个任务,忽略了这些实例经常出现在同一视频中以形成完整的视频内容的事实。在这项工作中,我们提出了UniAV,一个统一的视听感知网络,首次实现了TAL,SED和AVEL任务的联合学习。UniAV可以利用特定任务数据集中的各种数据,使模型能够跨任务和模式学习和共享互利知识。为了解决数据集(大小/域/持续时间)和不同任务特征的巨大变化所带来的挑战,我们建议对所有视频的视觉和音频模式进行统一编码,以获得通用表示,同时还设计特定于任务的专家来捕获每个任务的独特知识。此外,我们通过使用预训练的文本编码器开发了一个统一的语言感知分类器,使模型能够通过在推理过程中简单地改变提示来灵活地检测各种类型的实例和以前看不见的实例。UniAV在参数较少的情况下大幅优于其单任务同行,与ActivityNet 1.3,DESED和UnAV-100基准测试中最先进的特定任务方法相比,UniAV的性能达到了同等或更高的水平。摘要:Video localization tasks aim to temporally locate specific instances in videos, including temporal action localization (TAL), sound event detection (SED) and audio-visual event localization (AVEL). Existing methods over-specialize on each task, overlooking the fact that these instances often occur in the same video to form the complete video content. In this work, we present UniAV, a Unified Audio-Visual perception network, to achieve joint learning of TAL, SED and AVEL tasks for the first time. UniAV can leverage diverse data available in task-specific datasets, allowing the model to learn and share mutually beneficial knowledge across tasks and modalities. To tackle the challenges posed by substantial variations in datasets (size/domain/duration) and distinct task characteristics, we propose to uniformly encode visual and audio modalities of all videos to derive generic representations, while also designing task-specific experts to capture unique knowledge for each task. Besides, we develop a unified language-aware classifier by utilizing a pre-trained text encoder, enabling the model to flexibly detect various types of instances and previously unseen ones by simply changing prompts during inference. UniAV outperforms its single-task counterparts by a large margin with fewer parameters, achieving on-par or superior performances compared to state-of-the-art task-specific methods across ActivityNet 1.3, DESED and UnAV-100 benchmarks.
【7】 Mai Ho'omāuna i ka 'Ai: Language Models Improve Automatic Speech Recognition in Hawaiian标题:Mai Ho 'omāuna i ka'Ai:语言模型改善夏威夷语的自动语音识别链接:https://arxiv.org/abs/2404.03073作者:Kaavya Chaparala,Guido Zarrella,Bruce Torres Fischer,Larry Kimura,Oiwi Parker Jones摘要:在本文中,我们将大量的独立的文本数据整合到一个自动语音识别(ASR)的基础模型,耳语,以解决提高低资源的语言,夏威夷语的挑战。为了做到这一点,我们训练一个外部语言模型(LM)的夏威夷语文本约150万字。然后,我们使用LM对Whisper进行重新评分,并在手动策划的标记夏威夷数据测试集上计算单词错误率(WER)。作为基线,我们使用Whisper而不使用外部LM。实验结果表明,一个小的,但显着的改善WER时,ASR输出与夏威夷LM rescored。结果支持利用所有可用的数据在开发ASR系统的代表性不足的语言。摘要:In this paper we address the challenge of improving Automatic Speech Recognition (ASR) for a low-resource language, Hawaiian, by incorporating large amounts of independent text data into an ASR foundation model, Whisper. To do this, we train an external language model (LM) on ~1.5M words of Hawaiian text. We then use the LM to rescore Whisper and compute word error rates (WERs) on a manually curated test set of labeled Hawaiian data. As a baseline, we use Whisper without an external LM. Experimental results reveal a small but significant improvement in WER when ASR outputs are rescored with a Hawaiian LM. The results support leveraging all available data in the development of ASR systems for underrepresented languages.