今天跟大家分享一篇语音相关的论文合集:cs.SD语音6篇,eess.AS音频处理6篇。

cs.SD语音

【1】 Generative Extraction of Audio Classifiers for Speaker Identification

标题: 用于说话人识别的生成式音频分类器提取

链接:https://arxiv.org/abs/2207.12816

作者:Tejumade Afonja,Lucas Bourtoule,Varun Chandrasekaran,Sageev Oore,Nicolas Papernot
机构:Work done while an intern at the University of Toronto and Vector Institute, †Work done while a graduate student at the University of Toronto and Vector Institute, ‡University of Wisconsin-Madison, §Dalhousie University and Vector Institute
摘要:机器学习模型,尤其是深度神经网络,特别容易受到攻击,这也许不再令人惊讶,其中一个已经被充分研究的漏洞是模型提取:攻击者试图通过训练代理模型模仿受害者模型的决策边界来窃取受害者模型的一种现象.以前的工作已经证明了这种攻击的有效性及其破坏性后果,但是这些工作中的大部分主要是针对图像和文本处理任务进行的。我们的工作是对{\emaudio classification models}执行模型提取的首次尝试。我们的动机是攻击者的目标是模仿受害者模型的行为,该受害者模型被训练来识别说话人。这在安全敏感领域(如生物统计学认证)中尤其成问题。我们发现,先前的模型提取技术,其中攻击者\t {天真地}使用代理数据集来攻击潜在受害者的模型,失败。因此,我们建议使用生成式模型来创建足够大且多样的合成攻击查询池。我们发现,我们的方法能够使用基于\textttt {VoxCeleb}的代理数据集合成的查询来提取\textttt {LibriSpeech}上训练的受害者模型;我们在300万次查询的预算下获得了84.41%的测试准确率。
摘要:It is perhaps no longer surprising that machine learning models, especially deep neural networks, are particularly vulnerable to attacks. One such vulnerability that has been well studied is model extraction: a phenomenon in which the attacker attempts to steal a victim's model by training a surrogate model to mimic the decision boundaries of the victim model. Previous works have demonstrated the effectiveness of such an attack and its devastating consequences, but much of this work has been done primarily for image and text processing tasks. Our work is the first attempt to perform model extraction on {\em audio classification models}. We are motivated by an attacker whose goal is to mimic the behavior of the victim's model trained to identify a speaker. This is particularly problematic in security-sensitive domains such as biometric authentication. We find that prior model extraction techniques, where the attacker \textit{naively} uses a proxy dataset to attack a potential victim's model, fail. We therefore propose the use of a generative model to create a sufficiently large and diverse pool of synthetic attack queries. We find that our approach is able to extract a victim's model trained on \texttt{LibriSpeech} using queries synthesized with a proxy dataset based off of \texttt{VoxCeleb}; we achieve a test accuracy of 84.41\% with a budget of 3 million queries.


【2】 Distinguishing between pre- and post-treatment in the speech of patients  with chronic obstructive pulmonary disease

标题: 慢性阻塞性肺疾病患者治疗前后言语行为的差异

链接:https://arxiv.org/abs/2207.12784

作者:Andreas Triantafyllopoulos,Markus Fendler,Anton Batliner,Maurice Gerczuk,Shahin Amiriparian,Thomas M. Berghaus,Björn W. Schuller
机构:Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany, GLAM – Group on Language, Audio, & Music, Imperial College, UK, Department of Cardiology, Respiratory Medicine and Intensive Care, University Hospital
备注:Accepted in INTERSPEECH 2022
摘要:慢性阻塞性肺疾病(COPD)引起肺部炎症和气流阻塞,导致多种呼吸道症状; COPD也是导致死亡的主要原因,影响着全世界数百万人。患者经常需要治疗和住院,而目前尚无治愈方法。由于COPD主要影响呼吸系统,言语和非语言发声是衡量治疗效果的主要途径。在本研究中,我们展示了20例COPD患者的新数据集的结果,显示:通过使用经由说话者级特征归一化的个性化,我们可以在(嵌套)留一说话者排除交叉验证中以高达82%的未加权平均召回率(UAR)区分处理前和处理后语音。我们进一步识别最重要的特征并将它们与病态语音特性相关联,从而实现对治疗效果的听觉解释。基于这些方法的监测工具可帮助客观化COPD患者的临床状态,并促进个性化治疗计划。
摘要:Chronic obstructive pulmonary disease (COPD) causes lung inflammation and airflow blockage leading to a variety of respiratory symptoms; it is also a leading cause of death and affects millions of individuals around the world. Patients often require treatment and hospitalisation, while no cure is currently available. As COPD predominantly affects the respiratory system, speech and non-linguistic vocalisations present a major avenue for measuring the effect of treatment. In this work, we present results on a new COPD dataset of 20 patients, showing that, by employing personalisation through speaker-level feature normalisation, we can distinguish between pre- and post-treatment speech with an unweighted average recall (UAR) of up to 82\,\% in (nested) leave-one-speaker-out cross-validation. We further identify the most important features and link them to pathological voice properties, thus enabling an auditory interpretation of treatment effects. Monitoring tools based on such approaches may help objectivise the clinical status of COPD patients and facilitate personalised treatment plans.


【3】 An exhaustive variable selection study for linear models of soundscape  emotions: rankings and Gibbs analysis

标题: 音景情感线性模型的详尽变量选择研究:排序和吉布斯分析

链接:https://arxiv.org/abs/2207.12743

作者:R. San Millán-Castillo,L. Martino,E. Morgado,F. Llorente
机构:∗ Dep. of Signal Theory and Communications, Universidad Rey Juan Carlos (URJC), Madrid, Spain., † Dep. of Statistics, Universidad Carlos III de Madrid (UC,M), Madrid, Spain.
备注:None
摘要:近十年来,声景已成为声学研究中最活跃的课题之一,它提供了一种研究声环境的整体方法,涉及到人类的感知和语境,声景引发的情感是最核心的、最微妙的和不被注意的目前,声景情感识别是一个非常活跃的研究课题,我们提供了一个详尽的变量选择研究我们考虑两个声景描述符的线性声景情感模型:激发和效价。   在此基础上,提出了一种基于Gibbs抽样的特征选择方法,该方法能更全面、更清晰地反映出各变量之间的相关性.最后,通过实验验证了该方法的有效性我们还将我们的结果与基于p值的经典方法所得到的分析进行了比较,作为我们研究的结果,我们提出了两个简单而节约的线性模型,只有7个和16个变量(在122个可能的特征中),用于两个输出建议的线性模型提供了非常好的和有竞争力的性能,$R^2〉0.86$和$R^2〉0.63 $(交叉验证程序后获得的值)。
摘要:In the last decade, soundscapes have become one of the most active topics in Acoustics, providing a holistic approach to the acoustic environment, which involves human perception and context. Soundscapes-elicited emotions are central and substantially subtle and unnoticed (compared to speech or music). Currently, soundscape emotion recognition is a very active topic in the literature. We provide an exhaustive variable selection study (i.e., a selection of the soundscapes indicators) to a well-known dataset (emo-soundscapes). We consider linear soundscape emotion models for two soundscapes descriptors: arousal and valence.  Several ranking schemes and procedures for selecting the number of variables are applied. We have also performed an alternating optimization scheme for obtaining the best sequences keeping fixed a certain number of features. Furthermore, we have designed a novel technique based on Gibbs sampling, which provides a more complete and clear view of the relevance of each variable. Finally, we have also compared our results with the analysis obtained by the classical methods based on p-values. As a result of our study, we suggest two simple and parsimonious linear models of only 7 and 16 variables (within the 122 possible features) for the two outputs (arousal and valence), respectively. The suggested linear models provide very good and competitive performance, with $R^2>0.86$ and $R^2>0.63$ (values obtained after a cross-validation procedure), respectively.


【4】 Assessment of a cost-effective headphone calibration procedure for  soundscape evaluations

标题: 用于声景评价的成本有效的耳机校准程序的评估

链接:https://arxiv.org/abs/2207.12899

作者:Bhan Lam,Kenneth Ooi,Zhen-Ting Ong,Karn N. Watcharasupat,Trevor Wong,Woon-Seng Gan
机构:School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
备注:For 24th International Congress on Acoustics
摘要:为了提高声景标准的可用性和采用率,作为全球声景属性转换项目(SATP)的一部分,提出了一种通过耳机再现音频刺激的低成本校准方法,以验证ISO/TS12913-2:2018感知情感质量(PAQ)属性翻译。先前的初步研究显示,使用开路电压(OCV)校准程序,与预期的等效连续A加权声压级($L_{\text{A,eq}}$)存在显著偏差。为了更全面地以人为中心的观点,本文在心理声学参数方面进一步研究OCV方法,包括相关的超越水平来解释SATP中相同的27个刺激的时间效应,并通过36名被试的被试内实验来检验OCV校准对ISO/TS 12913-2:2018年,Bland-Altman对客观指标的分析显示,OCV方法在所有加权声级和响度指标上存在较大偏差;和\SI{10}{\%}超过水平下的粗糙度指标.在约\SI{20}{\%}的刺激中观察到由于OCV方法引起的显著知觉差异,这与有偏见的声学指标不明显对应.然而,对由于小样本和不成对样本引起的客观和知觉差异的谨慎解释为进一步研究提供了依据.
摘要:To increase the availability and adoption of the soundscape standard, a low-cost calibration procedure for reproduction of audio stimuli over headphones was proposed as part of the global ``Soundscape Attributes Translation Project'' (SATP) for validating ISO/TS~12913-2:2018 perceived affective quality (PAQ) attribute translations. A previous preliminary study revealed significant deviations from the intended equivalent continuous A-weighted sound pressure levels ($L_{\text{A,eq}}$) using the open-circuit voltage (OCV) calibration procedure. For a more holistic human-centric perspective, the OCV method is further investigated here in terms of psychoacoustic parameters, including relevant exceedance levels to account for temporal effects on the same 27 stimuli from the SATP. Moreover, a within-subjects experiment with 36 participants was conducted to examine the effects of OCV calibration on the PAQ attributes in ISO/TS~12913-2:2018. Bland-Altman analysis of the objective indicators revealed large biases in the OCV method across all weighted sound level and loudness indicators; and roughness indicators at \SI{5}{\%} and \SI{10}{\%} exceedance levels. Significant perceptual differences due to the OCV method were observed in about \SI{20}{\%} of the stimuli, which did not correspond clearly with the biased acoustic indicators. A cautioned interpretation of the objective and perceptual differences due to small and unpaired samples nevertheless provide grounds for further investigation.


【5】 Multimodal Speech Emotion Recognition using Cross Attention with Aligned  Audio and Text

标题: 基于交叉注意的多模态语音情感识别

链接:https://arxiv.org/abs/2207.12895

作者:Yoonhyung Lee,Seunghyun Yoon,Kyomin Jung
机构:Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea
备注:None
摘要:本文提出了一种新的语音情感识别模型——交叉注意网络(CAN),它使用对齐的音频和文本信号作为输入。它的灵感来自于人类将语音识别为同时产生的声音和文本信号的组合这一事实。首先,我们的方法以对齐的方式将音频和底层文本信号分割成相等数目的步骤,使得顺序信号的相同时间步骤在交叉注意中,通过对每个通道应用全局注意机制,独立地聚合每个通道的注意力,然后将每个通道的注意力权重以交叉的方式直接应用到另一个通道,在标准IEMOCAP数据集上进行的实验表明,该模型的加权和未加权准确率分别比现有系统提高了2.66%和3.18%.
摘要:In this paper, we propose a novel speech emotion recognition model called Cross Attention Network (CAN) that uses aligned audio and text signals as inputs. It is inspired by the fact that humans recognize speech as a combination of simultaneously produced acoustic and textual signals. First, our method segments the audio and the underlying text signals into equal number of steps in an aligned way so that the same time steps of the sequential signals cover the same time span in the signals. Together with this technique, we apply the cross attention to aggregate the sequential information from the aligned signals. In the cross attention, each modality is aggregated independently by applying the global attention mechanism onto each modality. Then, the attention weights of each modality are applied directly to the other modality in a crossed way, so that the CAN gathers the audio and text information from the same time steps based on each modality. In the experiments conducted on the standard IEMOCAP dataset, our model outperforms the state-of-the-art systems by 2.66% and 3.18% relatively in terms of the weighted and unweighted accuracy.


【6】 Implementation Of Tiny Machine Learning Models On Arduino 33 BLE For  Gesture And Speech Recognition

标题: 在Arduino 33 BLE上实现用于手势和语音识别的微型机器学习模型

链接:https://arxiv.org/abs/2207.12866

作者:Viswanatha V,Ramachandra A. C,Raghavendra Prasanna,Prem Chowdary Kakarla,Viveka Simha PJ,Nishant Mohan
机构:Asst.Professor, Electronics and Communication Engineering Department, Nitte Meenakshi Institute of Technology, Bangalore., Ramachandra A.C, Professor and Head, Electronics and Communication Engineering Department
摘要:在这篇文章中,手势识别和语音识别应用程序实现了嵌入式系统与微型机器学习(TinyML)。它具有3轴加速度计、3轴陀螺仪和3轴磁力计。手势识别,提供了一种创新的非语言交流方式。它在人机交互和手语方面有着广泛的应用。这里在实现手势识别时,利用EdgeImpulse框架训练和部署TinyML模型进行手势识别,Arduino Nano 33 BLE设备具有6轴IMU,可以确定手的移动方向。语音是一种通信模式。语音识别是计算机理解人类语音的语句或命令并作出相应反应的一种方法。语音识别的主要目的是实现人与机器之间的交流,在语音识别的实现中,从语音识别的EdgeImpulse框架出发,根据人发出的关键词训练和部署TinyML模型,Arduino Nano 33 BLE设备内置麦克风,可以根据发音关键字发出红色、绿色或蓝色的RGB LED发光。每个应用的结果都在结果部分列出,并对结果进行分析。
摘要:In this article gesture recognition and speech recognition applications are implemented on embedded systems with Tiny Machine Learning (TinyML). It features 3-axis accelerometer, 3-axis gyroscope and 3-axis magnetometer. The gesture recognition,provides an innovative approach nonverbal communication. It has wide applications in human-computer interaction and sign language. Here in the implementation of hand gesture recognition, TinyML model is trained and deployed from EdgeImpulse framework for hand gesture recognition and based on the hand movements, Arduino Nano 33 BLE device having 6-axis IMU can find out the direction of movement of hand. The Speech is a mode of communication. Speech recognition is a way by which the statements or commands of human speech is understood by the computer which reacts accordingly. The main aim of speech recognition is to achieve communication between man and machine. Here in the implementation of speech recognition, TinyML model is trained and deployed from EdgeImpulse framework for speech recognition and based on the keywords pronounced by human, Arduino Nano 33 BLE device having built-in microphone can make an RGB LED glow like red, green or blue based on keyword pronounced. The results of each application are obtained and listed in the results section and given the analysis upon the results.


eess.AS音频处理

【1】 Assessment of a cost-effective headphone calibration procedure for  soundscape evaluations

标题: 用于声景评价的成本有效的耳机校准程序的评估

链接:https://arxiv.org/abs/2207.12899

* 与cs.SD语音【4】为同一篇

作者:Bhan Lam,Kenneth Ooi,Zhen-Ting Ong,Karn N. Watcharasupat,Trevor Wong,Woon-Seng Gan
机构:School of Electrical and Electronic Engineering, Nanyang Technological University, Singapore
备注:For 24th International Congress on Acoustics
摘要:为了提高声景标准的可用性和采用率,作为全球声景属性转换项目(SATP)的一部分,提出了一种通过耳机再现音频刺激的低成本校准方法,以验证ISO/TS12913-2:2018感知情感质量(PAQ)属性翻译。先前的初步研究显示,使用开路电压(OCV)校准程序,与预期的等效连续A加权声压级($L_{\text{A,eq}}$)存在显著偏差。为了更全面地以人为中心的观点,本文在心理声学参数方面进一步研究OCV方法,包括相关的超越水平来解释SATP中相同的27个刺激的时间效应,并通过36名被试的被试内实验来检验OCV校准对ISO/TS 12913-2:2018年,Bland-Altman对客观指标的分析显示,OCV方法在所有加权声级和响度指标上存在较大偏差;和\SI{10}{\%}超过水平下的粗糙度指标.在约\SI{20}{\%}的刺激中观察到由于OCV方法引起的显著知觉差异,这与有偏见的声学指标不明显对应.然而,对由于小样本和不成对样本引起的客观和知觉差异的谨慎解释为进一步研究提供了依据.
摘要:To increase the availability and adoption of the soundscape standard, a low-cost calibration procedure for reproduction of audio stimuli over headphones was proposed as part of the global ``Soundscape Attributes Translation Project'' (SATP) for validating ISO/TS~12913-2:2018 perceived affective quality (PAQ) attribute translations. A previous preliminary study revealed significant deviations from the intended equivalent continuous A-weighted sound pressure levels ($L_{\text{A,eq}}$) using the open-circuit voltage (OCV) calibration procedure. For a more holistic human-centric perspective, the OCV method is further investigated here in terms of psychoacoustic parameters, including relevant exceedance levels to account for temporal effects on the same 27 stimuli from the SATP. Moreover, a within-subjects experiment with 36 participants was conducted to examine the effects of OCV calibration on the PAQ attributes in ISO/TS~12913-2:2018. Bland-Altman analysis of the objective indicators revealed large biases in the OCV method across all weighted sound level and loudness indicators; and roughness indicators at \SI{5}{\%} and \SI{10}{\%} exceedance levels. Significant perceptual differences due to the OCV method were observed in about \SI{20}{\%} of the stimuli, which did not correspond clearly with the biased acoustic indicators. A cautioned interpretation of the objective and perceptual differences due to small and unpaired samples nevertheless provide grounds for further investigation.


【2】 Multimodal Speech Emotion Recognition using Cross Attention with Aligned  Audio and Text

标题: 基于交叉注意的多模态语音情感识别

链接:https://arxiv.org/abs/2207.12895

* 与cs.SD语音【5】为同一篇

作者:Yoonhyung Lee,Seunghyun Yoon,Kyomin Jung
机构:Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea
备注:None
摘要:本文提出了一种新的语音情感识别模型——交叉注意网络(CAN),它使用对齐的音频和文本信号作为输入。它的灵感来自于人类将语音识别为同时产生的声音和文本信号的组合这一事实。首先,我们的方法以对齐的方式将音频和底层文本信号分割成相等数目的步骤,使得顺序信号的相同时间步骤在交叉注意中,通过对每个通道应用全局注意机制,独立地聚合每个通道的注意力,然后将每个通道的注意力权重以交叉的方式直接应用到另一个通道,在标准IEMOCAP数据集上进行的实验表明,该模型的加权和未加权准确率分别比现有系统提高了2.66%和3.18%.
摘要:In this paper, we propose a novel speech emotion recognition model called Cross Attention Network (CAN) that uses aligned audio and text signals as inputs. It is inspired by the fact that humans recognize speech as a combination of simultaneously produced acoustic and textual signals. First, our method segments the audio and the underlying text signals into equal number of steps in an aligned way so that the same time steps of the sequential signals cover the same time span in the signals. Together with this technique, we apply the cross attention to aggregate the sequential information from the aligned signals. In the cross attention, each modality is aggregated independently by applying the global attention mechanism onto each modality. Then, the attention weights of each modality are applied directly to the other modality in a crossed way, so that the CAN gathers the audio and text information from the same time steps based on each modality. In the experiments conducted on the standard IEMOCAP dataset, our model outperforms the state-of-the-art systems by 2.66% and 3.18% relatively in terms of the weighted and unweighted accuracy.


【3】 Implementation Of Tiny Machine Learning Models On Arduino 33 BLE For  Gesture And Speech Recognition

标题: 在Arduino 33 BLE上实现用于手势和语音识别的微型机器学习模型

链接:https://arxiv.org/abs/2207.12866

* 与cs.SD语音【6】为同一篇

作者:Viswanatha V,Ramachandra A. C,Raghavendra Prasanna,Prem Chowdary Kakarla,Viveka Simha PJ,Nishant Mohan
机构:Asst.Professor, Electronics and Communication Engineering Department, Nitte Meenakshi Institute of Technology, Bangalore., Ramachandra A.C, Professor and Head, Electronics and Communication Engineering Department
摘要:在这篇文章中,手势识别和语音识别应用程序实现了嵌入式系统与微型机器学习(TinyML)。它具有3轴加速度计、3轴陀螺仪和3轴磁力计。手势识别,提供了一种创新的非语言交流方式。它在人机交互和手语方面有着广泛的应用。这里在实现手势识别时,利用EdgeImpulse框架训练和部署TinyML模型进行手势识别,Arduino Nano 33 BLE设备具有6轴IMU,可以确定手的移动方向。语音是一种通信模式。语音识别是计算机理解人类语音的语句或命令并作出相应反应的一种方法。语音识别的主要目的是实现人与机器之间的交流,在语音识别的实现中,从语音识别的EdgeImpulse框架出发,根据人发出的关键词训练和部署TinyML模型,Arduino Nano 33 BLE设备内置麦克风,可以根据发音关键字发出红色、绿色或蓝色的RGB LED发光。每个应用的结果都在结果部分列出,并对结果进行分析。
摘要:In this article gesture recognition and speech recognition applications are implemented on embedded systems with Tiny Machine Learning (TinyML). It features 3-axis accelerometer, 3-axis gyroscope and 3-axis magnetometer. The gesture recognition,provides an innovative approach nonverbal communication. It has wide applications in human-computer interaction and sign language. Here in the implementation of hand gesture recognition, TinyML model is trained and deployed from EdgeImpulse framework for hand gesture recognition and based on the hand movements, Arduino Nano 33 BLE device having 6-axis IMU can find out the direction of movement of hand. The Speech is a mode of communication. Speech recognition is a way by which the statements or commands of human speech is understood by the computer which reacts accordingly. The main aim of speech recognition is to achieve communication between man and machine. Here in the implementation of speech recognition, TinyML model is trained and deployed from EdgeImpulse framework for speech recognition and based on the keywords pronounced by human, Arduino Nano 33 BLE device having built-in microphone can make an RGB LED glow like red, green or blue based on keyword pronounced. The results of each application are obtained and listed in the results section and given the analysis upon the results.


【4】 Generative Extraction of Audio Classifiers for Speaker Identification

标题: 用于说话人识别的生成式音频分类器提取

链接:https://arxiv.org/abs/2207.12816

* 与cs.SD语音【1】为同一篇

作者:Tejumade Afonja,Lucas Bourtoule,Varun Chandrasekaran,Sageev Oore,Nicolas Papernot
机构:Work done while an intern at the University of Toronto and Vector Institute, †Work done while a graduate student at the University of Toronto and Vector Institute, ‡University of Wisconsin-Madison, §Dalhousie University and Vector Institute
摘要:机器学习模型,尤其是深度神经网络,特别容易受到攻击,这也许不再令人惊讶,其中一个已经被充分研究的漏洞是模型提取:攻击者试图通过训练代理模型模仿受害者模型的决策边界来窃取受害者模型的一种现象.以前的工作已经证明了这种攻击的有效性及其破坏性后果,但是这些工作中的大部分主要是针对图像和文本处理任务进行的。我们的工作是对{\emaudio classification models}执行模型提取的首次尝试。我们的动机是攻击者的目标是模仿受害者模型的行为,该受害者模型被训练来识别说话人。这在安全敏感领域(如生物统计学认证)中尤其成问题。我们发现,先前的模型提取技术,其中攻击者\t {天真地}使用代理数据集来攻击潜在受害者的模型,失败。因此,我们建议使用生成式模型来创建足够大且多样的合成攻击查询池。我们发现,我们的方法能够使用基于\textttt {VoxCeleb}的代理数据集合成的查询来提取\textttt {LibriSpeech}上训练的受害者模型;我们在300万次查询的预算下获得了84.41%的测试准确率。
摘要:It is perhaps no longer surprising that machine learning models, especially deep neural networks, are particularly vulnerable to attacks. One such vulnerability that has been well studied is model extraction: a phenomenon in which the attacker attempts to steal a victim's model by training a surrogate model to mimic the decision boundaries of the victim model. Previous works have demonstrated the effectiveness of such an attack and its devastating consequences, but much of this work has been done primarily for image and text processing tasks. Our work is the first attempt to perform model extraction on {\em audio classification models}. We are motivated by an attacker whose goal is to mimic the behavior of the victim's model trained to identify a speaker. This is particularly problematic in security-sensitive domains such as biometric authentication. We find that prior model extraction techniques, where the attacker \textit{naively} uses a proxy dataset to attack a potential victim's model, fail. We therefore propose the use of a generative model to create a sufficiently large and diverse pool of synthetic attack queries. We find that our approach is able to extract a victim's model trained on \texttt{LibriSpeech} using queries synthesized with a proxy dataset based off of \texttt{VoxCeleb}; we achieve a test accuracy of 84.41\% with a budget of 3 million queries.


【5】 Distinguishing between pre- and post-treatment in the speech of patients  with chronic obstructive pulmonary disease

标题: 慢性阻塞性肺疾病患者治疗前后言语行为的差异

链接:https://arxiv.org/abs/2207.12784

* 与cs.SD语音【2】为同一篇

作者:Andreas Triantafyllopoulos,Markus Fendler,Anton Batliner,Maurice Gerczuk,Shahin Amiriparian,Thomas M. Berghaus,Björn W. Schuller
机构:Chair of Embedded Intelligence for Health Care and Wellbeing, University of Augsburg, Germany, GLAM – Group on Language, Audio, & Music, Imperial College, UK, Department of Cardiology, Respiratory Medicine and Intensive Care, University Hospital
备注:Accepted in INTERSPEECH 2022
摘要:慢性阻塞性肺疾病(COPD)引起肺部炎症和气流阻塞,导致多种呼吸道症状; COPD也是导致死亡的主要原因,影响着全世界数百万人。患者经常需要治疗和住院,而目前尚无治愈方法。由于COPD主要影响呼吸系统,言语和非语言发声是衡量治疗效果的主要途径。在本研究中,我们展示了20例COPD患者的新数据集的结果,显示:通过使用经由说话者级特征归一化的个性化,我们可以在(嵌套)留一说话者排除交叉验证中以高达82%的未加权平均召回率(UAR)区分处理前和处理后语音。我们进一步识别最重要的特征并将它们与病态语音特性相关联,从而实现对治疗效果的听觉解释。基于这些方法的监测工具可帮助客观化COPD患者的临床状态,并促进个性化治疗计划。
摘要:Chronic obstructive pulmonary disease (COPD) causes lung inflammation and airflow blockage leading to a variety of respiratory symptoms; it is also a leading cause of death and affects millions of individuals around the world. Patients often require treatment and hospitalisation, while no cure is currently available. As COPD predominantly affects the respiratory system, speech and non-linguistic vocalisations present a major avenue for measuring the effect of treatment. In this work, we present results on a new COPD dataset of 20 patients, showing that, by employing personalisation through speaker-level feature normalisation, we can distinguish between pre- and post-treatment speech with an unweighted average recall (UAR) of up to 82\,\% in (nested) leave-one-speaker-out cross-validation. We further identify the most important features and link them to pathological voice properties, thus enabling an auditory interpretation of treatment effects. Monitoring tools based on such approaches may help objectivise the clinical status of COPD patients and facilitate personalised treatment plans.


【6】 An exhaustive variable selection study for linear models of soundscape  emotions: rankings and Gibbs analysis

标题: 音景情感线性模型的详尽变量选择研究:排序和吉布斯分析

链接:https://arxiv.org/abs/2207.12743

* 与cs.SD语音【3】为同一篇

作者:R. San Millán-Castillo,L. Martino,E. Morgado,F. Llorente
机构:∗ Dep. of Signal Theory and Communications, Universidad Rey Juan Carlos (URJC), Madrid, Spain., † Dep. of Statistics, Universidad Carlos III de Madrid (UC,M), Madrid, Spain.
备注:None
摘要:近十年来,声景已成为声学研究中最活跃的课题之一,它提供了一种研究声环境的整体方法,涉及到人类的感知和语境,声景引发的情感是最核心的、最微妙的和不被注意的目前,声景情感识别是一个非常活跃的研究课题,我们提供了一个详尽的变量选择研究我们考虑两个声景描述符的线性声景情感模型:激发和效价。   在此基础上,提出了一种基于Gibbs抽样的特征选择方法,该方法能更全面、更清晰地反映出各变量之间的相关性.最后,通过实验验证了该方法的有效性我们还将我们的结果与基于p值的经典方法所得到的分析进行了比较,作为我们研究的结果,我们提出了两个简单而节约的线性模型,只有7个和16个变量(在122个可能的特征中),用于两个输出建议的线性模型提供了非常好的和有竞争力的性能,$R^2〉0.86$和$R^2〉0.63 $(交叉验证程序后获得的值)。
摘要:In the last decade, soundscapes have become one of the most active topics in Acoustics, providing a holistic approach to the acoustic environment, which involves human perception and context. Soundscapes-elicited emotions are central and substantially subtle and unnoticed (compared to speech or music). Currently, soundscape emotion recognition is a very active topic in the literature. We provide an exhaustive variable selection study (i.e., a selection of the soundscapes indicators) to a well-known dataset (emo-soundscapes). We consider linear soundscape emotion models for two soundscapes descriptors: arousal and valence.  Several ranking schemes and procedures for selecting the number of variables are applied. We have also performed an alternating optimization scheme for obtaining the best sequences keeping fixed a certain number of features. Furthermore, we have designed a novel technique based on Gibbs sampling, which provides a more complete and clear view of the relevance of each variable. Finally, we have also compared our results with the analysis obtained by the classical methods based on p-values. As a result of our study, we suggest two simple and parsimonious linear models of only 7 and 16 variables (within the 122 possible features) for the two outputs (arousal and valence), respectively. The suggested linear models provide very good and competitive performance, with $R^2>0.86$ and $R^2>0.63$ (values obtained after a cross-validation procedure), respectively.


机器翻译,仅供参考