本文经arXiv每日学术速递授权转载
【1】 Contrasting Deep Learning Models for Direct Respiratory Insufficiency Detection Versus Blood Oxygen Saturation Estimation
标题: 用于直接呼吸功能不全检测与血氧饱和度估计的深度学习模型对比
作者:Marcelo Matheus Gauy,Natalia Hitomi Koza,Ricardo Mikio Morita,Gabriel Rocha Stanzione,Arnaldo Candido Junior,Larissa Cristina Berti,Anna Sara Shafferman Levin,Ester Cerdeira Sabino,Flaviane Romani Fernandes Svartman,Marcelo Finger
备注:23 pages, 4 figures, in review at Journal of Biomedical Signal Processing and Control
链接:点击下载PDF文件
【2】 MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
标题: MMTrail:具有语言和音乐描述的多模式预告片视频数据集
作者:Xiaowei Chi,Yatian Wang,Aosong Cheng,Pengjun Fang,Zeyue Tian,Yingqing He,Zhaoyang Liu,Xingqun Qi,Jiahao Pan,Rongyu Zhang,Mengfei Li,Ruibin Yuan,Yanbing Jiang,Wei Xue,Wenhan Luo,Qifeng Chen,Shanghang Zhang,Qifeng Liu,Yike Guo
备注:15 Pages. Dataset report
链接:点击下载PDF文件
【3】 Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
标题: 通过两阶段解纠缠和功能表示的描述驱动的钢琴音乐生成
作者:Jingyue Huang,Ke Chen,Yi-Hsuan Yang
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
【4】 PiCoGen: Generate Piano Covers with a Two-stage Approach
标题: PiCoGen:用两阶段方法生成钢琴封面
作者:Chih-Pin Tan,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Published at ICMR 2024 (project page: this https URL)
链接:点击下载PDF文件
【5】 Abusive Speech Detection in Indic Languages Using Acoustic Features
标题: 利用声学特征进行印度语言中的辱骂语音检测
作者:Anika A. Spiesberger,Andreas Triantafyllopoulos,Iosif Tsangko,Björn W. Schuller
Journal-ref:Proc. INTERSPEECH 2023, 2683-2687
链接:点击下载PDF文件
【6】 Decoding Linguistic Representations of Human Brain
标题: 解码人脑的语言表达
作者:Yu Wang,Heyang Liu,Yuhao Wang,Chuan Xuan,Yixuan Hou,Sheng Feng,Hongcheng Liu,Yusheng Liao,Yanfeng Wang
链接:点击下载PDF文件
【7】 EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
标题: EgoSonics:为无声的以自我为中心的视频生成同步音频
作者:Aashish Rai,Srinath Sridhar
备注:preprint
链接:点击下载PDF文件
【8】 DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
标题: DeepSpeech模型展示了类似人类的性能和对Costa植入物输入的处理
作者:Cynthia R. Steinhardt,Menoua Keshishian,Nima Mesgarani,Kim Stachenfeld
备注:NEURIPS preprint
链接:点击下载PDF文件
【9】 SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
标题: SuperCodec:具有选择性反投影网络的神经语音编解码器
作者:Youqiang Zheng,Weiping Tu,Li Xiao,Xinmeng Xu
备注:Accepted by ICASSP 2024
链接:点击下载PDF文件
【10】 Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
标题: Futga:通过时间增强生成增强实现细粒度音乐理解
作者:Junda Wu,Zachary Novack,Amit Namburi,Jiaheng Dai,Hao-Wen Dong,Zhouhang Xie,Carol Chen,Julian McAuley
备注:6 pages
链接:点击下载PDF文件
【11】 Integrating audiological datasets via federated merging of Auditory Profiles
标题: 通过听觉配置文件的联邦合并来集成听力学数据集
作者:Samira Saak,Dirk Oetting,Birger Kollmeier,Mareike Buhl
链接:点击下载PDF文件
标题: $Tar{a}laGen:$ A自动$T系统ar{a}la$识别和生成
作者:Rahul Bapusaheb Kodag,Himanshu Jindal,Vipul Arora
链接:点击下载PDF文件
【2】 Contrasting Deep Learning Models for Direct Respiratory Insufficiency Detection Versus Blood Oxygen Saturation Estimation
标题: 用于直接呼吸功能不全检测与血氧饱和度估计的深度学习模型对比
作者:Marcelo Matheus Gauy,Natalia Hitomi Koza,Ricardo Mikio Morita,Gabriel Rocha Stanzione,Arnaldo Candido Junior,Larissa Cristina Berti,Anna Sara Shafferman Levin,Ester Cerdeira Sabino,Flaviane Romani Fernandes Svartman,Marcelo Finger
备注:23 pages, 4 figures, in review at Journal of Biomedical Signal Processing and Control
链接:点击下载PDF文件
【3】 MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
标题: MMTrail:具有语言和音乐描述的多模式预告片视频数据集
作者:Xiaowei Chi,Yatian Wang,Aosong Cheng,Pengjun Fang,Zeyue Tian,Yingqing He,Zhaoyang Liu,Xingqun Qi,Jiahao Pan,Rongyu Zhang,Mengfei Li,Ruibin Yuan,Yanbing Jiang,Wei Xue,Wenhan Luo,Qifeng Chen,Shanghang Zhang,Qifeng Liu,Yike Guo
备注:15 Pages. Dataset report
链接:点击下载PDF文件
【4】 Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
标题: 通过两阶段解纠缠和功能表示的描述驱动的钢琴音乐生成
作者:Jingyue Huang,Ke Chen,Yi-Hsuan Yang
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
【5】 PiCoGen: Generate Piano Covers with a Two-stage Approach
标题: PiCoGen:用两阶段方法生成钢琴封面
作者:Chih-Pin Tan,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Published at ICMR 2024 (project page: this https URL)
链接:点击下载PDF文件
【6】 Abusive Speech Detection in Indic Languages Using Acoustic Features
标题: 利用声学特征进行印度语言中的辱骂语音检测
作者:Anika A. Spiesberger,Andreas Triantafyllopoulos,Iosif Tsangko,Björn W. Schuller
Journal-ref:Proc. INTERSPEECH 2023, 2683-2687
链接:点击下载PDF文件
【7】 Integrating audiological datasets via federated merging of Auditory Profiles
标题: 通过听觉配置文件的联邦合并来集成听力学数据集
作者:Samira Saak,Dirk Oetting,Birger Kollmeier,Mareike Buhl
链接:点击下载PDF文件
【8】 Decoding Linguistic Representations of Human Brain
标题: 解码人脑的语言表达
作者:Yu Wang,Heyang Liu,Yuhao Wang,Chuan Xuan,Yixuan Hou,Sheng Feng,Hongcheng Liu,Yusheng Liao,Yanfeng Wang
链接:点击下载PDF文件
【9】 EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
标题: EgoSonics:为无声的以自我为中心的视频生成同步音频
作者:Aashish Rai,Srinath Sridhar
备注:preprint
链接:点击下载PDF文件
【10】 DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
标题: DeepSpeech模型展示了类似人类的性能和对Costa植入物输入的处理
作者:Cynthia R. Steinhardt,Menoua Keshishian,Nima Mesgarani,Kim Stachenfeld
备注:NEURIPS preprint
链接:点击下载PDF文件
【11】 SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
标题: SuperCodec:具有选择性反投影网络的神经语音编解码器
作者:Youqiang Zheng,Weiping Tu,Li Xiao,Xinmeng Xu
备注:Accepted by ICASSP 2024
链接:点击下载PDF文件
【12】 Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
标题: Futga:通过时间增强生成增强实现细粒度音乐理解
作者:Junda Wu,Zachary Novack,Amit Namburi,Jiaheng Dai,Hao-Wen Dong,Zhouhang Xie,Carol Chen,Julian McAuley
备注:6 pages
链接:点击下载PDF文件
标题: 用于直接呼吸功能不全检测与血氧饱和度估计的深度学习模型对比
作者:Marcelo Matheus Gauy,Natalia Hitomi Koza,Ricardo Mikio Morita,Gabriel Rocha Stanzione,Arnaldo Candido Junior,Larissa Cristina Berti,Anna Sara Shafferman Levin,Ester Cerdeira Sabino,Flaviane Romani Fernandes Svartman,Marcelo Finger
备注:23 pages, 4 figures, in review at Journal of Biomedical Signal Processing and Control
链接:点击下载PDF文件
摘要:我们对比了为一般音频分类任务设计的最先进的深度学习架构的高效率,通过自动音频分析对呼吸功能不全(RI)检测和血氧饱和度(SpO 2)估计和分类进行了改进。最近,已经提出了多种深度学习架构,通过音频分析来检测COVID患者的RI,实现了95%以上的准确率和0.93以上的F1分数。RI是一种与低SpO 2水平相关的疾病,通常定义为阈值SpO 2 <92%。虽然SpO 2是RI的关键决定因素,但医生的诊断通常依赖于多种因素。这些包括呼吸频率、心率、SpO 2水平等。在这里,我们研究了用于RI检测的预训练音频神经网络(CNN 6,CNN 10和CNN 14)和掩蔽自动编码器(Audio-MAE),这些模型实现了近乎完美的准确性,超过了以前的结果。然而,对于估计SpO 2水平的回归任务,模型实现的均方根误差值超过手指血氧计的可接受临床范围3.5%。Pearson相关系数不超过0.3。由于深度学习模型在分类方面的表现优于回归,因此我们将SpO 2回归转换为SpO 2阈值二进制分类问题,阈值为92%。然而,该任务仍然产生低于0.65的F1分数。因此,音频分析提供了对患者RI状态的有价值的见解,但不能提供关于实际SpO 2水平的准确信息,这表明在当前技术下语音和言语生物标志物在医学诊断中可能有用也可能无用的领域是分离的。摘要:We contrast high effectiveness of state of the art deep learning architectures designed for general audio classification tasks, refined for respiratory insufficiency (RI) detection and blood oxygen saturation (SpO2) estimation and classification through automated audio analysis. Recently, multiple deep learning architectures have been proposed to detect RI in COVID patients through audio analysis, achieving accuracy above 95% and F1-score above 0.93. RI is a condition associated with low SpO2 levels, commonly defined as the threshold SpO2 <92%. While SpO2 serves as a crucial determinant of RI, a medical doctor's diagnosis typically relies on multiple factors. These include respiratory frequency, heart rate, SpO2 levels, among others. Here we study pretrained audio neural networks (CNN6, CNN10 and CNN14) and the Masked Autoencoder (Audio-MAE) for RI detection, where these models achieve near perfect accuracy, surpassing previous results. Yet, for the regression task of estimating SpO2 levels, the models achieve root mean square error values exceeding the accepted clinical range of 3.5% for finger oximeters. Additionally, Pearson correlation coefficients fail to surpass 0.3. As deep learning models perform better in classification than regression, we transform SpO2-regression into a SpO2-threshold binary classification problem, with a threshold of 92%. However, this task still yields an F1-score below 0.65. Thus, audio analysis offers valuable insights into a patient's RI status, but does not provide accurate information about actual SpO2 levels, indicating a separation of domains in which voice and speech biomarkers may and may not be useful in medical diagnostics under current technologies.
【2】 MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
标题: MMTrail:具有语言和音乐描述的多模式预告片视频数据集
作者:Xiaowei Chi,Yatian Wang,Aosong Cheng,Pengjun Fang,Zeyue Tian,Yingqing He,Zhaoyang Liu,Xingqun Qi,Jiahao Pan,Rongyu Zhang,Mengfei Li,Ruibin Yuan,Yanbing Jiang,Wei Xue,Wenhan Luo,Qifeng Chen,Shanghang Zhang,Qifeng Liu,Yike Guo
备注:15 Pages. Dataset report
链接:点击下载PDF文件
摘要:大量的多模态数据集在促进大型视频语言模型的成功方面发挥着重要作用。然而,目前的视频语言数据集主要为视觉帧提供文本描述,认为音频是弱相关信息。他们往往忽视了对视听关联的挖掘,导致对每一种情态的注释单调,而不是全面、准确的描述。这种忽视导致了多通道研究的困难。为了弥补这一差距,我们提出了MMTrail,一个大规模的多模态视频语言数据集,包含超过2000万个带有视觉字幕的预告片剪辑,以及200万个带有多模态字幕的高质量剪辑。预告片预览完整长度的视频作品,并整合上下文,视觉框架和背景音乐。特别是预告片主要有两个优点:(1)题材多样,内容人物类型多样,例如,电影、新闻和游戏。(2)相应的背景音乐是定制设计的,使其与视觉环境更加一致。基于这些见解,我们提出了一个系统的字幕框架,实现了超过27.1k小时的拖车视频的各种模态注释。在这里,为了确保标题保留音乐视角,同时保留视觉上下文的权威性,我们利用高级LLM自适应地合并所有注释。以这种方式,我们的MMtrail数据集可能为细粒度的大型多模态语言模型训练铺平道路。在实验中,我们在数据集上提供了评估指标和基准测试结果,证明了我们的注释的高质量及其对模型训练的有效性。摘要:Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be weakly related information. They usually overlook exploring the potential of inherent audio-visual correlation, leading to monotonous annotation within each modality instead of comprehensive and precise descriptions. Such ignorance results in the difficulty of multiple cross-modality studies. To fulfill this gap, we present MMTrail, a large-scale multi-modality video-language dataset incorporating more than 20M trailer clips with visual captions, and 2M high-quality clips with multimodal captions. Trailers preview full-length video works and integrate context, visual frames, and background music. In particular, the trailer has two main advantages: (1) the topics are diverse, and the content characters are of various types, e.g., film, news, and gaming. (2) the corresponding background music is custom-designed, making it more coherent with the visual context. Upon these insights, we propose a systemic captioning framework, achieving various modality annotations with more than 27.1k hours of trailer videos. Here, to ensure the caption retains music perspective while preserving the authority of visual context, we leverage the advanced LLM to merge all annotations adaptively. In this fashion, our MMtrail dataset potentially paves the path for fine-grained large multimodal-language model training. In experiments, we provide evaluation metrics and benchmark results on our dataset, demonstrating the high quality of our annotation and its effectiveness for model training.
【3】 Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
标题: 通过两阶段解纠缠和功能表示的描述驱动的钢琴音乐生成
作者:Jingyue Huang,Ke Chen,Yi-Hsuan Yang
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
摘要:管理情感方面仍然是自动音乐生成中的一个挑战。以前的作品旨在一次学习各种情绪,导致建模不足。本文通过两个阶段的框架来探讨钢琴演奏生成过程中情感的分离。第一阶段着重于铅片的效价建模,第二阶段通过引入性能级属性来解决唤醒建模。为了进一步捕捉塑造效价的特征(这是以前的方法较少探讨的方面),我们引入了象征音乐的一种新颖的功能表示。这种表示法旨在捕捉大小调调性的情感影响,以及音符、和弦和键签名之间的相互作用。客观和主观实验验证了我们的框架在情绪效价和唤醒建模方面的有效性。我们进一步利用我们的框架在一个新的应用程序的情绪控制,显示了广泛的潜力,在情绪驱动的音乐生成。摘要:Managing the emotional aspect remains a challenge in automatic music generation. Prior works aim to learn various emotions at once, leading to inadequate modeling. This paper explores the disentanglement of emotions in piano performance generation through a two-stage framework. The first stage focuses on valence modeling of lead sheet, and the second stage addresses arousal modeling by introducing performance-level attributes. To further capture features that shape valence, an aspect less explored by previous approaches, we introduce a novel functional representation of symbolic music. This representation aims to capture the emotional impact of major-minor tonality, as well as the interactions among notes, chords, and key signatures. Objective and subjective experiments validate the effectiveness of our framework in both emotional valence and arousal modeling. We further leverage our framework in a novel application of emotional controls, showing a broad potential in emotion-driven music generation.
【4】 PiCoGen: Generate Piano Covers with a Two-stage Approach
标题: PiCoGen:用两阶段方法生成钢琴封面
作者:Chih-Pin Tan,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Published at ICMR 2024 (project page: this https URL)
链接:点击下载PDF文件
摘要:翻唱创作是音乐创作界一种流行的音乐创作方式。在这项研究中,我们介绍了钢琴封面生成(PiCoGen),自动封面歌曲生成的两个阶段的方法,转录的旋律线和和弦进行的歌曲给定的音频记录,然后使用产生的铅表作为条件,生成一个钢琴封面的符号域。这种方法的优点在于,它不需要用于训练的封面和它们的原始歌曲的配对数据。与需要这种配对数据的现有方法相比,我们的评估表明,PiCoGen在不同音乐类型的歌曲中表现出竞争力甚至更好的性能。摘要:Cover song generation stands out as a popular way of music making in the music-creative community. In this study, we introduce Piano Cover Generation (PiCoGen), a two-stage approach for automatic cover song generation that transcribes the melody line and chord progression of a song given its audio recording, and then uses the resulting lead sheet as the condition to generate a piano cover in the symbolic domain. This approach is advantageous in that it does not required paired data of covers and their original songs for training. Compared to an existing approach that demands such paired data, our evaluation shows that PiCoGen demonstrates competitive or even superior performance across songs of different musical genres.
【5】 Abusive Speech Detection in Indic Languages Using Acoustic Features
标题: 利用声学特征进行印度语言中的辱骂语音检测
作者:Anika A. Spiesberger,Andreas Triantafyllopoulos,Iosif Tsangko,Björn W. Schuller
Journal-ref:Proc. INTERSPEECH 2023, 2683-2687
链接:点击下载PDF文件
摘要:在线社交网络中的辱骂性内容是一个众所周知的问题,可造成严重的心理伤害并煽动仇恨。上传音频数据的能力增加了开发检测语音记录中滥用内容的方法的重要性。然而,简单地从书面虐待检测转移机制会忽略相关信息,如情绪和语气。此外,许多当前的算法需要在它们所用于的特定语言中进行训练。本文提出了使用声学和韵律特征来分类滥用内容。我们使用了ADIMA数据集,其中包含来自10种印度语言的录音,并在多语言和跨语言环境中训练了不同的模型。我们的研究结果表明,它是可能的分类滥用和非滥用内容仅使用声学和韵律特征。最重要和最有影响力的功能进行了讨论。摘要:Abusive content in online social networks is a well-known problem that can cause serious psychological harm and incite hatred. The ability to upload audio data increases the importance of developing methods to detect abusive content in speech recordings. However, simply transferring the mechanisms from written abuse detection would ignore relevant information such as emotion and tone. In addition, many current algorithms require training in the specific language for which they are being used. This paper proposes to use acoustic and prosodic features to classify abusive content. We used the ADIMA data set, which contains recordings from ten Indic languages, and trained different models in multilingual and cross-lingual settings. Our results show that it is possible to classify abusive and non-abusive content using only acoustic and prosodic features. The most important and influential features are discussed.
【6】 Decoding Linguistic Representations of Human Brain
标题: 解码人脑的语言表达
作者:Yu Wang,Heyang Liu,Yuhao Wang,Chuan Xuan,Yixuan Hou,Sheng Feng,Hongcheng Liu,Yusheng Liao,Yanfeng Wang
链接:点击下载PDF文件
摘要:语言作为高级生物体创造的信息媒介,一直是神经科学关注的问题,即它在大脑中是如何表现的。由于神经成像、医疗技术、生命科学和人工智能的快速发展,在诱发脑中解码语言表征已经取得了突破性的成就。在这项工作中,我们提出了一个分类的大脑语言解码的文本和语音格式。这项工作整合了两种类型的研究:专注于语言理解的神经科学和基于深度学习的大脑解码。从大脑活动中产生可辨别的语言信息不仅可以帮助那些发音受限的人,特别是肌萎缩侧索硬化症(ALS)患者,而且还为下一代脑机接口(BCI)开辟了新的途径。本文将帮助脑科学家和深度学习研究人员对细粒度的语言感知进行鸟瞰,从而促进他们对神经过程和语言解码的进一步调查和研究。摘要:Language, as an information medium created by advanced organisms, has always been a concern of neuroscience regarding how it is represented in the brain. Decoding linguistic representations in the evoked brain has shown groundbreaking achievements, thanks to the rapid improvement of neuroimaging, medical technology, life sciences and artificial intelligence. In this work, we present a taxonomy of brain-to-language decoding of both textual and speech formats. This work integrates two types of research: neuroscience focusing on language understanding and deep learning-based brain decoding. Generating discernible language information from brain activity could not only help those with limited articulation, especially amyotrophic lateral sclerosis (ALS) patients but also open up a new way for the next generation's brain-computer interface (BCI). This article will help brain scientists and deep-learning researchers to gain a bird's eye view of fine-grained language perception, and thus facilitate their further investigation and research of neural process and language decoding.
【7】 EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
标题: EgoSonics:为无声的以自我为中心的视频生成同步音频
作者:Aashish Rai,Srinath Sridhar
备注:preprint
链接:点击下载PDF文件
摘要:我们介绍了自我声波,一种方法来生成语义上有意义的和同步的音频轨道上无声的自我为中心的视频。为以自我为中心的无声视频生成音频可以在虚拟现实、辅助技术或增强现有数据集方面开辟新的应用。现有的工作仅限于语音、音乐或冲击声等领域,无法轻松捕获以自我为中心的视频中的广泛音频频率。EgoSonics通过建立条件音频合成的潜在扩散模型的强度来解决这些限制。我们首先将音频和视频数据编码和处理为适合生成的形式。编码后的数据用于训练我们的模型,以生成捕获输入视频语义的音轨。我们提出的SyncroNet建立在ControlNet之上,提供控制信号,使时间同步的合成音频。广泛的评估表明,我们的模型优于现有的工作在音频质量,并在我们新提出的同步评估方法。此外,我们展示了我们的模型在改善视频摘要的下游应用。摘要:We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality, assistive technologies, or for augmenting existing datasets. Existing work has been limited to domains like speech, music, or impact sounds and cannot easily capture the broad range of audio frequencies found in egocentric videos. EgoSonics addresses these limitations by building on the strength of latent diffusion models for conditioned audio synthesis. We first encode and process audio and video data into a form that is suitable for generation. The encoded data is used to train our model to generate audio tracks that capture the semantics of the input video. Our proposed SyncroNet builds on top of ControlNet to provide control signals that enables temporal synchronization to the synthesized audio. Extensive evaluations show that our model outperforms existing work in audio quality, and in our newly proposed synchronization evaluation method. Furthermore, we demonstrate downstream applications of our model in improving video summarization.
【8】 DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
标题: DeepSpeech模型展示了类似人类的性能和对Costa植入物输入的处理
作者:Cynthia R. Steinhardt,Menoua Keshishian,Nima Mesgarani,Kim Stachenfeld
备注:NEURIPS preprint
链接:点击下载PDF文件
摘要:耳蜗植入物(CI)可以说是最成功的神经植入物,已经恢复了全世界100多万人的听力。虽然CI研究的重点是模拟耳蜗激活响应于低水平的声学特征,我们假设这些植入物的成功在很大程度上是由于上游网络的作用,从降级的信号中提取有用的功能,并学习语言的统计数据来解决信号。在这项工作中,我们使用深度神经网络(DNN)DeepSpeech2作为一个范例来研究自然输入和基于人工耳蜗植入的输入如何随着时间的推移而被处理。我们从口语句子中生成自然主义和耳蜗植入式输入,并在类似的音素识别测试中测试模型性能与人类性能的相似性。我们的模型再现了正常听力和CI参与者研究中的噪声条件下的反应时间和音素混淆模式的错误模式。然后,我们使用可解释性技术来确定在处理自然主义和CI类输入时何时何地出现混淆。我们发现,随着时间的推移,在每一层的动态受到上下文以及输入类型的影响。在同一时间窗口内,所有音素的动力学在混淆和理解过程中发散,在网络的每一层中时间上向后移动。在CI的处理期间存在该信号的调制,其类似于听觉流中的人类EEG信号的变化。这种减少可能与编码音素身份的减少有关。这些研究结果表明,我们有一个可行的模型,可以及时探索语音相关信息的丢失,并且我们可以在优化人工耳蜗输入时使用它来寻找人群水平的编码信号,以改善基本语音相关信息的编码并改善感知。摘要:Cochlear implants(CIs) are arguably the most successful neural implant, having restored hearing to over one million people worldwide. While CI research has focused on modeling the cochlear activations in response to low-level acoustic features, we hypothesize that the success of these implants is due in large part to the role of the upstream network in extracting useful features from a degraded signal and learned statistics of language to resolve the signal. In this work, we use the deep neural network (DNN) DeepSpeech2, as a paradigm to investigate how natural input and cochlear implant-based inputs are processed over time. We generate naturalistic and cochlear implant-like inputs from spoken sentences and test the similarity of model performance to human performance on analogous phoneme recognition tests. Our model reproduces error patterns in reaction time and phoneme confusion patterns under noise conditions in normal hearing and CI participant studies. We then use interpretability techniques to determine where and when confusions arise when processing naturalistic and CI-like inputs. We find that dynamics over time in each layer are affected by context as well as input type. Dynamics of all phonemes diverge during confusion and comprehension within the same time window, which is temporally shifted backward in each layer of the network. There is a modulation of this signal during processing of CI which resembles changes in human EEG signals in the auditory stream. This reduction likely relates to the reduction of encoded phoneme identity. These findings suggest that we have a viable model in which to explore the loss of speech-related information in time and that we can use it to find population-level encoding signals to target when optimizing cochlear implant inputs to improve encoding of essential speech-related information and improve perception.
【9】 SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
标题: SuperCodec:具有选择性反投影网络的神经语音编解码器
作者:Youqiang Zheng,Weiping Tu,Li Xiao,Xinmeng Xu
备注:Accepted by ICASSP 2024
链接:点击下载PDF文件
摘要:神经语音编码是一个快速发展的课题,其中最先进的方法现在表现出优于传统方法的压缩性能。尽管取得了重大进展,但现有方法在保留和重建精细细节以实现最佳重建方面仍然存在局限性,特别是在低比特率下。在这项研究中,我们介绍了SuperCodec,一种神经语音编解码器,在低比特率下实现了最先进的性能。它采用了一种新的反投影方法与选择性特征融合增强表示。具体来说,我们建议使用选择性上采样反向投影(SUBP)和选择性下采样反向投影(SDBP)模块来分别替换编码器和解码器处的标准上采样层和下采样层。实验结果表明,我们的方法优于现有的神经语音编解码器在不同的比特率操作。具体来说,我们提出的方法可以实现更高质量的重建语音在1 kbps比天琴座V2在3.2 kbps和Encodec在6 kbps。摘要:Neural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction, especially at low bitrates. In this study, we introduce SuperCodec, a neural speech codec that achieves state-of-the-art performance at low bitrates. It employs a novel back projection method with selective feature fusion for augmented representation. Specifically, we propose to use Selective Up-sampling Back Projection (SUBP) and Selective Down-sampling Back Projection (SDBP) modules to replace the standard up- and down-sampling layers at the encoder and decoder, respectively. Experimental results show that our method outperforms the existing neural speech codecs operating at various bitrates. Specifically, our proposed method can achieve higher quality reconstructed speech at 1 kbps than Lyra V2 at 3.2 kbps and Encodec at 6 kbps.
【10】 Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
标题: Futga:通过时间增强生成增强实现细粒度音乐理解
作者:Junda Wu,Zachary Novack,Amit Namburi,Jiaheng Dai,Hao-Wen Dong,Zhouhang Xie,Carol Chen,Julian McAuley
备注:6 pages
链接:点击下载PDF文件
摘要:现有的音乐字幕方法仅限于生成简短音乐片段的简洁全局描述,无法捕获细粒度的音乐特征和时间感知的音乐变化。为了解决这些局限性,我们提出了FUTGA,一个通过学习生成增强与时间组成,配备了细粒度的音乐理解能力的模型。我们利用现有的音乐字幕数据集和大型语言模型(LLM)来合成具有结构描述和完整歌曲时间边界的细粒度音乐字幕。通过所提出的合成数据集的增强,FUTGA能够识别音乐在关键过渡点的时间变化及其音乐功能,并为每个音乐片段生成详细的描述。我们进一步介绍了FUTGA生成的全长音乐字幕数据集,作为MusicCaps和Song Describer数据集的增强。我们评估了多个下游任务中自动生成的字幕,包括音乐生成和检索。实验结果表明,所生成的字幕的质量和更好的性能,在各种下游任务所实现的建议的音乐字幕方法。我们的代码和数据集可以在 href{https: huggingface.co JoshuaW1997 FUTGA}{ textcolor{blue}{https: huggingface.co JoshuaW1997 FUTGA}}中找到。摘要:Existing music captioning methods are limited to generating concise global descriptions of short music clips, which fail to capture fine-grained musical characteristics and time-aware musical changes. To address these limitations, we propose FUTGA, a model equipped with fined-grained music understanding capabilities through learning from generative augmentation with temporal compositions. We leverage existing music caption datasets and large language models (LLMs) to synthesize fine-grained music captions with structural descriptions and time boundaries for full-length songs. Augmented by the proposed synthetic dataset, FUTGA is enabled to identify the music's temporal changes at key transition points and their musical functions, as well as generate detailed descriptions for each music segment. We further introduce a full-length music caption dataset generated by FUTGA, as the augmentation of the MusicCaps and the Song Describer datasets. We evaluate the automatically generated captions on several downstream tasks, including music generation and retrieval. The experiments demonstrate the quality of the generated captions and the better performance in various downstream tasks achieved by the proposed music captioning approach. Our code and datasets can be found in href{https: huggingface.co JoshuaW1997 FUTGA}{ textcolor{blue}{https: huggingface.co JoshuaW1997 FUTGA}}.
【11】 Integrating audiological datasets via federated merging of Auditory Profiles
标题: 通过听觉配置文件的联邦合并来集成听力学数据集
作者:Samira Saak,Dirk Oetting,Birger Kollmeier,Mareike Buhl
链接:点击下载PDF文件
摘要:听力学数据集包含有关患者听力损失的宝贵知识,可以使用数据驱动的联邦学习技术来发现这些知识。我们以前的方法将来自一个听力学数据集的患者信息总结为不同的听觉轮廓(AP)。然而,为了覆盖完整的听力学患者人群,必须在多个单独的数据集上分析患者模式,并最终将其集成到AP的组合集合中。本研究旨在通过AP合并步骤扩展现有的配置文件生成管道,从而能够基于听力学测量的相似性组合来自不同数据集的AP。将13个先前生成的AP(NA=595)与来自第二数据集(NB=1272)的31个新生成的AP合并,使用从两个数据集上的共同特征的重叠密度导出的相似性得分。为了确保临床适用性,针对各种场景创建了随机森林模型,包括听力学测量的不同组合。提出了一个新的13个组合的AP集,提供了良好的可分离的配置文件,仍然捕捉详细的患者信息,从各种测试结果组合。这些配置文件的分类性能是令人满意的。使用响度缩放、听力图和言语测试信息的组合可以获得最佳性能,而单一测量的性能最差。增强的配置文件生成管道证明了跨数据集组合AP的可行性,这应该推广到所有数据集,并可能在未来产生可解释的基于人口的配置文件集。分类模型保持临床适用性。因此,即使只有基于智能手机的措施可用,也可以将给定患者分类到适当的AP中。摘要:Audiological datasets contain valuable knowledge about hearing loss in patients, which can be uncovered using data-driven, federated learning techniques. Our previous approach summarized patient information from one audiological dataset into distinct Auditory Profiles (APs). To cover the complete audiological patient population, however, patient patterns must be analyzed across multiple, separated datasets, and finally, be integrated into a combined set of APs. This study aimed at extending the existing profile generation pipeline with an AP merging step, enabling the combination of APs from different datasets based on their similarity across audiological measures. The 13 previously generated APs (NA=595) were merged with 31 newly generated APs from a second dataset (NB=1272) using a similarity score derived from the overlapping densities of common features across the two datasets. To ensure clinical applicability, random forest models were created for various scenarios, encompassing different combinations of audiological measures. A new set with 13 combined APs is proposed, providing well-separable profiles, which still capture detailed patient information from various test outcome combinations. The classification performance across these profiles is satisfactory. The best performance was achieved using a combination of loudness scaling, audiogram and speech test information, while single measures performed worst. The enhanced profile generation pipeline demonstrates the feasibility of combining APs across datasets, which should generalize to all datasets and could lead to an interpretable population-based profile set in the future. The classification models maintain clinical applicability. Hence, even if only smartphone-based measures are available, a given patient can be classified into an appropriate AP.
eess.AS音频处理
【1】 $Tbar{a}laGen:$ A System for Automatic $Tbar{a}la$ Identification and Generation标题: $Tar{a}laGen:$ A自动$T系统ar{a}la$识别和生成
作者:Rahul Bapusaheb Kodag,Himanshu Jindal,Vipul Arora
链接:点击下载PDF文件
摘要:在印度斯坦古典音乐中,塔布拉作为节奏骨干和伴奏发挥着重要作用。在基于计算机的音乐分析、学习唱歌和学习乐器等应用中,tabla笔画转录、$t bar{a}la$识别和生成至关重要。本文提出了一个旨在应对这些挑战的综合系统。对于Tabla笔划转录,我们提出了一种基于模型不可知元学习(MAML)的新方法,该方法可以使用最少的数据准确识别Tabla笔划。利用这些transmittance,该系统介绍了两种新的$t bar{a}la$识别方法的基础上的序列分析的tabla笔划。 par此外,本文提出了一个框架,$t bar{a}la$生成传统和现代的学习方法的桥梁。该框架利用有限状态传感器(FST)和线性时不变(LTI)滤波器,通过用户交互,增强练习会话和音乐教育来生成具有实时节奏控制的$t bar{a}las$。Tabla独奏和音乐会数据集上的实验评估表明,该系统对真实世界数据的出色性能及其优于现有方法的能力。此外,所提出的$t bar{a}la$识别方法超越了最先进的技术。本文的贡献包括一个综合的方法,以Tabla笔画转录,创新的$t bar{a}la$识别技术,和一个强大的框架,为$t bar{a}la$代处理的节奏复杂性的印度斯坦音乐。摘要:In Hindustani classical music, the tabla plays an important role as a rhythmic backbone and accompaniment. In applications like computer-based music analysis, learning singing, and learning musical instruments, tabla stroke transcription, $t bar{a}la$ identification, and generation are crucial. This paper proposes a comprehensive system aimed at addressing these challenges. For tabla stroke transcription, we propose a novel approach based on model-agnostic meta-learning (MAML) that facilitates the accurate identification of tabla strokes using minimal data. Leveraging these transcriptions, the system introduces two novel $t bar{a}la$ identification methods based on the sequence analysis of tabla strokes. par Furthermore, the paper proposes a framework for $t bar{a}la$ generation to bridge traditional and modern learning methods. This framework utilizes finite state transducers (FST) and linear time-invariant (LTI) filters to generate $t bar{a}las$ with real-time tempo control through user interaction, enhancing practice sessions and musical education. Experimental evaluations on tabla solo and concert datasets demonstrate the system's exceptional performance on real-world data and its ability to outperform existing methods. Additionally, the proposed $t bar{a}la$ identification methods surpass state-of-the-art techniques. The contributions of this paper include a combined approach to tabla stroke transcription, innovative $t bar{a}la$ identification techniques, and a robust framework for $t bar{a}la$ generation that handles the rhythmic complexities of Hindustani music.
【2】 Contrasting Deep Learning Models for Direct Respiratory Insufficiency Detection Versus Blood Oxygen Saturation Estimation
标题: 用于直接呼吸功能不全检测与血氧饱和度估计的深度学习模型对比
作者:Marcelo Matheus Gauy,Natalia Hitomi Koza,Ricardo Mikio Morita,Gabriel Rocha Stanzione,Arnaldo Candido Junior,Larissa Cristina Berti,Anna Sara Shafferman Levin,Ester Cerdeira Sabino,Flaviane Romani Fernandes Svartman,Marcelo Finger
备注:23 pages, 4 figures, in review at Journal of Biomedical Signal Processing and Control
链接:点击下载PDF文件
摘要:我们对比了为一般音频分类任务设计的最先进的深度学习架构的高效率,通过自动音频分析对呼吸功能不全(RI)检测和血氧饱和度(SpO2)估计和分类进行了改进。最近,已经提出了多种深度学习架构,通过音频分析来检测COVID患者的RI,实现了95%以上的准确率和0.93以上的F1分数。RI是一种与低SpO2水平相关的疾病,通常定义为阈值SpO2 <92%。虽然SpO2是RI的关键决定因素,但医生的诊断通常依赖于多种因素。这些包括呼吸频率、心率、SpO2水平等。在这里,我们研究了用于RI检测的预训练音频神经网络(CNN6,CNN10和CNN14)和掩蔽自动编码器(Audio-MAE),这些模型实现了近乎完美的准确性,超过了以前的结果。然而,对于估计SpO2水平的回归任务,模型实现的均方根误差值超过手指血氧计的可接受临床范围3.5%。Pearson相关系数不超过0.3。由于深度学习模型在分类方面的表现优于回归,因此我们将SpO2回归转换为SpO2阈值二进制分类问题,阈值为92%。然而,该任务仍然产生低于0.65的F1分数。因此,音频分析为患者的RI状态提供了有价值的见解,但不能提供有关实际SpO 2水平的准确信息,这表明声音和言语生物标志物在当前技术下可能对医疗诊断有用也可能不有用的领域是分离的。摘要:We contrast high effectiveness of state of the art deep learning architectures designed for general audio classification tasks, refined for respiratory insufficiency (RI) detection and blood oxygen saturation (SpO2) estimation and classification through automated audio analysis. Recently, multiple deep learning architectures have been proposed to detect RI in COVID patients through audio analysis, achieving accuracy above 95% and F1-score above 0.93. RI is a condition associated with low SpO2 levels, commonly defined as the threshold SpO2 <92%. While SpO2 serves as a crucial determinant of RI, a medical doctor's diagnosis typically relies on multiple factors. These include respiratory frequency, heart rate, SpO2 levels, among others. Here we study pretrained audio neural networks (CNN6, CNN10 and CNN14) and the Masked Autoencoder (Audio-MAE) for RI detection, where these models achieve near perfect accuracy, surpassing previous results. Yet, for the regression task of estimating SpO2 levels, the models achieve root mean square error values exceeding the accepted clinical range of 3.5% for finger oximeters. Additionally, Pearson correlation coefficients fail to surpass 0.3. As deep learning models perform better in classification than regression, we transform SpO2-regression into a SpO2-threshold binary classification problem, with a threshold of 92%. However, this task still yields an F1-score below 0.65. Thus, audio analysis offers valuable insights into a patient's RI status, but does not provide accurate information about actual SpO2 levels, indicating a separation of domains in which voice and speech biomarkers may and may not be useful in medical diagnostics under current technologies.
【3】 MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
标题: MMTrail:具有语言和音乐描述的多模式预告片视频数据集
作者:Xiaowei Chi,Yatian Wang,Aosong Cheng,Pengjun Fang,Zeyue Tian,Yingqing He,Zhaoyang Liu,Xingqun Qi,Jiahao Pan,Rongyu Zhang,Mengfei Li,Ruibin Yuan,Yanbing Jiang,Wei Xue,Wenhan Luo,Qifeng Chen,Shanghang Zhang,Qifeng Liu,Yike Guo
备注:15 Pages. Dataset report
链接:点击下载PDF文件
摘要:大量的多模态数据集在促进大型视频语言模型的成功方面发挥着重要作用。然而,目前的视频语言数据集主要为视觉帧提供文本描述,认为音频是弱相关信息。他们往往忽视了对视听关联的挖掘,导致对每一种情态的注释单调,而不是全面、准确的描述。这种忽视导致了多通道研究的困难。为了弥补这一差距,我们提出了MMTrail,一个大规模的多模态视频语言数据集,包含超过2000万个带有视觉字幕的预告片剪辑,以及200万个带有多模态字幕的高质量剪辑。预告片预览完整长度的视频作品,并整合上下文,视觉框架和背景音乐。特别是预告片主要有两个优点:(1)题材多样,内容人物类型多样,例如,电影、新闻和游戏。(2)相应的背景音乐是定制设计的,使其与视觉环境更加一致。基于这些见解,我们提出了一个系统的字幕框架,实现了超过27.1k小时的拖车视频的各种模态注释。在这里,为了确保标题保留音乐视角,同时保留视觉上下文的权威性,我们利用高级LLM自适应地合并所有注释。以这种方式,我们的MMtrail数据集可能为细粒度的大型多模态语言模型训练铺平道路。在实验中,我们在数据集上提供了评估指标和基准测试结果,证明了我们的注释的高质量及其对模型训练的有效性。摘要:Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be weakly related information. They usually overlook exploring the potential of inherent audio-visual correlation, leading to monotonous annotation within each modality instead of comprehensive and precise descriptions. Such ignorance results in the difficulty of multiple cross-modality studies. To fulfill this gap, we present MMTrail, a large-scale multi-modality video-language dataset incorporating more than 20M trailer clips with visual captions, and 2M high-quality clips with multimodal captions. Trailers preview full-length video works and integrate context, visual frames, and background music. In particular, the trailer has two main advantages: (1) the topics are diverse, and the content characters are of various types, e.g., film, news, and gaming. (2) the corresponding background music is custom-designed, making it more coherent with the visual context. Upon these insights, we propose a systemic captioning framework, achieving various modality annotations with more than 27.1k hours of trailer videos. Here, to ensure the caption retains music perspective while preserving the authority of visual context, we leverage the advanced LLM to merge all annotations adaptively. In this fashion, our MMtrail dataset potentially paves the path for fine-grained large multimodal-language model training. In experiments, we provide evaluation metrics and benchmark results on our dataset, demonstrating the high quality of our annotation and its effectiveness for model training.
【4】 Emotion-driven Piano Music Generation via Two-stage Disentanglement and Functional Representation
标题: 通过两阶段解纠缠和功能表示的描述驱动的钢琴音乐生成
作者:Jingyue Huang,Ke Chen,Yi-Hsuan Yang
备注:Proceedings of the 25th International Society for Music Information Retrieval Conference, ISMIR 2024
链接:点击下载PDF文件
摘要:管理情感方面仍然是自动音乐生成中的一个挑战。以前的作品旨在一次学习各种情绪,导致建模不足。本文通过两个阶段的框架来探讨钢琴演奏生成过程中情感的分离。第一阶段着重于铅片的效价建模,第二阶段通过引入性能级属性来解决唤醒建模。为了进一步捕捉功能,形状价,一个方面较少探索以前的方法,我们介绍了一种新的功能表示的象征性音乐。这种表示法旨在捕捉大小调调性的情感影响,以及音符、和弦和键签名之间的相互作用。客观和主观实验验证了我们的框架在情绪效价和唤醒建模方面的有效性。我们进一步利用我们的框架在一个新的应用程序的情绪控制,显示了广泛的潜力,在情绪驱动的音乐生成。摘要:Managing the emotional aspect remains a challenge in automatic music generation. Prior works aim to learn various emotions at once, leading to inadequate modeling. This paper explores the disentanglement of emotions in piano performance generation through a two-stage framework. The first stage focuses on valence modeling of lead sheet, and the second stage addresses arousal modeling by introducing performance-level attributes. To further capture features that shape valence, an aspect less explored by previous approaches, we introduce a novel functional representation of symbolic music. This representation aims to capture the emotional impact of major-minor tonality, as well as the interactions among notes, chords, and key signatures. Objective and subjective experiments validate the effectiveness of our framework in both emotional valence and arousal modeling. We further leverage our framework in a novel application of emotional controls, showing a broad potential in emotion-driven music generation.
【5】 PiCoGen: Generate Piano Covers with a Two-stage Approach
标题: PiCoGen:用两阶段方法生成钢琴封面
作者:Chih-Pin Tan,Shuen-Huei Guan,Yi-Hsuan Yang
备注:Published at ICMR 2024 (project page: this https URL)
链接:点击下载PDF文件
摘要:翻唱创作是音乐创作界一种流行的音乐创作方式。在这项研究中,我们介绍了钢琴封面生成(PiCoGen),自动封面歌曲生成的两个阶段的方法,转录的旋律线和和弦进行的歌曲给定的音频记录,然后使用产生的铅表作为条件,生成一个钢琴封面的符号域。这种方法的优点在于,它不需要用于训练的封面和它们的原始歌曲的配对数据。与需要这种配对数据的现有方法相比,我们的评估表明,PiCoGen在不同音乐类型的歌曲中表现出竞争力甚至更好的性能。摘要:Cover song generation stands out as a popular way of music making in the music-creative community. In this study, we introduce Piano Cover Generation (PiCoGen), a two-stage approach for automatic cover song generation that transcribes the melody line and chord progression of a song given its audio recording, and then uses the resulting lead sheet as the condition to generate a piano cover in the symbolic domain. This approach is advantageous in that it does not required paired data of covers and their original songs for training. Compared to an existing approach that demands such paired data, our evaluation shows that PiCoGen demonstrates competitive or even superior performance across songs of different musical genres.
【6】 Abusive Speech Detection in Indic Languages Using Acoustic Features
标题: 利用声学特征进行印度语言中的辱骂语音检测
作者:Anika A. Spiesberger,Andreas Triantafyllopoulos,Iosif Tsangko,Björn W. Schuller
Journal-ref:Proc. INTERSPEECH 2023, 2683-2687
链接:点击下载PDF文件
摘要:在线社交网络中的辱骂性内容是一个众所周知的问题,可造成严重的心理伤害并煽动仇恨。上传音频数据的能力增加了开发检测语音记录中滥用内容的方法的重要性。然而,简单地从书面虐待检测转移机制会忽略相关信息,如情绪和语气。此外,许多当前的算法需要在它们所用于的特定语言中进行训练。本文提出了使用声学和韵律特征来分类滥用内容。我们使用了ADIMA数据集,其中包含来自10种印度语言的录音,并在多语言和跨语言环境中训练了不同的模型。我们的研究结果表明,它是可能的分类滥用和非滥用内容仅使用声学和韵律特征。最重要和最有影响力的功能进行了讨论。摘要:Abusive content in online social networks is a well-known problem that can cause serious psychological harm and incite hatred. The ability to upload audio data increases the importance of developing methods to detect abusive content in speech recordings. However, simply transferring the mechanisms from written abuse detection would ignore relevant information such as emotion and tone. In addition, many current algorithms require training in the specific language for which they are being used. This paper proposes to use acoustic and prosodic features to classify abusive content. We used the ADIMA data set, which contains recordings from ten Indic languages, and trained different models in multilingual and cross-lingual settings. Our results show that it is possible to classify abusive and non-abusive content using only acoustic and prosodic features. The most important and influential features are discussed.
【7】 Integrating audiological datasets via federated merging of Auditory Profiles
标题: 通过听觉配置文件的联邦合并来集成听力学数据集
作者:Samira Saak,Dirk Oetting,Birger Kollmeier,Mareike Buhl
链接:点击下载PDF文件
摘要:听力学数据集包含有关患者听力损失的宝贵知识,可以使用数据驱动的联邦学习技术来发现这些知识。我们以前的方法将来自一个听力学数据集的患者信息总结为不同的听觉轮廓(AP)。然而,为了覆盖完整的听力学患者人群,必须在多个单独的数据集上分析患者模式,并最终将其集成到AP的组合集合中。本研究旨在通过AP合并步骤扩展现有的配置文件生成管道,从而能够基于听力学测量的相似性组合来自不同数据集的AP。将13个先前生成的AP(NA=595)与来自第二数据集(NB=1272)的31个新生成的AP合并,使用从两个数据集上的共同特征的重叠密度导出的相似性得分。为了确保临床适用性,针对各种场景创建了随机森林模型,包括听力学测量的不同组合。提出了一个新的13个组合的AP集,提供了良好的可分离的配置文件,仍然捕捉详细的患者信息,从各种测试结果组合。这些配置文件的分类性能是令人满意的。最好的性能实现了响度缩放,听力图和语音测试信息的组合,而单一的措施表现最差。增强的配置文件生成管道证明了跨数据集组合AP的可行性,这应该推广到所有数据集,并可能在未来产生可解释的基于人口的配置文件集。分类模型保持临床适用性。因此,即使只有基于智能手机的措施可用,也可以将给定患者分类到适当的AP中。摘要:Audiological datasets contain valuable knowledge about hearing loss in patients, which can be uncovered using data-driven, federated learning techniques. Our previous approach summarized patient information from one audiological dataset into distinct Auditory Profiles (APs). To cover the complete audiological patient population, however, patient patterns must be analyzed across multiple, separated datasets, and finally, be integrated into a combined set of APs. This study aimed at extending the existing profile generation pipeline with an AP merging step, enabling the combination of APs from different datasets based on their similarity across audiological measures. The 13 previously generated APs (NA=595) were merged with 31 newly generated APs from a second dataset (NB=1272) using a similarity score derived from the overlapping densities of common features across the two datasets. To ensure clinical applicability, random forest models were created for various scenarios, encompassing different combinations of audiological measures. A new set with 13 combined APs is proposed, providing well-separable profiles, which still capture detailed patient information from various test outcome combinations. The classification performance across these profiles is satisfactory. The best performance was achieved using a combination of loudness scaling, audiogram and speech test information, while single measures performed worst. The enhanced profile generation pipeline demonstrates the feasibility of combining APs across datasets, which should generalize to all datasets and could lead to an interpretable population-based profile set in the future. The classification models maintain clinical applicability. Hence, even if only smartphone-based measures are available, a given patient can be classified into an appropriate AP.
【8】 Decoding Linguistic Representations of Human Brain
标题: 解码人脑的语言表达
作者:Yu Wang,Heyang Liu,Yuhao Wang,Chuan Xuan,Yixuan Hou,Sheng Feng,Hongcheng Liu,Yusheng Liao,Yanfeng Wang
链接:点击下载PDF文件
摘要:语言作为高级生物体创造的信息媒介,一直是神经科学关注的问题,即它在大脑中是如何表现的。由于神经成像、医疗技术、生命科学和人工智能的快速发展,在诱发脑中解码语言表征已经取得了突破性的成就。在这项工作中,我们提出了一个分类的大脑语言解码的文本和语音格式。这项工作整合了两种类型的研究:专注于语言理解的神经科学和基于深度学习的大脑解码。从大脑活动中产生可辨别的语言信息不仅可以帮助那些发音受限的人,特别是肌萎缩侧索硬化症(ALS)患者,而且还为下一代脑机接口(BCI)开辟了新的途径。本文将帮助脑科学家和深度学习研究人员对细粒度的语言感知进行鸟瞰,从而促进他们对神经过程和语言解码的进一步调查和研究。摘要:Language, as an information medium created by advanced organisms, has always been a concern of neuroscience regarding how it is represented in the brain. Decoding linguistic representations in the evoked brain has shown groundbreaking achievements, thanks to the rapid improvement of neuroimaging, medical technology, life sciences and artificial intelligence. In this work, we present a taxonomy of brain-to-language decoding of both textual and speech formats. This work integrates two types of research: neuroscience focusing on language understanding and deep learning-based brain decoding. Generating discernible language information from brain activity could not only help those with limited articulation, especially amyotrophic lateral sclerosis (ALS) patients but also open up a new way for the next generation's brain-computer interface (BCI). This article will help brain scientists and deep-learning researchers to gain a bird's eye view of fine-grained language perception, and thus facilitate their further investigation and research of neural process and language decoding.
【9】 EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos
标题: EgoSonics:为无声的以自我为中心的视频生成同步音频
作者:Aashish Rai,Srinath Sridhar
备注:preprint
链接:点击下载PDF文件
摘要:我们介绍了自我声波,一种方法来生成语义上有意义的和同步的音频轨道上无声的自我为中心的视频。为以自我为中心的无声视频生成音频可以在虚拟现实、辅助技术或增强现有数据集方面开辟新的应用。现有的工作仅限于语音,音乐或冲击声等领域,无法轻松捕获以自我为中心的视频中的广泛音频频率。EgoSonics通过建立条件音频合成的潜在扩散模型的强度来解决这些限制。我们首先将音频和视频数据编码和处理为适合生成的形式。编码后的数据用于训练我们的模型,以生成捕获输入视频语义的音轨。我们提出的SyncroNet建立在ControlNet之上,提供控制信号,使时间同步的合成音频。广泛的评估表明,我们的模型优于现有的工作在音频质量,并在我们新提出的同步评估方法。此外,我们展示了我们的模型在改善视频摘要的下游应用。摘要:We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality, assistive technologies, or for augmenting existing datasets. Existing work has been limited to domains like speech, music, or impact sounds and cannot easily capture the broad range of audio frequencies found in egocentric videos. EgoSonics addresses these limitations by building on the strength of latent diffusion models for conditioned audio synthesis. We first encode and process audio and video data into a form that is suitable for generation. The encoded data is used to train our model to generate audio tracks that capture the semantics of the input video. Our proposed SyncroNet builds on top of ControlNet to provide control signals that enables temporal synchronization to the synthesized audio. Extensive evaluations show that our model outperforms existing work in audio quality, and in our newly proposed synchronization evaluation method. Furthermore, we demonstrate downstream applications of our model in improving video summarization.
【10】 DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
标题: DeepSpeech模型展示了类似人类的性能和对Costa植入物输入的处理
作者:Cynthia R. Steinhardt,Menoua Keshishian,Nima Mesgarani,Kim Stachenfeld
备注:NEURIPS preprint
链接:点击下载PDF文件
摘要:耳蜗植入物(CI)可以说是最成功的神经植入物,已经恢复了全世界100多万人的听力。虽然CI研究的重点是模拟耳蜗激活响应于低水平的声学特征,我们假设这些植入物的成功在很大程度上是由于上游网络的作用,从降级的信号中提取有用的功能,并学习语言的统计数据来解决信号。在这项工作中,我们使用深度神经网络(DNN)DeepSpeech2作为一个范例来研究自然输入和基于人工耳蜗植入的输入如何随着时间的推移而被处理。我们从口语句子中生成自然主义和类似人工耳蜗植入的输入,并在类似音素识别测试中测试模型性能与人类性能的相似性。我们的模型再现了正常听力和CI参与者研究中的噪声条件下的反应时间和音素混淆模式的错误模式。然后,我们使用可解释性技术来确定在处理自然主义和CI类输入时何时何地出现混淆。我们发现,随着时间的推移,在每一层的动态受到上下文以及输入类型的影响。在同一时间窗口内,所有音素的动力学在混淆和理解过程中发散,在网络的每一层中时间上向后移动。在CI的处理期间存在该信号的调制,其类似于听觉流中的人类EEG信号的变化。这种减少可能与编码音素身份的减少有关。这些研究结果表明,我们有一个可行的模型,可以及时探索语音相关信息的丢失,并且我们可以在优化人工耳蜗输入时使用它来寻找人群水平的编码信号,以改善基本语音相关信息的编码并改善感知。摘要:Cochlear implants(CIs) are arguably the most successful neural implant, having restored hearing to over one million people worldwide. While CI research has focused on modeling the cochlear activations in response to low-level acoustic features, we hypothesize that the success of these implants is due in large part to the role of the upstream network in extracting useful features from a degraded signal and learned statistics of language to resolve the signal. In this work, we use the deep neural network (DNN) DeepSpeech2, as a paradigm to investigate how natural input and cochlear implant-based inputs are processed over time. We generate naturalistic and cochlear implant-like inputs from spoken sentences and test the similarity of model performance to human performance on analogous phoneme recognition tests. Our model reproduces error patterns in reaction time and phoneme confusion patterns under noise conditions in normal hearing and CI participant studies. We then use interpretability techniques to determine where and when confusions arise when processing naturalistic and CI-like inputs. We find that dynamics over time in each layer are affected by context as well as input type. Dynamics of all phonemes diverge during confusion and comprehension within the same time window, which is temporally shifted backward in each layer of the network. There is a modulation of this signal during processing of CI which resembles changes in human EEG signals in the auditory stream. This reduction likely relates to the reduction of encoded phoneme identity. These findings suggest that we have a viable model in which to explore the loss of speech-related information in time and that we can use it to find population-level encoding signals to target when optimizing cochlear implant inputs to improve encoding of essential speech-related information and improve perception.
【11】 SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
标题: SuperCodec:具有选择性反投影网络的神经语音编解码器
作者:Youqiang Zheng,Weiping Tu,Li Xiao,Xinmeng Xu
备注:Accepted by ICASSP 2024
链接:点击下载PDF文件
摘要:神经语音编码是一个快速发展的课题,其中最先进的方法现在表现出优于传统方法的压缩性能。尽管取得了重大进展,但现有方法在保留和重建精细细节以实现最佳重建方面仍然存在局限性,特别是在低比特率下。在这项研究中,我们介绍了SuperCodec,一种神经语音编解码器,在低比特率下实现了最先进的性能。它采用了一种新的反投影方法与选择性特征融合增强表示。具体来说,我们建议使用选择性上采样反向投影(SUBP)和选择性下采样反向投影(SDBP)模块来分别替换编码器和解码器处的标准上采样层和下采样层。实验结果表明,我们的方法优于现有的神经语音编解码器在不同的比特率操作。具体来说,我们提出的方法可以实现更高质量的重建语音在1 kbps比天琴座V2在3.2 kbps和Encodec在6 kbps。摘要:Neural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction, especially at low bitrates. In this study, we introduce SuperCodec, a neural speech codec that achieves state-of-the-art performance at low bitrates. It employs a novel back projection method with selective feature fusion for augmented representation. Specifically, we propose to use Selective Up-sampling Back Projection (SUBP) and Selective Down-sampling Back Projection (SDBP) modules to replace the standard up- and down-sampling layers at the encoder and decoder, respectively. Experimental results show that our method outperforms the existing neural speech codecs operating at various bitrates. Specifically, our proposed method can achieve higher quality reconstructed speech at 1 kbps than Lyra V2 at 3.2 kbps and Encodec at 6 kbps.
【12】 Futga: Towards Fine-grained Music Understanding through Temporally-enhanced Generative Augmentation
标题: Futga:通过时间增强生成增强实现细粒度音乐理解
作者:Junda Wu,Zachary Novack,Amit Namburi,Jiaheng Dai,Hao-Wen Dong,Zhouhang Xie,Carol Chen,Julian McAuley
备注:6 pages
链接:点击下载PDF文件
摘要:现有的音乐字幕方法仅限于生成简短音乐片段的简洁全局描述,无法捕获细粒度的音乐特征和时间感知的音乐变化。为了解决这些局限性,我们提出了FUTGA,一个通过学习生成增强与时间组成,配备了细粒度的音乐理解能力的模型。我们利用现有的音乐字幕数据集和大型语言模型(LLM)来合成具有结构描述和完整歌曲时间边界的细粒度音乐字幕。通过所提出的合成数据集的增强,FUTGA能够识别音乐在关键过渡点的时间变化及其音乐功能,并为每个音乐片段生成详细的描述。我们进一步介绍了FUTGA生成的全长音乐字幕数据集,作为MusicCaps和Song Describer数据集的增强。我们评估了多个下游任务中自动生成的字幕,包括音乐生成和检索。实验表明,所产生的字幕的质量和更好的性能,在各种下游任务所提出的音乐字幕的方法。我们的代码和数据集可以在 href{https: huggingface.co JoshuaW1997 FUTGA}{ textcolor{blue}{https: huggingface.co JoshuaW1997 FUTGA}}中找到。摘要:Existing music captioning methods are limited to generating concise global descriptions of short music clips, which fail to capture fine-grained musical characteristics and time-aware musical changes. To address these limitations, we propose FUTGA, a model equipped with fined-grained music understanding capabilities through learning from generative augmentation with temporal compositions. We leverage existing music caption datasets and large language models (LLMs) to synthesize fine-grained music captions with structural descriptions and time boundaries for full-length songs. Augmented by the proposed synthetic dataset, FUTGA is enabled to identify the music's temporal changes at key transition points and their musical functions, as well as generate detailed descriptions for each music segment. We further introduce a full-length music caption dataset generated by FUTGA, as the augmentation of the MusicCaps and the Song Describer datasets. We evaluate the automatically generated captions on several downstream tasks, including music generation and retrieval. The experiments demonstrate the quality of the generated captions and the better performance in various downstream tasks achieved by the proposed music captioning approach. Our code and datasets can be found in href{https: huggingface.co JoshuaW1997 FUTGA}{ textcolor{blue}{https: huggingface.co JoshuaW1997 FUTGA}}.
机器翻译,仅供参考
