今日论文合集:cs.SD语音7篇,eess.AS音频处理2篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Robust Neural Audio Fingerprinting using Music Foundation Models
标题:使用音乐基金会模型的稳健神经音频指纹识别
链接:https://arxiv.org/abs/2511.05399

作者:Shubhr Singh, Kiran Bhat, Xavier Riley, Benjamin Resnick, John Thickstun, Walter De Brouwer
摘要:在TikTok等现代媒体平台上,扭曲、压缩和操纵的音乐激增,促使人们开发更强大的音频指纹技术来识别音乐录音的来源。在本文中,我们开发和评估新的神经音频指纹技术,目的是提高其鲁棒性。我们对神经指纹识别方法做出了两个贡献:(1)我们使用预训练的音乐基础模型作为神经架构的骨干;(2)我们扩展了数据增强的使用,以在各种音频操作下训练指纹识别模型,包括时间拉伸,音调调制,压缩和滤波。我们系统地评估我们的方法相比,两个国家的最先进的神经指纹模型:NAFP和GraFPrint。结果表明,用音乐基础模型提取的指纹(例如,MuQ,MERT)始终优于从头开始训练或对非音乐音频进行预训练的模型。段级评估进一步揭示了它们准确定位指纹匹配的能力,这是目录管理的一个重要实用功能。
摘要:The proliferation of distorted, compressed, and manipulated music on modern media platforms like TikTok motivates the development of more robust audio fingerprinting techniques to identify the sources of musical recordings. In this paper, we develop and evaluate new neural audio fingerprinting techniques with the aim of improving their robustness. We make two contributions to neural fingerprinting methodology: (1) we use a pretrained music foundation model as the backbone of the neural architecture and (2) we expand the use of data augmentation to train fingerprinting models under a wide variety of audio manipulations, including time streching, pitch modulation, compression, and filtering. We systematically evaluate our methods in comparison to two state-of-the-art neural fingerprinting models: NAFP and GraFPrint. Results show that fingerprints extracted with music foundation models (e.g., MuQ, MERT) consistently outperform models trained from scratch or pretrained on non-musical audio. Segment-level evaluation further reveals their capability to accurately localize fingerprint matches, an important practical feature for catalog management.


【2】Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
标题:通过噪声增强的自编码器实现音乐的感知对齐表示
链接:https://arxiv.org/abs/2511.05350

作者:Mathias Rose Bjare, Giorgia Cantisani, Marco Pasini, Stefan Lattner, Gerhard Widmer
备注:Accepted at NeurIPS 2025 - AI for Music Workshop, 11 pages, 5 figures, 1 table
摘要:我们认为,训练自动编码器从其编码的噪声版本中重建输入,当与感知损失相结合时,会产生根据感知层次结构的编码。我们展示了这种分层结构的出现,表明,以这种方式训练音频自动编码器后,感知显着的信息被捕获在粗糙的表示结构比传统的训练。此外,我们还表明,在估计音乐音调中的惊喜和预测脑电对音乐聆听的反应的背景下,这种感知层次结构可以改善潜在的扩散解码。预先训练的权重可在github.com/CPJKU/pa-audioic上获得。
摘要:We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptual losses, yields encodings that are structured according to a perceptual hierarchy. We demonstrate the emergence of this hierarchical structure by showing that, after training an audio autoencoder in this manner, perceptually salient information is captured in coarser representation structures than with conventional training. Furthermore, we show that such perceptual hierarchies improve latent diffusion decoding in the context of estimating surprisal in music pitches and predicting EEG-brain responses to music listening. Pretrained weights are available on github.com/CPJKU/pa-audioic.


【3】Passive Acoustic Monitoring of Noisy Coral Reefs
标题:吵闹珊瑚礁的被动声学监测
链接:https://arxiv.org/abs/2511.05349

作者:Hari Vishnu, Yuen Min Too, Mandar Chitre, Danwei Huang, Teong Beng Koay, Sudhanshi S. Jain
摘要:被动声学监测有可能对珊瑚礁进行长期、空间广泛的评估。为了探索这种方法,我们在两年多的时间里在新加坡水域周围的十个珊瑚礁地点部署了水下声学记录仪。为了减轻持续的生物噪声掩盖低频珊瑚礁声景,我们训练了一个卷积神经网络去噪器。声学数据的分析揭示了不同的早晨和晚上的合唱。虽然与环境变量的相关性是模糊的噪声录音的低频部分,去噪数据显示的相关性的声学活动指数,如声压级和声学复杂性指数与潜水员为基础的评估珊瑚礁的健康,如活珊瑚的丰富度和覆盖,和藻类覆盖。此外,根据高频声波带计算的虾突变率与珊瑚礁参数在时间和空间上都具有鲁棒相关性。这项研究表明,被动声学拥有有价值的信息,可以帮助珊瑚礁监测,提供有效的数据去噪和解释。这种方法可以扩展到其他海洋环境中的声学监测受到持续的噪音。
摘要:Passive acoustic monitoring offers the potential to enable long-term, spatially extensive assessments of coral reefs. To explore this approach, we deployed underwater acoustic recorders at ten coral reef sites around Singapore waters over two years. To mitigate the persistent biological noise masking the low-frequency reef soundscape, we trained a convolutional neural network denoiser. Analysis of the acoustic data reveals distinct morning and evening choruses. Though the correlation with environmental variates was obscured in the low-frequency part of the noisy recordings, the denoised data showed correlations of acoustic activity indices such as sound pressure level and acoustic complexity index with diver-based assessments of reef health such as live coral richness and cover, and algal cover. Furthermore, the shrimp snap rate, computed from the high-frequency acoustic band, is robustly correlated with the reef parameters, both temporally and spatially. This study demonstrates that passive acoustics holds valuable information that can help with reef monitoring, provided the data is effectively denoised and interpreted. This methodology can be extended to other marine environments where acoustic monitoring is hindered by persistent noise.


【4】Model Merging Improves Zero-Shot Generalization in Bioacoustic Foundation Models
标题:模型合并改进了生物声学基础模型中的Zero-Shot概括
链接:https://arxiv.org/abs/2511.05171

作者:Davide Marincione, Donato Crisostomi, Roberto Dessi, Emanuele Rodolà, Emanuele Rossi
摘要:能够跨物种和任务推广的基础模型代表了生物声学领域一个有前途的新前沿,NatureLM是最突出的例子之一。虽然其特定领域的微调产生强大的性能生物声学基准,我们观察到,它也引入了权衡以下的灵活性。例如,NatureLM在分别提示通用名或学名时可以达到很高的准确性,但当在单个提示中同时请求两者时,其准确性会显著下降。我们通过应用一个简单的模型合并策略来解决这个问题,该策略将NatureLM与其基础语言模型进行插值,以最小的领域专业知识损失恢复自动跟踪功能。最后,我们表明,合并后的模型表现出显着更强的zero-shot泛化,实现了超过200%的相对改善,并设置一个新的国家的最先进的闭集zero-shot分类看不见的物种。
摘要:Foundation models capable of generalizing across species and tasks represent a promising new frontier in bioacoustics, with NatureLM being one of the most prominent examples. While its domain-specific fine-tuning yields strong performance on bioacoustic benchmarks, we observe that it also introduces trade-offs in instruction-following flexibility. For instance, NatureLM achieves high accuracy when prompted for either the common or scientific name individually, but its accuracy drops significantly when both are requested in a single prompt. We address this by applying a simple model merging strategy that interpolates NatureLM with its base language model, recovering instruction-following capabilities with minimal loss of domain expertise. Finally, we show that the merged model exhibits markedly stronger zero-shot generalization, achieving over a 200% relative improvement and setting a new state-of-the-art in closed-set zero-shot classification of unseen species.


【5】MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages
标题:MEaLiON-BER:针对英语和SEA语言的稳健语音情感识别模型
链接:https://arxiv.org/abs/2511.04914

作者:Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy, Chiu Ying Lay, Ding Yang, He Yingxu, Jiang Ridong, Li Jingtao, Liao Jingyi, Liu Zhuohan, Lu Yanfeng, Ma Yi, Manas Gupta, Muhammad Huzaifah Bin Md Shahrin, Nabilah Binte Md Johan, Nattadaporn Lertcheva, Pan Chunlei, Pham Minh Duc, Siti Maryam Binte Ahmad Subaidi, Siti Umairah Binte Mohammad Salleh, Sun Shuo, Tarun Kumar Vangani, Wang Qiongqiong, Won Cheng Yi Lewis, Wong Heng Meng Jeremy, Wu Jinyang, Zhang Huayun, Zhang Longyin, Zou Xunlong
备注:his https URL
摘要:我们提出了MERALiON-SER,一个强大的语音情感识别模型,设计的英语和东南亚语言。该模型使用混合目标结合加权分类交叉熵和一致性相关系数(CCC)损失联合离散和维度情感建模进行训练。这种双重方法使模型能够捕获不同类别的情绪(如快乐或愤怒)和细粒度的情绪,如唤醒(强度),效价(积极/消极)和优势(控制感),从而更全面和更强大地表示人类情感。对多语言新加坡语言(英语,中文,马来语和泰米尔语)和其他公共基准的广泛评估表明,MERaLiON-SER始终超过开源语音编码器和大型音频LLM。这些结果强调了专门的纯语音模型对于准确的平行语言理解和跨语言概括的重要性。此外,所提出的框架提供了一个基础,将情感感知感知到未来的代理音频系统,使更多的移情和上下文自适应多模态推理。
摘要:We present MERaLiON-SER, a robust speech emotion recognition model de- signed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weighted categorical cross-entropy and Concordance Correlation Coefficient (CCC) losses for joint discrete and dimensional emotion modelling. This dual approach enables the model to capture both the distinct categories of emotion (like happy or angry) and the fine-grained, such as arousal (intensity), valence (positivity/negativity), and dominance (sense of control), lead- ing to a more comprehensive and robust representation of human affect. Extensive evaluations across multilingual Singaporean languages (English, Chinese, Malay, and Tamil ) and other public benchmarks show that MERaLiON-SER consistently surpasses both open-source speech encoders and large Audio-LLMs. These results underscore the importance of specialised speech-only models for accurate paralin- guistic understanding and cross-lingual generalisation. Furthermore, the proposed framework provides a foundation for integrating emotion-aware perception into future agentic audio systems, enabling more empathetic and contextually adaptive multimodal reasoning.


【6】EMO100DB: An Open Dataset of Improvised Songs with Emotion Data
标题:EMO 100 DB:带有情感数据的即兴歌曲开放数据集
链接:https://arxiv.org/abs/2511.04755

作者:Daeun Hwang, Saebyul Park
备注:4 pages, 6 figures, International Conference on Music Perception and Cognition
摘要:在这项研究中,我们介绍了一个数据集,由即兴歌曲,记录和转录的情感数据的基础上罗素的回旋情感模型。该数据集是通过收集由20名年轻人演奏、演唱和录制的即兴歌曲而开发的,这些歌曲包括旋律、歌词和乐器伴奏。在录制每首歌之前,参与者被要求报告他们的情绪状态,轴代表基于罗素的情感循环模型的唤醒和效价。该数据集被组织成四个情感象限,它包括从参与者录音中提取的旋律的歌词文本和音频文件,以及WAV格式的原始音频。通过提供数据和分析的综合组合,本研究旨在提供一个全面的数据集,允许对音乐和情感之间的关系进行多样化的探索。
摘要:In this study, we introduce Emo100DB: a dataset consisting of improvised songs that were recorded and transcribed with emotion data based on Russell's circumplex model of emotion. The dataset was developed by collecting improvised songs that consist of melody, lyrics, and an instrumental accompaniment played, sung, and recorded by 20 young adults. Before recording each song, the participants were asked to report their emotional state, with the axes representing arousal and valence based on Russell's circumplex model of emotions. The dataset is organized into four emotion quadrants, and it includes the lyrics text and MIDI file of the melody extracted from the participant recordings, along with the original audio in WAV format. By providing an integrated composition of data and analysis, this study aims to offer a comprehensive dataset that allows for a diverse exploration of the relationship between music and emotion.


【7】A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
标题:为你的想法一分钱:从廉价的大脑信号中解码语音
链接:https://arxiv.org/abs/2511.04691

作者:Quentin Auster, Kateryna Shapovalenko, Chuang Ma, Demaio Sun
摘要:我们探索神经网络是否可以通过将EEG记录映射到音频表示来将大脑活动解码为语音。使用EEG数据记录为受试者听自然语音,我们训练一个模型与对比CLIP损失,以对齐EEG派生的嵌入嵌入从预训练的基于变换的语音模型。在Meta最先进的EEG解码器的基础上,我们引入了三个架构修改:(i)特定于主题的注意力层(+0.15%WER改进),(ii)个性化空间注意力(+0.45%),以及(iii)具有注意力的双路径RNN(-1.87%)。三个修改中的两个改进了性能,突出了个性化架构对脑-语音解码和脑-机接口应用的承诺。
摘要:We explore whether neural networks can decode brain activity into speech by mapping EEG recordings to audio representations. Using EEG data recorded as subjects listened to natural speech, we train a model with a contrastive CLIP loss to align EEG-derived embeddings with embeddings from a pre-trained transformer-based speech model. Building on the state-of-the-art EEG decoder from Meta, we introduce three architectural modifications: (i) subject-specific attention layers (+0.15% WER improvement), (ii) personalized spatial attention (+0.45%), and (iii) a dual-path RNN with attention (-1.87%). Two of the three modifications improved performance, highlighting the promise of personalized architectures for brain-to-speech decoding and applications in brain-computer interfaces.


eess.AS音频处理


【1】Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
标题:用选定的感知语音质量合成语音-吱吱作响的语音案例研究
链接:https://arxiv.org/abs/2511.05143

作者:Frederik Rautenberg, Fritz Seebauer, Jana Wiechmann, Michael Kuhlmann, Petra Wagner, Reinhold Haeb-Umbach
备注:Proceedings of Interspeech
摘要:文本到语音(TTS)系统中的感知语音质量的控制对于其中未操纵和操纵的语音探针可以用于说明否则难以掌握的语音概念的应用是感兴趣的。在这里,我们表明,一个TTS系统,这是一个全球性的扬声器属性操作块的基础上归一化flows1增强,是能够正确地操纵非持久性的,本地化质量的吱吱声的声音,从而避免了必要性,典型的不可靠的,帧明智的吱吱声预测。主观听音测试证实,与原始录音相比,在略微降低的MOS评分下,成功地进行了吱吱声操作。
摘要:The control of perceptual voice qualities in a text-to-speech (TTS) system is of interest for applications where unmanipu- lated and manipulated speech probes can serve to illustrate pho- netic concepts that are otherwise difficult to grasp. Here, we show that a TTS system, that is augmented with a global speaker attribute manipulation block based on normalizing flows1 , is capable of correctly manipulating the non-persistent, localized quality of creaky voice, thus avoiding the necessity of a, typi- cally unreliable, frame-wise creak predictor. Subjective listen- ing tests confirm successful creak manipulation at a slightly re- duced MOS score compared to the original recording.


【2】A Penny for Your Thoughts: Decoding Speech from Inexpensive Brain Signals
标题:为你的想法一分钱:从廉价的大脑信号中解码语音
链接:https://arxiv.org/abs/2511.04691

作者:Quentin Auster, Kateryna Shapovalenko, Chuang Ma, Demaio Sun
摘要:我们探索神经网络是否可以通过将EEG记录映射到音频表示来将大脑活动解码为语音。使用EEG数据记录为受试者听自然语音,我们训练一个模型与对比CLIP损失,以对齐EEG派生的嵌入嵌入从预训练的基于变换的语音模型。在Meta最先进的EEG解码器的基础上,我们引入了三个架构修改:(i)特定于主题的注意力层(+0.15%WER改进),(ii)个性化空间注意力(+0.45%),以及(iii)具有注意力的双路径RNN(-1.87%)。三个修改中的两个改进了性能,突出了个性化架构对脑-语音解码和脑-机接口应用的承诺。
摘要:We explore whether neural networks can decode brain activity into speech by mapping EEG recordings to audio representations. Using EEG data recorded as subjects listened to natural speech, we train a model with a contrastive CLIP loss to align EEG-derived embeddings with embeddings from a pre-trained transformer-based speech model. Building on the state-of-the-art EEG decoder from Meta, we introduce three architectural modifications: (i) subject-specific attention layers (+0.15% WER improvement), (ii) personalized spatial attention (+0.45%), and (iii) a dual-path RNN with attention (-1.87%). Two of the three modifications improved performance, highlighting the promise of personalized architectures for brain-to-speech decoding and applications in brain-computer interfaces.


机器翻译由腾讯交互翻译提供,仅供参考