微信公众号:arXiv_Daily
cs.SD语音
【1】A Novel CustNetGC Boosted Model with Spectral Features for Parkinson's Disease Prediction
标题:用于帕金森病预测的具有光谱特征的新型CustNetGC增强模型
链接:https://arxiv.org/pdf/2511.15485v1
摘要:帕金森病是一种神经退行性疾病,诊断和治疗非常棘手。这种早期症状包括震颤、喘息和音质变化,这些都是神经损伤的关键指标。值得注意的是,人们越来越感兴趣的是利用声音属性的变化作为标记,用于检测PD早期on.Based上的这种理解,本论文的目的是专注于声学特征分析的基础上诊断为PD患者和健康对照(HC)的语音记录。在本文中,我们介绍了一种新的分类和可视化模型称为CustNetGC,结合卷积神经网络(CNN)与自定义网络Grad-CAM和CatBoost,以提高PD诊断的效率。我们使用来自Figshare的公开数据集,包括81名参与者的语音记录:40名PD患者和41名健康对照。从这些记录中,我们提取了关键的光谱特征:L-mHP和光谱斜率。L-mHP特征组合了三种频谱图表示:对数梅尔频谱图、谐波频谱图和谐波频谱图,它们是使用谐波-冲击源分离(HPSS)导出的。Grad-CAM用于突出数据中的重要区域,从而使PD预测可解释且有效。我们提出的CustNetGC模型实现了99.06%的准确度和95.83%的精确度,PD类的ROC曲线下面积(AUC)为0.90,HC类为0.89。此外,CatBoost,梯度增强算法的组合,通过适当地分类PD和非PD样本,增强了鲁棒性和预测性能。因此,这些结果提供了CustNetGC系统在提高诊断准确性和帕金森病预测模型的可解释性方面的潜在改进。摘要:Parkinson's disease is a neurodegenerative disorder that can be very tricky to diagnose and treat. Such early symptoms can include tremors, wheezy breathing, and changes in voice quality as critical indicators of neural damage. Notably, there has been growing interest in utilizing changes in vocal attributes as markers for the detection of PD early on. Based on this understanding, the present paper was designed to focus on the acoustic feature analysis based on voice recordings of patients diagnosed with PD and healthy controls (HC). In this paper, we introduce a novel classification and visualization model known as CustNetGC, combining a Convolutional Neural Network (CNN) with Custom Network Grad-CAM and CatBoost to enhance the efficiency of PD diagnosis. We use a publicly available dataset from Figshare, including voice recordings of 81 participants: 40 patients with PD and 41 healthy controls. From these recordings, we extracted the key spectral features: L-mHP and Spectral Slopes. The L-mHP feature combines three spectrogram representations: Log-Mel spectrogram, harmonic spectrogram, and percussive spectrogram, which are derived using Harmonic-Percussive Source Separation (HPSS). Grad-CAM was used to highlight the important regions in the data, thus making the PD predictions interpretable and effective. Our proposed CustNetGC model achieved an accuracy of 99.06% and precision of 95.83%, with the area under the ROC curve (AUC) recorded at 0.90 for the PD class and 0.89 for the HC class. Additionally, the combination of CatBoost, a gradient boosting algorithm, enhanced the robustness and the prediction performance by properly classifying PD and non-PD samples. Therefore, the results provide the potential improvement in the CustNetGC system in enhancing diagnostic accuracy and the interpretability of the Parkinson's Disease prediction model.
【2】LargeSHS: A large-scale dataset of music adaptation
标题:LargeSHS:大规模音乐改编数据集
链接:https://arxiv.org/pdf/2511.15270v1
备注:submitted as an ISMIR 2025 late-breaking demo paper
摘要:基于人工智能的音乐生成的最新进展主要集中在文本条件模型上,而对基于参考的生成(如歌曲改编)的关注较少。为了支持这一系列研究,我们引入了LargeSHS,这是一个来自SecondHandSongs的大规模数据集,包含超过170万个元数据条目和大约90万个可公开访问的音频链接。与现有的数据集不同,LargeSHS包括音乐作品之间的结构化改编关系,从而能够构建代表翻唱歌曲家族的改编树和表演集群。我们提供了与现有数据集的全面统计和比较,突出了LargeSHS的独特规模和丰富性。该数据集为翻唱歌曲生成、基于参考的音乐生成和适应感知的MIR任务的新研究铺平了道路。摘要:Recent advances in AI-based music generation have focused heavily on text-conditioned models, with less attention given to reference-based generation such as song adaptation. To support this line of research, we introduce LargeSHS, a large-scale dataset derived from SecondHandSongs, containing over 1.7 million metadata entries and approximately 900k publicly accessible audio links. Unlike existing datasets, LargeSHS includes structured adaptation relationships between musical works, enabling the construction of adaptation trees and performance clusters that represent cover song families. We provide comprehensive statistics and comparisons with existing datasets, highlighting the unique scale and richness of LargeSHS. This dataset paves the way for new research in cover song generation, reference-based music generation, and adaptation-aware MIR tasks.【3】Aligning Generative Music AI with Human Preferences: Methods and Challenges
标题:将生成音乐人工智能与人类偏好保持一致:方法和挑战
链接:https://arxiv.org/pdf/2511.15038v1
备注:Accepted at the AAAI-2026 Senior Member Track
摘要:音乐生成人工智能的最新进展已经实现了显着的保真度和风格多样性,但由于它们使用的特定损失函数,这些系统往往无法与细微差别的人类偏好保持一致。本文主张系统地将偏好对齐技术应用于音乐生成,解决计算优化和人类音乐欣赏之间的根本差距。利用最近的突破,包括MusicRL的大规模偏好学习,多偏好对齐框架,如DiffRhythm+中基于扩散的偏好优化,以及Text 2 midi-InferAlign等推理时间优化技术,我们讨论了这些技术如何解决音乐的独特挑战:时间一致性,谐波一致性和主观质量评估。我们确定了主要的研究挑战,包括可扩展性,长期的组成,可靠性等偏好建模。展望未来,我们设想与偏好一致的音乐生成能够在交互式作曲工具和个性化音乐服务中实现变革性应用。这项工作需要持续的跨学科研究,结合机器学习和音乐理论的进步,以创建真正满足人类创造和体验需求的音乐AI系统。摘要:Recent advances in generative AI for music have achieved remarkable fidelity and stylistic diversity, yet these systems often fail to align with nuanced human preferences due to the specific loss functions they use. This paper advocates for the systematic application of preference alignment techniques to music generation, addressing the fundamental gap between computational optimization and human musical appreciation. Drawing on recent breakthroughs including MusicRL's large-scale preference learning, multi-preference alignment frameworks like diffusion-based preference optimization in DiffRhythm+, and inference-time optimization techniques like Text2midi-InferAlign, we discuss how these techniques can address music's unique challenges: temporal coherence, harmonic consistency, and subjective quality assessment. We identify key research challenges including scalability to long-form compositions, reliability amongst others in preference modelling. Looking forward, we envision preference-aligned music generation enabling transformative applications in interactive composition tools and personalized music services. This work calls for sustained interdisciplinary research combining advances in machine learning, music-theory to create music AI systems that truly serve human creative and experiential needs.
【4】Fine-tuning Pre-trained Audio Models for COVID-19 Detection: A Technical Report
标题:微调用于COVID-19检测的预训练音频模型:技术报告
链接:https://arxiv.org/pdf/2511.14939v1
备注:11 pages
摘要:本技术报告使用已建立的基准数据集调查了预训练音频模型在COVID-19检测任务中的性能。我们在Coswara和COUGHVID数据集上微调了Audio-MAE和三种PANN架构(CNN 6,CNN 10,CNN 14),评估了数据集内和跨数据集的泛化。我们按年龄和性别实施了严格的人口分层,以防止模型利用人口特征与COVID-19状态之间的虚假相关性。数据集内结果显示中等性能,Audio-MAE在Coswara上获得最强结果(0.82 AUC,0.76 F1评分),而所有模型在Coughvid上表现出有限性能(AUC 0.58-0.63)。交叉数据集评估显示所有模型均存在严重的泛化失败(AUC 0.43-0.68),Audio-MAE显示出强烈的性能下降(F1评分0.00-0.08)。我们的实验表明,人口统计平衡在降低表观模型性能的同时,通过消除人口统计泄漏--一个夸大性能指标的混杂因素--提供了对COVID-19检测能力的更现实的评估。此外,平衡后有限的数据集大小(1,219 - 2,160个样本)被证明不足以用于通常需要更大训练集的深度学习模型。这些发现突出了开发可推广的基于音频的COVID-19检测系统的根本挑战,并强调了严格的人口统计控制对临床稳健模型评估的重要性。摘要:This technical report investigates the performance of pre-trained audio models on COVID-19 detection tasks using established benchmark datasets. We fine-tuned Audio-MAE and three PANN architectures (CNN6, CNN10, CNN14) on the Coswara and COUGHVID datasets, evaluating both intra-dataset and cross-dataset generalization. We implemented a strict demographic stratification by age and gender to prevent models from exploiting spurious correlations between demographic characteristics and COVID-19 status. Intra-dataset results showed moderate performance, with Audio-MAE achieving the strongest result on Coswara (0.82 AUC, 0.76 F1-score), while all models demonstrated limited performance on Coughvid (AUC 0.58-0.63). Cross-dataset evaluation revealed severe generalization failure across all models (AUC 0.43-0.68), with Audio-MAE showing strong performance degradation (F1-score 0.00-0.08). Our experiments demonstrate that demographic balancing, while reducing apparent model performance, provides more realistic assessment of COVID-19 detection capabilities by eliminating demographic leakage - a confounding factor that inflate performance metrics. Additionally, the limited dataset sizes after balancing (1,219-2,160 samples) proved insufficient for deep learning models that typically require substantially larger training sets. These findings highlight fundamental challenges in developing generalizable audio-based COVID-19 detection systems and underscore the importance of rigorous demographic controls for clinically robust model evaluation.
【5】Voiced-Aware Style Extraction and Style Direction Adjustment for Expressive Text-to-Speech
标题:表达性文本到语音的语音感知风格提取和风格方向调整
链接:https://arxiv.org/pdf/2511.14824v1
备注:Master's thesis, Korea University, 2025
摘要:表达性文本到语音(TTS)的最新进展介绍了各种方法的基础上提取参考语音的风格嵌入。然而,合成高质量的表达语音仍然具有挑战性。我们提出了SpotlightTTS,它专门强调风格,通过风格感知的风格提取和风格方向调整。浊音感知风格提取关注与风格高度相关的浊音区域,同时保持不同语音区域之间的连续性以提高表现力。我们调整了提取的风格的方向,以最佳地整合到TTS模型中,从而提高了语音质量。实验结果表明,聚光灯TTS实现了卓越的表现力,整体语音质量和风格转移能力的基线模型相比。摘要:Recent advances in expressive text-to-speech (TTS) have introduced diverse methods based on style embedding extracted from reference speech. However, synthesizing high-quality expressive speech remains challenging. We propose SpotlightTTS, which exclusively emphasizes style via voiced-aware style extraction and style direction adjustment. Voiced-aware style extraction focuses on voiced regions highly related to style while maintaining continuity across different speech regions to improve expressiveness. We adjust the direction of the extracted style for optimal integration into the TTS model, which improves speech quality. Experimental results demonstrate that Spotlight-TTS achieves superior performance compared to baseline models in terms of expressiveness, overall speech quality, and style transfer capability.
【6】IHearYou: Linking Acoustic Features to DSM-5 Depressive Behavior Indicators
标题:IHearYou:将声学特征与DSM-5抑郁行为指标联系起来
链接:https://arxiv.org/pdf/2511.14801v1
摘要:抑郁症影响着全世界数百万人,但诊断仍然依赖于主观的自我报告和访谈,可能无法捕捉真实的行为。我们提出IHearYou,一种专注于语音声学的自动抑郁症检测方法。在家庭环境中使用被动传感,IHearYou提取语音特征,并通过针对重度抑郁症实例化的结构化链接框架将其与DSM-5(精神疾病诊断和统计手册)指标联系起来。该系统在本地运行以保护隐私,并包括一个持久性架构和仪表板,在商用笔记本电脑上显示实时吞吐量。为了确保可重复性,我们定义了一个配置驱动的协议,具有错误发现率(FDR)校正和性别分层测试。应用于DAIC-WOZ数据集,该协议揭示了方向一致的特征指标关联,而基于TESS的音频流实验验证了端到端的可行性。我们的研究结果表明,被动语音传感可以转化为可解释的DSM-5指标分数,弥合黑盒检测和临床可解释的设备上分析之间的差距。摘要:Depression affects over millions people worldwide, yet diagnosis still relies on subjective self-reports and interviews that may not capture authentic behavior. We present IHearYou, an approach to automated depression detection focused on speech acoustics. Using passive sensing in household environments, IHearYou extracts voice features and links them to DSM-5 (Diagnostic and Statistical Manual of Mental Disorders) indicators through a structured Linkage Framework instantiated for Major Depressive Disorder. The system runs locally to preserve privacy and includes a persistence schema and dashboard, presenting real-time throughput on a commodity laptop. To ensure reproducibility, we define a configuration-driven protocol with False Discovery Rate (FDR) correction and gender-stratified testing. Applied to the DAIC-WOZ dataset, this protocol reveals directionally consistent feature-indicator associations, while a TESS-based audio streaming experiment validates end-to-end feasibility. Our results show how passive voice sensing can be turned into explainable DSM-5 indicator scores, bridging the gap between black-box detection and clinically interpretable, on-device analysis.
【7】OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
标题:BOHS:一种用于实时音频压缩的优化块霍夫曼方案
链接:https://arxiv.org/pdf/2511.14793v1
备注:3 page, 2 figures, 2 tables
摘要:在本文中,我们介绍了OBHS(优化块霍夫曼方案),一种新的无损音频压缩算法,专为实时流媒体应用。OBHS利用具有规范代码表示和智能回退机制的逐块霍夫曼编码来实现高压缩比,同时保持低计算复杂度。我们的算法将音频数据划分为固定大小的块,为每个块构造最佳霍夫曼树,并采用规范码进行有效的存储和传输。实验结果表明,OBHS对静音丰富的音频的压缩率高达93.6%,并在各种音频类型(包括粉红噪声、音调和真实世界录音)中保持有竞争力的性能。对于n个音频样本,OBHS的线性时间复杂度为O(n),有效地平衡了压缩效率和计算需求,使其非常适合资源受限的实时音频流场景。摘要:In this paper, we introduce OBHS (Optimized Block Huffman Scheme), a novel lossless audio compression algorithm tailored for real-time streaming applications. OBHS leverages block-wise Huffman coding with canonical code representation and intelligent fallback mechanisms to achieve high compression ratios while maintaining low computational complexity. Our algorithm partitions audio data into fixed-size blocks, constructs optimal Huffman trees for each block, and employs canonical codes for efficient storage and transmission. Experimental results demonstrate that OBHS attains compression ratios of up to 93.6% for silence-rich audio and maintains competitive performance across various audio types, including pink noise, tones, and real-world recordings. With a linear time complexity of O(n) for n audio samples, OBHS effectively balances compression efficiency and computational demands, making it highly suitable for resource-constrained real-time audio streaming scenarios.【8】Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
标题:Auden-Voice:用于语音和语言理解的通用语音编码器
链接:https://arxiv.org/pdf/2511.15145v1
备注:Submitted to ICASSP2026
摘要:人类的声音编码身份和非语言线索,但编码器在大型音频语言模型(LALM)很少平衡这两个方面。在这项工作中,我们提出了一个研究建立一个通用的语音编码器,捕捉细微差别的语音线索。通过综合评估,我们发现,多任务训练产生最平衡的表示,而对比语言音频预训练(CLAP)主要提高检索,而不提高非语言理解。我们的最终编码器Auden-Voice在与LLM集成时也表现出强大的性能。代码和培训食谱将与音频理解工具包Auden一起发布。摘要:Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures nuanced voice cues. Through a comprehensive evaluation, we find that multi-task training yields the most balanced representations, whereas contrastive language-audio pretraining (CLAP) primarily improves retrieval without enhancing paralinguistic understanding. Our final encoder, Auden-Voice, also demonstrates strong performance when integrated with LLMs. The code and training recipes will be released with the audio understanding toolkit Auden.
【9】CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
标题:CASTELA:带有字幕和时间边界的长音频数据集
链接:https://arxiv.org/pdf/2511.15131v1
摘要:我们介绍CASTELLA,一个人类注释的音频基准音频时刻检索(AMR)的任务。尽管AMR具有各种有用的潜在应用,但仍然没有针对真实数据的既定基准。AMR的早期研究仅使用合成数据集训练模型。此外,评价基于少于100个样本的注释数据集。这导致报告的业绩不太可靠。为了确保应用程序在现实环境中的性能,我们提出了CASTELLA,一个大规模的手动注释AMR数据集。CASTELLA由分别用于训练、有效和测试分割的1,009、213和640个音频记录组成,比之前的数据集大24倍。我们还使用CASTELLA建立了AMR的基线模型。我们的实验表明,在对合成数据进行预训练后,在CASTELLA上进行微调的模型在Recall1@0.7中的表现优于仅在合成数据上训练的模型10.4个点。CASTELLA可在https: h-munakata.github.io CASTELLA-demo 上公开获取。摘要:We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The early study of AMR trained the model with solely synthetic datasets. Moreover, the evaluation is based on annotated dataset of fewer than 100 samples. This resulted in less reliable reported performance. To ensure performance for applications in real-world environments, we present CASTELLA, a large-scale manually annotated AMR dataset. CASTELLA consists of 1,009, 213, and 640 audio recordings for train, valid, and test split, respectively, which is 24 times larger than the previous dataset. We also establish a baseline model for AMR using CASTELLA. Our experiments demonstrate that a model fine-tuned on CASTELLA after pre-training on the synthetic data outperformed a model trained solely on the synthetic data by 10.4 points in Recall1@0.7. CASTELLA is publicly available in https: h-munakata.github.io CASTELLA-demo .
【1】Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
标题:Auden-Voice:用于语音和语言理解的通用语音编码器
链接:https://arxiv.org/pdf/2511.15145v1
备注:Submitted to ICASSP2026
摘要:人类的声音编码身份和非语言线索,但编码器在大型音频语言模型(LALM)很少平衡这两个方面。在这项工作中,我们提出了一个研究建立一个通用的语音编码器,捕捉细微差别的语音线索。通过综合评估,我们发现,多任务训练产生最平衡的表示,而对比语言音频预训练(CLAP)主要提高检索,而不提高非语言理解。我们的最终编码器Auden-Voice在与LLM集成时也表现出强大的性能。代码和培训食谱将与音频理解工具包Auden一起发布。摘要:Human voice encodes both identity and paralinguistic cues, yet encoders in large audio-language models (LALMs) rarely balance both aspects. In this work, we present a study toward building a general-purpose voice encoder that captures nuanced voice cues. Through a comprehensive evaluation, we find that multi-task training yields the most balanced representations, whereas contrastive language-audio pretraining (CLAP) primarily improves retrieval without enhancing paralinguistic understanding. Our final encoder, Auden-Voice, also demonstrates strong performance when integrated with LLMs. The code and training recipes will be released with the audio understanding toolkit Auden.
【2】CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
标题:CASTELA:带有字幕和时间边界的长音频数据集
链接:https://arxiv.org/pdf/2511.15131v1
摘要:我们介绍CASTELLA,一个人类注释的音频基准音频时刻检索(AMR)的任务。尽管AMR具有各种有用的潜在应用,但仍然没有针对真实数据的既定基准。AMR的早期研究仅使用合成数据集训练模型。此外,评价基于少于100个样本的注释数据集。这导致报告的业绩不太可靠。为了确保应用程序在现实环境中的性能,我们提出了CASTELLA,一个大规模的手动注释AMR数据集。CASTELLA由分别用于训练、有效和测试分割的1,009、213和640个音频记录组成,比之前的数据集大24倍。我们还使用CASTELLA建立了AMR的基线模型。我们的实验表明,在对合成数据进行预训练后,在CASTELLA上进行微调的模型在Recall1@0.7中的表现优于仅在合成数据上训练的模型10.4个点。CASTELLA可在https: h-munakata.github.io CASTELLA-demo 上公开获取。摘要:We introduce CASTELLA, a human-annotated audio benchmark for the task of audio moment retrieval (AMR). Although AMR has various useful potential applications, there is still no established benchmark with real-world data. The early study of AMR trained the model with solely synthetic datasets. Moreover, the evaluation is based on annotated dataset of fewer than 100 samples. This resulted in less reliable reported performance. To ensure performance for applications in real-world environments, we present CASTELLA, a large-scale manually annotated AMR dataset. CASTELLA consists of 1,009, 213, and 640 audio recordings for train, valid, and test split, respectively, which is 24 times larger than the previous dataset. We also establish a baseline model for AMR using CASTELLA. Our experiments demonstrate that a model fine-tuned on CASTELLA after pre-training on the synthetic data outperformed a model trained solely on the synthetic data by 10.4 points in Recall1@0.7. CASTELLA is publicly available in https: h-munakata.github.io CASTELLA-demo .
【3】Quality-Controlled Multimodal Emotion Recognition in Conversations with Identity-Based Transfer Learning and MAMBA Fusion
标题:基于身份的迁移学习和MAMBA融合在对话中进行质量控制的多模式情感识别
链接:https://arxiv.org/pdf/2511.14969v1
备注:8 pages, 14 images, 3 tables, Recognition Technologies, Inc. Technical Report RTI-20251118-01
Journal-ref:Recognition Technologies, Inc. Technical Reports, 2025
摘要:本文通过系统的质量控制和多阶段迁移学习来解决多模态会话情感识别(MERC)中的数据质量问题。我们为MELD和IEMOCAP数据集实现了一个质量控制管道,用于验证说话人身份、音频文本对齐和人脸检测。我们利用来自说话人和人脸识别的迁移学习,假设身份识别嵌入不仅捕获稳定的声学和面部特征,而且捕获特定于个人的情感表达模式。我们使用MMPadeEasy(R)引擎来提取512维说话人和人脸嵌入,微调MPNet-v2以实现情感感知的文本表示,并通过在单峰数据集上训练的情感特定MLP来调整这些特征。基于MAMBA的三峰融合在MELD上达到64.8%的准确率,在IEMOCAP上达到74.3%。这些结果表明,结合基于身份的音频和视频嵌入与情感调整的文本表示的质量控制的数据子集产生一致的竞争性能的多模态情感识别会话,并提供了一个基础上进一步改善具有挑战性的,低频的情感类。摘要:This paper addresses data quality issues in multimodal emotion recognition in conversation (MERC) through systematic quality control and multi-stage transfer learning. We implement a quality control pipeline for MELD and IEMOCAP datasets that validates speaker identity, audio-text alignment, and face detection. We leverage transfer learning from speaker and face recognition, assuming that identity-discriminative embeddings capture not only stable acoustic and Facial traits but also person-specific patterns of emotional expression. We employ RecoMadeEasy(R) engines for extracting 512-dimensional speaker and face embeddings, fine-tune MPNet-v2 for emotion-aware text representations, and adapt these features through emotion-specific MLPs trained on unimodal datasets. MAMBA-based trimodal fusion achieves 64.8% accuracy on MELD and 74.3% on IEMOCAP. These results show that combining identity-based audio and visual embeddings with emotion-tuned text representations on a quality-controlled subset of data yields consistent competitive performance for multimodal emotion recognition in conversation and provides a basis for further improvement on challenging, low-frequency emotion classes.【4】Aligning Generative Music AI with Human Preferences: Methods and Challenges
标题:将生成音乐人工智能与人类偏好保持一致:方法和挑战
链接:https://arxiv.org/pdf/2511.15038v1
备注:Accepted at the AAAI-2026 Senior Member Track
摘要:音乐生成人工智能的最新进展已经实现了显着的保真度和风格多样性,但由于它们使用的特定损失函数,这些系统往往无法与细微差别的人类偏好保持一致。本文主张系统地将偏好对齐技术应用于音乐生成,解决计算优化和人类音乐欣赏之间的根本差距。利用最近的突破,包括MusicRL的大规模偏好学习,多偏好对齐框架,如DiffRhythm+中基于扩散的偏好优化,以及Text 2 midi-InferAlign等推理时间优化技术,我们讨论了这些技术如何解决音乐的独特挑战:时间一致性,谐波一致性和主观质量评估。我们确定了主要的研究挑战,包括可扩展性,长期的组成,可靠性等偏好建模。展望未来,我们设想与偏好一致的音乐生成能够在交互式作曲工具和个性化音乐服务中实现变革性应用。这项工作需要持续的跨学科研究,结合机器学习和音乐理论的进步,以创建真正满足人类创造和体验需求的音乐AI系统。摘要:Recent advances in generative AI for music have achieved remarkable fidelity and stylistic diversity, yet these systems often fail to align with nuanced human preferences due to the specific loss functions they use. This paper advocates for the systematic application of preference alignment techniques to music generation, addressing the fundamental gap between computational optimization and human musical appreciation. Drawing on recent breakthroughs including MusicRL's large-scale preference learning, multi-preference alignment frameworks like diffusion-based preference optimization in DiffRhythm+, and inference-time optimization techniques like Text2midi-InferAlign, we discuss how these techniques can address music's unique challenges: temporal coherence, harmonic consistency, and subjective quality assessment. We identify key research challenges including scalability to long-form compositions, reliability amongst others in preference modelling. Looking forward, we envision preference-aligned music generation enabling transformative applications in interactive composition tools and personalized music services. This work calls for sustained interdisciplinary research combining advances in machine learning, music-theory to create music AI systems that truly serve human creative and experiential needs.
【5】Fine-tuning Pre-trained Audio Models for COVID-19 Detection: A Technical Report
标题:微调用于COVID-19检测的预训练音频模型:技术报告
链接:https://arxiv.org/pdf/2511.14939v1
备注:11 pages
摘要:本技术报告使用已建立的基准数据集调查了预训练音频模型在COVID-19检测任务中的性能。我们在Coswara和COUGHVID数据集上微调了Audio-MAE和三种PANN架构(CNN 6,CNN 10,CNN 14),评估了数据集内和跨数据集的泛化。我们按年龄和性别实施了严格的人口分层,以防止模型利用人口特征与COVID-19状态之间的虚假相关性。数据集内结果显示中等性能,Audio-MAE在Coswara上获得最强结果(0.82 AUC,0.76 F1评分),而所有模型在Coughvid上表现出有限性能(AUC 0.58-0.63)。交叉数据集评估显示所有模型均存在严重的泛化失败(AUC 0.43-0.68),Audio-MAE显示出强烈的性能下降(F1评分0.00-0.08)。我们的实验表明,人口统计平衡在降低表观模型性能的同时,通过消除人口统计泄漏--一个夸大性能指标的混杂因素--提供了对COVID-19检测能力的更现实的评估。此外,平衡后有限的数据集大小(1,219 - 2,160个样本)被证明不足以用于通常需要更大训练集的深度学习模型。这些发现突出了开发可推广的基于音频的COVID-19检测系统的根本挑战,并强调了严格的人口统计控制对临床稳健模型评估的重要性。摘要:This technical report investigates the performance of pre-trained audio models on COVID-19 detection tasks using established benchmark datasets. We fine-tuned Audio-MAE and three PANN architectures (CNN6, CNN10, CNN14) on the Coswara and COUGHVID datasets, evaluating both intra-dataset and cross-dataset generalization. We implemented a strict demographic stratification by age and gender to prevent models from exploiting spurious correlations between demographic characteristics and COVID-19 status. Intra-dataset results showed moderate performance, with Audio-MAE achieving the strongest result on Coswara (0.82 AUC, 0.76 F1-score), while all models demonstrated limited performance on Coughvid (AUC 0.58-0.63). Cross-dataset evaluation revealed severe generalization failure across all models (AUC 0.43-0.68), with Audio-MAE showing strong performance degradation (F1-score 0.00-0.08). Our experiments demonstrate that demographic balancing, while reducing apparent model performance, provides more realistic assessment of COVID-19 detection capabilities by eliminating demographic leakage - a confounding factor that inflate performance metrics. Additionally, the limited dataset sizes after balancing (1,219-2,160 samples) proved insufficient for deep learning models that typically require substantially larger training sets. These findings highlight fundamental challenges in developing generalizable audio-based COVID-19 detection systems and underscore the importance of rigorous demographic controls for clinically robust model evaluation.
【6】OBHS: An Optimized Block Huffman Scheme for Real-Time Audio Compression
标题:BOHS:一种用于实时音频压缩的优化块霍夫曼方案
链接:https://arxiv.org/pdf/2511.14793v1
备注:3 page, 2 figures, 2 tables
摘要:在本文中,我们介绍了OBHS(优化块霍夫曼方案),一种新的无损音频压缩算法,专为实时流媒体应用。OBHS利用具有规范代码表示和智能回退机制的逐块霍夫曼编码来实现高压缩比,同时保持低计算复杂度。我们的算法将音频数据划分为固定大小的块,为每个块构造最佳霍夫曼树,并采用规范码进行有效的存储和传输。实验结果表明,OBHS对静音丰富的音频的压缩率高达93.6%,并在各种音频类型(包括粉红噪声、音调和真实世界录音)中保持有竞争力的性能。对于n个音频样本,OBHS的线性时间复杂度为O(n),有效地平衡了压缩效率和计算需求,使其非常适合资源受限的实时音频流场景。摘要:In this paper, we introduce OBHS (Optimized Block Huffman Scheme), a novel lossless audio compression algorithm tailored for real-time streaming applications. OBHS leverages block-wise Huffman coding with canonical code representation and intelligent fallback mechanisms to achieve high compression ratios while maintaining low computational complexity. Our algorithm partitions audio data into fixed-size blocks, constructs optimal Huffman trees for each block, and employs canonical codes for efficient storage and transmission. Experimental results demonstrate that OBHS attains compression ratios of up to 93.6% for silence-rich audio and maintains competitive performance across various audio types, including pink noise, tones, and real-world recordings. With a linear time complexity of O(n) for n audio samples, OBHS effectively balances compression efficiency and computational demands, making it highly suitable for resource-constrained real-time audio streaming scenarios.机器翻译由腾讯交互翻译提供,仅供参考
