今日论文合集:CS.SD语音与音频 | 共 5 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 音乐信息检索与音乐生成 2 篇

2. 多模态音频与视听学习 1 篇

3. 数据集、基准与评测 2 篇

1. 音乐信息检索与音乐生成 | 2 篇

1. Extending Xenakis: From Architectural Geometry to Sonification of the Philips Pavilion

扩展克赛纳基斯:从建筑几何到飞利浦展馆的声波化

AI 总结:该研究重新审视克赛纳基斯的转换,颠倒工作流程将飞利浦展馆声波化。通过重建展馆为直纹面,提取母线和采样点,用Python实现系统生成MIDI并可视化,为建筑几何转化为可演奏音乐结构铺路,扩展其音乐思维至声波化等实践。

链接:https://arxiv.org/abs/2607.06589

作者:Changda Ma, Sunshiyu Wang, Canting Zhu, Alexandria Smith

英文摘要:Architecture and music have been linked through proportion and temporal structure, yet architectural geometry is rarely viewed as a source of generative music. Revisiting Xenakis' one-directional transformation from string glissandi in Metastaseis to the ruled surfaces of the Philips Pavilion, we invert this workflow and sonify the completed Pavilion as a temporal composition. We reconstruct the Pavilion as nine ruled surfaces, extract their governing ruling lines, and subdivide each surface into structural lines and spatial sampling points. Four evenly spaced ruling lines per surface generate continuous string glissandi, while 3357 sampled points develop five density-based energy blocks and a sparse brass and woodwind subsequence. Implemented in Python, the system produces MIDI rendered in Ableton Live, accompanied by a real-time 3D visualization that reveals architectural motion, stasis, and structural contrast through sound and image. In general, this work paves the way for the transfer of architectural geometry as a performable musical structure, extending Xenakis's architectural and musical thinking to sonification and interactive music practice.

2. Rag Classification of Tagore Songs using Symbolic Music Notation and Novel Weighted Distance Measures

使用符号音乐记谱法和新型加权距离度量对泰戈尔歌曲进行拉格分类

AI 总结:该研究针对罗宾德拉·桑吉特歌曲拉格识别难题,将其转化为监督分类问题,利用符号乐谱记谱法构建数据集,探讨多种距离度量,引入加权欧几里得距离,在k近邻框架下改进拉格分类,更好捕捉旋律特征。

链接:https://arxiv.org/abs/2607.07241

机构:XIM University(西姆大学)

作者:Chandan Misra, Swarup Chattopadhyay

英文摘要:Rabindra Sangeet, the body of songs written and composed by Rabindranath Tagore, occupies a distinctive position in Indian music by combining poetic expression with melodic ideas drawn from Hindustani rags, Bengali folk traditions, tappa, kırtan, Baul music, and Western tunes. Although many Tagore songs are associated with rag labels provided by Tagore himself or preserved in authoritative notational traditions, rag identification remains challenging because the songs often reflect creative freedom rather than strict adherence to classical rag grammar. This paper formulates rag identification in Rabindra Sangeet as a supervised classification problem using symbolic music-sheet notations from Swarabitan. Since large-scale annotated audio or music datasets for Rabindra Sangeet are not readily available, this study constructs a rag-labelled symbolic dataset from notated Tagore songs. The work investigates Euclidean distance and cosine similarity for rag classification and introduces a weighted Euclidean distance measure that assigns greater importance to notes belonging to characteristic rag sequences such as arohana and avarohana. Applied within a k-nearest-neighbour framework, the proposed measure improves rag classification by better capturing rag-specific melodic identity.

2. 多模态音频与视听学习 | 1 篇

3. EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments

EscFOA:通过360度教育环境中的生成性空间音频增强视障学习者的空间学习

AI 总结:研究针对360度教育环境中视障学习者空间学习受限问题,提出EscFOA框架,通过整合3DGS与条件扩散模型生成空间音频,在支持蒙眼参与者空间学习行为上优于传统音频,证明其能助力视障者访问复杂空间学习材料。

链接:https://arxiv.org/abs/2607.07015

作者:Ziyu Luo, Xiaowei Dai, Siying Zhu, Xiaoming Chen

英文摘要:Immersive 360-degree educational environments often lack accessible spatial structure, limiting visually impaired learners' ability to orient, explore, and construct mental representations. This paper proposes EscFOA, a geometry-aware spatial audio generation framework designed as an \emph{acoustic scaffolding} to support spatial cognition. By integrating 3D Gaussian Splatting (3DGS) with conditional diffusion models, EscFOA reconstructs scene geometry from 360-degree videos to synthesize high-fidelity spatial audio consistent with the environmental structure. Explicitly targeting learning outcomes like independent spatial orientation and reduced cognitive load, EscFOA significantly outperforms conventional monaural and stereo audio in supporting spatial learning behaviors among blindfolded sighted participants (simulating visually impaired learners). These findings demonstrate that geometry-consistent generative audio can effectively enable inclusive access to complex spatial learning materials.

3. 数据集、基准与评测 | 2 篇

4. MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

MADB:一个具有专业和多维度注释的大规模音乐美学数据集

AI 总结:研究音乐美学评估问题,引入含9999首曲目、由30名注释者标注的MADB数据集,通过多预训练模型建立统一评估框架,揭示模型与人类判断差距,为音乐理解提供新基准。

链接:https://arxiv.org/abs/2607.06929

作者:Sirui Zhang, Tianle Wang, Xinyi Tong, Peiyang Yu, Jishang Chen, Liangke Zhao, Haoxin Zhang, Duo Xu, Xin Jin, Feng Yu, Songchun Zhu

英文摘要:Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual judgments. Progress in this area has been limited by the lack of large-scale datasets with structured aesthetic annotations. We introduce MADB, a large-scale dataset and benchmark comprising 9,999 tracks annotated by 30 trained annotators. Each track is rated by around 10 annotators across 10 perceptual dimensions and one overall score, with additional textual comments for multimodal analysis. We establish a unified evaluation framework over multiple pretrained models. Results reveal substantial gaps between model predictions and human judgments, exposing key limitations of current approaches. MADB provides a new benchmark for human-aligned music understanding. Project page: this https URL

5. MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

MMGenre:跨多种音乐流派的歌声合成基准测试

AI 总结:研究跨多种音乐流派的歌声合成泛化能力,引入MMGenre基准及自动管道,涵盖多流派。评估显示流派辨别有限,零样本改进小,特定流派持续训练有显著提升,为多流派SVS评估提供框架,揭示关键挑战。

链接:https://arxiv.org/abs/2607.06986

机构:Renmin University of China(中国人民大学); Carnegie Mellon University(卡内基梅隆大学)

作者:Wenhao Feng, Yuxun Tang, Jiatong Shi, Qin Jin

英文摘要:Singing voice synthesis (SVS) has progressed rapidly, yet its ability to generalize across diverse musical genres remains underexplored. Existing benchmarks are heavily biased toward pop music, limiting systematic analysis of genre-dependent behavior. We introduce MMGenre, a benchmark for multi-genre SVS diagnosis, supported by an automatic pipeline for constructing genre-aligned music scores. MMGenre spans 10 major genres and 26 subgenres, enabling comprehensive analysis of genre-aware synthesis. Extensive evaluation of representative SVS models reveals limited genre discrimination: synthesized vocals across genres exhibit highly similar acoustic characteristics and weak separability. While zero-shot genre adaptation yields only marginal improvements, lightweight genre-specific continued training leads to substantial gains. MMGenre provides a standardized framework for multi-genre SVS evaluation and exposes critical challenges in achieving genre-aware singing voice synthesis.