今日论文合集:CS.SD语音与音频 | 共 6 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音合成与声音生成 1 篇

2. 音乐信息检索与音乐生成 2 篇

3. 语音翻译与语音语言模型 1 篇

4. 数据集、基准与评测 1 篇

5. 其他/综合语音音频 1 篇

1. 语音合成与声音生成 | 1 篇

1. Adapting a Diffusion-Based Music Synthesis Model to Human Voice Conversion

将基于扩散的音乐合成模型应用于人类语音转换

AI 总结:研究将基于扩散的音乐合成模型用于语音转换,通过扩展条件设定、重新解释音色条件等方法,使模型在自然度等方面表现良好,还发现纳入乐器数据的问题及现成特征提取器的作用,凸显跨域模型转移对统一音频生成系统的潜力。

链接:https://arxiv.org/abs/2607.13278

作者:Ben Maman, Frank Zalkow, Hans-Ulrich Berendes, Paolo Sani, Christian Dittmar, Meinard Müller

英文摘要:Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music. Qualitative samples can be found on our project page: this https URL

2. 音乐信息检索与音乐生成 | 2 篇

2. From Prediction to Collaboration: Interactive Symbolic Music Analysis

从预测到协作:交互式符号音乐分析

AI 总结:针对自动符号音乐分析系统单一模式问题,提出统一框架,结合预训练表示在准确性与交互响应性间权衡,支持多种分析操作,实验验证其有效性,为音乐分析交互式工具奠定基础。

链接:https://arxiv.org/abs/2607.13587

作者:Emmanouil Karystinaios, Johannes Hentschel, Markus Neuwirth, Gerhard Widmer

英文摘要:Automatic symbolic music analysis has made substantial progress, yet existing systems are typically designed for a single mode of use, such as full-score prediction, and therefore do not match the broader range of operations that arise in analysis workflows, including partial completion, local correction, and iterative refinement. As a result, there remains a gap between strong benchmark models and systems that can support interactive analytical use. We present a unified framework for symbolic Roman-numeral (RN) analysis that narrows this gap by combining strong predictive performance with direct support for constrained completion and revision. The method is designed to provide a practical trade-off between accuracy and interactive responsiveness by computing expensive pretrained representations once and reusing them during iterative refinement, making powerful pretrained models more amenable to interactive settings. It supports complete score analysis, targeted revision of existing labels, and inference of missing annotations from partial context through a shared modeling framework. Experiments on Dilemmadata, the largest and most heterogeneous benchmark of its kind, show that the proposed approach is a strong RN-analysis baseline while also supporting masked completion from partial labels. Together with a prototype interface for multi-level candidate inspection and editing, these results position automatic RN analysis not only as a prediction problem, but also as a foundation for future interactive tools for music analysis.

3. Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation

流派偏见还是审美感知?识别和减轻音乐评价中的捷径学习

AI 总结:研究音乐评价模型中流派诱导的捷径学习问题,提出联合重新加权硬样本并规范化组级性能的训练目标,减少流派相关偏差,提高与人类偏好的一致性。

链接:https://arxiv.org/abs/2607.13903

作者:Yizhou Zhang, Wangjin Zhou, Yi Zhao, Wei Tan, Keisuke Imoto, Zhi Gong

英文摘要:Music aesthetics scoring plays a critical role in applications such as dataset curation, generative model evaluation, and reward modeling for music generation. Recent approaches rely on deep neural networks trained on human-annotated ratings, but these models may exploit spurious correlations rather than capturing perceptually meaningful aesthetics. In this work, we identify a previously underexplored failure mode in music evaluation models: genre-induced shortcut learning. Through a systematic analysis of SongEval, we show that biases in training data lead to strong correlations between genre-related features and predicted scores, causing the model to use them as a proxy for aesthetics. This results in systematic overestimation of pop music and undervaluation of high-quality samples from other genres, leading to predictions that are inconsistent with human preferences. To address this issue, we propose a training objective that jointly reweights hard samples and regularizes group-level performance, encouraging the model to learn genre-invariant representations of musicality. Experimental results demonstrate that our method reduces genre-dependent bias and improves alignment with human preferences, as reflected by gains in both cross-genre and within-genre preference alignment.

3. 语音翻译与语音语言模型 | 1 篇

4. Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

大型音频语言模型语音评估中协议级捷径的审计

AI 总结:研究大型音频语言模型语音评估中协议级捷径,通过对三种部署协议审计发现多个LALM依赖此类捷径,如特征蓝图评判中错误标签影响准确率,拼接A/B比较有固定选择倾向,强调联合评估模型与协议及用匹配探针评估的重要性。

链接:https://arxiv.org/abs/2607.13477

作者:Joonyong Park, David M. Chan, Yuki Saito, Hiroshi Saruwatari

英文摘要:Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data supplied by the evaluation protocol itself, taking a shortcut in place of listening to the audio. In this paper, we audit such protocol-level ``shortcuts'' in LALM judges across three common deployment protocols: feature-blueprint judging, where the audio is replaced by a structured text description of acoustic features, reference-conditioned judging, and pairwise A/B comparison. Across six judges and four attributes, we find that several LALMs rely on protocol-level shortcuts. For example, in feature-blueprint judging, incorrect specialist labels reduce five judges' emotion accuracy to 0.10 or below, and in concatenated A/B comparisons, Qwen3-Omni-Thinking often picks the same slot regardless of order swaps. These results indicate that aggregate agreement can overstate the validity of LALM judges unless the model and the evaluation protocol are assessed jointly, and that each model-protocol pair should be evaluated with a matched shortcut probe.

4. 数据集、基准与评测 | 1 篇

5. From Continuous Deployment to Queryable Dataset: Terabyte-Scale AIS-Aligned Passive Acoustic Labelling

从持续部署到可查询数据集:PB 级 AIS 对齐的被动声学标记

AI 总结:研究针对长时间被动声学部署录音与船只轨迹等未关联问题,提出数据库原生工作流程,将水听器录音与 AIS 位置报告对齐生成距离解析数据,经处理大量数据生成结构化表,实现地理人工智能框架,支持相关海洋环境分析建模。

链接:https://arxiv.org/abs/2607.13840

作者:Wayne Renaud, Priyanka Aravindan, Gabriel Spadon

英文摘要:Long-duration passive acoustic deployments produce large archives of recordings that are not linked to vessel tracks or encounter structure, leaving range and contact conditions unavailable as variables and requiring manual selection for analysis. To address this limitation, we propose a database-native workflow that aligns hydrophone recordings with Automatic Identification System (AIS) position reports to produce distance-resolved data. Fixed-duration recording windows and AIS messages are stored as persistent geospatial tables and associated through an indexed spatiotemporal join, replacing in-memory nested iteration with a single scalable set-based database process capable of handling continuous, multi-year, million-window archival deployments without exhausting available memory. In this study, the approach processes approximately 9.5x10e5 recording windows and 6.9x10e6 AIS position reports, producing a structured table that separates no-contact, single-contact, and two-contact windows, with the closest point of approach computed directly where applicable and background conditions characterized via deterministic spectral ranking. This formulation enables a GeoAI framework in which spatially indexed, queryable data become directly usable for machine learning. The resulting data product reveals predominantly noise-dominated conditions, with vessel contributions emerging mainly at shorter ranges, indicating that the task lies in extracting structure under background-limited regimes. Spectrogram and quantitative analyses show weak tonal signatures embedded in noise and a consistent decay of signal-to-noise ratio with distance, supporting the use of this representation for scalable machine learning, similarity analysis, and predictive acoustic modelling in real maritime environments.

5. 其他/综合语音音频 | 1 篇

6. Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?

重新思考语音基础模型微调:更好的监督微调还是更好的匹配?

AI 总结:研究语音基础模型微调,通过对3个SUPERB分类任务、9个预训练检查点的8种SFT变体进行系统研究,发现SFT结果强烈依赖预训练实例,顶级SFT方法常因检查点而异,下游增益多为实例和种子依赖的匹配,非普遍性能提升。

链接:https://arxiv.org/abs/2607.13864

机构:Graduate School of Informatics, Kyoto University(京都大学信息学研究生院)

作者:Wangjin Zhou, Yizhou Zhang, Yichi Wang, Tatsuya Kawahara

英文摘要:Supervised fine-tuning (SFT) is widely used to adapt self-supervised speech representations to downstream classification tasks. Small gains observed under a single pretrained checkpoint are often interpreted as method-level improvements, i.e., a higher attainable performance ceiling. We show that such conclusions are not always reliable because SFT outcomes depend strongly on the specific pretrained instance. We conduct a systematic study on 3 SUPERB classification tasks, evaluating 8 SFT variants across 9 pretrained checkpoints from wav2vec~2.0, HuBERT, and WavLM, with multi-seed repetitions on representative base-scale models. We find that the identity of the statistically indistinguishable top-group SFT recipe is often checkpoint-dependent, with limited transferability across pretrained instances. These findings suggest that many reported downstream gains reflect instance and seed dependent elicitation match, rather than universally improving the attainable performance ceiling.