今日论文合集:CS.SD语音与音频 | 共 13 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音识别与关键词检测 1 篇

2. 语音合成与声音生成 2 篇

3. 音乐信息检索与音乐生成 4 篇

4. 语音翻译与语音语言模型 2 篇

5. 数据集、基准与评测 2 篇

6. 其他/综合语音音频 2 篇

1. 语音合成与声音生成 | 1 篇

1. Best-of-N TTS Evaluation is Confounded by ASR Family Alignment

N选最佳语音合成评估受自动语音识别家族对齐的影响

AI 总结:研究发现最佳N选语音合成评估受ASR家族对齐影响,同家族验证器-评估器对效果更好。提出两种跨家族排名集成方法,能降低平均词错误率,建议交叉评估器三角测量作为默认报告实践。

链接:https://arxiv.org/abs/2607.08256

作者:aehyung Yu, Seongjae Kang

英文摘要:

Best-of-N (BoN) inference improves content consistency in zero-shot text-to-speech by selecting among multiple candidates with an automatic speech recognition (ASR) verifier. We identify an evaluation confound: the apparent quality of a verifier depends strongly on the ASR family used for evaluation. On LibriSpeech-PC with F5-TTS, verifier rankings vary substantially across Whisper, wav2vec 2.0, and HuBERT evaluators, while same-family verifier and evaluator pairs recover considerably more oracle headroom than cross-family pairs despite highly similar representations. This pattern suggests identity- or lineage-level coupling rather than general representational similarity. To mitigate this bias, we propose two cross-family rank ensembles: rank averaging and conjunctive max-rank. Both improve mean word error rate across independent evaluators without degrading automatic similarity or quality metrics, and the best ensemble achieves a 12% relative WER reduction over F5-TTS at N=10. These findings motivate cross-evaluator triangulation as a more reliable default for reporting BoN TTS performance.

2. 语音合成与声音生成 | 2 篇

2. AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation

AudioScape-TTA:用于细粒度文本到音频评估的结构化声景基准

AI 总结:针对现有TTA评估基准缺乏细粒度语义洞察的问题,提出AudioScape-TTA基准及基于评分标准的评估框架,通过2258个音频-文本对等数据验证TTA模型的多项局限,其评估更贴合人类判断。

链接:https://arxiv.org/abs/2608.04479

作者:Jinting Wang, Yuguang Yang, Shengyu Li, Yan Rong, Shan Yang, Xiaoda Yang, Li Liu

英文摘要:Text-to-audio (TTA) generation has recently achieved remarkable progress in synthesizing realistic audio from natural language descriptions. However, determining whether generated audio faithfully satisfies complex textual instructions remains challenging. Existing benchmarks mainly rely on global similarity metrics, providing limited insight into fine-grained semantic failures. To address this limitation, we introduce \textbf{AudioScape-TTA}, a structured and complexity-aware benchmark for fine-grained TTA evaluation. AudioScape-TTA represents realistic soundscapes through modality-aware semantic structures and characterizes generation complexity using event density and structural complexity. Based on these annotations, we propose a rubric-based audio-grounded evaluation framework that verifies event realization, acoustic attributes, and speech content through fine-grained semantic criteria. The benchmark contains 2,258 audio-text pairs with 25,707 binary QA rubrics, enabling scalable and interpretable analysis of TTA systems. Experiments on 13 representative open-source TTA models reveal persistent limitations in fine-grained attribute control, speech-content preservation, and compositional soundscape generation. Human validation further demonstrates that our rubric-based evaluation achieves stronger alignment with human semantic judgments than conventional global similarity metrics.

3. Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation

视觉表示很重要:利用视频到音频生成中的时间差异

AI 总结:针对现有基于条件扩散的视频到音频(V2A)生成方法需额外网络或强归纳偏置的问题,提出TD-V2A框架,利用时间差异(TD)增强视觉条件,提升了V2A生成质量。

链接:https://arxiv.org/abs/2608.04902

作者:Zehua Chen, Junyou Wang, Yuxuan Jiang, Zhenying Fang, Yusheng Dai, Jianfei Chen, Ziwei Liu, Jun Zhu

英文摘要:Video-to-audio (V2A) generation extends image-to-audio generation (I2A) by introducing consecutive frames that provide essential temporal cues for audio synthesis. However, existing conditional diffusion-based V2A methods typically enhance visual conditioning with additional audio-visual supervision, acoustic structure prediction, or reasoning from large multimodal models, requiring extra networks or strong inductive biases. Inspired by recent advances in visual representation learning, we introduce TD-V2A, which leverages temporal differences (TD) as the key representation that distinguishes V2A from I2A, enriching visual conditioning with minimal architectural modification. We first investigate TD at both the frame and feature levels to identify the most effective representation level at which TD complements visual representations. Based on these findings, we develop a hierarchically continual learning strategy and an annealed temporal differences guidance method to progressively learn and exploit TD information during diffusion training and sampling process, respectively. Extensive experiments on benchmark datasets demonstrate that effectively exploiting TD through our proposed framework significantly improves end-to-end V2A generation quality, even outperforming dedicated V2A representations such as contrastive audio-visual pretraining.

3. 音乐信息检索与音乐生成  | 4 篇

4. InvFlowFD: Reference-Free and Background-Set-Free Perceptual Music Quality Metric with Flow Matching Inversion

InvFlowFD:基于流匹配逆变换的无参考且无背景集的感知音乐质量指标

AI 总结:本研究提出InvFlowFD,利用预训练流匹配骨干网络,通过流逆变换对比先验分布,实现无参考、无背景集的音乐感知质量评估,其与人类感知高度相关且更灵活。

链接:https://arxiv.org/abs/2608.04142

作者:Alon Ziv, Harel Pogoda, Yossi Adi

英文摘要:Existing reference-free methods for evaluating music perceptual quality alleviate the need for paired noisy-clean data, but they still rely on a background set, which is used to compute aggregated statistics of clean audio samples. In this work, we propose a novel approach that eliminates this requirement, achieving background-set-free and reference-free quality estimation using only a pre-trained Flow Matching backbone. We demonstrate that unconditional Flow Matching inversion via simple Euler integration is sufficient to detect various artificial distortions and accurately rank music generation models against human perceptual judgments. We introduce InvFlowFD, which performs flow inversion and compares a group of inverted samples to the prior distribution. We evaluate our method against prior work, quantitatively and with a thorough human study. Results suggest that InvFlowFD is highly correlated with human perception of sound distortions, as well as generative models' quality, while being more flexible and less restrictive than existing metrics.

5. Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

助力音乐协同创作智能体“良好聆听”:用于理解与生成的分层自监督世界模型

AI 总结:本研究提出分层自监督世界模型,通过Swin V2编码器与条件流匹配模型构建协同音乐创作智能体,提升了和弦与调式检测准确率,可快速生成音乐建议并支持交互演示。

链接:https://arxiv.org/abs/2608.04378

作者:Scott H. Hawley

英文摘要:Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 0.996, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in 2.8 s, or 0.6 s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.

6. A Dual Evaluation for Music Transcription

音乐转录的双重评估

AI 总结:该研究针对音乐转录系统提出记谱相似度与回放相似度的双重评估框架,发现CLEWS指标相关性佳且成本低,不同评估维度偏好不同系统,新系统Rubato记谱相似度提升且回放具竞争力。

链接:https://arxiv.org/abs/2608.04511

作者:Ping Wang, Guang Yang, Nazif Can Tamer, Victoria Ebert, Noah A. Smith

机构:University of Washington(华盛顿大学) ; Allen Institute for Artificial Intelligence(艾伦人工智能研究所)

英文摘要:Automatic music transcription systems produce sheet music that can be read and played back. We argue that these two targets call for complementary evaluations of notation similarity to a reference score and playback similarity to the original performance, respectively. Our study considers notation similarity metrics from the optical music recognition literature and a wide range of playback-similarity methods validated through a listening study across over 100 participants and 230 piano recordings covering 23 works, 30 performers, and six composers. We find, fortuitously, that the playback similarity metric that correlates best with human judgments, CLEWS, is also the cheapest to run. We also find that the two evaluation dimensions favor different systems among a collection of 24 pipelines formed by pairing eight audio-to-MIDI models with three MIDI-to-score converters, with the latter component systematically determining the favored objective. The complementarity between metrics also holds when adding to the pool Rubato, a new end-to-end system that offers substantially improved notation similarity while remaining competitive, though not the best, on playback similarity.

7. Masked diffusion enables coherent beat tracking

掩码扩散实现连贯节拍跟踪

AI 总结:该研究针对节拍跟踪神经网络的无效输出问题,提出经三项改进的掩码扩散方法,减少不稳定行为并提升了节拍跟踪性能。

链接:https://arxiv.org/abs/2608.04624

作者:Francesco Foscarin, Filip Korzeniowski, Richard Vogl

英文摘要:Current neural networks for beat tracking generate invalid outputs, such as consecutive downbeats and erratic tempo changes, even when these are not present in the training data. Heavy post-processing techniques can alleviate these problems, but the original cause of this inconsistent behaviour remains unknown. We hypothesise that it stems from inadequate modelling of multiple plausible output beat grids, resulting in an invalid mixture of competing interpretations. We propose a masked diffusion approach that properly models multiple outputs and enables the model to build coherent predictions through iterative inference. We devise three modifications to standard masked diffusion that enable its application to beat tracking: independent masking of beats and downbeats during training and inference, a balanced masking scheduler for inference, and peak-picking across inference steps. Our approach reduces erratic behaviours and improves beat-tracking performance.

4. 语音翻译与语音语言模型 | 2 篇

8. HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

HyPASE:用于大音频语言模型的参数高效语音情感微调框架的双曲几何方法

AI 总结:本文提出HyPASE双曲PEFT框架,利用庞加莱球模型及HGA、EMCA组件,在MELD、IEMOCAP等数据集上优于欧氏PEFT基线,实现高效的大音频语言模型语音情感微调。

链接:https://arxiv.org/abs/2608.04351

作者:Tian Jin, Ruikang Zhang, Zefeng Zhao, Ding Luo, Jin Zeng

机构:Tongji University(同济大学) ; The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Peking University(北京大学)

英文摘要:Large Audio-Language Models (LALMs) excel at general speech understanding; however, adapting them to fine-grained tasks like Speech Emotion Recognition (SER) remains a significant bottleneck. Current Parameter-Efficient Fine-Tuning (PEFT) methods typically operate in flat Euclidean space, and this geometry fails to capture the multi-granularity nature of emotion cues, which range from low-level prosody to high-level semantics. To address this, we propose HyPASE, a hyperbolic PEFT framework for LALM-based SER. HyPASE leverages the Poincare ball model, using the hyperbolic radius as an explicit proxy for representational granularity. The framework consists of two core components: a Hyperbolic Geometric Adapter (HGA) for layer-adaptive weight modulation, and an Emotion-aware Multi-capacity Cross-modal Aggregator (EMCA) that compresses multi-scale features into compact audio prefixes. Empirical results on standard benchmarks show that HyPASE outperforms Euclidean PEFT baselines across all metrics on MELD and achieves a notable Unweighted Accuracy gain on IEMOCAP, particularly in class-imbalanced emotion recognition, with the accompanying slight Weighted Accuracy trade-off reflecting hyperbolic space's geometric prioritization of minority-class representations; furthermore, HyPASE achieves robust zero-shot cross-dataset generalization within a constrained parameter budget. By grounding the adaptation process in hyperbolic geometry, HyPASE offers a highly efficient path for LALM fine-tuning.

9. AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models

AudioDER: 一种用于后训练大型音频语言模型的去重增强推理数据集

AI 总结:针对现有音频-语言数据集冗余导致后训练效果下降的问题,提出基于声学相似性去重的数据构建流程,生成包含191k样本的推理导向数据集AudioDER,显著提升LALM在多个音频推理基准上的性能。

链接:https://arxiv.org/abs/2606.14591

作者:Hui Geng, Yi Su, Zijian Gao, Tianjiao Wan, Qisheng Xu, Jiaxin Chen, Hengzhu Liu, Kele Xu

机构:College of Computer Science and Technology, National University of Defense Technology(国防科技大学计算机科学与技术学院) ; Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院) ; Shanghai Jiaotong University(上海交通大学)

英文摘要:Recent advances in pretrained large audio-language models (LALMs) have demonstrated strong capabilities across speech, sound, and music. To adapt these models to downstream tasks without the cost of pretraining from scratch, post-training has become a widely adopted paradigm. However, the effectiveness of post-training depends critically on the quality of the training corpus. We observe that existing post-training corpora, often constructed by aggregating public audio datasets, suffer from substantial acoustic redundancy, as many of these datasets are sourced from overlapping media platforms. Such redundancy leads to repeated exposure to similar acoustic patterns, causing diminishing returns in performance despite increased data volume. address this issue, we propose a three-stage data construction pipeline that performs acoustic redundancy filtering, converts retained samples into a unified multiple-choice question-answering format with chain-of-thought generation, and finally applies quality verification and filtering. Using this pipeline, we construct AudioRE, a post-training dataset of approximately 286k instances spanning sound, speech, and music. Supervised fine-tuning on AudioRE consistently improves the performance of Qwen2-Audio-7B-Instruct across diverse audio understanding and reasoning benchmarks, outperforming models trained on the unfiltered raw corpus with substantially more instances. These results validate the effectiveness of our redundancy-aware data construction pipeline and the resulting AudioRE dataset, and further highlight the importance of minimizing acoustic redundancy in audio-language post-training. To facilitate future research, we will release both the AudioRE and the fine-tuned Qwen2-AudioRE checkpoint.

5. 数据集、基准与评测 | 2 篇

10. Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study

面向真实场景下的鲁棒版本识别:数据集、基准及微调研究

AI 总结:该研究针对真实场景下音乐版本识别的领域不匹配问题,推出大规模DiVers数据集并开展基准测试,训练出的模型在多样含噪输入上鲁棒性显著提升且在干净基准上性能稳定,相关资源已开源支持可复现。

链接:https://arxiv.org/abs/2608.04543

作者:Simon Hachmeier, R. Oguz Araz, Dmitry Bogdanov, Robert Jäschke, Xavier Serra

英文摘要:Existing datasets for musical version identification (VI) are primarily derived from curated metadata sources such as SecondHandSongs and Discogs, and are therefore dominated by professionally recorded tracks. This leads to a domain mismatch with real-world scenarios, where amateur and user-generated content is prevalent. To address this limitation, we introduce DiVers, a large-scale VI dataset comprising over 1.1 million musical versions, with train-validation-test splits compatible with established datasets such as Discogs-VI-YT, SHS100K, and Da-TACOS. In addition to standard version-level annotations, DiVers provides automatically assigned tags (e.g., instrumental, live) and segment-level predictions indicating the presence or absence of music. We evaluate the proposed dataset by training state-of-the-art VI systems. Our results show that models trained on DiVers achieve substantially improved robustness to acoustically diverse and noisy inputs, while maintaining a stable performance on cleaner, studio-quality benchmarks. We release the dataset metadata, code for its construction, and all experimental pipelines to support reproducibility.

11. Leakage-Audited Benchmarking Reveals Limited Evidence for Cross-Subject Auditory-Evoked EEG Vowel Perception Decoding

我们能从听觉EEG解码元音有多好——一个严格的跨受试基准测试与诚实评估

AI 总结:本文提出一个跨受试基准测试,评估从听觉EEG解码五类元音的性能,比较了14种方法,发现XGBoost模型在低信号条件下表现最佳,且经典方法与深度学习模型竞争。

链接:https://arxiv.org/abs/2605.00865

作者:Xiaoyang Li, Zeyan Tao

机构:College of Medicine and Biological Information Engineering, Northeastern University(医学与生物信息工程学院,东北大学)

英文摘要:We tested whether auditory-evoked EEG supports subject-independent five-vowel perception decoding when trial identity, model identity, prediction provenance, and participant-level inference are controlled within a single benchmark. We reconstructed Study 2 event tables from OpenNeuro ds006104 version 1.0.1 and analyzed the consonant-vowel pair task. One-to-one marker-stimulus pairing yielded 3,840 independent trials; control-condition selection and artifact rejection retained 1,094 epochs from 16 participants and 61 EEG channels. Thirteen unique implementations were evaluated using leave-one-subject-out testing, with participant metrics reconstructed from 36,102 trial predictions across 33 complete prediction replicas. Random Forest was numerically highest at 21.474% balanced accuracy (95% participant-bootstrap interval, 19.526-23.482%; chance, 20%), but neither its participant-level tests nor any implementation survived correction across the 13-model family. Deep-model performance was close to chance, and several architectures showed substantial seed-dependent variation and low trial-label agreement. In a separate descriptive sensor-space representation, participant-associated effects accounted for 72.24% of the balanced standardized centroid sum of squares, compared with 2.04% for vowel-associated effects; between-participant same-vowel distances exceeded within-participant across-vowel distances for all 16 participants. An exploratory MDM analysis comprising 9,616 genuine refits across training cohorts of 3-15 participants showed no monotonic performance gain. Within this dataset and protocol, evidence for reliable cross-subject five-vowel decoding is limited. The benchmark provides a reproducible chain from source rows to retained epochs, predictions, participant-level metrics, multiplicity-adjusted inference, and bounded diagnostic analyses.


6. 其他/综合语音音频 | 2 篇


12. Deep Learning for Real-Time Sound Order Recognition in Human-Robot Interaction

面向人机交互中实时声音时序识别的深度学习方法

AI 总结:针对人机交互中重叠声音时序识别的难题,提出多分支CNN结合注意力融合的深度学习框架,在三类实验条件下取得较高准确率,验证了实时可行性。

链接:https://arxiv.org/abs/2608.04072

作者:Rezaul Tutul, Usaid Khan, Andre Jakob, Ilona Buchem

英文摘要:Recognizing the temporal order of overlapping sounds is an underexplored challenge in human-robot interaction (HRI), with direct relevance to applications such as first responder detection systems. This paper presents a deep learning framework for real-time sound order recognition using recordable buzzers that emit distinct non-verbal sounds (cat meows, dog barks, helicopter noises). A multi-branch convolutional neural network (CNN) processes Mel spectrograms, Mel-frequency cepstral coefficients (MFCCs), and short-time Fourier transform (STFT) features, with an attention-based fusion mechanism to emphasize critical temporal cues. Experiments were conducted under same-amplitude, varied-amplitude, and unseen sound conditions. The proposed system achieved 99% accuracy in balanced overlaps, 91% under amplitude variation, and 74% on unseen test data with normalization. These results demonstrate that deep learning can reliably recognize sound order in overlapping conditions, supporting practical HRI scenarios. While experiments were conducted on carefully controlled synthetic overlaps, we additionally report latency benchmarks demonstrating real-time feasibility and provide an extended discussion on generalization, ecological validity, and deployment challenges in real-room environments.

13. Smartphone Audio Based Distress Detection

基于智能手机音频的遇险检测

AI 总结:本研究提出名为Always Alert的智能手机音频遇险检测系统,采用两阶段SVM监督学习框架,经多环境音频训练后,可在日常环境中实现高遇险检测率且误报率低,平均开销仅约每3-4小时1条Facebook帖子。

链接:https://arxiv.org/abs/2608.04176

作者:Anil Sharma, Sarthak Ahuja, Mayank Gautam, Sanjit Kaul

机构:IIIT-Delhi(德里印度信息技术学院)

英文摘要:We investigate an unobtrusive and 24X7 human distress detection and signaling system, Always Alert, that requires the smartphone, and not its human owner, to be on alert. The system leverages the microphone sensor, at least one of which is available on every phone, and assumes the availability of a data network. We propose a novel two-stage supervised learning framework, using support vector machines (SVMs), that executes on a user's smartphone and monitors natural vocal expressions of fear---screaming and crying in our study---when a human being is in harm's way. The challenge is to achieve a high distress detection rate while ensuring that the false alarm rate is a manageable overhead, while a typical smartphone user goes about living life as usual. We train the learning framework with carefully selected audio fingerprints of distress and of varied environmental contexts. The audio is used to tune the learning framework to obtain a desirable distress detection rate and false alarm rate (FAR). The ability of the proposed framework to detect distress in rather challenging audio environments is demonstrated. Exploiting the time contiguous nature of false alarms further allows us to reduce the FAR. We show the feasibility of using our framework anytime and anywhere by testing it over many hours of audio fingerprints recorded by volunteers on their smartphones, as they went about their daily routines. We are able to achieve high distress detection rates at an average overhead that is equivalent to about 1 facebook post every 3 to 4 hours.