今日论文合集:CS.SD语音与音频 | 共 3 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. Towards Quantifying Benchmark Optimization in ASR Models
针对ASR模型中基准优化的量化研究
AI 总结:本文提出量化ASR模型基准优化的方法,识别三类行为探针,发现高分开源模型存在基准优化行为,该行为可被操控且会虚增基准性能。
链接:https://arxiv.org/abs/2608.19936
机构:Hume AI Research(休姆人工智能研究院)
作者:Theo Lebryk, David Ayllon, Alice Baird, Jakub Piotr Cłapa, Jens Madsen, Panagiotis Tzirakis
英文摘要:Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

2. Unified Music Identification for Tracks and Versions
曲目与版本的统一音乐识别
AI 总结:该研究针对曲目与版本识别分开处理的问题,提出统一基准评估7种模型的准确性与鲁棒性,训练出10秒TI查询的基线模型,验证了音乐识别可统一的可行性。
链接:https://arxiv.org/abs/2608.19919
作者:R. Oguz Araz, Joan Serrà, Yuki Mitsufuji, Xavier Serra, Dmitry Bogdanov
英文摘要:Given a music database, track identification (TI) retrieves the exact track matching an audio excerpt, whereas version identification (VI) retrieves its musical versions. Traditionally, the two tasks have been addressed separately. However, as every track is its own closest version, we investigate whether VI can subsume TI. This requires VI systems to be robust to both signal manipulation and audio degradation. We therefore propose a unified benchmark that evaluates accuracy and robustness on each task. Comparing seven existing models on this benchmark, we show that none of them are both accurate and robust on both tasks. We then train a baseline model targeting both tasks and show that a unified system is possible with 10 s TI queries. Lastly, we characterize the two retrieval constraints that limit our model's TI performance. We envision extending this unification to other music identification tasks.

3. Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
傅里叶是前沿:用于高保真音乐重建的频率感知自编码
AI 总结:该研究针对音频自编码器高压缩率下的失效问题,提出ear-VAE2复频谱自编码器及双工感知修正器,在音乐重建任务的多项指标上取得最优表现。
链接:https://arxiv.org/abs/2608.19843
机构:Qwen Team, Alibaba(通义千问团队,阿里巴巴); Monash University(莫纳什大学)
作者:Kangdi Wang, Yusheng Dai, Jin Xu
英文摘要:Continuous-latent audio autoencoders form the backbone of latent music generators, yet decoders at high compression rates commonly exhibit three failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These share a structural root: waveform autoencoders lack an explicit frequency axis, leaving no handle for targeted per-band correction. Among five matched-budget representations, the complex STFT achieves the lowest full-band and high-frequency spectral distances, providing direct access to magnitude and phase at every bin. Building on this, we present ear-VAE2, a complex-spectral autoencoder with cross-channel interaction. Spec-SnakeBeta learns a periodic activation per frequency bin with frequency-dependent initialization, outperforming other activation variants while using fewer parameters than the fully independent variant. Duplex-Aware Refiner applies band-specific corrections to magnitude and phase following duplex theory of sound localization. On the 546-track Song Describer Dataset, ear-VAE2 achieves the best point estimates on five of seven reconstruction metrics. The Duplex-Aware Refiner reduces Mel Distance by 19.4% and uses ~45% fewer residual-output dimensions than the Unconstrained Refiner, while also lowering spectral distances, spatial-cue errors, and receiving higher ratings from professional engineers. The downstream generator using ear-VAE2 latents achieves better point estimates on all 12 automatic this http URL page is available at this https URL.