今日论文合集:CS.SD语音与音频 | 共 5 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音识别与关键词检测 2 篇

2. 语音合成与声音生成 1 篇

3. 其他/综合语音音频 2 篇

1. 语音识别与关键词检测 | 2 篇

1. DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs

DoubleHelix:结合大语言模型的视听语音识别结构化跨模态融合方法

AI 总结:该研究针对视听语音识别跨模态融合缺乏结构化迭代优化的问题,提出DoubleHelix融合框架,通过三个组件实现迭代交互与退化感知增强,在LRS3数据集上大幅降低词错误率并提升噪声鲁棒性。

链接:https://arxiv.org/abs/2607.29112

作者:Ziwei Cheng, Zhenhua Tan, Zhuomin Zhu

英文摘要:Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement. We present DoubleHelix, a multimodal fusion framework that reformulates fusion as an iterative cross-modal interaction process with adaptive degradation-aware enhancement. The framework comprises three components including ReverseParallelHelix for multi-turn structured interaction with learned alignment constraints, QualitySensor for learning degradation-aware gating signals, and HelixReplication for consistency-guided conditional feature enhancement. Experiments on LRS3 demonstrate that DoubleHelix achieves 0.68% WER on clean audio, outperforming previous best results by 5.6% relative improvement under matched backbone settings. Comprehensive ablation studies validate each component contribution, including targeted analysis of design choices such as asymmetric pathway weighting. The framework shows improved robustness under evaluated babble-noise conditions, achieving 11.6% WER at SNR -5dB.

2. ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition

ParaASR:面向快速长上下文基于大语言模型的语音识别的多令牌预测

AI 总结:ParaASR是一款基于大语言模型的ASR系统,通过多令牌预测技术,在保持低错误率、低延迟的同时,支持32K长上下文及30分钟音频的快速转录,兼顾识别质量、速度与上下文长度。

链接:https://arxiv.org/abs/2607.29279

机构:StepFun(阶跃星辰); NTU(南洋理工大学); PKU(北京大学); UNSW(新南威尔士大学); SJTU(上海交通大学); USTC(中国科学技术大学)

作者:Qingjian Lin, Yuxin Li, Haoyang Zhang, Jun Chen, Yechang Huang, Feng Tian, Xie Li, Xiangyu Tony Zhang, Daijiao Liu, Yuxin Zhang, Jinglan Gong, Bo Zhao, Fei Tian, Xuerui Yang, Gang Yu, Xiangyu Zhang, Daxin Jiang

英文摘要:Audio-encoder-LLM-decoder architectures have become the dominant paradigm for modern automatic speech recognition (ASR), improving transcription quality through large-scale language modeling. However, the cost of autoregressive decoding scales with decoder size, creating a fundamental trade-off between recognition quality and serving latency. We argue this trade-off is not inherent: unlike open-ended text generation, ASR outputs are strongly anchored to the input speech signal, providing a natural inductive bias toward high-parallelism decoding. Building on this, we introduce ParaASR, an ASR system that leverages Multi-Token Prediction (MTP) to let a 4B LLM decoder emit multiple tokens per forward step. Starting from a publicly available audio-language foundation, the model first establishes a robust autoregressive recognizer and then aligns five future-token branches through a staged optimization recipe. At inference, it proposes a six-token continuation per step and admits only the verified prefix into the transcript, preserving the safety of standard autoregressive decoding. The average accepted length reaches 5.0 out of 6 proposed tokens, confirming that the deterministic structure of speech makes ASR an especially natural setting for multi-token decoding. ParaASR further retains a native 32K-context window and transcribes up to 30 minutes of audio in a single pass. Across diverse benchmarks, it attains average error rates of 2.97%, 3.68%, and 3.70% on Chinese, English, and long-form evaluations, respectively, while reaching a real-time factor (RTF) as low as 0.0053. These results show that decoder scaling, low-latency inference, and long-context transcription need not be competing goals when future-token proposals are anchored by the acoustic signal and guarded by autoregressive verification.

2. 语音合成与声音生成 | 1 篇

3. TORUS: A Test of Rendering-Understanding Self-Coherence for Unified Audio Models

TORUS:面向统一音频模型的渲染-理解自一致性测试

AI 总结:该研究提出首个针对音频原生统一模型的自一致性测试TORUS,评估发现现有统一音频模型自一致性有限,在音频编辑任务上表现不佳,级联基线的自一致性表现优于最佳统一模型。

链接:https://arxiv.org/abs/2607.28896

机构:Centific Global Solutions Inc.(森蒂菲克全球解决方案公司); University of Maryland(马里兰大学)

作者:Aryan Vijay Bhosale, Harshit Rajgarhia, Abhishek Mukherji, Dinesh Manocha

英文摘要:Unified audio models capable of audio understanding, audio generation and, increasingly, audio editing are proliferating rapidly. Yet a basic question about them remains unanswered: do the two heads of a unified model agree about the same audio? Current practice evaluates each capability in isolation on specialized benchmarks, and never asks whether a model can make sense of its own generations. We present TORUS, the first self-coherence test for audio-native unified models. TORUS comprises 48 three-stage self-coherence tests carrying 432 six-option questions spanning speech, sound and music across five task families. We holistically evaluate five open unified models alongside a Cascaded Baseline that combines state-of-the-art specialized generation, editing and understanding models. The best unified model answers 50.5% of questions against the Cascaded Baseline's 63.2% and a 16.7% chance floor. Models struggle on audio editing. Among the evaluated audio models (specialized and unified), we observe limited self-coherence, and thus position self-coherence as an essential test for future audio systems.

3. 其他/综合语音音频 | 2 篇

4. Learning to Predict Performance-induced Emotion Differences in Classical Piano Music

学习预测古典钢琴音乐中由演奏导致的情感差异

AI 总结:本研究针对古典钢琴音乐,提出Delta-VA相对回归框架,利用演奏特征预测演奏差异导致的情感偏差,经实验验证模型方向一致性高但存在幅度低估问题。

链接:https://arxiv.org/abs/2607.28876

作者:Joann Ching, Gerhard Widmer

英文摘要:Music is often used as a medium for communicating emotion, with performers shaping perceived affect through interpretation. This study addresses the challenge of identifying and predicting subtle changes in perceived emotion that are exclusively due to differences in performance. We focus on classical solo piano music, using a set of 6 commercial recordings of Bach's Well-Tempered Clavier Book I, annotated in terms of valence and arousal. By encoding the recordings through performance-specific features only, we isolate performance information from aspects of the composition itself, which tend to dominate the overall perceived emotional category. A preliminary analysis validates that these features vary meaningfully across performers. We then propose a relative regression framework, Delta-VA, to predict deviations in valence-arousal relative to an ``average'' performance, thereby focusing on the changes in emotion brought about by a specific way of playing a piece. In addition to the standard $R^2$ regression score, we introduce geometric evaluation metrics to assess the preservation of pairwise differences between performances. Results indicate high directional consistency with the ground truth, but also a compression in prediction magnitude, indicating that the model tends to underestimate expressive performance effects.

5. Do Music Foundation Models Embed Pitch in Helical Structure?

音乐基础模型是否以螺旋结构嵌入音高?

AI 总结:该研究分析音乐基础模型的中间表征,发现其音高表征形成反映八度周期性的螺旋结构,且结构特征随模型和输入声学特性变化,为阐明模型内部机制提供新方法。

链接:https://arxiv.org/abs/2607.29086

作者:Hayato Yagi, Shinnosuke Takamichi, Rin Sato, Keitaro Tanaka, Shigeo Morishima

英文摘要:This study analyzes the intermediate representations of music foundation models (MFMs) and reports the geometric structures used to represent pitch information. By inputting isolated musical notes into trained MFMs and analyzing their principal components, we reveal that the representations form a helical structure reflecting the octave periodicity of pitch. Furthermore, we show that the clarity and geometry of this helical structure vary not only across models but also with the acoustic properties of the input signals. Our analysis provides a novel approach for clarifying the internal mechanisms of MFMs.