今日论文合集:CS.SD语音与音频 | 共 4 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

1. Alignment Drift in Single-Model Speculative Decoding for ASR: Mechanism, Correction, and Cost
面向ASR的单模型推测解码中的对齐漂移:机制、修正与代价
AI 总结:该研究针对ASR单模型推测解码的对齐漂移问题,分析其机制,提出两种修正方法,其中AnchorDraft可提升端到端速度,揭示ASR自推测依赖的关键因素。
链接:https://arxiv.org/abs/2608.12703
作者:Xinyu Wang, Huapeng Zhou, Ziyu Zhao, Silin Meng, Ke Bai, Dongming Shen, Xiao-Wen Chang, Alex Smola
英文摘要:Speculative decoding speeds up generation by letting a cheap draft propose several tokens that a target model checks in one pass. In the single-model form, the draft is a lightweight module attached to the target rather than a separate model. Applying this design to Automatic Speech Recognition (ASR) introduces an extra problem. The draft can read the whole audio at every step, yet its proposals get worse as it runs on its own. Access is not localization. The accepted text keeps the transcript position explicit, but the draft must also track the changing audio position. In the primary matched comparison, per-step audio access changes the first proposal modestly but roughly doubles later-proposal acceptance. Fixed-width windows show that the audio position explains part of this gap. A correctly placed window recovers continuation, while an equally narrow window at the wrong position reduces it. Late-draft median error reaches 21 frames in the hardest reported condition, while target attention during verification stays within a 2-frame median. We test two ways to reduce this drift. The first reads the audio position from verification attention and uses it to guide the next draft round. It saves time only when the extra accepted tokens offset the readout cost. The second is AnchorDraft, which teaches the draft to track the audio position during training without changing the inference graph. The trained draft improves end-to-end speed at both tested target scales. These results show that ASR self-speculation depends on token prediction, audio-position tracking, and draft cost.

2. VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching
VoxAudio:基于多奖励自回归流匹配的发声音频合成
AI 总结:VoxAudio是一种多奖励自回归流匹配模型,通过架构、偏好、数据层面的优化,解决现有T2A系统发声控制不足的问题,在多基准上验证了有效性与效率。
链接:https://arxiv.org/abs/2608.12951
机构:Zhejiang University(浙江大学)
作者:Wenxiang Guo, Changhao Pan, Ziyue Jiang, Fei Wu, Zhou Zhao
英文摘要:Vocalized audio synthesis, the task of generating audio in which intelligible speech is embedded within an environmental soundscape, underpins applications such as podcast production and video dubbing. Existing Text-to-Audio (T2A) systems either reduce quoted speech to unintelligible vocal murmur or delegate it to a separate TTS model with post-hoc mixing, which forfeits control over when speech occurs and how it interacts with the scene. We present VoxAudio, a causal autoregressive flow matching model that addresses this problem from three complementary aspects. At the architecture level, chunk-wise causal factorization with independent per-chunk noise levels lets audio be emitted through sliding-window streaming inference with KV caching at variable target durations; to enable inference at arbitrary chunk granularities, we further pretrain the model with randomized chunk boundaries. At the preference level, multi-reward Negative-aware FineTuning (NFT) jointly optimizes semantic fidelity, linguistic accuracy, aesthetic quality, and temporal grounding At the data level, to supply the missing supervision for vocal content, we build VoxCorpus, a large-scale corpus whose captions quote the verbatim transcript of embedded speech with time intervals, and VoxBench, an interval-annotated benchmark with a temporal-grounding metric. Experiments on four benchmarks spanning general audio, speech, and unified vocalized audio validate the effectiveness and efficiency of VoxAudio. Our code and demos are available at this https URL.

3. HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement
HybridSB-MoE:用于语音增强的双域薛定谔桥与场景自适应专家路由
AI 总结:HybridSB-MoE是一种双域语音增强框架,通过非对称不确定性融合、异构MoE路由及离散化界保证,在VoiceBank+DEMAND数据集上实现了优于扩散与SB基线的性能。
链接:https://arxiv.org/abs/2608.12715
机构:Oakland University(奥克兰大学)
作者:Zhengyi Lu, Aswini Sivakumar, Jie Hu, Yao Qiang
英文摘要:Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schrödinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.

4. Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences
音乐驱动:面向车载体验的上下文感知生成音频
AI 总结:本研究提出Drive-to-Music系统,利用行车记录仪图像与车辆遥测数据,结合感知与生成组件实现低延迟实时上下文感知车载音乐生成,为个性化自适应车载音频体验奠定基础。
链接:https://arxiv.org/abs/2608.12615
机构:Mercedes-Benz Research & Development North America(梅赛德斯-奔驰北美研发中心)
作者:Cosmin Dragoiu, Nooshin Nabizadeh
英文摘要:In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. Using dashcam imagery and vehicle telemetry, the system extracts scene semantics and driving context, maps them to high-level musical descriptors, and conditions generative audio models to produce contextually aligned soundtracks. The architecture combines perception and generative components to translate visual and kinematic inputs into structured musical attributes and synthesize audio with low latency. It supports smooth transitions as driving conditions evolve, and to ensure robustness and deployment readiness, we incorporate constraint-based controls and safety checks across the generation pipeline. Our results demonstrate the feasibility of real-time, context-aware music generation in automotive settings, providing a foundation for personalized and adaptive in-vehicle audio experiences.