今日论文合集:CS.SD语音与音频 | 共 9 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 2 篇
2. 音乐信息检索与音乐生成 2 篇
3. 语音翻译与语音语言模型 1 篇
4. 数据集、基准与评测 2 篇
5. 安全、隐私与深度伪造音频 1 篇
6. 其他/综合语音音频 1 篇
1. 语音识别与关键词检测 | 2 篇
1. Voice-Light: A Full-Duplex Cascaded Voice Agent with Causal Turn-Taking and Speculative Generation
Voice-Light:具有因果轮换与投机式生成的级联全双工语音智能体
AI 总结:Voice-Light提出级联全双工语音智能体,结合因果轮换与投机式生成,实现低延迟交互,并通过混合控制器在真实场景中平衡误切断与轮次召回率。
链接:https://arxiv.org/abs/2609.20995
作者:Bertil Braun
英文摘要:Natural spoken interaction requires more than streaming ASR, language generation, and speech synthesis: a system must react to overlap without canceling on every acknowledgment, prepare a response before a turn is certain, and ensure canceled audio cannot enter conversation history. We present Voice-Light, a full-duplex cascaded voice agent that combines immediate acoustic onset, a causal adapter sharing a streaming ASR encoder, reversible playback control, and private speculative response generation. Structured tool calls execute concurrently with audible bridge speech, while browser acknowledgments make rendered audio authoritative for durable history. Locked evaluation on 1,673 real-conversation silence candidates found that an earlier learned completion checkpoint preserved a 2.70% false-cutoff rate but reached only 12.53% end-of-turn recall, compared with 95.60% for a Silero timing policy. The deployed system therefore retains a hybrid controller rather than claiming a learned-policy replacement. Across three unscripted operator-run microphone sessions, 36 measured response turns had a 758 ms median from final VAD endpoint to first server audio; 21 turns were below 800 ms. These sessions are an instrumented case study, not a controlled user evaluation. We release the synthetic data, model artifacts, evaluation code and summaries, source code, and deployment configuration supporting the result.
2. Online Algorithms for Independent Low-Rank Matrix Analysis and Rank-Constrained Spatial Covariance Matrix Estimation Based on Maximum Weighted Likelihood Estimation
基于最大加权似然估计的独立低秩矩阵分析与秩约束空间协方差矩阵估计的在线算法
AI 总结:本文提出基于最大加权似然估计的ILRMA和RCSCME在线算法,通过逐帧代价函数与辅助函数技术推导更新规则,并采用近似与加速技术,实现动态场景下实时多通道语音提取,性能优于传统方法。
链接:https://arxiv.org/abs/2609.21180
机构:The University of Tokyo(东京大学); National Institute of Technology, Kagawa College(国立工业高等专门学校香川高等专门学校); Yamaha Corporation(雅马哈公司)
作者:Yuto Ishikawa, Norihiro Takamune, Tomohiko Nakamura, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo
英文摘要:Real-time multichannel speech extraction (MSE) under diffuse noise conditions is an important task with a wide range of applications, such as speech recognition and hearing aids. In this paper, we propose online algorithms for independent low-rank matrix analysis (ILRMA) and rank-constrained spatial covariance matrix estimation (RCSCME). Previously, we proposed a real-time extension of the RCSCME-based method: an MSE method based on ILRMA and RCSCME using the blockwise batch algorithm. However, it assumes that the spatial characteristics are stationary within a single batch, and thus, in dynamic situations where the target speaker moves, its performance may degrade. To address this problem, we derive the online algorithms for ILRMA and RCSCME in the following three steps. First, we formulate framewise cost functions for ILRMA and RCSCME on the basis of maximum weighted likelihood estimation. Second, we derive the update rules for the framewise cost functions on the basis of auxiliary-function techniques. These naive update rules are computationally costly for real-time execution on a practical machine. Thus, we finally derive the online algorithms by approximating some intermediate parameters with their estimates. Furthermore, we propose stabilization and further acceleration techniques for these online algorithms. In experiments, we simulate situations where a target speaker is stationary or moves and show that the proposed method achieves superior speech extraction performance compared with conventional methods. In addition, using real-world recorded signals, we demonstrate the effectiveness of the proposed method in practical scenarios.
2. 音乐信息检索与音乐生成 | 2 篇
3. Composer2Vec: A Continuous Embedding Space of Composer Style Learned from Symbolic Melody Generation
Composer2Vec:从符号旋律生成中学习的作曲家风格连续嵌入空间
AI 总结:本文提出Composer2Vec,通过作曲家条件Transformer学习符号旋律生成中的作曲家嵌入,发现其形成可解释的连续潜在空间,捕捉音乐历史结构,且优于通用音频-文本嵌入。
链接:https://arxiv.org/abs/2609.20893
作者:Sakutaro Nishio, Osamu Ichikawa
英文摘要:We analyze the composer embeddings learned by a composer-conditioned Transformer as a continuous latent space of compositional style, rather than merely as an internal representation for generation. A model that recursively predicts melody continuations was trained on melodic sequences extracted from MIDI data, conditioned on composer identity (124 composers). Principal component analysis of the learned composer embedding matrix (124x128) shows that the first principal component correlates strongly with composer birth year (r = -0.884, p < 0.001, n = 123), a stronger correlation than we obtain by applying the same PC1-birth-year analysis to existing general-purpose audio-text embeddings (CLAP, MuQ-MuLan) trained on unrelated audio-text corpora, not on symbolic melody generation. A shuffle test (2,000 permutations) confirms that the Silhouette score for stylistic-period labels is statistically significant (0.0110, p < 0.001). We further show that vector arithmetic in the embedding space captures meaningful stylistic relationships between composers. These results suggest that composer embeddings, learned without supervision beyond composer identity, form an interpretable latent space that captures musical-historical structure.
4. Investigating the Performance and Energy Costs of Replicating Band-Split RNN for Music Source Separation
研究复现频带分割循环神经网络用于音乐源分离的性能与能耗成本
AI 总结:本文复现了频带分割循环神经网络(BSRNN)音乐源分离模型,通过实验研究设计选择并报告能耗成本,强调完整流程可用性对降低碳足迹的重要性,并公开代码与预训练模型。
链接:https://arxiv.org/abs/2609.21918
机构:Université de Lorraine, CNRS(洛林大学,国家科学研究中心); Inria, LORIA(法国国家信息与自动化研究所,洛林计算机科学实验室); Centrale Med, Aix Marseille Univ(中央梅德学院,艾克斯-马赛大学); CNRS, LIS(国家科学研究中心,信息与系统实验室)
作者:Paul Magron, Romain Serizel, Constance Douwes
英文摘要:Band-split recurrent neural network (BSRNN) is a popular music source separation model that yields close to state-of-the-art results using reasonable computational resources and public datasets. It is therefore interesting from a reproducible research perspective, but achieving its performance is not straightforward since its full code is not available. In this paper, we conduct a replication of BSRNN via implementing the full pipeline. We extend the original paper's analysis by experimentally studying various design choices about data preprocessing, the optimization protocol, and architectural parameters. We report and discuss this project's energy cost, and we underline how its footprint could have been substantial lower upon availability of the full pipeline, which advocates for more reproducible research practices. To comply with this objective, we publicly release our code and pre-trained models.
3. 语音翻译与语音语言模型 | 1 篇
5. I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance
我会留意外界声音:教 AudioLLM 主动音频辅助
AI 总结:针对音频大模型被动响应的问题,提出打断与静默建模(ISM)范式,通过特殊标记实现主动音频辅助,在ESC-50和Epic-Sounds上取得最优性能,平均延迟3.5秒。
链接:https://arxiv.org/abs/2609.21183
机构:Meta Reality Labs(Meta现实实验室)
作者:Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang, Seungwhan Moon, Shashank Jain, Pinar Donmez, Babak Damavandi
英文摘要:Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \texttt{<interrupt>} and \texttt{<silent>}, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication. Applied to Qwen2-Audio-7B, ISM achieves 99.6\% interrupt F1 and perfect de-duplication recall on ESC-50. On noisy Epic-Sounds kitchen audio, ISM achieves the highest interrupt F1 without domain-specific training, the only method maintaining strong onset detection without over-triggering or over-suppression. Streaming evaluation confirms real-time viability with 3.5-second average latency.
4. 数据集、基准与评测 | 2 篇
6. The Internet Archive Music Dataset
互联网档案馆音乐数据集
AI 总结:我们提出了互联网档案馆音乐数据集(IAMD),一个包含超过34,000小时音频的最大公开音乐-字幕数据集,并设计了一种自动字幕生成流程,经评估该流程可靠且可扩展。
链接:https://arxiv.org/abs/2609.20870
作者:Paraskevas Stamatiadis (S2A, LTCI, IDS), Bernardo V Miranda (S2A, LTCI, IDS), Clémentine Berger (LTCI, IP Paris, S2A, IDS, IMT), Gaël Richard (S2A, LTCI, IDS), Mathieu Fontaine (S2A, LTCI, IDS), Slim Essid (IDS, S2A, LTCI)
英文摘要:We introduce the Internet Archive Music Dataset (IAMD), a large-scale collection of captioned music segments derived from the Internet Archive. To the best of our knowledge, IAMD constitutes the largest publicly available music-caption dataset to date with over 34,000 hours of audio, providing a valuable benchmark for training and evaluating music understanding and generative models. The dataset is built from content declared to be distributed under Creative Commons licenses, and cross-referencing with MusicBrainz is done to improve license information reliability. To annotate IAMD, we present an automatic captioning pipeline that augments base captions produced by an audio-language model (ALM) with textual metadata sourced from the Internet Archive and imputed metadata obtained using audio classification models. Caption quality is assessed objectively and subjectively, and results indicate that the annotation pipeline is reliable and does not degrade caption quality with scaling.
7. Training Music Sample Identification Models on Real Sample Pairs
基于真实样本对的音乐样本识别模型训练
AI 总结:本文提出SI嵌入模型,利用真实样本对训练,在三个基准上达到最先进性能,并首次提供完全监督训练方法,为样本识别领域奠定基础。
链接:https://arxiv.org/abs/2609.21911
机构:Universitat Pompeu Fabra(庞培法布拉大学); Sony Europe(索尼欧洲); Sony CTC America(索尼CTC美洲); BMAT Licensing S.L.(BMAT许可公司)
作者:R. Oguz Araz, Joan Serrà, Xavier Lizarraga-Seijas, Emilio Molina, Xavier Serra, Yuki Mitsufuji, Dmitry Bogdanov
英文摘要:Sample identification (SI) is the task of matching pairs of tracks, where one track is created by musically transforming an element of the other. In the absence of sample annotations at scale, the dominant training paradigm has depended on artificially creating sample pairs. Although a recently released dataset provides annotations of real sample pairs at scale, an effective training recipe is missing. In this work, we present SI Embeddings (SIE), an SI model that achieves state-of-the-art results on three benchmarks, including a large-scale test set. We show that the previous state of the art trained on artificial pairs generalizes only partially to real pairs, and that its training data limits its performance. We also show that real pairs do not fully account for SIE's performance: its architecture and training recipe contribute substantially. We provide the first fully supervised training recipe for real-world SI, establishing a strong foundation for future research in the field.
5. 安全、隐私与深度伪造音频 | 1 篇
8. GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages
GenTraceBench:跨预训练与后训练阶段追踪音频深度伪造的基准
AI 总结:GenTraceBench基准系统评估TTS模型适配后音频深度伪造指纹的稳定性,发现DPO/GRPO保留指纹而部分SFT导致漂移,多镜头注册可显著降低验证错误率。
链接:https://arxiv.org/abs/2609.21738
作者:Li Wang, Kunyu Feng, Wan Lin, Dekun Chen, Qinke Ni, Xueyao Zhang, Lei Wang, Jie Shi, Haizhou Li, Zhizheng Wu
英文摘要:Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain valid after adaptation? We present GenTraceBench, a controlled benchmark spanning five TTS architectures, 16 pre-/post-training variants, and 49,728 utterances generated with fixed texts and speaker prompts. Under a train-on-foundation, test-on-adapted protocol, we evaluate binary detection, closed-set attribution, and open-set verification. DPO and GRPO generally preserve fingerprints, whereas some SFT and pre-training-data changes cause substantial drift; effect sizes vary across three forensic backbones. Repeated training runs confirm the largest W2V-BERT attribution drop, while a data-mixture control with comparable speech quality shows that composition change need not cause drift. In W2V-BERT verification, multi-shot enrollment reduces EER for the SFT condition from 44.4% to 11.0%, whereas the SingNet-only condition remains at or above 45% EER.
6. 其他/综合语音音频 | 1 篇
9. Spherical Harmonic Sliced Wasserstein Displacement Interpolation for Acoustic Source and Reflection Density Modeling
球谐切片Wasserstein位移插值用于声源与反射密度建模
AI 总结:本文提出球谐切片Wasserstein插值方法,用于声源与反射密度建模,通过新展开高效拟合密度并评估插值,实验证明优于线性与几何插值并实现模型降阶。
链接:https://arxiv.org/abs/2609.22028
机构:NuSpace Audio
作者:Yuancheng Luo
英文摘要:Spatial room impulse responses (SRIRs) capture directional distributions of acoustic sound-sources and their reflections. However, collecting SRIRs of moving sound-sources remains a challenge, requiring complex interpolations across measurements that account for multi-path spatial-temporal dynamics. This paper investigates the Wasserstein metric and displacement for evaluating interpolated SRIR echo densities in the spherical harmonic domain. We present novel sum-of-magnitude square expansions for efficiently fitting probability density functions, maximizing likelihood, inverse sampling, and computing spherical sliced Wasserstein interpolations. Experiments compare the Wasserstein displacements and metric to linear and geometric interpolations of SRIR image-source densities on a line-path, and demonstrate model-order reduction.
