微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 说话人识别、验证与分离 2 篇
3. 安全、隐私与深度伪造音频 2 篇
4. 其他/综合语音音频 1 篇
1. 语音识别与关键词检测 | 1 篇
1. VibeVoice-ASR-BitNet Technical Report
VibeVoice-ASR-BitNet技术报告
AI 总结:介绍VibeVoice-ASR-BitNet,针对边缘CPU实时推理优化。采用异构量化及渐进式量化感知训练,在ggml框架实现自定义内核和运算符,可比模型大小下速度快1.6 - 2.3倍,准确性略降。
链接:https://arxiv.org/abs/2607.21075
机构:Microsoft Research(微软研究院); Shanghai Jiao Tong University(上海交通大学); Fudan University(复旦大学); University of Chinese Academy of Sciences(中国科学院大学)
作者:Songchen Xu, Ting Song, Shaohan Huang, Zhiliang Peng, Yan Xia, Yujie Tu, Xin Huang, Jianwei Yu, Li Dong, Furu Wei
英文摘要:We present VibeVoice-ASR-BitNet, a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs. We apply heterogeneous quantization tailored to the computational characteristics of each stage: the VAE acoustic tokenizer uses full-pipeline INT8 quantization (I8_S) with kernel fusion and SIMD optimization, while the autoregressive language model adopts BitNet-style ternary weights (I2_S). To preserve accuracy under aggressive compression, we employ a progressive quantization-aware training strategy. For inference, we implement custom SIMD kernels and fused operators within the ggml framework targeting both ARM and x86 platforms, achieving real-time recognition with RTF < 1 using as few as 3 CPU threads. VibeVoice-ASR-BitNet is 1.6-2.3x faster than this http URL at comparable model sizes (~1.6 GB), with only modest accuracy degradation compared to the FP16 baseline.
2. 说话人识别、验证与分离 | 2 篇
2. Improving the performance of an ASV system using hybrid speech features
使用混合语音特征提高ASV系统的性能
AI 总结:研究旨在通过结合不同信号表示的混合特征集提升ASV系统性能,从MFCC、CQCC到RAB描述符,在谷歌语音命令数据集上于干净及有噪声场景实验,结果显示混合特征集(PNCC+RAB)能提高有噪声时的说话人验证性能。
链接:https://arxiv.org/abs/2607.20706
作者:Stanisław Ciszkiewicz, Artur Janicki
英文摘要:The growing need for secure and convenient authentication methods has led to the increasing popularity of biometric solutions. In addition to traditional and popular methods, such as fingerprint or iris scanning, voice-based approaches are also employed. User identity verification based on voice is conducted using Automatic Speaker Verification (ASV) systems. Despite their many advantages, these systems are sensitive to various types of attacks and acoustic noises, which can reduce verification accuracy. This work examines the potential to improve the performance of ASV systems by using hybrid feature sets that combine different signal representations, starting with widely-used Mel-Frequency Cepstral Coefficients (MFCC), through Constant Q Cepstral Coefficients (CQCC) and ending with the innovative RAB descriptor. Experiments were conducted on recordings from the Google Speech Commands dataset under two scenarios: in clean conditions and in the presence of acoustic noise. Finally, the systems' performance was compared using the EER metric to determine whether hybrid feature sets decrease verification error. The results show that using a hybrid feature set (PNCC+RAB) improves speaker verification performance under noisy conditions.
3. TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation
TF-MossFormer:集成卷积门控局部-全局注意力以增强时频域单声道语音分离
AI 总结:研究旨在改进单声道语音分离,提出TF-MossFormer,结合局部与全局注意力,利用内容感知滑动窗口注意力机制及卷积门控,在WSJ0-2Mix数据集上,凭借不同参数设置取得优异的SI-SDRi,性能超越先前方法。
链接:https://arxiv.org/abs/2607.21128
机构:Alibaba Group(阿里巴巴集团)
作者:Shengkui Zhao, Zexu Pan, Haoxu Wang, Biao Tian, Bin Ma, Xiangang Li
英文摘要:Transformers with global attention capture long-range dependencies but can miss the fine-grained local continuity crucial for speech separation. We propose TF-MossFormer, a time-frequency transformer that combines local and global attention to jointly model short- and long-range contexts for monaural speech separation. At its core is a content-aware sliding-window attention mechanism that dynamically adapts receptive fields for stronger local interactions, avoiding the rigidity of static convolutions. Unlike time-domain chunk-based methods, TF-MossFormer leverages the 2D spectrogram to model structure along both time and frequency axes. Convolutional gating between attention layers further improves feature selection and information flow. TF-MossFormer achieves SI-SDRi of 22.6, 24.0, and 24.4 dB on WSJ0-2Mix with 5.9M, 16.9M, and 25.4M parameters, respectively, outperforming prior approaches.
3. 安全、隐私与深度伪造音频 | 2 篇
4. Toward Interpretable Speech Deepfake Detection using Artifact-Specific Experts and Calibrated Detection Scores
使用特定伪迹专家和校准检测分数进行可解释的语音深度伪造检测
AI 总结:该研究针对语音深度伪造检测提出基于特定伪迹专家模型的框架,各专家检测特定伪迹并输出校准分数,经集成进行真假分类,保持可解释性,能捕捉合成语音可解释信号。
链接:https://arxiv.org/abs/2607.21127
机构:DEIB, Politecnico di Milano(米兰理工大学 电子、信息与生物工程系); National Institute of Informatics(国立情报学研究所)
作者:Viola Negroni, Xin Wang, Wanying Ge, Paolo Bestagini, Junichi Yamagishi, Stefano Tubaro
英文摘要:In this work, we propose an interpretable framework for speech deepfake detection based on artifact-specific expert models. Rather than relying on black-box decisions, the framework provides human-understandable evidence, which is critical in high-stakes settings. Each expert is trained to detect a specific speech synthesis artifact, and its output is calibrated into a log-likelihood ratio that serves as an interpretable evidence score. We evaluate five artifact-specific experts and show that, with proper calibration, they can capture their target artifacts and produce meaningful evidence. Importantly, each expert estimates only the presence of its assigned artifact rather than directly performing the final decision. Their outputs are aggregated into an ensemble to produce the actual real-versus-fake classification, while maintaining interpretability by indicating how strongly each expert supports or contradicts a fake classification. Results show that artifact-specific experts capture interpretable signals of synthetic speech across multiple generation pipelines.
5. Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness
研究用于神经编解码器鲁棒性的编解码器内部潜在音频水印
AI 总结:研究神经音频编解码器中用于鲁棒性的水印,将32位消息嵌入连续潜在表示,采用SEANet风格编码器 - 解码器等技术,表征水印载体移动时的权衡,在48kHz语音上提升了EnCodec - 24k比特准确率并降低了PESQ。
链接:https://arxiv.org/abs/2607.21132
作者:Zi Hu, Houmin Sun, Linxi Li, Yechen Wang, Liwei Jin, Carsten Maple, Ming Li
英文摘要:Neural audio codecs are challenging transformations for audio watermarking because they re-encode, quantize, and resynthesize speech. This paper investigates continuous latent-space watermarking for codec robustness. Instead of adding a watermark only to the waveform or spectrogram, we embed a 32-bit message into the continuous latent representation of a codec-like speech autoencoder. The pipeline uses a SEANet-style encoder-decoder, a Conformer-based message embedder, RVQ-guided latent decomposition, and a latent-domain detector trained under signal-processing and neural-codec transformations. Rather than proposing a final universal watermarking baseline, we characterize the trade-offs that appear when the watermark carrier is moved before neural decoding. On 48 kHz speech, EnCodec-aware training improves EnCodec-24k bit accuracy from 78.8% to 95.6% and 97.1%, while PESQ decreases from 3.727 to 3.514 and 3.427.
4. 其他/综合语音音频 | 1 篇
6. Spectrogram-Based Joint Detection, Localization, and Classification of Events in Continuously Recorded IBR Waveforms
基于频谱图的连续记录 IBR 波形中事件的联合检测、定位和分类
AI 总结:研究连续记录的 IBR 波形中事件的联合检测、定位和分类问题,提出基于频谱图的框架,将其转换为频谱图图像上的时间目标检测问题,通过短时傅里叶变换处理波形,实验证明该方法优于原始波形基线。
链接:https://arxiv.org/abs/2607.20817
机构:University of California, Riverside(加州大学河滨分校)
作者:Shivanshu Tripathi, Maziar Raissi, Hamed Mohsenian-Rad
英文摘要:Continuously recorded high-resolution waveform measurements provide rich information about fast power system dynamics. However, they require automated methods to identify events. This problem is addressed by developing a spectrogram-based framework to jointly detect, localize, and classify events in real-world continuously recorded waveforms at the terminal of an Inverter-Based Resource. We recast this problem as a temporal object detection problem on spectrogram images, as they capture the transient and harmonic signatures more explicitly than in raw waveform data. Each time-series waveform is transformed using the short-time Fourier transform, and the resulting per-channel spectrograms are stacked as a tensor for event detection. We benchmark this method against a detector operating directly on raw time-series measurements. Experiments on single-phase disturbances and three-phase faults demonstrate that the proposed spectrogram method consistently improves event detection, localization, and classification over the raw waveform baseline.
