今日论文合集:CS.SD语音与音频 | 共 12 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
1. Soft Posterior Speaker Injection for Multi-Talker Speech Recognition
面向多说话人语音识别的软后验说话人注入
AI 总结:提出软后验说话人注入(SPSI)方法,通过将帧级说话人后验注入 Whisper,在 LibriSpeech、LibriCSS 多说话人语音识别任务上降低了 cpWER,表现优于序列化输出训练(SOT)。
链接:https://arxiv.org/abs/2609.01287
机构:Zhejiang Lab(浙江实验室); Zhejiang International Studies University(浙江外国语大学)
作者:Jian Zhu, Cheng Luo
英文摘要:Multi-talker automatic speech recognition (MT-ASR) remains challenging under overlapping speech. Hard diarization-based segmentation introduces irreversible errors, whereas serialized output training (SOT) avoids explicit segmentation but does not condition a pretrained encoder on speaker activity. We propose Soft Posterior Speaker Injection (SPSI): a lightweight head predicts frame-level speaker posteriors $\hat{\mathbf{P}}$ and injects them into Whisper through multi-layer feature-wise linear modulation (FiLM) and decoder speaker-memory prompts. On controlled two-speaker LibriSpeech overlap, SPSI reduces utterance-mean constrained permutation word error rate (cpWER) from 50.7\% (SOT) to 49.6\% (one-sided paired bootstrap $p{\approx}0.006$), with a larger reduction in the high-overlap bin (60.4\%$\to$58.8\%). Same-backbone speaker-auxiliary objectives and voice activity detection (VAD) pipelines do not outperform SOT; zero-shot (ZS) LibriCSS is comparable. Freeze-posterior adaptation with overlap-heavy (OV-heavy) continuation reduces held-out LibriCSS cpWER (sessions 8--9) to 32.4\% (versus 37.5\% for SOT). Ablations indicate complementary encoder FiLM and decoder prompts, and that the effective signal is a \emph{soft} simplex-valued speaker share.
2. BiMTokenizer: Preserving Semantic-Acoustic Balance in Low-Bitrate Speech Tokenization via Bidirectional State-Space Modeling
BiMTokenizer:通过双向状态空间建模在低比特率语音分词中保持语义-声学平衡
AI 总结:该研究提出单塔低比特率语音编解码器BiMTokenizer,结合双向状态空间主干与RSLQ,在更少参数下实现了优于双塔基线的声学重建与语义保留性能。
链接:https://arxiv.org/abs/2609.00562
机构:Wuhan University of Technology(武汉理工大学); NEC Laboratories Asia Pacific(NEC亚太实验室); NEC Corporation(NEC公司); The Hong Kong Polytechnic University(香港理工大学)
作者:Xin Zhang, Lin Li, Chuanbo Liu, Jianquan Liu, Kong Aik Lee
英文摘要:Speech codecs serve as bridges between continuous speech signals and large language models, yet face an inherent conflict between acoustic fidelity and semantic preservation. To mitigate this conflict, recent works increasingly adopt dual-tower architectures to decouple semantic and acoustic modeling with separate encoders. However, these dual-tower designs incur substantial architectural overhead. To avoid such complexity, we revisit the single-tower paradigm and propose BiMTokenizer, a low-bitrate speech codec (around 1.1 kbps) combining a bidirectional state-space backbone with Residual Spherical Leech Quantization (RSLQ). The bidirectional backbone strengthens temporal modeling, while RSLQ offers a fixed, well-separated lattice bottleneck for robust semantic and acoustic tokenization without learned-codebook collapse. Experiments show that BiMTokenizer achieves superior acoustic reconstruction and the lowest WER among low-bitrate codec baselines across both clean and noisy environments, while using less than half the parameters of recent dual-tower baselines. Furthermore, its robust semantic representations yield strong performance on downstream speech understanding tasks, confirming that a well-designed single-tower codec can preserve the semantic-acoustic balance at low bitrates. The code and model weights are available at this https URL.
3. A Unified Uncertainty-Aware Back-End for Speaker Verification: Scoring, Normalization, and Calibration
面向说话人验证的统一不确定性感知后端:评分、归一化与校准
AI 总结:该研究针对说话人验证后端未传递不确定性的问题,提出含UAS-Norm、UQMF的统一不确定性感知后端,在ECAPA-TDNN与ResNet上降低了等错误率并提升了区分度。
链接:https://arxiv.org/abs/2609.01221
机构:The Hong Kong Polytechnic University(香港理工大学)
作者:Junjie Li, Kong Aik Lee
英文摘要:Speaker verification back-ends commonly combine similarity scoring, score normalization, and calibration. However, speaker embeddings extracted from real-world utterances have trial-dependent reliability because of factors such as duration, noise, and channel variation. Existing uncertainty-aware methods primarily improve the speaker encoder or the initial similarity score, while the estimated uncertainty is typically not propagated through subsequent normalization and calibration. We represent each utterance by a speaker embedding, interpreted as a posterior mean, together with its covariance as an uncertainty estimate. We present a unified uncertainty-aware back-end comprising uncertainty-aware cosine scoring, uncertainty-aware AS-Norm (UAS-Norm), and uncertainty-aware Quality Measure Function calibration (UQMF). Covariance information is incorporated throughout this pipeline to adjust score scaling, cohort statistics, normalized-score combination, and calibration features. Experiments with ECAPA-TDNN and ResNet show consistent EER reductions and improved target--non-target separation across both architectures.
4. ABSE-NET: A Lightweight Neural Model for Active Binaural Speech Enhancement in Open-Fit Hearing Aids
ABSE-NET:用于开放式助听器中主动双耳语音增强的轻量级神经模型
AI 总结:本文针对开放式助听器的声泄漏问题,提出将ANC与BSE结合的轻量级ABSE-NET框架,级联BMVDR与带特征融合模块的LNN,无需入耳麦克风,性能优于现有最先进方法。
链接:https://arxiv.org/abs/2609.00966
机构:College of Computer Science, Inner Mongolia University(内蒙古大学计算机学院); College of Electronic and Information Engineering, Inner Mongolia University(内蒙古大学电子与信息工程学院)
作者:De Hu, Xue Du, Qingying Zhao, Qintuya Si
英文摘要:Open-fit hearing aids have attracted growing attention due to their superior wearing comfort. However, the open-fit design inevitably causes acoustic leakage into the ear canal, degrading the performance of existing binaural speech enhancement (BSE). To this end, we propose ABSE-NET, an active BSE framework integrating active noise control (ANC) with BSE to jointly enhance target speech and suppress acoustic leakage. The ABSE-NET pipeline cascades a binaural MVDR (BMVDR) with a lightweight neural network (LNN). The former achieves a coarse BSE, whereas the latter simultaneously cancels acoustic leakage and compensates for BMVDR-induced distortion. The LNN uses an encoder-decoder with a feature fusion module, which includes frequency-time dependency learning and convolutional attention blocks. Unlike traditional BSE+ANC solutions via adaptive filtering, ABSE-NET needs no in-ear microphone in practical deployment. Experiments validate its superiority over state-of-the-art methods. Code repository: this https URL.
5. TUTTI: Toward generalizable audio-to-score transcription via fully synthesized data
TUTTI:基于完全合成数据实现可泛化的音频到乐谱转录
AI 总结:TUTTI是一种基于合成多乐器数据训练的A2S预训练范式,可提升模型泛化能力,在多乐器A2S任务上达新SOTA,且跨乐器迁移性优异。
链接:https://arxiv.org/abs/2609.00640
作者:Jianhuai Hu, Yashan Wang, Shangda Wu, Zhancheng Guo, Shijie Liang, Wuna Meng, Chuanqi Yang, Xiaobing Li, Feng Yu, Maosong Sun
英文摘要:Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often restricts the generalization of A2S models, limiting their efficacy primarily to single-instrumentation domains. To break this dependency on scarce real-world data, we introduce TUTTI (Transformer for Unified audio-To-score Transcription trained on Synthetic multi-Instrumentation Data), a pre-training paradigm driven by a purely synthetic, large-scale dataset. Rather than using human-composed scores, we leverage a symbolic music generation model to generate a massive, highly scalable multi-instrumentation corpus and create audio-score pairs with expressive acoustic characteristics. Capitalizing on the generated data, we employ a standard Transformer encoder-decoder architecture. We empirically demonstrate that pre-training a unified attention-based model on generated, multi-instrumentation data yields a consistently stronger foundational representation than single-instrumentation training. When fine-tuned with downstream real-world datasets, TUTTI outperforms previous approaches, establishing new overall state-of-the-art results across various A2S baselines. Notably, TUTTI shows remarkable cross-instrument transferability, effectively adapting to unseen instruments with highly competitive performance. The source code and the TuttiCorpus dataset will be made publicly available at this https URL.
6. Heard but Not Heeded: Paralinguistic Information Encoding and Loss in Audio-Language Models
听见却未留意:音频-语言模型中的副语言信息编码与丢失
AI 总结:本研究分析四种开源音频-语言模型的副语言信息编码与丢失,发现模型虽在音频编码器顶层强编码说话风格,但信息在输出前退化,模型分内容、声学驱动两类,揭示了模型编码与使用信息的差距。
链接:https://arxiv.org/abs/2609.00727
机构:Carnegie Mellon University(卡内基梅隆大学)
作者:Bhuvan Koduru, Dareen Safar B Alharthi, Rita Singh, Bhiksha Raj
英文摘要:Audio language models are designed to understand speech, yet it remains unclear whether they capture how something is said beyond what is said. We present a mechanistic analysis of paralinguistic information in four open source models, Whisper-large-v2, Qwen2-Audio-7B Instruct, Qwen2.5-Omni-7B, and Chroma-4B, using the Expresso dataset with controlled speaking styles. We combine centered kernel alignment, linear probing with leave one speaker out evaluation, open ended tone prediction, and a content prosody leakage metric to trace how style information moves from the audio encoder to the final output. All models strongly encode speaking style in the late encoder, that is, the top third of the audio encoder's layers, but this information is consistently degraded before reaching the output. The projector reshapes representation geometry without removing information, while decoders differ in how much style they preserve depending on architecture and training objective. At the output level, models fall into two behaviors. Some are content driven, where predictions depend mainly on text. Others are acoustic driven, where predictions vary with speaking style. The leakage metric quantifies this difference, and qualitative results confirm it. Overall, we identify a gap between what models encode and what they use, highlighting a key limitation in current audio language models.
7. Cleaner Speech, Weaker Generalization: Revisiting Pitt-Derived Benchmarks for Alzheimer's Disease Detection
更清晰的语音,更弱的泛化能力:重新审视用于阿尔茨海默病检测的皮特衍生基准
AI 总结:本研究重新审视语音预处理与数据集整理对阿尔茨海默病检测的影响,发现经语音增强的数据集虽提升域内性能但降低跨域鲁棒性,更清晰的语音数据集未必更可靠。
链接:https://arxiv.org/abs/2609.00276
机构:Center for Language and Speech Processing (CLSP), Johns Hopkins University(约翰斯·霍普金斯大学语言与语音处理中心); Research Center for Information Technology Innovation, Academia Sinica(中央研究院资讯科技创新研究中心); University of Michigan, Ann Arbor(密歇根大学安娜堡分校); Carnegie Mellon University(卡内基梅隆大学)
作者:Luqi Sun, Shreeram Suresh Chandra, Lin Zhang, You-Jin Li, Brian MacWhinney, Yu Tsao, Emily Mower Provost, Berrak Sisman
英文摘要:Speech-based Alzheimer's disease (AD) detection increasingly relies on speech-enhanced and curated versions of the Pitt Corpus, where speech enhancement, sample selection, and demographic balancing are often treated as beneficial preprocessing steps. However, whether these transformations improve real-world AD detection or instead affect model generalization and prediction behavior remains unclear. In this work, we revisit the role of speech preprocessing and dataset curation across widely used benchmarks for speech-based AD detection. We evaluate the speech quality of different datasets, the cross-dataset generalization of multiple deep learning models under matched and mismatched enhancement settings, and the behavior of several recent large audio-language models (LALMs). Experimental results show that across multiple supervised speech models, speech-enhanced datasets often improve in-domain performance while reducing robustness in cross-domain evaluation. Matched enhancement between training and test data alleviates, but does not eliminate, this degradation. LALMs show a similar sensitivity: enhanced datasets induce stronger class imbalance and prediction shifts than unprocessed data. These results suggest that speech preprocessing and dataset curation can substantially influence downstream AD detection behavior, indicating that ``cleaner'' speech datasets are not necessarily more reliable for real-world AD detection.
8. Perceptible or Not? Diagnosing Passive Fingerprints for Speech Deepfake Attribution
可感知还是不可感知?针对语音深度伪造归因的被动指纹诊断
AI 总结:本研究提出PIPDP协议区分语音深度伪造的可感知与不可感知被动指纹,实验显示不可感知指纹的归因线索更持久可靠,为语音深度伪造归因提供了新的诊断方法。
链接:https://arxiv.org/abs/2609.00765
机构:Imperial College London(帝国理工学院); Queen Mary University of London(伦敦玛丽女王大学); Johns Hopkins University(约翰斯·霍普金斯大学); Technical University of Munich(慕尼黑工业大学)
作者:Yupei Li, Qiyang Sun, Emmanouil Benetos, Berrak Sisman, Björn Schuller
英文摘要:Passive fingerprints (intrinsic traces naturally left by generators) have been shown to enable attribution in speech deepfake detection, yet their persistence, reproducibility, and content-independence remain unverified. Moreover, no prior work distinguishes perceptible from imperceptible fingerprints, although the two have very different implications for attribution reliability. Perceptible fingerprints, such as emotional expression, are shaped by perceptual quality objectives and may change across model updates, whereas imperceptible fingerprints are not explicitly optimised by current training objectives and are rarely considered in existing dataset design or training strategies, as they have limited influence on downstream applications. We therefore propose a Perceptible-Imperceptible Passive-fingerprint Diagnostic Protocol (PIPDP) to define and separately analyze these two fingerprint types. PIPDP comprises three complementary analyses: multi-evidence fingerprint verification through residual-energy, reproducibility, and saliency analyses, perceptually transparent perturbations preserving audio quality, and prompt-driven emotion change that modifies perceptible fingerprints without model retraining. Experiments across ten speech generators and three attribution detectors show that imperceptible fingerprints provide persistent attribution cues. Perceptually transparent perturbations reduce attribution accuracy by up to 48.2\% on HiggsAudioV3, whereas emotion-driven changes leave attribution largely unchanged, with only about a 1.0\% accuracy variation across emotions on CosyVoice2 using w2v-bert-MLP. These results suggest that imperceptible fingerprints are more reliable for trustworthy attribution.
9. XVAE-WMT: Explainable Wavelet-Temporal Variational Autoencoder for Blind Source Separation of Heart and Lung Sounds
XVAE-WMT:用于心肺音盲源分离的可解释小波-时间变分自编码器
AI 总结:本文提出XVAE-WMT算法,无需配对干净心肺音记录,结合VAE、XAI与CWT前端,经SHAP降维后在两个数据集上实现26.8 dB SDR等优异分离性能,提升心肺音盲源分离效果。
链接:https://arxiv.org/abs/2609.00238
机构:McMaster University(麦克马斯特大学); Bell Canada(贝尔加拿大公司)
作者:Yasaman Torabi, Shahram Shirani, James P. Reilly
英文摘要:The separation of cardiovascular sounds is a critical task in biomedical signal processing. In this paper, we introduce XVAE-WMT1, an unsupervised explainable generative AI algorithm combining a variational autoencoder (VAE) with explainable AI (XAI), wavelet-based inputs, a post-hoc output mask, and temporal consistency (TC) loss. Unlike existing supervised and VAE-based methods that rely on Short-Time Fourier Transform (STFT) and ignore latent interpretability, XVAE-WMT requires no paired clean recordings and integrates a Continuous Wavelet Transform (CWT) front-end for superior time-frequency localization. We assessed the latent space interpretability via different metrics, with SHAP (SHapley Additive exPlanations) enabling dimensionality reduction to the top 75% of latent features while preserving separation quality. Evaluated across two datasets using Signal-to-Distortion Ratio (SDR), Signal-to-Interference Ratio (SIR), and Signal-to-Artifacts Ratio (SAR), XVAE-WMT attains 26.8 dB SDR, 32.8 dB SIR, and 28.6 dB SAR.
10. MADS: A Multiview Acoustic Descriptor Set Beyond Standard Spectral Summaries
MADS:超越标准频谱摘要的多视图声学描述符集
AI 总结:提出19维物理信息描述符集MADS,在ESC-10、ESC-50、MSoS数据集上,其分类性能优于26维MFCC基线和38维频谱摘要基线,可作为音频建模的基础描述符层。
链接:https://arxiv.org/abs/2609.00792
机构:ABV-Indian Institute of Information Technology and Management(ABV-印度信息技术与管理学院)
作者:Utsab Ghosh, Roshni Chakraborty
英文摘要:Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. While highly successful, these representations do not explicitly expose the physical dynamics of the underlying sound-generating event. We introduce MADS (Multi-view Acoustic Descriptor Set), a compact 19-dimensional physics-informed descriptor set de- signed to capture complementary spectral, temporal, mechanical, and stochastic structure in audio signals. Rather than treating sound only as a spectral pattern, MADS encodes properties related to excitation, damping, periodicity, impulsiveness, and structural consistency within a unified multi-view representation. We evaluate MADS using standard classical machine learning models on ESC-10, ESC-50, and MSoS, and compare it against two conventional handcrafted baselines: a compact 26D MFCC- based baseline and an expanded 38D spectral-summary baseline. Across ESC-10 and ESC-50, MADS achieves the strongest peak results overall, reaching 81.00% and 52.78%, respectively, while using roughly half the dimensionality of the 38D baseline. On MSoS, MADS again delivers the strongest top-end performance, reaching 67.48%. These results establish MADS not merely as a competitive standalone descriptor set, but as the foundational descriptor layer of a broader acoustically grounded representation program for future frame-level and deep-learning-compatible audio modeling.
11. On the Human and Computer Alignment of Attribute-Based Music Matches
基于属性的音乐匹配中人与计算机的对齐研究
AI 总结:本研究针对音乐匹配开展感知实验,推出MATCHA数据集,发现人类判断与计算相似度度量存在部分对齐,强调生成式AI需采用领域特定的感知评估框架。
链接:https://arxiv.org/abs/2609.00987
机构:Music Technology Group, Universitat Pompeu Fabra(庞培法布拉大学音乐技术组); Sony AI(索尼人工智能公司); Joint Research Centre, European Commission(欧盟委员会联合研究中心); Sony Group Corporation(索尼集团公司)
作者:Roser Batlle-Roca, Woosung Choi, Joan Serrà, Fabio Morreale, Wei-Hsiang Liao, Xavier Serra, Emilia Gómez, Yuki Mitsufuji
英文摘要:Recent advances in generative AI are raising ethical concerns regarding the originality of generated content and the potential replication of training data, with further implications for transparency, attribution, and intellectual property. In music, several computational approaches have been proposed to identify potential replication, using audio-based similarity metrics. Yet, their alignment with human judgments across distinct musical attributes remains underexplored. To address this gap, we conduct a perceptual experiment on music matches, defined as strongly similar musical excerpts. We focus on five musical attributes: melody, harmony, rhythm, voice, and timbre. We design a triplet-based forced-choice task comprising 300 cases, including plagiarism examples, cover songs, and AI-generated music. From this experiment, we introduce the MATCHA (Musical Attribute-based Triplet Comparison with Human Annotations) dataset: a collection of 1105 perceptual assessments of attribute-based music matches from 83 expert participants. Our findings reveal measurable agreement among participants in identifying matches across attributes. We further observe partial alignment between human judgments and computational similarity measures. Overall, this work underscores the importance of domain-specific and perceptually grounded evaluation frameworks for generative AI in creative practice.
12. Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models
人工罗塞塔石碑:基于k阶马尔可夫模型的符号拉格序列约束最大后验概率(MAP)重建
AI 总结:该研究提出ARS框架,用k阶马尔可夫模型结合约束MAP问题重建符号拉格序列,经合成与真实音频实验验证了方法可行性,为音乐片段重建提供了概率建模方案。
链接:https://arxiv.org/abs/2609.01064
机构:Abstract Math Institute(抽象数学研究所)
作者:Saanvi Raghavendran, Abhishek Bhattacharjee (Abstract Math Institute)
英文摘要:Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial information, while a raga encodes constraints limiting allowable completions. This paper formalizes a mathematical framework for this, proposing the Artificial Rosetta Stone (ARS). We separate three claims often conflated: a symbolic sequence can be reconstructed probabilistically; a sequence can be consistent with an explicit grammar; and a historical performance can be authenticated. We only support the first two. We model a raga via a finite alphabet and constraint system, using an order-k Markov model for melodic probabilities. A symmetric Dirichlet prior yields a tractable posterior. We pose missing-note reconstruction as a constrained MAP problem. For fixed-length sequences and finite-order constraints, optimization admits an exact dynamic-programming solution with worst-case time complexity $O(TN^{k+1})$. We derive the parameter count $N^k(N - 1)$, prove a concentration bound under explicit mixing assumptions, and analyze estimation error propagation. A reproducible synthetic experiment uses six raga-inspired alphabets, orders $k \in \{1, 2, 3\}$, and masking rates up to 50%. This is a proof of concept, not historical reconstruction. A real-audio feasibility pilot evaluates 30 usable sequences from 42 Yaman clips via automated pitch extraction, segmentation, and quantization. Lacking documented provenance and relying on automated transcription, this is not expert-validated archival reconstruction. Claims are tied to stated conditions, not universal properties of Hindustani music. Code: this https URL.
