今日论文合集:cs.SD语音10篇,eess.AS音频处理9篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Efficient Vocal Source Separation Through Windowed Sink Attention
标题:基于窗口Sink Attention的有效声源分离
链接:https://arxiv.org/abs/2510.25745

作者:Christodoulos Benetatos, Yongyi Zang, Randal Leistikow
摘要:像Mel Band Roformer这样的最先进的声音分离模型依赖于完全的时间自注意机制,其中每个时间帧与每个其他帧相互作用。这导致了大量的计算成本,其与输入音频长度成二次方地缩放,从而激发了分块和窗口方法。通过分析预先训练的语音分离模型,我们发现时间注意模式是高度本地化的。在此基础上,我们用具有小时间注意窗口和注意汇的窗口化汇注意(WSA)取代了完全注意。我们的经验表明,从原来的检查点微调恢复92%的原始SDR性能,同时减少44.5倍的FLOPs。我们在https://github.com/smulelabs/windowed-roformer上发布了MIT许可证下的代码和检查点。
摘要:State-of-the-art vocal separation models like Mel-Band-Roformer rely on full temporal self-attention mechanisms, where each temporal frame interacts with every other frames. This incurs heavy computational costs that scales quadratically with input audio length, motivating chunking and windowing approaches. Through analysis of a pre-trained vocal separation model, we discovered that temporal attention patterns are highly localized. Building on this insight, we replaced full attention with windowed sink attention (WSA) with small temporal attention window and attention sinks. We show empirically that fine-tuning from the original checkpoint recovers 92% of the original SDR performance while reducing FLOPs by 44.5x. We release our code and checkpoints under MIT license at https://github.com/smulelabs/windowed-roformer.


【2】Binaspect -- A Python Library for Binaural Audio Analysis, Visualization & Feature Generation
标题:Binaspect --用于双耳音频分析、可视化和特征生成的Python库
链接:https://arxiv.org/abs/2510.25714

作者:Dan Barry, Davoud Shariat Panah, Alessandro Ragano, Jan Skoglund, Andrew Hines
摘要:我们提出了Binaspect,一个用于双耳音频分析,可视化和特征生成的开源Python库。Binaspect通过计算修改后的耳间时间和水平差频谱图,并将这些时间-频率(TF)箱聚类为稳定的时间-方位直方图表示来生成可解释的“方位图”。这使得多个活动源表现为不同的方位角集群,而退化则表现为加宽、扩散或移动的分布。至关重要的是,Binaspect对音频进行盲目操作,不需要预先了解头部模型。这些可视化使研究人员和工程师能够观察双耳线索如何被编解码器和渲染器设计选择以及其他下游过程所降级。我们展示了比特率阶梯,立体混响渲染和VBAP源定位,其中降级清楚地揭示了工具。除了它们的诊断价值之外,所提出的表示还可以导出为结构化特征,适用于在质量预测、空间音频分类和其他双耳任务中训练机器学习模型。Binaspect是在开放源代码许可下发布的,具有完整的可重复性脚本,网址为https://github.com/QxLabIreland/Binaspect。
摘要:We present Binaspect, an open-source Python library for binaural audio analysis, visualization, and feature generation. Binaspect generates interpretable "azimuth maps" by calculating modified interaural time and level difference spectrograms, and clustering those time-frequency (TF) bins into stable time-azimuth histogram representations. This allows multiple active sources to appear as distinct azimuthal clusters, while degradations manifest as broadened, diffused, or shifted distributions. Crucially, Binaspect operates blindly on audio, requiring no prior knowledge of head models. These visualizations enable researchers and engineers to observe how binaural cues are degraded by codec and renderer design choices, among other downstream processes. We demonstrate the tool on bitrate ladders, ambisonic rendering, and VBAP source positioning, where degradations are clearly revealed. In addition to their diagnostic value, the proposed representations can be exported as structured features suitable for training machine learning models in quality prediction, spatial audio classification, and other binaural tasks. Binaspect is released under an open-source license with full reproducibility scripts at https://github.com/QxLabIreland/Binaspect.


【3】Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
标题:用知识驱动的多重假设控制对比自我监督学习:应用于节拍跟踪
链接:https://arxiv.org/abs/2510.25560

作者:Antonin Gagnere, Slim Essid, Geoffroy Peeters
摘要:数据和问题约束中的模糊性可能会导致机器学习任务的多样化,同样合理的结果。例如,在节拍和强拍跟踪中,不同的听众可能会采用各种节奏解释,其中没有一个必然是不正确的。为了解决这个问题,我们提出了一种对比的自我监督预训练方法,该方法利用了关于数据中可能的阳性样本的多个假设。我们的模型经过训练,学习与不同假设兼容的表示,这些假设是用基于知识的评分函数选择的,以保留最合理的假设。当对标记数据进行微调时,我们的模型在标准基准测试中的表现优于现有方法,特别是在音乐表示学习中将领域知识与多假设选择相结合的优势。
摘要:Ambiguities in data and problem constraints can lead to diverse, equally plausible outcomes for a machine learning task. In beat and downbeat tracking, for instance, different listeners may adopt various rhythmic interpretations, none of which would necessarily be incorrect. To address this, we propose a contrastive self-supervised pre-training approach that leverages multiple hypotheses about possible positive samples in the data. Our model is trained to learn representations compatible with different such hypotheses, which are selected with a knowledge-based scoring function to retain the most plausible ones. When fine-tuned on labeled data, our model outperforms existing methods on standard benchmarks, showcasing the advantages of integrating domain knowledge with multi-hypothesis selection in music representation learning in particular.


【4】Studies for : A Human-AI Co-Creative Sound Artwork Using a Real-time Multi-channel Sound Generation Model
标题:研究:使用实时多通道声音生成模型的人类-AI合作创作声音艺术品
链接:https://arxiv.org/abs/2510.25228

作者:Chihiro Nagashima, Akira Takahashi, Zhi Zhong, Shusuke Takahashi, Yuki Mitsufuji
备注:Accepted at NeurIPS Creative AI Track 2025, 9 pages, 6 figures, 1 table, Demo page: this https URL
摘要:本文通过与声音艺术家Evala合作开发的生成声音装置Studies for(https://www.ntticc.or.jp/en/archive/works/studies-for/),探索人工智能技术与艺术工作流程的整合。该装置采用了轻量级但高质量的声音生成AI模型SpecMaskGIT,实时生成和播放八声道声音,在为期三个月的展览中创造出身临其境的听觉体验。该作品以“新形式的档案”为概念,旨在保留艺术家的艺术风格,同时通过不断产生新的声音元素来扩展艺术家过去的作品。这种推测性的档案保存方法是通过在由Evala过去200多个小时的声音艺术作品组成的数据集上训练AI模型来促进的。   通过解决使用人工智能共同创作艺术的关键要求,本研究强调了以下方面的价值:(1)整合艺术家反馈的必要性,(2)来自艺术家过去作品的数据集,以及(3)确保包含意想不到的新颖输出。在Studies for中,该模型的设计旨在反映艺术家的艺术身份,同时产生新的,以前从未听过的声音,使其成为“一种新形式的档案”概念的恰当实现。“我们提出了一个人类-人工智能共同创作框架,用于有效地将声音生成人工智能模型纳入声音艺术创作过程,并提出了创建和存档声音艺术的新可能性,这些声音艺术将艺术家的作品扩展到他们的物理存在之外。演示页面:https://sony.github.io/studies-for/
摘要:This paper explores the integration of AI technologies into the artistic workflow through the creation of Studies for, a generative sound installation developed in collaboration with sound artist Evala (https://www.ntticc.or.jp/en/archive/works/studies-for/). The installation employs SpecMaskGIT, a lightweight yet high-quality sound generation AI model, to generate and playback eight-channel sound in real-time, creating an immersive auditory experience over the course of a three-month exhibition. The work is grounded in the concept of a "new form of archive," which aims to preserve the artistic style of an artist while expanding beyond artists' past artworks by continued generation of new sound elements. This speculative approach to archival preservation is facilitated by training the AI model on a dataset consisting of over 200 hours of Evala's past sound artworks.   By addressing key requirements in the co-creation of art using AI, this study highlights the value of the following aspects: (1) the necessity of integrating artist feedback, (2) datasets derived from an artist's past works, and (3) ensuring the inclusion of unexpected, novel outputs. In Studies for, the model was designed to reflect the artist's artistic identity while generating new, previously unheard sounds, making it a fitting realization of the concept of "a new form of archive." We propose a Human-AI co-creation framework for effectively incorporating sound generation AI models into the sound art creation process and suggest new possibilities for creating and archiving sound art that extend an artist's work beyond their physical existence. Demo page: https://sony.github.io/studies-for/


【5】SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
标题:SFMS-ALR:具有自适应区域分辨率的脚本优先多语言语音合成
链接:https://arxiv.org/abs/2510.25178

作者:Dharma Teja Donepudi
备注:10 pages, 2 figures, 1 table. Demonstration prototype available at this https URL
摘要:句内多语言语音合成(代码转换TTS)仍然是一个主要的挑战,由于突然的语言转换,不同的脚本,和不匹配的韵律语言之间。传统的TTS系统通常是单语的,并且在混合语言环境中不能产生自然的、可理解的语音。我们介绍脚本优先多语言合成与自适应区域分辨率(SFMS-ALR),一个引擎不可知的框架流畅,实时代码切换语音生成。SFMS-ALR通过Unicode脚本对输入文本进行分段,应用自适应语言识别来确定每个分段的语言和区域,并使用情感感知调整来规范韵律以保持跨语言的表达连续性。该算法生成一个统一的SSML表示与适当的“lang”或“voice”跨度和合成的话语在一个单一的TTS请求。与端到端的多语言模型不同,SFMS-ALR不需要再培训,并且可以与Google、Apple、Amazon和其他提供商的现有语音无缝集成。与数据驱动管道(如Unicom和Mask LID)的比较分析证明了SFMS-ALR的灵活性,可解释性和可立即部署性。该框架建立了一个模块化的高质量,独立于引擎的多语言TTS的基线,并概述了可理解性,自然度和用户偏好的评估策略。
摘要:Intra-sentence multilingual speech synthesis (code-switching TTS) remains a major challenge due to abrupt language shifts, varied scripts, and mismatched prosody between languages. Conventional TTS systems are typically monolingual and fail to produce natural, intelligible speech in mixed-language contexts. We introduce Script-First Multilingual Synthesis with Adaptive Locale Resolution (SFMS-ALR), an engine-agnostic framework for fluent, real-time code-switched speech generation. SFMS-ALR segments input text by Unicode script, applies adaptive language identification to determine each segment's language and locale, and normalizes prosody using sentiment-aware adjustments to preserve expressive continuity across languages. The algorithm generates a unified SSML representation with appropriate "lang" or "voice" spans and synthesizes the utterance in a single TTS request. Unlike end-to-end multilingual models, SFMS-ALR requires no retraining and integrates seamlessly with existing voices from Google, Apple, Amazon, and other providers. Comparative analysis with data-driven pipelines such as Unicom and Mask LID demonstrates SFMS-ALR's flexibility, interpretability, and immediate deployability. The framework establishes a modular baseline for high-quality, engine-independent multilingual TTS and outlines evaluation strategies for intelligibility, naturalness, and user preference.


【6】Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
标题:基于带有部分标签的声音事件半监督训练的声音场景和声音事件联合分析
链接:https://arxiv.org/abs/2510.25075

作者:Keisuke Imoto
备注:Accepted to APSIPA Transactions on Signal and Information Processing
摘要:注释声音事件的时间边界是劳动密集型的,限制了音频检测中强监督学习的可扩展性。为了降低标注成本,仅使用剪辑级标签的弱监督学习已被广泛采用。作为替代方案,部分标签学习提供了一种具有成本效益的方法,其中提供了一组可能的标签,而不是精确的弱注释。然而,用于音频分析的部分标签学习在很大程度上仍未被探索。出于观察,声学场景提供上下文信息构建一组可能的声音事件,我们利用声学场景信息来构建部分标签的声音事件。在此基础上,本文提出了一个多任务学习框架,共同执行声学场景分类和声音事件检测与部分标签的声音事件。在降低标注成本的同时,弱监督和部分标签学习往往会由于缺乏精确的事件集及其时间标注而降低检测性能。为了更好地平衡注释成本和检测性能,我们还探索了一个利用强标签和部分标签的半监督框架。此外,为了细化部分标签,实现更好的模型训练,我们提出了一种基于自蒸馏的标签细化方法,所提出的方法与部分标签。
摘要:Annotating time boundaries of sound events is labor-intensive, limiting the scalability of strongly supervised learning in audio detection. To reduce annotation costs, weakly-supervised learning with only clip-level labels has been widely adopted. As an alternative, partial label learning offers a cost-effective approach, where a set of possible labels is provided instead of exact weak annotations. However, partial label learning for audio analysis remains largely unexplored. Motivated by the observation that acoustic scenes provide contextual information for constructing a set of possible sound events, we utilize acoustic scene information to construct partial labels of sound events. On the basis of this idea, in this paper, we propose a multitask learning framework that jointly performs acoustic scene classification and sound event detection with partial labels of sound events. While reducing annotation costs, weakly-supervised and partial label learning often suffer from decreased detection performance due to lacking the precise event set and their temporal annotations. To better balance between annotation cost and detection performance, we also explore a semi-supervised framework that leverages both strong and partial labels. Moreover, to refine partial labels and achieve better model training, we propose a label refinement method based on self-distillation for the proposed approach with partial labels.


【7】A Parameter-Efficient Multi-Scale Convolutional Adapter for Synthetic Speech Detection
标题:用于合成语音检测的参数高效多尺度卷积适配器
链接:https://arxiv.org/abs/2510.24852

作者:Yassine El Kheir, Fabian Ritter-Guttierez, Arnab Das, Tim Polzehl, Sebastian Möller
备注:6 pages
摘要:最近的合成语音检测模型通常通过微调来适应预先训练的SSL模型,这在计算上是有要求的。参数高效微调(PEFT)提供了一种替代方案。然而,现有的方法缺乏特定的归纳偏差所需的多尺度时间伪影特性的欺骗音频建模。本文介绍了多尺度卷积适配器(MultiConvAdapter),一个参数高效的架构,旨在解决这一限制。MultiConvAdapter在SSL编码器中集成了并行卷积模块,便于跨多个时间分辨率同时学习区分特征,捕获短期伪影和长期失真。MultiConvAdapter只需317万美元的可训练参数(SSL主干的1\%$),就大大减少了自适应的计算负担。在五个公共数据集上的评估表明,与完全微调和已建立的PEFT方法相比,MultiConvAdapter实现了卓越的性能。
摘要:Recent synthetic speech detection models typically adapt a pre-trained SSL model via finetuning, which is computationally demanding. Parameter-Efficient Fine-Tuning (PEFT) offers an alternative. However, existing methods lack the specific inductive biases required to model the multi-scale temporal artifacts characteristic of spoofed audio. This paper introduces the Multi-Scale Convolutional Adapter (MultiConvAdapter), a parameter-efficient architecture designed to address this limitation. MultiConvAdapter integrates parallel convolutional modules within the SSL encoder, facilitating the simultaneous learning of discriminative features across multiple temporal resolutions, capturing both short-term artifacts and long-term distortions. With only $3.17$M trainable parameters ($1\%$ of the SSL backbone), MultiConvAdapter substantially reduces the computational burden of adaptation. Evaluations on five public datasets, demonstrate that MultiConvAdapter achieves superior performance compared to full fine-tuning and established PEFT methods.


【8】Separating peripheral and higher-level effects on speech intelligibility using a hearing loss simulator and an objective intelligibility measure
标题:使用听力损失模拟器和客观可懂度测量分离对语音可懂度的外围和高层影响
链接:https://arxiv.org/abs/2510.25235

作者:Toshio Irino, Ayako Yamamoto, Fuki Miyazaki
备注:This is a manuscript that was submitted to Trends in Hearing on October 29, 2025
摘要:本文提出了一种新的方法来分离周边听力损失(HL)和更高级别的过程对语音清晰度(SI)的影响。在以前的研究中,我们进行了SI实验与14个老年人(OA)的听众,使用语音噪声的声音,无论是处理与一个理想的比率掩模(ESTA)增强技术或未经处理。目前的研究涉及一个SI实验与15个年轻的,听力正常(YNH)的听众。该实验使用了用WHIS模拟器处理的模拟HL声音,其反映了来自先前研究的特定OA的听力水平。结果表明,目标OA的SI评分高于平均YNH评分。这意味着目标OA的更高级别的过程可能比平均YNH更有效。为了了解其他OAs的特点,我们使用GESI客观可懂度测量来预测SI。首先,我们证实,GESI可以相当准确地预测的YNH和OA听众的SI分数。接下来,我们使用YNH实验中估计的参数预测了14名OA听众的SI分数。结果表明,一些OA的SI得分高于平均YNH,而一个OA的得分较低。这些差异可能反映了更高层次的加工效率的差异。这些结果表明,WHIS和GESI可以促进YNH和OA听众之间的对比实验,无论听力水平。这将使我们能够单独研究OA听众中更高级别过程的影响。
摘要:This paper presents a new method for separating the effects of peripheral hearing loss (HL) and higher-level processes on speech intelligibility (SI). In a previous study, we conducted an SI experiment with 14 older adult (OA) listeners, using speech-in-noise sounds that were either processed with an ideal ratio mask (IRM) enhancement technique or left unprocessed. The current study involved an SI experiment with 15 young, normal-hearing (YNH) listeners. This experiment used simulated HL sounds processed with the WHIS simulator that reflected the hearing level of a specific OA from the previous study. The results showed that the target OA's SI scores were higher than the average YNH scores. This implies that the target OA's higher-level processes may be more effective than those of the average YNH. To understand the characteristics of other OAs, we used the GESI objective intelligibility measure to predict SI. First, we confirmed that GESI could fairly accurately predict the SI scores for both the YNH and OA listeners. Next, we predicted the SI scores of the 14 OA listeners using the parameters estimated in the YNH experiment. The results showed that some OAs had higher SI scores than the average YNH, while one OA had lower scores. These differences in SI scores may reflect variations in the efficiency of higher-level processes.These results imply that WHIS and GESI could facilitate contrastive experiments between YNH and OA listeners, regardless of hearing level. This would allow us to study the effects of higher-level processes in OA listeners individually.


【9】State Space and Self-Attention Collaborative Network with Feature Aggregation for DOA Estimation
标题:状态空间和具有特征聚集的自注意协作网络用于到达方向估计
链接:https://arxiv.org/abs/2510.25193

作者:Qi You, Qinghua Huang, Yi-Cheng Lin
摘要:由于声源的声学特性在时间和频率上是连续变化的,因此精确的声源波达方向(DOA)估计具有挑战性。在这种情况下,准确的定位依赖于聚合相关特征和有效建模时间依赖性的能力。在时间序列建模中,实现模型性能和计算效率之间的平衡仍然是一个重大挑战。为了解决这个问题,我们提出了FA-Statemer,一个具有特征聚合的状态空间和自注意力协作网络。建议的网络首先采用一个功能聚合模块,以提高跨时间和频谱维度的信息功能。其次是一个轻量级的Conformer架构,灵感来自挤压和激励机制,其中前馈层被压缩,以减少冗余和参数开销。此外,还引入了时间移位机制来扩展卷积层的感受野,同时保持紧凑的内核大小。为了进一步增强序列建模能力,引入了双向Mamba模块,从而实现了前向和后向方向上的时间依赖性的基于状态空间的高效表示。剩余的自我注意层与Mamba块相结合,形成了一个协作建模框架,实现了表示能力和计算效率之间的平衡。大量的实验表明,FA-Stateformer实现了优越的性能和效率相比,传统的架构。
摘要:Accurate direction-of-arrival (DOA) estimation for sound sources is challenging due to the continuous changes in acoustic characteristics across time and frequency. In such scenarios, accurate localization relies on the ability to aggregate relevant features and model temporal dependencies effectively. In time series modeling, achieving a balance between model performance and computational efficiency remains a significant challenge. To address this, we propose FA-Stateformer, a state space and self-attention collaborative network with feature aggregation. The proposed network first employs a feature aggregation module to enhance informative features across both temporal and spectral dimensions. This is followed by a lightweight Conformer architecture inspired by the squeeze-and-excitation mechanism, where the feedforward layers are compressed to reduce redundancy and parameter overhead. Additionally, a temporal shift mechanism is incorporated to expand the receptive field of convolutional layers while maintaining a compact kernel size. To further enhance sequence modeling capabilities, a bidirectional Mamba module is introduced, enabling efficient state-space-based representation of temporal dependencies in both forward and backward directions. The remaining self-attention layers are combined with the Mamba blocks, forming a collaborative modeling framework that achieves a balance between representation capacity and computational efficiency. Extensive experiments demonstrate that FA-Stateformer achieves superior performance and efficiency compared to conventional architectures.


【10】Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
标题:保留混合表示以实现域广义异常声音检测
链接:https://arxiv.org/abs/2510.25182

作者:Phurich Saengthong, Tomoya Nishida, Kota Dohi, Natsuo Yamashita, Yohei Kawaguchi
备注:Submitted to ICASSP 2026
摘要:野外异常声音检测(ASD)需要对分布变化(例如机器和噪声类型的不可见低SNR输入混合)具有鲁棒性。最先进的系统从自适应音频编码器中提取嵌入,并通过最近邻搜索检测异常,但对嘈杂的机器声音进行微调通常会起到去噪的作用,在不匹配的混合或不一致的标记下抑制噪声并减少泛化。具有冻结自监督学习(SSL)编码器的免训练系统避免了这个问题,并显示出强大的第一次泛化能力,但当混合嵌入偏离清洁源嵌入时,它们的性能会下降。我们建议使用保留不去噪策略来改进SSL主干,该策略可以更好地保留来自混合声源的信息。该方法将多标签音频标记损失与混合物对齐损失相结合,该混合物对齐损失将学生混合嵌入与干净和噪声输入的凸教师嵌入对齐。对固定、非固定和不匹配噪声子集的控制实验表明,在分布变化下具有更好的鲁棒性,缩小了与Oracle混合表示的差距。
摘要:Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect anomalies via nearest-neighbor search, but fine tuning on noisy machine sounds often acts like a denoising objective, suppressing noise and reducing generalization under mismatched mixtures or inconsistent labeling. Training-free systems with frozen self-supervised learning (SSL) encoders avoid this issue and show strong first-shot generalization, yet their performance drops when mixture embeddings deviate from clean-source embeddings. We propose to improve SSL backbones with a retain-not-denoise strategy that better preserves information from mixed sound sources. The approach combines a multi-label audio tagging loss with a mixture alignment loss that aligns student mixture embeddings to convex teacher embeddings of clean and noise inputs. Controlled experiments on stationary, non-stationary, and mismatched noise subsets demonstrate improved robustness under distribution shifts, narrowing the gap toward oracle mixture representations.


eess.AS音频处理


【1】Lost in Phonation: Voice Quality Variation as an Evaluation Dimension for Speech Foundation Models
标题:迷失在发音中:作为语音基础模型评估维度的语音质量变化
链接:https://arxiv.org/abs/2510.25577

作者:Harm Lameris, Shree Harsha Bokkahalli Satish, Joakim Gustafson, Éva Székely
备注:8 pages, 3 figures, 4 tables, submitted to LREC 2026
摘要:语音基础模型(SFM)的最新进展已经使得能够从原始音频直接处理口语,绕过中间文本表示。这种能力允许SIM暴露于并潜在地响应于嵌入在输入语音信号中的丰富的语言学变化。一个未被充分研究的语言变异的维度是声音质量,包括发声类型,如吱吱作响的声音和呼吸声。这些发声类型被认为会影响听者如何推断言语中的情感状态、立场和社会意义。现有的语音理解基准在很大程度上依赖于多项选择问题回答(MCQA)格式,这是容易失败,因此在捕捉细微差别的方式不可靠的语言特征影响模型的行为。在本文中,我们通过开放式生成任务和语音情感识别来探测SFM,评估模型行为在不同的发声输入中是否一致。我们引入了一个新的并行数据集,其特征在于对语音质量的合成修改,旨在评估SFM对吱吱作响和呼吸声的响应。我们的工作提供了第一次检查SFM的敏感性,这些特定的非词汇方面的言语感知。
摘要:Recent advances in speech foundation models (SFMs) have enabled the direct processing of spoken language from raw audio, bypassing intermediate textual representations. This capability allows SFMs to be exposed to, and potentially respond to, rich paralinguistic variations embedded in the input speech signal. One under-explored dimension of paralinguistic variation is voice quality, encompassing phonation types such as creaky and breathy voice. These phonation types are known to influence how listeners infer affective state, stance and social meaning in speech. Existing benchmarks for speech understanding largely rely on multiple-choice question answering (MCQA) formats, which are prone to failure and therefore unreliable in capturing the nuanced ways paralinguistic features influence model behaviour. In this paper, we probe SFMs through open-ended generation tasks and speech emotion recognition, evaluating whether model behaviours are consistent across different phonation inputs. We introduce a new parallel dataset featuring synthesized modifications to voice quality, designed to evaluate SFM responses to creaky and breathy voice. Our work provides the first examination of SFM sensitivity to these particular non-lexical aspects of speech perception.


【2】PitchFlower: A flow-based neural audio codec with pitch controllability
标题:PitchFlower:具有音调可控性的基于流的神经音频编解码器
链接:https://arxiv.org/abs/2510.25566

作者:Diego Torres, Axel Roebel, Nicolas Obin
备注:5 pages, 5 figures
摘要:我们提出了PitchFlower,一个基于流的神经音频编解码器,具有显式的音高可控性。我们的方法通过一个简单的扰动来实现解纠缠:在训练过程中,F0轮廓被平坦化并随机移动,而真实的F0作为条件反射提供。矢量量化瓶颈阻碍了音高恢复,基于流的解码器生成高质量音频。实验表明,PitchFlower在更高的音频质量下实现了比WORLD更精确的音高控制,并且在可控性方面优于SiFiGAN,同时保持相当的质量。除了音高,这个框架提供了一个简单的和可扩展的路径,以解开其他语音属性。
摘要:We present PitchFlower, a flow-based neural audio codec with explicit pitch controllability. Our approach enforces disentanglement through a simple perturbation: during training, F0 contours are flattened and randomly shifted, while the true F0 is provided as conditioning. A vector-quantization bottleneck prevents pitch recovery, and a flow-based decoder generates high quality audio. Experiments show that PitchFlower achieves more accurate pitch control than WORLD at much higher audio quality, and outperforms SiFiGAN in controllability while maintaining comparable quality. Beyond pitch, this framework provides a simple and extensible path toward disentangling other speech attributes.


【3】Separating peripheral and higher-level effects on speech intelligibility using a hearing loss simulator and an objective intelligibility measure
标题:使用听力损失模拟器和客观可懂度测量分离对语音可懂度的外围和高层影响
链接:https://arxiv.org/abs/2510.25235

作者:Toshio Irino, Ayako Yamamoto, Fuki Miyazaki
备注:This is a manuscript that was submitted to Trends in Hearing on October 29, 2025
摘要:本文提出了一种新方法,用于分离外周听力损失(HL)和更高级别过程对语音清晰度(SI)的影响。在以前的研究中,我们进行了SI实验与14个老年人(OA)的听众,使用语音噪声的声音,无论是处理与一个理想的比率掩模(ESTA)增强技术或未经处理。目前的研究涉及一个SI实验与15个年轻的,听力正常(YNH)的听众。该实验使用了用WHIS模拟器处理的模拟HL声音,其反映了来自先前研究的特定OA的听力水平。结果表明,目标OA的SI评分高于平均YNH评分。这意味着目标OA的更高级别的过程可能比平均YNH更有效。为了了解其他OAs的特点,我们使用GESI客观可懂度测量来预测SI。首先,我们证实,GESI可以相当准确地预测的YNH和OA听众的SI分数。接下来,我们使用YNH实验中估计的参数预测了14名OA听众的SI分数。结果表明,一些OA的SI得分高于平均YNH,而一个OA的得分较低。这些差异可能反映了更高层次的加工效率的差异。这些结果表明,WHIS和GESI可以促进YNH和OA听众之间的对比实验,无论听力水平。这将使我们能够单独研究OA听众中更高级别过程的影响。
摘要:This paper presents a new method for separating the effects of peripheral hearing loss (HL) and higher-level processes on speech intelligibility (SI). In a previous study, we conducted an SI experiment with 14 older adult (OA) listeners, using speech-in-noise sounds that were either processed with an ideal ratio mask (IRM) enhancement technique or left unprocessed. The current study involved an SI experiment with 15 young, normal-hearing (YNH) listeners. This experiment used simulated HL sounds processed with the WHIS simulator that reflected the hearing level of a specific OA from the previous study. The results showed that the target OA's SI scores were higher than the average YNH scores. This implies that the target OA's higher-level processes may be more effective than those of the average YNH. To understand the characteristics of other OAs, we used the GESI objective intelligibility measure to predict SI. First, we confirmed that GESI could fairly accurately predict the SI scores for both the YNH and OA listeners. Next, we predicted the SI scores of the 14 OA listeners using the parameters estimated in the YNH experiment. The results showed that some OAs had higher SI scores than the average YNH, while one OA had lower scores. These differences in SI scores may reflect variations in the efficiency of higher-level processes.These results imply that WHIS and GESI could facilitate contrastive experiments between YNH and OA listeners, regardless of hearing level. This would allow us to study the effects of higher-level processes in OA listeners individually.


【4】Retaining Mixture Representations for Domain Generalized Anomalous Sound Detection
标题:保留混合表示以实现域广义异常声音检测
链接:https://arxiv.org/abs/2510.25182

作者:Phurich Saengthong, Tomoya Nishida, Kota Dohi, Natsuo Yamashita, Yohei Kawaguchi
备注:Submitted to ICASSP 2026
摘要:野外异常声音检测(ASD)需要对分布变化(例如机器和噪声类型的不可见低SNR输入混合)具有鲁棒性。最先进的系统从自适应音频编码器中提取嵌入,并通过最近邻搜索检测异常,但对嘈杂的机器声音进行微调通常会起到去噪的作用,在不匹配的混合或不一致的标记下抑制噪声并减少泛化。具有冻结自监督学习(SSL)编码器的免训练系统避免了这个问题,并显示出强大的第一次泛化能力,但当混合嵌入偏离清洁源嵌入时,它们的性能会下降。我们建议使用保留不去噪策略来改进SSL主干,该策略可以更好地保留来自混合声源的信息。该方法将多标签音频标记损失与混合物对齐损失相结合,该混合物对齐损失将学生混合嵌入与干净和噪声输入的凸教师嵌入对齐。对固定、非固定和不匹配噪声子集的控制实验表明,在分布变化下具有更好的鲁棒性,缩小了与Oracle混合表示的差距。
摘要:Anomalous sound detection (ASD) in the wild requires robustness to distribution shifts such as unseen low-SNR input mixtures of machine and noise types. State-of-the-art systems extract embeddings from an adapted audio encoder and detect anomalies via nearest-neighbor search, but fine tuning on noisy machine sounds often acts like a denoising objective, suppressing noise and reducing generalization under mismatched mixtures or inconsistent labeling. Training-free systems with frozen self-supervised learning (SSL) encoders avoid this issue and show strong first-shot generalization, yet their performance drops when mixture embeddings deviate from clean-source embeddings. We propose to improve SSL backbones with a retain-not-denoise strategy that better preserves information from mixed sound sources. The approach combines a multi-label audio tagging loss with a mixture alignment loss that aligns student mixture embeddings to convex teacher embeddings of clean and noise inputs. Controlled experiments on stationary, non-stationary, and mismatched noise subsets demonstrate improved robustness under distribution shifts, narrowing the gap toward oracle mixture representations.


【5】EasyEyes: Online hearing research using speakers calibrated by phones
标题:EasyEyes:使用手机校准的扬声器进行在线听力研究
链接:https://arxiv.org/abs/2510.25048

作者:Ivan Vican, Hugo De Moraes, Chongjun Liao, Nathnael H. Tsegaye, William O'Gara, Jasper Inamoto, Denis G. Pelli
摘要:听力研究需要校准的声源,传统上作为实验室设备。在线研究更快,更具包容性,但大多数参与者缺乏校准设备,他们的声源未经校准且多样化。本文介绍了开源网站EasyEyes.app如何在线校准扬声器。智能手机麦克风配置文件库允许EasyEyes使用参与者的手机在三分钟内校准他们的计算机扬声器。参与者选择他们的手机型号,并通过屏幕大小进行验证。校准采用Novak等人的非同步最大长度序列(MLS)算法。计算机的扬声器通过将其输入与其脉冲响应的倒数进行卷积来校正。研究人员可以通过使用测量麦克风校准手机来为开放获取库做出贡献。在库中,每个配置文件都链接回用于生成它的配置文件,即测量麦克风的制造商配置文件。校正精度使得通过校正的扬声器播放平坦频谱MLS产生几乎平坦的频谱,标准偏差小于3 dB。一项调查显示,来自主要品牌的94款手机型号将支持美国(87%)和英国(80%)的大多数参与者。这种方法有助于高效和包容性的在线听力研究。
摘要:Hearing research requires a calibrated sound source, traditionally as lab equipment. Online research is quicker and more inclusive, but most participants lack calibration equipment and their sound sources are uncalibrated and diverse. This article explains how the open-source EasyEyes.app calibrates loudspeakers online. A library of smartphone-microphone profiles allows EasyEyes to use the participant's phone to calibrate their computer's loudspeaker in three minutes. Participants select their phone model, which is verified by screen size. Calibration employs the Novak et al. nonsynchronous maximum-length-sequence (MLS) algorithm. The computer's loudspeaker is corrected by convolving its input with the inverse of its impulse response. Researchers can contribute to the open-access library by calibrating phones with a measurement microphone. In the library, each profile is linked back to the profile used to produce it, back to the manufacturer profile of a measurement microphone. Correction accuracy is such that playing the flat-spectrum MLS through the corrected loudspeaker produces a nearly flat spectrum, with standard deviation less than 3 dB. A survey shows that a library of 94 phone models from major brands will support most participants in the USA (87%) and UK (80%). This method facilitates efficient and inclusive online hearing research.


【6】Controlling Contrastive Self-Supervised Learning with Knowledge-Driven Multiple Hypothesis: Application to Beat Tracking
标题:用知识驱动的多重假设控制对比自我监督学习:应用于节拍跟踪
链接:https://arxiv.org/abs/2510.25560

作者:Antonin Gagnere, Slim Essid, Geoffroy Peeters
摘要:数据和问题约束中的模糊性可能会导致机器学习任务的多样化,同样合理的结果。例如,在节拍和强拍跟踪中,不同的听众可能会采用各种节奏解释,其中没有一个必然是不正确的。为了解决这个问题,我们提出了一种对比的自我监督预训练方法,该方法利用了关于数据中可能的阳性样本的多个假设。我们的模型经过训练,学习与不同假设兼容的表示,这些假设是用基于知识的评分函数选择的,以保留最合理的假设。当对标记数据进行微调时,我们的模型在标准基准测试中的表现优于现有方法,特别是在音乐表示学习中将领域知识与多假设选择相结合的优势。
摘要:Ambiguities in data and problem constraints can lead to diverse, equally plausible outcomes for a machine learning task. In beat and downbeat tracking, for instance, different listeners may adopt various rhythmic interpretations, none of which would necessarily be incorrect. To address this, we propose a contrastive self-supervised pre-training approach that leverages multiple hypotheses about possible positive samples in the data. Our model is trained to learn representations compatible with different such hypotheses, which are selected with a knowledge-based scoring function to retain the most plausible ones. When fine-tuned on labeled data, our model outperforms existing methods on standard benchmarks, showcasing the advantages of integrating domain knowledge with multi-hypothesis selection in music representation learning in particular.


【7】SFMS-ALR: Script-First Multilingual Speech Synthesis with Adaptive Locale Resolution
标题:SFMS-ALR:具有自适应区域分辨率的脚本优先多语言语音合成
链接:https://arxiv.org/abs/2510.25178

作者:Dharma Teja Donepudi
备注:10 pages, 2 figures, 1 table. Demonstration prototype available at this https URL
摘要:句内多语言语音合成(代码转换TTS)仍然是一个主要的挑战,由于突然的语言转换,不同的脚本,和不匹配的韵律语言之间。传统的TTS系统通常是单语的,并且在混合语言环境中不能产生自然的、可理解的语音。我们介绍脚本优先多语言合成与自适应区域分辨率(SFMS-ALR),一个引擎不可知的框架流畅,实时代码切换语音生成。SFMS-ALR通过Unicode脚本对输入文本进行分段,应用自适应语言识别来确定每个分段的语言和区域,并使用情感感知调整来规范韵律以保持跨语言的表达连续性。该算法生成一个统一的SSML表示与适当的“lang”或“voice”跨度和合成的话语在一个单一的TTS请求。与端到端的多语言模型不同,SFMS-ALR不需要再培训,并且可以与Google、Apple、Amazon和其他提供商的现有语音无缝集成。与数据驱动管道(如Unicom和Mask LID)的比较分析证明了SFMS-ALR的灵活性,可解释性和可立即部署性。该框架建立了一个模块化的高质量,独立于引擎的多语言TTS的基线,并概述了可理解性,自然度和用户偏好的评估策略。
摘要:Intra-sentence multilingual speech synthesis (code-switching TTS) remains a major challenge due to abrupt language shifts, varied scripts, and mismatched prosody between languages. Conventional TTS systems are typically monolingual and fail to produce natural, intelligible speech in mixed-language contexts. We introduce Script-First Multilingual Synthesis with Adaptive Locale Resolution (SFMS-ALR), an engine-agnostic framework for fluent, real-time code-switched speech generation. SFMS-ALR segments input text by Unicode script, applies adaptive language identification to determine each segment's language and locale, and normalizes prosody using sentiment-aware adjustments to preserve expressive continuity across languages. The algorithm generates a unified SSML representation with appropriate "lang" or "voice" spans and synthesizes the utterance in a single TTS request. Unlike end-to-end multilingual models, SFMS-ALR requires no retraining and integrates seamlessly with existing voices from Google, Apple, Amazon, and other providers. Comparative analysis with data-driven pipelines such as Unicom and Mask LID demonstrates SFMS-ALR's flexibility, interpretability, and immediate deployability. The framework establishes a modular baseline for high-quality, engine-independent multilingual TTS and outlines evaluation strategies for intelligibility, naturalness, and user preference.


【8】Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
标题:基于带有部分标签的声音事件半监督训练的声音场景和声音事件联合分析
链接:https://arxiv.org/abs/2510.25075

作者:Keisuke Imoto
备注:Accepted to APSIPA Transactions on Signal and Information Processing
摘要:注释声音事件的时间边界是劳动密集型的,限制了音频检测中强监督学习的可扩展性。为了降低注释成本,仅使用剪辑级标签的弱监督学习已被广泛采用。作为替代方案,部分标签学习提供了一种具有成本效益的方法,其中提供了一组可能的标签,而不是精确的弱注释。然而,用于音频分析的部分标签学习在很大程度上仍未被探索。出于观察,声学场景提供上下文信息构建一组可能的声音事件,我们利用声学场景信息来构建部分标签的声音事件。在此基础上,本文提出了一个多任务学习框架,共同执行声学场景分类和声音事件检测与部分标签的声音事件。在降低标注成本的同时,弱监督和部分标签学习往往会由于缺乏精确的事件集及其时间标注而降低检测性能。为了更好地平衡注释成本和检测性能,我们还探索了一个利用强标签和部分标签的半监督框架。此外,为了细化部分标签,实现更好的模型训练,我们提出了一种基于自蒸馏的标签细化方法,所提出的方法与部分标签。
摘要:Annotating time boundaries of sound events is labor-intensive, limiting the scalability of strongly supervised learning in audio detection. To reduce annotation costs, weakly-supervised learning with only clip-level labels has been widely adopted. As an alternative, partial label learning offers a cost-effective approach, where a set of possible labels is provided instead of exact weak annotations. However, partial label learning for audio analysis remains largely unexplored. Motivated by the observation that acoustic scenes provide contextual information for constructing a set of possible sound events, we utilize acoustic scene information to construct partial labels of sound events. On the basis of this idea, in this paper, we propose a multitask learning framework that jointly performs acoustic scene classification and sound event detection with partial labels of sound events. While reducing annotation costs, weakly-supervised and partial label learning often suffer from decreased detection performance due to lacking the precise event set and their temporal annotations. To better balance between annotation cost and detection performance, we also explore a semi-supervised framework that leverages both strong and partial labels. Moreover, to refine partial labels and achieve better model training, we propose a label refinement method based on self-distillation for the proposed approach with partial labels.


【9】Evaluating Emotion Recognition in Spoken Language Models on Emotionally Incongruent Speech
标题:评估情感不一致语音的口语模型中的情感识别
链接:https://arxiv.org/abs/2510.25054

作者:Pedro Corrêa, João Lima, Victor Moreno, Paula Dornhofer Paro Costa
备注:This work has been submitted to the IEEE for possible publication
摘要:口语处理的进步推动了口语模型(SLM)的发展,旨在通过联合学习文本和音频表示来实现广泛的音频理解。虽然已经取得了可喜的成果,有越来越多的讨论,这些模型的泛化能力,以及在何种程度上,他们真正整合音频和文本形式在其内部表示。在这项工作中,我们评估四个SLM的语音情感识别的任务,使用情绪不一致的语音样本的数据集,在这种情况下,口头话语的语义内容传达一种情感,而语音表达传达另一种。我们的研究结果表明,SLM主要依赖于文本语义,而不是语音情感来执行任务,这表明文本相关的表示在很大程度上占主导地位的声学表示。我们向社区发布了代码和非一致性合成语音数据集(EMIS)。
摘要:Advancements in spoken language processing have driven the development of spoken language models (SLMs), designed to achieve universal audio understanding by jointly learning text and audio representations for a wide range of tasks. Although promising results have been achieved, there is growing discussion regarding these models' generalization capabilities and the extent to which they truly integrate audio and text modalities in their internal representations. In this work, we evaluate four SLMs on the task of speech emotion recognition using a dataset of emotionally incongruent speech samples, a condition under which the semantic content of the spoken utterance conveys one emotion while speech expressiveness conveys another. Our results indicate that SLMs rely predominantly on textual semantics rather than speech emotion to perform the task, indicating that text-related representations largely dominate over acoustic representations. We release both the code and the Emotionally Incongruent Synthetic Speech dataset (EMIS) to the community.


机器翻译由腾讯交互翻译提供,仅供参考