微信公众号:arXiv_Daily
cs.SD语音
【1】Dynamic Fusion Multimodal Network for SpeechWellness Detection
标题:用于语音健康检测的动态融合多模式网络
链接:https://arxiv.org/abs/2508.18057
备注:6 pages, 5figures
摘要:自杀是青少年死亡的主要原因之一。以前的自杀风险预测研究主要集中在孤立的文本或声学信息,多模态信号的整合,如语音和文本,提供了一个更全面的了解个人的精神状态。出于这一动机,并在第一次SpeechWellness检测挑战的背景下,我们探索了一个轻量级的多分支多模式系统的基础上的动态融合机制的语音健康检测。为了解决依赖于时域波形进行声学分析的现有方法的局限性,我们的系统结合了时域和时频(TF)域声学特征以及语义表示。此外,我们引入了一个动态融合块,以自适应地整合来自不同模态的信息。具体而言,它在融合过程中将可学习的权重应用于每个模态,使模型能够调整每个模态的贡献。为了提高计算效率,我们设计了一个轻量级的结构,通过简化原始基线模型。实验结果表明,该系统具有优越的性能相比,挑战基线,实现了78%的减少模型参数和5%的精度提高。
摘要:Suicide is one of the leading causes of death among adolescents. Previous suicide risk prediction studies have primarily focused on either textual or acoustic information in isolation, the integration of multimodal signals, such as speech and text, offers a more comprehensive understanding of an individual's mental state. Motivated by this, and in the context of the 1st SpeechWellness detection challenge, we explore a lightweight multi-branch multimodal system based on a dynamic fusion mechanism for speechwellness detection. To address the limitation of prior approaches that rely on time-domain waveforms for acoustic analysis, our system incorporates both time-domain and time-frequency (TF) domain acoustic features, as well as semantic representations. In addition, we introduce a dynamic fusion block to adaptively integrate information from different modalities. Specifically, it applies learnable weights to each modality during the fusion process, enabling the model to adjust the contribution of each modality. To enhance computational efficiency, we design a lightweight structure by simplifying the original baseline model. Experimental results demonstrate that the proposed system exhibits superior performance compared to the challenge baseline, achieving a 78% reduction in model parameters and a 5% improvement in accuracy.
【2】Enhancing Speech Emotion Recognition with Multi-Task Learning and Dynamic Feature Fusion
标题:利用多任务学习和动态特征融合增强语音情感识别
链接:https://arxiv.org/abs/2508.17878
备注:accepted by interspeech2025
摘要:本研究探讨微调自我监督学习(SSL)模型,使用多任务学习(MTL),以提高 语音情感识别。该框架同时处理四个相关的任务:情感识别,性别 语音识别、说话人确认和自动语音识别。引入了一个创新的共同注意力模块,动态地捕捉来自 主要情感分类任务和辅助任务,使上下文感知融合。此外,我们引入了样本加权焦点对比损失函数(SWFC),通过调整困难样本和少数样本的样本权重来解决分类不平衡和语义混淆问题。该方法 在分类情绪识别任务中验证 自然条件下的语音情感识别挑战,表现出显着的性能改善。
摘要:This study investigates fine-tuning self-supervised learn ing (SSL) models using multi-task learning (MTL) to enhance speech emotion recognition (SER). The framework simultane ously handles four related tasks: emotion recognition, gender recognition, speaker verification, and automatic speech recog nition. An innovative co-attention module is introduced to dy namically capture the interactions between features from the primary emotion classification task and auxiliary tasks, en abling context-aware fusion. Moreover, We introduce the Sam ple Weighted Focal Contrastive (SWFC) loss function to ad dress class imbalance and semantic confusion by adjusting sam ple weights for difficult and minority samples. The method is validated on the Categorical Emotion Recognition task of the Speech Emotion Recognition in Naturalistic Conditions Chal lenge, showing significant performance improvements.
【3】Vocoder-Projected Feature Discriminator
标题:声码器投影特征鉴别器
链接:https://arxiv.org/abs/2508.17874
备注:Accepted to Interspeech 2024. Project page: this https URL
摘要:在文本到语音(TTS)和语音转换(VC)中,声学特征(诸如梅尔频谱图)由于其紧凑性和易于学习而通常用作合成或转换目标。然而,由于最终目标是生成高质量的波形,因此采用声码器将这些特征转换为波形并在时域中应用对抗训练是合理的。尽管如此,对波形进行上采样会引入大量时间和内存开销。为了解决这个问题,我们提出了一个声码器投影的特征提取(VPFD),它使用声码器的功能对抗训练。基于扩散的VC蒸馏实验表明,一个预训练和冻结的声码器特征提取器与一个单一的上采样步骤是必要的和足够的,以实现VC性能与波形鉴别器,同时减少训练时间和内存消耗的9.6和11.4倍,分别。
摘要:In text-to-speech (TTS) and voice conversion (VC), acoustic features, such as mel spectrograms, are typically used as synthesis or conversion targets owing to their compactness and ease of learning. However, because the ultimate goal is to generate high-quality waveforms, employing a vocoder to convert these features into waveforms and applying adversarial training in the time domain is reasonable. Nevertheless, upsampling the waveform introduces significant time and memory overheads. To address this issue, we propose a vocoder-projected feature discriminator (VPFD), which uses vocoder features for adversarial training. Experiments on diffusion-based VC distillation demonstrated that a pretrained and frozen vocoder feature extractor with a single upsampling step is necessary and sufficient to achieve a VC performance comparable to that of waveform discriminators while reducing the training time and memory consumption by 9.6 and 11.4 times, respectively.
【4】FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
标题:FasterEqualGrad:更快的基于扩散的一步语音转换,采用对抗扩散转换蒸馏
链接:https://arxiv.org/abs/2508.17868
备注:Accepted to Interspeech 2025. Project page: this https URL
摘要:基于扩散的语音转换(VC)模型(例如,VoiceGrad)可以实现较高的语音质量和说话人相似度,但是由于迭代采样,其转换过程较慢。FastVoiceGrad通过将VoiceGrad提炼为一步扩散模型来克服这一限制。然而,它仍然需要一个计算密集的内容编码器来解开说话者的身份和内容,这会减慢转换速度。因此,我们提出了FasterVoiceGrad,这是一种新的基于扩散的一步VC模型,通过使用对抗扩散转换蒸馏(ADCD)同时蒸馏扩散模型和内容编码器获得,其中蒸馏在转换过程中进行,同时利用对抗和分数蒸馏训练。对单次VC的实验评估表明,与FastVoiceGrad相比,FastVoiceGrad实现了具有竞争力的VC性能,在GPU和CPU上的速度分别提高了6.6-6.9和1.8倍。
摘要:A diffusion-based voice conversion (VC) model (e.g., VoiceGrad) can achieve high speech quality and speaker similarity; however, its conversion process is slow owing to iterative sampling. FastVoiceGrad overcomes this limitation by distilling VoiceGrad into a one-step diffusion model. However, it still requires a computationally intensive content encoder to disentangle the speaker's identity and content, which slows conversion. Therefore, we propose FasterVoiceGrad, a novel one-step diffusion-based VC model obtained by simultaneously distilling a diffusion model and content encoder using adversarial diffusion conversion distillation (ADCD), where distillation is performed in the conversion process while leveraging adversarial and score distillation training. Experimental evaluations of one-shot VC demonstrated that FasterVoiceGrad achieves competitive VC performance compared to FastVoiceGrad, with 6.6-6.9 and 1.8 times faster speed on a GPU and CPU, respectively.
【5】Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs
标题:语音离散令牌还是连续特征?SpeechLLM中口语理解的比较分析
链接:https://arxiv.org/abs/2508.17863
备注:Accepted to EMNLP 2025 Main Conference
摘要:随着语音大语言模型(SpeechLLM)的兴起,出现了两种主要的语音处理方法:离散标记和连续特征。每种方法都在音频相关处理任务中表现出强大的能力。然而,这两种范式之间的性能差距还没有得到彻底的探讨。为了解决这一差距,我们在相同的实验设置下对基于自监督学习(SSL)的离散和连续特征进行了公平的比较。我们使用小型和大型LLM(Qwen1.5-0.5B和Llama3.1-8B)评估了他们在六个口语理解相关任务中的表现。我们进一步进行了深入的分析,包括效率比较,SSL层分析,LLM层分析和鲁棒性比较。我们的研究结果表明,连续特征在各种任务中的表现通常优于离散标记。每种语音处理方法在如何学习和处理语音信息方面都表现出不同的特征和模式。我们希望我们的研究结果将提供有价值的见解,以促进口语理解SpeechLLM。
摘要:With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audio-related processing tasks. However, the performance gap between these two paradigms has not been thoroughly explored. To address this gap, we present a fair comparison of self-supervised learning (SSL)-based discrete and continuous features under the same experimental settings. We evaluate their performance across six spoken language understanding-related tasks using both small and large-scale LLMs (Qwen1.5-0.5B and Llama3.1-8B). We further conduct in-depth analyses, including efficient comparison, SSL layer analysis, LLM layer analysis, and robustness comparison. Our findings reveal that continuous features generally outperform discrete tokens in various tasks. Each speech processing method exhibits distinct characteristics and patterns in how it learns and processes speech information. We hope our results will provide valuable insights to advance spoken language understanding in SpeechLLMs.
【6】ClearMask: Noise-Free and Naturalness-Preserving Protection Against Voice Deepfake Attacks
标题:ClearMass:针对语音Deepfake攻击的无噪音和自然保护
链接:https://arxiv.org/abs/2508.17660
备注:14 Pages, Accepted by AsiaCCS 2025
摘要:语音deepfake攻击,人工模仿人类语音用于恶意目的,已经成为一个严重的威胁。现有的防御通常会将噪声注入到人类语音中,以损害语音合成模型中的语音编码器。然而,这些方法降低了音频质量,并且需要攻击方法的先验知识,从而限制了它们在各种情况下的有效性。此外,实时音频,例如虚拟会议中的讲话和语音消息,仍然受到语音deepfake的威胁。为了克服这些限制,我们提出了ClearMask,这是一种针对语音deepfake攻击的无噪声防御机制。与传统方法不同,ClearMask通过选择性地过滤某些频率来修改音频梅尔频谱图,从而在不注入噪声的情况下引起可转移的语音特征损失。然后,我们应用音频风格转移,以进一步欺骗语音解码器,同时保持感知的音质。最后,引入优化的混响来破坏语音生成模型的输出,而不影响语音的自然度。此外,我们还开发了LiveMask,通过通用频率滤波器和混响发生器实时保护流式语音。我们的实验结果表明,ClearMask和LiveMask有效地防止了语音deepfake攻击欺骗说话人验证模型和人类听众,即使是看不见的语音合成模型和黑盒API服务。此外,ClearMask还展示了针对自适应攻击者的弹性,这些攻击者试图从受保护的语音样本中恢复原始音频信号。
摘要:Voice deepfake attacks, which artificially impersonate human speech for malicious purposes, have emerged as a severe threat. Existing defenses typically inject noise into human speech to compromise voice encoders in speech synthesis models. However, these methods degrade audio quality and require prior knowledge of the attack approaches, limiting their effectiveness in diverse scenarios. Moreover, real-time audios, such as speech in virtual meetings and voice messages, are still exposed to voice deepfake threats. To overcome these limitations, we propose ClearMask, a noise-free defense mechanism against voice deepfake attacks. Unlike traditional approaches, ClearMask modifies the audio mel-spectrogram by selectively filtering certain frequencies, inducing a transferable voice feature loss without injecting noise. We then apply audio style transfer to further deceive voice decoders while preserving perceived sound quality. Finally, optimized reverberation is introduced to disrupt the output of voice generation models without affecting the naturalness of the speech. Additionally, we develop LiveMask to protect streaming speech in real-time through a universal frequency filter and reverberation generator. Our experimental results show that ClearMask and LiveMask effectively prevent voice deepfake attacks from deceiving speaker verification models and human listeners, even for unseen voice synthesis models and black-box API services. Furthermore, ClearMask demonstrates resilience against adaptive attackers who attempt to recover the original audio signal from the protected speech samples.
【7】Improving French Synthetic Speech Quality via SSML Prosody Control
标题:利用SSML韵律控制提高法语合成语音质量
链接:https://arxiv.org/abs/2508.17494
备注:13 pages, 9 figures, 6 tables. Accepted for presentation at ICNLSP 2025 (Odense, Denmark). Code and demo: this https URL. ACM Class: I.2.7; H.5.5
摘要:尽管最近的进展,合成语音往往缺乏表现力,由于有限的韵律控制,在商业文本到语音(TTS)系统。我们介绍了第一个端到端的管道,插入语音合成标记语言(SSML)标签到法语文本控制音高,语速,音量和暂停持续时间。我们采用了一个级联架构,其中包含两个QLoRA微调的Qwen 2.5- 7 B模型:一个预测短语中断位置,另一个对韵律目标进行回归,生成商业TTS兼容的SSML标记。在一个14小时的法语播客语料库上进行评估,我们的方法实现了99.2%F1的休息位置,并将音高,速率和音量的平均绝对误差降低了25-40%,与仅支持大型语言模型(LLM)和BiLSTM基线相比。在涉及18名参与者超过9小时的合成音频的感知评估中,我们的管道生成的SSL增强语音显着提高了自然度,平均意见得分从3.20增加到3.87(p < 0.005)。此外,18名听众中有15名更喜欢我们的增强合成。这些结果表明,在弥合合成和自然的法语语音之间的表达差距的实质性进展。我们的代码可在https://github.com/hi-paris/Prosody-Control-French-TTS上公开获取。
摘要:Despite recent advances, synthetic voices often lack expressiveness due to limited prosody control in commercial text-to-speech (TTS) systems. We introduce the first end-to-end pipeline that inserts Speech Synthesis Markup Language (SSML) tags into French text to control pitch, speaking rate, volume, and pause duration. We employ a cascaded architecture with two QLoRA-fine-tuned Qwen 2.5-7B models: one predicts phrase-break positions and the other performs regression on prosodic targets, generating commercial TTS-compatible SSML markup. Evaluated on a 14-hour French podcast corpus, our method achieves 99.2% F1 for break placement and reduces mean absolute error on pitch, rate, and volume by 25-40% compared with prompting-only large language models (LLMs) and a BiLSTM baseline. In perceptual evaluation involving 18 participants across over 9 hours of synthesized audio, SSML-enhanced speech generated by our pipeline significantly improves naturalness, with the mean opinion score increasing from 3.20 to 3.87 (p < 0.005). Additionally, 15 of 18 listeners preferred our enhanced synthesis. These results demonstrate substantial progress in bridging the expressiveness gap between synthetic and natural French speech. Our code is publicly available at https://github.com/hi-paris/Prosody-Control-French-TTS.
【8】DanceEditor: Towards Iterative Editable Music-driven Dance Generation with Open-Vocabulary Descriptions
标题:DanceEditor:迈向具有开放词汇描述的迭代编辑音乐驱动的舞蹈一代
链接:https://arxiv.org/abs/2508.17342
备注:None
摘要:从音乐信号生成连贯多样的人类舞蹈在虚拟化身动画方面取得了巨大进展。虽然现有的方法支持直接的舞蹈合成,他们没有认识到,使用户能够编辑舞蹈动作是更实际的在现实世界的编舞方案。此外,缺乏包含迭代编辑的高质量舞蹈数据集也限制了解决这一挑战。为了实现这一目标,我们首先构建了DanceRemix,这是一个大规模的多回合可编辑舞蹈数据集,包括超过2530万个舞蹈帧和84500对的提示。此外,我们提出了一个新的框架,迭代和可编辑的舞蹈生成连贯一致的音乐信号,即DanceEditor。考虑到舞蹈动作应该是音乐节奏,并使用户描述的迭代编辑,我们的框架是建立在一个预测,然后编辑范式统一多模态条件。在初始预测阶段,我们的框架通过直接从定制的对齐音乐中建模舞蹈动作来提高生成结果的权威性。此外,在随后的迭代编辑阶段,我们将文本描述作为条件信息,通过专门设计的跨模态编辑模块(CEM)绘制可编辑结果。具体而言,CEM自适应地将初始预测与音乐和文本提示集成为时间运动线索以指导合成序列。因此,结果显示音乐和声,同时保持与文本描述的细粒度语义对齐。大量的实验表明,我们的方法在我们新收集的DanceRemix数据集上的性能优于最先进的模型。代码可在https://lzvsdy.github.io/DanceEditor/上获得。
摘要:Generating coherent and diverse human dances from music signals has gained tremendous progress in animating virtual avatars. While existing methods support direct dance synthesis, they fail to recognize that enabling users to edit dance movements is far more practical in real-world choreography scenarios. Moreover, the lack of high-quality dance datasets incorporating iterative editing also limits addressing this challenge. To achieve this goal, we first construct DanceRemix, a large-scale multi-turn editable dance dataset comprising the prompt featuring over 25.3M dance frames and 84.5K pairs. In addition, we propose a novel framework for iterative and editable dance generation coherently aligned with given music signals, namely DanceEditor. Considering the dance motion should be both musical rhythmic and enable iterative editing by user descriptions, our framework is built upon a prediction-then-editing paradigm unifying multi-modal conditions. At the initial prediction stage, our framework improves the authority of generated results by directly modeling dance movements from tailored, aligned music. Moreover, at the subsequent iterative editing stages, we incorporate text descriptions as conditioning information to draw the editable results through a specifically designed Cross-modality Editing Module (CEM). Specifically, CEM adaptively integrates the initial prediction with music and text prompts as temporal motion cues to guide the synthesized sequences. Thereby, the results display music harmonics while preserving fine-grained semantic alignment with text descriptions. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected DanceRemix dataset. Code is available at https://lzvsdy.github.io/DanceEditor/.
【9】Modality-Specific Speech Enhancement and Noise-Adaptive Fusion for Acoustic and Body-Conduction Microphone Framework
标题:声学和体导麦克风框架中的特定模态语音增强和噪声自适应融合
链接:https://arxiv.org/abs/2508.17336
备注:None
摘要:体导麦克风信号(BMS)绕过空气传播的声音,提供了强大的抗噪能力。然而,需要补充模态来补偿高频信息的固有损失。在这项研究中,我们提出了一种新的多模态框架,结合BMS和声学麦克风信号(AMS),以实现噪声抑制和高频重建。与传统的多模态方法,简单地合并功能,我们的方法采用了两个专门的网络:基于映射的模型,以增强BMS和基于掩蔽的模型去噪AMS。这些网络通过动态融合机制进行集成,该机制适应当地的噪声条件,确保最佳利用每种模态的优势。我们使用客观的语音质量指标对TAPS数据集进行了评估,并使用DNS 2023噪声剪辑进行了增强。结果清楚地表明,我们的方法优于单模态的解决方案,在广泛的噪声环境。
摘要:Body\-conduction microphone signals (BMS) bypass airborne sound, providing strong noise resistance. However, a complementary modality is required to compensate for the inherent loss of high\-frequency information. In this study, we propose a novel multi\-modal framework that combines BMS and acoustic microphone signals (AMS) to achieve both noise suppression and high\-frequency reconstruction. Unlike conventional multi\-modal approaches that simply merge features, our method employs two specialized networks\: a mapping-based model to enhance BMS and a masking-based model to denoise AMS. These networks are integrated through a dynamic fusion mechanism that adapts to local noise conditions, ensuring the optimal use of each modality's strengths. We performed evaluations on the TAPS dataset, augmented with DNS\-2023 noise clips, using objective speech quality metrics. The results clearly demonstrate that our approach outperforms single\-modal solutions in a wide range of noisy environments.
【10】ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection
标题:ERF-BA-TFD+:用于视听深度造假检测的多模式模型
链接:https://arxiv.org/abs/2508.17282
摘要:Deepfake检测是识别被操纵的多媒体内容的关键任务。在现实世界中,deepfake内容可以在多种形式中表现出来,包括音频和视频。为了应对这一挑战,我们提出了ERF-BA-TFD+,这是一种新型的多模态深度伪造检测模型,它结合了增强的感受野(ERF)和视听融合。我们的模型同时处理音频和视频特征,利用它们的互补信息来提高检测的准确性和鲁棒性。ERF-BA-TFD+的关键创新在于它能够对视听输入中的长期依赖关系进行建模,从而更好地捕捉真实内容和虚假内容之间的细微差异。在我们的实验中,我们在DDL-AV数据集上评估ERF-BA-TFD+,该数据集由分段和全长视频剪辑组成。与以前的基准测试主要集中在孤立的部分不同,DDL-AV数据集允许我们在更全面和真实的环境中评估模型的性能。我们的方法在这个数据集上实现了最先进的结果,在准确性和处理速度方面都优于现有技术。ERF-BA-TFD+模型在“深度伪造检测、定位和可解释性研讨会”第二部分:视听检测和定位(DDL-AV)中展示了其有效性,并在本次比赛中获得第一名。
摘要:Deepfake detection is a critical task in identifying manipulated multimedia content. In real-world scenarios, deepfake content can manifest across multiple modalities, including audio and video. To address this challenge, we present ERF-BA-TFD+, a novel multimodal deepfake detection model that combines enhanced receptive field (ERF) and audio-visual fusion. Our model processes both audio and video features simultaneously, leveraging their complementary information to improve detection accuracy and robustness. The key innovation of ERF-BA-TFD+ lies in its ability to model long-range dependencies within the audio-visual input, allowing it to better capture subtle discrepancies between real and fake content. In our experiments, we evaluate ERF-BA-TFD+ on the DDL-AV dataset, which consists of both segmented and full-length video clips. Unlike previous benchmarks, which focused primarily on isolated segments, the DDL-AV dataset allows us to assess the model's performance in a more comprehensive and realistic setting. Our method achieves state-of-the-art results on this dataset, outperforming existing techniques in terms of both accuracy and processing speed. The ERF-BA-TFD+ model demonstrated its effectiveness in the "Workshop on Deepfake Detection, Localization, and Interpretability," Track 2: Audio-Visual Detection and Localization (DDL-AV), and won first place in this competition.
【11】Multi-Metric Preference Alignment for Generative Speech Restoration
标题:生成式语音恢复中的多度量偏好对齐
链接:https://arxiv.org/abs/2508.17229
备注:16 pages, 10 figures. demopage: this https URL
摘要:最近的生成模型大大提高了语音恢复任务,但它们的训练目标往往与人类的感知偏好不一致,导致质量不佳。虽然训练后对齐在其他生成领域(如文本和图像生成)中已被证明是有效的,但其在生成语音恢复中的应用在很大程度上仍未得到充分探索。这项工作研究了将基于偏好的后训练应用于这项任务的挑战,重点是如何定义一个强大的偏好信号和策划高质量的数据,以避免奖励黑客。为了解决这些挑战,我们提出了一种多指标偏好对齐策略。我们构建了一个新的数据集,GenSR-Pref,包括80 K的偏好对,其中每个选择的样本是一致赞成的一套互补的指标,涵盖感知质量,信号保真度,内容一致性,和音色保存。这种原则性的方法确保了整体的偏好信号。将直接偏好优化(DPO)应用于我们的数据集,我们在三种不同的生成范式中观察到一致且显着的性能提升:自回归模型(AR)、掩蔽生成模型(MGM)和流匹配模型(FM)在各种恢复基准上,在客观和主观评估中。消融研究证实了我们的多指标策略在减轻奖励黑客攻击方面优于单指标方法。此外,我们证明了我们的对齐模型可以作为强大的“数据注释器”,生成高质量的伪标签,作为传统判别模型在数据稀缺的情况下(如歌声恢复)的监督信号。演示页面:https://gensr-pref.github.io
摘要:Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven effective in other generative domains like text and image generation, its application to generative speech restoration remains largely under-explored. This work investigates the challenges of applying preference-based post-training to this task, focusing on how to define a robust preference signal and curate high-quality data to avoid reward hacking. To address these challenges, we propose a multi-metric preference alignment strategy. We construct a new dataset, GenSR-Pref, comprising 80K preference pairs, where each chosen sample is unanimously favored by a complementary suite of metrics covering perceptual quality, signal fidelity, content consistency, and timbre preservation. This principled approach ensures a holistic preference signal. Applying Direct Preference Optimization (DPO) with our dataset, we observe consistent and significant performance gains across three diverse generative paradigms: autoregressive models (AR), masked generative models (MGM), and flow-matching models (FM) on various restoration benchmarks, in both objective and subjective evaluations. Ablation studies confirm the superiority of our multi-metric strategy over single-metric approaches in mitigating reward hacking. Furthermore, we demonstrate that our aligned models can serve as powerful ''data annotators'', generating high-quality pseudo-labels to serve as a supervision signal for traditional discriminative models in data-scarce scenarios like singing voice restoration. Demo Page:https://gensr-pref.github.io
【12】Multi-scale Scanning Network for Machine Anomalous Sound Detection
标题:机器异常声音检测的多尺度扫描网络
链接:https://arxiv.org/abs/2508.17194
备注:Accepted by ICONIP 2025
摘要:机器声音在频域和时域中表现出一致和重复的模式,这些模式在不同机器类型的尺度上变化很大。例如,旋转机器通常在短时间间隔内显示出周期性特征,而往复式机器则表现出跨越时域的更广泛的模式。虽然之前的研究已经利用这些模式来改进异常声音检测(ASD),但模式在不同尺度上的变化仍然没有得到充分的探索。为了解决这个问题,我们引入了一个多尺度扫描网络(MSN),旨在捕捉在多个尺度的模式。MSN采用不同大小的内核框来扫描音频频谱图,并集成了一个具有共享权重的轻量级卷积网络,以实现高效和可扩展的特征表示。对DCASE 2020和DCASE 2023任务2数据集的实验评估表明,MSN实现了最先进的性能,突出了其在推进ASD系统方面的有效性。
摘要:Machine sounds exhibit consistent and repetitive patterns in both the frequency and time domains, which vary significantly across scales for different machine types. For instance, rotating machines often show periodic features in short time intervals, while reciprocating machines exhibit broader patterns spanning the time domain. While prior studies have leveraged these patterns to improve Anomalous Sound Detection (ASD), the variation of patterns across scales remains insufficiently explored. To address this gap, we introduce a Multi-scale Scanning Network (MSN) designed to capture patterns at multiple scales. MSN employs kernel boxes of varying sizes to scan audio spectrograms and integrates a lightweight convolutional network with shared weights for efficient and scalable feature representation. Experimental evaluations on the DCASE 2020 and DCASE 2023 Task 2 datasets demonstrate that MSN achieves state-of-the-art performance, highlighting its effectiveness in advancing ASD systems.
【13】Geolocation-Aware Robust Spoken Language Identification
标题:具有地理意识的稳健口语识别
链接:https://arxiv.org/abs/2508.17148
备注:Accepted to IEEE ASRU 2025. \c{opyright} 2025 IEEE. Personal use permitted. Permission from IEEE required for all other uses including reprinting/republishing, advertising, resale, redistribution, reuse, or creating collective works
摘要:虽然自监督学习(SSL)显著改善了口语识别(LID),但现有模型通常难以将同一语言的方言和口音一致地分类为统一的类别。为了应对这一挑战,我们提出了地理位置感知LID,一种新的方法,将语言级的地理位置信息融入到基于SSL的LID模型。具体来说,我们引入地理位置预测作为辅助任务,并将预测向量注入中间表示作为条件信号。这种明确的条件作用鼓励模型学习方言和重音变化的更统一的表示。在六个多语言数据集上的实验表明,我们的方法提高了对语言内变化和不可见域的鲁棒性,在FLEURS上实现了新的最先进的准确率(97.7%),在ML-SUPERB 2.0方言集上实现了9.7%的相对改进。
摘要:While Self-supervised Learning (SSL) has significantly improved Spoken Language Identification (LID), existing models often struggle to consistently classify dialects and accents of the same language as a unified class. To address this challenge, we propose geolocation-aware LID, a novel approach that incorporates language-level geolocation information into the SSL-based LID model. Specifically, we introduce geolocation prediction as an auxiliary task and inject the predicted vectors into intermediate representations as conditioning signals. This explicit conditioning encourages the model to learn more unified representations for dialectal and accented variations. Experiments across six multilingual datasets demonstrate that our approach improves robustness to intra-language variations and unseen domains, achieving new state-of-the-art accuracy on FLEURS (97.7%) and 9.7% relative improvement on ML-SUPERB 2.0 dialect set.
【14】SyncGuard: Robust Audio Watermarking Capable of Countering Desynchronization Attacks
标题:SyncGuard:能够对抗去序列化攻击的稳健音频水印
链接:https://arxiv.org/abs/2508.17121
摘要:音频水印技术在版权保护和来源追踪等方面有着广泛的应用。然而,由于音频信号的固有特性,水印的定位和抵抗去水印攻击仍然是重大的挑战。在本文中,我们提出了一个基于学习的计划名为SyncGuard来解决这些挑战。具体来说,我们设计了一个逐帧的广播嵌入策略嵌入水印在任意长度的音频,增强时间无关性,并消除了本地化水印提取过程中的需要。为了进一步增强鲁棒性,我们引入了精心设计的失真层。此外,我们采用扩张残差块结合扩张门控块,以有效地捕捉多分辨率的时频特征。大量的实验结果表明,SyncGuard有效地处理可变长度的音频段,在对各种攻击的鲁棒性方面优于最先进的方法,并提供卓越的听觉质量。
摘要:Audio watermarking has been widely applied in copyright protection and source tracing. However, due to the inherent characteristics of audio signals, watermark localization and resistance to desynchronization attacks remain significant challenges. In this paper, we propose a learning-based scheme named SyncGuard to address these challenges. Specifically, we design a frame-wise broadcast embedding strategy to embed the watermark in arbitrary-length audio, enhancing time-independence and eliminating the need for localization during watermark extraction. To further enhance robustness, we introduce a meticulously designed distortion layer. Additionally, we employ dilated residual blocks in conjunction with dilated gated blocks to effectively capture multi-resolution time-frequency features. Extensive experimental results show that SyncGuard efficiently handles variable-length audio segments, outperforms state-of-the-art methods in robustness against various attacks, and delivers superior auditory quality.
【15】RephraseTTS: Dynamic Length Text based Speech Insertion with Speaker Style Transfer
标题:RephraseTTS:基于动态长度文本的说话人风格转换语音插入
链接:https://arxiv.org/abs/2508.17031
摘要:我们提出了一种用于文本条件语音插入任务的方法,即,在相应的完整文本抄本的条件下,在输入语音样本中插入语音样本。该任务的一个示例用例将是在对相应的文本抄本进行校正时更新语音音频。所提出的方法遵循基于变换器的非自回归方法,该方法允许基于可用部分输入的文本转录和节奏在推断期间动态确定的可变长度的语音插入。它能够保持说话者的话音特性、韵律和可用语音输入的其它频谱特性。在LibriTTS上的实验和用户研究结果表明,该方法优于基于现有自适应文本到语音方法的基线。我们还提供了许多定性的结果,以欣赏所提出的方法的输出质量。
摘要:We propose a method for the task of text-conditioned speech insertion, i.e. inserting a speech sample in an input speech sample, conditioned on the corresponding complete text transcript. An example use case of the task would be to update the speech audio when corrections are done on the corresponding text transcript. The proposed method follows a transformer-based non-autoregressive approach that allows speech insertions of variable lengths, which are dynamically determined during inference, based on the text transcript and tempo of the available partial input. It is capable of maintaining the speaker's voice characteristics, prosody and other spectral properties of the available speech input. Results from our experiments and user study on LibriTTS show that our method outperforms baselines based on an existing adaptive text to speech method. We also provide numerous qualitative results to appreciate the quality of the output from the proposed method.
【16】MDD: A Dataset for Text-and-Music Conditioned Duet Dance Generation
标题:DDD:文本和音乐条件二重唱舞蹈生成的数据集
链接:https://arxiv.org/abs/2508.16911
备注:Accepted at ICCV 2025. Project page: this https URL
摘要:我们介绍多模态DuetDance(MDD),一个多样化的多模态基准数据集,专为文本控制和音乐调节的3D二重唱舞蹈动作生成。我们的数据集包括620分钟的高质量动作捕捉数据,由专业舞者表演,与音乐同步,并详细描述了超过10K的细粒度自然语言描述。这些注释捕获了丰富的运动词汇,详细描述了空间关系、身体运动和节奏,使MDD成为第一个无缝集成人体运动、音乐和文本以生成二重唱舞蹈的数据集。我们介绍了我们的数据集支持的两个新任务:(1)文本到二重唱,其中给定音乐和文本提示,生成领导者和追随者的舞蹈动作(2)文本到舞蹈伴奏,其中给定音乐,文本提示和领导者的动作,追随者的动作以内聚的,文本对齐的方式生成。我们对这两项任务进行了基线评估,以支持未来的研究。
摘要:We introduce Multimodal DuetDance (MDD), a diverse multimodal benchmark dataset designed for text-controlled and music-conditioned 3D duet dance motion generation. Our dataset comprises 620 minutes of high-quality motion capture data performed by professional dancers, synchronized with music, and detailed with over 10K fine-grained natural language descriptions. The annotations capture a rich movement vocabulary, detailing spatial relationships, body movements, and rhythm, making MDD the first dataset to seamlessly integrate human motions, music, and text for duet dance generation. We introduce two novel tasks supported by our dataset: (1) Text-to-Duet, where given music and a textual prompt, both the leader and follower dance motion are generated (2) Text-to-Dance Accompaniment, where given music, textual prompt, and the leader's motion, the follower's motion is generated in a cohesive, text-aligned manner. We include baseline evaluations on both tasks to support future research.
【17】WildSpoof Challenge Evaluation Plan
标题:WildSpoof挑战评估计划
链接:https://arxiv.org/abs/2508.16858
备注:ICASSP 2026 challenge
摘要:WildSpoof Challenge旨在促进在两个相互交织的语音处理任务中使用野外数据。它包括两个并行的轨道:(1)文本到语音(TTS)合成生成欺骗语音,(2)欺骗鲁棒自动说话人验证(SASV)检测欺骗语音。虽然组织者协调两个轨道并定义数据协议,但参与者将其视为单独和独立的任务。挑战的主要目标是:(i)促进TTS和SASV使用野外数据,超越传统的清洁和受控数据集,并考虑真实世界的场景;(ii)鼓励欺骗生成(TTS)和欺骗检测(SASV)社区之间的跨学科合作,从而促进开发更集成,更强大,更现实的系统。
摘要:The WildSpoof Challenge aims to advance the use of in-the-wild data in two intertwined speech processing tasks. It consists of two parallel tracks: (1) Text-to-Speech (TTS) synthesis for generating spoofed speech, and (2) Spoofing-robust Automatic Speaker Verification (SASV) for detecting spoofed speech. While the organizers coordinate both tracks and define the data protocols, participants treat them as separate and independent tasks. The primary objectives of the challenge are: (i) to promote the use of in-the-wild data for both TTS and SASV, moving beyond conventional clean and controlled datasets and considering real-world scenarios; and (ii) to encourage interdisciplinary collaboration between the spoofing generation (TTS) and spoofing detection (SASV) communities, thereby fostering the development of more integrated, robust, and realistic systems.
【18】TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
标题:TaDiCodec:用于语音语言建模的文本感知扩散语音令牌器
链接:https://arxiv.org/abs/2508.16790
摘要:语音标记器作为语音语言模型的基础组件,但目前的设计表现出几个限制,包括:1)依赖于多层残差矢量量化结构或高帧速率,2)依赖于辅助预训练模型进行语义提取,以及3)要求复杂的两阶段训练过程。在这项工作中,我们介绍了文本感知扩散Transformer语音编解码器(TaDiCodec),一种新的方法,旨在克服这些挑战。TaDiCodec通过扩散自动编码器对量化和重建进行端到端优化,同时将文本指导集成到扩散解码器中,以提高重建质量并实现最佳压缩。TaDiCodec实现了6.25 Hz的极低帧速率和0.0875 kbps的相应比特率,采用单层码本用于24 kHz语音,同时在关键语音生成评估指标(如字错误率(WER),说话人相似性(SIM)和语音质量(UTMOS))方面保持卓越的性能。值得注意的是,TaDiCodec采用单阶段、端到端的训练范式,并且不需要辅助的预训练模型。我们还验证了TaDiCodec在基于语言模型的zero-shot文本到语音与自回归建模和掩蔽生成建模的兼容性,证明了其语音语言建模的有效性和效率,以及显着小的重建代差。我们将开源我们的代码和模型检查点。音频示例可在https:/tadicodec.github.io/上获得。我们在https:/github.com/HeCheng0625/Diffusion-Speech-Tokenizer上发布代码和模型检查点。
摘要:Speech tokenizers serve as foundational components for speech language models, yet current designs exhibit several limitations, including: 1) dependence on multi-layer residual vector quantization structures or high frame rates, 2) reliance on auxiliary pre-trained models for semantic distillation, and 3) requirements for complex two-stage training processes. In this work, we introduce the Text-aware Diffusion Transformer Speech Codec (TaDiCodec), a novel approach designed to overcome these challenges. TaDiCodec employs end-to-end optimization for quantization and reconstruction through a diffusion autoencoder, while integrating text guidance into the diffusion decoder to enhance reconstruction quality and achieve optimal compression. TaDiCodec achieves an extremely low frame rate of 6.25 Hz and a corresponding bitrate of 0.0875 kbps with a single-layer codebook for 24 kHz speech, while maintaining superior performance on critical speech generation evaluation metrics such as Word Error Rate (WER), speaker similarity (SIM), and speech quality (UTMOS). Notably, TaDiCodec employs a single-stage, end-to-end training paradigm, and obviating the need for auxiliary pre-trained models. We also validate the compatibility of TaDiCodec in language model based zero-shot text-to-speech with both autoregressive modeling and masked generative modeling, demonstrating its effectiveness and efficiency for speech language modeling, as well as a significantly small reconstruction-generation gap. We will open source our code and model checkpoints. Audio samples are are available at https:/tadicodec.github.io/. We release code and model checkpoints at https:/github.com/HeCheng0625/Diffusion-Speech-Tokenizer.
【19】VGGSounder: Audio-Visual Evaluations for Foundation Models
标题:VGG Sounder:基础模型的视听评估
链接:https://arxiv.org/abs/2508.08237
备注:Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025
摘要:视听基础模型的出现强调了可靠评估其多模态理解的重要性。VGGSound数据集通常用作评估视听分类的基准。然而,我们的分析确定了VGGSound的几个局限性,包括不完整的标签,部分重叠的类别和未对齐的模式。这些导致对听觉和视觉能力的扭曲评估。为了解决这些限制,我们引入了VGGSounder,这是一个全面重新注释的多标签测试集,扩展了VGGSound,专门用于评估视听基础模型。VGGSounder具有详细的模态注释功能,可以精确分析特定模态的性能。此外,我们揭示了模型的局限性,通过分析性能下降时,添加另一个输入模态与我们的新的模态混淆度量。
摘要:The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.
【20】Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
标题:使用适配器实现轻量级文本到语音的隐形说话者和语言适应
链接:https://arxiv.org/abs/2508.18006
备注:Accepted at IEEE MLSP 2025
摘要:在本文中,我们研究跨语言的文本到语音(TTS)合成通过镜头的适配器,在轻量级的TTS系统的背景下。特别是,我们比较看不见的扬声器和语言适应的目标,合成目标语音在目标语言,其中目标语音没有录音的任务。客观评估的结果证明了适配器在学习特定语言和特定说话者信息方面的有效性,允许预先训练的模型学习看不见的说话者身份或语言,同时避免灾难性地忘记原始模型的说话者或语言信息。此外,要衡量如何本地生成的声音是口音方面,我们提出并验证了一个客观的度量的启发,在第二语言(L2)学习者的发音错误检测技术。该文件还提供了深入了解适配器的位置,配置和使用的扬声器数量的影响。
摘要:In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal of synthesising a target voice in a target language, in which the target voice has no recordings therein. Results from objective evaluations demonstrate the effectiveness of adapters in learning language-specific and speaker-specific information, allowing pre-trained models to learn unseen speaker identities or languages, while avoiding catastrophic forgetting of the original model's speaker or language information. Additionally, to measure how native the generated voices are in terms of accent, we propose and validate an objective metric inspired by mispronunciation detection techniques in second-language (L2) learners. The paper also provides insights into the impact of adapter placement, configuration and the number of speakers used.
【21】Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
标题:基于扩散的言语增强对发音障碍的客观和主观评价
链接:https://arxiv.org/abs/2508.17980
备注:Accepted to Interspeech 2025
摘要:构音障碍语音由于其高变异性和低可懂度,对自动语音识别(ASR)系统提出了重大挑战。在这项工作中,我们探讨了使用扩散模型的构音障碍语音增强,这是基于这样的假设,即使用基于扩散的语音增强移动构音障碍语音的分布更接近典型的语音,这可能会提高构音障碍语音识别性能。我们评估两个扩散为基础的语音增强算法和一个信号处理为基础的语音增强算法的可懂度和语音质量的两个英语构音障碍的语音语料库。我们将语音增强应用于典型和构音障碍语音,并使用Whisper-Turbo评估ASR性能,以及原始和增强的构音障碍语音的主观和客观语音质量。我们还对增强语音的Whisper-Turbo进行了微调,以评估其对识别性能的影响。
摘要:Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
【22】HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
标题:Hunyuan Video-Foley:高保真Foley音频生成的多模式扩散和表示对齐
链接:https://arxiv.org/abs/2508.16930
摘要:视频生成的最新进展产生了视觉上逼真的内容,但同步音频的缺乏严重损害了沉浸感。为了解决视频到音频生成中的关键挑战,包括现有方法中的多模态数据稀缺,模态不平衡和有限的音频质量,我们提出了浑源Video-Foley,一个端到端的文本视频到音频框架,该框架能够精确地合成与视觉动态和语义上下文相一致的高保真音频。我们的方法包含三个核心创新:(1)通过自动标注管理10万小时多模态数据集的可扩展数据管道;(2)使用自监督音频特征指导潜在扩散训练的表示对齐策略,有效提高音频质量和生成稳定性;(3)提出了一种新的多模态扩散Transformer,解决了模态竞争问题,包括通过联合注意实现的双流音视频融合和通过交叉注意实现的文本语义注入。综合评估表明,浑源Video-Foley在音频保真度、视觉语义对齐、时间对齐和分布匹配方面实现了新的最先进的性能。演示页面可以在https://szczesnys.github.io/hunyuanvideo-foley/上找到。
摘要:Recent advances in video generation produce visually realistic content, yet the absence of synchronized audio severely compromises immersion. To address key challenges in video-to-audio generation, including multimodal data scarcity, modality imbalance and limited audio quality in existing methods, we propose HunyuanVideo-Foley, an end-to-end text-video-to-audio framework that synthesizes high-fidelity audio precisely aligned with visual dynamics and semantic context. Our approach incorporates three core innovations: (1) a scalable data pipeline curating 100k-hour multimodal datasets through automated annotation; (2) a representation alignment strategy using self-supervised audio features to guide latent diffusion training, efficiently improving audio quality and generation stability; (3) a novel multimodal diffusion transformer resolving modal competition, containing dual-stream audio-video fusion through joint attention, and textual semantic injection via cross-attention. Comprehensive evaluations demonstrate that HunyuanVideo-Foley achieves new state-of-the-art performance across audio fidelity, visual-semantic alignment, temporal alignment and distribution matching. The demo page is available at: https://szczesnys.github.io/hunyuanvideo-foley/.
【23】Localization using Angle-of-Arrival Triangulation
标题:使用到达角三角测量的定位
链接:https://arxiv.org/abs/2508.16908
备注:6 pages, 5 figures, 1 table. Accepted at the ACM International Workshop on Environmental Sensing Systems for Smart Cities (EnvSys 2025). To appear in the MobiSys 2025 Proceedings
摘要:室内定位是移动计算中的一个长期挑战,对于在家庭、办公室和零售空间等智能环境中实现位置感知和智能应用具有重要意义。随着亚马逊Alexa和谷歌Nest等人工智能助手变得越来越普遍,配备麦克风的设备正在成为日常生活和家庭自动化的关键组成部分。本文介绍了一种被动的,基础设施光系统定位人类说话人使用两个或两个以上的空间分布的智能设备捕获的语音信号。所提出的方法,GCC+,扩展了广义互相关与相位变换(GCC-PHAT)的方法来估计在每个设备的音频信号的到达角(AoA),并应用强大的三角测量技术来推断扬声器的二维位置。为了进一步提高时间分辨率和定位精度,特征空间扩展和子样本插值技术被用于精确的到达时间差(TDoA)估计。该系统无需硬件修改、事先校准、明确的用户合作或扬声器信号内容的知识即可运行,从而为现实世界的部署提供了高度实用的解决方案。在真实家庭环境中的实验评估产生了2.2度的中值AoA估计误差和1.25米的中值定位误差,证明了基于音频的定位用于实现上下文感知、隐私保护的环境智能的可行性和有效性。
摘要:Indoor localization is a long-standing challenge in mobile computing, with significant implications for enabling location-aware and intelligent applications within smart environments such as homes, offices, and retail spaces. As AI assistants such as Amazon Alexa and Google Nest become increasingly pervasive, microphone-equipped devices are emerging as key components of everyday life and home automation. This paper introduces a passive, infrastructure-light system for localizing human speakers using speech signals captured by two or more spatially distributed smart devices. The proposed approach, GCC+, extends the Generalized Cross-Correlation with Phase Transform (GCC-PHAT) method to estimate the Angle-of-Arrival (AoA) of audio signals at each device and applies robust triangulation techniques to infer the speaker's two-dimensional position. To further improve temporal resolution and localization accuracy, feature-space expansion and subsample interpolation techniques are employed for precise Time Difference of Arrival (TDoA) estimation. The system operates without requiring hardware modifications, prior calibration, explicit user cooperation, or knowledge of the speaker's signal content, thereby offering a highly practical solution for real-world deployment. Experimental evaluation in a real-world home environment yields a median AoA estimation error of 2.2 degrees and a median localization error of 1.25 m, demonstrating the feasibility and effectiveness of audio-based localization for enabling context-aware, privacy-preserving ambient intelligence.
【1】Unseen Speaker and Language Adaptation for Lightweight Text-To-Speech with Adapters
标题:使用适配器实现轻量级文本到语音的隐形说话者和语言适应
链接:https://arxiv.org/abs/2508.18006
备注:Accepted at IEEE MLSP 2025
摘要:在本文中,我们研究跨语言的文本到语音(TTS)合成通过镜头的适配器,在轻量级的TTS系统的背景下。特别是,我们比较看不见的扬声器和语言适应的目标,合成目标语音在目标语言,其中目标语音没有录音的任务。客观评估的结果证明了适配器在学习特定语言和特定说话者信息方面的有效性,允许预先训练的模型学习看不见的说话者身份或语言,同时避免灾难性地忘记原始模型的说话者或语言信息。此外,要衡量如何本地生成的声音是口音方面,我们提出并验证了一个客观的度量的启发,在第二语言(L2)学习者的发音错误检测技术。该文件还提供了深入了解适配器的位置,配置和使用的扬声器数量的影响。
摘要:In this paper we investigate cross-lingual Text-To-Speech (TTS) synthesis through the lens of adapters, in the context of lightweight TTS systems. In particular, we compare the tasks of unseen speaker and language adaptation with the goal of synthesising a target voice in a target language, in which the target voice has no recordings therein. Results from objective evaluations demonstrate the effectiveness of adapters in learning language-specific and speaker-specific information, allowing pre-trained models to learn unseen speaker identities or languages, while avoiding catastrophic forgetting of the original model's speaker or language information. Additionally, to measure how native the generated voices are in terms of accent, we propose and validate an objective metric inspired by mispronunciation detection techniques in second-language (L2) learners. The paper also provides insights into the impact of adapter placement, configuration and the number of speakers used.
【2】Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
标题:基于扩散的言语增强对发音障碍的客观和主观评价
链接:https://arxiv.org/abs/2508.17980
备注:Accepted to Interspeech 2025
摘要:构音障碍语音由于其高变异性和低可懂度,对自动语音识别(ASR)系统提出了重大挑战。在这项工作中,我们探讨了使用扩散模型的构音障碍语音增强,这是基于这样的假设,即使用基于扩散的语音增强移动构音障碍语音的分布更接近典型的语音,这可能会提高构音障碍语音识别性能。我们评估两个扩散为基础的语音增强算法和一个信号处理为基础的语音增强算法的可懂度和语音质量的两个英语构音障碍的语音语料库。我们将语音增强应用于典型和构音障碍语音,并使用Whisper-Turbo评估ASR性能,以及原始和增强的构音障碍语音的主观和客观语音质量。我们还对增强语音的Whisper-Turbo进行了微调,以评估其对识别性能的影响。
摘要:Dysarthric speech poses significant challenges for automatic speech recognition (ASR) systems due to its high variability and reduced intelligibility. In this work we explore the use of diffusion models for dysarthric speech enhancement, which is based on the hypothesis that using diffusion-based speech enhancement moves the distribution of dysarthric speech closer to that of typical speech, which could potentially improve dysarthric speech recognition performance. We assess the effect of two diffusion-based and one signal-processing-based speech enhancement algorithms on intelligibility and speech quality of two English dysarthric speech corpora. We applied speech enhancement to both typical and dysarthric speech and evaluate the ASR performance using Whisper-Turbo, and the subjective and objective speech quality of the original and enhanced dysarthric speech. We also fine-tuned Whisper-Turbo on the enhanced speech to assess its impact on recognition performance.
【3】Optimal Pairwise Comparison Procedures for Subjective Evaluation
标题:主观评价的最佳成对比较程序
链接:https://arxiv.org/abs/2508.17840
备注:11th Convention of the European Acoustics Association, Forum Acusticum 2025, Málaga
摘要:音频信号处理算法经常通过主观听力测试进行评估,在主观听力测试中,参与者直接在一维数字尺度上对退化信号进行评分。然而,这种方法容易受到评估员之间尺度校准不一致的影响。降级信号之间的成对比较提供了一种更直观的替代方案,以较低的测量误差和降低的参与者疲劳度引出候选信号的相对分数。然而,由于必要的比较数量的二次增长,一个完整的成对比较集对于大型数据集变得不可行。本文比较了成对比较程序,以确定最有效的方法近似真实的质量分数与最小的比较。提出了一种新的采样方法,并在模拟数据集上对最先进的方法进行了基准测试。贝叶斯抽样产生最强大的得分估计在以前建立的方法,而建议的程序始终收敛最快的基本排名具有可比的得分精度。
摘要:Audio signal processing algorithms are frequently assessed through subjective listening tests in which participants directly score degraded signals on a unidimensional numerical scale. However, this approach is susceptible to inconsistencies in scale calibration between assessors. Pairwise comparisons between degraded signals offer a more intuitive alternative, eliciting the relative scores of candidate signals with lower measurement error and reduced participant fatigue. Yet, due to the quadratic growth of the number of necessary comparisons, a complete set of pairwise comparisons becomes unfeasible for large datasets. This paper compares pairwise comparison procedures to identify the most efficient methods for approximating true quality scores with minimal comparisons. A novel sampling procedure is proposed and benchmarked against state-of-the-art methods on simulated datasets. Bayesian sampling produces the most robust score estimates among previously established methods, while the proposed procedure consistently converges fastest on the underlying ranking with comparable score accuracy.
【4】Pinhole Effect on Linkability and Dispersion in Speaker Anonymization
标题:说话人模拟中的针孔效应对连通性和分散性
链接:https://arxiv.org/abs/2508.17134
备注:5 pages, 2 figures
摘要:说话人匿名化的目的是隐藏语音信号中的特定说话人属性,使匿名后的语音不受说话人身份的影响。最近的方法通过将语音分解为内容和扬声器组件来实现这一点,用伪扬声器代替后者。匿名语音可以映射到跨话语共享的公共伪说话者,也可以映射到每个话语所特有的不同伪说话者。本文研究了这些映射策略对三个关键维度的影响:说话人可链接性,在匿名说话人空间中的分散性,以及从原始身份的去识别。我们的研究结果表明,使用不同的伪扬声器增加扬声器分散和减少链接相比,共同的伪扬声器映射,从而提高隐私保护。这些意见解释通过建议针孔效应,概念框架介绍,以解释映射策略和匿名化性能之间的关系。通过实证分析验证了这一假设。
摘要:Speaker anonymization aims to conceal speaker-specific attributes in speech signals, making the anonymized speech unlinkable to the original speaker identity. Recent approaches achieve this by disentangling speech into content and speaker components, replacing the latter with pseudo speakers. The anonymized speech can be mapped either to a common pseudo speaker shared across utterances or to distinct pseudo speakers unique to each utterance. This paper investigates the impact of these mapping strategies on three key dimensions: speaker linkability, dispersion in the anonymized speaker space, and de-identification from the original identity. Our findings show that using distinct pseudo speakers increases speaker dispersion and reduces linkability compared to common pseudo-speaker mapping, thereby enhancing privacy preservation. These observations are interpreted through the proposed pinhole effect, a conceptual framework introduced to explain the relationship between mapping strategies and anonymization performance. The hypothesis is validated through empirical evaluation.
【5】HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
标题:Hunyuan Video-Foley:高保真Foley音频生成的多模式扩散和表示对齐
链接:https://arxiv.org/abs/2508.16930
摘要:视频生成的最新进展产生了视觉上逼真的内容,但同步音频的缺乏严重损害了沉浸感。为了解决视频到音频生成中的关键挑战,包括现有方法中的多模态数据稀缺,模态不平衡和有限的音频质量,我们提出了浑源Video-Foley,一个端到端的文本视频到音频框架,该框架能够精确地合成与视觉动态和语义上下文相一致的高保真音频。我们的方法包含三个核心创新:(1)通过自动标注管理10万小时多模态数据集的可扩展数据管道;(2)使用自监督音频特征指导潜在扩散训练的表示对齐策略,有效提高音频质量和生成稳定性;(3)提出了一种新的多模态扩散Transformer,解决了模态竞争问题,包括通过联合注意实现的双流音视频融合和通过交叉注意实现的文本语义注入。综合评估表明,浑源Video-Foley在音频保真度、视觉语义对齐、时间对齐和分布匹配方面实现了新的最先进的性能。演示页面可以在https://szczesnys.github.io/hunyuanvideo-foley/上找到。
摘要:Recent advances in video generation produce visually realistic content, yet the absence of synchronized audio severely compromises immersion. To address key challenges in video-to-audio generation, including multimodal data scarcity, modality imbalance and limited audio quality in existing methods, we propose HunyuanVideo-Foley, an end-to-end text-video-to-audio framework that synthesizes high-fidelity audio precisely aligned with visual dynamics and semantic context. Our approach incorporates three core innovations: (1) a scalable data pipeline curating 100k-hour multimodal datasets through automated annotation; (2) a representation alignment strategy using self-supervised audio features to guide latent diffusion training, efficiently improving audio quality and generation stability; (3) a novel multimodal diffusion transformer resolving modal competition, containing dual-stream audio-video fusion through joint attention, and textual semantic injection via cross-attention. Comprehensive evaluations demonstrate that HunyuanVideo-Foley achieves new state-of-the-art performance across audio fidelity, visual-semantic alignment, temporal alignment and distribution matching. The demo page is available at: https://szczesnys.github.io/hunyuanvideo-foley/.
【6】Localization using Angle-of-Arrival Triangulation
标题:使用到达角三角测量的定位
链接:https://arxiv.org/abs/2508.16908
备注:6 pages, 5 figures, 1 table. Accepted at the ACM International Workshop on Environmental Sensing Systems for Smart Cities (EnvSys 2025). To appear in the MobiSys 2025 Proceedings
摘要:室内定位是移动计算中的一个长期挑战,对于在家庭、办公室和零售空间等智能环境中实现位置感知和智能应用具有重要意义。随着亚马逊Alexa和谷歌Nest等人工智能助手变得越来越普遍,配备麦克风的设备正在成为日常生活和家庭自动化的关键组成部分。本文介绍了一种无源基础设施光系统,用于使用两个或更多空间分布的智能设备捕获的语音信号来定位人类说话者。所提出的方法,GCC+,扩展了广义互相关与相位变换(GCC-PHAT)的方法来估计在每个设备的音频信号的到达角(AoA),并应用强大的三角测量技术来推断扬声器的二维位置。为了进一步提高时间分辨率和定位精度,特征空间扩展和子样本插值技术被用于精确的到达时间差(TDoA)估计。该系统无需硬件修改、事先校准、明确的用户合作或扬声器信号内容的知识即可运行,从而为现实世界的部署提供了高度实用的解决方案。在真实家庭环境中的实验评估产生了2.2度的中值AoA估计误差和1.25米的中值定位误差,证明了基于音频的定位用于实现上下文感知、隐私保护的环境智能的可行性和有效性。
摘要:Indoor localization is a long-standing challenge in mobile computing, with significant implications for enabling location-aware and intelligent applications within smart environments such as homes, offices, and retail spaces. As AI assistants such as Amazon Alexa and Google Nest become increasingly pervasive, microphone-equipped devices are emerging as key components of everyday life and home automation. This paper introduces a passive, infrastructure-light system for localizing human speakers using speech signals captured by two or more spatially distributed smart devices. The proposed approach, GCC+, extends the Generalized Cross-Correlation with Phase Transform (GCC-PHAT) method to estimate the Angle-of-Arrival (AoA) of audio signals at each device and applies robust triangulation techniques to infer the speaker's two-dimensional position. To further improve temporal resolution and localization accuracy, feature-space expansion and subsample interpolation techniques are employed for precise Time Difference of Arrival (TDoA) estimation. The system operates without requiring hardware modifications, prior calibration, explicit user cooperation, or knowledge of the speaker's signal content, thereby offering a highly practical solution for real-world deployment. Experimental evaluation in a real-world home environment yields a median AoA estimation error of 2.2 degrees and a median localization error of 1.25 m, demonstrating the feasibility and effectiveness of audio-based localization for enabling context-aware, privacy-preserving ambient intelligence.
【7】Vocoder-Projected Feature Discriminator
标题:声码器投影特征鉴别器
链接:https://arxiv.org/abs/2508.17874
备注:Accepted to Interspeech 2024. Project page: this https URL
摘要:在文本到语音(TTS)和语音转换(VC)中,声学特征(诸如梅尔频谱图)由于其紧凑性和易于学习而通常用作合成或转换目标。然而,由于最终目标是生成高质量的波形,因此采用声码器将这些特征转换为波形并在时域中应用对抗训练是合理的。然而,对波形进行上采样引入了显著的时间和存储器开销。为了解决这个问题,我们提出了一个声码器投影的特征提取(VPFD),它使用声码器的功能对抗训练。基于扩散的VC蒸馏实验表明,一个预训练和冻结的声码器特征提取器与一个单一的上采样步骤是必要的和足够的,以实现VC性能与波形鉴别器,同时减少训练时间和内存消耗的9.6和11.4倍,分别。
摘要:In text-to-speech (TTS) and voice conversion (VC), acoustic features, such as mel spectrograms, are typically used as synthesis or conversion targets owing to their compactness and ease of learning. However, because the ultimate goal is to generate high-quality waveforms, employing a vocoder to convert these features into waveforms and applying adversarial training in the time domain is reasonable. Nevertheless, upsampling the waveform introduces significant time and memory overheads. To address this issue, we propose a vocoder-projected feature discriminator (VPFD), which uses vocoder features for adversarial training. Experiments on diffusion-based VC distillation demonstrated that a pretrained and frozen vocoder feature extractor with a single upsampling step is necessary and sufficient to achieve a VC performance comparable to that of waveform discriminators while reducing the training time and memory consumption by 9.6 and 11.4 times, respectively.
【8】FasterVoiceGrad: Faster One-step Diffusion-Based Voice Conversion with Adversarial Diffusion Conversion Distillation
标题:FasterEqualGrad:更快的基于扩散的一步语音转换,采用对抗扩散转换蒸馏
链接:https://arxiv.org/abs/2508.17868
备注:Accepted to Interspeech 2025. Project page: this https URL
摘要:基于扩散的语音转换(VC)模型(例如,VoiceGrad)可以实现较高的语音质量和说话人相似度,但是由于迭代采样,其转换过程较慢。FastVoiceGrad通过将VoiceGrad提炼为一步扩散模型来克服这一限制。然而,它仍然需要一个计算密集的内容编码器来解开说话者的身份和内容,这会减慢转换速度。因此,我们提出了FasterVoiceGrad,这是一种新的基于扩散的一步VC模型,通过使用对抗扩散转换蒸馏(ADCD)同时蒸馏扩散模型和内容编码器获得,其中蒸馏在转换过程中进行,同时利用对抗和分数蒸馏训练。对单次VC的实验评估表明,与FastVoiceGrad相比,FastVoiceGrad实现了具有竞争力的VC性能,在GPU和CPU上的速度分别提高了6.6-6.9和1.8倍。
摘要:A diffusion-based voice conversion (VC) model (e.g., VoiceGrad) can achieve high speech quality and speaker similarity; however, its conversion process is slow owing to iterative sampling. FastVoiceGrad overcomes this limitation by distilling VoiceGrad into a one-step diffusion model. However, it still requires a computationally intensive content encoder to disentangle the speaker's identity and content, which slows conversion. Therefore, we propose FasterVoiceGrad, a novel one-step diffusion-based VC model obtained by simultaneously distilling a diffusion model and content encoder using adversarial diffusion conversion distillation (ADCD), where distillation is performed in the conversion process while leveraging adversarial and score distillation training. Experimental evaluations of one-shot VC demonstrated that FasterVoiceGrad achieves competitive VC performance compared to FastVoiceGrad, with 6.6-6.9 and 1.8 times faster speed on a GPU and CPU, respectively.
【9】Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
标题:使用合成多发音的基于Trie的解码的Zero-Shot上下文偏置
链接:https://arxiv.org/abs/2508.17796
备注:Accepted to APSIPA ASC 2025
摘要:上下文自动语音识别(ASR)系统允许识别词汇表外(OOV)单词,诸如命名实体或稀有单词。然而,由于训练数据有限,发音模糊或不一致,它仍然具有挑战性。在本文中,我们提出了一种合成驱动的多发音上下文偏置方法,该方法在预训练的Whisper模型上执行zero-shot上下文ASR。具体来说,我们利用文本到语音(TTS)系统来合成包含每个目标稀有词的不同语音样本,然后使用预训练的Whisper模型来提取多个预测的发音变体。这些变体令牌序列被编译成前缀特里,其在波束搜索解码期间以浅融合方式向波束假设分配奖励。在此之后,任何识别出的变体都将在最终转录中映射回原始的罕见单词。Librispeech数据集上的评估结果表明,我们的方法将有偏的单词错误率(WER)降低了42%,在测试-干净和43%的测试-其他,同时保持无偏的WER基本不变。
摘要:Contextual automatic speech recognition (ASR) systems allow for recognizing out-of-vocabulary (OOV) words, such as named entities or rare words. However, it remains challenging due to limited training data and ambiguous or inconsistent pronunciations. In this paper, we propose a synthesis-driven multi-pronunciation contextual biasing method that performs zero-shot contextual ASR on a pretrained Whisper model. Specifically, we leverage text-to-speech (TTS) systems to synthesize diverse speech samples containing each target rare word, and then use the pretrained Whisper model to extract multiple predicted pronunciation variants. These variant token sequences are compiled into a prefix-trie, which assigns rewards to beam hypotheses in a shallow-fusion manner during beam-search decoding. After which, any recognized variant is mapped back to the original rare word in the final transcription. The evaluation results on the Librispeech dataset show that our method reduces biased word error rate (WER) by 42% on test-clean and 43% on test-other while maintaining unbiased WER essentially unchanged.
【10】EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Spoken Dialogue Systems
标题:情绪推理:口语对话系统中情绪推理能力的基准测试
链接:https://arxiv.org/abs/2508.17623
备注:Accepted at (ASRU 2025) 2025 IEEE Automatic Speech Recognition and Understanding Workshop
摘要:语音情感在人机交互、塑造参与和上下文感知通信中起着至关重要的作用。尽管口语对话系统最近取得了进展,但仍然缺乏评估情感推理的整体系统。为了解决这个问题,我们引入了EMO推理,这是一个评估对话系统中情感连贯性的基准。它利用通过文本到语音生成的策划数据集来模拟不同的情绪状态,克服了情感语音数据的稀缺性。我们进一步提出了跨话轮情感推理分数来评估多话轮对话中的情感转换。通过连续的,分类的和感知的指标评估七个对话系统,我们表明,我们的框架有效地检测情感不一致,为改善当前的对话系统提供见解。通过发布系统的评估基准,我们的目标是将情感感知的口语对话建模推向更自然和自适应的交互。
摘要:Speech emotions play a crucial role in human-computer interaction, shaping engagement and context-aware communication. Despite recent advances in spoken dialogue systems, a holistic system for evaluating emotional reasoning is still lacking. To address this, we introduce EMO-Reasoning, a benchmark for assessing emotional coherence in dialogue systems. It leverages a curated dataset generated via text-to-speech to simulate diverse emotional states, overcoming the scarcity of emotional speech data. We further propose the Cross-turn Emotion Reasoning Score to assess the emotion transitions in multi-turn dialogues. Evaluating seven dialogue systems through continuous, categorical, and perceptual metrics, we show that our framework effectively detects emotional inconsistencies, providing insights for improving current dialogue systems. By releasing a systematic evaluation benchmark, we aim to advance emotion-aware spoken dialogue modeling toward more natural and adaptive interactions.
【11】Multi-Metric Preference Alignment for Generative Speech Restoration
标题:生成式语音恢复中的多度量偏好对齐
链接:https://arxiv.org/abs/2508.17229
备注:16 pages, 10 figures. demopage: this https URL
摘要:最近的生成模型大大提高了语音恢复任务,但它们的训练目标往往与人类的感知偏好不一致,导致质量不佳。虽然训练后对齐在其他生成领域(如文本和图像生成)中已被证明是有效的,但其在生成语音恢复中的应用在很大程度上仍未得到充分探索。这项工作研究了将基于偏好的后训练应用于这项任务的挑战,重点是如何定义一个强大的偏好信号和策划高质量的数据,以避免奖励黑客。为了应对这些挑战,我们提出了一个多指标偏好对齐策略。我们构建了一个新的数据集GenSR-Pref,包括80 K个偏好对,其中每个选择的样本都受到一套互补指标的一致青睐,这些指标涵盖感知质量、信号保真度、内容一致性和音色保留。这种原则性的方法确保了整体的偏好信号。将直接偏好优化(DPO)应用于我们的数据集,我们在三种不同的生成范式中观察到一致和显着的性能增益:自回归模型(AR),掩蔽生成模型(MGM)和流匹配模型(FM)在各种恢复基准上,在客观和主观评估中。消融研究证实了我们的多指标策略在减轻奖励黑客攻击方面优于单指标方法。此外,我们证明了我们的对齐模型可以作为强大的“数据注释器”,生成高质量的伪标签,作为传统判别模型在数据稀缺的情况下(如歌声恢复)的监督信号。演示页面:https://gensr-pref.github.io
摘要:Recent generative models have significantly advanced speech restoration tasks, yet their training objectives often misalign with human perceptual preferences, resulting in suboptimal quality. While post-training alignment has proven effective in other generative domains like text and image generation, its application to generative speech restoration remains largely under-explored. This work investigates the challenges of applying preference-based post-training to this task, focusing on how to define a robust preference signal and curate high-quality data to avoid reward hacking. To address these challenges, we propose a multi-metric preference alignment strategy. We construct a new dataset, GenSR-Pref, comprising 80K preference pairs, where each chosen sample is unanimously favored by a complementary suite of metrics covering perceptual quality, signal fidelity, content consistency, and timbre preservation. This principled approach ensures a holistic preference signal. Applying Direct Preference Optimization (DPO) with our dataset, we observe consistent and significant performance gains across three diverse generative paradigms: autoregressive models (AR), masked generative models (MGM), and flow-matching models (FM) on various restoration benchmarks, in both objective and subjective evaluations. Ablation studies confirm the superiority of our multi-metric strategy over single-metric approaches in mitigating reward hacking. Furthermore, we demonstrate that our aligned models can serve as powerful ''data annotators'', generating high-quality pseudo-labels to serve as a supervision signal for traditional discriminative models in data-scarce scenarios like singing voice restoration. Demo Page:https://gensr-pref.github.io
【12】TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
标题:TaDiCodec:用于语音语言建模的文本感知扩散语音令牌器
链接:https://arxiv.org/abs/2508.16790
摘要:语音标记器作为语音语言模型的基础组件,但目前的设计表现出几个限制,包括:1)依赖于多层残差矢量量化结构或高帧速率,2)依赖于辅助预训练模型进行语义提取,以及3)要求复杂的两阶段训练过程。在这项工作中,我们介绍了文本感知扩散Transformer语音编解码器(TaDiCodec),一种新的方法,旨在克服这些挑战。TaDiCodec通过扩散自动编码器对量化和重建进行端到端优化,同时将文本指导集成到扩散解码器中,以提高重建质量并实现最佳压缩。TaDiCodec实现了6.25 Hz的极低帧速率和0.0875 kbps的相应比特率,采用单层码本用于24 kHz语音,同时在关键语音生成评估指标(如字错误率(WER),说话人相似性(SIM)和语音质量(UTMOS))方面保持卓越的性能。值得注意的是,TaDiCodec采用单阶段、端到端的训练范式,并且不需要辅助的预训练模型。我们还验证了TaDiCodec在基于语言模型的zero-shot文本到语音与自回归建模和掩蔽生成建模的兼容性,证明了其语音语言建模的有效性和效率,以及显着小的重建代差。我们将开源我们的代码和模型检查点。音频示例可在https:/tadicodec.github.io/上获得。我们在https:/github.com/HeCheng0625/Diffusion-Speech-Tokenizer上发布代码和模型检查点。
摘要:Speech tokenizers serve as foundational components for speech language models, yet current designs exhibit several limitations, including: 1) dependence on multi-layer residual vector quantization structures or high frame rates, 2) reliance on auxiliary pre-trained models for semantic distillation, and 3) requirements for complex two-stage training processes. In this work, we introduce the Text-aware Diffusion Transformer Speech Codec (TaDiCodec), a novel approach designed to overcome these challenges. TaDiCodec employs end-to-end optimization for quantization and reconstruction through a diffusion autoencoder, while integrating text guidance into the diffusion decoder to enhance reconstruction quality and achieve optimal compression. TaDiCodec achieves an extremely low frame rate of 6.25 Hz and a corresponding bitrate of 0.0875 kbps with a single-layer codebook for 24 kHz speech, while maintaining superior performance on critical speech generation evaluation metrics such as Word Error Rate (WER), speaker similarity (SIM), and speech quality (UTMOS). Notably, TaDiCodec employs a single-stage, end-to-end training paradigm, and obviating the need for auxiliary pre-trained models. We also validate the compatibility of TaDiCodec in language model based zero-shot text-to-speech with both autoregressive modeling and masked generative modeling, demonstrating its effectiveness and efficiency for speech language modeling, as well as a significantly small reconstruction-generation gap. We will open source our code and model checkpoints. Audio samples are are available at https:/tadicodec.github.io/. We release code and model checkpoints at https:/github.com/HeCheng0625/Diffusion-Speech-Tokenizer.
【13】VGGSounder: Audio-Visual Evaluations for Foundation Models
标题:VGG Sounder:基础模型的视听评估
链接:https://arxiv.org/abs/2508.08237
备注:Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2025
摘要:视听基础模型的出现强调了可靠评估其多模态理解的重要性。VGGSound数据集通常用作评估视听分类的基准。然而,我们的分析确定了VGGSound的几个局限性,包括不完整的标签,部分重叠的类别和未对齐的模式。这些导致对听觉和视觉能力的扭曲评估。为了解决这些限制,我们引入了VGGSounder,这是一个全面重新注释的多标签测试集,扩展了VGGSound,专门用于评估视听基础模型。VGGSounder具有详细的模态注释功能,可以精确分析特定模态的性能。此外,我们揭示了模型的局限性,通过分析性能下降时,添加另一个输入模态与我们的新的模态混淆度量。
摘要:The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchmark for evaluation audio-visual classification. However, our analysis identifies several limitations of VGGSound, including incomplete labelling, partially overlapping classes, and misaligned modalities. These lead to distorted evaluations of auditory and visual capabilities. To address these limitations, we introduce VGGSounder, a comprehensively re-annotated, multi-label test set that extends VGGSound and is specifically designed to evaluate audio-visual foundation models. VGGSounder features detailed modality annotations, enabling precise analyses of modality-specific performance. Furthermore, we reveal model limitations by analysing performance degradation when adding another input modality with our new modality confusion metric.
机器翻译由腾讯交互翻译提供,仅供参考
