微信公众号:arXiv_Daily
cs.SD语音
【1】Exploring Definitions of Quality and Diversity in Sonic Measurement Spaces
标题:探索声波测量空间中质量和多样性的定义
链接:https://arxiv.org/abs/2512.02783
摘要:数字声音合成提供了探索包含数百万种配置的巨大参数空间的机会。质量多样性(QD)进化算法提供了一个很有前途的方法来利用这种潜力,但他们的成功取决于适当的声音特征表示。现有的QD方法主要采用手工制作的描述符或监督分类器,可能会引入意想不到的探索偏差,并将发现限制在熟悉的声波区域。这项工作研究了无监督的降维方法,自动定义和动态重新配置声波行为空间在QD搜索。我们应用主成分分析(PCA)和自动编码器将高维音频特征投影到MAP-Elites的结构化网格上,通过定期进行模型再训练来实现动态重新配置。两个实验方案的比较表明,自动方法实现了显着更大的多样性比手工制作的行为空间,同时避免专家施加的偏见。动态行为空间重构保持进化压力,防止停滞,PCA证明是最有效的降维技术。这些结果有助于自动化的声波发现系统能够探索巨大的参数空间,而无需人工干预或监督训练的限制。
摘要:Digital sound synthesis presents the opportunity to explore vast parameter spaces containing millions of configurations. Quality diversity (QD) evolutionary algorithms offer a promising approach to harness this potential, yet their success hinges on appropriate sonic feature representations. Existing QD methods predominantly employ handcrafted descriptors or supervised classifiers, potentially introducing unintended exploration biases and constraining discovery to familiar sonic regions. This work investigates unsupervised dimensionality reduction methods for automatically defining and dynamically reconfiguring sonic behaviour spaces during QD search. We apply Principal Component Analysis (PCA) and autoencoders to project high-dimensional audio features onto structured grids for MAP-Elites, implementing dynamic reconfiguration through model retraining at regular intervals. Comparison across two experimental scenarios shows that automatic approaches achieve significantly greater diversity than handcrafted behaviour spaces while avoiding expert-imposed biases. Dynamic behaviour-space reconfiguration maintains evolutionary pressure and prevents stagnation, with PCA proving most effective among the dimensionality reduction techniques. These results contribute to automated sonic discovery systems capable of exploring vast parameter spaces without manual intervention or supervised training constraints.
【2】SAND Challenge: Four Approaches for Dysartria Severity Classification
标题:SAND挑战:味觉障碍严重程度分类的四种方法
链接:https://arxiv.org/abs/2512.02669
备注:7 pages, 5 figures
摘要:本文提出了一个统一的研究,四个不同的建模方法分类构音障碍的严重程度在语音分析神经退行性疾病(SAND)的挑战。所有模型都使用一个通用的语音记录数据集来处理相同的五类分类任务。我们调查:(1)在频谱图图像上利用Vision Transformer的ViT-OF方法,(2)使用具有多数投票融合的八个1-D CNN的1D-CNN方法,(3)使用具有多数投票融合的九个BiLSTM模型的BiLSTM-OF方法,以及(4)通过两阶段学习框架组合声门和共振峰特征的分层XGBoost集成。每种方法进行了描述,并对它们在53个扬声器的验证集上的性能进行了比较。结果表明,虽然经过特征设计的XGBoost集成达到了最高的宏F1(0.86),但深度学习模型(ViT,CNN,BiLSTM)达到了具有竞争力的F1分数(0.70),并提供了对问题的补充见解。
摘要:This paper presents a unified study of four distinct modeling approaches for classifying dysarthria severity in the Speech Analysis for Neurodegenerative Diseases (SAND) challenge. All models tackle the same five class classification task using a common dataset of speech recordings. We investigate: (1) a ViT-OF method leveraging a Vision Transformer on spectrogram images, (2) a 1D-CNN approach using eight 1-D CNN's with majority-vote fusion, (3) a BiLSTM-OF approach using nine BiLSTM models with majority vote fusion, and (4) a Hierarchical XGBoost ensemble that combines glottal and formant features through a two stage learning framework. Each method is described, and their performances on a validation set of 53 speakers are compared. Results show that while the feature-engineered XGBoost ensemble achieves the highest macro-F1 (0.86), the deep learning models (ViT, CNN, BiLSTM) attain competitive F1-scores (0.70) and offer complementary insights into the problem.
【3】Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
标题:钢琴家Transformer:通过可扩展的自我监督预训练实现表现力的钢琴演奏渲染
链接:https://arxiv.org/abs/2512.02652
摘要:现有的表现性音乐表现渲染方法依赖于对小标记数据集的监督学习,这限制了数据量和模型大小的缩放,尽管存在大量未标记的音乐,如视觉和语言。为了解决这一问题,我们引入了Pianist Transformer,它有四个主要贡献:1)统一的乐器数字接口(Musical Instrument Digital Interface,简称MIDI)数据表示,用于学习音乐结构和表达的共享原则,而无需显式注释; 2)高效的非对称架构,在不牺牲渲染质量的情况下实现更长的上下文和更快的推理; 3)具有10B令牌和135M参数模型的自监督预训练管道,解锁数据和模型缩放优势,以实现表现力表现渲染; 4)最先进的性能模型,实现强大的客观指标和人类主观评级。总的来说,钢琴家Transformer建立了一个可扩展的路径,向人类一样的性能合成在音乐领域。
摘要:Existing methods for expressive music performance rendering rely on supervised learning over small labeled datasets, which limits scaling of both data volume and model size, despite the availability of vast unlabeled music, as in vision and language. To address this gap, we introduce Pianist Transformer, with four key contributions: 1) a unified Musical Instrument Digital Interface (MIDI) data representation for learning the shared principles of musical structure and expression without explicit annotation; 2) an efficient asymmetric architecture, enabling longer contexts and faster inference without sacrificing rendering quality; 3) a self-supervised pre-training pipeline with 10B tokens and 135M-parameter model, unlocking data and model scaling advantages for expressive performance rendering; 4) a state-of-the-art performance model, which achieves strong objective metrics and human-level subjective ratings. Overall, Pianist Transformer establishes a scalable path toward human-like performance synthesis in the music domain.
【4】Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
标题:听听什么重要!文本条件选择性视频到音频生成
链接:https://arxiv.org/abs/2512.02650
摘要:这项工作介绍了一个新的任务,文本条件下的选择性视频到音频(V2 A)的生成,它只产生用户想要的声音从多对象视频。这种能力在多媒体制作中尤其重要,因为每个声源的音轨都要单独处理,以便进行精确的编辑、混音和创意控制。然而,目前的方法一次生成单源混合声音,主要是因为视觉特征纠缠在一起,区域提示或提示往往无法指定来源。我们提出了SelVA,一种新的文本条件V2 A模型,将文本提示作为目标源的显式选择器,并调制视频编码器以明显地提取与文本相关的视频特征。建议的补充令牌促进交叉注意,抑制文本无关的激活与有效的参数调整,产生强大的语义和时间接地。SelVA进一步采用自增强方案来克服单声道音轨监督的缺乏。我们在VGG-MONOAUDIO上评估SelVA,VGG-MONOAUDIO是一个针对此类任务的干净单源视频的策划基准。大量的实验和消融一致地验证了其在音频质量、语义对齐和时间同步方面的有效性。代码和演示可在https://jnwnlee.github.io/selva-demo/上获得。
摘要:This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio tracks are handled individually for each sound source for precise editing, mixing, and creative control. However, current approaches generate single source-mixed sounds at once, largely because visual features are entangled, and region cues or prompts often fail to specify the source. We propose SelVA, a novel text-conditioned V2A model that treats the text prompt as an explicit selector of target source and modulates video encoder to distinctly extract prompt-relevant video features. The proposed supplementary tokens promote cross-attention by suppressing text-irrelevant activations with efficient parameter tuning, yielding robust semantic and temporal grounding. SelVA further employs a self-augmentation scheme to overcome the lack of mono audio track supervision. We evaluate SelVA on VGG-MONOAUDIO, a curated benchmark of clean single-source videos for such a task. Extensive experiments and ablations consistently verify its effectiveness across audio quality, semantic alignment, and temporal synchronization. Code and demo are available at https://jnwnlee.github.io/selva-demo/.
【5】Spoken Conversational Agents with Large Language Models
标题:具有大型语言模型的口语对话代理
链接:https://arxiv.org/abs/2512.02593
备注:Accepted to EMNLP 2025 Tutorial
摘要:口语会话代理正在向语音本地LLM融合。本教程将从级联ASR/NLU到端到端、基于检索和视觉的系统的路径进行提炼。我们将文本LLM框架化为音频,跨模态对齐和联合语音文本训练;审查数据集,指标和跨口音的鲁棒性,并比较设计选择(级联与E2 E,ASR后校正,流)。我们将工业助理与当前的开放领域和面向任务的代理联系起来,突出可重复的基线,并概述隐私,安全和评估方面的开放问题。与会者带着实用的配方和清晰的系统级路线图离开。
摘要:Spoken conversational agents are converging toward voice-native LLMs. This tutorial distills the path from cascaded ASR/NLU to end-to-end, retrieval-and vision-grounded systems. We frame adaptation of text LLMs to audio, cross-modal alignment, and joint speech-text training; review datasets, metrics, and robustness across accents and compare design choices (cascaded vs. E2E, post-ASR correction, streaming). We link industrial assistants to current open-domain and task-oriented agents, highlight reproducible baselines, and outline open problems in privacy, safety, and evaluation. Attendees leave with practical recipes and a clear systems-level roadmap.
【6】Generative Multi-modal Feedback for Singing Voice Synthesis Evaluation
标题:用于歌唱声音合成评估的生成多模式反馈
链接:https://arxiv.org/abs/2512.02523
备注:16 pages, 5 figures
摘要:歌唱声音合成(SVS)已经取得了显著的进步,使模型能够生成具有准确音高和一致风格的人声。随着这些能力的提高,对可靠评估和优化的需求变得越来越重要。然而,目前的方法,如奖励系统,往往依赖于单一的数字分数,努力捕捉各种维度,如措辞或表达,并需要昂贵的注释,限制了可解释性和泛化。为了解决这些问题,我们提出了一种生成反馈(即,奖励模型)框架,为SVS评估提供多维语言和音频反馈。我们的方法利用音频语言模型来生成文本和音频评论-涵盖旋律、内容和听觉质量等方面。该模型在混合数据集上进行了微调,该数据集结合了人类音乐反应和MLLM的合成评论,增强了多样性和语言丰富性。定量实验验证了所提出的数据集和训练策略的有效性,表明该框架产生了适合指导生成模型改进的音乐准确和可解释的评价。代码在[https://github.com/opendilab/VocalCritic](https://github.com/opendilab/VocalCritic)
摘要:Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly critical. However, current methods like reward systems often rely on single numerical scores, struggle to capture various dimensions such as phrasing or expressiveness, and require costly annotations, limiting interpretability and generalization. To address these issues, we propose a generative feedback (i.e., reward model) framework that provides multi-dimensional language and audio feedback for SVS assessment. Our approach leverages an audio-language model to generate text and audio critiques-covering aspects such as melody, content, and auditory quality. The model is fine-tuned on a hybrid dataset combining human music reactions and synthetic critiques from a MLLMs, enhancing diversity and linguistic richness. Quantitative experiments validate the effectiveness of the proposed dataset and training strategy, demonstrating that the framework produces musically accurate and interpretable evaluations suitable for guiding generative model improvement. The code is at [https://github.com/opendilab/VocalCritic](https://github.com/opendilab/VocalCritic)
【7】VibOmni: Towards Scalable Bone-conduction Speech Enhancement on Earables
标题:VibOmni:迈向可扩展的Early骨导语音增强
链接:https://arxiv.org/abs/2512.02515
备注:Submitted to TMC
摘要:True Wireless Stereo耳机和VR/AR耳机等可穿戴设备越来越受欢迎,但其紧凑的设计对噪声环境中的电信和语音助理交互等强大的语音相关应用提出了挑战。现有的语音增强系统仅依赖于全向麦克风,与竞争扬声器一样与环境噪声作斗争。为了解决这些问题,我们提出了VibOmni,一个轻量级的,端到端的多模态语音增强系统,利用广泛使用的惯性测量单元(IMU)捕获的骨传导振动。VibOmni集成了一个双分支编码器-解码器深度神经网络,以融合音频和振动特征。为了克服配对音频振动数据集的稀缺性,我们引入了一种新的数据增强技术,该技术从有限的记录中对骨传导功能(BCF)进行建模,使合成振动数据生成仅具有4.5%的频谱图相似性误差。此外,多模态SNR估计器有助于持续学习和自适应推理,在动态噪声环境中优化性能,而无需设备上的反向传播。通过对32名志愿者使用不同设备的真实数据集进行评估,VibOmni在语音质量感知评估(PESQ)方面提高了21%,信噪比(SNR)提高了26%,WER降低了约40%,并且在移动设备上的延迟更少。一项有35名参与者的用户研究显示,87%的人更喜欢VibOmni,这证明了它在不同声学环境中的有效性。
摘要:Earables, such as True Wireless Stereo earphones and VR/AR headsets, are increasingly popular, yet their compact design poses challenges for robust voice-related applications like telecommunication and voice assistant interactions in noisy environments. Existing speech enhancement systems, reliant solely on omnidirectional microphones, struggle with ambient noise like competing speakers. To address these issues, we propose VibOmni, a lightweight, end-to-end multi-modal speech enhancement system for earables that leverages bone-conducted vibrations captured by widely available Inertial Measurement Units (IMUs). VibOmni integrates a two-branch encoder-decoder deep neural network to fuse audio and vibration features. To overcome the scarcity of paired audio-vibration datasets, we introduce a novel data augmentation technique that models Bone Conduction Functions (BCFs) from limited recordings, enabling synthetic vibration data generation with only 4.5% spectrogram similarity error. Additionally, a multi-modal SNR estimator facilitates continual learning and adaptive inference, optimizing performance in dynamic, noisy settings without on-device back-propagation. Evaluated on real-world datasets from 32 volunteers with different devices, VibOmni achieves up to 21% improvement in Perceptual Evaluation of Speech Quality (PESQ), 26% in Signal-to-Noise Ratio (SNR), and about 40% WER reduction with much less latency on mobile devices. A user study with 35 participants showed 87% preferred VibOmni over baselines, demonstrating its effectiveness for deployment in diverse acoustic environments.
【8】Continual Learning for Singing Voice Separation with Human in the Loop Adaptation
标题:通过人类在环适应来实现歌唱声音分离的持续学习
链接:https://arxiv.org/abs/2512.02432
备注:Proceedings of the 26th International Symposium on Frontiers of Research in Speech and Music, 2021
摘要:基于深度学习的歌唱声音分离工作在最近的过去表现得非常好。然而,这些作品中的大多数并没有关注允许用户与模型交互以提高性能。在现实世界中部署模型时,这一点至关重要,因为在现实世界中,音乐曲目可能与流派和乐器的原始训练数据不同。在本文中,我们提出了一个基于深度学习的交互式持续学习框架,用于歌声分离,允许用户微调声乐分离模型,使其符合新的目标歌曲。我们使用基于U-Net的基础模型架构,该架构产生用于将人声与声谱图分离的掩模,然后是人机交互任务,其中用户通过标记一些误报来提供反馈,即,提取的人声中应该是无声的区域。我们提出了两个持续学习算法。实验证明,在数据集内和数据集间的设置,所提出的算法在歌唱声分离性能的改善基础模型。
摘要:Deep learning-based works for singing voice separation have performed exceptionally well in the recent past. However, most of these works do not focus on allowing users to interact with the model to improve performance. This can be crucial when deploying the model in real-world scenarios where music tracks can vary from the original training data in both genre and instruments. In this paper, we present a deep learning-based interactive continual learning framework for singing voice separation that allows users to fine-tune the vocal separation model to conform it to new target songs. We use a U-Net-based base model architecture that produces a mask for separating vocals from the spectrogram, followed by a human-in-the-loop task where the user provides feedback by marking a few false positives, i.e., regions in the extracted vocals that should have been silence. We propose two continual learning algorithms. Experiments substantiate the improvement in singing voice separation performance by the proposed algorithms over the base model in intra-dataset and inter-dataset settings.
【9】WhAM: Towards A Translative Model of Sperm Whale Vocalization
标题:WhAM:迈向抹香鲸发声的翻译模型
链接:https://arxiv.org/abs/2512.02206
备注:NeurIPS 2025
摘要:抹香鲸通过一系列短的咔哒声进行交流,这些咔哒声被称为尾音。我们提出了WhAM(鲸鱼声学模型),第一个基于变压器的模型,能够从任何音频提示生成合成抹香鲸尾。WhAM是通过微调VampNet构建的,VampNet是一个在音乐音频上预训练的掩蔽声学令牌模型,使用了过去二十年收集的10k尾波录音。通过迭代掩蔽标记预测,WhAM生成高保真合成尾波,保留源录音的关键声学特征。我们使用Fréchet Audio Distance并通过与专家海洋生物学家的感知研究来评估WhAM的合成尾波。在包括节奏、社会单位和元音分类在内的下游分类任务中,WhAM的学习表征实现了强大的性能,尽管它是针对生成而不是分类进行训练的。我们的代码可在https://github.com/Project-CETI/wham上获得
摘要:Sperm whales communicate in short sequences of clicks known as codas. We present WhAM (Whale Acoustics Model), the first transformer-based model capable of generating synthetic sperm whale codas from any audio prompt. WhAM is built by finetuning VampNet, a masked acoustic token model pretrained on musical audio, using 10k coda recordings collected over the past two decades. Through iterative masked token prediction, WhAM generates high-fidelity synthetic codas that preserve key acoustic features of the source recordings. We evaluate WhAM's synthetic codas using Fréchet Audio Distance and through perceptual studies with expert marine biologists. On downstream classification tasks including rhythm, social unit, and vowel classification, WhAM's learned representations achieve strong performance, despite being trained for generation rather than classification. Our code is available at https://github.com/Project-CETI/wham
【10】Story2MIDI: Emotionally Aligned Music Generation from Text
标题:Story 2 Buttons:从文本生成情感一致的音乐
链接:https://arxiv.org/abs/2512.02192
备注:8 pages (6 pages of main text + 2 pages of references and appendices), 4 figures, 1 table. Presented at IEEE Big Data 2025 3rd Workshop on AI Music Generation (AIMG 2025)
摘要:在本文中,我们介绍了Story2Side,一个基于序列到序列转换器的模型,用于从给定的文本中生成情感对齐的音乐。为了开发这个模型,我们通过合并现有的数据集来构建Story2数据集,用于从文本和音乐中的情感分类进行情感分析。生成的数据集包含成对的文本简介和音乐片段,它们在读者或听众中唤起相同的情感。尽管我们的数据集规模小,计算资源有限,但我们的结果表明,我们的模型有效地学习了音乐中与情感相关的特征,并将其纳入其生成过程,生成了具有不同情感反应的样本。我们使用客观的音乐指标和人类听力研究来评估生成的输出,确认该模型捕捉预期情感线索的能力。
摘要:In this paper, we introduce Story2MIDI, a sequence-to-sequence Transformer-based model for generating emotion-aligned music from a given piece of text. To develop this model, we construct the Story2MIDI dataset by merging existing datasets for sentiment analysis from text and emotion classification in music. The resulting dataset contains pairs of text blurbs and music pieces that evoke the same emotions in the reader or listener. Despite the small scale of our dataset and limited computational resources, our results indicate that our model effectively learns emotion-relevant features in music and incorporates them into its generation process, producing samples with diverse emotional responses. We evaluate the generated outputs using objective musical metrics and a human listening study, confirming the model's ability to capture intended emotional cues.
【11】Dialect Identification Using Resource-Efficient Fine-Tuning Approaches
标题:使用资源高效的微调方法进行方言识别
链接:https://arxiv.org/abs/2512.02074
备注:Published in APSIPA ASC 2025
摘要:方言识别是从语音信号中识别出同一语言中不同方言的任务。DI可以帮助改善下游语音相关的任务,即使说话者有很强的方言。然而,微调语音模型的任务,如DI是昂贵的计算成本和内存需求方面。最近的研究探索了使用参数高效微调(PEFT)方法微调预训练语音模型,用于DI等任务,该方法提供了参数效率,但在内存效率和训练速度方面的改进有限。为了解决这些挑战,我们探索了最初为语言处理提出的内存高效微调(MEFT)方法,并将其应用于通用的预训练语音模型。然后,我们综合分析了GPU的内存使用和微调速度的基础上各种MEFT方法。作为一个案例研究,我们微调Whisper模型,从KeSpeech数据集中识别六种普通话方言,将GPU内存使用量减少了73.25%,并将训练速度提高了2.1倍,同时保持了与普通微调和PEFT方法相当的准确性。
摘要:Dialect Identification (DI) is a task to recognize different dialects within the same language from a speech signal. DI can help to improve the downstream speech related tasks even when speakers have a strong dialect. However, fine-tuning a speech model for tasks like DI is expensive in terms of computation cost and memory requirement. Recent studies have explored fine-tuning pre-trained speech models for tasks like DI using Parameter-Efficient Fine-Tuning (PEFT) methods, which offer parameter efficiency but limited improvement in memory efficiency and training speed. To address these challenges, we explore Memory-Efficient Fine-Tuning (MEFT) methods, originally proposed for language processing, and apply them to the general-purpose pre-trained speech model. We then comprehensively analyze the GPU memory usage and fine-tuning speed based on various MEFT methods. As a case study, we fine-tune the Whisper model to identify six Mandarin subdialects from the KeSpeech dataset, reducing GPU memory usage by up to 73.25% and accelerating training speed by a factor of 2.1, while maintaining accuracy comparable to vanilla fine-tuning and PEFT methods.
【12】Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
标题:利用多模式基础模型实现与年龄无关的面部声音关联
链接:https://arxiv.org/abs/2512.02759
备注:This paper presents the system description of the UZH-CL team for the FAME2026 Challenge at ICASSP 2026. Our model achieved second place in the final ranking
摘要:本文介绍了提交给FAME 2026挑战赛的UZH-CL系统。挑战的重点是在独特的多语言条件下进行跨模态验证,特别是看不见和听不见的语言。我们的方法研究了两种不同的架构,包括使用对比和正交投影损失从头开始训练的基线双编码器系统,以及利用ImageBind和LoRA的基础模型方法。为了解决数据稀缺和语言限制的挑战,我们从VoxBlink策划了一个外部阿拉伯语数据集。我们表现最好的系统ImageBind-LoRA展示了出色的跨语言泛化能力:尽管仅针对阿拉伯语音频进行了微调,但它在评估集(英语和德语)上的EER达到了24.73%,在比赛中获得了第二名。
摘要:This paper describes the UZH-CL system submitted to the FAME2026 Challenge. The challenge focuses on cross-modal verification under unique multilingual conditions, specifically unseen and unheard languages. Our approach investigates two distinct architectures, consisting of a baseline dual-encoder system trained from scratch using contrastive and orthogonal projection losses, and a foundation model approach leveraging ImageBind with LoRA. To address the data scarcity and language constraints of the challenge, we curated an external Arabic dataset from VoxBlink. Our best-performing system, ImageBind-LoRA, demonstrates remarkable cross-lingual generalization: despite being fine-tuned exclusively on Arabic audio, it achieved an EER of 24.73% on the evaluation set (English and German), securing 2nd place in the competition.
【1】Perceptual evaluation of Acoustic Level of Detail in Virtual Acoustic Environments
标题:虚拟声学环境中声学细节水平的感知评估
链接:https://arxiv.org/abs/2512.02891
备注:This work has been submitted to Acoustics for possible publication. Template provided by MDPI
摘要:虚拟声学环境使听觉研究和听力学至关重要的现实和生态有效的日常生活情况的创建和模拟。室内混响环境尤为重要。对于实时应用,房间声学仿真需要简化,然而,为了捕获所有感知相关的效果,必要的声学细节水平(ALOD)仍然不清楚。本研究探讨了不同的ALOD在模拟三个真实环境的影响:一个客厅与耦合的厨房,一个酒吧,和一个地铁站。通过生成不同数量的早期反射图像源,或通过排除每个环境特定的几何房间细节,ALOD是不同的。使用耳机与在相应的真实环境中使用假人头部或使用扬声器测量的双耳房间脉冲响应进行比较,对模拟进行感知评估。该研究评估了脉冲刺激,播放电贝司和语音令牌的感知总体差异。此外,对可解释性、语言清晰度和外化进行了评价。结果表明,一个强大的减少ALOD是可行的,同时保持类似的可扩展性,语音清晰度,和外部化与假人头部录音。早期反射的数量和精度似乎不太相关,提供了适当的漫反射后期混响。
摘要:Virtual acoustic environments enable the creation and simulation of realistic and eco-logically valid daily-life situations vital for hearing research and audiology. Reverberant indoor environments are particularly important. For real-time applications, room acous-tics simulation requires simplifications, however, the necessary acoustic level of detail (ALOD) remains unclear in order to capture all perceptually relevant effects. This study examines the impact of varying ALOD in simulations of three real environments: a living room with a coupled kitchen, a pub, and an underground station. ALOD was varied by generating different numbers of image sources for early reflections, or by excluding geo-metrical room details specific for each environment. Simulations were perceptually eval-uated using headphones in comparison to binaural room impulse responses measured with a dummy head in the corresponding real environments, or by using loudspeakers. The study assessed the perceived overall difference for a pulse stimulus, a played electric bass and a speech token. Additionally, plausibility, speech intelligibility, and externaliza-tion were evaluated. Results indicate that a strong reduction in ALOD is feasible while maintaining similar plausibility, speech intelligibility, and externalization as with dummy head recordings. The number and accuracy of early reflections appear less relevant, pro-vided diffuse late reverberation is appropriately represented.
【2】Towards Language-Independent Face-Voice Association with Multimodal Foundation Models
标题:利用多模式基础模型实现与年龄无关的面部声音关联
链接:https://arxiv.org/abs/2512.02759
备注:This paper presents the system description of the UZH-CL team for the FAME2026 Challenge at ICASSP 2026. Our model achieved second place in the final ranking
摘要:本文介绍了提交给FAME 2026挑战赛的UZH-CL系统。挑战的重点是在独特的多语言条件下进行跨模态验证,特别是看不见和听不见的语言。我们的方法研究了两种不同的架构,包括使用对比和正交投影损失从头开始训练的基线双编码器系统,以及利用ImageBind和LoRA的基础模型方法。为了解决数据稀缺和语言限制的挑战,我们从VoxBlink策划了一个外部阿拉伯语数据集。我们表现最好的系统ImageBind-LoRA展示了出色的跨语言泛化能力:尽管仅针对阿拉伯语音频进行了微调,但它在评估集(英语和德语)上的EER达到了24.73%,在比赛中获得了第二名。
摘要:This paper describes the UZH-CL system submitted to the FAME2026 Challenge. The challenge focuses on cross-modal verification under unique multilingual conditions, specifically unseen and unheard languages. Our approach investigates two distinct architectures, consisting of a baseline dual-encoder system trained from scratch using contrastive and orthogonal projection losses, and a foundation model approach leveraging ImageBind with LoRA. To address the data scarcity and language constraints of the challenge, we curated an external Arabic dataset from VoxBlink. Our best-performing system, ImageBind-LoRA, demonstrates remarkable cross-lingual generalization: despite being fine-tuned exclusively on Arabic audio, it achieved an EER of 24.73% on the evaluation set (English and German), securing 2nd place in the competition.
【3】On the Difficulty of Token-Level Modeling of Dysfluency and Fluency Shaping Artifacts
标题:关于流畅性障碍和流畅性塑造制品的代币级建模的困难
链接:https://arxiv.org/abs/2512.02027
备注:6 pages, 1 figure. Accepted to ASRU 2025. This is the arXiv preprint of the accepted paper
摘要:即使对于现代端到端(E2 E)自动语音识别(ASR)框架,口吃语音的自动转录仍然是一个挑战。不流畅和流畅性形成伪影往往被忽视,导致非逐字翻译,临床和研究价值有限。我们提出了一个参数有效的适应方法来解码不流畅和流畅的修改,作为特殊的令牌内transmittance,评估模拟(LibriStutter,英语)和自然(KSoF,德国)口吃语音数据集。为了减轻ASR表现差异和对英语的偏见,我们引入了一个多步微调策略与语言自适应预训练。标记化分析进一步突出了标记器以英语为中心的偏见,这对提高德语数据的性能提出了挑战。我们的研究结果表明,轻量级的适应技术的有效性dysfluency-aware ASR,同时暴露出多语言E2 E系统的关键限制。
摘要:Automatic transcription of stuttered speech remains a challenge, even for modern end-to-end (E2E) automatic speech recognition (ASR) frameworks. Dysfluencies and fluency-shaping artifacts are often overlooked, resulting in non-verbatim transcriptions with limited clinical and research value. We propose a parameter-efficient adaptation method to decode dysfluencies and fluency modifications as special tokens within transcriptions, evaluated on simulated (LibriStutter, English) and natural (KSoF, German) stuttered speech datasets. To mitigate ASR performance disparities and bias towards English, we introduce a multi-step fine-tuning strategy with language-adaptive pretraining. Tokenization analysis further highlights the tokenizer's English-centric bias, which poses challenges for improving performance on German data. Our findings demonstrate the effectiveness of lightweight adaptation techniques for dysfluency-aware ASR while exposing key limitations in multilingual E2E systems.
【4】Hear What Matters! Text-conditioned Selective Video-to-Audio Generation
标题:听听什么重要!文本条件选择性视频到音频生成
链接:https://arxiv.org/abs/2512.02650
摘要:这项工作介绍了一个新的任务,文本条件下的选择性视频到音频(V2 A)的生成,它只产生用户想要的声音从多对象视频。这种能力在多媒体制作中尤其重要,因为每个声源的音轨都要单独处理,以便进行精确的编辑、混音和创意控制。然而,当前的方法一次生成单一来源混合的声音,这主要是因为视觉特征相互纠缠,并且区域提示或提示通常无法指定来源。我们提出了SelVA,一种新的文本条件V2 A模型,将文本提示作为目标源的显式选择器,并调制视频编码器以明显地提取与文本相关的视频特征。建议的补充令牌促进交叉注意,抑制文本无关的激活与有效的参数调整,产生强大的语义和时间接地。SelVA进一步采用自增强方案来克服单声道音轨监督的缺乏。我们在VGG-MONOAUDIO上评估SelVA,VGG-MONOAUDIO是一个针对此类任务的干净单源视频的策划基准。大量的实验和消融一致地验证了其在音频质量、语义对齐和时间同步方面的有效性。代码和演示可在https://jnwnlee.github.io/selva-demo/上获得。
摘要:This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio tracks are handled individually for each sound source for precise editing, mixing, and creative control. However, current approaches generate single source-mixed sounds at once, largely because visual features are entangled, and region cues or prompts often fail to specify the source. We propose SelVA, a novel text-conditioned V2A model that treats the text prompt as an explicit selector of target source and modulates video encoder to distinctly extract prompt-relevant video features. The proposed supplementary tokens promote cross-attention by suppressing text-irrelevant activations with efficient parameter tuning, yielding robust semantic and temporal grounding. SelVA further employs a self-augmentation scheme to overcome the lack of mono audio track supervision. We evaluate SelVA on VGG-MONOAUDIO, a curated benchmark of clean single-source videos for such a task. Extensive experiments and ablations consistently verify its effectiveness across audio quality, semantic alignment, and temporal synchronization. Code and demo are available at https://jnwnlee.github.io/selva-demo/.
【5】Spoken Conversational Agents with Large Language Models
标题:具有大型语言模型的口语对话代理
链接:https://arxiv.org/abs/2512.02593
备注:Accepted to EMNLP 2025 Tutorial
摘要:口语会话代理正在向语音本地LLM融合。本教程将从级联ASR/NLU到端到端、基于检索和视觉的系统的路径进行提炼。我们将文本LLM框架化为音频,跨模态对齐和联合语音文本训练;审查数据集,指标和跨口音的鲁棒性,并比较设计选择(级联与E2 E,ASR后校正,流)。我们将工业助理与当前的开放领域和面向任务的代理联系起来,突出可重复的基线,并概述隐私,安全和评估方面的开放问题。与会者带着实用的配方和清晰的系统级路线图离开。
摘要:Spoken conversational agents are converging toward voice-native LLMs. This tutorial distills the path from cascaded ASR/NLU to end-to-end, retrieval-and vision-grounded systems. We frame adaptation of text LLMs to audio, cross-modal alignment, and joint speech-text training; review datasets, metrics, and robustness across accents and compare design choices (cascaded vs. E2E, post-ASR correction, streaming). We link industrial assistants to current open-domain and task-oriented agents, highlight reproducible baselines, and outline open problems in privacy, safety, and evaluation. Attendees leave with practical recipes and a clear systems-level roadmap.
机器翻译由腾讯交互翻译提供,仅供参考
