今日论文合集:cs.SD语音14篇,eess.AS音频处理8篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Contextualized Token Discrimination for Speech Search Query Correction
标题:用于语音搜索查询纠正的上下文化令牌辨别
链接:http://arxiv.org/pdf/2509.04393v1

作者:Junyu Lu, Di Jiang, Mengze Hong, Victor Junqiu Wei, Qintian Guo, Zhiyang Su
摘要:查询词拼写纠正是现代搜索引擎的一个重要功能,它可以有效地帮助用户清楚地表达自己的意图。随着自动语音识别(ASR)系统驱动的语音搜索的日益普及,本文提出了一种新的方法--语境化标记识别(CTD),用于进行有效的语音查询校正。在CTD中,我们首先使用BERT来生成令牌级的上下文表示,然后构造一个组合层来增强语义信息。最后,我们根据聚合的令牌表示产生正确的查询,通过比较原始令牌表示和上下文表示来纠正不正确的令牌。大量的实验表明,我们提出的方法在所有指标的优越性能,我们进一步提出了一个新的基准数据集与错误的ASR transmittance提供全面的评估音频查询校正。摘要:Query spelling correction is an important function of modern search engines since it effectively helps users express their intentions clearly. With the growing popularity of speech search driven by Automated Speech Recognition (ASR) systems, this paper introduces a novel method named Contextualized Token Discrimination (CTD) to conduct effective speech query correction. In CTD, we first employ BERT to generate token-level contextualized representations and then construct a composition layer to enhance semantic information. Finally, we produce the correct query according to the aggregated token representation, correcting the incorrect tokens by comparing the original token representations and the contextualized representations. Extensive experiments demonstrate the superior performance of our proposed method across all metrics, and we further present a new benchmark dataset with erroneous ASR transcriptions to offer comprehensive evaluations for audio query correction.


【2】Denoising GER: A Noise-Robust Generative Error Correction with LLM for Speech Recognition
标题:去噪GER:用于语音识别的LLM噪音鲁棒生成式错误纠正
链接:http://arxiv.org/pdf/2509.04392v1

作者:Yanyan Liu, Minqiang Xu, Yihao Chen, Liang He, Lei Fang, Sian Fang, Lin Liu
摘要:近年来,大语言模型(LLM)在自动语音识别(ASR)后处理的生成纠错(GER)任务中取得了重大进展。然而,在复杂的噪声环境中,它们仍然面临着适应性差、信息利用率低等挑战,导致GER的有效性有限。针对这些问题,提出了一种噪声鲁棒的多模态GER框架(Denoising GER)。该框架通过噪声自适应声学编码器增强了模型对不同噪声场景的适应性,并通过异构特征补偿动态融合(HFCDF)机制优化了多模态信息的集成,提高了LLM对多模态信息的利用率。此外,还引入了强化学习(RL)训练策略来增强模型的预测能力。实验结果表明,去噪GER显着提高精度和鲁棒性在噪声环境中,并表现出良好的泛化能力,在看不见的噪声场景。摘要:In recent years, large language models (LLM) have made significant progress in the task of generation error correction (GER) for automatic speech recognition (ASR) post-processing. However, in complex noisy environments, they still face challenges such as poor adaptability and low information utilization, resulting in limited effectiveness of GER. To address these issues, this paper proposes a noise-robust multi-modal GER framework (Denoising GER). The framework enhances the model's adaptability to different noisy scenarios through a noise-adaptive acoustic encoder and optimizes the integration of multi-modal information via a heterogeneous feature compensation dynamic fusion (HFCDF) mechanism, improving the LLM's utilization of multi-modal information. Additionally, reinforcement learning (RL) training strategies are introduced to enhance the model's predictive capabilities. Experimental results demonstrate that Denoising GER significantly improves accuracy and robustness in noisy environments and exhibits good generalization abilities in unseen noise scenarios.


【3】PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
标题:PARCO:通过对比实体歧义消除音素增强的鲁棒上下文ASB
链接:http://arxiv.org/pdf/2509.04357v1

作者:Jiajun He, Naoki Sawada, Koichi Miyazaki, Tomoki Toda

备注:Accepted by ASRU 2025

摘要:自动语音识别(ASR)系统与特定领域的命名实体,特别是同音异义词斗争。上下文ASR提高了识别,但由于有限的实体多样性,通常无法捕获细粒度的音素变化。此外,现有方法将实体视为独立的令牌,导致不完全的多令牌偏置。为了解决这些问题,我们提出了音素增强的鲁棒上下文ASR通过对比实体消歧(PARCO),它集成了音素感知编码,对比实体消歧,实体级监督和分层实体过滤。这些组件增强了语音识别,确保完整的实体检索,并减少不确定性下的误报。实验结果表明,PARCO在1,000个干扰项下对中文AISHELL-1的CER为4.22%,对英文DATA 2的WER为11.14%,显著优于基线。PARCO还在域外数据集(如THCHS-30和LibriSpeech)上展示了强大的增益。摘要:Automatic speech recognition (ASR) systems struggle with domain-specific named entities, especially homophones. Contextual ASR improves recognition but often fails to capture fine-grained phoneme variations due to limited entity diversity. Moreover, prior methods treat entities as independent tokens, leading to incomplete multi-token biasing. To address these issues, we propose Phoneme-Augmented Robust Contextual ASR via COntrastive entity disambiguation (PARCO), which integrates phoneme-aware encoding, contrastive entity disambiguation, entity-level supervision, and hierarchical entity filtering. These components enhance phonetic discrimination, ensure complete entity retrieval, and reduce false positives under uncertainty. Experiments show that PARCO achieves CER of 4.22% on Chinese AISHELL-1 and WER of 11.14% on English DATA2 under 1,000 distractors, significantly outperforming baselines. PARCO also demonstrates robust gains on out-of-domain datasets like THCHS-30 and LibriSpeech.


【4】AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
标题:AUDETER:开放世界中用于Deepfake音频检测的大规模数据集
链接:http://arxiv.org/pdf/2509.04345v1

作者:Qizhou Wang, Hanxun Huang, Guansong Pang, Sarah Erfani, Christopher Leckie
摘要:语音生成系统可以产生非常逼真的发声,这些发声通常与人类语音无法区分,这对真实性提出了重大挑战。尽管已经开发了许多deepfake检测方法,但由于不同的人类语音和快速发展的语音合成系统产生的训练和测试样本之间的域转移,它们在现实环境中的有效性仍然是不可靠的。目前的数据集没有充分解决这一问题,这些数据集缺乏真实世界的应用挑战,在真实和深度伪造类别中都有多样化和最新的音频。为了填补这一空白,我们引入了AUDETER(AUdio DEepfake TEst Range),这是一个大规模、高度多样化的deepfake音频数据集,用于全面评估和稳健开发deepfake音频检测的通用模型。它由11个最新的TTS模型和10个具有广泛TTS 声码器模式的声码器生成的超过4,500小时的合成音频组成,总计300万个音频片段,使其成为规模最大的deepfake音频数据集。通过对AUDETER的广泛实验,我们发现:i)在现有数据集上训练的最先进(SOTA)方法很难推广到新的deepfake音频样本,并且在看不见的人类声音上存在很高的误报率,这强调了对全面数据集的需求;在AUDETER上训练的方法具有很高的泛化检测性能,检测错误率降低了44.1%到51.6%,在流行的In-the-Wild数据集中,不同的跨域样本的错误率仅为4.17%,为训练通用的deepfake音频检测器铺平了道路。AUDETER在GitHub上可用。摘要:Speech generation systems can produce remarkably realistic vocalisations that are often indistinguishable from human speech, posing significant authenticity challenges. Although numerous deepfake detection methods have been developed, their effectiveness in real-world environments remains unrealiable due to the domain shift between training and test samples arising from diverse human speech and fast evolving speech synthesis systems. This is not adequately addressed by current datasets, which lack real-world application challenges with diverse and up-to-date audios in both real and deep-fake categories. To fill this gap, we introduce AUDETER (AUdio DEepfake TEst Range), a large-scale, highly diverse deepfake audio dataset for comprehensive evaluation and robust development of generalised models for deepfake audio detection. It consists of over 4,500 hours of synthetic audio generated by 11 recent TTS models and 10 vocoders with a broad range of TTS vocoder patterns, totalling 3 million audio clips, making it the largest deepfake audio dataset by scale. Through extensive experiments with AUDETER, we reveal that i) state-of-the-art (SOTA) methods trained on existing datasets struggle to generalise to novel deepfake audio samples and suffer from high false positive rates on unseen human voice, underscoring the need for a comprehensive dataset; and ii) these methods trained on AUDETER achieve highly generalised detection performance and significantly reduce detection error rate by 44.1% to 51.6%, achieving an error rate of only 4.17% on diverse cross-domain samples in the popular In-the-Wild dataset, paving the way for training generalist deepfake audio detectors. AUDETER is available on GitHub.


【5】PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
标题:PianoBind:一个流行钢琴音乐的多模态联合嵌入模型
链接:http://arxiv.org/pdf/2509.04215v1

作者:Hayeon Bang, Eunjin Choi, Seungheon Doh, Juhan Nam
备注:Accepted for publication at the 26th International Society for Music Information Retrieval Conference (ISMIR 2025)
摘要:钢琴独奏音乐,尽管是一种单一的乐器媒介,具有显着的表达能力,传达丰富的语义信息,跨流派,情绪和风格。然而,目前的通用音乐表示模型,主要是在大规模数据集上训练的,往往难以捕捉同质独奏钢琴音乐中的微妙语义差异。此外,现有的钢琴特定的表示模型通常是单峰的,无法捕捉钢琴音乐的固有多模态性质,通过音频,符号和文本形式表达。为了解决这些限制,我们提出了PianoBind,一个钢琴特定的多模态联合嵌入模型。我们系统地研究了多源训练和模态利用的策略,该策略在一个联合嵌入框架内进行了优化,用于捕获(1)小规模和(2)同质钢琴数据集中的细粒度语义差异。我们的实验结果表明,PianoBind学习多模态表示,有效地捕捉钢琴音乐的细微差别,与通用音乐联合嵌入模型相比,在域内和域外钢琴数据集上实现了卓越的文本到音乐检索性能。此外,我们的设计选择为钢琴音乐之外的同质数据集的多模态表示学习提供了可重用的见解。摘要:Solo piano music, despite being a single-instrument medium, possesses significant expressive capabilities, conveying rich semantic information across genres, moods, and styles. However, current general-purpose music representation models, predominantly trained on large-scale datasets, often struggle to captures subtle semantic distinctions within homogeneous solo piano music. Furthermore, existing piano-specific representation models are typically unimodal, failing to capture the inherently multimodal nature of piano music, expressed through audio, symbolic, and textual modalities. To address these limitations, we propose PianoBind, a piano-specific multimodal joint embedding model. We systematically investigate strategies for multi-source training and modality utilization within a joint embedding framework optimized for capturing fine-grained semantic distinctions in (1) small-scale and (2) homogeneous piano datasets. Our experimental results demonstrate that PianoBind learns multimodal representations that effectively capture subtle nuances of piano music, achieving superior text-to-music retrieval performance on in-domain and out-of-domain piano datasets compared to general-purpose music joint embedding models. Moreover, our design choices offer reusable insights for multimodal representation learning with homogeneous datasets beyond piano music.


【6】Crossing the Species Divide: Transfer Learning from Speech to Animal Sounds
标题:跨越物种鸿沟:从语音到动物声音的迁移学习
链接:http://arxiv.org/pdf/2509.04166v1

作者:Jules Cauzinille, Marius Miron, Olivier Pietquin, Masato Hagiwara, Ricard Marxer, Arnaud Rey, Benoit Favre
备注:5 pages, 3 figures, uses this http URL, submitted to DCASE 2025
摘要:自监督语音模型在语音处理中表现出了令人印象深刻的性能,但它们对非语音数据的有效性仍然有待探索。我们研究了这种模型在生物声学检测和分类任务上的迁移学习能力。我们发现,HuBERT,WavLM和XEUS等模型可以生成丰富的动物声音的潜在表示。我们分析了模型的性质与线性探测的时间平均表示。然后,我们扩展的方法来考虑与其他下游架构的时间明智的信息的影响。最后,我们研究了频率范围和噪声对性能的影响。值得注意的是,我们的结果与微调的生物声学预训练模型具有竞争力,并显示了噪声鲁棒性预训练设置的影响。这些发现突出了基于语音的自我监督学习作为推进生物声学研究的有效框架的潜力。摘要:Self-supervised speech models have demonstrated impressive performance in speech processing, but their effectiveness on non-speech data remains underexplored. We study the transfer learning capabilities of such models on bioacoustic detection and classification tasks. We show that models such as HuBERT, WavLM, and XEUS can generate rich latent representations of animal sounds across taxa. We analyze the models properties with linear probing on time-averaged representations. We then extend the approach to account for the effect of time-wise information with other downstream architectures. Finally, we study the implication of frequency range and noise on performance. Notably, our results are competitive with fine-tuned bioacoustic pre-trained models and show the impact of noise-robust pre-training setups. These findings highlight the potential of speech-based self-supervised learning as an efficient framework for advancing bioacoustic research.


【7】Wav2DF-TSL: Two-stage Learning with Efficient Pre-training and Hierarchical Experts Fusion for Robust Audio Deepfake Detection
标题:Wave 2DF-TSL:具有高效预训练和分层专家融合的两阶段学习,用于鲁棒的音频深度伪造检测
链接:http://arxiv.org/pdf/2509.04161v1

作者:Yunqi Hao, Yihao Chen, Minqiang Xu, Jianbo Zhan, Liang He, Lei Fang, Sian Fang, Lin Liu
摘要:近年来,自监督学习(SSL)模型在音频深度伪造检测(ADD)任务中取得了重大进展。然而,现有的SSL模型主要依赖于大规模的真实语音进行预训练,缺乏对欺骗样本的学习,这导致在ADD任务的微调过程中容易受到领域偏差的影响。为此,我们提出了一种基于预训练和分层专家融合的两阶段学习策略(Wav 2DF-TSL),用于鲁棒的音频深度伪造检测。在预训练阶段,我们使用适配器从3000小时的未标记欺骗语音中有效地学习工件,提高前端功能的适应性,同时减轻灾难性遗忘。在微调阶段,我们提出了分层自适应混合专家(HA-MoE)的方法,动态融合多级欺骗线索,通过多专家合作与门控路由。实验结果表明,该方法在所有四个基准数据集上的性能都明显优于基线系统,尤其是在跨域In-the-wild数据集上,等错误率(EER)相对提高了27.5%,优于现有的最先进的系统。索引术语:音频deepfake检测,自监督学习,参数高效微调,专家混合摘要:In recent years, self-supervised learning (SSL) models have made significant progress in audio deepfake detection (ADD) tasks. However, existing SSL models mainly rely on large-scale real speech for pre-training and lack the learning of spoofed samples, which leads to susceptibility to domain bias during the fine-tuning process of the ADD task. To this end, we propose a two-stage learning strategy (Wav2DF-TSL) based on pre-training and hierarchical expert fusion for robust audio deepfake detection. In the pre-training stage, we use adapters to efficiently learn artifacts from 3000 hours of unlabelled spoofed speech, improving the adaptability of front-end features while mitigating catastrophic forgetting. In the fine-tuning stage, we propose the hierarchical adaptive mixture of experts (HA-MoE) method to dynamically fuse multi-level spoofing cues through multi-expert collaboration with gated routing. Experimental results show that the proposed method significantly outperforms the baseline system on all four benchmark datasets, especially on the cross-domain In-the-wild dataset, achieving a 27.5% relative improvement in equal error rate (EER), outperforming the existing state-of-the-art systems. Index Terms: audio deepfake detection, self-supervised learning, parameter-efficient fine-tuning, mixture of experts


【8】Enhancing Self-Supervised Speaker Verification Using Similarity-Connected Graphs and GCN
标题:使用相似性连接图和GCN增强自我监督说话人验证
链接:http://arxiv.org/pdf/2509.04147v1

作者:Zhaorui Sun, Yihao Chen, Jialong Wang, Minqiang Xu, Lei Fang, Sian Fang, Lin Liu
摘要:随着语音识别技术的不断发展,说话人确认已经成为身份认证的重要手段。传统的SV方法依赖于手工特征提取,而深度学习显著提高了系统性能。然而,标记数据的稀缺性仍然限制了深度学习在SV中的广泛应用。自监督学习通过挖掘大型未标记数据集中的潜在信息,增强了模型的泛化能力,是解决这一问题的关键技术。 DINO是一种高效的自监督学习方法,通过聚类从未标记的语音数据中生成伪标签,支持后续训练。然而,聚类可能会产生嘈杂的伪标签,这可能会降低整体识别性能。 针对这一问题,提出了一种基于相似连接图和图卷积网络的改进聚类框架。通过利用GCN对结构化数据建模的能力并将节点之间的关系信息并入相似性连接图中,优化了聚类过程,提高了伪标签准确性并增强了自监督说话人确认系统的鲁棒性和性能。实验结果表明,该方法显著提高了系统性能,为自监督说话人确认提供了一种新的方法。 索引词:说话人确认,自监督学习,DINO,聚类算法,图卷积网络,相似连接图摘要:With the continuous development of speech recognition technology, speaker verification (SV) has become an important method for identity authentication. Traditional SV methods rely on handcrafted feature extraction, while deep learning has significantly improved system performance. However, the scarcity of labeled data still limits the widespread application of deep learning in SV. Self-supervised learning, by mining latent information in large unlabeled datasets, enhances model generalization and is a key technology to address this issue. DINO is an efficient self-supervised learning method that generates pseudo-labels from unlabeled speech data through clustering, supporting subsequent training. However, clustering may produce noisy pseudo-labels, which can reduce overall recognition performance. To address this issue, this paper proposes an improved clustering framework based on similarity connection graphs and Graph Convolutional Networks. By leveraging GCNs' ability to model structured data and incorporating relational information between nodes in the similarity connection graph, the clustering process is optimized, improving pseudo-label accuracy and enhancing the robustness and performance of the self-supervised speaker verification system. Experimental results show that this method significantly improves system performance and provides a new approach for self-supervised speaker verification. Index Terms: Speaker Verification, Self-Supervised Learning, DINO, Clustering Algorithm, Graph Convolutional Network, Similarity Connection Graph


【9】Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis
标题:用于自然和交互式语音合成的开源全Duplex对话数据集
链接:http://arxiv.org/pdf/2509.04093v1

作者:Zhitong Zhou, Qingqing Zhang, Lei Luo, Jiechen Liu, Ruohua Zhou
摘要:在会话式TTS系统中,全双工、自发的会话数据对于增强合成语音的自然性和交互性是必不可少的。我们提出了两个开源的双轨会话语音数据集,一个在中文和英文,旨在通过提供更真实的会话数据,以提高合成语音的自然度。这两个数据集包含在隔离房间中记录的总共15小时的自然,自发的对话,为每个扬声器产生单独的高质量音轨。这些对话涵盖了各种日常话题和领域,捕捉到了现实的互动模式,包括频繁的重叠、反向通道反应、笑声和其他非语言发声。我们介绍了数据收集程序,转录和注释方法。我们展示了这些语料库的效用,微调基线TTS模型与建议的数据集。与基线相比,微调的TTS模型实现了更高的主观和客观评价指标,表明合成语音中的自然度和会话真实性得到了改善。所有的数据,注释和支持代码的微调和评估,以促进进一步研究会话语音合成。摘要:Full-duplex, spontaneous conversational data are essential for enhancing the naturalness and interactivity of synthesized speech in conversational TTS systems. We present two open-source dual-track conversational speech datasets, one in Chinese and one in English, designed to enhance the naturalness of synthesized speech by providing more realistic conversational data. The two datasets contain a total of 15 hours of natural, spontaneous conversations recorded in isolated rooms, which produces separate high-quality audio tracks for each speaker. The conversations cover diverse daily topics and domains, capturing realistic interaction patterns including frequent overlaps, backchannel responses, laughter, and other non-verbal vocalizations. We introduce the data collection procedure, transcription and annotation methods. We demonstrate the utility of these corpora by fine-tuning a baseline TTS model with the proposed datasets. The fine-tuned TTS model achieves higher subjective and objective evaluation metrics compared to the baseline, indicating improved naturalness and conversational realism in synthetic speech. All data, annotations, and supporting code for fine-tuning and evaluation are made available to facilitate further research in conversational speech synthesis.


【10】WenetSpeech-Yue: A Large-scale Cantonese Speech Corpus with Multi-dimensional Annotation
标题:WenetSpeech-Yue:具有多维注释的大型粤语语音库
链接:http://arxiv.org/pdf/2509.03959v1

作者:Longhao Li, Zhao Guo, Hongjie Chen, Yuhang Dai, Ziyu Zhang, Hongfei Xue, Tianlun Zuo, Chengyou Wang, Shuiyuan Wang, Jie Li, Xin Xu, Hui Bu, Binbin Zhang, Ruibin Yuan, Ziya Zhou, Wei Xue, Lei Xie
摘要:大规模、高质量语音数据集的出现大大加速了语音理解和生成的发展。其中,语音合成和语音合成被认为是最确定和最基本的任务。然而,对于全球约有8490万母语人士使用的粤语(粤语),有限的注释资源阻碍了进展,并导致ASR和TTS性能不佳。为了应对这一挑战,我们提出了WenetSpeech-Pipe,这是一个集成管道,用于构建大规模语音语料库,并为语音理解和生成量身定制多维注释。它包括六个模块:音频采集、说话人属性标注、语音质量标注、自动语音识别、文本后处理和识别器输出投票,实现丰富和高质量的标注。基于这一管道,我们发布了WenetSpeech-Yue,这是第一个针对ASR和TTS进行多维标注的大规模粤语语音语料库,涵盖10个领域的21,800小时,标注包括ASR转录,文本置信度,说话人身份,年龄,性别,语音质量分数等。我们还发布了WSYue-eval,一个全面的粤语基准测试,有两个组件:WSYue-ASR-eval,一个手动注释集,用于评估短和长话语,代码切换和各种声学条件下的ASR,以及WSYue-TTS-eval,用于标准和泛化测试的基础和覆盖子集。实验结果表明,在WenetSpeech-Yue上训练的模型与最先进的(SOTA)粤语ASR和TTS系统(包括商业和基于LLM的模型)相比,取得了有竞争力的结果,突出了我们的数据集和管道的价值。摘要:The development of speech understanding and generation has been significantly accelerated by the availability of large-scale, high-quality speech datasets. Among these, ASR and TTS are regarded as the most established and fundamental tasks. However, for Cantonese (Yue Chinese), spoken by approximately 84.9 million native speakers worldwide, limited annotated resources have hindered progress and resulted in suboptimal ASR and TTS performance. To address this challenge, we propose WenetSpeech-Pipe, an integrated pipeline for building large-scale speech corpus with multi-dimensional annotation tailored for speech understanding and generation. It comprises six modules: Audio Collection, Speaker Attributes Annotation, Speech Quality Annotation, Automatic Speech Recognition, Text Postprocessing and Recognizer Output Voting, enabling rich and high-quality annotations. Based on this pipeline, we release WenetSpeech-Yue, the first large-scale Cantonese speech corpus with multi-dimensional annotation for ASR and TTS, covering 21,800 hours across 10 domains with annotations including ASR transcription, text confidence, speaker identity, age, gender, speech quality scores, among other annotations. We also release WSYue-eval, a comprehensive Cantonese benchmark with two components: WSYue-ASR-eval, a manually annotated set for evaluating ASR on short and long utterances, code-switching, and diverse acoustic conditions, and WSYue-TTS-eval, with base and coverage subsets for standard and generalization testing. Experimental results show that models trained on WenetSpeech-Yue achieve competitive results against state-of-the-art (SOTA) Cantonese ASR and TTS systems, including commercial and LLM-based models, highlighting the value of our dataset and pipeline.


【11】VoxRole: A Comprehensive Benchmark for Evaluating Speech-Based Role-Playing Agents
标题:VoxRole:评估基于语音的角色扮演代理的综合基准
链接:http://arxiv.org/pdf/2509.03940v1

作者:Weihao Wu, Liang Cao, Xinyu Wu, Zhiwei Lin, Rui Niu, Jingbei Li, Zhiyong Wu
摘要:大型语言模型(LLM)的最新进展极大地推动了角色扮演会话代理(RPCA)的发展。这些系统旨在通过一致的角色采用来创建沉浸式用户体验。然而,目前RPCA的研究面临着双重限制。首先,现有的工作主要集中在语篇模态,完全忽略了关键的非语言学特征,包括语调,韵律和节奏的讲话,这是必不可少的传达人物情感和塑造生动的身份。其次,基于语音的角色扮演领域长期缺乏标准化的评估基准。大多数当前的口语对话数据集只针对基本的能力评估,其特征在于粗略描绘或定义不清的角色轮廓。因此,他们无法有效地量化模型在核心能力(如长期角色一致性)方面的表现。为了解决这一关键差距,我们引入了VoxRole,这是第一个专门为评估基于语音的RPCA而设计的综合基准。该基准测试包括13335个多回合对话,总计65.6小时的演讲,来自261部电影中的1228个独特角色。为了构建这个资源,我们提出了一个新的两阶段自动化管道,首先将电影音频与脚本对齐,然后采用LLM系统地为每个角色构建多维配置文件。利用VoxRole,我们对当代口语对话模型进行了多维度的评估,揭示了它们在保持人物角色一致性方面各自的优势和局限性。摘要:Recent significant advancements in Large Language Models (LLMs) have greatly propelled the development of Role-Playing Conversational Agents (RPCAs). These systems aim to create immersive user experiences through consistent persona adoption. However, current RPCA research faces dual limitations. First, existing work predominantly focuses on the textual modality, entirely overlooking critical paralinguistic features including intonation, prosody, and rhythm in speech, which are essential for conveying character emotions and shaping vivid identities. Second, the speech-based role-playing domain suffers from a long-standing lack of standardized evaluation benchmarks. Most current spoken dialogue datasets target only fundamental capability assessments, featuring thinly sketched or ill-defined character profiles. Consequently, they fail to effectively quantify model performance on core competencies like long-term persona consistency. To address this critical gap, we introduce VoxRole, the first comprehensive benchmark specifically designed for the evaluation of speech-based RPCAs. The benchmark comprises 13335 multi-turn dialogues, totaling 65.6 hours of speech from 1228 unique characters across 261 movies. To construct this resource, we propose a novel two-stage automated pipeline that first aligns movie audio with scripts and subsequently employs an LLM to systematically build multi-dimensional profiles for each character. Leveraging VoxRole, we conduct a multi-dimensional evaluation of contemporary spoken dialogue models, revealing crucial insights into their respective strengths and limitations in maintaining persona consistency.


【12】SwinSRGAN: Swin Transformer-based Generative Adversarial Network for High-Fidelity Speech Super-Resolution
标题:SwinSRGAN:基于Swin变换器的高保真语音超分辨率生成对抗网络
链接:http://arxiv.org/pdf/2509.03913v1
作者:Jiajun Yuan, Xiaochen Wang, Yuhang Xiao, Yulin Wu, Chenhao Hu, Xueyang Lv
备注:5 pages
摘要:语音超分辨率(SR)从低分辨率语音信号中重建高频内容。现有的系统通常遭受两级梅尔声码器流水线中的表示失配和仅CNN生成器对超高频带内容的过度平滑。扩散和流动模型的计算成本很高,并且它们在域和采样率上的鲁棒性仍然有限。我们提出了SwinSRGAN,一个端到端的框架上修改离散余弦变换(MDCT)幅度。它是一个基于Swin变换器的U-Net,通过混合对抗方案捕获长距离谱时依赖关系,将时域MPD MSD鉴别器与专门用于高频带的多带MDCT鉴别器相结合。我们采用了稀疏感知正则化arcsinh压缩的MDCT,以更好地保留瞬态分量。该系统以各种采样率对输入进行单次上采样至48 kHz,并实时运行。在标准基准测试中,SwinSRGAN减少了客观误差,提高了ABX偏好评分。在没有微调的HiFi-TTS zero-shot测试中,它的性能优于NVSR和mdctGAN,证明了跨数据集的强大泛化能力摘要:Speech super-resolution (SR) reconstructs high-frequency content from low-resolution speech signals. Existing systems often suffer from representation mismatch in two-stage mel-vocoder pipelines and from over-smoothing of hallucinated high-band content by CNN-only generators. Diffusion and flow models are computationally expensive, and their robustness across domains and sampling rates remains limited. We propose SwinSRGAN, an end-to-end framework operating on Modified Discrete Cosine Transform (MDCT) magnitudes. It is a Swin Transformer-based U-Net that captures long-range spectro-temporal dependencies with a hybrid adversarial scheme combines time-domain MPD MSD discriminators with a multi-band MDCT discriminator specialized for the high-frequency band. We employs a sparse-aware regularizer on arcsinh-compressed MDCT to better preserve transient components. The system upsamples inputs at various sampling rates to 48 kHz in a single pass and operates in real time. On standard benchmarks, SwinSRGAN reduces objective error and improves ABX preference scores. In zero-shot tests on HiFi-TTS without fine-tuning, it outperforms NVSR and mdctGAN, demonstrating strong generalization across datasets


【13】Accelerated Interactive Auralization of Highly Reverberant Spaces using Graphics Hardware
标题:使用图形硬件加速高度回响空间的交互式听觉
链接:http://arxiv.org/pdf/2509.04390v1
作者:Hannes Rosseel, Toon van Waterschoot
备注:8 pages, 6 figures, submitted to Journal of the Audio Engineering Society
摘要:

交互式声学可听化允许用户实时探索虚拟声学环境,使音乐厅或历史崇拜空间(HWS)的声学重建不再可访问,声学改变或不切实际的访问。交互式声学合成需要输入信号与一组合成滤波器的实时卷积,所述合成滤波器对空间的时空声学响应进行建模。音乐厅和HWS中的声学都以长混响时间为特征,从而导致包含许多滤波器抽头的合成滤波器。因此,卷积过程可能在计算上要求很高,从而引入了限制可听化系统的实时交互性的显著延迟。本文介绍了一个实时多通道语音可听化系统的实现。该系统能够使用GPU加速实时合成高度混响空间的声学效果。传统的基于CPU的卷积和GPU加速的卷积之间的比较,表明后者可以实现实时性能,显着降低延迟。此外,该系统在GPU上集成了声学合成和声反馈消除,创建了一个统一的基于扬声器的可听化框架,最大限度地减少了处理延迟。摘要:Interactive acoustic auralization allows users to explore virtual acoustic environments in real-time, enabling the acoustic recreation of concert hall or Historical Worship Spaces (HWS) that are either no longer accessible, acoustically altered, or impractical to visit. Interactive acoustic synthesis requires real-time convolution of input signals with a set of synthesis filters that model the space-time acoustic response of the space. The acoustics in concert halls and HWS are both characterized by a long reverberation time, resulting in synthesis filters containing many filter taps. As a result, the convolution process can be computationally demanding, introducing significant latency that limits the real-time interactivity of the auralization system. In this paper, the implementation of a real-time multichannel loudspeaker-based auralization system is presented. This system is capable of synthesizing the acoustics of highly reverberant spaces in real-time using GPU-acceleration. A comparison between traditional CPU-based convolution and GPU-accelerated convolution is presented, showing that the latter can achieve real-time performance with significantly lower latency. Additionally, the system integrates acoustic synthesis with acoustic feedback cancellation on the GPU, creating a unified loudspeaker-based auralization framework that minimizes processing latency.


【14】LibriQuote: A Speech Dataset of Fictional Character Utterances for Expressive Zero-Shot Speech Synthesis
标题:LibriQuote:用于表达性Zero-Shot语音合成的虚构字符话语的语音数据集
链接:http://arxiv.org/pdf/2509.04072v1
作者:Gaspard Michel, Elena V. Epure, Christophe Cerisara
摘要:文本到语音(TTS)系统最近通过扩展到大的语音数据集实现了更有表现力和更自然的语音合成。然而,在如此大规模的语料库中,表达性言语的比例往往是不清楚的。此外,现有的表达性语音语料库通常规模较小,主要用于测试TTS系统。在本文中,我们介绍了LibriQuote数据集,一个英语语料库,来自阅读有声读物,设计用于微调和基准测试表达zero-shot TTS系统。训练数据集包括12.7K小时的阅读,非表达性语音和5.3K小时的主要表达性语音,这些语音来自字符引用。表达子集中的每个话语都补充了它所写的上下文,以及用于描述引用的言语动词和副词的伪标签( textit{e.g.“he whisperly softly”})。此外,我们提供了一个具有挑战性的7.5小时的测试集,旨在为基准TTS系统:给定一个中性的参考语音作为输入,我们评估系统的能力,合成一个富有表现力的话语,同时保留参考音色。我们验证定性的测试集显示,它涵盖了广泛的情感相比,非表达性的讲话,随着各种口音。广泛的主观和客观的评价表明,微调的基准TTS系统LibriQuote显着提高其合成语音的可懂度,最近的系统无法合成的表达和自然的地面真理话语的语音。数据集和评估代码是免费提供的。音频样本可以在https: libriquote.github.io 上找到。摘要:Text-to-speech (TTS) systems have recently achieved more expressive and natural speech synthesis by scaling to large speech datasets. However, the proportion of expressive speech in such large-scale corpora is often unclear. Besides, existing expressive speech corpora are typically smaller in scale and primarily used for benchmarking TTS systems. In this paper, we introduce the LibriQuote dataset, an English corpus derived from read audiobooks, designed for both fine-tuning and benchmarking expressive zero-shot TTS system. The training dataset includes 12.7K hours of read, non-expressive speech and 5.3K hours of mostly expressive speech drawn from character quotations. Each utterance in the expressive subset is supplemented with the context in which it was written, along with pseudo-labels of speech verbs and adverbs used to describe the quotation ( textit{e.g. he whispered softly''}). Additionally, we provide a challenging 7.5 hour test set intended for benchmarking TTS systems: given a neutral reference speech as input, we evaluate system's ability to synthesize an expressive utterance while preserving reference timbre. We validate qualitatively the test set by showing that it covers a wide range of emotions compared to non-expressive speech, along with various accents. Extensive subjective and objective evaluations show that fine-tuning a baseline TTS system on LibriQuote significantly improves its synthesized speech intelligibility, and that recent systems fail to synthesize speech as expressive and natural as the ground-truth utterances. The dataset and evaluation code are freely available. Audio samples can be found at https: libriquote.github.io .


eess.AS音频处理


【1】Accelerated Interactive Auralization of Highly Reverberant Spaces using Graphics Hardware
标题:使用图形硬件加速高度回响空间的交互式听觉
链接:http://arxiv.org/pdf/2509.04390v1
作者:Hannes Rosseel, Toon van Waterschoot
备注:8 pages, 6 figures, submitted to Journal of the Audio Engineering Society
摘要:

交互式声学可听化允许用户实时探索虚拟声学环境,使音乐厅或历史崇拜空间(HWS)的声学重建不再可访问,声学改变或不切实际的访问。交互式声学合成需要输入信号与一组合成滤波器的实时卷积,所述合成滤波器对空间的时空声学响应进行建模。音乐厅和HWS中的声学都以长混响时间为特征,从而导致包含许多滤波器抽头的合成滤波器。因此,卷积过程可能在计算上要求很高,从而引入了限制可听化系统的实时交互性的显著延迟。本文介绍了一个实时多通道语音可听化系统的实现。该系统能够使用GPU加速实时合成高度混响空间的声学效果。传统的基于CPU的卷积和GPU加速的卷积之间的比较,表明后者可以实现实时性能,显着降低延迟。此外,该系统在GPU上集成了声学合成和声反馈消除,创建了一个统一的基于扬声器的可听化框架,最大限度地减少了处理延迟。摘要:Interactive acoustic auralization allows users to explore virtual acoustic environments in real-time, enabling the acoustic recreation of concert hall or Historical Worship Spaces (HWS) that are either no longer accessible, acoustically altered, or impractical to visit. Interactive acoustic synthesis requires real-time convolution of input signals with a set of synthesis filters that model the space-time acoustic response of the space. The acoustics in concert halls and HWS are both characterized by a long reverberation time, resulting in synthesis filters containing many filter taps. As a result, the convolution process can be computationally demanding, introducing significant latency that limits the real-time interactivity of the auralization system. In this paper, the implementation of a real-time multichannel loudspeaker-based auralization system is presented. This system is capable of synthesizing the acoustics of highly reverberant spaces in real-time using GPU-acceleration. A comparison between traditional CPU-based convolution and GPU-accelerated convolution is presented, showing that the latter can achieve real-time performance with significantly lower latency. Additionally, the system integrates acoustic synthesis with acoustic feedback cancellation on the GPU, creating a unified loudspeaker-based auralization framework that minimizes processing latency.


【2】Test-Time Adaptation for Speech Enhancement via Domain Invariant Embedding Transformation
标题:通过域不变嵌入变换实现语音增强的测试时自适应
链接:http://arxiv.org/pdf/2509.04280v1

作者:Tobias Raichle, Niels Edinger, Bin Yang
备注:This work has been submitted to the IEEE for possible publication
摘要:当测试分布与训练条件匹配时,基于深度学习的语音增强模型可以实现卓越的性能,但当部署在具有域偏移的不可预测的现实环境中时,性能往往会下降。为了应对这一挑战,我们提出了LaDen(潜在去噪),这是第一个专门为语音增强设计的测试时间自适应方法。我们的方法利用强大的预训练语音表示来执行潜在的去噪,通过噪声嵌入的线性变换来近似干净的语音表示。我们表明,这种转换推广以及跨域,使有效的伪标记的目标域没有标记的目标数据。由此产生的伪标签,使有效的测试时间适应语音增强模型在不同的声学环境。我们提出了一个全面的基准跨越多个数据集与各种域的变化,包括噪声类型,扬声器特性和语言的变化。我们广泛的实验表明,LaDen在感知指标上始终优于基线方法,特别是对于说话者和语言域的变化。摘要:Deep learning-based speech enhancement models achieve remarkable performance when test distributions match training conditions, but often degrade when deployed in unpredictable real-world environments with domain shifts. To address this challenge, we present LaDen (latent denoising), the first test-time adaptation method specifically designed for speech enhancement. Our approach leverages powerful pre-trained speech representations to perform latent denoising, approximating clean speech representations through a linear transformation of noisy embeddings. We show that this transformation generalizes well across domains, enabling effective pseudo-labeling for target domains without labeled target data. The resulting pseudo-labels enable effective test-time adaptation of speech enhancement models across diverse acoustic environments. We propose a comprehensive benchmark spanning multiple datasets with various domain shifts, including changes in noise types, speaker characteristics, and languages. Our extensive experiments demonstrate that LaDen consistently outperforms baseline methods across perceptual metrics, particularly for speaker and language domain shifts.

【3】LibriQuote: A Speech Dataset of Fictional Character Utterances for Expressive Zero-Shot Speech Synthesis
标题:LibriQuote:用于表达性Zero-Shot语音合成的虚构字符话语的语音数据集
链接:http://arxiv.org/pdf/2509.04072v1
作者:Gaspard Michel, Elena V. Epure, Christophe Cerisara
摘要:文本到语音(TTS)系统最近通过扩展到大的语音数据集实现了更有表现力和更自然的语音合成。然而,在如此大规模的语料库中,表达性言语的比例往往是不清楚的。此外,现有的表达性语音语料库通常规模较小,主要用于测试TTS系统。在本文中,我们介绍了LibriQuote数据集,一个英语语料库,来自阅读有声读物,设计用于微调和基准测试表达zero-shot TTS系统。训练数据集包括12.7K小时的阅读,非表达性语音和5.3K小时的主要表达性语音,这些语音来自字符引用。表达子集中的每个话语都补充了它所写的上下文,以及用于描述引用的言语动词和副词的伪标签( textit{e.g.“he whisperly softly”})。此外,我们提供了一个具有挑战性的7.5小时的测试集,旨在为基准TTS系统:给定一个中性的参考语音作为输入,我们评估系统的能力,合成一个富有表现力的话语,同时保留参考音色。我们验证定性的测试集显示,它涵盖了广泛的情感相比,非表达性的讲话,随着各种口音。广泛的主观和客观的评价表明,微调的基准TTS系统LibriQuote显着提高其合成语音的可懂度,最近的系统无法合成的表达和自然的地面真理话语的语音。数据集和评估代码是免费提供的。音频样本可以在https: libriquote.github.io 上找到。摘要:Text-to-speech (TTS) systems have recently achieved more expressive and natural speech synthesis by scaling to large speech datasets. However, the proportion of expressive speech in such large-scale corpora is often unclear. Besides, existing expressive speech corpora are typically smaller in scale and primarily used for benchmarking TTS systems. In this paper, we introduce the LibriQuote dataset, an English corpus derived from read audiobooks, designed for both fine-tuning and benchmarking expressive zero-shot TTS system. The training dataset includes 12.7K hours of read, non-expressive speech and 5.3K hours of mostly expressive speech drawn from character quotations. Each utterance in the expressive subset is supplemented with the context in which it was written, along with pseudo-labels of speech verbs and adverbs used to describe the quotation ( textit{e.g. he whispered softly''}). Additionally, we provide a challenging 7.5 hour test set intended for benchmarking TTS systems: given a neutral reference speech as input, we evaluate system's ability to synthesize an expressive utterance while preserving reference timbre. We validate qualitatively the test set by showing that it covers a wide range of emotions compared to non-expressive speech, along with various accents. Extensive subjective and objective evaluations show that fine-tuning a baseline TTS system on LibriQuote significantly improves its synthesized speech intelligibility, and that recent systems fail to synthesize speech as expressive and natural as the ground-truth utterances. The dataset and evaluation code are freely available. Audio samples can be found at https: libriquote.github.io .


【4】Hierarchical Sparse Sound Field Reconstruction with Spherical and Linear Microphone Arrays
标题:利用球形和线性麦克风阵列进行分层稀疏声学重建
链接:http://arxiv.org/pdf/2509.03902v1

作者:Shunxi Xu, Craig T. Jin
备注:Accepted by APSIPA ASC 2025
摘要:球形麦克风阵列(SMA)被广泛用于声场分析,稀疏恢复(SR)技术可以通过将声场建模为主导平面波的稀疏叠加来显着提高其空间分辨率。然而,SMA的空间分辨率从根本上受到它们的球谐阶数的限制,并且它们的性能在混响环境中经常下降。本文提出了一个两阶段的SR框架与残余细化,集成了从中央SMA和四个周围的线性麦克风阵列(LMA)的意见。其核心思想是利用互补的空间特性,将SMA作为一个主要的估计和LMA作为一个空间互补的细化。仿真结果表明,所提出的SMA-LMA方法显着提高空间能量图重建在不同的混响条件下,相比,无论是SMA-only和直接一步联合处理。这些结果证明了所提出的框架在增强复杂声学环境中的空间保真度和鲁棒性方面的有效性。摘要:Spherical microphone arrays (SMAs) are widely used for sound field analysis, and sparse recovery (SR) techniques can significantly enhance their spatial resolution by modeling the sound field as a sparse superposition of dominant plane waves. However, the spatial resolution of SMAs is fundamentally limited by their spherical harmonic order, and their performance often degrades in reverberant environments. This paper proposes a two-stage SR framework with residue refinement that integrates observations from a central SMA and four surrounding linear microphone arrays (LMAs). The core idea is to exploit complementary spatial characteristics by treating the SMA as a primary estimator and the LMAs as a spatially complementary refiner. Simulation results demonstrate that the proposed SMA-LMA method significantly enhances spatial energy map reconstruction under varying reverberation conditions, compared to both SMA-only and direct one-step joint processing. These results demonstrate the effectiveness of the proposed framework in enhancing spatial fidelity and robustness in complex acoustic environments.


【5】SwinSRGAN: Swin Transformer-based Generative Adversarial Network for High-Fidelity Speech Super-Resolution
标题:SwinSRGAN:基于Swin变换器的高保真语音超分辨率生成对抗网络
链接:http://arxiv.org/pdf/2509.03913v1
作者:Jiajun Yuan, Xiaochen Wang, Yuhang Xiao, Yulin Wu, Chenhao Hu, Xueyang Lv
备注:5 pages
摘要:语音超分辨率(SR)从低分辨率语音信号中重建高频内容。现有的系统通常遭受两级梅尔声码器流水线中的表示失配和仅CNN生成器对超高频带内容的过度平滑。扩散和流动模型的计算成本很高,并且它们在域和采样率上的鲁棒性仍然有限。我们提出了SwinSRGAN,一个端到端的框架上修改离散余弦变换(MDCT)幅度。它是一个基于Swin变换器的U-Net,通过混合对抗方案捕获长距离谱时依赖关系,将时域MPD MSD鉴别器与专门用于高频带的多带MDCT鉴别器相结合。我们采用了稀疏感知正则化arcsinh压缩的MDCT,以更好地保留瞬态分量。该系统以各种采样率对输入进行单次上采样至48 kHz,并实时运行。在标准基准测试中,SwinSRGAN减少了客观误差,提高了ABX偏好评分。在没有微调的HiFi-TTS zero-shot测试中,它的性能优于NVSR和mdctGAN,证明了跨数据集的强大泛化能力摘要:Speech super-resolution (SR) reconstructs high-frequency content from low-resolution speech signals. Existing systems often suffer from representation mismatch in two-stage mel-vocoder pipelines and from over-smoothing of hallucinated high-band content by CNN-only generators. Diffusion and flow models are computationally expensive, and their robustness across domains and sampling rates remains limited. We propose SwinSRGAN, an end-to-end framework operating on Modified Discrete Cosine Transform (MDCT) magnitudes. It is a Swin Transformer-based U-Net that captures long-range spectro-temporal dependencies with a hybrid adversarial scheme combines time-domain MPD MSD discriminators with a multi-band MDCT discriminator specialized for the high-frequency band. We employs a sparse-aware regularizer on arcsinh-compressed MDCT to better preserve transient components. The system upsamples inputs at various sampling rates to 48 kHz in a single pass and operates in real time. On standard benchmarks, SwinSRGAN reduces objective error and improves ABX preference scores. In zero-shot tests on HiFi-TTS without fine-tuning, it outperforms NVSR and mdctGAN, demonstrating strong generalization across datasets


【6】Multimodal Proposal for an AI-Based Tool to Increase Cross-Assessment of Messages
标题:关于基于人工智能的工具以增加消息交叉评估的多模式提案
链接:http://arxiv.org/pdf/2509.03529v1

作者:Alejandro Alvarez Castro, Joaquín Ordieres-Meré
备注:Presented at NLMLT2025 (this https URL), 15 pages, 5 figures
摘要:盈利电话会议代表了一种独特的丰富和半结构化的财务沟通来源,混合了照本宣科的管理评论和脱稿的分析师对话。尽管金融情绪分析的最新进展已经集成了多模态信号,如文本内容和语调,但大多数系统依赖于平面文档级或文档级模型,无法捕获这些交互的分层话语结构。本文介绍了一种新的多模态框架,旨在生成语义丰富,结构意识嵌入的盈利电话,通过编码它们作为层次话语树。每个节点,包括独白或问答对,都富含来自文本,音频和视频的情感信号,以及包括连贯性分数,主题标签和答案覆盖评估的结构化元数据。提出了一种两阶段的Transformer架构:第一阶段使用对比学习在节点级别对多模态内容和话语元数据进行编码,而第二阶段为整个会议合成全局嵌入。实验结果表明,由此产生的嵌入形成稳定的,语义上有意义的表示,反映情感基调,结构逻辑和主题对齐。除了财务报告,拟议的系统推广到其他高风险的脱稿交际领域,如远程医疗,教育和政治话语,提供了一个强大的和可解释的方法,多模态话语表示。这种方法提供了实际效用的下游任务,如财务预测和话语评价,同时也提供了一个可推广的方法适用于其他领域涉及高风险的通信。摘要:Earnings calls represent a uniquely rich and semi-structured source of financial communication, blending scripted managerial commentary with unscripted analyst dialogue. Although recent advances in financial sentiment analysis have integrated multi-modal signals, such as textual content and vocal tone, most systems rely on flat document-level or sentence-level models, failing to capture the layered discourse structure of these interactions. This paper introduces a novel multi-modal framework designed to generate semantically rich and structurally aware embeddings of earnings calls, by encoding them as hierarchical discourse trees. Each node, comprising either a monologue or a question-answer pair, is enriched with emotional signals derived from text, audio, and video, as well as structured metadata including coherence scores, topic labels, and answer coverage assessments. A two-stage transformer architecture is proposed: the first encodes multi-modal content and discourse metadata at the node level using contrastive learning, while the second synthesizes a global embedding for the entire conference. Experimental results reveal that the resulting embeddings form stable, semantically meaningful representations that reflect affective tone, structural logic, and thematic alignment. Beyond financial reporting, the proposed system generalizes to other high-stakes unscripted communicative domains such as tele-medicine, education, and political discourse, offering a robust and explainable approach to multi-modal discourse representation. This approach offers practical utility for downstream tasks such as financial forecasting and discourse evaluation, while also providing a generalizable method applicable to other domains involving high-stakes communication.


【7】Enhancing Speech Large Language Models through Reinforced Behavior Alignment
标题:通过强化行为对齐增强语音大型语言模型
链接:http://arxiv.org/pdf/2509.03526v1

作者:Yansong Liu, Jiateng Li, Yuan Liu
摘要:大型语言模型(LLM)的最新进展激发了人们对将其语言能力从文本扩展到其他形式的研究兴趣,这导致了基于语音的LLM(SpeechLM)的出现,该LLM具有处理语音或文本格式的用户请求的能力。然而,由于模态间的差异,这些SpeechLM仍然表现出显着的性能差距相比,他们的基于文本的LLM同行在解释以下,特别是当面对用户语音的动态和可变的性质。为了应对这一挑战,本文介绍了一个框架,被称为强化行为对齐(RBA),旨在加强语言生成能力的SpeechLM。RBA不依赖于人类注释的监督微调,而是采用自我合成方法,通过强大的教师LLM生成广泛的高保真比对数据。然后,SpeechLMs使用基于强化学习的方法将其行为与教师的行为对齐。实验结果表明,该方法有效地提高了SpeechLM的蒸馏跟踪能力,优于传统的蒸馏基线。至关重要的是,我们证明了RBA可以无缝扩展到包括口语问答和语音到文本翻译在内的任务,仅使用自我生成的数据就可以在开放基准测试中获得最先进的性能。摘要:The recent advancements of Large Language Models (LLMs) have spurred considerable research interest in extending their linguistic capabilities beyond text to other modalities, which leads to emergence of speech-based LLMs (SpeechLMs) with capability of processing user request in either speech or textual formats. However, owing to inter-modal discrepancies, these SpeechLMs still exhibit a significant performance gap compared to their text-based LLM counterparts in instruction-following, particularly when confronted with the dynamic and variable nature of user speech. To address this challenge, this paper introduces a framework termed Reinforced Behavior Alignment (RBA), designed to bolster the language generation proficiency of SpeechLMs. Instead of relying on supervised fine-tuning from human annotations, RBA employs a self-synthesis methodology to generate extensive, high-fidelity alignment data by a powerful teacher LLM. Then SpeechLMs is aligned its behavior with that of a teacher using a reinforcement learning-based approach. Experimental results demonstrate that this method effectively enhances the instruction-following capabilities of SpeechLMs that outperform conventional distillation baselines. Crucially, we demonstrate that RBA can be seamlessly extended to tasks such including spoken question answering and speech-to-text translation, attaining state-of-the-art performance on open benchmarks with only self-generated data.


【8】Speech-Based Cognitive Screening: A Systematic Evaluation of LLM Adaptation Strategies
标题:基于言语的认知筛查:LLM适应策略的系统评估
链接:http://arxiv.org/pdf/2509.03525v1

作者:Fatemeh Taherinezhad, Mohamad Javad Momeni Nezhad, Sepehr Karimi, Sina Rashidi, Ali Zolnour, Maryam Dadkhah, Yasaman Haghbin, Hossein AzadMaleki, Maryam Zolnoori
摘要:超过一半患有阿尔茨海默病和相关痴呆症的美国成年人仍未被诊断出来,基于语音的筛查提供了一种可扩展的检测方法。我们比较了使用DementiaBank语音语料库进行痴呆症检测的大型语言模型适应策略,评估了DementiaBank语音语料库记录的9个纯文本模型和3个多模态音频文本模型。适应包括在上下文学习与不同的示范选择政策,推理增强提示,参数有效的微调,和多模态集成。结果表明,类质心演示实现了最高的上下文学习性能,推理改善了较小的模型,标记级微调通常产生最好的成绩。添加分类头大大改善了表现不佳的模型。在多模态模型中,微调的音频文本系统表现良好,但没有超过顶级的纯文本模型。这些发现强调了模型适应策略,包括演示选择,推理设计和调整方法,对基于语音的痴呆症检测产生了至关重要的影响,并且适当适应的开放权重模型可以匹配或超过商业系统。摘要:Over half of US adults with Alzheimer disease and related dementias remain undiagnosed, and speech-based screening offers a scalable detection approach. We compared large language model adaptation strategies for dementia detection using the DementiaBank speech corpus, evaluating nine text-only models and three multimodal audio-text models on recordings from DementiaBank speech corpus. Adaptations included in-context learning with different demonstration selection policies, reasoning-augmented prompting, parameter-efficient fine-tuning, and multimodal integration. Results showed that class-centroid demonstrations achieved the highest in-context learning performance, reasoning improved smaller models, and token-level fine-tuning generally produced the best scores. Adding a classification head substantially improved underperforming models. Among multimodal models, fine-tuned audio-text systems performed well but did not surpass the top text-only models. These findings highlight that model adaptation strategies, including demonstration selection, reasoning design, and tuning method, critically influence speech-based dementia detection, and that properly adapted open-weight models can match or exceed commercial systems.


机器翻译由腾讯交互翻译提供,仅供参考