今日论文合集:cs.SD语音9篇,eess.AS音频处理16篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


cs.SD语音


【1】Audio-Vision Contrastive Learning for Phonological Class Recognition
标题:音类识别的视听对比学习
链接:https://arxiv.org/abs/2507.17682

作者:Daiqi Liu, Tomás Arias-Vergara, Jana Hutter, Andreas Maier, Paula Andrea Pérez-Toro
备注:conference to TSD 2025
摘要:发音-语音特征的准确分类在理解人类语音产生和开发强大的语音技术方面起着至关重要的作用,特别是在临床环境中,有针对性的音素分析和治疗可以提高疾病诊断的准确性和个性化康复。在这项工作中,我们提出了一个多模态深度学习框架,该框架结合了实时磁共振成像(rtMRI)和语音信号,对三个关键的发音维度进行分类:发音方式、发音位置和发声。我们对从上述发音维度导出的15个语音类别进行分类,并使用四种音频/视觉配置评估系统:单模态rtMRI、单模态音频信号、多模态中间融合和基于对比学习的视听融合。在USC-TIMIT数据集上的实验结果表明,我们基于对比学习的方法实现了最先进的性能,平均F1分数为0.81,比单峰基线绝对增加了0.23。实验结果证实了对比表征学习在多模态发音分析中的有效性。我们的代码和处理后的数据集将在https://github.com/DaE-plz/AC_Contrastive_Phonology上公开,以支持未来的研究。
摘要:Accurate classification of articulatory-phonological features plays a vital role in understanding human speech production and developing robust speech technologies, particularly in clinical contexts where targeted phonemic analysis and therapy can improve disease diagnosis accuracy and personalized rehabilitation. In this work, we propose a multimodal deep learning framework that combines real-time magnetic resonance imaging (rtMRI) and speech signals to classify three key articulatory dimensions: manner of articulation, place of articulation, and voicing. We perform classification on 15 phonological classes derived from the aforementioned articulatory dimensions and evaluate the system with four audio/vision configurations: unimodal rtMRI, unimodal audio signals, multimodal middle fusion, and contrastive learning-based audio-vision fusion. Experimental results on the USC-TIMIT dataset show that our contrastive learning-based approach achieves state-of-the-art performance, with an average F1-score of 0.81, representing an absolute increase of 0.23 over the unimodal baseline. The results confirm the effectiveness of contrastive representation learning for multimodal articulatory analysis. Our code and processed dataset will be made publicly available at https://github.com/DaE-plz/AC_Contrastive_Phonology to support future research.


【2】BoSS: Beyond-Semantic Speech
标题:BoSS:超越语义的言语
链接:https://arxiv.org/abs/2507.17563

作者:Qing Wang, Zehan Li, Hang Lv, Hongjie Chen, Yaodong Song, Jian Kang, Jie Lian, Jie Li, Yongxiang Li, Zhongjiang He, Xuelong Li
摘要:人类的交流不仅仅涉及明确的语义,隐含的信号和上下文线索在塑造意义方面发挥着关键作用。然而,现代语音技术,如自动语音识别(ASR)和文本到语音(TTS)往往无法捕捉这些超越语义的维度。为了更好地表征和基准的语音智能的进展,我们介绍了口语交互系统能力水平(L1-L5),一个分层的框架说明了口语对话系统的演变从基本的命令识别类人的社会互动。为了支持这些先进的功能,我们提出超越语义语音(BoSS),它是指在语音通信中的信息集,包括但超越明确的语义。它通过情感线索、语境动态和隐含语义等多维特征传达情感、语境,并修改或扩展意义,从而增强对交际意图和情景的理解。我们提出了一个形式化的框架BoSS,利用认知相关理论和机器学习模型来分析时间和上下文的语音动态。我们在五个不同的维度上评估了与口语相关的属性,发现当前的口语模型(SLM)很难完全解释语义之外的信号。这些发现强调了推进BoSS研究的必要性,以实现更丰富,更具上下文感知的人机通信。
摘要:Human communication involves more than explicit semantics, with implicit signals and contextual cues playing a critical role in shaping meaning. However, modern speech technologies, such as Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) often fail to capture these beyond-semantic dimensions. To better characterize and benchmark the progression of speech intelligence, we introduce Spoken Interaction System Capability Levels (L1-L5), a hierarchical framework illustrated the evolution of spoken dialogue systems from basic command recognition to human-like social interaction. To support these advanced capabilities, we propose Beyond-Semantic Speech (BoSS), which refers to the set of information in speech communication that encompasses but transcends explicit semantics. It conveys emotions, contexts, and modifies or extends meanings through multidimensional features such as affective cues, contextual dynamics, and implicit semantics, thereby enhancing the understanding of communicative intentions and scenarios. We present a formalized framework for BoSS, leveraging cognitive relevance theories and machine learning models to analyze temporal and contextual speech dynamics. We evaluate BoSS-related attributes across five different dimensions, reveals that current spoken language models (SLMs) are hard to fully interpret beyond-semantic signals. These findings highlight the need for advancing BoSS research to enable richer, more context-aware human-machine communication.


【3】Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
标题:Seed LiveInterpret 2.0:使用语音进行端到端同步语音翻译
链接:https://arxiv.org/abs/2507.17527

备注:Seed-LiveInterpret 2.0 Technical Report
摘要:同声传译(SI)是翻译行业最令人生畏的前沿领域之一,产品级自动化系统长期以来一直受到棘手挑战的困扰:低于标准的转录和翻译质量,缺乏实时语音生成,多说话者混淆以及翻译语音膨胀,特别是在长篇话语中。在这项研究中,我们介绍了Seed-LiveInterpret 2.0,这是一种端到端的SI模型,可提供高保真,超低延迟的语音到语音生成,并具有语音克隆功能。作为一个完全可操作的产品级解决方案,Seed-LiveInterpret 2.0通过我们新颖的双工语音到语音理解生成框架正面应对这些挑战。实验结果表明,通过大规模的预训练和强化学习,该模型在翻译准确性和延迟之间实现了更好的平衡,经人工口译员验证,在复杂场景中的正确率超过70%。值得注意的是,Seed-LiveInterpret 2.0在翻译质量方面明显优于商业SI解决方案,同时将克隆语音的平均延迟从近10秒降至近实时的3秒,减少了近70%,大大提高了实际可用性。
摘要:Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.


【4】Application of Whisper in Clinical Practice: the Post-Stroke Speech Assessment during a Naming Task
标题:Whisper在临床实践中的应用:语音任务期间中风后言语评估
链接:http://arxiv.org/pdf/2507.17326v1

作者:Milena Davudova, Ziyuan Cai, Valentina Giunchiglia, Dragos C. Gruia, Giulia Sanguedolce, Adam Hampshire, Fatemeh Geranmayeh
摘要:中风后语言障碍的详细评估仍然是一项认知复杂和临床医生密集的任务,限制了及时和可扩展的诊断。自动语音识别(ASR)基础模型为通过智能系统增强人类评估提供了一条有前途的途径,但它们在语音和语言障碍背景下的有效性仍然不确定。在这项研究中,我们评估了最先进的ASR基础模型Whisper是否可以用于转录和分析中风患者在常用的图片命名任务中的语音。我们评估了逐字转录的准确性和模型支持下游语言功能预测的能力,这对中风后的结果有重要意义。我们的研究结果表明,基线耳语模型表现不佳的单字语音话语。然而,微调Whisper显著提高了转录准确性(在健康语音中将单词错误率降低了87.72%,在患者语音中降低了71.22%)。此外,从模型中学习的表示能够准确预测语音质量(健康人的平均F1 Macro为0.74,患者为0.75)。然而,对一个看不见的(TORGO)数据集的评估显示了有限的普遍性,突出了Whisper无法对域外临床语音执行单字话语的zero-shot转录,并强调了需要使模型适应特定的临床人群。虽然在跨领域泛化方面仍然存在挑战,但这些发现突出了基础模型在适当微调时的潜力,以推进自动语音和语言评估以及中风相关障碍的康复。
摘要:Detailed assessment of language impairment following stroke remains a cognitively complex and clinician-intensive task, limiting timely and scalable diagnosis. Automatic Speech Recognition (ASR) foundation models offer a promising pathway to augment human evaluation through intelligent systems, but their effectiveness in the context of speech and language impairment remains uncertain. In this study, we evaluate whether Whisper, a state-of-the-art ASR foundation model, can be applied to transcribe and analyze speech from patients with stroke during a commonly used picture-naming task. We assess both verbatim transcription accuracy and the model's ability to support downstream prediction of language function, which has major implications for outcomes after stroke. Our results show that the baseline Whisper model performs poorly on single-word speech utterances. Nevertheless, fine-tuning Whisper significantly improves transcription accuracy (reducing Word Error Rate by 87.72% in healthy speech and 71.22% in speech from patients). Further, learned representations from the model enable accurate prediction of speech quality (average F1 Macro of 0.74 for healthy, 0.75 for patients). However, evaluations on an unseen (TORGO) dataset reveal limited generalizability, highlighting the inability of Whisper to perform zero-shot transcription of single-word utterances on out-of-domain clinical speech and emphasizing the need to adapt models to specific clinical populations. While challenges remain in cross-domain generalization, these findings highlight the potential of foundation models, when appropriately fine-tuned, to advance automated speech and language assessment and rehabilitation for stroke-related impairments.


【5】On Temporal Guidance and Iterative Refinement in Audio Source Separation

标题:音频源分离中的时间引导和迭代细化
链接:https://arxiv.org/abs/2507.17297

作者:Tobias Morocutti, Jonathan Greif, Paul Primus, Florian Schmid, Gerhard Widmer
摘要:声音场景的空间语义分割(S5)涉及到活动声音类别的准确识别以及从复杂的声学混合物中精确分离它们的来源。传统的系统依赖于一个两阶段的流水线-音频标记,然后是标签条件源分离-但往往受到限制,缺乏细粒度的时间信息的有效分离的关键。在这项工作中,我们通过引入一种新的S5方法来解决这个限制,该方法增强了事件检测和源分离阶段之间的协同作用。我们的主要贡献有三个方面。首先,我们微调预训练的Transformer以检测活动声音类别。其次,我们利用一个单独的实例,这个微调Transformer执行声音事件检测(SED),提供详细的,随时间变化的指导分离模块。第三,我们实现了一个迭代的细化机制,逐步提高分离质量递归重用分离器的输出从以前的迭代。这些进步导致音频标记和源分离性能的显着改进,正如我们的系统在DCASE挑战赛2025的任务4中获得第二名所证明的那样。我们的实现和模型检查点可在我们的GitHub存储库中找到:https://github.com/theMoro/dcase25task4。
摘要:Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a two-stage pipeline - audio tagging followed by label-conditioned source separation - but are often constrained by the absence of fine-grained temporal information critical for effective separation. In this work, we address this limitation by introducing a novel approach for S5 that enhances the synergy between the event detection and source separation stages. Our key contributions are threefold. First, we fine-tune a pre-trained Transformer to detect active sound classes. Second, we utilize a separate instance of this fine-tuned Transformer to perform sound event detection (SED), providing the separation module with detailed, time-varying guidance. Third, we implement an iterative refinement mechanism that progressively enhances separation quality by recursively reusing the separator's output from previous iterations. These advancements lead to significant improvements in both audio tagging and source separation performance, as demonstrated by our system's second-place finish in Task 4 of the DCASE Challenge 2025. Our implementation and model checkpoints are available in our GitHub repository: https://github.com/theMoro/dcase25task4 .


【6】Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems
标题:行业级CRM系统中增强型ASB模型的弱监督技术
链接:https://arxiv.org/abs/2507.16843

作者:Zhongsheng Wang, Sijie Wang, Jia Wang, Yung-I Liang, Yuxi Zhang, Jiamou Liu
备注:Accepted by ICONIP 2024
摘要:在客户关系管理(CRM)系统设计中,准确识别客户类型并提供个性化服务是提高客户满意度和忠诚度的关键。然而,这一过程面临着识别客户声音和意图的挑战,而一般的预训练自动语音识别(ASR)模型很难有效地解决特定行业的语音识别任务。针对这一问题,我们创新性地提出了针对特定行业的ASR模型微调解决方案,显著提高了微调后的ASR模型在行业应用中的性能。实验结果表明,该方法大大提高了ASR模型在行业CRM系统中的重要辅助作用,并在实际行业应用中得到了应用。
摘要:In the design of customer relationship management (CRM) systems, accurately identifying customer types and offering personalized services are key to enhancing customer satisfaction and loyalty. However, this process faces the challenge of discerning customer voices and intentions, and general pre-trained automatic speech recognition (ASR) models make it difficult to effectively address industry-specific speech recognition tasks. To address this issue, we innovatively proposed a solution for fine-tuning industry-specific ASR models, which significantly improved the performance of the fine-tuned ASR models in industry applications. Experimental results show that our method substantially improves the crucial auxiliary role of the ASR model in industry CRM systems, and this approach has also been adopted in actual industrial applications.


【7】Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
标题:使用具有非并行数据的自监督离散令牌进行口音规范化
链接:https://arxiv.org/abs/2507.17735

作者:Qibing Bai, Sho Inoue, Shuai Wang, Zhongjie Jiang, Yannan Wang, Haizhou Li
备注:Accepted to INTERSPEECH 2025
摘要:口音规范化将外国口音的语音转换为类似母语的语音,同时保留说话者身份。我们提出了一种使用自监督离散令牌和非并行训练数据的新型管道。该系统从源语音中提取标记,通过专用模型转换它们,并使用流匹配合成输出。我们的方法表现出优越的性能,在一个帧到帧的基线自然,重音减少,并在多个英语口音的音色保存。通过分词级的语音分析,我们验证了我们的基于分词的方法的有效性。我们还开发了两种持续时间保存方法,适用于配音等应用。
摘要:Accent normalization converts foreign-accented speech into native-like speech while preserving speaker identity. We propose a novel pipeline using self-supervised discrete tokens and non-parallel training data. The system extracts tokens from source speech, converts them through a dedicated model, and synthesizes the output using flow matching. Our method demonstrates superior performance over a frame-to-frame baseline in naturalness, accentedness reduction, and timbre preservation across multiple English accents. Through token-level phonetic analysis, we validate the effectiveness of our token-based approach. We also develop two duration preservation methods, suitable for applications such as dubbing.


【8】Clustering-based hard negative sampling for supervised contrastive speaker verification
标题:基于触发的硬负采样用于监督对比说话人验证
链接:https://arxiv.org/abs/2507.17540

作者:Piotr Masztalski, Michał Romaniuk, Jakub Żak, Mateusz Matuszewski, Konrad Kowalczyk
备注:Accepted to INTERSPEECH 2025
摘要:在说话人确认中,对比学习作为传统的基于分类的方法的替代方案越来越受欢迎。对比方法可以从有效使用硬否定对中受益,硬否定对是不同类别的样本,由于它们的相似性,对于验证模型来说特别具有挑战性。在本文中,我们提出了CHNS -基于聚类的硬负采样方法,专用于监督对比说话人表示学习。我们的方法聚类嵌入相似的扬声器,并调整批次组成,以获得最佳比例的硬和容易的负面对比损失计算。实验评估表明,CHNS优于基线监督对比方法,有和没有基于丢失的硬负采样,以及最先进的基于分类的方法,以多达18%的相对EER和minDCF的VoxCeleb数据集上使用两个轻量级模型架构的说话人验证。
摘要:In speaker verification, contrastive learning is gaining popularity as an alternative to the traditionally used classification-based approaches. Contrastive methods can benefit from an effective use of hard negative pairs, which are different-class samples particularly challenging for a verification model due to their similarity. In this paper, we propose CHNS - a clustering-based hard negative sampling method, dedicated for supervised contrastive speaker representation learning. Our approach clusters embeddings of similar speakers, and adjusts batch composition to obtain an optimal ratio of hard and easy negatives during contrastive loss calculation. Experimental evaluation shows that CHNS outperforms a baseline supervised contrastive approach with and without loss-based hard negative sampling, as well as a state-of-the-art classification-based approach to speaker verification by as much as 18 % relative EER and minDCF on the VoxCeleb dataset using two lightweight model architectures.


【9】Enhancing Lung Disease Diagnosis via Semi-Supervised Machine Learning
标题:通过半监督机器学习增强肺部疾病诊断
链接:https://arxiv.org/abs/2507.16845

作者:Xiaoran Xu, In-Ho Rab, Ravi Sankarc
摘要:包括肺癌和COPD在内的肺部疾病是全球重大的健康问题。传统的诊断方法可能是昂贵的,耗时的和侵入性的。本研究使用MFCC+CNN的模型组合来研究半监督学习方法在肺音信号检测中的应用。通过引入半监督学习模块,如Mix Match,Co-Refinement和Co Refurbishing,我们的目标是提高检测性能,同时减少对手动注释的依赖。通过添加半监督模块,MFCC+CNN模型的准确率为92.9%,比基线模型提高了3.8%。该研究通过解决个体差异、特征标记数据不足等挑战,为肺部疾病声音检测领域做出了贡献。
摘要:Lung diseases, including lung cancer and COPD, are significant health concerns globally. Traditional diagnostic methods can be costly, time-consuming, and invasive. This study investigates the use of semi supervised learning methods for lung sound signal detection using a model combination of MFCC+CNN. By introducing semi supervised learning modules such as Mix Match, Co-Refinement, and Co Refurbishing, we aim to enhance the detection performance while reducing dependence on manual annotations. With the add-on semi-supervised modules, the accuracy rate of the MFCC+CNN model is 92.9%, an increase of 3.8% to the baseline model. The research contributes to the field of lung disease sound detection by addressing challenges such as individual differences, feature insufficient labeled data.


eess.AS音频处理


【1】Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
标题:使用具有非并行数据的自监督离散令牌进行口音规范化
链接:https://arxiv.org/abs/2507.17735

作者:Qibing Bai, Sho Inoue, Shuai Wang, Zhongjie Jiang, Yannan Wang, Haizhou Li
备注:Accepted to INTERSPEECH 2025
摘要:口音规范化将外国口音的语音转换为类似母语的语音,同时保留说话者身份。我们提出了一种使用自监督离散令牌和非并行训练数据的新型管道。该系统从源语音中提取标记,通过专用模型转换它们,并使用流匹配合成输出。我们的方法表现出优越的性能,在一个帧到帧的基线自然,重音减少,并在多个英语口音的音色保存。通过分词级的语音分析,我们验证了我们的基于分词的方法的有效性。我们还开发了两种持续时间保存方法,适用于配音等应用。
摘要:Accent normalization converts foreign-accented speech into native-like speech while preserving speaker identity. We propose a novel pipeline using self-supervised discrete tokens and non-parallel training data. The system extracts tokens from source speech, converts them through a dedicated model, and synthesizes the output using flow matching. Our method demonstrates superior performance over a frame-to-frame baseline in naturalness, accentedness reduction, and timbre preservation across multiple English accents. Through token-level phonetic analysis, we validate the effectiveness of our token-based approach. We also develop two duration preservation methods, suitable for applications such as dubbing.


【2】Clustering-based hard negative sampling for supervised contrastive speaker verification
标题:基于触发的硬负采样用于监督对比说话人验证
链接:https://arxiv.org/abs/2507.17540

作者:Piotr Masztalski, Michał Romaniuk, Jakub Żak, Mateusz Matuszewski, Konrad Kowalczyk
备注:Accepted to INTERSPEECH 2025
摘要:在说话人确认中,对比学习作为传统的基于分类的方法的替代方案越来越受欢迎。对比方法可以从有效使用硬否定对中受益,硬否定对是不同类别的样本,由于它们的相似性,对于验证模型来说特别具有挑战性。在本文中,我们提出了CHNS -基于聚类的硬负采样方法,专用于监督对比说话人表示学习。我们的方法聚类嵌入相似的扬声器,并调整批次组成,以获得最佳比例的硬和容易的负面对比损失计算。实验评估表明,CHNS优于基线监督对比方法,有和没有基于丢失的硬负采样,以及最先进的基于分类的方法,以多达18%的相对EER和minDCF的VoxCeleb数据集上使用两个轻量级模型架构的说话人验证。
摘要:In speaker verification, contrastive learning is gaining popularity as an alternative to the traditionally used classification-based approaches. Contrastive methods can benefit from an effective use of hard negative pairs, which are different-class samples particularly challenging for a verification model due to their similarity. In this paper, we propose CHNS - a clustering-based hard negative sampling method, dedicated for supervised contrastive speaker representation learning. Our approach clusters embeddings of similar speakers, and adjusts batch composition to obtain an optimal ratio of hard and easy negatives during contrastive loss calculation. Experimental evaluation shows that CHNS outperforms a baseline supervised contrastive approach with and without loss-based hard negative sampling, as well as a state-of-the-art classification-based approach to speaker verification by as much as 18 % relative EER and minDCF on the VoxCeleb dataset using two lightweight model architectures.


【3】SLASH: Self-Supervised Speech Pitch Estimation Leveraging DSP-derived Absolute Pitch
标题:SLASH:利用DSP推导的绝对音调的自监督语音音调估计
链接:https://arxiv.org/abs/2507.17208

作者:Ryo Terashima, Yuma Shirahata, Masaya Kawamura
备注:Accepted to INTERSPEECH 2025
摘要:提出了一种基于自监督学习(SSL)的语音信号基音周期估计方法SLASH。为了提高传统的基于SSL的方法,主要取决于从音高偏移的相对音高差的性能,我们的方法结合绝对音高值1)引入一个先验的音高分布从数字信号处理(DSP),和2)通过梯度下降与目标和可微分DSP导出的频谱图之间的损失优化绝对音高。为了稳定优化,使用了一种新的频谱图生成方法,该方法跳过了复杂的波形生成。此外,语音中的非周期成分通过可微DSP精确预测,增强了该方法在语音信号处理中的适用性。实验结果表明,该方法优于基线DSP和基于SSL的基音周期估计方法,归因于SSL和DSP的有效集成。
摘要:We present SLASH, a pitch estimation method of speech signals based on self-supervised learning (SSL). To enhance the performance of conventional SSL-based approaches that primarily depend on the relative pitch difference derived from pitch shifting, our method incorporates absolute pitch values by 1) introducing a prior pitch distribution derived from digital signal processing (DSP), and 2) optimizing absolute pitch through gradient descent with a loss between the target and differentiable DSP-derived spectrograms. To stabilize the optimization, a novel spectrogram generation method is used that skips complicated waveform generation. In addition, the aperiodic components in speech are accurately predicted through differentiable DSP, enhancing the method's applicability to speech signal processing. Experimental results showed that the proposed method outperformed both baseline DSP and SSL-based pitch estimation methods, attributed to the effective integration of SSL and DSP.


【4】Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages
标题:技术报告:持续时间预测对印度语言特定于说话者的TTC的影响
链接:https://arxiv.org/abs/2507.16875

作者:Isha Pandey, Pranav Gaikwad, Amruta Parulekar, Ganesh Ramakrishnan
摘要:由于数据有限和语言结构多样,低资源语言(如许多印度语言)的高质量语音生成仍然是一个重大挑战。时长预测是许多语音生成流程中的关键组成部分,在韵律和语音节奏建模中起着关键作用。而最近的一些生成方法选择省略显式持续时间建模,通常以更长的训练时间为代价。我们保留和探索这个模块,以更好地了解其在印度语言丰富和数据稀缺的景观的影响。我们训练一个非自回归连续归一化流(CNF)的语音模型使用公开的印度语言数据和评估多个持续时间预测策略zero-shot,扬声器特定的一代。我们对语音填充任务的比较分析揭示了细微的权衡:基于填充的预测器提高了某些语言的可理解性,而说话者提示的预测器更好地保留了其他语言的说话者特征。这些发现为针对特定语言和任务的持续时间策略的设计和选择提供了信息,强调了持续时间预测等可解释组件在适应低资源多语言环境的高级生成架构方面的持续价值。
摘要:High-quality speech generation for low-resource languages, such as many Indian languages, remains a significant challenge due to limited data and diverse linguistic structures. Duration prediction is a critical component in many speech generation pipelines, playing a key role in modeling prosody and speech rhythm. While some recent generative approaches choose to omit explicit duration modeling, often at the cost of longer training times. We retain and explore this module to better understand its impact in the linguistically rich and data-scarce landscape of India. We train a non-autoregressive Continuous Normalizing Flow (CNF) based speech model using publicly available Indian language data and evaluate multiple duration prediction strategies for zero-shot, speaker-specific generation. Our comparative analysis on speech-infilling tasks reveals nuanced trade-offs: infilling based predictors improve intelligibility in some languages, while speaker-prompted predictors better preserve speaker characteristics in others. These findings inform the design and selection of duration strategies tailored to specific languages and tasks, underscoring the continued value of interpretable components like duration prediction in adapting advanced generative architectures to low-resource, multilingual settings.


【5】Enhancing Lung Disease Diagnosis via Semi-Supervised Machine Learning
标题:通过半监督机器学习增强肺部疾病诊断
链接:https://arxiv.org/abs/2507.16845

作者:Xiaoran Xu, In-Ho Rab, Ravi Sankarc
摘要:包括肺癌和COPD在内的肺部疾病是全球重大的健康问题。传统的诊断方法可能是昂贵的,耗时的和侵入性的。本研究使用MFCC+CNN的模型组合来研究半监督学习方法在肺音信号检测中的应用。通过引入半监督学习模块,如Mix Match,Co-Refinement和Co Refurbishing,我们的目标是提高检测性能,同时减少对手动注释的依赖。通过添加半监督模块,MFCC+CNN模型的准确率为92.9%,比基线模型提高了3.8%。该研究通过解决个体差异、特征标记数据不足等挑战,为肺部疾病声音检测领域做出了贡献。
摘要:Lung diseases, including lung cancer and COPD, are significant health concerns globally. Traditional diagnostic methods can be costly, time-consuming, and invasive. This study investigates the use of semi supervised learning methods for lung sound signal detection using a model combination of MFCC+CNN. By introducing semi supervised learning modules such as Mix Match, Co-Refinement, and Co Refurbishing, we aim to enhance the detection performance while reducing dependence on manual annotations. With the add-on semi-supervised modules, the accuracy rate of the MFCC+CNN model is 92.9%, an increase of 3.8% to the baseline model. The research contributes to the field of lung disease sound detection by addressing challenges such as individual differences, feature insufficient labeled data.


【6】Segmentation-free Goodness of Pronunciation
标题:无音段的语音美
链接:https://arxiv.org/abs/2507.16838

作者:Xinwei Cao, Zijian Fan, Torbjørn Svendsen, Giampiero Salvi
备注:This work has been submitted to the IEEE for possible publication
摘要:发音错误检测与诊断是现代计算机辅助语言学习系统的重要组成部分。在MDD中,音素水平的发音评估是帮助二语学习者提高发音的关键。然而,大多数系统是基于一种形式的良好的发音(GOP),这需要预先分割语音成语音单元。这限制了这些方法的准确性以及使用现代基于CTC的声学模型进行评估的可能性。在这项研究中,我们首先提出了自对准GOP(GOP-SA),它可以使用CTC训练的ASR模型进行MDD。接下来,我们定义了一个更通用的无干扰方法,该方法考虑了目标音素的所有可能的对齐(GOP-AF)。我们给出了我们的GOP-AF的定义,实现解决潜在的数值问题,以及适当的规范化,使该方法适用于随着时间的推移具有不同峰值的声学模型的理论说明。我们在CMU Kids和Speechocean 762数据集上提供了大量的实验结果,比较了我们方法的不同定义,估计了GOP-AF对声学模型峰值和目标音素周围上下文量的依赖性。最后,我们将我们的方法与最近对Speechocean 762数据的研究进行了比较,结果表明,从所提出的方法中推导出的特征向量在音素级发音评估方面取得了最先进的结果。
摘要:Mispronunciation detection and diagnosis (MDD) is a significant part in modern computer aided language learning (CALL) systems. Within MDD, phoneme-level pronunciation assessment is key to helping L2 learners improve their pronunciation. However, most systems are based on a form of goodness of pronunciation (GOP) which requires pre-segmentation of speech into phonetic units. This limits the accuracy of these methods and the possibility to use modern CTC-based acoustic models for their evaluation. In this study, we first propose self-alignment GOP (GOP-SA) that enables the use of CTC-trained ASR models for MDD. Next, we define a more general alignment-free method that takes all possible alignments of the target phoneme into account (GOP-AF). We give a theoretical account of our definition of GOP-AF, an implementation that solves potential numerical issues as well as a proper normalization which makes the method applicable with acoustic models with different peakiness over time. We provide extensive experimental results on the CMU Kids and Speechocean762 datasets comparing the different definitions of our methods, estimating the dependency of GOP-AF on the peakiness of the acoustic models and on the amount of context around the target phoneme. Finally, we compare our methods with recent studies over the Speechocean762 data showing that the feature vectors derived from the proposed method achieve state-of-the-art results on phoneme-level pronunciation assessment.


【7】From Black Box to Biomarker: Sparse Autoencoders for Interpreting Speech Models of Parkinson's Disease
标题:从黑匣子到生物标志物:用于解释帕金森病语音模型的稀疏自动编码器
链接:https://arxiv.org/abs/2507.16836

作者:Peter Plantinga, Jen-Kai Chen, Roozbeh Sattari, Mirco Ravanelli, Denise Klein
备注:14 pages, 5 figures, submitted to NeurIPS 2025
摘要:语音有望成为帕金森病(PD)等神经系统疾病的成本效益和非侵入性生物标志物。虽然在原始音频上训练的深度学习系统可以找到手工制作的功能无法获得的微妙信号,但它们的黑盒性质阻碍了临床应用。为了解决这个问题,我们应用稀疏自动编码器(SAE)来揭示基于语音的PD检测系统的可解释的内部表示。我们引入了一种新的基于掩码的激活,用于使SAE适应小型生物医学数据集,创建稀疏的解纠缠字典表示。这些字典条目被发现有很强的关联与PD语音中的特征发音缺陷,如减少频谱通量和增加频谱平坦度的低能量区域突出的模型注意。我们进一步表明,光谱通量与MRI扫描的壳核体积测量相关,证明SAE揭示疾病监测和诊断的临床相关生物标志物的潜力。
摘要:Speech holds promise as a cost-effective and non-invasive biomarker for neurological conditions such as Parkinson's disease (PD). While deep learning systems trained on raw audio can find subtle signals not available from hand-crafted features, their black-box nature hinders clinical adoption. To address this, we apply sparse autoencoders (SAEs) to uncover interpretable internal representations from a speech-based PD detection system. We introduce a novel mask-based activation for adapting SAEs to small biomedical datasets, creating sparse disentangled dictionary representations. These dictionary entries are found to have strong associations with characteristic articulatory deficits in PD speech, such as reduced spectral flux and increased spectral flatness in the low-energy regions highlighted by the model attention. We further show that the spectral flux is related to volumetric measurements of the putamen from MRI scans, demonstrating the potential of SAEs to reveal clinically relevant biomarkers for disease monitoring and diagnosis.


【8】Evaluating Speech-to-Text x LLM x Text-to-Speech Combinations for AI Interview Systems
标题:评估人工智能面试系统的语音转文本x LLM x文本转语音组合
链接:https://arxiv.org/abs/2507.16835

作者:Ali Ansari, Ali Ansari, Aruj Mahajan, Amirhossein Afsharrad, Seyed Shahabeddin Mousavi
摘要:基于语音的会话AI系统越来越依赖于级联架构,这些架构结合了语音到文本(STT)、大型语言模型(LLM)和文本到语音(TTS)组件。然而,在生产环境中对不同成分组合的系统评价仍然研究不足。我们使用来自30多万次人工智能面试的数据,对STT x LLM x TTS堆栈进行了大规模的实证比较。我们开发了一个自动化的评估框架,使用法学硕士作为一个法官来评估会话质量,技术准确性和技能评估能力。我们对四种生产配置的分析表明,Google STT与GPT-4.1搭配使用,在会话和技术质量指标方面都明显优于其他产品。令人惊讶的是,我们发现客观的质量指标与用户满意度得分的相关性很弱,这表明基于语音的AI系统中的用户体验取决于技术性能之外的因素。我们的研究结果为选择多模态会话AI系统中的组件提供了实用指导,并为基于语音的交互提供了经过验证的评估方法。
摘要:Voice-based conversational AI systems increasingly rely on cascaded architectures combining speech-to-text (STT), large language models (LLMs), and text-to-speech (TTS) components. However, systematic evaluation of different component combinations in production settings remains understudied. We present a large-scale empirical comparison of STT x LLM x TTS stacks using data from over 300,000 AI-conducted job interviews. We develop an automated evaluation framework using LLM-as-a-Judge to assess conversational quality, technical accuracy, and skill assessment capabilities. Our analysis of four production configurations reveals that Google STT paired with GPT-4.1 significantly outperforms alternatives in both conversational and technical quality metrics. Surprisingly, we find that objective quality metrics correlate weakly with user satisfaction scores, suggesting that user experience in voice-based AI systems depends on factors beyond technical performance. Our findings provide practical guidance for selecting components in multimodal conversational AI systems and contribute a validated evaluation methodology for voice-based interactions.


【9】Towards Robust Speech Recognition for Jamaican Patois Music Transcription
标题:牙买加方言音乐转录的稳健语音识别
链接:https://arxiv.org/abs/2507.16834

作者:Jordan Madden, Matthew Stone, Dimitri Johnson, Daniel Geddez
摘要:虽然牙买加方言是一种广泛使用的语言,但目前的语音识别系统在方言音乐上表现不佳,产生不准确的字幕,限制了可访问性并阻碍了下游应用。在这项工作中,我们采取了以数据为中心的方法来解决这个问题,通过策划超过40小时的手动转录的Patois音乐。我们使用该数据集来微调最先进的自动语音识别(ASR)模型,并使用结果来开发Whisper模型在牙买加方言音频上的性能的缩放律。我们希望这项工作将对牙买加方言音乐的可访问性和牙买加方言语言建模的未来产生积极的影响。
摘要:Although Jamaican Patois is a widely spoken language, current speech recognition systems perform poorly on Patois music, producing inaccurate captions that limit accessibility and hinder downstream applications. In this work, we take a data-centric approach to this problem by curating more than 40 hours of manually transcribed Patois music. We use this dataset to fine-tune state-of-the-art automatic speech recognition (ASR) models, and use the results to develop scaling laws for the performance of Whisper models on Jamaican Patois audio. We hope that this work will have a positive impact on the accessibility of Jamaican Patois music and the future of Jamaican Patois language modeling.


【10】Does Language Matter for Early Detection of Parkinson's Disease from Speech?
标题:语言对于通过言语早期检测帕金森病很重要吗?
链接:https://arxiv.org/abs/2507.16832

作者:Peter Plantinga, Briac Cordelle, Dominique Louër, Mirco Ravanelli, Denise Klein
备注:Accepted to IEEE Workshop on Machine Learning for Signal Processing (MLSP) 2025
摘要:使用语音样本作为生物标志物是检测和监测帕金森病(PD)进展的一种有前途的途径,但在文献中关于如何最好地收集和分析这些数据存在相当大的分歧。早期的研究在检测PD从语音使用持续元音发声(SVP)的任务,而最近的一些研究探索了更认知要求的任务的录音。为了评估语言在PD检测中的作用,我们测试了具有不同数据类型和预训练目标的预训练模型,发现(1)纯文本模型与语音特征模型的性能相匹配,(2)多语言Whisper优于自监督模型,而单语Whisper表现更差,(3)AudioSet预训练提高了SVP的性能,但没有提高自发语音的性能。这些发现共同强调了语言在帕金森病早期检测中的关键作用。
摘要:Using speech samples as a biomarker is a promising avenue for detecting and monitoring the progression of Parkinson's disease (PD), but there is considerable disagreement in the literature about how best to collect and analyze such data. Early research in detecting PD from speech used a sustained vowel phonation (SVP) task, while some recent research has explored recordings of more cognitively demanding tasks. To assess the role of language in PD detection, we tested pretrained models with varying data types and pretraining objectives and found that (1) text-only models match the performance of vocal-feature models, (2) multilingual Whisper outperforms self-supervised models whereas monolingual Whisper does worse, and (3) AudioSet pretraining improves performance on SVP but not spontaneous speech. These findings together highlight the critical role of language for the early detection of Parkinson's disease.


【11】Audio-Vision Contrastive Learning for Phonological Class Recognition
标题:音类识别的视听对比学习
链接:https://arxiv.org/abs/2507.17682

作者:Daiqi Liu, Tomás Arias-Vergara, Jana Hutter, Andreas Maier, Paula Andrea Pérez-Toro
备注:conference to TSD 2025
摘要:发音-语音特征的准确分类在理解人类语音产生和开发强大的语音技术方面起着至关重要的作用,特别是在临床环境中,有针对性的音素分析和治疗可以提高疾病诊断的准确性和个性化康复。在这项工作中,我们提出了一个多模态深度学习框架,该框架结合了实时磁共振成像(rtMRI)和语音信号,对三个关键的发音维度进行分类:发音方式、发音位置和发声。我们对从上述发音维度导出的15个语音类别进行分类,并使用四种音频/视觉配置评估系统:单模态rtMRI、单模态音频信号、多模态中间融合和基于对比学习的视听融合。在USC-TIMIT数据集上的实验结果表明,我们基于对比学习的方法实现了最先进的性能,平均F1分数为0.81,比单峰基线绝对增加了0.23。实验结果证实了对比表征学习在多模态发音分析中的有效性。我们的代码和处理后的数据集将在https://github.com/DaE-plz/AC_Contrastive_Phonology上公开,以支持未来的研究。
摘要:Accurate classification of articulatory-phonological features plays a vital role in understanding human speech production and developing robust speech technologies, particularly in clinical contexts where targeted phonemic analysis and therapy can improve disease diagnosis accuracy and personalized rehabilitation. In this work, we propose a multimodal deep learning framework that combines real-time magnetic resonance imaging (rtMRI) and speech signals to classify three key articulatory dimensions: manner of articulation, place of articulation, and voicing. We perform classification on 15 phonological classes derived from the aforementioned articulatory dimensions and evaluate the system with four audio/vision configurations: unimodal rtMRI, unimodal audio signals, multimodal middle fusion, and contrastive learning-based audio-vision fusion. Experimental results on the USC-TIMIT dataset show that our contrastive learning-based approach achieves state-of-the-art performance, with an average F1-score of 0.81, representing an absolute increase of 0.23 over the unimodal baseline. The results confirm the effectiveness of contrastive representation learning for multimodal articulatory analysis. Our code and processed dataset will be made publicly available at https://github.com/DaE-plz/AC_Contrastive_Phonology to support future research.


【12】BoSS: Beyond-Semantic Speech
标题:BoSS:超越语义的言语
链接:https://arxiv.org/abs/2507.17563

作者:Qing Wang, Zehan Li, Hang Lv, Hongjie Chen, Yaodong Song, Jian Kang, Jie Lian, Jie Li, Yongxiang Li, Zhongjiang He, Xuelong Li
摘要:人类的交流不仅仅涉及明确的语义,隐含的信号和上下文线索在塑造意义方面发挥着关键作用。然而,现代语音技术,如自动语音识别(ASR)和文本到语音(TTS)往往无法捕捉这些超越语义的维度。为了更好地表征和基准的语音智能的进展,我们介绍了口语交互系统能力水平(L1-L5),一个分层的框架说明了口语对话系统的演变从基本的命令识别类人的社会互动。为了支持这些先进的功能,我们提出超越语义语音(BoSS),它是指在语音通信中的信息集,包括但超越明确的语义。它通过情感线索、语境动态和隐含语义等多维特征传达情感、语境,并修改或扩展意义,从而增强对交际意图和情景的理解。我们提出了一个形式化的框架BoSS,利用认知相关理论和机器学习模型来分析时间和上下文的语音动态。我们在五个不同的维度上评估了与口语相关的属性,发现当前的口语模型(SLM)很难完全解释语义之外的信号。这些发现强调了推进BoSS研究的必要性,以实现更丰富,更具上下文感知的人机通信。
摘要:Human communication involves more than explicit semantics, with implicit signals and contextual cues playing a critical role in shaping meaning. However, modern speech technologies, such as Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) often fail to capture these beyond-semantic dimensions. To better characterize and benchmark the progression of speech intelligence, we introduce Spoken Interaction System Capability Levels (L1-L5), a hierarchical framework illustrated the evolution of spoken dialogue systems from basic command recognition to human-like social interaction. To support these advanced capabilities, we propose Beyond-Semantic Speech (BoSS), which refers to the set of information in speech communication that encompasses but transcends explicit semantics. It conveys emotions, contexts, and modifies or extends meanings through multidimensional features such as affective cues, contextual dynamics, and implicit semantics, thereby enhancing the understanding of communicative intentions and scenarios. We present a formalized framework for BoSS, leveraging cognitive relevance theories and machine learning models to analyze temporal and contextual speech dynamics. We evaluate BoSS-related attributes across five different dimensions, reveals that current spoken language models (SLMs) are hard to fully interpret beyond-semantic signals. These findings highlight the need for advancing BoSS research to enable richer, more context-aware human-machine communication.


【13】Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
标题:Seed LiveInterpret 2.0:使用语音进行端到端同步语音翻译
链接:https://arxiv.org/abs/2507.17527

备注:Seed-LiveInterpret 2.0 Technical Report
摘要:同声传译(SI)是翻译行业最令人生畏的前沿领域之一,产品级自动化系统长期以来一直受到棘手挑战的困扰:低于标准的转录和翻译质量,缺乏实时语音生成,多说话者混淆以及翻译语音膨胀,特别是在长篇话语中。在这项研究中,我们介绍了Seed-LiveInterpret 2.0,这是一种端到端的SI模型,可提供高保真,超低延迟的语音到语音生成,并具有语音克隆功能。作为一个完全可操作的产品级解决方案,Seed-LiveInterpret 2.0通过我们新颖的双工语音到语音理解生成框架正面应对这些挑战。实验结果表明,通过大规模的预训练和强化学习,该模型在翻译准确性和延迟之间实现了更好的平衡,经人工口译员验证,在复杂场景中的正确率超过70%。值得注意的是,Seed-LiveInterpret 2.0在翻译质量方面明显优于商业SI解决方案,同时将克隆语音的平均延迟从近10秒降至近实时的3秒,减少了近70%,大大提高了实际可用性。
摘要:Simultaneous Interpretation (SI) represents one of the most daunting frontiers in the translation industry, with product-level automatic systems long plagued by intractable challenges: subpar transcription and translation quality, lack of real-time speech generation, multi-speaker confusion, and translated speech inflation, especially in long-form discourses. In this study, we introduce Seed-LiveInterpret 2.0, an end-to-end SI model that delivers high-fidelity, ultra-low-latency speech-to-speech generation with voice cloning capabilities. As a fully operational product-level solution, Seed-LiveInterpret 2.0 tackles these challenges head-on through our novel duplex speech-to-speech understanding-generating framework. Experimental results demonstrate that through large-scale pretraining and reinforcement learning, the model achieves a significantly better balance between translation accuracy and latency, validated by human interpreters to exceed 70% correctness in complex scenarios. Notably, Seed-LiveInterpret 2.0 outperforms commercial SI solutions by significant margins in translation quality, while slashing the average latency of cloned speech from nearly 10 seconds to a near-real-time 3 seconds, which is around a near 70% reduction that drastically enhances practical usability.


【14】On Temporal Guidance and Iterative Refinement in Audio Source Separation
标题:音频源分离中的时间引导和迭代细化
链接:https://arxiv.org/abs/2507.17297

作者:Tobias Morocutti, Jonathan Greif, Paul Primus, Florian Schmid, Gerhard Widmer
摘要:声音场景的空间语义分割(S5)涉及到活动声音类别的准确识别以及从复杂的声学混合物中精确分离它们的来源。传统的系统依赖于一个两阶段的流水线-音频标记,然后是标签条件源分离-但往往受到限制,缺乏细粒度的时间信息的有效分离的关键。在这项工作中,我们通过引入一种新的S5方法来解决这个限制,该方法增强了事件检测和源分离阶段之间的协同作用。我们的主要贡献有三个方面。首先,我们微调预训练的Transformer以检测活动声音类别。其次,我们利用一个单独的实例,这个微调Transformer执行声音事件检测(SED),提供详细的,随时间变化的指导分离模块。第三,我们实现了一个迭代的细化机制,逐步提高分离质量递归重用分离器的输出从以前的迭代。这些进步导致音频标记和源分离性能的显着改进,正如我们的系统在DCASE挑战赛2025的任务4中获得第二名所证明的那样。我们的实现和模型检查点可在我们的GitHub存储库中找到:https://github.com/theMoro/dcase25task4。
摘要:Spatial semantic segmentation of sound scenes (S5) involves the accurate identification of active sound classes and the precise separation of their sources from complex acoustic mixtures. Conventional systems rely on a two-stage pipeline - audio tagging followed by label-conditioned source separation - but are often constrained by the absence of fine-grained temporal information critical for effective separation. In this work, we address this limitation by introducing a novel approach for S5 that enhances the synergy between the event detection and source separation stages. Our key contributions are threefold. First, we fine-tune a pre-trained Transformer to detect active sound classes. Second, we utilize a separate instance of this fine-tuned Transformer to perform sound event detection (SED), providing the separation module with detailed, time-varying guidance. Third, we implement an iterative refinement mechanism that progressively enhances separation quality by recursively reusing the separator's output from previous iterations. These advancements lead to significant improvements in both audio tagging and source separation performance, as demonstrated by our system's second-place finish in Task 4 of the DCASE Challenge 2025. Our implementation and model checkpoints are available in our GitHub repository: https://github.com/theMoro/dcase25task4 .


【15】Triple X: A LLM-Based Multilingual Speech Recognition System for the INTERSPEECH2025 MLC-SLM Challenge
标题:Triple X:一个基于LLM的多语言语音识别系统,用于INTERSPEECT 2025 MLC-LAM挑战赛
链接:https://arxiv.org/abs/2507.17288

作者:Miaomiao Gao, Xiaoxiao Xiang, Yiwen Guo
摘要:本文介绍了我们的三重X语音识别系统提交的任务1的多语言会话语音语言建模(MLC-SLM)的挑战。我们的工作重点是通过创新的编码器-适配器-LLM架构来优化多语言会话场景中的语音识别准确性。该框架利用了基于文本的大型语言模型的强大推理能力,同时结合了特定领域的适应性。为了进一步提高多语言识别性能,我们采用了精心设计的多阶段训练策略,利用广泛的多语言音频数据集。实验结果表明,我们的方法在开发和测试集上都取得了有竞争力的字错误率(WER)性能,在挑战排名中获得第二名。
摘要:This paper describes our Triple X speech recognition system submitted to Task 1 of the Multi-Lingual Conversational Speech Language Modeling (MLC-SLM) Challenge. Our work focuses on optimizing speech recognition accuracy in multilingual conversational scenarios through an innovative encoder-adapter-LLM architecture. This framework harnesses the powerful reasoning capabilities of text-based large language models while incorporating domain-specific adaptations. To further enhance multilingual recognition performance, we adopted a meticulously designed multi-stage training strategy leveraging extensive multilingual audio datasets. Experimental results demonstrate that our approach achieves competitive Word Error Rate (WER) performance on both dev and test sets, obtaining second place in the challenge ranking.


【16】Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems
标题:行业级CRM系统中增强型ASB模型的弱监督技术
链接:https://arxiv.org/abs/2507.16843

作者:Zhongsheng Wang, Sijie Wang, Jia Wang, Yung-I Liang, Yuxi Zhang, Jiamou Liu
备注:Accepted by ICONIP 2024
摘要:在客户关系管理(CRM)系统设计中,准确识别客户类型并提供个性化服务是提高客户满意度和忠诚度的关键。然而,这一过程面临着识别客户声音和意图的挑战,而一般的预训练自动语音识别(ASR)模型很难有效地解决特定行业的语音识别任务。针对这一问题,我们创新性地提出了针对特定行业的ASR模型微调解决方案,显著提高了微调后的ASR模型在行业应用中的性能。实验结果表明,该方法大大提高了ASR模型在行业CRM系统中的重要辅助作用,并在实际行业应用中得到了应用。
摘要:In the design of customer relationship management (CRM) systems, accurately identifying customer types and offering personalized services are key to enhancing customer satisfaction and loyalty. However, this process faces the challenge of discerning customer voices and intentions, and general pre-trained automatic speech recognition (ASR) models make it difficult to effectively address industry-specific speech recognition tasks. To address this issue, we innovatively proposed a solution for fine-tuning industry-specific ASR models, which significantly improved the performance of the fine-tuned ASR models in industry applications. Experimental results show that our method substantially improves the crucial auxiliary role of the ASR model in industry CRM systems, and this approach has also been adopted in actual industrial applications.


机器翻译由腾讯交互翻译提供,仅供参考