微信公众号:arXiv_Daily
cs.SD语音
【1】Identifying Hearing Difficulty Moments in Conversational Audio
标题:识别对话音频中的听力困难时刻
链接:https://arxiv.org/abs/2507.23590
摘要:人们在日常谈话中经常遇到听力困难的时刻。识别这些听力困难的时刻在听力辅助技术领域具有特殊意义,其中及时干预是实时听力辅助的关键。在本文中,我们提出并比较了机器学习解决方案,用于连续检测识别会话音频中这些特定时刻的话语。我们表明,音频语言模型通过其多模态推理能力,在这项任务中表现出色,显着优于简单的ASR热词启发式算法和使用Wav 2 Vec(一种最先进的纯音频输入架构)的更传统的微调方法。自动语音识别(ASR)的最新技术。
摘要:Individuals regularly experience Hearing Difficulty Moments in everyday conversation. Identifying these moments of hearing difficulty has particular significance in the field of hearing assistive technology where timely interventions are key for realtime hearing assistance. In this paper, we propose and compare machine learning solutions for continuously detecting utterances that identify these specific moments in conversational audio. We show that audio language models, through their multimodal reasoning capabilities, excel at this task, significantly outperforming a simple ASR hotword heuristic and a more conventional fine-tuning approach with Wav2Vec, an audio-only input architecture that is state-of-the-art for automatic speech recognition (ASR).
【2】"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
链接:https://arxiv.org/abs/2507.23365
摘要:我回顾了我在创作两张音乐专辑时的经历,这两张专辑都是以最先进的基于人工智能的音乐生成平台为中心的。第一张专辑明确提出了一个问题:当我的垃圾邮件与这些平台发生冲突时会发生什么?第二张专辑是对第一张专辑的直接回应,并与最先进的基于人工智能的音乐生成平台无法生成未经“练习”、“打磨”和“制作”的音乐的情况作了比较。我在一个大型语言模型(LLM)中植入了关于这些专辑的信息,并让它采访我,这导致了对几个更深层次问题的探索:我在多大程度上是作者?我在音乐中的位置?当我面对在某些方面比我更有天赋的机器时,我的音乐身份是如何改变的?我的作品为我或其他人/事物打开了什么新的音乐空间?最后,我反思我的反思,以及法学硕士介导的自我反思的方法。
摘要:I reflect on my experience creating two music albums centered on state-of-the-art prompt-based AI music generation platforms. The first album explicitly poses the question: What happens when I collide my junk mail with these platforms? The second album is a direct response to the first, and toys with the inability of state-of-the-art prompt-based AI music generation platforms to generate music that is not ``practiced'', ``polished'', and ``produced''. I seed a large language model (LLM) with information about these albums and have it interview me, which results in the exploration of several deeper questions: To what extent am I the author? Where am I in the resulting music? How is my musical identity changing as I am faced with machines that are in some ways far more talented than I? What new musical spaces does my work open, for me or anyone/thing else? I conclude by reflecting on my reflections, as well as LLM-mediated self-reflection as method.
【3】Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
标题:阿凡达专注聆听系统的各种类型点头的实时生成
链接:https://arxiv.org/abs/2507.23298
备注:Accepted by 27th ACM International Conference on Multimodal Interaction (ICMI '25), Long paper
摘要:在人类对话中,非语言信息(如点头和面部表情)与语言信息一样重要,口语对话系统也被期望表达此类非语言行为。我们专注于点头,这是一个专注的倾听系统的关键,并提出了一个模型,预测其时间和类型的实时。该模型建立在语音活动投影(VAP)模型的基础上,该模型可以从听者和扬声器音频中预测语音活动。我们将其扩展到预测不同类型的点头在一个连续的和实时的方式不同于传统的模型。此外,所提出的模型将多任务学习与言语反向通道预测和一般对话数据的预训练结合起来。在时间和类型预测任务中,多任务学习的有效性得到了显著的体现。我们证实,降低处理速率可以实现实时操作,而不会大幅降低准确性,并将该模型集成到一个化身专注倾听系统中。主观评价表明,它优于传统的方法,这总是点头同步与口头反向通道。代码和训练模型可在www.example.com上获得。
摘要:In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical in an attentive listening system, and propose a model that predicts both its timing and type in real time. The proposed model builds on the voice activity projection (VAP) model, which predicts voice activity from both listener and speaker audio. We extend it to prediction of various types of nodding in a continuous and real-time manner unlike conventional models. In addition, the proposed model incorporates multi-task learning with verbal backchannel prediction and pretraining on general dialogue data. In the timing and type prediction task, the effectiveness of multi-task learning was significantly demonstrated. We confirmed that reducing the processing rate enables real-time operation without a substantial drop in accuracy, and integrated the model into an avatar attentive listening system. Subjective evaluations showed that it outperformed the conventional method, which always does nodding in sync with verbal backchannel. The code and trained models are available at https://github.com/MaAI-Kyoto/MaAI.
【4】Moravec's Paradox: Towards an Auditory Turing Test
标题:莫拉韦茨悖论:走向听觉图灵测试
链接:https://arxiv.org/abs/2507.23091
摘要:这项研究工作表明,目前的人工智能系统在人类毫不费力地完成的听觉任务上失败了。从Moravec的悖论(即,对于人类来说简单的任务往往对机器来说很难,反之亦然),我们介绍了一个听觉图灵测试,包括七个类别的917个挑战:重叠语音,噪声中的语音,时间失真,空间音频,咖啡店噪声,电话失真和感知错觉。我们对最先进的音频模型(包括GPT-4的音频功能和OpenAI的Whisper)进行了评估,结果显示失败率超过93%,即使是性能最好的模型,在人类解决任务时的准确率也只有6.9%,而人类解决任务的成功率要高出7.5倍(52%)。这些结果暴露了人工智能系统如何处理复杂听觉场景的聚焦失败,特别是在选择性注意、噪声鲁棒性和上下文适应方面。我们的基准测试不仅量化了人机听觉差距,还提供了为什么会发生这些故障的见解,这表明当前的架构缺乏类似人类的听觉场景分析的基本机制。音频CAPTCHA的传统设计突出了人类进化而机器无法在多模态语言模型中选择的常见过滤器。这项工作建立了一个诊断框架,用于衡量人类水平的机器听力的进展,并强调了将选择性注意力、基于物理的音频理解和上下文感知集成到多模态AI系统中的新方法的必要性。
摘要:This research work demonstrates that current AI systems fail catastrophically on auditory tasks that humans perform effortlessly. Drawing inspiration from Moravec's paradox (i.e., tasks simple for humans often prove difficult for machines, and vice versa), we introduce an auditory Turing test comprising 917 challenges across seven categories: overlapping speech, speech in noise, temporal distortion, spatial audio, coffee-shop noise, phone distortion, and perceptual illusions. Our evaluation of state-of-the-art audio models including GPT-4's audio capabilities and OpenAI's Whisper reveals a striking failure rate exceeding 93%, with even the best-performing model achieving only 6.9% accuracy on tasks that humans solved at 7.5 times higher success (52%). These results expose focusing failures in how AI systems process complex auditory scenes, particularly in selective attention, noise robustness, and contextual adaptation. Our benchmark not only quantifies the human-machine auditory gap but also provides insights into why these failures occur, suggesting that current architectures lack fundamental mechanisms for human-like auditory scene analysis. The traditional design of audio CAPTCHAs highlights common filters that humans evolved but machines fail to select in multimodal language models. This work establishes a diagnostic framework for measuring progress toward human-level machine listening and highlights the need for novel approaches integrating selective attention, physics-based audio understanding, and context-aware perception into multimodal AI systems.
【5】Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
标题:研究多峰潜空间的可逆性:基于优化的方法的局限性
链接:https://arxiv.org/abs/2507.23010
摘要:本文研究了任务特定AI(人工智能)模型中多模态潜在空间的逆能力和更广泛的实用性。虽然这些模型在其设计的前瞻性任务(例如,文本到图像生成、音频到文本转录),但是它们用于逆映射的潜力仍然在很大程度上未被探索。我们提出了一个基于优化的框架,从所需的输出中推断输入特征,将其双向应用于文本图像(BLIP,Flux.1-dev)和文本音频(Whisper-Large-V3,Chatterbox-TTS)模式。 我们的中心假设是,虽然优化可以引导模型实现反向任务,但它们的多模态潜在空间不会始终如一地支持语义上有意义和感知上连贯的反向映射。实验结果一致验证了这一假设。我们证明,虽然优化可以迫使模型产生与目标文本一致的输出(例如,生成图像字幕模型正确描述的图像的文本到图像模型,或者准确转录优化音频的ASR模型),这些反转的感知质量是混乱和不连贯的。此外,当试图从生成模型中推断原始语义输入时,重建的潜在空间嵌入通常缺乏语义可解释性,与无意义的词汇标记对齐。 这些发现凸显了一个关键的局限性。主要针对特定前向任务优化的多模态潜在空间并不固有地具有鲁棒的和可解释的逆映射所需的结构。我们的工作强调了需要进一步研究开发真正语义丰富和可逆的多模态潜在空间。
摘要:This paper investigates the inverse capabilities and broader utility of multimodal latent spaces within task-specific AI (Artificial Intelligence) models. While these models excel at their designed forward tasks (e.g., text-to-image generation, audio-to-text transcription), their potential for inverse mappings remains largely unexplored. We propose an optimization-based framework to infer input characteristics from desired outputs, applying it bidirectionally across Text-Image (BLIP, Flux.1-dev) and Text-Audio (Whisper-Large-V3, Chatterbox-TTS) modalities. Our central hypothesis posits that while optimization can guide models towards inverse tasks, their multimodal latent spaces will not consistently support semantically meaningful and perceptually coherent inverse mappings. Experimental results consistently validate this hypothesis. We demonstrate that while optimization can force models to produce outputs that align textually with targets (e.g., a text-to-image model generating an image that an image captioning model describes correctly, or an ASR model transcribing optimized audio accurately), the perceptual quality of these inversions is chaotic and incoherent. Furthermore, when attempting to infer the original semantic input from generative models, the reconstructed latent space embeddings frequently lack semantic interpretability, aligning with nonsensical vocabulary tokens. These findings highlight a critical limitation. multimodal latent spaces, primarily optimized for specific forward tasks, do not inherently possess the structure required for robust and interpretable inverse mappings. Our work underscores the need for further research into developing truly semantically rich and invertible multimodal latent spaces.
【6】Balancing Information Preservation and Disentanglement in Self-Supervised Music Representation Learning
标题:自我监督音乐表示学习中的信息保存和解开平衡
链接:https://arxiv.org/abs/2507.22995
备注:In proceedings of WASPAA 2025. 4 pages, 4 figures, 1 table
摘要:自监督学习(SSL)方法的最新进展提供了一系列策略,用于从音乐音频中捕获有用的表示,而无需标记数据。虽然一些技术专注于通过重建来保留全面的细节,但其他技术则倾向于通过对比目标来实现语义结构。很少有作品在一个统一的SSL框架研究这些范例之间的相互作用。在这项工作中,我们提出了一个多视图的SSL框架解开音乐音频表示相结合的对比和重建的目标。该体系结构的目的是促进信息保真度和结构化的语义因素在解开子空间。我们进行了广泛的评价对比策略的设计选择,在受控设置中使用音乐音频表示。我们发现,虽然重建和对比策略表现出一致的权衡,当有效地结合起来,他们相辅相成,这使得解开的音乐属性,而不损害信息的完整性。
摘要:Recent advances in self-supervised learning (SSL) methods offer a range of strategies for capturing useful representations from music audio without the need for labeled data. While some techniques focus on preserving comprehensive details through reconstruction, others favor semantic structure via contrastive objectives. Few works examine the interaction between these paradigms in a unified SSL framework. In this work, we propose a multi-view SSL framework for disentangling music audio representations that combines contrastive and reconstructive objectives. The architecture is designed to promote both information fidelity and structured semantics of factors in disentangled subspaces. We perform an extensive evaluation on the design choices of contrastive strategies using music audio representations in a controlled setting. We find that while reconstruction and contrastive strategies exhibit consistent trade-offs, when combined effectively, they complement each other; this enables the disentanglement of music attributes without compromising information integrity.
【7】MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
标题:MECAT:针对细粒度音频理解任务的多专家构建的基准
链接:https://arxiv.org/abs/2507.23511
备注:9 main pages, 5 figures, 3 tables, and 14 appendix pages
摘要:虽然大型音频语言模型具有先进的开放式音频理解,但它们仍然无法达到人类水平的细微差别理解。这一差距之所以持续存在,主要是因为当前的基准测试受到数据注释和评估指标的限制,无法可靠地区分通用和高度详细的模型输出。为此,本文介绍了MECAT,一个用于细粒度音频理解任务的多专家构建基准。MECAT通过一个管道生成,该管道将专业专家模型的分析与思想链大型语言模型推理相集成,提供多视角,细粒度的标题和开放式问答对。该基准由一个新的度量标准补充:日期(歧视性增强音频文本评估)。该度量通过将单样本语义相似性与跨样本可区分性相结合来惩罚通用术语并奖励详细描述。本文还对最先进的音频模型进行了全面评估,为它们当前的能力和局限性提供了新的见解。数据和代码可在https://github.com/xiaomi-research/mecat上获得
摘要:While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation metrics, fail to reliably distinguish between generic and highly detailed model outputs. To this end, this work introduces MECAT, a Multi-Expert Constructed Benchmark for Fine-Grained Audio Understanding Tasks. Generated via a pipeline that integrates analysis from specialized expert models with Chain-of-Thought large language model reasoning, MECAT provides multi-perspective, fine-grained captions and open-set question-answering pairs. The benchmark is complemented by a novel metric: DATE (Discriminative-Enhanced Audio Text Evaluation). This metric penalizes generic terms and rewards detailed descriptions by combining single-sample semantic similarity with cross-sample discriminability. A comprehensive evaluation of state-of-the-art audio models is also presented, providing new insights into their current capabilities and limitations. The data and code are available at https://github.com/xiaomi-research/mecat
【8】CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
标题:CUHK-EE Systems应对NCMMSC 2025的虚拟挑战
链接:https://arxiv.org/abs/2507.23266
备注:Under review
摘要:本文介绍了香港中文大学电子工程系数字信号处理与语音技术实验室(DSP&STL)为参加第20届全国人机语音通信会议(NCMMSC 2025)语音挑战赛而开发的语音音色属性检测系统。所提出的系统利用WavLM-Large嵌入和精心的统计池来提取鲁棒的说话人表示,然后是Diff-Net的两个变体,即,前馈神经网络(FFN)和挤压和激励增强的残差FFN(SE-ResFFN),以比较发声对之间的音色属性强度。实验结果表明,WavLM-Large+FFN系统更好地推广到看不见的说话人,达到77.96%的准确率和21.79%的EER,而WavLM-Large+SE-ResFFN模型在“看到”设置中表现出色,准确率为94.42%,EER为5.49%。这些发现突出了模型复杂性和泛化之间的权衡,并强调了细粒度扬声器建模中架构选择的重要性。我们的分析还揭示了说话人身份,注释主观性和数据不平衡对系统性能的影响,指出了未来的方向,提高音色属性检测的鲁棒性和公平性。
摘要:This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Electronic Engineering (EE) at The Chinese University of Hong Kong (CUHK) for the 20th National Conference on Human-Computer Speech Communication (NCMMSC 2025) vTAD Challenge. The proposed systems leverage WavLM-Large embeddings with attentive statistical pooling to extract robust speaker representations, followed by two variants of Diff-Net, i.e., Feed-Forward Neural Network (FFN) and Squeeze-and-Excitation-enhanced Residual FFN (SE-ResFFN), to compare timbre attribute intensities between utterance pairs. Experimental results demonstrate that the WavLM-Large+FFN system generalises better to unseen speakers, achieving 77.96% accuracy and 21.79% EER, while the WavLM-Large+SE-ResFFN model excels in the 'Seen' setting with 94.42% accuracy and 5.49% EER. These findings highlight a trade-off between model complexity and generalisation, and underscore the importance of architectural choices in fine-grained speaker modelling. Our analysis also reveals the impact of speaker identity, annotation subjectivity, and data imbalance on system performance, pointing to future directions for improving robustness and fairness in timbre attribute detection.
【9】Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
标题:跨领域特征对改善助听器非侵入性语音可理解度预测的重要性
链接:https://arxiv.org/abs/2507.23223
备注:Accepted to Interspeech 2025
摘要:鉴于非侵入式语音清晰度评估在助听器(HA)中的关键作用,本文通过引入跨域特征重要性(FiDo)来提高其性能。我们估计频谱和时域声学特征以及耳语的潜在表示的特征重要性。每帧计算重要性权重,并根据这些权重,将特征投影到新的空间中,使模型能够尽早关注重要区域。接下来,在评估模块处理特征之前,执行特征串联以组合特征。实验结果表明,当FiDo被纳入改进的多分支语音可懂度模型MBI-Net+,RMSE可以降低7.62%(从26.10到24.11)。与2023年清晰度预测挑战赛中的最佳系统相比,具有FiDo的MBI-Net+还实现了3.98%的相对RMSE降低。这些结果验证了FiDo在增强HA神经言语评估中的有效性。
摘要:Given the critical role of non-intrusive speech intelligibility assessment in hearing aids (HA), this paper enhances its performance by introducing Feature Importance across Domains (FiDo). We estimate feature importance on spectral and time-domain acoustic features as well as latent representations of Whisper. Importance weights are calculated per frame, and based on these weights, features are projected into new spaces, allowing the model to focus on important areas early. Next, feature concatenation is performed to combine the features before the assessment module processes them. Experimental results show that when FiDo is incorporated into the improved multi-branched speech intelligibility model MBI-Net+, RMSE can be reduced by 7.62% (from 26.10 to 24.11). MBI-Net+ with FiDo also achieves a relative RMSE reduction of 3.98% compared to the best system in the 2023 Clarity Prediction Challenge. These results validate FiDo's effectiveness in enhancing neural speech assessment in HA.
【10】Exploring Dynamic Parameters for Vietnamese Gender-Independent ASR
标题:探索越南与性别无关的ASB的动态参数
链接:https://arxiv.org/abs/2507.22964
备注:None
摘要:语音信号的动态特性提供了语音信号的时域信息,对提高语音识别能力具有重要作用。在这项工作中,我们的特点是在一个比率平面的频谱子带质心频率(SSCFs)使用极参数捕捉语音的动态特性,并尽量减少频谱变化的声学过渡。这些动态参数结合梅尔频率倒谱系数(MFCC)在越南ASR捕捉更详细的光谱信息。SSCF 0被用作基频(F0)的伪特征,以鲁棒地描述音调信息。研究结果表明,建议的参数显着降低字错误率,并表现出更大的性别独立性比基线MFCC。
摘要:The dynamic characteristics of speech signal provides temporal information and play an important role in enhancing Automatic Speech Recognition (ASR). In this work, we characterized the acoustic transitions in a ratio plane of Spectral Subband Centroid Frequencies (SSCFs) using polar parameters to capture the dynamic characteristics of the speech and minimize spectral variation. These dynamic parameters were combined with Mel-Frequency Cepstral Coefficients (MFCCs) in Vietnamese ASR to capture more detailed spectral information. The SSCF0 was used as a pseudo-feature for the fundamental frequency (F0) to describe the tonal information robustly. The findings showed that the proposed parameters significantly reduce word error rates and exhibit greater gender independence than the baseline MFCCs.
【1】MECAT: A Multi-Experts Constructed Benchmark for Fine-Grained Audio Understanding Tasks
标题:MECAT:针对细粒度音频理解任务的多专家构建的基准
链接:https://arxiv.org/abs/2507.23511
备注:9 main pages, 5 figures, 3 tables, and 14 appendix pages
摘要:虽然大型音频语言模型具有先进的开放式音频理解,但它们仍然无法达到人类水平的细微差别理解。这一差距之所以持续存在,主要是因为当前的基准测试受到数据注释和评估指标的限制,无法可靠地区分通用和高度详细的模型输出。为此,本文介绍了MECAT,一个用于细粒度音频理解任务的多专家构建基准。MECAT通过一个管道生成,该管道将专业专家模型的分析与思想链大型语言模型推理相集成,提供多视角,细粒度的标题和开放式问答对。该基准由一个新的度量标准补充:日期(歧视性增强音频文本评估)。该度量通过将单样本语义相似性与跨样本可区分性相结合来惩罚通用术语并奖励详细描述。本文还对最先进的音频模型进行了全面评估,为它们当前的能力和局限性提供了新的见解。数据和代码可在https://github.com/xiaomi-research/mecat上获得
摘要:While large audio-language models have advanced open-ended audio understanding, they still fall short of nuanced human-level comprehension. This gap persists largely because current benchmarks, limited by data annotations and evaluation metrics, fail to reliably distinguish between generic and highly detailed model outputs. To this end, this work introduces MECAT, a Multi-Expert Constructed Benchmark for Fine-Grained Audio Understanding Tasks. Generated via a pipeline that integrates analysis from specialized expert models with Chain-of-Thought large language model reasoning, MECAT provides multi-perspective, fine-grained captions and open-set question-answering pairs. The benchmark is complemented by a novel metric: DATE (Discriminative-Enhanced Audio Text Evaluation). This metric penalizes generic terms and rewards detailed descriptions by combining single-sample semantic similarity with cross-sample discriminability. A comprehensive evaluation of state-of-the-art audio models is also presented, providing new insights into their current capabilities and limitations. The data and code are available at https://github.com/xiaomi-research/mecat
【2】CUHK-EE Systems for the vTAD Challenge at NCMMSC 2025
标题:CUHK-EE Systems应对NCMMSC 2025的虚拟挑战
链接:https://arxiv.org/abs/2507.23266
备注:Under review
摘要:本文介绍了香港中文大学电子工程系数字信号处理与语音技术实验室(DSP&STL)为参加第20届全国人机语音通信会议(NCMMSC 2025)语音挑战赛而开发的语音音色属性检测系统。所提出的系统利用WavLM-Large嵌入和统计注意池来提取鲁棒的说话者表示,然后是Diff-Net的两个变体,即,前馈神经网络(FFN)和挤压和激励增强的残差FFN(SE-ResFFN),以比较发声对之间的音色属性强度。实验结果表明,WavLM-Large+FFN系统更好地推广到看不见的说话人,达到77.96%的准确率和21.79%的EER,而WavLM-Large+SE-ResFFN模型在“看到”设置中表现出色,准确率为94.42%,EER为5.49%。这些发现突出了模型复杂性和泛化之间的权衡,并强调了细粒度扬声器建模中架构选择的重要性。我们的分析还揭示了说话人身份,注释主观性和数据不平衡对系统性能的影响,指出了未来的方向,提高音色属性检测的鲁棒性和公平性。
摘要:This paper presents the Voice Timbre Attribute Detection (vTAD) systems developed by the Digital Signal Processing & Speech Technology Laboratory (DSP&STL) of the Department of Electronic Engineering (EE) at The Chinese University of Hong Kong (CUHK) for the 20th National Conference on Human-Computer Speech Communication (NCMMSC 2025) vTAD Challenge. The proposed systems leverage WavLM-Large embeddings with attentive statistical pooling to extract robust speaker representations, followed by two variants of Diff-Net, i.e., Feed-Forward Neural Network (FFN) and Squeeze-and-Excitation-enhanced Residual FFN (SE-ResFFN), to compare timbre attribute intensities between utterance pairs. Experimental results demonstrate that the WavLM-Large+FFN system generalises better to unseen speakers, achieving 77.96% accuracy and 21.79% EER, while the WavLM-Large+SE-ResFFN model excels in the 'Seen' setting with 94.42% accuracy and 5.49% EER. These findings highlight a trade-off between model complexity and generalisation, and underscore the importance of architectural choices in fine-grained speaker modelling. Our analysis also reveals the impact of speaker identity, annotation subjectivity, and data imbalance on system performance, pointing to future directions for improving robustness and fairness in timbre attribute detection.
【3】Feature Importance across Domains for Improving Non-Intrusive Speech Intelligibility Prediction in Hearing Aids
标题:跨领域特征对改善助听器非侵入性语音可理解度预测的重要性
链接:https://arxiv.org/abs/2507.23223
备注:Accepted to Interspeech 2025
摘要:鉴于非侵入式语音清晰度评估在助听器(HA)中的关键作用,本文通过引入跨域特征重要性(FiDo)来提高其性能。我们估计频谱和时域声学特征以及耳语的潜在表示的特征重要性。每帧计算重要性权重,并根据这些权重,将特征投影到新的空间中,使模型能够尽早关注重要区域。接下来,在评估模块处理特征之前,执行特征级联以组合特征。实验结果表明,当FiDo被纳入改进的多分支语音可懂度模型MBI-Net+,RMSE可以降低7.62%(从26.10到24.11)。与2023年清晰度预测挑战赛中的最佳系统相比,具有FiDo的MBI-Net+还实现了3.98%的相对RMSE降低。这些结果验证了FiDo在增强HA神经言语评估中的有效性。
摘要:Given the critical role of non-intrusive speech intelligibility assessment in hearing aids (HA), this paper enhances its performance by introducing Feature Importance across Domains (FiDo). We estimate feature importance on spectral and time-domain acoustic features as well as latent representations of Whisper. Importance weights are calculated per frame, and based on these weights, features are projected into new spaces, allowing the model to focus on important areas early. Next, feature concatenation is performed to combine the features before the assessment module processes them. Experimental results show that when FiDo is incorporated into the improved multi-branched speech intelligibility model MBI-Net+, RMSE can be reduced by 7.62% (from 26.10 to 24.11). MBI-Net+ with FiDo also achieves a relative RMSE reduction of 3.98% compared to the best system in the 2023 Clarity Prediction Challenge. These results validate FiDo's effectiveness in enhancing neural speech assessment in HA.
【4】Full-Duplex-Bench v1.5: Evaluating Overlap Handling for Full-Duplex Speech Models
标题:全速工作台v1.5:评估全速语音模型的重叠处理
链接:https://arxiv.org/abs/2507.23159
备注:Work in Progress
摘要:虽然全双工语音代理承诺自然,低延迟的人机交互,同时处理输入和输出语音,重叠管理仍然低估。我们介绍了Full-Duplex-Bench v1.5,这是一个模块化的,全自动的基准测试,它模拟了四种重叠的场景:用户中断,听众反向通道,侧对话和环境语音。我们的框架支持开源和商业模型,提供了一个全面的,可扩展的度量套件-分类对话行为,停止和响应延迟,韵律适应和感知语音质量-可以根据应用程序特定的标准进行定制。基准五个国家的最先进的代理人揭示了两个主要的战略:修复第一快速收益率与连续性第一持续流,并突出了依赖于网络的性能趋势。这种开源设计可以无缝扩展新的音频资产、语言和部署环境,使从业人员能够自定义和加速对强大的全双工语音系统的评估。
摘要:While full-duplex speech agents promise natural, low-latency human--machine interaction by concurrently processing input and output speech, overlap management remains under-evaluated. We introduce Full-Duplex-Bench v1.5, a modular, fully automated benchmark that simulates four overlap scenarios: user interruption, listener backchannel, side conversation, and ambient speech. Our framework supports both open-sourced and commercial models, offering a comprehensive, extensible metric suite -- categorical dialogue behaviors, stop and response latency, prosodic adaptation, and perceived speech quality -- that can be tailored to application-specific criteria. Benchmarking five state-of-the-art agents reveals two principal strategies: repair-first rapid yielding versus continuity-first sustained flow, and highlights scenario-dependent performance trends. The open-sourced design enables seamless extension with new audio assets, languages, and deployment contexts, empowering practitioners to customize and accelerate the evaluation of robust full-duplex speech systems.
【5】Exploring Dynamic Parameters for Vietnamese Gender-Independent ASR
标题:探索越南与性别无关的ASB的动态参数
链接:https://arxiv.org/abs/2507.22964
备注:None
摘要:语音信号的动态特性提供了语音信号的时域信息,对提高语音识别能力具有重要作用。在这项工作中,我们的特点是在一个比率平面的频谱子带质心频率(SSCFs)使用极参数捕捉语音的动态特性,并尽量减少频谱变化的声学过渡。这些动态参数结合梅尔频率倒谱系数(MFCC)在越南ASR捕捉更详细的光谱信息。SSCF 0被用作基频(F0)的伪特征,以鲁棒地描述音调信息。研究结果表明,与基线MFCC相比,拟议的参数显着降低了单词错误率,并表现出更大的性别独立性。
摘要:The dynamic characteristics of speech signal provides temporal information and play an important role in enhancing Automatic Speech Recognition (ASR). In this work, we characterized the acoustic transitions in a ratio plane of Spectral Subband Centroid Frequencies (SSCFs) using polar parameters to capture the dynamic characteristics of the speech and minimize spectral variation. These dynamic parameters were combined with Mel-Frequency Cepstral Coefficients (MFCCs) in Vietnamese ASR to capture more detailed spectral information. The SSCF0 was used as a pseudo-feature for the fundamental frequency (F0) to describe the tonal information robustly. The findings showed that the proposed parameters significantly reduce word error rates and exhibit greater gender independence than the baseline MFCCs.
【6】Identifying Hearing Difficulty Moments in Conversational Audio
标题:识别对话音频中的听力困难时刻
链接:https://arxiv.org/abs/2507.23590
摘要:人们在日常谈话中经常遇到听力困难的时刻。识别这些听力困难的时刻在听力辅助技术领域具有特殊意义,其中及时干预是实时听力辅助的关键。在本文中,我们提出并比较了机器学习解决方案,用于连续检测识别会话音频中这些特定时刻的话语。我们表明,音频语言模型通过其多模态推理能力,在这项任务中表现出色,显着优于简单的ASR热词启发式算法和使用Wav 2 Vec(一种最先进的纯音频输入架构)的更传统的微调方法。自动语音识别(ASR)的最新技术。
摘要:Individuals regularly experience Hearing Difficulty Moments in everyday conversation. Identifying these moments of hearing difficulty has particular significance in the field of hearing assistive technology where timely interventions are key for realtime hearing assistance. In this paper, we propose and compare machine learning solutions for continuously detecting utterances that identify these specific moments in conversational audio. We show that audio language models, through their multimodal reasoning capabilities, excel at this task, significantly outperforming a simple ASR hotword heuristic and a more conventional fine-tuning approach with Wav2Vec, an audio-only input architecture that is state-of-the-art for automatic speech recognition (ASR).
【7】"I made this (sort of)": Negotiating authorship, confronting fraudulence, and exploring new musical spaces with prompt-based AI music generation
链接:https://arxiv.org/abs/2507.23365
摘要:我回顾了我在创作两张音乐专辑时的经历,这两张专辑都是以最先进的基于人工智能的音乐生成平台为中心的。第一张专辑明确提出了一个问题:当我的垃圾邮件与这些平台发生冲突时会发生什么?第二张专辑是对第一张专辑的直接回应,并与最先进的基于人工智能的音乐生成平台无法生成未经“练习”、“打磨”和“制作”的音乐的情况作了比较。我在一个大型语言模型(LLM)中植入了关于这些专辑的信息,并让它采访我,这导致了对几个更深层次问题的探索:我在多大程度上是作者?我在音乐中的位置?当我面对在某些方面比我更有天赋的机器时,我的音乐身份是如何改变的?我的作品为我或其他人/事物打开了什么新的音乐空间?最后,我反思我的反思,以及法学硕士介导的自我反思的方法。
摘要:I reflect on my experience creating two music albums centered on state-of-the-art prompt-based AI music generation platforms. The first album explicitly poses the question: What happens when I collide my junk mail with these platforms? The second album is a direct response to the first, and toys with the inability of state-of-the-art prompt-based AI music generation platforms to generate music that is not ``practiced'', ``polished'', and ``produced''. I seed a large language model (LLM) with information about these albums and have it interview me, which results in the exploration of several deeper questions: To what extent am I the author? Where am I in the resulting music? How is my musical identity changing as I am faced with machines that are in some ways far more talented than I? What new musical spaces does my work open, for me or anyone/thing else? I conclude by reflecting on my reflections, as well as LLM-mediated self-reflection as method.
【8】Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System
标题:阿凡达专注聆听系统的各种类型点头的实时生成
链接:https://arxiv.org/abs/2507.23298
备注:Accepted by 27th ACM International Conference on Multimodal Interaction (ICMI '25), Long paper
摘要:在人类对话中,非语言信息(如点头和面部表情)与语言信息一样重要,口语对话系统也被期望表达此类非语言行为。我们专注于点头,这是一个专注的倾听系统的关键,并提出了一个模型,预测其时间和类型的实时。该模型建立在语音活动投影(VAP)模型的基础上,该模型可以从听者和扬声器音频中预测语音活动。我们将其扩展到预测不同类型的点头在一个连续的和实时的方式不同于传统的模型。此外,该模型将多任务学习与口头反向通道预测和一般对话数据的预训练相结合。在时间和类型预测任务中,多任务学习的有效性得到了显著的体现。我们证实,降低处理速率可以实现实时操作,而不会大幅降低准确性,并将该模型集成到一个化身专注倾听系统中。主观评价表明,它优于传统的方法,这总是点头同步与口头反向通道。代码和训练模型可在https://github.com/MaAI-Kyoto/MaAI上获得。
摘要:In human dialogue, nonverbal information such as nodding and facial expressions is as crucial as verbal information, and spoken dialogue systems are also expected to express such nonverbal behaviors. We focus on nodding, which is critical in an attentive listening system, and propose a model that predicts both its timing and type in real time. The proposed model builds on the voice activity projection (VAP) model, which predicts voice activity from both listener and speaker audio. We extend it to prediction of various types of nodding in a continuous and real-time manner unlike conventional models. In addition, the proposed model incorporates multi-task learning with verbal backchannel prediction and pretraining on general dialogue data. In the timing and type prediction task, the effectiveness of multi-task learning was significantly demonstrated. We confirmed that reducing the processing rate enables real-time operation without a substantial drop in accuracy, and integrated the model into an avatar attentive listening system. Subjective evaluations showed that it outperformed the conventional method, which always does nodding in sync with verbal backchannel. The code and trained models are available at https://github.com/MaAI-Kyoto/MaAI.
【9】SequenceLayers: Sequence Processing and Streaming Neural Networks Made Easy
标题:SequenceLayers:简化序列处理和流神经网络
链接:https://arxiv.org/abs/2507.23292
摘要:我们引入了一个神经网络层API和序列建模库,旨在轻松创建可以逐层执行的序列模型(例如,教师强制培训)和逐步(例如,自回归采样)。为了实现这一点,层定义它们随时间的状态的显式表示(例如,一个Transformer KV缓存,一个卷积缓冲区,一个RNN隐藏状态),以及一个演化该状态的step方法,测试为给无状态逐层调用提供相同的结果。SequenceLayers合约的这一点和其他方面使复杂模型能够立即流式传输,减轻了流式传输和并行序列处理中出现的各种常见错误,并且可以在任何深度学习库中实现。可组合和声明式API以及一套全面的层和组合子,简化了从简单的流组件构建生产规模模型的过程,同时保持了强大的正确性保证。我们目前的SequenceLayers实现(JAX,TensorFlow 2)可在https://github.com/google/sequence-layers上获得。
摘要:We introduce a neural network layer API and library for sequence modeling, designed for easy creation of sequence models that can be executed both layer-by-layer (e.g., teacher-forced training) and step-by-step (e.g., autoregressive sampling). To achieve this, layers define an explicit representation of their state over time (e.g., a Transformer KV cache, a convolution buffer, an RNN hidden state), and a step method that evolves that state, tested to give identical results to a stateless layer-wise invocation. This and other aspects of the SequenceLayers contract enables complex models to be immediately streamable, mitigates a wide range of common bugs arising in both streaming and parallel sequence processing, and can be implemented in any deep learning library. A composable and declarative API, along with a comprehensive suite of layers and combinators, streamlines the construction of production-scale models from simple streamable components while preserving strong correctness guarantees. Our current implementations of SequenceLayers (JAX, TensorFlow 2) are available at https://github.com/google/sequence-layers.
【10】Moravec's Paradox: Towards an Auditory Turing Test
标题:莫拉韦茨悖论:走向听觉图灵测试
链接:https://arxiv.org/abs/2507.23091
摘要:这项研究工作表明,目前的人工智能系统在人类毫不费力地完成的听觉任务上失败了。从Moravec的悖论(即,对于人类来说简单的任务往往对机器来说很难,反之亦然),我们介绍了一个听觉图灵测试,包括七个类别的917个挑战:重叠语音,噪声中的语音,时间失真,空间音频,咖啡店噪声,电话失真和感知错觉。我们对最先进的音频模型(包括GPT-4的音频功能和OpenAI的Whisper)进行了评估,结果显示失败率超过93%,即使是性能最好的模型,在人类解决任务时的准确率也只有6.9%,而人类解决任务的成功率要高出7.5倍(52%)。这些结果暴露了人工智能系统如何处理复杂听觉场景的聚焦失败,特别是在选择性注意、噪声鲁棒性和上下文适应方面。我们的基准测试不仅量化了人机听觉差距,还提供了为什么会发生这些故障的见解,这表明当前的架构缺乏类似人类的听觉场景分析的基本机制。音频CAPTCHA的传统设计突出了人类进化而机器无法在多模态语言模型中选择的常见过滤器。这项工作建立了一个诊断框架,用于衡量人类水平的机器听力的进展,并强调了将选择性注意力、基于物理的音频理解和上下文感知集成到多模态AI系统中的新方法的必要性。
摘要:This research work demonstrates that current AI systems fail catastrophically on auditory tasks that humans perform effortlessly. Drawing inspiration from Moravec's paradox (i.e., tasks simple for humans often prove difficult for machines, and vice versa), we introduce an auditory Turing test comprising 917 challenges across seven categories: overlapping speech, speech in noise, temporal distortion, spatial audio, coffee-shop noise, phone distortion, and perceptual illusions. Our evaluation of state-of-the-art audio models including GPT-4's audio capabilities and OpenAI's Whisper reveals a striking failure rate exceeding 93%, with even the best-performing model achieving only 6.9% accuracy on tasks that humans solved at 7.5 times higher success (52%). These results expose focusing failures in how AI systems process complex auditory scenes, particularly in selective attention, noise robustness, and contextual adaptation. Our benchmark not only quantifies the human-machine auditory gap but also provides insights into why these failures occur, suggesting that current architectures lack fundamental mechanisms for human-like auditory scene analysis. The traditional design of audio CAPTCHAs highlights common filters that humans evolved but machines fail to select in multimodal language models. This work establishes a diagnostic framework for measuring progress toward human-level machine listening and highlights the need for novel approaches integrating selective attention, physics-based audio understanding, and context-aware perception into multimodal AI systems.
【11】Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods
标题:研究多峰潜空间的可逆性:基于优化的方法的局限性
链接:https://arxiv.org/abs/2507.23010
摘要:本文研究了任务特定AI(人工智能)模型中多模态潜在空间的逆能力和更广泛的实用性。虽然这些模型在其设计的前瞻性任务(例如,文本到图像生成、音频到文本转录),但是它们用于逆映射的潜力仍然在很大程度上未被探索。我们提出了一个基于优化的框架,从所需的输出中推断输入特征,将其双向应用于文本图像(BLIP,Flux.1-dev)和文本音频(Whisper-Large-V3,Chatterbox-TTS)模式。 我们的中心假设是,虽然优化可以引导模型实现反向任务,但它们的多模态潜在空间不会始终如一地支持语义上有意义和感知上连贯的反向映射。实验结果一致验证了这一假设。我们证明,虽然优化可以迫使模型产生与目标文本一致的输出(例如,生成图像字幕模型正确描述的图像的文本到图像模型,或者准确转录优化音频的ASR模型),这些反转的感知质量是混乱和不连贯的。此外,当试图从生成模型中推断原始语义输入时,重建的潜在空间嵌入通常缺乏语义可解释性,与无意义的词汇标记对齐。 这些发现突出了一个关键的局限性。主要针对特定前向任务优化的多模态潜在空间并不固有地具有鲁棒的和可解释的逆映射所需的结构。我们的工作强调了需要进一步研究开发真正语义丰富和可逆的多模态潜在空间。
摘要:This paper investigates the inverse capabilities and broader utility of multimodal latent spaces within task-specific AI (Artificial Intelligence) models. While these models excel at their designed forward tasks (e.g., text-to-image generation, audio-to-text transcription), their potential for inverse mappings remains largely unexplored. We propose an optimization-based framework to infer input characteristics from desired outputs, applying it bidirectionally across Text-Image (BLIP, Flux.1-dev) and Text-Audio (Whisper-Large-V3, Chatterbox-TTS) modalities. Our central hypothesis posits that while optimization can guide models towards inverse tasks, their multimodal latent spaces will not consistently support semantically meaningful and perceptually coherent inverse mappings. Experimental results consistently validate this hypothesis. We demonstrate that while optimization can force models to produce outputs that align textually with targets (e.g., a text-to-image model generating an image that an image captioning model describes correctly, or an ASR model transcribing optimized audio accurately), the perceptual quality of these inversions is chaotic and incoherent. Furthermore, when attempting to infer the original semantic input from generative models, the reconstructed latent space embeddings frequently lack semantic interpretability, aligning with nonsensical vocabulary tokens. These findings highlight a critical limitation. multimodal latent spaces, primarily optimized for specific forward tasks, do not inherently possess the structure required for robust and interpretable inverse mappings. Our work underscores the need for further research into developing truly semantically rich and invertible multimodal latent spaces.
机器翻译由腾讯交互翻译提供,仅供参考
