微信公众号:arXiv_Daily
cs.SD语音
标题:MerIT:学习分解的音乐表示以获得音频相似性
链接:https://arxiv.org/abs/2605.27346
摘要:当前的音乐相似性模型通常计算单个整体乐谱,纠缠不同的音乐维度,如旋律,节奏和音色。这限制了用户的控制和可解释性,使其无法执行细微的查询。我们介绍MERIT,一个框架,学习解开,特定因素的音乐表示量身定制的这三个核心尺寸。为了克服现实世界音频中缺乏孤立的音乐变化,我们使用了一种新的训练策略,该策略使用条件音频生成和源分离的茎来强烈鼓励训练数据中的单因素变化。我们的评估表明了强有力的因素分解。每个头部强烈响应其预期的感知维度,同时保持对其他头部的接近机会,这是一种在合成训练域和独立的真实世界音频中保持的代表性属性。
摘要:Current music similarity models typically compute a single, monolithic score, entangling distinct musical dimensions like melody, rhythm, and timbre. This limits user control and interpretability, making it impossible to execute nuanced queries. We introduce MERIT, a framework for learning disentangled, factor-specific music representations tailored to these three core dimensions. To overcome the lack of isolated musical variations in real-world audio, we use a novel training strategy that uses conditional audio generation and source-separated stems to strongly encourage single-factor variation in training data. Our evaluations demonstrate strong factor-wise disentanglement. Each head responds strongly to its intended perceptual dimension while remaining near chance on the others, a representational property that holds across both the synthetic training domain and independent real-world audio.
【2】PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis
标题:PilotTTC:一种用于竞争性语音合成的定制模块化方案链接:https://arxiv.org/abs/2605.27258
摘要:
摘要:
【3】Learning When to Think While Listening in Large Audio-Language Models
标题:在大型音频语言模型中聆听时学习何时思考链接:https://arxiv.org/abs/2605.27190
备注:19 pages, 4 figures, 6 tables
摘要:
摘要:
【4】Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy
标题:超越二进制:跨认知分数层次的语音表示链接:https://arxiv.org/abs/2605.27189
摘要:
摘要:
【5】An investigation of AI integration in sound designer workflows and experiences
标题:声音设计师工作流程和体验中人工智能集成的调查链接:https://arxiv.org/abs/2605.27174
摘要:
摘要:
【6】PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech
标题:PashtoTTS-Bench:低资源非拉丁字母文本到语音的自动筛选链接:https://arxiv.org/abs/2605.26978
摘要:
摘要:
【7】Can We Hear from Events? Generating Speech from Event Camera
标题:我们能听到事件的消息吗?从事件摄像机生成语音链接:https://arxiv.org/abs/2605.26672
摘要:
摘要:
【8】LongAV-Compass: Towards Unified Evaluation of Minute-Scale Audio-Visual Generation Across T2AV, I2AV, and V2AV
标题:LongAV-Compass:迈向T2 AV、I2 AV和V2 AV分钟级视听生成的统一评估链接:https://arxiv.org/abs/2605.26244
摘要:
摘要:
【9】DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture Generation
标题:DuoGesture:受神经启发和生物力学启发的双流联合语音手势生成链接:https://arxiv.org/abs/2605.26236
摘要:
摘要:
【10】PitchBench: Measuring Pitch Hearing in Audio-Language Models
标题:PitchBench:测量音频语言模型中的音调听力链接:https://arxiv.org/abs/2605.26176
备注:Preprint
摘要:
摘要:
【11】Eroding Trust in Real Speech: A Large-Scale Study of Human Audio Deepfake Perception
标题:侵蚀对真实语音的信任:人类音频Deepfake感知的大规模研究链接:https://arxiv.org/abs/2605.26136
摘要:
摘要:
【12】Why Can't They Remember? Uncovering Representation and Retrieval Bottlenecks in Multi-Turn Acoustic Memory
标题:为什么他们记不住了?揭开多圈声记忆的表达和检索瓶颈链接:https://arxiv.org/abs/2605.27039
摘要:
摘要:
标题:为什么他们记不住了?揭开多圈声记忆的表达和检索瓶颈
链接:https://arxiv.org/abs/2605.27039
摘要:大型音频语言模型(LALM)处理语音和环境声学线索,但难以在多轮交互中保留非语音信息。语义(语音)和声学(非语音)理解之间的性能差距仍然知之甚少,表示和检索的潜在机制仍然不清楚。这项工作介绍了EnvMem,一个受控的多圈基准,旨在研究这一差距,并确定故障的根本原因在代表(即,潜在嵌入)和检索级别(即,注意分配)。我们进一步进行事后干预,以探测表征结构和注意力动态。我们的结果揭示了代表性轨迹漂移是关键的失效模式,同时表明注意力分配在解释观察到的退化方面发挥的作用有限。总的来说,我们提供了一个系统的框架,用于分析和改善长上下文LALM中的非语言记忆,揭示未来的数据和训练设计,以实现强大的声学记忆建模。
摘要:Large audio language models (LALMs) process both speech and environmental acoustic cues, yet struggle to retain non-speech information across multi-turn interactions. The performance gap between semantic (speech) and acoustic (non-speech) understanding remains poorly understood, and the underlying mechanisms of representation and retrieval are still unclear. This work introduces EnvMem, a controlled multi-turn benchmark designed to study this gap and identify the root causes of failures at the representation (i.e., latent embeddings) and retrieval levels (i.e., attention allocation). We further conduct post-hoc interventions to probe representational structure and attention dynamics. Our results reveal representational trajectory drift as the key failure mode, while showing that attention allocation plays a limited role in explaining the observed degradation. Overall, we provide a systematic framework for analyzing and improving non-linguistic memory in long-context LALMs, shedding light on future data and training design for robust acoustic memory modeling.
【2】CFMDCTCodec: A Low-Bitrate Neural Speech Codec with Noise-Prior-aware Conditional Flow Matching for MDCT-Spectral Enhancement
标题:CFMDCTCodec:一种具有噪音先验感知条件流匹配的低比特率神经语音编解码器,用于IDT频谱增强链接:https://arxiv.org/abs/2605.26812
备注:Accepted by IEEE Transactions on Audio, Speech and Language Processing
摘要:低比特率下的高质量语音编码对于带宽受限的应用至关重要,但由于高度压缩表示中的质量关键信息的严重丢失,因此仍然具有挑战性。为了克服这一挑战,我们提出CFMDCTCodec,一个低比特率的神经语音编解码器,完全在修改后的离散余弦变换(MDCT)域。CFMDCTCodec集成了一个轻量级的编码器-量化器-解码器风格的MDCT频谱编解码器与噪声先验感知的、基于条件流匹配(CFM)的MDCT频谱增强器。在这个框架内,编解码器作为一个基本模块,competencydiscretizes的MDCT频谱提取的语音,并产生一个初始的粗重建,而增强器进一步恢复细粒度的频谱细节。增强器通过将条件MDCT速度场滤波器与常微分方程(ODE)求解器集成来改进解码MDCT频谱,在MDCT导出的幅度自适应噪声先验的指导下,旨在强调感知上显著的高能量区域,同时稳定低能量和无声区域。最后,增强的MDCT频谱重建到解码语音使用逆MDCT。在优化CFMDCT Codec时,我们采用了统一的非对抗性训练策略,将重建、量化和CFM目标结合起来。客观和主观评估都表明,CFMDCT编解码器在低比特率制度中优于竞争基线,例如,0.65 kbps,同时接近具有显著更少的参数和计算的大规模编解码器的感知质量。
摘要:High-quality speech coding at low bitrates is crucial for bandwidth-constrained applications, yet remains challenging due to the severe loss of quality-critical information in highly compressed representations. To overcome this challenge, we propose CFMDCTCodec, a low-bitrate neural speech codec that operates entirely in the modified discrete cosine transform (MDCT) domain. CFMDCTCodec integrates a lightweight encoder-quantizer-decoder-style MDCT-spectral codec with a noise-prior-aware, conditional-flow-matching (CFM)-based MDCT-spectral enhancer. Within this framework, the codec serves as a base module that compactly discretizes the MDCT spectrum extracted from speech and produces an initial coarse reconstruction, while the enhancer further restores fine-grained spectral details. The enhancer improves the decoded MDCT spectrum by integrating a conditional MDCT velocity-field filter with an ordinary differential equation (ODE) solver, under the guidance of an MDCT-derived magnitude-adaptive noise prior, aiming to emphasize perceptually significant high-energy regions while stabilizing low-energy and silent regions. Finally, the enhanced MDCT spectrum is reconstructed into the decoded speech using the inverse MDCT. When optimizing CFMDCTCodec, we adopt a unified non-adversarial training strategy that jointly combines reconstruction, quantization and CFM objectives. Both objective and subjective evaluations show that CFMDCTCodec outperforms competitive baselines in low-bitrate regimes, e.g., 0.65 kbps, while approaching the perceptual quality of large-scale codecs with significantly fewer parameters and computations.
【3】Beyond Binary: Speech Representations Across the Cognitive Score Hierarchy
标题:超越二进制:跨认知分数层次的语音表示链接:https://arxiv.org/abs/2605.27189
摘要:本研究旨在探讨轻度认知功能障碍患者言语表征与认知功能评定等级结构之间的关系。利用5,754个德国神经心理学评估记录,我们在三个评分水平上评估了六个认知任务:任务,域和全局水平。我们将手工制作的声学特征与自监督学习(SSL)嵌入进行了比较。结果表明,虽然SSL表示一般优于手工制作的功能在较低的水平,MCI分类的趋势逆转。此外,特定任务的约束影响性能:具有更大响应自由度的任务表现出性能稀释的层次结构的水平增加,建议“专家”的代表,而性能的高度结构化的任务增加向更高的水平,建议“通才”的代表。这些研究结果表明,在自动临床语音分析的任务约束和评估层次之间的联系。
摘要:This study examines the relationship between speech representations and the hierarchical structure of cognitive assessment in mild cognitive impairment. Utilizing 5,754 German neuropsychological assessment recordings, we evaluate six cognitive tasks across three score levels: task, domain, and global levels. We compare hand-crafted acoustic features with self-supervised learning (SSL) embeddings. Results show that although SSL representations generally outperform hand-crafted features at lower levels, this trend reverses for MCI classification. Furthermore, task-specific constraints influence performance: tasks with greater response freedom exhibit performance dilution as hierarchical levels increase, suggesting ``specialist'' representations, whereas the performance of highly structured tasks increases toward higher levels, suggesting ``generalist'' representations. These findings show links between task constraints and assessment hierarchy in automated clinical speech analysis.
机器翻译由腾讯交互翻译提供,仅供参考
