微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 语音合成与声音生成 2 篇
3. 语音增强、降噪与音频修复 1 篇
4. 安全、隐私与深度伪造音频 1 篇
5. 其他/综合语音音频 3 篇
1. 语音识别与关键词检测 | 1 篇
1. An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge
用于第二届MLC-SLM挑战赛的基于全语言自动语音识别的语音语言模型系统
AI 总结:该研究针对第二届MLC-SLM挑战赛任务1,提出级联分帧识别系统,结合多种技术,测试无先验信息。分析工程选择影响,如基于嵌入的聚类更优,重叠感知分割虽提召回率但增加tcpMER,在开发集和评估集有相应表现。
链接:https://arxiv.org/abs/2607.12468
作者:Shuming Fang, Shuifei Zeng
英文摘要:We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time. On the official Development set (150 conversations, 21 language/accent categories) the system attains a macro tcpMER of 29.27%, versus 79.15% for the official baseline; on the Evaluation set it scores 50.23%. We also analyze two engineering choices that substantially affect tcpMER. First, embedding-based speaker clustering outperforms an end-to-end-style alternative that assigns speakers from ASR
2. 语音合成与声音生成 | 2 篇
2. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
AutoSIFT:用于可控语音生成的自动风格筛选与任意风格填充
AI 总结:研究针对TTS模型难以细粒度控制说话风格的问题,提出AutoSIFT框架,将风格分解为可描述和残余两类,通过广义风格解缠器和任意风格填充器,可在保留残余风格时替换指定风格类别,实现高度可定制的语音生成。
链接:https://arxiv.org/abs/2607.12706
机构:UNSW Sydney(新南威尔士大学悉尼分校); University of California San Diego(加利福尼亚大学圣地亚哥分校); Macquarie University(麦考瑞大学); Adobe Research(Adobe 研究院)
作者:Haowei Lou, Junda Wu, Chengkai Huang, Tong Yu, Hye-young Paik, Wen Hu, Lina Yao
英文摘要:State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.
3. Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
神经变形:神经音频编解码器中序列优化的令牌级变形
AI 总结:研究神经音频编解码器中令牌级变形,提出神经变形方法,结合RVQ组转移策略与连续性约束序列匹配器,实现可控音频转换,专注于可部署系统的实现及实时行为。
链接:https://arxiv.org/abs/2607.12725
作者:Emmanouil Karystinaios
英文摘要:Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and realtime behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.
3. 语音增强、降噪与音频修复 | 1 篇
4. Low-Latency Neural Models for Real-Time Music Enhancement
用于实时音乐增强的低延迟神经模型
AI 总结:研究严格因果和低延迟约束下的实时音乐增强,采用紧凑因果网络并与多种模型比较,结果显示改进因多种因素而异,主要贡献是给出基准和分析,强调实时音乐增强虽可行,但稳健改进需多方面考量。
链接:https://arxiv.org/abs/2607.12872
机构:JKU(林茨约翰内斯开普勒大学)
作者:Emmanouil Karystinaios, Jonathan Greif, David Nadrchal, Paul Primus, Gerhard Widmer
英文摘要:Music recordings and live streams are often affected by noise, reverberation, spectral imbalances, or artifacts that degrade listening quality. While speech enhancement has matured into a well-defined research area, music enhancement is less established because musical signals combine overlapping sources, wide bandwidths, strong dynamics, and intentional production effects. We study real-time music enhancement under strict causal and low-latency constraints. We formulate the task around recovery of the intended produced mix from acoustic and production-oriented degradations, adapt compact causal networks to music, and compare speech-derived real-time baselines, an external music-denoising model, an offline restoration reference, and a music-specific MusicFilterNet-MS variant. On the tested hardware, all causal models run faster than real time, but improvements depend strongly on the dataset, degradation type, and metric family; under several objective criteria, indiscriminate enhancement can worsen the degraded input. The main contribution is therefore a benchmark and an analysis rather than a universal best model: real-time music enhancement is feasible, but robust improvement requires degradation-aware modeling, stereo-aware processing, identity-preserving correction, and evaluation beyond a single objective score.
4. 安全、隐私与深度伪造音频 | 1 篇
5. Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
通过维纳-霍普夫线性预测实现可设计解释的音频深度伪造检测
AI 总结:针对音频深度伪造检测难题,提出基于维纳-霍普夫线性预测和轻量级二维卷积神经网络的可设计解释框架,实验显示其检测性能优、计算复杂度低,解释性分析揭示其关注要点,鲁棒性实验表明微调可应对后处理退化。
链接:https://arxiv.org/abs/2607.12584
机构:Amped Software(Amped软件公司)
作者:Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani
英文摘要:The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transitional regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.
5. 其他/综合语音音频 | 3 篇
6. UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection
UD-ASD:一种用于异常声音检测的统一扩散模型
AI 总结:研究针对异常声音检测,提出含轻量级模块的统一扩散模型,先将音频转对数梅尔频谱图,通过嵌入机器ID引导模型为特定机器重建数据,经高斯混合模型拟合误差分布,实验验证该模型在DCASE2022任务2中相比基线有显著提升。
链接:https://arxiv.org/abs/2607.12576
机构:University of Science and Technology of China(中国科学技术大学)
作者:Pengxiang Gao, Yu Qiu, Yanzhi Song
英文摘要:Anomalous Sound Detection (ASD) aims to determine whether faults have occurred by monitoring sounds. Existing methods detect a limited range of anomalies, exhibit poor generalization, or train a separate model for each machine. Diffusion models possess strong generalization and can generate specific data with condition guidance. We propose a unified diffusion model only with a small module. The audio is first transformed into log-Mel spectrograms. The lightweight module embeds machine IDs into condition embeddings, guiding the model to reconstruct data for specific machines. Then diffusion model reconstructs data with condition, using Gaussian Mixture Models to fit the distributions of reconstruction errors. Our unified model could monitor multiple machine types and learn more fundamental feature spaces with cross-domain learning. Experiments on DCASE2022 Challenge Task 2 show that our model achieves 3.44% AUC and 2.52% pAUC improvements over baseline, validating its effectiveness.
7. What is a Musical Scale? Regularity and Convention in the Organization of Pitch
什么是音阶?音高组织中的规律性与惯例
AI 总结:研究探讨音阶定义缺乏共识的问题,采用相对于主音的音高组织统计规律这一经验性定义,结合原型理论将音阶分组到命名类别,为跨文化重新审视音阶提供了新视角和基础。
链接:https://arxiv.org/abs/2607.12596
机构:University of Vienna(维也纳大学); Austrian Academy of Science(奥地利科学院)
作者:John M McBride
英文摘要:Musical scales are near-universal in human music, and most readers will feel they already know what a scale is. On closer inspection, however, the literature lacks a consensus definition: which conditions are necessary and sufficient shifts across disciplines and traditions, and the term turns out to cover several distinct objects. I argue this is less a failure of rigour than a sign that ``scale'' names several related objects: prescriptive abstractions, instrument tunings, statistical regularities in performed pitch, perceptual categories, social conventions. I adopt an empirical definition -- a scale as a statistical regularity in pitch organisation relative to a tonic -- that is portable across traditions and computable from recordings, and situate it alongside the other senses of the term. Even this empirical core is not purely observational, as convention enters in deciding which pitches belong to a scale. And a further step of grouping scales into named categories is a separate convention, which I approach through prototype theory and illustrate with examples from Irish music. Separating these layers provides a basis from which scales can be re-examined empirically and cross-culturally.
8. ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
ChartGenEval:用于节奏游戏图表生成的经过损坏测试的多维反馈
AI 总结:研究针对节奏游戏图表生成,引入ChartGenEval评估框架,通过自动且经损坏测试的核心,在开放音符选择时锚定歌曲节奏,经多组测试及压力测试,能提供多维反馈,为生成器比较与迭代提供依据。
链接:https://arxiv.org/abs/2607.12857
机构:National Yang Ming Chiao Tung University(国立阳明交通大学)
作者:Jhen-Ke Lin
英文摘要:A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a six-question evaluation framework with an automatic, corruption-tested core. It leaves note choice open while anchoring timing to the song: the matched official chart supplies only its authored timing map, never target notes. We test each core output with dose-controlled failures rather than assume that a familiar statistic measures chart quality. Across 80 held-out song groups, seven output axes satisfy prespecified sensitivity and invariance criteria in nine nonredundant tests. Complementary stress tests on the 40-song development panel expose two broader lessons. A chart-wide phase estimate recovers injected shifts of 15, 30, and 60 ms while chart-only outputs remain essentially unchanged. Common-pattern rewriting lowers mean language-model perplexity by 37%, and loop collapse raises mean self-similarity by 62%. ChartGenEval therefore reports separate, role-specific signals instead of one proxy or total score. This profile provides automatic feedback for comparing and iterating generators; selected outputs are candidate optimization targets or constraints after task-specific stress testing.
