今日论文合集:CS.SD语音与音频 | 共 6 篇。
本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
1. A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study
用于信号分类的无乘法特征提取器:关键词检索案例研究
AI 总结:该研究针对关键词检索问题提出无乘法的next iRDT特征提取器,在Google KWS 12类数据集上获94.7%验证准确率,处理时间远短于MFCC,适合超低功耗边缘设备。
链接:https://arxiv.org/abs/2608.17108
作者:Radu Dogaru, Ioana Dogaru
英文摘要:A very low complexity feature extractor called next iRDT is proposed and evaluated for the problem of keyword spotting (KWS). Unlike any other types of feature extractors including the widely used MFCC, or adaptive, CNN-based ones, our algorithm is multiplier-free and it employs only simple, energy-efficient arithmetic operators. Since keyword-spotting of speech commands (KWS) is a typical application for TinyML platforms requiring low complexity for the signal classification chain, we consider it as a case study to evaluate complexity and functional performance. If properly tuned, iRDT demonstrates similar accuracy to solutions based on MFCC or CNN-based extractors using baseline classifiers on Google's KWS 12-classes dataset. With a different classifier the system achieved 94.7% validation accuracy. Processing times on CPU for the proposed feature extractor, are at least one order of magnitude smaller than for the MFCC. The proposed algorithm has a very low hardware footprint, making it ideal for ultra-low power edge devices. Code and demo are available [18].
2. FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
FireRedTTS3:基于语义丰富的语音表征的统一语音生成与编辑
AI 总结:本研究提出FireRedTTS3框架,利用冻结音频编码器正则化特征空间缓解自回归误差,其两个变体在对应任务数据集上均优于竞争系统,实现稳定可控的高保真语音生成与编辑。
链接:https://arxiv.org/abs/2608.17492
机构:Xiaohongshu(小红书)
作者:Feiyu Shen, Kun Xie, Yichen Wu, Ziqi Dai, Yichen Han, Junjie Li, Xuelong Geng, Fenglong Xie, Lei Xie, Xu Tang, Yao Hu
英文摘要:Recent continuous autoregressive TTS models operate directly on continuous speech representations, preserving rich acoustic details while leveraging the instruction-following capabilities of text LLMs. This paradigm opens new possibilities for voice cloning, instruction-controlled voice design, and speech editing, but remains susceptible to error accumulation during autoregressive generation. Existing solutions often require additional semantic modules, multi-stage tokenizer training pipelines, or complex autoregressive architectures. In this work, we propose FireRedTTS3, a simple yet effective speech generation and editing framework that mitigates error accumulation at the representation level. Specifically, we leverage a frozen Audio Encoder trained on diverse speech understanding tasks as a semantic teacher to regularize the audio feature space. This improves text-speech alignment and stabilizes autoregressive generation while keeping the overall system simple. FireRedTTS3 provides two variants: FireRedTTS3-Base for multilingual and multi-dialect zero-shot voice cloning, and FireRedTTS3-Instruct for unified voice cloning, instruction-controlled voice design, and speech editing. Experiments show that FireRedTTS3-Base achieves the best average speech intelligibility and speaker similarity among compared systems on Seed-TTS-Eval and MiniMax-MLS-Test, while FireRedTTS3-Instruct outperforms competing systems on InstructTTSEval and Ming-Freeform-Audio-Edit. These results demonstrate that semantically enriched continuous speech representations, combined with a simple architecture, enable stable, controllable, and high-fidelity speech generation and editing. Code and models are available at this https URL.
3. Target Speaker Identification: A Low-Latency Streaming Pipeline
目标说话人识别:一种低延迟流式处理管道
AI 总结:针对助听器低延迟需求,提出含流式说话人分割与验证的两步法管道,经多工具基准测试与参数优化,17个评估片段中中位数准确率超0.90,为低延迟选择性放大提供概念验证。
链接:https://arxiv.org/abs/2608.17972
机构:Arizona State University(亚利桑那州立大学)
作者:Patrick S. Burke (Children's National Hospital), Satyam Raj (Arizona State University), Sean Kinahan (Arizona State University)
英文摘要:We present a real-time pipeline of open source, pretrained models for streaming identification of a target speaker, motivated by hearing-aid applications where latency as low as 10 ms can be perceptible. We formulate a two-step approach in which incoming audio is first segmented by speaker using low-latency streaming diarization, followed by speaker verification against a registered target speaker. To emulate conversational speech while minimizing overlap, we use the This American Life Podcast Transcripts dataset and select the host as a consistent target speaker. We benchmark offline diarization with Pyannote and LIUM using diarization error rate (DER) and select Pyannote based on baseline performance and compatibility with streaming. We then evaluate speaker verification using Pyannote and TitaNet-Large and generate ROC curves to select an operating region. We integrate Diart and tune clustering parameters to reduce DER while maintaining real-time operation. We pair Diart with Pyannote verification and evaluate system-level performance by converting predicted and ground-truth speech regions into 100 ms binary masks. Across 17 evaluation episodes, the system achieves greater than 0.90 median accuracy with high specificity (0.95-0.98) at cosine distance thresholds of 0.7-0.75, demonstrating a practical proof of concept for downstream low-latency selective amplification.
4. Automatic Transcription of Microtonal Free-Rhythm Vocal Music: A Case Study in Iranian Classical Music
微分音自由节奏声乐的自动转录:以伊朗古典音乐为例
AI 总结:本文以伊朗古典音乐为案例,提出结合音高直方图与DTW的自动转录工作流程,采用music21库实现,可转录含tahrir等装饰音的微分音声乐,配套可视化与编辑工具,推进计算民族音乐学发展。
链接:https://arxiv.org/abs/2608.17114
机构:Cu Test Inc.(Cu Test公司)
作者:Sepideh Shafiei, Shapour Hakam, Harsh Dange, Joel Rodriguez Caraballo
英文摘要:This paper introduces a computational workflow for automatically transcribing microtonal, free-rhythm vocal music, with Iranian classical music as a case study. Our approach is based on performances by the renowned vocalist Karimi and ground truth transcriptions by the prominent ethnomusicologist Masoudieh [14], which were subsequently incorporated into the IRMA Audio-MIDI dataset [20]. To accurately extract melodies, we employ pitch histograms in conjunction with Dynamic Time Warping (DTW). Additionally, we introduce specialized musical notations to capture the intricate ornamentations characteristic of the genre, with particular emphasis on the vocal technique tahrir. The transcription process is implemented in Python using the music21 library for symbolic music representation [5]. This study not only advances the field of computational ethnomusicology but also highlights the potential of computational methods in preserving and analyzing complex musical traditions. The transcription system also generates a combined visualization of the audio pitch contour and the DTW-aligned MIDI representation, enabling users to inspect the correspondence between the performance and the generated transcription. A companion visual editor supports expert-in-the-loop correction of the resulting notation.
5. UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding
UniVerse:针对文化包容性低资源音乐理解的基准测试与增强
AI 总结:该研究针对低资源音乐理解问题,推出UniVerse基准与数据集,训练LALMs并研究多模态不平衡学习策略,实现性能提升但仍存在深层音乐理解的差距。
链接:https://arxiv.org/abs/2608.17852
机构:Sogang University(西江大学); KAIST(韩国科学技术院); NYU Shanghai(上海纽约大学); JIUTIAN Research, China Mobile(中国移动九天研究院); The State Key Laboratory of Multimedia Information Processing, Peking University(北京大学多媒体信息处理国家重点实验室); China Mobile (Hong Kong) Innovation Research Institute(中国移动(香港)创新研究院); Central Conservatory of Music(中央音乐学院)
作者:Ziya Zhou, Shangda Wu, Shenyang Xu, Yutong Zheng, Dafang Liang, Suin Chung, Danbinaerin Han, Junyan Jiang, Yongyi Zang, Ruibin Yuan, Rongxiu Zhong, Shilei Zhang, Junlan Feng, Jinglei Liu, Haotian Zhou, Zijin Li, Dasaem Jeong, Wei Xue, Yike Guo
英文摘要:Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. However, limited attention has been paid to improving their adaptability across diverse musical traditions, particularly folk music rooted in distinct cultural contexts. Folk-music traditions are typically resource-scarce, unevenly represented across regions, and poorly documented. Even when such samples appear in large-scale pre-training, LALMs often fail to capture their structural and stylistic characteristics, partly due to the absence of dedicated evaluation protocols and training solutions. To address these limitations, we introduce UniVerse, a reproducible solution for low-resource music understanding. Specifically, we propose UniVerseBench, a benchmark of 5,042 Q&A pairs across more than 38 cultural and linguistic entities, constructed via an expert-guided yet highly automated pipeline. In parallel, we construct a fully automated, model-generated multi-turn dialogue training dataset UniVerseSet. By training LALMs on UniVerseSet, we systematically adapt and investigate representative multimodal imbalance learning strategies across both dense and Mixture-of-Experts (MoE) architectures. Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
6. The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report
Deepfake语音检测的最后一公里:产学研经验报告
AI 总结:本文结合与Phonexia的三年合作经验,指出Deepfake语音检测落地的障碍,提出商业可用数据集标准等研究与协调建议,为该领域产学研衔接提供参考。
链接:https://arxiv.org/abs/2608.17585
机构:Brno University of Technology(布尔诺理工大学); Phonexia(丰西亚公司)
作者:Anton Firc, Kamil Malinka, Vojtěch Staněk, Miroslav Hlaváček, Marek Bartoň
英文摘要:Synthetic speech detection benchmarks now report sub-1% error rates on some in-domain evaluations, yet performance degrades under unseen attacks, channel mismatch, and distribution shift. Based on a three-year effort with Phonexia, a commercial speaker-recognition vendor, we report barriers encountered while building and deploying a detector. Many public benchmarks are not licensed for commercial model development. Real inputs are not four-second clean clips but long, codec-degraded, sometimes partially synthetic recordings. And when a calibrated system returns a
语音与音频学术速递[8.19]
评论 0
文明发言,友善讨论
还没有评论,发表你的看法,来抢沙发~
