今日论文合集:CS.SD语音与音频 | 共 10 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


快速导航

1. 语音识别与关键词检测 2 篇

2. 语音合成与声音生成 1 篇

3. 语音翻译与语音语言模型 2 篇

4. 安全、隐私与深度伪造音频 1 篇

5. 其他/综合语音音频 4 篇


1. 语音识别与关键词检测 | 2 篇

1. Orukeet: Multilingual ASR with Frozen Gabor Kernels

Orukeet:具有冻结Gabor核的多语言自动语音识别

AI 总结:Orukeet用冻结的Gabor核替换Parakeet一半时间滤波器,在多语言数据上训练,使25种语言平均WER从11.01%降至9.85%,相对降低10.6%,并在多数子集上表现更优。

链接:https://arxiv.org/abs/2609.10054


机构:Oruk AI; Stanford University(斯坦福大学); University of Cambridge(剑桥大学); OpenWhispr; Hoid

作者:Nathan Roll (1 and 2), Irene Yi (1 and 2), Büşra Marşan (1 and 2), Vianney Grenez (1), Gabriel Stein (4), Momcilo Mrkaic (5), Pavle Padjin (5), Vladimir Zeljkovic (5), Calbert Graham (1 and 3) ((1) Oruk AI, (2) Stanford University, (3) University of Cambridge, (4) OpenWhispr, (5) Hoid)

英文摘要:Orukeet replaces half of an adapted Parakeet encoder's temporal filters with 12,288 fitted Gabor kernels, freezes these replacements, and trains the remaining parameters on multilingual and multi-accent data. Final adaptation and checkpoint selection use LibriSpeech test-other. Across 20,146 FLEURS recordings in 25 languages, pooled word error rate (WER) falls from Parakeet's 11.01% to Orukeet's 9.85%, a 10.6% relative reduction. Orukeet has lower WER on 23 of the 25 languages. Orukeet outperforms Parakeet on 61 out of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). All comparisons decode the same audio with matched NeMo settings. The fitted kernels are stored as ordinary convolution weights, retaining Parakeet's architecture and inference operators.


2. NOPE-HYPE: A Structured Simulation Workflow for Robust Speech-to-Text Across Diverse Acoustic Environments

NOPE-HYPE:一种面向多样声学环境的鲁棒语音转文本的结构化仿真工作流

AI 总结:NOPE-HYPE提出结合可控环境模拟器、PSD覆盖缩减和超参数搜索的结构化训练工作流,提升语音转文本在多样声学环境中的鲁棒性。

链接:https://arxiv.org/abs/2609.10058

机构:IISER Bhopal(印度科学教育与研究院博帕尔分院); IIT Bombay(印度理工学院孟买分校); IIT Roorkee(印度理工学院鲁尔基分校)

作者:Niramay M. Patel, Bibek Behera, Raksha Sharma

英文摘要:Robust speech-to-text translation systems should perform reliably across diverse acoustic conditions, yet practical pipelines lack controllable tools for systematic environment exploration. Large speech models remain sensitive to unseen acoustic conditions, as training data rarely cover the full range of real this http URL present NOPEHYPE, a structured training workflow that combines a controllable environment simulator, coverage-optimal environment reduction on Power Spectral Density (PSD) templates, and a small, interpretable hyperparameter search over simulator knobs. We show that simulator-generated noise achieves performance comparable to balanced realnoise training across Whisper and SeamlessM4T models, provide principled environment prototype sets, and identify practical default simulator configurations from a structured 27-run hyperparameter sweep.


2. 语音合成与声音生成 | 1 篇

3. Deterministic Prompting for Speaker-Stable Low-Resource Greek TTS

确定性提示实现说话人稳定的低资源希腊语文本转语音

AI 总结:提出数据整理与确定性提示及LoRA微调方案,在3.5小时数据上实现低资源希腊语TTS,达到WER 10.7%和接近人类的说话人一致性。

链接:https://arxiv.org/abs/2609.10022

机构:Institute for Language and Speech Processing, Athena R.C.(雅典娜研究与创新中心语言与语音处理研究所); Department of Digital Medicine, University of Bern(伯尔尼大学数字医学系); School of Electrical and Computer Engineering, National Technical University of Athens(雅典国家技术大学电气与计算机工程学院)

作者:Georgios Syllas, Efthymios Georgiou, Kosmas Kritsis, Alexandros Potamianos

英文摘要:Modern TTS systems approach human quality for high-resource languages but degrade when clean speech data is scarce. Modern Greek exemplifies this, lacking the curated corpora behind state-of-the-art synthesis. We propose a data curation recipe that transforms audiobook recordings into TTS-ready data via WhisperX alignment and filtering. Then we fine-tune Parler-TTS (880M), a prompt-based multilingual model whose pre-training encodes phonetic priors transferable to Greek. During development, we find that LLM-generated style prompts introduce speaker drift at inference. Replacing them with deterministic prompts resolves this, and a speaker-specific LoRA stage trained on 3.5 h of single-speaker data anchors identity while updating ~5% of parameters. Our system achieves WER 10.7% (2.9 above the ASR floor), MOS-I 4.00 (vs. 4.36 human speech), and near-human speaker consistency (MOS-C 4.24 vs. 4.30), showing that robust single-speaker Greek TTS is achievable with limited curated data.


3. 语音翻译与语音语言模型 | 2 篇

4. Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models

超越准确率:用于评估大型音频语言模型音频推理能力的ARIA-Rubrics

AI 总结:针对大型音频语言模型推理评估难题,提出ARIA-Rubrics框架,利用六个指标和思维链提示,实现无需标注的自动透明评估,并识别出三种推理模式。

链接:https://arxiv.org/abs/2609.09681

机构:Imperial College London(帝国理工学院); Technische Universität München(慕尼黑工业大学); Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学); Shanghai Jiao Tong University(上海交通大学); Johns Hopkins University(约翰斯·霍普金斯大学)

作者:Yupei Li, Qiyang Sun, Mohamed Mady, Chenxi Wang, Zhengwei Gong, Berrak Sisman, Björn Schller

英文摘要:Large Audio Language Models (LALMs) have shown strong performance on audio reasoning benchmarks, but accuracy alone cannot distinguish true reasoning from superficial pattern matching, often overestimating reasoning ability since high scores may result from guessing rather than genuine audio understanding. Evaluating the reasoning process itself is essential for improving LALMs' reasoning ability, yet remains challenging. Existing methods either rely on costly human annotation or opaque LLM-as-judge approaches, making them impractical, biased, and lacking transparency. Moreover, audio reasoning introduces unique challenges absent in text-based settings, perceptual hallucination and cross-modal alignment between audio understanding and textual inference, hence text-based evaluation frameworks cannot be directly applied. Therefore, we propose ARIA-Rubrics (Audio Reasoning Integrity Assessment), a lightweight, annotation-free gold reasoning chains, automatic and transparent framework comprising six complementary metrics that evaluate audio reasoning quality across perceptual grounding, reasoning coherence, and answer consistency. We use Chain-of-Thought prompting as an externalization mechanism to make the reasoning process observable. Experiments on 9 models across 2 benchmarks identify three reasoning modes of current LALMs with actionable directions for future development, with ARIA-Rubrics achieving high correlation with human judgments. The code is available at the Github Repository.


5. Unifying Score and Performance for Fine-Grained Music Understanding in Audio-Language Models

统一乐谱与演奏:面向音频语言模型的细粒度音乐理解

AI 总结:针对音频语言模型在细粒度音乐理解上的不足,提出统一乐谱与演奏的MuNo-SP表示及自动数据生成流程,构建MAESTROCaps数据集,经人工评估和基准测试验证其优于MIDI基线。

链接:https://arxiv.org/abs/2609.10351

机构:Bryel Labs(Bryel 实验室); UC Berkeley(加州大学伯克利分校)

作者:Milan Liessens Dujardin, Song-Ze Yu, Kevin Miao

英文摘要:Large audio language models (LALMs) have shown promising progress in broad music-understanding tasks such as tagging, retrieval, and captioning. Music understanding that requires finer hearing over both the content and how it is realized within a performance through dynamics, phrasing, articulation, time, and other performance techniques, however, remains at an earlier stage. Existing audio-language model (ALM) training pipelines typically rely on coarse, weakly grounded captions and therefore provide little support for learning these subtle nuances in music, limiting their ability to serve real-world applications in education or artistic practice. We therefore introduce MuNo-SP (Music Notation unifying Score and Performance), a text-based representation that jointly encodes score content and performance information. Building on MuNo-SP, we develop an automatic training-data generation pipeline that uses aligned scores and performances to produce long-form auditory analyses and musically informed question-answer pairs. We use this pipeline to construct MAESTROCaps, a classical piano dataset comprising 148 long-form performance analyses and 31,080 question-answer pairs derived from 148 aligned score-performance pairs. In a human evaluation, MuNo-SP analyses were preferred by majority vote over MIDI-only analyses for eight of nine excerpts. MuNo-SP also performed strongly on a benchmark of score-performance understanding, suggesting that integrating score and performance information enables more reliable and musically informative LALM supervision than a MIDI-only baseline.


4. 安全、隐私与深度伪造音频 | 1 篇

6. Zero-Shot Temporal Localisation of Audio Deepfakes in Multi-Speaker Conversations

多说话人对话中音频深度伪造的零样本时间定位

AI 总结:针对多说话人对话中合成语音注入的定位问题,提出无需训练的流水线,利用冻结检测器与迟滞解码器,在ASVspoof 5上实现高时间定位精度,并建立零样本基线。

链接:https://arxiv.org/abs/2609.10051

机构:Institute for Advancing Intelligence (IAI), TCG CREST, Kolkata(加尔各答TCG CREST推进智能研究所(IAI))

作者:Soumyadeep Roy

英文摘要:Voice-cloning fraud increasingly relies on surgical injection: a genuine conversation in which only one or two sentences are replaced by synthetic speech. Utterance-level deepfake detectors emit a single real/fake label per clip and cannot report where the synthetic speech lies. We formalise this as Temporal Deepfake Localisation in Multi-Speaker Conversations (TDLMC), show that equal error rate and min-DCF are ill-posed once a file contains both classes, and propose temporal metrics for this regime. Our contribution is a training-free five-stage pipeline that wraps a frozen binary detector and adds segment-level output with no retraining, using a two-threshold hysteresis finitestate-machine decoder to turn noisy window scores into coherent intervals. On 180 constructed multi-speaker conversations from ASVspoof 5, the system attains temporal intersection-over-union 0.90, temporal detection rate 0.95, and MS-DCF 0.26 with a strong backbone, and its false-alarm rate on genuine speech is below 6%, falling under 2% on genuine real multi-speaker dialogue (AMI). Under an identical pipeline, a trained localiser improves temporal IoU by only about 0.04, bounding the cost of forgoing supervision. Evaluated across three frozen detectors under one decoder whose constants are selected on a held-out calibration split, and with a controlled analysis attributing the residual false-alarm rate to a backbone domain gap rather than to the decoder, this provides the first zero-shot baseline and a reusable benchmark for TDLMC.


5. 其他/综合语音音频 | 4 篇

7. Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

声音还是刻板印象?语音到语音模型中的声学性别与内容性别解耦

AI 总结:通过受控实验解耦声音与内容,发现S2S模型在性别归因上受内容刻板印象主导,而非声音,且固定声音评估无法察觉此偏见。

链接:https://arxiv.org/abs/2609.09263

机构:Centific Research(Centific研究院); University of Washington(华盛顿大学)

作者:Xiaoqun Liu, Tanu Mitra, Harshit Rajgarhia, Abhishek Mukherji

英文摘要:Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the words shift the perceived gender of the output voice (voice rendering)? And when the model states the speaker's gender, does it follow the voice or the content (gender attribution)? We answer both with one controlled experiment crossing male and female voices with masculine-, neutral-, and feminine-stereotyped passages, on five open- and closed-source models in English, Spanish, and Mandarin. The rendered voice shows no stereotype drift. But every model decides the speaker's gender from the content, not the voice. Making the content one step more feminine (masculine -> neutral -> feminine) multiplies the odds of a "female" judgment by 1.7-24. When the content clashes with the voice, the worst model misgenders the speaker in 90% of cases. When they agree, it misgenders in only 2%. The bias thus hides in gender attribution, where fixed-voice evaluation cannot see, and where audits must look as S2S systems increasingly speak for real people.


8. Arti-JEPA: Adapting Video World Model to Real-Time MRI of the Vocal Tract for Speech-Production Analysis

Arti-JEPA:将视频世界模型适配到声道实时MRI以进行语音产生分析

AI 总结:Arti-JEPA通过自监督学习适配视频世界模型至声道实时MRI,在音素预测、口吃分类和术后语音分析中验证了冻结编码器的有效性,揭示了域适配的任务依赖性及跨说话者转移差距。

链接:https://arxiv.org/abs/2609.09757

机构:University of Southern California(南加州大学); College of Staten Island, City University of New York(纽约城市大学史泰登岛学院); University of Potsdam(波茨坦大学)

作者:Hong Nguyen, Sean Foley, Christina Hagedorn, Yijing Lu, Sudarsana Reddy Kadiri, Dani Byrd, Shrikanth Narayanan

英文摘要:Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, grayscale, low-resolution - differs substantially from the natural videos that video foundation models are trained on. We introduce Arti-JEPA, a joint embedding predictive architecture to model vocal tract rtMRI by continuing its self-supervised objective on about 62h of unlabelled vocal-tract videos, and evaluate the frozen representation on three tasks: cross-domain phoneme prediction (on typical speakers), fluent-vs-disfluent classification (a corpus containing stuttered speech), and characterizing pre/post-operative transfer (after partial glossectomy). Three key findings emerge. (1) A temporal video prior decisively outperforms per-frame image encoders, and latent prediction (V-JEPA) is at least as strong as pixel reconstruction (VideoMAE), with the edge on fine-grained phonemes. (2) Domain adaptation is \emph{task-dependent}: it roughly doubles cross-domain phoneme prediction $\kappa$ (to 0.352) but does not help binary stuttering classification. (3) Arti-JEPA was able to recover phoneme signal from pre/post glossectomy speech --- an in-domain probe decodes patients at least as well as a typical speaker, indicating that the residual transfer gap is cross-speaker/domain misalignment, not surgical signal loss, and post-operative decoding does not fall below performance on pre-operative speech. Together, these position a frozen, domain-adapted rtMRI encoder as a reusable measurement tool for articulatory and clinical speech science.


9. Robust Rank Aggregation for Multimodal Speech-Based Alzheimer's Disease Detection

基于多模态语音的阿尔茨海默病检测的鲁棒排名聚合

AI 总结:针对多模态语音AD检测中概率平均的尺度不匹配问题,提出基于排名聚合的鲁棒框架,结合置信门控随机森林,在ADReSS2020和ADReSSo2021上分别取得95.83%和90.14%的准确率。

链接:https://arxiv.org/abs/2609.09948

机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)); Center for Language, Intelligence and Machines (LIMA), Shenzhen Loop Area Institute (SLAI)(深圳环区研究所(SLAI)语言、智能与机器中心(LIMA))

作者:Zemin Jin, Tomoko Matsui

英文摘要:Speech-based Alzheimer's disease (AD) detection has recently benefited from multimodal foundation-model representations that integrate complementary acoustic and linguistic information. However, conventional probability averaging over these complementary classifiers is unreliable, because their posterior probabilities exhibit mismatched scales: identical values may reflect different confidence levels across models. We propose a robust rank aggregation framework that aggregates normalized prediction ranks instead of posterior probabilities. Each subject is scored by its percentile within a fixed training-cohort distribution of out-of-fold predictions; since rank ordering is invariant to monotonic transformations, this avoids probability-scale mismatch while preserving classifier confidence ordering. A confidence-gated Random Forest further corrects residual errors using clinically interpretable linguistic features, overriding the rank prediction only when the two disagree and the RF is highly confident, without additional deep model training or explicit posterior-probability calibration. On ADReSS2020 and ADReSSo2021, the method achieves accuracies of 95.83% and 90.14%, respectively, comparing favorably with previously reported results.


10. TimeCues Studio: A Workspace for Music Annotation and Algorithm Prototyping

TimeCues Studio:一个用于音乐标注与算法原型开发的工作空间

AI 总结:TimeCues Studio是一个开源工作空间,支持团队对音乐语料库进行歧义感知标注、比较检测算法并原型化新算法,通过统一时间线集成标注与开发,采用MIT许可并支持Docker Compose部署。

链接:https://arxiv.org/abs/2609.10338

机构:Bar-Ilan University(巴伊兰大学)

作者:Sapir Caduri, Yoav Goldberg

英文摘要:Multimedia applications require precise music annotation-labeled positions, segments, or loops-placed by hand or algorithmically. Machine-learning algorithms are scalable and effective but need annotated training data, scarce for many tasks. TimeCues Studio is an open-source workspace where algorithm-development teams annotate a music corpus, compare detection algorithms against those annotations, and prototype new ones. Unlike existing tools built for a single track at a time, TimeCues targets teams annotating whole collections, tightly integrated with algorithm development. Annotators place several marker types-each supporting ambiguity-aware labeling-on a grid-locked timeline that visualizes many music features, including separated audio stems. The same timeline drives an algorithm-comparison engine with bundled baselines, a Python sandbox for prototyping new models, and an ambiguity-aware evaluator that honors the structured fields. The same visualization suits solo annotators on music-sync projects. TimeCues is MIT-licensed and deploys via one Docker Compose command.