今日论文合集:CS.SD语音与音频 | 共 12 篇。

本文经arXiv每日学术速递授权转载,微信公众号:arXiv_Daily

[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准


快速导航

1. 语音识别与关键词检测 3 篇

2. 语音合成与声音生成 1 篇

3. 语音翻译与语音语言模型 2 篇

4. 数据集、基准与评测 2 篇

5. 安全、隐私与深度伪造音频 1 篇

6. 其他/综合语音音频 3 篇


1. 语音识别与关键词检测 | 3 篇

1. Look Less, Hear Better: Jointly Rewarded GRPO for Streaming ASR

看得更少,听得更好:联合奖励的 GRPO 用于流式语音识别

AI 总结:针对流式ASR中结构延迟不能反映用户感知延迟的问题,提出AWED指标和基于GRPO的延迟奖励后训练方法,在多个前瞻预算下同时降低词错误率和延迟,推进准确率-延迟帕累托前沿。

链接:https://arxiv.org/abs/2609.18333

机构:University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

作者:Xiuwen Zheng

英文摘要:Streaming automatic speech recognition (ASR) must be judged jointly on what it transcribes and on how quickly it commits each word. Delayed streams modeling (DSM) has become the dominant paradigm for streaming large audio-language models, exposing a structural delay $\tau$ that bounds the decoder's lookahead. We show that $\tau$ is a poor proxy for user-perceived latency, and that the alignment-based supervision of DSM leaves latency on the table: the same forced-aligned transcript is used at every $\tau$, forcing the model to withhold words it could already commit. We introduce AWED, a word-level emission-delay metric defined relative to the acoustic end of each word, and post-train a DSM recognizer with GRPO under a reward that scores transcription accuracy and measured delay jointly. Trained at a single operating point ($\tau=6$ frames), our model dominates both its supervised fine-tuning initialization and the Voxtral Realtime backbone across all evaluated lookahead budgets: it cuts WER by 30.8\% relative at an 80\,ms structural delay, and by 5.7\% relative at 480\,ms while lowering median AWED from 1.17\,s to 1.04\,s. Latency-rewarded post-training thus advances the accuracy--latency Pareto frontier of streaming ASR without architectural change.


2. Beyond EER: Multi-Dimensional Evaluation of Information Leakage in Speaker De-Identification

超越EER:说话人去标识化中信息泄漏的多维评估

AI 总结:本文提出一个包含五个互补指标的多维评估框架,用于全面衡量说话人去标识化系统中的信息泄漏,并证明单一指标(如EER)不足以准确反映隐私保护效果。

链接:https://arxiv.org/abs/2609.18673

机构:National Institute of Standards and Technology(美国国家标准与技术研究院); Chakra Consulting Inc.(Chakra咨询公司)

作者:Seungmin Seo, Oleg Aulov, P. Jonathon Phillips, Kevin Mangold, Jonathan Eskin

英文摘要:Speaker de-identification (SDID) aims to preserve privacy by concealing speaker identity while maintaining speech utility. However, current evaluations often reduce privacy to a single dimension - biometric verification performance - typically measured by Equal Error Rate (EER). This narrow focus ignores critical leakage channels, such as soft biometric inference, embedding-level re-identification, and structural template similarity, which threaten the unlinkability and irreversibility of biometric references. We propose a holistic evaluation framework across five complementary metrics: (i) EER, (ii) soft biometric leakage score, (iii) cumulative match characteristic re-identification analysis, (iv) canonical correlation analysis and Procrustes embedding alignment, and (v) intelligibility via word error rate and semantic similarity. Evaluating five SDID systems from the IARPA ARTS program, we demonstrate that these metrics capture independent dimensions of information leakage. Our results indicate that reliance on a single metric can misrepresent the privacy properties of an SDID system.


3. Multi-Teacher Distillation for Cross-Domain Streaming Electrolaryngeal Speech Encoding

跨域流式电喉语音编码的多教师蒸馏

AI 总结:提出多教师蒸馏框架训练轻量流式编码器,融合SSL离散目标和EL微调连续特征,将EL词错误率从39.3%降至21.2%,Mel-Conformer实现最佳效率与精度平衡。

链接:https://arxiv.org/abs/2609.18686

机构:Graz University of Technology(格拉茨技术大学); Medical University of Vienna(维也纳医科大学)

作者:Benedikt Mayrhofer, Enrique Orozco Olivares, Franz Pernkopf, Philipp Aichinger, Martin Hagmüller

英文摘要:Self-supervised learning (SSL) has improved speech representations, yet performance degrades in pathological domains such as electrolaryngeal (EL) speech, and the computational footprint of SSL models limits their applicability in real-time, on-device deployment. We propose a multi-teacher knowledge distillation framework to train a lightweight, streaming content encoder that generalizes across healthy (HE) and EL speech. Two teachers are distilled progressively: a frozen SSL model providing discrete phonetic cluster targets from HE speech, and an EL-fine-tuned speech recognition model supplying continuous bottleneck feature targets. Evaluated via downstream speech recognition, our approach reduces the EL word error rate to 21.2%, compared to 39.3% for the strongest zero-shot SSL baseline. Among causal convolutional, Transformer, Conformer, and Mamba-based student architectures, a Mel-Conformer achieves the best combination of EL accuracy and computational efficiency. The final encoder contains 21.9,M parameters and runs at a real-time factor of 0.30 under ONNX Runtime on a single CPU core.


2. 语音合成与声音生成 | 1 篇

4. Variable-Rate Harmonic-Percussive Time-Scale Modification with Real-Time Playback in Python

可变速率谐波-打击乐时间尺度修改及其Python实时播放

AI 总结:本文提出一种Python实现的谐波-打击乐时间尺度修改方法,支持可变速率实时播放,通过预计算表优化减少约一半运行时间,且感知质量不变。

链接:https://arxiv.org/abs/2609.18999

机构:Harvey Mudd College(哈维穆德学院); University of California, Los Angeles(加州大学洛杉矶分校)

作者:Sayema Lubis, Clark Peng, Jared Carreño, TJ Tsai

英文摘要:Time-scale modification (TSM) has a number of open-source implementations, but these are designed almost exclusively for offline use, in which a recording is processed at a fixed rate and written out ahead of time. Applications such as automatic musical accompaniment require a different setting, which we call variable-rate playback: the recording to be stretched is known in advance, but the playback rate is not, and must change continuously in response to a live performer. The few implementations that generate output in real time are written in C++ and optimized for speed rather than for ease of modification, experimentation, and integration with the primarily Python-based research ecosystem. This paper describes a Python implementation of the widely used harmonic-percussive TSM method for the variable-rate playback setting. Harmonic-percussive separation is performed offline as a preprocessing step on the known input recording, while synthesis and playback are carried out in real time with a time-scale factor that may change at every frame. We further propose a family of variants that reduce runtime by replacing the phase vocoder's analysis-stage FFT and instantaneous frequency calculations with lookups into precomputed tables. Subjective listening tests with 24 participants (1114 pairwise ratings) show that these approximations become perceptually indistinguishable from the exact implementation once the precomputed hop size is sufficiently small, while reducing total runtime by roughly half. We characterize the resulting tradeoffs among precomputation, runtime, memory, and perceptual quality to guide algorithm selection, and we release our implementation as an open-source package.


3. 语音翻译与语音语言模型 | 2 篇

5. VoiceTrace: A Benchmark and Retrieval Framework for Who-Said-What Speech Retrieval

VoiceTrace:面向“谁说了什么”语音检索的基准与检索框架

AI 总结:提出VoiceTrace基准与两阶段检索框架,联合文本与参考语音实现“谁说了什么”的混合语音检索,在语义和混合检索任务上均达最优性能。

链接:https://arxiv.org/abs/2609.18521

机构:Zhejiang University(浙江大学); Johns Hopkins University(约翰斯·霍普金斯大学)

作者:Aaron Yee, Fengjie Lu, Jiarui Hai, Chenang Jiang, Helin Wang, Siwei Tu, Weitao You, Lingyun Sun

英文摘要:Speech retrieval has become increasingly important as spoken content continues to grow across meetings, lectures, podcasts, and videos. Existing benchmarks and models have advanced semantic search over spoken content, but largely focus on \emph{what} is said while overlooking \emph{who} says it. In many real-world scenarios, however, users need to retrieve speech based jointly on semantic content and a target speaker, where the speaker may be specified naturally through a reference speech utterance rather than a predefined identity. To address this gap, we introduce \textbf{VoiceTrace-Bench}, a benchmark for hybrid speech retrieval in which each query combines text specifying \emph{what} to retrieve with reference speech specifying \emph{who} to retrieve. This setting requires models to integrate complementary semantic and speaker information directly from heterogeneous query inputs. Motivated by the joint audio-text modeling capabilities of audio-language models (ALMs), we develop \textbf{VoiceTrace}, a two-stage retrieval framework consisting of \textbf{VoiceTrace-Emb}, an embedding model that learns unified representations for efficient large-scale retrieval, and \textbf{VoiceTrace-Reranker}, a reranking model that jointly examines each query--candidate pair for fine-grained relevance estimation. Experiments show that VoiceTrace achieves state-of-the-art performance on established semantic speech retrieval benchmarks, while substantially outperforming cascade-based approaches on VoiceTrace-Bench, demonstrating its effectiveness for both conventional semantic retrieval and the new hybrid retrieval setting.


6. FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection

FRAUDSkill:面向音频反欺诈检测的结构化冻结权重技能优化

AI 总结:针对音频反欺诈检测中模型难以适应演变的问题,提出FRAUDSkill框架,通过冻结权重并优化外部技能层,在TeleAntiFraud基准上实现73.50% Macro-F1,较基线提升31.96%。

链接:https://arxiv.org/abs/2609.18766

机构:JD Technology(京东科技); Meta; University of Science and Technology of China(中国科学技术大学); Northeastern University(东北大学); Peking University, Shenzhen(北京大学深圳研究生院)

作者:Chengxian Hu, Zhiming Ma, Mingjun Pan, Yifan Wang, Shun Zhang, Qifan Wang, Zhilei Zhao, Yijin Zhou, Yuxi Zhao, Huiyuan Liu, Peidong Wang, Peng Chen

英文摘要:Large audio-language models have shown promise for anti-fraud detection by directly processing speech and reasoning over fraud-related evidence. Their deployment, however, requires predictions to follow a predefined label space and a structured decision protocol consisting of service-scenario identification, fraud detection, and conditional fraud-type classification. Existing fine-tuning and prompt-based approaches typically encode task knowledge, constraints, and decision rules into model parameters or manually maintained prompts, making them difficult to adapt as fraud patterns and labeling policies evolve. To this end, we propose FRAUDSkill, a structured frozen-weight adaptation framework that leaves the underlying audio-language model unchanged while optimizing an external layer of skill programs, route-specific policies, and decision rules. We further combine structured output control with validation-guided multi-path inference to ensure protocol-compliant predictions. On the TeleAntiFraud benchmark, FRAUDSkill achieves 73.50% Macro-F1, outperforming the shared frozen-model baseline by 31.96% while reducing invalid outputs to 1.94%. Extensive experiments demonstrate that external skill optimization provides an effective and adaptable solution for structured audio anti-fraud detection without modifying the underlying model. The source code is available at this https URL.


4. 数据集、基准与评测 | 2 篇

7. TTM-Bench: A Framework for Text-to-Music System Performance Benchmarking

TTM-Bench:一种文本到音乐系统性能基准测试框架

AI 总结:针对文本到音乐系统性能比较困难的问题,提出TTM-Bench框架,从音乐内容对齐和计算效率两维度评估,并通过案例研究证明多维度评估的必要性。

链接:https://arxiv.org/abs/2609.18585

机构:Institute of Information Systems and Networking (ISIN), SUPSI(信息系统与网络研究所(ISIN),瑞士南部应用科学与艺术大学); t2b AG(t2b股份公司)

作者:Giorgia Adorni, Michela Papandrea, Battista Rimoldi, Tiziano Leidi

英文摘要:Text-to-music (TTM) systems are increasingly used to generate musical audio from natural-language descriptions. Robust evaluation is therefore essential, yet reliable performance comparison remains challenging. This difficulty stems from differences in system architecture, supported conditioning information, and access mode, as well as heterogeneous and fragmented metrics that cannot be applied uniformly across systems. To address these challenges, we introduce TTM-Bench, a framework that defines a common protocol for systematic, reproducible performance benchmarking of contemporary TTM systems. It evaluates performance along two dimensions: musical-content alignment, quantified by interpretable semantic, genre, and musical-descriptor agreement scores against a common musical specification and summarized by an aggregate score; and computational efficiency, characterized by generation latency and real-time factor, alongside resource use for local models and cost for hosted services. We demonstrate the framework through a preliminary comparative case study, illustrating the complementary evidence captured by these dimensions. The results show that higher musical-content alignment does not systematically coincide with lower computational demands, highlighting the importance of assessing TTM performance through distinct, interpretable measures rather than a reductive overall indicator.


8. TeleAntiFraud 2.0: A Refreshable, Profile-Grounded, and Audio-Based Benchmark for Telecom Fraud Detection

TeleAntiFraud 2.0:一个可刷新、基于画像和音频的电信诈骗检测基准

AI 总结:TeleAntiFraud 2.0提出可刷新、基于画像的音频基准,通过混合树流水线生成近域负例,揭示分类器在近域场景下性能显著下降,确立近域构建和崩溃感知报告为评估核心要求。

链接:https://arxiv.org/abs/2609.18748

机构:JD Technology(京东科技); Chongqing Ant Consumer Finance Co., Ltd., Ant Group(重庆蚂蚁消费金融有限公司,蚂蚁集团); Meta; University of Science and Technology of China(中国科学技术大学); Northeastern University(东北大学)

作者:Huiyuan Liu, Zhiming Ma, Yanxing Liu, Shun Zhang, Qifan Wang, Di Liu, Yifan Wang, Yuyang Deng, Haoyang Meng, Yijin Zhou, Yuxi Zhao, Chengxian Hu, Peidong Wang, Peng Chen

英文摘要:Telecom fraud scripts evolve rapidly and are often designed to resemble routine service conversations, creating two key requirements for audio-based telecom-fraud evaluation. First, benchmarks must incorporate newly observed scam patterns without overwriting previously established test sets. Second, they must distinguish fraud from lawful, near-domain calls rather than relying on topic-separated negative examples. We present TeleAntiFraud 2.0, constructed with our Mixed-Tree Anti-Fraud Generation Pipeline and evaluated under a monthly frozen evaluation protocol. The pipeline transforms online fraud-case abstracts into profile-grounded scenarios, expands them through mixed-tree generation, realizes fraud and non-fraud dialogue paths under shared contexts, renders validated dialogues as role-matched speech, and freezes the resulting audio, labels, prompts, manifests, and provenance records for each monthly evaluation set. Each frozen set contains 900 Chinese calls, comprising 600 fraud and 300 near-domain non-fraud cases. Controlled text experiments show that three classifiers achieve perfect macro-averaged F1 (Macro-F1) when evaluated against unrelated or ordinary negatives, but drop to 0.65-0.68 with near-domain sibling negatives. Full-set audio and automatic-speech-recognition plus large-language-model (ASR+LLM) evaluations further reveal class-prior shortcuts, prediction collapse, and snapshot sensitivity. Together, these findings establish near-domain construction and collapse-aware reporting as core requirements for evaluating audio-based telecom-fraud models under realistic confusable conditions. The accompanying research artifact includes the construction code, evaluation scripts, manifests, and documentation. Our dataset and code are available at this https URL.


5. 安全、隐私与深度伪造音频 | 1 篇


9. What Affects the Performance of Fake Audio Detection? Analyzing Factors in a Continual Learning Setting

什么影响伪造音频检测的性能?在持续学习设置中分析因素

AI 总结:本研究在持续学习环境下分析攻击者架构、训练数据集、说话人多样性和任务顺序对伪造音频检测性能的影响,发现数据集伪影和任务顺序等因素显著影响系统鲁棒性。

链接:https://arxiv.org/abs/2609.19067

作者:Yixuan Xiao, Ngoc Thang Vu

英文摘要:The increasing sophistication of deepfake audio generation technologies makes it important to develop robust fake audio detection systems that can adapt over time. This study examines how various factors impact the performance of detection systems in a continual learning setting. We focus on factors such as attacker architectures, attackers' training datasets, speaker diversity, and task order. We evaluate the performance of three detection models trained with four different strategies, including direct fine-tuning, one-class classification, random replay, and Learning without Forgetting. Results show that artifacts from the fake audios might arise from the attackers' training datasets, and simply changing attacker architectures does not sufficiently challenge detection systems. Moreover, task order and speaker diversity can significantly influence performance, with varying degrees of sensitivity across different detection models and training strategies. These insights underline the need for careful consideration of these factors when developing robust detection systems.


6. 其他/综合语音音频 | 3 篇

10. Timbre Analysis of the Hulusi, a Southwestern Chinese Free-Reed Instrument, using Machine Learning

使用机器学习对云南葫芦丝(一种中国西南部自由簧管乐器)的音色分析

AI 总结:本研究使用机器学习(COMSAR框架中的SOM)分析云南葫芦丝的音色,发现频谱质心、锐度和分形关联维数可形成音高聚类,其中分形维数聚类效果最佳,且最高音高混沌性最低。

链接:https://arxiv.org/abs/2609.17612

机构:College of Art, Zhejiang Normal University(浙江师范大学艺术学院); Institute of Systematic Musicology, University of Hamburg(汉堡大学系统音乐学研究所)

作者:Yang Xia, Rolf Bader

英文摘要:The hulusi is a wind instrument that was invented in Yunnan Province, China, and has become tremendously popular in recent years. It consists of a mouthpiece, a gourd, and three bamboo tubes, all with free reeds made of copper. The main bamboo tube in the middle has seven finger holes. In this instrument, the pipe length, not the free reed's eigenfrequency, determines the instrument's pitch, unlike, for example, with the Western accordion or the blues harp. In this study, a machine learning model implemented in the COMSAR framework ( this https URL ) was used to investigate the timbre characteristics of the \emph{hulusi} to cluster different instruments and pitches. The measured \emph{hulusi} pitches C, B, A, G, and F were analyzed according to seven psychoacoustic features, among which only the spectral centroid, sharpness, and fractal correlation dimension are shown to form pitch clusters. These timbre features were used to train Kohonen self-organizing maps (SOMs) for clustering. Brightness and sharpness analysis revealed that the highest pitches were less bright and less sharp than mid- and low-range pitches were. Furthermore, the fractal correlation dimension, which mainly determines the chaoticity of the initial transients, was the best-clustering timbre feature for the hulusi, with the highest pitches showing the least chaoticity. This result is supported by defining a cluster quality index for the SOMs.


11. The Unbearable Weight: Scaling Models and Methods for UAV Audio Classification

难以承受之重:无人机音频分类的模型与方法缩放

AI 总结:针对无人机音频分类数据稀缺问题,系统比较多种模型与微调方法,发现轻量CNN结合选择性批归一化微调(更新<0.5%参数)达97.65%准确率,缩放方法优于缩放模型。

链接:https://arxiv.org/abs/2609.17884

机构:College of Charleston(查尔斯顿学院)

作者:Andrew P. Berg, Qian Zhang, Mia Y. Wang

英文摘要:As unmanned aerial vehicles (UAVs) become increasingly prevalent in consumer and defense settings, classifying them reliably from limited, modality-specific data is an urgent challenge. The dominant approach, large pretrained networks fully fine-tuned on task data, carries a substantial computational and memory weight that is hard to bear in resource-constrained UAV deployments, where edge inference and rapid retraining for emerging platforms are both required. This paper systematically scales across both model architectures and fine-tuning methods for UAV audio classification, asking when that weight is justified and when lighter alternatives prevail. Using a custom dataset of 3,100 audio clips spanning 31 drone classes, we evaluate transformer (ViT, AST) and convolutional (custom CNN, ResNet-18/152, MobileNet-V3-S/L, EfficientNet-B0/B7) backbones under full fine-tuning, classifier-only fine-tuning, and four parameter-efficient fine-tuning (PEFT) methods: SSF, IA3, OFT, and selective batch-norm tuning. All configurations are evaluated with 5-fold cross-validation across accuracy, training time, trainable-parameter share, and inference-time memory footprint. Selective batch-norm fine-tuning of EfficientNet-B7 with three-fold augmentations achieves the highest validation accuracy (97.65% +- 0.30) while updating under 0.5% of model parameters. Across the sweep, lightweight CNNs consistently outperform transformers on both accuracy and efficiency. For UAV audio classification under data scarcity, scaling the method outperforms scaling the model.


12. CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

CPR:在连续自回归建模中结合全局作曲、局部演奏与全序列精修的钢琴渲染

AI 总结:CPR框架结合连续自回归建模、局部流匹配和上采样精修,实现高质量钢琴MIDI渲染,并引入BREPA和MT-RoPE增强语义与对齐。

链接:https://arxiv.org/abs/2609.18216

机构:The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))

作者:Chong Jing, Junan Zhang, Zhizheng Wu

英文摘要:Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording. Existing approaches primarily follow two paradigms: autoregressive (AR) modeling and flow matching (or diffusion). Discrete-codec AR models provide causal temporal modeling, but quantization can discard acoustic detail. Flow matching better preserves acoustic structure in the cost of full-sequence attention costs and worse semantic structure. Continuous autoregressive models operate directly on continuous representations. It not only combines the condition-following ability of AR models and distribution-modeling capacity of flow matching but also bypasses the quantization bottleneck with lower computational costs. Building on this principle, we present Composer--Performer--Refiner (CPR) framework. Composer autoregressively predicts continuous hidden states, Performer generates 24kHz acoustic latents through local flow matching and Refiner then upsamples the waveform to 48 kHz. We further introduce Bottlenecked Representation Alignment (BREPA) and Modality--Time RoPE (MT-RoPE) to strengthen musical semantic structure in Composer hidden states and temporal alignments across modalities. Codes are available at this https URL