今日论文合集:CS.SD语音与音频 | 共 7 篇。


本文经arXiv每日学术速递授权转载

微信公众号:arXiv_Daily


[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准

快速导航

1. 语音识别与关键词检测 1 篇

2. 语音合成与声音生成 3 篇

3. 其他/综合语音音频 3 篇

1. 语音识别与关键词检测 | 1 篇

1. What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

波形知道什么:使用字幕工作室实现透明优先的语音和音频智能

AI 总结:研究借助字幕工作室平台实现透明优先的语音和音频智能。核心方法是构建基于FastAPI的三层架构系统。主要贡献为提出透明优先框架,明确指标性质,提升语音分析的可追溯性、可解释性与可靠性,还介绍了相关系统架构等内容。

链接:https://arxiv.org/abs/2607.18704

机构:Faculty of Science, Agriculture, and Engineering, Newcastle University Singapore(新加坡纽卡斯尔大学科学、农业与工程学院); School of Information and Control Engineering, Qingdao University of Technology(青岛理工大学信息与控制工程学院); Department of EEE, Amrita School of Engineering, Amrita Vishwa Vidyapeetham(阿姆瑞塔工程学院电气与电子工程系,阿姆瑞塔大学)

作者:Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar

英文摘要:Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searchable content through automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle generation. The system is built on a FastAPI backend with a real-time dashboard and adopts a three-layer architecture comprising (i) a transcription and diarization core based on Whisper-class automatic speech recognition and pyannote speaker diarization, (ii) an audio intelligence layer that extracts acoustic and linguistic features, including waveforms, spectrograms, pitch, speaking rate, silence, filler-word frequency, and sentiment, directly from the audio signal, and (iii) an integration layer that supports data export and downstream workflow integration. A principal contribution of this work is the transparency-first framework, in which every reported metric is explicitly identified as measured, derived, or unavailable, thereby improving the traceability, interpretability, and reliability of speech analytics. The paper presents the system architecture, benchmarking methodology, explainability and uncertainty framework, and key considerations for enterprise-scale deployment.

2. 语音合成与声音生成 | 3 篇

2. A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

约鲁巴语情境语音合成器:系统设计、音系规则架构及轮廓的正字法扩展

AI 总结:介绍用于约鲁巴语的基于规则的双音素拼接语音合成器TTSYoruba,阐述其音系架构,包括声调处理等,还介绍正字法扩展,通过听众研究评估性能,为约鲁巴语语音合成及相关研究提供支持。

链接:https://arxiv.org/abs/2607.18317

作者:Kola Tubosun, Adedayo Oluokun, Hafiz Adewuyi, Dadepo Aderemi

英文摘要:We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the this http URL open dictionary of Yoruba personal names. The system takes tone-marked Yoruba text as input and produces audio output by applying a hand-crafted phonological rule system to a recorded inventory of 651 diphone units spanning five tonal variants of every consonant-vowel combination in the language. We describe the phonological architecture of the system in detail, including our complete tonal file-selection logic, our treatment of the three-way nasal disambiguation problem (oral /n/, nasalized vowel, and syllabic nasal), and the derivation of contextual rising and falling tones from level-tone input. We also present, as an orthographic contribution, the adoption of the caron and circumflex, which are symbols with prior standing in Yoruba phonological transcription, as standard single-vowel contour tone markers, integrated into the TTS normalization pipeline and the WriteYoruba keyboard input tool. The system's performance was evaluated through a listener study (N=50), with detailed results on Mean Opinion Scores (MOS) presented in Section 6. Keywords: Yoruba, text-to-speech, low-resource languages, diphone synthesis, contour tones, African language NLP, rule-based synthesis

3. CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

CS-ETS:基于混沌启发的基于桑巴的肌电到语音合成与非线性混沌损失

AI 总结:该研究提出CS-ETS架构用于肌电到语音合成,结合基于桑巴的编码器与LER、MSDFA两个混沌启发的损失函数,参数数量降低40.79%,新方法提升了相关指标,还减少计算量,首次实现用非线性混沌物理监督ETS,获更小高性能模型。

链接:https://arxiv.org/abs/2607.18629

作者:Sajid Fardin Dipto, Tarikul Islam Tamiti, David Vergano, Luke Baja-Ricketts, Anomadarshi Barua

英文摘要:We propose a chaos-inspired new architecture for EMG-to-Speech (ETS) synthesis called CS-ETS, which combines a Samba-based encoder with two novel chaos-inspired loss functions -- Lyapunov Exponent Regularization (LER) and Multi-Scale Detrended Fluctuation Analysis (MSDFA). LER is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. MSDFA exploits detrended fluctuation analysis to quantify fractal-like, long-range temporal chaotic correlation. CS-ETS surpasses prior work with a 40.79\% lower parameter count (32M vs 54.1M) and introduces a new Post-Vocoder Alignment approach that improves LSD by 2.1x, STOI by 4.7x, and SI-SDR by 1.25x. CS-ETS reduces computation by 13.33\% while maintaining improved performance. To the best of our knowledge, for the first time, we show how ETS can be supervised by the subtle non-linear chaotic physics with Samba attention to achieve a significantly smaller model with superior performance.

4. Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer

基于流匹配文本到语音教师的分层深度剪枝蒸馏:一个紧凑的印地语语音合成器

AI 总结:研究在数据量严重不足时构建紧凑印地语TTS模型的方法,通过分层深度剪枝蒸馏大型流匹配教师模型热启动学生模型,经多步剪枝与微调,得到不同参数的模型,在特定参数下效果良好,还解决了一些模型问题并进行了基准测试。

链接:https://arxiv.org/abs/2607.18662

作者:Sivateja Trikutam

英文摘要:We present a practical recipe for building a compact Hindi text-to-speech (TTS) model by distilling a large flow-matching teacher (IndicF5, 337M-parameter DiT) under a severe data budget (~17.6 hours). Training a small model from scratch on this much data fails outright. Instead we warm-start the student from the teacher by pruning depth only: keeping the teacher's width, text dimension, attention heads, and mel/text I/O fixed so all non-block tensors copy one-to-one, and retaining an evenly-spaced subset of transformer blocks. We first measure how much depth the teacher tolerates (it remains near-functional at -27% blocks but collapses past -50%), then descend gradually (22 -> 16 -> 12 -> 8 -> 6 blocks), re-fine-tuning after each prune, with each step gated by an objective ASR word-error-rate (WER) check. The resulting students reach WER 0.00 on unseen sentences at 249M and 190M parameters, and remain robust down to 131M; at 102M we observe a clear capacity cliff that we attribute to the data budget rather than the recipe. We also document two train/inference feature- and library-parity failures (mel filterbank and rotary-embedding library versions) that silently degrade audio, and a version-independent fix. The method yields a high-quality Hindi voice that runs in real time on a 6 GB laptop GPU. An independent 50-sentence FLEURS benchmark compares the released 190M student against its teacher and MMS-TTS-hin.

3. 其他/综合语音音频 | 3 篇

5. Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation

Fretiq:通过工程频谱特征和留出的自由演奏评估实现浏览器原生电吉他弦分类

AI 总结:研究针对单音电吉他音频弦分类难题,提出Fretiq系统,基于26维特征表示及比较训练方法,在平衡帧验证中达97.1%准确率,留出自由演奏评估总体准确率87.8%,还描述了特征提取管道及实现失败模式,且系统可在浏览器内运行。

链接:https://arxiv.org/abs/2607.18303

作者:Aadi Garg

英文摘要:Identifying which string produces a given pitch in monophonic electric guitar audio is a fundamental classification challenge: a single pitch can often be produced on multiple strings at different fret positions, with timbral differences that prior listening studies confirm are largely imperceptible to untrained humans. Existing approaches using support vector machines and spectral envelope features have achieved F-measures of 0.90 for six-string electric guitar classification, while String-Inverse Frequency features in earlier work achieved F1 scores up to 0.72. We present Fretiq, a preliminary single-instrument, single-player browser-based string classification system built on a 26-dimensional feature representation integrating frequency band energies, spectral statistics, and 13 Mel-Frequency Cepstral Coefficients, achieving 97.1% shuffled frame-level validation accuracy across 322,215 balanced frames. An ablation study identifies MFCCs as the primary accuracy driver (92.2% to 97.1%). We additionally introduce Comparison Training -- a data collection methodology in which adjacent open-string and fifth-fret string pairs are recorded in deliberate alternation -- and evaluate its contribution via confusion matrix analysis. Comparison Training reduces the D3 to A2 frame-level confusion rate by 44% but shows mixed results on other targeted pairs. A held-out free-play evaluation on 103,000 frames yields 87.8% overall accuracy. We describe the feature extraction pipeline in both Python and TypeScript to guarantee training-inference parity, and document two critical implementation failure modes. The system runs entirely in-browser with no hexaphonic pickup, fretboard sensor, camera, or multi-microphone setup required.

6. Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models

用扩散生成模型解决听觉注意力解码中的数据有限问题

AI 总结:研究针对助听器中听觉注意力解码因数据有限面临的挑战,利用扩散概率模型生成合成语音诱发EEG数据进行数据增强,实验证明该方法能显著提高AAD性能,凸显其减轻训练数据限制及增强模型鲁棒性的潜力。

链接:https://arxiv.org/abs/2607.18345

作者:David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic

英文摘要:Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephalogram (EEG) data to decode listener&#x27;s attention, enabling real-time tracking of specific sound sources. However, achieving high AAD performance with short time windows typical in HAs (<=1s) is challenging due to the scarcity of real-world speech-evoked EEG data. To address this issue, we investigate diffusion probabilistic models (DPMs) for generating synthetic speech-evoked EEG data. DPMs learn the underlying complex data structure through a denoising process and can generate realistic samples suitable for data augmentation. We evaluate the use of synthetic EEG data for augmenting datasets in locus-of-attention (LoA) classification tasks. Our experiments demonstrate that DPMs can generate realistic EEG signals and that incorporating synthetic data significantly improves AAD performance compared to models trained solely on measured EEG data (p<0.05). These results highlight the potential of diffusion-based data augmentation to mitigate training data limitations and improve the robustness of short-window AAD models in HA applications.

7. End-to-End Markov State Sequence Learning for Auditory Attention Decoding

用于听觉注意力解码的端到端马尔可夫状态序列学习

AI 总结:研究针对听觉注意力解码中多数模型为独立短窗口分类器的问题,提出基于条件随机场的端到端马尔可夫AAD框架,引入ESCNet,通过联合优化目标学习注意力状态序列,实验证明该方法在多个数据集上优于传统孤立窗口分类。

链接:https://arxiv.org/abs/2607.18614

机构:NERC-SLIP, University of Science and Technology of China(中国科学技术大学NERC-SLIP); School of Information and Communication Engineering, Guangzhou Maritime University(广州航海学院信息与通信工程学院); Artificial Intelligence Research Institute, iFLYTEK Company, Ltd.(科大讯飞股份有限公司人工智能研究院); School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学信息科学技术学院)

作者:Yushan Yashengjiang, Jie Zhang, Miao Sun, Huadong Liang, Xin Li, Zhen-hua Ling

英文摘要:Auditory attention decoding (AAD) identifies the speaker a listener attends to from neural responses like electroencephalography (EEG), making it a key algorithm in neuro-steered hearing aids. However, most neural AAD models are trained as independent short-window classifiers, despite auditory attention being a temporally persistent cognitive state and short-window EEG--audio evidence often being noisy and ambiguous. We propose an end-to-end Markov AAD framework based on conditional random field (CRF) that trains window-level neural emissions under a two-state attention prior. The framework treats the logits of any AAD backbone as Markov emissions, learns the transition rate from a standard HMM initialization, and jointly optimizes cross-entropy and CRF objectives, allowing temporal continuity to guide representation learning rather than merely smoothing predictions after training. We also introduce ESCNet, an EEG--speech correlation backbone that preserves time-aligned features and converts the difference between two mean Pearson correlations into state logits. We evaluate the framework with four emission backbones spanning correlation-based, convolutional, recurrent, and attention-based designs. On the dynamic AVGC dataset, CRF training generally outperforms post-hoc HMM smoothing; with ESCNet, it achieves $86.5\%$ causal and $92.4\%$ non-causal accuracy using $1$s windows. On the static KUL and USTC datasets, it improves causal decoding over fixed-rate post-hoc HMM baselines by $5.6\%$ and $2.0\%$, respectively, showing the superiority of learning AAD as attention state sequence over isolated-window classification.