微信公众号:arXiv_Daily
[机构]信息由AI分析生成,可能存在错误,仅供参考,以论文实际显示为准
快速导航
1. 语音识别与关键词检测 1 篇
2. 音频事件检测与场景理解 1 篇
3. 音乐信息检索与音乐生成 1 篇
4. 语音翻译与语音语言模型 1 篇
5. 其他/综合语音音频 1 篇
1. 语音识别与关键词检测 | 1 篇
1. RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
RW-Voice-EQ基准:评估语音人工智能系统的真实世界基准
AI 总结:研究提出真实世界语音EQ基准,用于跨TTS、STS、SU和ASR评估语音人工智能,指出当前基准多评估孤立能力,新基准能考量声学等信息,评估表明性能依赖维度,语音人工智能应综合多方面能力评估而非单一总分。
链接:https://arxiv.org/abs/2607.14846
机构:Hume AI Research(休谟人工智能研究)
作者:David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr Cłapa, Theo Lebryk, Jens Madsen, Olya Ossipova, Sharath Rao, Hoon Shin, Tigran Soghbatyan, Georg Streich, Rashish Tandon, Panagiotis Tzirakis
英文摘要:Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.
2. 音频事件检测与场景理解 | 1 篇
2. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
令牌能竞争吗?针对BirdCLEF+ 2026的与监督式CNN主干竞争的令牌表示
AI 总结:针对BirdCLEF+ 2026中动物发声多标签检测任务,先构建监督基线,后对比神经音频编解码器的编解码表示与基础嵌入的语义表示,比较两个生物声学专家模型和四个基于令牌的编码器,探究基于令牌的表示能否竞争。
链接:https://arxiv.org/abs/2607.14474
机构:Georgia Institute of Technology(佐治亚理工学院)
作者:Anthony Miyaguchi, Murilo Gustineli, Adrian Cheung
英文摘要:This paper details the DS@GT ARC team's approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the Pantanal wetlands. The 2026 edition adds about an hour of labeled soundscapes, shifting the task toward supervised pipelines fit to the labeled set. First, we build a competitive supervised baseline that ensembles a frozen Perch v2 backbone, a trained HGNetV2-B0 sound-event-detection network, and a non-bird prototypical head, reaching a private leaderboard score of 0.936 at rank 1894 within a 90-minute CPU budget. Second, we ask whether token-based representations can compete, contrasting codec representations from neural audio codecs against semantic representations from foundational embeddings. We compare two bioacoustic specialist models against four token-based encoders trained on AudioSet. The repository for this work can be found at this https URL.
3. 音乐信息检索与音乐生成 | 1 篇
3. MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
MIDI-RAE-JEPA:用于符号音乐的分层表示学习与生成
AI 总结:研究针对符号音乐表示自监督方法探索不足的问题,提出MIDI-RAE-JEPA,结合音高和时间移位等方差目标、LeJEPA及Swin Transformer V2编码器学习分层表示,经实验验证该方法在多方面表现良好,为符号音乐表示提供了可行途径。
链接:https://arxiv.org/abs/2607.14537
作者:Scott H. Hawley
英文摘要:Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives -- including a masked embedding predictor (MEP) -- with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.
4. 语音翻译与语音语言模型 | 1 篇
4. Large Audio Language Models for Spoofing-Aware Speaker Verification
用于防欺骗语音识别的大型音频语言模型
AI 总结:研究用于防欺骗语音识别的大型音频语言模型,通过零样本提示等多种方式进行系统评估,发现预训练模型零样本时不佳,特定任务适应可改善,还找到多种实现竞争力性能的途径,为统一防欺骗语音识别提供基础。
链接:https://arxiv.org/abs/2607.14753
机构:Applied AI Institute(应用人工智能研究所); MIRAI(未来人工智能研究机构); HSE(高等经济学院); MTUCI(莫斯科国立通信与信息技术大学)
作者:Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh, Oleg Y. Rogov
英文摘要:Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.
5. 其他/综合语音音频 | 1 篇
5. ITGPT: A Transformer Based Architecture for the Generation of Dance Dance Revolution and In the Groove Charts
ITGPT:一种基于Transformer的用于生成《劲舞革命》和《狂热节拍》谱面的架构
AI 总结:研究针对《劲舞革命》和《狂热节拍》谱面生成难题,提出基于Transformer的ITGPT架构,相比前人工作,该架构在生成准确性和计算成本上有显著提升。
链接:https://arxiv.org/abs/2607.14148
作者:Miguel O'Malley
英文摘要:Dance Dance Revolution and In the Groove are rhythm games consisting of songs and accompanying choreography, referred to as charts. Players press arrows on a device referred to as a dance pad in time with steps determined by the song's chart. The process of manual chart generation is timestaking and difficult, motivating interest in automation. We propose ITGPT, a new transformer based architecture for the generation of DDR/ITG charts, and demonstrate significant improvements to generation accuracy and computational cost in comparison to predecessor work.
