INTERSPEECH 2025 论文预讲会由CCF语音对话与听觉专委会、语音之家主办,旨在为学者们提供更多的交流机会,更方便、快捷地了解领域前沿。活动将邀请 INTERSPEECH 2025 录用论文的作者进行报告交流。

INTERSPEECH 2025 论文预讲会第五期邀请到香港中文大学(深圳)的李珈祺做本次会议的专场分享,欢迎大家观看。

第五期 - 个人专场

时间:7月29日(周二)19:00 ~ 19:30

形式:线上

议程:每位嘉宾分享30分钟(含5分钟QA)

嘉宾&主题

嘉宾简介:李珈祺,香港中文大学(深圳)一年级博士生,导师为香港中文大学(深圳)教授武执政,研究方向包括音频编解码器、大语言TTS模型。曾在Interspeech、ICASSP、SLT等会议发表学术论文。

分享主题:DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation

摘要:Neural audio codecs serve as the fundamental building blocks for speech language model-based speech generation. To improve speech generation performance, recent neural audio codec SpeechTokenizer proposed to distill the first-layer codec tokens from semantic-rich self-supervised (SSL) representations. Our work improves on this idea of semantically-enhanced audio codec, making it more usable for speech generation. 1) We increase the semantic information accuracy in first-layer codec tokens by proposing a dual encoding method. Dual encoding replaces semantic distillation with a two-stream encoding of SSL and waveform in an end-to-end codec framework. 2) We achieve outstanding low bitrate audio reconstruction quality with DAC-based waveform encoding and adopting larger codebooks. 3) We adopt low frame rates of 25Hz and 12.5Hz to improve speech generation efficiency. The resulting codec model, DualCodec, outperforms existing audio codecs including SpeechTokenizer and Mimi in both audio reconstruction and text-to-speech performance, making it ideal for efficient speech synthesis. We open-source our models and codes.

论文链接:http://arxiv.org/abs/2505.13000

代码链接:https://github.com/jiaqili3/dualcodec

Demo链接:https://dualcodec.github.io/

参与方式

直播将通过语音之家微信视频号进行直播

手机端、PC端可同步观看

👇👇👇

预讲会征集

INTERSPEECH 2025 论文预讲会面向全球线上招募,结合定向邀请与征集报名的方式来选择预讲会的嘉宾。

为了共创高质量的论文预讲会,我们诚挚邀请所有 INTERSPEECH 2025 作者参与到此次预讲会活动中来,也欢迎大家推荐适合此次预讲会活动的学者。

预讲会报名方式

联系人邮箱:bd@speechhome.com

联系人微信