由内蒙古大学、中国计算机学会语音对话与听觉专委会主办的语音语言技术分享会将于 2024年8月8日9:00 - 10:00 线上线下同步开启。

内蒙古大学丨语音语言技术分享会【第四期】

时间:8月8日(周四)9:00 ~ 10:00

形式:线上直播 & 线下参会

线下:内蒙古大学计算机学院119会议室


嘉宾&主题


Prof. Eng Siong Chng
Nanyang Technological University (NTU), Singapore
嘉宾介绍:Prof. Eng Siong Chng is currently an Associate Professor in the College of Computing and Data Science (CCDS) at Nanyang Technological University (NTU) in Singapore. Prior to joining NTU in 2003, he worked at Knowles Electronics (USA), Lernout and Hauspie (Belgium), the Institute of Infocomm Research (I2R) in Singapore, and RIKEN in Japan. He received both a PhD and a BEng (Hons) from the University of Edinburgh, U.K., in 1996 and 1991, respectively, specializing in digital signal processing. His areas of expertise include speech research, Large Language Models, machine learning, and speech enhancement.
报告题目:Enabling LLM for ASR
摘要:The decoder-only LLM, such as ChatGPT, was originally developed to accept only text input. Recent advances have enabled it to handle other modalities, such as audio, video, and images. This talk focuses on integrating speech modality into LLMs. The research community has proposed various innovative approaches for this task, including applying discrete representations, integrating pre-trained encoders with existing LLM decoder architectures (e.g., Qwen), multitask learning, and multimodal pretraining. In the talk, I will review recent approaches to the ASR task using LLMs and introduce two works from NTU’s Speech Lab: (i) “Hyporadise,” which applies LLMs to N-best hypotheses generated by traditional ASR models to improve the top-1 transcription result, demonstrating that LLMs not only exceed the performance of traditional language model rescoring but also recover and generate correct words not found in the N-best hypothesis—an ability we call Generative Error Correction (GER); and (ii) leveraging LLMs for ASR and noise-robust ASR by extending the Hyporadise approach to include noisy language embeddings, capturing the diversity of N-best hypotheses under low SNR conditions, and showing improved GER performance with fine-tuning.
线上参与

直播将通过语音之家微信视频号进行直播

手机端、PC端可同步观看
👇👇👇

线下参与
内蒙古大学计算机学院119会议室