第21届国际语音通讯协会年会INTERSPEECH 2020将于2020年10月25日- 30日全程在线举行。本次大会共邀请到4位重磅嘉宾做大会主题报告(Keynote Talk),涉及的演讲内容包括语言建模、语音听觉感知、口语技术和面向会话应用的底层语音交互技术等四个方面。今天给大家介绍的演讲嘉宾是:Amazon研发主管,Alexa技术负责人Shehzad Mevawalla,将给大家带来题为“Successes, Challenges and Opportunities for Speech Technology in Conversational Agents”的大会报告。

北京时间10月29日18:00到19:00
Shehzad Mevawalla将发表题为“Successes, Challenges and Opportunities for Speech Technology in Conversational Agents”的主题演讲,欢迎大家在线观看。
Keynotes:
Title: Successes, Challenges and Opportunities for Speech Technology in Conversational Agents
Time: Thursday, 29 October, 18:00-19:00 (GMT+8)
Speaker: Shehzad Mevawalla, Amazon Alexa
Abstract:
From the early days of modern ASR research in the 1990s, one of the driving visions of the field has been a computer-based assistant that could accomplish tasks for the user, simply by being spoken to. Today, we are close to achieving that vision, with a whole array of speech-enabled AI agents eager to help users. Amazon’s Alexa pioneered the AI assistant concept for smart speaker devices enabled by far-field ASR. It currently supports billions of customer interactions per week, on over 100 million devices across multiple languages. This keynote will give an overview of the interplay between underlying speech technologies, including wakeword detection, endpointing, speaker identification, and speech recognition that enable Alexa. We highlight the complexities of combining these technologies into a seamless and robust speech-enabled user experience under large production load and real-time constraints. Interesting algorithmic and engineering challenges arise from choices between deployment in the cloud versus on edge devices, and from constraints on latency and memory versus trade-offs in accuracy. Adapting recognition systems to trending topics, changing domain knowledge bases, and to the customer’s personal catalogs adds additional complexity, as does the need to support adaptive conversational behavior (such as normal versus whispered speech). We also dive into the unique data aspects of large-scale deployments like Alexa, where a continuous stream of unlabeled data enables successful applications of weakly supervised learning. Finally, we highlight problems for the speech research community that remain to be solved before the promise of a fully natural, conversational assistant is fully realized.
