“CCF走进高校”是CCF组织的系列演讲公益活动,中国计算机学会(CCF)致力于为计算机领域专业人士的职业发展提供服务。CCF语音专委走进高校第一站是上海交通大学,本次活动由CCF主办,由CCF语音对话与听觉专业委员会、上海交通大学跨媒体语言智能实验室承办。活动将于2022年10月22日线上举行,欢迎感兴趣的老师同学在线参加! 活动日程活动时间:2022年10月22日9:30-11:30主持人:钱彦旻 教授参会方式:腾讯会议会议ID:477-7360-5980会议链接:https://meeting.tencent.com/dm/r7LUoKqUTLyR 主持人 钱彦旻 教授 个人介绍:上海交通大学计算机科学与工程系教授,博士生导师。清华大学博士,英国剑桥大学工程系博士后。国家优秀青年基金、上海市青年英才扬帆计划、吴文俊人工智能自然科学奖一等奖(第一完成人)获得者。现为IEEE高级会员、ISCA会员,同时也是国际开源项目Kaldi语音识别工具包的13位创始成员之一。担任InterSpeech, ISCSLP等国际会议的领域主席和TPC委员;IEEE T-ASLP, IEEE J-STSP, IEEE SPL, ICASSP, InterSpeech等期刊和国际会议审稿人。有10余年从事智能语音及语言处理、人机交互、模式识别及机器学习的研究和产业化工作经验。在本领域的一流国际期刊和会议上发表学术论文200余篇,Google Scholar引用总数10000余次,申请60余项中美专利,合作撰写和翻译多本外文书籍。3次获得领域内国际权威期刊和会议的最优论文奖,3次带队获得国际评测冠军。作为负责人和主要参与者参加了包括国家自然科学基金、国家脑科学计划、国家重点研发计划、国防JKW、国家863、英国EPSRC等多个项目。目前的研究领域包括:语音识别,说话人和语种识别,语音抗噪与分离,语音情感感知,自然语言理解,深度学习建模,多媒体信号处理等。 特邀嘉宾 俞凯 教授 个人介绍:俞凯,上海交通大学计算机系教授,清华大学本科、硕士,剑桥大学博士,中组部“青年千人”“万人计划”,国家自然科学基金委“优秀青年科学基金”,上海市“东方学者”特聘教授,IEEE高级会员,中国大陆高校首个IEEE Speech and Language Processing Technical Committee成员,中国计算机学会人机交互专委会委员,中国声学学会语音语言、听觉及音乐分会执委会委员,中国语音产业联盟技术工作组副组长。在国际权威期刊及会议发表论文200余篇,其中两篇论文获得InterSpeech2010最佳论文,曾经担任InterSpeech等国际会议语音及对话系统领域主席。他在2011年获对话系统国际挑战赛可控测试冠军,2013年获国际语音通信联盟(ISCA)2008-2012 Computer Speech Language 最佳论文奖,2014年获得中国人工智能学会吴文俊奖,2016年ISCSLP最佳论文奖、获评《科学中国人》年度人物,2018年中国计算机学会“青竹奖”。此外,他还是苏州思必驰公司创始人兼首席科学家,思必驰入选2016 高盛全球人工智能报告《AI, Machine Learning and Data Fuel the Future of Productivity》“Key AI Players”及2017年Gartner “Cool Vendors for AI”。
主题报告
张王优 博士研究生 讲者介绍:Wangyou Zhang is a fifth-year Ph.D. student in X-LANCE Lab, Department of Computer Science and Engineering, Shanghai Jiao Tong University. His supervisor is Prof. Yanmin Qian. His research interest includes speech signal processing and robust speech recognition.报告题目:End-to-End Dereverberation, Beamforming, and Speech Recognition in a Cocktail Party 报告简介:Humans are known to have the capability of listening to, following and recognizing one target speaker in a "cocktail party" scenario where multiple talkers speak simultaneously with the presence of background noise and reverberation. While humans can easily focus their auditory attention on one sound source in the cocktail party scenario, it is very difficult for machines to separate and pay attention to one or two sounds of interest in such complex auditory conditions. In the past few decades, researchers have tried to develop algorithms for machines to mimic humans' capability, but the performance is still far from satisfactory. In this report, I will present our recent progress efforts on multi-channel speech processing in the cocktail party problem, and also discuss the new challenges and exciting directions to solve the cocktail party problem. 陈正阳 博士研究生 讲者介绍:Zhengyang Chen is now a fourth-year Ph.D. student in Shanghai Jiao Tong University, Shanghai, China, under the supervision of Professor Yanmin Qian. His current research interests include speaker recognition, speaker diarization and deep learning.报告题目:The SJTU X-LANCE System for Speaker Recognition Challenge VoxSRC2022 and CNSRC2022 报告简介:The speaker recognition task has received a lot of attention due to its efforts to use human voice to identify people. The general research papers often only focus on a certain technical problem and the speaker recognition challenge requires combining different advanced techniques. The Voxceleb Speaker Recognition Challenge (VoxSRC) is the most popular speaker recognition challenge, which is based on the Voxceleb audio dataset downloaded from Youtube. The CN-Celeb Speaker Recognition Challenge (CNSRC) is held for the first time this year, which is based on the CN-Celeb audio dataset downloaded from Chinese social medial websites. We participated in both challenges this year, achieving 1st place in CNSRC2022 Track1 and 3rd place in VoxSRC2022 Track1&3. In this report, we will show how to build a strong speaker recognition challenge system and which technology should be used in each module. 徐薛楠 博士研究生 讲者介绍:Xuenan Xu is currently a second-year PhD candidate at X-Lance Lab, Shanghai Jiao Tong University, under supervision from Prof. Mengyue Wu and Kai Yu. His main research interests lie on general audio processing, including detection and classification of acoustic scenes and events, and the interaction between audio signal and natural language processing.报告题目:Connecting sounds and natural language: recent advances in audio-text research 报告简介:Speech and language processing has attracted much attention from previous research while non-speech sounds (e.g., environmental sounds, music, animal and human non-vocal sounds) also plays an important role in our daily life. Humans can comprehensively perceive and describe these sounds with natural language, yet it is still a challenging task for machine perception and cognition. Research bridging general sounds understanding and natural language processing remains under-investigated. In this talk, we mainly focus on one audio-text task: automatic audio captioning (AAC). We will summarize recent advances in AAC and introduce two recent works from our team: diversity-controllable captioning and temporal relation-controllable captioning, where a control signal is used to control the output format while maintaining its semantic content. Finally we will give a brief introduction to other audio-text tasks, including text-to-audio grounding, audio-text retrieval and audio question answering. 陈 志 博士研究生 讲者介绍:Zhi Chen is now a five-year Ph.D. student in Shanghai Jiao Tong University, Shanghai, China, under the supervision of Professor Kai Yu. His current research interests include dialogue system, pre-training language model and reinforcement learning.报告题目:Unified Generative Model for Knowledge-Grounding Dialogue 报告简介:Conversation based on natural language may be the most important way to exchange knowledge for human beings. Except for dialogue content, grounding knowledge is also essential in helping us understand the dialogue better. There are two core dialogue research topics: dialogue understanding and dialogue generation. They both rely on external knowledge, like the ontology for task-oriented dialogue generation and the database for NL-to-SQL. However, there are different task definitions for various dialogue-oriented tasks, like slot filling with sequence labeling method and intent detection with classification method.Building a universal conversational agent has been a long-standing goal of the dialogue research community. Most previous works only focus on a small set of dialogue tasks. In this talk, We aim to build a unified generative model which can be used to solve massive dialogue tasks. To achieve this goal, a large-scale well-annotated dialogue dataset with rich task diversity is collected. We introduce a framework to unify all dialogue tasks and propose novel auxiliary self-supervised tasks to achieve stable training of the pre-trained dialogue model on the highly diverse large-scale dialogue corpus.