日程安排
09:05-09:40 Junichi Yamagishi (National Institute of Informatics, Japan)
Building privacy-aware large-scale speech datasets through generative modelling
09:40-10:15 Kong Aik Lee (The Hong Kong Polytechnic University)
Uncertainty in Speaker Recognition: What it Represents and How to Handle It
10:15-10:35 茶歇
10:35-11:10 范存航 (安徽大学) Audio Deepfake Detection: Features, Models and Robustness
11:10-11:45 刘瑞 (内蒙古大学) Generative Conversational Speech Synthesis
14:00 -14:35 Tomoki Toda (Nagoya University, Japan)Voice conversion techniques to separately control static and dynamic speech characteristics
14:35-15:10 凌震华 (中国科学技术大学)Multimodal-Description-Driven and Machine-Perception-Oriented Voice Transformation
15:10-15:45 王龙标 (天津大学)Speech Decoupling-style Pre-training Method for Speech Representation Learning
15:45-16:05 茶歇
16:05-16:40 卢恒(阿里巴巴)Introducing Large Language Model based Speech Synthesis Project CosyVoice and Its Open Source Implementation
16:40-17:15 高莹莹 (中国移动研究院)Neural Network Disentanglement under Holistic Artificial Intelligence
17:15-17:20 闭幕
欢迎感兴趣的各界人士参加!

中国科学技术大学校内师生无需报名

Title:Voice conversion techniques to separately control static and dynamic speech characteristics
Abstract: In this talk, I will introduce our recent research on voice conversion (VC) to separately control diverse types of information embedded into a speech waveform. A seq-to-seq parallel VC framework is capable of jointly converting both static speech characteristics, such as speaker-specific voice quality, and dynamic speech characteristics, such as a speaker-specific speaking style. On the other hand, a frame-wise non-parallel VC framework converts the static speech characteristics while keeping the dynamic speech characteristics. By combining these two, we have developed a new VC framework capable of separately converting the static and dynamic speech characteristics, making it possible to apply VC to ground-truth free tasks where any target voices are unavailable. In this talk, I will review these VC frameworks and present specific examples of the ground-truth free tasks.
Biography: Tomoki Toda received his B.E. degree from Nagoya University, Japan, in 1999 and his M.E. and D.E. degrees from Nara Institute of Science and Technology (NAIST), Japan, in 2001 and 2003, respectively. He was a Research Fellow of the Japan Society for the Promotion of Science from 2003 to 2005. He was then an Assistant Professor (2005-2011) and an Associate Professor (2011-2015) at NAIST. From 2015, he has been a Professor in the Information Technology Center at Nagoya University. His research interests include information processing for sound media, such as speech, music, and environmental sounds.

Title: Building privacy-aware large-scale speech datasets through generative modelling
Abstract: The success of deep learning in speech and speaker recognition relies heavily on the use of large datasets. However, ethical, privacy and legal concerns also arise when using large datasets of speech collected from real human speech data. In particular, there are significant concerns in this regard when collecting large number of speaker’s speech data from the web. On the other hand, the quality of synthesised speech produced by recent generative models is very high. Is it possible to 'generate' large, privacy-aware, unbiased and fair datasets with speech generative models? Such studies have started not only for speech datasets but also for facial image datasets. In this talk, I will introduce our efforts to construct a synthetic VoxCeleb2 dataset called SynVox2 that is speaker-anonymised and privacy-aware. In addition to the procedures and methods used in the construction, the challenges and problems of using synthetic data will be discussed by showing the performance and fairness of a speaker verification system built using the SynVox2 database.
Biography: Junichi Yamagishi received a Ph.D. degree from the Tokyo Institute of Technology (Tokyo Tech), Tokyo, Japan, in 2006. From 2007 to 2013 he was a research fellow in the Centre for Speech Technology Research, University of Edinburgh, U.K. He became an associate professor with the National Institute of Informatics, Japan, in 2013, where he is currently a professor. His research interests include speech processing, machine learning, signal processing, biometrics, digital media cloning, and media forensics. He served as a co-organizer for the bi-annual ASVspoof Challenge and the bi-annual Voice Conversion Challenge. He also served as a member on the IEEE Speech and Language Technical Committee from 2013 to 2019, as an Associate Editor for IEEE/ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING (TASLP) from 2014 to 2017, as an Senior Area Editor for IEEE/ACM TASLP from 2019 to 2023, as the chairperson for ISCA SynSIG from 2017 to 2021, and as a member at large of IEEE Signal Processing Society Education Board from 2019 to 2023.

Title: Uncertainty in Speaker Recognition: What it Represents and How to Handle It
Abstract: Speaker identity information is among the most critical elements conveyed by speech signals. With accurate modeling, this information can be used for speaker recognition, speaker diarization, speech synthesis, and target speaker extraction, just to name a few. In modern speaker recognition systems, speaker identity information is captured in the form of compressed utterance-level representations, referred to as speaker embedding vectors, to recognize voices. The extraction of speaker embedding vectors is typically accomplished using deep neural networks trained on large speech corpora. Learning from data is inseparably connected with uncertainty. This talk will discuss the applicability of data uncertainty in speaker recognition. The so-called xi-vector approach developed for speaker embedding will be introduced as an example and case studies on speaker verification tasks show the way in which uncertainty modeling could be useful.
Biography: Kong Aik Lee is currently an Associate Professor at the Hong Kong Polytechnic University, Hong Kong. Before joining PolyU, he was a Principal Scientist and a Group Leader with the Agency for Science, Technology and Research (A*STAR), Singapore, while holding a joint appointment as an Associate Professor at the Singapore Institute of Technology, Singapore. From 2018 to 2020, he was a Senior Principal Researcher at the Data Science Research Laboratories, NEC Corporation, Tokyo, Japan. He received his Ph.D. from Nanyang Technological University, Singapore, in 2006. After this, he joined the Institute for Infocomm Research, Singapore, as a Research Scientist and then a Strategic Planning Manager (concurrent appointment). He was the recipient of the Singapore IES Prestigious Engineering Achievement Award 2013 and the Outstanding Service Award by IEEE ICME 2020. Since 2016, he has been an Editorial Board Member of Elsevier Computer Speech and Language. From 2017 to 2021, he was an Associate Editor for IEEE/ACM Transactions on Audio, Speech, and Language Processing. He was the General Chair of the Speaker Odyssey 2020 Workshop. He is currently serving as a Senior AE for IEEE Signal Processing Letters and as an elected Member of the IEEE Speech and Language Processing Technical Committee. His research interests include the automatic and para-linguistic analysis of speaker characteristics, ranging from speaker recognition, language and accent recognition, voice biometrics, spoofing, and countermeasures.

Title: Multimodal-Description-Driven and Machine-Perception-Oriented Voice Transformation
Abstract: Voice conversion is a task that has attracted a lot of research attentions in the last several decades, which aims to convert the speech of a source speaker to make it sound like uttered by a target speaker, while keeping linguistic contents unchanged. In this talk, we will introduce some of our recent studies on extending the scopes of conventional voice conversion tasks from two aspects. First, we extend the reference speech in voice conversion to other modalities by building multimodal-description-driven voice transformation tasks, including face-driven zero-shot voice conversion and voice attribute editing with text prompt. Second, we shift the focus of voice conversion from subjective perception to machine perception by defining machine-perception-oriented voice transformation tasks, including speaker adversarial speech generation and asynchronous voice anonymization. Some of our preliminary attempts on these tasks will also be discussed.
Biography: Zhen-Hua Ling is a Professor at the department of Electronic Engineering and Information Science of University of Science and Technology of China (USTC) and the vice director of National Engineering Research Centre of Speech and Language Information Processing (NERC-SLIP). He received the B.E. degree in electronic information engineering, the M.S. and Ph.D. degree in signal and information processing from USTC, in 2002, 2005, and 2008, respectively. From October 2007 to March 2008, he was a Marie Curie Fellow with the Centre for Speech Technology Research (CSTR), University of Edinburgh, U.K. From July 2008 to February 2011, he was a joint Postdoctoral Researcher with the University of Science and Technology of China, and iFLYTEK Co., Ltd., China. He also worked at the University of Washington, USA, as a Visiting Scholar from August 2012 to August 2013. His research interest covers speech synthesis and natural language processing. He is currently a member of the Speech and Language Technical Committee of IEEE Signal Processing Society and a committee member of ISCA SynSIG. He was the recipient of the IEEE Signal Processing Society Young Author Best Paper Award in 2010, and an Associate Editor of IEEE/ACM Transactions on Audio, Speech, and Language Processing from 2014 to 2018.

Title: Speech Decoupling-style Pre-training Method for Speech Representation Learning
Abstract:Self-supervised learning (SSL) has attracted significant attention in speech processing. Jointly improving the performance of pre-trained models across various downstream tasks, each requiring different speech information, poses significant challenges. In this talk, we’ll provide a brief overview of SSL’s background and recent advances firstly. Next, we’ll introduce a decoupling-style (progressive residual extraction) based pre-training method for speech representation learning. This method can progressively extract pitch variation, speaker, and content representations from the input speech. Finally, we’ll compare the experimental results between our proposed method and baseline SSL models, indicating joint performance improvements across various tasks.
Biography:Longbiao Wang received his Dr. Eng. degree from Toyohashi University of Technology, Japan, in 2008. He was an assistant professor in the faculty of engineering at Shizuoka University, Japan from April 2008 to September 2012. From October 2012 to August 2016 he was an associate professor at Nagaoka University of Technology, Japan. Currently he is a Professor at the Tianjin University, China and a visiting professor of Japan Advanced Institute of Science and Technology (JAIST), Japan. His research interests include acoustic signal processing, robust speech recognition, speaker recognition, speech emotion recognition and speech generation.

Title: Generative Conversational Speech Synthesis
Abstract:Conversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy understanding and expression. However, they often need to design complex network architectures and meticulously optimize the modules within them. In addition, due to the limitations of small-scale datasets containing scripted recording styles, they often fail to simulate real natural conversational styles. This report will focus on the difficulties and challenges faced by CSS, introduce the cutting-edge and the latest research progress of deep learning technology, especially large language model (LLM) technology, in CSS.
Biography:Rui Liu is currently a Professor in National and Local Joint Engineering Research Center of Mongolian Intelligent Information Processing, Inner Mongolia University. Rui Liu received Ph.D degree from Inner Mongolia University, China in 2020 and Bachelor degree in Taiyuan University of Technology, ShanXi, China in 2014. From 2019 to 2020, he has been an exchange PhD candidate at the Department of Electrical \& Computer Engineering of National University of Singapore (NUS), funded by China Scholarship Council (CSC). From 2020 to 2022, he worked as a Research Fellow at the Department of Electrical and Computer Engineering, National University of Singapore, Singapore. He was the recipient of the ``Best Paper Award'' at the 2021 International Conference on Asian Language Processing (IALP). His main research interests include speech synthesis, multimodal human-machine dialogue and the publications include top-tier NLP/ML/AI conferences and journals, including IEEE/ACM Transactions on Audio, Speech, and Language Processing, IEEE Transactions on Affective Computing, Neural Networks, AAAI, ICASSP, INTERSPEECH, etc. He is a member of IEEE, ISCA, CAAI, CIPS and CCF, and serves as the reviewer for many major referred journal and conference papers. His research has been supported by several research funds, including National Natural Science Foundation of China Young Scientists Fund and Guangdong Provincial Key Laboratory of Human Digital Twin, etc. (Website: ttslr.github.io)

Title: Audio Deepfake Detection: Features, Models and Robustness
Abstract:Speech synthesis and conversion technologies have obtained significant improvement in recent years and the generated speech is very realistic. The improper use of these technologies seriously threatens social stability and national security. Audio deepfake detection is proposed to solve this problem. In this report, I will share our recent works on audio deepfake detection from three aspects: features, models, and robustness.
Biography:Cunhang Fan, associate professor, School of Computer Science and Technology, Anhui University, master's supervisor, graduated from Institute of Automation, Chinese Academy of Sciences with a doctor's degree. I have been engaged in research on speech signal processing and brain computer interfaces. As the project leader, I have led and undertaken five national/provincial level projects, including the sub projects of National Key Research and Development Program of China, the National Natural Science Foundation of China, and the Open Research Projects of Zhejiang Lab. Received awards such as Outstanding Young Scientist of Anhui Province Computer Society in 2023 (two in total), the first prize of China Artificial Intelligence Society Teaching Achievement, and Second Prize of Anhui Province Teaching Achievement. Published over 30 academic papers in international conferences and journals such as CCF-A Conference and IEEE/ACM Transactions. Served as the area and publication chair for international conferences such as APSIPA ASC 2023, IEEE SLT 2024, and IALP 2024.

Title:CosyVoice 生成式语音大模型及其开源项目介绍
Abstract:最近随着神经网络以及生成式大模型技术的发展, 语音以及音频生成领域的技术发展日新月异。和传统的基于小模型的语音生成方案相比,基于生成式大模型的语音及音频生成在发音自然度、相似度、发音准确性上都有极大提升,并且涌现出小模型不具有的比如: zero-shot 零样本声音复刻 , 以及基于 Instruct Learning 的副语言学习能力等等。
本次报告主要介绍CosyVoice 的基本原理框架 。CosyVoice 是通义实验室语音团队全新推出生成式语音大模型,在语调、韵律、情感表达等方面获得极大提升,提供超自然拟人的语音合成能力,并已在通义APP语音通话等功能中率先落地。不同于其他基于声学或弱语义token的生成模型,CosyVoice基于有监督训练得到的强语义token,能够很好的将语义、声学和韵律进行解耦,同时具备很好的data scaling-up特性,在不同规模的训练数据下均具有较好的表现。通过在海量的多语种、多情感数据上进行训练,CosyVoice具备很强的zero-shot音色克隆、跨语言合以及副语言carry-on的能力,适用于广泛的语音合成场景。
本次报告同介绍CosyVoice 开源项目。CosyVoice 开源项目也是业界首次开源10w+ 小时训练语料训练的工业级别的语音生成式大模型。并且除了中文以外,还拥有中、英、日、韩多语种生成能力 ;高兴,悲伤,害怕等情感合成 ;以及语速、音调等细粒度控制的副语言合成能力。同时基于开源模型,用户能够实现zero-shot 零样本声音复刻,利用5-10秒钟的目标发音人声音即可定制目标说话人音色和风格。
Biography:卢恒,2011年博士毕业于中国科学技术大学语音及语言信息处理国家工程实验室。先后于美国Nuance Communications,腾讯 AI lab西雅图分部任研究员,曾任喜马拉雅首席科学家。目前担任阿里通义语音实验室生成式大模型负责人,主导语音生成式大模型的研发项目。目前同时担任CCF中国计算机学会语音对话与听觉专委会委员,任多项国际会议以及期刊的审稿人。研究方向主要包括多模态语音合成,说话人转换,歌声的生成及转换,语音识别以及语音评测,虚拟人等。在各类国际会议和刊物中发表论文以及专利60篇。

Title: Neural Network Disentanglement under Holistic Artificial Intelligence
Abstract:In the past few decades, a large number of AI models have been trained and supported thousands of industrial applications. It is often necessary to combine multiple models or capabilities, including foundation models, industry models, or small models for specific tasks, while meeting multiple constraints such as compute, transport, security, and controllability, and be able to optimize end-to-end to serve business objectives. To this end, China Mobile proposes a new AI technology framework, Holistic AI (HAI). The central control unit of HAI powered by foundation models maps the user’s requirement to an execution plan consisting of a specific model or a cascaded process of multiple models, under the constraints of computing and network facilities for efficient deployment. Under HAI, the foundation models or small models from model zoo are disassembled and recombined for dedicated task, based on the principles of high reuse, easy scheduling, self-closed loop, and easy adaptation. With the success of generic large models, the work related to network disentanglement has gradually increased. We classify the relevant methods into three types: 1) model-zoo-based, 2) task-specific distillation or pruning, and 3) sparse training, and introduce our research on network disentanglement around these three aspects.
Biography:Yingying Gao, Ph.D., graduated from Beijing Jiaotong University and currently works at the Artificial Intelligence and Intelligent Operations Center of China Mobile Research Institute. Her research focuses on speech language understanding and Holistic Artificial Intelligence.
