COMMUNITY

BOARD

News

Accepted at INTERSPEECH, a Globally Prestigious Conference in Speech AI

Author
College of Software Convergence
Date
2023-05-23
View
21

Files

Two papers by master's students of the Intelligent Spoken Dialogue Systems Laboratory (ISDS) research team

Accepted at INTERSPEECH, a Globally Prestigious Conference in Speech AI


▲ (Top row, from left) Professor Koo Myoung-wan, Department of Computer Science & Engineering; Choi Ye-rin, master's student, Department of Artificial Intelligence

(Bottom row, from left) Yeon Hee-yeon and Kim Min-ju, master's students, Department of Artificial Intelligence

 

Text-to-speech (TTS) and spoken language understanding (SLU) research by master's students of the University's Department of Artificial Intelligence will be presented at a globally prestigious conference on spoken language processing.

 

Two papers submitted by master's students of the Intelligent Spoken Dialogue Systems Laboratory (ISDS), supervised by Professor Koo Myoung-wan of the Department of Computer Science & Engineering, have been accepted at INTERSPEECH 2023, held in Dublin, Ireland, from 20 to 24 August (local time).

 

'DC CoMix TTS: An End-to-End Expressive TTS with Discrete Code Collaborated with Mixer', speech synthesis research by master's student Choi Ye-rin of the graduate Department of Artificial Intelligence (3rd semester) and Professor Koo Myoung-wan (corresponding author), presents a technique for richer emotional expression in synthesized speech. The team proposed using quantized vectors extracted from a neural audio codec model as model inputs so that acoustic information such as prosody can be richly captured during speech synthesis. With this method they demonstrated improved performance in speaker-independent prosody expression in both qualitative and quantitative evaluation.


 

▲ Overview of the DcCoMix TTS architecture proposed by Choi Ye-rin, master's student, Department of Artificial Intelligence

 

'I Learned Error, I Can Fix It! : A Detector-Corrector Structure for ASR Error Calibration', by master's students of the same department Yeon Hee-yeon (3rd semester) and Kim Min-ju (2nd semester) (joint first authors) and Professor Koo Myoung-wan (corresponding author), proposes a model that automatically corrects speech recognition errors in spoken utterances. Speech recognition errors are a representative problem that undermines language understanding performance in the spoken dialogue pipeline. The proposed model consists of a detector trained on a wide variety of simulated speech recognition errors and a corrector trained to predict the correct phrase for the segments identified as errors. The team experimentally demonstrated that the proposed model improves performance on downstream spoken language understanding tasks.

 

▲ Overview of the detector-corrector structure proposed by Yeon Hee-yeon and Kim Min-ju, master's students, Department of Artificial Intelligence

 

Both pieces of research are achievements at an international conference by teams made up entirely of master's students with relatively short engineering research careers. The students who wrote the papers belong to the first and second cohorts of the Department of Artificial Intelligence, established last year. All majored as undergraduates in the humanities, social sciences or convergence disciplines — economics (Choi Ye-rin), Korean literature (Kim Min-ju) and Art & Technology (Yeon Hee-yeon).

 

Professor Koo said: "INTERSPEECH is where the world's finest researchers in speech share their work, so this means our master's students are doing research on a par with them," adding, "I will continue to guide my students so that they can find creative ways of solving problems in speech and language processing research."