Don't Just Scratch the Surface: Enhancing Word Representations for Korean with Hanja
Information
Title
Don't Just Scratch the Surface: Enhancing Word Representations for Korean with Hanja
Authors
Kang Min Yoo*, Taeuk Kim*, Sang-goo Lee
(*equal contribution)
Year
2019 / 11
Keywords
natural language processing, word representations, korean processing
Acknowledgement
BK21 Plus for Pioneers in Innovative Computing
Publication Type
International Conference
Publication
2019 Conference on Empirical Methods in Natural Language Processing (EMNLP 2019)
Link
url
Abstract
We propose a simple yet effective approach for improving Korean word representations using additional linguistic annotation (i.e. Hanja). We employ cross-lingual transfer learning in training word representations by leveraging the fact that Hanja is closely related to Chinese. We evaluate the intrinsic quality of representations learned through our approach using the word analogy and similarity tests. In addition, we demonstrate their effectiveness on several downstream tasks, including a novel Korean news headline generation task.