透過您的圖書館登入
IP:3.129.23.30
  • 期刊
  • OpenAccess

Histogram Equalization on Statistical Approaches for Chinese Unknown Word Extraction

並列摘要


With the evolution of human lives and the spread of information, new things emerge quickly and new terms are created every day. Therefore, it is important for natural language processing systems to extract new words in progression with time. Due to the broad areas of applications, however, there might exist the mismatch of statistical characteristics between the training domain and the testing domain, which inevitably degrades the performance of word extraction. This paper proposes a scheme of word extraction in which histogram equalization for feature normalization is used. Through this scheme, the mismatch of the feature distributions due to different corpus sizes or changes of domain can be compensated for appropriately such that unknown word extraction becomes more reliable and applicable to novice domains. The scheme was initially evaluated on the corpora announced in SIGHAN2. 68.43% and 71.40% F-measures for word identification, which correspond to 66.72%/32.94% and 75.99%/58.39% recall rates for IV/OOV, respectively, were achieved for the CKIP and the CUHK test sets, respectively, using four combined features with equalization. When applied to unknown word extraction for a novice domain, this scheme can identify such pronouns as ”海角七號” (Cape No. 7, the name of a film), ”蠟筆小新” (Crayon Shinchan, the name of a cartoon figure), ”金融海嘯” (Financial Tsunami) and so on, which cannot be extracted reliably with rule-based approaches, although the approach appears not so good at identifying such terms as the names of humans, places, or organizations, for which the semantic structure is prominent. This scheme is complementary with the outcomes of two word segmentation systems, and is promising if other rule-based approaches could be further integrated.

參考文獻


Sun, M.S.,Huang, C. N.,Gao, H.Y.,Fang, J.(1994).Identifying Chinese Name in Unrestricted Texts.Journal of Chinese Language and Computing.4(2),113-122.
Lu, X. Q.,Zhang, L.,Hu, J. F.(2005).Statistical Substring Reduction in Linear Time.Lecture Notes in Computer Science.3248,320-327.
Zhao, H.,Kit, C. Y.(2008).An Empirical Comparison of Goodness Measures for Unsupervised Chinese Word Segmentation with a Unified Framework.Proceedings of The 3nd International Joint Conference on Natural Language Processing(IJCNLP).(Proceedings of The 3nd International Joint Conference on Natural Language Processing(IJCNLP)).
Chen, K. J.,Ma, W. Y.(2002).Unknown Word Extraction for Chinese Documents.Proceedings of The 19nd International Conference on Computational Linguistics (COLING).(Proceedings of The 19nd International Conference on Computational Linguistics (COLING)).
梁婷、葉大榮()。

延伸閱讀