Histogram Equalization on Statistical Approaches for Chinese Unknown Word Extraction

With the evolution of human lives and the spread of information, new things emerge quickly and new terms are created every day. Therefore, it is important for natural language processing systems to extract new words in progression with time. Due to the broad areas of applications, however, there might exist the mismatch of statistical characteristics between the training domain and the testing domain, which inevitably degrades the performance of word extraction. This paper proposes a scheme of word extraction in which histogram equalization for feature normalization is used. Through this scheme, the mismatch of the feature distributions due to different corpus sizes or changes of domain can be compensated for appropriately such that unknown word extraction becomes more reliable and applicable to novice domains. The scheme was initially evaluated on the corpora announced in SIGHAN2. 68.43% and 71.40% F-measures for word identification, which correspond to 66.72%/32.94% and 75.99%/58.39% recall rates for IV/OOV, respectively, were achieved for the CKIP and the CUHK test sets, respectively, using four combined features with equalization. When applied to unknown word extraction for a novice domain, this scheme can identify such pronouns as ”海角七號” (Cape No. 7, the name of a film), ”蠟筆小新” (Crayon Shinchan, the name of a cartoon figure), ”金融海嘯” (Financial Tsunami) and so on, which cannot be extracted reliably with rule-based approaches, although the approach appears not so good at identifying such terms as the names of humans, places, or organizations, for which the semantic structure is prominent. This scheme is complementary with the outcomes of two word segmentation systems, and is promising if other rule-based approaches could be further integrated.

並列關鍵字

Unknown Word Extraction ； Word Identification ； Machine Learnin ； Multilayer Perceptrons ； Histogram Equalization

參考文獻

Sun, M.S.,Huang, C. N.,Gao, H.Y.,Fang, J.(1994).Identifying Chinese Name in Unrestricted Texts.Journal of Chinese Language and Computing.4(2),113-122.

Google Scholar

Lu, X. Q.,Zhang, L.,Hu, J. F.(2005).Statistical Substring Reduction in Linear Time.Lecture Notes in Computer Science.3248,320-327.

Google Scholar

Zhao, H.,Kit, C. Y.(2008).An Empirical Comparison of Goodness Measures for Unsupervised Chinese Word Segmentation with a Unified Framework.Proceedings of The 3nd International Joint Conference on Natural Language Processing(IJCNLP).(Proceedings of The 3nd International Joint Conference on Natural Language Processing(IJCNLP)).

Google Scholar

Chen, K. J.,Ma, W. Y.(2002).Unknown Word Extraction for Chinese Documents.Proceedings of The 19nd International Conference on Computational Linguistics (COLING).(Proceedings of The 19nd International Conference on Computational Linguistics (COLING)).

Google Scholar

梁婷、葉大榮()。

Google Scholar

國際替代計量

Histogram Equalization on Statistical Approaches for Chinese Unknown Word Extraction

全文下載

主題瀏覽