透過您的圖書館登入
IP:216.73.216.200
  • 期刊
  • OpenAccess

Noisy Channel Models for Corrupted Chinese Text Restoration and GB-to-Big5 Conversion

並列摘要


In this article, we propose a noisy channel/information restoration model for error recovery problems in Chinese natural language processing. A language processing system is considered as an information restoration process executed through a noisy channel. By feeding a large-scale standard corpus C into a simulated noisy channel, we can obtain a noisy version of the corpus N. Using N as the input to the language processing system (i.e., the information restoration process), we can obtain the output results C'. After that, the automatic evaluation module compares the original corpus C and the output results C', and computes the performance index (i.e., accuracy) automatically. The proposed model has been applied to two common and important problems related to Chinese NLP for the Internet: corrupted Chinese text restoration and GB-to-BIG5 conversion. Sinica Corpora version 1.0 and 2.0 are used in the experiment. The results show that the proposed model is useful and practical.

並列關鍵字

無資料

參考文獻


Chang, C. H.(1992).Proceedings of ICCPCOL-92.
Chang, C. H.(1993).Proceedings of Workshop on Very Large Corpora.
Chang, C. H.(1996).Simulated Annealing Clustering of Chinese Words for Contextual Text Recognition.Pattern Recognition Letters.17,57-66.
Chang, C. H.(1994).Proceedings of COLING-94.
Chen, C. D.,Chang, C. H.(1996).Application Issues of SA-class Bigram Language Models.Computer Processing of Oriental Languages.10(1),1-15.

延伸閱讀