  • 期刊
  • OpenAccess

Design and Evaluation of Approaches to Automatic Chinese Text Categorization


In this paper, we propose and evaluate approaches to categorizing Chinese texts, which consist of term extraction, term selection, term clustering and text classification. We propose a scalable approach which uses frequency counts to identify left and right boundaries of possibly significant terms. We used the combination of term selection and term clustering to reduce the dimension of the vector space to a practical level. While the huge number of possible Chinese terms makes most of the machine learning algorithms impractical, results obtained in an experiment on a CAN news collection show that the dimension could be dramatically reduced to 1200 while approximately the same level of classification accuracy was maintained using our approach. We also studied and compared the performance of three well known classifiers, the Rocchio linear classifier, naive Bayes probabilistic classifier and k-nearest neighbors (kNN) classifier, when they were applied to categorize Chinese texts. Overall, kNN achieved the best accuracy, about 78.3%, but required large amounts of computation time and memory when used to classify new texts. Rocchio was very time and memory efficient, and achieved a high level of accuracy, about 75.4%. In practical implementation, Rocchio may be a good choice.


Baker, Douglas,McCallum, Kachites(1998).Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR'98).
簡立峰 Lee-Feng, Lee-Feng(1999).The Fourth International Workshop on Information Retrieval with Asian Languages (IRAL'99).
Ferragina, Paolo,Grossi, Roberto(1999).The String B-tree: A New Data Structure for String Search in External Memory and its Application.Journal of ACM.46(2),236-280.
Frobelink, Marko,Mladenic, Dunja(1998).Proceedings of the 13th European Conference on Artifical Intelligence.
Furnas, G. W.,Dumais, S. T.,Landauer, T. K.,Deerwester, S. C.,Harshamn, R. A.(1990).Indexing by Latent Semantic Analysis.Journal of the American Society for Information Science.41(6),391-407.


Chen, Y. C. (2005). 中文零代詞解析與應用 [doctoral dissertation, Tatung University]. Airiti Library. https://www.airitilibrary.com/Article/Detail?DocID=U0081-0607200917233352
Budiansyah, A. (2010). Text Trend Analysis via Significant Term A Based on Indonesia News [master's thesis, Asia University]. Airiti Library. https://www.airitilibrary.com/Article/Detail?DocID=U0118-1511201215465544
