透過您的圖書館登入
IP:18.118.200.86
  • 期刊
  • OpenAccess

Design and Evaluation of Approaches to Automatic Chinese Text Categorization

並列摘要


In this paper, we propose and evaluate approaches to categorizing Chinese texts, which consist of term extraction, term selection, term clustering and text classification. We propose a scalable approach which uses frequency counts to identify left and right boundaries of possibly significant terms. We used the combination of term selection and term clustering to reduce the dimension of the vector space to a practical level. While the huge number of possible Chinese terms makes most of the machine learning algorithms impractical, results obtained in an experiment on a CAN news collection show that the dimension could be dramatically reduced to 1200 while approximately the same level of classification accuracy was maintained using our approach. We also studied and compared the performance of three well known classifiers, the Rocchio linear classifier, naive Bayes probabilistic classifier and k-nearest neighbors (kNN) classifier, when they were applied to categorize Chinese texts. Overall, kNN achieved the best accuracy, about 78.3%, but required large amounts of computation time and memory when used to classify new texts. Rocchio was very time and memory efficient, and achieved a high level of accuracy, about 75.4%. In practical implementation, Rocchio may be a good choice.

參考文獻


卜小蝶,簡立峰 Lee-Feng, Lee-Feng(1996).Important Issues on Chinese Information Retrieval.中文計算語言學期刊.1(1),205-221.
Baker, Douglas,McCallum, Kachites(1998).Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR'98).
簡立峰 Lee-Feng, Lee-Feng(1999).The Fourth International Workshop on Information Retrieval with Asian Languages (IRAL'99).
Ferragina, Paolo,Grossi, Roberto(1999).The String B-tree: A New Data Structure for String Search in External Memory and its Application.Journal of ACM.46(2),236-280.
Frobelink, Marko,Mladenic, Dunja(1998).Proceedings of the 13th European Conference on Artifical Intelligence.

被引用紀錄


Chen, Y. C. (2005). 中文零代詞解析與應用 [doctoral dissertation, Tatung University]. Airiti Library. https://www.airitilibrary.com/Article/Detail?DocID=U0081-0607200917233352
曾冠燕(2008)。生物資訊文獻查詢-利用文件相似度〔碩士論文,亞洲大學〕。華藝線上圖書館。https://www.airitilibrary.com/Article/Detail?DocID=U0118-0807200916283006
劉宣榮(2010)。中華民國專利之關鍵字歷史資料查詢系統〔碩士論文,亞洲大學〕。華藝線上圖書館。https://www.airitilibrary.com/Article/Detail?DocID=U0118-1511201215465542
Budiansyah, A. (2010). Text Trend Analysis via Significant Term A Based on Indonesia News [master's thesis, Asia University]. Airiti Library. https://www.airitilibrary.com/Article/Detail?DocID=U0118-1511201215465544

延伸閱讀