Tree-Based Missing Value Imputation Using Feature Selection

Researchers and practitioners of many areas of knowledge frequently struggle with missing data. Missing data is a problem because almost all standard statistical methods assume that the information is complete. Consequently, missing value imputation offers a solution to this problem. The main contribution of this paper lies on the development of a random forest-based imputation method (TI-FS) that can handle any type of data, including high-dimensional data with nonlinear complex interactions. The premise behind the proposed scheme is that a variable can be imputed considering only those variables that are related to it using feature selection. This work compares the performance of the proposed scheme with other two imputation methods commonly used in literature: KNN and missForest. The results suggest that the proposed method can be useful in complex scenarios with categorical variables and a high volume of missing values, while reducing the amount of variables used and their corresponding preliminary imputations.

關鍵字

K nearest neighbor ； missForest ； missing data ； random forest

國際替代計量

全文下載

主題瀏覽

Tree-Based Missing Value Imputation Using Feature Selection

摘要

關鍵字

延伸閱讀

國際替代計量

本網站使用Cookies