International Journal of Mathematics and Statistics Studies (IJMSS)

Enhancing Predictive Accuracy in Incomplete Datasets: A Comparative Evaluation of Missing-Data Imputation Methods

Abstract

Missing data remains one of the most persistent challenges in statistical modelling and machine learning, often degrading predictive accuracy and leading to biased inference when handled inappropriately. Although numerous imputation techniques have been developed, their comparative effectiveness under different missingness mechanisms remains an active area of research. This study comparatively evaluates four widely used imputation techniques: Mean Imputation, k-nearest neighbours (KNN), Multiple Imputation by Chained Equations (MICE), and Autoencoder-based Imputation, with respect to their ability to preserve predictive performance across multiple benchmark datasets. Seventeen datasets representing diverse application domains were employed. Artificial missingness was introduced under three missing-data mechanisms: Missing Completely at Random (MCAR), Missing at Random (MAR), and Missing Not at Random (MNAR). Following imputation, Random Forest models were fitted to the reconstructed datasets, and predictive performance was evaluated using the Root Mean Squared Error (RMSE). The original complete datasets served as the benchmark against which the performance of each imputation method was assessed. Comparative evaluation was conducted using absolute deviation from the benchmark RMSE, paired hypothesis tests, Bonferroni-adjusted multiple comparisons, and non-parametric tests where appropriate. The results demonstrate that Multiple Imputation consistently produced RMSE values closest to those obtained from the original complete datasets under all three missingness mechanisms. KNN imputation generally ranked second, whereas Mean Imputation and Autoencoder-based Imputation exhibited larger deviations from the benchmark, particularly under more challenging missingness conditions. Statistical analyses further showed that Multiple Imputation most consistently preserved the predictive characteristics of the original datasets, while the remaining methods displayed varying levels of performance depending on the missingness mechanism and dataset characteristics. The findings highlight the importance of selecting imputation methods according to both the missing-data mechanism and the intended predictive modelling objective. Overall, Multiple Imputation provides the most reliable and robust performance across heterogeneous datasets, making it an appropriate general-purpose imputation strategy for predictive analytics involving incomplete data.

Keywords: K-nearest neighbours, Random Forest, autoencoder, missing data, multiple imputation, predictive modelling, root mean squared error

cc logo

This work by European American Journals is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 4.0 Unported License

 

Recent Publications

Email ID: editor.ijmss@ea-journals.org
Impact Factor: 7.80
Print ISSN: 2053-2229
Online ISSN: 2053-2210
DOI: https://doi.org/10.37745/ijmss.13

Author Guidelines
Submit Papers
Review Status

 

Scroll to Top

Don't miss any Call For Paper update from EA Journals

Fill up the form below and get notified everytime we call for new submissions for our journals.