September 2026

Conference Paper

A Comparative Analysis of Machine Learning Models and Resulting Gridded Population Distribution Datasets

By:
Tsybina, Evgeniya ; Lebakula, Viswadeep ; Stipek, Clinton W; Nukavarapu, Nivedita ; Gonzales, Jack J; Urban, Marie L
Page Number:
84-91
Book Title:
2026 11th International Conference on Machine Learning Technologies (ICMLT)
Publication Date:
September 25, 2026
Publisher Location:
IEEE, New Jersey, United States of America
Conference Name:
International Conference on Machine Learning Technologies (ICMLT)
Conference Location:
Berlin, Germany
Conference Sponsor:
IEEE
View DOI Listing:
https://doi.org/10.1109/ICMLT69916.2026.11689306

Abstract

Human populations are projected to reach 10 billion by 2050, underscoring the growing importance of accurate population modeling. As modeling complexity increases, it necessitates the integration of a variety of data sources and methods. The resulting gridded population datasets show different levels of performance for different regions, possibly due to population characteristics of those regions or due to varying underlying methodologies. These datasets are developed for different purposes, complicating their direct comparison. However, we can understand error better if we test different methodologies while keeping original data and research objectives constant. We compare three commonly used machine learning methods - decision trees, random forests and XGBoost, across all countries in Africa. All three achieve high accuracy (R2=0.90−0.98). The approaches show midline convergence, performing better in moderately populated areas and showing diminished accuracy in sparsely populated and densely populated areas. No approach is overall best. Random forests tend to perform better in low density population areas and worse in very densely populated areas. XGBoost demonstrates the best edge performance, while decision trees show average performance in all types of areas.