- By:
- Tsybina, Evgeniya ; Lebakula, Viswadeep ; Stipek, Clinton W; Nukavarapu, Nivedita ; Gonzales, Jack J; Urban, Marie L
- Page Number:
- 84-91
- Book Title:
- 2026 11th International Conference on Machine Learning Technologies (ICMLT)
- Publication Date:
- September 25, 2026
- Publisher Location:
- IEEE, New Jersey, United States of America
- Conference Name:
- International Conference on Machine Learning Technologies (ICMLT)
- Conference Location:
- Berlin, Germany
- Conference Sponsor:
- IEEE
- View DOI Listing:
- https://doi.org/10.1109/ICMLT69916.2026.11689306
Abstract
Human populations are projected to reach 10 billion by 2050, underscoring the growing importance of accurate population modeling. As modeling complexity increases, it necessitates the integration of a variety of data sources and methods. The resulting gridded population datasets show different levels of performance for different regions, possibly due to population characteristics of those regions or due to varying underlying methodologies. These datasets are developed for different purposes, complicating their direct comparison. However, we can understand error better if we test different methodologies while keeping original data and research objectives constant. We compare three commonly used machine learning methods - decision trees, random forests and XGBoost, across all countries in Africa. All three achieve high accuracy (R2=0.90−0.98). The approaches show midline convergence, performing better in moderately populated areas and showing diminished accuracy in sparsely populated and densely populated areas. No approach is overall best. Random forests tend to perform better in low density population areas and worse in very densely populated areas. XGBoost demonstrates the best edge performance, while decision trees show average performance in all types of areas.