Evaluating machine learning models for credit risk prediction across retail segments of South Africa
Loading...
Date
Authors
Researcher ID
Supervisors
Journal Title
Journal ISSN
Volume Title
Publisher
North-West University (South Africa).
Record Identifier
Abstract
The current study evaluated the relevance of product segmentation in credit application scorecards within the retail credit industry of South Africa. It specifically tested whether developing distinct models for different product populations yields a better predictive
performance on a single model trained on a combined dataset. A quantitative experimental design was employed utilising the AutoGluon automated machine learning (AutoML) framework to train and evaluate competing models (including neural networks and gradient boosting ensembles). The study compared the Area Under the Curve (AUC) performance of the models trained on the segmented product data against the single generalised model. The results indicated that segmenting by credit product types, using the same target definition, over the same cross sectional time frame did not improve the overall model performance. Contrary to industry norms, the single generalised model achieved a higher AUC to that of the segmented models across all three product categories. The study concludes that training models on combined datasets resulted in superior risk differentiation compared to segmented datasets. This suggests that modern AutoML frameworks leverages increased data volume more effectively than traditional segmentation in the retail credit environment.
Sustainable Development Goals
Industry, Innovation and Infrastructure
Description
Thesis, Master of Business Administration -- North-West University, Potchefstroom Campus
