Search for collections on FTS Digilib

Effects of Data Resampling on Predicting Customer Churn via a Comparative Tree-based Random Forest and XGBoost

Ako, Rita Erhovwo and Aghware, Fidelis Obukohwo and Okpor, Margaret Dumebi and Akazue, Maureen Ifeanyi and Yoro, Rume Elizabeth and Ojugo, Arnold Adimabua and Setiadi, De Rosal Ignatius Moses and Odiakaose, Chris Chukwufunaya and Abere, Reuben Akporube and Emordi, Frances Uche and Geteloma, Victor Ochuko and Ejeh, Patrick Ogholuwarami (2024) Effects of Data Resampling on Predicting Customer Churn via a Comparative Tree-based Random Forest and XGBoost. Journal of Computing Theories and Applications, 2 (1). pp. 86-101. ISSN 3024-9104

[thumbnail of 10562-Article Text-34506-1-10-20240627.pdf]
Preview
Text
10562-Article Text-34506-1-10-20240627.pdf - Published Version

Download (446kB) | Preview

Abstract

Customer attrition has become the focus of many businesses today – since the online market space has continued to proffer customers, various choices and alternatives to goods, services, and products for their monies. Businesses must seek to improve value, meet customers' teething demands/needs, enhance their strategies toward customer retention, and better monetize. The study compares the effects of data resampling schemes on predicting customer churn for both Random Forest (RF) and XGBoost ensembles. Data resampling schemes used include: (a) default mode, (b) random-under-sampling RUS, (c) synthetic minority oversampling technique (SMOTE), and (d) SMOTE-edited nearest neighbor (SMOTEEN). Both tree-based ensembles were constructed and trained to assess how well they performed with the chi-square feature selection mode. The result shows that RF achieved F1 0.9898, Accuracy 0.9973, Precision 0.9457, and Recall 0.9698 for the default, RUS, SMOTE, and SMOTEEN resampling, respectively. Xgboost outperformed Random Forest with F1 0.9945, Accuracy 0.9984, Precision 0.9616, and Recall 0.9890 for the default, RUS, SMOTE, and SMOTEEN, respectively. Studies support that the use of SMOTEEN resampling outperforms other schemes; while, it attributed XGBoost enhanced performance to hyper-parameter tuning of its decision trees. Retention strategies of recency-frequency-monetization were used and have been found to curb churn and improve monetization policies that will place business managers ahead of the curve of churning by customers.

Item Type: Article
Subjects: Q Science > QA Mathematics > QA75 Electronic computers. Computer science
Depositing User: Unnamed user with email cute.moses89@gmail.com
Date Deposited: 17 Nov 2024 16:59
Last Modified: 17 Nov 2024 16:59
URI: https://dl.futuretechsci.org/id/eprint/16

Actions (login required)

View Item
View Item