Developing a risk profile for at risk student(s) at a university in South Africa
Loading...
Files
Date
Authors
Researcher ID
Supervisors
Journal Title
Journal ISSN
Volume Title
Publisher
North-West University (South Africa)
Record Identifier
Abstract
Higher education authorities continue to be concerned about dropout rates among university students. Dropping out affects cost efficiency and tarnishes the reputation of the institution. It is therefore crucial to identify at risk students of dropout.
This study aims to develop a risk profile for at risk student(s) at a university in South Africa. The risk profile will allow the university to identify at risk student(s) of dropout and put measures that could prevent students from dropping out. The successful intervention could increase the retention and graduation rates, while minimising the dropout rate. The research questions to achieve the primary objective of the study were:
- How can a profile of at risk students of dropout from the university be developed?
- What are the features that can assist to identify at risk students?
- Is it possible to predict at risk students from administrative data using applied statistical learning (machine learning) techniques?
- What recommendations can be made to the university to address at risk students intrying to lower dropout rate?
The research was conducted by means of a literature study and secondary data analysis. The literature study reviewed feature selection methods, machine learning algorithms for classification problems. The article format was adopted for this study which culminated in three research articles covering the objectives of the study. The articles took the formatting required by the journals where they were submitted.
The first article focused on identifying the features that can assist to identify at risk students of dropout using administrative university data. Feature selection has many potential benefits such as facilitating data visualisation, data understanding and improving prediction performance of identifying at risk students of dropout. Machine learning
methods were used to pre-process data and select relevant features that contribute to student dropout. In particular, the methods of interest to this work include weight of evidence (WOE) and information value (IV), Sequential Feature Selection Method (SFE). The findings of the article highlighted that the main features to student dropout were participation average marks, number of modules registered, and number of modules failed.
The second article dealt with evaluation of machine learning algorithms in predicting at risk students of dropout from university using administrative data. Successful prediction of at risk students of dropout can assist higher education institutions in implementing preventative measures to retain students. As a result, the retention rate will rise, and the image of the higher education institution will be preserved.
The study used machine learning techniques on imbalanced administrative data to predict at risk students of dropout using suitable selected variables obtained in article one. The study focused on four supervised machine learning algorithms: the random forest, decision tree, kNN, and logistic regression. The random forest algorithm was found to be the best at predicting the students at risk of dropout after the imbalanced dataset was resampled using oversampling and synthetic minority over-sampling technique (SMOTE) methods.
The third article aimed at building a profile of at risk students of dropout using administrative university data. The risk profile may assist the university in identifying students who are at risk of dropout and give necessary support to prevent them from dropping out. The researcher employed weight of evidence (WOE) and information value (IV) to build risk profile of students at risk of dropout. The predictors were chosen based on their weight of evidence (WOE) and information value (IV) and were then used to build a profile of at risk students. The study found that a student is at risk of dropout has the
following characteristics: student was born between year (1931,1967] and (1994, 2001]; failed more than four modules in an academic year; obtained participation average mark of 43 percent or less on modules registered in an academic year; and entered university at second entry level.
This study recommends first that the risk profile be used to identify at risk students of dropout. Secondly, the risk profile of at risk students of dropout be developed in each faculty/school. Thirdly, appropriate preventative measures and suitable intervention programs be designed to mitigate the risk faced by students identified at risk of dropout.
Keywords: machine learning, imbalanced data, student dropout, feature selection, algorithms, resampling, weight of evidence, information value.
Sustainable Development Goals
Description
PhD (Operational Research), North-West University, Vanderbijlpark Campus
