All projects
Study Project2026-02-14

Data Science with R: Student Performance Prediction

Student-habit analysis and regression benchmarking; Linear Regression led exam-score prediction with an R² of 0.897.

PythonRJupyterRegressionMachine Learning
Data Science with R: Student Performance Prediction

Overview

DSR301m examines the relationships between student habits, mental health, and academic performance. The team developed analysis and prediction workflows for three targets: exam score, mental-health rating, and daily study hours.

Problem

The student dataset contains a range of academic and lifestyle factors, but their contribution to each outcome was unclear. The project required data cleaning, feature preparation, and consistent model comparisons on held-out data to identify a suitable approach for each target.

Role

I was responsible for the exam-score analysis and prediction work, as well as consolidating the team's code and results in the shared repository. My scope covered data preparation, model training, evaluation, and presentation of the model comparison.

Solution

  • Explored and preprocessed data describing student behaviour, study patterns, and outcomes.
  • Created train/test splits so that models were evaluated under the same conditions.
  • Compared Linear Regression, Decision Tree, and Random Forest for exam-score prediction.
  • Extended the experiments to mental health and study hours using regression and ensemble models.
  • Preserved notebooks, RDS models, and result images so the analyses could be reviewed.

Technical decisions

Linear Regression served as an interpretable baseline before comparison with Decision Tree, Random Forest, and additional models. MSE quantified prediction error, while R² measured the proportion of variance explained. A shared data split and evaluation criteria made the model results directly comparable.

Results

On the exam-score test set, Linear Regression achieved an MSE of 26.50 and an R² of 0.897, outperforming Decision Tree and Random Forest in the recorded experiment. For daily study-hour prediction, Linear Regression also ranked first among the four compared models, with an MSE of 0.244 and an R² of 0.876.

Lessons learned

  • A more complex model does not automatically outperform a well-matched linear baseline.
  • A consistent evaluation workflow matters more than an isolated model score.
  • Relative data paths and a clear repository structure are essential for reproducibility.

Keep exploring

Explainable Google Stock Forecasting