Articles

From idea to execution: a laptop price prediction model

Comparing Elastic Net, SVR and Random Forest regression on 1,721 laptops; Random Forest won with an R² of 88.24%. Code, data and report on GitHub.

Published
Type
Project · 3 min read
Originally on
LinkedIn
Code
GitHub

One of my most rewarding projects during my MSc was laptop price prediction using machine learning. It combines my interest in machine learning with solving a real-world problem, and the full project is open on GitHub.

The problem

Laptops are indispensable, and their prices vary widely by configuration, brand and features. Estimating a fair price is hard for buyers. So I built a model that predicts laptop prices from key specifications such as processor, RAM and GPU.

Overview

The goal was to build regression models that predict price accurately, and compare:

  • linear regression with L1/L2 regularisation (Elastic Net)
  • Support Vector Regression (SVR)
  • Random Forest regression

Random Forest performed best, with an R² of 88.24%, because it captures the non-linear relationships in the data.

How I approached it

Data preparation

The dataset had 1,721 records and 23 attributes: numerical, categorical and binary features.

  • Encoding: target, one-hot and ordinal encoding, depending on the feature type.
  • Scaling: StandardScaler to normalise the data so the models performed well.
  • Dimensionality reduction: Principal Component Analysis (PCA) to reduce dimensions while keeping most of the variance.

Model development

  • Elastic Net combined Lasso (L1) and Ridge (L2) regularisation to balance performance and interpretability.
  • SVR captured non-linear patterns with an RBF kernel and L2 regularisation for stability.
  • Random Forest outperformed both by handling complex relationships and interactions between features.

Evaluation and insights

  • Random Forest gave the best result: R² of 88.24%.
  • The biggest drivers of price were processor type, RAM, GPU and screen quality.
  • Using the trained model, I predicted prices for new, unseen configurations.

Technologies

Python, with Pandas, NumPy, Scikit-learn, Matplotlib and Seaborn, in Jupyter Notebook and Google Colab.

What I learnt

  • Data preprocessing (encoding, scaling and dimensionality reduction) does more for accuracy than model choice alone.
  • Regularisation is how you balance interpretability against performance.
  • Why models like Random Forest suit non-linear relationships, and what you trade for it.

Explore the code

The GitHub repository includes a documented Jupyter Notebook (preprocessing, model building and evaluation), the dataset for reproducibility, and the project report with the full analysis.

Want to talk about this, or something like it for your team?

Email me