From idea to execution: a laptop price prediction model
Comparing Elastic Net, SVR and Random Forest regression on 1,721 laptops; Random Forest won with an R² of 88.24%. Code, data and report on GitHub.
One of my most rewarding projects during my MSc was laptop price prediction using machine learning. It combines my interest in machine learning with solving a real-world problem, and the full project is open on GitHub.
The problem
Laptops are indispensable, and their prices vary widely by configuration, brand and features. Estimating a fair price is hard for buyers. So I built a model that predicts laptop prices from key specifications such as processor, RAM and GPU.
Overview
The goal was to build regression models that predict price accurately, and compare:
- linear regression with L1/L2 regularisation (Elastic Net)
- Support Vector Regression (SVR)
- Random Forest regression
Random Forest performed best, with an R² of 88.24%, because it captures the non-linear relationships in the data.
How I approached it
Data preparation
The dataset had 1,721 records and 23 attributes: numerical, categorical and binary features.
- Encoding: target, one-hot and ordinal encoding, depending on the feature type.
- Scaling:
StandardScalerto normalise the data so the models performed well. - Dimensionality reduction: Principal Component Analysis (PCA) to reduce dimensions while keeping most of the variance.
Model development
- Elastic Net combined Lasso (L1) and Ridge (L2) regularisation to balance performance and interpretability.
- SVR captured non-linear patterns with an RBF kernel and L2 regularisation for stability.
- Random Forest outperformed both by handling complex relationships and interactions between features.
Evaluation and insights
- Random Forest gave the best result: R² of 88.24%.
- The biggest drivers of price were processor type, RAM, GPU and screen quality.
- Using the trained model, I predicted prices for new, unseen configurations.
Technologies
Python, with Pandas, NumPy, Scikit-learn, Matplotlib and Seaborn, in Jupyter Notebook and Google Colab.
What I learnt
- Data preprocessing (encoding, scaling and dimensionality reduction) does more for accuracy than model choice alone.
- Regularisation is how you balance interpretability against performance.
- Why models like Random Forest suit non-linear relationships, and what you trade for it.
Explore the code
The GitHub repository includes a documented Jupyter Notebook (preprocessing, model building and evaluation), the dataset for reproducibility, and the project report with the full analysis.
Want to talk about this, or something like it for your team?
Email me