Skip to content

Latest commit

 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Linear Models from Scratch

From mathematics to production code – building linear regression algorithms with NumPy, step‑by‑step.

This repository is my personal deep‑dive into the foundations of machine learning. Instead of treating algorithms as black boxes, I derive the math, implement the logic from scratch, and validate every line of code against industry‑standard libraries like Scikit‑Learn.


✅ Current Status

Model Implementation Validated against sklearn
Simple Linear Regression ✅ Done ✅ Pass
Multiple Linear Regression ✅ Done ✅ Pass
Polynomial Regression ⏳ Planned ⏳ Planned
Ridge (L2) Regression ⏳ Planned ⏳ Planned
Lasso (L1) Regression ⏳ Planned ⏳ Planned
ElasticNet Regression ⏳ Planned ⏳ Planned

🧠 Why This Project?

Most ML courses teach you to call model.fit() and model.predict() without explaining what happens inside. This project flips that approach:

  • Derive the math – OLS formulas, gradient descent, and regularization written out in detail.
  • Code from scratch – pure NumPy implementations with clean OOP design.
  • Prove correctness – every implementation is unit‑tested against Scikit‑Learn’s production‑grade models.
  • Build for reuse – scikit‑learn compatible API so others can use your code with zero learning curve.

📦 Installation

Clone the repository and install in development mode:

git clone https://github.com/FaraAbbasi/linear-models-from-scratch.git
cd linear-models-from-scratch
pip install -e .

Dependencies

  • Python 3.8+
  • NumPy
  • pandas (for the example notebook)
  • scikit-learn (for validation & comparison)

You can install all dependencies with:

pip install -r requirements.txt

🚀 Quick Start

Simple Linear Regression (1 feature)

import numpy as np
from linear_models.simple_linear_regression import SimpleLinearRegression

X = 2 * np.random.rand(100)
y = 4 + 3 * X + np.random.randn(100) * 0.5

model = SimpleLinearRegression()
model.fit(X, y)

print(f"Coefficient: {model.m:.4f}")  # ~3.0
print(f"Intercept:   {model.b:.4f}")  # ~4.0

Multiple Linear Regression (multiple features)

from linear_models.multiple_linear_regression import MultipleLinearRegression

# X is 2D: (n_samples, n_features)
X = np.random.rand(100, 3)
y = 2 + 0.5 * X[:, 0] + 1.2 * X[:, 1] - 0.8 * X[:, 2] + np.random.randn(100) * 0.1

model = MultipleLinearRegression()
model.fit(X, y)

print(f"Coefficients: {model.coef_}")   # array([0.5, 1.2, -0.8])
print(f"Intercept:    {model.intercept_}")  # ~2.0

Both models provide the same fit() and predict() methods as Scikit‑Learn.


📊 Validation & Comparison

Every model in this repository is mathematically equivalent to Scikit‑Learn’s implementation. For example, the results for Multiple Linear Regression on a test dataset:

Metric Scikit‑Learn Our Model
R² Score 0.9897 0.9897
MAE $15,787.48 $15,787.48
RMSE $19,813.29 $19,813.29

The same rigorous validation applies to all implemented models – the coefficients, intercept, and all evaluation metrics are identical to those produced by sklearn.linear_model.


📂 Repository Structure

linear-models-from-scratch/
│
├── src/
│   └── linear_models/
│       ├── __init__.py
│       ├── simple_linear_regression.py    # ✅ Single feature OLS
│       └── multiple_linear_regression.py  # ✅ Multiple features (Normal Equation)
│
├── notebooks/
│   ├── simple-linear-regression.ipynb     # EDA, sklearn comparison, custom model
│   └── multiple-linear-regression.ipynb   # EDA, sklearn comparison, custom model
│
├── requirements.txt
├── pyproject.toml                         # (soon) pip-installable package
├── LICENSE
└── README.md

📓 Notebook Walkthroughs

Each notebook guides you through the full journey for a given model:

  1. Data Loading & Cleaning – handling missing values, duplicates, type conversion.
  2. Exploratory Visualisation – scatter plots, histograms, boxplots to confirm assumptions.
  3. Scikit‑Learn Baseline – training, evaluating, and plotting the "official" model.
  4. From‑Scratch Implementation – deriving the math and building the custom class.
  5. Side‑by‑Side Comparison – proving that both models produce identical results.

Check out the notebooks to see the derivation and validation in action.


🛤️ Roadmap

  • Simple Linear Regression (OLS closed‑form)
  • Multiple Linear Regression (Normal Equation + Matrix approach)
  • Gradient Descent (Batch, SGD, Mini‑batch)
  • Polynomial Regression (Feature engineering)
  • Ridge Regression (L2 regularization)
  • Lasso Regression (L1 regularization)
  • ElasticNet (Combined L1+L2)
  • Model persistence (.save() / .load())
  • PyPI publication

🧪 Testing (Coming Soon)

Although the tests are not yet pushed, they are ready and fully pass. They cover:

  • Coefficient and intercept matching with sklearn
  • R² score consistency
  • Edge cases (zero variance, single sample, invalid input shapes)
  • Clear error messages for misuse

Once pushed, you can run them with:

pytest tests/

🤝 Contributing

This is a personal learning project, but feedback and suggestions are always welcome! Feel free to open an issue or a pull request.


📄 License

Distributed under the MIT License. See the LICENSE file for more details.


Made with ❤️ and NumPy.

About

Implementation of Linear, Ridge, and Lasso regression from scratch using NumPy, validated against Scikit-Learn.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages