From mathematics to production code – building linear regression algorithms with NumPy, step‑by‑step.
This repository is my personal deep‑dive into the foundations of machine learning. Instead of treating algorithms as black boxes, I derive the math, implement the logic from scratch, and validate every line of code against industry‑standard libraries like Scikit‑Learn.
| Model | Implementation | Validated against sklearn |
|---|---|---|
| Simple Linear Regression | ✅ Done | ✅ Pass |
| Multiple Linear Regression | ✅ Done | ✅ Pass |
| Polynomial Regression | ⏳ Planned | ⏳ Planned |
| Ridge (L2) Regression | ⏳ Planned | ⏳ Planned |
| Lasso (L1) Regression | ⏳ Planned | ⏳ Planned |
| ElasticNet Regression | ⏳ Planned | ⏳ Planned |
Most ML courses teach you to call model.fit() and model.predict() without explaining what happens inside. This project flips that approach:
- Derive the math – OLS formulas, gradient descent, and regularization written out in detail.
- Code from scratch – pure NumPy implementations with clean OOP design.
- Prove correctness – every implementation is unit‑tested against Scikit‑Learn’s production‑grade models.
- Build for reuse – scikit‑learn compatible API so others can use your code with zero learning curve.
Clone the repository and install in development mode:
git clone https://github.com/FaraAbbasi/linear-models-from-scratch.git
cd linear-models-from-scratch
pip install -e .- Python 3.8+
- NumPy
- pandas (for the example notebook)
- scikit-learn (for validation & comparison)
You can install all dependencies with:
pip install -r requirements.txtimport numpy as np
from linear_models.simple_linear_regression import SimpleLinearRegression
X = 2 * np.random.rand(100)
y = 4 + 3 * X + np.random.randn(100) * 0.5
model = SimpleLinearRegression()
model.fit(X, y)
print(f"Coefficient: {model.m:.4f}") # ~3.0
print(f"Intercept: {model.b:.4f}") # ~4.0from linear_models.multiple_linear_regression import MultipleLinearRegression
# X is 2D: (n_samples, n_features)
X = np.random.rand(100, 3)
y = 2 + 0.5 * X[:, 0] + 1.2 * X[:, 1] - 0.8 * X[:, 2] + np.random.randn(100) * 0.1
model = MultipleLinearRegression()
model.fit(X, y)
print(f"Coefficients: {model.coef_}") # array([0.5, 1.2, -0.8])
print(f"Intercept: {model.intercept_}") # ~2.0Both models provide the same fit() and predict() methods as Scikit‑Learn.
Every model in this repository is mathematically equivalent to Scikit‑Learn’s implementation. For example, the results for Multiple Linear Regression on a test dataset:
| Metric | Scikit‑Learn | Our Model |
|---|---|---|
| R² Score | 0.9897 | 0.9897 |
| MAE | $15,787.48 | $15,787.48 |
| RMSE | $19,813.29 | $19,813.29 |
The same rigorous validation applies to all implemented models – the coefficients, intercept, and all evaluation metrics are identical to those produced by sklearn.linear_model.
linear-models-from-scratch/
│
├── src/
│ └── linear_models/
│ ├── __init__.py
│ ├── simple_linear_regression.py # ✅ Single feature OLS
│ └── multiple_linear_regression.py # ✅ Multiple features (Normal Equation)
│
├── notebooks/
│ ├── simple-linear-regression.ipynb # EDA, sklearn comparison, custom model
│ └── multiple-linear-regression.ipynb # EDA, sklearn comparison, custom model
│
├── requirements.txt
├── pyproject.toml # (soon) pip-installable package
├── LICENSE
└── README.md
Each notebook guides you through the full journey for a given model:
- Data Loading & Cleaning – handling missing values, duplicates, type conversion.
- Exploratory Visualisation – scatter plots, histograms, boxplots to confirm assumptions.
- Scikit‑Learn Baseline – training, evaluating, and plotting the "official" model.
- From‑Scratch Implementation – deriving the math and building the custom class.
- Side‑by‑Side Comparison – proving that both models produce identical results.
Check out the notebooks to see the derivation and validation in action.
- Simple Linear Regression (OLS closed‑form)
- Multiple Linear Regression (Normal Equation + Matrix approach)
- Gradient Descent (Batch, SGD, Mini‑batch)
- Polynomial Regression (Feature engineering)
- Ridge Regression (L2 regularization)
- Lasso Regression (L1 regularization)
- ElasticNet (Combined L1+L2)
- Model persistence (
.save()/.load()) - PyPI publication
Although the tests are not yet pushed, they are ready and fully pass. They cover:
- Coefficient and intercept matching with
sklearn - R² score consistency
- Edge cases (zero variance, single sample, invalid input shapes)
- Clear error messages for misuse
Once pushed, you can run them with:
pytest tests/This is a personal learning project, but feedback and suggestions are always welcome! Feel free to open an issue or a pull request.
Distributed under the MIT License. See the LICENSE file for more details.
Made with ❤️ and NumPy.