Open to → Quant Research · SWE · Machine Learning · Data Science Internships
Pairs-trading engine with Kalman-filter dynamic hedge ratios and walk-forward OOS validation over ~14 years — 1.02 out-of-sample Sharpe, survived Deflated Sharpe Ratio correction for multiple-testing bias.
Python statsmodels pytest
Retrieval-then-ranking recommender on MovieLens-25M (1.48M interactions): ALS + two-tower retrieval feeding LightGBM LambdaMART and neural rankers, served behind a FastAPI endpoint — NDCG@10 of 0.622, with two-tower lifting Recall@200 ~7% over ALS (0.455 vs. 0.425). Four-loss ablation shows listwise consistently beating pairwise.
Python PyTorch LightGBM
LoRA fine-tune of Qwen2.5-3B for a low-resource language — +8pp on Belebele (22→30%) and 4× SIB-200 macro-F1 vs. few-shot. A native-authored eval harness exposes chrF++ as a poor quality proxy (Cohen's κ = 0.000).
Python PyTorch LoRA
S/T/X/R/DR meta-learners for treatment-effect (CATE) estimation on Criteo's 14M-row benchmark, with Qini/AUUC metrics from scratch. The X-learner beats a response-model baseline by 28% on top-decile targeting, exposing outcome ROC-AUC as the wrong objective for incrementality.
Python EconML scikit-uplift
3-state Gaussian HMM (EM) driving VaR/CVaR estimation with 94.7% empirical coverage on SPY. Performance-critical path implemented in C++.
Python C++ NumPy
End-to-end origination pipeline over SEC EDGAR XBRL filings: ranks acquisition targets by embedding-based strategic adjacency and segment complementarity, runs a full DCF/WACC/CAPM stack with trading comps, precedent transactions, and EPS accretion/dilution, and ships a one-command CLI generating deal teasers with football-field charts and premium × synergy sensitivity heatmaps. 87% top-10 / 100% top-20 hit-rate vs. real deal pairings, 33 pytest cases.
Python scikit-learn pandas
Languages
Machine Learning / AI
Data Science / Statistics
Serving / MLOps / Tooling
Visualization
Frontend / Web



