
E-commerce Repeat Buyer Prediction
LiveOverview
This project predicts repeat buyers in e-commerce using Spark MLlib on 20,692,840 clickstream events from the REES46 dataset. It combines big data processing, model selection, leakage-safe temporal splitting, RFM segmentation, and campaign targeting analysis.
The workflow benchmarks five ML model families, identifies the best performer (Gradient-Boosted Trees at AUC 0.9177), and translates model rankings into business-actionable campaign targeting strategy.
Architecture
Challenges & Solutions
20.7M clickstream events cannot be processed efficiently with local-only pandas workflows.
Used PySpark and Databricks for distributed processing, enabling the full pipeline to run on the complete dataset without sampling.
Repeat-buyer prediction can easily leak future purchase behavior into features.
Enforced a strict temporal split - October–December activity as features, January–February purchases as labels - to prevent data leakage.
Results
AUC-ROC
0.9177
Gradient-Boosted Trees
F1 Score
0.8155
Precision 0.7988, Recall 0.8329
Campaign Lift
1.96x
vs. random baseline targeting
Dataset Scale
20.7M
Clickstream events processed
Features
- PySpark ETL pipeline on 20,692,840 clickstream events from REES46 dataset
- Databricks distributed compute workflow
- Leakage-safe temporal split: Oct–Dec features, Jan–Feb labels
- Spark MLlib model comparison: Logistic Regression, Decision Tree, Random Forest, Weighted Ensemble, GBT
- RFM customer segmentation (Recency, Frequency, Monetary)
- Feature importance analysis with top-15 GBT features
- Campaign targeting interpretation: ~5,426 targeted users, ~4,334 expected returners