Command Palette

Search for a command to run...

Command Palette

Search for a command to run...

Projects

E-commerce Repeat Buyer Prediction
Github
Website
Post

E-commerce Repeat Buyer Prediction

Live
RoleData Engineering Lead / ML Engineer
TeamTeam of 4
TimelineMar-Apr 2026

Overview

This project predicts repeat buyers in e-commerce using Spark MLlib on 20,692,840 clickstream events from the REES46 dataset. It combines big data processing, model selection, leakage-safe temporal splitting, RFM segmentation, and campaign targeting analysis.

The workflow benchmarks five ML model families, identifies the best performer (Gradient-Boosted Trees at AUC 0.9177), and translates model rankings into business-actionable campaign targeting strategy.

Architecture

Rendering diagram...
pinch · drag

Challenges & Solutions

1

20.7M clickstream events cannot be processed efficiently with local-only pandas workflows.

Used PySpark and Databricks for distributed processing, enabling the full pipeline to run on the complete dataset without sampling.

2

Repeat-buyer prediction can easily leak future purchase behavior into features.

Enforced a strict temporal split - October–December activity as features, January–February purchases as labels - to prevent data leakage.

Results

AUC-ROC

0.9177

Gradient-Boosted Trees

F1 Score

0.8155

Precision 0.7988, Recall 0.8329

Campaign Lift

1.96x

vs. random baseline targeting

Dataset Scale

20.7M

Clickstream events processed

Features

  • PySpark ETL pipeline on 20,692,840 clickstream events from REES46 dataset
  • Databricks distributed compute workflow
  • Leakage-safe temporal split: Oct–Dec features, Jan–Feb labels
  • Spark MLlib model comparison: Logistic Regression, Decision Tree, Random Forest, Weighted Ensemble, GBT
  • RFM customer segmentation (Recency, Frequency, Monetary)
  • Feature importance analysis with top-15 GBT features
  • Campaign targeting interpretation: ~5,426 targeted users, ~4,334 expected returners