Analytics··11 min read
Synthetic Control: Pitfalls, Limitations, & Metrics to Know Before the Test
Part 3 of 3: how to choose the test city, donor pool, time grain and metric for Synthetic Control, with industry examples, limitations and a leadership summary.
#Experimentation#synthetic control#geo testing#experiment design#marketing analytics
Analytics··12 min read
Synthetic Control: Result Validation with RMSPE and Placebos
Part 2 of 3: validate a Synthetic Control result with RMSPE, the RMSPE ratio, placebo tests, plus the Python code that fits the weights.
#Experimentation#synthetic control#placebo test#RMSPE#python#causal inference
Analytics··12 min read
Synthetic Control: The Core Idea with a Easy to understand Example
Part 1 of 3: what Synthetic Control is, why it beats before-vs-after and single control-city tests, with a hand-checkable free-delivery example.
#Experimentation#synthetic control#geo testing#counterfactual#marketing analytics
What's New··3 min read
AI Referred Shoppers Are Up 130%: Does Your Analytics Tool Track Them?
Adobe expects traffic from AI chat tools and AI browsers to US retail sites to grow 130% this holiday season. Here's how to make AI referrals visible in GA4 and Microsoft Clarity before the peak.
#Web Analytics#GA4#Microsoft Clarity#Ecommerce#Attribution#Web & Customer Analytics
Data Science··7 min read
Opening the black box: explaining the model with SHAP
What SHAP actually measures, a real customer's score broken down feature by feature, and three different views of what matters most.
#SHAP#explainable AI#model interpretability#machine learning#Model Evaluation
Data Science··6 min read
The Brier score: one number for whether your probabilities can be trusted
A worked example of the Brier score, why it punishes confident wrong answers, and how a badly calibrated model can lose to a trivial guess.
#Brier score#model evaluation#calibration#probability#machine learning#Model Evaluation
Data Science··6 min read
Calibration: turning a ranking into a probability you can actually trust
Why class-weighted scores lie about probability, and how Platt scaling and isotonic regression fix it.
#calibration#Platt scaling#isotonic regression#probability#machine learning#Model Evaluation
Projects··10 min read
Beyond Propensity Scores: uplift, production, and keeping a model honest over time
Part 6 of 6. Propensity versus uplift on a simulated randomised campaign, a production architecture, and how to monitor a model for drift.
#uplift modeling#model monitoring#MLOps#drift detection#purchase propensity
Projects··10 min read
Model Testing and Targeting: Lift, Calibration, and the Campaign Audience
Part 5 of 6. Judging the model on the untouched test set, then turning its scores into a real campaign audience with business rules and a holdout.
#purchase propensity#lift and gain#decile analysis#campaign targeting#marketing analytics
Projects··10 min read
Model Training: Logistic Regression, Random Forest, and LightGBM
Part 4 of 6. Splitting the data by time, training three models of increasing complexity, and reading the first comparison table.
#purchase propensity#machine learning#LightGBM#random forest#model training
Projects··10 min read
Data Preparation: Cleaning, Sessionizing, and Feature Engineering
Part 3 of 6. How to clean web clickstream data, build sessions, create the target and engineer features that carry purchase intent.
#purchase propensity#marketing analytics#machine learning#e-commerce
Projects··10 min read
Model Designing: Observation Window,Populations and Data Leakage
Part 2 of 6 series to explain the problem statement and concept behind Purchase Propensity Model.
#purchase propensity#marketing analytics#machine learning#e-commerce
Projects··10 min read
Purchase Propensity Modeling: A Beginner’s Guide to Predicting Customer Purchase Probabilities
Part 1 of 6 series to explain the problem statement and concept behind Purchase Propensity Model.
#purchase propensity#marketing analytics#machine learning#e-commerce
Running··5 min read
What Running Has Taught Me About Consistency
A personal reflection on repetition, patience, feedback loops, and long-term progress.
#Running#Consistency#Performance
Data Science··6 min read
Context Windows Demystified: How Much Can Your AI Actually Remember?
The finale of our LLM tokenization series: what context windows are, how they relate to token limits, and how real AI products manage them.
#Tokenisation#NLP#Text Processing
Analytics··9 min read
From Data to Decisions: Building Better Analytics
A framework for connecting measurement to decisions that people can actually act on.
#Analytics#Decision Making#Data
Data Science··6 min read
Token Limits: Why Your AI Chatbot Suddenly Forgets Everything
A deep dive into how token limits affect the performance of large language models and what you can do about it.
#Tokenisation#NLP#Text Processing
Data Science··8 min read
Inside the Tokeniser: BPE, WordPiece, SentencePiece & Unigram Explained
An example-driven walkthrough of the four major tokenisation algorithms behind modern LLMs, with a step-by-step explanation.
#Tokenisation#NLP#Text Processing
Data Science··8 min read
Tokens 101: The Secret Language Your AI Actually Speaks
A beginner-friendly introduction to what tokens are, how they differ from words and characters, and why they matter when using LLMs like ChatGPT and Claude.
#Tokenisation#NLP#Text Processing
Data Science··12 min read
Clustering in Machine Learning: Finding Hidden Groups in Your Data
A simple introduction to clustering. What it is, how it differs from classification, and real retail examples.
#Clustering#Machine Learning#Unsupervised Learning
Analytics··10 min read
What is Open Graph and How Does It Work?
A simple introduction to Open Graph protocol and its role in social media sharing.
#Open Graph#Social Media#SEO
Analytics··8 min read
A Practical Introduction to Web Analytics
How to move from reports and dashboards toward questions, decisions, and measurable outcomes.
#Web Analytics#KPIs#Measurement
Analytics··7 min read
Understanding GA4 Event Tracking
A practical mental model for designing useful GA4 events instead of collecting noise.
#GA4#Events#Measurement Strategy