Analytics
Synthetic Control: Pitfalls, Limitations, & Metrics to Know Before the Test
Part 3 of 3: how to choose the test city, donor pool, time grain and metric for Synthetic Control, with industry examples, limitations and a leadership summary.
Jargon check
- Counterfactual: What would have happened without the campaign.
- Donor pool: The set of untreated cities you allow the model to borrow from.
- Spillover: When the campaign affects cities you thought were untouched, for example shoppers ordering from a nearby city.
- Power analysis: A check done before launch that tells you whether the test can detect the effect size you expect.
- Convex hull: The range of results you can build by blending donors. A city outside it cannot be matched.
- Overfitting: Matching past noise so closely that the model predicts new weeks poorly.
- Net sales: Sales after returns and discounts.
- ROI: Return on investment: (incremental margin minus campaign cost) divided by campaign cost.
Up till now we now know how Synthetic Control works and how to validate it. This last part covers the practical choices that make or break a real test, the limits of the method, and how to present the result.
Consider these three fictional examples
Retail: StyleLane runs a billboard campaign in one metro. Two competitors ran their own sales that month. The team removes every city with a known competitor event from the donor pool before fitting. A clean donor pool matters as much as the algorithm.
Finance: BrightBank opens a branch cluster in one city. The team wants to know if a regional ad push lifted new savings accounts. They build a synthetic city from eight cities with no new branches and no ads. If the gap appears only after launch and no placebo city shows a similar jump, the push worked.
Healthcare: MediQuick launches tele-consult in one state. Bookings rise, but flu season started too. A synthetic state built from similar states carries the seasonal pattern, so the feature is not credited with the season.
How to choose cities
The test city is often a business decision. Avoid the extreme city (biggest or fastest-growing), and run a power analysis first to confirm the test can detect the effect you expect.
The donor pool:
- No exposure to the same campaign, other big campaigns or competitor events.
- No spillover. If Pune is treated, exclude neighbours such as Pimpri-Chinchwad.
- Similar structure: country, currency, city tier, category mix.
- Complete, stable data with no missing weeks or tracking changes.
- Correlated pre-trends. It is suggested to keep donors whose pre-period correlation with the test city is above about 0.5 to 0.7, and the top 10 to 30.
- Not too many donors. Most get a weight near zero, and typically 3 to 6 cities form the synthetic city.
Time grain and metric
- Weekly is the standard. Weekly data is preferable as it smooths weekday noise and payday effects. A typical setup is 26 to 104 pre-weeks and 4 to 12 post-weeks.
- Daily data is too short and have lot of noise.
- Monthly data is avoided as it leaves with too few data points.
##Match the metric to the goal:
| Campaign | Primary metric | Secondary metrics |
|---|---|---|
| Promotion or discount | Net sales | Units, margin |
| Free delivery | Orders | Average order value, new customers |
| Brand TV or outdoor | Sales, app installs | Site sessions, branded search |
| Loyalty launch | Repeat purchase rate | Sales per customer |
In order to be right, use net sales (after returns and discounts), additionally, check margin too, because a promotion can raise sales and still lose money. For cities of very different sizes, use sales per store or indexed sales.
Limitations
- Convex hull problem. Weights are positive and sum to 1, so the synthetic city can only sit between donors. If the test city is bigger than all of them, it cannot be matched. This is the main reason why we do SDID. This is covered in next series.
- Short pre-period. In case of short pre period, let's say about 20 to 30 weeks, there is risk of fitting noise.
- Contaminated donors. A donor with its own campaign, or spillover from the test city, biases the answer.
- Structural breaks. A shifting festival or a store opening mid-test breaks the pre-launch relationship.
- Weak inference with few donors. With 10 donors, the best p-value is about 0.09.
- No built-in confidence interval. Uncertainty comes from placebos, Bayesian or conformal methods.
- Overfitting. Many donors and few weeks can give a perfect fit that does not generalise. Always hold out some weeks.
- Small effects. A 1% lift in a noisy city is often undetectable. In order to be sure do a power analysis first.
Test Summary : How to show the result to leadership
| Question | Answer |
|---|---|
| What did we test? | Free delivery in Pune, 3 weeks |
| How did we measure? | A "virtual Pune" (50% Nagpur, 30% Indore, 20% Surat) that matched real Pune within about 1.2% before launch |
| What was the impact? | +Rs 32 lakh incremental sales, +8.8% |
| Why not before vs. after? | It shows +22%, but roughly 13 points were market growth every city saw |
| How confident are we? | Post-launch gap is 8.5 times normal noise, and Pune ranks first in the placebo test |
| Why not one control city? | About 40% lower error before launch (1.26 vs. 2.14) |
| Next step | Compare the Rs 32 lakh (and its margin) with the cost of free delivery |
The business translation: ROI = (incremental margin minus campaign cost) divided by campaign cost.
Key takeaways
- Clean donors beat clever algorithms. Remove contaminated and neighbouring cities.
- Use weekly data, 26 or more pre-weeks, and a metric that matches the campaign goal.
- Know the limits: Oversized test cities, short histories and few donors usually leads to failures.
- Always translate lift into incremental margin and ROI.
Full code: All the code behind this series is available on GitHub.
What is the hardest part of your data to keep clean: donors, metrics or history length? Share your experience in the comments.