STUDY 06 · SYNTHETIC DATA
Can You Tell Which Customer Is Fake?
A probabilistic model trained on real online shopping behaviour, then I generated an artificial customer population, and then asked a machine-learning model to tell the real shoppers from the synthetic ones.
THE QUESTION
What makes synthetic data realistic?
A synthetic dataset can reproduce simple statistics such as purchase rate or visitor type while still failing to reproduce the relationships that exist between customer behaviours.
This experiment tests both. We compare the distributions directly and train a classifier whose only job is to identify whether a shopping session came from the real or synthetic population.
EXPERIMENT
Real data in. Artificial customers out.
Start with observed online shopping sessions.
Estimate conditional behavioural relationships.
Create new artificial shopping sessions.
Ask ML to identify which sessions are synthetic.
THE DETECTOR
Can the machine spot the fake?
Loading experiment results...
REAL VS SYNTHETIC
Do the populations look alike?
DISTRIBUTION TEST
Similarity by variable
100% means the synthetic distribution closely matches the observed distribution.
WHY THIS MATTERS
Matching averages is not enough.
Two datasets can have almost identical averages while containing very different relationships between variables. Synthetic-data quality therefore has to be tested at more than one statistical level.
METHODOLOGY
How the experiment works
Source
Online Shoppers Purchasing Intention dataset containing observed web sessions.
Generation
Conditional probability distributions are estimated from behavioural relationships in the original data.
Validation
Real and synthetic distributions are compared using total variation distance.
ML Challenge
A random forest receives mixed real and synthetic observations and attempts to identify their origin.
INTERPRETATION
Similar does not mean identical.
A low detection accuracy does not prove that synthetic data is equivalent to real data, anonymous, or safe for every use. It only indicates that the classifier used in this experiment had difficulty distinguishing the two populations using the selected variables.
The dependency structure used here is probabilistic. It should not be interpreted as proof that one customer behaviour causes another.
VIEW PYTHON ANALYSIS →