STUDY 06 · SYNTHETIC DATA

Can You Tell Which Customer Is Fake?

A probabilistic model trained on real online shopping behaviour, then I generated an artificial customer population, and then asked a machine-learning model to tell the real shoppers from the synthetic ones.

THE QUESTION

What makes synthetic data realistic?

A synthetic dataset can reproduce simple statistics such as purchase rate or visitor type while still failing to reproduce the relationships that exist between customer behaviours.

This experiment tests both. We compare the distributions directly and train a classifier whose only job is to identify whether a shopping session came from the real or synthetic population.

— REAL SESSIONS
— SYNTHETIC SESSIONS
— DETECTION ACCURACY
— AVG. DISTRIBUTION SIMILARITY

EXPERIMENT

Real data in. Artificial customers out.

01 Real shoppers

Start with observed online shopping sessions.

→
02 Learn dependencies

Estimate conditional behavioural relationships.

→
03 Generate

Create new artificial shopping sessions.

→
04 Challenge

Ask ML to identify which sessions are synthetic.

THE DETECTOR

Can the machine spot the fake?

Loading experiment results...

— classification accuracy
50%
random guessing

REAL VS SYNTHETIC

Do the populations look alike?

METRIC REAL SYNTHETIC DIFFERENCE
Purchase rate — — —
Returning visitors — — —
Weekend sessions — — —

DISTRIBUTION TEST

Similarity by variable

100% means the synthetic distribution closely matches the observed distribution.

WHY THIS MATTERS

Matching averages is not enough.

Two datasets can have almost identical averages while containing very different relationships between variables. Synthetic-data quality therefore has to be tested at more than one statistical level.

METHODOLOGY

How the experiment works

01

Source

Online Shoppers Purchasing Intention dataset containing observed web sessions.

02

Generation

Conditional probability distributions are estimated from behavioural relationships in the original data.

03

Validation

Real and synthetic distributions are compared using total variation distance.

04

ML Challenge

A random forest receives mixed real and synthetic observations and attempts to identify their origin.

INTERPRETATION

Similar does not mean identical.

A low detection accuracy does not prove that synthetic data is equivalent to real data, anonymous, or safe for every use. It only indicates that the classifier used in this experiment had difficulty distinguishing the two populations using the selected variables.

The dependency structure used here is probabilistic. It should not be interpreted as proof that one customer behaviour causes another.

VIEW PYTHON ANALYSIS →