Live model · generated subscriber table · computed in your browser

Data that can leave the building: keep the structure, copy no one

A table of subscribers is useful to a partner only if its relationships survive and none of its people do. Shuffling each column — a common first instinct — keeps every distribution and destroys every relationship. A generative model fitted to the table produces new rows that keep the correlations, train a working model and match no real record. The same QA report is run on both.

Synthetic data QA report

Fitting the generator…

Source table · first rows of generated subscribers
source rowssynthetic rowsheatmaps: Pearson correlation, orange + / blue −
1 · Fit

Marginals and a dependence structure

Each column is mapped to normal scores through its own distribution — categories through their frequency intervals — and the correlation of those scores is estimated. This Gaussian copula is the simplest member of the family evaluated in the project, which ran from copulas to CTGAN, CTAB-GAN and diffusion models.

2 · Generate

New rows, mapped back

Correlated normal draws are pushed back through each column's inverse distribution, so values stay in range, categories stay valid and the joint pattern — who pays what, who leaves — is reproduced without any real row being reused.

3 · Verify before release

Utility and privacy, measured

Marginals, correlations, a model trained on synthetic rows and tested on real ones, exact-copy counts and nearest-record distances against a real holdout. In the project this evidence supported certification of the reproduced data as anonymous before it left the company.