Data that can leave the building
A telecom's viewing logs, location patterns and usage histories are among the most valuable datasets in the country — and among the most sensitive. This is the synthetic-data programme I led as technical lead: evaluating generative methods for tabular and time-series data, building the tools to measure utility and privacy, and producing reproduced datasets that were certified as anonymous and provided to a public-sector data platform.
Valuable, sensitive, stuck
Telecommunications companies hold data such as viewing logs, location patterns, payment behaviour and service-usage history. Privacy, re-identification and legal risk limit its external use and foreclose data-partnership opportunities — the data is most valuable exactly where it cannot go. Masking and aggregation are the traditional way out, and they discard most of what made the data valuable in the first place.
Synthetic data takes a different route: learn the statistical structure of a table, then generate new rows from that structure. If the generated rows preserve the relationships an analyst needs and reproduce no real person, the data can be shared.
What was built
- Method research and prototyping. As technical lead I researched and prototyped synthetic-data generation for tabular and time-series data, evaluating Wasserstein GAN, Conditional GAN, CTGAN, CTAB-GAN and CTAB-GAN Plus, TabFair GAN, MTcopula, TABDDPM, TTS-GAN and TTS-CGAN.
- A generation and evaluation tool. A web application for generating synthetic tabular and time-series data while evaluating privacy–utility trade-offs, so that a dataset's fitness for release could be read from a report rather than argued.
- Public-sector data provision. Through the KISA data-platform project, reproduced datasets were generated from company data, evaluated for utility and anonymity, and provided externally in anonymized form.
Live model, computed in your browser. A table of generated "subscribers" is the source; a naive de-identification (shuffle each column independently) and a Gaussian-copula generator each produce a synthetic table, and both get the same QA report: marginals, correlation matrices, a churn model trained on synthetic rows and tested on real ones, exact-copy counts and nearest-record distances against a holdout. The copula is the simplest member of the family evaluated in the project. Open the live model on its own page ↗
Architecture, stack and core formulation
A generation and evaluation toolkit: profile a table, train a generator, sample new rows, and release only what passes a utility and privacy report — packaged as a web application.
Columns and conditions
Upload a table or time series, set column types and the variables to reproduce or condition on.
Pick a generator
Tabular GANs, conditional GANs, diffusion and copula models; transformer GANs for time series.
Sample on demand
Draw as many rows as needed, optionally under conditions.
Utility and privacy
Distributions and relationships compared with the source; models trained on synthetic data tested on real; re-identification risk checked.
Certified anonymous
Datasets reviewed and certified as anonymous through the KISA data-platform program before leaving the company.
| Layer | Technology | What it does here |
|---|---|---|
| Tabular generators | Wasserstein GAN, conditional GAN, CTGAN, CTAB-GAN, CTAB-GAN+, TabFairGAN | Mixed continuous and categorical tables |
| Diffusion & copulas | TabDDPM, MTcopula, Gaussian copulas | Alternative generators with different strengths |
| Time series | TTS-GAN, TTS-CGAN (transformer-based) | Viewing logs and other sequences |
| Evaluation | Marginal and correlation fidelity, train-on-synthetic / test-on-real, privacy checks | The evidence for release decisions |
| Application | Web MVP on Google Cloud Run | Upload → configure → train → QA report → download |
WGAN min_G max_(‖D‖_L ≤ 1) E_x[ D(x) ] − E_z[ D(G(z)) ]
CTGAN continuous column → mode-specific normalization (variational Gaussian mixture)
discrete columns → conditional vector + training-by-sampling
copula z_j = Φ⁻¹( F̂_j(x_j) ), z ~ N(0, Σ̂), x̃_j = F̂_j⁻¹( Φ(z_j) )
TabDDPM q(x_t | x_(t−1)) adds noise step by step; learn p_θ(x_(t−1) | x_t) to reverse it
TSTR utility = AUC( model trained on synthetic, evaluated on real holdout )- Utility is a relationship. Matching each column's histogram is easy; the report leads with correlations and train-on-synthetic, test-on-real performance.
- Privacy against a holdout. How close synthetic rows come to real training rows is compared with how close new real rows come — memorization shows up as rows that are too close.
- One framework, many models. Every generator was measured the same way before deciding what could be released.
| Component | In production | In the live model above |
|---|---|---|
| Generators | GAN, diffusion and copula families above | A Gaussian copula, written in JavaScript |
| Baseline | — | Column shuffling, to show why marginals are not enough |
| Evaluation | Utility and anonymity review, KISA certification | KS / TV, correlation error, TSTR AUC, exact copies, nearest-record distance |
| Data | IPTV viewing, delivery-app usage, origin–destination data | A generated subscriber table |
Design notes
Utility is a relationship, not a histogram
Column shuffling is a useful foil because it is perfect on the easiest check — every marginal distribution is exactly preserved — and useless on the one that matters: a model trained on shuffled rows learns nothing about real customers. Any evaluation that stops at histograms will approve it. The report therefore leads with dependence: correlation error and train-on-synthetic, test-on-real performance.
Privacy against a holdout, not against intuition
"No exact copies" is necessary and far from sufficient. The more informative test compares how close synthetic rows come to real training rows with how close genuinely new real rows — a holdout the generator never saw — come to them. Synthetic rows that sit no nearer than strangers do not single anyone out; rows that cluster tightly around training records are memorization, however the model was trained.
Why the family matters
A copula captures monotone dependence cleanly and is nearly free to fit; GANs and diffusion models capture multi-modal and conditional structure a copula misses, at the cost of training stability and tuning. The programme's contribution was less any one model than the discipline of measuring each one the same way before deciding what could leave.
Outcome
The project converted sensitive datasets — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — into externally usable synthetic data, and provided an early practical case of using synthetic data to unlock data value while addressing privacy constraints.
Limitations
- Synthetic data cannot contain more information than its source; rare segments and subtle interactions are the first to blur.
- Privacy metrics are evidence, not proof: distance- and copy-based checks address memorization, and release decisions also rested on a formal anonymity review.
- The live model fits a Gaussian copula on a seven-column table; the project's datasets were wider, included time series, and used the GAN and diffusion models listed above.
About the demo and confidentiality
The "source" subscribers in the embedded model are themselves generated for this page from a few fixed rules. No customer data, schema, evaluation result or provided dataset from the project appears here.