Synthetic data · Privacy · Generative modelsLG Uplus, technical lead project · KISA data platformWrite-up October 2026 · 7 min read

Data that can leave the building

A telecom's viewing logs, location patterns and usage histories are among the most valuable datasets in the country — and among the most sensitive. This is the synthetic-data programme I led as technical lead: evaluating generative methods for tabular and time-series data, building the tools to measure utility and privacy, and producing reproduced datasets that were certified as anonymous and provided to a public-sector data platform.

Built withPythonWGANCTGANCTAB-GAN+TabFairGANTTS-GANTabDDPMCopulasCloud Run
SOURCE · REAL PEOPLE identifiable · cannot be shared generative modelmarginals + dependence SYNTHETIC · NO ONE correlation error0.04model trained on it · AUC0.79 vs 0.80rows copied from source0nearest real recordno closer than a stranger
The shape of the programme: a model learns a table's distributions and dependencies, generates rows that belong to no one, and a QA report proves both halves before anything is shared. Numbers illustrative.

Valuable, sensitive, stuck

Telecommunications companies hold data such as viewing logs, location patterns, payment behaviour and service-usage history. Privacy, re-identification and legal risk limit its external use and foreclose data-partnership opportunities — the data is most valuable exactly where it cannot go. Masking and aggregation are the traditional way out, and they discard most of what made the data valuable in the first place.

Synthetic data takes a different route: learn the statistical structure of a table, then generate new rows from that structure. If the generated rows preserve the relationships an analyst needs and reproduce no real person, the data can be shared.

What was built

Live model, computed in your browser. A table of generated "subscribers" is the source; a naive de-identification (shuffle each column independently) and a Gaussian-copula generator each produce a synthetic table, and both get the same QA report: marginals, correlation matrices, a churn model trained on synthetic rows and tested on real ones, exact-copy counts and nearest-record distances against a holdout. The copula is the simplest member of the family evaluated in the project. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

A generation and evaluation toolkit: profile a table, train a generator, sample new rows, and release only what passes a utility and privacy report — packaged as a web application.

1 · Profile

Columns and conditions

Upload a table or time series, set column types and the variables to reproduce or condition on.

web app
2 · Train

Pick a generator

Tabular GANs, conditional GANs, diffusion and copula models; transformer GANs for time series.

CTGANCTAB-GAN+TabDDPMTTS-GAN
3 · Generate

Sample on demand

Draw as many rows as needed, optionally under conditions.

conditional sampling
4 · QA report

Utility and privacy

Distributions and relationships compared with the source; models trained on synthetic data tested on real; re-identification risk checked.

evaluation
5 · Release

Certified anonymous

Datasets reviewed and certified as anonymous through the KISA data-platform program before leaving the company.

KISA program
Stack
LayerTechnologyWhat it does here
Tabular generatorsWasserstein GAN, conditional GAN, CTGAN, CTAB-GAN, CTAB-GAN+, TabFairGANMixed continuous and categorical tables
Diffusion & copulasTabDDPM, MTcopula, Gaussian copulasAlternative generators with different strengths
Time seriesTTS-GAN, TTS-CGAN (transformer-based)Viewing logs and other sequences
EvaluationMarginal and correlation fidelity, train-on-synthetic / test-on-real, privacy checksThe evidence for release decisions
ApplicationWeb MVP on Google Cloud RunUpload → configure → train → QA report → download
Core formulation
WGAN      min_G  max_(‖D‖_L ≤ 1)   E_x[ D(x) ] − E_z[ D(G(z)) ]

CTGAN     continuous column → mode-specific normalization (variational Gaussian mixture)
          discrete columns → conditional vector + training-by-sampling

copula    z_j = Φ⁻¹( F̂_j(x_j) ),    z ~ N(0, Σ̂),    x̃_j = F̂_j⁻¹( Φ(z_j) )

TabDDPM   q(x_t | x_(t−1)) adds noise step by step;    learn p_θ(x_(t−1) | x_t) to reverse it

TSTR      utility = AUC( model trained on synthetic, evaluated on real holdout )
  • Utility is a relationship. Matching each column's histogram is easy; the report leads with correlations and train-on-synthetic, test-on-real performance.
  • Privacy against a holdout. How close synthetic rows come to real training rows is compared with how close new real rows come — memorization shows up as rows that are too close.
  • One framework, many models. Every generator was measured the same way before deciding what could be released.
In production vs in the live model
ComponentIn productionIn the live model above
GeneratorsGAN, diffusion and copula families aboveA Gaussian copula, written in JavaScript
Baseline—Column shuffling, to show why marginals are not enough
EvaluationUtility and anonymity review, KISA certificationKS / TV, correlation error, TSTR AUC, exact copies, nearest-record distance
DataIPTV viewing, delivery-app usage, origin–destination dataA generated subscriber table

Design notes

Utility is a relationship, not a histogram

Column shuffling is a useful foil because it is perfect on the easiest check — every marginal distribution is exactly preserved — and useless on the one that matters: a model trained on shuffled rows learns nothing about real customers. Any evaluation that stops at histograms will approve it. The report therefore leads with dependence: correlation error and train-on-synthetic, test-on-real performance.

Privacy against a holdout, not against intuition

"No exact copies" is necessary and far from sufficient. The more informative test compares how close synthetic rows come to real training rows with how close genuinely new real rows — a holdout the generator never saw — come to them. Synthetic rows that sit no nearer than strangers do not single anyone out; rows that cluster tightly around training records are memorization, however the model was trained.

Why the family matters

A copula captures monotone dependence cleanly and is nearly free to fit; GANs and diffusion models capture multi-modal and conditional structure a copula misses, at the cost of training stability and tuning. The programme's contribution was less any one model than the discipline of measuring each one the same way before deciding what could leave.

Outcome

Certifiedanonymous reproduced data provided through the KISA program
Threesensitive sources unlocked: IPTV real-time viewing, delivery-app usage, origin–destination data
Tengenerative methods evaluated on one privacy–utility framework

The project converted sensitive datasets — IPTV real-time viewing logs, delivery-app usage records and origin–destination information — into externally usable synthetic data, and provided an early practical case of using synthetic data to unlock data value while addressing privacy constraints.

Limitations

About the demo and confidentiality

The "source" subscribers in the embedded model are themselves generated for this page from a few fixed rules. No customer data, schema, evaluation result or provided dataset from the project appears here.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Technical lead: method research, prototyping, evaluation framework and the KISA data-provision work. Demo re-implemented on generated data for this site.