Network analysis · Community detection · Customer dataLG Uplus, mobile and home servicesWrite-up October 2026 · 6 min read

Who lives with whom: household inference from network signals

An operator sells mobile lines to individuals and internet, IPTV and bundles to homes — yet its data knows only individuals, and the billing account is a poor stand-in for the family. This is the model that treated households as small communities in a graph of calls, locations and subscriptions, found them with Parallel Louvain, and gave each one an identifier that lasts.

Built withPythonGraph constructionParallel LouvainModularityLocation inferenceID tracking
on the blue account · lives with the orange household signals: calls · nights on the same cell · shared accounthouseholds = communities · heads carry the ID
Three households as communities in the signal graph. The outlined line sits on another family's billing account but calls, sleeps and lives with the orange one.

Individuals in the data, households in the business

LG Uplus provides not only mobile services but also home-based services such as internet and IPTV. Understanding customers only as individuals misses family and household-level service patterns that matter for bundled products, targeting, retention and service recommendations: who should be offered a family plan, which home is one cancellation away from losing three lines, where a bundle discount would land on people who already live together. The obvious proxy — the billing account — splits families across accounts and carries lines that are not family at all.

What was built

  1. Relationships from indirect signals. Call history, location information and service-subscription data were combined to infer relationships between lines. None of the signals is reliable alone; together they are.
  2. Households as communities. Treating families and households as small communities, network analysis and community detection were applied: Parallel Louvain to identify tightly connected groups of lines, combined with location-based household inference to turn groups into homes.
  3. An identity that persists. An ID structure tracks family and household changes over time — members join, leave, move — so that analysis and campaigns address the same home across months rather than a fresh clustering each run.

Live model, computed in your browser on a generated city of about a thousand lines in 420 planted households. Left: lines grouped by billing account. Right: a weighted graph of calls, night-time co-location and shared accounts, Louvain communities, strong-tie households inside them, and a size split — scored against the planted households and rerun a month later to see how many household IDs survive. Try the dense-apartment setting, where neighbours share a cell at night. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

Signals into a weighted graph, communities by parallel Louvain, households from location inference inside them, and an identifier that follows each household over time.

1 · Signals

Calls, location, subscriptions

Call history, location information and service-subscription data turned into pairwise relationship evidence.

Python
2 · Graph

Weighted ties

Lines as nodes, combined evidence as edge weights.

graph
3 · Communities

Parallel Louvain

Modularity-maximizing community detection over the full graph, in parallel.

Parallel Louvain
4 · Households

Location-based inference

Groups turned into households using location-based inference.

location inference
5 · Identity

Head of household

An ID keyed to a head of household tracks joins, departures and moves over time.

ID structure
Stack
LayerTechnologyWhat it does here
InputsCall history, location information, subscription dataIndirect evidence of family relationships
GraphWeighted relationship graphLines connected by combined signal strength
Community detectionParallel Louvain, modularity optimizationTightly connected groups at full subscriber scale
Household inferenceLocation-based household assignmentFrom social groups to homes
TrackingHead-of-household ID structureThe same household across months
Core formulation
modularity   Q = (1 / 2m) · Σ_ij [ A_ij − k_i·k_j / 2m ] · δ(c_i, c_j)

Louvain      phase 1: move node i to the neighbour community with the largest ΔQ
             phase 2: collapse communities into nodes; repeat until Q stops rising

edge weight  w_ij = f(calls_ij, night co-location_ij, shared account_ij)        (live model: interaction terms)
household    connected groups of strong ties inside a community; split groups too large for one home
ID           head = longest-tenured member;   ID kept while most members stay together
  • Resolution limit. Modularity favours communities large relative to the graph, so tiny households need a second, local step inside each community.
  • Signals fail differently. Night co-location is strong in suburbs and weak in apartment blocks; calls separate family from neighbours most of the time. Combining them degrades gracefully.
  • The ID is the product. A stable household ID is what lets other teams join their data and run family-level campaigns.
In production vs in the live model
ComponentIn productionIn the live model above
ScaleFull subscriber graph, parallel LouvainAbout a thousand lines, Louvain in JavaScript
HouseholdsLocation-based inferenceStrong ties (two signals coinciding) inside communities, with size splitting
SignalsReal call, location and subscription dataGenerated calls, night co-location and accounts with planted households

Design notes

Communities, then homes

Modularity-based community detection is built for groups that are large relative to the graph; households are tiny, and any single tie between two of them can be enough for Louvain to merge them. The practical answer was to use Louvain for what it is good at — carving a city-scale graph into social neighbourhoods in parallel — and then to recognise homes inside each neighbourhood from ties where several signals coincide, splitting any group too large to be one home. The live model shows the same two-stage logic.

Where the signals fail

Night-time location is the strongest signal in a suburb and a weak one in an apartment block, where a cell covers dozens of homes. Calls distinguish family from neighbours most of the time and fail for close friends next door. The combination degrades gracefully instead of collapsing, which is why the dense setting in the live model is harder but still ahead of the paperwork.

The identifier is the product

A clustering that changes every month is unusable by a campaign team. Keying each household to a head — its longest-tenured line — and carrying the ID forward while most of the household stays together turned a monthly analysis into a stable dimension other teams could join their data to.

Outcome

Household viewof customers who had only been visible as lines
Segmentation & bundlesplanned on real homes rather than billing accounts
LongitudinalIDs that track family changes over time

The model created a foundation for customer understanding at the family and household level, enabling more realistic segmentation, bundled-product strategy, family-level campaigns and long-term relationship analysis.

Limitations

About the demo and confidentiality

Lines, calls, locations, accounts and households in the embedded model are generated from a planted structure. No subscriber, call, location or account data from any operator appears here, and the real signal set and thresholds are not described.

Taehee Lee · Data Scientist / Technical Lead, LG Uplus (2020 – 2023)Signal design, graph construction, community detection, household inference and the ID structure. Demo re-implemented on generated data for this site.