Who lives with whom: household inference from network signals
An operator sells mobile lines to individuals and internet, IPTV and bundles to homes — yet its data knows only individuals, and the billing account is a poor stand-in for the family. This is the model that treated households as small communities in a graph of calls, locations and subscriptions, found them with Parallel Louvain, and gave each one an identifier that lasts.
Individuals in the data, households in the business
LG Uplus provides not only mobile services but also home-based services such as internet and IPTV. Understanding customers only as individuals misses family and household-level service patterns that matter for bundled products, targeting, retention and service recommendations: who should be offered a family plan, which home is one cancellation away from losing three lines, where a bundle discount would land on people who already live together. The obvious proxy — the billing account — splits families across accounts and carries lines that are not family at all.
What was built
- Relationships from indirect signals. Call history, location information and service-subscription data were combined to infer relationships between lines. None of the signals is reliable alone; together they are.
- Households as communities. Treating families and households as small communities, network analysis and community detection were applied: Parallel Louvain to identify tightly connected groups of lines, combined with location-based household inference to turn groups into homes.
- An identity that persists. An ID structure tracks family and household changes over time — members join, leave, move — so that analysis and campaigns address the same home across months rather than a fresh clustering each run.
Live model, computed in your browser on a generated city of about a thousand lines in 420 planted households. Left: lines grouped by billing account. Right: a weighted graph of calls, night-time co-location and shared accounts, Louvain communities, strong-tie households inside them, and a size split — scored against the planted households and rerun a month later to see how many household IDs survive. Try the dense-apartment setting, where neighbours share a cell at night. Open the live model on its own page ↗
Architecture, stack and core formulation
Signals into a weighted graph, communities by parallel Louvain, households from location inference inside them, and an identifier that follows each household over time.
Calls, location, subscriptions
Call history, location information and service-subscription data turned into pairwise relationship evidence.
Weighted ties
Lines as nodes, combined evidence as edge weights.
Parallel Louvain
Modularity-maximizing community detection over the full graph, in parallel.
Location-based inference
Groups turned into households using location-based inference.
Head of household
An ID keyed to a head of household tracks joins, departures and moves over time.
| Layer | Technology | What it does here |
|---|---|---|
| Inputs | Call history, location information, subscription data | Indirect evidence of family relationships |
| Graph | Weighted relationship graph | Lines connected by combined signal strength |
| Community detection | Parallel Louvain, modularity optimization | Tightly connected groups at full subscriber scale |
| Household inference | Location-based household assignment | From social groups to homes |
| Tracking | Head-of-household ID structure | The same household across months |
modularity Q = (1 / 2m) · Σ_ij [ A_ij − k_i·k_j / 2m ] · δ(c_i, c_j)
Louvain phase 1: move node i to the neighbour community with the largest ΔQ
phase 2: collapse communities into nodes; repeat until Q stops rising
edge weight w_ij = f(calls_ij, night co-location_ij, shared account_ij) (live model: interaction terms)
household connected groups of strong ties inside a community; split groups too large for one home
ID head = longest-tenured member; ID kept while most members stay together- Resolution limit. Modularity favours communities large relative to the graph, so tiny households need a second, local step inside each community.
- Signals fail differently. Night co-location is strong in suburbs and weak in apartment blocks; calls separate family from neighbours most of the time. Combining them degrades gracefully.
- The ID is the product. A stable household ID is what lets other teams join their data and run family-level campaigns.
| Component | In production | In the live model above |
|---|---|---|
| Scale | Full subscriber graph, parallel Louvain | About a thousand lines, Louvain in JavaScript |
| Households | Location-based inference | Strong ties (two signals coinciding) inside communities, with size splitting |
| Signals | Real call, location and subscription data | Generated calls, night co-location and accounts with planted households |
Design notes
Communities, then homes
Modularity-based community detection is built for groups that are large relative to the graph; households are tiny, and any single tie between two of them can be enough for Louvain to merge them. The practical answer was to use Louvain for what it is good at — carving a city-scale graph into social neighbourhoods in parallel — and then to recognise homes inside each neighbourhood from ties where several signals coincide, splitting any group too large to be one home. The live model shows the same two-stage logic.
Where the signals fail
Night-time location is the strongest signal in a suburb and a weak one in an apartment block, where a cell covers dozens of homes. Calls distinguish family from neighbours most of the time and fail for close friends next door. The combination degrades gracefully instead of collapsing, which is why the dense setting in the live model is harder but still ahead of the paperwork.
The identifier is the product
A clustering that changes every month is unusable by a campaign team. Keying each household to a head — its longest-tenured line — and carrying the ID forward while most of the household stays together turned a monthly analysis into a stable dimension other teams could join their data to.
Outcome
The model created a foundation for customer understanding at the family and household level, enabling more realistic segmentation, bundled-product strategy, family-level campaigns and long-term relationship analysis.
Limitations
- Inferred households are probabilistic; they are a basis for offers and analysis, not a record of who lives where, and the use of location and call signals was governed by the operator's privacy policy.
- Single-person households and close neighbours are the hard cases; the live model's dense setting shows the error rate rising there.
- The embedded model uses a plain Louvain on a thousand lines with three planted signals; the production system ran a parallel implementation over a far larger graph with richer signals.
About the demo and confidentiality
Lines, calls, locations, accounts and households in the embedded model are generated from a planted structure. No subscriber, call, location or account data from any operator appears here, and the real signal set and thresholds are not described.