Data Scientist · Applied AI Scientist · Ph.D. Statistics

Turning ambiguous business problems into production decision systems.

I lead applied AI at a group holding company's AI center, across manufacturing, logistics, bio/healthcare, retail and media. LLM/RAG and agentic systems, mathematical optimization, causal inference and forecasting — owned end to end, from problem definition to deployment and validation.

Seoul, Korea7 patents · 8 publicationsTechnical lead, CJ AI Center
7patents (lead or named inventor)
$1.86Mconfirmed bio-process cost savings
+20%logistics productivity, field-validated
95%recall, ad legal-risk review in production
8peer-reviewed papers & posters
Featured work

Systems that run in real operations

Three recent builds at CJ AI Center, each patented and in use by the business. Every one started as a problem nobody had framed yet. All 19 write-ups →

Capabilities

Four ways I turn a question into a system

Methodological depth from a statistics Ph.D., applied by hand: each capability below is listed with the production or research work it was used in.

LLM/RAG & agentic systems

LangGraph workflows, hybrid search and re-ranking, verification agents, prompt and context engineering, self-hosted model internalization, FM evaluation.

Optimization, simulation & decision systems

MILP, metaheuristics, reinforcement learning, discrete-event simulation, routing, bottleneck and inventory models, constraint modeling.

Bio/healthcare AI & synthetic data

Protein embeddings (ESM2, gLM), structure and clustering tools, microbiome modeling, biostatistics; GAN/diffusion/copula synthetic data with privacy-utility evaluation.

Live demos

The core models, rebuilt on synthetic data

Each write-up embeds a re-implementation of the project's core: a baseline and the model side by side, with the result in one line. Optimization and causal models are solved live on a Python server; forecasting, causal-discovery, synthetic-data and experimentation demos run in the browser; GenAI demos compute retrieval and checks in the browser and replay pre-computed LLM steps. No company data, identifiers or prompts are used. Nine are shown here; every project has one.

07 · GenAI · regulated domain

Ad copy legal-risk review

Keyword filter vs retrieval + LLM judgment vs an independent verifier, on invented ads.

OCRRAGVerification
Read with the live model →
08 · GenAI · multi-agent

Recruitment assistant

A first draft with an invented claim, caught by the verification layer and regenerated.

AgentsGroundingSelf-hosted
Read with the live model →
17 · Bio AI · in the browser

Immunotherapy response

Species-only vs protein-family clusters on held-out cohorts, and a cluster index over a protein catalogue.

EmbeddingsVector search
Read with the live model →
02 · Logistics · Python server

QPS allocation & injection order

Two ring conveyors replay the same batch: arrival order vs simulation + Tabu Search.

SimulationTabu Search
Read with the live model →
04 · Media · Python server

Cinema scheduling

A rule-of-thumb day vs a CP-SAT schedule with zero rule violations and on-target seat share.

CP-SAT30+ rules
Read with the live model →
05 · Manufacturing · Python server

Process causal model

The variable correlation ranks #3 has no causal effect; do-interventions find the real levers.

Causal discoverydo-operator
Read with the live model →
18 · Retail · in the browser

Bakery demand & production

Last week plus 10% vs a global forecast across 300 stores, baked at the profit-maximizing quantile.

Global modelNewsvendor
Read with the live model →
11 · Telecom · in the browser

Churn leading indicators

The indicators most correlated with churn follow it; PCMCI finds the ones that lead it, with lags and effect sizes.

PCMCIG-formula
Read with the live model →
13 · Experimentation · in the browser

A/B testing platform

A before/after launch fooled by a promotion week vs a randomized test read at its planned sample.

Sample sizeBayesianBandit
Read with the live model →
Selected work

Experience, project by project

Every linked title opens a write-up with its live model.

CJ AI Center

2023 — present · Data Scientist / Applied AI Scientist, Technical Lead

The group holding company's central AI organization, taking on cross-affiliate problems that individual companies cannot solve alone. I set technical direction, review architectures, mentor data scientists and work directly with affiliate executives and domain experts.

GenAI · Patent · Production

Advertising copy legal-risk review AI →

OCR with coordinate grounding, RAG over an 11,000-item per-category violation knowledge store, an independent verification unit; lead inventor.

Precision 81% · Recall 95%499 ad images
Multi-agent · Patent

Group-wide HR recruitment assistant →

Parallel summary, evidence and interview-question agents with a self-correcting verification layer on internalized open-weight models.

QA-FactEval 83.45%review time −50%
GenAI · Production

Dr. SmileGut personal AI advisor →

Vector DB, LangGraph workflow, hybrid search and re-ranking, guardrails for a microbiome-health advisor.

Commercial launch 2025-04-09
Agentic · Deployed

Logistics QPS real-time decision agent →

Workload allocation, task priority and bottleneck response, plus an LLM-agent AI Floor Supervisor in operation; multi-center rollout.

+20% productivityfield-validated
Analytics · LLM

Integrated data analytics & persona analysis →

Cross-affiliate data connected with LLM-based persona estimation and an insight dashboard.

Drove platform investment
Bio AI · Research

Immunotherapy response prediction & biomarker discovery →

ESM2/gLM embeddings, AlphaFold/ESMFold, MMseqs2/Foldseek clustering; vector search over ~6M sequences in 1–2 s.

+15% over ML baselineSITC 2024 poster
Causal · 5 patents

AI-based bio-process optimization →

Causal discovery and inference over fragmented process data across three global plants, and an in-silico simulator for testing hypotheses.

≈ USD 1.86M savingstechnology licensing
Optimization

Cart picking system optimization →

Order-box mapping, cart allocation, routing and bottleneck waiting reformulated as one problem, with standardized inputs for multi-center rollout.

≈20% productivityin simulation
MILP · Forecasting

Cinema scheduling optimization →

30+ scheduling rules as MILP constraints with tractability work; occupancy prediction by site and time slot.

≈10% MAEoccupancy model
Feasibility

Inventory optimization →

MILP, search, RL and simulation evaluated; demand-based inventory policy proposed.

Capacity & transport savings quantified
Forecasting · Licensed

Store demand forecasting & production recommendation →

Re-scoped a seasonal forecast into year-round store/category forecasting for 300+ stores (TFT, PatchTST, DeepAR) with a Streamlit dashboard.

+5.7% revenuetechnology licensing
Forecasting

Predictive management signals →

Daily sales forecasting and a traffic-light system that flags monthly-target risk about a month early.

97.6% accuracy

LG Uplus

2020 — 2023 · Data Scientist / Technical Lead

Customer retention, customer value, experimentation and synthetic data on telecom-scale data across mobile, IPTV and subscription services.

Synthetic data

Synthetic data generation & public-sector data sharing →

WGAN, CTGAN, CTAB-GAN, TabFair GAN, TTS-GAN, TABDDPM, MTcopula; certified anonymized IPTV, delivery-app and mobility data through the KISA program.

External data provision unlocked
Causal inference

Causal drivers of mobile churn →

~400 variables; PC, PCMCI, Bayesian networks, PSM, IPTW, G-formula.

Churn-gap vs. leader narrowed
Experimentation

A/B testing platform →

Metric definition, sample size, randomization, testing, Bayesian and bandit extensions for IPTV and mobile TV.

Standardized experiments
Forecasting

IPTV movie revenue forecasting →

Title totals with tree ensembles, post-release trajectories with the Bass diffusion model, the whole market with Prophet.

Acquisition & promotion decisionson expected revenue
Customer value

Customer lifetime value →

Temporal Fusion Transformer for revenue trajectories; MTLR and Cox with elastic net for survival; combined into customer-level long-term value.

Future-value lensfor retention resources
Network analysis

Family & household inference →

Call, location and subscription signals; Parallel Louvain community detection with location-based household inference; an ID structure over time.

Household-level customer view
Customer analytics

Engagement score & rank →

RFM usage variables, optimal binning with a genetic algorithm, least-squares service weights against ARPU and churn.

Operational fan score
Service & marketing

Additional analytics projects

Preferred-team prediction for a baseball streaming service with BYOL and label-imbalance methods; process mining of subscription flows for UI/UX fixes; interest indices for LG Hausys target marketing combined by rank aggregation.

Personalization, UX, targeting

CCNI Research · Hanmi Pharmaceutical

2017 — 2020 · Statistician / Biostatistician

Healthcare prediction and clinical-trial statistics, where survival analysis, hypothesis testing and sample-size work met real medical and pharmaceutical studies.

Healthcare prediction

Disease exacerbation prediction

Environmental and patient data for asthma, COPD and pediatric-cancer exacerbation risk; SMOTE and ADASYN for rare events and deep neural networks, connected to a patient-facing app.

CCNI Research2018 — 2020
Clinical trials

Statistical analysis and analysis plans

Sample-size calculation, analysis-method review, significance testing and interpretation for clinical trials and their statistical analysis plans.

CCNI Research
Biostatistics

Early-phase dose finding

Bayesian logistic regression, the continual reassessment method and overdose control for dose finding; graphical test procedures for multiple comparisons.

Hanmi Pharmaceutical2017 — 2018
Research & IP

Publications and patents

Statistical assessment of biosimilarity based on the relative distance between follow-on biologics for time-to-event endpointsStatistics in Biopharmaceutical Research (2020) · Lee & Kang · first-author methodology
Alertness during working hours among eight-hour rotating-shift nurses: an observational studyJournal of Nursing Scholarship 54(4), 2022
Effects of low-dose pirfenidone on survival and lung function decline in patients with IPF: results from a real-world studyPLoS ONE 16(12), 2021
Sleep, fatigue and alertness during working hours among rotating-shift nurses in KoreaJournal of Nursing Management 29(8), 2021
Readmission of high-risk discharged patients at a tertiary hospital in KoreaThe Journal for Healthcare Quality 41(4), 2019
Combined effects of diabetes and low household income on mortality: a 12-year follow-up of 505,677 Korean adultsDiabetic Medicine 35(10), 2018
Diabetes, frequency of exercise, and mortality over 12 years: analysis of the NHIS-HEALS databaseJournal of Korean Medical Science 33(8), 2018
Patents
7
Also
  • Technology-licensing agreements (demand forecasting, process optimization)
  • Internal technical documents, RL hands-on workshop, trend reports, LLM reading group
  • Teaching at Yonsei University: R & Python programming, Introduction to Statistics
About

Statistician by training, builder by habit

I hold a Ph.D. in Statistics from Yonsei University, where my dissertation developed biosimilarity assessment methodology for time-to-event endpoints. That training — survival analysis, hypothesis testing, experimental design — is still the backbone of how I evaluate every model I ship.

My core strength is redefining a business problem into an analyzable one: translating implicit domain knowledge, operational constraints and business rules into objective functions, evaluation metrics and production-ready systems. I own projects end to end, from problem definition and data pipelines through deployment and performance validation, and I prefer systems that business teams can keep using without me.

LanguagesKorean (native) · English (conversational; peer-reviewed technical writing)
ToolsPython, R, SQL, SAS, SPSS · LangGraph, vector DBs, Streamlit · MILP solvers, simulation
HonorsBrain Korea 21 Plus research excellence scholarship and fellowship · Korea Weather Big Data Contest, agency president award (2016)
2023 —
CJ AI CenterData Scientist / Applied AI Scientist, Technical Lead
2020 — 2023
LG UplusData Scientist / Technical Lead
2018 — 2020
CCNI ResearchStatistician
2017 — 2018
Hanmi PharmaceuticalBiostatistician
2015 — 2020
Yonsei UniversityPh.D. in Statistics · GPA 4.2 / 4.3
Contact

Working on a hard, unframed problem? Let's talk.

thstar.ai@gmail.com

Seoul, Korea · Always happy to talk about applied AI, optimization and causal inference problems.