I lead applied AI at a group holding company's AI center, across manufacturing, logistics, bio/healthcare, retail and media. LLM/RAG and agentic systems, mathematical optimization, causal inference and forecasting — owned end to end, from problem definition to deployment and validation.
Three recent builds at CJ AI Center, each patented and in use by the business. Every one started as a problem nobody had framed yet. All 19 write-ups →
OCR with coordinate grounding, RAG over an 11,000-item per-category violation knowledge store, an independent verification unit and retrain-free policy evolution. Flags risky copy with its judgment basis and on-screen location.
Parallel summarization, evidence-highlighting and interview-question agents with an independent self-correcting verification layer. Core models internalized on an open GPT-OSS base; ~73,000 applications across three cycles.
Causal discovery and inference over fragmented process data to find actionable operating conditions, plus an in-silico simulator process experts use to test hypotheses before touching the line. Some conditions became plant standards.
Methodological depth from a statistics Ph.D., applied by hand: each capability below is listed with the production or research work it was used in.
LangGraph workflows, hybrid search and re-ranking, verification agents, prompt and context engineering, self-hosted model internalization, FM evaluation.
MILP, metaheuristics, reinforcement learning, discrete-event simulation, routing, bottleneck and inventory models, constraint modeling.
TFT, PatchTST, DeepAR, Prophet, tree ensembles; causal discovery (PC, PCMCI, Bayesian networks), PSM/IPTW/G-formula, A/B platforms, survival models.
Protein embeddings (ESM2, gLM), structure and clustering tools, microbiome modeling, biostatistics; GAN/diffusion/copula synthetic data with privacy-utility evaluation.
Each write-up embeds a re-implementation of the project's core: a baseline and the model side by side, with the result in one line. Optimization and causal models are solved live on a Python server; forecasting, causal-discovery, synthetic-data and experimentation demos run in the browser; GenAI demos compute retrieval and checks in the browser and replay pre-computed LLM steps. No company data, identifiers or prompts are used. Nine are shown here; every project has one.
Keyword filter vs retrieval + LLM judgment vs an independent verifier, on invented ads.
A first draft with an invented claim, caught by the verification layer and regenerated.
Species-only vs protein-family clusters on held-out cohorts, and a cluster index over a protein catalogue.
Two ring conveyors replay the same batch: arrival order vs simulation + Tabu Search.
A rule-of-thumb day vs a CP-SAT schedule with zero rule violations and on-target seat share.
The variable correlation ranks #3 has no causal effect; do-interventions find the real levers.
Last week plus 10% vs a global forecast across 300 stores, baked at the profit-maximizing quantile.
The indicators most correlated with churn follow it; PCMCI finds the ones that lead it, with lags and effect sizes.
A before/after launch fooled by a promotion week vs a randomized test read at its planned sample.
Every linked title opens a write-up with its live model.
The group holding company's central AI organization, taking on cross-affiliate problems that individual companies cannot solve alone. I set technical direction, review architectures, mentor data scientists and work directly with affiliate executives and domain experts.
OCR with coordinate grounding, RAG over an 11,000-item per-category violation knowledge store, an independent verification unit; lead inventor.
Parallel summary, evidence and interview-question agents with a self-correcting verification layer on internalized open-weight models.
Vector DB, LangGraph workflow, hybrid search and re-ranking, guardrails for a microbiome-health advisor.
Workload allocation, task priority and bottleneck response, plus an LLM-agent AI Floor Supervisor in operation; multi-center rollout.
Cross-affiliate data connected with LLM-based persona estimation and an insight dashboard.
ESM2/gLM embeddings, AlphaFold/ESMFold, MMseqs2/Foldseek clustering; vector search over ~6M sequences in 1–2 s.
Causal discovery and inference over fragmented process data across three global plants, and an in-silico simulator for testing hypotheses.
Order-box mapping, cart allocation, routing and bottleneck waiting reformulated as one problem, with standardized inputs for multi-center rollout.
30+ scheduling rules as MILP constraints with tractability work; occupancy prediction by site and time slot.
MILP, search, RL and simulation evaluated; demand-based inventory policy proposed.
Re-scoped a seasonal forecast into year-round store/category forecasting for 300+ stores (TFT, PatchTST, DeepAR) with a Streamlit dashboard.
Daily sales forecasting and a traffic-light system that flags monthly-target risk about a month early.
Customer retention, customer value, experimentation and synthetic data on telecom-scale data across mobile, IPTV and subscription services.
WGAN, CTGAN, CTAB-GAN, TabFair GAN, TTS-GAN, TABDDPM, MTcopula; certified anonymized IPTV, delivery-app and mobility data through the KISA program.
~400 variables; PC, PCMCI, Bayesian networks, PSM, IPTW, G-formula.
Metric definition, sample size, randomization, testing, Bayesian and bandit extensions for IPTV and mobile TV.
Title totals with tree ensembles, post-release trajectories with the Bass diffusion model, the whole market with Prophet.
Temporal Fusion Transformer for revenue trajectories; MTLR and Cox with elastic net for survival; combined into customer-level long-term value.
Call, location and subscription signals; Parallel Louvain community detection with location-based household inference; an ID structure over time.
RFM usage variables, optimal binning with a genetic algorithm, least-squares service weights against ARPU and churn.
Preferred-team prediction for a baseball streaming service with BYOL and label-imbalance methods; process mining of subscription flows for UI/UX fixes; interest indices for LG Hausys target marketing combined by rank aggregation.
Healthcare prediction and clinical-trial statistics, where survival analysis, hypothesis testing and sample-size work met real medical and pharmaceutical studies.
Environmental and patient data for asthma, COPD and pediatric-cancer exacerbation risk; SMOTE and ADASYN for rare events and deep neural networks, connected to a patient-facing app.
Sample-size calculation, analysis-method review, significance testing and interpretation for clinical trials and their statistical analysis plans.
Bayesian logistic regression, the continual reassessment method and overdose control for dose finding; graphical test procedures for multiple comparisons.
I hold a Ph.D. in Statistics from Yonsei University, where my dissertation developed biosimilarity assessment methodology for time-to-event endpoints. That training — survival analysis, hypothesis testing, experimental design — is still the backbone of how I evaluate every model I ship.
My core strength is redefining a business problem into an analyzable one: translating implicit domain knowledge, operational constraints and business rules into objective functions, evaluation metrics and production-ready systems. I own projects end to end, from problem definition and data pipelines through deployment and performance validation, and I prefer systems that business teams can keep using without me.
Seoul, Korea · Always happy to talk about applied AI, optimization and causal inference problems.