GenAI · RAG · Regulated domainCJ AI Center, with CJ Olive YoungWrite-up October 2026 · 8 min read

Reviewing ad copy for legal risk — with the reason and the location

Every phrase in a beauty or food advertisement has to clear advertising and labelling law, and the same words can be fine in one category and a violation in another. This is the OCR-and-RAG system built to flag risky copy, explain why, point to where it sits on the ad, and check its own work.

Built withPythonCloud document OCRManaged LLMText embeddingsChromaDBAho-CorasickFastAPIPostgreSQLDocker
Barrier Cream Clears acne in 7 days The No.1 cream Ceramide formula for dry skin Naturally derived ingredients* *92% of ingredients by weight HIGH · MEDICAL EFFICACY CLAIM Claims to clear acne — a treatment claim not permitted for a cosmetic. nearest case 0.46 · verifier: kept CLEARED BY VERIFIER Footnote qualifies "naturally derived".
Every flag comes back with a category, a risk level, the reason, the case it resembles, and a box on the original image. Some flags are withdrawn by the verifier before anyone sees them.

Why keywords and generic LLMs both fall short

A large retailer runs ad copy on product pages, in-store displays and flyers, and every phrase must be checked against advertising and labelling law. This is not keyword detection: the same expression is judged differently by product category and context, and the decision depends on regulation, risk level and prior violation cases. Missing risky copy is costly — sales suspension, fines, damage to brand trust.

Review had relied on manual reading, keyword filters and OCR extraction, with violation cases and policy living in documents and people's experience. General-purpose LLMs reflect industry-specific regulation poorly, do not show their basis or the detection location, and cannot absorb new rules continuously — which limits both trust and scale.

System design

The system is a modular pipeline — preprocessing, a reference store, a judgment engine, an independent verification unit and output — with a policy layer that generates and updates review policies.

ModuleWhat it does
PreprocessingOCR extracts copy from images and documents, splits it into paragraphs, sentences and blocks, normalizes it, and keeps coordinates for every phrase so a flag can be drawn on the original and traced back.
Reference storePer-category violation cases (food, cosmetics, functional products), prohibited and recommended expressions, risk-scoring criteria and correction guides — a registry of about 11,000 risk items — used as the retrieval ground.
Judgment engineCombines the OCR text with retrieved references to judge copy in context, against category rules and past cases rather than keyword lists; outputs a risk score, the basis and a suggested correction.
Verification unitKept separate from judgment. Cross-checks criteria conformance, that the cited evidence exists, and score consistency; triggers re-judgment when a check fails.
Policy layerGenerates per-category review policies from the case store and updates them from reviewer outcomes and new violation cases, so a new category is covered by adding references — no retraining.

Live model on two invented ads. The three stages are compared on the same lines: a keyword filter, retrieval over a small case store plus the LLM judgment, and the independent verifier. Retrieval and keyword matching run in your browser; the LLM judgments and verifier notes were produced in advance for these ads and are replayed. Click a line to see its evidence. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

An offline knowledge-base build and an online review path — OCR with coordinates, retrieval over policy examples, an LLM judge, an independent critic and a keyword path — merged into a located, explained report.

1 · Knowledge base

Policies, enriched

Policy examples per risk category enriched offline with definitions, boundaries, core expressions and keywords, then embedded; updated incrementally every day.

LLMembeddingsChromaDB
2 · OCR

Text with coordinates

Images, PDFs and captured web pages go through document OCR; words rebuilt into blocks, nearby blocks merged, footnotes linked to the lines they qualify.

cloud OCRPyMuPDFPlaywright
3 · Retrieve

Recall first

Dense search over blocks and sentence spans plus substring matching against examples; best score per example, then top categories and evidence.

ChromaDBcosine HNSW
4 · Judge

Precision by LLM

A temperature-0 LLM decides per category with the retrieved examples, a few sentences per call.

LLM
5 · Verify

Independent critic

A second pass audits non-obvious violation flags and can only clear them — for an unsupported reason, inconsistency with sibling lines or a topic-only match.

LLM
6 · Serve

Report and API

Block-level JSON and an HTML overlay; async, batch and sync APIs with persistent jobs.

FastAPIPostgreSQLDocker
Stack
LayerTechnologyWhat it does here
IngestionPyMuPDF (200 dpi), Playwright full-page capture, image preprocessingImages, PDFs and web pages turned into OCR-ready images
OCRManaged document-text OCR, union-find block mergingText blocks with coordinates; footnotes attached
IndexManaged text embeddings, ChromaDB (cosine HNSW)Per-category policy examples and enriched expressions
JudgmentManaged lightweight LLM, temperature 0Per-category decision with a written reason
KeywordsAho-Corasick automata (pyahocorasick), Unicode normalizationProhibited-keyword path in parallel, with one context check per image
ServiceFastAPI, PostgreSQL (JSONB), object storage, Docker ComposeAsync / batch / sync APIs, job recovery, callbacks
Core formulation
score(e | s) = 1.0                  substring match of example e in sentence s
             = 1 − cos_dist(s, e)   otherwise                     keep the max per example
evidence     = top-3 categories by max score, ≤ 3 examples each, score ≥ τ

judge        LLM(s, category rules, evidence) → { violation, reason }      temperature 0
critic       audit only violation flags that are not high-confidence
             clear  if  reason unsupported ∨ inconsistent with siblings ∨ topic-only
             never turns a clear line into a violation;  on error keep the decision
  • Retrieval for recall, LLM for precision. Retrieval keeps every plausible match; the judge decides with the category's own rules in front of it.
  • A one-way critic. The verifier can only remove flags, which bounds how much a second model can cost in recall.
  • Cost scales with content. LLM calls per image ≈ one probe + ⌈sentences / batch⌉ per category + audited categories + one keyword check.
  • Evaluation. Ground-truth phrases are matched to OCR blocks and categories by Hungarian assignment; micro and macro precision, recall and F1.
In production vs in the live model
ComponentIn productionIn the live model above
OCRManaged document OCR with block rebuildingText blocks drawn on two invented ads
RetrievalEmbeddings + ChromaDB over the per-category policy storeCharacter-trigram TF-IDF over 25 invented cases
Judge & criticLive LLM callsPre-computed decisions, replayed
KeywordsAho-Corasick over a daily-refreshed listA 15-word keyword filter as the baseline

Explainability by construction

Two design choices carry most of the trust. First, coordinates survive the whole pipeline: a reviewer sees the exact phrase highlighted on the ad, not a sentence in a log. Second, judgment and verification are separate components. The verifier looks for a wrongly cited basis, a missing detection location, or a score inconsistent with similar lines, and sends the case back. Hallucination is handled structurally rather than by hoping the judge does not make mistakes.

The policy layer turns the case store into reusable review policies — scope, conditions, exceptions, relevant statutes, representative cases, risk criteria — and keeps them current from review outcomes. Covering a new product group means adding reference information, not changing the model or the architecture.

Results

81.14%precision on 499 ad images (KPI ≥ 70%)
94.96%recall on the same set (KPI ≥ 90%)
Leadinventor on the resulting patent

The quality-management team reviewed the AI-detected copy against a KPI of recall ≥ 90% and precision ≥ 70%; both were met. Across 499 images (299 food, 200 cosmetics), overall precision was 81.14% and recall 94.96% — food 82.88% / 93.83%, cosmetics 76.00% / 98.82%.

The system is in production on non-product display copy — in-store copy and flyers — with product-page copy moving onto a new system from October 2026. It removes the variance and workload of subjective manual review and turns every judgment into a reusable case with a standardized guide. Operation and category expansion are run by the group IT affiliate, and extension to other affiliates' label pre-checks and live-commerce ad monitoring is under review.

Limitations

About the demo and confidentiality

Products, brands, ad copy, violation cases and keywords in the embedded model are invented. No customer policy data, keyword list, ad image, prompt or system detail from the production service appears here.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterEnd-to-end design and development; lead inventor on the patent. Demo re-implemented on fictional ads for this site.