Reviewing ad copy for legal risk — with the reason and the location
Every phrase in a beauty or food advertisement has to clear advertising and labelling law, and the same words can be fine in one category and a violation in another. This is the OCR-and-RAG system built to flag risky copy, explain why, point to where it sits on the ad, and check its own work.
Why keywords and generic LLMs both fall short
A large retailer runs ad copy on product pages, in-store displays and flyers, and every phrase must be checked against advertising and labelling law. This is not keyword detection: the same expression is judged differently by product category and context, and the decision depends on regulation, risk level and prior violation cases. Missing risky copy is costly — sales suspension, fines, damage to brand trust.
Review had relied on manual reading, keyword filters and OCR extraction, with violation cases and policy living in documents and people's experience. General-purpose LLMs reflect industry-specific regulation poorly, do not show their basis or the detection location, and cannot absorb new rules continuously — which limits both trust and scale.
System design
The system is a modular pipeline — preprocessing, a reference store, a judgment engine, an independent verification unit and output — with a policy layer that generates and updates review policies.
| Module | What it does |
|---|---|
| Preprocessing | OCR extracts copy from images and documents, splits it into paragraphs, sentences and blocks, normalizes it, and keeps coordinates for every phrase so a flag can be drawn on the original and traced back. |
| Reference store | Per-category violation cases (food, cosmetics, functional products), prohibited and recommended expressions, risk-scoring criteria and correction guides — a registry of about 11,000 risk items — used as the retrieval ground. |
| Judgment engine | Combines the OCR text with retrieved references to judge copy in context, against category rules and past cases rather than keyword lists; outputs a risk score, the basis and a suggested correction. |
| Verification unit | Kept separate from judgment. Cross-checks criteria conformance, that the cited evidence exists, and score consistency; triggers re-judgment when a check fails. |
| Policy layer | Generates per-category review policies from the case store and updates them from reviewer outcomes and new violation cases, so a new category is covered by adding references — no retraining. |
Live model on two invented ads. The three stages are compared on the same lines: a keyword filter, retrieval over a small case store plus the LLM judgment, and the independent verifier. Retrieval and keyword matching run in your browser; the LLM judgments and verifier notes were produced in advance for these ads and are replayed. Click a line to see its evidence. Open the live model on its own page ↗
Architecture, stack and core formulation
An offline knowledge-base build and an online review path — OCR with coordinates, retrieval over policy examples, an LLM judge, an independent critic and a keyword path — merged into a located, explained report.
Policies, enriched
Policy examples per risk category enriched offline with definitions, boundaries, core expressions and keywords, then embedded; updated incrementally every day.
Text with coordinates
Images, PDFs and captured web pages go through document OCR; words rebuilt into blocks, nearby blocks merged, footnotes linked to the lines they qualify.
Recall first
Dense search over blocks and sentence spans plus substring matching against examples; best score per example, then top categories and evidence.
Precision by LLM
A temperature-0 LLM decides per category with the retrieved examples, a few sentences per call.
Independent critic
A second pass audits non-obvious violation flags and can only clear them — for an unsupported reason, inconsistency with sibling lines or a topic-only match.
Report and API
Block-level JSON and an HTML overlay; async, batch and sync APIs with persistent jobs.
| Layer | Technology | What it does here |
|---|---|---|
| Ingestion | PyMuPDF (200 dpi), Playwright full-page capture, image preprocessing | Images, PDFs and web pages turned into OCR-ready images |
| OCR | Managed document-text OCR, union-find block merging | Text blocks with coordinates; footnotes attached |
| Index | Managed text embeddings, ChromaDB (cosine HNSW) | Per-category policy examples and enriched expressions |
| Judgment | Managed lightweight LLM, temperature 0 | Per-category decision with a written reason |
| Keywords | Aho-Corasick automata (pyahocorasick), Unicode normalization | Prohibited-keyword path in parallel, with one context check per image |
| Service | FastAPI, PostgreSQL (JSONB), object storage, Docker Compose | Async / batch / sync APIs, job recovery, callbacks |
score(e | s) = 1.0 substring match of example e in sentence s
= 1 − cos_dist(s, e) otherwise keep the max per example
evidence = top-3 categories by max score, ≤ 3 examples each, score ≥ τ
judge LLM(s, category rules, evidence) → { violation, reason } temperature 0
critic audit only violation flags that are not high-confidence
clear if reason unsupported ∨ inconsistent with siblings ∨ topic-only
never turns a clear line into a violation; on error keep the decision- Retrieval for recall, LLM for precision. Retrieval keeps every plausible match; the judge decides with the category's own rules in front of it.
- A one-way critic. The verifier can only remove flags, which bounds how much a second model can cost in recall.
- Cost scales with content. LLM calls per image ≈ one probe + ⌈sentences / batch⌉ per category + audited categories + one keyword check.
- Evaluation. Ground-truth phrases are matched to OCR blocks and categories by Hungarian assignment; micro and macro precision, recall and F1.
| Component | In production | In the live model above |
|---|---|---|
| OCR | Managed document OCR with block rebuilding | Text blocks drawn on two invented ads |
| Retrieval | Embeddings + ChromaDB over the per-category policy store | Character-trigram TF-IDF over 25 invented cases |
| Judge & critic | Live LLM calls | Pre-computed decisions, replayed |
| Keywords | Aho-Corasick over a daily-refreshed list | A 15-word keyword filter as the baseline |
Explainability by construction
Two design choices carry most of the trust. First, coordinates survive the whole pipeline: a reviewer sees the exact phrase highlighted on the ad, not a sentence in a log. Second, judgment and verification are separate components. The verifier looks for a wrongly cited basis, a missing detection location, or a score inconsistent with similar lines, and sends the case back. Hallucination is handled structurally rather than by hoping the judge does not make mistakes.
The policy layer turns the case store into reusable review policies — scope, conditions, exceptions, relevant statutes, representative cases, risk criteria — and keeps them current from review outcomes. Covering a new product group means adding reference information, not changing the model or the architecture.
Results
The quality-management team reviewed the AI-detected copy against a KPI of recall ≥ 90% and precision ≥ 70%; both were met. Across 499 images (299 food, 200 cosmetics), overall precision was 81.14% and recall 94.96% — food 82.88% / 93.83%, cosmetics 76.00% / 98.82%.
The system is in production on non-product display copy — in-store copy and flyers — with product-page copy moving onto a new system from October 2026. It removes the variance and workload of subjective manual review and turns every judgment into a reusable case with a standardized guide. Operation and category expansion are run by the group IT affiliate, and extension to other affiliates' label pre-checks and live-commerce ad monitoring is under review.
Limitations
- Recall was prioritized over precision by design; reviewers still see some false alarms, by intent, rather than miss violations.
- OCR quality bounds everything downstream; very stylized type and low-resolution images still cause misses.
- The live model's ads, case store and keyword list are invented and tiny; its precision and recall illustrate the stages, not the production system.
About the demo and confidentiality
Products, brands, ad copy, violation cases and keywords in the embedded model are invented. No customer policy data, keyword list, ad image, prompt or system detail from the production service appears here.