GenAI · Multi-agent · Self-hosted modelsCJ AI Center, with CJ Group HRWrite-up October 2026 · 8 min read

A recruitment assistant that checks its own work

Tens of thousands of application essays a year, read by people at fifteen to twenty minutes each. This is the multi-agent assistant built to summarize them, ground every claim in the applicant's own sentences, prepare interview questions — and verify itself before a recruiter sees anything.

Built withPythonLangGraphLangChainvLLMGPT-OSS · self-hostedpandasRegex sentence index
APPLICATION ESSAY Summary agentEvidence agentQuestion agent Verification layertone · hallucination · fidelity regenerate until it passes
Three agents work on the same essay at once; a separate layer scores what they wrote and sends failures back. Highlights are rebuilt from the essay itself.

A high-stakes, auditable reading task

Group-wide open recruitment brings roughly 73,000 applications across three recent cycles, and recruiters read well over 60,000 essays a year. External tools rely on general datasets and quantitative scoring, so they do not reflect the group's management philosophy, talent model or values — and buying separate vendor tools for summarizing, question generation and chat adds up at this scale.

Hiring is also auditable. Free-form generation can make claims the applicant never made, and without traceable evidence an output is unusable for fair evaluation or a later appeal. The goal was a culture-fit competency model on internalized models, with outputs that are explainable and leave the decision to a person.

Three agents in parallel

The assistant runs inside the open-recruitment system. It takes each applicant's per-question essays and the talent model as criteria and runs three main agents in parallel:

Live model on an invented essay. The agent outputs — including a deliberately flawed first summary — were produced in advance and are replayed; the support check (does each summary line's wording appear in the essay, do its numbers match) runs live in your browser. Switch between the first draft and the verified version, and pick a competency to see its evidence. Open the live model on its own page ↗

How it's built

Architecture, stack and core formulation

A batch pipeline of three engines on a self-hosted open-weight model: summaries with a verify-and-regenerate loop, evidence returned as sentence indices and rebuilt from the source, and interview questions checked for coverage and policy.

1 · Ingest

Essays, aligned

Per-question essays loaded and checked so questions and answers stay paired; applicants processed in parallel.

pandas
2 · Summary engine

Generate → verify → regenerate

Job keywords per answer, then whole and per-question summaries, each through a pass/fail verifier with the failure reason fed back.

LangGraphLLM
3 · Evidence engine

Indices, not quotes

Sentences split with character offsets into a global index; one call per competency returns indices and the quote is rebuilt from the essay.

regex splitterLLM
4 · Question engine

Ask, check, filter

Question sets per behavioural area, a coverage judge with cited evidence, and a policy filter on follow-ups.

LangGraphLLM
5 · Serving

Self-hosted model

An open-weight GPT-OSS model served by vLLM behind an OpenAI-compatible API; one JSON result per applicant.

vLLMGPT-OSS
Stack
LayerTechnologyWhat it does here
OrchestrationLangGraph state graphs, LangChain output parsersParallel fan-out, verification loops, structured outputs
Model servingvLLM, OpenAI-compatible API, GPT-OSS open-weight modelSelf-hosted inference — no external commercial API
GroundingRegex sentence splitter with character offsetsEvidence returned as positions and rebuilt from the essay
RobustnessJSON repair chain, retries with backoffValid structured output at batch scale
EvaluationQA-FactEval on 380 essays; recruiter pilotFactual consistency against commercial and open-weight baselines
Core formulation
summary loop     for attempt = 1 … 5
                     s ← LLM(essay, persona, feedback)
                     if verifier(s) = PASS and |s| ≤ 1.1·L:  break
                     feedback ← verifier reason

evidence         sentences σ₁ … σₙ with offsets;     LLM(essay, competency) → { i }
                 quote_i = essay[offset_i]               never the model's own wording

coverage judge   (question, essay) → { answered ∈ {Y, N}, evidence: 2–4 indices, follow-ups }
policy filter    follow-up on a prohibited topic → delete or rewrite
  • Verification is structural. Tone and completeness by a pass/fail verifier, invented claims by a coverage judge that must cite indices, quotes by rebuilding from the source, and policy by a filter.
  • Deterministic overrides. Code, not the model, enforces consistency rules between main and follow-up questions and range-checks every evidence index.
  • Self-hosted. Running an open-weight model under vLLM kept applicant data in-house and made the QA-FactEval benchmark against commercial models possible.
In production vs in the live model
ComponentIn productionIn the live model above
AgentsThree engines on a self-hosted GPT-OSS model (vLLM)Agent outputs pre-written and replayed
VerificationPass/fail verifier loop, coverage judge, index-based evidence, policy filterWord-coverage and number checks, run live in the browser
DataReal applicant essaysOne invented essay

The verification layer

The core idea is a self-correcting verification layer kept independent of generation. Separate checks score every output on three criteria — tone (style and inappropriate language), hallucination (claims not in the document, such as an experience attributed to the wrong company) and content fidelity (key experiences left out) — and regenerate through a feedback loop until the output passes. The pipeline is input normalization, the parallel agents, verification with self-correction, and structured output parsed for the recruiter's screen.

Grounding is enforced structurally as well: evidence is returned as positions in the source and re-assembled from the original text, and follow-up questions are filtered against a policy list of topics an interviewer must not ask. The recruiter remains the judge; there is no black-box score.

The core models were internalized on an open-weight base and self-hosted rather than called through external commercial APIs, which also made the benchmarking below possible.

Results

83.45%QA-FactEval of the internalized summarizer, 380 essays (75.95% commercial small model, 68.69% open-weight)
−50%first-pass essay review time, 15 min → under 7
−66%interview preparation time, 30 min → 10

Reproducibility across repeated runs was stable (standard deviation 0.09), and job-core keyword extraction reached 85.36% versus 76.79% for an external model. In a pilot with 31 recruiters and the recruitment task force on 490 samples, positive-plus-neutral responses exceeded 94%; across 21,142 evaluator responses, 87.14% rated outputs helpful. The verification loop improved quality on 21.6% of whole-summary cases, 18.7% of per-question cases and 37.3% of interview-question cases on the issue set.

The work is filed as a patent and moved into production for the first-half 2026 open recruitment. Recruiters spend their time on judgment rather than reading, explainable outputs ease audit and appeals, and external vendor cost does not accumulate.

Limitations

About the demo and confidentiality

The applicant, essay, competencies, questions and outputs in the embedded model are invented. No applicant data, talent-model wording, prompt or policy list from the production system appears here.

Taehee Lee · Data Scientist / Applied AI Scientist, CJ AI CenterLed design and development; model internalization and benchmarking; patent inventor. Demo built on an invented essay for this site.