A recruitment assistant that checks its own work
Tens of thousands of application essays a year, read by people at fifteen to twenty minutes each. This is the multi-agent assistant built to summarize them, ground every claim in the applicant's own sentences, prepare interview questions — and verify itself before a recruiter sees anything.
A high-stakes, auditable reading task
Group-wide open recruitment brings roughly 73,000 applications across three recent cycles, and recruiters read well over 60,000 essays a year. External tools rely on general datasets and quantitative scoring, so they do not reflect the group's management philosophy, talent model or values — and buying separate vendor tools for summarizing, question generation and chat adds up at this scale.
Hiring is also auditable. Free-form generation can make claims the applicant never made, and without traceable evidence an output is unusable for fair evaluation or a later appeal. The goal was a culture-fit competency model on internalized models, with outputs that are explainable and leave the decision to a person.
Three agents in parallel
The assistant runs inside the open-recruitment system. It takes each applicant's per-question essays and the talent model as criteria and runs three main agents in parallel:
- Summary agent — whole-document and per-question summaries in a recruitment-analyst persona, extracting job-core keywords and the experiences that matter.
- Evidence agent — maps each competency in the talent model back to its supporting sentences and their positions in the original text.
- Interview-question agent — a structured probing set across eight behavioural areas: a main question, a predicted answer, follow-up and probe questions.
Live model on an invented essay. The agent outputs — including a deliberately flawed first summary — were produced in advance and are replayed; the support check (does each summary line's wording appear in the essay, do its numbers match) runs live in your browser. Switch between the first draft and the verified version, and pick a competency to see its evidence. Open the live model on its own page ↗
Architecture, stack and core formulation
A batch pipeline of three engines on a self-hosted open-weight model: summaries with a verify-and-regenerate loop, evidence returned as sentence indices and rebuilt from the source, and interview questions checked for coverage and policy.
Essays, aligned
Per-question essays loaded and checked so questions and answers stay paired; applicants processed in parallel.
Generate → verify → regenerate
Job keywords per answer, then whole and per-question summaries, each through a pass/fail verifier with the failure reason fed back.
Indices, not quotes
Sentences split with character offsets into a global index; one call per competency returns indices and the quote is rebuilt from the essay.
Ask, check, filter
Question sets per behavioural area, a coverage judge with cited evidence, and a policy filter on follow-ups.
Self-hosted model
An open-weight GPT-OSS model served by vLLM behind an OpenAI-compatible API; one JSON result per applicant.
| Layer | Technology | What it does here |
|---|---|---|
| Orchestration | LangGraph state graphs, LangChain output parsers | Parallel fan-out, verification loops, structured outputs |
| Model serving | vLLM, OpenAI-compatible API, GPT-OSS open-weight model | Self-hosted inference — no external commercial API |
| Grounding | Regex sentence splitter with character offsets | Evidence returned as positions and rebuilt from the essay |
| Robustness | JSON repair chain, retries with backoff | Valid structured output at batch scale |
| Evaluation | QA-FactEval on 380 essays; recruiter pilot | Factual consistency against commercial and open-weight baselines |
summary loop for attempt = 1 … 5
s ← LLM(essay, persona, feedback)
if verifier(s) = PASS and |s| ≤ 1.1·L: break
feedback ← verifier reason
evidence sentences σ₁ … σₙ with offsets; LLM(essay, competency) → { i }
quote_i = essay[offset_i] never the model's own wording
coverage judge (question, essay) → { answered ∈ {Y, N}, evidence: 2–4 indices, follow-ups }
policy filter follow-up on a prohibited topic → delete or rewrite- Verification is structural. Tone and completeness by a pass/fail verifier, invented claims by a coverage judge that must cite indices, quotes by rebuilding from the source, and policy by a filter.
- Deterministic overrides. Code, not the model, enforces consistency rules between main and follow-up questions and range-checks every evidence index.
- Self-hosted. Running an open-weight model under vLLM kept applicant data in-house and made the QA-FactEval benchmark against commercial models possible.
| Component | In production | In the live model above |
|---|---|---|
| Agents | Three engines on a self-hosted GPT-OSS model (vLLM) | Agent outputs pre-written and replayed |
| Verification | Pass/fail verifier loop, coverage judge, index-based evidence, policy filter | Word-coverage and number checks, run live in the browser |
| Data | Real applicant essays | One invented essay |
The verification layer
The core idea is a self-correcting verification layer kept independent of generation. Separate checks score every output on three criteria — tone (style and inappropriate language), hallucination (claims not in the document, such as an experience attributed to the wrong company) and content fidelity (key experiences left out) — and regenerate through a feedback loop until the output passes. The pipeline is input normalization, the parallel agents, verification with self-correction, and structured output parsed for the recruiter's screen.
Grounding is enforced structurally as well: evidence is returned as positions in the source and re-assembled from the original text, and follow-up questions are filtered against a policy list of topics an interviewer must not ask. The recruiter remains the judge; there is no black-box score.
The core models were internalized on an open-weight base and self-hosted rather than called through external commercial APIs, which also made the benchmarking below possible.
Results
Reproducibility across repeated runs was stable (standard deviation 0.09), and job-core keyword extraction reached 85.36% versus 76.79% for an external model. In a pilot with 31 recruiters and the recruitment task force on 490 samples, positive-plus-neutral responses exceeded 94%; across 21,142 evaluator responses, 87.14% rated outputs helpful. The verification loop improved quality on 21.6% of whole-summary cases, 18.7% of per-question cases and 37.3% of interview-question cases on the issue set.
The work is filed as a patent and moved into production for the first-half 2026 open recruitment. Recruiters spend their time on judgment rather than reading, explainable outputs ease audit and appeals, and external vendor cost does not accumulate.
Limitations
- QA-FactEval measures factual consistency of summaries, not the quality of a hiring decision.
- Grounding checks catch claims unsupported by the essay; they cannot tell whether the essay itself is truthful.
- The live model's support check is a simple word-coverage and number test standing in for the production verifiers, and its agent outputs are pre-written for one invented essay.
About the demo and confidentiality
The applicant, essay, competencies, questions and outputs in the embedded model are invented. No applicant data, talent-model wording, prompt or policy list from the production system appears here.