GenAI · On-premise RAG · LangGraphCJ AI CenterOctober 2026 · 9 min read

In-house LLM/RAG platform: a LangGraph agent, Dynamic Alpha Tuning hybrid search and every model on-premise

How we moved every model of a document-chat platform onto the company's own GPUs, let a LangGraph agent decide whether and where to search, and let each question set its own balance between keyword and vector retrieval.

Built withLangGraphvLLMgpt-ossQwen3 Embedding & RerankerRolmOCRPostgreSQL + pgvectorText Embeddings InferenceFastAPINext.jsSurfSense
COMPANY NETWORK POLICY TR-07 keyword vector α = 0.3 CITED ANSWER Overseas trips use the regional rate in Annex B [TR-07, p.3] gpt-ossQwen3RolmOCRvLLM all on in-house GPUs external model APIs
Everything happens inside the company network: locked and scanned documents are read on in-house GPUs, two retrievers are mixed by a weight that each question sets for itself, and the answer cites its page. Invented document, for illustration.

Much of what a large company knows lives in documents: policies with clause numbers, operating playbooks, product specification sheets, meeting notes. Many are DRM-protected office files or scanned PDFs, the working language is Korean, and employees wanted to ask them questions and get answers that show where each claim came from. The documents are confidential, and so are the questions and the answers; none of it could go to an external LLM, embedding or OCR API.

Open-source RAG chat applications solve the general problem, but not this one. They assume external model APIs. Their keyword search is set up for English — an English full-text configuration stems words and drops tokens it does not recognise, which quietly breaks Korean. And they apply one retrieval recipe to every question, although clause numbers, product codes and figures need exact keyword matching while conceptual questions need semantic search. The goal was a platform on which every model runs in-house, which reads DRM-protected and scanned files, answers in the language of the question with page-level citations and adapts its retrieval to each question — and which can also draft market briefings in a fixed house format.

I designed the platform and led its development as architecture and development lead, extending the open-source SurfSense framework: we rebuilt the chat agent, replaced the whole model stack with self-hosted models, and added OCR ingestion, DRM integration, web search, market data, Korean support and a redesigned interface. It runs in production as an in-house generative-AI service with no external model API in the path. This post covers the self-hosted model stack, a LangGraph agent that decides whether and where to search, Dynamic Alpha Tuning hybrid search, and the pipeline that turns locked and scanned files into citable chunks.

How it worksOne invented question through the platform: documents ingested on in-house GPUs, the question rewritten, expanded and routed by a LangGraph agent, a hybrid search whose weight the question sets through Dynamic Alpha Tuning, and a reranked answer with citations.
One invented question through the platform: documents ingested on in-house GPUs, the question rewritten, expanded and routed by a LangGraph agent, a hybrid search whose weight the question sets through Dynamic Alpha Tuning, and a reranked answer with citations.

Every model in-house: the vLLM serving stack

Removing external APIs meant replacing every model a RAG system calls — chat, embedding, reranking and OCR — with open-weight models on one in-house GPU server. vLLM, the open-source inference server, makes this practical: it pages the attention cache in GPU memory, batches requests continuously as they arrive, and exposes an OpenAI-compatible API, so code written for a hosted model calls a self-hosted one unchanged.

JobModelServed with
Every chat step: rewriting, routing, grading, answeringgpt-oss-120bvLLM
Secondary model, selectable through the APIgpt-oss-20bvLLM
A summary of each document at ingestionLong-context Gemma modelvLLM
Embeddings (1,024 dimensions)Qwen3-Embedding-0.6BText Embeddings Inference
RerankingQwen3-Reranker-0.6BvLLM
OCR of page imagesRolmOCRvLLM

The models share the server's GPUs; the small embedding and reranking models sit beside the language models, the embedding model on Hugging Face's Text Embeddings Inference server. Each user's model roles — fast, strategic, long-context — are set automatically at sign-up, and a LiteLLM client layer addresses every model the same way.

Around the models sit a FastAPI backend, a Next.js front end that streams the agent's progress and then the answer with its sources, and one PostgreSQL database with pgvector that holds both search indexes — HNSW for vectors, GIN on a generated tsvector column for full text — so a hybrid query never leaves one system. Because every model already sits behind an OpenAI-compatible endpoint, the same stack is also offered to other internal systems through an API with API-key authentication.

LangGraph describes an LLM application as a graph: nodes are steps, edges — including conditional ones — decide what runs next, and a typed state object travels from node to node. The chat flow is eight nodes sharing one state.

One turn of the agent. The router sends each question down one of four branches; green tabs mark the rule-based fallback on every step that must return JSON.
One turn of the agent. The router sends each question down one of four branches; green tabs mark the rule-based fallback on every step that must return JSON.

reformulate_query rewrites a follow-up into a standalone question — "And for overseas trips?" becomes a question that names the policy — then expands it for search: up to three semantic queries, two keyword queries in web-search syntax, and a HyDE passage, a short hypothetical answer without specific numbers or file names, embedded as an extra vector query because text shaped like an answer tends to sit near real answers. Optional web search runs here. check_finance_query is plain rules: market questions trigger market-data look-ups.

route_document_strategy makes the decision that matters most. An LLM picks one of four strategies and returns it as JSON with a reason and a confidence:

The default is simple: when unsure, search. The router also switches on report mode for market-research requests.

Both search branches end in rerank_documents, where Qwen3-Reranker reorders up to 20 candidates against the original and the standalone question. generate_answer receives the top ten chunks as source blocks with title, page and section, plus any web results, market data and recent turns, and cites documents by title and web pages by site, in the language of the question. In report mode, validate_report_format fixes line breaks and structure without changing content. A document question normally costs five LLM calls — rewrite, expansion, strategy, α judgment and answer — and up to seven with web search and report mode.

Every step that has to return JSON has a rule-based fallback: the original question, a duplicated query, a rule-chosen strategy, α = 0.5. If the reranker is unreachable, the fused order stands. A malformed output from an open-weight model costs some answer quality, never the turn.

Dynamic Alpha Tuning: letting each question set its own hybrid mix

Hybrid search combines two retrievers with opposite strengths. Keyword search matches exact tokens — a clause number, a product code, a figure — and misses paraphrases; vector search matches meaning and blurs exact identifiers. Most systems fuse them with a fixed weight α, a compromise for both kinds of question. We used an approach called Dynamic Alpha Tuning (DAT), in which every question sets its own weight.

Dynamic Alpha Tuning on an invented question: two retrievers, one grading call, the α rule and min-max fusion — and how the ranking would differ at a fixed α = 0.5.
Dynamic Alpha Tuning on an invented question: two retrievers, one grading call, the α rule and min-max fusion — and how the ranking would differ at a fixed α = 0.5.

The keyword side is PostgreSQL full-text search: each keyword query is parsed with websearch_to_tsquery and ranked with ts_rank_cd. This is where Korean had broken, and we switched the tsvector configuration from english to the language-neutral simple, which keeps every token instead of applying English stemming and stop-word rules.

The vector side embeds each semantic query and the HyDE passage with Qwen3 and searches by cosine similarity in pgvector.

The weight. One short LLM call reads the opening of the top keyword hit and the top vector hit and grades each from 0 to 5 on directness, completeness, accuracy and context. A rule turns the grades into α: both 0 gives 0.5; a vector 5 against a lower keyword grade gives 1.0, and the mirror case 0.0; otherwise α is the vector grade's share of the two, rounded to 0.1. A question whose keyword hit is the very clause it names leans toward keywords; a conceptual one leans toward vectors.

The fusion. Full-text ranks and cosine similarities live on different scales, so each side is min-max normalised to 0–1, and a candidate scores α times its vector score plus 1 − α times its keyword score; a chunk found by only one side scores 0 on the other. The top ten go to the reranker, and α is cached for the question.

The appeal is the price: each retriever is most confident at its top hit, so grading those two shows which one understood the question, for one short call. Inside picked documents the candidates are already few, so that branch simply averages a vector top 20 and a keyword top 20.

From DRM-protected scans to searchable chunks

Nothing can be cited until it has been read, and much of this corpus resists reading. Ingestion takes every format down one path:

  1. DRM-protected uploads are decrypted through the in-house DRM decryption service.
  2. Office, HTML and text files are converted to PDF, with headless LibreOffice among the converters.
  3. Every page is rendered to a 200-DPI image with PyMuPDF.
  4. RolmOCR, a multimodal OCR model, reads the page images in parallel batches.
  5. The long-context Gemma model summarises each document.
  6. The text is cut into fixed-size chunks with chonkie, embedded and stored in pgvector with page number, section path and chunk index.

Because scans, slides and office files all become page images, they take one path, and page numbers survive to the end — which is what makes page-level citations possible. Content hashes keep a document from being ingested twice.

Briefings in a house format

In report mode, web search runs through Perplexity and Tavily in parallel, with page text enriched by trafilatura, and rules pull market data through yfinance — indices, exchange rates, interest rates, commodities and sector ETFs. The briefing follows a fixed structure — title, one-line implication, background and current situation, market reaction, assessment and outlook, takeaways — in a terse report style, and the format check tidies it.

Results

The platform runs in production as an in-house generative-AI service. Employees upload documents, DRM-protected and scanned files included, and get answers in Korean with page-level citations, while documents, questions and answers stay on the company network. Running every model in-house removed the dependence on external model APIs; Korean keyword retrieval works because the full-text configuration no longer drops Korean tokens; and an OpenAI-compatible endpoint gives other internal systems the same self-hosted model stack.

It is operated as a service, with containerised deployment, health monitoring with Teams alerts, daily database backups and a rebuild-and-handover guide with a verification checklist.

Try it

The live model below is a small, working copy of the chat flow over fictional company documents — policies with clause numbers, playbooks, spec sheets, meeting notes, a few documents in Korean and two market notes. Ask a question, or pick documents first, and follow the agent: the rewritten question and queries, the chosen strategy, the two graded top hits and the α they set, the fused list beside a fixed α = 0.5, and the cited answer. Ask about a clause number, then a conceptual question, and watch α move.

Live model on fictional documents, pre-split into page-tagged chunks. The flow is a LangGraph graph: rewrite, multi-query and HyDE, routing among the four strategies (picked documents enable the two selected-document ones), Dynamic Alpha Tuning hybrid search — BM25-style keyword scoring over language-neutral tokens plus cosine vector search, an LLM grade of the two top hits, the α rule and min-max fusion, shown beside a fixed α = 0.5 — then reranking, a streamed cited answer and, in report mode, a format check. Production runs gpt-oss, Qwen3 and RolmOCR on in-house GPUs; this demo uses gemini-3.8-flash for rewriting and answers, gemini-3.5-flash-lite for expansion, routing, grading, reranking (an LLM relevance call stands in for Qwen3-Reranker) and the format check, and gemini-embedding-2 for vectors. There is no web search, live market data, OCR or DRM step: the documents count as already ingested, and report mode draws only on the two fictional market notes. Questions are limited per visitor and per day, and nothing is stored. Open the live model on its own page ↗

What we learned

Give every JSON step a way out. An open-weight model served in-house will occasionally return output that does not parse. With a rule-based fallback on each such step, a bad output lowers the quality of one answer instead of stopping the conversation.

Look at the tokens before tuning the ranking. Korean keyword retrieval was not a ranking problem. The English configuration altered or dropped Korean tokens before any ranking happened, and a language-neutral configuration fixed it without a new model or a new kind of index.

A cheap signal is enough to adapt retrieval. Grading only the two top hits costs one short call per question, yet it lets a clause-number question lean on keywords and a conceptual one on vectors — which no fixed weight can do for both.

Own the stack behind a standard interface. With every model behind an OpenAI-compatible endpoint, changing a model is a serving change rather than an application change, and other internal systems can use the same self-hosted models through one API.

Limitations

The documents, questions, grades, scores and answers in the figures and the live model are invented. No internal product name, document, prompt, keyword list, server or network detail, DRM configuration or usage figure from the production platform appears in this post.

Taehee Lee · Data Scientist / Applied AI Scientist, Technical Lead, CJ AI CenterArchitecture and development lead — overall system design, and the agent, retrieval, model serving and ingestion. Live demo built on fictional documents for this site.