A small, working copy of an in-house RAG platform's chat flow: conversational query rewriting, multi-query expansion with a hypothetical answer (HyDE), routing between four document strategies, Dynamic Alpha Tuning hybrid search, reranking, a streamed answer that cites document and page, and a format check for market briefings. The documents belong to Larkmoor Creamery, an invented ice-cream maker.
Left: chat with 14 documents of an invented company, in English or Korean — answers cite the document and page. Right: the eight steps of the pipeline, and the mix chosen for this question: orange is keyword search (BM25), blue is vector search. An LLM grades each side's top hit, and the grades set α.
Pick up to 4 to unlock the two selected-document strategies: an instruction about them (“summarise these”) uses the whole documents; a question searches inside them only.
Ask a question to run the pipeline.
The grades of the two top hits will set α here.
Details of the last turn appear here.
Production ran every model in-house — gpt-oss and Gemma LLMs and the Qwen3 embedding and reranker models, served with vLLM — so no document, question or answer left the company network. This copy runs the same LangGraph flow on Gemini: gemini-3.8-flash for the rewrite and the answer, gemini-3.5-flash-lite for expansion, routing, the α grade, reranking (an LLM relevance call standing in for Qwen3-Reranker) and the briefing format check, and gemini-embedding-2 for vectors.
The 14 documents, their codes and every figure are invented. No web search, live market data, OCR or DRM step runs here — briefings use the two fictional market notes only. Questions are limited per visitor and per day, conversations expire after 30 minutes, and nothing is stored.