A personal health advisor built on validated knowledge
People who receive a gut-microbiome report want to know what it means for them and what to do next. A generic chatbot cannot answer that. This is the retrieval-augmented system behind an advisor that answers from the customer's own results and from validated guidance — launched commercially in April 2025.
Why a generic chatbot is not enough
SmileGut is a microbiome-based diagnostic and health service. Customers need personalized explanations of their results and practical recommendations. But microbiome, diet and nutrition knowledge is complex and spread across reports, diagnostic criteria, analysis documents and dietary guides, so an FAQ page or a general chatbot falls short. Health-related answers also have to be accurate, current and stable, which ruled out letting a model answer from its own memory.
What was built
- Domain data, structured. Working with microbiome, diet and nutrition experts, data extracted from health reports, diagnostic criteria, analysis documents and dietary guides was structured into data marts and a vector database for retrieval.
- Retrieval designed, not assumed. Hybrid search, re-ranking and knowledge-graph-based structures were evaluated, and the pipeline retrieves on the question together with the person's diagnostic context.
- A workflow around the model. LangGraph controls the conversation: which step runs next, how errors are avoided, how personalized answers are generated, and how answers stay stable from one session to the next.
- Hardened for operation. Prompts, guardrails and error handling were tuned for a live service rather than a demo.
Live model on a fictional report and a small library of general-wellness text. Pick a question: the report fields used, the two retrievers' scores, the fused and re-ranked passages and the guardrail all run live in your browser; the answers were written in advance and are replayed, with citation numbers bound to whatever was actually retrieved. The last question shows the guardrail path. Not medical advice. Open the live model on its own page ↗
Architecture, stack and core formulation
A LangGraph conversation graph over a knowledge graph built from the service's domain tables: conversation memory, question rewriting, graph-plus-vector retrieval and streamed answers.
Tables to documents
Report guides, a microbe dictionary, ingredient and nutrient tables, recipes and FAQ rendered into documents by templates.
Knowledge graph + vectors
Entities and relations extracted by an LLM, embedded and stored with the source chunks; rebuilds are idempotent by content hash.
Memory and routing
A LangGraph state graph loads the conversation from MongoDB; exact FAQ matches return the curated answer.
Rewrite, then search
The question is rewritten with recent turns, keywords extracted, and entities, relations and source chunks retrieved together.
Stream and save
The answer streams from the LLM; a summary of the turn is saved for the next one.
| Layer | Technology | What it does here |
|---|---|---|
| Orchestration | LangGraph StateGraph with a checkpointer | Load memory → FAQ match or rewrite → retrieve → generate → save |
| Knowledge | LightRAG knowledge graph (entities + relations), FAISS vectors | Graph-aware retrieval over validated domain knowledge |
| Models | Amazon Bedrock — Claude 3.5 Sonnet (rewrite, answer), Nova Micro (extraction, summaries), Titan Text Embeddings v2 | Cost-tiered: a small model for bulk work, a strong one per query |
| Memory | MongoDB | Conversation history and turn summaries |
| Serving | FastAPI streaming endpoint, Streamlit front end, cloud VM | Streamed answers; timing recorded per node |
| Earlier version | Self-hosted EXAONE 3.5 + BGE-m3 + FAISS with corrective / self-RAG grader nodes | Replaced after latency profiling: fewer LLM calls, streaming answers |
graph START → load_memory → ( faq_hit ? curated_answer : rewrite → retrieve → generate ) → save → END
q′ = rewrite(question, last 4 turns, diagnostic context)
retrieve(q′) = entities(q′) ∪ relations(q′) ∪ graph neighbours ∪ source chunks top-k
answer = LLM(q′, retrieved context) streamed
branches inside the answer step: on-topic · unrelated · ambiguous (ask back) · general chat- Fewer calls per question. Per-node profiling showed graders and classifiers cost more latency than they added; the graph went from about fifteen LLM calls per question to four.
- Model tiering. A small model does bulk extraction and turn summaries; a strong model handles the per-query rewrite and answer.
- Two-track memory. Prompts see turn summaries; full text is archived separately.
| Component | In production | In the live model above |
|---|---|---|
| Retrieval | LightRAG knowledge graph + dense vectors | BM25 + character n-gram retrievers fused by RRF, in the browser |
| Answer | Claude on Bedrock, streamed | Pre-written answers replayed, with citations bound to the retrieved passages |
| Scope handling | Branches inside the answer step | A rule-based guardrail on treatment questions |
| Data | Validated domain knowledge and customer reports | A fictional report and 14 general-wellness passages |
Design notes
Personal context first
The most useful answers start from the person's own numbers — a diversity index below reference, a fibre intake half the goal. The pipeline retrieves on the question together with the person's diagnostic context, so guidance is chosen for this person rather than for the question in general.
Retrieval was evaluated, not assumed
Hybrid keyword-plus-vector search, re-ranking and knowledge-graph retrieval were all evaluated. A knowledge graph built from the domain tables captures how a microbe, a nutrient, a food and a test result connect — relations that flat text chunks lose — and graph retrieval was combined with dense vector search. The live model illustrates the idea of combining two views with a keyword retriever and a character n-gram retriever fused by rank.
Stability as a requirement
For a health product, the same question should get the same substance of answer. Controlling the flow as a graph — load the conversation, rewrite the question, retrieve, generate, save — makes behaviour predictable and testable. Timing every node showed where latency went, and steps that cost more than they added were removed; unrelated, ambiguous and treatment questions are handled by explicit branches in the answer step.
Outcome
The advisor established a model that combines domain knowledge with personal health data, strengthened the customer experience of the diagnostic service, and laid a technical foundation for follow-on services in diet, probiotics, supplements and personalized health management.
Limitations
- The advisor provides general wellness information; it does not diagnose or treat, and it says so.
- Answer quality is bounded by the knowledge base; topics it does not cover should produce a careful non-answer rather than a guess.
- The live model's report, passages and answers are invented, and its retrievers are simple in-browser stand-ins for the production knowledge-graph and vector retrieval.
About the demo and confidentiality
The user, report values, knowledge passages and answers in the embedded model are invented general-wellness text. No customer data, report format, knowledge content, prompt or system detail from the production service appears here.