Python, LangGraph, Ollama, Chroma, Streamlit, Docker
Agentic RAG Chatbot
Creator · Proof of concept · Oct 2026
A proof of concept I designed and built alone: a chatbot that answers web frontend questions from the official documentation and computes the facts a language model would otherwise guess. General-purpose LLMs answer these questions fluently but unreliably — they invent APIs, mix up the Next.js Pages and App Router, confuse React’s useState hook with Nuxt’s useState composable, and cannot tell whether :has() works in Safari 15. Here a LangGraph workflow routes every message, splits a comparison into one search per framework and runs them in parallel, writes an answer with numbered citations, and checks it against the retrieved text before it reaches the user. It runs entirely on local hardware — Qwen3.5-4B on Ollama, multilingual E5 embeddings on the CPU, Chroma, and a Streamlit UI that shows every step — with no paid APIs.
Question (HU / EN) → analyze_request
direct → finalize_response
single → run_rag_subtask
tool → call_tool
complex → plan_subtasks → up to 5 sub-tasks in parallel (Send)
→ synthesize_answer → verify_answer (skipped for an exact tool answer)
→ re-plan what is missing (max 2) | finalize_response
RAG subgraph, inside run_rag_subtask:
rewrite_query → retrieve (Chroma + BM25, RRF) → grade_documents → build_context
Tools: check_contrast · css_specificity · browser_support
Local: Qwen3.5-4B on Ollama · multilingual-e5-small on CPU · ChromaDeep dive
Why agentic, not retrieve-then-generate
Real questions mix explanation with verification and often span frameworks, so one retrieve-then-generate pass falls short. “How do I fetch data on the server in Next.js and in Nuxt?” becomes one retrieval per framework, run in parallel and combined into a single comparison. Facts that can be computed are computed: #777777 on white is 4.47:1, just under the 4.5:1 that WCAG AA requires — a margin a model easily gets wrong. And the verification step checks each draft against the retrieved documentation and re-plans when it is not supported, which catches invented APIs before they reach the user.
A corpus pinned to exact commits
A source list fetches MDN, React, Vue, Next.js, Nuxt and the TypeScript Handbook by sparse git checkout at pinned commits, so every run indexes the same versions and no share-alike text enters the repository. Loaders turn each dialect — MDN macros, React and Next.js JSX components, Vue’s VitePress containers, Nuxt’s MDC — into plain Markdown, keep code verbatim, and leave out legacy surfaces such as the Pages Router and Nuxt Bridge. Chunks follow the H2/H3 structure, keep code blocks whole, and open with a context line of title and heading path, so “useState – React” and “useState – Nuxt” stay apart in retrieval and in citations. 1,160 pages become 18,654 chunks; a re-ingest embeds only what changed.
Hybrid retrieval, Hungarian questions
The corpus is English, the questions may be Hungarian, and multilingual embeddings alone missed those. The RAG subgraph therefore rewrites every question into an English search query first. Chroma and a SQLite FTS5 BM25 index then rank the same chunks, reciprocal-rank fusion keeps 20 candidates, and one structured-output call grades them together — one LLM call per retrieval, not one per chunk — before up to four become a context with numbered citations. The keyword query carries both the original question and the rewrite, so API names such as :has or $fetch survive the translation. That keyword ranking found the one page the vectors missed: hit@4 went from 0.90 to 1.00.
Tools that compute, a verifier that checks
Three deterministic tools answer what has an exact answer: WCAG 2.x contrast with AA and AAA verdicts, Selectors Level 4 specificity with the winning selector, and browser support looked up in MDN’s browser-compat-data at a pinned commit. When a single tool answers a question on its own, its output is the answer, verbatim — no synthesis call, no verification. Everything else goes through verify_answer, which re-plans what the draft does not support, at most twice; a verification it cannot read withholds the draft rather than claiming grounding. The planner emits typed sub-tasks as JSON instead of native tool calls, because small local models produce structured output more reliably.
A UI that shows its work
The Streamlit UI streams the run: each step appears as its node finishes, parallel steps are grouped, and every search lists its RAG steps underneath — the English query, the retrieved chunks and the ones kept. Citations link to their source pages, and the retrieved-context panel shows each numbered source with the commit it was indexed at. A missing index, an index built with other embeddings, an unreachable Ollama or a model that is not pulled is explained in the chat with the fix, and a run the user stops still closes its turn. An empty chat offers one example question per route.
Measured, not asserted
A 17-question set — searches across all six sources, comparisons, tool questions, a greeting and an off-topic question, five of them in Hungarian — scores routing and retrieval directly, and correctness and faithfulness through a local LLM judge whose verdicts were reviewed by hand. The model was chosen this way: Qwen3.5-4B with thinking off beat Qwen2.5-7B-Instruct on routing (1.00 vs 0.88) and on Hungarian answers (0.90 vs 0.80), and thinking mode made every call about ten times slower without better answers. Current scores: routing 1.00, hit@4 1.00, faithfulness 1.00, correctness 0.91. A separate 24-question holdout stays out of development and is reported as it came out — correctness 0.67, or 0.79 without three judge errors — with failures at version boundaries, on an ambiguous question, and in the answer language.
Where the time goes
The load test drives the compiled graph with 100 requests from parallel threads and reports latency per node. LLM inference is 99% of a request — about 4.7 model calls, against 20 ms of retrieval — and Ollama serves one request at a time, so going from one user to four barely moves throughput while latency queues up. One user gets a 3.9 s median and a 12.9 s p95 on a laptop GPU. The measured fix is parallel decoding slots, which doubled the 7B model’s throughput (9.5 → 20.0 requests a minute) and cut its p95 by 57% — but Ollama does not batch Qwen3.5’s hybrid architecture, so a multi-user deployment has to choose the model and the serving stack together.
One command from a fresh clone
From a fresh clone, docker compose up --build pulls the model, downloads the corpus, builds the index and serves the UI with no other step — 25 minutes the first time, mostly downloads, and 11 s on a restart. The two-stage image installs exactly what uv.lock pins, so a code change rebuilds only two small layers (7 s); it runs as a non-root user with read-only code, and Ollama is not published on the host. A scripted fake LLM and hashed embeddings keep 800+ tests offline, and CI runs ruff, the test suite and the image build on every pull request and push to main.