Client-side · direct prompting vs. RAG
Same model,
same question, one document.
Upload a PDF and ask it something. A small model (Qwen2.5-0.5B-Instruct) answers twice: once from its own parametric knowledge alone, and once after retrieving the most relevant passages from your actual PDF and reading them first. Both runs happen in this tab — the model, the PDF text extraction, and the retrieval never leave your device.
Step 1
Upload a PDF
Anything with real text works best — a paper, a manual, a report. Scanned image-only PDFs won't have extractable text.
No document loaded yet.
Step 2
Ask it something
Loading the on-device model…
Direct prompting
No document context — just the model's own knowledge.
Retrieval-augmented
Grounded in the passages retrieved below, from your PDF.
Retrieved passages
The actual top-scoring chunks of your PDF for this question (BM25), in the order the RAG run read them.
- None yet.
How this works
How it works.
Retrieval is real BM25 — the same ranking function search engines used before embeddings, run over the words your PDF actually contains, chunked into ~140-word overlapping windows. It's lexical, not semantic: it matches on shared words, not meaning, so a question phrased very differently from the document's own wording will retrieve less relevant passages. The original project used Supermemory's embedding-based retrieval server-side; this port trades that for something that needs no server and no second model download, so it stays fast in a browser tab.
Generation is a real small model — Qwen2.5-0.5B-Instruct, via WebLLM, running on your GPU through WebGPU (falling back to a slower CPU path if that's unavailable). It downloads once (a few hundred MB, cached after) and every answer after that streams entirely on-device.
Source: amanyagami/SLM-based-QA.
A model this small will still get things wrong in both modes — that is expected. The comparison is meant to show is the shape of the difference: direct prompting has to answer from whatever it already knows, right or wrong, with nothing to check itself against, while the RAG run can only work with what's actually printed in the passages below it — including admitting the document doesn't say, if it doesn't.
If a question is answered wrong in the RAG panel, look at the retrieved passages: either BM25 missed the right chunk (a retrieval failure — most fixable), or the chunk is there but the small model still lost the thread reading it (a generation failure — orthogonal to retrieval quality).