applied ml
MediQuery
Medical Q&A that shows its work: hybrid retrieval over 505K documents with per-answer confidence and source attribution.

about
My capstone. Most medical chatbots answer confidently and cite nothing, which is the exact failure mode that makes them unusable in practice. MediQuery retrieves with dense embeddings and BM25 in parallel, fuses the rankings with Reciprocal Rank Fusion, and attaches an explainability score built from retrieval confidence, generation confidence, and source attribution, so a wrong answer is visibly a low-confidence answer. It is backed by a 20-file pytest suite and a 6-stage CI/CD pipeline covering linting, testing, and security scanning.
505K+
documents indexed
3-way
hybrid retrieval
20
test files
stack
- FastAPI
- LangChain
- Qdrant
- ChromaDB
- Next.js
- Docker
how it works
- 1
Safety check
Every query is screened for emergency keywords, dangerous content, and known drug-interaction patterns before retrieval runs.
- 2
Hybrid retrieval
Dense search over 505,584 chunks in Qdrant Cloud (all-MiniLM-L6-v2 embeddings, from MedQuAD, MedQA, and clinical guidelines) runs alongside BM25 keyword search. Results are fused and re-ranked.
- 3
Generation
The answer is generated from the retrieved passages only, through a LangChain pipeline with a self-correcting LangGraph RAG graph.
- 4
Explainability
A confidence score built from retrieval score, entity coverage, source agreement, and answer length, plus the source passages highlighted where they were used.
- 5
Safety check
The response is screened again before it is returned.
engineering notes
Confidence you can inspect
The score is a breakdown, not one opaque number, so a user can see why an answer is low-confidence rather than being told it is.
A fully on-device Android port
No server and no internet permission: Gemma 3 1B through LiteRT-LM, hand-written BM25 + dense retrieval (ONNX Runtime) fused with RRF, and the same confidence scorer, Platt-calibrated against a 97-question on-device evaluation run.
CI that deploys
Every push runs ruff, pytest, a frontend type-check and build, gitleaks, and a Trivy scan. Pushes to main build a Docker image to GHCR and deploy the backend to Hugging Face Spaces.
known limits
- Educational use only - it does not replace clinical advice, and the app says so.