Skip to content
projects

applied ml

MediQuery

Medical Q&A that shows its work: hybrid retrieval over 505K documents with per-answer confidence and source attribution.

MediQuery interface

about

My capstone. Most medical chatbots answer confidently and cite nothing, which is the exact failure mode that makes them unusable in practice. MediQuery retrieves with dense embeddings and BM25 in parallel, fuses the rankings with Reciprocal Rank Fusion, and attaches an explainability score built from retrieval confidence, generation confidence, and source attribution, so a wrong answer is visibly a low-confidence answer. It is backed by a 20-file pytest suite and a 6-stage CI/CD pipeline covering linting, testing, and security scanning.

505K+

documents indexed

3-way

hybrid retrieval

20

test files

stack

  • FastAPI
  • LangChain
  • Qdrant
  • ChromaDB
  • Next.js
  • Docker

how it works

  1. 1

    Safety check

    Every query is screened for emergency keywords, dangerous content, and known drug-interaction patterns before retrieval runs.

  2. 2

    Hybrid retrieval

    Dense search over 505,584 chunks in Qdrant Cloud (all-MiniLM-L6-v2 embeddings, from MedQuAD, MedQA, and clinical guidelines) runs alongside BM25 keyword search. Results are fused and re-ranked.

  3. 3

    Generation

    The answer is generated from the retrieved passages only, through a LangChain pipeline with a self-correcting LangGraph RAG graph.

  4. 4

    Explainability

    A confidence score built from retrieval score, entity coverage, source agreement, and answer length, plus the source passages highlighted where they were used.

  5. 5

    Safety check

    The response is screened again before it is returned.

engineering notes

Confidence you can inspect

The score is a breakdown, not one opaque number, so a user can see why an answer is low-confidence rather than being told it is.

A fully on-device Android port

No server and no internet permission: Gemma 3 1B through LiteRT-LM, hand-written BM25 + dense retrieval (ONNX Runtime) fused with RRF, and the same confidence scorer, Platt-calibrated against a 97-question on-device evaluation run.

CI that deploys

Every push runs ruff, pytest, a frontend type-check and build, gitleaks, and a Trivy scan. Pushes to main build a Docker image to GHCR and deploy the backend to Hugging Face Spaces.

known limits

  • Educational use only - it does not replace clinical advice, and the app says so.