Welcome to Ruthvik Nath Bandari's portfolio

All projects
Clinical NLP · Information Retrieval

Blood Report Analyzer and Recommendation Engine

AAI 6620 NLP final project · team with Om Patel and Yash Jain

A course prototype for extracting biomarkers from blood reports. I built the extraction, NER, retrieval, and training stages.

  • PubMedBERT
  • PyMuPDF
  • FastAPI
  • FAISS
  • TF-IDF
  • SLURM
Abstract document with highlighted text spans and a droplet, on a red fieldIllustration

Highlights

  • PubMedBERT biomedical NER (biomarker, value, unit, reference-range) at entity-token F1 0.8902 +/- 0.0034 over 5-fold cross-validation, measured on synthetic, OCR-noise-augmented data rather than on clinical documents
  • Chunked inference with overlap and span aggregation, token accuracy 0.923, precision 0.956 and recall 0.833
  • Document extraction via PyMuPDF with an OCR fallback behind a router
  • Hybrid recommendation retrieval designed as TF-IDF plus MiniLM/FAISS semantic search with 0.6/0.4 score fusion; the recorded run fell back to the lexical branch over an 8-document corpus
  • NER training on Northeastern's Explorer cluster (SLURM, NVIDIA H200)

How it was built

Our course pipeline reads a blood-report PDF or image. My share covered extraction, entity recognition, retrieval and training. A router picks PyMuPDF text extraction or an OCR fallback. I fine-tuned a PubMedBERT token classifier to tag biomarker, value, unit and reference range, using synthetic training reports with OCR-style noise, on Explorer H200 GPUs. Five-fold cross-validation gave entity-token F1 of 0.8902 +/- 0.0034 on that synthetic data. I also built hybrid retrieval for recommendations; teammates built interpretation, templating, the API and the interface. Limits: no clinical documents; the recorded retrieval run used only the lexical branch over eight documents.