Blood Report Analyzer and Recommendation Engine
AAI 6620 NLP final project · team with Om Patel and Yash Jain
A course prototype for extracting biomarkers from blood reports. I built the extraction, NER, retrieval, and training stages.
IllustrationHighlights
- PubMedBERT biomedical NER (biomarker, value, unit, reference-range) at entity-token F1 0.8902 +/- 0.0034 over 5-fold cross-validation, measured on synthetic, OCR-noise-augmented data rather than on clinical documents
- Chunked inference with overlap and span aggregation, token accuracy 0.923, precision 0.956 and recall 0.833
- Document extraction via PyMuPDF with an OCR fallback behind a router
- Hybrid recommendation retrieval designed as TF-IDF plus MiniLM/FAISS semantic search with 0.6/0.4 score fusion; the recorded run fell back to the lexical branch over an 8-document corpus
- NER training on Northeastern's Explorer cluster (SLURM, NVIDIA H200)
How it was built
Our course pipeline reads a blood-report PDF or image. My share covered extraction, entity recognition, retrieval and training. A router picks PyMuPDF text extraction or an OCR fallback. I fine-tuned a PubMedBERT token classifier to tag biomarker, value, unit and reference range, using synthetic training reports with OCR-style noise, on Explorer H200 GPUs. Five-fold cross-validation gave entity-token F1 of 0.8902 +/- 0.0034 on that synthetic data. I also built hybrid retrieval for recommendations; teammates built interpretation, templating, the API and the interface. Limits: no clinical documents; the recorded retrieval run used only the lexical branch over eight documents.