Welcome to Ruthvik Nath Bandari's portfolio

My roles across applied AI, research, and healthcare innovation.

  1. 2026

    —

    Lahey Clinic

    Current

    Student Intern, Innovation Hub

    Beth Israel Lahey Health

    Burlington, MA (remote)

    Lung cancer screening

    Building a lung cancer screening eligibility prototype using synthetic clinical data, with a focus on low-dose CT screening and patient outreach.

    Explore the work at Beth Israel Lahey Health, Student Intern, Innovation Hub · includes planned deliverables

    Leading research, workflow analysis, and quality auditing within a three-person team; coordinating baseline research and identification of workflow gaps.

    Scope

    • Develop synthetic-data methods for identifying patients eligible for low-dose CT lung cancer screening.
    • Analyze and organize the clinical data elements that bear on screening eligibility, outreach, and completion.

    Baseline phase

    • Workflow research, analysis of competing solutions, and identification of gaps.
    • Own research, workflow analysis, quality audits, and debugging for the three-person team.

    Working practice

    • Pull-request review on GitHub, team coordination in Microsoft Teams, and decision records in Notion.

    Planned work

    Sep 2026 – Nov 2026

    The engagement plan calls for screening-rule implementation, prototype validation, a total-cost-of-ownership assessment, and an executive handoff. These are planned deliverables for the ongoing internship.

    Tools

    • Python
    • Synthetic clinical data
    • Workflow analysis
    • GitHub
    • Microsoft Teams
    • Notion
  2. —

    Northeastern University

    Research Assistant

    CHEW · Making Universities Work

    Boston, MA

    Higher-education compliance software

    Built, from scratch, a multi-tenant higher-education civil-rights compliance analysis platform end to end: React 19 and TypeScript frontend, FastAPI backend, PostgreSQL with pgvector, hybrid retrieval, Clerk authentication, and production deployment.

    Explore the work at CHEW · Making Universities Work, Research Assistant

    Frontend

    • 18 route modules and about 36,700 lines of TypeScript on React 19, Vite, TanStack Router and Query, Zustand, Zod, Tailwind CSS 4, shadcn/ui, Radix, and motion.
    • Zod-validated URL search params make every view state a shareable link; includes a side-by-side deep compare mode and audience-specific view modes.
    • Charts are hand-built SVG rather than a chart library.
    • Design system with all colour as CSS custom properties, light and dark themes on every surface, and 42 recorded design decisions.

    Backend and data

    • FastAPI service with 34 endpoints across 10 route modules and 16 service modules, where services never import the web framework.
    • Clerk authentication with RS256 and JWKS verification and a pinned algorithm, role-based admin access, rate limiting, RFC 9457 problem+json errors, request and statement timeouts, and 8 startup health checks that fail the boot.
    • PostgreSQL 17 with pgvector: 7 schemas, 95 tables, 16 Alembic migrations, and about 4,100 lines of SQL.
    • Multi-tenant isolation with 64 row-level-security policies using FORCE ROW LEVEL SECURITY, an API role without BYPASSRLS, a SECURITY DEFINER tenant-context function, and 13 invariants enforced by 8 CI views.
    • Tenant isolation is layered four ways (namespace assignment, separate index files, deny-by-default searchers, row-level security) and checked by a script that exits non-zero on any leak.

    Retrieval

    • Hybrid retrieval: a legal-hierarchy chunking engine with offset-traceable chunks, BGE-M3 1024-dimension embeddings with a pinned revision and integrity guards, an in-repo BM25 index, reciprocal rank fusion, and a cross-encoder reranker.
    • Document text extracted with MinerU, with OCR retry and SHA-256 document identity.
    • Schema-validated intake so new institutional records onboard without component code changes, plus a publish-blocking validation pipeline over a unified JSON schema.

    Quality and accessibility

    • Audited per route against WCAG 2.1 AA with axe-core, at 8 viewport widths and in both themes.
    • Layered automated testing: frontend unit tests with Vitest, Playwright end-to-end specs across routes, viewports, and themes, plus backend, retrieval, and corpus suites.
    • Tracked defects each paired with a regression test.

    Deployment and tooling

    • Deployed to production with Vercel configuration, Redis, Bun, oxlint, and TypeScript typechecking.
    • Password-gateable auth server and a SLURM-compatible startup script for shared and HPC environments.
    • Consolidated a codebase fragmented across several workspaces into one repository with git history preserved.
    • Operational documentation maintained alongside the code.

    Tools

    • React
    • TypeScript
    • Vite
    • TanStack
    • Tailwind CSS
    • FastAPI
    • PostgreSQL
    • pgvector
    • Alembic
    • Clerk
    • Redis
    • BGE-M3
    • Vitest
    • Playwright
    • axe-core
    • Bun
    • Vercel
    • SLURM
  3. —

    Northeastern University

    Current

    Research Assistant

    Dr. Rominder Singh · Regulatory Affairs

    Boston, MA

    Regulatory data and AI tutoring

    Built a regulatory-data preparation pipeline that cleans, validates, deduplicates, and exports source documents for retrieval-augmented generation.

    Explore the work at Dr. Rominder Singh · Regulatory Affairs, Research Assistant

    Developing RegRAGA Tutor, a text-and-voice prototype for Regulatory Affairs, with retrieval, answer verification, and course-aware responses.

    Zebrafish regulatory data pipeline

    Feb 2026 – May 2026

    Built solo. A recorded FDA/EMA export accepted 8,250 documents and produced 38,447 chunks. Its registry lists 202 regulatory authority entries across 199 national ISO codes; that registry is broader than the acquired data in this recorded export.

    Coverage

    • Registry configuration of 110 per-authority scrape configs, 9 regional bodies, 4 APIs, and 88 gap-filled jurisdictions. This describes registry breadth, not acquired data.

    Pipeline

    • Five stages: clean, normalize, enrich, validate, and export to JSONL, Markdown, and chunked JSONL.
    • Deterministic weighted-regex classifier with 129 signals and traceable decisions.
    • SHA-256 and SimHash-LSH deduplication, and 512-token chunks with 64-token overlap and context prefixes.

    Resilience and quality

    • Per-domain politeness delays, backoff, robots.txt handling, per-document error isolation, checkpoint resume, and snapshot versioning.
    • Python codebase of about 28,000 lines under strict mypy.

    RegRAGA Tutor

    Jun 2026 – Present

    A working prototype co-built with Om Patel, who owns the Canvas and LTI registration, while the running system is Ruthvik's. It is a text-and-voice, Canvas-embeddable RAG tutor for Regulatory Affairs built on FastAPI, pgvector, and self-hosted models. Canvas integration, FERPA review, and deployment remain work in progress, and LTI/OIDC currently uses mocks.

    Architecture

    • Hexagonal ports-and-adapters design with interfaces for the retriever, LLM, verifier, scope gate, transcriber, and synthesizer; runs with no GPU and no database.
    • Five-stage answer pipeline with five short-circuits that skip the LLM entirely.
    • Courses are data, so two courses share one React interface.

    Models and retrieval

    • Self-hosted Qwen behind an OpenAI-compatible adapter, bge-small embeddings, cross-encoder reranking, and an answer verifier combining an LLM judge with NLI.
    • Boot guard that refuses a non-loopback inference host.
    • CPU-only voice input that transcribes and discards the audio.

    Measured improvements (prototype measurements)

    • Groundedness gate: 15 of 18 answers passing the gate under a pre-registered decision rule, up from 7 of 18, with zero false positives (prototype measurement).
    • Typo-query top-1 retrieval: 19 of 25, up from 11 of 25.
    • First-request retrieval latency cut from 18.6 s to 27.6 ms with a pre-warm pass, at the cost of about 15 s of boot time.
    • Two experiments with negative results were kept and shipped off by default.

    Scale and operations

    • About 52,000 lines of Python with Prometheus and OpenTelemetry instrumentation, and about 34,000 lines of documentation.

    Tools

    • Python
    • RAG
    • FastAPI
    • pgvector
    • Qwen
    • BGE
    • React
    • Prometheus
    • OpenTelemetry
  4. 2023

    —

    Heuristers Technology Solutions

    Machine Learning Intern

    Sathyabama Incubation Center

    Chennai, India

    Applied machine learning

    Built classification and regression models in scikit-learn across 13 notebooks, including linear regression, Random Forest tasks, and a diabetes-prediction exercise.

    Explore the work at Sathyabama Incubation Center, Machine Learning Intern

    Practiced data preparation, visualization, train/test splitting, and model evaluation with Pandas and Matplotlib.

    Notebooks

    • Linear regression with scikit-learn.
    • Four Matplotlib and Pandas visualization notebooks: bar, scatter, line, and combined charts.
    • Eight Random Forest classification tasks and a Random Forest diabetes-prediction project on the Pima dataset with an 80/20 split.
    • One standalone script. This was introductory work, and no headline accuracy is claimed.

    Tools

    • Python
    • scikit-learn
    • Pandas
    • NumPy
    • Matplotlib
    • Seaborn
    • Jupyter
  5. 2022

    —

    Leadership

    ACM Student Chapter

    Chapter Lead

    Sathyabama University

    Chennai, India

    Technical workshops and mentoring

    Ran hands-on workshops on Git, Python, and applied AI/ML, including computer vision, deep learning, and neural networks.

    Explore the work at Sathyabama University, Chapter Lead

    Mentored junior members from foundational concepts through building working AI models.

    Workshop topics

    • Git, Python, computer vision, deep learning, and neural networks.

    Tools

    • Teaching
    • Python
    • Git
    • AI/ML

More of my engineering work. Explore selected projects