Welcome to Ruthvik Nath Bandari's portfolio

All projects
Data Engineering · NLPPreprint

Research Intelligence Pipeline

TechRxiv (IEEE) preprint · co-first author · team project

A multi-source research pipeline connected to a TechRxiv preprint. I contributed document clustering and LinkedIn collection.

  • Python
  • scikit-learn
  • K-means
  • TF-IDF
  • NLTK
Abstract Y-shaped antibody among scattered points, on an aqua fieldIllustration

Highlights

  • Built the clustering stage: TF-IDF at 500 max features with unigrams and bigrams, document-frequency filters, and cluster-count selection by silhouette scoring across 2 to 20 configurations
  • Clustering converged on 18 clusters over 5,028 scientific items, selected by silhouette score
  • Core analysis completed in 343 seconds within a 33-minute total session
  • Built the LinkedIn collection feeding a five-source corpus, and documented free-tier API quota exhaustion as a measured collection limit rather than hiding it

How it was built

This was a team study of whether AI-built pipelines can match hand-built ones; I worked on the research-aggregation pipeline for antibody machine-learning literature. I built the clustering stage, which vectorises text with TF-IDF (500 features, unigrams and bigrams) and picks the number of clusters by silhouette score across 2 to 20 options. It converged on 18 clusters over 5,028 items from five sources. I also built the LinkedIn collection, and we recorded where the free-tier search quota ran out as a measured limit. I am a co-first author of the TechRxiv preprint. Limits: not peer reviewed.

Preprint. TechRxiv (IEEE). Preprint v1, posted 6 Feb 2026 under CC-BY 4.0. Not peer reviewed; no journal or conference acceptance is recorded as of 1 Oct 2026.