Research Intelligence Pipeline
TechRxiv (IEEE) preprint · co-first author · team project
A multi-source research pipeline connected to a TechRxiv preprint. I contributed document clustering and LinkedIn collection.
IllustrationHighlights
- Built the clustering stage: TF-IDF at 500 max features with unigrams and bigrams, document-frequency filters, and cluster-count selection by silhouette scoring across 2 to 20 configurations
- Clustering converged on 18 clusters over 5,028 scientific items, selected by silhouette score
- Core analysis completed in 343 seconds within a 33-minute total session
- Built the LinkedIn collection feeding a five-source corpus, and documented free-tier API quota exhaustion as a measured collection limit rather than hiding it
How it was built
This was a team study of whether AI-built pipelines can match hand-built ones; I worked on the research-aggregation pipeline for antibody machine-learning literature. I built the clustering stage, which vectorises text with TF-IDF (500 features, unigrams and bigrams) and picks the number of clusters by silhouette score across 2 to 20 options. It converged on 18 clusters over 5,028 items from five sources. I also built the LinkedIn collection, and we recorded where the free-tier search quota ran out as a measured limit. I am a co-first author of the TechRxiv preprint. Limits: not peer reviewed.
Preprint. TechRxiv (IEEE). Preprint v1, posted 6 Feb 2026 under CC-BY 4.0. Not peer reviewed; no journal or conference acceptance is recorded as of 1 Oct 2026.