DiaFoot.AI
Research project · training-data composition study
My contribution
DiaFoot.AI is my repository and my experiment. I built the data pipeline, the split and leakage tooling, the training and evaluation code, and ran the 75-model study. The related abstract is first-authored with Prof. Mohammad Eslami of Harvard Medical School.

U-Net++ DFU Dice for five training-set compositions, one slice of the 75-model study (5 compositions x 3 architectures x 5 folds, single seed). Points are 5-fold CV means; whiskers are fold standard deviation, not a confidence interval. Every model was scored on one fixed test split of 1,215 foot photographs, with DFU Dice measured on its 318 ulcer images. DFU + healthy beat DFU-only by about 0.007, significant in 2 of 5 folds. Ordinary photographs; research prototype, not a clinical product.
Problem
Diabetic foot ulcer segmentation is usually improved by collecting more images. It is rarely checked whether the extra images help, or whether the benchmark itself is clean. I wanted to measure what the composition of the training set does when the architecture and settings stay fixed.
Approach
Three architectures (U-Net++, SegFormer-B0, DINOv2 with UPerNet) were each trained on five compositions of the same pool: DFU only, DFU plus healthy feet, DFU plus other wounds, everything, and a random mix sized to match DFU only. Each cell used 5-fold cross-validation with one seed and was scored on one fixed held-out test split.
How it was built
I built a data pipeline that assembles foot photographs into train, validation and test splits, with checks for path, ID, content-hash and near-duplicate overlap. The first audit found 96,829 near-duplicate train-test pairs, so I rebuilt the splits to 5,674, 1,216 and 1,215 images with zero overlap and reran everything. I then trained U-Net++, SegFormer-B0 and DINOv2 on five training-set compositions with 5-fold cross-validation, running the 75 jobs as a SLURM array on an HPC cluster. Each result file records the commit and checksums. Limits: one seed, one test split, research code only; the SPIE abstract has no recorded decision.
Evidence and limitations
Auditing the first split found 96,829 near-duplicate train-test image pairs. I rebuilt every split to 5,674 train, 1,216 validation and 1,215 test images with zero path, ID, content-hash or near-duplicate overlap, then reran the study.
On U-Net++, DFU plus healthy reached a mean DFU Dice of 0.879 against 0.872 for DFU only. That gap is small and was significant in 2 of 5 folds, so the safer reading is that adding other wounds or a random mix lowers DFU Dice, while healthy images mainly cut false positives on intact skin. The ordering held across all three architectures.
Limits: one seed, one test split, spread shown as fold standard deviation rather than a confidence interval, and ordinary foot photographs only. This is research code, not a medical device. The SPIE Medical Imaging 2027 abstract has been submitted and no decision is recorded.
Links
Abstract submitted. SPIE Medical Imaging 2027. Abstract submitted to SPIE Medical Imaging; no decision recorded as of October 2026, so this is neither accepted nor published.