Building DiaFoot.AI: what a leaky benchmark taught me
Published · Medical imaging · Segmentation · Data leakage
DiaFoot.AI asks a narrow question about diabetic foot ulcer (DFU) segmentation: does the kind of images you train on matter more than how many you use? This is how the project got to that question, the mistakes in my own benchmark along the way, and what I can and cannot claim.
The first version worked, and was useless
Version one was a U-Net++ trained only on ulcer images. It reported a Dice score of about 92%. It had also never seen a healthy foot, so it predicted an ulcer on every input, including healthy skin. The score measured how well it outlined wounds in a set where every image contained one. In March 2026 I rebuilt the project as version two, adding healthy-foot and other-wound photographs as negative classes and adding a triage classifier stage before segmentation (architecture only; its earlier numbers are withdrawn).
The leak in my own splits
The numbers still looked better than they should. A perceptual-hash audit of my splits found 96,829 near-duplicate image pairs between train and test, plus 13,557 between train and validation and 2,375 between validation and test. Many images were crops or close variants of one photograph, so models were being tested on pictures they had effectively seen. Internal scores near 0.98 to 1.0, against about 0.33 on an external set, were the symptom.
The fix was to group every crop of one source image together, merge near-duplicates before partitioning, and audit again. The rebuilt splits hold 5,674 training, 1,216 validation and 1,215 test images, with zero path, ID, content-hash and near-duplicate overlap. I retrained everything on them, and the honest numbers were lower. One lesson from the audit tool itself: it skips images it cannot find, so it can report zero overlap on a machine that does not hold the images. It has to run where the data lives.
My own documentation also disagreed with itself. Some older tables still quoted test-set sizes from the leaky split, and the numbers behind my earlier cascade results, a classifier-plus-segmenter pipeline, are not saved as result files in the repository. I do not quote those figures anywhere.
No patient identifiers
The four public datasets I used publish no patient IDs, so I cannot guarantee one patient's photographs stay inside one cross-validation fold. I grouped folds by the strongest image-provenance identifier available and say plainly that patient-level independence is not claimed. To keep the headline number honest anyway, every model is scored on the same fixed, leakage-audited test split that no fold touches.
Separating what from how much
Comparing a large mixed dataset with a small disease-specific one confuses two things, so I added a control: a random mix with the same number of images as the DFU-only set. I also measured how often a model predicts a wound on skin whose ground-truth mask is empty, because a single stray pixel there scores zero Dice and disappears inside an average.
What 75 models showed
I trained five training-set compositions on three architectures (U-Net++, SegFormer-B0 and DINOv2 with UPerNet) across five folds, as a SLURM array on an HPC cluster, with seed 42. All 75 were scored on the 1,215 test photographs, 318 of which contain an ulcer. For U-Net++:
| Training set | DFU Dice | False positives on empty skin |
|---|---|---|
| DFU + healthy feet | 0.879 | 0.009 |
| DFU only | 0.872 | 0.083 |
| All three classes | 0.846 | 0.015 |
| DFU + other wounds | 0.834 | 0.221 |
| Random mix, size of DFU only | 0.795 | 0.016 |

The ranking was the same on all three architectures. The size-matched random mix did worst, so the gap is not just dataset size. Adding other wound types made models over-segment healthy skin; the worst case was SegFormer-B0 at 0.448. Adding healthy feet barely moved Dice but cut false positives roughly ninefold on U-Net++, from 0.083 to 0.009.
What I would not claim
The Dice gain from healthy images, 0.879 against 0.872, is small. A paired bootstrap against DFU-only was significant in 2 of 5 folds for U-Net++; across all three architectures the difference was positive in every fold comparison and significant in 9 of 15. The false-positive reduction is the stronger result. Everything is a single seed. Folds are provenance-grouped, not patient-grouped. The data come from four public sources, so transfer to other hospitals or cameras is unshown. The test set is effectively one skin-tone group. About 16% of the other-wounds pool is itself diabetic, which if anything makes the harm estimate conservative.
The inputs are ordinary foot photographs, not thermal scans. This is research code, not a medical device. An abstract based on the work was submitted to SPIE Medical Imaging 2027 with Prof. Mohammad Eslami of Harvard Medical School; as of October 2026 no decision is recorded.