"Two courses, one idea: a cross-validation score is a statistic with its own sampling distribution, the same way a sample mean is"
STAT 3130 (survey sampling) and STAT 4630 (machine learning) look like they have nothing in common, but several of their core ideas are the same concept in different vocabulary. The clearest one: a sample mean has a sampling distribution, because a different random sample would have given a different number. A cross-validation error estimate is exactly the same kind of object. Re-split the data and you get a different score, so a single CV number says as little about its own reliability as a single survey estimate does without its standard error.
Data leakage is a sampling-design mistake. When a dataset has repeated visits per patient, splitting train and test by row lets the model recognize patients it already saw. In survey terms, the sampling unit (the patient) is coarser than the observation unit (the visit), and the split ignored that. Grouped cross-validation is the fix, and it is the same instruction survey design gives: sample at the level of the true independent unit. Stratified k-fold works the same way as stratified sampling, with one telling difference: it uses proportional allocation rather than the variance-minimizing Neyman allocation, because every fold has to stand on its own as a representative test set.
The same question sits under the garment-capture pilot: does a sample represent what I want to draw conclusions about? Three garments can show that the pipeline works, but not that the clustering results would hold across the collection. That is why the next milestone is 15 to 20 garments, and why I describe the current results as suggestive rather than conclusive.