We explain how choosing a small, representative dataset from a large population can improve model training reliability.