Dataset Diversity Metrics and Impact on Classification Models
Diversity is commonly viewed as a good and important feature for a dataset, as a large and diverse dataset should lead to better results and generalization. However, while datasets are often claimed to be diverse, what “diverse” is, is not clearly defined and may change from paper to paper. It could, for example, refer to the demographics of the patients, the scanner being used, the annotators, etc.
There are some metrics to quantitatively measure the diversity of a set, they are however mostly used in a generative context such as image generation or answers from LLMs. Their usage on real datasets and their correlation with a downstream performance task is therefore unclear. Finally, while datasets are increasingly multimodal, for example containing chest X-rays and radiology reports with metadata, most studies only assess the diversity of a single modality for a dataset.