Artificial intelligence and machine learning are increasingly used in medicine and public health to detect disease earlier, predict risk, monitor patients, and tailor treatments.
These tools depend on training datasets that should reflect the populations where healthcare decisions will be applied.
South Asians are described as underrepresented in genomic and multi-modal health datasets used for research, including large biobanks and genome-wide association study (GWAS) catalogues.
In GWAS participation patterns, most participants are described as having European ancestry, while South Asian participation is described as a very small fraction.
Underrepresentation is also described in single-cell atlas resources, which are increasingly used as reference maps and to train AI models.
Because European-heavy training data may not match South Asian genetic patterns, genetic risk prediction tools can be less accurate for South Asian groups.
