South Asians are missing from global health databases: why this matters and what needs to change
Account subscription benefits alongside Premium Stories, Editorials, Opinions and more. Unlock these with Subscription
Globally, more than one in 10 adults now live with diabetes. If you have South Asian roots, that risk is even higher, and it hits earlier than it does in many other populations. In India alone, the number of people with diabetes is projected to reach 125 million by 2045 .
Yet, when scientists try to understand why diseases such as diabetes and cardiovascular disease affect South Asians differently, they often have to rely on genetic data drawn from European populations.
Advances in artificial intelligence and machine learning are allowing scientists to mine vast amounts of genomic and health data to detect disease earlier, predict risk, monitor patients and tailor treatments to individuals. But at the heart of this changing landscape lies an old, constant problem – the data used to build these tools lack diversity.
Integrated biobanks such as the U.K. Biobank, which combine participants’ genomic information with electronic health records, environmental exposures and lifestyle data, have transformed biomedical research. These repositories have accelerated drug development, informed clinical guidelines and helped shape public health policy across the world.
But South Asians remain largely absent from these datasets. The NHGRI-EBI GWAS Catalogue , an online database of human genome-wide association studies, shows that between 2005 and 2025, more than 86% of participants in these studies were of European ancestry, while South Asians accounted for less than 1%.
“More than 20% of the world is being neglected in multi-modal data integration,” said Bhramar Mukherjee, senior associate dean of public health data science and data equity, Yale School of Public Health, United States. “It denies [them] the human right and opportunity to attain the maximal possible health”.
The pattern repeats in newer tools, too. A recent study published in Cell Genomics reviewed more than 13,500 samples across three major single-cell resources—the Human Cell Atlas, the Human Tumour Atlas Network and the PsychAD Consortium and found found a “striking, pervasive European overrepresentation and underrepresentation of Asian and Latino individuals.”
“These single-cell atlases are becoming the reference maps for biology and medicine, and they are increasingly used to train the AI models that will shape future research and care,” said Kuan-lin Huang, senior author of the study, who is an associate professor of genetics and genomic sciences, and AI in human health, at the Icahn School of Medicine, United States.
These data points are more than just statistics. South Asians face higher rates of type 2 diabetes, cardiovascular disease and asthma than people of European ancestry, which means the tools built on European-heavy data are less accurate for the population that needs them most.
Take polygenic risk scores, which combine the effects of many genetic variants associated with a disease to estimate a person’s overall genetic risk.
5News aggregated this summary from the outlet’s public feed. The full article, with all the context, is on www.thehindu.com — the content belongs to The Hindu - Health.