Introduction

79% of Genome-Wide Association Study (GWAS) participants are of European descent, despite Europeans accounting for just 16% of the world population. Non-European participation in such studies both stagnated and declined after 2014, implying this issue is not on par with correcting itself (Martin et al.). GWAS began as an experimental research tool into shaping the backbone of modern life, medicine, and healthcare; commercial DNA testing; and dictates how common diseases are predicted, how everyday medications are dosed, and how new drugs are discovered. Beyond diagnosis and disease prevention, GWAS also plays a critical role in understanding complex trait heritability for conditions such as heart disease, diabetes, and psychiatric disorders; tracing human migration patterns, ancestry, and ancestral relationships; and measuring how much of the variation in a trait within a population comes from genetics. As a South Asian student researcher with experience in computational protein engineering in the field of genetics, specifically, with such applications in mind, I found myself pondering how this could impact people like me.

Prediction accuracy decays as a function of genetic distance from the discovery group. This bias has been baked into any model. Most GWAS arrays genotype variants selected for being common among European populations, and reference panels used to input additional variants are overwhelmingly Eurocentric, too. This causes the input features to be skewed even before training begins (Martin et al.). Empirical data have shown that fewer than 40% of European bone mineral density variants replicate in Asian and African populations, showing how poorly these estimates apply (Wu et al.).

So What?

Polygenic risk scores (PRS) are several times more accurate for people of European ancestry than in others. Martin et al. argue that this makes PRS systematically deliver greater benefits to European-descent patients if clinically deployed today. Manrai et al. identified benign variants that had been classified as pathogenic variants for hypertrophic cardiomyopathy — a genetic heart condition where the muscle in the heart becomes abnormally thick, making it harder to pump blood.

Every single patient that received a false-positive was of African or unspecified ancestry. Simulations showed that including a small number of Black Americans in control cohorts could have prevented this (a frequency of 0.0157 in Black Americans versus 0.000122 in white Americans) (Manrai et al.).

Exacerbations by Advances in ML

We live in a time where technology advances faster than most of us can process. As frightening as it is, previously tedious, expensive, and difficult research has been streamlined by advancements in deep learning, machine learning, and bioinformatics. Despite this efficiency, standard deep learning approaches fail to learn unbiased representations under population-skewed training data, similar to PRS (Gyawali et al.). Generalized population labels fail to capture the complexity of human genetic variants, especially for underrepresented populations. Moreover, the “fat data” problem compounds it. Genetic variants outnumber samples exponentially, making deep models prone to overfitting (Boulanger et al.). This oversimplification propagates into downstream health predictions, risking the reinforcement of discriminatory interpretations.

Conclusion

As deep learning gets folded into genomics, the risk isn’t that bias disappears with more sophisticated tools; it is that it becomes harder to see. Neural networks that are trained on skewed data cannot announce the blind spots it has, the way an allele frequency table can, it just delivers opaque predictions that inherit the gap in the data it was fed to train on. Creating truly diverse discovery and control cohorts, building non-European reference panels with the same rigor as European ones, and reporting PRS accuracy stratified by ancestry is not impossible. As someone who falls into a category outside of what these models were built around, this is a critical difference between a risk score that protects some and reassures others incorrectly. Genomic and precision medicine are moving toward a future of more individualized and predictive care, and whether that future is actually equitable doesn’t depend on the advancements of algorithms, but on the decisions being made in cohort design.

References

Martin, Alicia R., et al. “Current Clinical Use of Polygenic Scores Will Risk Exacerbating Health Disparities.” bioRxiv, 1 Feb. 2019, www.biorxiv.org/content/10.1101/441261v3, doi:10.1101/441261.

Martin, Alicia R., et al. “Low-Coverage Sequencing Cost-Effectively Detects Known and Novel Variation in Underrepresented Populations.” bioRxiv, 28 Apr. 2020, www.biorxiv.org/content/10.1101/2020.04.27.064832v1, doi:10.1101/2020.04.27.064832.

Wu, Qiong, et al. “Bridging Genomic Research Disparities in Osteoporosis GWAS: Insights for Diverse Populations.” Current Osteoporosis Reports, vol. 23, no. 1, 24 May 2025, article 24, doi:10.1007/s11914-025-00917-2.

Martin, Alicia R., et al. “Clinical Use of Current Polygenic Risk Scores May Exacerbate Health Disparities.” Nature Genetics, vol. 51, no. 4, 29 Mar. 2019, pp. 584–591, doi:10.1038/s41588-019-0379-x.

Manrai, Arjun K., et al. “Genetic Misdiagnoses and the Potential for Health Disparities.” The New England Journal of Medicine, vol. 375, no. 7, 18 Aug. 2016, pp. 655–665, doi:10.1056/NEJMsa1507092.

Ding, Yi, et al. “Improving Genetic Risk Prediction across Diverse Population by Disentangling Ancestry Representations.” Nature Communications, vol. 13, 2022, article 6455, doi:10.1038/s41467-022-34363-2.

Rochefort-Boulanger, Camille, et al. “A Transparent and Generalizable Deep Learning Framework for Genomic Ancestry Prediction.” bioRxiv, 31 Aug. 2025, www.biorxiv.org/content/10.1101/2025.08.26.672448v1, doi:10.1101/2025.08.26.672448.