The Filter That Fails Nigerian Genomes
Clinical genetics decides whether a variant matters partly by asking how common it is. The databases that answer that question barely contain Nigerians, and the result is a test that works less well for Nigerian patients for reasons that have nothing to do with their biology.
A Nigerian woman with a family history of breast cancer has a panel test. It finds a variant in a gene that matters. The laboratory has to decide whether that variant is dangerous, harmless, or unknown.
One of the first things it checks is how often the variant appears in people without the disease. Something present in five percent of a healthy population is not causing a rare cancer syndrome. Something absent from every control database is more suspicious.
This is reasonable, it is standard, and for this patient it is close to useless. The control databases contain almost no Nigerians.
How frequency became formal evidence
The 2015 ACMG and AMP guidelines, still the backbone of clinical variant interpretation worldwide, made population frequency an explicit scoring criterion. Three rules do most of the work.
BA1 is standalone: a variant above five percent frequency is benign, full stop, no further analysis. BS1 is strong benign evidence, applied when frequency exceeds what the disorder's prevalence could support. PM2 is moderate evidence of pathogenicity, applied when a variant is absent from controls.
Read PM2's original wording closely, because it contains the whole problem. The criterion asks whether the variant is absent from a large control group, of more than a thousand individuals, of the same population as the patient. The guideline authors understood that frequency is only meaningful against a matched reference. They wrote the requirement in.
They just could not make the reference exist.
What "African" means in the database everyone uses
gnomAD is the reference in practice. Nearly every clinical laboratory checks it, and many national guidelines name it directly.
Its African-ancestry grouping is labelled African/African American, and in version 4.0 it holds around 20,800 samples. gnomAD's own analysis of that group found that only about 32 percent of those individuals have more than 95 percent African ancestry. The other 68 percent are substantially two-way admixed, carrying a majority of West African ancestry alongside variable European ancestry.
That is a description of African American population history, and it is accurate. It is not a description of Nigerian population history. A Yoruba woman in Ibadan and an African American woman in Baltimore share deep ancestry and differ considerably in the details, and the details are exactly what allele frequency is measuring.
gnomAD knows this. In October 2024 it released local ancestry inference for that group specifically to separate African and European haplotypes, on the stated grounds that aggregate frequencies could lead to inaccurate variant classification in underrepresented populations. That is a real improvement. It also does not add a single Nigerian to the database.
Nigeria's actual representation
The 1000 Genomes Project remains the foundational reference panel for human variation. Nigeria appears in it twice: YRI, Yoruba in Ibadan, at 108 individuals, and ESN, Esan in Nigeria, at roughly 100.
So approximately two hundred people, drawn from two ethnic groups, stand in for a country of more than two hundred million people and several hundred ethnolinguistic groups.
Coriell, which distributes those samples, publishes a warning alongside them. It states that the full population name matters because the sample set does not represent all Yoruba people, whose population history is complex, and that the population should not be described merely as African, Sub-Saharan African, West African, or Nigerian, since each of those terms covers many communities.
The custodians of the reference data explicitly warn against the generalisation that clinical pipelines perform on it every day.
And the warning is not academic. When 54gene sequenced 449 Nigerians across 47 ethnolinguistic groups, published in Cell Genomics in 2023, population structure analysis found genetic differentiation between groups within Nigeria. Yoruba and Esan are not interchangeable with Igbo, Ibibio, Bini, or Izon, let alone with the Hausa and Fulani populations of the north. Treating "Nigerian" as one frequency bucket is already too coarse.
What this does to results
The effect is measurable, and it lands in one direction.
In a multi-centre study using a hereditary cancer panel across 2,571 patients of European ancestry and 110 of African ancestry, pathogenic variant detection was almost identical: 13.3 percent against 12.7 percent. The variants of uncertain significance were not. They ran at 46.1 percent in the European ancestry group and 65.5 percent in the African ancestry group.
The disease is equally detectable. The certainty is not.
Across the wider literature, individuals of non-European ancestry are estimated to be roughly 1.5 to 2 times more likely to receive a VUS result on a hereditary cancer panel. In a South African study of early-onset colorectal cancer in Indigenous African patients, 47 percent carried variants of uncertain significance judged to lean pathogenic, against 16 percent with confirmed pathogenic findings, and the authors attributed the imbalance directly to limited representation of African genomes in reference databases.
A VUS is not a diagnosis. Guidelines are explicit that it cannot be used as the basis for clinical decisions. So the patient receives a result that changes nothing, having been through the counselling, the blood draw, the wait and the cost. Reported alongside anxiety, altered self-perception, and occasionally unnecessary intervention when a non-specialist misreads it.
A test that returns uncertainty at one and a half times the rate for one group of patients is not a neutral test.
Why this is not an algorithm problem
There is a persistent hope that better computational methods will close the gap. They will not, and the reason is structural rather than technical.
Allele frequency is not inferred. It is counted. If a variant appears in six percent of Igbo people and nobody has sequenced enough Igbo people, no model recovers that number, because the information was never collected. Machine learning trained predominantly on European-ancestry data reproduces the imbalance rather than correcting it: polygenic risk scores derived from European cohorts are estimated to lose between 39 and 73 percent of their predictive accuracy when applied to African-ancestry cohorts.
There is a second-order problem too. African genomes carry more variation than any others, because human genetic diversity is greatest in Africa. In gnomAD the African-ancestry cohort showed roughly a 1.8-fold enrichment of common missense variants compared to the non-Finnish European cohort, from a far smaller sample. More variation observed against thinner reference data produces more uncertainty per patient, not less. Underrepresentation and high diversity compound.
A study of childhood hearing loss among Yoruba families near Ibadan illustrates what that looks like at the bedside. No biallelic pathogenic variants were found in GJB2, the most common cause of inherited deafness in many populations. Likely causal variants were identified in about a third of independent cases, none of them recurring in more than one family, and 77 percent had not previously been associated with hearing loss at all. The authors described an unusually high level of genetic heterogeneity and flagged the consequences for screening and counselling.
That is not an interpretation failure. That is a population whose genetics have not been described yet.
What actually fixes it
Only one thing does, and it is unglamorous: count the variants in the population you intend to serve, at sufficient scale, with sufficient granularity, and make the resulting frequencies available to the laboratories doing the interpreting.
That means a national reference resource with real ethnolinguistic breadth rather than two communities standing in for a country. It means those frequencies reaching clinical pipelines in the formats laboratories already consume, which is an infrastructure question as much as a scientific one. And it means the resource being governed durably enough that laboratories can rely on it for the decades over which variant classifications get revised.
Every one of those is a reason for national genomic infrastructure that has nothing to do with national prestige. A Nigerian allele frequency database is not a symbolic asset. It is the missing input to a clinical decision that is currently being made badly, thousands of times a year, in Nigerian hospitals and in laboratories abroad reading Nigerian samples.
There is an uncomfortable postscript. The most detailed catalogue of Nigerian genomic variation published to date, those 449 genomes across 47 ethnolinguistic groups, came out of 54gene. The company no longer exists, and its biorepository has been the subject of court proceedings in Lagos over an attempted sale.
The work was started. What was missing was never the science.
Contact