Skip to content

What Preventive Genetic Testing Can and Cannot Tell You

Oak
A dark glowing forest of fiber-like trunks, with only a small cluster of bright amber seed pods lit in the foreground.

Preventive genetic testing means sequencing a healthy person’s DNA to find variants that raise their risk of a condition before symptoms appear. The useful part of it is narrow and well defined: roughly 80 genes with strong evidence linking a pathogenic variant to a condition where surveillance or intervention changes outcomes, plus pharmacogenes that predict how you metabolize certain drugs. Between one and three percent of unselected adults carry a pathogenic variant in that actionable set. Almost everything else in your genome is either uninterpretable today, weakly predictive, or interesting but not actionable, and a test that presents all of it with equal confidence is doing you a disservice.

The actionable set is small and the boundary is explicit

The American College of Medical Genetics and Genomics maintains a list of genes for which it recommends reporting secondary findings when clinical sequencing is done for any reason. The current version, ACMG SF v3.2, contains 81 genes. The list exists because these are the cases where the argument is strongest: a pathogenic variant confers substantial risk, the condition is serious, and there is an established management pathway a clinician can follow.

The composition tells you what kinds of evidence clear that bar. Hereditary cancer predisposition dominates: BRCA1 and BRCA2, the Lynch syndrome mismatch repair genes (MLH1, MSH2, MSH6, PMS2, EPCAM), APC, TP53, PTEN, RB1, RET, the SDHx paraganglioma genes. Cardiovascular disease is the second cluster: hypertrophic and dilated cardiomyopathy genes (MYH7, MYBPC3, TNNT2, LMNA), the long QT and Brugada channel genes (KCNQ1, KCNH2, SCN5A), the aortopathy genes (FBN1, TGFBR1/2, ACTA2), and LDLR, APOB, and PCSK9 for familial hypercholesterolemia. Then a short tail: HFE for hereditary hemochromatosis, RYR1 and CACNA1S for malignant hyperthermia susceptibility, TTR for hereditary transthyretin amyloidosis, OTC, GLA, ATP7B.

Everything on that list has the same structure. A rare variant with a large effect, a phenotype with a known natural history, and something a clinician does differently if they know. A pathogenic LDLR variant is worth finding because familial hypercholesterolemia is common (roughly 1 in 250) and frequently undiagnosed until a first cardiac event. Malignant hyperthermia susceptibility matters because the trigger is a specific class of anesthetic, and knowing beforehand is the entire intervention. If you find something in this set, you take it to a clinical geneticist or a genetic counselor. Confirmation in an accredited clinical laboratory is the standard next step, because research-grade calls are not diagnoses.

Public health genomics has drawn this line deliberately. The case for population-level screening rests on validated tests with demonstrated clinical utility, and most genomic associations do not meet that standard even when the statistical association is real.1

Why we recommend whole genome over a targeted panel

A hereditary cancer panel from a clinical lab covers 30 to 90 genes, is CLIA-certified, and costs a few hundred dollars. Whole genome sequencing at 30x coverage costs more and gives you three billion base pairs. If your only question is “do I carry a pathogenic BRCA1 variant,” the panel is the better instrument: better validated, ordered by a clinician, and reportable.

We still prefer whole genome for a baseline, for three reasons. First, the panel answers the question you asked in 2026 and nothing else. Genome data is reinterpretable as ClinVar and gnomAD grow. A variant of uncertain significance in your 2026 VCF can be reclassified in 2031 without drawing blood again, and reclassification is common. Second, genomes cover structural variation, intronic regions, and pharmacogenomic haplotypes that panels drop by design, including the CYP2D6 copy-number and hybrid alleles that most targeted assays call poorly. Third, exome capture has uneven coverage at GC-rich first exons, and whole genome sequencing gives more uniform depth across the same coding regions. Payers have noticed the difference in what the two assays deliver but generally still treat genome sequencing as investigational for asymptomatic adults, which is why preventive sequencing is usually self-funded.2

The tradeoff to name plainly: a genome hands you far more variants of uncertain significance, and the failure mode of preventive testing is not missing a variant but over-reading one.

What you receive, and what to do with it

Insist on raw data. The deliverables that matter are a FASTQ or aligned CRAM file and a joint-called VCF, not a PDF.

  • FASTQ, or better, CRAM aligned to GRCh38 with an explicit reference (GRCh38_full_analysis_set_plus_decoy_hla.fa). CRAM is roughly 60 percent smaller than BAM for the same data. A 30x human genome is about 15 GB as CRAM.
  • A VCF with per-sample genotype quality (GQ), depth (DP), and allele balance, so you can filter. bcftools view -i 'FILTER="PASS" && FMT/DP>=10 && FMT/GQ>=20' removes most of the noise.
  • A gVCF if you care about reference confidence, which you do if you want to distinguish “no variant” from “no coverage.” A homozygous-reference call in a region with DP=3 is not evidence of absence.

For interpretation, the pipeline we would run is annotation with Ensembl VEP (--everything --assembly GRCh38 --cache), then a filter down to the 81 ACMG genes by coordinate, then an intersection with ClinVar restricted to variants with a review status of “reviewed by expert panel” or “criteria provided, multiple submitters, no conflicts.” Anything labeled pathogenic or likely pathogenic in that filtered set is worth a clinical conversation. Anything of uncertain significance is not, and should be recorded rather than acted on.

Population frequency is the strongest single filter you have. gnomAD v4 covers more than 800,000 exomes and genomes. A variant claimed to cause a dominant, severe, early-onset condition cannot appear at 0.5 percent allele frequency in a matched population. When you see a scary-sounding annotation, check the frequency first, and check it in the population that matches your ancestry, since a variant common in Finnish or Ashkenazi cohorts and absent from European reference sets is a recurring source of false alarms.

Two technical failure modes deserve specific attention. PMS2 has a pseudogene, PMS2CL, with high sequence similarity, and short-read alignment produces both false positives and false negatives there. SMN1 and SMN2 differ at a handful of positions, so standard pipelines call carrier status for spinal muscular atrophy unreliably. Both require targeted assays such as MLPA or long-read sequencing to resolve. If your report is silent on these genes, that silence may reflect a pipeline limitation rather than a clean result.

Polygenic scores, nutrigenomics, and the weakly predictive middle

Beyond the rare high-penetrance variants sits a large middle ground where associations are real but individual prediction is poor. Polygenic risk scores for coronary artery disease or type 2 diabetes sum thousands of small effects and do separate populations into risk strata, but they transfer badly across ancestries because they were derived mostly from European cohorts, and for any single person the score shifts a probability rather than settling a question.

Nutritional genomics is the clearest example of a field where the biology is genuine and the clinical translation lags. Gene-diet interactions like MTHFR and folate metabolism or APOE genotype and dietary fat response are documented, but the guidance for practice remains that evidence supporting individualized dietary recommendations from genotype is limited.3 The professional position statements are explicit that most gene-diet findings are not yet actionable at the individual level.4 Treat nutrigenomic reports as hypotheses to test against your own biomarkers, not as prescriptions.

The general caution is that a genotype constrains rather than determines, and the strength of that constraint varies enormously across traits.5 Where sequencing is genuinely diagnostic, it tends to be for conditions with clear Mendelian architecture: inherited retinal dystrophies, where identifying the causative gene determines eligibility for gene-specific therapy and trials, are a good illustration of the ceiling for that kind of result.6 At the other extreme, traits like bronchopulmonary dysplasia have substantial heritability estimates with almost no replicated variants, which shows what the middle ground looks like when you probe it hard.7

Questions people also ask

What is preventive genetic testing? Sequencing a person without symptoms or a specific clinical indication, to find variants that raise disease risk before disease appears. In practice, the useful output is a short list from the ACMG secondary findings genes plus pharmacogenomic haplotypes, not a comprehensive risk forecast.

Does insurance usually cover genetic testing? Diagnostic testing with a clear indication (a suspicious family history, an existing cancer diagnosis, a symptomatic child) is often covered. Sequencing an asymptomatic adult for screening generally is not, and US private payers have treated genome sequencing in particular as insufficiently supported by evidence for coverage.2 Expect to self-fund preventive sequencing.

Who qualifies for genetic testing for cancer risk? Clinical criteria center on family history: multiple relatives with the same or related cancers, unusually early age at diagnosis, a known familial variant, certain ancestries with founder variants, or specific tumor features such as a mismatch-repair-deficient colorectal tumor. If you meet those criteria, a clinician-ordered panel is the right route, since you need a reportable result and a genetic counselor to interpret it.

What should I do if my sequencing report flags a pathogenic variant? Take it to a clinical geneticist or genetic counselor and expect confirmatory testing in an accredited clinical laboratory before anything follows from it. Research-grade and direct-to-consumer calls have real false positive rates, and management decisions require a clinician.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Muin J. Khoury, Michael S. Bowen, Wylie Burke, et al. Current Priorities for Public Health Practice in Addressing the Role of Human Genomics in Improving Population Health. American Journal of Preventive Medicine, 2011. https://doi.org/10.1016/j.amepre.2010.12.009 ↩

  2. Kathryn A. Phillips, Julia R. Trosman, Michael P. Douglas, et al. US private payers’ perspectives on insurance coverage for genome sequencing versus exome sequencing: A study by the Clinical Sequencing Evidence-Generating Research Consortium (CSER). Genetics in Medicine, 2022. https://doi.org/10.1016/j.gim.2021.08.009 ↩ ↩2

  3. Ruth M. Debusk, Colleen P. Fogarty, José M. Ordovas, et al. Nutritional genomics in practice: Where do we begin?. Journal of the American Dietetic Association, 2005. https://doi.org/10.1016/j.jada.2005.01.002 ↩

  4. Kathryn M. Camp, Elaine Trujillo. Position of the Academy of Nutrition and Dietetics: Nutritional Genomics. Journal of the Academy of Nutrition and Dietetics, 2014. https://doi.org/10.1016/j.jand.2013.12.001 ↩

  5. Evelyn Fox Keller. Nature, Nurture, and the Human Genome Project. The Ethics of Biotechnology, 2021. https://doi.org/10.4324/9781003075035-23 ↩

  6. Megan E. Soucy. Genetic Testing for Inherited Retinal Dystrophies: Basic Understanding. Advances in Experimental Medicine and Biology, 2025. https://doi.org/10.1007/978-3-031-72230-1_56 ↩

  7. Gary M. Shaw, Hugh M. O’Brodovich. Progress in understanding the genetics of bronchopulmonary dysplasia. Seminars in Perinatology, 2013. https://doi.org/10.1053/j.semperi.2013.01.004 ↩