Skip to content

Proactive Genetic Testing: What to Sequence, What It Misses, and How to Read Your Own Data

Oak
A glowing filament runs tree to tree through a misty bioluminescent forest, mostly dim grey-green with a few bright amber segments ringed by moss.

Proactive genetic testing means sequencing someone who has no symptoms and no family history that triggered the order, then looking for variants that predict risk before anything shows up in a clinic. The short answer on what to do: get whole-genome sequencing at 30x or better, insist on the raw CRAM/BAM and a gVCF rather than a PDF, and treat the report you get back as one interpretation of that data at one point in time, not as the result. Yield in unselected adults is real but modest, and most of what comes back will be variants of uncertain significance you cannot act on. The value is in owning a substrate you can reinterpret for the rest of your life.

Yield: what a proactive test returns

A Mayo Clinic cohort offered predictive genomic testing inside ordinary medical practice and found clinically actionable results in a minority of apparently healthy adults, with the bulk concentrated in cancer predisposition and cardiovascular genes 1. That is the shape of the prior: most people get nothing that changes anything, a small fraction get something that changes surveillance, and nearly everyone gets a pile of uncertain calls.

The comparison that matters more is panel versus genome. When a medically actionable gene panel was run against genome sequencing in the same proactive pediatric screening population, the genome surfaced at-risk findings the panel did not 2. The panel is not wrong, it is bounded by its design, and the design was frozen before you were sequenced. A genome is not frozen.

The sharpest illustration is familial hypercholesterolemia. Consumer-style limited-variant screening tests a fixed set of roughly two dozen known pathogenic variants. Compared against comprehensive sequencing of LDLR, APOB, and PCSK9, limited-variant screening detects only a small fraction of carriers, and the miss rate is worst in people of non-European ancestry because the variant list was built from European cohorts 3. A negative from that kind of test carries almost no information. This is the single best argument for sequencing the gene rather than genotyping a list.

Panel, exome, or genome

We would pay for a genome. Reasons, in order of weight:

  1. Reinterpretation. ClinVar moves. A variant classified VUS in 2023 may be likely pathogenic in 2027. If you hold a CRAM, you re-run the annotation. If you hold a panel report, you re-order the test.
  2. Non-coding and structural coverage. Exomes give you roughly 1–2% of the genome and poor, uneven coverage of deep intronic splice regions and promoters.
  3. Copy number and repeats come along for free if the pipeline is set up for them.

The tradeoff is real: a genome produces 4–5 million variant calls against GRCh38, and the false positive burden in low-complexity and segmental duplication regions is where you will waste your time. A clinical panel is enriched for signal per unit of analyst attention. If you have no intention of touching the data yourself, a curated panel with a real clinical lab behind it is a defensible choice.

The variant classes a naive SNV pipeline drops

This is where most self-analysis goes wrong. A standard DeepVariant or GATK HaplotypeCaller run followed by VEP and a ClinVar join will silently return “nothing found” for entire categories of pathogenic variation.

  • Large copy number events. CMT1A, the most common inherited peripheral neuropathy, is a 1.4 Mb duplication at 17p11.2 spanning PMP22, not a point mutation 4. You need read-depth CNV calling (GATK gCNV, CNVkit) or a structural variant caller (Manta, GRIDSS) to see it.
  • Short tandem repeat expansions. Run ExpansionHunter with a catalog covering the known pathogenic loci. Short reads give you a noisy size estimate, not a clean answer, and long expansions beyond read length are systematically underestimated.
  • Segmental duplication blind spots. PMS2 exons 11–15 are near-identical to the pseudogene PMS2CL. SMN1 versus SMN2 differs at a handful of paralogous bases. CYP2D6 has hybrid alleles with CYP2D7. Use purpose-built callers: SMNCopyNumberCaller for SMN, Cyrius or Aldy for CYP2D6.
  • LPA KIV-2 copy number, which drives lipoprotein(a) levels, sits in a tandem repeat array that short reads cannot phase or count reliably. Measure Lp(a) in plasma instead.

Published best-practice guidance for clinical WGS interpretation is explicit that a genome test is a collection of analyses (SNV, indel, CNV, SV, repeat, mitochondrial) with separate sensitivities, and that reporting should state which were performed 5. Ask your provider which of these they ran. If they cannot answer, they ran the first one.

The stack we’d run on our own CRAM

Assume you have sample.cram aligned to GRCh38 and sample.g.vcf.gz.

# 1. Sanity: coverage and contamination before anything else
mosdepth --by 500 --fast-mode sample sample.cram
verifybamid2 --SVDPrefix 1000g.phase3.100k.b38 --BamFile sample.cram

# 2. Annotate
vep -i sample.vcf.gz --cache --assembly GRCh38 --offline \
    --everything --pick_allele_gene --vcf --compress_output bgzip \
    --plugin dbNSFP,dbNSFP4.7a.gz,REVEL_score,CADD_phred,AlphaMissense \
    --plugin SpliceAI,snv=spliceai_snv.hg38.vcf.gz,indel=spliceai_indel.hg38.vcf.gz \
    -o sample.vep.vcf.gz

# 3. Join ClinVar, keep 2-star and above
bcftools annotate -a clinvar.vcf.gz \
    -c INFO/CLNSIG,INFO/CLNREVSTAT,INFO/CLNDN sample.vep.vcf.gz -Oz -o sample.cv.vcf.gz

# 4. Filter: rare, high-quality, in a gene you care about
slivar expr --vcf sample.cv.vcf.gz --pass-only \
  --info 'INFO.gnomad_popmax_af < 0.001 && variant.FILTER == "PASS" && variant.QUAL > 30' \
  --region acmg_sf_v3.2.bed -o candidates.vcf

Two details that matter more than the tool choice. First, restrict your initial read to a defined gene list, and be deliberate about which one. The ACMG secondary findings list (v3.2, on the order of eighty genes) is the conservative default: those genes were selected because a finding changes surveillance and the evidence is strong. Looking genome-wide on your first pass produces hundreds of ClinVar “pathogenic” hits in genes where the annotation is wrong, stale, or recessive-only and you are a carrier.

Second, filter on ClinVar review status, not just CLNSIG. A single-submitter, no-assertion-criteria “Pathogenic” is a claim, not a classification. Keep CLNREVSTAT at criteria_provided,_multiple_submitters,_no_conflicts or better. Then check zygosity, then pull the reads in IGV and look at the actual alignment before you believe anything. Introductory WGS analysis walkthroughs cover the mechanics of the alignment-to-variant path if you want the underlying steps 6.

VUS is the default state, and it is getting better

Most missense variants in most genes have no classification. The population-scale fix is multiplexed assays of variant effect: measure every possible missense substitution in a gene functionally, then use that map as evidence for every carrier who ever appears. A recent systematic effort did exactly this for AIRE, producing effect estimates for missense variants ahead of anyone needing them 7. As these maps accumulate, VUS rates in covered genes fall, which is the concrete reason to keep your CRAM rather than a report. Re-annotation is a weekend, re-sequencing is not.

Data security, plainly

Your genome is the one identifier you cannot rotate, and it partially discloses your relatives without their consent. The legal and technical literature on genomic data protection treats it as a distinct category for that reason: re-identification from a handful of markers is straightforward, and consent frameworks written for de-identified clinical data do not transfer 8. Practical posture: keep the primary copy of the CRAM on hardware you control, encrypted at rest (LUKS or age for the file itself), with an offline backup. Read the provider’s policy on secondary research use and on what happens to your data if the company is acquired or liquidated. Sequencing providers are ordinary companies and some go bankrupt, at which point the customer database is an asset. The broader analysis of next-generation sequencing data governance makes the same point about custody chains and third-party transfer 9. Separately, patent and licensing practice has historically constrained who may run a given test and reinterpret it, which is a reason to prefer providers that return raw data over ones that return only an interpretation 10.

Nothing here is a diagnosis. A pathogenic finding in a cancer predisposition or cardiomyopathy gene needs confirmation in a CLIA/CAP laboratory on a fresh sample, and needs a genetic counselor or medical geneticist to place it against your family history and decide on surveillance. Do not change anything you are doing based on a variant you found in your own VCF.

Questions people also ask

What is the downside to genetic testing? Three real ones. The dominant output is uncertain variants that cannot guide anything and do generate anxiety and downstream imaging. Absence of a finding is weakly informative, especially with limited-variant products 3. And the data is permanently identifying, for you and partly for your relatives 8.

What is the most secure genetic testing? The arrangement where you hold the raw files and the provider holds as little as possible after delivery. Look for a written deletion policy with a timeline, no default enrollment in research or third-party sharing, and delivery of CRAM plus gVCF so you are not dependent on the vendor’s continued existence for reinterpretation.

Is it worth doing PGT testing on embryos, and what are the downsides? Preimplantation genetic testing is a different procedure from proactive adult sequencing and sits squarely with a reproductive endocrinologist and a genetic counselor. The known limits include biopsy of a few trophectoderm cells as a proxy for the whole embryo, mosaicism producing discordant results, and amplification artifacts. We do not give a view on whether to do it.

Which DNA testing company got sued? Several, over data breaches, marketing claims, and patent disputes. Rather than track the litigation, evaluate any provider on two mechanical questions: do they return raw data, and what is written in the contract about secondary use.

Should I use a panel or a genome? Genome, if you intend to hold and reanalyze the data. Genome sequencing surfaced at-risk findings that a curated actionable-gene panel missed in the same proactive screening population 2. If you want a clean clinical answer and nothing else, a CLIA panel with a counselor attached is the lower-variance path.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Jennifer L. Anderson, Teresa M. Kruisselbrink, Emily C. Lisi, et al. Clinically Actionable Findings Derived From Predictive Genomic Testing Offered in a Medical Practice Setting. Mayo Clinic Proceedings, 2021. https://doi.org/10.1016/j.mayocp.2020.08.051 ↩

  2. Jorune Balciuniene, Ruby Liu, Lora Bean, et al. At-Risk Genomic Findings for Pediatric-Onset Disorders From Genome Sequencing vs Medically Actionable Gene Panel in Proactive Screening of Newborns and Children. JAMA Network Open, 2023. https://doi.org/10.1001/jamanetworkopen.2023.26445 ↩ ↩2

  3. Amy C. Sturm, Rebecca Truty, Thomas E. Callis, et al. Limited-Variant Screening vs Comprehensive Genetic Testing for Familial Hypercholesterolemia Diagnosis. JAMA Cardiology, 2021. https://doi.org/10.1001/jamacardio.2021.1301 ↩ ↩2

  4. Vincent Timmerman, Alleene Strickland, Stephan Züchner. Genetics of Charcot-Marie-Tooth (CMT) Disease within the Frame of the Human Genome Project Success. Genes, 2014. https://doi.org/10.3390/genes5010013 ↩

  5. Christina A. Austin-Tse, Vaidehi Jobanputra, Denise L. Perry, et al. Best practices for the interpretation and reporting of clinical whole genome sequencing. npj Genomic Medicine, 2022. https://doi.org/10.1038/s41525-022-00295-z ↩

  6. Alexis N. Burian, Wufan Zhao, Te‐Wen Lo, et al. Genome sequencing guide: An introductory toolbox to whole‐genome analysis methods. Biochemistry and Molecular Biology Education, 2021. https://doi.org/10.1002/bmb.21561 ↩

  7. Anna Axakova, Amund H. Berger, Warren van Loggerenberg, et al. Systematic and proactive evaluation of AIRE missense variant effects. The American Journal of Human Genetics, 2026. https://doi.org/10.1016/j.ajhg.2026.07.008 ↩

  8. Marlena Szalata, Mikołaj Danielewski, Karolina Wielgus, et al. Why Should a Genome Be Protected? Ethical, Legal, and Security Challenges in the Protection of Genomic Data. Biology, 2026. https://doi.org/10.3390/biology15090726 ↩ ↩2

  9. Abhisikta Basu, Keya De Mukhopadhyay. Next-Generation Sequencing of Human Genomes: Ethico-Legal Challenges with Regard to Data Security. Lecture Notes in Networks and Systems, 2026. https://doi.org/10.1007/978-981-96-8632-2_9 ↩

  10. the members of the Public and Professional Policy Committee (PPPC) and Patenting and Licensing Committee (PLC), on behalf of the ESHG, Sirpa Soini, Ségolène Aymé, et al. Patenting and licensing in genetic testing: ethical, legal and social issues. European Journal of Human Genetics, 2008. https://doi.org/10.1038/ejhg.2008.37 ↩