Skip to content

What Genetic Testing for Wellness Tells You

Oak
A giant tree stump with glowing continuous rings beside a smaller stump lit by only a few scattered blue points, in a misty glowing forest.

Genetic testing for wellness is worth doing once, at whole-genome depth, for a narrow set of findings: carrier status, pharmacogenomic haplotypes, and a short list of actionable dominant disease genes with real penetrance. The “wellness” panels that sell you caffeine metabolism, muscle fiber type, and optimal macronutrient ratio are reporting single SNPs with effect sizes too small to change any decision you would make. The sequencing is cheap and accurate. The interpretation layer is where almost all consumer products fail, and it fails in a predictable direction: it converts weak associations into confident-sounding instructions.

If you want to work with your own data rather than read someone’s PDF, the practical question is which files you get and whether you can reprocess them. That determines everything downstream.

The difference between an array and a genome

A $99 consumer DNA test is a genotyping array. It interrogates roughly 600,000 to 1,000,000 pre-selected positions with allele-specific probes. It does not read your genome. It answers yes/no at sites the chip designer chose, mostly common variants with population frequency above 1%, selected for ancestry inference and GWAS tag coverage.

Consequences you should internalize before buying anything:

  • Array calls at rare variants have poor positive predictive value. A pathogenic BRCA1 variant with allele frequency 1 in 20,000 that the chip happens to probe will produce false positives at a rate that swamps true positives, because the probe error rate is on the order of 0.1% and the prior is 0.005%. This is why every reputable lab requires confirmatory Sanger or targeted NGS before a clinical report.
  • Arrays cannot call indels reliably, cannot resolve copy number at exon resolution, and cannot phase haplotypes without imputation.
  • Imputed genotypes, which is what most third-party interpretation tools are quietly using, are statistical estimates conditioned on a reference panel. They are fine for polygenic score arithmetic and useless for single-variant clinical questions.

Whole-genome sequencing at 30x mean coverage reads the actual bases. You get roughly 4 to 5 million variants relative to GRCh38, of which about 4 to 5 million are SNVs and indels and a few thousand are structural. Short-read WGS at 30x reaches high sensitivity for SNVs in mappable regions, and the clinical-grade pipelines that hit those numbers are well documented 1. It still misses things: GC-extreme regions, segmental duplications, repeat expansions like the FMR1 CGG tract and HTT CAG tract, and the pseudogene-shadowed exons of PMS2 and SMN1. If you have a family history pointing at one of those, short-read WGS is the wrong instrument and you need a targeted assay ordered through a clinician.

What a healthy person’s genome returns

The MedSeq Project sequenced healthy adults and reported the yield. Roughly one in five participants carried a monogenic disease variant, and reanalysis over time changed classifications in a meaningful fraction of cases 2. That last point matters more than the headline number: a variant called “likely pathogenic” in 2019 may be reclassified to VUS or benign by 2026 as ClinVar accumulates evidence. Your report has a shelf life. Your VCF does not.

The broader literature on sequencing healthy people is cautious for good reason. Lindor, Thibodeau, and Burke laid out the core problem: in an unselected person, the prior probability of any specific pathogenic variant is low, penetrance estimates derived from ascertained families overstate risk in the general population, and the dominant output of the test is variants of uncertain significance 3. Population screening programs implemented in primary care with large panels confirm the operational shape of this: most findings are carrier states and pharmacogenomic variants, actionable dominant findings are a few percent, and the bottleneck is clinician time for follow-up rather than sequencing capacity 4.

The early argument for integrating WGS into wellness care was built on combining genome data with longitudinal phenotype, not on the genome alone 5. That is still the right framing. A genome is a static prior. Your blood panel, your proteomics, and your glucose trace are the evidence. The genome tells you where to look harder.

The report you get is one interpretation of many

Same FASTQ, different pipelines, different output. The variance is real and it comes from specific choices:

  • Reference build. GRCh37 versus GRCh38 changes coordinates and fixes several hundred assembly errors. Liftover between them is lossy at problem loci.
  • Aligner and caller. BWA-MEM plus GATK HaplotypeCaller in GVCF mode is the conventional baseline. DeepVariant generally gives better indel precision on 30x Illumina data. Different callers disagree on thousands of indels per genome.
  • Annotation. VEP with --everything --assembly GRCh38 --cache versus ANNOVAR versus snpEff will pick different canonical transcripts and therefore different HGVS nomenclature for the same variant. Use --pick_allele_gene deliberately and know which transcript set you are on (MANE Select is the current sensible default).
  • Classification. ACMG criteria applied by two analysts to the same variant diverge often enough that ClinVar shows conflicting submissions on thousands of entries.

Platform-level tooling for this exists and is improving. Recent work on integrated WGS analysis for both clinical and wellness use describes the pieces you need in one place: variant calling, ACMG-oriented classification, pharmacogenomic haplotype calling, and polygenic scoring 6. If you are building your own stack, that is a reasonable blueprint.

The minimum file set to insist on: the BAM or CRAM aligned to GRCh38 (CRAM at roughly 15 to 20 GB for a 30x genome, BAM closer to 60 to 90 GB), the gVCF, a hard-filtered VCF, and the FASTQs if you can get them. With the BAM you can re-call anything later. With only a PDF you own nothing.

A concrete starting point once you have the VCF:

# normalize and left-align against the reference
bcftools norm -f GRCh38.fa -m -both -Oz -o sample.norm.vcf.gz sample.vcf.gz
bcftools index sample.norm.vcf.gz

# annotate
vep -i sample.norm.vcf.gz --cache --assembly GRCh38 \
    --everything --pick_allele_gene --vcf \
    --plugin CADD,whole_genome_SNVs.tsv.gz \
    -o sample.vep.vcf

# pharmacogenomic star alleles (needs the BAM, not just the VCF)
pharmcat_vcf_preprocessor -vcf sample.norm.vcf.gz -refFna GRCh38.fa
java -jar pharmcat.jar -vcf sample.preprocessed.vcf -reporterJson

PharmCAT output is one of the few genuinely high-value wellness returns, because haplotypes in CYP2C19, CYP2D6, DPYD, TPMT, and SLCO1B1 have CPIC guidelines behind them with graded evidence. Do not act on a pharmacogenomic result yourself. Bring the report to the clinician who prescribes for you. Star-allele calling for CYP2D6 in particular depends on copy-number and hybrid-allele resolution that short reads handle imperfectly.

Where wellness genomics goes wrong

The failure mode is not bad sequencing. It is a product incentive to return findings. Fiala, Taher, and Diamandis catalogued the harms of wellness testing initiatives generally: cascades of follow-up imaging and labs triggered by findings with no demonstrated benefit, cost shifted onto the health system, and anxiety in people who were well 7. Sequencing amplifies this because the per-person VUS count is high and every VUS is an invitation to investigate.

There is also the question of what the result is worth to you, which economics has trouble measuring. Standard cost-effectiveness frameworks capture health outcomes and miss the things people buy sequencing for: reproductive planning information, the value of knowing, the end of diagnostic uncertainty 8. If you want the information for its own sake, that is a legitimate reason. Just be clear that is the reason, and do not let a report convert it into a medical plan. The ethical and legal issues around genomic testing across the life cycle, including insurance implications and what you learn about relatives who did not consent, are documented and worth reading before you sequence 9.

The consumer industry’s older version of this problem was ancestry-flavored overreach, and the critique from that era, that the market sells statistical certainty it cannot support, has held up 10.

What we would do

Sequence once at 30x, keep the CRAM and gVCF, run your own annotation and re-run it yearly against fresh ClinVar and gnomAD releases. Read out three things: carrier status for recessive conditions relevant to reproductive planning, PharmCAT haplotypes, and the ACMG secondary findings gene list. Treat polygenic scores as interesting and poorly calibrated outside the ancestry they were trained in. Ignore the trait panel.

Then spend your attention on longitudinal measurement, because that is where change is visible. A genome you sequenced in 2026 reads the same in 2036. Your apoB, HbA1c, and inflammatory proteomics do not.

Anything that looks like a pathogenic finding in a dominant disease gene requires a genetic counselor and confirmatory clinical-grade testing before you believe it. Research-grade WGS is not a diagnosis.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Neil A. Miller, Emily G. Farrow, Margaret Gibson, et al. A 26-hour system of highly sensitive whole genome sequencing for emergency management of genetic diseases. Genome Medicine, 2015. https://doi.org/10.1186/s13073-015-0221-8 ↩

  2. Kalotina Machini, Ozge Ceyhan-Birsoy, Danielle R. Azzariti, et al. Analyzing and Reanalyzing the Genome: Findings from the MedSeq Project. The American Journal of Human Genetics, 2019. https://doi.org/10.1016/j.ajhg.2019.05.017 ↩

  3. Noralane M. Lindor, Stephen N. Thibodeau, Wylie Burke. Whole-Genome Sequencing in Healthy People. Mayo Clinic Proceedings, 2017. https://doi.org/10.1016/j.mayocp.2016.10.019 ↩

  4. Robert S. Wildin, Christine A. Giummo, Aaron W. Reiter, et al. Primary Care Implementation of Genomic Population Health Screening Using a Large Gene Sequencing Panel. Frontiers in Genetics, 2022. https://doi.org/10.3389/fgene.2022.867334 ↩

  5. Chirag J Patel, Ambily Sivadas, Rubina Tabassum, et al. Whole genome sequencing in support of wellness and health maintenance. Genome Medicine, 2013. https://doi.org/10.1186/gm462 ↩

  6. Xiya Song, Xinmeng Liao, Emre Green, et al. GenRiskPro: A Comprehensive Whole-Genome Sequencing Analysis Platform for Clinical and Wellness Applications. Computational and Structural Biotechnology Journal, 2026. https://doi.org/10.34133/csbj.0011 ↩

  7. Clare Fiala, Jennifer Taher, Eleftherios P. Diamandis. Benefits and harms of wellness initiatives. Clinical Chemistry and Laboratory Medicine (CCLM), 2019. https://doi.org/10.1515/cclm-2019-0122 ↩

  8. Dean A. Regier, Deirdre Weymann, James Buchanan, et al. Valuation of Health and Nonhealth Outcomes from Next-Generation Sequencing: Approaches, Challenges, and Solutions. Value in Health, 2018. https://doi.org/10.1016/j.jval.2018.06.010 ↩

  9. Gemma A. Bilkey, Belinda L. Burns, Emily P. Coles, et al. Genomic Testing for Human Health and Disease Across the Life Cycle: Applications and Ethical, Legal, and Social Challenges. Frontiers in Public Health, 2019. https://doi.org/10.3389/fpubh.2019.00040 ↩

  10. Hans‐Jürgen Bandelt, Yong‐Gang Yao, Martin B. Richards, et al. The brave new era of human genetic testing. BioEssays, 2008. https://doi.org/10.1002/bies.20837 ↩