The Best Health DNA Test Is Whole-Genome Sequencing With Raw Data Access
If you want a DNA test for health and you are comfortable working with data, buy whole-genome sequencing at 30x coverage or higher, delivered as FASTQ and BAM plus a called VCF, from a provider that lets you download all of it. Everything else in the consumer market, including 23andMe, AncestryDNA, and the “comprehensive” kits that advertise hundreds of traits, is a genotyping array: a fixed panel of roughly 600,000 to 750,000 pre-selected positions out of three billion base pairs. An array is a good instrument for the question it was designed to answer (which of these known common variants do you carry) and a poor instrument for almost any other genetic question, because a position that is not on the chip does not exist as far as the test is concerned.
The distinction matters more than marketing copy suggests. Arrays report common single-nucleotide polymorphisms with high per-probe accuracy, but rare variants, which are where most of the clinically meaningful signal in a single person’s genome lives, are either absent from the panel or poorly genotyped when present. Array probes for rare pathogenic variants have a well-documented false-positive problem, which is why array-based reports carry a confirmatory-testing instruction. Sequencing takes a different approach: read the whole thing, then decide afterward what to ask of it. That flexibility is the main reason to pay for it.
What the three tiers of testing measure
A genotyping array hybridizes your fragmented DNA to oligonucleotide probes fixed on a chip. The output is a set of genotype calls at designed positions, typically shipped as a tab-delimited text file with rsID, chromosome, position, and a two-letter genotype. Nothing is measured between the probes. Consumer “polygenic risk” and nutrition reports are computed from a few dozen of these calls each, often on variants with odds ratios near 1.1.
Whole-exome sequencing captures the roughly 1 to 2 percent of the genome that codes for protein, which is where a large majority of known Mendelian disease variants sit. It is cheaper than a genome and reads at high depth, but capture is uneven, some exons drop out reproducibly, and it sees almost nothing of the regulatory genome or structural variation.
Whole-genome sequencing shotgun-sequences the entire genome, typically as 150 bp paired-end reads on an Illumina instrument, at an average depth of 30x or more. Depth is the number of independent reads covering an average base, and 30x is the conventional floor for confident diploid genotyping because it keeps the chance of missing a heterozygous allele acceptably low across most of the genome 1. A genome covers coding regions more evenly than exome capture does, detects copy-number and structural variants, resolves mitochondrial DNA at high depth as a by-product, and can be reanalyzed against next year’s variant databases without a new sample 2. As a diagnostic instrument in clinical settings, genome sequencing has displaced staged panel testing for exactly these reasons 3.
If a sales page says “99.9% analytical accuracy” without naming the platform, the depth, or the reference genome build, assume an array until told otherwise. Ask three questions before buying: Is this sequencing or genotyping? What is the mean depth and what fraction of the genome is covered at 20x or better? Do I get FASTQ and BAM, or only a PDF?
Files you should insist on, and what to do with them
The deliverable determines what you can do later. We would not buy a test that does not hand over all four of these.
FASTQ files hold the raw reads and base quality scores, roughly 100 GB for a 30x human genome as gzipped paired files. Keep them. They are the only artifact that lets you realign to a newer reference (GRCh38 today, T2T-CHM13 increasingly) or rerun with a different caller when the tooling improves.
A BAM or CRAM file is the alignment against a named reference. CRAM is reference-compressed and runs about a third the size, which matters when you are paying for storage. Check the header before anything else:
samtools view -H sample.cram | grep -E '@RG|@PG|AS:'
samtools coverage sample.cram
mosdepth --by 1000 --fast-mode sample sample.cram
samtools coverage gives per-chromosome mean depth, and mosdepth gives you the distribution so you can see what fraction of callable bases sit above 20x. Coverage below roughly 20x in a region means genotypes there are provisional.
The VCF is the variant call file, listing positions where you differ from the reference, with genotype, depth, and quality annotations. A 30x human genome yields roughly 4 to 5 million variants, the overwhelming majority of them common and benign. Ask whether a gVCF is included, since it records confident reference calls too and tells you the difference between “you don’t carry this variant” and “we couldn’t see this position”.
Annotation is where the work is. We use Ensembl VEP with the --everything flag plus the gnomAD and ClinVar plugins, or bcftools csq for fast consequence calls. Filtering the file down to something a person can read is mostly a population-frequency exercise:
bcftools view -i 'FILTER="PASS" && FMT/DP>=20' sample.vcf.gz \
| bcftools +split-vep -c gnomADg_AF:Float,CLIN_SIG,IMPACT \
| bcftools view -i 'gnomADg_AF<0.001 & IMPACT="HIGH"'
That returns a few hundred candidates, and this is the point where the interpretive difficulty begins rather than ends. Most rare high-impact variants in a healthy person’s genome are not causing disease 4. Pathogenicity requires evidence of segregation, functional data, or curated clinical assertions, and even ClinVar submissions conflict. Sequencing healthy individuals reliably produces variants of uncertain significance, and the rate at which incidental findings turn out to be actionable is low 5. Any variant you are tempted to act on needs confirmation in a clinical laboratory and interpretation by a genetics professional, because orthogonal confirmation and family context change the conclusion often enough to matter.
Where DNA alone stops being useful
A genome is a fixed text, and most questions about your health concern present state. Nutrition and fitness reports built on a handful of common variants (MTHFR, ACTN3, CYP1A2, FTO) have small effect sizes and poor predictive value for an individual, whatever the interface suggests. If you want to know your vitamin D status, measure 25-hydroxyvitamin D. If you want to know how you respond to a given carbohydrate load, wear a continuous glucose monitor. Genotype sets priors, and measurement of transcripts, proteins, and metabolites tells you what is happening now.
That is the boundary of DNA testing for health, and it has been the boundary since the early expectations of the genome project were set out 6. Sequence is most valuable where penetrance is high and monitoring is well established: pathogenic variants in cancer-predisposition genes, hereditary cardiac conditions, pharmacogenes with strong prescribing evidence, and recessive carrier status relevant to reproductive planning 7. Those are also exactly the domains where the report should send you to a clinician rather than to a supplement page.
Two practical notes on process. Sample type affects yield and quality, and buccal swabs and saliva vary considerably in human DNA content across extraction protocols, which is one reason blood remains preferable for sequencing when logistics permit 8. Mitochondrial genome coverage comes free with whole-genome data at depths in the thousands, and heteroplasmy quantification is possible from the same BAM, though calling low-level heteroplasmy well requires attention to alignment artifacts around the mtDNA control region 9.
Questions people also ask
How much does a DNA health analysis cost? Genotyping arrays run $100 to $200. Clinical-grade panels for specific gene sets cost several hundred to a few thousand dollars and are usually ordered by a physician. Consumer whole-genome sequencing at 30x with raw data typically lands between $400 and $1,500 depending on turnaround and whether interpretation is bundled. Costs have fallen steeply and continue to, which is part of why national sequencing programs are now feasible at scale 10.
Does AncestryDNA show diseases? No. AncestryDNA reports ancestry composition and relative matching, and its health-adjacent offerings have been limited and intermittent. You can download the raw array file and annotate it yourself, but you are still limited to the positions on the chip.
What are the downsides? Variants of uncertain significance that generate anxiety without action, false positives from array probes, incidental findings you did not consent to think about, privacy exposure through terms of service and corporate bankruptcy, and the tendency of a genetic result to feel more deterministic than it is.
Is there a test that doesn’t sell my information? Look for providers that do not run a research-consent business model, allow deletion of both data and biological sample, and let you export everything so your copy is not dependent on their continued existence. Read the secondary-use clause, not the privacy summary.
How accurate are DNA tests for health? Analytically, a 30x genome calls single-nucleotide variants in well-covered regions with accuracy above 99.5 percent, with structural variants and repeat expansions considerably worse. Clinically, accuracy is about interpretation, and the same variant file can support very different conclusions depending on the evidence base applied to it 4.
Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Footnotes
-
Rachel L Goldfeder, Dennis P Wall, Muin J Khoury, et al. Human Genome Sequencing at the Population Scale: A Primer on High-Throughput DNA Sequencing and Analysis. American Journal of Epidemiology, 2017. https://doi.org/10.1093/aje/kww224 ↩
-
Hans Lehrach. DNA sequencing methods in human genetics and disease research. F1000Prime Reports, 2013. https://doi.org/10.12703/p5-34 ↩
-
Gregory Costain, Ronald D. Cohn, Stephen W. Scherer, et al. Genome sequencing as a diagnostic test. Canadian Medical Association Journal, 2021. https://doi.org/10.1503/cmaj.210549 ↩
-
David B. Goldstein, Andrew Allen, Jonathan Keebler, et al. Sequencing studies in human genetics: design and interpretation. Nature Reviews Genetics, 2013. https://doi.org/10.1038/nrg3455 ↩ ↩2
-
Noralane M. Lindor, Stephen N. Thibodeau, Wylie Burke. Whole-Genome Sequencing in Healthy People. Mayo Clinic Proceedings, 2017. https://doi.org/10.1016/j.mayocp.2016.10.019 ↩
-
Francis S. Collins. Medical and Societal Consequences of the Human Genome Project. New England Journal of Medicine, 1999. https://doi.org/10.1056/nejm199907013410106 ↩
-
Katherine A. Johansen Taber, Barry D. Dickinson, Modena Wilson. The Promise and Challenges of Next-Generation Genome Sequencing for Clinical Care. JAMA Internal Medicine, 2014. https://doi.org/10.1001/jamainternmed.2013.12048 ↩
-
Latifa El Bali, Aurélie Diman, Alfred Bernard, et al. Comparative Study of Seven Commercial Kits for Human DNA Extraction from Urine Samples Suitable for DNA Biomarker-Based Public Health Studies. Journal of Biomolecular Techniques : JBT, 2014. https://doi.org/10.7171/jbt.14-2504-002 ↩
-
Yue Yao, Motoi Nishimura, Kei Murayama, et al. A simple method for sequencing the whole human mitochondrial genome directly from samples and its application to genetic testing. Scientific Reports, 2019. https://doi.org/10.1038/s41598-019-53449-y ↩
-
Annie T. W. Chu, Amy H. Y. Tong, Desiree M. S. Tse, et al. The Hong Kong genome project: building genome sequencing capacity and capability for advancing genomic science in Hong Kong. Journal of Translational Genetics and Genomics, 2023. https://doi.org/10.20517/jtgg.2023.22 ↩