What to look for in a genome kit
If you want to work with your own molecular data, the only genome kit worth buying is 30x (or deeper) PCR-free whole-genome sequencing on an Illumina instrument, delivered as FASTQ or CRAM plus a per-sample gVCF, with no restriction on downloading and re-analyzing the raw files. Everything else on the shelf at a pharmacy is a genotyping array: roughly 600,000 to 900,000 pre-chosen positions out of 3.1 billion, sold with a consumer report attached. Arrays are fine for ancestry and for a handful of well-studied variants. They cannot find anything rare, and rare is where most of the interesting personal signal lives.
What you are buying is a file, not a report
Judge a kit by the artifacts it hands you. The minimum set:
- Raw reads: paired FASTQ (
_R1.fastq.gz,_R2.fastq.gz), roughly 2 × 50–60 GB at 30x with 150 bp reads. Or CRAM aligned to a named reference, ~15–20 GB per genome. - Alignment: CRAM/BAM plus index (
.crai/.bai), with the reference build stated explicitly (GRCh38 with ALT contigs is the practical default; T2T-CHM13 is better for segmental duplications but has thinner tooling and annotation support). - Small variants: a gVCF and a filtered VCF. A single 30x human genome against GRCh38 yields on the order of 4–5 million variants, mostly SNVs with several hundred thousand indels.
- QC: the caller’s metrics files, or at least the flagstat and coverage summaries.
If a vendor will only give you a PDF and a “raw data” download that turns out to be a 15 MB tab-separated array export, you bought a different product.
The spec sheet that matters
Coverage. “30x” should mean mean mapped, deduplicated depth on the autosomes, not raw yield divided by 3.1 Gb. Ask which one they quote. The practical target is ≥90% of the callable genome at ≥20x. Below about 20x mean depth, heterozygote sensitivity drops fast because you start seeing 0-of-8-read support for the alternate allele. Deep coverage is also what makes singleton and low-frequency variant discovery meaningful at population scale 1.
Library prep. PCR-free if you can get it. PCR amplification introduces duplicate reads, GC bias that thins coverage in high-GC promoters, and indel stutter in homopolymers. PCR-free needs more input DNA (typically ≥500 ng of high molecular weight), which in practice means a blood draw rather than a cheek swab.
Read length. 2 × 150 bp is standard. Anything shorter hurts mapping in repeats. If a vendor offers long reads (ONT or PacBio HiFi) at 15–30x, that is a genuinely different and better product for structural variants and phasing, at two to four times the price.
Reference and pipeline. You want to know the aligner and caller: BWA-MEM2 or DRAGEN for alignment, DeepVariant or GATK HaplotypeCaller for small variants. This matters because you will eventually re-call from the CRAM yourself, and you need to reproduce or beat their baseline.
Data rights. Read the terms. Confirm you can download everything, that deletion is honored, and that re-identification risk is disclosed. A whole genome is not anonymizable and carries information about relatives who never consented, which is the substance of most of the ethics literature on consumer WGS 2.
Saliva, swab, or blood
Sample type is the failure mode nobody advertises. Saliva and buccal kits collect human epithelial cells along with oral bacteria. Non-human DNA fractions of 10–30% are common, and we have seen worse from people who ate before collecting. The vendor sequences the tube, so a 30x order can land as 22x of human coverage. Blood (EDTA tube, leukocyte DNA) is cleaner, gives higher molecular weight DNA suitable for PCR-free prep, and is what we would choose every time.
If you are stuck with saliva: no food, drink, or tobacco for 30 minutes, collect first thing in the morning, fill to the line, cap firmly so the stabilizing buffer mixes.
QC to run the day the files land
Do not trust the delivery report. Five checks, in order:
# 1. Integrity
md5sum -c checksums.md5
# 2. Mapping rate and duplicates
samtools flagstat -@ 8 sample.cram > flagstat.txt
# expect >99% mapped, >98% properly paired,
# duplicates <2% for PCR-free, <10% for PCR+
# 3. Coverage distribution
mosdepth --by 1000 --fast-mode -t 4 sample sample.cram
# check sample.mosdepth.summary.txt: autosomal mean,
# and the global distribution for fraction >= 20x
# 4. Contamination / sample mixture
verifyBamID2 --SVDPrefix resource/1000g.phase3.100k.b38.vcf.gz.dat \
--Reference GRCh38.fa --BamFile sample.cram
# FREEMIX above ~0.02 means a mixed or swapped sample
# 5. Sanity on the variant file
bcftools stats sample.vcf.gz | grep -E "^SN|^TSTV"
# ~4-5M records, Ts/Tv around 2.0-2.1 genome-wide
A Ts/Tv well below 2.0 means the filtering is loose and you are looking at sequencing noise. Chromosome X and Y depth ratios are a free sample-identity check: compare chrX mean depth to autosomal mean, and confirm it matches what you expect.
What the genome will and will not tell you
Your germline genome is fixed at conception and is the cheapest layer per unit of information, because you sequence once and re-annotate forever as databases improve. That is the real argument for buying WGS rather than an array: the array is frozen at the set of probes the manufacturer chose in a particular year, while the genome can be re-queried against next year’s ClinVar.
The limits are worth stating plainly. Interpretation is bounded by reference panels, and those panels are still skewed toward European ancestry. Population-specific sequencing efforts keep finding large numbers of variants absent from existing catalogs, which is exactly why allele frequency estimates for your ancestry may be wrong or missing 3. The original common-variant maps built the scaffolding for association studies, but common SNPs explain only part of the picture and polygenic scores transfer poorly across populations 4.
Also, the genome is static and you are not. It will not tell you what your immune system is doing this month, how your lipids are trending, or what your glucose does after dinner. Those need RNA, protein, blood chemistry, and CGM. If you are considering a methylation array as an add-on, know that the EPIC BeadChip covers a fixed probe set with known cross-reactive and polymorphism-affected probes that need filtering before you interpret anything 5.
One line that is not negotiable: a research-grade WGS result is not a clinical diagnosis. If a variant in your file looks pathogenic, the next step is a confirmatory test in a CLIA-certified lab ordered through a physician or genetic counselor, not a decision about medication or surveillance made from a VCF on your laptop.
Questions people also ask
Which WGS test is best? The one that gives you ≥30x PCR-free short-read sequencing from a blood sample, FASTQ or CRAM plus gVCF, GRCh38 with a named pipeline, and unrestricted download. Rank vendors on those five attributes, not on the number of “traits” in their report.
Can you buy a DNA kit at Walgreens or CVS? Yes. Retail pharmacies stock saliva collection kits for ancestry and paternity, and those are genotyping arrays or short STR panels. They are not whole-genome sequencing and will not produce a file you can align or re-call.
How much does a genome test cost? Consumer arrays run $50–$150. Consumer 30x WGS is roughly $300–$900 depending on promotion and whether it is PCR-free. Clinical diagnostic WGS ordered through a physician, with interpretation and CLIA confirmation, is typically several thousand dollars.
What is the most accurate at-home test? Accuracy depends on the assay, not on where you spit. For SNVs and small indels, 30x short-read WGS from blood with a PCR-free library is the most accurate broadly available option. Long-read sequencing is more accurate for structural variants and repeat expansions.
Do I need a clinician? For any finding you would act on, yes. Use your own analysis to generate questions and bring the file to a genetic counselor or physician who can order confirmatory testing.
Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Footnotes
-
Masao Nagasaki, Jun Yasuda, Fumiki Katsuoka, et al. Rare variant discovery by deep whole-genome sequencing of 1,070 Japanese individuals. Nature Communications, 2015. https://doi.org/10.1038/ncomms9018 ↩
-
Wim Pinxten, Heidi Howard. Ethical issues raised by whole genome sequencing. Best Practice & Research Clinical Gastroenterology, 2014. https://doi.org/10.1016/j.bpg.2014.02.004 ↩
-
Israel Aguilar-Ordóñez, Eugenio Guzman-Cerezo, David Torres-Treviño, et al. Whole genome sequencing of 1427 Mexican individuals from the oriGen cohort. Nature Communications, 2026. https://doi.org/10.1038/s41467-026-77389-0 ↩
-
The International SNP Map Working Group, Cold Spring Harbor Laboratories:, Ravi Sachidanandam, et al. A map of human genome sequence variation containing 1.42 million single nucleotide polymorphisms. Nature, 2001. https://doi.org/10.1038/35057149 ↩
-
Ruth Pidsley, Elena Zotenko, Timothy J. Peters, et al. Critical evaluation of the Illumina MethylationEPIC BeadChip microarray for whole-genome DNA methylation profiling. Genome biology, 2016. https://doi.org/10.1186/s13059-016-1066-1 ↩