Skip to content

What to Ask For From a Whole Genome Sequencing Service

Oak
A small feathered reptile specimen on a steel post, its plumage in tight overlapping layers with two bare patches, lit against black.

Buy whole-genome sequencing on four criteria: PCR-free library prep, ≥30x mean aligned coverage with a stated fraction of the genome at ≥20x, delivery of raw FASTQ or lossless CRAM plus a gVCF, and a lab you can email with a question about a specific position. Everything else (the portal, the “wellness report,” the ancestry map) is disposable. If a service will not hand you the reads, you are renting an interpretation, not buying data. WGS stands for whole-genome sequencing: short reads tiled across all ~3.1 billion bases rather than the ~1-2% of the genome that is protein-coding exons.

What the deliverables should be

Ask for this list by name before you pay:

  • FASTQ (gzipped, paired, _R1/_R2), or CRAM 3.0 aligned to a named reference with original quality scores retained. A 30x human genome is roughly 90-110 GB as gzipped FASTQ and 15-20 GB as CRAM. If the CRAM has been quality-binned or reference-compressed lossily, you cannot fully rederive the FASTQ, so ask which.
  • The exact reference build. GRCh38 with alts and decoys (the GCA_000001405.15_GRCh38_no_alt_analysis_set convention) is the common denominator. T2T-CHM13v2.0 resolves centromeres, segmental duplications, and the short arms of acrocentric chromosomes, but most annotation databases still key on GRCh38. Our preference: get GRCh38 as the primary, and realign to CHM13 yourself with bwa-mem2 if you care about hard regions.
  • gVCF, not just a filtered VCF. The gVCF records reference-confidence blocks, so you can distinguish “this position is reference” from “we had no reads here.” That distinction matters constantly.
  • A structural variant VCF from at least one caller (Manta or DRAGEN SV), plus a CNV call set.
  • Per-base or per-window depth as a mosdepth output or BigWig.
  • The QC report: duplicate rate, insert size distribution, contamination estimate, Ti/Tv, sex check.

If the vendor’s answer to “can I have the FASTQ” is “we can export a 23andMe-style text file,” walk. That file holds ~600k-1M genotyped positions. A 30x genome yields roughly 4-5 million SNVs and indels relative to GRCh38, plus thousands of structural variants.

Wet-lab parameters that change your results

PCR-free libraries. PCR amplification introduces GC bias that thins coverage in promoters and first exons, and it inflates duplicate rates. On a PCR-free 30x library you should see optical/PCR duplicates under about 5% and even coverage across GC fractions from 30-60%.

Read length and platform. Illumina 2x150bp is the default and is what almost every clinical pipeline is validated against. Sequencing chemistry has driven per-genome cost down by orders of magnitude since capillary sequencing, which is why 30x WGS is now a consumer-priced product at all 1. Long reads (PacBio HiFi, Oxford Nanopore) are a different instrument class: dramatically better for structural variants, repeat expansions, and phasing across tens of kilobases, worse per-base cost. If you are buying one genome and want maximum interpretability today, buy short reads at 30x. If you specifically want SV resolution or full-length phasing, buy HiFi at 15-20x instead, and know you are trading away some of the annotation ecosystem.

Coverage target. 30x mean is the floor for confident heterozygous calling. At 20x you start losing het sensitivity, particularly for indels. Mean coverage also hides the tail: ask what fraction of the genome, and of a clinically curated gene list, sits at ≥20x. Dewey and colleagues found that a meaningful fraction of inherited-disease genes went inadequately covered even on nominally complete genomes, and that indel concordance between sequencing platforms was poor 2. Nothing since has made that problem vanish.

QC you should run on your own data

Do not trust the vendor’s PDF. Run this yourself, it takes an hour:

# depth and evenness
mosdepth --by 1000 --fast-mode sample sample.cram

# contamination: FREEMIX should be < 0.02
verifybamid2 --SVDPrefix resource/1000g.phase3.100k.b38.vcf.gz.dat \
  --Reference GRCh38.fa --BamFile sample.cram

# sex, relatedness, ancestry sanity check
somalier extract -d extracted/ --sites sites.hg38.vcf.gz -f GRCh38.fa sample.cram
somalier relate extracted/*.somalier

# variant-level sanity
bcftools stats -F GRCh38.fa sample.vcf.gz | grep -E "^SN|^TSTV"

Expected values for a good 30x human genome: FREEMIX < 0.02, genome-wide Ti/Tv near 2.0-2.1 (exonic subset closer to 3.0), het/hom-alt ratio around 1.4-2.0 depending on ancestry, median insert size 300-450bp. A Ti/Tv of 1.6 means your variant set is full of false positives and someone skipped filtering. Het/hom ratios far outside the range usually mean contamination or a sample swap.

The pipeline we would run

Align with bwa-mem2 mem -K 100000000 -Y (the -K fixes chunk size so output is deterministic), sort and mark duplicates with samtools and sambamba or GATK MarkDuplicatesSpark. Call small variants with DeepVariant (--model_type=WGS) rather than GATK HaplotypeCaller if you only care about accuracy on a single sample; DeepVariant’s indel precision on PCR-free 30x Illumina is better out of the box and requires no BQSR. Call SVs with Manta plus smoove, and CNVs with CNVpytor on read depth. Phase with WhatsHap, then extend statistically with SHAPEIT5 against a reference panel if you want haplotypes for imputation or PGS.

Annotate with Ensembl VEP, offline, cached:

vep -i sample.vcf.gz --cache --offline --assembly GRCh38 \
  --everything --vcf --fork 8 \
  --plugin CADD,cadd.tsv.gz --plugin SpliceAI,spliceai.vcf.gz \
  --custom gnomad.v4.genomes.sites.vcf.gz,gnomADg,vcf,exact,0,AF,nhomalt \
  --custom clinvar.vcf.gz,ClinVar,vcf,exact,0,CLNSIG,CLNREVSTAT

Then filter hard. Population frequency is the single most effective filter: a variant at 3% allele frequency in gnomAD is not the cause of a rare disease. Rank by ClinVar review status (CLNREVSTAT), not just CLNSIG, because a large share of single-submitter “pathogenic” entries do not survive reanalysis. Keep pharmacogenomic star-allele calling separate: run PharmCAT or PyPGx, because CYP2D6 in particular has copy-number and hybrid alleles that a generic VCF pipeline will silently get wrong.

Interpretation is the expensive part, not sequencing 3. Published best-practice guidance for clinical WGS exists precisely because turning 4.5 million variants into a defensible statement about one person requires structured evidence review, not a scoring script 4. If anything in your results touches a clinical question (a ClinVar pathogenic variant in a cancer predisposition or cardiomyopathy gene, a suspected repeat expansion, a variant in a gene on the ACMG secondary findings list), that goes to a clinical geneticist or genetic counselor working from a CLIA/CAP-validated report. A research-grade VCF is not a diagnosis, and a variant called once at 30x has not been confirmed.

What WGS does not tell you

WGS is a static readout of germline sequence. It does not measure whether a gene is being expressed, what your proteins are doing, or what your glucose did after lunch. It is also incomplete in specific, known ways: short reads struggle in segmental duplications and homologous pseudogene regions (SMN1/SMN2, PMS2, CYP21A2), tandem repeat expansions need targeted callers like ExpansionHunter, and mitochondrial heteroplasmy needs its own pipeline. Interpretation is also time-dependent, since variant classifications change as evidence accumulates, which is why reanalysis of an old genome yields new findings 5. And most of the genome is non-coding, where our ability to predict functional consequence remains weak compared to coding variants 6.

Questions people also ask

Is WGS worth it? As a data purchase for someone who will analyze it, yes, once, at 30x PCR-free with raw files delivered. The genome does not change, so it is a one-time cost against a lifetime of reanalysis. As a medical purchase for an asymptomatic adult, the evidence on cost-effectiveness is thinner than marketing implies, with published economic evaluations varying widely in method and conclusion 7. In a diagnostic setting with a suspected genetic condition, WGS has a much clearer case and is increasingly a first-line test 3.

What’s the difference between WGS and WES? Whole-exome sequencing captures only protein-coding exons, roughly 1-2% of the genome, using hybridization probes. WGS sequences everything without capture, so it has more uniform coverage, better detection of structural and copy-number variants, and access to non-coding and mitochondrial sequence. WES is cheaper per sample and produces smaller files. Given how much 30x WGS costs now, we would buy the genome.

What does the acronym WGS stand for? Whole-genome sequencing. Note that the same acronym is used for bacterial and viral genome sequencing in public health labs, which is why generic searches return outbreak-tracking pages.

Who should see my results? Anything with clinical implications goes to a clinician. Professional guidance has been explicit for over a decade that whole-genome sequencing in a health-care context requires pre- and post-test counseling and a policy for unsolicited findings 8.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Sang Tae Park, Jayoung Kim. Trends in Next-Generation Sequencing and a New Era for Whole Genome Sequencing. International Neurourology Journal, 2016. https://doi.org/10.5213/inj.1632742.371 ↩

  2. Frederick E. Dewey, Megan E. Grove, Cuiping Pan, et al. Clinical Interpretation and Implications of Whole-Genome Sequencing. JAMA, 2014. https://doi.org/10.1001/jama.2014.1717 ↩

  3. Frederik Otzen Bagger, Line Borgwardt, Andreas Sand Jespersen, et al. Whole genome sequencing in clinical practice. BMC Medical Genomics, 2024. https://doi.org/10.1186/s12920-024-01795-w ↩ ↩2

  4. Christina Austin‐Tse, Vaidehi Jobanputra, Denise Perry, et al. Best practices for the interpretation and reporting of clinical whole genome sequencing. npj Genomic Medicine, 2022. https://doi.org/10.1038/s41525-022-00295-z ↩

  5. Petar Brlek, Luka Bulić, Matea Bračić, et al. Implementing Whole Genome Sequencing (WGS) in Clinical Practice: Advantages, Challenges, and Future Perspectives. Cells, 2024. https://doi.org/10.3390/cells13060504 ↩

  6. Tuuli Lappalainen, Alexandra J. Scott, Margot Brandt, et al. Genomic Analysis in the Age of Human Genome Sequencing. Cell, 2019. https://doi.org/10.1016/j.cell.2019.02.032 ↩

  7. Katharina Schwarze, James Buchanan, Jenny C. Taylor, et al. Are whole-exome and whole-genome sequencing approaches cost-effective? A systematic review of the literature. Genetics in Medicine, 2018. https://doi.org/10.1038/gim.2017.247 ↩

  8. Carla van El, Martina C. Cornel, Pascal Borry, et al. Whole-genome sequencing in health care. European Journal of Human Genetics, 2013. https://doi.org/10.1038/ejhg.2013.46 ↩