Skip to content

What WES Testing Is and When It Falls Short

Oak
A long coral strand under dark water with only scattered glowing segments, gathered by drifting jellyfish with fine filaments.

Whole exome sequencing (WES) is a DNA test that sequences only the protein-coding portion of your genome, roughly 1-2% of the 3.1 billion bases, plus a small margin around splice junctions. It works by hybridization capture: biotinylated RNA or DNA baits are hybridized to your fragmented library, pulled down with streptavidin beads, and everything not captured is washed away. The result is deep coverage (typically 100x mean) over about 30-60 Mb of target, and essentially nothing elsewhere. It finds single-nucleotide variants and small indels in coding sequence well, finds structural and copy-number variants poorly, and by construction cannot see the 98% of your genome that sits outside the baits.

That design choice is the whole story. Everything useful and everything frustrating about WES follows from it.

What the test indicates

Diagnostically, WES is the standard first-tier test for suspected monogenic disease: a child with an unexplained syndrome, a family with a pattern of early-onset disease, a fetus with structural anomalies on ultrasound. In a prospective cohort of infants with suspected monogenic disorders, exome sequencing as a first-tier test produced a diagnosis in about 57% of cases, versus 14% with standard investigations.1 In fetuses with structural anomalies and normal karyotype and microarray, WES gave a molecular diagnosis in 10.3% of cases in a prospective study of 234 fetus-parent trios.2

Those numbers depend heavily on phenotype. A well-phenotyped infant with a clear syndromic presentation is a different population from an adult who wants to know about their genome. If you have no clinical indication, the base rate of finding something actionable is low, and the rate of finding variants of uncertain significance is high. Interpretation of any result in a disease context belongs with a clinical geneticist or genetic counselor, not with you and a Jupyter notebook.

The key diagnostic gain comes from trio sequencing: you, mother, father. De novo variants (present in you, absent in both parents) are enormously enriched for pathogenicity in severe early-onset disease, and the trio design lets you filter a candidate list from thousands to a handful. Singleton WES on an adult with a vague complaint rarely resolves anything.

The coverage problem

WES marketing says “the complete coding region.” It is not complete. Capture efficiency varies with GC content, repeat structure, and bait design, and the variance is large enough to matter. Comparisons of exome and genome sequencing on the same samples find that a meaningful fraction of coding exons are covered poorly or not at all by exome capture, while genome sequencing gives far more uniform coverage across the same targets.3 Genome sequencing detects exome variants with better sensitivity and better genotype quality than exome sequencing does, including within the exome regions both technologies nominally cover.4 Coverage distribution is the reason: WGS at 30x gives you a tight, near-Poisson coverage distribution, while WES at 100x mean has a long left tail where some exons sit at 5x or zero.

Practically, this means a negative WES is a weak negative. If the gene you care about has a GC-rich first exon (very common: promoters and 5’ exons are CpG-rich), your read depth there may be too low to call a heterozygous variant with confidence. Before you trust a negative, check depth at the specific locus.

# per-base depth over a region of interest, from the BAM you were given
samtools depth -a -r chr17:43044295-43125483 -q 20 -Q 20 sample.bam \
  | awk '$3 < 10 {n++} END {print n" bases under 10x"}'

# or callable regions genome-wide, using mosdepth
mosdepth --by exome_targets.bed --thresholds 1,10,20,30 -Q 20 sample sample.bam
# then read sample.thresholds.bed.gz: fraction of each target above 10x/20x

If your lab will not give you the BAM or CRAM, ask. A VCF alone cannot distinguish “reference at this position” from “no data at this position,” and that distinction is exactly what you need. A gVCF (GATK HaplotypeCaller -ERC GVCF) carries reference confidence blocks and solves this, but most clinical labs return only the filtered VCF.

Copy number and structural variants

WES can detect CNVs, but through an indirect route: read-depth ratios across capture targets, compared against a panel of normals sequenced on the same platform with the same bait set. Tools like ExomeDepth, CoNIFER, XHMM, and CNVkit all work this way. This can work: applying read-depth CNV calling to exome data recovered clinically relevant copy number variants that would otherwise have required a separate microarray.5 But the resolution floor is roughly one exon, breakpoints are unlocalized, and balanced events (inversions, translocations) are invisible because they do not change read depth. Repetitive regions, which is where a lot of structural variation lives, are not captured.

Bait-to-bait depth normalization also means your CNV calls are only as good as your reference panel. A different capture kit version, a different sequencer, a different library prep batch, and the depth ratios drift. If a lab reports an exome CNV without confirming it orthogonally (microarray, MLPA, or long-read), treat it as a hypothesis.

Variant calling choices, if you are doing your own

Assume you get a CRAM aligned to GRCh38. The defensible pipeline for germline short variants:

# 1. align (if starting from FASTQ)
bwa-mem2 mem -t 32 -R '@RG\tID:1\tSM:sample\tPL:ILLUMINA\tLB:lib1' \
  GRCh38_full_analysis_set_plus_decoy_hla.fa r1.fq.gz r2.fq.gz \
  | samtools sort -@ 8 -o sample.bam

# 2. mark duplicates (important for capture data; PCR duplication is higher than WGS)
gatk MarkDuplicates -I sample.bam -O sample.md.bam -M dup.txt

# 3. call, restricted to targets plus 100bp padding
gatk HaplotypeCaller -R GRCh38.fa -I sample.md.bam \
  -L exome_targets.bed --interval-padding 100 \
  -O sample.g.vcf.gz -ERC GVCF

Use the reference with decoy and HLA alt contigs. Omitting decoys drives spurious mappings into real genes. Callers disagree more than you would expect: head-to-head comparisons of GATK, DeepVariant, and other callers on the same WGS data show substantial non-overlap, particularly for indels.6 For anything you intend to act on, call with two independent callers and take the intersection, or use DeepVariant, which generally has better indel precision. Benchmark against Genome in a Bottle (NA12878/HG001 or HG002) with hap.py restricted to the high-confidence BED. Small differences in pipeline configuration produce clinically meaningful differences in which variants get called at known pathogenic sites.7

For annotation, VEP with --everything --pick_allele_gene plus gnomAD allele frequencies and ClinVar. Filter: allele frequency, consequence class, and then read the evidence for each surviving variant yourself.

When to choose WGS instead

We would sequence the genome. Cost difference has narrowed to the point where the capture reagents and the extra analyst time on coverage gaps eat most of the savings, and WGS gives you uniform coverage, better exome variant calls, structural variants from paired-end and split-read signal, mitochondrial DNA, and non-coding regulatory variation you can revisit as annotation improves.48 WES remains sensible when you need very deep coverage of coding regions on a budget, when you are sequencing hundreds or thousands of samples for a research question confined to coding variation, or when an existing clinical workflow is validated on exome and switching means revalidation.9 Large exome cohorts have produced real findings: gene-microbiota interaction analyses in inflammatory bowel disease used WES across thousands of participants to link host coding variation to gut microbial composition.10

One more practical reason to prefer WGS: the data is future-proof. A WES file is fixed to the bait set used in 2026. A WGS file is re-analyzable forever, against annotation databases that did not exist when you sequenced.

Questions people also ask

What does a whole exome sequencing test indicate? It reports single-nucleotide variants and small insertions/deletions in protein-coding exons and immediately flanking splice sites, classified against ClinVar and ACMG criteria. A “positive” result names a variant with evidence of disease association. A negative result means nothing was found in the captured, sufficiently covered regions, which is a narrower statement than most people assume.

What are the disadvantages of whole exome sequencing? Uneven coverage with GC-dependent dropout, so some coding exons are under-called or missed entirely.3 Poor detection of structural and copy-number variation, especially balanced events. No visibility into introns, promoters, enhancers, or repeat expansions. High PCR duplicate rates from capture library prep. And lower sensitivity than genome sequencing even inside the target regions.4

What diseases can be diagnosed with whole exome sequencing? Primarily monogenic conditions with coding-sequence causes: inborn errors of metabolism, many neurodevelopmental and mitochondrial disorders, skeletal dysplasias, primary immunodeficiencies, some cardiomyopathies and channelopathies. It also contributes to prenatal evaluation of structural anomalies.2 It does not diagnose polygenic conditions, and a diagnosis requires a clinician to correlate the variant with your phenotype.

Is WES the same as a clinical exome or gene panel? No. A clinical exome (CES) restricts analysis to a curated set of disease-associated genes, often a few thousand, which lowers the rate of uncertain findings at the cost of missing novel gene-disease associations.9 A gene panel is narrower still, typically tens to hundreds of genes, usually with better and more uniform depth per gene.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Zornitza Stark, Tiong Y. Tan, Belinda Chong, et al. A prospective evaluation of whole-exome sequencing as a first-tier molecular test in infants with suspected monogenic disorders. Genetics in Medicine, 2016. https://doi.org/10.1038/gim.2016.1 ↩

  2. Slavé Petrovski, Vimla Aggarwal, Jessica L Giordano, et al. Whole-exome sequencing in the evaluation of fetal structural anomalies: a prospective cohort study. The Lancet, 2019. https://doi.org/10.1016/s0140-6736(18)32042-7 ↩ ↩2

  3. Stefan H. Lelieveld, Malte Spielmann, Stefan Mundlos, et al. Comparison of Exome and Genome Sequencing Technologies for the Complete Capture of Protein‐Coding Regions. Human Mutation, 2015. https://doi.org/10.1002/humu.22813 ↩ ↩2

  4. Aziz Belkadi, Alexandre Bolze, Yuval Itan, et al. Whole-genome sequencing is more powerful than whole-exome sequencing for detecting exome variants. Proceedings of the National Academy of Sciences, 2015. https://doi.org/10.1073/pnas.1418631112 ↩ ↩2 ↩3

  5. Joep de Ligt, Philip M. Boone, Rolph Pfundt, et al. Detection of Clinically Relevant Copy Number Variants with Whole-Exome Sequencing. Human Mutation, 2013. https://doi.org/10.1002/humu.22387 ↩

  6. Anna Supernat, Oskar Valdimar Vidarsson, Vidar M. Steen, et al. Comparison of three variant callers for human whole genome sequencing. Scientific Reports, 2018. https://doi.org/10.1038/s41598-018-36177-7 ↩

  7. Rachel L. Goldfeder, James R. Priest, Justin M. Zook, et al. Medical implications of technical accuracy in genome sequencing. Genome Medicine, 2016. https://doi.org/10.1186/s13073-016-0269-0 ↩

  8. Tuuli Lappalainen, Alexandra J. Scott, Margot Brandt, et al. Genomic Analysis in the Age of Human Genome Sequencing. Cell, 2019. https://doi.org/10.1016/j.cell.2019.02.032 ↩

  9. Hanns-Georg Klein, Peter Bauer, Tina Hambuch. Whole genome sequencing (WGS), whole exome sequencing (WES) and clinical exome sequencing (CES) in patient care. LaboratoriumsMedizin, 2014. https://doi.org/10.1515/labmed-2014-0025 ↩ ↩2

  10. Shixian Hu, Arnau Vich Vila, Ranko Gacesa, et al. Whole exome sequencing analyses reveal gene–microbiota interactions in the context of IBD. Gut, 2020. https://doi.org/10.1136/gutjnl-2019-319706 ↩