How to Get Your Genome Sequenced, and What You Receive
Getting your genome sequenced means giving a lab a blood or saliva sample, having it read on a short-read or long-read sequencer, and receiving files: a FASTQ or BAM/CRAM of your aligned reads and a VCF of the positions where you differ from a reference genome. Consumer whole-genome sequencing currently runs roughly $300 to $1,000 for 30x short-read coverage, and clinical-grade sequencing ordered through a physician runs several thousand dollars with a written interpretive report. The decision that matters is not which vendor has the nicest dashboard but whether you get the raw data, at what depth and read length, and whether the variants you care about are in regions the technology can see. Most of what follows is about making those choices deliberately.
What the sequencing produces
A short-read instrument does not read your chromosomes end to end. It fragments your DNA into pieces of a few hundred base pairs, amplifies them, and reads 100 to 150 bases from each end of each fragment, producing hundreds of millions of short reads that software then aligns to a reference genome 1. The reference itself is a mosaic assembly, originally finished in 2004 for the euchromatic portion of the genome, with persistent gaps in centromeres, segmental duplications, and other repetitive regions 2. Your “sequenced genome” is therefore a list of differences from that mosaic, not an independent assembly of your own chromosomes.
That distinction determines what you can and cannot learn. Single-nucleotide variants and small insertions and deletions in unique, well-mapped sequence are called reliably at 30x coverage. Structural variants, repeat expansions, and anything in a segmental duplication are called poorly or not at all with 150 bp reads, because a read that could have come from three near-identical places carries little positional information. Long-read platforms and the recent shift toward complete telomere-to-telomere assemblies and pangenome references directly address these regions, and the field is moving there for good reasons 3. A reference built from a single mosaic also misrepresents variation in populations underrepresented in the original project, which is the motivation for the ongoing work to build references spanning human diversity 4.
The specification we would insist on
If you are buying sequencing for yourself, treat it as a data purchase and write down the specification before you look at prices.
Ask for 30x mean coverage as the floor for short-read WGS. Thirty-fold is the conventional threshold at which heterozygous SNV sensitivity plateaus above roughly 99 percent in callable regions, and going to 15x to save money roughly doubles the rate of missed heterozygous calls because you need multiple independent reads supporting each allele. Ask what “30x” means in the contract: mean coverage over the whole genome, mean over callable regions, or a promised number of gigabases. A promise of 90 Gb of raw data is not the same as 30x aligned, deduplicated coverage after low-quality reads are removed.
Ask for 2x150 bp paired-end reads rather than 2x100. Longer reads improve mapping in repetitive sequence and give better indel sensitivity at no extra cost on current instruments. Ask whether library preparation is PCR-free. PCR amplification introduces duplicate reads and GC bias that degrade coverage uniformity, and PCR-free libraries measurably improve calling in GC-rich promoters. Ask whether the lab is CLIA-certified and CAP-accredited if you intend to show the result to a clinician, because most clinicians will not act on research-grade data and will order a confirmatory targeted test regardless.
Finally, ask in writing which files you receive and for how long they are hosted. The answer you want is FASTQ or CRAM plus a gVCF, downloadable, with checksums. A vendor that gives you only a curated report and a proprietary browser has sold you an interpretation, not a genome.
Checking the data yourself
The first thing to do with a delivery is verify it, before reading any variant. Confirm integrity with md5sum -c, then look at the raw reads with fastqc and aggregate across lanes with multiqc. You are checking for per-base quality dropping below Q30 near read ends, adapter contamination, and a duplicate rate above roughly 10 to 15 percent, which indicates an over-amplified library.
For alignment-level quality, run samtools stats and mosdepth on the BAM or CRAM. Mosdepth gives you the coverage distribution quickly, and the number to look at is the fraction of the genome at or above 10x and at or above 20x, not the mean. A mean of 30x with 8 percent of the genome under 10x tells a different story from a mean of 30x with 2 percent under 10x. Also check the reported reference build, because a CRAM is meaningless without the exact FASTA it was compressed against. GRCh38 is the sensible default today. If a vendor ships GRCh37/hg19 coordinates, you will spend time with CrossMap or bcftools +liftover, and liftover fails silently in exactly the regions where the two builds disagree most.
If you want to re-call variants yourself rather than trust the vendor’s pipeline, the route we would take is bwa-mem2 or minimap2 for alignment, then DeepVariant for small variants, then bcftools norm -m -both -f ref.fa to left-align and split multiallelic sites so your variants have canonical representations. Annotate with Ensembl VEP or snpEff, and check calls against the Genome in a Bottle benchmark using hap.py if you want a sensitivity and precision number for your own pipeline rather than a vendor’s marketing claim. The reason to re-call is reproducibility: you will know exactly which version of which caller produced each genotype, which matters when you revisit the data in three years.
What interpretation is and is not available to you
A VCF from a 30x genome typically contains four to five million variants, of which the overwhelming majority are common, benign, and uninformative about you specifically. Useful interpretation comes from narrowing to a small set of well-characterized loci. ClinVar gives you curated clinical assertions, with a review-status field that distinguishes expert-panel classifications from single-submitter guesses, and the distinction matters enormously. gnomAD population frequencies let you discard candidate variants that are too common to cause a rare condition. PharmGKB and CPIC document pharmacogenomic star alleles, though the important caveat is that CYP2D6 in particular involves copy-number variation and hybrid alleles that short-read data frequently miscalls.
Even with good annotation, the interpretive gap is the hard part. A large fraction of rare coding variants are classified as variants of uncertain significance, which means the evidence is insufficient to call them benign or pathogenic, and that category resolves slowly. The clinical literature has been direct about this: sequencing generates findings whose meaning and actionability are genuinely unclear, and the burden of that ambiguity falls on patients and families 5. The value of a genome is real but narrower than the marketing implies, and thoughtful reviews have made that point since consumer sequencing became affordable 6.
This is the point at which a clinician is not optional. If your data suggests a pathogenic variant in a cancer predisposition gene, a cardiomyopathy gene, or a familial hypercholesterolemia gene, the correct next step is a genetic counselor and a confirmatory clinical test in an accredited laboratory, not a decision made from a VCF. Research-grade calls have false-positive rates high enough that acting on an unconfirmed result is a mistake, and the family implications of a true positive require someone trained to discuss them.
Questions people also ask
Is genome sequencing worth the money? At $300 to $1,000 for 30x data you own permanently, the cost-per-decision is reasonable if you intend to work with the files yourself, because reanalysis against an improved ClinVar in five years costs nothing additional. If you want a single answer to a specific clinical question, a targeted panel ordered by a physician is cheaper, more sensitive for that gene, and interpretable.
How much does it cost to get your entire genome sequenced? Roughly $300 to $1,000 for consumer 30x short-read WGS, $1,000 to $3,000 for long-read sequencing that resolves structural variation, and $3,000 to $10,000 for clinical WGS with a written report and counseling. The reduction from the multi-billion-dollar cost of the original public project reflects two decades of instrument and chemistry development rather than any change in what a genome is 7.
Will insurance pay for genome sequencing? Usually only when a physician documents medical necessity: an undiagnosed condition with a suspected genetic basis, a critically ill infant, or a cancer case where tumor sequencing guides management. Elective sequencing of a healthy adult is nearly always out of pocket, and getting coverage means working through a clinician who can supply the diagnostic indication and appropriate procedure codes.
Is genome sequencing good or bad? The data is neutral and the risks are informational: uncertain findings that cause anxiety without changing anything, incidental discoveries about family relationships, and privacy exposure if you upload your files to services with weak or changeable terms. These concerns have been part of the debate since the project’s funding was first contested 8. Read the data policy, keep your own encrypted copy, and decide in advance which categories of result you want to look at.
What does it mean to get your genome sequenced? It means you hold a file describing where your DNA differs from a reference assembly, at the resolution the sequencing technology allows, annotated with whatever the literature currently says about those positions. It is a measurement with known blind spots, not a complete readout of you.
Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Footnotes
-
Ayman Grada, Kate Weinbrecht. Next-Generation Sequencing: Methodology and Application. Journal of Investigative Dermatology, 2013. https://doi.org/10.1038/jid.2013.248 ↩
-
International Human Genome Sequencing Consortium. Finishing the euchromatic sequence of the human genome. Nature, 2004. https://doi.org/10.1038/nature03001 ↩
-
Dylan J. Taylor, Jordan M. Eizenga, Qiuhui Li, et al. Beyond the Human Genome Project: The Age of Complete Human Genome Sequences and Pangenome References. Annual Review of Genomics and Human Genetics, 2024. https://doi.org/10.1146/annurev-genom-021623-081639 ↩
-
Roxanne Khamsi. A more-inclusive genome project aims to capture all of human diversity. Nature, 2022. https://doi.org/10.1038/d41586-022-00726-y ↩
-
Danton S Char, Mildred Cho, David Magnus. Whole genome sequencing in critically ill children. The Lancet Respiratory Medicine, 2015. https://doi.org/10.1016/s2213-2600(15)00006-5 ↩
-
A.J. Marian. Sequencing Your Genome: What Does it Mean?. Methodist DeBakey Cardiovascular Journal, 2014. https://doi.org/10.14797/mdcj-10-1-3 ↩
-
Elaine R. Mardis. The impact of next-generation sequencing technology on genetics. Trends in Genetics, 2008. https://doi.org/10.1016/j.tig.2007.12.007 ↩
-
Leslie Roberts. Genome Backlash Going Full Force. Science, 1990. https://doi.org/10.1126/science.11642769 ↩