Skip to content

What Genome Mapping Companies Sell, and Which One to Use

Oak
Three identical lab instruments in a row, their open hatches showing scattered dots, short glowing fragments, and one long lit filament.

If you want your own genome and you want the data, order 30x whole-genome sequencing on an Illumina short-read instrument from a provider that delivers FASTQ plus a BAM or CRAM and a GVCF, and confirm in writing that you get the FASTQ. That costs roughly $300 to $1,000 in 2026 depending on provider and turnaround. “Genome mapping” as a phrase splits into two unrelated things: consumer short-read sequencing sold under that name, and optical genome mapping, a cytogenetics technique that images long DNA molecules to find large structural rearrangements and does not give you base-level sequence. If you searched for genome mapping companies expecting the first thing, the optical mapping vendors are not competitors to what you want.

The three categories, and which one matches your intent

Genotyping arrays (23andMe, AncestryDNA, most $99 kits) measure 600,000 to 1.2 million pre-selected sites. That is under 0.05 percent of your 3.1 billion base pairs. The raw download is a tab-delimited text file of rsIDs and genotype calls, useful for ancestry and for a handful of well-characterized common variants, useless for anything rare. Arrays interrogate only sites the chip designer chose in advance, so a novel variant in your own family is invisible by construction. Imputation can fill in correlated sites statistically, but imputed genotypes carry error rates that make single-variant interpretation unwise.

Short-read whole-genome sequencing reads the whole genome in 100 to 150 base fragments, aligns them to a reference, and calls variants. This is what Nebula, Dante Labs, Sequencing.com’s partner labs, Veritas-style offerings, and most academic core facilities sell. At 30x mean coverage you get reliable single-nucleotide variants and small indels genome-wide, roughly 4 to 5 million variants per person against GRCh38, with several hundred thousand that are rare or absent in public databases. Deep WGS of 1,070 Japanese individuals found the large majority of discovered variants were novel relative to existing catalogs, which is the practical argument for sequencing rather than genotyping: population-specific and family-specific variation lives outside array content.1

Long-read sequencing (PacBio HiFi, Oxford Nanopore) reads 10 to 100 kilobase molecules. It resolves structural variants, repeat expansions, segmental duplications, and phasing that short reads miss, and it reads CpG methylation directly on nanopore without a separate assay. Consumer-facing long-read options exist but are priced at $1,500 to $3,000 and the interpretation tooling is thinner. Get short reads first.

Optical genome mapping (Bionano Saphyr) labels a specific 6-mer motif across long DNA molecules and images them, producing a fluorescent barcode you align to a reference map. It detects insertions, deletions, inversions, and translocations above roughly 500 base pairs to multiple megabases, and it is genuinely better than short reads at balanced translocations. It returns no sequence. There is no variant you can look up, no FASTQ, no pharmacogenomic call. It is ordered by clinical cytogenetics labs, not by individuals.

How to evaluate an offer

Ask for five things in writing before you pay.

Mean coverage and the definition used. “30x” should mean 30x mean depth of aligned, deduplicated reads across the callable genome, not 30x of raw bases sequenced. Duplicate rates of 10 to 20 percent are normal, so 30x raw becomes 25x aligned, and callable fraction falls. Ask for the percentage of the genome at ≥10x and ≥20x. Good short-read runs hit above 95 percent at 10x.

Reference build. GRCh38, not hg19. If a provider still ships hg19 BAMs, their pipeline is old. T2T-CHM13 alignments are better in the hard regions but ecosystem support for annotation is still catching up, so GRCh38 remains the pragmatic default.

Deliverables, by file type. FASTQ (two gzipped files per lane for paired-end reads, R1 and R2), CRAM or BAM plus index, and a GVCF or VCF from a named caller with a named version. A VCF alone is a dead end: you cannot re-call, cannot check depth at a site of interest, cannot run a structural variant caller later. Expect roughly 40 to 60 GB for gzipped FASTQ at 30x, 15 to 25 GB for CRAM, 1 to 2 GB for a GVCF.

Delivery mechanism and retention window. S3 presigned URLs or an SFTP endpoint, not a portal that streams a PDF. Ask how long the files stay available. Several providers delete after 30 to 90 days.

Lab accreditation and jurisdiction. CLIA-certified and CAP-accredited labs run documented QC; research-grade labs may not. Research-grade is fine for exploration. Any result you would take to a physician needs confirmation in a clinical lab, and interpretation belongs with a genetic counselor or physician.

What we would order

Short-read 30x WGS on a NovaSeq X or equivalent, paired-end 150, GRCh38, DRAGEN or a GATK best-practices pipeline, FASTQ plus CRAM plus GVCF. Price anchor: $300 to $700 from the cheapest credible providers, up to $1,000 with clinical accreditation and faster turnaround. If you care about structural variation specifically, add nanopore later rather than substituting optical mapping.

Then re-call it yourself, because the vendor’s VCF is one set of choices:

bwa-mem2 mem -t 32 -R '@RG\tID:1\tSM:you\tPL:ILLUMINA' \
  GRCh38.fa R1.fq.gz R2.fq.gz \
| samtools sort -@ 8 -o you.bam
samtools markdup -@ 8 you.bam you.md.bam
samtools index you.md.bam
samtools view -T GRCh38.fa -C -o you.cram you.md.bam

deepvariant --model_type=WGS --ref=GRCh38.fa \
  --reads=you.md.bam --output_vcf=you.dv.vcf.gz --num_shards=32

manta configManta.py --bam you.md.bam --referenceFasta GRCh38.fa \
  --runDir manta_out

Then annotate:

vep -i you.dv.vcf.gz --cache --assembly GRCh38 --everything \
  --plugin CADD,whole_genome_SNVs.tsv.gz \
  --vcf -o you.vep.vcf.gz --compress_output bgzip

Check the QC before you read a single variant: samtools stats for insert size distribution and error rate, mosdepth for coverage uniformity, and bcftools stats for the het/hom ratio and Ti/Tv. Ti/Tv near 2.0 to 2.1 genome-wide is expected. Values below 1.8 suggest false positives leaking in.

Known failure modes

Coverage is not uniform. GC-rich promoters, the HLA region, and segmental duplications drop out. A “no variant found” call in PMS2, SMN1, or the pseudoautosomal regions frequently means the reads never mapped there. Short-read platforms also differ in error profile and throughput in ways that affect what you can call, and those tradeoffs are documented in the platform comparison literature.2

Structural variants are undercalled. Short reads find deletions well, tandem duplications adequately, inversions and balanced translocations poorly. This matters for anyone interested in large rearrangements: whole-genome analysis of pancreatic tumors found structural variation patterns that exome or panel data would have missed entirely, and that reclassified a meaningful fraction of cases.3 The lesson for a personal genome is that the mutational classes you can see are bounded by the assay.

Interpretation is where most of the money is wasted. A 30x genome produces around 4.5 million variants. Filtering to rare, protein-altering, and in a gene with credible evidence leaves tens to low hundreds. Almost all of those are variants of uncertain significance. The ACMG secondary findings list (currently 81 genes) is the narrow, well-curated subset where a positive call has established clinical meaning. Anything you find there needs orthogonal confirmation in a clinical lab and a conversation with a genetic counselor before it means anything about you.

Finally, the genome is static. It does not tell you what your body is doing this month. Sequence once, then measure the things that change.

Questions people also ask

What is the best genome sequencing company? For data you can work with, pick on deliverables rather than brand: 30x short-read WGS on GRCh38 with FASTQ, CRAM, and GVCF returned over S3 or SFTP. Nebula, Dante Labs, and university core facilities all clear that bar at different price and turnaround points. Illumina makes the instruments and does not sell to individuals. Bionano sells optical mapping to clinical labs, a different assay with no sequence output.

How much does it cost to map a genome? In 2026, $300 to $1,000 for 30x short-read WGS, $1,500 to $3,000 for long-read, $99 to $200 for a genotyping array that reads well under 1 percent of your genome. Clinical-grade sequencing ordered through a physician bills at $2,000 to $5,000 because it includes accredited interpretation and a signed report.

Does insurance cover genome mapping? Generally not for healthy adults ordering out of curiosity. Coverage is routine when a physician documents a diagnostic indication, such as an undiagnosed condition in a child or tumor profiling. Consumer WGS is a cash purchase, and whether it is HSA- or FSA-eligible depends on your plan administrator.

How much does genetic mapping cost compared to genotyping? A $99 array and a $500 genome differ by about 2,000-fold in the number of positions measured. If you intend to do your own analysis, the array will run out of information within a week.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Masao Nagasaki, Jun Yasuda, Fumiki Katsuoka, et al. Rare variant discovery by deep whole-genome sequencing of 1,070 Japanese individuals. Nature Communications, 2015. https://doi.org/10.1038/ncomms9018 ↩

  2. Sang Tae Park, Jayoung Kim. Trends in Next-Generation Sequencing and a New Era for Whole Genome Sequencing. International Neurourology Journal, 2016. https://doi.org/10.5213/inj.1632742.371 ↩

  3. Australian Pancreatic Cancer Genome Initiative, Nicola Waddell, Marina Pajic, et al. Whole genomes redefine the mutational landscape of pancreatic cancer. Nature, 2015. https://doi.org/10.1038/nature14169 ↩