All of Us DNA Results Are Gone. Here Is What You Can Still Get.
If you enrolled in the All of Us Research Program hoping to receive your DNA results, that door has closed. The program returned research DNA results to eligible participants from 2020 through 2024, reaching over 220,000 people with genetic ancestry and trait reports, and has since stopped returning them. Ancestry and trait results were subsequently deleted from participant accounts, so if you did not save the PDFs, they are not recoverable through the portal. Health-related results that were delivered with genetic counseling, hereditary disease risk and medicine-and-genes reports, may persist in your own records or your clinician’s chart if you shared them, but the program is not issuing new ones.
The more useful point for a technical reader is one that was true the entire time: All of Us never gave participants raw genomic data. No FASTQ, no CRAM, no VCF. The reports were interpretation layers rendered as consumer-facing documents. If your goal is to run your own variant calls, compute polygenic scores, or hand a genome to an agent for analysis, participation in All of Us was never going to get you there.
What the program returned and what it kept
Three categories of result existed. Genetic ancestry and traits was the light-touch report, a set of ancestry proportions against reference populations plus a handful of trait calls of the earwax-and-cilantro variety. Hereditary disease risk covered pathogenic and likely pathogenic variants in a curated gene list derived from the American College of Medical Genetics secondary-findings recommendations, and was returned only with genetic counseling. Medicine-and-genes covered pharmacogenes, again with counseling attached.
The ancestry component deserves a note on how to read it. All of Us ancestry inference is a projection of your genotypes onto principal components computed from reference panels, then assigned to continental-scale categories. Work on population structure within the cohort shows the resulting groupings are coarse relative to the actual genetic variation present, and that a large fraction of participants have substantial admixture that a small set of labels compresses badly 1. A percentage in an ancestry pie chart is a model output conditioned on the reference panel, not a measurement of your heritage.
The underlying genomic data was never the participant’s to hold. It went into the Researcher Workbench, a controlled-access cloud environment where credentialed researchers query short-read whole-genome sequences and genotyping arrays under a data use agreement. The data is de-identified, which means you cannot locate your own sample in it even if you have access. You cannot join the workbench as a member of the public, download your own CRAM, and walk away.
There is also a governance argument worth understanding before you decide how you feel about all of this. Critics have pointed out that a program built on the promise of inclusion collected DNA from communities, including Indigenous peoples, under consent structures that do not give those communities ongoing control over how their genomic data is used or by whom 2. Whatever you conclude, the structural fact is the same: you contributed data, and the institution retains custody of it.
Accuracy, and why the array-versus-sequencing distinction is the one that matters
People ask which DNA test is most accurate, and the answer depends almost entirely on the assay, not the brand. Consumer genotyping arrays interrogate roughly 500,000 to 900,000 preselected sites. At those sites, genotype concordance is high, typically above 99 percent. Everything else in the file is either absent or imputed by statistical inference against a reference panel, a method that works well for common variants and degrades sharply for rare ones 3. The failure mode is specific and well documented: rare pathogenic variants reported by array-based platforms have high false-positive rates, because a rare allele call rests on a few probes with little statistical support.
Short-read whole-genome sequencing at 30x mean coverage measures roughly 3.1 billion positions directly, with per-base quality scores you can inspect and re-analyze against updated variant databases later 4. That is the difference that matters. The cost curve that made population-scale sequencing feasible is the same one that makes an individual genome tractable for a single person to own and process 5.
Getting raw data you control, and what to do with it
If you want to do the analysis yourself, buy sequencing from a provider that contractually returns raw files. The minimum acceptable deliverable is FASTQ or an aligned CRAM against GRCh38 plus a gVCF. Ask specifically for the reference build, the aligner, and the caller. A 30x human genome is roughly 90 to 120 GB as gzipped FASTQ and 40 to 60 GB as CRAM, so plan storage before the drive arrives.
Check the data before you interpret any of it. Run samtools stats and mosdepth on the alignment and confirm mean coverage near 30x with at least 90 percent of the callable genome above 20x. Run VerifyBamID2 and look for an estimated contamination (FREEMIX) below 0.02. On the VCF, bcftools stats should show a transition-to-transversion ratio around 2.0 to 2.1 genome-wide, with roughly 3.0 to 3.3 in coding regions. A Ti/Tv well below those values means your call set is noisy and everything downstream inherits the noise.
For annotation, we would use Ensembl VEP with the --cache --offline --everything flags plus a ClinVar VCF as a custom annotation source, rather than a web uploader, because you want the annotation database version pinned and reproducible. For polygenic scores, PLINK 2 with --score and weights pulled from the PGS Catalog works, with the standing caveat that scores derived in one ancestry group transfer poorly to others and that the ancestry adjustment matters more than the score itself.
Expect the interpretation to be harder than the computation. A 30x genome yields roughly 4 to 5 million variants relative to the reference, the overwhelming majority of them common and benign, and the large noncoding fraction remains difficult to interpret even with the regulatory element catalogs built for exactly that purpose 6. Anything you find in a medically actionable gene needs orthogonal clinical confirmation in a CLIA-certified laboratory and a conversation with a genetic counselor or physician before it means anything about your health. Research-grade calls are not diagnostic, and we would not treat them as such.
Questions people also ask
Is the All of Us Research Program legitimate? Yes. It is run by the National Institutes of Health and funded by the federal government, with data collection through a network of academic medical centers and community partners. Its scientific output is real and peer-reviewed, including analyses of population structure across the cohort 1.
Does All of Us pay you? Participants have received modest compensation for completing certain steps, typically in the range of a $25 gift card for an initial visit with biosample donation. There is no ongoing payment and no fee to participate.
How much does All of Us cost? Nothing to the participant. Enrollment, surveys, and the physical measurements and biosample collection are free, and the DNA results, while they were being returned, were free as well.
Does the government have access to everyone’s DNA? The government holds DNA from All of Us participants who consented to give it, not from the general population. The data sits in a controlled-access environment under a Certificate of Confidentiality that limits compelled disclosure. That protection is legal rather than technical, which is a reasonable thing to weigh when you decide where your samples go.
Can I still get my results if I saved nothing? No. The ancestry and trait reports were deleted from participant accounts, and the program is not reissuing them. If your health-related results were returned through a counseling session and entered a clinical record, request them from that provider.
Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.
Footnotes
-
Shivam Sharma, Shashwat Deepali Nagar, Priscilla Pemu, et al. Genetic ancestry and population structure in the All of Us Research Program cohort. 2024. https://doi.org/10.1101/2024.12.21.629909 ↩ ↩2
-
Keolu Fox. The Illusion of Inclusion — The “All of Us” Research Program and Indigenous Peoples’ DNA. New England Journal of Medicine, 2020. https://doi.org/10.1056/nejmp1915987 ↩
-
The 1000 Genomes Project Consortium. A map of human genome variation from population-scale sequencing. Nature, 2010. https://doi.org/10.1038/nature09534 ↩
-
Elaine R. Mardis. A decade’s perspective on DNA sequencing technology. Nature, 2011. https://doi.org/10.1038/nature09796 ↩
-
Rachel L Goldfeder, Dennis P Wall, Muin J Khoury, et al. Human Genome Sequencing at the Population Scale: A Primer on High-Throughput DNA Sequencing and Analysis. American Journal of Epidemiology, 2017. https://doi.org/10.1093/aje/kww224 ↩
-
The ENCODE Project Consortium. An integrated encyclopedia of DNA elements in the human genome. Nature, 2012. https://doi.org/10.1038/nature11247 ↩