Skip to content

What a DNA Clock Test Measures, and How to Run One on Your Own Data

Oak
A feathered reptile-bird specimen on a black plinth, its quills marked with rows of lit and unlit beads and banded tail rings.

A “DNA clock test” almost always means a DNA methylation age test: you measure the fraction of cells methylated at a few hundred thousand CpG sites, feed a specific subset of those sites into a published linear model, and get back a number in years (or in years-per-year, for pace-of-aging clocks). The measurement is real and reproducible to a point. The interpretation is where most of the value gets lost, because second-generation clocks are trained to predict mortality and morbidity risk in cohorts, not to tell an individual how old their body is. If you want to run one, the practical path is an Illumina EPIC v2 array on a whole-blood or buffy-coat sample, IDAT files in hand, and the scoring done yourself in R.

Three different things are called a DNA clock

Disambiguate before you buy anything.

The molecular clock of evolutionary genetics is the observation that neutral substitutions accumulate at a roughly constant rate across lineages, used to date divergence times. It has been tested extensively against mammalian sequence data, with clear rate heterogeneity between lineages and between site classes 1. Non-coding, low-recombination regions have been used to estimate effective population size and coalescent times in humans 2, and population-genomic studies of wild microbial isolates use the same machinery to date local adaptation 3. This has nothing to do with your personal aging.

The circadian clock is a transcriptional feedback loop, not a DNA measurement at all. Roughly 43% of protein-coding genes show circadian rhythmicity in at least one mouse tissue, which is why sampling time matters for any expression assay 4. The human clock genes and their architecture were mapped out in the post-genome era 5, microarray studies defined the output programs tissue by tissue 6, and a substantial share of the rhythm is generated post-transcriptionally rather than at the promoter 7. Clock gene expression also carries clinical signal in some contexts: BMAL1 expression correlates with antitumor immune infiltration and survival in metastatic melanoma 8. If you want to know your circadian phase, you measure dim-light melatonin onset or a timed transcriptome, not methylation.

The epigenetic clock is the rest of this page.

What the epigenetic assay produces

Bisulfite (or enzymatic) conversion turns unmethylated cytosine into uracil and leaves 5-methylcytosine intact. An array or sequencer then reads out, per CpG site, a beta value between 0 and 1: the estimated fraction of DNA molecules methylated at that position in that sample. The clock is a weighted sum of a few hundred of those betas plus an intercept.

Three assay options, with the tradeoffs we would weigh:

Illumina EPIC v2 array. ~935,000 probes, raw output is a pair of IDAT files per sample (Grn and Red). This is what every published clock was trained on, or on its predecessors (450K, EPIC v1). Per-sample cost is low, turnaround is a couple of weeks, and the coefficients transfer directly. The catch with v2 is probe renaming and dropped probes: some Horvath 353 and PhenoAge 513 sites are missing or have changed suffixes, so you need an imputation or mapping step and you should report how many clock CpGs were missing.

Targeted sequencing. TIME-seq multiplexes bisulfite libraries across samples with a transposase-based protocol and cuts both the cost and the hands-on time of methylation clock construction by a large factor relative to arrays 9. If you are building or recalibrating a clock rather than running one, this is the direction the field is moving.

Long-read WGS with native methylation calling. Oxford Nanopore duplex or PacBio HiFi gives you variants and 5mC in one run. Basecall with dorado basecaller sup,5mCG_5hmCG, then modkit pileup --cpg --ref GRCh38.fa --combine-strands to get a bedMethyl file. Per-site methylation estimates are binomial in read depth, so at 30x you are estimating a beta from about 30 observations and the standard error at beta=0.5 is roughly 0.09. That is far noisier per site than an array, which is why we do not recommend deriving a Horvath-style score from 30x WGS without site-level pooling. Population-scale sequencing primers are a good grounding in how depth and error profiles propagate into downstream calls 10.

Computing the score yourself

Given IDATs:

library(sesame)
sdfs  <- openSesame("idats/", func = NULL)        # SigDF objects
betas <- openSesame(sdfs, prep = "QCDPB")          # mask, dye bias, nonlinear bg, pOOBAH

QCDPB applies pOOBAH detection-p masking, which is stricter than minfi’s default and matters because clock CpGs sitting behind failed probes will otherwise be silently imputed to a cohort mean.

Then score:

library(methylclock)
ages <- DNAmAge(betas, clocks = c("Horvath", "Hannum", "Levine", "skinHorvath"))

For DunedinPACE use the DunedinPACE package (PACEProjector()), which normalizes to its own reference distribution rather than to your batch. Report, alongside the number, the count of clock CpGs that were present. If you lost more than about 5% of the Horvath 353, treat the point estimate as unreliable.

Cell composition is the single largest confound in blood. Naive T cell fraction declines with age, and several clocks partly read that shift rather than anything intrinsic to a cell. Run EpiDISH with the Houseman reference and report the six or twelve estimated fractions next to the age estimate. If you draw twice and the granulocyte fraction moved 15 points because you had a cold, the age delta means less than it looks like.

How accurate is it

Two very different accuracies get conflated.

Correlation with chronological age across a cohort is high: first-generation clocks reach median absolute errors of a few years in adult blood. That statistic tells you almost nothing about you, because it is a property of a regression line fit to hundreds of people spanning fifty years of age.

Test-retest reliability on the same person is the number you care about, and it is much weaker for the older elastic-net clocks. Two aliquots of the same blood draw, run on different array positions, can differ by a year or more in Horvath age. Pace-of-aging measures built explicitly for reliability (DunedinPACE) and principal-component versions of the classic clocks (PC-Horvath, PC-GrimAge) have substantially better intraclass correlation, and those are what we would use for any longitudinal tracking. If a vendor reports a single age with no confidence interval and no technical replicate, you have no way to tell a real 2-year change from array noise.

Practical rule: never compare a clock value from one vendor to a value from another, and never compare across array versions. Compare only within one assay, one lab, one pipeline, ideally with a stored reference aliquot re-run in each batch.

Failure modes worth knowing before you draw blood

  • Tissue mismatch. Saliva clocks are dominated by buccal epithelium and neutrophil fractions that vary with brushing technique. Blood clocks and saliva clocks are not interchangeable.
  • Pre-analytic delay. Time from draw to freeze changes the leukocyte mix. Standardize it, record it.
  • Batch and position. Randomize sample position on the chip if you are running several timepoints at once. Never run baseline and follow-up on separate slides at separate times if you can avoid it.
  • Regression to the mean. A first measurement at the tail of the distribution will tend to move toward the center on repeat for purely statistical reasons.
  • Over-interpretation. An epigenetic age above chronological age is not a diagnosis of anything. If a result concerns you, the useful next step is standard clinical workup with a physician, not a supplement protocol.

Questions people also ask

How do I test my epigenetic clock? Get an EPIC v2 methylation array run on a whole-blood or buffy-coat sample from a lab that will return the raw IDAT files, then score it yourself with sesame plus methylclock or DunedinPACE. Vendors that return only a number and no raw data leave you unable to check probe coverage, cell composition, or batch effects.

How do I figure out my biological age? There is no single measurement of it. Methylation clocks are one estimator, trained on cohort outcomes. Pairing a clock with cell-fraction deconvolution, standard blood chemistry, and repeated measurements over years gives you a slope, which is more informative than any single cross-sectional number.

Are epigenetic clocks accurate? Accurate at ranking large groups by age and mortality risk, considerably less precise for one person at one timepoint. The second-generation and PC-based clocks have meaningfully better test-retest reliability than the original 2013 models, and that reliability, not the cohort correlation, is what limits what you can learn about yourself.

How reliable is epigenetics as a measurement? The chemistry is reliable. Beta values from arrays replicate well site by site. Most of the instability in a reported age comes from the model on top, from cell composition shifts, and from batch handling, all of which you can inspect if you hold the raw data.

Can I get my epigenetic age from a 23andMe file or a standard WGS? No. Genotyping chips and short-read WGS report sequence, not methylation state. You need bisulfite or enzymatic conversion, or a long-read platform that calls modified bases natively.

Oak builds longitudinal molecular profiles of individuals: whole-genome sequencing, RNA sequencing, proteomics, blood biomarkers, and continuous glucose data, integrated into one model of you. Build your profile.

Footnotes

  1. Wen-Hsiung Li, Masako Tanimura, Paul M. Sharp. An evaluation of the molecular clock hypothesis using mammalian DNA sequences. Journal of Molecular Evolution, 1987. https://doi.org/10.1007/bf02603118 ↩

  2. Henrik Kaessmann, Florian Heißig, Arndt von Haeseler, et al. DNA sequence variation in a non-coding region of low recombination on the human X chromosome. Nature Genetics, 1999. https://doi.org/10.1038/8785 ↩

  3. Christopher E. Ellison, Charles Hall, David Kowbel, et al. Population genomics and local adaptation in wild isolates of a model microbial eukaryote. Proceedings of the National Academy of Sciences, 2011. https://doi.org/10.1073/pnas.1014971108 ↩

  4. Ray Zhang, Nicholas F. Lahens, Heather I. Ballance, et al. A circadian gene expression atlas in mammals: Implications for biology and medicine. Proceedings of the National Academy of Sciences, 2014. https://doi.org/10.1073/pnas.1408886111 ↩

  5. Jonathan D. Clayton, Charalambos P. Kyriacou, Steven M. Reppert. Keeping time with the human genome. Nature, 2001. https://doi.org/10.1038/35057006 ↩

  6. G. E. Duffield. DNA Microarray Analyses of Circadian Timing: The Genomic Basis of Biological Time. Journal of Neuroendocrinology, 2003. https://doi.org/10.1046/j.1365-2826.2003.01082.x ↩

  7. Shihoko Kojima, Carla B. Green. Circadian Genomics Reveal a Role for Post-transcriptional Regulation in Mammals. Biochemistry, 2014. https://doi.org/10.1021/bi500707c ↩

  8. Leonardo Vinícius Monteiro de Assis, Gabriela Sarti Kinker, Maria Nathália Moraes, et al. Expression of the Circadian Clock Gene BMAL1 Positively Correlates With Antitumor Immunity and Patient Survival in Metastatic Melanoma. Frontiers in Oncology, 2018. https://doi.org/10.3389/fonc.2018.00185 ↩

  9. Patrick T. Griffin, Alice E. Kane, Alexandre Trapp, et al. TIME-seq reduces time and cost of DNA methylation measurement for epigenetic clock construction. Nature Aging, 2024. https://doi.org/10.1038/s43587-023-00555-2 ↩

  10. Rachel L Goldfeder, Dennis P Wall, Muin J Khoury, et al. Human Genome Sequencing at the Population Scale: A Primer on High-Throughput DNA Sequencing and Analysis. American Journal of Epidemiology, 2017. https://doi.org/10.1093/aje/kww224 ↩