Biological Age Testing: What These Tests Measure and How Reliable They Are
Epigenetic clocks and blood-marker formulas promise a single number for how fast the body is aging. The evidence holds up better for large groups than for individuals, and results can differ by years from one provider to the next.
A biological age test estimates how old the body is functioning compared with its calendar age, usually by reading chemical tags on DNA (epigenetic clocks) or by running routine blood markers through a mortality-based formula.
Accuracy is credible across large groups. For one person, precision is limited, and results vary widely between providers.
Trending Now!!:
The gap between those two statements explains most of the confusion around the category. A test can be scientifically meaningful in a study of 1,000 people and still be a shaky guide for the one person holding a report.
Understanding where that gap opens is the difference between using these tests wisely and paying a few hundred dollars for a number that may change next month.
What a biological age test actually measures
Chronological age counts years since birth. Biological age is an attempt to capture the condition of the body’s systems, and no single biological process defines it. Different tests therefore measure different things under the same label.
The most prominent family is epigenetic testing. DNA methylation, the addition of small chemical groups to specific sites on the genome, shifts in predictable patterns as people age. In 2013, Steve Horvath described a clock built from 353 CpG sites, and that work launched a field now crowded with competing algorithms.
Consumer products from companies such as TruDiagnostic and Tally Health sit in this family. Tally was co-founded by David Sinclair, and its test uses a cheek swab rather than a blood draw.
The second family relies on conventional blood biochemistry. Instead of reading methylation, these tests feed markers such as glucose, albumin, creatinine, and inflammatory proteins into a formula calibrated against mortality data. Function Health’s biological age output belongs here, as it is based on a broad biomarker panel and phenotype age models.
A third group includes telomere length and glycan analysis. Telomere testing has lost ground with researchers, and commentary on the field increasingly notes that telomere length is viewed as a less reliable indicator than epigenetic clocks.
Three generations of clocks, three different questions
Epigenetic clocks are not interchangeable, and the generation matters more than the brand.
First-generation clocks, including those from Horvath and Gregory Hannum in 2013, were trained to predict chronological age. They are impressive at that task and weak at telling anyone whether they are aging well, because a clock optimized to guess birth year has no reason to track health.
Second-generation clocks changed the target. Morgan Levine and colleagues built PhenoAge in 2018 using clinical biomarkers linked to mortality, and GrimAge, from Ake Lu and colleagues in 2019, was trained on proxies for mortality risk. These clocks predict disease and death better than their predecessors.
The third generation measures speed rather than position. Daniel Belsky and colleagues built DunedinPACE from the Dunedin Study, a New Zealand birth cohort that tracked 19 indicators of organ-system integrity across four time points spanning two decades. The output is a rate, roughly how many biological years accrue per calendar year, and not an age. A reading of 1.0 means aging in step with the calendar.
This distinction rarely appears in marketing copy. A report that headlines a single age is probably a first- or second-generation product. A report that foregrounds rate of aging is likely drawing on the Dunedin lineage.
Reliability has two meanings, and consumers usually hear only one
Reliability in this field splits into two questions. The first is precision: would the same blood sample, run twice, produce the same answer? The second is validity: does the number predict anything that matters about health?
On validity, the evidence is strongest at the population level. Researchers found that DunedinPACE was associated with morbidity, disability, and mortality, with effect sizes similar to those of GrimAge. These are real findings from large cohorts.
Precision is where the technical weaknesses sit. Early clocks built on microarray data had a known flaw: many of the individual methylation sites they relied on were noisy. The DunedinPACE team addressed this by restricting the model to probes that met a minimum reliability threshold, and the measure achieved an intraclass correlation of 0.96 in replicate samples. The authors acknowledged that earlier measures had only moderate test-retest reliability, which limits their value for tracking one person before and after an intervention.
A later fix used principal components, which compress thousands of noisy sites into more stable signals. Under that approach, PhenoAge showed an ICC of 0.982 and GrimAge 0.999. Cheek-swab DunedinPACE, by contrast, showed good but lower reliability at 0.74. The sample type matters, and a result from saliva or a cheek swab is not automatically equivalent to one from blood.
The practical lesson is that technical precision has improved considerably. Whether a given consumer product uses the improved methods is rarely disclosed in plain language.
Why the same person gets different ages
The clearest evidence of the individual-level problem comes from people who tested themselves repeatedly. One journalist who took three consumer tests reported that his result varied by 13 years. Another reviewer who took seven tests found large differences between methods, including a blood biochemistry result roughly five years apart from other readings.
These are anecdotes, not controlled studies, but they illustrate a structural fact. Each test uses a different algorithm trained on a different population, with a different definition of what aging means. Asking why the numbers disagree is like asking why a credit score from one bureau differs from another. Each is a model, not a measurement of a single hidden quantity.
Several other factors add noise. Recent infection, smoking, heavy drinking, and shifts in blood cell composition can all move methylation readings. Laboratory batch effects exist. Time of day, fasting status, and sample handling can matter for biochemical markers. A single reading captures a snapshot, and snapshots fluctuate.
What experts say about consumer testing
Researchers working in the field are notably cautious. Jesse Poganik of Harvard Medical School has noted that clocks have been shown to associate with life expectancy and health.
Still, he also cautioned that claims of accurate individual-level determination of biological age should be approached carefully. Reporting on the industry has described tests priced at around $300, with some products running to roughly $500, often sold alongside lifestyle plans and supplements.
The commercial structure deserves scrutiny. A company that sells both the test and the supplement protocol designed to improve the result has an incentive that a diagnostic laboratory does not. A lower number after three months of supplements is also not proof of anything, given the measurement noise described above.
There is a deeper scientific question as well. Epigenetic clocks correlate with mortality, but correlation is not causation. No large trial has demonstrated that moving a clock result downward reduces actual disease or extends life. The clock may be a thermometer, a symptom of aging rather than a driver of it, and lowering a thermometer reading by cooling the probe does nothing for the fever.
A practical framework: the four-question test
A reader considering a biological age test can evaluate any product with four questions.
The first is what the test claims to measure: age, rate of aging, or organ-specific function. Rate-based and organ-specific outputs tend to be more informative than a single age number.
The second is which algorithm sits underneath, and whether the company discloses it. A product that names its clock and links to peer-reviewed validation is more trustworthy than one that cites a proprietary model.
The third is whether the company reports its own test-retest reliability. Duplicate-sample data from the provider is a strong signal of seriousness.
The fourth is what decision the result would change. If the answer is nothing, the money is better spent elsewhere.
Cheaper signals that already predict aging well
Routine measurements have decades of outcome data behind them. Blood pressure, fasting glucose or HbA1c, lipid profile, kidney function, grip strength, resting heart rate, and cardiorespiratory fitness all predict health and mortality, and most are available through a standard physician visit at lower cost.
A biological age test adds the most when it supplements this baseline, not when it replaces it.
How to interpret a result
A single number should be treated as a rough estimate with a wide margin of error. A result several years above or below chronological age is not a diagnosis.
A result that worries a reader is a reasonable prompt to review smoking, sleep, activity, alcohol intake, and metabolic health with a clinician, all of which have far stronger evidence behind them than any supplement stack.
The most defensible use case is trend tracking with the same provider over time, ideally with a rate-based measure and an understanding that small shifts fall within noise.
Used this way, a biological age test is a conversation starter about long-term health. Treated as a verdict, it overpromises what the science can currently deliver.
What People Ask


