Methodology

Data Sources

Percentile data are drawn from peer-reviewed studies, large-scale population health surveys, authoritative technical reports, and documented analyses of open datasets. Primary sources include:

All sources are listed on the references index.

Population Context

Several metrics, including BMI and body fat percentage, use data from the US population (NHANES). NHANES is one of the largest, most rigorously collected, and most frequently updated population health datasets available, with measured (not self-reported) values and nationally representative sampling.

However, the US has one of the highest obesity rates among major countries. US percentile distributions therefore skew higher than most other populations; the 50th percentile reflects the US median, not a health threshold.

Where this matters, context is added directly on the metric pages:

  • BMI: WHO category bands are overlaid on charts so bars visually land in Underweight/Normal/Overweight/Obese zones. An international comparator note references mean BMI by country from the NCD-RisC (Lancet, 2024).
  • Body fat percentage: A caveat notes that Americans carry 5 to 10 percentage points more body fat than European populations at the same age and BMI.
  • Calf circumference: The historical US NHANES distribution is paired with a multinational Asian comparator covering cohorts in Japan, Malaysia, and Taiwan (Chen et al. 2025). The metric page treats this as population context rather than a conversion because body size, sampling, and measurement protocols differ.
  • Blood pressure: Systolic and diastolic BP percentiles are from NHANES 2001-2008. Age-standardized mean BP and hypertension prevalence vary substantially across countries (NCD-RisC, Lancet 2021), so the upper percentiles may not generalize globally. Each overview page notes this caveat.
  • 30-Second Chair Stand: Normative data is from US community-dwelling older adults (Rikli & Jones, n=7,183). A German study (n=1,657, ages 65-75) found lower scores in comparable age groups, partly attributed to higher body weight.

Where international comparator data exists, it is linked on each metric's reference page.

Percentile Calculation

Five percentile points are reported for each metric: 5th, 25th, 50th (median), 75th, and 95th. Depending on the source study, these values are derived by one of the following methods:

Method How it works Where used
Directly reported The source study publishes the exact percentile values shown. Most studies
LMS curve fitting The source provides L (skewness), M (median), and S (coefficient of variation) parameters by age; percentiles follow from the standard LMS formula. Lean mass index, appendicular lean mass index (Kelly et al. 2009)
Normal distribution Percentiles are calculated from reported means and standard deviations, assuming a normal distribution. Sources reporting summary statistics but no percentile tables
Percentile proxy When P5, P25, P75, or P95 is unavailable, the nearest reported value is substituted (e.g. P10 for P5, P20 for P25). Documented per metric
Interpolation When reported percentiles bracket the target (e.g. P20 and P30 but not P25), linear interpolation is applied: P25 = (P20 + P30) / 2. Tomkinson 2017/2018 youth norms
Calculator-derived Where the authors publish an interactive calculator exposing the full distribution, quartile values are read directly from it. Powerlifting lifts (van den Hoek 2024, via thestrengthinitiative.com)
Category-boundary estimate For tests scored in discrete performance categories, category boundaries are mapped to approximate percentile positions. Push-up norms by fitness category
Derived from microdata Where no published percentile table exists, percentiles are computed from publicly available raw survey data. NHANES ratios and calf circumference (see WHtR and calf derivations)
Equation-derived Where the source publishes regression equations rather than tables, those equations are evaluated at representative age, height, and sex values. FEV1, GLI-2012 (see FEV1 derivation)

Each metric's reference page documents which method applies. When multiple studies report data for the same demographic, priority is given to larger sample sizes and more recent publication dates.

Derived Datasets

For some metrics, no published study provides pre-tabulated age- and sex-stratified percentile tables. Percentiles are instead computed from published regression equations or derived from public microdata. Full derivation details are documented on each metric's methodology page:

Metric Source Method
Body Roundness Index NHANES 2021-2023 Single cycle; multi-cycle pool rejected as primary
Calf Circumference NHANES 1999-2006 public microdata Complete-case DXA cohort, eight-year MEC weights, weighted empirical quantiles
Conicity Index NHANES 2015-2023 pooled All-quantile discontinuity gate; population context
FEV1 (Lung Function) GLI-2012 equations GLI-2012 "Caucasian" coefficient regression at NHANES median heights
FEV1/FVC Ratio (Tiffeneau Index) GLI-2012 equations Dedicated ratio regression; sex-specific L-equations
FVC (Forced Vital Capacity) GLI-2012 equations GLI-2012 "Caucasian" coefficient regression at NHANES median heights
Marathon Finish Time London Marathon open data (Zenodo, CC-BY-4.0) Per-runner finish times, derived
Waist-to-Height Ratio NHANES 2015-2023 pooled Waist and height, pooled cycles
Waist-to-Hip Ratio NHANES 2017-2023 pooled Waist and hip, pooled cycles

Age Brackets

Age brackets match the source study's native groupings to preserve data fidelity. Most metrics use decade brackets (20-29, 30-39, ... 80+), while some use 5-year brackets (20-24, 25-29, ...) when the source data supports finer granularity. The number of brackets varies by metric depending on the age range covered by the primary study.

For trend curves showing how metrics change with age, we use the midpoint of each bracket, treating it as a continuous interval (e.g., 25 for a 20-29 bracket, 22.5 for a 20-24 bracket, 85 for the open-ended 80+ bracket).

Percentile Approximation

Source studies rarely report all five target percentiles (P5, P25, P50, P75, P95) directly. The approximation method varies by source:

  • Tomkinson 2017/2018 (youth norms) reports P10, P20, P30, P70, P80, and P90. P25 is interpolated as (P20 + P30) / 2 and P75 as (P70 + P80) / 2, while P5 uses P10 as a proxy and P95 uses P90.
  • van den Hoek 2024 (powerlifting) reports deciles P10 to P90, so P5 uses P10 and P95 uses P90. P25 and P75 are read from the authors' online calculator, which exposes the full distribution.
  • All other sources fall back to the nearest available percentile where P25 or P75 is absent and no interpolation bracket exists. Each metric page notes which of its percentiles are approximated.

Rating System

A five-tier rating system is used based on percentile rankings:

Rating Percentile Interpretation
Excellent 95th Top 5% of the reference population
Above Average 75th Higher than 75% of the reference population
Average 50th Median of the reference population
Below Average 25th Higher than 25% of the reference population
Poor 5th Bottom 5% of the reference population

For lower-is-better metrics (e.g. resting heart rate, reaction time), the scale is inverted: the 5th percentile is rated Excellent and the 95th percentile Poor.

Some clinical and body-composition metrics do not use the performance scale above: waist-to-height ratio, waist-to-hip ratio, body fat percentage, mid-upper arm circumference, calf circumference, Body Roundness Index, Conicity Index, and blood pressure. These report where a value sits in the population distribution using a neutral five-tier label: Very low, Low, Average, High, Very high (mapped 5th to 95th percentile). The labels describe population position only; they are not performance ratings or clinical classifications. For clinical measurements such as blood pressure, a value at either extreme may be medically significant, so the label is not a health judgment.

Limitations

  • Population representation: Several metrics use US-only data (NHANES), and the US has unusually high obesity rates. Percentiles for body composition metrics will skew higher than in most other countries. See Population Context above for details and the specific caveats we add to affected pages.
  • Measurement methods: Different studies may use slightly different protocols. For example, VO2 max can be measured directly via metabolic cart or estimated from submaximal tests.
  • Selection bias: Clinical and fitness registry data may over-represent healthier, more active individuals who seek testing.
  • Temporal changes: Population fitness levels change over time. The most recent available data are used, but some studies may be several years old.
  • Special-population norms: Some metrics are norms for a specific subpopulation rather than the general public. The powerlifting lifts (squat, bench press, deadlift) are based on drug-tested competitive powerlifters, a highly trained group whose values are substantially higher than recreational gym-goers. Where this applies, we add a prominent caveat on the metric page and note it on the reference page.

Updates and Evidence Review

Three dates can describe different parts of a metric. The source year is when the underlying research was published. Page updated is the most recent substantive change to the FitnessNorms metric family. Evidence reviewed by FitnessNorms is when that metric family most recently completed our internal editorial and data review.

A substantive update includes a change to percentile values, age brackets, source selection, derivation methods, test protocol, population limitations, or material interpretation. Layout, styling, search terms, related links, and generated build files do not change the page-update date.

The evidence review checks data provenance, copied values across overview, trend, and age pages, derivation notes, source links, protocol details, population caveats, and numerical claims in the prose. A page does not receive the label until all applicable checks are complete. If a substantive update is made, the review must be completed again before a current review date is shown.

A recent review date does not mean the source itself is recent. Older research can remain the best available evidence for a particular test or population. The source year remains visible so readers can judge its age directly.

This is an internal FitnessNorms review, not external peer review, clinical review, or endorsement by the source authors. Population norms describe reference distributions; they are not medical advice, diagnostic thresholds, or a substitute for individual assessment.