Where one ancient genome sits
among living people

GJM is an individual sequenced from ancient remains. This page shows where that genome falls within two present-day reference populations across 46 polygenic scores. Not what this person had, but how their inherited variants compare with ours.

European reference, 503 people Czech cohort, 273 people ↕ scored in the opposite direction

Highest placed

    Lowest placed

      Percentiles above are within the European panel. The two references agree closely. Most traits sit within a few points of each other, which is what you would expect for a European individual compared against a European and a Czech panel.

      A second, independent set of scores

      The same genome was run again through the same pipeline against a different set of scores: six personality indices published in 2025, built from a much larger study and by a different statistical method. Nothing about the genome or the reference panels changed, only the scores applied to them. This is the closest thing available here to a replication check.

      European reference, 503 people Czech cohort, 273 people

      Two of these are worth pointing at. The neuroticism placement, near the bottom of both panels, agrees with the two neuroticism scores in the first chart, which come from different studies and a different method entirely. Agreement like that across independent scores is more informative than any single placement. And the two neuroticism rows differ only in whether one large biobank was included in the study behind them, which moves the placement by six percentile points. That is a fair measure of how much these numbers depend on choices made upstream of this analysis.

      What a percentile here does and does not mean

      A polygenic score adds up the small effects of many inherited variants. Placing this genome at the 90th percentile means its combination of variants sums higher than 90% of the people in that reference panel. It does not mean the person had the trait, and it is not a risk figure or a diagnosis. Nothing here is calibrated to an absolute probability, and none of it is clinical.

      Three cautions travel with these numbers.

      One score was removed. A 2021 kidney disease score is not usable here. Its value barely varies between people, so small technical changes swing it wildly. Re-processing this genome moved it from the bottom of the distribution to the middle. It is excluded from the chart.

      One score runs backwards. The 2022 kidney score, marked ↕, measures kidney function rather than disease, so a high placement there means lower liability, not higher. Read in a consistent direction, all three kidney scores agree that this genome carries low genetic liability for kidney disease.

      Some scores rest on fewer variants than published. All three genomes have to be compared on exactly the same set of variants, so every score is restricted to variants present in all of them. Most lose about 5% of their weight. Three lose much more, up to 83%, and are correspondingly less reliable.

      How the comparison is kept fair

      An ancient genome is incomplete. Comparing it naively against modern data measures the gaps in the sequencing rather than the biology, because a missing variant is silently scored as an absence. Every score here is therefore computed on the intersection: only variants successfully observed in the ancient genome, the European panel and the Czech cohort. The pipeline refuses to emit a results table unless that intersection is verified identical across all three, variant by variant.

      Missing positions in the ancient genome were reconstructed statistically from the sequencing reads against a reference panel of 3,202 people. Scoring used pgsc_calc, the PGS Catalog's own tool, against published scoring files, so the weights are the original authors' and not ours.