2.58 million race results turned into age-group statistics

Triathletes ask a simple question: what finish time would put me in the top 20, 10, 5 or 1 percent of my age group at this race? Answering it takes millions of results and some care with percentiles.

Start with what you are allowed to use

The obvious data sources were the race organiser’s own results pages. We read their terms first. They forbid automated collection and storing results in a database, so we didn’t. We found an open research dataset under a CC BY 4.0 licence instead: about 2.7 million anonymised results from 2002 to 2026. No athlete names, and the product didn’t need any.

This step costs a day and saves a legal letter. We do it on every data project.

What we did

Loaded the whole dataset in one pass. 2.58 million results across 250 races and about 1,500 editions, bulk-loaded with PostgreSQL COPY and ranked in SQL. The full load takes about 25 minutes and can be re-run from scratch.

Cleaned up the messy parts. The same race appears under different names over the years, some events split men’s and women’s races across days, and distances vary. A small registry of rules maps all of that to one consistent list of races and series, documented next to the code.

Defined the statistic precisely. A “top 10% time” is the finish time of the athlete at position ceil(n x 0.10) among finishers in that age group. It is written down once, so every page means the same thing.

Made it fast. Computing percentiles across a whole race series live took 25 seconds. The thresholds are now precomputed after each data load, so pages render instantly.

What it shows

Data sourcing with the licence checked, a reproducible pipeline, a statistic defined once and used everywhere, and the performance work to make it usable.

Tell us what's broken.

A few sentences is enough. We reply within one working day and the first call is free.