Start with what you are allowed to use
The obvious data sources were the race organiser’s own results pages. We read their terms first. They forbid automated collection and storing results in a database, so we didn’t. We found an open research dataset under a CC BY 4.0 licence instead: about 2.7 million anonymised results from 2002 to 2026. No athlete names, and the product didn’t need any.
This step costs a day and saves a legal letter. We do it on every data project.
What we did
Loaded the whole dataset in one pass. 2.58 million results across 250 races and about 1,500 editions, bulk-loaded with PostgreSQL COPY and ranked in SQL. The full load takes about 25 minutes and can be re-run from scratch.
Cleaned up the messy parts. The same race appears under different names over the years, some events split men’s and women’s races across days, and distances vary. A small registry of rules maps all of that to one consistent list of races and series, documented next to the code.
Defined the statistic precisely. A “top 10% time” is the finish time of the athlete at position ceil(n x 0.10) among finishers in that age group. It is written down once, so every page means the same thing.
Made it fast. Computing percentiles across a whole race series live took 25 seconds. The thresholds are now precomputed after each data load, so pages render instantly.
What it shows
Data sourcing with the licence checked, a reproducible pipeline, a statistic defined once and used everywhere, and the performance work to make it usable.