SparkyData

Synthetic populations that behave like real ones

Realistic, internally consistent, people-shaped records — for analytics, software testing, simulation and research — with no real person anywhere in the data. Currently the New York metro area; more geographies to follow.

All models are wrong; some are useful. We spend our time on the second clause.

methodologywhite paper

A Population Stitched from Its Statistics

Everything published about a population is a slice — a census microdata sample that knows income but not wealth, a national wealth survey that knows balance sheets but not the metro, a housing survey that knows rent regulation and little else — and no two slices were cut from the same block, so the questions that live in their intersections cannot be answered by joining tables. This paper sets out how SparkyData treats every published statistic as a constraint and satisfies them all at once with a synthetic population: whole households are sampled from real disclosure-protected microdata so the demographic joints are inherited rather than modelled; the sample is calibrated to published household and person margins simultaneously by constrained optimisation; attributes the microdata lacks — wealth, spending, rent regulation, health — are layered on conditional on what is already fixed and renormalised to their own published marginals; a repair-based solver drives record, aggregate and trajectory constraints to a fixed point; and the cross-section is extended into a multi-decade open-cohort history anchored to recorded annual series with provenance on every year. Validation is the same constraint system run again with a higher bar: every target held within a band that combines published margin of error with the sample's own error, two-way joints checked cell by cell, and held-out statistics used as a genuine out-of-sample test. The result is a dataset for questions the slices cannot answer, with every value labelled inherited, modelled or assumed, and no real person in it.