Academic Journal
A large language model framework for sample-free population synthesis.
| Title: | A large language model framework for sample-free population synthesis. |
|---|---|
| Authors: | Jones M; School of Engineering, Newcastle University, Newcastle, United Kingdom., Dawson R; School of Engineering, Newcastle University, Newcastle, United Kingdom., Mills J; School of Engineering, Newcastle University, Newcastle, United Kingdom. |
| Source: | PloS one [PLoS One] 2026 Jun 02; Vol. 21 (6), pp. e0341704. Date of Electronic Publication: 2026 Jun 02 (Print Publication: 2026). |
| Publication Type: | Journal Article |
| Language: | English |
| Journal Info: | Publisher: Public Library of Science Country of Publication: United States NLM ID: 101285081 Publication Model: eCollection Cited Medium: Internet ISSN: 1932-6203 (Electronic) Linking ISSN: 19326203 NLM ISO Abbreviation: PLoS One Subsets: MEDLINE |
| Imprint Name(s): | Original Publication: San Francisco, CA : Public Library of Science |
| MeSH Terms: | Demography*/methods , Large Language Models*, Humans ; Family Characteristics |
| Abstract: | Synthetic populations provide the demographic foundations for agent-based models in transport, public health, disaster management and other sectors, enabling credible representations of individual characteristics and behaviours. Many established synthesis methods rely on census microdata; however, such data are infrequently collected, privacy-restricted, and usually available only as small public-use samples at coarse geographic scales. This paper introduces a sample-free framework that uses a large language model (LLM) to generate complete, household-structured populations directly from aggregate demographic data. The framework is LLM agnostic and follows a multi-step process: objective definition, input preparation, LLM selection, and synthetic household generation. No model fine-tuning is required, meaning that data requirements are low and the framework is easily accessible. Population synthesis is formulated as an iterative prompting process in which an LLM generates households guided by the discrepancies between synthetic and target distributions. The model draws on prior knowledge encoded during pre-training to propose plausible attribute combinations, resulting in both statistical alignment and structural feasibility. In a global evaluation covering 109 countries, the framework achieved very close alignment on simpler marginals such as gender (SRMSE: 0.003) and household size (SRMSE: 0.026), while more structurally complex attributes such as household composition (SRMSE: 0.062) and age (SRMSE: 0.128) were also reproduced with good accuracy. These results were supported by detailed case studies in Newcastle upon Tyne (UK) and Dar es Salaam (Tanzania). The principal contribution of the framework is to enable the construction of coherent household-structured populations when detailed microdata are unavailable, expanding the applicability of agent-based modelling in data-constrained settings. (Copyright: © 2026 Jones et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.) |
| Competing Interests: | The authors have declared that no competing interests exist. |
| Entry Date(s): | Date Created: 20260602 Date Completed: 20260604 Latest Revision: 20260726 |
| Update Code: | 20260726 |
| PubMed Central ID: | PMC13229344 |
| DOI: | 10.1371/journal.pone.0341704 |
| PMID: | 42228755 |
| Database: | MEDLINE |
| ISSN: | 1932-6203 |
|---|---|
| DOI: | 10.1371/journal.pone.0341704 |