Report

From Plausible Agents to Accountable Simulation: A Technical and Governance Framework for LLM-Enabled Agent-Based Modelling

This report examines how LLM-enabled social simulations can become more valid, transparent and accountable for public decision-making.

Large language models (LLMs) have made agent-based social simulation dramatically more plausible. Agents in these systems hold memories, adopt personas, deliberate in natural language, and interact in ways that read as recognisably human. This fluency has arrived at the moment such simulations are being proposed as instruments of public reasoning: emergency preparedness, urban and transport planning, public health, and economic and platform policy.

This report argues that plausibility and validity are being conflated, and that the conflation is consequential. A simulation can be entirely plausible — its agents can speak and act like people — and still be empirically unrealistic; recent evaluations against ground-truth human data confirm that today’s systems frequently are. The analysis treats the simulation stack — the parsers, defaults, schedulers and environment machinery — as a major source of observed plausibility and as the unit of validation and accountability.

Beyond the survey and validation literature, the analysis reviews fourteen public LLM-agent projects and nine traditional agent-based modelling platforms, including direct source-code inspection where available. The projects were reviewed in July 2026.

Model outputs enter the simulation through parsing, validation or constrained-action interfaces. Across the runnable systems examined, unusable outputs are handled through repair, retries, defaults or failure handling. In one widely cited macroeconomic simulator, each household agent decides how much labour to supply and how much income to consume; the model’s answer is evaluated as a literal Python expression and, if that fails, replaced by the constant “supply all labour, consume half of income.”

Macro-level dynamics are jointly generated by LLM-mediated behaviour and the environment machinery that translates behaviour into simulated consequences. In the systems examined, gravity models, recommender systems, resolution loops and inherited economic components strongly structure mobility, congestion, virality and market outcomes. The relative contribution of the language model and the environment therefore has to be established empirically rather than inferred from the apparent autonomy of the agents.

Researcher choices across the stack are rarely disclosed and can materially affect results: base model and version, prompt wording, memory and retrieval weights, scheduling and exposure rules, sampling temperature, and the parser and fallback logic itself. Most of the LLM-driven systems examined leave at least some controllable randomness unseeded, while provider-side model behaviour may also vary across runs. Reproducibility therefore depends on which layers of the stack are controlled, recorded or replayed.

These findings reframe the field’s central problem. Validating an agent-based model has been the discipline’s core methodological difficulty for fifty years, and language models inherit it intact. By making agents more expressive and more opaque at once, they can make it harder. This report proposes making LLM-enabled simulation accountable: specifying, for a given use, what the model is for, which quantities must be validated against which evidence, how robust the conclusions are, and what claims they license. This extends to the language-model layer the validation practices (fitness-for-purpose, pattern-oriented modelling, and transparent documentation) that the agent-based modelling community developed long before LLMs.

For a policy audience the stakes are specific. A simulation that looks scientific can influence decisions about real populations, and the populations most often modelled (residents of a district, users of a platform, recipients of a health intervention) are frequently those least able to audit or contest the model. The report closes with technical and governance recommendations for ensuring that these tools support public decisions without lending them unearned authority.