Model outputs enter the simulation through parsing, validation or constrained-action interfaces. Across the runnable systems examined, unusable outputs are handled through repair, retries, defaults or failure handling. In one widely cited macroeconomic simulator, each household agent decides how much labour to supply and how much income to consume; the model’s answer is evaluated as a literal Python expression and, if that fails, replaced by the constant “supply all labour, consume half of income.”
Macro-level dynamics are jointly generated by LLM-mediated behaviour and the environment machinery that translates behaviour into simulated consequences. In the systems examined, gravity models, recommender systems, resolution loops and inherited economic components strongly structure mobility, congestion, virality and market outcomes. The relative contribution of the language model and the environment therefore has to be established empirically rather than inferred from the apparent autonomy of the agents.
Researcher choices across the stack are rarely disclosed and can materially affect results: base model and version, prompt wording, memory and retrieval weights, scheduling and exposure rules, sampling temperature, and the parser and fallback logic itself. Most of the LLM-driven systems examined leave at least some controllable randomness unseeded, while provider-side model behaviour may also vary across runs. Reproducibility therefore depends on which layers of the stack are controlled, recorded or replayed.
These findings reframe the field’s central problem. Validating an agent-based model has been the discipline’s core methodological difficulty for fifty years, and language models inherit it intact. By making agents more expressive and more opaque at once, they can make it harder. This report proposes making LLM-enabled simulation accountable: specifying, for a given use, what the model is for, which quantities must be validated against which evidence, how robust the conclusions are, and what claims they license. This extends to the language-model layer the validation practices (fitness-for-purpose, pattern-oriented modelling, and transparent documentation) that the agent-based modelling community developed long before LLMs.
For a policy audience the stakes are specific. A simulation that looks scientific can influence decisions about real populations, and the populations most often modelled (residents of a district, users of a platform, recipients of a health intervention) are frequently those least able to audit or contest the model. The report closes with technical and governance recommendations for ensuring that these tools support public decisions without lending them unearned authority.