Summary

A bioRxiv preprint introduces MindEvolve, an autonomous workflow that uses large language models to generate interpretable symbolic models of human decision-making in social interactions. The models performed most reliably on simpler economic and social preferences, while some advanced systems showed promise on theory of mind and recursive planning.

A bioRxiv preprint posted on September 13, 2026, describes MindEvolve, an autonomous workflow designed to make large language models (LLMs) generate interpretable models of human decision-making in social interactions.

The work, by researchers from East China Normal University, Peking University, the Hebrew University of Jerusalem, the Austrian Institute of Technology and the Chinese Academy of Sciences, evaluates multiple LLMs across socioeconomic games. Rather than focusing only on whether a model can imitate human choices, the researchers asked whether it could produce a symbolic account of the cognitive processes behind those choices.

MindEvolve turns predictions into cognitive models

An ordinary prediction system can associate a situation with a likely action without exposing a useful explanation of how the decision was reached. MindEvolve is intended to generate a symbolic model instead: an explicit representation of the preferences, reasoning steps or planning processes that could produce observed behavior.

Such a model can be inspected by people and compared with existing theories of human decision-making. In the study, expert human evaluators assessed the proposed models for interpretability and theoretical coherence. This made the evaluation broader than a prediction-accuracy test: a model also had to provide an account that could be understood as a theory of behavior.

The researchers examined four areas of social cognition through a battery of socioeconomic games:

  • Economic preferences, such as how people value different outcomes.
  • Social preferences, which involve decisions affected by the outcomes of other people.
  • Social reasoning, including theory of mind—the ability to reason about another person's beliefs, intentions or knowledge.
  • Recursive planning, in which an individual reasons through several levels of decisions or anticipated reasoning.

These categories move from relatively direct preferences toward increasingly complex forms of strategic reasoning.

Performance varied with the complexity of the task

The preprint reports that most of the evaluated LLMs could robustly capture economic and social preferences in relatively simple strategic settings. In some cases, the models generated new explanations that incorporated broader knowledge than the models proposed by human experts.

The results were less consistent for more complex psychological processes. A subset of state-of-the-art models showed promising performance in modelling higher-order reasoning, including theory of mind and recursive planning, but the abstract describes this as an emerging capability rather than a general solution to cognitive modelling.

That distinction is important. The study is not presenting LLMs as established psychological theories or as replacements for human researchers. Instead, it presents a workflow for using them as candidates for theory construction: the models propose structured explanations, and human experts judge whether those explanations are interpretable and theoretically coherent.

The authors describe the findings as a roadmap toward LLM-based cognitive models that could approach human-expert-level theory construction. For now, the strongest results appear to be concentrated in simpler preference-based tasks, with more demanding forms of social reasoning remaining a harder test.

The work is a bioRxiv preprint and has not undergone peer review. The supplied abstract also does not identify the evaluated model versions, provide task-level performance measures or describe the detailed symbolic models, so the relative performance of individual systems cannot be assessed from the available source text.

Sources