Research Questions
- To what extent can LLMs simulate human behaviour across different populations in experimental settings?
- Is it possible to reproduce classic behavioural experiments (Ultimatum Game, Garden Path Sentences, Milgram Shock, Wisdom of Crowds) using LLMs?
- How realistically can demographic variation (e.g., name, gender) be reflected in LLM outputs?
Results
- LLMs successfully replicated several known patterns of human behaviour.
- Larger models produced responses that were more human-like.
- Gender-based behavioural differences emerged in LLM outputs (e.g., the chivalry effect: men accepted unfair offers more often when the proposer was female).
- Newer and more aligned LLMs exhibited hyper-accuracy distortion, performing unrealistically well on certain knowledge tasks.
Findings
- Behavioural Imitation:
- Large language models reflected established human behavioural patterns in experiments such as the Ultimatum Game and Garden Path Sentences.
- Demographic Variation:
- LLMs were able to simulate behavioural differences based on demographic cues such as name and gender.
- Hyper-Accuracy:
- Some models produced answers that were unrealistically accurate relative to typical human performance, failing to mirror real-world human knowledge distributions.
- Data Contamination Concerns:
- Because LLMs may have been exposed to these classic experiments during training, questions arise regarding the originality and validity of the reproduced behaviours.
- Ethical Risks:
- Simulating harmful experiments such as the Milgram Shock Experiment raises significant ethical concerns.
Scores
- LLM Models: 5
- Synthetic Data: 4
- Method: 4
- Speed: 3
- Ethics: 4
- Accuracy: 3
- Demographics: 4
If you would like to explore this research in more detail, click here to read the full paper.