Research Questions
- Can LLMs directly imitate human preferences and decision patterns?
- If perfect imitation is not feasible, can they reflect the diversity across customer segments (especially those differing by language)?
- Does the chain-of-thought conjoint approach make LLM decisions more human-like and help explain preference heterogeneity?
Results
- LLMs (GPT-3.5 and GPT-4) were more impatient than humans; GPT-4’s discount rates were far above human levels.
- GPT-3.5 displayed a lexicographic preference structure, producing outcomes inconsistent with human behaviour.
- The chain-of-thought method reduced GPT-4’s impatience, yet it remained more impatient than human respondents.
- LLMs were able to reflect language-driven differences; models were more patient in weak-FTR (Future Time Reference) languages.
- LLMs are not reliable for direct preference measurement, but they can be valuable tools for hypothesis generation and analysing heterogeneity across segments.
Findings
- Impatience:
- GPT-4’s discount rate (δ) was far higher than that of humans; GPT-3.5 and GPT-4 rarely chose the larger, later reward (22% and 16% of the time, respectively).
- Lexicographic Behaviour (GPT-3.5):
- Decisions were insensitive to the interest rate and could not be explained by any valid utility function.
- Effect of Chain-of-Thought:
- Chain-of-thought raised GPT-4’s rate of choosing the larger, later reward from 16% to 34.5%. It also revealed the themes behind the decisions (risk, uncertainty, opportunity cost).
- Language Differences:
- In weak-FTR languages (e.g. German, Mandarin), the GPT models were more patient, which is consistent with the literature.
- Thematic Analysis:
- As delay length increased, references to “risk and uncertainty” rose systematically across model outputs.
Scores
- LLM Models: 5
- Synthetic Data: 2
- Method: 5
- Speed: 1
- Ethics: 1
- Accuracy: 4
- Demographics: 3
If you would like to explore this research in more detail, click here to read the full paper.