Research Questions
- How can LLM outputs be calibrated across demographic groups to better approximate human responses?
- Can this approach be made model-agnostic (applicable to any LLM)?
- To what extent can calibration be transferred across domains or geographic regions?
Results
- Human Mimicry Calibration (HMC) significantly improved alignment between LLM responses and human data.
- HMC outperformed traditional weighting methods based on population density.
- Geographic transfer (Texas → New York, California, Florida) was successful, while transfer across topics, especially politics and sensitive subjects, was limited.
- The highest weights were assigned to young (18–34) and low-income demographics, which produced outputs most similar to human behaviour.
Findings
- Calibration Technique:
- HMC improves LLM-human alignment by reweighting demographic persona outputs towards groups that better mimic human data.
- Performance:
- HMC achieved 20–30% higher accuracy than uniform and population-weighted baselines.
- Transferability:
- Geographic transfer worked effectively.
- Cross-domain transfer, particularly in sensitive or political topics, showed reduced accuracy.
- Evaluation Metrics:
- Differences between human and LLM responses were measured using Kendall’s Tau and Wasserstein Distance, and both metrics showed the gap narrowing after calibration.
- Key Insight:
- Younger and lower-income groups emerged as the most representative demographic segments in terms of producing human-like LLM outputs.
Scores
- LLM Models: 5
- Synthetic Data: 4
- Method: 5
- Speed: 3
- Ethics: 3
- Accuracy: 5
- Demographics: 5
If you would like to explore this research in more detail, click here to read the full paper.