Research Questions
- To what extent do LLMs show bias towards specific brands?
- Which metrics can be used to measure this bias?
- How can an appropriate brand comparison test be designed?
- What results emerge when these metrics are applied across different LLMs and brands?
Results
- All LLMs showed consistent (transitive) brand preferences, producing the ordering Apple > Samsung > Huawei.
- ChatGPT-4 was the most decisive model, whereas Gemma produced the most uncertain responses.
- ChatGPT models displayed a strong positive bias towards Apple and Samsung, while Gemma showed a weaker positive bias and higher uncertainty.
- The methodology can be applied not only to brands but also to a wide range of human-centred concepts.
Findings
-
Preference Dimension:
- All models demonstrated logical consistency; the aggregate preference ordering was: Apple > Samsung > Huawei
-
Sentiment Dimension:
- Gemma was more consistent but exhibited lower bias.
- ChatGPT-3.5 and ChatGPT-4 showed a strong positive bias towards Apple and Samsung (e.g. a 93.6% positive response rate for Apple).
- Responses about Huawei differed between models.
-
Consistency Metrics:
- ChatGPT models: low entropy, high decisiveness.
- Gemma: entropy levels 4–5× higher, indicating greater uncertainty.
- Results broadly agreed across question formats; Gemma showed the largest deviations.
Scores
- LLM Models: 5
- Synthetic Data: 1
- Method: 5
- Speed: 2
- Ethics: 1
- Accuracy: 4
- Demographics: 0
If you would like to explore this research in more detail, click here to read the full paper.