A research team conducted experiments on facial appearance using large language models (LLMs) such as OpenAI's GPT-4o, Google's Gemini 3 Flash Preview, and Anthropic's Claude Sonnet 4.5, revealing that AI exhibits stronger biases based on facial appearance than humans.
Study suggests AI may hold stronger facial biases than humans, with advanced models magnifying prejudice
This article is a translation. Read the Japanese original
The study utilized pairs of facial photographs where facial features perceived by humans as "competent" or "trustworthy" were manipulated, and asked the AI to choose between them. As a result, GPT-4o selected faces that humans judged as competent with a probability of 87.83%. This significantly exceeds the approximately 62.65% mathematically estimated as the human selection probability.
The results of the experiment confirmed a trend where these biases tend to strengthen as AI models become more sophisticated. It was reported that GPT-5 shows greater bias than GPT-4o, reproducing human bias with a probability of 94.33% in competence judgments and 97.04% in tasks related to real-world decision-making.
Stephen Lear, a researcher at Cangrade and the lead author of the paper, pointed out that because AI has higher response consistency than humans, the reduced variability in facial judgments may be amplifying human biases. Lear stated that when utilizing AI in influential fields such as recruitment and legal affairs, it is necessary to always conduct thorough testing for potential bias.
Sources
- AI時代にはこれまで以上に「顔の良さ」が重要になる可能性がある、AIは人間以上に顔の偏見が強いため (GIGAZINE、2026-09-29)