Artificial intelligence models trained to process language and images tend to judge a person’s character based entirely on their facial features, mirroring a common human bias. Recent experiments suggest that these language models use arbitrary facial characteristics to make weighty decisions about whether someone is competent or even likely to commit crimes. The findings were published in PNAS Nexus.
People frequently look at a face and instantly guess whether that person is trustworthy or competent. This tendency, known as face-to-character inference, relies on subtle variations in facial shape and texture to make judgments about personality. Evidence indicates this is a widespread cognitive error, as physical facial structure provides no real information about someone’s actual traits. For instance, a 2014 study of children and adults found that even three-year-olds consistently judge character based on facial features, showing how deeply ingrained this habit is in human psychology.
“Faces are both pervasive and significant social stimuli,” Mahzarin R. Banaji, the Richard Clarke Cabot Research Professor of Social Ethics at Harvard University and external faculty at the Santa Fe Institute, told PsyPost. “So much so that the human brain has a dedicated region that responds to faces—we are, each and every one of us, ‘face experts.’ Face-based judgment permeate so many decisions we make and so getting it right is important. So there was a pragmatic reason to focus on face-based judgments.”
The research, led by Steven A. Lehr of Cangrade, Inc. alongside Banaji and Yash Lothe, sought to determine if artificial intelligence models share this specific human quirk. Large language models are increasingly multimodal, meaning they can analyze and respond to both text and images. Because their training data contains vast amounts of human writing, which is full of human biases, the models might learn to associate certain facial types with certain character traits.
“My own operating theory is that LLMs are a mirror that will reflect just about any human characteristic in surprisingly high fidelity,” Lehr said. “For this reason, I’m on the lookout for humanlike characteristics that would be surprising in a machine. Face-to-character biases fit the bill.”
There was also a question of whether these text-based systems would be immune to visual errors. “LLMs have been trained on language and we hoped that LLMs may not have learned biases that emanate from images,” Banaji explained. “Would AI save us from our own face-based errors of judgment (see the work of the brilliant psychologist, Alex Todorov).”
At the same time, modern models undergo extensive alignment training designed to prevent them from making harmful or prejudiced statements. The scientists wanted to find out if language models would remain neutral when asked to judge a face, or if they would replicate the human error of assessing character from physical appearance.
“We wanted to see if the machines that successfully avoid explicit bias in publicized domains would also act ethically in a less-publicized one,” Lehr noted. “This would have been a marker of more generalized egalitarianism. But, unfortunately, the models did show the bias, and strongly.”
To test this, the researchers conducted 13 experiments totaling nearly 8,000 trials across four different language models. In the first two experiments, the researchers presented the GPT-4o model with pairs of computer-generated human faces. These faces were digitally altered to vary in specific physical features that humans typically associate with either competence or trustworthiness. The visual differences between the faces in each pair were measured in standard deviations, allowing the researchers to test pairs that looked very different alongside pairs that looked nearly identical.
In 600 trials focused on competence, GPT-4o was asked to select the more competent or incompetent face from a pair. The model chose the face that humans typically rate as more competent 87.83 percent of the time. When asked to judge trustworthiness across another 600 trials, the model chose the expected face 72.67 percent of the time.
The model’s bias grew stronger when the faces were more physically distinct. For the pairs separated by six standard deviations in competence-related features, GPT-4o chose the expected face 98 percent of the time. When the faces were separated by only two standard deviations, making them look visually similar, the model’s selection rate dropped to 70 percent, though it still reliably favored the expected face.
Interestingly, the model appeared to amplify human biases. Based on mathematical estimates of human behavior on the same competence task, humans would be expected to choose the more competent-looking face roughly 62.65 percent of the time, well below the model’s rate of 87.83 percent.
“Readers should notice that these were large and practically meaningful effects,” Lehr said. “Indeed, according to standard effect size measures, the bias appeared to be notably amplified relative to that of humans. It should be noted that this is partly because the LLMs were so consistent in showing the bias, so this may reflect partly low response variance as opposed to just truly greater essential levels of bias. But of course, consistency matters in contexts like selection: it makes the bias more reliable.”
The researchers then tested whether this bias generalized to related personality traits. Across 2,160 trials, GPT-4o evaluated the same faces on traits related to competence, such as being smart or lazy, and traits related to trustworthiness, such as being warm or selfish. The model consistently generalized its judgments, choosing the expected face 74.63 percent of the time for competence-related traits and 73.33 percent of the time for trustworthiness-related traits.
In a fifth experiment, the scientists tested the model on 150 pairs of macaque monkey faces. These faces had previously been rated by humans as looking either mean or nice. GPT-4o chose the “nice” monkeys as more trustworthy in 66 percent of the trials. This suggests the model did not just memorize human faces from its training data but instead developed a generalized concept of facial trustworthiness that it applies even to nonhuman primates.
“The first was that these LLMs were even capable of showing these biases,” Lehr said regarding the models’ behavior. “Where do they come from? It shows that language (and models built on language) can pick up a surprising array of characteristics, some of which you might not intuitively expect.”
The researchers also explored whether the model would apply these arbitrary facial judgments to extreme behaviors and high-stakes scenarios. In an experiment with 540 trials, GPT-4o was asked which of two faces was more likely to be a serial killer, engage in human trafficking, or commit financial fraud. The model selected the less trustworthy-looking face 68.70 percent of the time.
A separate experiment tested positive real-world decisions, such as hiring a university president or funding a technology startup. Across 540 trials, GPT-4o recommended the more competent-looking individual 75.19 percent of the time. The language model readily incorporated groundless facial biases into its recommendations for both highly negative and highly positive outcomes.
“GPT-4o did not show any real reluctance to say that one person was more likely to be a serial killer, human trafficker, or Ponzi schemer,” Lehr pointed out. “And the models readily told us we should hire the person with the more competent-looking face.”
To see if this issue was specific to GPT-4o, the scientists replicated the basic competence and real-world decision tasks using three newer models. They tested GPT-5, Gemini 3 Flash Preview, and Claude Sonnet 4.5. All three models demonstrated massive face-to-character biases, sometimes outperforming older models in their levels of bias.
GPT-5 exhibited an even larger bias than its predecessor, choosing the expected face 94.33 percent of the time on basic competence judgments and 97.04 percent of the time when advising on real-world decisions. Gemini 3 and Claude Sonnet 4.5 also showed highly elevated levels of bias, often matching or exceeding the bias seen in GPT-4o. This indicates that as these language models become more advanced, their tendency to judge character based on physical appearance might actually be increasing.
“One might have thought that the three tested reasoning models would think through the questions more and avoid this bias, but in fact, they showed it to a greater degree,” Lehr said.
There are a few things to keep in mind when interpreting these findings. The experiments relied heavily on forced-choice scenarios, which required the artificial intelligence to pick one face over another. In more natural settings where a model is not strictly forced to make a direct comparison, its behavior might differ. Additionally, the study provided the models with two-dimensional static images. It is unknown if the models would respond differently to video inputs or more varied photographic angles.
The researchers also cautioned against dismissing these results as merely regurgitated patterns. “It is not actually immediately obvious that face biases should be reflected in the training data, since they are biases of vision rather than language,” Lehr explained. “Second, the ‘just in the training data’ critique doesn’t hold up well when we see bias amplified relative to humans and increasing across model generations. LLMs are not reflecting the training data—they are exaggerating it. If LLMs are truly expected to pick up anything that’s been expressed in language, this seems to me like a very dangerous proposition.”
Another detail to consider is the focus on trustworthiness and competence. While these are foundational social traits, it remains to be seen how language models evaluate faces based on other variables, such as gender or race. Researchers usually implement specific safety guardrails to stop models from showing obvious gender or racial prejudice, but these protections do not seem to cover subtler biases based on facial shape.
Future studies might explore how language models actually learn these visual associations from text-based training data. Discovering the exact origin of this bias could help developers create broader safety measures. If models are deployed to assist with human resources or legal decisions, their tendency to favor certain facial structures could lead to highly unfair outcomes.
“The training designed to align AI models with human values has achieved what looks like surface egalitarianism, but something deeper is needed,” Lehr concluded. “We should be cautious about using AI models in high-impact domains, such as hiring and law, and if we do use them, we should be certain to always carefully test them for biases.”
The study, “Like humans, language models demonstrate face-to-character biases,” was authored by Steven A. Lehr, Yash Lothe, and Mahzarin R. Banaji.
Leave a comment
You must be logged in to post a comment.