The model exhibits a stronger tendency to stereotype individuals based on demographic groups compared to human participants in the original study. In terms of segregation, where a score of 2 indicates complete confinement of groups to their specialized areas, human participants achieved a score of 0.84. In stark contrast, OpenAI’s o3 model scores approximately 65% higher, at 1.83, nearing the maximum value possible.
According to Ryan Liu, a doctoral student at Princeton University and co-author of the study presented at ICML in Seoul this July, LLMs excel at “generating generalizations from limited data.” This inclination is precisely what they are optimized for. Decision-makers, whether human or machine, grapple with the “exploitation-exploration dilemma,” akin to the choice between trying a new restaurant versus sticking with a reliable favorite.
Given that LLMs are trained on mathematics, coding, and scientific problems—tasks that reward generalization from minimal examples—they tend to adopt assumptions rapidly. The same instincts that aid LLMs in solving logical problems can also lead them into rigid mindsets. Our experiments demonstrated that newer models with advanced inference capabilities, such as OpenAI’s o3 and DeepSeek’s R1, displayed even greater biases. “When LLMs attempt to generalize too swiftly in social contexts, that’s when errors often occur,” Liu notes. OpenAI and Anthropic did not respond to requests for comment.
Angelina Wang, a computer scientist at Cornell University and not directly involved in the study, remarked that these findings are particularly significant now that chatbots have enhanced memory and personalization capabilities. As chatbots leverage previous conversation history, they can “over-index on the same types of behavior they’ve encountered,” leading to biases. However, users desire chatbots to retain conversational context, indicating that merely reducing the memory capacity won’t resolve the issue. “We’re still working to find the optimal balance—neither too much nor too little,” Wang explains.
Simply instructing the model to be fair did not significantly alter its behavior. “You either fail to implement these values, or the process is hindered by a tendency to optimize for the most suitable hire,” asserts Liu. However, by incentivizing models with bonuses for diverse recruitment, the bias was markedly diminished. The key lies in crafting objectives that “integrate desirable social values to guide large-scale language models toward socially favorable behavior,” concludes Liu.
Source: www.technologyreview.com


