HeadlinesBriefing favicon HeadlinesBriefing.com

AI Biases in Hiring Outpace Humans

MIT Technology Review •
×

The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s good reason to question whether AI will judge you fairly. Researchers already know that LLMs pick up human biases from their training data. New research suggests that LLMs can also develop their own biases from experience—and stereotype job applicants more than humans do.

Researchers at Princeton University and the University of Chicago ran LLMs, including Chat GPT, Claude, and Gemini, through a simulated hiring game. Each model was told it had been hired as a consultant by the mayor of a fictional city and was then asked to help hire people for 20 jobs. Candidates came from four fictional ethnic groups: Tufa, Aima, Reku, and Weki. After each hire, the model learned whether the candidate succeeded, and the task was to make as many successful hires as possibleراط over 40 rounds. Unbeknownst to the models, all candidates were equally likely to succeed. The models quickly began segregating candidates by group.

On the study’s segregation scale, where 2 means every group has been confined to its own job niche, human participants scored 0.84. The models scored roughly 65% higher, with Open AI’s reasoning model o3 scoring 1.83, close to the maximum possible. Newer models with higher reasoning capabilities, such as Deep Seek’s R1, showed even stronger biases. The finding is especially relevant now that chatbots are gaining improved memory and personalization features, says Angelina Wang.

Telling the model to be fair didn’t change its behavior much, but promising an additional bonus for diverse hiring made them far less biased. The trick is to design goals that incorporate desirable social values to make large language models act in socially desirable ways.