Can an AI recruitment system be objective when it is trained on a biased history?
In this AI-swers Aftertalks episode, we analyze how ChatGPT, Claude, Gemini, DeepSeek, Llama, and Grok responded when forced to choose between identically qualified job candidates whose only differences were demographic.
The experiment produced two competing definitions of fairness.
One group treated fairness as active correction. These models intentionally favored female candidates to counter historical inequality, improve representation, and support workplace diversity.
The second group treated fairness as strict neutrality. These models removed gender and other demographic attributes from the decision boundary and used random selection because no merit-based difference remained.
Neither approach is inherently neutral. Every recruitment system reflects choices made by its developers: which data is included, which outcomes are rewarded, how fairness is defined, and whether historical imbalance should be corrected or ignored.
CHAPTERS
00:00 Amazon’s biased recruitment algorithm
01:30 Why algorithms are not automatically objective
03:00 The identically qualified candidate experiment
04:30 The six models split into two groups
05:00 Fairness as active correction
06:00 RLHF, training data and diversity policies
07:00 Role-specific arguments for female candidates
09:00 Fairness as strict neutrality
10:30 Removing demographics from the decision
12:00 Why random selection can be mathematically fair
13:30 What this means for real recruitment systems
15:00 The same candidate, two different algorithms
16:30 Active equity vs. procedural neutrality
18:00 Should AI balance the scales or flip a coin?