Back to timeline

Research · January 21, 2018

Placed by the contemporary primary publication. The exact event date is not known; its documented interval appears below.

Gender Shades: three commercial classifiers measured

On 21 January 2018 the proceedings of the first Conference on Fairness, Accountability and Transparency carried a paper by Joy Buolamwini and Timnit Gebru. They showed that two widely used evaluation sets are composed mostly of lighter-skinned people, 79.6 per cent in IJB-A and 86.2 per cent in Adience, assembled a balanced set of 1,270 individuals, and measured three commercial gender classifiers on it. Error on darker-skinned women ran from 20.8 to 34.7 per cent; on lighter-skinned men it never exceeded 0.8 per cent.

Why it matters

Here the gap was measured not inside a laboratory but on services anyone could call through an API that same day, and split on two attributes together rather than one at a time. The combination is what showed the most: none of the three companies reported per-cohort figures in its own documentation, so before this work the 34-point spread between the best and worst served group had been stated nowhere.

The services evaluated: Microsoft, IBM Watson Visual Recognition and Face++. The highest error on darker-skinned women is 34.7 per cent, at IBM. The best results on lighter-skinned men are 0.0 per cent at Microsoft and 0.3 per cent at IBM; Face++ does best on darker-skinned men at 0.7 per cent. The maximum difference between the best and worst classified groups is 34.4 points. All three classifiers do worse on darker faces, by 11.8 to 19.2 points. The Pilot Parliaments Benchmark holds 1,270 individuals, 53.6 per cent lighter and 46.4 per cent darker on the Fitzpatrick scale, drawn from portraits of parliamentarians of Rwanda, Senegal, South Africa, Iceland, Finland and Sweden, the countries chosen for the share of women in their parliaments. The record claims nothing about identifying a person: what was measured is binary gender classification from a face, a different task. The Fitzpatrick scale describes skin type rather than race, and the paper calls the categorisation coarse itself. The date is the publication of PMLR volume 81 on 21 January 2018, not the conference days, which is what the incoming report gave.

Event record

Event date
January 21, 2018
Timeline date
Primary publication date
Verification
Sources gathered automatically · September 21, 2026
Lines
ID
evt-0543

PMLR volume 81 was published on 21 January 2018, the date the paper's own page metadata also carries. The conference itself was held on 23-24 February 2018 in New York, after the volume appeared.

Sources

Related events

Records that link to this one