How a risk scale's errors fell by race
On 23 May 2016 ProPublica published an audit of COMPAS, the recidivism risk scale that United States courts used in bail and parole decisions. The outlet took the scores of more than 7,000 people arrested in Broward County, Florida in 2013 and 2014 and checked them against new charges over the following two years. Of those labelled higher risk who did not re-offend, the share was 44.9 per cent among black defendants against 23.5 per cent among white ones.
Why it matters
Until then the argument about algorithms in the courts was about intent and disclosure. Here a number that could be recomputed took the place of a general claim: the outlet released the data, and the argument that followed was conducted on that same set. The atlas holds no other record in which a working system for deciding about people was measured from outside rather than described by the party that built it.
The figures come from the table Prediction Fails Differently for Black Defendants: labelled higher risk but did not re-offend, 23.5 per cent of white defendants against 44.9 per cent of black ones; labelled lower risk yet did re-offend, 47.7 per cent of white defendants against 28.0 per cent of black ones. Overall the tool predicted recidivism correctly 61 per cent of the time. A statistical test isolating race from criminal history, recidivism, age and gender left black defendants 77 per cent more likely to draw a higher score on the violent recidivism scale and 45 per cent more likely on the general one. The record does not claim that these numbers establish that the scale was biased. The developer replied on 8 July 2016 that the outlet had compared the wrong quantities: the table holds the share of non-recidivists with a positive result, the complement of specificity, while the caption describes it as the share of those labelled higher risk who did not offend, the complement of predictive value. Both sides computed different things correctly; why both could not be right at once was shown by a paper four months later. The outlet is not a research institution and the work was not peer reviewed. It published its method and the data themselves, which is why it could be recomputed at all.