Both discrimination statistics fall straight out of a score band table, provided you count the ties. Here is the full calculation, and what the bands cost you.
From a score band table, the Gini is 2 × AUC − 1, where the AUC counts, for every good, the bads in lower bands plus half the bads in its own band; KS is the largest gap between the cumulative share of bads and of goods. On a ten-band table of 6,000 small business loans this gives a Gini of 45.61 per cent and a KS of 32.47 per cent. The same card measured case by case reads 47.07 and 34.12: banding costs about a point and a half of each.
Worked in full in Machine Learning for Finance by Julian R. Sterling, with every figure reproduced in a free workbook.See the book on Amazon →
The data are the development sample of the scorecard built in Machine Learning for Finance: 6,000 applications, 379 of them bad (6.32 per cent), cut into ten score bands of about 600 cases each. Band 1 holds the lowest scores. The counts are the book's; the banded calculation is this article's.
| Band | Cases | Goods | Bads | Bad rate | Cum. bads | Cum. goods | Gap |
|---|---|---|---|---|---|---|---|
| 1 | 600 | 467 | 133 | 22.17% | 35.09% | 8.31% | 26.78% |
| 2 | 599 | 542 | 57 | 9.52% | 50.13% | 17.95% | 32.18% |
| 3 | 601 | 562 | 39 | 6.49% | 60.42% | 27.95% | 32.47% |
| 4 | 600 | 564 | 36 | 6.00% | 69.92% | 37.98% | 31.94% |
| 5 | 600 | 565 | 35 | 5.83% | 79.16% | 48.03% | 31.12% |
| 6 | 598 | 574 | 24 | 4.01% | 85.49% | 58.25% | 27.24% |
| 7 | 602 | 582 | 20 | 3.32% | 90.77% | 68.60% | 22.17% |
| 8 | 600 | 587 | 13 | 2.17% | 94.20% | 79.04% | 15.15% |
| 9 | 600 | 584 | 16 | 2.67% | 98.42% | 89.43% | 8.98% |
| 10 | 600 | 594 | 6 | 1.00% | 100.00% | 100.00% | 0.00% |
| Total | 6,000 | 5,621 | 379 | 6.32% |
The area under the ROC curve is the probability that a randomly chosen good scores above a randomly chosen bad. There are 5,621 × 379 = 2,130,359 good-bad pairs. A good in band i beats every bad in bands below it, and ties with the bads in its own band; a tie counts as half a win.
AUC = Σi goodsi × (bads below band i + 0.5 × badsi) / (goods × bads)
Band 2: 542 × (133 + 0.5 × 57) = 87,533.0 pairs won. Band 10: 594 × (373 + 3) = 223,344.0
All ten bands: 1,550,971.5 / 2,130,359 = 0.7280
Gini = 2 × AUC − 1 = 2 × 0.7280 − 1 = 0.4561, or 45.61 per cent
Excel, cumulative shares in columns: =SUMPRODUCT(CumGood-CumGoodPrev, CumBad+CumBadPrev)-1
The Excel line is the trapezoid form of the same quantity, the area under the ROC curve plotted as cumulative bads against cumulative goods; it returns the identical figure, which makes it a good cross-check. The Gini used in credit scoring is the accuracy ratio: zero for a random score, one for a score that puts every bad below every good.
The Kolmogorov-Smirnov statistic is the widest separation between the two cumulative distributions. Read down the gap column: it peaks at the top of band 3, where 60.42 per cent of bads but only 27.95 per cent of goods have been passed.
KS = maxi (cum. share of badsi − cum. share of goodsi) = 60.42% − 27.95% = 32.47%
Excel: =MAX(CumBad-CumGood) entered over the ten rows
The KS sits at a band boundary because that is the only place a band table can measure it. Case by case, the book finds the maximum at a score of 565.22, a few points above the top of band 3; the banded figure is the best the table can see.
A band table treats every case inside a band as tied, and ties throw away ranking information. The coarser the bands, the more the metrics fall, even though the model has not changed.
| Resolution | AUC | Gini | KS |
|---|---|---|---|
| Case level (book) | 0.7353 | 47.07% | 34.12% |
| Ten bands | 0.7280 | 45.61% | 32.47% |
| Five bands | 0.7130 | 42.61% | 32.18% |
| Two bands | 0.6556 | 31.12% | 31.12% |
Ten bands lose 1.46 points of Gini, about 3.1 per cent of it. Two bands lose a third; with only one cut, Gini and KS become the same number. The out-of-time sample tells the same story: banded, 42.54 per cent Gini and 31.17 per cent KS against 43.70 and 32.43 case by case. The fall from development to out of time is almost identical on either basis, 3.06 points banded and 3.36 case level, so a monitoring report can track the change on bands provided it never compares a banded figure with a case-level one.
The first mistake is ignoring ties. Counting only the bads in strictly lower bands gives an AUC of 0.6809 and a Gini of 36.17 per cent, more than nine points too low, and the error grows with band width. The second is comparing a validation team's banded Gini with a developer's case-level Gini and calling the 1.46-point gap deterioration. The third is reading either figure as a measure of whether the model is right. Gini and KS measure ranking only. The same card under-predicts the out-of-time bad rate in nine bands out of ten while its Gini barely moves; that is a calibration failure, and it needs its own test.
Fix the band boundaries on the development sample and pour every later sample into the same boundaries. Recutting deciles on each new sample changes the ties and moves the banded metrics for reasons unrelated to the model. For the drift check on the same fixed bands, see how to calculate the population stability index.
The 6,000 development cases, the ten bands and the case-level AUC, Gini and KS are all live in the free workbook for this case.
It depends on the portfolio and the data, so treat any benchmark as illustrative. Application cards on thin small-business data often sit between 40 and 60 per cent; behavioural cards with payment history usually score higher. The card here reads 47.07 per cent case by case. A Gini far above peers is a reason to look for leakage, such as a field observed after the loan was made.
Gini equals 2 x AUC minus 1. An AUC of 0.7353 is a Gini of 47.07 per cent, and an AUC of 0.5, a random score, is a Gini of zero. The identity holds whenever both are computed on the same sample with ties treated the same way, and the Gini here is the accuracy ratio of the cumulative accuracy profile.
Because a band table can only measure the gap between cumulative bads and goods at its boundaries, and the true maximum usually falls between two of them. On the card here the case-level KS is 34.12 per cent at a score of 565.22; the nearest band boundary sits slightly below it, and the banded KS is 32.47 per cent.
This article is one calculation from Machine Learning for Finance. The book takes the same case from first principles to the decision, chapter by chapter, and every figure it prints is a live formula in the free companion workbooks.
Get the book on Amazon →Free companion files
Also on Amazon UK · Amazon Germany · Amazon France · Amazon Canada
Reading guide: corporate finance, valuation and markets → · All 453 articles →
If this book helped, or didn’t, a few lines on Amazon are worth more than they look: they are what the next reader goes on. Write a review. The workbook stays free either way.