The PSI is ten subtractions, ten logarithms and a sum. The bands it is computed on, and what it cannot see, matter more than the threshold it is compared with.
Cut the development scores into bands, pour the new population into the same bands, and sum over the bands (new share − old share) × ln(new share / old share). On a small-business scorecard moving from its 2015-2019 development sample to 2020-2021 applications, that gives a PSI of 0.090517, under the usual 0.10 threshold, while the bad rate of the same applications doubled.
Worked in full in Machine Learning for Finance by Julian R. Sterling, with every figure reproduced in a free workbook.See the book on Amazon →
The card is the Brechin scorecard from Machine Learning for Finance. Its development sample is 6,000 applications decided between 2015 and 2019; the out-of-time sample is 2,400 applications from 2020 and 2021. The ten bands are cut once, on the development scores, so that each holds about 600 cases, and those boundaries are never moved. All figures are illustrative.
| Input | Value |
|---|---|
| Development applications, 2015-2019 | 6,000 |
| Out-of-time applications, 2020-2021 | 2,400 |
| Bands, cut on development scores | 10 |
| Rule-of-thumb thresholds | 0.10 and 0.25 |
For each band, compute its share of the development population and its share of the new population. Each band's term is the difference times the log of the ratio. Both factors have the same sign, so every term is positive and the index only rises with any movement, in either direction.
PSI = Σ (new% − old%) × ln(new% / old%)
Band 10: (5.33% − 10.00%) × ln(5.33 / 10.00) = −0.0467 × −0.6286 = 0.029335
Excel, per band: =(New-Old)*LN(New/Old), then =SUM() the column
| Band | Development n | Out-of-time n | Development share | Out-of-time share | ln(ratio) | PSI term |
|---|---|---|---|---|---|---|
| 1 | 600 | 355 | 10.00% | 14.79% | 0.3915 | 0.0188 |
| 2 | 599 | 315 | 9.98% | 13.12% | 0.2736 | 0.0086 |
| 3 | 601 | 308 | 10.02% | 12.83% | 0.2478 | 0.0070 |
| 4 | 600 | 285 | 10.00% | 11.88% | 0.1719 | 0.0032 |
| 5 | 600 | 246 | 10.00% | 10.25% | 0.0247 | 0.0001 |
| 6 | 598 | 203 | 9.97% | 8.46% | −0.1641 | 0.0025 |
| 7 | 602 | 219 | 10.03% | 9.12% | −0.0949 | 0.0009 |
| 8 | 600 | 179 | 10.00% | 7.46% | −0.2933 | 0.0075 |
| 9 | 600 | 162 | 10.00% | 6.75% | −0.3930 | 0.0128 |
| 10 | 600 | 128 | 10.00% | 5.33% | −0.6286 | 0.0293 |
| Total | 6,000 | 2,400 | 100% | 100% | 0.0905 |
The index is 0.090517. On the usual reading, under 0.10 means no material shift, and a monitoring report would mark the card green. The table says more than the total. The population has moved down the scale: the bottom band grew from 10.00 to 14.79 per cent of applications and the top band shrank from 10.00 to 5.33. The two end bands contribute 53.1 per cent of the index, and band 10 alone 32.4 per cent.
It helps to know what the number is. The PSI is a symmetrised Kullback-Leibler divergence: the divergence of the new distribution from the old, 0.0441, plus the divergence of the old from the new, 0.0464. That is why it treats a band that empties and a band that fills alike, and why it has no natural unit beyond the rule-of-thumb thresholds attached to it.
The PSI measures the distribution of scores, not outcomes. Over the same period the bad rate went from 6.32 per cent in development to 12.63 per cent out of time, a doubling, and in the lowest band from 22.17 to 31.83 per cent. A stable input distribution is not a stable model. The index is an early, cheap signal that the applicant mix has changed; it is not evidence that the predictions still hold.
| Variant | PSI | Reading |
|---|---|---|
| Ten fixed development bands | 0.0905 | Under 0.10 |
| Five fixed bands, neighbours merged | 0.0836 | Under 0.10 |
| Deciles recut on the new population | 0 | Nothing measured |
| Same direction of shift, 1.5 times larger | 0.2221 | 0.10 to 0.25 |
| Same direction of shift, twice as large | 0.5040 | Above 0.25 |
Coarser bands hide movement within them: merging pairs keeps 92.4 per cent of the index here, but the loss grows when the shift is concentrated inside a band. The index is also strongly non-linear: half as much shift again takes it from 0.0905 to 0.2221, and doubling the shift takes it to 0.5040, when the top band would hold only 0.67 per cent of applications.
The fatal error is recutting the bands on the new population. Deciles of the new scores hold 10 per cent each by definition, so the PSI is 0 whatever has happened, and the tool erases exactly what it is meant to detect. The bands must be the development boundaries, stored with the model. The second error is a zero count. A band with no new cases makes the logarithm infinite, and the usual patch, flooring the share at a small number, then decides the answer: emptying band 10 with a floor of 0.0001 gives 0.7611, while a single case in it gives 0.6112. Merge sparse bands instead. The third is reading 0.10 and 0.25 as statistical tests. They are conventions, and with 2,400 cases the index carries sampling noise of its own.
The cost of leaving the cut-off where development put it, on this same out-of-time population, is worked in what leaving a credit score cut-off unchanged actually costs.
Every term of the index appears as its own row, on the fixed boundaries, in the free workbook for this case.
The usual rule of thumb reads under 0.10 as no material shift, 0.10 to 0.25 as a moderate shift worth investigating, and above 0.25 as a significant change. These are conventions, not tests. A scorecard here scores 0.090517, comfortably green, while its bad rate doubles from 6.32 to 12.63 per cent.
No. The bands must be the development boundaries, fixed once. Recut deciles on the new population and every band holds 10 per cent by construction, so the PSI is 0 whatever has happened. On the fixed boundaries the same data gives 0.090517, with the bottom band growing to 14.79 per cent and the top band shrinking to 5.33.
A zero share makes the logarithm infinite, so floor each share at a small value or merge the band with its neighbour. The floor drives the answer: emptying the top band of 2,400 out-of-time cases with a 0.0001 floor gives a PSI of 0.7611; one case in it gives 0.6112. Merging is usually the safer fix.
This article is one calculation from Machine Learning for Finance. The book takes the same case from first principles to the decision, chapter by chapter, and every figure it prints is a live formula in the free companion workbooks.
Get the book on Amazon →Free companion files
Also on Amazon UK · Amazon Germany · Amazon France · Amazon Canada
Reading guide: corporate finance, valuation and markets → · All 453 articles →
If this book helped, or didn’t, a few lines on Amazon are worth more than they look: they are what the next reader goes on. Write a review. The workbook stays free either way.