Attribute sampling when the sample is not clean: the upper error bound, the ladder of sample sizes, and why two out of 59 is not 3.4 per cent.
Two exceptions in a 59-item sample prove only that the error rate is below 10.3 per cent at 95 per cent confidence, not below 5 per cent. The 59 supports the 5 per cent claim only if every item is clean. To make the claim again with two exceptions already found, the sample has to grow to 124 items, 65 more, all clean.
Worked in full in The Private Fund Compliance Officer by Julian R. Sterling, with every figure reproduced in a free workbook.See the book on Amazon →
Compliance testing of fee calculations, expense allocations or trade approvals is attribute sampling: each item is either right or wrong. The standard zero-error formula gives the sample that supports a claim if nothing is found.
n = ln(1 − confidence) / ln(1 − tolerable error rate) = ln(0.05) / ln(0.95) = 58.40, rounded up to 59
The derivation of that figure, and what twenty clean items prove (an error rate below 13.9 per cent), are set out in how many expense allocations you have to test. The formula has a condition built in: zero exceptions. Most testing programmes do not stop to ask what happens when the condition fails.
A compliance officer tests 59 expense allocations, selected at random from the year, at 95 per cent confidence against a tolerable error rate of 5 per cent. Two allocations are wrong. The draft report reads "two exceptions out of 59, an error rate of 3.4 per cent, within tolerance". The arithmetic is right and the conclusion is not.
The question a sample answers is not "what is the error rate?" but "how high could the error rate be, given what we found?". With exceptions in the sample, that upper bound is the error rate p at which finding two or fewer exceptions in 59 items has only a 5 per cent probability. It comes from the binomial distribution.
Find p such that P(X ≤ 2 | n = 59, p) = 5%
∑ C(59, i) pi (1 − p)59 − i for i = 0, 1, 2 = 0.05 → p = 10.3%
Excel: =BETA.INV(0.95,2+1,59-2) returns 10.29%
Check it the other way. If the true error rate were exactly 5 per cent, a sample of 59 would find two or fewer exceptions 42.9 per cent of the time. Finding two is therefore entirely consistent with a population that is at, or well above, tolerance. At 10.3 per cent the same result would occur only 5.0 per cent of the time, which is where the bound sits.
| Exceptions found | Upper bound on the error rate | Sample needed for a 5% claim | Additional clean items | Hours at 30 minutes |
|---|---|---|---|---|
| 0 | 4.95% | 59 | 0 | 29.5 |
| 1 | 7.79% | 93 | 34 | 46.5 |
| 2 | 10.29% | 124 | 65 | 62.0 |
| 3 | 12.62% | 153 | 94 | 76.5 |
There are two honest options. The first is to report what the sample proves: with two exceptions, the error rate is below 10.3 per cent at 95 per cent confidence, which is above tolerance, so the control is not shown to be working. The second is to extend the sample until the claim holds. With two exceptions in hand the total sample has to reach 124, so 65 more items must be tested, and every one of them must be clean. A third exception moves the target to 153.
At 30 minutes an item the extension costs 32.5 hours, about 4.3 working days. That is real, but it is the price of a sentence that can be defended to an examiner.
Do not draw a fresh 59 and report that one if it comes back clean. The two exceptions happened; discarding the sample that found them is the one move that turns a testing gap into a conduct issue.
The binomial figures assume a large population. When the sample is a meaningful share of the items, sampling without replacement helps. On a quarter's population of 240 allocations, with 12 errors as the tolerance boundary, the hypergeometric sample sizes are 52, 80, 104 and 126 for none to three exceptions. With two exceptions the extension shrinks from 124 to 104 items.
| Sample | Upper bound |
|---|---|
| 40 | 14.9% |
| 59 | 10.3% |
| 93 | 6.6% |
| 124 | 5.0% |
| 200 | 3.1% |
Lowering the confidence is the other lever, and the weaker one: at 90 per cent, two exceptions in 59 still only prove a rate below 8.8 per cent.
The mistake is reading the point estimate as the conclusion: 3.4 per cent sounds like a pass against a 5 per cent tolerance, and it is reported as one. Sampling conclusions are about the upper bound, never the observed rate. The second mistake is treating the exceptions only as items to fix. Each one should be corrected, with any amount refunded to the fund, and its cause traced, because two wrong allocations in a random sample are evidence about the other allocations nobody tested. Where the cause is a single rule applied wrongly, test that rule across the whole population rather than sampling further.
The sampling sheet and the case of the sample that found two are in the free workbook for this case, ready for your own population.
3.4 per cent is the point estimate, two divided by 59, but a compliance conclusion needs an upper bound. At 95 per cent confidence, two exceptions in 59 are consistent with a true error rate as high as 10.3 per cent. The test therefore no longer supports a statement that fewer than 5 per cent of items are wrong.
Enough that the total sample, exceptions included, supports the original claim. To show an error rate below 5 per cent at 95 per cent confidence with two exceptions takes 124 items, so 65 more, all of them clean. At 30 minutes an item that is 32.5 hours, about 4.3 working days.
Yes, when the sample is a meaningful share of the population. Sampling without replacement from 240 items, the sample needed with two exceptions falls from 124 to 104, and a clean sample from 59 to 52. For populations in the thousands the reduction is negligible and the binomial figures apply.
This article is one calculation from The Private Fund Compliance Officer. The book takes the same case from first principles to the decision, chapter by chapter, and every figure it prints is a live formula in the free companion workbooks.
Get the book on Amazon →Free companion files
Also on Amazon UK · Amazon Germany · Amazon France · Amazon Canada
Reading guide: regulation, compliance and banking → · All 453 articles →
If this book helped, or didn’t, a few lines on Amazon are worth more than they look: they are what the next reader goes on. Write a review. The workbook stays free either way.