Articles

What does a compliance test sample with two exceptions prove?

Attribute sampling when the sample is not clean: the upper error bound, the ladder of sample sizes, and why two out of 59 is not 3.4 per cent.

Two exceptions in a 59-item sample prove only that the error rate is below 10.3 per cent at 95 per cent confidence, not below 5 per cent. The 59 supports the 5 per cent claim only if every item is clean. To make the claim again with two exceptions already found, the sample has to grow to 124 items, 65 more, all clean.

Worked in full in The Private Fund Compliance Officer by Julian R. Sterling, with every figure reproduced in a free workbook.See the book on Amazon →

Where 59 comes from

Compliance testing of fee calculations, expense allocations or trade approvals is attribute sampling: each item is either right or wrong. The standard zero-error formula gives the sample that supports a claim if nothing is found.

n = ln(1 − confidence) / ln(1 − tolerable error rate) = ln(0.05) / ln(0.95) = 58.40, rounded up to 59

The derivation of that figure, and what twenty clean items prove (an error rate below 13.9 per cent), are set out in how many expense allocations you have to test. The formula has a condition built in: zero exceptions. Most testing programmes do not stop to ask what happens when the condition fails.

The case

A compliance officer tests 59 expense allocations, selected at random from the year, at 95 per cent confidence against a tolerable error rate of 5 per cent. Two allocations are wrong. The draft report reads "two exceptions out of 59, an error rate of 3.4 per cent, within tolerance". The arithmetic is right and the conclusion is not.

The calculation, step by step

The question a sample answers is not "what is the error rate?" but "how high could the error rate be, given what we found?". With exceptions in the sample, that upper bound is the error rate p at which finding two or fewer exceptions in 59 items has only a 5 per cent probability. It comes from the binomial distribution.

Find p such that P(X ≤ 2 | n = 59, p) = 5%

∑ C(59, i) pi (1 − p)59 − i for i = 0, 1, 2 = 0.05 → p = 10.3%

Excel: =BETA.INV(0.95,2+1,59-2) returns 10.29%

Check it the other way. If the true error rate were exactly 5 per cent, a sample of 59 would find two or fewer exceptions 42.9 per cent of the time. Finding two is therefore entirely consistent with a population that is at, or well above, tolerance. At 10.3 per cent the same result would occur only 5.0 per cent of the time, which is where the bound sits.

What a sample of 59 proves, by number of exceptions, at 95 per cent confidence
Exceptions foundUpper bound on the error rateSample needed for a 5% claimAdditional clean itemsHours at 30 minutes
04.95%59029.5
17.79%933446.5
210.29%1246562.0
312.62%1539476.5

Restoring the claim

There are two honest options. The first is to report what the sample proves: with two exceptions, the error rate is below 10.3 per cent at 95 per cent confidence, which is above tolerance, so the control is not shown to be working. The second is to extend the sample until the claim holds. With two exceptions in hand the total sample has to reach 124, so 65 more items must be tested, and every one of them must be clean. A third exception moves the target to 153.

At 30 minutes an item the extension costs 32.5 hours, about 4.3 working days. That is real, but it is the price of a sentence that can be defended to an examiner.

Do not draw a fresh 59 and report that one if it comes back clean. The two exceptions happened; discarding the sample that found them is the one move that turns a testing gap into a conduct issue.

What if the population is small?

The binomial figures assume a large population. When the sample is a meaningful share of the items, sampling without replacement helps. On a quarter's population of 240 allocations, with 12 errors as the tolerance boundary, the hypergeometric sample sizes are 52, 80, 104 and 126 for none to three exceptions. With two exceptions the extension shrinks from 124 to 104 items.

Upper bound with two exceptions, by sample size, 95 per cent confidence
SampleUpper bound
4014.9%
5910.3%
936.6%
1245.0%
2003.1%

Lowering the confidence is the other lever, and the weaker one: at 90 per cent, two exceptions in 59 still only prove a rate below 8.8 per cent.

The common mistake

The mistake is reading the point estimate as the conclusion: 3.4 per cent sounds like a pass against a 5 per cent tolerance, and it is reported as one. Sampling conclusions are about the upper bound, never the observed rate. The second mistake is treating the exceptions only as items to fix. Each one should be corrected, with any amount refunded to the fund, and its cause traced, because two wrong allocations in a random sample are evidence about the other allocations nobody tested. Where the cause is a single rule applied wrongly, test that rule across the whole population rather than sampling further.

Takeaway

The sampling sheet and the case of the sample that found two are in the free workbook for this case, ready for your own population.

Questions readers ask

If a sample of 59 finds two errors, is the error rate 3.4 per cent?

3.4 per cent is the point estimate, two divided by 59, but a compliance conclusion needs an upper bound. At 95 per cent confidence, two exceptions in 59 are consistent with a true error rate as high as 10.3 per cent. The test therefore no longer supports a statement that fewer than 5 per cent of items are wrong.

How many more items do you need to test after finding exceptions?

Enough that the total sample, exceptions included, supports the original claim. To show an error rate below 5 per cent at 95 per cent confidence with two exceptions takes 124 items, so 65 more, all of them clean. At 30 minutes an item that is 32.5 hours, about 4.3 working days.

Does a small population reduce the sample size?

Yes, when the sample is a meaningful share of the population. Sampling without replacement from 240 items, the sample needed with two exceptions falls from 124 to 104, and a clean sample from 59 to 52. For populations in the thousands the reduction is negligible and the binomial figures apply.

Read the whole case

This article is one calculation from The Private Fund Compliance Officer. The book takes the same case from first principles to the decision, chapter by chapter, and every figure it prints is a live formula in the free companion workbooks.

Get the book on Amazon →Free companion files

Also on Amazon UK · Amazon Germany · Amazon France · Amazon Canada

Also on this site

Reading guide: regulation, compliance and banking → · All 453 articles →

If this book helped, or didn’t, a few lines on Amazon are worth more than they look: they are what the next reader goes on. Write a review. The workbook stays free either way.