SOX control testing sample sizes: a practical reference
Every SOX team eventually writes the same table on a whiteboard: control operates this often, so we test this many. The table isn't in the standards — neither the PCAOB nor the SEC prescribes sample sizes — but decades of firm methodologies have converged on ranges that external auditors recognize on sight. If your sizes sit inside those ranges and your rationale is written down, sample size is one of the few SOX arguments you can settle in advance.
Here is the reference version of that table, with the reasoning behind it and the edge cases that actually cause trouble.
The conventional frequency table
For manual controls tested for operating effectiveness, with no deviations expected and a full-year coverage period:
| Control frequency | Population (approx.) | Typical sample size |
|---|---|---|
| Annual | 1 | 1 |
| Quarterly | 4 | 2 |
| Monthly | 12 | 2–4 (3 is common) |
| Weekly | 52 | 5–10 (5 is common) |
| Daily | ~250 | 15–25 (25 is common) |
| Many times per day / recurring | thousands+ | 25–60 (25 baseline, 40–60 for higher risk) |
Three things to understand about this table before you rely on it:
It's a floor for low-risk, not a universal answer. These sizes trace back to attribute-sampling math: a sample of 25 with zero deviations supports roughly 90% confidence that the true deviation rate is below 9–10%; a sample of 60 tightens that materially. The small-population rows (annual, quarterly, monthly) aren't statistical at all — with a population of 4 or 12, you're doing a haphazard or complete examination, and the conventional sizes simply reflect what firms accept as sufficient coverage.
Every firm's table differs slightly. One methodology says monthly = 2, another says 3, another scales by risk within a range. If your external auditor plans to rely on management's testing, align your table with theirs at planning — it costs one email and prevents a season of re-performance.
The table assumes zero expected deviations. Conventional SOX samples are designed so that any deviation means the sample no longer supports effectiveness at the planned level. That has consequences when you find one — more below.
Adjusting for risk
Most methodologies scale within the ranges based on the risk associated with the control. Factors that push you toward the top of a range (or beyond it):
- the control mitigates a risk of material misstatement directly, with no compensating or redundant control behind it;
- it's a key control over a significant, judgmental, or fraud-susceptible area (revenue cutoff, reserves, journal entries);
- history: deviations found in prior years, control performers changed mid-year, the process was modified;
- complexity: the control requires judgment (review of an estimate) rather than a mechanical check (matching two numbers).
Conversely, a low-risk control with a clean multi-year history sits comfortably at the bottom of the range. Whichever way you go, write the risk assessment down next to the sample size — a reviewer who can see "daily control, higher risk, sample 40" reasoned in one line won't reopen the question. A bare "40" invites it.
Automated controls: test of one
An automated control that performs the same way every time doesn't need a sample of 25 — it needs a test of one, plus evidence that it kept performing the same way. The "kept performing" half is the part small teams skip: a test of one is only valid if the relevant IT general controls (change management and access over that application) were effective for the period. If ITGCs failed, the benchmark breaks, and your elegant sample of one silently becomes no evidence at all for the rest of the year.
The same logic applies to IT-dependent manual controls in reverse: the manual part gets the frequency-table sample, and the system-generated report the reviewer relied on needs its own IPE evidence. Sampling the review 25 times while never once validating the report it reviews is the most common structural gap we see in small-team SOX files.
When you find an exception
This is where sample-size discipline pays for itself, because the wrong response is instinctive: "one exception in 25 — let's pull 15 more and see." Extending the sample hoping for clean results is testing until you like the answer, and external auditors are rightly allergic to it.
The defensible sequence:
- Understand the deviation. Was it a control failure, a documentation gap, or a population error (an item that shouldn't have been in scope)? An item wrongly included in the population is removable with documentation; a control that didn't operate is not.
- Assess it as a deficiency first. A zero-deviation design means one true deviation fails the sample as planned. The honest path is usually to conclude the control did not operate effectively for that period, evaluate severity (could the deficiency result in a material misstatement, considering compensating controls?), and look at remediation plus re-testing of the remediated control for the remaining period.
- Extend only with a rationale. There are legitimate extension designs — if your methodology pre-defines that one deviation triggers an expanded sample at a defined size with defined acceptance criteria, decided before looking at results. What you cannot do is invent the extension after seeing the exception.
Document the exception, the evaluation, and the conclusion in the test itself, not in a side memo. When the same finding is re-examined months later during the deficiency evaluation, the trail should already be there.
Selection: the part reviewers actually challenge
Sample size debates are rare once your table matches your auditor's. Sample selection challenges are common, because selection is where bias creeps in:
- Define the population first, and keep evidence of it: the report name, parameters, date run, and row count. "Selected 25 invoices" from an undefined population is untestable — completeness of the population is the first thing a reviewer asks about.
- Select randomly, or at least demonstrably without bias. Haphazard selection is accepted by most methodologies for small populations, but "the tester picked 25 items" ages badly when three of them turn out to be from the same easy week. Random selection with a documented method costs nothing and removes the argument.
- Make it reproducible. The gold standard is a selection someone else can re-run and get the identical items: population reference, selection method, and — if random — the seed. This matters more than most teams realize, because sample selections get questioned at year-end, months after the tester made them, and "I used Excel RAND() and didn't save it" means the selection is unverifiable by anyone, including you.
- Cover the period. For a full-year conclusion, the sample should span the year, not cluster in Q1 when the tester had time. If you tested at interim, the roll-forward period needs its own coverage — either additional selections or an update procedure proportionate to the remaining months.
This is one of the places tooling genuinely helps. SoxDesk, for what it's worth, generates suggested sizes from the control's frequency and does seeded random selection — the population reference, seed, and selected items are stored on the test itself, so the selection is reproducible by anyone and freezes once the test is submitted. No more archaeology on how a sample was picked.
A worked example
Monthly bank reconciliation review, key control, low risk, no prior deviations:
- Population: 12 monthly reconciliations, Jan–Dec.
- Size: 3 (mid-range monthly; low risk documented).
- Selection: random, seed recorded; months drawn: March, July, November — reasonable spread across the year, no reselection needed.
- Interim/roll-forward: tested through September at interim (March, July); November tested in January as the roll-forward selection.
- Result: no deviations; conclude effective for the year, referencing the interim work plus roll-forward.
Ten minutes of documentation, and every future question — why 3, why those months, why that's enough for December — has a written answer.
The one-paragraph policy worth writing
If your SOX program doesn't have a sampling policy, write one this week; it can be a page. State your frequency table, your risk-adjustment logic, your selection method and reproducibility requirement, your interim/roll-forward approach, and your exception-handling sequence. Share it with your external auditor at planning and get their nod in writing. From then on, every sample-size conversation for the rest of the year is a reference to the policy instead of a negotiation.
SoxDesk is an on-premise audit workflow app for SOX and internal audit teams — enforced tester/reviewer sign-offs, seeded reproducible sampling, ITGC scoping linkage, all on your own network. If sample selection archaeology sounds familiar, the 60-day free trial is the full product: soxdesk.com.