01 · Interval and start
A frame has 360 units and a sample of 30 is required. Find the interval and the range for a random start.
Hint
Use N/n.
Worked solution
The interval is 12. Choose the start uniformly from integers 1–12.
Understand · explore · practise
Calculate a sampling interval, choose a random start and select a fixed-size sample. Explore how repeating patterns in a list affect systematic samples.
Before you startSampling frames, simple random sampling and arithmetic sequences.
01 / Regular spacing, random start
For N = nk, choose a random start r from 1 to k.
Select labels r, r + k, r + 2k, …, r + (n − 1)k.
For 24 records and a sample of 6, the interval is 24 ÷ 6 = 4. Choose the first label uniformly from 1, 2, 3 and 4, then take every fourth record. There are four possible systematic samples under this rule.
Try all four starts. Compare the repeating values with the rearranged list.
Selected labels: 1, 5, 9, 13, 17, 21.
Observed values: 4, 4, 4, 4, 4, 4.
Sample mean = 4; population mean = 10.
Full list: 4, 8, 12, 16, 4, 8, 12, 16, 4, 8, 12, 16, 4, 8, 12, 16, 4, 8, 12, 16, 4, 8, 12, 16.
The labels are positions in the list. This is a manual comparison of starts, not a random draw. For actual sampling, choose the start randomly. Both lists contain the same invented values in different orders.
02 / Calculate the interval
k = 120 ÷ 15 = 8
The interval is eight positions.
Choose r uniformly at random from 1–8
The start must be chosen from one full interval.
If r = 6, take 6, 14, 22, …, 118
There are 15 labels, since the final label is 6 + 14 × 8.
A frame has 360 units and a sample of 30 is required. Find the interval and the range for a random start.
Use N/n.
The interval is 12. Choose the start uniformly from integers 1–12.
With interval 12, start 5 and sample size 30, find the last label.
There are 29 jumps after the first selected unit.
5 + 29 × 12 = 353.
03 / Write a complete procedure
“Take every tenth person” leaves the first selection unspecified. A reproducible method identifies the eligible list, numbers it in its existing or stated order, chooses a random start from the first interval and stops after the required number of units.
Describe a systematic sample of 20 from 200 numbered customer records.
There should be ten equally possible starts.
Check the frame contains the 200 eligible records once. Choose a uniform random integer r from 1–10. Select r, r + 10, …, r + 190, giving 20 distinct records.
A researcher always starts at record 1 and takes every fifth record. What part of a random-start systematic design is missing?
The start has been predetermined.
The start should be chosen randomly from 1–5. Always starting at 1 gives only one fixed sample under the stated list order.
04 / Not a simple random sample
In the 24-record model, only four six-record subsets are possible. A simple random sample would allow every six-record subset. A random-start systematic design can give every individual the same inclusion probability while still assigning probability zero to many subsets.
For N = 24, n = 6 and k = 4 with a fair start, what is the probability that record 10 is selected?
Which start produces 2, 6, 10, …?
Only start 2, so the probability is 1/4. Every record belongs to exactly one of the four possible samples.
Could a sample containing both records 1 and 2 arise under the interval-four rule?
Selected labels all have the same remainder modulo four.
No. Record 1 requires start 1 and record 2 requires start 2. One sample uses only one start.
05 / Watch for repeating patterns
The population mean is (4 + 8 + 12 + 16)/4 = 10
Each value occurs six times.
Start 1 selects six 4s; start 4 selects six 16s
The sample means are 4 and 16.
All four possible means are 4, 8, 12 and 16
The realised estimate may be far from 10.
Their average over equally likely starts is 10
In this complete divisible example, the issue is high sampling variability, not a systematic shift of the average estimator.
Describe the actual pattern risk rather than claiming any ordered list must be biased. Random start is important, but it does not ensure a balanced realised sample.
Pause, replay or seek freely. The notes explain the same idea and stay in view.
Four machines contribute one product each in a repeating order. A sampler takes every fourth product. What could happen?
The interval matches the cycle length.
The sample can contain products from only one machine. If machines produce different results, the sample may poorly represent overall production.
06 / Use context to assess the order
Systematic sampling spreads selections through a list and is easy to carry out. Its quality depends on the relationship between list order, interval and measured variable. A pattern aligned with the interval can matter; the mere fact that a list is alphabetical does not establish bias.
A simple random design avoids locking every selection to one fixed interval. Randomly shuffling a complete frame first can also remove an existing order pattern, although it adds practical work.
Improve the claim “The sample is biased because the list is ordered.”
Name a possible relationship between the ordering and interval.
For example: “If the list repeats groups every k positions and the sampling interval is k, one start may repeatedly select the same group, producing an unbalanced sample.” Without such context, ordering alone is not enough to establish bias.
07 / When N/n is not an integer
If N = 50 and n = 8, N/n = 6.25. Simply using interval 6 and starts 1–6 with eight observations covers only labels 1–48; labels 49 and 50 cannot be selected. State how the design handles the remainder.
One fixed-size alternative is a fractional interval: choose u uniformly in (0, 6.25], then take the ceilings of u, u + 6.25, …, u + 7 × 6.25. This gives eight distinct valid labels and treats all 50 labels symmetrically in inclusion probability. Use a method appropriate to the question and explain it clearly.
Positions: 2.5, 8.75, 15, 21.25, 27.5, 33.75, 40, 46.25
Add 6.25 each time.
Ceilings: 3, 9, 15, 22, 28, 34, 40, 47
The ceiling is the smallest integer at least as large as the position.
What is wrong with claiming the rounded-interval-six method above gives every unit a chance?
Find the largest possible final label.
The largest final label is 6 + 7 × 6 = 48. Labels 49 and 50 have zero chance.
08 / Selected records may lack measurements
A weather table may contain a missing gust reading on a selected date. Selecting 20 distinct dates does not guarantee 20 usable gust values. Report missing records and follow a defensible preplanned rule; do not quietly substitute convenient dates or treat missing values as zero.
A systematic sample selects 18 dates; 3 have no reading for the variable studied. How many usable values remain?
Count only records with a reading.
15 usable values. The requested sample contained 18 dates; its usable measurement count is smaller.
Why should “not available” not automatically become a zero?
Missingness does not state the measured quantity.
Zero is an actual value. Replacing missing readings by zero invents observations and can distort a mean or other statistic.
09 / Interval, start and ordering
A frame contains 480 invoices and you need 40. Describe the selection, including the last label if the random start is 9.
There are 39 jumps after the start.
Use interval 480/40 = 12. Randomly choose a start from 1–12, then take every twelfth invoice until 40 are selected. For start 9, the last label is 9 + 39 × 12 = 477. Check that invoice order has no problematic cycle aligned with 12.
Section 1 of 9 · Regular spacing, random start