01 · Train journey
Find P(train).
Hint
Add both train cells.
Worked solution
(20+5)/50 = 1/2.
Understand · explore · practise
Choose the correct population and denominator, calculate event probabilities from category tables and distinguish exact grouped counts from histogram estimates.
Before you startFractions, frequency density and reading grouped data.
01 / Define what is chosen
For uniform random selection of one record: P(event) = qualifying records ÷ all eligible records.
State the population and the selection rule before calculating.
This can be exact for the recorded finite population. Using its proportion to predict a future wait is a separate modelling step: representativeness and stable conditions matter.
Forty recorded waits: 0 ≤ t < 10: 8; 10 ≤ t < 20: 12; 20 ≤ t < 40: 20. Select one record uniformly at random.
At c = 0, no recorded wait is below 0: probability 0, exactly from these class boundaries.
Inside a class, the shaded proportion estimates the count by assuming constant density within that class.
02 / Count joint categories
Original records for 50 journeys: bus/on-time 18, bus/late 7, train/on-time 20, train/late 5. Select one of these 50 records uniformly.
The event specifies both bus and late
Only 7 records meet both conditions.
P(bus and late) = 7/50
All 50 records remain eligible for selection.
Find P(train).
Add both train cells.
(20+5)/50 = 1/2.
Find P(on time).
Include both types of transport.
(18+20)/50 = 38/50 = 19/25.
Why is 7/25 not the answer for a uniformly chosen journey being both bus and late?
Which records are eligible?
The experiment chooses from all 50 journeys, not only the 25 bus journeys. The joint probability is 7/50. Selecting only from bus records would define a different experiment.
03 / Avoid counting an overlap twice
Find P(bus or late), including journeys that satisfy both.
Count bus/on-time, bus/late and train/late once each.
(18+7+5)/50 = 30/50 = 3/5. Equivalently (25+12−7)/50.
Find P(neither bus nor late).
Which single cell remains?
Only train/on-time qualifies: 20/50 = 2/5. This complements the previous union.
04 / Use histogram area
Pause, replay or seek freely. The notes explain the same idea and stay in view.
Frequency densities are 0.8, 1.2 and 1
Each density is frequency divided by class width.
Total histogram area is 8 + 12 + 20 = 40
The widest bar contains the most records despite not being the tallest.
P(t < 20) = 20/40 = 1/2
The first two complete classes determine this exactly for the recorded population.
Find P(20 ≤ t < 40).
Use frequency, not bar height.
20/40 = 1/2. Dividing its height by the sum of the heights is invalid with unequal widths.
Which class has the highest frequency density, and which has the most records?
Compare density and frequency separately.
The 10–20 class has the highest density, 1.2. The 20–40 class has the greatest frequency, 20.
05 / A threshold inside a class
Estimate a partial count using the fraction of the class width.
Assume observations are spread uniformly within that class.
All 8 records in 0–10 qualify
The first class is complete.
Estimate half of the 12 records in 10–20 below 15
Estimated partial count = (5/10)×12 = 6.
Estimated probability = (8+6)/40 = 7/20
The exact count below 15 cannot be recovered from these groups.
Estimate P(t < 30).
Use half of the final class.
Estimated count = 8+12+(10/20)×20 = 30, giving 30/40 = 3/4.
Estimate P(t < 5).
Use half the first class.
Estimated count = (5/10)×8 = 4, giving 4/40 = 1/10. This depends on the within-class model.
06 / Say what is known exactly
What bounds on P(t < 15) follow from the groups alone?
All 8 first-class observations qualify; anywhere from 0 to 12 of the next class may qualify.
The count lies from 8 to 20 inclusive, so 1/5 ≤ P(t < 15) ≤ 1/2. The estimate 7/20 is within these bounds.
Estimate P(15 ≤ t < 30).
Take half the second class and half the third.
Estimated count = 6+10 = 16, giving 16/40 = 2/5.
Does the exact proportion of on-time journeys in these 50 records prove the probability for next month is 19/25?
The population has changed from a fixed collection to future journeys.
No. It provides an estimate under assumptions about how representative these records are and whether conditions remain similar.
07 / Read the interval convention
Does a recorded wait of exactly 20 minutes contribute to P(t < 20)?
The event uses a strict inequality.
No. It belongs to 20 ≤ t < 40 and is excluded from t < 20. For rounded records, the exact probability of t ≤ 20 cannot be found unless the count equal to 20 is known.
Five additional journeys have unknown punctuality. Can you silently count them as late?
Unknown is not a response category you observed.
No. Preserve the missing status. State whether the calculation concerns only the 50 complete records, and explain that extending it to all 55 requires more information or assumptions.
08 / Counts, areas and assumptions
Identify who or what is selected. Count joint categories carefully, subtract overlap when needed and use histogram areas for unequal class widths. Distinguish exact probabilities for a recorded finite population, within-class estimates and predictions about future observations.
Section 1 of 8 · Define what is chosen