Hersi Maths WhatsApp me

Understand · explore · practise

Data representations: mixed practice

Independent mixed questions on outliers, box plots, cumulative frequency, histograms, scale calibration and contextual comparisons, with worked solutions and assumption checks.

Before you startThe preceding Representations of data lessons.

01 / Choose the representation

Decide what the supplied information can answer.

Identify the variable, units, grouping, sample size and requested quantity.

Write down any interpolation or scale assumption before calculating.

Try each question before opening its hint. These original examples mix methods so that recognising what to do is part of the practice. An unsupported exact answer is not improved by extra decimal places.

Audit the methodExplore

Q₁ = 12, Q₃ = 20. Is 32 an outlier under a strict 1.5 IQR rule?

Show method and limitation

IQR 8 gives upper fence 32. Equality is not beyond the fence, so 32 is not flagged. A flag would require investigation, not automatic deletion.

02 / Flag, then investigate

An unusual observation is not automatically an error.

01 · Strict IQR fences

Q₁ = 12 and Q₃ = 20. Using 1.5 IQR fences, classify observations 0, 32 and 33.

Hint

Compute the IQR before applying the multiplier.

Worked solution

IQR = 8, fences 0 and 32. Under a strict beyond-fence rule, 0 and 32 are not flagged; 33 is flagged.

02 · SD rule

Mean 50, variance 9. Flag values strictly more than 2 SD from the mean among 44, 43, 56 and 57.

Hint

SD is √9 = 3, not 9.

Worked solution

Limits 44 and 56. Flag 43 and 57; equality at 44 or 56 is not flagged.

03 · Correction evidence

An unusually high measurement appears in the source log with matching units and a documented unusual event. Should it be replaced by the mean?

Hint

Distinguish genuine unusual values from verified recording errors.

Worked solution

No. Retain a verified in-scope observation and document the context. Replacing it by the mean would alter valid data without justification.

03 / Separate fences and whiskers

Draw observed extrema according to the convention.

04 · Modified box plot

Ordered data: 2,4,6,8,10,12,14,30. Use medians of the lower and upper halves for quartiles and 1.5 IQR fences. Give the box, whiskers and any outlier.

Hint

Q₁ is the mean of 4 and 6; Q₃ the mean of 12 and 14.

Worked solution

Median 9; Q₁ = 5; Q₃ = 13; IQR 8; fences −7 and 25. Box from 5 to 13, median at 9. Whiskers at observed 2 and 14; 30 is separately plotted.

05 · Compare

Sample A has median 18 minutes and IQR 6 minutes; B has median 21 and IQR 4. Give two comparisons.

Hint

There is a trade-off in centre and spread.

Worked solution

A has the shorter median journey time by 3 minutes. B’s middle half is less spread out, with IQR 2 minutes smaller. Neither group is uniformly better on both summaries.

06 · Missing extrema

A grouped cumulative graph begins at (0,0) and ends at (40,80). Must 0 and 40 be observed values?

Hint

They may be class boundaries.

Worked solution

No. They bound the grouped data under its class convention. Exact raw minimum and maximum are not recovered from these points.

04 / Read counts and values

Use the appropriate direction through the graph.

07 · Construct

Classes [0,10), [10,20), [20,40) have counts 5, 9 and 6. Give cumulative points including the starting point.

Hint

Accumulate at upper boundaries.

Worked solution

(0,0), (10,5), (20,14), (40,20).

08 · Estimate median

For question 7, estimate the median using straight within-class interpolation.

Hint

Target cumulative count 10 lies in [10,20).

Worked solution

10 + (10 − 5)/9×10 = 140/9 ≈ 15.56. It is a grouped estimate under uniformity within the class.

09 · Above a threshold

For question 7, estimate the count at least 30.

Hint

Half of the last class remains above 30 under uniformity.

Worked solution

6×10/20 = 3. The actual count is not fixed by the grouped table.

05 / Add areas with the right widths

A partial class requires an assumption.

Watch: half a class contributes half its count

Pause, replay or seek freely. The notes explain the same idea and stay in view.

10 · Unequal bars

Classes [0,10), [10,30), [30,40) have counts 6, 12 and 8. Find their densities.

Hint

Divide by widths 10, 20 and 10.

Worked solution

0.6, 0.6 and 0.8. The first two bars have equal height but different areas and frequencies.

11 · Count below 20

For question 10, estimate the count below 20, then its percentage of the sample.

Hint

Include all of the first class and half of the second.

Worked solution

Estimated count 6 + 12×10/20 = 12. Total 26; estimated percentage 100×12/26 ≈ 46.15%. State uniformity within the second class.

12 · Polygon points

For question 10, give points joining histogram-top centres with a density axis.

Hint

Use midpoints 5, 20 and 35.

Worked solution

(5,0.6), (20,0.6), (35,0.8). A frequency-height polygon would instead use 6, 12, 8 as heights.

06 / Recover missing scale information

Calibrate before treating area as count.

13 · Paper area

A 2 cm by 4 cm bar represents 32 observations. How many does a 3 cm by 5 cm bar represent on the same histogram?

Hint

Find observations per cm².

Worked solution

32/8 = 4 observations/cm². The second area is 15 cm², so its frequency is 60.

14 · Missing height

A width-5 class with frequency 20 has displayed height 2 cm. Find the displayed height of a width-10 class with frequency 30.

Hint

The known density 4 corresponds to 2 cm.

Worked solution

One cm represents density 2. Unknown density 30/10 = 3, so height = 1.5 cm.

07 / Keep conclusions proportionate

Use denominators and respect the study design.

15 · Different samples

A has 15 of 60 values above a threshold; B has 24 of 120. Compare the sample proportions.

Hint

Divide by each sample’s own total.

Worked solution

A: 25%; B: 20%. B has more observations above the threshold but the smaller sample proportion.

16 · Mean and SD

A report supplies only mean 20 and SD 4. A student claims exactly 16% of observations exceed 24. Is that established?

Hint

No distribution or raw observations have been supplied.

Worked solution

No. Mean and SD alone do not determine the percentage. The student is importing an unstated distributional assumption, and even a model probability would not guarantee an exact sample percentage.

08 / Review the source of each error

Revisit the method, not just the numerical answer.

If you missed a question, identify whether the cause was a boundary, quartile convention, scale, denominator, interpolation assumption or overclaim. Use the chapter links to revisit that method, then retry without the solution.

Section 1 of 8 · Choose the representation