Hersi Maths WhatsApp me

Understand · explore · practise

Statistics review: displays and regression

Combine histogram area, box plots, percentiles, outlier checks and regression interpretation in original cumulative practice.

Before you startRepresentations of data and correlation.

01 / Read what the diagram actually represents

Use area, position or trend as appropriate.

Histogram frequency is area; box-plot comparisons need context.

A regression line models one response from one explanatory variable. None of these displays alone proves causation. The numerical data in this review are constructed.

Area beyond a thresholdExplore

Constructed grouped times: 0–10 minutes has frequency 8; 10–20 has 12; 20–40 has 20. Assume uniform spread within each class.

02 / Use density when widths differ

The tallest bar need not contain the most observations.

01 · Heights

For the three classes in the model, calculate the frequency densities.

Hint

Divide each frequency by its class width.

Worked solution

8/10=0.8, 12/10=1.2 and 20/20=1.0 observations per minute.

02 · Largest group

Which class contains the most observations, and which has the tallest bar?

Hint

Compare area and height separately.

Worked solution

The 20–40 class contains 20 observations, the largest frequency. The 10–20 class is tallest, with density 1.2.

03 · Partial class

Estimate the number with time above 15 minutes.

Hint

Take the remaining half of the middle class plus the last class.

Worked solution

0.5×12+20=26. Estimated probability is 26/40=0.65, using uniform spread within the middle class.

04 · Another threshold

Estimate the proportion above 30 minutes.

Hint

Only half the final class remains.

Worked solution

10/40=0.25. This is an estimate from grouped data, not an exact count from raw times.

Watch: a threshold cuts through a histogram bar

Pause, replay or seek freely. The notes explain the same idea and stay in view.

03 / Connect area to cumulative positions

Read the class before interpolating.

Use cumulative frequencies 8, 20 and 40 at upper boundaries 10, 20 and 40.Worked example

Median position = 40/2=20

This is the boundary 20, so the grouped estimate is 20 minutes.

Lower quartile position = 10

Q₁≈10+((10−8)/12)×10=11.6667 minutes.

Upper quartile position = 30

Q₃≈20+((30−20)/20)×20=30 minutes. IQR≈18.3333 minutes.

05 · Cumulative endpoints

Why include a point at the initial lower boundary with cumulative frequency zero?

Hint

No observations have accumulated yet.

Worked solution

It anchors the graph at the start of the covered range. Here use (0,0), followed by (10,8), (20,20) and (40,40).

06 · Interpretation

Does Q₃=30 mean exactly 30 observations equal 30 minutes?

Hint

A percentile is a position, not a repeated value count.

Worked solution

No. It estimates the value below which roughly 75% of observations lie.

04 / Compare centre and spread with context

Use both a location measure and a variability measure.

Two constructed completion-time samples have (Q₁, median, Q₃) of (18,24,31) for group A and (20,27,30) for group B, all in minutes.

07 · Centre comparison

Which group typically completes sooner?

Hint

Compare medians and keep the time context.

Worked solution

Group A has the lower median, 24 rather than 27 minutes, suggesting typically quicker completion.

08 · Spread comparison

Which group has more consistent middle-half times?

Hint

Compare IQRs.

Worked solution

A has IQR 31−18=13 minutes; B has 30−20=10 minutes. Group B has less spread in its middle half.

05 / Use the stated rule and then investigate

An outlier is not automatically an error.

09 · IQR fence

Using group A and the 1.5×IQR rule, find the upper fence and classify a 52-minute result.

Hint

Upper fence = Q₃+1.5IQR.

Worked solution

31+1.5×13=50.5. Since 52>50.5, it is flagged as an upper outlier under this rule.

10 · Retain or remove

Should that 52-minute observation automatically be deleted?

Hint

It may be a valid unusual outcome.

Worked solution

No. Check the record and context. Correct a verified error with documentation; retain a valid observation unless the analysis has a justified exclusion rule.

06 / Interpret units and stay near supported data

A slope is a predicted change per unit of the explanatory variable.

A constructed regression of weekly orders y, measured in hundreds, on advertising spend x, measured in tens of pounds, is y=1.5+0.8x. The observed x values ranged from 2 to 12.

11 · Slope units

Interpret the slope 0.8 in ordinary units.

Hint

One x-unit is £10; one y-unit is 100 orders.

Worked solution

An additional £10 of advertising spend is associated with a predicted increase of 80 weekly orders within this model.

12 · Supported prediction

Predict weekly orders at spend £60.

Hint

x=6, then convert hundreds to orders.

Worked solution

y=1.5+0.8×6=6.3 hundreds, or 630 orders. This is interpolation within the observed spend range.

13 · Extrapolation

Why is predicting at £300 less secure?

Hint

Compare x=30 with the observed range.

Worked solution

It is far beyond x=2 to 12. The fitted linear relationship may not continue there.

14 · Reverse prediction

Can the y-on-x equation automatically serve as the least-squares x-on-y regression?

Hint

The residual directions differ.

Worked solution

No. Solving the equation algebraically does not generally produce the regression fitted in the reverse direction.

07 / Interpret patterns without overclaiming

A useful model still has limitations.

15 · Causal claim

Does a positive fitted slope prove advertising caused extra orders?

Hint

Could another variable affect both?

Worked solution

No. Association alone does not establish causation; seasonality or other changes might influence both variables.

16 · Intercept

Must 1.5 mean that zero advertising causes exactly 150 orders?

Hint

Zero is outside the observed range, and a fitted line is a model.

Worked solution

No. It is the model intercept, but x=0 is outside the observed range and the equation is not a deterministic or causal law.

08 / Check the scale before the calculation

State what is exact, estimated or modelled.

Histogram areas determine frequencies; grouped percentiles assume within-class spread. Compare box plots in context and investigate outliers. Interpret regression units carefully, distinguish interpolation from extrapolation and keep association separate from causation.

Section 1 of 8 · Read what the diagram actually represents