Hersi Maths WhatsApp me

Understand · explore · practise

Least squares and residuals

Calculate signed vertical residuals, compare sums of squared residuals and understand what a y-on-x least-squares fit minimises, including outlier influence and model limits.

Before you startScatter diagrams, straight-line equations and squaring signed numbers.

01 / Measure a prediction error

Compare an observation with the line at the same x.

Residual e = observed y − predicted ŷ.

For regression of y on x, measure the difference vertically, holding x fixed.

A positive residual places the observation above the fitted line; a negative residual places it below. The fitted value ŷ is a model prediction, not a replacement for the observed y.

Compare candidate linesExplore

Constructed data: (0,1), (1,2), (2,2), (3,5). Candidate prediction ŷ = a + bx.

For a = 0 and b = 0, predictions are all 0; residuals are 1, 2, 2, 5; squared residuals sum to 34.

This menu compares selected candidates. The global least-squares solution is a = 0.7, b = 1.2, with squared residual total 1.8.

02 / Keep the sign before squaring

Overprediction and underprediction are different.

The candidate line is ŷ = 1 + x. At x = 2 the observed response is y = 2.Worked example

Predicted response = 1 + 2 = 3

Substitute the observation’s x.

Residual = 2 − 3 = −1

The model overpredicts by 1.

Squared residual = (−1)² = 1

Squaring removes the sign for the objective.

01 · Above the line

At x = 4, the model ŷ = 2 + 3x predicts a response. The observed y is 17. Find the residual.

Hint

Observed minus predicted.

Worked solution

Prediction 14; residual 17 − 14 = 3. The model underpredicts by 3.

02 · Below the line

A model predicts 12 but the observed response is 8. Give the residual and its square.

Hint

Square the whole signed residual.

Worked solution

Residual −4; square 16. It is not −16.

03 / Add squared vertical errors

Signed errors can cancel even for a poor fit.

Watch: opposite errors both contribute

Pause, replay or seek freely. The notes explain the same idea and stay in view.

For the model data and candidate ŷ = 1 + x:Worked example

Predictions: 1, 2, 3, 4

One fitted value for each original x.

Residuals: 0, 0, −1, 1

Their signed sum is zero.

Squared total: 0 + 0 + 1 + 1 = 2

Zero signed total does not mean every prediction is correct.

03 · Cancellation

Two residuals are −5 and 5. Find their signed total and squared total.

Hint

Add squares, not the square of the sum.

Worked solution

Signed total 0; squared total 25 + 25 = 50. Squaring the signed total would wrongly give 0.

04 · Compare candidates

On the same observations, line A has squared residual total 12 and line B has 9. Which is preferred by this criterion?

Hint

Smaller is better for this objective on the same response scale.

Worked solution

B. This comparison alone does not establish that B is a suitable model, a causal relationship or the best line among all possibilities.

04 / Understand the fitted line

Least squares selects the smallest squared total.

For the four constructed points, the y-on-x least-squares line is ŷ = 0.7 + 1.2x.Worked example

Predictions: 0.7, 1.9, 3.1, 4.3

Keep enough precision.

Residuals: 0.3, 0.1, −1.1, 0.7

Observed minus predicted, in the original row order.

Squared total = 0.09 + 0.01 + 1.21 + 0.49 = 1.8

This improves on the candidate total 2.

05 · Through the mean

The four x values have mean 1.5 and the four y values mean 2.5. Verify the fitted line passes through their mean point.

Hint

Substitute x = 1.5.

Worked solution

0.7 + 1.2×1.5 = 2.5. An ordinary unweighted least-squares line with a fitted intercept passes through (x̄,ȳ).

06 · Does it pass through all points?

Does least squares require every observation to be on the fitted line?

Hint

The objective minimises the total; it need not make it zero.

Worked solution

No. The example’s minimum is 1.8, with nonzero residuals. A line through every point exists only when the paired observations are exactly collinear in the relevant direction.

05 / Name the residual direction

The criterion depends on what is being predicted.

y on x minimises Σ[y − (a + bx)]².

It does not minimise perpendicular distances to the line.

07 · Horizontal error

A student measures each point’s horizontal gap from a line and calls it the y-on-x residual. What is wrong?

Hint

Hold the explanatory x fixed when predicting y.

Worked solution

The y-on-x residual is the vertical difference. Horizontal residuals belong to a different prediction direction.

08 · Reversing roles

Can you assume the fitted x-on-y line is just the same fitted y-on-x equation rearranged?

Hint

The minimised errors change direction.

Worked solution

No, not in general. A separate fit minimises horizontal rather than vertical errors. The two coincide for a perfect nonvertical, nonhorizontal linear relationship.

06 / Inspect large residuals and distant x values

Squaring makes large errors contribute strongly.

09 · Squared contributions

Compare the contributions of residuals 2 and 6 to the squared total.

Hint

Compare 2² with 6².

Worked solution

4 and 36: the second contributes nine times as much, although its magnitude is only three times as large.

10 · Delete to improve?

Removing an unusual point lowers the fitted squared total. Is that alone a valid reason to delete it?

Hint

Fewer observations and a different sample change the comparison.

Worked solution

No. Investigate errors, scope and influence. Do not discard genuine observations merely to make the fit look better; document any justified exclusion.

07 / Check more than the fitted total

A calculated optimum may still be a poor model.

11 · Curvature

A curved scatter pattern has a computed best straight line. Does minimising the squared total prove the straight-line model is appropriate?

Hint

Best within one family is not the same as suitable.

Worked solution

No. Systematic curvature in the residuals or scatter may suggest a different model or a restricted range.

12 · Units

Responses measured in metres are converted to centimetres and the corresponding predictions are also multiplied by 100. What happens to each residual and the squared total?

Hint

Residuals scale once; squares scale twice.

Worked solution

Residuals multiply by 100; the squared total multiplies by 10,000. Do not compare raw totals across response scales as if the units were unchanged.

08 / Observe, predict, subtract, square

Keep the data, prediction direction and objective explicit.

Find ŷ at each observed x, calculate y − ŷ and add the squares. A least-squares fit minimises this total within its model family. Inspect unusual observations, residual patterns and context before trusting predictions.

Section 1 of 8 · Measure a prediction error