01 · One row
A value x = 4 has frequency f = 3. Calculate fx² and (fx)².
Hint
The brackets change what is squared.
Worked solution
fx² = 3×16 = 48. (fx)² = 12² = 144. The weighted squared-sum column needs 48.
Understand · explore · practise
Calculate weighted variance and standard deviation for exact frequency data, estimate them from grouped midpoints, and explain what the summaries cannot reveal.
Before you startVariance, weighted means and grouped midpoints.
01 / Weight every squared value
Descriptive variance = Σfx²/Σf − (Σfx/Σf)²
For exact values x, the result is exact apart from arithmetic rounding.
Use n = Σf, not the number of rows. In fx², only x is squared: multiply the frequency by the squared data value.
Exact values 1, 3, 5 have frequencies 2, k, 2.
| x | f | fx | fx² |
|---|---|---|---|
| 1 | 2 | 2 | 2 |
| 3 | 0 | 0 | 0 |
| 5 | 2 | 10 | 50 |
n = 4; Σfx = 12; Σfx² = 52.
Mean = 3; variance = 52/4 − 9 = 4; SD = 2.
Adding observations equal to the mean leaves total squared deviation 16 unchanged, but spreads it across a larger count. The descriptive variance falls.
02 / Build the four columns
fx: 2, 12, 10 → Σfx = 24
The total count is eight.
fx²: 2, 36, 50 → Σfx² = 88
Square each value, then multiply by its frequency.
Mean = 24/8 = 3
The symmetric frequency pattern centres the data at three.
A value x = 4 has frequency f = 3. Calculate fx² and (fx)².
The brackets change what is squared.
fx² = 3×16 = 48. (fx)² = 12² = 144. The weighted squared-sum column needs 48.
Values 2, 6 have frequencies 5, 3. Find n, Σfx and Σfx².
A two-row table contains eight observations.
n = 8; Σfx = 5×2 + 3×6 = 28; Σfx² = 5×4 + 3×36 = 128.
03 / Calculate the descriptive spread
Pause, replay or seek freely. The notes explain the same idea and stay in view.
Variance = 88/8 − 3² = 2
Use the weighted sums.
SD = √2 ≈ 1.414
Take the positive square root.
Sxx = 2×(−2)² + 4×0² + 2×2² = 16
Direct deviations give the same variance 16/8.
Use the summaries from question 2 to find mean and descriptive variance.
Mean = 28/8 = 3.5.
Variance = 128/8 − 3.5² = 16 − 12.25 = 3.75; SD = √3.75 ≈ 1.936.
Value 100 appears in a table with frequency zero. What does it contribute to the sums?
There is no observation equal to 100.
It contributes zero to n, Σfx and Σfx². It also does not define an observed maximum.
04 / Substitute class midpoints
For intervals, replace x by each class midpoint m. The resulting variance Σfm²/Σf − (Σfm/Σf)² and its square root are estimates. They describe the midpoint-substituted data, not the exact unknown observations.
Classes [0,2), [2,4), [4,6) have frequencies 2, 4, 2. Estimate mean, variance and SD.
The midpoints are 1, 3, 5.
Estimated mean = 3; estimated variance = 2; estimated SD = √2 ≈ 1.414. These match the exact table above only because we substituted its values as midpoints.
All observations lie in [10,20). Midpoint substitution gives estimated SD zero. Does that prove the observations were identical?
All values were replaced by 15, losing within-class variation.
No. For example, observations 10 and 19 lie in the same class but have positive exact SD. Zero midpoint SD can hide substantial within-class spread.
05 / Use frequency weights, not interval widths
Classes [0,10) and [10,30) have frequencies 4 and 6. Estimate variance and SD.
Midpoints 5 and 20; n = 10; estimated mean 14.
Σfm² = 4×25 + 6×400 = 2500. Estimated variance = 2500/10 − 14² = 54; SD ≈ 7.348.
An occupied class is “30 or more”. Can you form its midpoint-square term without assumptions?
The upper boundary is not supplied.
No. A midpoint is not defined. Obtain more information or clearly justify an additional model before estimating mean or SD.
06 / Solve for an unknown frequency
Exact values 1, 3, 5 have frequencies 2, k, 2. Their descriptive variance is 2. Find k.
The mean is three and Sxx = 16 for every non-negative k.
16/(4 + k) = 2 gives 4 + k = 8, so k = 4.
For that table with k = 4, Sxx = 16 and n = 8. Find the variance using denominator n − 1.
This question explicitly asks for the alternative denominator.
s² = 16/7, not 16/8. Its SD is √(16/7) ≈ 1.512. Identify the convention when comparing answers.
07 / Do not invent threshold counts
The same mean and SD can arise from datasets with different shapes and different counts above a threshold. Statements such as “about 68% within one SD” require an appropriate distributional model; they are not automatic properties of every dataset.
Compare A = −1, −1, 1, 1 and B = −√3, 1/√3, 1/√3, 1/√3. Both have mean 0 and descriptive SD 1. Do they have the same number strictly above zero?
Count the positive values directly.
No. A has two; B has three. For B, the sum is −√3 + 3/√3 = 0 and squared sum is 3 + 3×(1/3) = 4, giving variance 1. The two summaries cannot recover the count.
Classes [0,10), [10,20), [20,30) have frequencies 4, 8, 3. How many observations are definitely below 15, and what is the possible total below 15?
The first class is wholly below 15; the second straddles it.
Four are definitely below 15. Between zero and eight from the middle class may also be below it, so the exact count can be any integer from 4 to 12. A uniformity assumption would estimate 8, but grouping alone does not prove that count.
08 / Calculate and qualify
Check the frequency total, square the value rather than its weighted total, use the requested denominator, and label midpoint results as estimates. Do not infer unobserved threshold counts from a mean and SD alone.
Section 1 of 8 · Weight every squared value