Hersi Maths WhatsApp me

Understand · explore · practise

Hypothesis-test conclusions and limitations

Write justified contextual conclusions, distinguish statistical evidence from proof or causation and evaluate sampling and repeated-testing limitations.

Before you startComplete one-tailed and two-tailed tests.

01 / A decision needs an evidence statement

Match the original alternative and the stated level.

Reject: sufficient evidence for the alternative under the model.

Do not reject: insufficient evidence for that alternative at the stated level. Neither decision proves a hypothesis.

Name the population probability in context. Include the test level and any important limitations. A bare “reject” leaves the reader without the meaning of the result.

Audit the claim before revealing the repairExplore

Reveal the issue and a better conclusion

02 / Insufficient evidence does not establish equality

The sample may be too small to detect a real difference.

01 · Repair an overclaim

A test of H₀:p=0.7 against p<0.7 does not reject at 5%. Repair “the success rate is definitely still 70%”.

Hint

State what the data do not establish.

Worked solution

There is insufficient evidence at 5%, under the stated model, that the population success probability has fallen below 0.7.

02 · Choosing words

Why is “accept H₀” potentially misleading?

Hint

A failure to show a difference is not a proof.

Worked solution

It can suggest the null has been established. “Do not reject H₀” makes the limited evidential conclusion clearer.

A non-significant sample can occur both when H₀ is true and when it is false. This chapter does not require calculating the probability of detecting every possible alternative.

03 / Rejection is not certainty or a measured effect

A probability estimate and a test answer different questions.

03 · Exact-value claim

A sample proportion is 0.82 and H₀:p=0.7 is rejected against p>0.7. Has p=0.82 been proved?

Hint

Separate estimate from evidence against the null.

Worked solution

No. The test supports an increase from 0.7 under its assumptions. The sample proportion estimates p but does not fix its exact population value.

04 · Practical importance

Does statistical significance alone establish that an improvement is large enough to matter?

Hint

What information is missing?

Worked solution

No. Consider the size of the estimated change, uncertainty and practical context. A test decision alone does not measure importance.

Watch: the same change described by opposite events

Pause, replay or seek freely. The notes explain the same idea and stay in view.

04 / Keep the conditional probability in the correct direction

The test asks about possible data under a hypothesis.

05 · Probability reversal

Repair “a p-value of 0.03 means a 3% chance H₀ is true”.

Hint

Identify what is assumed in the calculation.

Worked solution

The evidence probability is calculated assuming H₀ and the model. It is not a probability assigned to H₀ being true.

06 · Actual level

An actual significance level is 0.04. Does this mean every rejection has a 4% chance of being wrong?

Hint

The conditioning is important.

Worked solution

No. It is the probability of rejection when H₀ is true for that test and model. It does not directly give the probability H₀ is true after a rejection.

05 / Define success before stating the direction

Working and faulty probabilities move oppositely.

Let p be the probability an item is faulty and q the probability it works, with q=1−p.Worked example

H₀:p=0.08 is equivalent to H₀:q=0.92

The same baseline is expressed using complementary events.

H₁:p<0.08 is equivalent to H₁:q>0.92

Fewer faults means more working items.

For 30 items, Y=30−X

X≤1 faults is the same event as Y≥29 working items.

07 · Reverse the claim

Translate p>0.08 into a claim about q.

Hint

Subtract from one and reverse the inequality.

Worked solution

q<0.92: an increased fault probability corresponds to a decreased working probability.

08 · Equivalent evidence

Why must P(X≤1) equal P(Y≥29) when Y=30−X?

Hint

The two descriptions select identical outcomes.

Worked solution

Every sample with at most one fault has at least 29 working items, and conversely. They are the same event.

06 / A precise answer can rest on weak assumptions

Check independence, common probability and relevance to the target population.

09 · Population mismatch

A school wants a conclusion about all pupils but samples only volunteers from the maths club. What limitation should be stated?

Hint

Who had a chance to enter the sample?

Worked solution

The volunteers may not represent the whole school. A correct calculation does not remove selection bias.

10 · Dependence

Why might 30 consecutive rainy-day indicators be poorly modelled as independent trials?

Hint

Weather persists across days.

Worked solution

Successive conditions may be related, so the binomial variation need not fit. Seasonal changes can also undermine a common probability.

11 · Cause

A new teaching method is followed by a significant rise in success probability. Does the test alone prove the method caused it?

Hint

Other differences may coincide with the change.

Worked solution

No. Statistical evidence of a difference is not causal proof. The study design and possible confounding factors matter.

07 / Several opportunities to reject increase false-alarm risk

Multiplication needs explicitly independent tests.

Suppose three independent tests all have true nulls and actual level 0.04.Worked example

Probability none rejects = 0.96³ = 0.884736

Multiply the three independent nonrejection probabilities.

Probability at least one rejects = 1−0.884736 = 0.115264

This is 11.5264%, not 4% for the whole set.

Probability all three reject = 0.04³ = 0.000064

This answers a different question.

12 · Dependence limitation

Can the same products be assumed if all tests reuse strongly overlapping data?

Hint

Independence has not been established.

Worked solution

No. Shared data can make the decisions dependent; the products are not justified without independence or another suitable model.

13 · Selective reporting

Why is repeatedly trying tests and reporting only a significant one misleading?

Hint

How many opportunities were there for an extreme result?

Worked solution

It hides the repeated opportunities for rejection and can exaggerate the strength of evidence. The planned analysis and all relevant tests should be transparent.

14 · Final audit

What should accompany a contextual conclusion?

Hint

Think of the model, rule and scope.

Worked solution

The defined parameter and hypotheses, model assumptions, test level and method, relevant probability or critical-region membership, and important sampling limitations.

08 / Say what the evidence supports

Keep uncertainty visible.

State sufficient or insufficient evidence for the original alternative at the chosen level. Avoid claims of certainty, exact parameter values or causation. Define complementary events consistently and mention material model and sampling limitations. Treat repeated testing explicitly.

Section 1 of 8 · A decision needs an evidence statement