Exam code: H240
1/120Still learning
Know0
Why is the population mean usually estimated from a sample rather than measured directly?
Because taking a census is often impractical or impossible: the population may simply be too large to collect data from every member.
Or collecting the data may destroy or compromise what is being measured, so that testing every item would leave nothing behind: the lifetime of a light bulb can only be found by running it until it fails.
So the mean of a sample is used as an estimate of , and the point of this subtopic is working out how good an estimate that is.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
A sample of size is taken from a population
. Complete the distribution of the sample mean:
The completed distribution is:
The mean is unchanged and the variance is divided by the sample size.
is the distribution of all the values the sample mean could take, over every sample of size
that might have been drawn. It is a normal distribution in its own right, narrower than the population it came from.
Why does taking a sample mean leave the mean unchanged but reduce the variance?
The mean is unchanged because a sample is as likely to come out above as below it, so sample means centre on the same value the population does.
The variance is reduced because averaging dilutes extreme values. One unusually large value in a sample of 20 is divided by 20 before it reaches the sample mean, so it moves the mean far less than it would move a single observation.
Sample means are therefore clustered more tightly around than individual values are, which is exactly what makes a sample mean a useful estimate.
Was this flashcard helpful?
Why is the population mean usually estimated from a sample rather than measured directly?
Because taking a census is often impractical or impossible: the population may simply be too large to collect data from every member.
Or collecting the data may destroy or compromise what is being measured, so that testing every item would leave nothing behind: the lifetime of a light bulb can only be found by running it until it fails.
So the mean of a sample is used as an estimate of , and the point of this subtopic is working out how good an estimate that is.
A sample of size is taken from a population
. Complete the distribution of the sample mean:
The completed distribution is:
The mean is unchanged and the variance is divided by the sample size.
is the distribution of all the values the sample mean could take, over every sample of size
that might have been drawn. It is a normal distribution in its own right, narrower than the population it came from.
Why does taking a sample mean leave the mean unchanged but reduce the variance?
The mean is unchanged because a sample is as likely to come out above as below it, so sample means centre on the same value the population does.
The variance is reduced because averaging dilutes extreme values. One unusually large value in a sample of 20 is divided by 20 before it reaches the sample mean, so it moves the mean far less than it would move a single observation.
Sample means are therefore clustered more tightly around than individual values are, which is exactly what makes a sample mean a useful estimate.
What is the standard deviation of the distribution of the sample mean, and how does it depend on the sample size?
It is
which comes from square rooting the variance .
It is inversely proportional to the square root of the sample size, not to the sample size itself. So the larger the sample, the narrower the distribution of the sample means and the more reliable the estimate, but the improvement slows down as grows.
True or False?
Doubling the sample size halves the standard deviation of the sample mean.
False.
The standard deviation is , so doubling
divides it by
, which is about 1.41.
To halve it you have to quadruple the sample size, since .
That is the practical cost of precision: each further halving of the spread needs four times as much data as the last one.
A random sample of 10 observations is taken from . What is the distribution of the sample mean?
Divide the variance by the sample size:
The 25 in the original bracket is the variance, so the population standard deviation is 5 and the standard deviation of the sample mean is to 3 significant figures.
Check which one a question has given you before dividing: a distribution written as means the same thing, but
would not.
In a hypothesis test on the mean of a normal distribution, what is the test statistic?
The sample mean rather than a single observation from the population.
Everything after that is judged against the distribution the sample mean follows, which has the same centre as the population but a smaller spread.
A machine is set to fill bottles with a mean of 500 ml, and a test is carried out on whether that mean has changed. Complete the hypotheses:
The completed hypotheses are:
Both are written in terms of the population mean even though the test is carried out on a sample mean, because the claim being tested is about the population.
Define in words what stands for before writing them down, unless the question has already done it.
True or False?
The critical value for a test on a population mean is found from the population's own distribution.
False.
The test is on the sample mean, so the critical value comes from the distribution of the sample mean, which has the same centre but a smaller spread.
Using the population's standard deviation instead makes the critical region far too wide, and a genuinely significant result can then be missed.
What has to be known about the population before a hypothesis test on its mean can be carried out?
Its variance, and that it is normally distributed.
The variance is needed because the spread of the sample mean's distribution is worked out from it, and that spread is what fixes the critical values.
Estimating the population variance from the sample is outside this course, so a question always gives the variance or says that it may be assumed unchanged.
How do you carry out a two-tailed test on the mean of a normal distribution?
Find two critical values, one in each tail of the sample mean's distribution, splitting the significance level between them.
The observed sample mean is compared with both, and the null hypothesis is rejected if it lies beyond either one.
Because a normal distribution is symmetrical, the two critical values sit the same distance either side of the assumed mean.
In a test on a normal mean, which probability is the p-value?
The probability that the sample mean would be at least as far from the assumed mean as the one actually observed, worked out from the sample mean's distribution on the assumption that the null hypothesis is true.
It comes straight from the normal cumulative function, so this route needs no critical value at all, which is what makes it quick when there is a single observed mean to judge.
By signing up you agree to our Terms and Privacy Policy