Hypothesis Testing (Normal Distribution) (AQA A Level Maths: Statistics): Flashcards

Exam code: 7357

1/12

0Still learning

Know0

  • Why is the population mean usually estimated from a sample rather than measured directly?

Cards in this collection (12)

  • Why is the population mean usually estimated from a sample rather than measured directly?

    Because taking a census is often impractical or impossible: the population may simply be too large to collect data from every member.

    Or collecting the data may destroy or compromise what is being measured, so that testing every item would leave nothing behind: the lifetime of a light bulb can only be found by running it until it fails.

    So the mean of a sample is used as an estimate of \mu, and the point of this subtopic is working out how good an estimate that is.

  • A sample of size n is taken from a population X \sim \text{N} \left(\mu , \sigma^{2}\right). Complete the distribution of the sample mean:

    \bar{X} \sim \text{N} \left(\_\_\_\_\_\_ , \_\_\_\_\_\_\right)

    The completed distribution is:

    \bar{X} \sim \text{N} \left(\mu , \frac{\sigma^{2}}{n}\right)

    The mean is unchanged and the variance is divided by the sample size.

    \bar{X} is the distribution of all the values the sample mean could take, over every sample of size n that might have been drawn. It is a normal distribution in its own right, narrower than the population it came from.

  • Why does taking a sample mean leave the mean unchanged but reduce the variance?

    The mean is unchanged because a sample is as likely to come out above \mu as below it, so sample means centre on the same value the population does.

    The variance is reduced because averaging dilutes extreme values. One unusually large value in a sample of 20 is divided by 20 before it reaches the sample mean, so it moves the mean far less than it would move a single observation.

    Sample means are therefore clustered more tightly around \mu than individual values are, which is exactly what makes a sample mean a useful estimate.

  • What is the standard deviation of the distribution of the sample mean, and how does it depend on the sample size?

    It is

    \frac{\sigma}{\sqrt{n}}

    which comes from square rooting the variance \frac{\sigma^{2}}{n}.

    It is inversely proportional to the square root of the sample size, not to the sample size itself. So the larger the sample, the narrower the distribution of the sample means and the more reliable the estimate, but the improvement slows down as n grows.

  • True or False?

    Doubling the sample size halves the standard deviation of the sample mean.

    False.

    The standard deviation is \frac{\sigma}{\sqrt{n}}, so doubling n divides it by \sqrt{2}, which is about 1.41.

    To halve it you have to quadruple the sample size, since \sqrt{4} = 2.

    That is the practical cost of precision: each further halving of the spread needs four times as much data as the last one.

  • A random sample of 10 observations is taken from X \sim \text{N} \left(30 , 25\right). What is the distribution of the sample mean?

    Divide the variance by the sample size:

    \bar{X} \sim \text{N} \left(30 , \frac{25}{10}\right) = \text{N} \left(30 , 2 . 5\right)

    The 25 in the original bracket is the variance, so the population standard deviation is 5 and the standard deviation of the sample mean is \frac{5}{\sqrt{10}} = 1 . 58 to 3 significant figures.

    Check which one a question has given you before dividing: a distribution written as \text{N} \left(30 , 5^{2}\right) means the same thing, but \text{N} \left(30 , 5\right) would not.

  • A test is carried out on whether a mean has decreased from 204. Complete the hypotheses:

    \text{H}_{0} : \mu = \_\_\_\_\_\_ \text{ and } \text{H}_{1} : \mu \_\_\_\_\_\_ 204

    The completed hypotheses are:

    \text{H}_{0} : \mu = 204 \text{ and } \text{H}_{1} : \mu < 204

    Both are written in terms of \mu, the population mean, because that is the quantity whose value is in question.

    Define in words what \mu measures before writing them down, since a mean has no meaning until you say a mean of what.

  • What is the test statistic in a normal hypothesis test?

    The test statistic is the sample mean \bar{x}, the single number that the sample produces.

    It is that value which gets compared with a critical value, or turned into a probability, and the whole test is an argument about how surprising a sample mean like it would be.

  • You are testing a mean using a sample of 12. Which distribution do you work with, and why not the population one?

    Work with the distribution of the sample mean, not with the distribution of a single observation, because the test statistic is a sample mean and the test asks how likely a value like it would be.

    Reaching for the population distribution instead is the commonest error here, and it does not merely shift the answer slightly: it uses a spread that is far too wide, so real evidence gets reported as unremarkable.

    Assuming \text{H}_{0} is true is what fixes the centre of that distribution, and everything after it is an ordinary normal probability calculation.

  • How do you find the critical value for a normal hypothesis test?

    Apply the inverse normal function to the distribution of the sample mean, taken with the value that \text{H}_{0} assumes for the mean.

    Give it the area in the tail being tested and it returns the sample mean at the boundary, so any observed mean beyond that value lies in the critical region.

    A two-tailed test needs this done twice, once at each end, because the two critical values are different numbers.

  • True or False?

    A hypothesis test for the mean of a normal distribution can be carried out without knowing the population variance.

    False.

    The probability the test rests on comes from the distribution of the sample mean, and that distribution needs a spread as well as a centre, which only the population variance can supply.

    That is why such a question either states the variance or says it may be assumed unchanged: without it there is nothing to calculate a probability from.

  • Puzzle times are \text{N} \left(204 , 81\right) minutes. Twelve puzzles average 201 minutes. Is there evidence of a decrease at the 5% level?

    No. With \text{H}_{0} : \mu = 204 and \text{H}_{1} : \mu < 204, and assuming \text{H}_{0}, the mean of 12 puzzles follows \bar{X} \sim \text{N} \left(204 , \frac{9^{2}}{12}\right).

    The probability of a mean at least as low as the one observed is \text{P} \left(\bar{X} \le 201\right) = 0 . 1241, which is greater than 5%.

    So there is insufficient evidence at the 5% level to say her time has decreased: a mean 3 minutes below the claim is not unusual when sample means of 12 have a standard deviation of about 2.6 minutes.

Sign up to unlock flashcards

or