Exam code: YMA01
1/210Still learning
Know0
Fill in the two missing sampling terms:
A is the whole set of things you are interested in.
A is a subset of it that you actually collect data from.
The completed statements are:
A population is the whole set of things you are interested in.
A sample is a subset of it that you actually collect data from.
If a vet wants to know how long a typical French bulldog sleeps, the population is every French bulldog there is, and a sample might be the bulldogs from a few cities.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
What is a sampling unit, and what is a sampling frame?
A sampling unit is one individual member of the population, such as a single French bulldog.
A sampling frame is a list of all the members of the population, such as a company's list of its employees' names.
What is the difference between a population parameter and a sample statistic?
A population parameter is a numerical value describing the whole population, such as the mean height of all 16-year-olds in the UK, and it is usually unknown.
A sample statistic is a value computed from the data in a sample, such as the mean height of 200 of them.
Statistics exist in order to estimate parameters, which is the whole reason for sampling at all.
Was this flashcard helpful?
Fill in the two missing sampling terms:
A is the whole set of things you are interested in.
A is a subset of it that you actually collect data from.
The completed statements are:
A population is the whole set of things you are interested in.
A sample is a subset of it that you actually collect data from.
If a vet wants to know how long a typical French bulldog sleeps, the population is every French bulldog there is, and a sample might be the bulldogs from a few cities.
What is a sampling unit, and what is a sampling frame?
A sampling unit is one individual member of the population, such as a single French bulldog.
A sampling frame is a list of all the members of the population, such as a company's list of its employees' names.
What is the difference between a population parameter and a sample statistic?
A population parameter is a numerical value describing the whole population, such as the mean height of all 16-year-olds in the UK, and it is usually unknown.
A sample statistic is a value computed from the data in a sample, such as the mean height of 200 of them.
Statistics exist in order to estimate parameters, which is the whole reason for sampling at all.
What does a census collect, and why is sampling usually preferred?
A census collects data about every member of the population, which is what makes its results fully accurate.
It is slow and expensive, and where the members are consumables it destroys the population, as testing every firework a company makes would.
Sampling is quicker, cheaper and leaves less data to analyse, at the cost of possibly not representing the population and of introducing bias.
Why is not a statistic, when
is?
Because a statistic has to be calculable from the sample data alone, and is an unknown population parameter.
is the mean of the sample, so the second expression can actually be worked out once the sample is in hand, while the first cannot.
That is the test to apply: if an unknown population value appears anywhere in it, it is not a statistic.
Define the sampling distribution of a statistic.
The sampling distribution of a statistic gives all the possible values that the statistic can take, together with the probability of each one.
A statistic is itself a random variable, because its value depends on which sample happens to be drawn, so it has a distribution like any other random variable.
How do you find the sampling distribution of a statistic such as the sample mean?
List every possible sample, work out the value of the statistic for each one, and work out the probability of each sample being drawn.
Then build the distribution table by adding together the probabilities of all the samples that give the same value of the statistic.
Grouping samples with the same combination of members as you list them makes that last step much easier.
True or False?
When listing the possible samples, AAB and ABA count as the same sample.
False.
They are listed separately, because they are different samples even though they contain the same members.
That is why a sample of 3 taken from a population with 2 types of member has possible samples rather than fewer, and grouping them by combination afterwards is a shortcut rather than a reason to leave any out.
The population is described as large. What does that let you assume when working out each sample's probability?
That the sampling can be treated as though it were with replacement.
The probability of drawing each type of member then stays constant however many have already been taken, so the selections are independent and their probabilities simply multiply.
Without that assumption every draw would change the make-up of what was left, and the probabilities would have to be recalculated at each step.
Define the null hypothesis and the alternative hypothesis.
The null hypothesis states the value of the population parameter on the assumption that nothing has changed.
The alternative hypothesis states how that parameter might have changed instead.
A test assumes is true the whole way through, and ends by either rejecting it or failing to reject it.
A hypothesis test concerns a population parameter . Fill in the two missing relation symbols:
testing for a change,
The completed hypotheses are:
testing for a change,
The null hypothesis always takes the equals form, whatever kind of test is being run.
A two-tailed test uses because it claims only that the parameter has moved, while a one-tailed test uses
or
and claims a direction as well.
Define the significance level of a hypothesis test.
The significance level is the probability threshold the test is judged against: it sets how unlikely the observed result has to be, assuming is true, before
is rejected.
It is usually 1%, 5% or 10%, although a question may set some other value.
A smaller significance level demands stronger evidence before the null hypothesis is given up.
Why must the significance level be fixed before the test is carried out?
Because choosing it afterwards would let you pick whichever threshold gives the answer you wanted.
The significance level is a statement about how much evidence you are demanding, so it has to be settled independently of what the data turns out to say.
The same observed result can be significant at 5% and not significant at 1%, which is exactly why the level cannot be chosen once the result is known.
Define the critical region and the critical value.
The critical region is the range of values the test statistic could take that would lead to being rejected.
The critical value is the boundary of that region, the least extreme value that still causes rejection.
Where the region sits is fixed by the significance level, and the test statistic itself is also called the observed value.
True or False?
The critical value lies just outside the critical region.
False.
The critical value is inside the region: it is the least extreme value that still leads to being rejected.
So if the critical region is then 3 is the critical value, an observed value of 3 rejects
, and an observed value of 4 does not.
For a discrete distribution, why is the actual significance level usually below the stated one?
Because the critical value has to be one of the whole values the variable can actually take, so the region cannot be trimmed to land on the stated level exactly.
The probability of falling inside it, assuming is true, therefore comes out at or below the level that was asked for.
That probability is the actual significance level, and it is the probability of rejecting a null hypothesis that was in fact true.
In a two-tailed test, what is a one-tail probability compared with?
Half the significance level, since the level has to be split between the two tails.
At the 5% level each tail carries 2.5%, so a one-tail probability of 0.02 is significant while one of 0.03 is not.
Testing a single tail against the full 5% would reject twice as readily as the test is meant to.
A test is carried out without finding a critical region. What probability do you work out, and what do you compare it with?
The p-value: the probability, assuming is true, of getting a value at least as extreme as the one observed.
Testing for an increase makes the extreme values the large ones, so you want , while testing for a decrease makes them the small ones and you want
.
Reject if that probability is less than the significance level.
How should the conclusion of a hypothesis test be worded?
In the context of the question, reusing its own wording, and never as a definite statement.
Rejecting gives sufficient evidence to suggest that the alternative hypothesis is true at that significance level.
Failing to reject it gives insufficient evidence to suggest the alternative hypothesis, which is not at all the same as having shown the null hypothesis to be true.
True or False?
A hypothesis test carried out perfectly correctly can still reach the wrong conclusion.
True.
The test weighs how probable one particular sample was, and an unusual sample can point the wrong way through no fault of the method.
That is why a conclusion is written as evidence at a stated significance level rather than as proof, and why a different sample might well have given a different outcome.
A two-tailed test rejects . What may the conclusion claim, and what may it not?
It may claim there is evidence that the population parameter has changed.
It may not say in which direction, so a conclusion that the parameter has increased or decreased is wrong even when the observed value sat in the upper tail.
Only a one-tailed test gives evidence about direction, and it can do so because the direction was named in before any data was seen.
By signing up you agree to our Terms and Privacy Policy