Population, Sampling & Collecting Data (AQA GCSE Statistics: Foundation): Flashcards

Exam code: 8382

1/25

0Still learning

Know0

  • What is the difference between a population, a sampling frame and a sample?

Cards in this collection (25)

  • What is the difference between a population, a sampling frame and a sample?

    The population is the whole set of things you are interested in.

    A sampling frame is a list of all the members of that population, which not every population has.

    A sample is a subset of the population that data is actually collected from.

  • True or False?

    In a study of dentists in the UK, the population is everyone living in the UK.

    False.

    The study is only interested in dentists, so the population is all the dentists in the UK and nobody else.

    The word 'population' means different things in different contexts, so decide what it is from what the study is about rather than assuming it means everybody.

  • What is a census, and why is one rarely carried out?

    A census collects data from every member of the population, as the UK government does every ten years.

    It is time consuming and expensive, and it can use up or destroy the whole population: a company cannot test every firework it makes and still have any left to sell.

  • What is opportunity (convenience) sampling, and what is its main drawback?

    In opportunity (convenience) sampling the sample is made up of whichever members of the population happen to be available and fit the study, such as the first 50 people to walk past you.

    It is quick and needs no list of the population, but the sample is unlikely to be representative, so the results can be badly biased.

  • Define a simple random sample.

    In a simple random sample, every member of the population has an equal probability of being selected.

    This makes the selection fair and unbiased, so the sample is likely to represent the population well.

    It does need a numbered list of the whole population, so it cannot be used where one is impossible to make, such as for the fish in a lake.

  • When choosing a simple random sample with random numbers, what do you do with a number that is out of range, or one that has already come up?

    Ignore it, and keep generating numbers until you have enough usable ones.

    A number larger than any on your list matches nobody, and a repeat would put the same member into the sample twice, so both are skipped.

  • How would you choose a systematic sample of 40 students from a list of 480?

    Divide the size of the population by the size of the sample to get the interval, \frac{480}{40} = 12.

    Choose one student at random as the starting point, then take every 12th student along the list after that, wrapping back round to the start if you reach the end.

  • Why can systematic sampling go wrong if the list has a repeating pattern in it?

    Because the interval you count in can line up with the pattern, so the same kind of member gets picked every time.

    In a list of five-person teams with the captain written first, taking every 5th name gives you either all captains or no captains at all.

  • True or False?

    In a stratified sample, the members taken from each stratum are chosen at random.

    True.

    Once you know how many to take from a stratum, those members are picked by taking a simple random sample from within that stratum.

    That is what keeps the method unbiased, since the choice within each group is not left to anyone's judgement.

  • Why is a stratified sample more likely than a simple random sample to reflect the structure of a population?

    Because every stratum is guaranteed to appear in the sample, in the same proportion as it appears in the population.

    A simple random sample could, just by chance, miss out a whole group altogether.

  • Complete the formula for the number of members to take from a stratum.

    \text{number from stratum} = \frac{\_\_\_\_\_\_}{\text{size of population}} \times \_\_\_\_\_\_

    The completed formula is:

    \text{number from stratum} = \frac{\text{size of stratum}}{\text{size of population}} \times \text{size of sample}

  • In stratified sampling, why must every member of the population belong to exactly one stratum?

    Because the strata have to account for the whole population once each, so that their proportions add up to 1.

    If the groups overlapped, some members could be counted twice and the numbers taken from the strata would not add up to the sample size you wanted.

  • What is the difference between a laboratory, a field and a natural experiment?

    A laboratory experiment is run in an environment that the researcher controls.

    A field experiment is run in the subject's usual environment, but with the researcher still controlling the situation and some of the variables.

    A natural experiment is run in the subject's usual environment with the researcher controlling nothing at all.

  • A laboratory experiment controls extraneous variables better than any other kind. What is the cost of that control?

    Test subjects may not behave naturally in an environment that is not their own, so what the experiment measures may not be what they would really do.

    A field or natural experiment is more likely to show usual behaviour, but it gives up some or all of that control.

  • Complete the two statements about collected data.

    Data is \_\_\_\_\_\_ if repeated measurements give similar results, and data is \_\_\_\_\_\_ if it measures what it was intended to measure.

    The completed statements are:

    Data is reliable if repeated measurements give similar results, and data is valid if it measures what it was intended to measure.

  • What is the difference between an open question and a closed question on a questionnaire?

    An open question lets the respondent answer anything at all, so every answer can be different.

    A closed question offers a fixed set of answers to choose from, so you can count how many people chose each one and summarise the data easily.

  • What is wrong with asking 'How delighted are you with our excellent new product?'

    It is a leading question: the wording pushes the respondent towards a positive answer before they have given one.

    The responses collected will be biased, so the question should be worded neutrally, for example 'What do you think of our new product?'

  • A questionnaire asks 'How many times a day do you use our app?' and offers the options 1 time, 2 times, and 3 to 5 times. What is wrong with these options?

    They do not cover every possibility, so a respondent who never uses the app, or who uses it more than 5 times a day, has nothing to tick.

    The options on a closed question must cover every answer a respondent might give, and they must not overlap, so 'never' and 'more than 5 times' are both needed here.

  • Give one advantage of using interviews and one advantage of using anonymous questionnaires.

    In an interview the response rate is higher, and the interviewer can explain a question while the respondent can explain an answer.

    An anonymous questionnaire can be sent to a much larger sample, and people are more likely to answer sensitive questions honestly because there is no interviewer bias.

  • What is a pilot survey, and why is one run before the real survey?

    A pilot survey gives the questionnaire to a small sample of people first.

    It reveals problems with the design, such as questions people misunderstand or refuse to answer, so they can be fixed before the questionnaire is used for real.

    The same idea is used before an experiment, where it is called a pre-test.

  • True or False?

    Any value that is much larger or much smaller than the rest should be removed before the data is analysed.

    False.

    An anomalous value may be a genuine unusual value that has to be kept, such as one very high salary in a list of company salaries that belongs to the chief executive.

    Only values judged to be mistakes should be removed, and any removal has to be justified.

  • Why do units and symbols such as 'kg' or '£' need removing from data before it goes into a spreadsheet?

    Because a spreadsheet or statistics package reads '12 kg' as text rather than as a number, so it cannot calculate with it.

    Cleaning the data first puts every value into a consistent numerical format, and all the final calculations should then be done on the cleaned data set.

  • Why can a low response rate make a sample unrepresentative, even when the sample was chosen fairly?

    Because the members who fail to respond may have something in common, so a whole kind of member ends up under-represented.

    If questionnaires are sent to every business on a list, the ones that have closed down cannot reply, so struggling businesses go missing from the results.

  • Define an extraneous variable.

    An extraneous variable is a variable you are not interested in, but which can still affect the results of your experiment.

    Extraneous variables should be identified before the experiment starts and then controlled, so that their effect on the data is removed or at least reduced.

  • In a study of whether height affects 100 metre sprint times, half the students are timed before school and half after school. Why is this a problem, and how could it be fixed?

    Students timed at the end of the day may run more slowly because they are tired, so a difference in times could be caused by tiredness rather than by height.

    The time of day should be controlled by timing every student at the same time of day.

Sign up to unlock flashcards

or