Population, Sampling & Collecting Data (Edexcel GCSE Statistics: Higher): Flashcards

Exam code: 1ST0

1/28

0Still learning

Know0

  • Complete the sentence about the two ways of collecting data.

    A \_\_\_\_\_\_ collects data from every member of the population, whereas a \_\_\_\_\_\_ collects data from only a subset of it.

Cards in this collection (28)

  • Complete the sentence about the two ways of collecting data.

    A \_\_\_\_\_\_ collects data from every member of the population, whereas a \_\_\_\_\_\_ collects data from only a subset of it.

    The completed sentence is:

    A census collects data from every member of the population, whereas a sample collects data from only a subset of it.

  • True or False?

    In a statistical investigation, 'the population' always means all the people living in a country.

    False.

    The population is the whole set of things being studied, and what that is depends entirely on the investigation.

    It could be all the dentists in the UK, all the French bulldogs in the world, or every item produced in a factory.

  • Define sampling frame.

    A sampling frame is a list of all the members of the population, such as a company's list of its employees.

    Not every population has one that can easily be obtained, which rules out some sampling methods.

  • A firework manufacturer wants to know what proportion of its fireworks work properly. Why can it not simply test them all?

    Because testing a firework uses it up, so testing every one would destroy the entire stock.

    Any test that damages or consumes the items being tested has to be carried out on a sample.

  • What has to be true of a simple random sample, and what do you need in order to take one?

    Every member of the population must have an equal probability of being selected, which is what makes it the best method for avoiding bias.

    You need a list of the whole population so that every member can be numbered, and the numbers are then chosen at random.

  • You are picking a random sample from a numbered list of 80 people, and your random numbers include 93 and a repeat of 27. What should you do with each?

    Ignore both of them: 93 does not match anyone on the list, and the second 27 would select the same person twice.

    Keep generating random numbers until you have enough different ones that do match items on the list.

  • Describe how to take a 10% systematic sample of the 60 employees in a factory.

    Number all 60 employees in a list. A 10% sample is 6 people, so you take every 10th name.

    Choose the starting point at random, then take every 10th name after it, wrapping back to the start of the list if you reach the end.

  • Why can systematic sampling go wrong when the list is arranged in repeating groups?

    Because the sampling interval can line up with the pattern in the list, so the same kind of member is picked every time.

    If names are listed in five-person teams with the captain first, taking every 5th name gives you either all captains or no captains.

  • Quota and stratified sampling both split the population into groups. What is the crucial difference between them?

    In stratified sampling the members taken from each group are chosen at random, which is what makes it a random method.

    In quota sampling they are not: you simply keep asking people until each quota is filled, so anyone who refuses is replaced and the sample can be biased.

  • A school has 636 students, 36 teachers and 48 non-teaching staff. How do you work out how many students belong in a stratified sample of 60?

    The population is 636 + 36 + 48 = 720 people, so you divide the size of the stratum by the size of the population and multiply by the size of the sample.

    \frac{636}{720} \times 60 = 53 \text{ students}

  • What must be true of the groups used as strata in a stratified sample?

    Every member of the population must belong to exactly one group.

    The groups cannot overlap, and none of the population can be left out of them, or the sample will not reflect the structure of the population.

  • Opportunity sampling and judgement sampling are both quick and need no sampling frame. What makes them risky?

    Neither of them selects its members at random, so the sample is unlikely to represent the population and the results can be biased.

    With judgement sampling the person choosing may introduce bias without meaning to, which is why it is rarely a preferred method.

  • What is cluster sampling, and why is it cheaper than a simple random sample of the same size?

    The population is divided into sensible clusters, such as schools, and a number of whole clusters are chosen at random to form the sample.

    It is cheaper because you only collect data from a few clusters, instead of reaching a few people in every school in the country.

  • A researcher studies dogs' obedience by visiting each dog at its own home, but choosing which commands each dog is asked to perform. Which of the three kinds of experiment is this?

    It is a field experiment.

    It takes place in the subject's usual environment, which rules out a laboratory experiment, but the researcher controls part of what happens, which rules out a natural experiment.

  • True or False?

    A laboratory experiment has to be carried out in an actual laboratory.

    False.

    A laboratory experiment simply means one carried out in an environment the researcher controls.

    Studying people's sleep in a specially arranged room, with the lighting, temperature and bedding all set by the researchers, counts as one.

  • Complete the two statements about collected data.

    Data is \_\_\_\_\_\_ when repeated measurements give similar results, and it is \_\_\_\_\_\_ when it measures what it was intended to measure.

    The completed statements are:

    Data is reliable when repeated measurements give similar results, and it is valid when it measures what it was intended to measure.

  • A model represents a population in which 23% of people carry a genetic marker, using random two-digit numbers from 00 to 99. How is the marker built into the model?

    The 100 possible two-digit numbers stand for 100 people, so the 23 numbers from 00 to 22 represent someone who carries the marker.

    The remaining numbers, 23 to 99, represent someone who does not, which reproduces the 23% rate in the model.

  • What is the difference between an open and a closed question, and which is easier to analyse?

    A closed question offers a fixed set of answers to choose from, while an open question lets the respondent say anything at all.

    Closed questions are far easier to summarise and analyse, because you can count how many people chose each response.

  • Define interviewer bias.

    Interviewer bias is when the opinions or expectations of the interviewer affect the answers a respondent gives.

    It can come through the interviewer's tone or manner as well as their wording, and it is one reason an anonymous questionnaire may get more honest answers.

  • Why is 'How delighted are you with our excellent new product?' a poor questionnaire question?

    It is a leading question: describing the product as 'excellent' pushes the respondent towards a positive answer.

    The responses collected are then biased, so the data does not show what people really think.

  • What is a pilot survey, and what is it for?

    A pilot survey is a trial run of a questionnaire, given to a smaller sample before the real survey.

    It shows up problems with the design, such as questions people misunderstand or refuse to answer, so they can be fixed before the questionnaire is used for real.

    The same idea applied to an experiment is called a pre-test.

  • A survey asks: 'Have you ever driven while using a handheld phone? Flip a coin first: if you get heads, answer Yes; if you get tails, answer honestly.' Why ask it in this odd way?

    It is the random response method, used for questions people would not otherwise answer honestly.

    Because nobody can tell whether a Yes came from the coin or from the truth, a respondent can answer honestly without revealing anything about themselves.

  • In a random response survey, heads means answer Yes and tails means answer honestly. Of 1200 replies, 700 are Yes. Estimate the percentage who would genuinely have answered Yes.

    About 600 of the 1200 said Yes only because they got heads, leaving 700 - 600 = 100 genuine Yes answers.

    Those 100 sit among the 500 + 100 = 600 genuine replies, and \frac{100}{600} is about 17%.

  • True or False?

    An anomalous value in a data set should sometimes be kept rather than removed.

    True.

    An anomalous value is only removed if you decide it is a mistake; if it is a genuine reading it stays in.

    One very high salary in a list of company salaries may really belong to the chief executive, and removing it would distort the data.

  • A questionnaire is sent to every business on a government list, and some do not reply because they have gone out of business. Why does that matter for the results?

    Because struggling or unsuccessful businesses are then under-represented among the replies.

    The non-responses are not random, so the results end up biased towards businesses that are doing well.

  • Eight towels are put in each hotel room. In the record of how many went missing from each room, one entry is 15. Why must that be an error, and what must you do if you remove it?

    Only 8 towels are put in each room, so at most 8 can go missing and 15 is impossible.

    Any value taken out of a data set has to be justified, and all the final calculations are then done on the cleaned data.

  • Define extraneous variable.

    An extraneous variable is one you are not interested in, but which can still affect the results of your experiment.

    They should be identified before the experiment starts and then controlled, so their effect on the data is eliminated or at least minimised.

  • What is a matched pairs approach, and what problem can it run into?

    Each person in the test group is paired with someone in the control group who is as similar as possible, so the only difference between them is the variable being studied.

    The difficulty is finding enough matched pairs to give a large enough sample.

Sign up to unlock flashcards

or