Working with Data (Cambridge (CIE) A Level Maths: Probability & Statistics 1): Flashcards

Exam code: 9709

1/14

0Still learning

Know0

  • What two things must you always comment on when describing or comparing a data set?

Cards in this collection (14)

  • What two things must you always comment on when describing or comparing a data set?

    A measure of location, such as the mean or the median, and a measure of spread, such as the range, the interquartile range or the standard deviation.

    One without the other leaves half the picture, because two data sets can share the same mean and still be spread out completely differently.

  • Complete the pairings used when describing a data set:

    The mean is paired with the \_\_\_\_\_\_, and the median is paired with the \_\_\_\_\_\_.

    The completed pairings are:

    The mean is paired with the standard deviation, and the median is paired with the interquartile range.

    Each pair is matched because its two members are built the same way: the mean and the standard deviation both use every value, while the median and the quartiles both depend only on position in the order.

  • When should you use the median and interquartile range rather than the mean and standard deviation?

    When the data contains outliers or extreme values.

    Extreme values distort the mean and the standard deviation, because both are calculated from every value, but they barely move the median and the quartiles, which depend only on position.

  • A new value is added to a data set. Why can you not say in advance whether the median will change?

    Because the median is whichever value sits in the middle position, so whether it moves depends entirely on where the new value falls in the order.

    Each case has to be checked individually, unlike the mean, which changes unless the new value happens to equal it exactly.

  • True or False?

    Adding one extremely large value to a data set changes the median by as much as it changes the mean.

    False.

    The mean is calculated from every value, so an extreme one drags it a long way, while the median simply shifts by one place in the order and so moves no more than any other new value would move it.

    This is the whole reason the median is preferred for data containing extreme values.

  • Two groups are timed completing a puzzle, and one group has the lower median. What does that say about them?

    That group was quicker on average, since a lower time means a faster finish.

    Whether a lower value is better always depends on what is being measured: it is good for a time taken, and bad for a score on a test, so a comparison has to be written in the context of the data.

  • What does a larger interquartile range tell you about a data set?

    That its middle 50% is more spread out, so the data is less consistent.

    A smaller interquartile range means the values are clustered more tightly together, and the same reading applies to any measure of spread: the larger it is, the more varied the data.

  • Define skewness.

    Skewness describes the way the data in a non-symmetrical distribution leans, with one side stretched further out than the other.

    Every distribution is either symmetrical or skewed, and the skew is what a symmetrical distribution has none of.

  • Which way is a distribution skewed if its tail stretches out to the right?

    It has positive skew, because the skew is named after the side the tail is on rather than the side the peak is on.

    A tail stretching out to the left is negative skew, so the name follows the direction along the number line in which the data has been drawn out.

  • An estimate of a distribution's mean turns out to be larger than its median. What feature of the distribution accounts for the two being different?

    The distribution is not symmetrical, which is to say it is skewed.

    In a perfectly symmetrical distribution the mean and the median fall in the same place, so any difference between them is telling you that one side of the distribution is stretched out further than the other.

  • On a box plot, how do the gaps between the quartiles tell you which way the data is skewed?

    Compare Q_{2} - Q_{1} with Q_{3} - Q_{2}: the larger gap lies on the side the data is stretched out towards.

    So if Q_{2} - Q_{1} < Q_{3} - Q_{2} the median sits closer to the lower quartile and the skew is positive, and if Q_{2} - Q_{1} > Q_{3} - Q_{2} the median sits closer to the upper quartile and the skew is negative.

  • Complete the order of the three averages in a positively skewed distribution:

    \text{mode} < \_\_\_\_\_\_ < \_\_\_\_\_\_

    The completed order is:

    \text{mode} < \text{median} < \text{mean}

    The mode stays at the peak, while the mean is pulled furthest along the tail because it is calculated from every value, and the median finishes between them.

    For a negatively skewed distribution the whole order reverses, giving \text{mean} < \text{median} < \text{mode}.

  • True or False?

    A distribution with two separate peaks can still be perfectly symmetrical, with no skew at all.

    True.

    Symmetry is a question of whether the two halves of the distribution are mirror images of each other, and has nothing to do with how many peaks it has.

    Two equal peaks placed either side of the centre give a symmetrical distribution, and so does a flat uniform distribution with no peak at all.

  • A data set has Q_{1} = 12, Q_{2} = 14 and Q_{3} = 22. Which way is it skewed, and how can you tell?

    It is positively skewed.

    The gap below the median is 14 - 12 = 2 and the gap above it is 22 - 14 = 8, so the median sits far closer to the lower quartile and the data is stretched out towards the higher values.

Sign up to unlock flashcards

or