Statistical Measures (Edexcel International A Level (IAL) Maths: Statistics 1): Flashcards

Exam code: YMA01

1/34

0Still learning

Know0

  • Define a measure of location.

Cards in this collection (34)

  • Define a measure of location.

    A measure of location gives information about where data sits in the number system.

    The mean, median and mode are measures of central tendency, a particular kind of location measure describing where the centre of the data is.

    The quartiles and percentiles are measures of location too, but they describe other positions in the data rather than the centre.

  • Why must a data set be put in order of size before you can find its median, but not before you can find its mean?

    Because the median is defined by its position: it is the middle value once the data is in order, so putting the data in order is what identifies it.

    The mean is a total divided by how many values there are, and that total comes out the same whatever order the values are added in.

    With an even number of values there is no single middle one, so the median is the midpoint of the two middle values.

  • In summary-statistic notation, what do \Sigma x and \bar{x} each stand for, and how is each read aloud?

    \Sigma x is the sum of all the data values, x_{1} + x_{2} + \ldots + x_{n}, and is read as sigma x.

    \bar{x} is the mean, \frac{\Sigma x}{n}, and is read as x bar.

  • A data set contains one or two extreme values. Which average is affected most, and which average would you report instead?

    The mean is affected most, because it is worked out from every value in the set, so one extreme value drags the whole total with it.

    Report the median instead: it depends only on which value sits in the middle position, so an extreme value at one end shifts it by at most one place.

    The mode is not affected either, but it can sit nowhere near the centre of the data, and a set can have several modes or none at all.

  • The 70th percentile of a data set is 42. What does that tell you about the data?

    70% of the data lies below 42, and the remaining 30% lies above it.

    Percentiles divide the data into 100 parts, in the same way that the quartiles divide it into four: Q_{1} is the 25th percentile, the median is the 50th and Q_{3} is the 75th.

  • For a list of n raw data values written in order, \frac{n}{4} is worked out in order to locate the lower quartile. Why does what you do next depend on whether \frac{n}{4} is a whole number?

    If \frac{n}{4} is not a whole number, round it up, and Q_{1} is the value in that position.

    If \frac{n}{4} is a whole number, the quarter-way point falls exactly between two data values, so Q_{1} is the midpoint of that value and the one above it.

    The upper quartile is found the same way, using \frac{3 n}{4}.

  • Complete the formula for the interquartile range by filling in the two missing subscripts:

    \text{IQR} = Q_{\_\_\_\_\_\_} - Q_{\_\_\_\_\_\_}

    The completed formula is:

    \text{IQR} = Q_{3} - Q_{1}

    It measures the width of the stretch of the number line covered by the central half of the data, so it is a measure of spread rather than of location.

    Like the range, the largest value minus the smallest, it carries the same units as the original data.

  • True or False?

    The 25th to 75th interpercentile range is the same thing as the interquartile range.

    True.

    The 25th percentile is Q_{1} and the 75th percentile is Q_{3}, so that particular interpercentile range is Q_{3} - Q_{1}, which is exactly the interquartile range.

    Any other pair of percentiles gives something different, which is why you must always say which two you mean, as in the 10th to 90th interpercentile range.

  • True or False?

    Putting data into a frequency table loses the individual data values.

    False.

    An ungrouped frequency table keeps every value: it records how many times each individual value occurred, so exact averages, ranges and summary statistics can still be calculated from it.

    It is grouping the data that loses the values. Once a table records only that eight items fell between 240 and 260, those eight values are gone, and every statistic calculated from it is an estimate.

  • For discrete data in an ungrouped frequency table, which position gives the median, and how do you find the value there?

    The median is the \frac{n + 1}{2} th value, where n is the total frequency.

    Add the frequencies up in order until the running total first reaches that position, and the median is the data value you are on when it does.

    So, for example, with n = 60 the median is the 30.5th value, meaning halfway between the 30th and the 31st; grouped data uses a different position, because the individual values are no longer known.

  • Complete the formula for the mean of data given in a frequency table, where x is a data value and f is its frequency:

    \bar{x} = \frac{\Sigma \_\_\_\_\_\_}{\Sigma \_\_\_\_\_\_}

    The completed formula is:

    \bar{x} = \frac{\Sigma x f}{\Sigma f}

    Multiply each value by its frequency and add the results, then divide by the total frequency.

    \Sigma x f is the sum of all the data values, exactly as \Sigma x would be if they were listed out one by one, and \Sigma f is how many values there are.

  • A grouped frequency table for continuous data has the classes 10 \le x \le 19 and 20 \le x \le 29. What is wrong with them, and what do you do about it?

    There is a gap: nothing between 19 and 20 has a class to go in, and a continuous quantity can take those values.

    Close the gap before doing any calculation: for data that has been rounded, the boundaries become

    9 . 5 \leq x < 19 . 5 \textrm{ }\text{and}\textrm{ } 19 . 5 \leq x < 29 . 5

    For data that has been truncated, such as an age counted in whole years, they become 10 \leq x < 20 and 20 \leq x < 30 instead, since someone is 19 right up until their 20th birthday.

  • For grouped data you give the modal class rather than the mode. Why can you not give a mode?

    Grouping the data throws away the individual values, so there is no longer any way to tell which single value occurred most often.

    What survives is which class the data is most concentrated in, so that is what you report.

    The same limitation applies to the mean and the median from a grouped table: everything you can calculate is an estimate.

  • What do you use in place of the actual data values when estimating the mean from a grouped frequency table?

    The midpoint of each class, which stands in for every value in that class; the midpoint is the mean of the class's lower and upper boundaries, so 150 \leq h < 155 has midpoint 152.5.

    This is why the answer is only an estimate: it assumes the values in each class average out at its midpoint, which they will not do exactly.

    Watch for a class of a different width from the others, since its midpoint still has to be worked out from its own boundaries.

  • Complete the assumption that linear interpolation makes about grouped data:

    The data values are assumed to be \_\_\_\_\_\_ spread throughout each class.

    The completed assumption is:

    The data values are assumed to be evenly spread throughout each class.

    That assumption is what makes the method work. If the values are evenly spread, then a position one third of the way through a class's frequency corresponds to a data value one third of the way through that class's range, so proportion can be used.

    It is also why an interpolated median or quartile is an estimate: real data is rarely spread evenly.

  • Q_{1} is the 6.25th value, and it lies in the class 155 \le h < 160, which runs from a cumulative frequency of 3 to a cumulative frequency of 8. How does linear interpolation give Q_{1}?

    Set the fraction of the way through the cumulative frequency equal to the fraction of the way through the class boundaries:

    \frac{6.25 - 3}{8 - 3} = \frac{Q_{1} - 155}{160 - 155}

    \frac{3.25}{5} = \frac{Q_{1} - 155}{5}

    so Q_{1} = 155 + 3.25 = 158.25.

    The part that is easiest to forget is adding the lower class boundary back on at the end. The proportion gives you how far into the class the value sits, not the value itself.

  • Define the variance of a set of data.

    The variance measures how spread out a set of data is: it is the mean of the squared deviations from the mean.

    Widely spread data has a large variance; data clustered close together has a small one.

    The standard deviation is the square root of the variance and is written \sigma, which is why the variance itself is written \sigma^{2}.

  • Complete the version of the variance formula that is usually quickest to use:

    \sigma^{2} = \frac{\Sigma \_\_\_\_\_\_}{n} - \left(\_\_\_\_\_\_\right)^{2}

    The completed formula is:

    \sigma^{2} = \frac{\Sigma x^{2}}{n} - \left(\bar{x}\right)^{2}

    In words, it is the mean of the squares minus the square of the mean.

    It needs only \Sigma x^{2}, \Sigma x and n, which is why it is quicker than working out every deviation x - \bar{x} one at a time.

  • The summary statistic S_{x x} = \Sigma \left(x - \bar{x}\right)^{2} is given in the formula booklet. How do you get the variance from it?

    Divide it by n:

    \sigma^{2} = \frac{S_{x x}}{n}

    S_{x x} is the total of the squared deviations, so dividing by how many values there are turns that total into their mean, which is what the variance is.

    The formulae for the variance and the standard deviation are not in the booklet, so those two you have to know.

  • True or False?

    A set of data can have a variance of zero.

    True.

    The variance is zero exactly when every deviation x - \bar{x} is zero, which happens when every value in the set is the same.

    It can never be negative, because it is built out of squares, so zero is the smallest value a variance can take.

  • The times in a data set are measured in minutes. What are the units of the variance, and what are the units of the standard deviation?

    The variance is in minutes squared, because it is built out of squared deviations from the mean.

    The standard deviation is in minutes, the same units as the data itself, because taking the square root undoes that squaring.

    That is what makes the standard deviation the more natural of the two to quote when describing a set of data.

  • How does the variance formula change when the data is given in a grouped frequency table?

    Every value is weighted by its frequency, and the class midpoints are used in place of the data values x:

    \sigma^{2} = \frac{\Sigma f x^{2}}{\Sigma f} - \left(\frac{\Sigma f x}{\Sigma f}\right)^{2}

    The total frequency \Sigma f replaces n, because that is how many values there are.

    Since the midpoints only stand in for the real values, the answer is an estimate.

  • Define coding a set of data.

    Coding is applying the same formula to every value in a data set, to simplify the numbers you have to work with.

    It is most useful when the data consists of very large or very small numbers, so coding such as x = \frac{h - 250}{20} turns awkward values into small, convenient ones.

    The coding must be applied to every value, or the coded data no longer represents the original.

  • Data x is coded using y = a x + b. What happens to the mean?

    The mean is coded in exactly the same way as the data:

    \bar{y} = a \bar{x} + b

    The reason is that the mean is a measure of location.

    Changing every value changes where the data sits, and the mean follows it: multiply everything by a and the centre is multiplied by a; add b to everything and the centre moves up by b.

  • Data x is coded using y = a x + b. What happens to the standard deviation?

    Only the multiplier affects it:

    \sigma_{y} = \left|a\right| \sigma_{x}

    Adding or subtracting b makes no difference at all, because the standard deviation is a measure of spread: sliding every value along by the same amount leaves every gap between values exactly as it was.

    Multiplying by a does stretch those gaps, so it stretches the standard deviation with them.

  • Data x was coded using y = a x + b, and the mean of the coded data is \bar{y}. Complete the way the original mean is recovered:

    \bar{x} = \frac{\bar{y} - \_\_\_\_\_\_}{\_\_\_\_\_\_}

    The completed formula is:

    \bar{x} = \frac{\bar{y} - b}{a}

    This is just \bar{y} = a \bar{x} + b rearranged, which is worth doing rather than memorising: whatever the coding formula was, write it down with the means in it and solve for the one you want.

    Undo the operations in reverse order, so subtract before dividing.

  • In \sigma_{y} = \left|a\right| \sigma_{x}, why is the modulus of a taken?

    Because a standard deviation can never be negative: it measures how spread out data is, and a spread of less than nothing means nothing.

    Coding with a negative multiplier, such as y = - 3 x, reflects the data as well as stretching it, but reflecting does not change how spread out it is.

    So the standard deviation is multiplied by the positive value 3, not by - 3.

  • Heights h, in cm, are coded using x = \frac{h - 250}{20}, and the coded data has standard deviation 1.3102. What is the standard deviation of the heights?

    Only the division by 20 has to be reversed:

    \sigma_{h} = 20 \times 1.3102 = 26.2 \text{ cm}

    Subtracting 250 slid every height down the scale without changing how spread out they were, so it has no part in the calculation.

    Note that the answer carries the units of the original data, cm, not of the coded values, which have no units at all.

  • Define a statistical model.

    A statistical model uses mathematics to represent a real-life situation, such as the temperature in a city across a month or the sleeping times of a baby.

    It is a deliberately simplified version of that situation rather than the situation itself, which is what makes it something you can calculate with.

  • Give two advantages of using a statistical model rather than working with the real-life situation directly.

    It simplifies a complicated situation down to something that can be handled mathematically, and it can be built quickly and easily compared with observing the real situation.

    Once it exists it can also be used to make predictions, which is usually the reason for building one in the first place.

  • True or False?

    A good statistical model is one that includes as many features of the real situation as possible.

    False.

    Simplifying is the whole point: leaving features out is what turns an unmanageable situation into something that can be built quickly and calculated with.

    A model is judged on whether its predictions match what is actually observed closely enough to be useful, not on how much of the real situation it contains.

  • A model of a shop's daily sales was built using data collected in December. Why might it predict badly in June?

    A statistical model is often only applicable to the particular circumstances it was built for, because it carries the pattern of the data it was built from.

    December sales follow the run-up to Christmas, so a model fitted to them describes that period rather than a quiet summer month.

    Using a model outside the situation it was designed for is one of the main reasons predictions stop being accurate.

  • Why must real-life data be collected after a statistical model has been used to make its predictions?

    Because the model is judged by comparing the values it predicted with the values actually observed, and that comparison only means anything if the predictions came first.

    Predictions made after seeing the data could simply be fitted to it, and would say nothing about whether the model works.

    It is that comparison which decides whether the model is accepted or has to be improved.

  • The values a statistical model predicts turn out not to match the data that is then collected. What happens to the model?

    It is adjusted and improved, then used again, rather than thrown away.

    The way the data was selected and collected is examined as well, since a mismatch can say as much about the data as about the model.

    That is what makes statistical modelling a cycle: predict, collect, compare, refine, and round again.

Sign up to unlock flashcards

or