Exam code: 9MA0
1/290Still learning
Know0
Define measure of central tendency.
A measure of central tendency is an average that describes where the centre of the data is; the mean, the median and the mode are all examples.
They are all measures of location, the wider term for a statistic saying where data sits in the number system; quartiles and percentiles are measures of location too, but they are not averages.
Because there are three of them, "the average" is not specific enough in statistics: always say which one you mean.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
Which of the three averages is affected by extreme values, and which are not?
The mean is affected, because it is calculated from every value in the data set, so one very large or very small value pulls it towards itself.
The median is not, because it is a position in the ordered data rather than a calculation; an extreme value at one end moves it by at most one place.
The mode is not affected either, but it has its own problems: a data set can have more than one mode, which makes it bimodal, or no mode at all, and a mode can sit nowhere near the centre of the data.
What do the symbols and
stand for, and how is each read aloud?
is the sum of all the data values, read "sigma
", and written out in full it is
is the mean, read "
bar", and the two are linked by
Statistics written this way are called summary statistics, because they summarise a whole data set in one number.
Was this flashcard helpful?
Define measure of central tendency.
A measure of central tendency is an average that describes where the centre of the data is; the mean, the median and the mode are all examples.
They are all measures of location, the wider term for a statistic saying where data sits in the number system; quartiles and percentiles are measures of location too, but they are not averages.
Because there are three of them, "the average" is not specific enough in statistics: always say which one you mean.
Which of the three averages is affected by extreme values, and which are not?
The mean is affected, because it is calculated from every value in the data set, so one very large or very small value pulls it towards itself.
The median is not, because it is a position in the ordered data rather than a calculation; an extreme value at one end moves it by at most one place.
The mode is not affected either, but it has its own problems: a data set can have more than one mode, which makes it bimodal, or no mode at all, and a mode can sit nowhere near the centre of the data.
What do the symbols and
stand for, and how is each read aloud?
is the sum of all the data values, read "sigma
", and written out in full it is
is the mean, read "
bar", and the two are linked by
Statistics written this way are called summary statistics, because they summarise a whole data set in one number.
What do the three quartiles divide a data set into, and what does each one split?
The quartiles divide the data into four equal sections.
The lower quartile splits the lowest 25% from the highest 75%
The median is the value 50% of the way through the data
The upper quartile splits the lowest 75% from the highest 25%
Quartiles are measures of location, not averages: they say where a point in the data sits, not what is typical.
You have raw data values written in order. How do you find the position of the lower quartile?
Calculate , then look at what kind of number it is.
If is not an integer, round it up to the next integer and take that value
If is an integer, take the midpoint of that value and the one above it
So, for example, with ,
is not an integer, so
is the 2nd value; the upper quartile works the same way from
.
This is the rule for raw values in order; estimating a quartile from grouped data uses a different rule, because the individual values are no longer known.
What is the difference between the range and the interquartile range?
Both are measures of spread, telling you how spread out the data is rather than where it is, and both carry the same units as the original data.
The range is the largest value minus the smallest value, so every data point contributes to it, including any extreme values.
The interquartile range is , so it covers only the middle 50% of the data and is not affected by extreme values.
True or False?
A data set with a larger range must also have a larger interquartile range.
False.
The two measure different things and can move independently.
A single extreme value at one end makes the range enormous while leaving the interquartile range completely unchanged, because that value lies outside the middle 50% of the data.
This is exactly why the interquartile range is the more useful measure of spread when a data set contains extreme values.
What is the 70th percentile, and what is an interpercentile range?
Percentiles divide the data into 100 parts, so the 70th percentile is the value seven tenths of the way through the data: 70% of the data lies below it and 30% above it.
An interpercentile range is the difference between two given percentiles, so, for example, the 20th to 80th interpercentile range is the 80th percentile minus the 20th percentile.
The units are the same as the units of the original data, and the interquartile range is itself the 25th to 75th interpercentile range.
True or False?
Putting data into a frequency table loses the individual data values.
False.
An ungrouped frequency table keeps every value: it records how many times each individual value occurred, so exact averages, ranges and summary statistics can still be calculated from it.
It is grouping the data that loses the values. Once a table records only that eight items fell between 240 and 260, those eight values are gone, and every statistic calculated from it is an estimate.
For discrete data in an ungrouped frequency table, which position gives the median, and how do you find the value there?
The median is the th value, where
is the total frequency.
Add the frequencies up in order until the running total first reaches that position, and the median is the data value you are on when it does.
So, for example, with the median is the 30.5th value, meaning halfway between the 30th and the 31st; grouped data uses a different position, because the individual values are no longer known.
Complete the formula for the mean of data given in a frequency table, where is a data value and
is its frequency:
The completed formula is:
Multiply each value by its frequency and add the results, then divide by the total frequency.
is the sum of all the data values, exactly as
would be if they were listed out one by one, and
is how many values there are.
A grouped frequency table for continuous data has the classes and
. What is wrong with them, and what do you do about it?
There is a gap: nothing between 19 and 20 has a class to go in, and a continuous quantity can take those values.
Close the gap before doing any calculation: for data that has been rounded, the boundaries become
For data that has been truncated, such as an age counted in whole years, they become and
instead, since someone is 19 right up until their 20th birthday.
For grouped data you give the modal class rather than the mode. Why can you not give a mode?
Grouping the data throws away the individual values, so there is no longer any way to tell which single value occurred most often.
What survives is which class the data is most concentrated in, so that is what you report.
The same limitation applies to the mean and the median from a grouped table: everything you can calculate is an estimate.
What do you use in place of the actual data values when estimating the mean from a grouped frequency table?
The midpoint of each class, which stands in for every value in that class; the midpoint is the mean of the class's lower and upper boundaries, so has midpoint 152.5.
This is why the answer is only an estimate: it assumes the values in each class average out at its midpoint, which they will not do exactly.
Watch for a class of a different width from the others, since its midpoint still has to be worked out from its own boundaries.
Complete the assumption that linear interpolation makes about grouped data:
The data values are assumed to be spread throughout each class.
The completed assumption is:
The data values are assumed to be evenly spread throughout each class.
That assumption is what makes the method work. If the values are evenly spread, then a position one third of the way through a class's frequency corresponds to a data value one third of the way through that class's range, so proportion can be used.
It is also why an interpolated median or quartile is an estimate: real data is rarely spread evenly.
is the 6.25th value, and it lies in the class
, which runs from a cumulative frequency of 3 to a cumulative frequency of 8. How does linear interpolation give
?
Set the fraction of the way through the cumulative frequency equal to the fraction of the way through the class boundaries:
so .
The part that is easiest to forget is adding the lower class boundary back on at the end. The proportion gives you how far into the class the value sits, not the value itself.
Define variance.
The variance is a measure of spread: it measures how varied a set of data is about its mean.
Data that is spread out has a greater variance, and data whose values are close together has a smaller variance.
The symbol for the variance of a population is , using the lower-case Greek letter sigma.
How are the standard deviation and the variance related, and why is the standard deviation usually the one quoted?
The standard deviation is the square root of the variance, so , and the reason for preferring it is units.
Squaring the deviations to find the variance squares the units as well, so a variance of times in minutes comes out in minutes squared, which cannot be compared with the data.
Taking the square root puts them back, leaving the standard deviation in the same units as the original data; the two are used interchangeably in this course, so check which one a question gives you and which it wants.
Complete the version of the variance formula that is quickest to use in most questions:
The completed formula is:
An easy way to remember it is "the mean of the squares minus the square of the mean".
The order matters: subtracting the other way round would give a negative answer, and a variance can never be negative.
is the definition of the variance, so why is it rarely the version you use?
It is slow: you have to find the mean first, then go back through the data working out a separate squared deviation for every single value, and only then add them up.
The equivalent version needs only two totals,
and
, and those are usually exactly what a question hands you as summary statistics, or what a calculator produces from the data.
The two forms always give the same answer: one is a rearrangement of the other.
How does the variance formula change when the data is given in a frequency table?
Every value is weighted by its frequency, and the total frequency replaces
:
It is the same formula, with the second bracket just written out from the table; for a grouped table,
is the midpoint of each class, so the answer is an estimate.
Be consistent about which statistics you use: if the you put in is calculated from coded or midpoint values, the mean you square must come from the same values.
What is the summary statistic , and how does it give the variance?
is the total squared deviation from the mean:
The variance is then , and the standard deviation is
.
It is worth recognising, because it is the form the formula booklet uses and it turns up again in the formula for the product moment correlation coefficient.
True or False?
Two data sets with the same mean must have the same standard deviation.
False.
The mean is a measure of location and the standard deviation is a measure of spread, and they carry completely independent information.
So, for example, 49, 50, 51 and 10, 50, 90 both have a mean of 50, but the second is far more spread out and has a much larger standard deviation.
This is why a description of a data set should always give one of each: an average on its own says nothing about how varied the data is.
Define coding a set of data.
Coding is applying the same formula to every value in a data set, to simplify the numbers you have to work with.
It is most useful when the data consists of very large or very small numbers, so coding such as turns awkward values into small, convenient ones.
The coding must be applied to every value, or the coded data no longer represents the original.
Data is coded using
. What happens to the mean?
The mean is coded in exactly the same way as the data:
The reason is that the mean is a measure of location.
Changing every value changes where the data sits, and the mean follows it: multiply everything by and the centre is multiplied by
; add
to everything and the centre moves up by
.
Data is coded using
. What happens to the standard deviation?
Only the multiplier affects it:
Adding or subtracting makes no difference at all, because the standard deviation is a measure of spread: sliding every value along by the same amount leaves every gap between values exactly as it was.
Multiplying by does stretch those gaps, so it stretches the standard deviation with them.
Data was coded using
, and the mean of the coded data is
. Complete the way the original mean is recovered:
The completed formula is:
This is just rearranged, which is worth doing rather than memorising: whatever the coding formula was, write it down with the means in it and solve for the one you want.
Undo the operations in reverse order, so subtract before dividing.
In , why is the modulus of
taken?
Because a standard deviation can never be negative: it measures how spread out data is, and a spread of less than nothing means nothing.
Coding with a negative multiplier, such as , reflects the data as well as stretching it, but reflecting does not change how spread out it is.
So the standard deviation is multiplied by the positive value 3, not by .
Heights , in cm, are coded using
, and the coded data has standard deviation 1.3102. What is the standard deviation of the heights?
Only the division by 20 has to be reversed:
Subtracting 250 slid every height down the scale without changing how spread out they were, so it has no part in the calculation.
Note that the answer carries the units of the original data, cm, not of the coded values, which have no units at all.
By signing up you agree to our Terms and Privacy Policy