Exam code: 9709
1/280Still learning
Know0
Define a measure of central tendency.
A measure of central tendency is a single value describing where the centre of a data set lies, which is what is usually meant by an average.
The three used in this course are the mean, the median and the mode.

Join for free to unlock a full flashcard set, track what you know,
and turn revision into real progress.
What is the median of a data set, and how do you find it when there is an even number of values?
The median is the middle value once the data has been put in order of size.
With an even number of values there is no single middle one, so the median is the midpoint of the two middle values: for 14, 19, 19, 23, 27, 28 it is .
Complete the formula for the mean of data values:
The completed formula is:
is the total of all the data values and
is how many of them there are, so the mean shares that total out equally between them.
Was this flashcard helpful?
Define a measure of central tendency.
A measure of central tendency is a single value describing where the centre of a data set lies, which is what is usually meant by an average.
The three used in this course are the mean, the median and the mode.
What is the median of a data set, and how do you find it when there is an even number of values?
The median is the middle value once the data has been put in order of size.
With an even number of values there is no single middle one, so the median is the midpoint of the two middle values: for 14, 19, 19, 23, 27, 28 it is .
Complete the formula for the mean of data values:
The completed formula is:
is the total of all the data values and
is how many of them there are, so the mean shares that total out equally between them.
Why can a single extreme value make the mean a poor description of a data set?
Because the mean is worked out from every value, so one unusually large or small value changes the total and drags the mean towards it.
The mean can then sit above almost all of the data, which is why the median is preferred for data sets containing extreme values.
True or False?
Every data set has exactly one mode.
False.
A data set has two modes if two values tie for the highest frequency, and no mode at all if every value occurs the same number of times.
A data set with two modes is described as bimodal.
What do the three quartiles divide an ordered data set into, and what does each one mark?
They divide it into four equal sections, each holding a quarter of the data.
The lower quartile marks the point 25% of the way through, the median
the point 50% of the way through, and the upper quartile
the point 75% of the way through.
In an ordered list of values, how do you find the position of the lower quartile?
Work out , which gives the position of the lower quartile within the ordered list.
If that comes out as a whole number the lower quartile is the value in that position, and if it does not you take the midpoint of the two values on either side of it.
The upper quartile is found the same way from .
Why is the interquartile range often a better measure of spread than the range?
Because the range is worked out from only the largest and smallest values, which are the two most likely to be unusual, so a single odd value can change it completely.
The interquartile range, , covers only the middle half of the data, so values at either extreme do not affect it at all.
True or False?
The range and the interquartile range are measured in the same units as the data itself.
True.
Both are found by subtracting one data value from another, so the answer carries whatever unit the data was measured in: a range of 67 for data measured in metres is a range of 67 metres.
Why is a mean found from a grouped frequency table only an estimate?
Because grouping records only how many values fell in each class, so the individual values are lost and cannot be added up.
An ungrouped table records how many times each individual value occurred, so nothing is thrown away and the mean it gives is the true one.
In a frequency table, why is each data value multiplied by its frequency before the total is divided?
Because a value listed once with a frequency of stands for
separate data values, so
is their combined contribution.
Adding these products and dividing by the total frequency therefore gives the mean:
In an ungrouped frequency table of values, how do you find which data value is the median?
Work out to get the median's position in the ordered data, then add the frequencies down the table until the running total first reaches that position.
The median is the value in the row where the running total arrives, because the table already lists the values in order and the running total says how far through them you have reached.
True or False?
In a table with classes and
, a height of exactly 160 cm belongs to the first class.
False.
The first class stops below 160, because excludes 160 itself, while the second class starts at 160, because
includes it.
So 160 cm goes into , and it is the inequality signs rather than the numbers that settle every case like this.
Grouping a frequency table hides the individual data values. Why does that mean you can only give a modal class?
Because the table records only how many values fell in each class and never which values they were, so there is no way to tell which single value occurred most often.
The most that can be said is which class the data is concentrated in, and that class is called the modal class.
When estimating the mean from a grouped frequency table, what value do you use for each class, and what does that assume?
Use the midpoint of each class, which is the mean of its lower and upper boundaries, and then work exactly as you would for an ungrouped table.
This assumes that the values within each class average out at its midpoint, which they need not do.
A grouped table has classes and
, leaving a gap between them. Complete the boundaries once the gap has been closed:
The completed boundaries are:
The boundaries move to the halfway points because the data has been rounded, so a value recorded as 19 could really be anything up to 19.5.
Closing the gap matters because a value of 19.4 would otherwise belong to no class at all.
Define the standard deviation of a set of data.
The standard deviation, written , measures how spread out the values in a data set are about the mean.
A data set whose values cluster close together has a small standard deviation, and one whose values are widely scattered has a large one.
Complete the formula for the variance of data values with mean
:
The completed formula is:
is the sum of the squares of the data values, which is not the same thing as
, the square of their sum.
This is the useful form because the two totals and
are enough to find the standard deviation without knowing a single individual value.
What are the units of the variance, and what are the units of the standard deviation?
The standard deviation is in the same units as the data itself, while the variance is in those units squared.
So a set of times measured in minutes has, for example, a variance of and a standard deviation of
.
True or False?
A data set with a variance of 16 has a standard deviation of 16.
False.
The standard deviation is the square root of the variance, so a variance of 16 gives a standard deviation of .
This is why the two are written and
: the symbols themselves record that one is the square of the other.
The variance can also be written as . What does this form show about what the variance measures?
It shows that the variance is the mean of the squared deviations from the mean: each value's distance from the mean is squared, and those squares are then averaged.
Written this way it is slow to calculate, because it needs the mean first and then a separate subtraction for every value in the data set.
The variance has more than one formula. What decides which one you can use?
What the question supplies decides it. If you are given the individual data values then either form will work.
If you are given only the totals and not the values themselves, the form built from those totals is the only one available.
Define an assumed mean.
An assumed mean is a convenient value that is subtracted from every value in a data set before the statistics are calculated.
It is usually chosen as an easy round number close to where the true mean is expected to lie, and the data is then said to have been coded.
What is the point of coding a data set before calculating with it?
Coding replaces awkward numbers with much smaller and simpler ones, which makes the arithmetic quicker and less error-prone.
It is most useful when the data consists of very large or very small numbers, such as volumes all close to 150 ml, where the coded values come out as small numbers instead.
True or False?
Subtracting the same number from every value in a data set leaves the standard deviation unchanged.
True.
Subtracting a constant slides every value along by the same amount, so the values sit exactly as far apart from one another as they did before, and a measure of spread cannot detect a shift.
In a sample of 20 cups of coffee, , where
ml is the volume of a cup. Find the mean volume.
The mean of the coded values is , so on average a cup holds 0.8 ml less than the assumed mean of 150 ml.
Adding the 150 back gives a mean volume of
A data set is coded using . How is the standard deviation of the coded data related to the standard deviation of the original?
They are related by
so the coded standard deviation is the original one multiplied by the size of the multiplier .
The modulus is needed because a negative multiplier reverses the order of the data, but a measure of spread can never come out negative.
The mean of some coded data is
. How do you get back to the mean
of the original data?
Reverse the coding, undoing the addition first and then the multiplication:
The mean is a measure of location, so whatever was done to every data value has been done to the mean as well, and can simply be undone.
By signing up you agree to our Terms and Privacy Policy